Highly multiplexed PCR methods and compositions
A method using a library of test primers to minimize primer dimers in multiplex PCR improves the accuracy and efficiency of nucleic acid analysis, particularly in prenatal diagnosis, by amplifying thousands to millions of target loci with high specificity and sensitivity.
Patent Information
- Application Number
- JP2024101181
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2012-11-21
- Filing Date
- 2024-06-24
- Publication Date
- 2026-02-19
- Estimated Expiration
- 2032-11-21
AI Technical Summary
Current multiplex PCR methods generate non-target amplification products such as primer dimers, limiting the use of amplified products in subsequent analyses, particularly in applications like noninvasive prenatal genetic diagnosis, which are either inaccurate or risky.
A method involving a library of test primers that simultaneously hybridize to thousands to millions of target loci, with specific conditions to minimize primer dimer formation, allowing for high specificity and sensitivity in amplifying target amplicons.
The method significantly reduces primer dimers, enhancing the accuracy and efficiency of nucleic acid analysis, particularly in prenatal diagnosis, by enabling the detection of chromosomal abnormalities with high sensitivity and specificity.
Smart Images

Figure 0007818039000040 
Figure 0007818039000041 
Figure 0007818039000042
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Utility Application No. 13 / 683,604, filed November 21, 2012, and U.S. Provisional Patent Application No. 61 / 675,020, filed July 24, 2012. U.S. Utility Application No. 13 / 683,604 is a continuation-in-part of U.S. Utility Application No. 13 / 300,235, filed November 18, 2011, which is a continuation-in-part of U.S. Utility Application No. 13 / 110,685, filed May 18, 2011, and claims the benefit of U.S. Provisional Patent Application No. 61 / 675,020, filed July 24, 2012. U.S. Utility Model Application No. 13 / 110,685 claims the benefit of U.S. Provisional Patent Application No. 61 / 395,850, filed May 18, 2010; U.S. Provisional Patent Application No. 61 / 398,159, filed June 21, 2010; U.S. Provisional Patent Application No. 61 / 462,972, filed February 9, 2011; U.S. Provisional Patent Application No. 61 / 448,547, filed March 2, 2011; and U.S. Provisional Patent Application No. 61 / 516,996, filed April 12, 2011. U.S. Utility Model Application No. 13 / 300,235 claims the benefit of U.S. Provisional Patent Application No. 61 / 571,248, filed June 23, 2011. The entirety of all of these patent applications is incorporated herein by reference for all their teachings.
[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made with support from the National Institutes of Health under Grant No. 5R44HD60423-3. The United States government may reserve rights to any patents originating from this application.
[0003] FIELD OF THE INVENTION The present invention generally relates to methods and compositions for simultaneously amplifying multiple nucleic acid regions of interest in a single reaction volume. [Background technology]
[0004] To increase assay throughput and enable more efficient use of nucleic acid samples, simultaneous amplification of many target nucleic acids in a sample of interest can be performed by mixing many oligonucleotide primers with the sample and then subjecting the sample to polymerase chain reaction (PCR) conditions in a process known in the art as multiplex PCR. The use of multiplex PCR is a very simple experimental procedure and can reduce the time required for nucleic acid analysis and detection. However, when multiple pairs are added to the same PCR reaction, non-target amplification products, such as amplified primer dimers, can be generated. The likelihood of such product generation increases with increasing number of primers. These non-target amplification products significantly limit the use of the amplified products in subsequent analyses and / or assays. Therefore, improved methods are needed to reduce the formation of non-target amplification products during multiplex PCR.
[0005] Improved multiplex PCR methods would be useful for a variety of applications, including noninvasive prenatal genetic diagnosis (NPD). Specifically, current methods of prenatal diagnosis can alert physicians and parents to abnormalities in the developing fetus. Without prenatal diagnosis, one in 50 infants is born with significant physical or mental disabilities, and as many as one in 30 have some form of congenital malformation. Unfortunately, standard methods involve invasive procedures that are either inaccurate or carry the risk of miscarriage. Methods based on maternal blood hormone levels or ultrasound measurements are noninvasive but similarly inaccurate. Methods such as amniocentesis, chorionic villus biopsy, and fetal blood sampling are highly accurate but invasive and carry significant risks. Amniocentesis was performed in approximately 3% of all pregnancies in the United States, but its frequency of use has declined over the past 15 years.
[0006] Normal humans have two sets of 23 chromosomes in every healthy diploid cell, one copy from each parent. Aneuploidy, a condition in which a nuclear cell has too many and / or too few chromosomes, is thought to be responsible for the majority of implantation failures, miscarriages, and genetic diseases. Detecting chromosomal abnormalities can, among other things, increase the chances of successful pregnancy and identify individuals or embryos with conditions such as Down syndrome, Klinefelter syndrome, and Turner syndrome. Testing for chromosomal abnormalities is particularly important because it is estimated that at least 40% of embryos are abnormal when the mother is between 35 and 40 years old, and more than half of embryos are abnormal when the mother is over 40 years old.
[0007] Recently, it has been discovered that cell-free fetal DNA and intact fetal cells can enter the maternal blood circulation. Therefore, analyzing this genetic material may enable early NPD. Improved methods would be desirable to increase sensitivity and specificity and reduce the time and cost required for NPD. Summary of the Invention [Means for solving the problem]
[0008] In one aspect, the invention features a method for amplifying target loci in a nucleic acid sample. In some embodiments, the method includes (i) contacting the nucleic acid sample with a library of test primers that simultaneously hybridize to at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci to generate a reaction mixture, and (ii) subjecting the reaction mixture to primer extension reaction conditions to generate amplification products that include target amplicons. In some embodiments, the method also includes determining the presence or absence of at least one target amplicon (e.g., at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target amplicons). In some embodiments, the method also includes determining the sequence of at least one target amplicon (eg, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target amplicon).
[0009] In various embodiments of any aspect of the invention, at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci are amplified. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the amplification products are target amplification products. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci are amplified. In various embodiments, less than 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.5, 0.25, 0.1, or 0.05% of the amplification products are primer dimers. In some embodiments, the library of test primers includes at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 test primer pairs, each primer pair including a forward test primer and a reverse test primer that hybridize to the same target locus. In some embodiments, the library of test primers includes at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 individual test primers that hybridize to different target loci, wherein the individual primers are not part of a primer pair.
[0010] In various embodiments of any aspect of the invention, the concentration of each test primer is less than 100, 75, 50, 25, 10, 5, 2, or 1 nM. In various embodiments, the GC content of the test primers is 30-80%, e.g., 40-70% or 50-60%. In some embodiments, the range of GC content of the test primers (e.g., maximum GC content - minimum GC content, e.g., 80%-60% = 20% range) is less than 30, 20, 10, or 5%. In some embodiments, the melting temperature (T m) is 40 to 80°C, e.g., 50 to 70°C, 55 to 65°C, or 57 to 60.5°C. In some embodiments, the melting temperature range of the test primers is less than 20, 15, 10, 5, 3, or 1°C. In some embodiments, the length of the test primers is 15 to 100 nucleotides, e.g., 15 to 75 nucleotides, 15 to 40 nucleotides, 17 to 35 nucleotides, 18 to 30 nucleotides, or 20 to 65 nucleotides. In some embodiments, the test primers include a tag that is not target-specific, e.g., a tag that forms an internal loop structure. In some embodiments, the tag is located between two DNA-binding regions. In various embodiments, the test primers include a 5' region specific to the target locus, an internal region that forms a loop structure that is not specific to the target locus, and a 3' region specific to the target locus. In various embodiments, the length of the 3' region is at least 7 nucleotides. In some embodiments, the 3' region is 7 to 20 nucleotides in length, e.g., 7 to 15 nucleotides, or 7 to 10 nucleotides in length. In various embodiments, the test primer includes a 5' region that is not specific to the target locus (e.g., a tag or universal primer binding site), followed by a region that is specific to the target locus, an internal region that forms a loop structure that is not specific to the target locus, and a 3' region that is specific to the target locus. In some embodiments, the length of the test primer ranges from 50, 40, 30, 20, 10, or less than 5 nucleotides. In some embodiments, the length of the target amplicon is 50 to 100 nucleotides, e.g., 60 to 80 nucleotides, or 60 to 75 nucleotides. In some embodiments, the length of the target amplicon ranges from 50, 25, 15, 10, or less than 5 nucleotides.
[0011] In various embodiments of any aspect of the invention, the primer extension reaction conditions are polymerase chain reaction (PCR) conditions. In various embodiments, the length of the annealing step is greater than 3, 5, 8, 10, or 15 minutes. In various embodiments, the length of the extension step is greater than 3, 5, 8, 10, or 15 minutes.
[0012] In various embodiments of any aspect of the invention, test primers are used to simultaneously amplify at least 1,000 different target loci in a sample containing maternal DNA from the pregnant mother of a fetus and fetal DNA to determine the presence or absence of a fetal chromosomal abnormality. In various embodiments, the method includes ligating a universal primer binding site to DNA molecules in the sample, amplifying the ligated DNA molecules with at least 1,000 specific primers and a universal primer to generate a first set of amplicons, and amplifying the first set of amplicons with at least 1,000 pairs of specific primers to generate a second set of amplicons. In various embodiments, at least 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different primer pairs are used.
[0013] In various embodiments of any of the aspects of the invention, the test primers are used to simultaneously amplify at least 1,000 different target loci in a sample containing DNA from the alleged father of a fetus, and also to simultaneously amplify target loci in a sample containing maternal DNA from the pregnant mother of the fetus and fetal DNA to determine whether the alleged father is the biological father of the fetus.
[0014] In various embodiments of any of the aspects of the invention, the test primers are used to simultaneously amplify at least 1,000 different target loci in a cell or cells from an embryo to determine the presence or absence of chromosomal abnormalities, hi various embodiments, cells from a set of two or more embryos are analyzed and one embryo is selected for in vitro fertilization.
[0015] In various embodiments of any of the aspects of the invention, the test primers are used to simultaneously amplify at least 1,000 different target loci in a forensic nucleic acid sample, hi various embodiments, the length of the annealing step is greater than 3, 5, 8, 10, or 15 minutes.
[0016] In various embodiments of any of the aspects of the invention, the method includes using test primers to simultaneously amplify at least 1,000 different target loci in a control nucleic acid sample to produce a first set of target amplicons and simultaneously amplify the target loci in a test nucleic acid sample to produce a second set of target amplicons, and comparing the first and second sets of target amplicons to determine whether a target locus is present in one sample but absent in the other sample, or whether a target locus is present at different levels in the control sample and the test sample. In various embodiments, the test sample is from an individual suspected of having a disease or phenotype of interest (e.g., cancer) or at increased risk of having the disease or phenotype of interest, and further wherein one or more target loci comprise a sequence (e.g., a polymorphism or other mutation) associated with an increased risk of, or associated with, the disease or phenotype of interest. In various embodiments, the method includes using test primers to simultaneously amplify 1,000 different target loci in a control sample comprising RNA to produce a first set of target amplicons and simultaneously amplify target loci in a test sample comprising RNA to produce a second set of target amplicons, and comparing the first and second sets of target amplicons to determine whether there is a difference in RNA expression levels between the control sample and the test sample. In various embodiments, the RNA is mRNA. In various embodiments, the test sample is from an individual suspected of having or at increased risk of a disease or phenotype of interest (e.g., cancer), and one or more target loci contain a sequence (e.g., a polymorphism or other mutation) associated with an increased risk of or associated with the disease or phenotype of interest. In some embodiments, the test sample is from an individual diagnosed with a disease or phenotype of interest (e.g., cancer), where a difference in RNA expression levels between the control sample and the test sample indicates that the target locus contains a sequence (e.g., a polymorphism or other mutation) associated with an increased or decreased risk of the disease or phenotype of interest.
[0017] In some embodiments of any of the aspects of the invention, the test primers are selected from a library of candidate primers based on one or more parameters, such as the selection of primers using any of the methods of the invention, hi some embodiments, the test primers are selected from the library of candidate primers based at least in part on the ability of the candidate primers to form primer-dimers.
[0018] In one aspect, the invention features a method for selecting test primers from a library of candidate primers. In various embodiments, the selection includes (i) using a computer to calculate undesirability scores for most or all possible combinations of two candidate primers from the library, where each undesirability score is based at least in part on the likelihood of dimer formation between the two candidate primers; (ii) removing the candidate primer from the library of candidate primers with the highest undesirability score; (iii) if the candidate primer removed in step (ii) is a member of a primer pair, removing the other member of the primer pair from the library of candidate primers; and (iv) optionally, repeating steps (ii) and (iii) to select a library of test primers. In some embodiments, the selection method is performed until all undesirability scores for the candidate primer combinations remaining in the library are below a minimum threshold. In some embodiments, the selection method is performed until the number of candidate primers remaining in the library is reduced to a desired number. In various embodiments, the undesirability score is calculated for at least 80, 90, 95, 98, 99, or 99.5% of the possible combinations of candidate primers in the library. In various embodiments, the candidate primers remaining in the library are capable of simultaneously amplifying at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci. In various embodiments, the method also includes (v) contacting a nucleic acid sample containing the target loci with the candidate primers remaining in the library to produce a reaction mixture; and (vi) subjecting the reaction mixture to primer extension reaction conditions to produce amplification products containing the target amplicons.
[0019] In one aspect, the invention features a method for selecting test primers from a library of candidate primers. In various embodiments, the method for selecting test primers from a library of candidate primers includes: (i) calculating, using a computer, undesirability scores for most or all possible combinations of two candidate primers from the library, where each undesirability score is based at least in part on the likelihood of dimer formation between the two candidate primers; (ii) removing from the library of candidate primers those candidate primers that are part of the greatest number of combinations of two candidate primers that have undesirability scores above a first minimum threshold; (iii) if the candidate primer removed in step (ii) is a member of a primer pair, removing the remaining members of the primer pair from the library of candidate primers; and (iv) optionally, repeating steps (ii) and (iii) to select a library of test primers. In some embodiments, the selection method is performed until the undesirability scores for the candidate primer combinations remaining in the library are all below the first minimum threshold. In some embodiments, the selection method is performed until the number of candidate primers remaining in the library is reduced to a desired number. In various embodiments, the undesirability score is calculated for at least 80, 90, 95, 98, 99, or 99.5% of the possible candidate primer combinations in the library. In various embodiments, the candidate primers remaining in the library are capable of simultaneously amplifying at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci. In various embodiments, the method also includes (v) contacting a nucleic acid sample containing the target loci with the candidate primers remaining in the library to produce a reaction mixture; and (vi) subjecting the reaction mixture to primer extension reaction conditions to produce amplification products containing the target amplicons.
[0020] In various embodiments of any of the aspects of the invention, the selection method includes further reducing the number of candidate primers remaining in the library by lowering the first minimum threshold used in step (ii) to a lower second minimum threshold, and optionally repeating (ii) and (iii). In some embodiments, the selection method includes increasing the first minimum threshold used in step (ii) to a higher second minimum threshold, and optionally repeating (ii) and (iii). In some embodiments, the selection method is carried out until the undesirability scores for the candidate primer combinations remaining in the library are all less than or equal to the second minimum threshold, or until the number of candidate primers remaining in the library has been reduced to a desired number.
[0021] In various embodiments of any of the aspects of the invention, the method includes, prior to step (i), identifying or selecting primers that hybridize to the target locus. In some embodiments, multiple primers (or primer pairs) hybridize to the same target locus, but one primer (or primer pair) is selected for that target locus using a selection method based on one or more parameters. In various embodiments, the method includes, prior to step (ii), removing from the library primer pairs that generate target amplicons that overlap with target amplicons generated by another primer pair. In various embodiments, candidate primers are selected from a group of two or more candidate primers with comparable undesirability scores based on one or more other parameters for removal from the candidate primer library. In some embodiments, the candidate primers remaining in the library are used as a test primer library in any of the methods of the invention. In some embodiments, the resulting test primer library comprises any of the primer libraries of the invention.
[0022] In various embodiments of any of the aspects of the invention, the undesirability score is based at least in part on one or more parameters selected from the group consisting of: heterozygosity rate of the target locus, prevalence associated with the sequence (e.g., polymorphism) of the target locus, disease penetrance associated with the sequence (e.g., polymorphism) of the target locus, specificity of the candidate primer for the target locus, size of the candidate primer, melting temperature of the target amplicon, GC content of the target amplicon, amplification efficiency of the target amplicon, and size of the target amplicon.
[0023] In various embodiments of any of the aspects of the invention, the undesirability score is based at least in part on one or more parameters selected from the group consisting of heterozygosity rate of the target locus, specificity of the candidate primer for the target locus, size of the candidate primer, melting temperature of the target amplicon, GC content of the target amplicon, amplification efficiency of the target amplicon, and size of the target amplicon, and the test primers are used to simultaneously amplify at least 1,000 different target loci in a sample containing maternal DNA from the pregnant mother of the fetus and fetal DNA to determine the presence or absence of a fetal chromosomal abnormality. In various embodiments, the method includes ligating universal primer binding sites to DNA molecules in the sample, amplifying the ligated DNA molecules with at least 1,000 specific primers and the universal primer to generate a first set of amplicons, and amplifying the first set of amplicons with at least 1,000 pairs of specific primers to generate a second set of amplicons. In various embodiments, at least 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different primer pairs are used. In various embodiments, at least 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci are amplified.
[0024] In various embodiments of any aspect of the invention, the undesirability score is based at least in part on one or more parameters selected from the group consisting of the heterozygosity rate of the target locus, the specificity of the candidate primer for the target locus, the size of the candidate primer, the melting temperature of the target amplicon, the GC content of the target amplicon, the amplification efficiency of the target amplicon, and the size of the target amplicon, and the test primers are used to amplify at least 1,000 different target loci in a sample containing DNA from the alleged father of the fetus and simultaneously amplify target loci in a sample containing maternal DNA from the pregnant mother of the fetus and fetal DNA to determine whether the alleged father is the biological father of the fetus. In various embodiments, at least 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci are amplified.
[0025] In various embodiments of any aspect of the invention, the undesirability score is based at least in part on one or more parameters selected from the group consisting of the heterozygosity rate of the target locus, the specificity of the candidate primer for the target locus, the size of the candidate primer, the melting temperature of the target amplicon, the GC content of the target amplicon, the amplification efficiency of the target amplicon, and the size of the target amplicon, and the test primers are used to simultaneously amplify at least 1,000 different target loci in one or more cells from an embryo to determine the presence or absence of chromosomal abnormalities. In various embodiments, cells from a set of two or more embryos are analyzed, and one embryo is selected for in vitro fertilization. In various embodiments, at least 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci are amplified.
[0026] In various embodiments of any aspect of the invention, the undesirability score is based at least in part on one or more parameters selected from the group consisting of the heterozygosity rate of the target locus, the specificity of the candidate primer for the target locus, the size of the candidate primer, the melting temperature of the target amplicon, the GC content of the target amplicon, the amplification efficiency of the target amplicon, and the size of the target amplicon, and at least 1,000 different target loci are simultaneously amplified in the forensic nucleic acid sample using the test primers. In various embodiments, the length of the annealing step is greater than 3, 5, 8, 10, or 15 minutes. In various embodiments, at least 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci are amplified.
[0027] In various embodiments of any of the aspects of the invention, the undesirability score is based at least in part on one or more parameters selected from the group consisting of: heterozygosity rate of the target locus, prevalence associated with the sequence (e.g., polymorphism) of the target locus, disease penetrance associated with the sequence (e.g., polymorphism) of the target locus, specificity of the candidate primer for the target locus, size of the candidate primer, melting temperature of the target amplicon, GC content of the target amplicon, amplification efficiency of the target amplicon, and size of the target amplicon; and the method includes using the test primers to simultaneously amplify at least 1,000 different target loci in a control nucleic acid sample to produce a first set of target amplicons and simultaneously amplifying the target loci in a test nucleic acid sample to produce a second set of target amplicons; and comparing the first and second sets of target amplicons to determine whether the target loci are present in one sample and absent in the other, or whether the target loci are present at different levels in the control sample and the test sample. In various embodiments, the test sample is from an individual suspected of having a disease or phenotype of interest, or an increased risk of the disease or phenotype of interest, and wherein one or more target loci comprise a sequence (e.g., a polymorphism) associated with an increased risk of, or associated with, the disease or phenotype of interest at the target locus. In various embodiments, at least 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci are amplified.
[0028] In various embodiments of any of the aspects of the invention, the undesirability score is based at least in part on one or more parameters selected from the group consisting of the heterozygosity rate of the target locus, the prevalence associated with the sequence (e.g., polymorphism) of the target locus, the disease penetrance associated with the sequence (e.g., polymorphism) of the target locus, the specificity of the candidate primer for the target locus, the size of the candidate primer, the melting temperature of the target amplicon, the GC content of the target amplicon, the amplification efficiency of the target amplicon, and the size of the target amplicon; and the method includes using the test primers to simultaneously amplify 1,000 different target loci in a control nucleic acid sample comprising RNA to produce a first set of target amplicons and simultaneously amplifying the target loci in a test sample comprising RNA to produce a second set of target amplicons, and comparing the first and second sets of target amplicons to determine whether there is a difference in RNA expression levels between the control sample and the test sample. In various embodiments, the RNA is mRNA. In various embodiments, the test sample is from an individual suspected of having a disease or phenotype of interest (e.g., cancer) or an increased risk of a disease or phenotype of interest (e.g., cancer), where one or more target loci contain a sequence (e.g., a polymorphism or other mutation) associated with an increased risk of or associated with the disease or phenotype of interest. In some embodiments, the test sample is from an individual diagnosed with a disease or phenotype of interest (e.g., cancer), where a difference in RNA expression levels between the control sample and the test sample indicates that the target locus contains a sequence (e.g., a polymorphism or other mutation) associated with an increased or decreased risk of the disease or phenotype of interest. In various embodiments, at least 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci are amplified.
[0029] In one aspect, the invention features a primer library. In some embodiments, primers are selected from a library of candidate primers using any of the methods of the invention. In some embodiments, the library includes primers that simultaneously hybridize to at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci. In some embodiments, the library includes primers that simultaneously amplify at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci. In some embodiments, the library contains primers that simultaneously amplify at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci, such that less than 60, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.5, 0.25, 0.1, or 0.05% of the amplification products are primer dimers. In some embodiments, the library comprises primers that simultaneously amplify 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci, such that at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the amplified products are target amplified products. In some embodiments, the library comprises primers that simultaneously amplify target loci such that at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci from 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci are amplified.In some embodiments, the primer library includes at least 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 primer pairs, each primer pair including a forward test primer and a reverse test primer, and each test primer pair hybridizes to a target locus. In some embodiments, the primer library includes at least 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 individual primers that each hybridize to a different target locus, and the individual primers are not part of a primer pair.
[0030] In various embodiments of any aspect of the invention, the concentration of each primer is less than 100, 75, 50, 25, 10, 5, 2, or 1 nM. In various embodiments, the GC content of the primers is 30-80%, e.g., 40-70% or 50-60%. In some embodiments, the GC content range of the primers is less than 30, 20, 10, or 5%. In some embodiments, the melting temperature of the primers is 40-80°C, e.g., 50-70°C, 55-65°C, or 57-60.5°C. In some embodiments, the melting temperature range of the primers is less than 15, 10, 5, 3, or 1°C. In some embodiments, the length of the primers is 15-100 nucleotides, e.g., 15-75 nucleotides, 15-40 nucleotides, 17-35 nucleotides, 18-30 nucleotides, or 20-65 nucleotides. In some embodiments, the primer includes a tag that is not target-specific, e.g., a tag that forms an internal loop structure. In some embodiments, the tag is located between two DNA-binding regions. In various embodiments, the primer includes a 5' region that is specific for the target locus, an internal region that is not specific for the target locus and forms a loop structure, and a 3' region that is specific for the target locus. In various embodiments, the 3' region is at least 7 nucleotides in length. In some embodiments, the 3' region is 7 to 20 nucleotides in length, e.g., 7 to 15 nucleotides, or 7 to 10 nucleotides in length. In various embodiments, the primer includes a 5' region that is not specific for the target locus (e.g., another tag or universal primer binding site), followed by a region that is specific for the target locus, an internal region that is not specific for the target locus and forms a loop structure, and a 3' region that is specific for the target locus. In some embodiments, the length of the primer ranges from 50, 40, 30, 20, 10, or less than 5 nucleotides. In some embodiments, the length of the target amplicon is between 50 and 100 nucleotides, e.g., between 60 and 80 nucleotides, or between 60 and 75 nucleotides, In some embodiments, the length range of the target amplicon is less than 50, 25, 15, 10, or 5 nucleotides.
[0031] In one aspect, the invention provides kits comprising any of the primer libraries of the invention for amplifying target loci in a nucleic acid sample, hi some embodiments, the kits include instructions for using the libraries to amplify target loci.
[0032] In one aspect, the invention features a method for determining the chromosomal ploidy state of a gestating fetus. In some embodiments, the method includes contacting a nucleic acid sample with a library of primers that simultaneously hybridize to at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different polymorphic loci to generate a reaction mixture, wherein the nucleic acid sample includes maternal DNA from the mother of the fetus and fetal DNA from the fetus. In some embodiments, the reaction mixture is subjected to primer extension reaction conditions to generate amplification products, sequencing data is generated from the amplification products using a high-throughput sequencer, the allele counts at the polymorphic loci are calculated by a computer based on the sequencing data, multiple ploidy hypotheses are generated by a computer, each associated with a different possible ploidy state of the chromosome, a joint distribution model is constructed by a computer for the predicted allele counts at the polymorphic loci of the chromosome for each ploidy hypothesis, the relative probability of each ploidy hypothesis is calculated by a computer using the joint distribution model and the allele counts, and the ploidy state of the fetus is called by selecting the ploidy state corresponding to the hypothesis with the greatest probability.
[0033] In one aspect, the invention features a method for determining the ploidy state of chromosomes in a gestating fetus. In one embodiment, the method for determining the ploidy state of chromosomes in a gestating fetus includes obtaining a first DNA sample containing maternal DNA from the fetus's mother and fetal DNA from the fetus, preparing the first sample by isolating the DNA to obtain a prepared sample, measuring the DNA in the prepared sample at multiple polymorphic loci on a chromosome, calculating the allele counts at the multiple polymorphic loci from the DNA measurements made on the prepared sample, generating multiple ploidy hypotheses on a computer, each of which relates to a different possible ploidy state of the chromosome, constructing a joint distribution model for each ploidy hypothesis on a computer for the expected allele counts at the multiple polymorphic loci on the chromosome, using the joint distribution model and the allele counts measured in the prepared sample to determine the relative probability of each ploidy hypothesis on a computer, and calling the ploidy state of the fetus by selecting the ploidy state corresponding to the hypothesis with the greatest probability.
[0034] In one aspect, the invention features a method for testing for abnormal chromosomal distribution in a sample containing a mixture of maternal and fetal DNA. In some embodiments, the method includes the steps of (i) contacting the sample with a primer library that simultaneously hybridizes to at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci to generate a reaction mixture, where the target loci are from a plurality of different chromosomes, and the plurality of different chromosomes includes at least one first chromosome suspected of having an abnormal distribution in the sample and at least one second chromosome presumed to be normally distributed in the sample; and (ii) contacting the sample with a primer library that simultaneously hybridizes to at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci from a plurality of different chromosomes to generate a reaction mixture, where the target loci are from a plurality of different chromosomes, and the plurality of different chromosomes includes at least one first chromosome suspected of having an abnormal distribution in the sample and at least one second chromosome presumed to be normally distributed in the sample. (iii) subjecting the amplified product to primer extension reaction conditions to produce an amplified product; (iii) sequencing the amplified product to obtain a plurality of sequence tags aligned to the target loci, wherein the sequence tags are of sufficient length to be assigned to specific target loci; (iv) assigning, on a computer, the plurality of sequence tags to their corresponding target loci; (v) determining, on a computer, the number of sequence tags to assign to the target loci of the first chromosome and the number of sequence tags to assign to the target loci of the second chromosome; and (vi) comparing, on a computer, the numbers from step (v), to determine the presence or absence of an abnormal distribution of the first chromosome.
[0035] In one aspect, the present invention provides a method for detecting the presence or absence of fetal aneuploidy. In some embodiments, the method includes the steps of: (i) contacting a sample comprising a mixture of maternal and fetal DNA with a library of primers that simultaneously hybridize to at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different non-polymorphic target loci from a plurality of different chromosomes to generate a reaction mixture; (ii) subjecting the reaction mixture to primer extension reaction conditions to generate amplification products comprising target amplicons; (iii) quantifying on a computer the relative frequencies of the target amplicons from the first and second chromosomes of interest; (iv) comparing on a computer the relative frequencies of the target amplicons from the first and second chromosomes of interest; and (v) identifying the presence or absence of aneuploidy based on the compared relative frequencies of the first and second chromosomes of interest. In some embodiments, the first chromosome is a chromosome suspected to be euploid. In some embodiments, the second chromosome is a chromosome suspected to be aneuploid.
[0036] In one aspect, a method for determining the presence or absence of fetal aneuploidy in a maternal tissue sample comprising fetal genomic DNA and maternal genomic DNA, the method comprising: (a) obtaining a mixture of fetal and maternal genomic DNA from the maternal tissue sample; (b) performing massively parallel DNA sequencing of DNA fragments randomly selected from the mixture of fetal and maternal genomic DNA of step (a) to determine the sequences of the DNA fragments; (c) identifying the chromosomes to which the sequences obtained in step (b) belong; (d) using the data from step (c) to determine the amount of at least one first chromosome in the mixture of maternal and fetal genomic DNA, wherein the at least one first chromosome is presumed to be euploid in the fetus; and (e) using the data from step (c) to determine the amount of at least one first chromosome in the mixture of maternal and fetal genomic DNA. (f) calculating the proportion of fetal DNA in the mixture of fetal and maternal DNA; (g) if the second target chromosome is euploid, calculating a predicted distribution of the amounts of the second target chromosome using the number in step (d); (h) if the second target chromosome is aneuploid, calculating a predicted distribution of the amounts of the second target chromosome using the first number in step (d) and the proportion of fetal DNA in the mixture of fetal and maternal DNA calculated in step (f); and (i) determining whether the amount of the second chromosome determined in step (e) is more likely to be part of the distribution calculated in step (g) or the distribution calculated in step (h), thereby indicating the presence or absence of fetal aneuploidy.
[0037] In various embodiments of any of the aspects of the invention, the method also includes obtaining genotype data from one or both parents of the fetus. In some embodiments, obtaining genotype data from one or both parents of the fetus includes preparing DNA from the parents, the preparing DNA comprising preferentially enriching DNA at a plurality of polymorphic loci to obtain prepared parental DNA, optionally amplifying the prepared parental DNA, and measuring the parental DNA in the prepared sample at the plurality of polymorphic loci.
[0038] In various embodiments of any of the aspects of the invention, constructing a joint distribution model for the expected allele count probabilities at multiple polymorphic loci on a chromosome is performed using genetic data obtained from one or both parents. In some embodiments, a sample (e.g., a first sample) is isolated from maternal plasma, and obtaining genotype data from the mother is performed by estimating the maternal genotype data from DNA measurements performed on the prepared sample.
[0039] In one aspect, a diagnostic box is disclosed that aids in determining the ploidy status of chromosomes in a gestating fetus, and that is capable of carrying out any of the preparation and measurement steps of the methods of the present invention.
[0040] In various embodiments of any of the aspects of the invention, the allele counts are probabilistic rather than binary. In some embodiments, measurements of DNA in the prepared sample at multiple polymorphic loci are also used to determine whether the fetus has one or more disease-linked haplotypes.
[0041] In various embodiments of any of the aspects of the invention, constructing a joint distribution model for allele count probabilities is performed by modeling the dependence between polymorphic alleles on a chromosome using data on the probability of chromosomal crossovers at different locations within the chromosome. In some embodiments, constructing a joint distribution model for allele counts and determining the relative probability of each hypothesis is performed using a method that does not require the use of a reference chromosome.
[0042] In various embodiments of any of the aspects of the invention, the step of determining the relative probability of each hypothesis uses an estimated fraction of fetal DNA in the prepared sample. In some embodiments, the DNA measurements from the prepared sample used in calculating allele count probabilities and determining the relative probability of each hypothesis comprise primary genetic data. In some embodiments, the step of selecting the ploidy state corresponding to the hypothesis with the greatest probability is performed using maximum likelihood estimation or maximum a posteriori estimation.
[0043] In various embodiments of any of the aspects of the invention, calling the ploidy state of the fetus also includes combining the relative probabilities of each of the ploidy hypotheses determined using a joint distribution model and allele count probabilities with the relative probabilities of each of the ploidy hypotheses calculated using a statistical technique selected from the group consisting of read count analysis, comparison of heterozygosity rates, statistics available only using parental genetic information, genotype signal probabilities normalized to a particular parental status, statistics calculated using an estimated fetal fraction of the sample (e.g., the first sample) or prepared sample, and combinations thereof.
[0044] In various embodiments of any of the aspects of the invention, a confidence estimate is calculated for the called ploidy state. In some embodiments, the method also includes taking a clinical action selected from one of terminating the pregnancy or maintaining the pregnancy based on the called fetal ploidy state.
[0045] In various embodiments of any aspect of the invention, the method may be performed on a fetus between 4 and 5 weeks of gestation; between 5 and 6 weeks of gestation; between 6 and 7 weeks of gestation; between 7 and 8 weeks of gestation; between 8 and 9 weeks of gestation; between 9 and 10 weeks of gestation; between 10 and 12 weeks of gestation; between 12 and 14 weeks of gestation; between 14 and 20 weeks of gestation; between 20 and 40 weeks of gestation; in the first trimester; in the second trimester; in the third trimester; or combinations thereof.
[0046] In various embodiments of any of the aspects of the invention, the method is used to generate a report showing the determined ploidy state of the chromosome in the gestating fetus. In some embodiments, a kit for determining the ploidy state of a target chromosome in a gestating fetus, designed for use with any of the methods of the invention, is disclosed, the kit including a plurality of internal forward primers and, optionally, a plurality of internal reverse primers, each of which is designed to hybridize to a region of DNA immediately upstream and / or downstream of one of the polymorphic sites on the target chromosome, and optionally, an additional chromosome, where the hybridizing region is separated from the polymorphic site by a small number of bases, the small number being selected from the group consisting of 1, 2, 3, 4, 5, 6 to 10, 11 to 15, 16 to 20, 21 to 25, 26 to 30, 31 to 60, and combinations thereof.
[0047] In one aspect, the invention features a method of determining whether an alleged father is the biological father of a fetus carried by a pregnant mother. In some embodiments, the method includes the steps of: (i) simultaneously amplifying a plurality of polymorphic loci, including polymorphic loci on at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different genetic material from the alleged father, to generate a first set of amplification products; (ii) simultaneously amplifying a corresponding plurality of polymorphic loci in a mixed sample of DNA, including fetal DNA and maternal DNA, from a blood sample from the pregnant mother to generate a second set of amplification products; (iii) determining, on a computer, the probability that the alleged father is the biological father of the fetus using genotypic measurements based on the first and second sets of amplification products; and (iv) establishing whether the alleged father is the biological father of the fetus using the determined probability that the alleged father is the biological father of the fetus. In various embodiments, the method further comprises simultaneously amplifying polymorphic loci on corresponding genetic material from the mother to generate a third set of amplification products, wherein the probability that the alleged father is the biological father of the fetus is determined using genotype measurements based on the first, second, and third sets of amplification products.
[0048] In one aspect, the present invention provides methods for estimating the relative likelihood of each embryo from a set of embryos to develop as desired. In some embodiments, the method includes contacting a sample from each embryo with a library of primers that simultaneously hybridize to at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci to generate a reaction mixture for each embryo, wherein the samples are each obtained from cells from one or more embryos. In some embodiments, each reaction mixture is subjected to primer extension reaction conditions to generate amplification products. In some embodiments, the method includes determining, on a computer, one or more characteristics of at least one cell from each embryo based on the amplification products; and estimating, on a computer, the relative likelihood of each embryo to develop as desired based on the one or more characteristics of the at least one cell of each embryo.
[0049] In one aspect, the invention features a method for measuring the amount of two or more target loci in a nucleic acid sample. In some embodiments, the method includes (i) using PCR to amplify a nucleic acid sample containing a first reference locus, a second reference locus, a first target locus, and a second target locus to form amplification products, where the first reference locus and the first target locus have the same number of nucleotides but have different sequences at one or more nucleotide positions, and the second reference locus and the second target locus have the same number of nucleotides but have different sequences at one or more nucleotide positions; and (ii) sequencing the amplification products to determine the abundance of the amplified loci. (iii) determining a reference ratio comparing the relative amount of the amplified first target locus compared to the amplified second reference locus, where the reference ratio indicates a difference in PCR efficiency for amplification of the first reference locus and the second reference locus; (iv) determining a target ratio comparing the relative amount of the amplified first target locus compared to the amplified second target locus; and (iv) adjusting the target ratio from step (iii) based on the reference ratio from step (ii) to determine the relative amounts of the first target locus and the second target locus in the sample. In various embodiments, the method includes determining the absolute amounts of the first target locus and the second target locus in the sample. In various embodiments, the method further comprises determining the presence or absence of target loci (e.g., at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci) in the sample. In various embodiments, the method comprises using any of the primer libraries of the invention. In various embodiments, the method comprises simultaneously amplifying 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci.
[0050] In one aspect, the invention features a method for quantitatively measuring multiple genetic targets in an analytical sample. In some embodiments, the method includes: (i) mixing genetic material from the analytical sample with multiple target-specific amplification reagents and multiple standard sequences corresponding to the targets of the target-specific amplification reagents; (ii) amplifying target regions of the genetic material and the standard sequences to generate target amplicons and standard sequence amplicons; and (iii) measuring the amounts of the generated target amplicons and standard sequence amplicons. In some embodiments, the genetic material is present in a genetic library. In some embodiments, the genetic targets are polymorphic loci (e.g., SNPs). In some embodiments, measuring the amount is achieved by counting sequences. In some embodiments, the method further includes measuring an estimated copy number of at least one chromosome in the sample from which the genetic library was derived, wherein the measuring includes comparing the number of sequence reads of the target amplicons with the number of sequence reads of the standard amplicons. In some embodiments, the standard sequences and the genetic library contain universal priming sites that allow priming by the same primer. In some embodiments, the mixing step comprises at least 10, 100, 500, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different target-specific amplification reagents and at least 10, 100, 500, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 standard sequences. In various embodiments, the method comprises using any of the primer libraries of the invention. In various embodiments, the method comprises simultaneously amplifying 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target regions. In some embodiments, the relative amounts of each of the reference sequences are known.In some embodiments, the relative amounts of each sequence are calibrated to a reference genome. In some embodiments, the analytical sample comprises a mixture of fetal and maternal genomes. In some embodiments, the analytical sample is derived from the blood or plasma of a pregnant woman. In some embodiments, the reference genome has at least one aneuploidy, such as an aneuploidy of chromosome 13, 18, 21, X, or Y. In some embodiments, the reference genome is diploid.
[0051] In one aspect, the invention features a mixture comprising a plurality of genetic standard sequences, wherein the relative amounts of each genetic standard sequence in the mixture are determined by calibration to a reference genome. In various embodiments, the mixture comprises at least 10, 100, 500, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 genetic standard sequences. In various embodiments, the genetic standard sequences comprise a first universal priming site, a second universal priming site, a first target-specific priming site, a second target-specific priming site, and a marker sequence located between the first and second target-specific priming sites, wherein the first target-specific site and the second target-specific priming site are located between the first and second universal priming sites. In various embodiments, calibration comprises using any of the primer libraries of the present invention.In various embodiments, calibration comprises simultaneously amplifying 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target regions.In some embodiments, the reference genome has at least one aneuploidy, such as the aneuploidy of chromosome 13, 18, 21, X, or Y.In some embodiments, the reference genome is diploid.
[0052] In one aspect, the invention features a method for generating a set of calibrated genetic standard sequences. In some embodiments, the method includes: (i) forming an amplification reaction mixture including a genetic library prepared from a reference genome, a set of target-specific amplification primer reagents, and a set of genetic standard sequences corresponding to the set of target-specific amplification reagents; (ii) amplifying the genetic library and the genetic standard sequences to generate target sequence-derived amplification products and genetic standard sequence-derived amplification products; (iii) measuring the amounts of the target sequence-derived amplification products and the genetic standard sequence-derived amplification products; and (iv) determining the relative amounts of each genetic standard sequence to each other, thereby calibrating the plurality of genetic standard sequences. In various embodiments, at least 10; 100; 500; 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 genetic standard sequences are used. In various embodiments, the method comprises using any of the primer libraries of the present invention. In various embodiments, the present invention comprises simultaneously amplifying 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different sequences. In some embodiments, the reference genome has at least one aneuploidy, such as an aneuploidy of chromosome 13, 18, 21, X, or Y. In some embodiments, the reference genome is diploid.
[0053] In one aspect, the invention provides a set of genetic standard sequences that have been calibrated by any of the methods of the invention. In one aspect, the invention provides a set of genetic standard sequences that can be calibrated before, during or after the method is performed.
[0054] In one aspect, the invention features a method for determining the copy number of a gene of interest containing at least one allele with a deletion. In some embodiments, the method includes: (i) mixing genetic material from a sample to be analyzed with an amplification reagent specific to the gene of interest and incapable of significantly amplifying the deletion-containing allele of the gene of interest, a standard sequence corresponding to the gene of interest, an amplification reagent specific to the reference sequence, and a standard sequence corresponding to the reference sequence; (ii) amplifying the gene sequence of interest, the standard sequence corresponding to the gene of interest, the reference sequence, and the standard sequence corresponding to the reference sequence to generate gene of interest amplicons, reference sequence amplicons, and standard sequence amplicons; and (iii) measuring the quantity of the generated target amplicons and standard sequence amplicons. In some embodiments, measuring the quantity is achieved by counting sequence reads. In some embodiments, the method further includes calculating an estimated copy number of a chromosome in a sample from which at least one genetic library is derived, wherein the calculating step includes comparing the number of sequences in the target amplicons with the number of sequences in the standard amplicons. In some embodiments, the standard sequence and the genetic library contain universal priming sites that allow priming with the same primer. In some embodiments, the relative abundance of each sequence is calibrated to a reference genome. In various embodiments, at least 10; 100; 500; 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 gene reference sequences are used. In various embodiments, the method includes using any of the primer libraries of the present invention. In various embodiments, the method includes simultaneously amplifying 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target regions. In some embodiments, the reference genome is diploid. In some embodiments, the sample for analysis is derived from blood.
[0055] In various embodiments of any of the aspects of the invention, preferentially enriching DNA in a sample (e.g., a first sample) at target loci (e.g., a plurality of polymorphic loci) includes using a plurality of pre-circularized probes, each probe targeting one of the loci (e.g., polymorphic loci), preferably designed such that its 3′ and 5′ ends hybridize to a region of DNA separated from the polymorphic site of the locus by a small number of bases, wherein the small number is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, The method includes the steps of obtaining a probe that is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21-25, 26-30, 31-60, or a combination thereof; hybridizing the pre-circularized probe with DNA from a sample (e.g., a first sample); filling the gap between the hybridized probe ends using a DNA polymerase; circularizing the pre-circularized probe; and amplifying the circularized probe.
[0056] In various embodiments of any of the aspects of the invention, preferentially enriching DNA at a target locus (e.g., a plurality of polymorphic loci) includes using a plurality of ligation-mediated PCR probes, each PCR probe targeting one of the target loci (e.g., polymorphic loci), with upstream and downstream PCR probes designed to hybridize to a region of DNA on one strand of DNA that is separated from the polymorphic site of the locus by a small number of bases, preferably, where the small number is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, , 13, 14, 15, 16, 17, 18, 19, 20, 21-25, 26-30, 31-60, or a combination thereof; hybridizing the ligation-mediated PCR probe with DNA from a sample (e.g., a first sample); filling the gap between the ends of the ligation-mediated PCR probe with a DNA polymerase; ligating the ligation-mediated PCR probe; and amplifying the ligated ligation-mediated PCR probe.
[0057] In some embodiments of various aspects of the invention, preferentially enriching DNA at target loci (e.g., a plurality of polymorphic loci) includes obtaining a plurality of hybrid capture probes that target the loci (e.g., polymorphic loci), hybridizing the hybrid capture probes to DNA in a sample (e.g., a first sample), and physically removing some or all of the unhybridized DNA from the DNA-related sample (e.g., the first sample).
[0058] In some embodiments of any aspect of the present invention, the hybrid capture probe is designed to hybridize to a region adjacent to but not overlapping with the polymorphic site. In some embodiments, the hybrid capture probe is designed to hybridize to a region adjacent to but not overlapping with the polymorphic site, and the length of the adjacent capture probe can be selected from the group consisting of less than about 120 bases, less than about 110 bases, less than about 100 bases, less than about 90 bases, less than about 80 bases, less than about 70 bases, less than about 60 bases, less than about 50 bases, less than about 40 bases, less than about 30 bases, and less than about 25 bases. In some embodiments, the hybrid capture probe is designed to hybridize to a region overlapping with the polymorphic site, and the plurality of hybrid capture probes includes at least two hybrid capture probes for each polymorphic locus, each hybrid capture probe designed to be complementary to a different allele at one polymorphic locus.
[0059] In some embodiments of any of the aspects of the invention, preferentially enriching DNA at a plurality of polymorphic loci includes obtaining a plurality of inner forward primers, each primer targeting one of the polymorphic loci, and a 3' end of the inner forward primer designed to hybridize to a region of DNA upstream of the polymorphic site and separated from the polymorphic site by a small number of bases, the small number being selected from the group consisting of 1 base pair, 2 base pairs, 3 base pairs, 4 base pairs, 5 base pairs, 6-10 base pairs, 11-15 base pairs, 16-20 base pairs, 21-25 base pairs, 26-30 base pairs, or 31-60 base pairs; and, optionally, providing a plurality of inner forward primers. The method includes the steps of obtaining inner reverse primers, each primer targeting one of the polymorphic loci, and the 3' end of the inner reverse primers designed to hybridize to a region of DNA upstream of the polymorphic site and separated from the polymorphic site by a small number of bases, the small number being selected from the group consisting of 1 base pair, 2 base pairs, 3 base pairs, 4 base pairs, 5 base pairs, 6-10 base pairs, 11-15 base pairs, 16-20 base pairs, 21-25 base pairs, 26-30 base pairs, or 31-60 base pairs; hybridizing the inner primers to the DNA; and amplifying the DNA using polymerase chain reaction to form an amplification product.
[0060] In some embodiments of any of the aspects of the invention, the method further includes obtaining a plurality of outer forward primers, each of which targets one of the target loci (e.g., polymorphic loci) and is designed to hybridize to a region of DNA upstream of the inner forward primer; optionally obtaining a plurality of outer reverse primers, each of which targets one of the target loci (e.g., polymorphic loci) and is designed to hybridize to a region of DNA immediately downstream of the inner reverse primer; hybridizing the first primer to the DNA; and amplifying the DNA using polymerase chain reaction.
[0061] In some embodiments of any of the aspects of the invention, the method further includes obtaining a plurality of outer reverse primers, each of which targets one of the target loci (e.g., polymorphic loci) and is designed to hybridize to a region of DNA immediately downstream of the inner reverse primer, and optionally obtaining a plurality of outer forward primers, each of which targets one of the polymorphic loci and is designed to hybridize to a region of DNA upstream of the inner forward primer, hybridizing the first primer to the DNA, and amplifying the DNA using polymerase chain reaction.
[0062] In some embodiments of any of the aspects of the invention, preparing the sample (e.g., the first sample) further comprises adding universal adaptors to DNA in the sample (e.g., the first sample) and amplifying the DNA in the sample (e.g., the first sample) using polymerase chain reaction. In some embodiments, at least a portion of the amplified amplicons is less than 100 bp, less than 90 bp, less than 80 bp, less than 70 bp, less than 65 bp, less than 60 bp, less than 55 bp, less than 50 bp, or less than 45 bp, where the portion is 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 99%.
[0063] In some embodiments of any of the aspects of the invention, the step of amplifying the DNA is carried out in one or more individual reaction volumes, each of which contains more than 100 different forward and reverse primer pairs, more than 200 different forward and reverse primer pairs, more than 500 different forward and reverse primer pairs, more than 1,000 different forward and reverse primer pairs, more than 2,000 different forward and reverse primer pairs, more than 5,000 different forward and reverse primer pairs, more than 10,000 different forward and reverse primer pairs, more than 20,000 different forward and reverse primer pairs, more than 50,000 different forward and reverse primer pairs, or more than 100,000 different forward and reverse primer pairs.
[0064] In some embodiments of any aspect of the invention, preparing the sample (e.g., a first sample) further comprises dividing the sample (e.g., a first sample) into multiple portions, wherein DNA within each portion is preferentially enriched at a subset of target loci (e.g., a plurality of polymorphic loci). In some embodiments, the inner primers are selected by identifying primer pairs that are likely to form undesired primer duplexes and removing at least one of the identified primer pairs that are likely to form undesired primer duplexes from the plurality of primers. In some embodiments, the inner primers contain a region designed to hybridize either upstream or downstream of the target locus (e.g., a polymorphic locus) and, optionally, a universal priming sequence designed to enable PCR amplification. In some embodiments, at least some of the primers further contain a random region that is different for each individual primer molecule. In some embodiments, at least some of the primers further contain a molecular barcode.
[0065] In some embodiments of any aspect of the present invention, the preferential enrichment results in an average allelic bias between the prepared sample and the sample (e.g., the first sample) of a factor selected from the group consisting of 2-fold or less, 1.5-fold or less, 1.2-fold or less, 1.1-fold or less, 1.05-fold or less, 1.02-fold or less, 1.01-fold or less, 1.005-fold or less, 1.002-fold or less, 1.001-fold or less, and 1.0001-fold or less. In some embodiments, the plurality of polymorphic loci are SNPs. In some embodiments, measuring the DNA in the prepared sample is performed by sequencing.
[0066] In some embodiments of any aspect of the invention, the target loci are present on the same nucleic acid (e.g., the same chromosome or the same region of a chromosome) of the control. In some embodiments, at least some target loci are present on different nucleic acids (e.g., different chromosomes) of the subject. In some embodiments, the nucleic acid sample comprises fragmented or digested nucleic acids. In some embodiments, the nucleic acid sample comprises genomic DNA, cDNA, or mRNA. In some embodiments, the nucleic acid sample comprises DNA from a single cell. In some embodiments, the nucleic acid sample is a substantially cell-free blood or plasma sample. In some embodiments, the nucleic acid sample comprises or is derived from blood, plasma, saliva, semen, sperm, cell culture supernatant, mucus secretion, dental plaque, gastrointestinal tissue, stool, urine, hair, bone, body fluids, tears, tissue, skin, nail, blastomere, embryo, amniotic fluid, chorionic villus sample, bile, lymph, cervical mucus, or a forensic sample. In some embodiments, the target loci are segments of human nucleic acid. In some embodiments, the target locus comprises or consists of a single nucleotide polymorphism (SNP). In some embodiments, the primer is a DNA molecule.
[0067] In some embodiments of any aspect of the invention, the DNA in the sample (e.g., the first sample) is derived from maternal plasma. In some embodiments, preparing the sample (e.g., the first sample) further comprises amplifying the DNA. In some embodiments, preparing the sample (e.g., the first sample) further comprises preferentially enriching DNA in the sample (e.g., the first sample) at target loci (e.g., a plurality of polymorphic loci).
[0068] In various embodiments, the primer extension reaction or polymerase chain reaction involves the addition of one or more nucleotides by a polymerase. In various embodiments, the primer extension reaction or polymerase chain reaction does not involve ligation-mediated PCR. In various embodiments, the primer extension reaction or polymerase chain reaction does not involve the ligation of two primers by a ligase. In various embodiments, the primer does not include a ligated inverse probe (LIP), which is also called a pre-circularized probe, pre-circularizing probe or circularization probe, circularization probe, Padlock probe, or molecular inversion probe (MIP).
[0069] All aspects and embodiments of the invention described herein will be understood to include "comprising," "consisting," and "consisting essentially of" aspects and embodiments.
[0070] definition A single nucleotide polymorphism (SNP) refers to a single nucleotide that may differ between the genomes of two members of the same species. The use of this term should not imply any limitation on the frequency with which each variant occurs.
[0071] Sequence refers to a DNA sequence or gene sequence. Sequence can refer to the primary physical structure of an individual's DNA molecule or strand. Sequence can refer to the sequence of nucleotides found in a DNA molecule or a complementary strand of a DNA molecule. Sequence can refer to the information contained in a DNA molecule, represented in silico.
[0072] A locus refers to a specific region of interest on an individual's DNA and may refer to a SNP, which is a potential site of insertion or deletion or some other relevant genetic variation. A disease-linked SNP may also refer to a disease-linked locus.
[0073] A polymorphic allele, also a "polymorphic locus," refers to an allele or locus whose genotype varies among individuals within a given species. Some examples of polymorphic alleles include single nucleotide polymorphisms, short tandem repeats, deletions, duplications, and inversions.
[0074] A polymorphic site refers to a specific nucleotide found in a polymorphic region that varies between individuals.
[0075] An allele refers to a gene that occupies a particular locus.
[0076] Genetic data, also "genotype data," refers to data describing aspects of the genome of one or more individuals. It can refer to a locus or set of loci, a partial sequence or the entire sequence, a portion of a chromosome or an entire chromosome, or the entire genome. It can refer to the identity of one or more nucleotides, it can refer to a sequential set of nucleotides or nucleotides from different locations in the genome, or a combination thereof. Genotype data is generally in silico, but it is also possible to think of the physical nucleotides in a sequence as chemically encoded genetic data. Genotype data can be referred to as "on," "of," "at," "from," or "on" an individual(s). Genotype data can refer to output measurements from a genotyping platform when these measurements are made on genetic material.
[0077] Genetic material, also "genetic sample," refers to physical material, such as tissue or blood, from one or more individuals that contains DNA or RNA.
[0078] Noisy genetic data refers to genetic data that has any of the following: allele dropouts, uncertain base pair measurements, inaccurate base pair measurements, missing base pair measurements, uncertain insertion or deletion measurements, uncertain chromosome segment copy number measurements, false signals, missing measurements, other errors, or a combination thereof.
[0079] Confidence refers to the statistical likelihood that a called SNP, allele, set of alleles, ploidy call, or determined number of chromosome segment copies accurately represents an individual's true genetic state.
[0080] Ploidy calling, also "chromosome copy number calling" or "copy number calling" (CNC), can refer to the act of determining the amount and / or identity of one or more chromosomes present in a cell.
[0081] Aneuploidy refers to the state in which there is an incorrect number of chromosomes in cells (for example, there is an incorrect number of complete chromosomes or an incorrect number of chromosome segments, for example, there is the deletion or duplication of chromosome segments).In the case of human somatic cells, aneuploidy can refer to the case in which cells do not contain 22 pairs of autosomes and one pair of sex chromosomes.In the case of human gametes, aneuploidy can refer to the case in which cells do not contain one of each of the 23 chromosomes.In the case of single chromosome type, aneuploidy can refer to the case in which there are roughly two homologous but not identical chromosome copies, or the case in which there are two chromosome copies originating from the same parent.In some embodiments, the deletion of chromosome segments is microdeletion.
[0082] Ploidy state refers to the amount and / or chromosomal identity of one or more chromosome types in a cell.
[0083] Chromosome may refer to a single chromosome copy, meaning a single molecule of DNA present in 46 copies in normal somatic cells, an example of which is "maternally derived chromosome 18." Chromosome may also refer to the chromosome type present in 23 copies in normal human somatic cells, an example of which is "chromosome 18."
[0084] Chromosomal identity can refer to the number of chromosomes referred to, i.e., chromosome type. A normal human has 22 numbered autosome types and two sex chromosomes. Chromosomal identity can also refer to the chromosomes of parental origin. Chromosomal identity can also refer to the specific chromosomes inherited from a parent. Chromosomal identity can also refer to other identifying features of a chromosome.
[0085] The state of genetic material, or simply "genetic state," can refer to the identity of a set of SNPs on DNA, the haplotype of phased genetic material, and the sequence of DNA, including insertions, deletions, repeats, and mutations. It can also refer to the ploidy state of one or more chromosomes, a chromosome segment, or a set of chromosome segments.
[0086] Allele data refers to a set of genotype data for a set of one or more alleles. Allele data may refer to phased haplotype data. Allele data may refer to SNP identity, and allele data may refer to DNA sequence data, including insertions, deletions, repeats, and mutations. Allele data may include each allele of parental origin.
[0087] An allelic state refers to the actual state of a gene within a set of one or more alleles. An allelic state may refer to the actual state of a gene as described in the allele data.
[0088] Allelic ratio or allele ratio refers to the ratio between the amount of each allele at a locus present in a sample or individual.When a sample is measured by sequencing, allele ratio can refer to the ratio of sequence reads that map to each allele at a locus.When a sample is measured by intensity-based measurement method, allele ratio can refer to the ratio of the amount of each allele present at a locus estimated by measurement method.
[0089] The allele count refers to the number of sequences that map to a particular locus; if the locus is polymorphic, the allele count refers to the number of sequences that map to each of the alleles. If each allele is counted in a binary fashion, the allele count will be an integer. If alleles are counted probabilistically, the allele count can be a fraction.
[0090] Allele count probability refers to the number of sequences that may be mapped to a set of alleles at a particular locus or polymorphic locus, combined with the probability of mapping. Note that the allele count, where the probability of mapping for each counted sequence is binary (0 or 1), is equal to the allele count probability. In some embodiments, the allele count probability can be binary. In some embodiments, the allele count probability can be set equal to the DNA measurement.
[0091] Allele distribution or "allele number distribution" refers to the relative amount of each allele present at each locus within a set of loci. Allele distribution can refer to an individual, a sample, or a set of measurements taken for a sample. In the context of sequencing, allele distribution refers to the number of reads or likelihoods that map to a particular allele for each allele within a set of polymorphic loci. Allele measurements can be treated probabilistically, i.e., the likelihood that a given allele represents a given sequence read is a fraction between 0 and 1, or allele measurements can be treated in a binary manner, i.e., any given read is considered to have exactly 0 or 1 copies of a particular allele.
[0092] Allele distribution pattern refers to the collection of different allele distributions for different parental situations. A particular allele distribution pattern can indicate a particular ploidy state.
[0093] Allelic bias refers to the degree to which the measured allele ratio at heterozygous loci differs from the ratio present in the original DNA sample.The degree of allelic bias at a particular locus is equal to the allele ratio observed at that locus divided by the allele ratio in the original DNA sample at that locus.Allelic bias can be defined as greater than 1, and therefore, when calculating the degree of allelic bias, if a value x is less than 1, the degree of allelic bias can be expressed as 1 / x.Allelic bias can be due to amplification bias, purification bias, or some other phenomenon that affects different alleles differently.
[0094] A primer, also a "PCR probe," refers to a single DNA molecule (DNA oligomer) or a population of DNA molecules (DNA oligomers), where the DNA molecules are identical or nearly identical, and the primer contains a region designed to hybridize with a target locus (e.g., a target polymorphic locus or a non-polymorphic locus) and may contain a priming sequence designed to enable PCR amplification. The primer may also contain a molecular barcode. The primer may contain a random region that is different for each individual molecule. The terms "test primer" and "candidate primer" are not meant to be limiting and may refer to any primer disclosed herein.
[0095] A primer library refers to a collection of two or more primers. In various embodiments, the library comprises at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different primers. In various embodiments, the library comprises at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different primer pairs, each primer pair comprising a forward test primer and a reverse test primer, and each test primer pair hybridizes to a target locus. In some embodiments, the primer library includes at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different individual primers that each hybridize to a different target locus, and the individual primers are not part of a primer pair. In some embodiments, the library includes both (i) primer pairs and (ii) individual primers that are not part of a primer pair (e.g., universal primers).
[0096] Hybrid capture probe refers to any nucleic acid sequence, which may be modified, that is generated by various methods, such as PCR or direct synthesis, and is intended to be complementary to one strand of specific target DNA sequence in sample.Exogenous hybrid capture probe can be added to prepared sample, and hybridized through denaturation-reannealing process to form exogenous fragment-endogenous fragment double strand.These double strands can then be physically separated from sample by various means.
[0097] Sequence read refers to the data that shows the sequence of nucleotide bases measured by clonal sequencing.Clonal sequencing can produce the sequence data that shows a single original DNA molecule or a clone of an original DNA molecule or a cluster of an original DNA molecule.Sequence read can also have an associated quality score that indicates the probability that the nucleotide is correctly called at each base position of the sequence.
[0098] Mapping a sequence read is the process of determining the starting location of a sequence read within the genome sequence of a particular organism, based on the nucleotide sequence similarity between the read and the genome sequence.
[0099] Matching copy error, also known as "matching chromosome aneuploidy" (MCA), refers to the aneuploid state in which one cell contains two identical or nearly identical chromosomes. This type of aneuploidy can occur during gamete formation in meiosis, and can be referred to as meiotic non-disjunction error. This type of error can occur during mitosis. Matching trisomy can refer to the case where an individual has three copies of a given chromosome, and two of the copies are identical.
[0100] Mismatched copy errors, also known as "unique chromosome aneuploidy" (UCA), refer to a state of aneuploidy in which a cell contains two chromosomes that are from the same parent and may be homologous but not identical. This type of aneuploidy can occur during meiosis and can be referred to as a meiotic error. Mismatched trisomy can refer to the case where an individual has three copies of a given chromosome, two of which are from the same parent and are homologous but not identical. Note that mismatched trisomy can refer to the case where two homologous chromosomes from one parent are present, with some segments of the chromosomes being identical while other segments are merely homologous.
[0101] Homologous chromosomes refer to chromosome copies that contain the same set of genes that normally pair up during meiosis.
[0102] Identical chromosomes refer to chromosome copies that contain the same set of genes, and for each gene, identical chromosomes have the same set of alleles that are identical or nearly identical.
[0103] Allelic dropout (ADO) refers to the situation where at least one of the base pairs in a set of base pairs from the homologous chromosome at a given allele is not detected.
[0104] Locus dropout (LDO) refers to the situation where both base pairs within a set of base pairs from homologous chromosomes at a given allele are not detected.
[0105] Homozygous refers to having similar alleles at corresponding chromosomal loci.
[0106] Heterozygous refers to having dissimilar alleles at corresponding chromosomal loci.
[0107] Heterozygosity rate refers to the rate at which individuals in a population have heterozygous alleles at a given locus. Heterozygosity rate can also refer to the expected or measured allele ratio at a given locus in an individual or DNA sample.
[0108] Highly informative single nucleotide polymorphisms (HISNPs) refer to SNPs in which the fetus has an allele that is not present in the mother's genotype.
[0109] A chromosomal region refers to a segment of a chromosome or a complete chromosome.
[0110] A chromosomal segment refers to a section of a chromosome that can range in size from a single base pair to an entire chromosome.
[0111] Chromosome refers to either a complete chromosome or a segment or section of a chromosome.
[0112] Copy refers to the number of copies of chromosome segment.It can refer to the same copy of chromosome segment or the non-identical homologous copy, and different copies of chromosome segment contain substantially similar sets of loci, and one or more of alleles are different.Note that in some cases of aneuploidy, such as M2 copy error, there may be some copies of given chromosome segment that are the same and some copies of the same chromosome segment that are not the same.
[0113] Haplotype generally refers to the combination of alleles at multiple loci that are inherited together on the same chromosome.Haplotype can refer to just two loci or the entire chromosome, depending on the number of recombination events that occur between a given set of loci.Haplotype can also refer to a set of statistically related single nucleotide polymorphisms (SNPs) on a single chromatid.
[0114] Haplotype data, also "phased data" or "ordered genetic data," refers to data from a single chromosome of a diploid or polyploid genome, i.e., from either the separated maternal or paternal copy of a chromosome of a diploid genome.
[0115] Phasing refers to the act of determining an individual's haplotype genetic data given unordered, diploid (or polyploid) genetic data. Phasing can refer to the act of determining, for a set of alleles found on a chromosome, which of the two genes at the allele is associated with each of the individual's two homologous chromosomes.
[0116] Phased data refers to genetic data for which one or more haplotypes have been determined.
[0117] A hypothesis refers to a set of possible ploidy states at a given set of chromosomes or possible allelic states at a given set of loci. A set of possibilities may contain one or more elements.
[0118] Copy number hypothesis, also known as "ploidy state hypothesis," refers to a hypothesis about the number of copies of an individual's chromosomes. It can also refer to a hypothesis about the identity of each chromosome, including the parent of origin of each chromosome and which of the two parental chromosomes is present in the individual. It can also refer to a hypothesis about which, if any, chromosomes or chromosome segments from related individuals genetically correspond to a given chromosome from an individual.
[0119] A target individual refers to an individual whose genetic state is to be determined. In some embodiments, only a limited amount of DNA is available from the target individual. In some embodiments, the target individual is a fetus. In some embodiments, there may be more than one target individual. In some embodiments, each fetus born to a pair of parents may be considered a target individual. In some embodiments, the genetic data to be determined is an allele call or a collection of allele calls. In some embodiments, the genetic data to be determined is a ploidy call.
[0120] A related individual refers to any individual who is genetically related to the target individual and therefore shares a haplotype block with the target individual. In some situations, a related individual may be the genetic parent of the target individual, or any genetic material derived from a parent, such as a sperm, polar body, embryo, fetus, or child. A related individual may also refer to a sibling, parent, or grandparent.
[0121] A sibling refers to any individual whose genetic parents are the same as the individual in question. In some embodiments, a sibling can refer to a born child, embryo, or fetus, or one or more cells derived from a born child, embryo, or fetus. A sibling can also refer to a haploid individual, such as a sperm, polar body, or any other collection of haplotype genetic material, originating from one parent. An individual can be considered a sibling.
[0122] Fetal refers to "of the fetus" or "of the region of the placenta that is genetically similar to the fetus." In pregnant women, parts of the placenta are genetically similar to the fetus, and free-floating fetal DNA found in maternal blood may originate from parts of the placenta that genotype the fetus. Note that the genetic information of half of a fetal chromosomes is inherited from the fetus's mother. In some embodiments, DNA from these maternally inherited chromosomes derived from fetal cells is considered to be "of fetal origin" rather than "of maternal origin."
[0123] DNA of fetal origin refers to DNA that was originally part of a cell whose genotype was essentially identical to that of the fetus.
[0124] DNA of maternal origin refers to DNA that was originally part of a cell whose genotype was essentially identical to that of the mother.
[0125] Offspring may refer to an embryo, a blastomere, or a fetus. It should be noted that in the presently disclosed embodiments, the concepts described apply equally well to an individual that is a born child, fetus, embryo, or a collection of cells derived therefrom. The use of the term offspring is simply meant to connote that the individual referred to as offspring is the genetic descendant of the parent.
[0126] Parent refers to an individual's genetic mother or father. An individual generally has two parents, a mother and a father, although this is not necessarily the case, for example, in genetic or chromosomal chimerism. A parent can be considered an individual.
[0127] Parental context refers to the genetic state of a given SNP on each of the two relevant chromosomes for one or both of the two parents of the target.
[0128] Desired development, as well as "normal development," refers to implantation of a viable embryo in the uterus, resulting in pregnancy, and / or continuation of the pregnancy, resulting in birth, and / or the absence of chromosomal abnormalities in the born offspring, and / or the absence of other undesirable genetic conditions in the born offspring, such as disease-linked genes. The term "desired development" is meant to encompass anything that may be desired by parents or healthcare providers. In some cases, "desired development" can refer to a non-viable or viable embryo that is useful for medical research or other purposes.
[0129] Uterine insertion refers to the process of transferring an embryo into the uterine cavity in the context of in vitro fertilization.
[0130] Maternal plasma refers to the plasma portion of blood from a pregnant woman.
[0131] A clinical decision refers to any decision to take or not take an action that has an outcome that affects the health or survival of an individual. In the context of prenatal diagnosis, a clinical decision can refer to a decision to abort or not abort a fetus. A clinical decision can also refer to a decision to conduct further testing, to take action to reduce an undesirable phenotype, or to take action to prepare for the birth of a child with an abnormality.
[0132] A diagnostic box refers to a machine or combination of machines designed to perform one or more aspects of the methods disclosed herein. In some embodiments, the diagnostic box can be located at a patient care location. In some embodiments, the diagnostic box can perform targeted amplification followed by sequencing. In some embodiments, the diagnostic box can function independently or with the assistance of a technician.
[0133] Informatics-based methods refer to methods that rely heavily on statistics to interpret large amounts of data. In the context of prenatal diagnosis, informatics-based methods refer to methods designed to determine the ploidy state of one or more chromosomes or the allelic state of one or more alleles, not by directly physically measuring the state, but by statistically inferring the most likely state, taking into account large amounts of genetic data, for example, from molecular arrays or sequencing. In some embodiments of the present disclosure, the informatics-based technique may be that disclosed in this patent. In some embodiments of the present disclosure, the informatics-based technique may be PARENTAL SUPPORT™.
[0134] Primary genetic data refers to the analog intensity signals output from a genotyping platform. In the context of SNP arrays, primary genetic data refers to the intensity signals before any genotype calls are made. In the context of sequencing, primary genetic data refers to the analog measurements, similar to a chromatogram, that come from a sequencer before the identity of any base pairs is determined and before the sequence is mapped to a genome.
[0135] Secondary genetic data refers to the processed genetic data output from genotyping platform.In the context of SNP array, secondary genetic data refers to the allele call made by the software associated with SNP array reader, which calls whether a given allele is present or absent in a sample.In the context of sequencing, secondary genetic data refers to determining the identity of the base pair of the sequence, and in some cases, also refers to the sequence being mapped to genome.
[0136] Non-invasive prenatal diagnosis (NPD) or, similarly, "non-invasive prenatal screening" (NPS) refers to a method of determining the genetic status of a fetus during a mother's pregnancy using genetic material found in the mother's blood, which is obtained by drawing the mother's intravenous blood.
[0137] Preferential enrichment of DNA corresponding to a locus, or preferential enrichment of DNA at a locus, refers to any method that results in a higher percentage of DNA molecules in the enriched DNA mixture corresponding to that locus than the percentage of DNA molecules in the pre-enriched DNA mixture corresponding to that locus. The method may involve selective amplification of DNA molecules corresponding to the locus. The method may involve removing DNA molecules that do not correspond to the locus. The method may involve a combination of methods. The degree of enrichment is defined as the percentage of DNA molecules in the enriched mixture corresponding to the locus divided by the percentage of DNA molecules in the pre-enriched mixture corresponding to that locus. Preferential enrichment can be performed at multiple loci. In some embodiments of the present disclosure, the degree of enrichment is greater than 20. In some embodiments of the present disclosure, the degree of enrichment is greater than 200. In some embodiments of the present disclosure, the degree of enrichment is greater than 2,000. When preferential enrichment is performed at multiple loci, the degree of enrichment can refer to the average degree of enrichment of all loci in the set of loci.
[0138] Amplification refers to a method for increasing the number of copies of a DNA molecule.
[0139] Selective amplification can refer to a method that increases the number of copies of specific DNA molecules or DNA molecules corresponding to specific regions of DNA.Selective amplification can also refer to a method that increases the number of copies of specific target DNA molecules or target DNA regions more than increasing the number of unlabeled molecules or regions of DNA.Selective amplification can be a method of preferential enrichment.
[0140] Universal priming sequence refers to the DNA sequence that can be added to a group of target DNA molecules, for example, by ligation, PCR or ligation-mediated PCR.After being added to a group of target molecules, a single pair of amplification primers can be used to amplify the target group using the primer that is specific to universal priming sequence.Universal priming sequence is generally not related to target sequence.
[0141] Universal Adapters or "Ligation Adapters" or "Library Tags" is a DNA molecule containing a universal priming sequence that can be covalently linked to the 5' and 3' ends of a population of target double-stranded DNA molecules. The addition of adapters provides the universal priming sequences at the 5' and 3' ends of the target population from which PCR amplification can be performed, amplifying all molecules from the target population using a single pair of amplification primers.
[0142] Targeting refers to methods used to selectively amplify or otherwise preferentially enrich DNA molecules corresponding to a set of loci in a mixture of DNA.
[0143] A joint distribution model refers to a model that defines the probability of a defined event with respect to multiple random variables, considering multiple random variables defined over the same probability space, where the probabilities of the variables are related. In some embodiments, a degenerate case can be used, where the probabilities of the variables are not related.
[0144] The presently disclosed embodiments will be further described with reference to the accompanying drawings, in which like structure is referenced by like numerals throughout the several views. The drawings shown are not necessarily to scale, emphasis generally being placed upon illustrating the principles of the presently disclosed embodiments. [Brief explanation of the drawings]
[0145] [Figure 1] Schematic diagram of direct multiplex mini-PCR method. [Figure 2] FIG. 1 is an illustration of the semi-nested mini-PCR method. [Figure 3] FIG. 1 is a diagram illustrating the fully nested mini-PCR method. [Figure 4] FIG. 1 is an illustration of the hemi-nested mini-PCR method. [Figure 5] FIG. 1 is an illustration of triple hemi-nested mini-PCR. [Figure 6] FIG. 1 is a diagram illustrating one-sided nested mini-PCR. [Figure 7] FIG. 1 is an explanatory diagram of one-sided mini-PCR. [Figure 8] FIG. 1 is a diagram illustrating the reverse semi-nested mini-PCR method. [Figure 9] FIG. 1 is a diagram of some possible workflows for semi-nested methods. [Figure 10] FIG. 1 is an illustration of a loop ligation adaptor. [Figure 11] FIG. 1 is an illustration of an internally tagged primer. [Figure 12] 1 is an example of some primers with internal tags. [Figure 13] FIG. 1 is an illustration of a method using primers with ligation adaptor binding regions. [Figure 14] 1 is a graph showing the accuracy of simulated ploidy calls for counting methods using two different analytical techniques. [Figure 15]FIG. 1 shows the ratio of the two alleles for multiple SNPs in cell lines from experiment 4. [Figure 16] FIG. 1 shows the ratio of the two alleles for multiple SNPs in the cell lines of experiment 4, separated by chromosome. [Figure 17A] 17A-D show the ratios of the two alleles for multiple SNPs in plasma samples from four pregnant women, separated by chromosome. [Figure 17B] 17A-D show the ratios of the two alleles for multiple SNPs in plasma samples from four pregnant women, separated by chromosome. [Figure 17C] 17A-D show the ratios of the two alleles for multiple SNPs in plasma samples from four pregnant women, separated by chromosome. [Figure 17D] 17A-D show the ratios of the two alleles for multiple SNPs in plasma samples from four pregnant women, separated by chromosome. [Figure 18] 1 is a graph showing the proportion of data that can be explained by binomial variance before and after data correction. [Figure 19] 1 is a graph showing the relative enrichment of fetal DNA in samples after a short library preparation protocol. [Figure 20] This is a graph comparing direct PCR and semi-nested methods in terms of read depth. [Figure 21] 1 is a graph showing a comparison of direct PCR read depth for three genomic samples. [Figure 22] 1 is a graph showing a comparison of read depth of semi-nested mini-PCR for three samples. [Figure 23] 1 is a graph showing a comparison of read depth for 1,200-plex and 9,600-plex reactions. [Figure 24] This is a graph showing the read number ratios for three chromosomes for six types of cells. [Figure 25]Figure 1 shows allele ratios for two 3-cell reactions and for a third reaction performed on 1 ng of genomic DNA for three chromosomes. [Figure 26] FIG. 1 shows allelic ratios for two single-cell reactions across three chromosomes. [Figure 27] FIG. 1 shows the number of loci with a particular minor allele frequency targeted by each of the two primer libraries. [Figure 28-1] Figure 28A: Electrophoresis graph of PCR products. [Figure 28-2] 28B to 28M are electrophoretic diagrams of lanes 1 to 12 in FIG. 28A, respectively. [Figure 28-3] 28B to 28M are electrophoretic diagrams of lanes 1 to 12 in FIG. 28A, respectively. [Figure 28-4] 28B to 28M are electrophoretic diagrams of lanes 1 to 12 in FIG. 28A, respectively. [Figure 29] Figures 29A-29E: Image representation of a method of the present invention for determining fetal aneuploidy (Figure 29A). Maternal and paternal genotype data (from blood or cheek swabs) and crossover frequency data from the HapMap database are used to generate multiple independent hypotheses for each possible fetal ploidy state in silico (Figure 29B). Each of these hypotheses is expanded to include sub-hypotheses that account for different possible crossover points. A data model predicts the likely sequencing data (predicted allele distribution) for each hypothetical fetal genotype and different fetal cfDNA percentages and compares them with the actual sequencing data (Figure 29C). Bayesian statistics is used to determine the likelihood for each hypothesis. In this hypothetical example, the hypothesis with the highest likelihood (euploidy) is determined (Figure 29D). The individual likelihoods for each copy number hypothesis family (monosomy, disomy, or triploidy) in Figure 29C are summed. The maximum likelihood hypothesis was called as the ploidy state, revealing the fetal fraction and demonstrating sample-specific calculation accuracy (Figure 29E). [Figure 30-1]Figures 30A-30H: Representative graphical representations of euploidy (Figures 30A-30C), monosomy (Figure 30D), and trisomy (Figures 30E-30H). In all plots, the x-axis represents the linear position of the individual polymorphic locus along each chromosome (shown at the bottom of the plot), and the y-axis represents the number of A allele reads as a percentage of the total (A+B) allele reads. Maternal and fetal genotypes and the y-axis location of the band centers are shown to the right of the plots. For ease of visualization, plots may be color-coded according to maternal genotype, with red representing an AA maternal genotype, blue representing a BB maternal genotype, and green representing an AB maternal genotype. If desired, maternal allele contributions may be indicated by color in the "Fetal Genotype" column. Allelic contributions are expressed in the format Maternal|Fetal, such as AA|AB for AA maternal and AB fetal alleles. Figure 30A shows a plot generated when two chromosomes are present and the fetal cfDNA fraction is 0%. This plot is from a non-pregnant woman and therefore represents a pattern when the genotype is entirely maternal. Thus, allele clusters are centered around 1 (AA allele), 0.5 (AB allele), and 0 (BB allele). Figure 30B shows a plot generated when two chromosomes are present and the fetal fraction is 12%. The contribution of fetal alleles to the proportion of A allele reads causes the positions of some allele spots to shift up and down along the y-axis. Therefore, bands are centered around 1 (AA|AA allele), 0.94 (AA|AB allele), 0.56 (AB|AA allele), 0.50 (AB|AB allele), 0.44 (AB|BB allele), 0.06 (BB|AB allele), and 0 (BB|BB allele). Figure 30C shows the plot generated when two chromosomes are present and the fetal fraction is 26%. A pattern is readily visible, including two red and two blue peripheral bands and a central green band of the triplet (colors not shown).The bands are centered at 1 (AA|AA alleles), 0.87 (AA|AB alleles), 0.63 (AB|AA alleles), 0.50 (AB|AB alleles), 0.37 (AB|BB alleles), 0.13 (BB|AB alleles), and 0 (BB|BB alleles). Figure 30D shows the plot generated when one chromosome is present and the fetal fraction is 26%. The hallmark pattern of one outer red and one outer blue peripheral band and two central green bands indicates maternally inherited monosomy (colors not shown). Because the fetus contributes only a single allele (A or B) to the allelic reads, the inner peripheral red and blue bands are absent, and the central triplet band is compressed to two bands (colors not shown). Bands are centered at 1 (AA|A allele), 0.57 (AB|A allele), 0.43 (AB|B allele), and 0 (BB|B allele). Figure 30E is a plot generated when three chromosomes are present and the fetal fraction is 27%. This pattern of two red and two blue peripheral bands and two central green bands indicates maternally inherited meiotic trisomy (colors not shown). Bands are centered at 1 (AA|AAA allele), 0.88 (AA|AAB allele), 0.56 (AB|AAB allele), 0.44 (AB|ABB allele), 0.12 (BB|ABB allele), and 0 (BB|BBB allele). Figure 30F is a plot generated when three chromosomes are present and the fetal fraction is 14%. This pattern of three red and three blue peripheral bands, as well as two central green bands, indicates a paternally inherited meiotic trisomy (colors not shown). The bands are centered at 1 (AA|AAA alleles), 0.93 (AA|AAB alleles), 0.87 (AA|ABB alleles), 0.60 (AB|AAA alleles), 0.53 (AB|AAB alleles), 0.47 (AB|ABB alleles), 0.40 (AB|BBB alleles), 0.13 (BB|AAB alleles), 0.07 (BB|ABB alleles), and 0 (BB|BBB alleles). Figure 30G shows a plot generated when three chromosomes are present and the fetal fraction is 35%.This pattern of two red and two blue peripheral bands and four green bands indicates a maternally inherited mitotic trisomy (colors not shown). The bands are centered at 1 (AA|AAA alleles), 0.85 (AA|AAB alleles), 0.72 (AB|AAA alleles), 0.57 (AB|AAB alleles), 0.43 (AB|ABB alleles), 0.28 (AB|BBB alleles), 0.15 (BB|ABB alleles), and 0 (BB|BBB alleles). Figure 30H shows a plot generated when three chromosomes are present and the fetal fraction is 25%. This pattern of two red and two blue peripheral bands and four central green bands indicates a paternally inherited mitotic trisomy (colors not shown). This pattern can be distinguished from a maternally inherited mitotic trisomy (as in Figure 30G) by the location of the inner peripheral bands. Specifically, the bands were centered at 1 (AA|AAA alleles), 0.78 (AA|ABB alleles), 0.67 (AB|AAA alleles), 0.56 (AB|AAB alleles), 0.44 (AB|ABB alleles), 0.33 (AB|BBB alleles), 0.22 (BB|AAB alleles), and 0 (BB|BBB alleles). [Figure 30-2]Figures 30A-30H: Representative graphical representations of euploidy (Figures 30A-30C), monosomy (Figure 30D), and trisomy (Figures 30E-30H). In all plots, the x-axis represents the linear position of the individual polymorphic locus along each chromosome (shown at the bottom of the plot), and the y-axis represents the number of A allele reads as a percentage of the total (A+B) allele reads. Maternal and fetal genotypes and the y-axis location of the band centers are shown to the right of the plots. For ease of visualization, plots may be color-coded according to maternal genotype, with red representing an AA maternal genotype, blue representing a BB maternal genotype, and green representing an AB maternal genotype. If desired, maternal allele contributions may be indicated by color in the "Fetal Genotype" column. Allelic contributions are expressed in the format Maternal|Fetal, such as AA|AB for AA maternal and AB fetal alleles. Figure 30A shows a plot generated when two chromosomes are present and the fetal cfDNA fraction is 0%. This plot is from a non-pregnant woman and therefore represents a pattern when the genotype is entirely maternal. Thus, allele clusters are centered around 1 (AA allele), 0.5 (AB allele), and 0 (BB allele). Figure 30B shows a plot generated when two chromosomes are present and the fetal fraction is 12%. The contribution of fetal alleles to the proportion of A allele reads causes the positions of some allele spots to shift up and down along the y-axis. Therefore, bands are centered around 1 (AA|AA allele), 0.94 (AA|AB allele), 0.56 (AB|AA allele), 0.50 (AB|AB allele), 0.44 (AB|BB allele), 0.06 (BB|AB allele), and 0 (BB|BB allele). Figure 30C shows the plot generated when two chromosomes are present and the fetal fraction is 26%. A pattern is readily visible, including two red and two blue peripheral bands and a central green band of the triplet (colors not shown).The bands are centered at 1 (AA|AA alleles), 0.87 (AA|AB alleles), 0.63 (AB|AA alleles), 0.50 (AB|AB alleles), 0.37 (AB|BB alleles), 0.13 (BB|AB alleles), and 0 (BB|BB alleles). Figure 30D shows the plot generated when one chromosome is present and the fetal fraction is 26%. The hallmark pattern of one outer red and one outer blue peripheral band and two central green bands indicates maternally inherited monosomy (colors not shown). Because the fetus contributes only a single allele (A or B) to the allelic reads, the inner peripheral red and blue bands are absent, and the central triplet band is compressed to two bands (colors not shown). Bands are centered at 1 (AA|A allele), 0.57 (AB|A allele), 0.43 (AB|B allele), and 0 (BB|B allele). Figure 30E is a plot generated when three chromosomes are present and the fetal fraction is 27%. This pattern of two red and two blue peripheral bands and two central green bands indicates maternally inherited meiotic trisomy (colors not shown). Bands are centered at 1 (AA|AAA allele), 0.88 (AA|AAB allele), 0.56 (AB|AAB allele), 0.44 (AB|ABB allele), 0.12 (BB|ABB allele), and 0 (BB|BBB allele). Figure 30F is a plot generated when three chromosomes are present and the fetal fraction is 14%. This pattern of three red and three blue peripheral bands, as well as two central green bands, indicates a paternally inherited meiotic trisomy (colors not shown). The bands are centered at 1 (AA|AAA alleles), 0.93 (AA|AAB alleles), 0.87 (AA|ABB alleles), 0.60 (AB|AAA alleles), 0.53 (AB|AAB alleles), 0.47 (AB|ABB alleles), 0.40 (AB|BBB alleles), 0.13 (BB|AAB alleles), 0.07 (BB|ABB alleles), and 0 (BB|BBB alleles). Figure 30G shows a plot generated when three chromosomes are present and the fetal fraction is 35%.This pattern of two red and two blue peripheral bands and four green bands indicates a maternally inherited mitotic trisomy (colors not shown). The bands are centered at 1 (AA|AAA alleles), 0.85 (AA|AAB alleles), 0.72 (AB|AAA alleles), 0.57 (AB|AAB alleles), 0.43 (AB|ABB alleles), 0.28 (AB|BBB alleles), 0.15 (BB|ABB alleles), and 0 (BB|BBB alleles). Figure 30H shows a plot generated when three chromosomes are present and the fetal fraction is 25%. This pattern of two red and two blue peripheral bands and four central green bands indicates a paternally inherited mitotic trisomy (colors not shown). This pattern can be distinguished from a maternally inherited mitotic trisomy (as in Figure 30G) by the location of the inner peripheral bands. Specifically, the bands were centered at 1 (AA|AAA alleles), 0.78 (AA|ABB alleles), 0.67 (AB|AAA alleles), 0.56 (AB|AAB alleles), 0.44 (AB|ABB alleles), 0.33 (AB|BBB alleles), 0.22 (BB|AAB alleles), and 0 (BB|BBB alleles). [Figure 30-3]Figures 30A-30H: Representative graphical representations of euploidy (Figures 30A-30C), monosomy (Figure 30D), and trisomy (Figures 30E-30H). In all plots, the x-axis represents the linear position of the individual polymorphic locus along each chromosome (shown at the bottom of the plot), and the y-axis represents the number of A allele reads as a percentage of the total (A+B) allele reads. Maternal and fetal genotypes and the y-axis location of the band centers are shown to the right of the plots. For ease of visualization, plots may be color-coded according to maternal genotype, with red representing an AA maternal genotype, blue representing a BB maternal genotype, and green representing an AB maternal genotype. If desired, maternal allele contributions may be indicated by color in the "Fetal Genotype" column. Allelic contributions are expressed in the format Maternal|Fetal, such as AA|AB for AA maternal and AB fetal alleles. Figure 30A shows a plot generated when two chromosomes are present and the fetal cfDNA fraction is 0%. This plot is from a non-pregnant woman and therefore represents a pattern when the genotype is entirely maternal. Thus, allele clusters are centered around 1 (AA allele), 0.5 (AB allele), and 0 (BB allele). Figure 30B shows a plot generated when two chromosomes are present and the fetal fraction is 12%. The contribution of fetal alleles to the proportion of A allele reads causes the positions of some allele spots to shift up and down along the y-axis. Therefore, bands are centered around 1 (AA|AA allele), 0.94 (AA|AB allele), 0.56 (AB|AA allele), 0.50 (AB|AB allele), 0.44 (AB|BB allele), 0.06 (BB|AB allele), and 0 (BB|BB allele). Figure 30C shows the plot generated when two chromosomes are present and the fetal fraction is 26%. A pattern is readily visible, including two red and two blue peripheral bands and a central green band of the triplet (colors not shown).The bands are centered at 1 (AA|AA alleles), 0.87 (AA|AB alleles), 0.63 (AB|AA alleles), 0.50 (AB|AB alleles), 0.37 (AB|BB alleles), 0.13 (BB|AB alleles), and 0 (BB|BB alleles). Figure 30D shows the plot generated when one chromosome is present and the fetal fraction is 26%. The hallmark pattern of one outer red and one outer blue peripheral band and two central green bands indicates maternally inherited monosomy (colors not shown). Because the fetus contributes only a single allele (A or B) to the allelic reads, the inner peripheral red and blue bands are absent, and the central triplet band is compressed to two bands (colors not shown). Bands are centered at 1 (AA|A allele), 0.57 (AB|A allele), 0.43 (AB|B allele), and 0 (BB|B allele). Figure 30E is a plot generated when three chromosomes are present and the fetal fraction is 27%. This pattern of two red and two blue peripheral bands and two central green bands indicates maternally inherited meiotic trisomy (colors not shown). Bands are centered at 1 (AA|AAA allele), 0.88 (AA|AAB allele), 0.56 (AB|AAB allele), 0.44 (AB|ABB allele), 0.12 (BB|ABB allele), and 0 (BB|BBB allele). Figure 30F is a plot generated when three chromosomes are present and the fetal fraction is 14%. This pattern of three red and three blue peripheral bands, as well as two central green bands, indicates a paternally inherited meiotic trisomy (colors not shown). The bands are centered at 1 (AA|AAA alleles), 0.93 (AA|AAB alleles), 0.87 (AA|ABB alleles), 0.60 (AB|AAA alleles), 0.53 (AB|AAB alleles), 0.47 (AB|ABB alleles), 0.40 (AB|BBB alleles), 0.13 (BB|AAB alleles), 0.07 (BB|ABB alleles), and 0 (BB|BBB alleles). Figure 30G shows a plot generated when three chromosomes are present and the fetal fraction is 35%.This pattern of two red and two blue peripheral bands and four green bands indicates a maternally inherited mitotic trisomy (colors not shown). The bands are centered at 1 (AA|AAA alleles), 0.85 (AA|AAB alleles), 0.72 (AB|AAA alleles), 0.57 (AB|AAB alleles), 0.43 (AB|ABB alleles), 0.28 (AB|BBB alleles), 0.15 (BB|ABB alleles), and 0 (BB|BBB alleles). Figure 30H shows a plot generated when three chromosomes are present and the fetal fraction is 25%. This pattern of two red and two blue peripheral bands and four central green bands indicates a paternally inherited mitotic trisomy (colors not shown). This pattern can be distinguished from a maternally inherited mitotic trisomy (as in Figure 30G) by the location of the inner peripheral bands. Specifically, the bands were centered at 1 (AA|AAA alleles), 0.78 (AA|ABB alleles), 0.67 (AB|AAA alleles), 0.56 (AB|AAB alleles), 0.44 (AB|ABB alleles), 0.33 (AB|BBB alleles), 0.22 (BB|AAB alleles), and 0 (BB|BBB alleles). [Figure 30-4]Figures 30A-30H: Representative graphical representations of euploidy (Figures 30A-30C), monosomy (Figure 30D), and trisomy (Figures 30E-30H). In all plots, the x-axis represents the linear position of the individual polymorphic locus along each chromosome (shown at the bottom of the plot), and the y-axis represents the number of A allele reads as a percentage of the total (A+B) allele reads. Maternal and fetal genotypes and the y-axis location of the band centers are shown to the right of the plots. For ease of visualization, plots may be color-coded according to maternal genotype, with red representing an AA maternal genotype, blue representing a BB maternal genotype, and green representing an AB maternal genotype. If desired, maternal allele contributions may be indicated by color in the "Fetal Genotype" column. Allelic contributions are expressed in the format Maternal|Fetal, such as AA|AB for AA maternal and AB fetal alleles. Figure 30A shows a plot generated when two chromosomes are present and the fetal cfDNA fraction is 0%. This plot is from a non-pregnant woman and therefore represents a pattern when the genotype is entirely maternal. Thus, allele clusters are centered around 1 (AA allele), 0.5 (AB allele), and 0 (BB allele). Figure 30B shows a plot generated when two chromosomes are present and the fetal fraction is 12%. The contribution of fetal alleles to the proportion of A allele reads causes the positions of some allele spots to shift up and down along the y-axis. Therefore, bands are centered around 1 (AA|AA allele), 0.94 (AA|AB allele), 0.56 (AB|AA allele), 0.50 (AB|AB allele), 0.44 (AB|BB allele), 0.06 (BB|AB allele), and 0 (BB|BB allele). Figure 30C shows the plot generated when two chromosomes are present and the fetal fraction is 26%. A pattern is readily visible, including two red and two blue peripheral bands and a central green band of the triplet (colors not shown).The bands are centered at 1 (AA|AA alleles), 0.87 (AA|AB alleles), 0.63 (AB|AA alleles), 0.50 (AB|AB alleles), 0.37 (AB|BB alleles), 0.13 (BB|AB alleles), and 0 (BB|BB alleles). Figure 30D shows the plot generated when one chromosome is present and the fetal fraction is 26%. The hallmark pattern of one outer red and one outer blue peripheral band and two central green bands indicates maternally inherited monosomy (colors not shown). Because the fetus contributes only a single allele (A or B) to the allelic reads, the inner peripheral red and blue bands are absent, and the central triplet band is compressed to two bands (colors not shown). Bands are centered at 1 (AA|A allele), 0.57 (AB|A allele), 0.43 (AB|B allele), and 0 (BB|B allele). Figure 30E is a plot generated when three chromosomes are present and the fetal fraction is 27%. This pattern of two red and two blue peripheral bands and two central green bands indicates maternally inherited meiotic trisomy (colors not shown). Bands are centered at 1 (AA|AAA allele), 0.88 (AA|AAB allele), 0.56 (AB|AAB allele), 0.44 (AB|ABB allele), 0.12 (BB|ABB allele), and 0 (BB|BBB allele). Figure 30F is a plot generated when three chromosomes are present and the fetal fraction is 14%. This pattern of three red and three blue peripheral bands, as well as two central green bands, indicates a paternally inherited meiotic trisomy (colors not shown). The bands are centered at 1 (AA|AAA alleles), 0.93 (AA|AAB alleles), 0.87 (AA|ABB alleles), 0.60 (AB|AAA alleles), 0.53 (AB|AAB alleles), 0.47 (AB|ABB alleles), 0.40 (AB|BBB alleles), 0.13 (BB|AAB alleles), 0.07 (BB|ABB alleles), and 0 (BB|BBB alleles). Figure 30G shows a plot generated when three chromosomes are present and the fetal fraction is 35%.This pattern of two red and two blue peripheral bands and four green bands indicates a maternally inherited mitotic trisomy (colors not shown). The bands are centered at 1 (AA|AAA alleles), 0.85 (AA|AAB alleles), 0.72 (AB|AAA alleles), 0.57 (AB|AAB alleles), 0.43 (AB|ABB alleles), 0.28 (AB|BBB alleles), 0.15 (BB|ABB alleles), and 0 (BB|BBB alleles). Figure 30H shows a plot generated when three chromosomes are present and the fetal fraction is 25%. This pattern of two red and two blue peripheral bands and four central green bands indicates a paternally inherited mitotic trisomy (colors not shown). This pattern can be distinguished from a maternally inherited mitotic trisomy (as in Figure 30G) by the location of the inner peripheral bands. Specifically, the bands were centered at 1 (AA|AAA alleles), 0.78 (AA|ABB alleles), 0.67 (AB|AAA alleles), 0.56 (AB|AAB alleles), 0.44 (AB|ABB alleles), 0.33 (AB|BBB alleles), 0.22 (BB|AAB alleles), and 0 (BB|BBB alleles). [Figure 31-1] Figure 31: Graphical representation of euploid (Figure 31A), T13 (Figure 31B), T18 (Figure 31C), T21 (Figure 31D), 45,X (Figure 31E), and 47,XXY (Figure 31F) test samples. Each chromosome is shown at the top of the plot, and fetal and maternal genotypes are shown to the right of the plot. The x-axis represents the linear position of the SNP along each chromosome, and the y-axis represents the number of A allele reads as a percentage of total reads. Note that cluster positions have been changed based on fetal fraction, as described herein. Each spot represents a single SNP locus. Fetal and maternal genotypes are shown to the right of the plot, and chromosome identity is shown at the top of the plot. [Figure 31-2]Figure 31: Graphical representation of euploid (Figure 31A), T13 (Figure 31B), T18 (Figure 31C), T21 (Figure 31D), 45,X (Figure 31E), and 47,XXY (Figure 31F) test samples. Each chromosome is shown at the top of the plot, and fetal and maternal genotypes are shown to the right of the plot. The x-axis represents the linear position of the SNP along each chromosome, and the y-axis represents the number of A allele reads as a percentage of total reads. Note that cluster positions have been changed based on fetal fraction, as described herein. Each spot represents a single SNP locus. Fetal and maternal genotypes are shown to the right of the plot, and chromosome identity is shown at the top of the plot. [Figure 31-3] Figure 31: Graphical representation of euploid (Figure 31A), T13 (Figure 31B), T18 (Figure 31C), T21 (Figure 31D), 45,X (Figure 31E), and 47,XXY (Figure 31F) test samples. Each chromosome is shown at the top of the plot, and fetal and maternal genotypes are shown to the right of the plot. The x-axis represents the linear position of the SNP along each chromosome, and the y-axis represents the number of A allele reads as a percentage of total reads. Note that cluster positions have been changed based on fetal fraction, as described herein. Each spot represents a single SNP locus. Fetal and maternal genotypes are shown to the right of the plot, and chromosome identity is shown at the top of the plot. [Figure 31-4] Figure 31: Graphical representation of euploid (Figure 31A), T13 (Figure 31B), T18 (Figure 31C), T21 (Figure 31D), 45,X (Figure 31E), and 47,XXY (Figure 31F) test samples. Each chromosome is shown at the top of the plot, and fetal and maternal genotypes are shown to the right of the plot. The x-axis represents the linear position of the SNP along each chromosome, and the y-axis represents the number of A allele reads as a percentage of total reads. Note that cluster positions have been changed based on fetal fraction, as described herein. Each spot represents a single SNP locus. Fetal and maternal genotypes are shown to the right of the plot, and chromosome identity is shown at the top of the plot. [Figure 32]1 is a graph showing that the combined birth prevalence of sex chromosome aneuploidies is higher than that of autosomal aneuploidies. DETAILED DESCRIPTION OF THE INVENTION
[0146] While the above figures illustrate embodiments disclosed herein, other embodiments are contemplated, as noted in the discussion. This disclosure presents embodiments by way of example and not limitation. Those skilled in the art will be able to devise numerous other modifications and embodiments which fall within the scope and spirit of the principles of the embodiments disclosed herein.
[0147] The present invention is based, in part, on the unexpected discovery that a relatively small number of primers in a primer library are often responsible for the formation of a significant amount of amplification primer dimers during a multiplex PCR reaction. Methods were developed to select the most undesirable primers for removal from a candidate primer library. By reducing the amount of primer dimers to negligible amounts (approximately 0.1% of PCR product), these methods enable the resulting primer library to be used to simultaneously amplify multiple target loci in a single multiplex PCR reaction. Primers hybridize to target loci and amplify them without hybridizing to other primers to form amplified primer dimers, thereby increasing the number of different target loci that can be amplified. It was also discovered that using lower-than-normal primer concentrations and significantly longer annealing times increases the likelihood that primers will hybridize to target loci rather than hybridizing to each other to form primer dimers.
[0148] During PCR amplification and sequencing of 19,488 target loci in genomic samples, 99.4–99.7% of these sequencing reads were mapped to the genome, and 99.99% were mapped to the target loci. For plasma samples with 10 million sequencing reads, at least 19,350 (99.3%) of the 19,488 target loci were typically amplified and sequenced. The ability to simultaneously amplify such a large number of target loci at once significantly reduces the time and amount of DNA required to analyze thousands of target loci. For example, DNA from a single cell is sufficient to simultaneously analyze thousands of target loci, which is important for applications with low DNA quantities, such as genetic testing of single cells derived from embryos prior to in vitro fertilization or genetic testing of forensic samples containing small amounts of DNA. Furthermore, the ability to analyze target loci in a single reaction volume (e.g., a single vessel or well) without splitting the sample into multiple separate reactions reduces potential variability during the reaction. Furthermore, the method has been developed to use a reference standard to correct for possible amplification biases between different target loci. For example, differences in amplification efficiency between target loci due to factors such as GC content may result in different amounts of PCR product for target loci that should actually be produced in equal amounts. The use of a reference standard similar to the target loci allows such amplification biases to be detected and corrected for during quantification of the target loci.
[0149] During sequencing of PCR products, artifacts such as primer dimers can be detected, thereby interfering with the detection of target amplification products. Due to this limitation, microarrays with hybridization probes are often used for detection because of their low sensitivity to primer dimer interference. The high level of multiplexing currently achievable, with minimal non-target amplification products, allows PCR followed by sequencing to be used as an alternative to microarrays.
[0150] The multiplex PCR method of the present invention can be used for various purposes. For example, it can be used for genotyping analysis, detection of chromosomal abnormalities (e.g., fetal chromosomal aneuploidy), gene mutation and polymorphism (e.g., single nucleotide polymorphism, SNP) analysis, gene deletion analysis, paternity testing, analysis of genetic differences within populations, forensic analysis, disease predisposition measurement, mRNA quantitative analysis, and detection and identification of infectious pathogens (e.g., bacteria, parasites, and viruses). The multiplex PCR method can also be used for non-invasive prenatal genetic testing, for example, paternity testing or detection of fetal chromosomal abnormalities.
[0151] Representative primer design methods Highly multiplexed PCR can often produce a very high percentage of product DNA resulting from non-productive side reactions, such as primer dimer formation. In one embodiment, specific primers most likely to cause non-productive side reactions can be removed from a primer library to obtain a primer library that produces a high percentage of amplified DNA that maps to the genome. The step of removing problematic primers, i.e., primers that may particularly stabilize dimers, unexpectedly enabled very high PCR multiplexing levels for subsequent analysis by sequencing. In systems such as sequencing, where performance is significantly reduced by primer dimers and / or other adverse products, multiplexing levels greater than 10-fold, 50-fold, and even 100-fold higher than other described multiplexing levels have been achieved. This contrasts with probe-based detection methods, e.g., microarrays, TAQMAN, and PCR, where excess primer dimers do not appreciably affect results. It should also be noted that the general consensus in the art is that multiplexed PCR for sequencing is limited to approximately 100 assays in the same well. Fluidigm and Rain Dance provide platforms for performing 48 or 1000 PCR assays in parallel reactions on one sample.
[0152] There are several methods for selecting primers for libraries that minimize the amount of non-mapping primer dimers or other adverse primer products. Empirical data has shown that a small number of "bad" primers are responsible for a large number of non-mapping primer dimer side reactions. Removing these "bad" primers can increase the percentage of sequence reads that map to the target locus. One method for identifying "bad" primers is to examine the sequencing data of DNA amplified by targeted amplification, and remove the most frequently occurring primer dimers to produce a primer library that is significantly less likely to produce by-product DNA that does not map to the genome. Publicly available programs also exist that can calculate the binding energy of various primer combinations, and removing primer combinations with the highest binding energies similarly produces a primer library that is significantly less likely to produce by-product DNA that does not map to the genome.
[0153] In some embodiments for primer selection, an initial candidate primer library is generated by designing one or more primers or primer pairs for candidate target loci. The set of candidate target loci (e.g., SNPs) can be selected based on publicly available information regarding desired parameters of the target loci, such as the frequency of the SNP or the heterozygosity rate of the SNP within a target population. In one embodiment, PCR primers can be designed using the Primer3 program (available worldwide on the primer3.sourceforge.net web; libprimer3 release 2.2.3, incorporated herein by reference in its entirety). If desired, primers can be designed to anneal within a specific annealing temperature range, to have a specific range of GC content, to fall within a specific size range, to produce target amplicons within a specific size range, and / or to have other parameter characteristics. Starting with multiple primers or primer pairs per candidate target locus increases the likelihood that primers or primer pairs for most or all target loci will remain in the library. In one embodiment, the selection criterion may be that at least one primer pair per target locus must remain in the library. In this way, most or all of the target loci will be amplified using the final primer library. This is desirable for applications such as screening for deletions, or duplications at multiple sites in the genome, or screening for multiple sequences (e.g., polymorphisms or other mutations) associated with disease or increased risk of disease. If a primer pair from the library produces a target amplicon that overlaps with a target amplicon produced by another primer pair, one of the primer pairs can be removed from the library to prevent interference.
[0154] In some embodiments, an "undesirability score" (with higher scores indicating least desirability) is calculated (e.g., computationally) for most or all possible combinations of two primers from a library of candidate primers. In various embodiments, an undesirability score is calculated for at least 80, 90, 95, 98, 99, or 99.5% of the possible combinations of candidate primers in the library. Each undesirability score is based, at least in part, on the likelihood of dimer formation between the two candidate primers. Optionally, the undesirability score may also be based on one or more other parameters selected from the group consisting of the heterozygosity rate of the target locus, the prevalence of a disease associated with a sequence (e.g., a polymorphism) at the target locus, the disease penetrance associated with a sequence (e.g., a polymorphism) at the target locus, the specificity of the candidate primer for the target locus, the size of the candidate primer, the melting temperature of the target amplicon, the GC content of the target amplicon, the amplification efficiency of the target amplicon, and the size of the target amplicon. When multiple factors are considered, the undesirability score may be calculated based on a weighted average of various parameters. Parameters can be assigned different weights based on their importance to the particular application for which the primer is being used. In some embodiments, the primer with the highest undesirability score is removed from the library. If the removed primer is a member of a primer pair that hybridizes to one target locus, the other member of the primer pair can also be removed from the library. This primer removal process can be repeated as necessary. In some embodiments, the selection method is performed until the undesirability scores of all candidate primer combinations remaining in the library are below a minimum threshold. In some embodiments, the selection method is performed until the number of candidate primers remaining in the library is reduced to a desired number.
[0155] In various embodiments, after the undesirability scores are calculated, candidate primers that are part of the greatest number of combinations of two candidate primers with undesirability scores above a first minimum threshold are removed from the library. This step ignores interactions below the first minimum threshold because these interactions are less important. If the removed primer is a member of a primer pair that hybridizes to one target locus, the other member of the primer pair can be removed from the library. The primer removal process can be repeated as necessary. In some embodiments, the selection method is performed until all undesirability scores for the candidate primer combinations remaining in the library are below the first minimum threshold. If the number of candidate primers remaining in the library is greater than desired, the number of primers can be reduced by lowering the first minimum threshold to a second minimum threshold and repeating the primer removal process. If the number of candidate primers remaining in the library is less than desired, the method can continue by increasing the first minimum threshold to a higher second minimum threshold and repeating the primer removal process using the original candidate primer library, thereby leaving more candidate primers in the library. In some embodiments, the selection method is carried out until the undesirability scores of all complementary primer combinations remaining in the library are below a second minimum threshold, or until the number of candidate primers remaining in the library is reduced to a desired number.
[0156] If desired, primer pairs that produce target amplicons that overlap with those produced by other primer pairs can be split into separate amplification reactions. Multiplex PCR amplification reactions may be preferred for applications where it is desirable to analyze all candidate target loci (rather than excluding candidate target loci from analysis due to overlapping target amplicons).
[0157] These selection methods minimize the number of candidate primers that need to be removed from the library to achieve the desired reduction in primer dimers. By removing fewer candidate primers from the library, the resulting primer library can be used to amplify more (or all) of the target loci.
[0158] Multiplexing a large number of primers places significant constraints on the assays that can be included. Assays that unintentionally interact will result in spurious amplification products. The size constraints of miniPCR can pose further constraints. In one embodiment, it is possible to start with a very large number of potential SNP targets (between approximately 500 and over 1 million) and attempt to design primers to amplify each SNP. If primers can be designed, it is possible to attempt to identify primer pairs that are likely to form spurious products by evaluating the likelihood of spurious primer duplex formation between all possible primer pairs using published thermodynamic parameters for DNA duplex formation. Primer interactions can be ranked by a scoring function related to the interaction, and primers with the worst interaction scores can be eliminated until the desired number of primers is met. If potentially heterozygous SNPs are most useful, it is possible to similarly rank the list of assays and select the assay that best matches heterozygosity. Experiments have verified that primers with high interaction scores are most likely to form primer dimers. While it is impossible to eliminate all spurious interactions at high multiplexing levels, it is essential to remove primers or primer pairs with the highest in silico interaction scores, as they dominate the overall reaction and significantly limit amplification from the intended target. This procedure has been implemented to generate multiplexed primer sets up to, and in some cases exceed, 10,000 primers. The improvement achieved by this procedure is substantial, allowing for amplification of over 80%, 90%, 95%, 98%, and even 99% of target products, as determined by sequencing all PCR products, compared to 10% from reactions in which the worst primers were not removed. When combined with the previously described partial semi-nested approach, over 90% and even 95% of the amplification products can be mapped to the target sequence.
[0159] It should be noted that there are other methods for determining which PCR probes may form dimers. In some embodiments, analysis of a pool of DNA amplified using a set of non-optimized primers may be sufficient to determine the problematic primers. For example, analysis can be performed using sequencing, and the dimers present in the greatest number can be determined to be the ones most likely to form dimers and can be removed.
[0160] This method has several potential applications, such as SNP genotyping, heterozygosity rate determination, copy number measurement, and other targeted sequencing applications. In certain embodiments, the primer design method can be used in combination with the mini-PCR method described elsewhere in this document. In some embodiments, the primer design method can be used as part of a large-scale multiplex PCR method.
[0161] The use of tags in primers can reduce amplification and sequencing of primer dimer products. In some embodiments, a primer includes an internal region that forms a loop structure with the tag. In certain embodiments, a primer includes a 5' region specific to a target locus, an internal region that forms a loop structure that is not specific to the target locus, and a 3' region that is specific to the target locus. In some embodiments, the loop region may be located between two binding regions that are designed to bind to adjacent or neighboring regions of the template DNA. In various embodiments, the 3' region is at least 7 nucleotides in length. In some embodiments, the 3' region is 7 to 20 nucleotides in length, e.g., 7 to 15 nucleotides, or 7 to 10 nucleotides in length. In various embodiments, a primer includes a 5' region that is not specific to a target locus (e.g., a tag or universal primer binding site), followed by a region specific to the target locus, an internal region that forms a loop structure that is not specific to the target locus, and a 3' region that is specific to the target locus. Using tag-primers, the required target-specific sequence can be shortened to less than 20 base pairs, less than 15 base pairs, less than 12 base pairs, or even less than 10 base pairs. This can be discovered by chance when designing standard primers, when the target sequence is fragmented within the primer binding site, or it can be designed into primer design. The advantages of this method include increasing the number of assays that can be designed for a specific maximum amplification product length, and reducing the sequencing of "uninformative" primer sequences. This method can also be used in combination with internal tagging (see elsewhere in this document).
[0162] In some embodiments, the relative amount of non-productive products in multiplex target PCR amplification can be reduced by increasing the annealing temperature. When amplifying libraries with the same tag as the target-specific primers, the annealing temperature can be increased compared to genomic DNA because the tag contributes to primer binding. In some embodiments, significantly lower primer concentrations than previously reported are used, along with longer annealing times than reported elsewhere. In some embodiments, annealing times can be greater than 3 minutes, greater than 5 minutes, greater than 8 minutes, greater than 10 minutes, greater than 15 minutes, greater than 20 minutes, greater than 30 minutes, greater than 60 minutes, greater than 120 minutes, greater than 240 minutes, greater than 480 minutes, and even greater than 960 minutes. In some embodiments, longer annealing times than previously reported are used, allowing for lower primer concentrations. In various embodiments, longer extension times than usual, for example, greater than 3, 5, 8, 10, or 15 minutes, are used. In some embodiments, the primer concentration is as low as 50 nM, 20 nM, 10 nM, 5 nM, 1 nM, and even less than 1 μM. Surprisingly, this results in robust performance for highly multiplexed reactions, such as 1,000-plex reactions, 2,000-plex reactions, 5,000-plex reactions, 10,000-plex reactions, 20,000-plex reactions, 50,000-plex reactions, and even 100,000-plex reactions. In some embodiments, amplification uses 1, 2, 3, 4, or 5 cycles with long annealing times, followed by PCR cycles with tagged primers and more than the usual number of annealing times.
[0163] To select target locations, one can start with a pool of candidate primer pair designs, create a thermodynamic model of potentially deleterious interactions between the primer pairs, and then use this model to eliminate designs that are incompatible with other designs in the pool.
[0164] After the selection process, the primers remaining in the library can be used in any of the methods of the invention.
[0165] Representative primer library In one aspect, the invention features a library of primers, e.g., primers selected using any of the methods of the invention from a library of candidate primers. In some embodiments, the library includes primers that simultaneously hybridize to (or are capable of simultaneously hybridizing to) or simultaneously amplify (or are capable of simultaneously amplifying) at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci in a single reaction volume. In various embodiments, the library comprises primers that simultaneously amplify (or are capable of simultaneously amplifying) between 1,000 and 2,000; between 2,000 and 5,000; between 5,000 and 7,500; between 7,500 and 10,000; between 10,000 and 20,000; between 20,000 and 25,000; between 25,000 and 30,000; between 30,000 and 40,000; between 40,000 and 50,000; between 50,000 and 75,000; between 75,000 and 100,000 different target loci in a single reaction volume. In various embodiments, the library comprises primers that simultaneously amplify (or are capable of simultaneously amplifying) 1,000 to 100,000 different target loci in one reaction volume, e.g., 1,000 to 50,000; 1,000 to 30,000; 1,000 to 20,000; 1,000 to 10,000; 2,000 to 30,000; 2,000 to 20,000; 2,000 to 10,000; 5,000 to 30,000; 5,000 to 20,000; or 5,000 to 10,000 different target loci. In some embodiments, the library includes primers that simultaneously amplify (or are capable of simultaneously amplifying) target loci in a single reaction volume, such that less than 60, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.5, 0.25, 0.1, or 0.5% of the amplification products are primer dimers. In various embodiments, the amount of primer dimer amplification product is between 0.5 and 60%, e.g., between 0.1 and 40%, 0.1 and 20%, 0.25 and 20%, 0.25 and 10%, 0.5 and 20%, 0.5 and 10%, 1 and 20%, or 1 and 10%.In some embodiments, the primers simultaneously amplify (or are capable of simultaneously amplifying) target loci in a single reaction volume, such that at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the amplified products are target amplicons. In various embodiments, the amount of amplified product that is target amplicons is between 50 and 99.5%, e.g., between 60 and 99%, 70 and 98%, 80 and 98%, 90 and 99.5%, or 95 and 99.5%. In some embodiments, the primers simultaneously amplify (or are capable of simultaneously amplifying) target loci in a single reaction volume, such that at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci are amplified. In various embodiments, the amount of amplified target loci is 50-99.5%, e.g., 60-99%, 70-98%, 80-99%, 90-99.5%, 95-99.9%, or 98-99.99%. In some embodiments, the primer library includes at least 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 primer pairs, each pair of primers including a forward test primer and a reverse test primer, and each test primer pair hybridizes to a target locus. In some embodiments, the primer library comprises at least 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 individual primers that each hybridize to a different target locus, and the individual primers are not part of a primer pair.
[0166] In various embodiments, the concentration of each primer is less than 100, 75, 50, 25, 20, 10, 5, 2, or 1 nM, or less than 500, 100, 10, or 1 uM. In various embodiments, the concentration of each primer is between 1 uM and 100 nM, e.g., between 1 uM and 1 nM, 1 to 75 nM, 2 to 50 nM, or 5 to 50 nM. In various embodiments, the GC content of the primers is between 30 and 80%, e.g., between 40 and 70%, or between 50 and 60%. In some embodiments, the GC content of the primers is less than 30, 20, 10, or 5%. In some embodiments, the GC content of the primers is between 5 and 30%, e.g., between 5 and 20% or between 5 and 10%. In some embodiments, the melting temperature (T m ) is 40 to 80°C, for example, 50 to 70°C, 55 to 65°C, or 57 to 60.5°C. In some embodiments, T mis calculated with the primer3 program (libprimer3 release 2.2.3) using the built-in SantaLucia parameters (available worldwide at primer3.sourceforge.net). In some embodiments, the melting temperature range of the primers is less than 15, 10, 5, 3, or 1°C. In some embodiments, the melting temperature range of the primers is 1 to 15°C, e.g., 1 to 10°C, 1 to 5°C, or 1 to 3°C. In some embodiments, the length of the primers is 15 to 100 nucleotides, e.g., 15 to 75 nucleotides, 15 to 40 nucleotides, 17 to 35 nucleotides, 18 to 30 nucleotides, or 20 to 65 nucleotides. In some embodiments, the length of the primers is less than 50, 40, 30, 20, 10, or 5 nucleotides. In some embodiments, the length of the primers is 5 to 50 nucleotides, e.g., 5 to 40 nucleotides, 5 to 20 nucleotides, or 5 to 10 nucleotides. In some embodiments, the length of the target amplicon is 50-100 nucleotides, e.g., 60-80 nucleotides, or 60-75 nucleotides. In some embodiments, the length range of the target amplicon is less than 50, 25, 15, 10, or 5 nucleotides. In some embodiments, the length range of the target amplicon is 5-50 nucleotides, e.g., 5-25 nucleotides, 5-15 nucleotides, or 5-10 nucleotides.
[0167] These primer libraries can be used in any of the methods of the present invention.
[0168] Representative primer kits In one aspect, the present invention features a kit (e.g., a kit for amplifying a target locus in a nucleic acid sample) containing any of the primer libraries of the present invention. In some embodiments, a kit can be formulated that includes multiple primers designed to perform the methods described herein. The primers can be outer forward and reverse primers, inner forward and reverse primers, as disclosed herein, or primers designed to have low binding affinity to other primers in the kit as disclosed in the primer design section, or hybrid capture probes or pre-circularization probes, or some combination thereof, as described in the relevant sections. In one embodiment, a kit for determining the ploidy state of a target chromosome in a gestating fetus, designed for use in the methods disclosed herein, can be constructed that includes multiple inner forward primers, optionally multiple inner reverse primers, and optionally outer forward and outer reverse primers, each designed to hybridize to a region of DNA immediately upstream and / or downstream of one of the target sites (e.g., polymorphic sites) on the target chromosome and, optionally, another chromosome. In one embodiment, the primer kit can be used in combination with a diagnostic box, as described elsewhere in this document. In some embodiments, the kit includes instructions for using the library to amplify target loci.
[0169] Typical multiplex PCR method In one aspect, the invention features a method for amplifying target loci in a nucleic acid sample, the method including: (i) contacting the nucleic acid sample with a library of primers that simultaneously hybridize to at least 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different target loci to generate a reaction mixture; and (ii) subjecting the reaction mixture to primer extension reaction conditions (e.g., PCR conditions) to generate amplification products that include target amplicons. In some embodiments, the method also includes determining the presence or absence of at least one target amplicon (e.g., at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target amplicons). In some embodiments, the method also includes determining the sequence of at least one target amplicon (e.g., at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target amplicons). In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci are amplified. In various embodiments, less than 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.5, 0.25, 0.1, or 0.05% of the amplicons are primer dimers.
[0170] In an embodiment, the method disclosed herein uses highly efficient, highly multiplexed target PCR to amplify DNA, followed by high-throughput sequencing to determine the allele frequency at each target locus.It is novel and non-obvious that more than about 50 or 100 PCR primers can be multiplexed in one reaction volume, so that the majority of the resulting sequence reads map to the target locus.One technique that enables highly multiplexed target PCR to be performed in a highly efficient manner involves designing primers that are unlikely to hybridize with each other. PCR probes, commonly referred to as primers, are selected by creating a thermodynamic model of potentially harmful interactions between at least 500, at least 1,000, at least 2,000, at least 5,000, at least 7,500, at least 10,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000 potential primer pairs, or unintended interactions between primers and sample DNA, and then using this model to eliminate designs that are incompatible with other designs in the pool. Another technique that allows highly multiplexed target PCR to be performed in a very efficient manner is to perform target PCR using a partial or full nesting approach. Using one or a combination of these techniques, it is possible to multiplex at least 300, at least 800, at least 1,200, at least 4,000, or at least 10,000 primers in a single pool, with the resulting amplified DNA containing a majority of DNA molecules that, when sequenced, map to the targeted locus. Using one or a combination of these techniques, it is possible to multiplex a large number of primers in a single pool, with the resulting amplified DNA containing more than 50%, more than 60%, more than 67%, more than 80%, more than 90%, more than 95%, more than 96%, more than 97%, more than 98%, more than 99%, or more than 99.5% of the DNA molecules that map to the targeted locus.
[0171] In some embodiments, detection of target genetic material can be performed in a multiplexed manner. The number of genetic target sequences that can be run in parallel can range from 1 to 10, 10 to 100, 100 to 1,000, 1,000 to 10,000, 10,000 to 100,000, 100,000 to 1,000,000, or 1,000,000 to 10,000,000. Previous attempts to multiplex more than 100 primers per pool have resulted in significant problems with undesirable side reactions, such as primer dimerization.
[0172] Targeted PCR In some embodiments, PCR can be used to target specific locations in the genome. In plasma samples, the original DNA is highly fragmented (generally less than 500 bp, with an average length of less than 200 bp). In PCR, both the forward and reverse primers anneal to the same fragment to enable amplification. Therefore, if the fragment is short, the PCR assay must also amplify a relatively short region. As in MIPS, if the polymorphism location is too close to the polymerase binding site, amplification from different alleles will be biased. Currently, PCR primers targeting polymorphic regions, such as those containing SNPs, are generally designed so that the 3' end of the primer hybridizes to the base immediately adjacent to one or more polymorphic bases. In certain embodiments of the present disclosure, the 3' ends of both the forward and reverse PCR primers are designed to hybridize to a base that is one or a few positions away from the mutation location (polymorphic site) of the targeted allele. The number of bases between the polymorphic site (SNP or other type) and the base to which the 3' end of the primer is designed to hybridize may be 1, 2, 3, 4, 5, 6, 7 to 10, 11 to 15, or 16 to 20. The forward primer and reverse primer can be designed to hybridize with different numbers of bases away from the polymorphic site.
[0173] While it is possible to generate a large number of PCR assays, interactions between different PCR assays make it difficult to multiplex them beyond about 100 assays. While various complex molecular techniques can be used to increase the level of multiplexing, it may still be limited to less than 100, perhaps 200, or even 500 assays per reaction. Samples with large amounts of DNA can be split into multiple subreactions and then recombined before sequencing. For samples in which either the overall sample or a subset of DNA molecules is limited, splitting the sample will introduce statistical noise. In certain embodiments, a low or limited amount of DNA may refer to an amount less than 10 pg, between 10 pg and 100 pg, between 100 pg and 1 ng, between 1 ng and 10 ng, or between 10 ng and 100 ng. While this method is particularly useful for small amounts of DNA, where other methods involving splitting into multiple pools can pose significant problems related to the introduction of stochastic noise, this method offers the benefit of minimizing bias when performed on samples of any amount of DNA. In these situations, a universal pre-amplification step can be used to increase the overall sample volume, which ideally should not appreciably alter the allele distribution.
[0174] In certain embodiments, the disclosed methods enable the generation of PCR products specific to a large number of target loci, specifically 1,000-5,000 loci, 5,000-10,000 loci, or more than 10,000 loci, from a limited sample, such as a single cell or DNA from a bodily fluid, for genotyping by sequencing or some other genotyping method. Currently, performing multiplex PCR reactions for more than 5-10 targets presents significant challenges, often hindered by primer by-products, such as primer dimers and other artifacts. When detecting target sequences using microarrays with hybridization probes, primer dimers and other artifacts are undetectable and can be ignored. However, when using sequencing as a detection method, the majority of sequencing reads sequence such artifacts and not the desired target sequence in the sample. Methods described in the prior art used to multiplex greater than 50 or 100 reactions in one reaction volume followed by sequencing typically result in greater than 20%, and often greater than 50%, often greater than 80%, and in some cases greater than 90% off-target sequence reads.
[0175] Typically, to perform targeted sequencing on a large number (n) of targets (>50, >100, >500, or >1,000) in a sample, the sample can be split into several parallel reactions, amplifying a single individual target. This can be performed in PCR multi-well plates or on commercial platforms, such as the Fluidigm Access Array (48 reactions per sample in a microfluidic chip) or Droplet PCR from Rain Dance Technology (100s to thousands of targets). Unfortunately, these split-and-pool methods are problematic for samples with limited amounts of DNA, as there are often insufficient copies of the genome to ensure one copy of each region of the genome is present in each well. This is particularly significant when targeting polymorphic loci, where the relative proportions of alleles at the polymorphic locus are required, as the stochastic noise introduced by splitting and pooling results in a poorly accurate measurement of the proportions of alleles present in the original DNA sample. Described herein are methods for effective and efficient amplification of many PCR reactions that are applicable when only limited amounts of DNA are available. In certain embodiments, the methods can be applied to analyze single cells, body fluids, mixtures of DNA, such as free-floating DNA found in maternal plasma, biopsies, environmental samples, and / or forensic samples.
[0176] In some embodiments, targeted sequencing may include one, several, or all of the following steps: a) generating and amplifying libraries with adapter sequences at both ends of DNA fragments; b) library amplification followed by partitioning into multiple reactions; c) generating and optionally amplifying libraries with adapter sequences at both ends of DNA fragments; d) performing 1000-10,000-plex amplification of selected targets using one target-specific "forward" primer and one tag-specific primer per target; e) performing a second amplification of this product using a "reverse" target-specific primer and one (or more) primers specific to the universal tag introduced as part of the target-specific forward primer in the first round; f) performing 1000-plex preamplification of selected targets for a limited number of cycles; g) dividing the product into multiple aliquots and amplifying subpools of targets in separate reactions (e.g., 50-500-plex (this method can be used with any plex up to single plex)); h) pooling the products of the parallel subpool reactions. i) During these amplifications, the primers can carry sequencing-compatible tags (partial or full length), allowing the products to be sequenced.
[0177] Advanced multiplex PCR Disclosed herein are methods that enable targeted amplification of hundreds to tens of thousands of target sequences (e.g., SNP loci) from nucleic acid samples, such as genomic DNA obtained from plasma. The amplified samples are relatively free of primer-dimer products and exhibit low allelic bias at the target loci. Analysis of these products can be performed by sequencing if sequencing-compatible adapters are added to the products during or after amplification.
[0178] By performing highly multiplexed PCR amplification using methods known in the art, the desired amplification products are in excess, and primer dimer products that are not suitable for sequencing are generated. These can be reduced by empirically eliminating the primers that form these products or by performing in silico primer selection. However, the greater the number of assays, the more difficult this problem becomes.
[0179] One solution is to split the 5,000-plex reaction into several lower-plex amplifications, e.g., 100 50-plex reactions or 50 100-plex reactions, or to use microfluidics, or even to split the sample into individual PCR reactions. However, when sample DNA is limited, such as in non-invasive prenatal testing from pregnancy plasma, splitting the sample among multiple reactions should be avoided because this creates bottlenecks.
[0180] Described herein is a method for first globally amplifying the plasma DNA of a sample and then dividing the sample into multiple multiplexed target enrichment reactions with a more moderate number of target sequences per reaction. In an embodiment, the disclosed method can be used to preferentially enrich a DNA mixture at multiple loci, and the method includes one or more of the following steps: generating and amplifying a library from the DNA mixture, where molecules in the library have adapter sequences ligated to both ends of the DNA fragments; dividing the amplified library into multiple reactions; and performing a first round of multiplex amplification of selected targets using one target-specific "forward" primer per target and one or more adapter-specific universal "reverse" primers. In an embodiment, the disclosed method further includes performing a second amplification using a "reverse" target-specific primer and one or more primers specific to the universal tag introduced as part of the target-specific forward primer in the first round. In some embodiments, the method may involve a fully nested PCR approach, a hemi-nested PCR approach, a semi-nested PCR approach, a one-sided fully nested PCR approach, a one-sided hemi-nested PCR approach, or a one-sided semi-nested PCR approach. In some embodiments, the disclosed method is used to preferentially enrich a DNA mixture at multiple loci, comprising multiplexed pre-amplification of selected targets for a limited number of cycles, dividing the product into multiple aliquots, amplifying subpools of targets in individual reactions, and pooling the products of the parallel subpool reactions. Note that this approach can be used to perform targeted amplification with low levels of allelic bias for 50-500 loci, 500-5,000 loci, 5,000-50,000 loci, or even 50,000-500,000 loci. In some embodiments, the primers carry partial-length or full-length sequencing-compatible tags.
[0181] A workflow may involve (1) extracting DNA, such as plasma DNA, (2) preparing a fragment library with universal adapters at both ends of the fragments, (3) amplifying the library using adapter-specific universal primers, (4) dividing the amplified sample "library" into multiple aliquots, (5) performing multiplexed amplification (e.g., about 100-plex, 1,000-plex, or 10,000-plex using one target-specific primer per target and tag-specific primers) on the aliquots, (6) pooling the aliquots of one sample, (7) barcoding the sample, (8) mixing the sample and adjusting the concentration, and (9) sequencing the sample. A workflow may include multiple substeps containing one of the listed steps (e.g., the library preparation step in step (2) may involve three enzymatic steps (blunting, dA tailing, and adapter ligation) and three purification steps). Workflow steps can be combined, separated, or performed in a different order (e.g., barcoding and pooling samples).
[0182] It is important to note that library amplification can be biased to amplify shorter fragments more efficiently. In this way, it is possible to preferentially amplify shorter sequences, such as mononucleosomal DNA fragments, as cell-free fetal DNA (of placental origin) found in the circulation of pregnant women. It is important to note that PCR assays may contain tags, such as sequencing tags (usually 15-25 base truncated forms). After multiplexing, the PCR multiplexed products of the samples are pooled, and then tagging (including barcoding) is completed by tag-specific PCR (which can also be performed by ligation). Alternatively, the complete sequencing tag can be added to the same reaction as multiplexing. In the first cycle, the target can be amplified using target-specific primers, followed by the tag-specific primers dominating to complete the complete SQ adapter sequence. The PCR primers do not need to carry tags. Sequencing tags can be added to the amplification products by ligation.
[0183] In some embodiments, highly multiplexed PCR can be used to evaluate the amplified material, followed by clonal sequencing, for various applications, such as detecting fetal aneuploidy. Conventional multiplexed PCR simultaneously evaluates up to 50 loci, but the techniques described herein can be used to simultaneously evaluate more than 50 loci, more than 100 loci, more than 500 loci, more than 1,000 loci, more than 5,000 loci, more than 10,000 loci, more than 50,000 loci, and more than 100,000 loci. Experiments have shown that up to, including, and more than 10,000 distinct loci can be simultaneously evaluated in a single reaction with sufficient efficiency and specificity to enable non-invasive prenatal aneuploidy diagnosis and / or copy number calling with high accuracy. The assay can be combined in a single reaction with the whole sample, such as a cfDNA sample isolated from maternal plasma, a fraction thereof, or a further processed derivative of the cfDNA sample. The sample (e.g., cfDNA or derivatives) can also be split into multiple parallel multiplex reactions. Optimal sample splitting and multiplexing is determined by trading off various performance specifications. Due to limited material quantities, splitting the sample into multiple fractions can result in increased sampling noise, handling time, and potential for error. Conversely, high multiplexing can result in increased amounts of spurious amplification and increased amplification inequality, both of which reduce test performance.
[0184] Two crucial considerations in applying the methods described herein are the limited amount of original sample (e.g., plasma) and the number of original molecules within the material from which allele frequencies or other measurements are obtained. Below a certain level, random sampling noise can become significant and affect the accuracy of the test. Generally, performing measurements on samples containing the equivalent of 500–1000 original molecules per target locus can yield data of sufficient quality for noninvasive prenatal aneuploidy diagnosis. Several methods exist for increasing the number of distinct measurements, e.g., increasing the sample volume. Each manipulation applied to the sample also potentially results in material loss. To avoid losses that could compromise test performance, it is essential to characterize the losses incurred by various manipulations and avoid or, if necessary, improve the yield of certain manipulations.
[0185] In some embodiments, amplifying all or a percentage of the original sample (e.g., a cfDNA sample) can reduce potential losses in subsequent steps. Various methods are available for amplifying all of the genetic material in a sample, thereby increasing the amount available for downstream procedures. In some embodiments, ligation-mediated PCR (LM-PCR) amplifies DNA fragments by PCR after ligating one, two, or many separate adapters. In some embodiments, multiple displacement amplification (MDA) phi-29 polymerase is used to amplify all DNA isothermally. In DOP-PCR and its variants, random priming is used to amplify the original DNA material. Each method has specific characteristics, such as uniformity of amplification across all representative genomic regions, efficiency of capturing and amplifying the original DNA, and amplification performance depending on fragment length.
[0186] In some embodiments, LM-PCR can be performed with a single heteroduplex adapter containing a 3' tyrosine. The heteroduplex adapter allows for the use of a single adapter molecule that can be converted into two distinct sequences on the 5' and 3' ends of the original DNA fragment during the first round of PCR. In some embodiments, it is possible to fractionate the amplified library by size separation, or the products using methods such as AMPURE, TASS, or other similar methods. Prior to ligation, the sample DNA is blunt-ended and then a single adenosine base is added to the 3' end. Prior to ligation, the DNA can be cleaved using a restriction enzyme or some other cleavage method. During ligation, the 3' adenosine of the sample fragment and the complementary 3' tyrosine overhang of the adapter can enhance ligation efficiency. The extension step of PCR amplification can be time-limited, reducing amplification from fragments longer than about 200 bp, 300 bp, 400 bp, 500 bp, or 1,000 bp. Because the longer DNA found in maternal plasma is almost exclusively maternal, this could result in a 10-50% enrichment of fetal DNA and improved test performance. Several reactions performed using conditions specified by the commercially available kit resulted in successful ligation of less than 10% of the sample DNA molecules. Further optimization of the reaction conditions improved ligation to approximately 70%.
[0187] Mini-PCR The mini-PCR method described below is suitable for samples containing short, digested, or fragmented nucleic acids, such as cfDNA. While traditional PCR assay design results in significant loss of characteristic fetal molecules, this loss can be significantly reduced by designing a very short PCR assay, referred to as a mini-PCR assay. Fetal cfDNA in maternal serum is highly fragmented, with fragment sizes distributed in a roughly Gaussian fashion, with a mean of 160 bp, a standard deviation of 15 bp, a minimum size of approximately 100 bp, and a maximum size of approximately 220 bp. The distribution of fragment start and end positions for target polymorphisms is not necessarily random but varies widely between individual targets and among all targets collectively, and polymorphic sites at a particular target locus can occupy any position, from start to end, among the various fragments originating from that locus. Note that the term mini-PCR can equally well refer to conventional PCR without further restriction or limitation.
[0188] During PCR, amplification occurs only from template DNA fragments that contain both forward and reverse primer sites.Because fetal cfDNA fragments are short, the likelihood of both primer sites being present, and the likelihood of a fetal fragment of length L containing both forward and reverse primer sites, is the ratio of the length of the amplification product to the length of the fragment.Under ideal conditions, assays in which the amplification product is 45bp, 50bp, 55bp, 60bp, 65bp or 70bp will successfully amplify 72%, 69%, 66%, 63%, 59% or 56% of the available template fragment molecules, respectively.The length of the amplification product is the distance between the 5' ends of the forward priming site and the reverse priming site.Amplification products of shorter length than those commonly used by those skilled in the art can result in more efficient measurement of desired polymorphic loci by only requiring short sequence reads. In certain embodiments, a substantial fraction of the amplification products should be less than 100 bp, less than 90 bp, less than 80 bp, less than 70 bp, less than 65 bp, less than 60 bp, less than 55 bp, less than 50 bp, or less than 45 bp.
[0189] It should be noted that in methods known in the prior art, short assays such as those described herein are not necessary and are usually avoided because they impose significant constraints on primer design due to limited primer length, annealing properties, and distance between forward and reverse primers.
[0190] It should also be noted that amplification bias potentially exists if the 3' end of either primer is within approximately 1 to 6 bases of the polymorphic site. This single-base difference in the initial polymerase binding site can result in preferential amplification of one allele, which can alter the observed allele frequency and reduce performance. All of these constraints make it very difficult to identify primers that successfully amplify specific loci and, in addition, to design a large collection of primers that are compatible with the same multiplex reaction. In one embodiment, the 3' ends of the inner forward and reverse primers are designed to hybridize to a region of DNA upstream of the polymorphic site and separated from it by a small number of bases. Ideally, the number of bases can be between 6 and 10 bases, but it can equally well be between 4 and 15 bases, 3 and 20 bases, 2 and 30 bases, or 1 and 60 bases, and essentially the same results can be achieved.
[0191] Multiplex PCR may involve a single round of PCR in which all targets are amplified, or it may involve one round of PCR followed by one or more rounds of nested PCR or some variants of nested PCR.Nested PCR consists of one or more rounds of PCR amplification using one or more new primers that bind at least one base pair closer to the interior than the primer used in the previous round.Nested PCR reduces the number of false amplification targets by amplifying only the amplification products from the previous reaction that have the correct internal sequence in the subsequent reaction.Reducing false amplification targets improves the number of useful measurements that can be obtained, especially in sequencing.Nested PCR generally involves designing primers that are completely internal to the previous primer binding site, which necessarily increases the minimum DNA segment size required for amplification.For samples such as maternal plasma cfDNA, where DNA is highly fragmented, a larger assay size reduces the number of distinct cfDNA molecules that can be measured. In one embodiment, to counteract this effect, a partial nesting approach can be used in which one or both of the second round primers overlap the first binding site extending into it by several bases, achieving additional specificity while minimizing the increase in overall assay size.
[0192] In some embodiments, multiplexed pools of PCR assays are designed to amplify SNPs or other polymorphic or non-polymorphic loci on one or more chromosomes that are potentially heterozygous, and these assays are used in a single reaction to amplify DNA. The number of PCR assays can be between 50 and 200 PCR assays, between 200 and 1,000 PCR assays, between 1,000 and 5,000 PCR assays, or between 5,000 and 20,000 PCR assays (50-200plex, 200-1,000plex, 1,000-5,000plex, 5,000-20,000plex, and over 20,000plex, respectively). In one embodiment, a multiplex pool of approximately 10,000 PCR assays (10,000-plex) is designed to amplify potentially heterozygous SNP loci on chromosomes X, Y, 13, 18, and 21, and 1 or 2. These assays are used in a single reaction to amplify cfDNA obtained from plasma samples, chorionic villus samples, amniocentesis samples, single or small numbers of cells, other body fluids or tissues, cancer, or other genetic material. The SNP frequency of each locus can be determined by clonal sequencing of the amplified product or by some other method. Statistical analysis of allele frequency distribution or the ratio of all assays can be used to determine whether the sample contains one or more trisomies of the chromosomes included in the test. In another embodiment, the original cfDNA sample is divided into two samples, and parallel 5,000-plex assays are performed. In another embodiment, the original cfDNA sample is divided into n samples, and parallel (approximately 10,000 / n)plex assays are performed, where n is between 2 and 12, or between 12 and 24, or between 24 and 48, or between 48 and 96. Data is collected and analyzed as previously described. Note that this method is equally applicable to detecting translocations, deletions, duplications, and other chromosomal abnormalities.
[0193] In some embodiments, tails that have no homology to the target genome can be added to the 3' or 5' end of any of the primers. These tails facilitate subsequent manipulations, procedures, or measurements. In some embodiments, the tail sequence can be the same for the target-specific forward primer and the target-specific reverse primer. In some embodiments, different tails can be used for the target-specific forward primer and the target-specific reverse primer. In some embodiments, multiple different tails can be used for different loci or sets of loci. Specific tails can be shared among all loci or among subsets of loci. For example, forward and reverse tails corresponding to the forward and reverse sequences required for any current sequencing platform can be used to enable direct sequencing after amplification. In some embodiments, the tails can be used as a general priming site among all amplified targets that can be used to add other useful sequences. In some embodiments, the inner primers can contain regions designed to hybridize either upstream or downstream of the target locus (e.g., polymorphic locus). In some embodiments, the primers can contain molecular barcodes. In some embodiments, the primers may contain universal priming sequences designed to allow for PCR amplification.
[0194] In one embodiment, a 10,000-plex PCR assay pool is generated in which the forward and reverse primers have tails corresponding to the required forward and reverse sequences required for a high-throughput sequencing instrument, such as HISEQ, GAIIX, or MYSEQ, available from ILLUMINA. Additionally, an additional sequence is included 5' to the sequencing tails that can be used as a priming site for adding a nucleotide barcode sequence to the amplification product in a subsequent PCR, thereby enabling multiplexed sequencing of multiple samples on a single lane of the high-throughput sequencing instrument.
[0195] In one embodiment, a 10,000-plex PCR assay pool is created in which the reverse primer has a tail corresponding to the required reverse sequence required for high-throughput sequencing instruments. After amplification in the first 10,000-plex assay, subsequent PCR amplification can be performed using another 10,000-plex pool with a partially nested forward primer (e.g., 6-base nested) for all targets and a reverse primer corresponding to the reverse sequencing tail included in the first round. This subsequent round of partially nested amplification using only one target-specific primer and a universal primer limits the size required for the assay, reducing sampling noise but significantly reducing the number of spurious amplification products. Sequencing tags can be added to the added ligation adapter and / or as part of the PCR probe, so that the tag becomes part of the final amplification product.
[0196] The fetal fraction affects test performance. Several methods exist for enriching the fetal fraction of DNA found in maternal plasma. The fetal fraction can be increased by the LM-PCR method already discussed above, as well as by targeted removal of long maternal fragments. In one embodiment, prior to multiplex PCR amplification of the target loci, an additional multiplex PCR reaction can be performed to selectively remove long, large maternal fragments corresponding to the loci targeted in the subsequent multiplex PCR. Additional primers are designed to anneal to sites that are further away from the polymorphism than would be expected to exist among cell-free fetal DNA fragments. These primers can be used in a single-cycle multiplex PCR reaction prior to multiplex PCR of the target polymorphic loci. These distal primers are tagged with a molecule or moiety that can enable selective recognition of the tagged DNA fragment. In one embodiment, these DNA molecules can be covalently modified with biotin molecules, which allows for removal of newly formed double-stranded DNA containing these primers after one cycle of PCR. The double-stranded DNA formed during the first round is likely to be of maternal origin. Removal of hybrid materials can be achieved by using magnetic streptavidin beads. There are other tagging methods that can work equally well. In some embodiments, size selection methods can be used to enrich samples for shorter DNA, such as DNA less than about 800 bp, less than about 500 bp, or less than about 300 bp. Amplification of short fragments can then proceed as usual.
[0197] The mini-PCR method described in this disclosure enables highly multiplexed amplification and analysis of hundreds, thousands, or even millions of loci from a single sample in a single reaction. Simultaneous detection of amplified DNA can be multiplexed, and by using barcoding PCR, tens to hundreds of samples can be multiplexed in a single sequencing lane. This multiplexed detection has been successfully tested up to 49-plex, allowing for much higher degrees of multiplexing. Effectively, this allows for genotyping of hundreds of samples at thousands of SNPs in a single sequencing run. For these samples, the method allows for the determination of genotype and heterozygosity rate, as well as simultaneous copy number determination, both of which can be used to detect aneuploidy. This method is particularly useful for detecting aneuploidy in a gestational fetus from free-floating DNA found in maternal plasma. This method can be used as part of a method for fetal sexing and / or fetal paternity prediction. This method can also be used as part of a method for mutation burden determination. This method can be used on any amount of DNA or RNA, and the targeted regions can be SNPs, other polymorphic regions, non-polymorphic regions, and combinations thereof.
[0198] In some embodiments, ligation-mediated universal PCR amplification of fragmented DNA can be used. Ligation-mediated universal PCR amplification can be used to amplify plasma DNA, which can then be divided into multiple parallel reactions. Ligation-mediated universal PCR amplification can also be used to preferentially amplify short fragments, thereby enriching the fetal fraction. In some embodiments, adding tags to fragments by ligation can allow for the detection of shorter fragments, the use of shorter target sequence-specific portions of primers, and / or annealing at higher temperatures to reduce non-specific reactions.
[0199] The methods described herein can be used for a number of purposes when a collection of target DNA is present that is mixed with a certain amount of contaminating DNA. In some embodiments, the target DNA and contaminating DNA can be from genetically related individuals. For example, genetic abnormalities in a fetus (target) can be detected from maternal plasma containing fetal (target) DNA and also maternal (contaminating) DNA, including whole chromosomal abnormalities (e.g., aneuploidies), partial chromosomal abnormalities (e.g., deletions, duplications, inversions, translocations), polynucleotide polymorphisms (e.g., STRs), single nucleotide polymorphisms, and / or other genetic abnormalities or variations. In some embodiments, the target DNA and contaminating DNA can be from the same individual, although, for example, in the case of cancer, the target DNA and contaminating DNA differ by one or more mutations (see, e.g., H. Mamon et al. Preferential Amplification of Apoptotic DNA from Plasma: Potential for Enhancing Detection of Minor DNA Alterations in Circulating DNA. Clinical Chemistry 54:9 (2008)). In some embodiments, DNA can be found in cell culture (apoptotic) supernatants. In some embodiments, apoptosis can be induced in biological samples (e.g., blood) for subsequent library preparation, amplification, and / or sequencing. Several possible workflows and protocols for achieving this goal are presented elsewhere in this disclosure.
[0200] In some embodiments, the target DNA may be from a single cell, a DNA sample consisting of less than one copy of the target genome, small amounts of DNA, DNA from mixed sources (e.g., pregnancy plasma: placenta and maternal DNA; cancer patient plasma and tumor: a mixture of healthy and cancer DNA, transplants, etc.), other bodily fluids, cell cultures, culture supernatants, forensic DNA samples, ancient DNA samples (e.g., insects trapped in amber), other DNA samples, and combinations thereof.
[0201] In some embodiments, short amplicon sizes can be used, which are particularly suitable for fragmented DNA (see, e.g., A. Sikora et al., Detection of increased amounts of cell-free fetal DNA with short PCR amplicons. Clin Chem. 2010 January;56(1):136-8).
[0202] Using short amplicon sizes can provide several important benefits. A short amplicon size can result in optimized amplification efficiency. A short amplicon size generally results in shorter products, and therefore less chance of nonspecific priming. Shorter products can be clustered more densely on a sequencing flow cell, resulting in smaller clusters. Note that the methods described herein may work equally well with longer PCR amplicons. The amplicon length can be increased, if necessary, for example, when sequencing a larger sequence range. Experiments using 146-plex targeted amplification with 100- to 200-bp-long assays as the first step of a nested PCR protocol have been performed on single cells and genomic DNA with positive results.
[0203] In some embodiments, the methods described herein can be used to amplify and / or detect SNPs, copy number, nucleotide methylation, mRNA levels, expression levels of other types of RNA, other genetic and / or epigenetic features. The mini-PCR methods described herein can be used in conjunction with next-generation sequencing, and the methods can be used in conjunction with other downstream methods, such as microarrays, digital PCR counting, real-time PCR, mass spectrometry, etc.
[0204] In some embodiments, the mini-PCR amplification method described herein can be used as part of a method for accurately quantifying minority populations. The method can be used for absolute quantification using spiked calibrators. The method can be used for quantification of variants / low-abundance alleles by ultra-deep sequencing and can be performed in a highly multiplexed manner. The method can be used for standard paternity and identity testing of relatives or ancestry in humans, animals, plants, or other organisms. The method can be used for forensic testing. The method can be used for rapid genotyping and copy number analysis (CN) of any type of material, such as amniotic fluid and CVS, sperm, and products of conception (POC). The method can be used for single-cell analysis, such as genotyping biopsy samples from embryos. The method can be used for rapid embryo analysis (less than one day, one day, or two days after biopsy) by targeted sequencing using min-PCR.
[0205] In some embodiments, mini-PCR amplification methods can be used for tumor analysis: tumor biopsies are often a mixture of healthy and tumor cells. Targeted PCR allows deep sequencing of SNPs and loci without nearby background sequences. The method can be used for copy number and loss of heterozygosity analysis of tumor DNA. The tumor DNA can be present in many different body fluids or tissues of tumor patients. The method can be used for detecting tumor recurrence and / or tumor screening. The method can be used for seed quality control testing. The method can be used for breeding or fishing. Note that any of these methods can be used equally well to target non-polymorphic loci for ploidy calling.
[0206] Some references describing some of the basic methods underlying this disclosure include: (1) Wang HY, Luo M, Tereshchenko IV, Frikker DM, Cui X, LiJ Y, Hu G, Chu Y, Azaro MA, Lin Y, Shen L, Yang Q, Kambouris ME, Gao R, Shih W, Li H. Genome Res. 2005 Feb;15(2):276-83. Department of Molecular Genetics, Microbiology and Immunology / The Cancer Institute of New Jersey, Robert Wood Johnson Medical School, New Brunswick, New Jersey 08903, USA. (2) High-throughput genotyping of single nucleotide polymorphisms with high sensitivity. Li H, Wang HY, Cui X, Luo M, Hu G, Greenawalt DM, Tereshchenko IV, Li JY, Chu Y, Gao R. Methods Mol Biol. 2007;396 - PubMed PMID:18025699. (3) Nested Patch PCR enables highly multiplexed mutation discovery in candidate genes. Varley KE, Mitra RD. Genome Res. 2008 Nov;18(11):1844-50. Epub 2008 Oct 10 (which describes a method involving multiplexing an average of nine assays for sequencing). Note that the method disclosed herein allows for orders of magnitude more multiplexing than in the above references.
[0207] Variations of Targeted PCR - Nesting There are many possible workflows for performing PCR, and several workflows typical of the methods disclosed herein are described. The steps outlined herein are not intended to exclude other possible steps, nor are any of the steps described herein required for the method to function properly. Numerous parameter variations or other modifications are known in the literature and can be made without affecting the core of the invention. One particular general workflow is shown below, followed by several possible variants. Variants generally refer to possible secondary PCR reactions, such as different types of nesting (step 3), that can be performed. It is important to note that variants can be performed at different times or in a different order than those explicitly described herein. Examples using polymorphic loci for illustration can be easily adapted to the amplification of non-polymorphic loci, if desired. 1. Ligation adaptors, often referred to as library tags or ligation adaptor tags (LT), can be added to the DNA in the sample, where the ligation adaptors contain universal priming sequences followed by universal amplification. In some embodiments, this can be done using standard protocols designed for generating sequencing libraries after fragmentation. In some embodiments, the DNA sample can be blunt-ended, and then an A can be added to the 3' end. A Y-adapter with a T-overhang can be added and ligated. In some embodiments, other sticky ends besides A or T overhangs can be used. In some embodiments, other adaptors, such as loop ligation adaptors, can be added. In some embodiments, the adaptors can have tags designed for PCR amplification. 2. Specific Target Amplification (STA): Hundreds, thousands, tens of thousands, or even hundreds of thousands of targets can be multiplexed with pre-amplification in a single reaction volume. STA is typically performed for 10-30 cycles, but can also be performed for 5-40, 2-50, and even 1-100 cycles. For example, primers can be tailed for a simpler workflow or to avoid sequencing most dimers. Note that dimers of both primers carrying the same tag generally will not be efficiently amplified or sequenced. In some embodiments, PCR can be performed for 1 to 10 cycles, in some embodiments, PCR can be performed for 10 to 20 cycles, in some embodiments, PCR can be performed for 20 to 30 cycles, in some embodiments, PCR can be performed for 30 to 40 cycles, and in some embodiments, PCR can be performed for more than 40 cycles. Amplification can be linear. The number of PCR cycles can be optimized to result in an optimal depth-of-read (DOR) profile. Different DOR profiles may be desired for different purposes. In some embodiments, a more even distribution of reads across all assays is desirable; if the DOR for some assays is very small, the stochastic noise may be too high for the data to be very useful, but if the read depth is very deep, the marginal usefulness of each additional read is relatively small. Primer tails can improve detection of fragmented DNA from universally tagged libraries. When the library tag and primer tail contain homologous sequences, hybridization can be improved (e.g., melting temperature (T M)), the primer can be extended only if a portion of the primer target sequence is within the DNA fragment of the sample. In some embodiments, 13 or more target-specific base pairs can be used. In some embodiments, 10-12 target-specific base pairs can be used. In some embodiments, 8-9 target-specific base pairs can be used. In some embodiments, 6-7 target-specific base pairs can be used. In some embodiments, STA can be performed on pre-amplified DNA, such as MDA, RCA, other whole genome amplification, or adapter-mediated universal PCR. In some embodiments, STA can be performed on samples in which specific sequences and populations have been enriched or depleted, for example, by size selection, target capture, or directed degradation. 3. In some embodiments, secondary multiplex PCR or primer extension reactions can be performed to increase specificity and reduce unwanted products. For example, full nesting, semi-nesting, hemi-nesting, and / or subdividing smaller assay pools into parallel reactions are all techniques that can be used to increase specificity. Experiments have shown that splitting a sample into three 400-plex reactions yields product DNA with higher specificity than a single 1,200-plex reaction using the exact same primers. Similarly, experiments have shown that splitting a sample into four 2,400-plex reactions yields product DNA with higher specificity than a single 9,600-plex reaction using the exact same primers. In certain embodiments, target-specific and tag-specific primers of the same and opposite orientations can be used. 4. In some embodiments, the DNA sample (diluted, purified, or otherwise) produced by the STA reaction can be amplified using tag-specific primers and "universal amplification," i.e., amplifying many or all of the pre-amplified, tagged targets. The primers may contain additional functional sequences, such as barcodes or complete adapter sequences, required for sequencing on high-throughput sequencing platforms.
[0208] These methods can be used to analyze any DNA sample, and are particularly useful when the DNA sample is particularly small or when the DNA originates from more than one individual, such as maternal plasma. These methods can be used with DNA samples such as single or small numbers of cells, genomic DNA, plasma DNA, amplified plasma libraries, amplified apoptotic supernatant libraries, or other mixed DNA samples. In some embodiments, these methods can be used when cells with different genetic makeups may be present in a single individual, such as cancer or transplants.
[0209] Protocol Variations (variations and / or additions to the workflow above) Direct multiplex mini-PCR: Specific target amplification (STA) of multiple target sequences using tagged primers is shown in Figure 1. 101 denotes double-stranded DNA with the polymorphic locus of interest at X. 102 denotes double-stranded DNA with ligation adapters added for universal amplification. 103 denotes universally amplified single-stranded DNA hybridized with PCR primers. 104 denotes the final PCR product. In some embodiments, STA can be performed on more than 100, more than 200, more than 500, more than 1,000, more than 2,000, more than 5,000, more than 10,000, more than 20,000, more than 50,000, more than 100,000, or more than 200,000 targets. In a subsequent reaction, tag-specific primers amplify all target sequences and extend the tags to include all sequences required for sequencing, including sampling indexes. In some embodiments, primers may not be tagged, or only certain primers may be tagged. Sequencing adapters can be added by conventional adapter ligation. In some embodiments, the first primer may carry a tag.
[0210] In some embodiments, primers are designed to amplified DNA of unexpectedly short length. Prior art demonstrates that those skilled in the art typically design amplification products of 100+ bp. In some embodiments, amplification products can be designed to be less than 80 bp. In some embodiments, amplification products can be designed to be less than 70 bp. In some embodiments, amplification products can be designed to be less than 60 bp. In some embodiments, amplification products can be designed to be less than 50 bp. In some embodiments, amplification products can be designed to be less than 45 bp. In some embodiments, amplification products can be designed to be less than 40 bp. In some embodiments, amplification products can be designed to be less than 35 bp. In some embodiments, amplification products can be designed to be between 40 bp and 65 bp.
[0211] Experiments were performed using this protocol with 1200-plex amplification. Both genomic DNA and pregnancy plasma were used; approximately 70% of sequence reads mapped to target sequences. Details are provided elsewhere in this document. 1042-plex sequencing without assay design and selection resulted in >99% of sequences being primer dimer products.
[0212] After sequential PCR:STA1, multiple aliquots of the product can be amplified in parallel using pools of decreasing complexity with the same primers. The first amplification can generate enough material for splitting. This method is particularly effective for small samples, e.g., about 6-100 pg, about 100 pg-1 ng, about 1 ng-10 ng, or about 10 ng-100 ng. The 1200-plex protocol was expanded three times to 400-plex. The mapping of sequencing reads increased from about 60-70% with 1200-plex alone to over 95%.
[0213] Semi-nested mini-PCR: (See Figure 2) After STA1, a second STA consisting of a multiplex set of inner nested forward primers (103B, 105b) and one (or a few) tag-specific reverse primers (103A) is performed. 101 represents double-stranded DNA with the polymorphic locus of interest at X. 102 represents double-stranded DNA with a ligation adapter added for universal amplification. 103 represents universally amplified single-stranded DNA with hybridized forward primer B and reverse primer A. 104 represents the PCR product from 103. 105 represents the product from 104 with hybridized nested forward primer b and reverse tag A, which is already part of the molecule from the PCR generated between 103 and 104. 106 represents the final PCR product. Using this workflow, typically more than 95% of the sequences map to the intended target. Nested primers may overlap the outer forward primer sequence, but introduce additional 3' terminal bases. In some embodiments, between 1 and 20 extra 3' bases can be used. Experiments have shown that using 9 or more extra 3' bases works well in 1200-plex designs.
[0214] Fully nested mini-PCR: (See FIG. 3) After STA step 1, a second multiplex PCR (or parallel, reduced complexity mpPCR) can be performed using two nested primers carrying tags (A, a, B, b). 101 denotes double-stranded DNA with the polymorphic locus of interest at X. 102 denotes double-stranded DNA with ligation adapters added for universal amplification. 103 denotes universally amplified single-stranded DNA hybridized with forward primer B and reverse primer A. 104 denotes the PCR product from 103. 105 denotes the product from 104 hybridized with nested forward primer b and nested reverse primer a. 106 denotes the final PCR product. In some embodiments, a complete set of two primers can be used. Experiments using a fully nested mini-PCR protocol were used to perform 146-plex amplification on single cells and triplicates without step 102, which adds universal ligation adapters and amplifies.
[0215] Hemi-nested mini-PCR: (See Figure 4) It is possible to use target DNA with adapters at the ends of the fragments. STA is performed consisting of a multiplex set of forward primers (B) and one (or a few) tag-specific reverse primers (A). A second STA can be performed using a universal tag-specific forward primer and a target-specific reverse primer. 101 indicates double-stranded DNA with the polymorphic locus of interest at X. 102 indicates double-stranded DNA with a ligation adapter added for universal amplification. 103 indicates universally amplified single-stranded DNA hybridized with reverse primer A. 104 indicates the PCR product from 103 amplified using reverse primer A and ligation adapter tag primer LT. 105 indicates the product from 104 hybridized with forward primer B. 106 indicates the final PCR product. In this workflow, target-specific forward and reverse primers are used in separate reactions, thereby reducing reaction complexity and preventing forward and reverse primer dimer formation. Note that in this example, primers A and B can be considered first primers, and primers "a" and "b" can be considered inner primers. This method is comparable to direct PCR but is a significant improvement over direct PCR because it avoids primer dimers. After the first round of the hemi-nested protocol, approximately 99% of non-target DNA is typically observed, but after the second round, this is typically significantly improved.
[0216] Triple hemi-nested mini-PCR: (See Figure 5) It is possible to use target DNA with adapters at the ends of the fragments. STA is performed using a multiplex set of forward primers (B) and one (or a few) tag-specific reverse primers (A) and (a). A second STA can be performed using a universal tag-specific forward primer and a target-specific reverse primer. 101 indicates double-stranded DNA with the polymorphic locus of interest at X. 102 indicates double-stranded DNA with a ligation adapter added for universal amplification. 103 indicates universally amplified single-stranded DNA hybridized with reverse primer A. 104 indicates the PCR product from 103 amplified using reverse primer A and ligation adapter tag primer LT. 105 indicates the product from 104 hybridized with forward primer B. 106 indicates the PCR product from 105 amplified using reverse primer A and forward primer B. 107 indicates the product from 106 to which the reverse primer "a" hybridized. 108 indicates the final PCR product. Note that in this example, primers "a" and B can be considered inner primers, and A can be considered the first primer. If desired, both A and B can be considered first primers, and "a" can be considered the inner primer. The names of the reverse and forward primers can be switched. In this workflow, target-specific forward and reverse primers are used in separate reactions, thereby reducing reaction complexity and preventing dimer formation between the forward and reverse primers. This method is comparable to direct PCR but is a significant improvement over direct PCR because it avoids primer dimers. After the first round of the hemi-nested protocol, approximately 99% non-target DNA is typically observed, but after the second round, this is typically greatly improved.
[0217] One-sided nested mini-PCR: (See Figure 6) It is possible to use target DNA with adapters at the ends of the fragments. STA can also be performed using a multiplex set of nested forward primers and a ligation adapter tag as the reverse primer. A second STA can then be performed using a set of nested forward primers and a universal reverse primer. 101 indicates double-stranded DNA with a polymorphic locus of interest at X. 102 indicates double-stranded DNA with a ligation adapter added for universal amplification. 103 indicates universally amplified single-stranded DNA hybridized with forward primer A. 104 indicates the PCR product from 103 amplified using forward primer A and ligation adapter tag reverse primer LT. 105 indicates the product from 104 hybridized with the nested forward primer. 106 indicates the final PCR product. This method allows for detection of shorter target sequences than standard PCR by using overlapping primers in the first and second STAs. The method is generally performed minus a sample of DNA that has already undergone STA step 1 above—universal tagging and amplification—using two nested primers on only one side and a library tag on the other. The method was performed on libraries of apoptotic supernatant and pregnancy plasma. Using this workflow, approximately 60% of sequences were mapped to the intended target. Note that reads that contained reverse adapter sequences were not mapped; therefore, this number would be expected to be higher if reads containing reverse adapter sequences were mapped.
[0218] Single-sided mini-PCR: It is possible to use target DNA with adapters at the ends of the fragments (see Figure 7). STA can be performed using a multiplex set of forward primers and one (or a few) tag-specific reverse primers. 101 indicates double-stranded DNA with the polymorphic locus of interest at X. 102 indicates double-stranded DNA with a ligation adapter added for universal amplification. 103 indicates single-stranded DNA hybridized with forward primer A. 104 indicates the PCR product from 103 amplified using forward primer A and ligation adapter tag reverse primer LT, which is the final PCR product. This method allows for the detection of shorter target sequences than standard PCR. However, since only one target-specific primer is used, it can be relatively nonspecific. The efficiency of this protocol is half that of single-sided nested mini-PCR.
[0219] Reverse semi-nested mini-PCR: Target DNA with adapters at the ends of the fragments can be used (see Figure 8). STA can be performed using a multiplex set of forward primers and one (or a few) tag-specific reverse primers. 101 denotes double-stranded DNA with the polymorphic locus of interest at X. 102 denotes double-stranded DNA with a ligation adapter added for universal amplification. 103 denotes single-stranded DNA hybridized with reverse primer B. 104 denotes the PCR product from 103 amplified using reverse primer B and ligation adapter tag forward primer LT. 105 denotes PCR product 104 hybridized with forward primer A and inner reverse primer "b." 106 denotes the PCR product amplified from 105 using forward primer A and reverse primer "b," which is the final PCR product. This method allows for the detection of shorter target sequences than standard PCR.
[0220] There may be further variations that are simply repetitions or combinations of the above methods, such as double-nested PCR using three sets of primers. Another variation is one-sided semi-nested mini-PCR, in which STA can be performed using multiple sets of nested forward primers and one (or a few) tag-specific reverse primers.
[0221] Note that in all of these variations, the identities of the forward and reverse primers can be swapped. Note that in some embodiments, the nested variations can be performed equally well without the addition of adapter tags and the initial library preparation, which includes a universal amplification step. Note that in some embodiments, additional rounds of PCR can be included, with additional forward and / or reverse primers and amplification steps; these additional steps can be particularly useful when it is desirable to further increase the percentage of DNA molecules that correspond to the targeted locus.
[0222] Nesting Workflow There are many ways to perform amplification with different degrees of nesting and different degrees of multiplexing. In Figure 9, a flowchart is shown with some of the possible workflows. Note that the use of 10,000-plex PCR is only an example, and these flowcharts work equally well for other degrees of multiplexing.
[0223] Loop Ligation Adapter For example, when adding adapters with universal tags to generate libraries for sequencing, there are several ways to ligate adapters. One method is to blunt-end sample DNA, perform A-tailing, and ligate with adapters that have T-overhangs. There are several other ways to ligate adapters. There are also several adapters that can be ligated. For example, a Y-adapter can be used, which consists of two strands of DNA, one strand having a double-stranded region and a region designated by a forward primer region, and the other strand having a double-stranded region that is complementary to the double-stranded region on the first strand and a region with a reverse primer. When annealed, the double-stranded region can contain a T-overhang to ligate with double-stranded DNA that has A-overhangs.
[0224] In one embodiment, the adaptor may be a loop of DNA with complementary terminal regions containing a forward primer-tagged region (LFT), a reverse primer-tagged region (LRT), and a cleavage site between the two (see FIG. 10 ). 101 refers to a double-stranded, blunt-ended target DNA. 102 refers to an A-tailed target DNA. 103 refers to a loop ligation adaptor with a T-overhang "T" and a cleavage site "Z." 104 refers to a target DNA with a loop ligation adaptor appended. 105 refers to a target DNA with a ligation adaptor appended, cleaved at the cleavage site. LFT refers to a ligation adaptor forward tag, and LRT refers to a ligation adaptor reverse tag. The complementary region may end in a T-overhang or other feature that can be used to ligate to the target DNA. The cleavage site may be a series of uracils for cleavage along UNG, or may be a sequence that can be recognized and cleaved by a restriction enzyme or other cleavage method, or simply basic amplification. These adapters can be used, for example, to prepare any library for sequencing. These adapters can be used in combination with any of the other methods described herein, for example, mini-PCR amplification methods.
[0225] Internally tagged primers When using sequencing to determine the allele present at a given polymorphic locus, the sequence read typically begins upstream of the primer binding site (a) and then reads the polymorphic site (X). Tags are typically arranged as shown on the left side of Figure 11. 101 refers to the single-stranded target DNA bearing the polymorphic locus of interest "X" and primer "a" with tag "b" attached. To avoid nonspecific hybridization, the primer binding site (the region of the target DNA complementary to "a") is typically 18-30 bp in length. Sequence tag "b" is typically approximately 20 bp; in theory, these can be any length longer than approximately 15 bp, but many people use primer sequences sold by sequencing platform companies. The distance "d" between "a" and "X" should be at least 2 bp to avoid allelic bias. When performing multiplex PCR amplification using the methods disclosed herein or other methods, careful primer design is required to avoid excessive primer-primer interactions. The window of acceptable distance "d" between "a" and "X" can vary considerably: from 2 bp to 10 bp, from 2 bp to 20 bp, from 2 bp to 30 bp, or even from 2 bp to more than 30 bp. Thus, using the primer configuration shown on the left side of Figure 11, sequence reads must be a minimum of 40 bp to obtain reads long enough to measure polymorphic loci, and depending on the lengths of "a" and "d," sequence reads of up to 60 bp or 75 bp may be required. Longer sequence reads typically increase the cost and time required to sequence a given number of reads; therefore, minimizing the required read length can save both time and money. Furthermore, because, on average, base reads earlier in the read are more accurate than those later in the read, reducing the required sequence read length can also increase the accuracy of measuring polymorphic regions.
[0226] In one embodiment, as shown in 103 of FIG. 11 , a primer binding site (a), referred to as an internally tagged primer, is divided into multiple segments (a′, a″, a′″...), and a sequence tag (b) is placed on the segment of DNA in the middle of the two primer binding sites. This arrangement allows the sequencer to generate shorter sequence reads. In one embodiment, a′+a″ should be at least about 18 bp, and can be 30 bp, 40 bp, 50 bp, 60 bp, 80 bp, 100 bp, or longer. In one embodiment, a″ should be at least about 6 bp, and in one embodiment, between about 8 bp and 16 bp. All other factors being equal, the use of internally tagged primers can reduce the length of the required sequence read to at least as much as 6 bp, 8 bp, 10 bp, 12 bp, 15 bp, or even as much as 20 bp or 30 bp. This can result in significant financial, time, and accuracy advantages. An example of an internally tagged primer is shown in FIG.
[0227] Primer with ligation adaptor binding region One problem with fragmented DNA is that, due to its short length, polymorphisms are more likely to be located near the ends of the DNA strands than longer strands (e.g., 101, Figure 10). Because PCR capture of polymorphisms requires primer binding sites of appropriate length on both sides of the polymorphism, a significant number of DNA strands bearing the target polymorphism will be missed due to insufficient overlap between the primer and target binding sites. In one embodiment, a ligation adaptor 102 can be added to the target DNA 101, and the target primer 103 can have a region (cr) complementary to a ligation adaptor tag (lt) added upstream of the designed binding region (a) (see Figure 13). Thus, if the binding region (the region of 101 complementary to a) is shorter than the 18 bp typically required for hybridization, the region of the primer (cr) complementary to the library tag can increase the binding energy to the point where PCR can proceed. Note that any specificity lost due to the shorter binding region can be compensated for by other PCR primers with appropriately long target binding regions. Note that this embodiment can be used in combination with direct PCR or any of the other methods described herein, such as nested PCR, semi-nested PCR, hemi-nested PCR, one-sided nested or semi-nested or hemi-nested PCR, or other PCR protocols.
[0228] When determining ploidy using sequencing data in combination with analytical methods that involve comparing observed allele data with predicted allele distributions for various hypotheses, each additional read from an allele with a low read depth provides more information than a read from an allele with a high read depth. Ideally, therefore, a uniform depth of read (DOR) is achieved, where each locus has a similar number of representative sequence reads. Therefore, it is desirable to minimize the variance of the DOR. In some embodiments, increasing the annealing time can reduce the coefficient of variation of the DOR (which can be defined as the standard deviation of the DOR / average DOR). In some embodiments, the annealing temperature can be longer than 2 minutes, 4 minutes, 10 minutes, 30 minutes, and even longer than 1 hour. Because annealing is an equilibrium process, there is no limit to the improvement in DOR variance with increasing annealing time. In some embodiments, increasing the primer concentration reduces the DOR variance.
[0229] Representative whole genome amplification methods In some embodiments, DNA amplification, such as whole genome applications, can be included to amplify nucleic acid samples before amplifying only target loci. DNA amplification is the process of converting small amounts of genetic material into larger amounts containing a collection of similar genetic data and can be performed by a variety of methods, including, but not limited to, polymerase chain reaction (PCR). One method for amplifying DNA is whole genome amplification (WGA). Several methods for WGA are available: ligation-mediated PCR (LM-PCR), degenerate oligonucleotide primer PCR (DOP-PCR), and multiple displacement amplification (MDA). In LM-PCR, short DNA sequences called adapters are ligated to blunt ends of DNA. These adapters contain universal amplification sequences and are used to amplify DNA by PCR. In DOP-PCR, random primers, also containing universal amplification sequences, are used in the first round of annealing and PCR. A second round of PCR is then used to further amplify the sequences using universal primer sequences. MDA uses phi-29 polymerase, a highly processive and nonspecific enzyme that replicates DNA and has been used for single-cell analysis. The major limitations to amplifying material from single cells are (1) the need to use extremely dilute DNA concentrations or very small volumes of reaction mixtures, and (2) the difficulty of reliably dissociating DNA from proteins across the entire genome. Nevertheless, single-cell whole genome amplification has been used successfully for a variety of applications for many years. There are other methods for amplifying DNA from a sample of DNA. DNA amplification converts the initial DNA sample into a sample of DNA with a similar set of sequences but in much larger quantities. In some cases, amplification may not be necessary.
[0230] In some embodiments, DNA can be amplified using universal amplification, such as WGA or MDA. In some embodiments, DNA can be amplified by targeted amplification, such as targeted PCR or circularization probes. In some embodiments, DNA can be preferentially enriched using targeted amplification methods or methods that result in complete or partial separation of desired and undesired DNA, such as capture by hybridization techniques. In some embodiments, DNA can be amplified by using a combination of universal amplification and preferential enrichment methods. A more complete description of some of these methods can be found elsewhere in this document.
[0231] Representative enrichment and sequencing methods In certain embodiments, the methods disclosed herein use selective enrichment techniques that preserve the relative allele frequencies present at each target locus (e.g., each polymorphic locus) from a set of target loci (e.g., polymorphic loci) in the original DNA sample. While enrichment is particularly useful for analyzing polymorphic loci, these enrichment methods can be easily adapted for non-polymorphic loci, if desired. In some embodiments, amplification and / or selective enrichment techniques may involve PCR, such as ligation-mediated PCR, fragment capture by hybridization, molecular inversion probes, or other circularization probes. In some embodiments, amplification or selective enrichment methods may involve the use of probes such that, when hybridized correctly to the target sequence, the 3' or 5' end of the nucleotide probe is separated from the polymorphic site of the allele by a small number of nucleotides. This separation reduces preferential amplification of one allele, referred to as allelic bias. This is an improvement over methods involving the use of probes in which the 3' or 5' end of the correctly hybridized probe is directly adjacent to or very close to the polymorphic site of the allele. In some embodiments, probes whose hybridization region may or certainly contains a polymorphic site are excluded. The presence of a polymorphic site at the hybridization site may cause unequal hybridization in some alleles or may inhibit hybridization altogether, resulting in preferential amplification of specific alleles. These embodiments are an improvement over other methods involving targeted amplification and / or selective enrichment in that they better preserve the original allele frequency at each polymorphic locus in the sample, whether the sample is a pure genomic sample from a single individual or a mixture of individuals.
[0232] The use of techniques to enrich a DNA sample for a set of targeted loci followed by sequencing as part of a method for non-invasive prenatal allele or ploidy calling can confer several unexpected advantages. In some embodiments of the present disclosure, the method includes measuring genetic data for use in informatics-based methods, such as PARENTAL SUPPORT™ (PS). The end result of some embodiments is actionable genetic data for an embryo or fetus. There are many methods that can be used to measure genetic data for an individual and / or related individuals as part of an embodied method. In certain embodiments, methods for enriching the concentration of a set of targeted alleles are disclosed herein, the methods comprising one or more of the following steps: targeted amplification of genetic material, addition of locus-specific oligonucleotide probes, ligation of specific DNA strands, isolation of the desired set of DNA, removal of undesired components of the reaction, detection of specific DNA sequences by hybridization, and detection of the sequences of one or more DNA strands by DNA sequencing methods. In some cases, the DNA strand may refer to the target genetic material, in some cases the DNA strand may refer to a primer, in some cases the DNA strand may refer to a synthesized sequence, or a combination thereof. These steps can be performed in a number of different orders.
[0233] For example, a universal amplification step of DNA prior to targeted amplification can provide several advantages, such as eliminating the risk of bottlenecks and reducing allele bias. DNA can be mixed with oligonucleotide probes that can hybridize with two adjacent regions on either side of the target sequence. After hybridization, the ends of the probes can be linked by adding a polymerase, which is a means of ligation, and any necessary reagents to enable circularization of the probe. After circularization, exonuclease can be added to digest the non-circularized genetic material, and then the circularized probe can be detected. DNA can be mixed with PCR primers that can hybridize with two adjacent regions on either side of the target sequence. After hybridization, the ends of the probes can be linked by adding a polymerase, which is a means of ligation, and any necessary reagents to complete PCR amplification. Amplified or unamplified DNA can be targeted by a hybrid capture probe that targets a set of loci, and after hybridization, the probe can be localized and separated from the mixture to produce a mixture of DNA enriched with target sequences.
[0234] The use of methods to target and subsequently sequence specific loci as part of an allele or ploidy calling method can confer several unexpected advantages. Some methods by which DNA can be targeted or preferentially enriched include the use of circularized probes, linked inverted probes (LIPs, MIPs), hybridization capture methods such as SURESELECT, and targeted PCR or ligation-mediated PCR amplification strategies.
[0235] In some embodiments, the methods of the present disclosure involve measuring genetic data for use in informatics-based methods, such as PARENTAL SUPPORT™ (PS), further described herein. PARENTAL SUPPORT™ is an informatics-based approach for manipulating genetic data, aspects of which are described herein. The end outcome of some embodiments is actionable genetic data for an embryo or fetus, followed by a clinical decision based on the actionable data. The algorithm behind the PS method obtains measured genetic data for a target individual, often an embryo or fetus, and measured genetic data from related individuals, which can increase the accuracy with which the genetic status of the target individual is known. In some embodiments, the measured genetic data is used in the context of determining ploidy during prenatal genetic diagnosis. In some embodiments, the measured genetic data is used in the context of determining ploidy or allele calling on an embryo during in vitro fertilization. There are many methods that can be used to measure genetic data for an individual and / or related individuals in the above-mentioned situations. Different methods involve a number of steps, often involving amplifying genetic material, adding oligonucleotide probes, ligating specific DNA strands, isolating the desired DNA population, removing unwanted components of the reaction, detecting specific DNA sequences by hybridization, and detecting the sequence of one or more DNA strands by DNA sequencing methods. In some cases, DNA strand refers to the target genetic material, in other cases to primers, in other cases to synthesized sequences, or a combination thereof. These steps can be performed in a number of different orders.
[0236] It should be noted that theoretically, any number of loci within a genome can be targeted, from one to over one million loci. When a DNA sample is targeted and then sequenced, the proportion of alleles read by the sequencer is enriched relative to the amount that naturally occurs in the sample. The degree of enrichment can be anything from 1 percent (or even lower) to 10-fold, 100-fold, 1,000-fold, or even up to one million-fold. The human genome contains approximately 3 billion base pairs and nucleotides, including approximately 75 million polymorphic loci. The more loci that are targeted, the less enrichment is possible. The fewer loci that are targeted, the greater the degree of enrichment is possible, and the greater read depth can be achieved at those loci for a given number of sequence reads.
[0237] In some embodiments of the present disclosure, targeting or preference can be entirely focused on SNPs. In some embodiments, targeting or preference can be focused on any polymorphic site. Several commercial targeting products for enriching exons are available. Surprisingly, targeting exclusively SNPs or exclusively polymorphic loci is particularly advantageous when using methods for NPD that rely on allele distribution. There are also published methods for NPD that use sequencing, such as U.S. Patent No. 7,888,017, which involves read count analysis, in which the number of reads is focused on counting the number of reads that map to a given chromosome, and the analyzed sequence reads are not focused on polymorphic genomic regions. These types of methodologies that do not focus on polymorphic alleles are not as useful as targeting or preferentially enriching a set of alleles.
[0238] In some embodiments of the present disclosure, it is possible to use a targeting method that focuses on SNPs to enrich genetic samples in polymorphic regions of genome.In some embodiments, it is possible to focus on a small number of SNPs, for example, between 1 and 100 SNPs, or a larger number, for example, between 100 and 1,000, between 1,000 and 10,000, between 10,000 and 100,000, or more than 100,000 SNPs.In some embodiments, it is possible to focus on one or a small number of chromosomes that correlate with live birth with trisomy, for example, chromosome 13, chromosome 18, chromosome 21, chromosome X and chromosome Y, or some combinations thereof.In some embodiments, it is possible to enrich target SNPs by a small factor, for example, between 1.01 times and 100 times, or by a larger factor, for example, between 100 times and 1,000,000 times, or more than 1,000,000 times. In some embodiments of the present disclosure, a targeting method can be used to generate a sample of DNA that is preferentially enriched in polymorphic regions of the genome. In some embodiments, this method can be used to generate a mixture of DNA with any of these characteristics, where the mixture of DNA also contains maternal DNA and free-floating fetal DNA. In some embodiments, this method can be used to generate a mixture of DNA with any combination of these factors. For example, the method described herein can be used to generate a mixture of DNA that contains maternal DNA and fetal DNA and is preferentially enriched in DNA corresponding to 200 SNPs, all of which are located on either chromosome 18 or 21 and are enriched by an average of 1,000 times. In another example, the method can be used to generate a mixture of DNA that is preferentially enriched in 10,000 SNPs, all or most of which are located on chromosomes 13, 18, 21, X, and Y, with an average enrichment per locus of more than 500 times. Any of the targeting methods described herein can be used to generate a mixture of DNA that is preferentially enriched at a particular locus.
[0239] In some embodiments, the methods of the present disclosure further include measuring the DNA in the mixed fraction using a high-throughput DNA sequencer, wherein the DNA in the mixed fraction contains a disproportionate number of sequences from one or more chromosomes, wherein the one or more chromosomes are selected from the group consisting of chromosome 13, chromosome 18, chromosome 21, chromosome X, chromosome Y, and combinations thereof.
[0240] Three methods are described herein: multiplex PCR, targeted capture by hybridization, and linked inverse probe (LIP), which can be used to obtain and analyze measurements from a sufficient number of polymorphic loci from maternal plasma samples to detect fetal aneuploidy. This does not exclude other methods for selectively enriching target loci. Other methods can be used equally well without changing the core of the method. In each case, the polymorphisms assayed may include single nucleotide polymorphisms (SNPs), small indels, or STRs. A preferred method involves the use of SNPs. Each method generates allele frequency data, and the allele frequency data for each target locus and / or the joint allele frequency distribution from these loci can be analyzed to determine fetal ploidy. Each method has its own considerations due to the limited source material and the fact that maternal plasma consists of a mixture of maternal and fetal DNA. This method can be combined with other methods to provide more accurate determinations. In some embodiments, this method can be combined with a sequence counting method, such as that described in U.S. Patent No. 7,888,017. The described method can also be used to non-invasively detect fetal paternity from maternal plasma samples. Furthermore, each method can be applied to other DNA mixtures or pure DNA samples to detect the presence or absence of aneuploid chromosomes, genotype multiple SNPs from degraded DNA samples, detect segmental copy number variations (CNV), detect other target genotype states, or some combination thereof.
[0241] Accurate measurement of allele distribution in a sample Current sequencing methods can be used to estimate the distribution of alleles in samples.One of such methods involves randomly sampling sequences from pooled DNA, which is called shotgun sequencing.The proportion of specific alleles in sequencing data is generally very low and can be determined by simple statistics.Human genome contains approximately 3 billion base pairs.Therefore, if the sequencing method used generates 100bp reads, specific alleles will be measured approximately once every 30 million sequence reads.
[0242] In an embodiment, the method of the present disclosure is used to determine the presence or absence of two or more different haplotypes that contain the same set of loci in a DNA sample from the measured allele distribution of the loci from that chromosome.Different haplotypes can represent two different homologous chromosomes from one individual, three different homologous chromosomes from a trisomic individual, three different homologous haplotypes from a mother and a fetus, where one of the haplotypes is shared between the mother and the fetus, three or four haplotypes from a mother and a fetus, where one or two of the haplotypes is shared between the mother and the fetus, or other combinations.While polymorphic alleles between haplotypes tend to be more informative, any alleles for which both the mother and the father are not homozygous for the same allele can provide useful information through the measured allele distribution beyond the information that can be obtained from simple read count analysis.
[0243] However, shotgun sequencing of such samples is very inefficient, because it will result in many sequences for regions that are not polymorphic between different haplotypes in samples or for chromosomes that are not of interest, and therefore will not show information about the proportion of target haplotypes.Described herein is a method for specifically targeting and / or preferentially enriching the DNA segments in samples that are more likely to be polymorphic in genome, thereby increasing the yield of allele information obtained by sequencing.It should be noted that for the allele distribution measured in enriched samples to truly represent the actual amount present in target individuals, it is important that the preferential enrichment of one allele is slight or non-existent compared to other alleles at a given locus in the target segment.Currently known methods for targeting polymorphic alleles in the art are designed to ensure that at least some of any alleles present are detected.However, these methods have not been designed to measure the unbiased allele distribution of polymorphic alleles present in original mixtures. It is not obvious that any particular method of target enrichment can produce an enriched sample in which the measured allele distribution better accurately represents the allele distribution present in the original unamplified sample than any other method. While many enrichment methods can be expected to achieve this goal in theory, those skilled in the art are well aware that current amplification, targeting, and other preferential enrichment methods have a significant amount of stochastic or deterministic bias. One embodiment of the method described herein allows multiple alleles found in a mixture of DNA corresponding to a given locus in a genome to be amplified or preferentially enriched so that the degree of enrichment of each allele is approximately the same. In other words, the method allows the ratio between alleles corresponding to each locus to remain essentially the same as the ratio in the original DNA mixture, while increasing the relative amount of alleles present in the mixture as a whole. Some reported methods can result in allele biases of more than 1%, more than 2%, more than 5%, or even more than 10%.This preferential enrichment may be due to capture bias when using capture by hybridization techniques, or amplification bias, which may be small for each cycle but may become significant when compounded over 20, 30, or 40 cycles. For purposes of this disclosure, the ratio remains essentially the same means that the allele ratio in the original mixture divided by the allele ratio in the resulting mixture is between 0.95 and 1.05, between 0.98 and 1.02, between 0.99 and 1.01, between 0.995 and 1.005, between 0.998 and 1.002, between 0.999 and 1.001, or between 0.9999 and 1.0001. Note that the allele ratio calculations presented herein cannot be used to determine the ploidy state of the target individual, but may simply be a metric used to measure allele bias.
[0244] In one embodiment, once the mixture is preferentially enriched for a set of target loci, it can be sequenced using any one of the previous, current, or next-generation sequencing instruments that sequence clonal samples (samples generated from single molecules; examples include ILLUMINA GAIIx, ILLUMINA HiSeq, Life Technologies SOLiD, and 5500XL). The ratio can be assessed by sequencing through specific alleles within the targeted region. These sequencing reads can be analyzed and counted according to the allele type and, therefore, the assignment of different alleles. For variations that are one to several bases in length, allele detection is performed by sequencing, and it is essential that the sequencing reads span the allele in question to assess the allele composition of the captured molecules. The total number of captured molecules assayed for genotype can be increased by increasing the length of the sequencing read. Complete sequencing of all molecules ensures the collection of the maximum amount of data available in the enriched pool. However, sequencing is currently expensive, and methods that can measure allele distributions using a small number of sequence reads are extremely valuable. Furthermore, as read lengths increase, technical and accuracy limitations emerge regarding the maximum possible read length. While alleles of greatest utility are those one to several base pairs long, theoretically any allele shorter than the length of a sequencing read can be used. While allelic variation occurs in all forms, the examples provided herein focus on SNPs or variants contained within only a few adjacent base pairs. Larger variants, such as segmental copy number variants, can often be detected by summing these smaller variations, since the entire population of SNPs within a segment overlaps. Variants larger than a few bases, such as STRs, require special consideration and some targeted approach studies, while others do not.
[0245] There are several targeting methods that can be used to specifically isolate and enrich one or more variant locations within a genome. Generally, these methods rely on utilizing non-mutated sequences adjacent to the variant sequence. Other researchers have reported on targeting in sequencing when the substrate is maternal plasma (see, for example, Liao et al., Clin. Chem. 2011;57(1):pp.92-101). However, these methods use targeting probes that target exons and do not focus on targeting polymorphic regions of the genome. In some embodiments, the disclosed method involves using targeting probes that focus exclusively or nearly exclusively on polymorphic regions. In some embodiments, the disclosed method involves using targeting probes that focus exclusively or nearly exclusively on SNPs. In some embodiments of the present disclosure, the targeted polymorphic sites consist of at least 10% SNPs, at least 20% SNPs, at least 30% SNPs, at least 40% SNPs, at least 50% SNPs, at least 60% SNPs, at least 70% SNPs, at least 80% SNPs, at least 90% SNPs, at least 95% SNPs, at least 98% SNPs, at least 99% SNPs, at least 99.9% SNPs, or exclusively SNPs.
[0246] In an embodiment, the disclosed methods can be used to determine genotypes (the base composition of DNA at specific loci) and the relative proportions of these genotypes from a mixture of DNA molecules, which may originate from one or several genetically distinct individuals. In an embodiment, the disclosed methods can be used to determine genotypes at a set of polymorphic loci and the relative ratios of the amounts of different alleles present at those loci. In an embodiment, the polymorphic loci may consist entirely of SNPs. In an embodiment, the polymorphic loci may include SNPs, single tandem repeats, and other polymorphisms. In an embodiment, the disclosed methods can be used to determine the relative distribution of alleles at a set of polymorphic loci in a mixture of DNA, which mixture of DNA includes DNA originating from the mother and DNA originating from the fetus. In an embodiment, the joint allele distribution can be determined for a mixture of DNA isolated from the blood of a pregnant woman. In an embodiment, the allele distribution at a set of loci can be used to determine the ploidy state of one or more chromosomes for a gestating fetus.
[0247] In some embodiments, the mixture of DNA molecules may be derived from DNA extracted from multiple cells of a single individual. In some embodiments, if an individual is mosaic (germline or somatic), the original population of cells from which the DNA is derived may contain a mixture of diploid or haploid cells of the same or different genotypes. In some embodiments, the mixture of DNA molecules may be derived from DNA extracted from a single cell. In some embodiments, the mixture of DNA molecules may be derived from DNA extracted from a mixture of two or more cells from the same individual or two or more cells from different individuals. In some embodiments, the mixture of DNA molecules may be derived from DNA isolated from biological material already freed from cells, such as plasma, which is known to contain cell-free DNA. In some embodiments, the biological material may be a mixture of DNA from one or more individuals, as in the case of pregnancy, where fetal DNA has been shown to be present in the mixture. In some embodiments, the biological material may be derived from a mixture of cells found in maternal blood, some of which are of fetal origin. In some embodiments, the biological material may be cells from blood during pregnancy enriched for fetal cells.
[0248] Circularization probe Some embodiments of the present disclosure include using previously described "ligated inverted probes" (LIPs) to amplify target loci before or after amplification with non-LIP primers in the multiplex PCR methods of the present invention. LIP is a generic term meant to encompass techniques that involve creating circular DNA molecules, where the probes are designed to hybridize to regions of target DNA on either side of the target allele. Thus, with the addition of appropriate polymerase and / or ligase and appropriate conditions, buffers, and other reagents, complementary inverted regions of DNA spanning the target allele are completed, creating a circular loop of DNA that captures the information found in the target allele. LIPs are also referred to as pre-circularized probes, pre-circularizing probes, or circularization probes. LIP probes can be linear DNA molecules between 50 and 500 nucleotides in length, and in some embodiments, between 70 and 100 nucleotides in length, and in some embodiments, can be longer or shorter than described herein. Other embodiments of the present disclosure involve different implementations of LIP technology, such as Padlock probes and molecular inverse probes (MIPs).
[0249] One method for targeting specific locations for sequencing is to synthesize probes whose 3' and 5' ends anneal to the target DNA in an inverted fashion, adjacent to and flanking the target region. Adding DNA polymerase and DNA ligase then results in extension from the 3' end, adding bases to the single-stranded probe complementary to the target molecule (gap filling), and then ligating the new 3' end to the 5' end of the original probe, resulting in a circular DNA molecule that can then be isolated from background DNA. The probe ends are designed to flank the target region of interest. One aspect of this approach, commonly referred to as MIPS, has been used in conjunction with array technology to determine the nature of the filled sequence. One drawback of using MIPs in the context of measuring allele ratios is that the hybridization, circularization, and amplification steps do not occur at equivalent rates for different alleles at the same locus. This results in measured allele ratios that do not represent the actual allele ratios present in the original mixture.
[0250] In some embodiments, circularization probes are constructed so that the region of the probe designed to hybridize upstream of the target polymorphic locus and the region of the probe designed to hybridize downstream of the target polymorphic locus are covalently connected through a non-nucleic acid backbone. This backbone can be any biocompatible molecule or combination of biocompatible molecules. Some examples of potential biocompatible molecules are poly(ethylene glycol), polycarbonate, polyurethane, polyethylene, polypropylene, sulfone polymers, silicone, cellulose, fluoropolymers, acrylic compounds, styrene block copolymers, and other block copolymers.
[0251] In certain embodiments of the present disclosure, this technique is modified to facilitate sequencing as a means of examining intrasequence filling. To preserve the original allele ratios of the original sample, at least one important consideration must be taken into account: The variable positions between different alleles in the gap-fill region should not be too close to the probe binding site, as this may result in biased initiation by the DNA polymerase, leading to differentiation of the variants. Another consideration is the possibility of additional variation in the probe binding site that correlates with variants in the gap-fill region, which may result in unequal amplification from different alleles. In certain embodiments of the present disclosure, the 3' and 5' ends of the pre-circularization probe are designed to hybridize to bases that are one or a few positions away from the variant position (polymorphic site) of the target allele. The number of bases between the polymorphic site (SNP or other type) and the base to which the 3' and / or 5' ends of the pre-circularization probe are designed to hybridize can be 1, 2, 3, 4, 5, 6, 7-10, 11-15, 16-20, 20-30, or 30-60 bases. Forward and reverse primers can be designed to hybridize with different numbers of bases away from the polymorphic site. Current DNA synthesis techniques can be used to generate large numbers of circularization probes, potentially allowing for the generation and pooling of very large numbers of probes, enabling the simultaneous interrogation of many loci. Work with over 300,000 probes has been reported. Two articles discussing methods involving circularized probes that can be used to measure genomic data of a target individual include Porreca et al., Nature Methods, 2007, Vol. 4(11), pp. 931-936; and Turner et al., Nature Methods, 2009, Vol. 6(5), pp. 315-316. The methods described in these articles can be used in combination with other methods described herein.Certain steps of the methods from these two articles can be used in combination with other steps from other methods described herein.
[0252] In some embodiments of the methods disclosed herein, the genetic material of a target individual is optionally amplified, then hybridized with a pre-circularized probe, gap-filled to fill in the bases between the two ends of the hybridized probe, and ligated to form a circularized probe, which is then amplified using, for example, rolling circle amplification. Once the genetic information of the desired target allele is captured using a properly designed circularized oligonucleotide probe, such as a LIP system, the genetic sequence of the circularized probe can be measured to provide the desired sequence data. In some embodiments, a properly designed oligonucleotide probe can be directly circularized on the unamplified genetic material of a target individual and then amplified. It should be noted that several amplification procedures, including rolling circle amplification, MDA, or other amplification protocols, can be used to amplify the original genetic material or circularize the LIP. Different methods can be used to measure the genetic information on the target genome, for example, high-throughput sequencing, Sanger sequencing, other sequencing methods, hybridization capture, circularization capture, multiplex PCR, other hybridization methods, and combinations thereof.
[0253] After using one or a combination of the above methods, informatics-based methods, such as the PARENTAL SUPPORT™ method, together with appropriate genetic measurement to measure the genetic material of an individual, it can then be used to determine the ploidy state of one or more chromosomes in the individual, and / or the genetic state of one or a set of alleles (particularly, the alleles that correlate with the target disease or genetic state).It should be noted that the use of LIP has been reported to multiplex capture genetic sequences and then use sequencing to determine genotype.However, the sequencing data generated by LIP-based strategies has not been used to amplify the genetic material found in single cells, a small number of cells, or extracellular DNA to determine the ploidy state of the target individual.
[0254] The application of informatics-based methods for determining the ploidy status of individuals from genetic data measured by hybridization arrays, such as ILLUMINA INFINIUM arrays or AFFYMETRIX gene chips, has been described in other references in this document.However, the method described herein represents an improvement over the methods previously described in the literature.For example, the LIP-based method followed by high-throughput sequencing unexpectedly provides better genotype data because this method has better multiplexing capacity, better capture specificity, better uniformity, and less allele bias.Higher multiplexing allows more alleles to be targeted, resulting in more accurate results.Higher uniformity allows more target alleles to be measured, resulting in more accurate results.Lower allele bias rate reduces the rate of false calls and provides more accurate results.More accurate results improve clinical outcomes and provide better medical care.
[0255] It is important to note that LIP can be used as a method to target specific loci in a sample of DNA for genotyping by methods other than sequencing. For example, LIP can be used to target DNA for genotyping using SNP arrays or other DNA or RNA-based microarrays.
[0256] Ligation-mediated PCR Ligation-mediated PCR can be used to amplify target loci before or after PCR amplification using unligated primers. Ligation-mediated PCR is a PCR method used to preferentially enrich a sample of DNA by amplifying one or more loci in a mixture of DNA, the method comprising the steps of obtaining a set of primer pairs, each primer of the pair containing a target-specific sequence and a non-target sequence, preferably the target-specific sequences are designed to anneal to target regions, one upstream of the polymorphic site and one downstream of the polymorphic site, and the target-specific sequences are within 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, The steps may be separated by 1 to 30, 31 to 40, 41 to 50, 51 to 100, or more than 100, and include polymerizing DNA from the 3' end of the upstream primer to fill in the single-stranded region between it and the 5' end of the downstream primer with a nucleotide complementary to the target molecule; ligating the last polymerized base of the upstream primer to the adjacent 5' base of the downstream primer; and amplifying only the polymerized and ligated molecules using a non-target sequence containing the 5' end of the upstream primer and the 3' end of the downstream primer. Primer pairs for different targets can be mixed in the same reaction. The non-target sequence serves as a universal sequence, so all successfully polymerized and ligated primer pairs can be amplified using a single pair of amplification primers.
[0257] Hybridization capture In some embodiments, the disclosed methods can include using any of the following hybridization capture methods in addition to amplifying target loci using multiplex PCR: Preferential enrichment of a specific set of sequences in a target genome can be achieved in a number of ways. Elsewhere in this document, we describe how LIP can be used to target a specific set of sequences; however, in all of these applications, other targeting and / or preferential enrichment methods can be used equally well for the same purpose. One example of another targeting method is hybridization capture. Some examples of commercial hybridization capture technologies include AGILENT's SURE SELECT and ILLUMINA's TruSeq. Hybridization capture involves hybridizing a set of oligonucleotides complementary or nearly complementary to the desired target sequences with a mixture of DNA and then physically separating them from the mixture. Once the desired sequences hybridize to the targeting oligonucleotides, physically removing the targeting oligonucleotides also removes the target sequences. After hybridized oligos are removed, they can be heated above their melting temperature and amplified.Some methods for physically removing targeting oligonucleotides are by covalently binding targeting oligos to solid support, for example, magnetic beads or chip.Another method for physically removing targeting oligonucleotides is by covalently binding targeting oligonucleotides to molecular moieties that have strong affinity for other molecular moieties.An example of such molecular pair is, for example, biotin and streptavidin, which are used in SURE SELECT.Therefore, the sequence of the target can be covalently attached to biotin molecule, and after hybridization, the biotinylated oligonucleotide that the sequence of the target is hybridized with can be pulled down using the solid support that streptavidin is added to.
[0258] Hybrid capture involves hybridizing a probe complementary to the target of interest with a target molecule. Hybrid capture probes were originally developed to target and enrich large portions of the genome with relative uniformity between targets. In this application, it was important that all targets amplified have sufficient uniformity to allow all regions to be detected by sequencing, but attention was not paid to maintaining the allele ratio in the original sample. After capture, the alleles present in the sample can be determined by direct sequencing of the captured molecules. These sequencing reads can be analyzed and counted according to the type of allele. However, using current technology, the allele distribution of the measured captured sequences generally does not represent the original allele distribution.
[0259] In some embodiments, allele detection is carried out by sequencing.In order to capture the identity of alleles at polymorphic sites, it is essential that sequencing reads cover the allele in question in order to evaluate the allele composition of captured molecules.Because the length of captured molecules often varies, when sequencing, it is necessary to sequence the entire molecule to ensure that the mutation position overlaps.However, due to cost considerations and technical limitations on the maximum possible length and accuracy of sequencing reads, sequencing the entire molecule is not feasible.In some embodiments, the read length can be increased from about 30 bases to about 50 bases or about 70 bases, which can significantly increase the number of reads that overlap with the mutation position in the target sequence.
[0260] Another way to increase the number of reads that examine the position of interest is to reduce the length of the probe, as long as it does not cause bias in the underlying enriched allele.The length of the synthesized probe should be long enough so that two probes designed to hybridize with two different alleles found in one locus can hybridize with the various alleles in the original sample with approximately equal affinity.Currently, the method known in the art generally describes a probe longer than 120 bases.In the current embodiment, when the allele is one or a few bases, the capture probe can be less than about 110 bases, less than about 100 bases, less than about 90 bases, less than about 80 bases, less than about 70 bases, less than about 60 bases, less than about 50 bases, less than about 40 bases, less than about 30 bases, and less than about 25 bases, and this amount is sufficient to ensure equal enrichment from all alleles. When the mixture of DNA to be enriched using hybrid capture techniques is a mixture containing free-floating DNA isolated from blood, e.g., maternal blood, the average length of the DNA is fairly short, generally less than 200 bases. Using shorter probes increases the likelihood that the hybrid capture probe will capture the desired DNA fragment. Larger variations may require longer probes. In some embodiments, the variation of interest is one (SNP) to several bases in length. In some embodiments, targeted regions within the genome can be preferentially enriched using hybrid capture probes, where the hybrid capture probes are less than 90 bases in length, and may be less than 80 bases, 70 bases, 60 bases, 50 bases, 40 bases, 30 bases, or 25 bases in length. In certain embodiments, to increase the likelihood that the desired allele will be sequenced, the length of the probe designed to hybridize to the region adjacent to the location of the polymorphic allele can be reduced from more than 90 bases to about 80 bases, or to about 70 bases, or to about 60 bases, or to about 50 bases, or to about 40 bases, or to about 30 bases, or to about 25 bases.
[0261] To enable capture, there is a minimum overlap between the synthesized probe and the target molecule. This synthesized probe can be as short as possible, but still be larger than this minimum required overlap. The effect of using a shorter probe length to target a polymorphic region is that more molecules overlap with the target allele region. The fragmentation state of the original DNA molecule also affects the number of reads that overlap with the target allele. Some DNA samples, such as plasma samples, are already fragmented due to biological processes occurring in vivo. However, samples with longer fragments benefit from fragmentation prior to sequencing library preparation and enrichment. When both the probe and fragments are short (approximately 60–80 bp), maximum specificity can be achieved with relatively few sequence reads that cannot overlap with the critical region of interest.
[0262] In some embodiments, hybridization conditions can be adjusted to maximize the uniformity of capture of different alleles present in the original sample. In some embodiments, the hybridization temperature is lowered to minimize differences in hybridization bias between alleles. Methods known in the art avoid using lower temperatures for hybridization because lowering the temperature has the effect of increasing hybridization between the probe and unintended targets. However, if the goal is to preserve the allele ratio with maximum fidelity, a method using a lower hybridization temperature will result in the most accurate allele ratio, despite the fact that current technical teachings avoid this approach. The hybridization temperature can also be increased to require greater overlap between the target and the synthesized probe, so that only targets with substantial overlap of the target region are captured. In some embodiments of the present disclosure, the hybridization temperature is reduced from the normal hybridization temperature to about 40°C, to about 45°C, to about 50°C, to about 55°C, to about 60°C, to about 65°C, or to about 70°C.
[0263] In some embodiments, hybrid capture probes can be designed so that the region of the capture probe that has DNA complementary to the DNA found in the region adjacent to the polymorphic allele is not immediately adjacent to the polymorphic site.Instead, the capture probe can be designed so that the region of the capture probe that is designed to hybridize with the DNA adjacent to the target polymorphic site is separated from the part of the capture probe that contacts the polymorphic site by van der Waals by a small distance equal to one or a few bases in length.In some embodiments, hybrid capture probes are designed to hybridize with the region adjacent to the polymorphic allele but not across it, and are referred to as adjacent capture probes.The length of the adjacent capture probe can be less than about 120 bases, less than about 110 bases, less than about 100 bases, less than about 90 bases, and can be less than about 80 bases, less than about 70 bases, less than about 60 bases, less than about 50 bases, less than about 40 bases, less than about 30 bases, or less than about 25 bases. The region of the genome targeted by the flanking capture probe may be separated from the polymorphic locus by 1 base pair, 2 base pairs, 3 base pairs, 4 base pairs, 5 base pairs, 6 base pairs, 7 base pairs, 8 base pairs, 9 base pairs, 10 base pairs, 11-20 base pairs, or more than 20 base pairs.
[0264]
[0013] A description of targeted capture-based disease screening tests using targeted sequence capture. Custom targeted sequence capture, such as those currently offered by AGILENT (SURE SELECT), ROCHE-NIMBLEGEN, or ILLUMINA, can be used. Capture probes can be custom designed to ensure capture of various types of mutations. For point mutations, one or more probes overlapping the point mutation should be sufficient to capture and sequence the mutation.
[0265] For small insertions or deletions, one or more probes that overlap the mutation may be sufficient to capture and sequence the fragment containing the mutation.Hybridization may generally be less efficient than the probe-limited capture efficiency designed to reference genome sequences.To ensure the capture of fragments containing mutations, two probes can be designed, one that matches the normal allele and one that matches the mutant allele.A longer probe can enhance hybridization.Multiple overlapping probes can enhance capture.Finally, by placing probes immediately adjacent but not overlapping, mutations may allow for relatively similar capture efficiencies for the normal allele and the mutant allele.
[0266] For simple tandem repeats (STRs), probes that overlap these highly variable sites are unlikely to successfully capture the fragments. To enhance capture, probes can be placed close to, but not overlapping, the variable sites. The fragments can then be sequenced as usual to reveal the length and composition of the STRs.
[0267] For large deletions, a series of overlapping probes, a common approach currently used in exon capture systems, can work. However, using this approach, it can be difficult to determine whether an individual is heterozygous. Targeting and evaluating SNPs within the captured region can potentially indicate loss of heterozygosity across that region, indicating that the individual is a carrier. In one embodiment, non-overlapping or singleton probes can be placed across the potentially deleted region, and the number of captured fragments can be used as a measure of heterozygosity. If an individual has a large deletion, it is expected that half the number of fragments will be available for capture compared to a non-deleted (diploid) reference locus. Therefore, the number of reads obtained from the deleted region should be approximately half the number of reads obtained from a normal diploid locus. By summing and averaging the sequencing read depth from multiple singleton probes across the potentially deleted region, the signal can be enhanced and diagnostic confidence can be improved. The two approaches can also be combined: targeting SNPs to identify loss of heterozygosity and using multiple singleton probes to obtain a quantitative measure of the amount of underlying fragment from that locus. Either or both of these strategies can be combined with other strategies to better achieve the same results.
[0268] During the test, the detection of male fetus cfDNA, which is indicated by the presence of Y chromosome fragments that are captured and sequenced in the same test, and either X-linked dominant mutations that mother and father are not affected or dominant mutations that mother is not affected, indicates that the risk to the fetus is increased.The detection of two mutant recessive alleles in the same gene in unaffected mother means that the fetus inherits a mutant allele from the father and potentially a second mutant allele from the mother through inheritance.In all cases, follow-up testing by amniocentesis or chorionic villus sampling can also be indicated.
[0269] Targeted capture-based disease screening tests can be combined with targeted capture-based non-invasive prenatal diagnostic tests for aneuploidy. There are several ways to reduce the variability of depth of read (DOR): for example, primer concentration can be increased, longer targeted amplification probes can be used, or more STA cycles can be performed (e.g., more than 25, more than 30, more than 35, or even more than 40).
[0270] Representative methods for determining the number of DNA molecules in a sample Described herein is a method for determining the number of DNA molecules in a sample by generating a uniquely identified molecule for each original DNA molecule in the sample during a first round of DNA amplification. Described herein is a procedure for achieving the above objectives, followed by single molecule or clonal sequencing.
[0271] This method involves targeting one or more specific loci, and generating tagged copies of original molecules, so that each target locus has a unique tag, and when this barcode is sequenced using clonal sequencing or single molecule sequencing, it can be distinguished from each other.Each unique sequenced barcode represents a unique molecule in original sample.At the same time, sequencing data is used to identify the loci from which molecules originate.This information can be used to determine the number of unique molecules in original sample for each locus.
[0272] This method can be used for any application that requires quantitative evaluation of the number of molecules in original sample.In addition, the number of unique molecules of one or more targets can be related to the number of unique molecules of one or more other targets to determine relative copy number, allele distribution or allele ratio.Alternatively, the number of copies detected from various targets can be modeled by distribution to identify the most likely number of copies of original target.Applications include but are not limited to: detection of insertions and deletions, such as those found in carriers of Duchenne muscular dystrophy; quantification of deletions or duplications of chromosome segments, such as those observed in copy number variants; chromosome copy number of samples from born individuals; chromosome copy number of samples from unborn individuals, such as embryos or fetuses.
[0273] Said method can be combined with the evaluation of the simultaneous variation contained in sequence targeting.This can be used to determine the number of molecules that represent each allele in original sample.This copy number method can be combined with the evaluation of SNP or other sequence variations to determine the chromosome copy number of born and unborn individuals;Distinguish and quantify the copies from the locus that has short sequence variations but can be amplified from multiple target regions by PCR, for example, for the purpose of detecting carriers of spinal muscular atrophy;Determine the copy number of various sources of molecules from the sample that consists of a mixture of different individuals, for example, for the purpose of detecting fetal aneuploidy from the floating DNA obtained from maternal plasma.
[0274] In one embodiment, a method involving a single target locus may include one or more of the following steps: (1) designing a standard pair of oligomers for PCR amplification of a specific locus; (2) adding a specific base sequence with no or minimal complementarity to the target locus or genome to the 5' end of one of the target-specific oligomers during synthesis. This sequence, called the tail, is a known sequence used for subsequent amplification and is followed by a sequence of random nucleotides. These random nucleotides comprise a random region. The random region comprises randomly generated nucleic acid sequences that stochastically differ between each probe molecule. Thus, after synthesis, the tailed oligomer pool consists of a population of oligomers starting with a known sequence, followed by an unknown sequence that differs between molecules, followed by the target-specific sequence; and (3) performing one round of amplification (denaturation, annealing, extension) using only the tailed oligomers. (4) adding an exonuclease to the reaction to effectively terminate the PCR reaction and incubating the reaction at an appropriate temperature to remove any forward single-stranded oligos that did not anneal to the template and extend them to form double-stranded products; (5) incubating the reaction at an elevated temperature to denature the exonuclease and eliminate its activity; (6) adding a new oligonucleotide complementary to the tail of the oligomer used in the first reaction along with other target-specific oligomers to the reaction to allow PCR amplification of the products generated in the first round of PCR; (7) continuing amplification to generate sufficient product for downstream clonal sequencing; and (8) measuring the amplified PCR products, which have a sufficient number of bases in sequence, by a number of methods, such as clonal sequencing.
[0275] In some embodiments, the disclosed method involves targeting multiple loci in parallel or otherwise. Primers for different target loci can be independently generated and mixed to create multiplex PCR pools. In some embodiments, the original sample can be divided into subpools, and different loci can be targeted in each subpool, followed by recombination and sequencing. In some embodiments, before dividing the pool, tagging and multiple amplification cycles can be performed to ensure efficient targeting of all targets, followed by division and improvement, and then amplification can be continued by using a smaller set of primers in the divided pools.
[0276] One example of the application that this technology is particularly useful for is non-invasive prenatal aneuploidy diagnosis, and the allele ratio at a given locus or the allele distribution at several loci can be used to determine the number of copies of chromosomes present in fetus.In this situation, it is desirable to amplify the DNA present in initial sample while maintaining the relative amount of various alleles.In some cases, particularly when the amount of DNA present is very small, for example, less than 5,000 copies of genome, less than 1,000 copies of genome, less than 500 copies of genome, and less than 100 copies of genome, a phenomenon called bottlenecking may occur.This means that there are a small number of copies of any given allele in initial sample, and as a result of amplification bias, the ratio of these alleles that the amplified DNA pool has is significantly different from the ratio of these alleles that the initial DNA mixture has. By applying a unique or nearly unique set of barcodes to each strand of DNA prior to standard PCR amplification, it is possible to eliminate n-1 copies of DNA from a set of n identical molecules of sequenced DNA that originate from the same original molecule.
[0277] For example, consider a heterozygous SNP in an individual's genome and a mixture of DNA from that individual where 10 molecules of each allele are present in the original DNA sample. After amplification, there may be 100,000 molecules of DNA corresponding to that locus. Due to stochastic processes, the ratio of DNA may be anywhere from 1:2 to 2:1, but because each original molecule was tagged with a unique tag, it is possible to determine that the DNA in the amplified pool originated from exactly 10 molecules of DNA from each allele. This method therefore provides a more accurate measure of the relative abundance of each allele than methods that do not use this technique. This method provides accurate data for methods where it is desirable to minimize the relative abundance of alleles.
[0278] Association of sequenced fragments with target loci can be achieved in several ways. In one embodiment, a molecular barcode corresponding to the target sequence, as well as a sufficient number of unique bases, are linked to the target fragment to obtain a sequence of sufficient length to allow unambiguous identification of the target locus. In another embodiment, the molecular barcoding primer containing the randomly generated molecular barcode can also contain a locus-specific barcode (locus barcode) that identifies the target with which it is associated. This locus barcode is identical among all molecular barcoding primers for each individual target, and therefore among all resulting amplification products, but is different from all other targets. In one embodiment, the tagging method described herein can be combined with a one-sided nesting protocol.
[0279] In one embodiment, the design and generation of molecular barcoding primers can be implemented as follows: the molecular barcoding primer may consist of a sequence that is not complementary to the target sequence, followed by a random molecular barcode region, followed by a target-specific sequence. The sequence 5' of the molecular barcode can be used for partial sequence PCR amplification and may contain sequences useful in converting the amplified products into a library for sequencing. Random molecular barcode sequences can be generated in a number of ways. In a preferred method, molecularly tagged primers are synthesized to contain all four bases for the reaction during synthesis of the barcode region. All of the bases or various combinations of bases can be specified using the IUPAC DNA ambiguity code. In this way, the population of synthesized molecules contains a random mixture of sequences within the molecular barcode region. The length of the barcode region determines how many primers contain unique barcodes. The number of unique sequences is determined by N L The length of the barcode region is related to the length of the barcode region as N, where N is the number of bases, typically 4, and L is the length of the barcode. A 5-base barcode can result in up to 1024 unique sequences, and an 8-base barcode can result in 65536 unique barcodes. In some embodiments, the DNA can be measured by a sequencing method, with the sequence data representing the sequence of a single molecule. This can include direct sequencing of a single molecule, or a method referred to herein as clonal sequencing, in which a single molecule is amplified to form a clone detectable by a sequencing instrument, but still representing a single molecule.
[0280] Representative Methods and Reagents for Quantifying Amplification Products Quantitation of specific nucleic acid sequences of interest is typically performed by quantitative real-time PCR techniques such as TAQMAN (LIFE TECHNOLOGIES), INVADER probes (THIRD WAVE TECHNOLOGIES), etc. Such techniques have many drawbacks, including a limited ability for parallel simultaneous analysis of multiple sequences (multiplexing) and the ability to generate accurate quantitative data only within a narrow range of amplification cycles (e.g., the logarithm of PCR amplification yield versus cycle number is within the linear range). DNA sequencing technologies, particularly high-throughput next-generation sequencing technologies (often referred to as massively parallel sequencing technologies), such as those employed by MYSEQ (ILLUMINA), HISEQ (ILLUMINA), ION TORRENT (LIFE TECHNOLOGIES), GENOME ANALYZERILX (ILLUMINA), and GSFLEX+ (ROCHE454), can be used to quantitatively measure the number of copies of a sequence of interest present in a sample, thereby providing quantitative information about the starting material, such as copy number or transcription level. High-throughput gene sequencers use barcoding (i.e., tagging samples with distinctive nucleic acid sequences) to identify specific samples from distinct populations, thereby enabling the simultaneous analysis of multiple samples in a single run of the DNA sequencer. The number of times a given region of genome in a library preparation (or other nucleic acid preparation of interest) is sequenced (number of reads) will be proportional to the copy number (or, in the case of cDNA-containing preparations, expression level) of that sequence in the genome of interest. However, genetic library preparation and sequencing (and similar genome-derived preparations) can introduce many biases that interfere with obtaining accurate quantitative reads of nucleic acid sequences of interest. For example, different nucleic acid sequences may amplify with different efficiencies during the nucleic acid amplification step that occurs during genetic library preparation or sample preparation.
[0281] Problems associated with varying amplification efficiencies can be alleviated by using certain embodiments of the present invention. The present invention includes various methods and compositions related to the use of content criteria during the amplification process, which can be used to improve quantification accuracy. As described herein and, inter alia, in U.S. Pat. Nos. 8,008,018; 7,332,277; WO 2012 / 078792 A2; and WO 2011 / 146632 A1, the present invention is particularly useful in the area of detecting fetal aneuploidy by analyzing free-floating fetal DNA in maternal blood. These patents are incorporated herein by reference in their entireties. Embodiments of the present invention are also useful for detecting aneuploidy in in vitro-generated embryos. Commercially important detectable aneuploidies include aneuploidy of human chromosomes 13, 18, 21, X, and Y.
[0282] Embodiments of the invention can be used with human or non-human nucleic acids and are applicable to both animal and plant-derived nucleic acids. Embodiments of the invention can also be used to detect and / or quantitate alleles associated with other genetic disorders characterized by deletions or insertions. Deletion-containing alleles may be detected in individuals suspected of being carriers of the allele of interest.
[0283] One embodiment of the present invention includes standards present in known amounts (relative or absolute). For example, consider a genetic library created from a genetic source in which chromosome 8 (containing locus A) is diploid and chromosome 21 (containing locus B) is triploid. A genetic library can be generated from a sample containing sequences in amounts that are a function of the number of chromosomes present in the sample, e.g., 200 copies of locus A and 300 copies of locus B. However, if locus A amplifies with a much higher efficiency than locus B, there may be 60,000 copies of A amplicon and 30,000 copies of B amplicon after PCR, thus obscuring the true chromosome copy number of the initial genomic sample when analyzed by high-throughput DNA sequencing (or other quantitative nucleic acid detection technology). To alleviate this problem, a standard sequence for locus A is employed, where the standard sequence is amplified with substantially the same efficiency as locus A. Similarly, a standard sequence for locus B is formed, where the standard sequence is amplified with substantially the same efficiency as locus B. Prior to PCR (or other amplification technique), a reference sequence for locus A and a reference sequence for locus B are added to the mixture. These reference sequences are present in known amounts, either relative or absolute. Thus, under the same set of conditions, if a 1:1 mixture of reference sequence A and reference sequence B were added to the mixture in the previous example (prior to amplification), 3,000 copies of reference A amplicon would be produced and 1,000 copies of reference B amplicon would be produced, indicating that locus A is amplified three times more efficiently than locus B.
[0284] One or more selected regions of the genome containing SNPs (or other polymorphisms) of interest can be specifically amplified and subsequently sequenced. This target-specific amplification can be performed during the creation of a genetic library for sequencing. The library can contain many target amplified regions. In some embodiments, the library contains at least 10; 100; 500; 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000; 75,000; or 100,000 regions of interest. Examples of such libraries are described herein and can also be found in U.S. Patent Publication No. 2012 / 0270212, filed November 18, 2011, which is incorporated herein by reference in its entirety.
[0285] Many high-throughput DNA sequencing techniques require modification of the starting genetic material for library formation, such as the ligation of universal priming sites and / or barcodes, to facilitate the clonal amplification of small nucleic acid fragments prior to subsequent sequencing reactions. In some embodiments, one or more reference sequences are added during gene library formation or to precursor components of the gene library prior to library amplification. Reference sequences can be selected to mimic (yet be distinguishable based on nucleotide base sequence) the target genomic fragments prepared for sequencing by high-throughput gene sequencing techniques. In one embodiment, the reference sequence may be identical to the target genomic fragment except for 1, 2, 3, 4-10, or 11-20 nucleotides. In some embodiments, if the target gene sequence contains a SNP, the reference sequence may be identical to the SNP except for the nucleotide at the polymorphic base, which can be selected to be one of four nucleotides not observed at the natural site. Reference sequences can be used for highly multiplexed analysis of multiple target loci (e.g., polymorphic loci). Standard sequences can be added in known amounts (relative or absolute) during the library formation (pre-amplification) process to provide a reference metric for greater accuracy in determining the amount of target sequences of interest in an analytical sample. The combination of information about the known amounts of standard sequences used, along with information about the ploidy level of a sequencing library formed from a genome with previously characterized ploidy levels, e.g., where all autosomes are known to be diploid, can be used to calibrate the amplification characteristics of each standard sequence for its corresponding target sequence and to account for variations between batch mixtures containing multiple standard sequences. Given that it is often necessary to simultaneously analyze a large number of loci, it is useful to generate mixtures containing many sets of standard sequences. Embodiments of the present invention include mixtures containing multiple standard sequences. Ideally, the amount of each standard sequence in a mixture would be known with high precision. However, this ideal is extremely difficult to achieve.This is because, in practice, there is significant variation in the amount of each standard sequence in a mixture, especially in mixtures containing many different synthetic oligonucleotides. This variation can arise from many sources, including batch-to-batch variation in in vitro oligonucleotide synthesis efficiency, volumetric inaccuracies, and pipetting variations. Furthermore, this variation can occur even between batches that theoretically contain the exact same set of standard sequences in the exact same amounts. Therefore, it is useful to independently calibrate each batch of standard sequences. Standard sequence batches can be calibrated against a reference genome of known chromosomal composition. Standard sequence batches can be calibrated by sequencing the batch with minimal or no amplification steps included in the sequencing protocol. Embodiments of the present invention include calibrated mixtures of different standard sequences. Other embodiments of the present invention include methods for calibrating mixtures of different standard sequences and calibrated mixtures of different standard sequences produced by these methods.
[0286] Various embodiments of the target standard sequence mixture and their methods of use may include at least 10, 100, 500, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 or more standard sequences, as well as various intermediate numbers. The number of standard sequences may be the same as the number of target sequences selected for analysis during the generation of a target library for DNA sequencing. However, in some embodiments, it may be advantageous to use a number of standard sequences that is less than the number of target regions in the library to be constructed. It may be advantageous to use a smaller number to avoid reaching the limit of the sequencing capacity of the high-throughput DNA sequencer used. The number of standard sequences can be 50% or less of the number of target regions, 40% or less of the number of target regions, 30% or less of the number of target regions, 20% or less of the number of target regions, 10% or less of the number of target regions, 5% or less of the number of target regions, 1% or less of the number of target regions, and various intermediate numbers. For example, if a gene library is generated using 15,000 pairs of primers targeting specific SNP-containing loci, an appropriate mixture containing 1500 standard sequences corresponding to 1500 of the 15,000 target loci can be added prior to the amplification step of library construction.
[0287] The amount of standard sequence added during library construction can vary widely between individual embodiments. In some embodiments, the amount of each standard sequence can be approximately the same as the predicted amount of target sequence present in the genomic material sample used for library preparation. In other embodiments, the amount of each standard sequence can be greater or less than the predicted amount of target sequence present in the genomic material sample used for library preparation. The initial relative amounts of target and standard sequences are not critical for purposes of the present invention, but are preferably in the range of 100-fold greater to 100-fold less than the amount of target sequence present in the genomic material sample used for library preparation. Excessive amounts of standard sequence can use too much of the sequencing capacity of a DNA sequencer in a given number of runs of the instrument. Using too little standard sequence will result in insufficient data to be useful in analyzing variations in amplification efficiency.
[0288] The reference sequence can be selected to have a nucleotide base sequence highly similar to the region of interest to be amplified; preferably, the reference sequence has exactly the same primer binding sites as the genomic region being analyzed, i.e., the "target sequence." The reference sequence must be distinguishable from the corresponding target sequence at a given locus. For convenience, this distinguishable region of the reference sequence is referred to as the "marker sequence." In some embodiments, the marker sequence region of the target sequence contains a polymorphic region, e.g., a SNP, and primer binding regions may be located on either side. The reference sequence can be selected to closely match the GC content of the corresponding target sequence. In some embodiments, the primer binding regions of the reference sequence are flanked by universal priming sites. These universal priming sites are selected to match the universal priming sites used in the genomic library being analyzed. In other embodiments, the reference sequence lacks universal priming sites, and universal priming sites are added during library formation. The reference sequence is typically provided in single-stranded form. The reference sequence is defined relative to the corresponding target sequence, and the target sequence is amplified using sequence-specific reagents. In some embodiments, the target sequence contains a polymorphism of interest, e.g., a SNP, deletion, or insertion, present in the nucleic acid sample for analysis. The reference sequence is a synthetic polynucleotide that is similar in nucleotide base sequence to the target sequence but is still distinguishable from the target sequence due to at least one nucleotide base difference, thereby providing a mechanism for distinguishing an amplification product sequence derived from the reference sequence from an amplification product sequence derived from the target sequence. The reference sequence is selected to have substantially the same amplification characteristics as the corresponding target sequence when amplified with the same set of amplification reagents, e.g., PCR primers. In some embodiments, the reference sequence can have the same primer sequence binding site as the corresponding target sequence. In other embodiments, the reference sequence can have a different primer sequence binding site than the corresponding target sequence. In some embodiments, the reference sequence can be selected to generate an amplification product of the same length as the length of the amplification product derived from the corresponding target sequence.In other embodiments, the reference sequence can be selected to generate an amplification product of a length that differs slightly from the length of the amplification product from the corresponding target sequence.
[0289] After the amplification reaction is complete, the library is sequenced using a high-throughput DNA sequencer, where individual molecules are clonally amplified and sequenced. The number of sequence reads for each allele of the target sequence is counted, as well as the number of sequence reads for the reference sequence corresponding to the target sequence. This process is also repeated for at least another pair of target and reference sequences. For example, for locus A, allele 1 of locus A is X 1 , and so on. A1 Read, X for allele 2 at locus A A2 Reads are generated and X is identified against the reference sequence A. AC Suppose reads are generated. For each target locus, X AC for (X A1 Plus X A2 ) is determined. As discussed previously, this process can be performed on a reference genome, e.g., a genome in which all chromosomes are known to be diploid. This process can be repeated multiple times to obtain a large number of read values and measure the average read count and standard deviation of the read count. This process is performed on a mixture containing many different reference sequences corresponding to different loci. (1)X A1 Plus X A2 corresponds to a known number of chromosomes, e.g., 2 for a normal human female genome, and (2) the standard sequences have similar amplification (and detection) properties to their corresponding native loci, so the relative amounts of different standard sequences in the multiplex standard mixture can be determined. The calibrated multiplex standard sequence mixture can then be used to adjust for variations in amplification efficiency between different loci in a multiplex amplification reaction.
[0290] Other embodiments of the present invention include methods and compositions for measuring the copy number of a specific gene of interest, such as a mutant gene characterized by large deletions that can interfere with duplication and quantification by sequencing. Sequencing can have difficulty detecting alleles with such deletions. The use of an amplification process that includes a reference sequence can reduce this problem.
[0291] In one embodiment of the present invention, the target sequence for analysis is a wild-type (i.e., functional) and mutant form of a gene characterized by a deletion. A representative example of such a gene is SMN1, a deletion-carrying allele responsible for the genetic disease spinal muscular atrophy (SMA). High-throughput gene sequencing techniques are useful for detecting individuals carrying mutant forms of the gene. Application of such techniques to the detection of deletion mutations can be problematic, particularly because of the lack of sequence observed during sequencing (as opposed to the detection of simple point mutations or SNPs). In such an embodiment, the following are used: (1) a pair of amplification primers specific to the gene of interest that amplify the gene of interest (or a portion thereof) but do not significantly amplify the mutant allele; (2) a reference sequence that corresponds to the wild-type allele of the gene of interest (i.e., the target sequence) but differs by at least one detectable nucleotide base; (3) a pair of amplification primers specific to a second target sequence that serves as a reference sequence; and (4) a reference sequence that corresponds to the reference sequence.
[0292] In one embodiment of the present invention, a method for determining the copy number of a gene of interest is provided, where the gene of interest has one intended allele containing a deletion. The method can employ amplification reagents, e.g., PCR primers, specific to the gene of interest in that they amplify at least a portion of the gene of interest, or the entire gene of interest, or a region adjacent to the gene of interest, without amplifying the deletion-containing allele of the gene of interest. Furthermore, the method employs a reference sequence corresponding to the gene of interest, where the reference sequence differs from the gene of interest by at least one nucleotide base (so that the sequence of the reference sequence is easily distinguishable from the native gene of interest). Typically, the reference sequence contains the same primer binding site as the gene of interest, thereby minimizing any amplification differences between the gene of interest and the reference sequence corresponding to the gene of interest. The reaction also includes amplification reagents specific to the reference sequence. The reference sequence is a sequence of known (or at least assumed to be known) copy number in the genome being analyzed. The reaction further includes a reference sequence corresponding to the reference sequence. Typically, a standard sequence corresponding to a reference sequence contains the same primer binding sites as the reference sequence, thereby minimizing any amplification differences between the reference sequence and the standard sequence corresponding to the reference sequence.
[0293] Representative nucleic acid samples In some embodiments, the genetic sample may be prepared and / or purified. There are several standard procedures known in the art to accomplish such ends. In some embodiments, the sample may be centrifuged to separate the various layers. In some embodiments, filtration may be used to isolate the DNA. In some embodiments, the preparation of the DNA may involve amplification, separation, chromatographic purification, liquid-liquid separation, isolation, preferential enrichment, preferential amplification, targeted amplification, or any of several other techniques known in the art or described herein.
[0294] In some embodiments, the methods disclosed herein can be used in situations where very little DNA is present, such as in in vitro fertilization or forensic situations where one or a few cells (typically fewer than 10, 20, or 40 cells) are available. In these embodiments, the methods disclosed herein are useful for making ploidy calls from small amounts of DNA where other DNA is not contaminating but the small amount of DNA makes ploidy calling very difficult. In some embodiments, the methods disclosed herein can be used in situations where the target DNA is contaminated with DNA from another individual, such as maternal blood in the context of prenatal testing, paternity testing, or the product of conception testing. Some other situations where these methods are particularly advantageous are cancer testing, where only one or a few cells are present among a larger number of normal cells. The genetic measurements used as part of these methods can be performed on any sample containing DNA or RNA, including, but not limited to, blood, plasma, bodily fluids, urine, hair, tears, saliva, tissue, skin, fingernails, blastomeres, embryos, amniotic fluid, chorionic villus samples, feces, bile, lymph, cervical mucus, semen, or other cells or materials containing nucleic acids. In certain embodiments, the methods disclosed herein can be performed in conjunction with nucleic acid detection methods, such as sequencing, microarrays, qPCR, digital PCR, or other methods used to measure nucleic acids. If desired for any reason, the ratio of allele count probabilities at loci can be calculated, and the allele ratios can be used in combination with some of the methods described herein to determine ploidy state, as long as they are compatible with those methods. In some embodiments, the methods disclosed herein include a step of calculating allele ratios at multiple polymorphic loci by computer from DNA measurements performed on the processed sample. In some embodiments, the methods disclosed herein, together with any combination of other improvements described in this disclosure, comprise calculating by computer allele ratios at a pl...
Claims
1. 1. A method for preparing a DNA fraction useful for analyzing genetic or epigenetic characteristics associated with cancer from a biological sample of a subject, comprising: (a) extracting cell-free DNA from the biological sample; (b) generating a DNA-enriched fraction, which comprises: (1) introducing at least one adaptor comprising a universal priming sequence into the extracted cell-free DNA to generate a plurality of adapted DNA sequences comprising the universal priming sequence; (2) performing universal amplification on the plurality of adapted DNA sequences using the universal priming sequence to generate a plurality of amplified adapted DNA sequences; (3) selectively enriching a subset of the plurality of amplified, adapted DNA sequences that includes one or more preselected loci, thereby generating enriched DNA sequences. This includes: (c) performing massively parallel sequencing on the enriched DNA sequences to obtain sequence reads that include at least a portion of one or more of the preselected loci, and identify one or more cancer-associated genetic or epigenetic signatures. A method comprising:
2. 10. The method of claim 1, wherein the biological sample is a blood, plasma, serum, or urine sample.
3. 10. The method of claim 1, wherein the cancer-associated genetic or epigenetic signature comprises a single nucleotide polymorphism or variant, a copy number variation, an insertion, a deletion, or differential methylation.
4. 10. The method of claim 1, wherein step (b) comprises selectively enriching between 1,000 and 500,000 preselected loci.
5. 10. The method of claim 1, wherein step (b) comprises selectively enriching between 10,000 and 200,000 preselected loci.
6. 10. The method of claim 1, wherein the selective enrichment comprises targeted multiplex amplification.
7. 2. The method of claim 1, wherein said selective enrichment comprises capturing DNA sequences from said plurality of amplified, adapted DNA sequences comprising one or more preselected loci using hybrid capture probes.
8. 10. The method of claim 1, wherein the adapters further comprise molecular barcodes, and the method comprises identifying sequence reads that originate from the same cell-free DNA molecule based on the molecular barcodes.
9. 2. The method of claim 1, wherein the universal amplification introduces sample-specific barcodes and the method comprises pooling the enriched DNA sequences from multiple samples together and sequencing them simultaneously in a single sequencing run.
10. 2. The method of claim 1, wherein the cell-free DNA comprises cancer DNA, and the method further comprises estimating a cancer DNA fraction in the cell-free DNA based on the sequence reads.
11. 1. A method for preparing a DNA fraction useful for analyzing genetic or epigenetic characteristics associated with cancer from a biological sample of a subject, comprising: (a) extracting cell-free DNA from the biological sample; (b) preparing a DNA fraction, this step introducing at least one adaptor comprising a universal priming sequence into at least a subset of the extracted cell-free DNA to generate a population of adapted DNA comprising one or more target loci of interest; performing universal amplification on at least a portion of the population of adapted DNA using the universal priming sequence to generate an amplified population of adapted DNA; Selectively enriching at least a subset of the amplified, adapted DNA population to produce enriched DNA. This includes: (c) performing massively parallel sequencing on the enriched DNA to obtain sequence data from at least a portion of the one or more target loci of interest; (d) obtaining sequence data-derived identities of one or more genetic or epigenetic features associated with cancer; A method comprising:
12. 12. The method of claim 11, wherein the biological sample is a blood, plasma, serum, or urine sample.
13. 12. The method of claim 11, wherein the genetic or epigenetic feature associated with cancer comprises a single nucleotide polymorphism or variant, a copy number variation, an insertion, a deletion, or a base methylation.
14. 12. The method of claim 11, wherein step (b) comprises selectively enriching between 1,000 and 500,000 loci.
15. 12. The method of claim 11, wherein step (b) comprises selectively enriching between 10,000 and 200,000 loci.
16. 12. The method of claim 11 , wherein the selective enrichment comprises targeted multiplex amplification.
17. 12. The method of claim 11, wherein said selective enrichment comprises capturing DNA comprising one or more loci from said amplified, adapted DNA using a hybrid capture probe.
18. 18. The method of claim 17, wherein the selective enrichment further comprises amplifying the captured DNA by a second universal amplification.
19. 20. The method of claim 18, wherein the second universal amplification introduces a sample-specific barcode.
20. 12. The method of Claim 11, wherein the adapters further comprise molecular barcodes, and the method comprises identifying sequence reads that originate from the same cell-free DNA molecule based on the molecular barcodes.
21. 12. The method of claim 11, wherein the universal amplification introduces sample-specific barcodes and the method comprises pooling the enriched DNA sequences from multiple samples together and sequencing them simultaneously in a single sequencing run.
22. 12. The method of claim 11, wherein the cell-free DNA comprises cancer DNA, and the method further comprises estimating a cancer DNA fraction in the cell-free DNA based on the sequence reads.
23. 12. The method of claim 11, wherein at least one of the one or more target loci of interest is selected from the group consisting of a single nucleotide variant (SNV), a single nucleotide polymorphism (SNP), an indel, a methylation site, and a copy number variant (CNV).
24. 12. The method of claim 11, wherein the one or more target loci of interest include one target locus containing a single base variant (SNV) and one target locus containing a base methylation site.
Citation Information
Patent Citations
Tools and methods for genetic tests using next generation sequencing
US20100227329A1
Identification of polymorphic sequences in mixtures of genomic DNA by whole genome sequencing
US20110230358A1
Multi-primer amplification method for barcoding of target nucleic acids
WO2010115154A1
Breast cancer associated circulating nucleic acid biomarkers
WO2011130751A1
Methods and kits to analyze microrna by nucleic acid sequencing
WO2011146942A1