Methods for non-invasive prenatal testing
Patent Information
- Application Number
- JP2024513727
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-01
- Filing Date
- 2022-08-24
- Publication Date
- 2025-07-16
AI Technical Summary
Current non-invasive prenatal testing methods are inadequate for identifying pregnancies at high risk for adverse perinatal outcomes such as preterm delivery, preeclampsia, fetal growth retardation, spontaneous abortion, and non-live birth, beyond aneuploidy risk.
A method involving the extraction and amplification of cell-free DNA from maternal blood samples to analyze 200-20,000 SNP loci, followed by high-throughput sequencing to determine the ploidy status of specific chromosomes, identifying pregnancies at risk through fetal fraction analysis and ploidy status determination.
Accurately identifies pregnancies at high risk for preterm delivery, preeclampsia, fetal growth retardation, spontaneous abortion, and non-live birth, providing clinicians with early intervention opportunities.
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 239,901, filed September 1, 2021, which is incorporated by reference in its entirety. [Background technology]
[0002] There is a need for non-invasive prenatal testing methods that can not only identify aneuploidy risk, but also identify pregnancies at high risk for other adverse perinatal outcomes such as preterm birth, pre-eclampsia, fetal growth restriction, spontaneous abortion, and non-live birth. Summary of the Invention
[0003] One aspect of the present disclosure is a method for preparing a preparation of amplified DNA from a first blood sample or a fraction thereof from a pregnant woman, useful for identifying pregnancies at high risk for preterm birth, preeclampsia, fetal growth restriction, spontaneous abortion, and / or non-live birth, comprising: (a) extracting cell-free DNA from the first blood sample or a fraction thereof to obtain a first extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; and (b) subjecting the first preparation of amplified DNA to targeted multiplex amplification of the first extracted DNA to amplify between 200 and 20,000 SNP loci in a single reaction volume to obtain amplified DNA. (c) analyzing a first preparation of amplified DNA by performing high throughput sequencing on the amplified DNA to obtain sequence reads and using the sequence reads to determine a ploidy state of the one or more chromosomes of interest, wherein a fetal fraction of less than 2.8% and / or no call of the ploidy state of the one or more chromosomes of interest indicates a pregnancy having a high risk of preterm birth, pre-eclampsia, fetal growth restriction, spontaneous abortion, and / or non-live birth.
[0004] Another aspect of the present disclosure is a method of preparing a preparation of amplified DNA useful for identifying pregnancies at high risk of preterm birth, pre-eclampsia, fetal growth restriction, spontaneous abortion, and / or non-live birth, comprising: (a) extracting cell-free DNA from a first blood sample or a fraction thereof of a pregnant woman to obtain a first extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; (b) preparing a first preparation of amplified DNA by performing targeted multiplex amplification on the first extracted DNA to amplify 200-20,000 SNP loci in a single reaction volume to obtain amplified DNA, wherein the 200-20,000 SNP loci are located on one or more chromosomes of interest; (c) analyzing the first preparation of amplified DNA by performing high throughput sequencing on the amplified DNA to obtain sequence reads and using the sequence reads to determine the ploidy state of the one or more chromosomes of interest; and (d) extracting cell-free DNA from a first blood sample or a fraction thereof of a pregnant woman to obtain a first extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; or a fraction thereof to obtain a second extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; (e) preparing a second preparation of amplified DNA by performing targeted multiplex amplification on the second extracted DNA to amplify 200-20,000 SNP loci in a single reaction volume to obtain amplified DNA, wherein the 200-20,000 SNP loci are located on one or more chromosomes of interest; and (f) analyzing the second preparation of amplified DNA by performing high throughput sequencing on the amplified DNA to obtain sequence reads and using the sequence reads to determine the ploidy state of the one or more chromosomes of interest, wherein for each of the first and second blood samples, a fetal fraction of less than 2.8% and / or no call of the ploidy state of the one or more chromosomes of interest indicates a pregnancy at high risk of preterm birth, pre-eclampsia, fetal growth restriction, spontaneous abortion, and / or non-live birth.
[0005] Another aspect of the present disclosure is a method for preparing a preparation of amplified DNA useful for identifying pregnancies at high risk for preterm birth, preeclampsia, fetal growth restriction, spontaneous abortion, and / or non-live birth, comprising: (a) extracting cell-free DNA from a first blood sample or a fraction thereof from a pregnant woman to obtain a first extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; and (b) subjecting the first preparation of amplified DNA to targeted multiplex amplification of the first extracted DNA to amplify between 200 and 20,000 SNP loci in a single reaction volume. (c) analyzing the first preparation of amplified DNA by performing high throughput sequencing on the amplified DNA to obtain sequence reads and using the sequence reads to determine the ploidy state of the one or more chromosomes of interest; (d) extracting cell-free DNA from a second longitudinally collected blood sample of the pregnant woman or a fraction thereof to obtain a second extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; (e) preparing a second preparation of amplified DNA by performing targeted multiplex amplification on the second extracted DNA to amplify 200-20,000 SNP loci in a single reaction volume to obtain amplified DNA, wherein the 200-20,000 SNP loci are located on one or more chromosomes of interest; and (f) analyzing the second preparation of amplified DNA by performing high throughput sequencing on the amplified DNA to obtain sequence reads. and analyzing by determining a ploidy state of one or more chromosomes of interest using the sequence reads, wherein a fetal fraction of less than 2.8%, or less than 2.7%, or less than 2.6%, or less than 2.5%, or less than 2.4%, or less than 2.3%, or less than 2.2%, or less than 2.1%, or less than 2.0% for each of the first and second blood samples indicates a pregnancy having a high risk of preterm birth, pre-eclampsia, fetal growth restriction, spontaneous abortion, and / or non-live birth. In some embodiments, the fetal fraction is quantified using the sequence reads.In some embodiments, the fetal fraction is quantified using methylation-based multiplex ddPCR. In some embodiments, the fetal fraction is quantified using fragment length and fragment number.
[0006] A further aspect of the present disclosure is a method of preparing a preparation of amplified DNA useful for identifying pregnancies at high risk of preterm birth, pre-eclampsia, fetal growth restriction, spontaneous abortion, and / or non-live birth, comprising: (a) extracting cell-free DNA from a first blood sample or a fraction thereof of a pregnant woman to obtain a first extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; (b) preparing a first preparation of amplified DNA by performing targeted multiplex amplification on the first extracted DNA to amplify 200-20,000 SNP loci in a single reaction volume to obtain amplified DNA, wherein the 200-20,000 SNP loci are located on one or more chromosomes of interest; (c) analyzing the first preparation of amplified DNA by performing high throughput sequencing on the amplified DNA to obtain sequence reads and using the sequence reads to determine the ploidy state of the one or more chromosomes of interest; and (d) analyzing a first preparation of amplified DNA by performing targeted multiplex amplification on the amplified DNA to obtain sequence reads and using the sequence reads to determine the ploidy state of the one or more chromosomes of interest. (e) preparing a second preparation of amplified DNA by performing targeted multiplex amplification on the second extracted DNA to amplify 200-20,000 SNP loci in a single reaction volume to obtain amplified DNA, wherein the 200-20,000 SNP loci are located on one or more chromosomes of interest; and (f) analyzing the second preparation of amplified DNA by performing high throughput sequencing on the amplified DNA to obtain sequence reads and using the sequence reads to determine a ploidy state of the one or more chromosomes of interest, wherein for each of the first and second blood samples, no call of the ploidy state of the one or more chromosomes of interest indicates a pregnancy having a high risk of preterm birth, pre-eclampsia, fetal growth restriction, spontaneous abortion, and / or non-live birth. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0007] WO2011 / 041485, filed September 30, 2010 as PCT / US2010 / 050824, is incorporated herein by reference in its entirety. WO2011 / 146632, filed May 18, 2011 as PCT / US2011 / 037018, is incorporated herein by reference in its entirety. WO2012 / 108920, filed November 18, 2011 as PCT / US2011 / 061506, is incorporated herein by reference in its entirety. WO2012 / 088456, filed December 22, 2011 as PCT / US2011 / 066938, is incorporated herein by reference in its entirety. WO2014 / 018080, filed November 21, 2012 as PCT / US2012 / 066339, is incorporated herein by reference in its entirety. WO2014 / 028778, filed August 15, 2013 as PCT / US2013 / 055205, is incorporated herein by reference in its entirety. WO2015 / 164432, filed April 21, 2015 as PCT / US2015 / 026957, is incorporated herein by reference in its entirety. WO2016 / 183106, filed May 10, 2016 as PCT / US2016 / 031686, is incorporated herein by reference in its entirety. US2016 / 0371428, filed June 20, 2016 as US15 / 186,774, is incorporated herein by reference in its entirety. US2018 / 0173845, filed February 2, 2018 as US15 / 887,864, is incorporated herein by reference in its entirety.
[0008] 1. A method for preparing a preparation of amplified DNA from a first blood sample or a fraction thereof of a pregnant woman, useful for identifying pregnancies at high risk of preterm birth, preeclampsia, and / or fetal growth retardation, comprising: (a) extracting cell-free DNA from the first blood sample or a fraction thereof to obtain a first extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; and (b) preparing a first preparation of amplified DNA by performing targeted multiplex amplification on the first extracted DNA to amplify 200 to 20,000 SNP loci in a single reaction volume to obtain amplified DNA, wherein the 200 to 20,000 SNP loci are amplified in a single reaction volume. Disclosed herein is a method comprising: (a) preparing a first preparation of amplified DNA, the first preparation comprising 0,000 SNP loci located on one or more chromosomes of interest; and (b) analyzing the first preparation of amplified DNA by performing high throughput sequencing on the amplified DNA to obtain sequence reads and using the sequence reads to quantify the fetal fraction in the first blood sample or a fraction thereof to determine the ploidy state of the one or more chromosomes of interest, wherein a fetal fraction of less than 2.8% and / or no call of the ploidy state of the one or more chromosomes of interest indicates a pregnancy having a high risk of preterm birth, pre-eclampsia, and / or fetal growth retardation.
[0009] In some embodiments, the method includes (d) extracting cell-free DNA from a second longitudinally collected blood sample of the pregnant woman, or a fraction thereof, to obtain a second extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; and (e) preparing a second preparation of amplified DNA by performing targeted multiplex amplification on the second extracted DNA to amplify 200-20,000 SNP loci in a single reaction volume to obtain amplified DNA, wherein the 200-20,000 SNP loci are associated with one or more chromosomes of interest. and (f) analyzing a second preparation of the amplified DNA by performing high throughput sequencing on the amplified DNA to obtain sequence reads and using the sequence reads to determine the ploidy state of the one or more chromosomes of interest, wherein for each of the first and second blood samples, a fetal fraction of less than 2.8% and / or no call of the ploidy state of the one or more chromosomes of interest further indicates a pregnancy having a high risk of preterm birth, pre-eclampsia, fetal growth restriction, spontaneous abortion, and / or non-live birth.
[0010] In some embodiments, the method further includes identifying pregnant women without a call of the ploidy state of one or more chromosomes of interest for each of the first and second blood samples as having at least a 30%, or at least a 35%, or at least a 40%, or at least a 45%, or at least a 50% risk of preterm birth before 37 weeks, pre-eclampsia, and / or fetal growth restriction.
[0011] In some embodiments, the method further includes identifying pregnant women without a call of the ploidy state of one or more chromosomes of interest for each of the first and second blood samples as having at least a 30%, or at least a 35%, or at least a 40%, or at least a 45%, or at least a 50% risk of preterm birth before 37 weeks, pre-eclampsia, stillbirth, and / or fetal growth restriction.
[0012] In some embodiments, the method further includes identifying pregnant women without a call of the ploidy state of one or more chromosomes of interest for each of the first and second blood samples as having at least a 12%, or at least a 13%, or at least a 14%, or at least a 15%, or at least a 16%, or at least a 17%, or at least a 18% risk of pre-eclampsia.
[0013] In some embodiments, the method further includes identifying pregnant women without a call of the ploidy state of one or more chromosomes of interest for each of the first and second blood samples as having at least a 10%, or at least a 12%, or at least a 14%, or at least a 16%, or at least a 18%, or at least a 20%, or at least a 22% risk of preterm birth before 28 weeks.
[0014] In some embodiments, the method further includes identifying pregnant women without a call of the ploidy state of one or more chromosomes of interest for each of the first and second blood samples as having a risk of preterm birth before 34 weeks of at least 16%, or at least 18%, or at least 20%, or at least 22%, or at least 24%, or at least 26%, or at least 28%.
[0015] In some embodiments, the method further includes identifying pregnant women without a call of the ploidy state of one or more chromosomes of interest for each of the first and second blood samples as having a risk of preterm birth before 37 weeks of at least 24%, or at least 28%, or at least 32%, or at least 36%, or at least 40%, or at least 44%.
[0016] In some embodiments, the method further includes identifying pregnant women without a call of the ploidy state of one or more chromosomes of interest for each of the first and second blood samples as having at least a 10%, or at least a 10.5%, or at least a 11%, or at least a 11.5%, or at least a 12%, or at least a 12.5%, or at least a 13%, or at least a 13.5% risk of fetal growth restriction.
[0017] In some embodiments, the fetal fraction is quantified using sequence reads. In some embodiments, the fetal fraction is quantified using methylation-based multiplex ddPCR. In some embodiments, the fetal fraction is quantified using fragment length and fragment number.
[0018] In some embodiments, the method includes identifying pregnant women having a fetal fraction of less than 2.8%, or less than 2.7%, or less than 2.6%, or less than 2.5%, or less than 2.4%, or less than 2.3%, or less than 2.2%, or less than 2.1%, or less than 2.0% for the first blood sample. In some embodiments, the method includes identifying pregnant women having a fetal fraction of less than 2.8%, or less than 2.7%, or less than 2.6%, or less than 2.5%, or less than 2.4%, or less than 2.3%, or less than 2.2%, or less than 2.1%, or less than 2.0% for the second blood sample. In some embodiments, the method includes identifying pregnant women having a fetal fraction of less than 2.8%, or less than 2.7%, or less than 2.6%, or less than 2.5%, or less than 2.4%, or less than 2.3%, or less than 2.2%, or less than 2.1%, or less than 2.0% for each of the first and second blood samples.
[0019] In some embodiments, the method includes identifying pregnant women with a percentile of fetal fraction less than the 3rd percentile, or less than the 2nd percentile, or less than the 1st percentile, or less than the 0.5th percentile, or less than the 0.2th percentile, or less than the 0.1th percentile for the first blood sample, optionally adjusted for maternal weight and gestational age. In some embodiments, the method includes identifying pregnant women with a percentile of fetal fraction less than the 3rd percentile, or less than the 2nd percentile, or less than the 1st percentile, or less than the 0.5th percentile, or less than the 0.2th percentile, or less than the 0.1th percentile for the second blood sample, optionally adjusted for maternal weight and gestational age. In some embodiments, the methods include identifying pregnant women having a fetal fraction percentile less than the 3rd percentile, or less than the 2nd percentile, or less than the 1st percentile, or less than the 0.5th percentile, or less than the 0.2nd percentile, or less than the 0.1th percentile for each of the first and second blood samples, optionally adjusted for maternal weight and gestational age.
[0020] In some embodiments, the method includes identifying pregnant women having a fetal fraction of less than 2.8%, or less than 2.7%, or less than 2.6%, or less than 2.5%, or less than 2.4%, or less than 2.3%, or less than 2.2%, or less than 2.1%, or less than 2.0% for each of the first and second blood samples as having a high risk of pre-eclampsia (e.g., a risk of pre-eclampsia of at least 12%, or at least 13%, or at least 14%, or at least 15%, or at least 16%, or at least 17%, or at least 18%).
[0021] In some embodiments, the methods include identifying pregnant women having a fetal fraction percentile below the 3rd percentile, or below the 2nd percentile, or below the 1st percentile, or below the 0.5th percentile, or below the 0.2nd percentile, or below the 0.1th percentile for each of the first and second blood samples as having a high risk of pre-eclampsia (e.g., a risk of pre-eclampsia of at least 12%, or at least 13%, or at least 14%, or at least 15%, or at least 16%, or at least 17%, or at least 18%), optionally adjusted for maternal weight and gestational age.
[0022] In some embodiments, the method includes identifying pregnant women having a fetal fraction of less than 2.8%, or less than 2.7%, or less than 2.6%, or less than 2.5%, or less than 2.4%, or less than 2.3%, or less than 2.2%, or less than 2.1%, or less than 2.0% for each of the first and second blood samples as having an elevated risk of preterm birth before 28 weeks (e.g., a risk of at least 10%, or at least 12%, or at least 14%, or at least 16%, or at least 18%, or at least 20%, or at least 22%).
[0023] In some embodiments, the methods include identifying pregnant women having a fetal fraction percentile below the 3rd percentile, or below the 2nd percentile, or below the 1st percentile, or below the 0.5th percentile, or below the 0.2nd percentile, or below the 0.1th percentile for each of the first and second blood samples as having an elevated risk of preterm birth before 28 weeks (e.g., a risk of preterm birth before 28 weeks of at least 10%, or at least 12%, or at least 14%, or at least 16%, or at least 18%, or at least 20%, or at least 22%), optionally adjusted for maternal weight and gestational age.
[0024] In some embodiments, the method includes identifying pregnant women having a fetal fraction of less than 2.8%, or less than 2.7%, or less than 2.6%, or less than 2.5%, or less than 2.4%, or less than 2.3%, or less than 2.2%, or less than 2.1%, or less than 2.0% for each of the first and second blood samples as having an elevated risk of preterm birth before 34 weeks (e.g., a risk of at least 16%, or at least 18%, or at least 20%, or at least 22%, or at least 24%, or at least 26%, or at least 28%).
[0025] In some embodiments, the methods include identifying pregnant women having a fetal fraction percentile below the 3rd percentile, or below the 2nd percentile, or below the 1st percentile, or below the 0.5th percentile, or below the 0.2nd percentile, or below the 0.1th percentile for each of the first and second blood samples as having an elevated risk of preterm birth before 34 weeks (e.g., a risk of preterm birth before 34 weeks of at least 16%, or at least 18%, or at least 20%, or at least 22%, or at least 24%, or at least 26%, or at least 28%), optionally adjusted for maternal weight and gestational age.
[0026] In some embodiments, the method includes identifying pregnant women having a fetal fraction of less than 2.8%, or less than 2.7%, or less than 2.6%, or less than 2.5%, or less than 2.4%, or less than 2.3%, or less than 2.2%, or less than 2.1%, or less than 2.0% for each of the first and second blood samples as having an elevated risk of preterm birth before 37 weeks (e.g., a risk of at least 24%, or at least 28%, or at least 32%, or at least 36%, or at least 40%, or at least 44%).
[0027] In some embodiments, the methods include identifying pregnant women having a fetal fraction percentile below the 3rd percentile, or below the 2nd percentile, or below the 1st percentile, or below the 0.5th percentile, or below the 0.2nd percentile, or below the 0.1th percentile for each of the first and second blood samples as having an elevated risk of preterm birth before 37 weeks (e.g., a risk of preterm birth before 37 weeks of at least 24%, or at least 28%, or at least 32%, or at least 36%, or at least 40%, or at least 44%), optionally adjusted for maternal weight and gestational age.
[0028] In some embodiments, the methods include identifying pregnant women having a fetal fraction of less than 2.8%, or less than 2.7%, or less than 2.6%, or less than 2.5%, or less than 2.4%, or less than 2.3%, or less than 2.2%, or less than 2.1%, or less than 2.0% for each of the first and second blood samples as having an elevated risk of fetal growth restriction (e.g., a risk of at least 10%, or at least 10.5%, or at least 11%, or at least 11.5%, or at least 12%, or at least 12.5%, or at least 13%, or at least 13.5% fetal growth restriction).
[0029] In some embodiments, the methods include identifying pregnant women having a percentile of the fetal fraction below the 3rd percentile, or below the 2nd percentile, or below the 1st percentile, or below the 0.5th percentile, or below the 0.2nd percentile, or below the 0.1th percentile for each of the first and second blood samples as having an elevated risk of fetal growth retardation (e.g., a risk of fetal growth retardation of at least 10%, or at least 10.5%, or at least 11%, or at least 11.5%, or at least 12%, or at least 12.5%, or at least 13%, or at least 13.5%), optionally adjusted for maternal weight and gestational age.
[0030] In some embodiments, the methods include identifying samples that have an abnormally high fetal fraction.
[0031] In some embodiments, the method includes using details of cfDNA fragments as part of an algorithm to predict preterm birth, preeclampsia, fetal growth retardation, spontaneous abortion, and / or non-live birth. In some embodiments, the method includes using fragment length to predict preterm birth, preeclampsia, fetal growth retardation, spontaneous abortion, and / or non-live birth. In some embodiments, the method includes using details of fragments, such as location in the genome or start and stop points, to predict preterm birth, preeclampsia, fetal growth retardation, spontaneous abortion, and / or non-live birth.
[0032] In some embodiments, the method further comprises repeating steps (d)-(f) with a third, fourth, or further blood sample or fraction thereof collected longitudinally.
[0033] In some embodiments, step (a) comprises extracting cell-free DNA from the plasma fraction of the blood sample. In some embodiments, step (a) further comprises ligating at least one adaptor to the extracted DNA, the adaptor comprising a universal priming sequence. In some embodiments, step (a) further comprises performing a universal PCR amplification using at least one primer that binds to the universal priming sequence.
[0034] In some embodiments, step (b) comprises PCR amplification of 200-20,000 SNP loci using 200-20,000 pairs of target-specific PCR primers in one reaction mixture, or using a universal primer and 200-20,000 target-specific primers in one reaction mixture. In some embodiments, step (b) comprises PCR amplification of 500-20,000 SNP loci using 500-20,000 pairs of target-specific PCR primers in one reaction mixture, or using a universal primer and 500-20,000 target-specific primers in one reaction mixture. In some embodiments, step (b) comprises PCR amplification of 1,000-20,000 SNP loci using 1,000-20,000 pairs of target-specific PCR primers in one reaction mixture, or using a universal primer and 1,000-20,000 target-specific primers in one reaction mixture. In some embodiments, step (b) comprises PCR amplification of 2,000-20,000 SNP loci using 2,000-20,000 pairs of target-specific PCR primers in one reaction mixture, or using a universal primer and 2,000-20,000 target-specific primers in one reaction mixture. In some embodiments, step (b) comprises PCR amplification of 5,000-20,000 SNP loci using 5,000-20,000 pairs of target-specific PCR primers in one reaction mixture, or using a universal primer and 5,000-20,000 target-specific primers in one reaction mixture. In some embodiments, step (b) comprises PCR amplification of 10,000-20,000 SNP loci using 10,000-20,000 pairs of target-specific PCR primers in one reaction mixture, or using a universal primer and 10,000-20,000 target-specific primers in one reaction mixture.In some embodiments, step (b) comprises PCR amplification of between 20,000 and 50,000 SNP loci using between 20,000 and 50,000 pairs of target-specific PCR primers in one reaction mixture, or using a universal primer and between 20,000 and 50,000 target-specific primers in one reaction mixture.
[0035] In some embodiments, the amplified DNA in step (b) each comprises 100 bp or less amplified from the extracted DNA. In some embodiments, the amplified DNA in step (b) each comprises 90 bp or less amplified from the extracted DNA. In some embodiments, the amplified DNA in step (b) each comprises 80 bp or less amplified from the extracted DNA. In some embodiments, the amplified DNA in step (b) each comprises 80 bp or less amplified from the extracted DNA. In some embodiments, the amplified DNA in step (b) each comprises 70 bp or less amplified from the extracted DNA. In some embodiments, the amplified DNA in step (b) each comprises 50 to 100 bp amplified from the extracted DNA. In some embodiments, the amplified DNA in step (b) each comprises 60 to 80 bp amplified from the extracted DNA. In some embodiments, the amplified DNA in step (b) each comprises 65 to 80 bp amplified from the extracted DNA.
[0036] In some embodiments, step (b) further comprises barcode PCR after targeted multiplex amplification. In some embodiments, the barcode PCR introduces a sample-specific barcode or sample-specific identifier sequence. In some embodiments, the barcode PCR introduces a sequencing tag for subsequent high-throughput sequencing.
[0037] In some embodiments, the ploidy state of one or more chromosomes of interest is determined by: calculating allele counts at SNP loci based on the sequence reads; generating multiple ploidy hypotheses each associated with a different possible ploidy state of the chromosome of interest; building a joint distribution model of the expected allele counts at the SNP loci on the chromosome of interest for each ploidy hypothesis; determining the relative probability of each of the ploidy hypotheses using the joint distribution model and the allele counts; and calling the ploidy state of the fetus by selecting the ploidy state corresponding to the hypothesis with the greatest probability.
[0038] In one embodiment, the present disclosure provides an ex vivo method for determining the ploidy state of chromosomes of a gestating fetus from genotype data measured from a mixed sample of DNA (i.e., DNA from the fetus's mother and DNA from the fetus) and, optionally, from samples of genetic material from the mother and possibly the father, by using a joint distribution model to generate a set of predicted allele distributions for different possible fetal ploidy states given the parental genotype data, comparing the predicted allele distributions to the actual allele distributions measured in the mixed sample, and selecting the ploidy state in which the predicted allele distribution pattern most closely matches the observed allele distribution pattern. In one embodiment, the mixed sample is derived from maternal blood, or maternal serum or plasma. In one embodiment, the mixed sample of DNA may be preferentially enriched at multiple polymorphic loci. In one embodiment, the preferential enrichment is performed in a manner that minimizes allelic bias. In one embodiment, the present disclosure relates to a composition of DNA that is preferentially enriched at multiple loci such that allelic bias is low. In one embodiment, the distribution(s) of alleles are measured by sequencing DNA from the mixed sample. In one embodiment, the joint distribution model assumes that alleles are binomial distributed. In one embodiment, a set of predicted joint allele distributions is created for genetically linked loci, taking into account existing recombination frequencies from various sources, for example, using data from the International HapMap Consortium.
[0039] In one embodiment, the present disclosure provides a method for non-invasive prenatal diagnosis (NPD), specifically determining the aneuploidy status of a fetus by observing allele measurements at multiple polymorphic loci in genotype data measured on a DNA mixture, where measurements of certain alleles indicate an aneuploid fetus, while measurements of other alleles indicate a euploid fetus. In one embodiment, the genotype data is measured by sequencing a DNA mixture derived from maternal plasma. In one embodiment, the DNA sample may be preferentially enriched with molecules of DNA corresponding to multiple loci for which the allele distribution is being calculated. In one embodiment, a sample of DNA containing only or almost only genetic material from the mother, and possibly also a sample of DNA containing only or almost only genetic material from the father, is measured. In one embodiment, the genetic measurements of one or both parents, along with the estimated fetal fraction, are used to generate multiple predicted allele distributions corresponding to different possible underlying genetic states of the fetus, where the predicted allele distributions may be referred to as hypotheses. In one embodiment, the genetic data of the mother is not determined by measuring genetic material that is essentially exclusive or nearly exclusive, but rather is estimated from genetic measurements made on maternal plasma that contains a mixture of maternal and fetal DNA. In some embodiments, the hypotheses may include the ploidy of the fetus at one or more chromosomes, which segments of that chromosome were inherited by which chromosomes of the fetus from which parents, and combinations thereof. In some embodiments, the ploidy state of the fetus is determined by comparing the observed allele measurements with different hypotheses, at least some of which hypotheses correspond to different ploidy states, and selecting the ploidy state that corresponds to the hypothesis that is most likely to be true given the observed allele measurements. In one embodiment, the method involves using allele measurement data from some or all of the measured SNPs, regardless of whether the locus is homozygous or heterozygous, and therefore does not involve the use of alleles at loci that are only heterozygous. This method may not be suitable for situations where the genetic data relates to only one polymorphic locus.This method is particularly advantageous when the genetic data includes data for more than 10 polymorphic loci or more than 20 polymorphic loci for the target chromosome. This method is particularly advantageous when the genetic data includes data for more than 50 polymorphic loci for the target chromosome, more than 100 polymorphic loci for the target chromosome, or more than 200 polymorphic loci for the target chromosome. In some embodiments, the genetic data may include data for more than 500 polymorphic loci for the target chromosome, more than 1,000 polymorphic loci for the target chromosome, more than 2,000 polymorphic loci, or more than 5,000 polymorphic loci for the target chromosome.
[0040] In one embodiment, the methods disclosed herein use selective enrichment techniques that preserve the relative allele frequencies present in the original sample of DNA at each polymorphic locus from the set of polymorphic loci. In some embodiments, the amplification and / or selective enrichment techniques may involve PCR (e.g., ligation-mediated PCR), fragment capture by hybridization, molecular inversion probes, or other circularization probes. In some embodiments, the methods for amplification or selective enrichment may involve using probes that, upon correct hybridization to a target sequence, separate the 3' or 5' end of the nucleotide probe from the polymorphic site of the allele by a small number of nucleotides. This separation reduces the preferential amplification of one allele, referred to as allele bias. This is an improvement over methods that involve using probes in which the 3' or 5' end of a correctly hybridized probe is located directly adjacent to or very close to the polymorphic site of the allele. In one embodiment, probes whose hybridizing region may or certainly does contain the polymorphic site are excluded. Polymorphic sites at the hybridization sites may cause unequal hybridization or completely inhibit hybridization at some alleles, resulting in preferential amplification of certain alleles. These embodiments are improvements over other methods involving targeted amplification and / or selective enrichment in that they better preserve the original allele frequencies of the sample at each polymorphic locus, and the sample may be a pure genomic sample from a single individual or a mixture of individuals.
[0041] In one embodiment, the method disclosed herein uses highly efficient, highly multiplexed targeted PCR to amplify DNA, followed by high-throughput sequencing to determine the allele frequency at each target locus. The ability to multiplex more than about 50 or 100 PCR primers in one reaction in such a way that most of the resulting sequence reads map to the target loci is novel and non-obvious. One technique that allows highly multiplexed targeted PCR to be performed in a highly efficient manner involves designing primers that are unlikely to hybridize with each other. PCR probes, typically called primers, are selected by creating a thermodynamic model of potentially harmful interactions between at least 500, at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 50,000, or at least 100,000 potential primer pairs or unintended interactions between primers and sample DNA, and then using this model to filter out designs that are incompatible with other designs in the pool. Another technique that allows highly multiplexed targeted PCR to be performed in a highly efficient manner is to use a partial or complete nesting approach to targeted PCR. Using one or a combination of these approaches, multiplexing of at least 300, at least 800, at least 1,200, at least 4,000, or at least 10,000 primers in a single pool is possible, and the resulting amplified DNA, which contains the majority of DNA, maps to the target locus when sequenced. Using one or a combination of these approaches, multiplexing of many primers in a single pool is possible, and the resulting amplified DNA contains more than 50%, more than 80%, more than 90%, more than 95%, more than 98%, or more than 99% of DNA molecules that map to the target locus.
[0042] In one embodiment, the method disclosed herein provides a quantitative measure of the number of independent observations of each allele at a polymorphic locus. This differs from most methods, such as microarrays or qualitative PCR, which provide information about the ratio of the two alleles but do not quantify the number of independent observations of either allele. In methods that provide quantitative information about the number of independent observations, only the ratio is utilized for fold calculations, and the quantitative information itself is not useful. To illustrate the importance of retaining information about the number of independent observations, consider a sample locus with two alleles, A and B. In a first experiment, 20 A alleles and 20 B alleles are observed, and in a second experiment, 200 A alleles and 200 B alleles are observed. In both experiments, the ratio (A / (A+B)) is equal to 0.5, but the second experiment conveys more information about the certainty of the frequency of the A or B allele than the first. Some methods known in the prior art provide allelic ratios (channel ratios) from individual alleles (i.e., x i / y i ) and compare this ratio to a reference chromosome or analyze this ratio using rules on how this ratio is expected to behave in a particular situation. Allele weighting is not suggested in methods known in the art, which can ensure approximately the same amount of PCR product for each allele, and it is assumed that all alleles should behave in the same way. Such methods have several drawbacks and, more importantly, they preclude the use of several improvements described elsewhere in this disclosure.
[0043] In one embodiment, the method disclosed herein explicitly models the allele frequency distributions expected for disomy, as well as the multiple allele frequency distributions that may be expected for trisomy resulting from nondisjunction during meiosis I, nondisjunction during meiosis II, and / or nondisjunction during mitosis early in fetal development. To explain why this is important, imagine the case where there was no crossover: nondisjunction during meiosis I results in a trisomy with two different homologs inherited from one parent, in contrast to nondisjunction during meiosis II or mitosis early in fetal development resulting in two copies of the same homolog from one parent. Each scenario would result in different allele frequencies expected at each polymorphic locus due to genetic linkage, and would also be different for all loci considered jointly. Crossover resulting in exchange of genetic material between homologs makes inheritance patterns more complex, and in one embodiment, the method addresses this by using recombination rate information in addition to the physical distance between loci. In one embodiment, to improve the distinction between meiosis I nondisjunction and meiosis II or mitotic nondisjunction, the method of the invention incorporates into the model an increasing probability of crossing over with increasing distance from the centromere. Meiosis II nondisjunction and mitotic nondisjunction can be distinguished by the fact that mitotic nondisjunction typically results in identical or near-identical copies of one homolog, whereas the two homologs present following a meiosis II nondisjunction event often differ by one or more crossing overs during gamete formation.
[0044] In some embodiments, the methods disclosed herein involve comparing the observed allele measurements with theoretical hypotheses corresponding to possible fetal genetic aneuploidies, and do not involve quantifying the allele ratio at heterozygous loci. When the number of loci is less than about 20, the multiple determination performed using a method involving quantifying the allele ratio at heterozygous loci and the multiple determination performed using a method involving comparing the observed allele measurements with theoretical allele distribution hypotheses corresponding to possible fetal genetic states may yield similar results. However, when the number of loci is more than 50, these two methods are likely to yield significantly different results, and when the number of loci is more than 400, 1,000, or 2,000, these two methods are very likely to yield increasingly significantly different results. These differences are due to the fact that methods that involve quantifying allele ratios at heterozygous loci without measuring the magnitude of each allele independently and aggregating or averaging the ratios preclude the use of techniques including the use of joint distribution models, performing joint analyses, using binomial distribution models, and / or other advanced statistical techniques, whereas the use of methods that involve comparing observed allele measurements to theoretical allele distribution hypotheses corresponding to possible fetal genetic states allows the use of these techniques that can substantially increase the accuracy of the determination.
[0045] In one embodiment, the method disclosed herein involves using a joint distribution model to determine whether the distribution of observed allele measurements indicates a euploid or aneuploid fetus. The use of a joint distribution model is a significant improvement over methods of determining heterozygosity rates by independently treating polymorphic loci in that the resulting determinations are significantly more accurate. Without being bound to any particular theory, it is believed that one reason they have higher accuracy is that the joint distribution model takes into account the linkage between SNPs and the likelihood of crossovers that occurred during meiosis that gave rise to the gametes that formed the embryo that developed into the fetus. The purpose of using the concept of linkage when creating a predicted distribution of allele measurements for one or more hypotheses is to allow the creation of predicted allele measurement distributions that correspond much better to reality than if linkage is not used. For example, assume that there are two SNPs, 1 and 2, located close to each other, and the mother is A at SNP1 and A at SNP2 on one homolog, and B at SNP1 and B at SNP2 on homolog2. If the father is A for both SNPs on both homologs and B is measured for fetal SNP1, this indicates that homolog 2 has been inherited by the fetus and therefore that B is much more likely to be present in the fetus at SNP2. Models that take linkage into account predict this, but models that do not take linkage into account do not. Alternatively, if the mother is AB at SNP1 and AB at nearby SNP2, two hypotheses corresponding to maternal trisomy at that position can be used, one with a matching copy error (nondisjunction in meiosis II or mitosis early in fetal development) and one with a non-matching copy error (nondisjunction in meiosis I). In the case of a matching copy error trisomy, if the fetus inherits AA from the mother at SNP1, the fetus is much more likely to inherit either AA or BB from the mother at SNP2, rather than AB. In the case of a non-matching copy error, the fetus will inherit AB from the mother at both SNPs.The allele distribution hypotheses made by ploidy calling that takes linkage into account make these predictions and therefore correspond to the actual allele measurements to a significantly greater extent than ploidy calling that does not take linkage into account. Note that the linkage approach is not possible when using methods that rely on calculating allele ratios and aggregating those allele ratios.
[0046] One reason why fold determination using a method involving comparing observed allele measurements to theoretical hypotheses corresponding to possible fetal genetic states is believed to be more accurate is that when sequencing is used to measure alleles, the method can glean more information from data from alleles with lower total read counts than other methods. For example, methods that rely on calculating and aggregating allele ratios generate disproportionately weighted stochastic noise. For example, imagine a set of loci that involves measuring alleles using sequencing, where only five sequence reads were detected for each locus. In one embodiment, for each of the alleles, the data may be compared to the hypothesized allele distribution and weighted according to the number of sequence reads, so that the data from these measurements will be appropriately weighted and incorporated into the overall determination. This is in contrast to methods that involve quantifying allele ratios at heterozygous loci, which can only calculate 0%, 20%, 40%, 60%, 80%, or 100% ratios as possible allele ratios, none of which are close to the expected allele ratios. In this latter case, the calculated allele ratios must be discarded due to insufficient reads or otherwise have disproportionate weighting, introducing stochastic noise into the determination, thereby reducing the accuracy of the determination. In one embodiment, the measurements of individual alleles may be treated as independent measurements, and the relationship between measurements made on alleles at the same locus is no different from the relationship between measurements made on alleles at different loci.
[0047] In one embodiment, the method disclosed herein involves determining whether the distribution of observed allele measurements indicates a euploid or aneuploid fetus without comparing any metric with the observed allele measurements on a reference chromosome predicted to be disomic (referred to as the RC method). This is a significant improvement over methods such as those using shotgun sequencing to detect aneuploidy by assessing the proportion of randomly sequenced fragments from a suspect chromosome relative to one or more presumed disomic reference chromosomes. This RC method will give erroneous results if the presumed disomic reference chromosome is not actually disomic. This can occur when the aneuploidy is more substantial than a single chromosome trisomy or when the fetus is triploid and all autosomes are trisomic. In the case of a female triploid (69,XXX) fetus, there are in fact no disomic chromosomes at all. The method described herein does not require a reference chromosome and will be able to accurately identify trisomic chromosomes in a female triploid fetus. For each chromosome, hypothesis, daughter fraction and noise level, a joint distribution model can be fitted without any reference chromosome data, estimates of the overall daughter fraction, or a fixed reference hypothesis.
[0048] In one embodiment, the method disclosed herein shows how the observation of allele distribution at polymorphic loci can be used to determine the ploidy state of the fetus with higher accuracy than prior art methods. In one embodiment, the method uses targeted sequencing to obtain maternal-fetal mixed genotypes, and optionally maternal and / or paternal genotypes, at multiple SNPs to first establish various expected allele frequency distributions under different hypotheses, then observes the quantitative allele information obtained at the maternal-fetal mixtures, evaluates which hypothesis best fits the data, and the genetic state corresponding to the hypothesis that best fits the data is called the correct genetic state. In one embodiment, the method disclosed herein also uses a measure of fit to generate confidence that the called genetic state is the correct genetic state. In one embodiment, the method disclosed herein includes using an algorithm to analyze the distribution of alleles found for loci with different parent contexts, and comparing the observed distribution of alleles to the expected distribution of alleles for different ploidy states for different parent contexts (different parent genotype patterns). This is different from or an improvement over methods that do not use methods that allow for estimation of the number of independent instances of each allele at each locus in a maternal-fetal mixed sample. In one embodiment, the methods disclosed herein involve determining whether the distribution of observed allele measurements indicates a euploid or aneuploid fetus using the distribution of observed alleles measured at loci where the mother is heterozygous. This is different from or an improvement over methods that do not use the observed allele distribution at loci where the mother is heterozygous, because it allows for the use of approximately twice as much genetic measurement data from a set of sequence data in the ploidy determination, resulting in a more accurate determination, when DNA is not preferentially enriched or is preferentially enriched for loci that are not known to be highly informative for that particular target individual.
[0049] In one embodiment, the method disclosed herein uses a joint distribution model that assumes that the allele frequency at each locus is essentially polynomial (and therefore binomial if the SNP is biallelic). In some embodiments, the joint distribution model uses a beta-binomial distribution. If a measurement technique such as sequencing is used to provide a quantitative measure for each allele present at each locus, the binomial model can be applied to each locus, and the underlying degree of allele frequency and the reliability of that frequency can be ascertained. Methods known in the art that generate ploidy calls from allele ratios, or methods in which quantitative allele information is discarded, cannot ascertain the certainty in the observed ratios. This method is different from or an improvement over methods that calculate allele ratios and aggregate those ratios to make ploidy calls, because any method that involves calculating allele ratios at a particular locus and then aggregating those ratios necessarily assumes that the measured intensities or numerical values, which are indicative of the amount of DNA from any given allele or locus, are distributed in a Gaussian manner. The methods disclosed herein do not involve calculating allele ratios. In some embodiments, the methods disclosed herein may involve incorporating the number of observations of each allele at multiple loci into the model. In some embodiments, the methods disclosed herein may involve calculating the predictive distribution itself, which allows the use of a joint binomial distribution model that may be more accurate than any model that assumes a Gaussian distribution of allele measurements. As the number of loci increases, the likelihood that the binomial distribution model is significantly more accurate than the Gaussian distribution increases. For example, when less than 20 loci are examined, the likelihood that the binomial distribution model is significantly superior is low. However, when more than 100 loci, or particularly more than 400 loci, or particularly more than 1,000 loci, or particularly more than 2,000 loci are used, the binomial distribution model has a very high likelihood of being significantly more accurate than the Gaussian distribution model, thereby resulting in more accurate multiple determination.The likelihood that the binomial distribution model is significantly more accurate than the Gaussian distribution increases as the number of observations at each locus increases.For example, when less than 10 different sequences are observed at each locus, the likelihood that the binomial distribution model is significantly superior is low.However, when more than 50 sequence reads, or particularly more than 100 sequence reads, or particularly more than 200 sequence reads, or particularly more than 300 sequence reads are used for each locus, the binomial distribution model is highly likely to be significantly more accurate than the Gaussian distribution model, resulting in more accurate multiple determination.
[0050] In one embodiment, the method disclosed herein uses sequencing to measure the number of instances of each allele at each locus in a DNA sample. Each sequencing read can be mapped to a specific locus and treated as a binary sequence read; alternatively, the probability of the identity of the read and / or mapping can be incorporated as part of the sequence read, resulting in a probabilistic sequence read, i.e., the total or fractional number of possible sequence reads that map to a given locus. Using the binary counts or the probability of the counts, it is possible to use a binomial distribution for each set of measurements and calculate confidence intervals for the number of counts. This ability to use a binomial distribution allows for more accurate fold estimates and more accurate confidence intervals to be calculated. This is different from or an improvement over methods that use intensity to measure the amount of alleles present, such as using microarrays, or using a fluorescent reader to measure the intensity of fluorescently tagged DNA in electrophoretic bands.
[0051] In one embodiment, the methods disclosed herein use aspects of the instant dataset to determine the parameters of the estimated allele frequency distribution of the dataset. This is an improvement over methods that utilize a training set of data or a previous set of data to set the parameters of the current expected allele frequency distribution, or perhaps the expected allele ratios. This is because the collection and measurement of each genetic sample involves a different set of conditions, so methods that use data from an instant dataset to determine the parameters of the joint distribution model used to determine the folds of that sample tend to be more accurate.
[0052] In one embodiment, the method disclosed herein involves determining whether the distribution of observed allele measurements indicates a euploid or aneuploid fetus using a maximum likelihood technique. The use of the maximum likelihood technique is a significant improvement over methods using single hypothesis rejection techniques in that the resulting decisions are made with significantly greater accuracy. One reason is that the single hypothesis rejection technique sets the cutoff threshold based on only one measurement distribution instead of two, which means that the threshold is usually not optimal. Another reason is that the maximum likelihood technique allows the cutoff threshold to be optimized for each individual sample, instead of determining a cutoff threshold to be used for all samples regardless of the specific characteristics of the individual sample. Another reason is that the use of the maximum likelihood technique allows the calculation of the confidence of each ploidy call. The confidence calculation can be made for each call, so that the practitioner knows which calls are likely to be correct and which calls are likely to be incorrect. In some embodiments, a wide variety of methods may be combined with the maximum likelihood estimation technique to increase the accuracy of the ploidy calls. In one embodiment, the maximum likelihood technique may be used in combination with the method described in U.S. Pat. No. 7,888,017. In one embodiment, the maximum likelihood technique may be used in combination with a method of amplifying the DNA in a mixed sample using targeted PCR amplification, followed by sequencing and analysis using a read counting method such as that used by Tandem Diagnostics, as presented at the International Society of Human Genetics 2011, Montreal, October 2011. In one embodiment, the method disclosed herein includes estimating the fetal fraction of the DNA in the mixed sample and using that estimate to calculate both the ploidy call and the confidence of the ploidy call. Note that this is different from a method that uses the estimated fetal fraction as a screen for sufficient fetal fraction, followed by a ploidy call made using a single hypothesis rejection technique that does not consider the fetal fraction and does not generate a confidence calculation for the call.
[0053] In one embodiment, the method disclosed herein takes into account the tendency of data to be noisy and erroneous by attaching a probability to each measurement. Using maximum likelihood techniques to select the correct hypothesis from a set of hypotheses made using measurement data with attached probability estimates increases the likelihood that erroneous measurements are discounted and the correct measurements are used in the calculations that lead to the ploidy call. More precisely, the method systematically reduces the impact of erroneously measured data on ploidy determination. This is an improvement over methods in which all data are assumed to be equally correct or outlying data are arbitrarily excluded from the calculations that lead to the ploidy call. Existing methods that use channel ratio measurements claim to extend the method to multiple SNPs by averaging individual SNP channel ratios. Not weighting individual SNPs by expected measurement variance based on SNP quality and observed read depth reduces the accuracy of the resulting statistics and significantly reduces the accuracy of ploidy calls, especially in borderline cases.
[0054] In one embodiment, the method disclosed herein does not presuppose knowledge of which SNPs or other polymorphic loci are heterozygous in the fetus. This method allows ploidy calls to be made when paternal genotype information is not available. This is an improvement over methods that require prior knowledge of which SNPs are heterozygous in order to appropriately select loci to target or to interpret genetic measurements made on mixed fetal / maternal DNA samples.
[0055] The methods described herein are particularly advantageous when used on samples where a small amount of DNA is available or where the percentage of fetal DNA is low. This is due to the corresponding higher allele dropout rate and / or the corresponding higher fetal allele dropout rate that occurs when only a small amount of DNA is available when the percentage of fetal DNA in a mixed sample of fetal and maternal DNA is low. A high allele dropout rate means that a large proportion of alleles were not measured for the target individual, resulting in an inaccurate calculation of fetal fraction, an inaccurate ploidy determination. The methods disclosed herein may use a joint distribution model that takes into account the linkage in the inheritance pattern between SNPs, so that a significantly more accurate ploidy determination may be made. The methods described herein allow for accurate ploidy determination when the percentage of molecules of DNA that are fetal in the mixture is less than 40%, less than 30%, less than 20%, less than 10%, less than 8%, and less than 6%.
[0056] In one embodiment, it is possible to determine the ploidy state of an individual based on measurements when the individual's DNA is mixed with the DNA of a related individual. In one embodiment, the mixture of DNA is free floating DNA found in maternal plasma, which may include DNA from the mother with known karyotype and known genotype, and may be mixed with the DNA of the fetus with unknown karyotype and unknown genotype. Using known genotype information from one or both parents, it is possible to predict multiple potential genetic states of the DNA in the mixed sample for different ploidy states, different chromosomal contributions from each parent to the fetus, and optionally different fetal DNA fractions in the mixture. Each potential composition may be referred to as a hypothesis. The ploidy state of the fetus can then be determined by looking at the actual measurements and determining which potential composition is most likely given the observed data.
[0057] In some embodiments, the methods disclosed herein can be used in situations where very little DNA is present, such as in vitro fertilization, or in forensic situations where one or a few cells are available (typically less than 10 cells, less than 20 cells, or less than 40 cells). In these embodiments, the methods disclosed herein are useful for making ploidy calls from small amounts of DNA that are not contaminated by other DNA, but where it is very difficult to call small amounts of DNA. In some embodiments, the methods disclosed herein can be used in situations where the target DNA is contaminated with another individual's DNA, for example, in maternal blood in the context of prenatal testing, paternity testing, or product of conception testing. Some other situations where these methods would be particularly advantageous are in the case of cancer testing, where only one or a few cells are present among many more normal cells. The genetic measurements used as part of these methods may be performed on any sample containing DNA or RNA, including, but not limited to, blood, plasma, bodily fluids, urine, hair, tears, saliva, tissue, skin, fingernails, blastomeres, embryos, amniotic fluid, chorionic villus samples, feces, bile, lymph, cervical mucus, semen, or other cells or materials containing nucleic acids. In one embodiment, the methods disclosed herein can be performed with nucleic acid detection methods such as sequencing, microarray, qPCR, digital PCR, or other methods used to measure nucleic acids. If it is found to be desirable for any reason, the ratio of allele count probabilities at loci can be calculated, and the allele ratios can be used to determine ploidy state in combination with some of the methods described herein, provided that the methods are compatible. In some embodiments, the methods disclosed herein involve computationally calculating the allele ratios at multiple polymorphic loci from DNA measurements made on processed samples. In some embodiments, the methods disclosed herein involve computationally calculating allele ratios at multiple polymorphic loci from DNA measurements performed on the processed samples, along with any combination of other improvements described in this disclosure.
[0058] Non-Invasive Prenatal Testing (NPD) The process of non-invasive prenatal diagnosis involves several steps. Some of the steps may include: (1) obtaining genetic material from a fetus; (2) enriching the fetal genetic material, which may be in a mixed sample, ex vivo; (3) amplifying the genetic material ex vivo; (4) preferentially enriching specific loci in the genetic material ex vivo; (5) measuring the genetic material ex vivo; and (6) analyzing the genotype data on a computer and ex vivo. Methods are described herein that are reduced to carrying out these six and other related steps. At least some of the method steps are not directly applied to the body. In one embodiment, the present disclosure relates to therapeutic and diagnostic methods applied to tissues and other biological materials isolated and separated from the body. At least some of the method steps are performed on a computer.
[0059] Some embodiments of the present disclosure allow clinicians to determine the genetic status of a fetus gestating in a mother in a non-invasive manner, such that the baby's health is not endangered by collection of the fetal genetic material and such that the mother does not have to undergo an invasive procedure. Furthermore, in certain aspects, the present disclosure allows for the determination of fetal genetic status with greater accuracy, and significantly greater accuracy, than non-invasive maternal serum analyte-based screens, such as the triple test, that are widely used in prenatal care.
[0060] The high accuracy of the methods disclosed herein is the result of an informatics approach to the analysis of genotype data, as described herein. Modern technological advances have provided the ability to measure large amounts of genetic information from genetic samples using methods such as high-throughput sequencing and genotyping arrays. The methods disclosed herein allow clinicians to take greater advantage of the large amounts of data available and make more accurate diagnoses of fetal genetic conditions. Details of some embodiments are provided below. Different embodiments may involve different combinations of the aforementioned steps. Various combinations of different embodiments of different steps may be used interchangeably.
[0061] In one embodiment, a blood sample is taken from a pregnant mother and the free floating DNA in the plasma of the mother's blood, containing a mixture of both maternal and fetal DNA, is isolated and used to determine the ploidy state of the fetus. In one embodiment, the method disclosed herein involves preferential enrichment of those DNA sequences in a mixture of DNA corresponding to polymorphic alleles in such a way that the allele ratio and / or allele distribution remains largely consistent upon enrichment. In one embodiment, the method disclosed herein involves highly efficient targeted PCR-based amplification such that a very high percentage of the resulting molecules correspond to the target locus. In one embodiment, the method disclosed herein involves sequencing a mixture of DNA containing both maternal and fetal DNA. In one embodiment, the method disclosed herein involves using the measured allele distribution to determine the ploidy state of the fetus carried by the mother. In one embodiment, the method disclosed herein involves reporting the determined ploidy state to a clinician. In one embodiment, the methods disclosed herein include preparation for clinical action, e.g., performance of follow-up invasive testing such as chorionic villus sampling or amniocentesis, birth of a trisomic individual or selective termination of a trisomic fetus.
[0062] Screening maternal blood for free floating fetal DNA The methods described herein may be used to help determine the genotype of a child, fetus, or other target individual in which the target genetic material is found in the presence of a certain amount of other genetic material. In some embodiments, the genotype may refer to the ploidy state of one or more chromosomes, may refer to one or more disease-linked alleles, or some combination thereof. In this disclosure, the discussion focuses on determining the genetic status of a fetus in which fetal DNA is found in maternal blood, but this example is not intended to be limiting to the possible contexts in which the method may be applied. In addition, the method may be applied when the amount of target DNA is in any ratio with non-target DNA, for example, the target DNA may occupy anywhere from 0.000001 to 99.999999% of the DNA present. In addition, the non-target DNA does not necessarily have to be from one individual, or related individuals, so long as genetic data from some or all of the related non-target individual(s) is known. In one embodiment, the methods disclosed herein may be used to determine fetal genotype data from maternal blood containing fetal DNA. This may also be used when there are multiple fetuses in a pregnant woman's uterus, or when there may be other contaminating DNA in the sample, for example from other already-born siblings.
[0063] This technique may exploit the phenomenon that fetal blood cells access the maternal circulation through the placental villi. Usually, only a very small number of fetal cells enter the maternal circulation in this way (not enough to produce a positive Kleihauer-Betke test for fetal-maternal bleeding). Fetal cells can be sorted and analyzed by various techniques to look for specific DNA sequences, but without the risks that invasive procedures inherently have. This technique may also exploit the phenomenon of free floating fetal DNA, which accesses the maternal circulation by DNA release following apoptosis of placental tissue, where the placental tissue in question contains DNA of the same genotype as the fetus. Free floating DNA found in maternal plasma has been shown to contain fetal DNA at rates of 30-40% fetal DNA.
[0064] In one embodiment, blood may be collected from a pregnant woman. Research has shown that maternal blood may contain small amounts of free-floating DNA from fetus in addition to free-floating DNA from maternal origin. In addition, in addition to many blood cells of maternal origin that do not typically contain nuclear DNA, there may also be enucleated fetal blood cells that contain DNA from fetal origin. There are many methods known in the art for isolating fetal DNA or making fractions that are enriched in fetal DNA. For example, chromatography has been shown to make certain fractions that are enriched in fetal DNA.
[0065] When a sample of maternal blood, plasma, or other fluid containing an amount of fetal DNA, either cellular or free-floating, extracted in a relatively non-invasive manner and enriched in its proportion to maternal DNA or in its original proportion, is in hand, the DNA found in the sample may be genotyped. In some embodiments, blood may be collected using a needle to draw blood from a vein, for example, the basilica vein. The methods described herein may be used to determine fetal genotype data. For example, it may be used to determine the ploidy state at one or more chromosomes, and it may be used to determine the identity of one or a set of SNPs, including insertions, deletions, and translocations. It may be used to determine one or more haplotypes, including the parent of origin of one or more genotypic characteristics.
[0066] It should be noted that this method works with any nucleic acid that can be used for any genotyping and / or sequencing method, such as the ILLUMINA INFINIUM array platform, AFFYMETRIX GENECHIP, ILLUMINA genome analyzer, or LIFE TECHNOLGIES solid-state system. This includes free floating DNA extracted from plasma, or its amplification (e.g., whole genome amplification, PCR), genomic DNA from other cell types (e.g., human lymphocytes from whole blood), or its amplification. For preparation of DNA, any extraction or purification method that produces genomic DNA suitable for one of these platforms will work as well. This method may work well with RNA samples as well. In one embodiment, storage of the sample may be done in a way that minimizes degradation (e.g., under freezing, at about -20C, or at lower temperatures).
[0067] definition A single nucleotide polymorphism (SNP) refers to a single nucleotide that may differ between the genomes of two members of the same species. The use of this term should not imply any restriction on the frequency with which each variant occurs.
[0068] Sequence refers to a DNA sequence or gene sequence. It may refer to the primary physical structure of a DNA molecule or strand in an individual. It may refer to the sequence of nucleotides found within that DNA molecule, or a complementary strand to a DNA molecule. It may refer to the information contained in a DNA molecule as its representation in a computer.
[0069] A locus refers to a particular region of interest on an individual's DNA, which may refer to a SNP, a site of possible insertion or deletion, or the site of some other associated genetic variation. A disease-associated SNP may also refer to a disease-associated locus.
[0070] Polymorphic allele, also known as "polymorphic locus", refers to an allele or locus whose genotype varies among individuals within a given species. Some examples of polymorphic alleles include single nucleotide polymorphisms, short tandem repeats, deletions, duplications, and inversions.
[0071] A polymorphic site refers to the particular nucleotides found at a polymorphic region that differ between individuals.
[0072] An allele refers to a gene occupying a particular locus.
[0073] Genetic data, also known as "genotype data", refers to data describing aspects of the genome of one or more individuals. It may refer to one or a set of loci, partial or entire sequences, partial or entire chromosomes, or the entire genome. It may refer to the identity of one or more nucleotides, which may refer to a series of contiguous nucleotides, or nucleotides from different locations in the genome, or a combination thereof. Although genotype data is computational, it is also possible to consider the physical nucleotides in a sequence as chemically encoded genetic data. Genotype data may be said to be "on" an individual(s), "of" an individual(s), "at" an individual(s), "from" an individual(s), or "about" an individual. Genotype data may refer to output measurements from a genotyping platform where these measurements are made on genetic material.
[0074] Genetic material, also known as a "genetic sample," refers to physical material, such as tissue or blood, from one or more individuals that contains DNA or RNA.
[0075] Noisy genetic data refers to genetic data that has any of the following: allele dropouts, uncertain base pair measurements, inaccurate base pair measurements, missing base pair measurements, uncertain measurements of insertions or deletions, uncertain measurements of chromosomal segment copy numbers, false signals, missing measurements, other errors, or a combination thereof.
[0076] Confidence refers to the statistical likelihood that a determined number of so-called SNPs, alleles, sets of alleles, ploidy calls, or chromosome segment copies correctly represents an individual's actual genetic state.
[0077] Ploidy calling, also called "chromosome copy number calling" or "copy number calling" (CNC), can refer to the act of determining the amount and / or chromosomal identity of one or more chromosomes present in a cell.
[0078] Aneuploidy refers to the state that the wrong number of chromosomes exists in a cell.For human somatic cells, this can refer to the case where the cell does not contain 22 pairs of autosomes and one pair of sex chromosomes.For human gametes, this can refer to the case where the cell does not contain one of each of the 23 chromosomes.For single chromosome types, this can refer to the case where there are more than two or fewer homologous but not identical chromosome copies, or the case where there are two chromosome copies that come from the same parent.
[0079] Ploidy state refers to the amount and / or chromosomal identity of one or more chromosomes in a cell.
[0080] Chromosome can refer to a single chromosome copy, meaning a single molecule of DNA of which there are 46 in normal somatic cells, an example of which is "maternally derived chromosome 18." Chromosome can also refer to a chromosome type of which there are 23 in normal human somatic cells, an example of which is "chromosome 18."
[0081] Chromosomal identity may refer to the chromosome number, or chromosome type, of a subject. In a normal human, there are 22 numbered autosomal types and two types of sex chromosomes. It may also refer to the parental origin of a chromosome. It may also refer to a particular chromosome inherited from a parent. It may also refer to other identifying features of a chromosome.
[0082] The state of genetic material, or simply "genetic state," can refer to the identity of a set of SNPs on the DNA, the graded haplotype of the genetic material, and the sequence of the DNA, including insertions, deletions, repeats, and mutations. It can also refer to the ploidy state of one or more chromosomes, chromosomal segments, or sets of chromosomal segments.
[0083] Allelic data refers to a set of genotype data for a set of one or more alleles. It may refer to graded haplotype data. It may refer to SNP identities, and may refer to DNA sequence data, including insertions, deletions, repeats, and mutations. It may include the parental origin of each allele.
[0084] An allelic state refers to the actual state of a gene in a set of one or more alleles. It may refer to the actual state of a gene as described by the allelic data.
[0085] Allele ratio or allele ratio refers to the ratio between the amount of each allele at a locus present in a sample or individual.If a sample is measured by sequencing, the allele ratio can refer to the ratio of sequence reads that map to each allele at a locus.If a sample is measured by an intensity-based measurement method, the allele ratio can refer to the ratio of the amount of each allele present at that locus as estimated by the measurement method.
[0086] The allele count refers to the number of sequences that map to a particular locus, or, if the locus is polymorphic, the number of sequences that map to each of the alleles. If each allele is counted in a binary fashion, the allele count will be an integer. If alleles are counted probabilistically, the allele count may be a fraction.
[0087] Allele count probability, in combination with the probability of mapping, refers to the number of sequences that are likely to map to a particular locus or set of alleles at a polymorphic locus. Note that allele count is equivalent to allele count probability, where the probability of mapping for each counted sequence is binary (zero or one). In some embodiments, the allele count probability may be binary. In some embodiments, the allele count probability may be set equal to the DNA measurement.
[0088] Allele distribution, or "allele number distribution", refers to the relative amount of each allele present for each locus in a set of loci. Allele distribution can refer to an individual, a sample, or a set of measurements made on a sample. In the context of sequencing, allele distribution refers to the number of reads or possible numbers that map to a particular allele for each allele in a set of polymorphic loci. Allele measurements may be treated probabilistically, i.e., the likelihood that a given allele is present for a given sequence read is a fraction between 0 and 1, or they may be treated in a binary manner, i.e., a given read is considered to be exactly zero or one copy of a particular allele.
[0089] An allele distribution pattern refers to a set of different allele distributions for different parental contexts. A particular allele distribution pattern may indicate a particular ploidy state.
[0090] Allelic bias refers to the degree to which the ratio of alleles measured at a heterozygous locus differs from the ratio present in the original sample of DNA. The degree of allelic bias at a particular locus is the observed allele ratio at that locus at the time of measurement divided by the ratio of alleles in the original DNA sample at that locus. Allelic bias may be defined as greater than 1, such that if a calculation of the degree of allelic bias returns a value x less than 1, the degree of allelic bias may be rewritten as 1 / x. Allelic bias may be due to amplification bias, purification bias, or some other phenomenon that affects different alleles differently.
[0091] A primer, also known as a "PCR probe," refers to a single DNA molecule (DNA oligomer) or a collection of DNA molecules (DNA oligomers) where the DNA molecules are identical or nearly identical, the primer contains a region designed to hybridize to a target polymorphic locus, and m contains a priming sequence designed to allow PCR amplification. A primer may also contain a molecular barcode. A primer may contain a random region that is different for each individual molecule.
[0092] Hybrid capture probe refers to any nucleic acid sequence, optionally modified, that is generated by various methods, such as PCR or direct synthesis, and is intended to be complementary to one strand of a specific target DNA sequence in a sample. Exogenous hybrid capture probes may be added to the prepared sample and hybridized through a denaturation-annealing process to form exogenous-endogenous fragment duplexes. These duplexes may then be physically separated from the sample by various means.
[0093] A sequence read refers to data representing a sequence of nucleotide bases measured using a clonal sequencing method. Clonal sequencing may generate sequence data representing a single, or a clone, or a cluster of original DNA molecules. A sequence read may also have an associated quality score for each base position of the sequence that indicates the likelihood that the nucleotide was called correctly.
[0094] Mapping a sequence read is the process of determining the location of the origin of a sequence read in the genomic sequence of a particular organism. The location of the origin of a sequence read is based on the nucleotide sequence similarity of the read and the genomic sequence.
[0095] Matched copy errors, also called "matching chromosome aneuploidy" (MCA), refer to a state of aneuploidy in which one cell contains two identical or nearly identical chromosomes. This type of aneuploidy may arise during the formation of gametes in meiosis and may be called meiotic non-disjunction errors. This type of error may occur in mitosis. Matched trisomy may refer to the case where three copies of a given chromosome are present in an individual and two of the copies are identical.
[0096] Unmatched copy error, also called "unique chromosome anomaly" (UCA), refers to a state of aneuploidy in which one cell contains two chromosomes from the same parent that may be homologous but not identical. This type of aneuploidy may occur during meiosis and may be called meiotic error. Unmatched trisomy may refer to the case where three copies of a given chromosome are present in an individual, two of the copies are from the same parent and are homologous but not identical. Note that unmatched trisomy may refer to the case where two homologous chromosomes from one parent are present, some segments of the chromosome are identical and other segments are simply homologous.
[0097] Homologous chromosomes refer to chromosome copies that contain the same set of genes that normally pair up during meiosis.
[0098] Identical chromosomes refer to chromosome copies that contain the same set of genes, and for each gene, they have the same set of identical or nearly identical alleles.
[0099] Allelic dropout (ADO) refers to the situation where at least one of the base pairs in a set of base pairs from the homologous chromosome at a given allele is not detected.
[0100] Locus dropout (LDO) refers to the situation where both base pairs within a set of base pairs from the homologous chromosome at a given allele are not detected.
[0101] Homozygous refers to having similar alleles at corresponding chromosomal loci.
[0102] Heterozygosity refers to the possession of different alleles at corresponding chromosomal loci.
[0103] Heterozygosity rate refers to the proportion of individuals in a group that have heterozygous alleles at a given locus, and may refer to the expected or measured proportion of alleles at a given locus in an individual or sample of DNA.
[0104] A highly informative single nucleotide polymorphism (HISNP) refers to a SNP in which the fetus carries an allele that is not present in the maternal genotype.
[0105] A chromosomal region refers to a segment of a chromosome or a complete chromosome.
[0106] A chromosomal segment refers to a section of a chromosome, which can range in size from a single base pair to an entire chromosome.
[0107] Chromosome refers to either a complete chromosome or a segment or section of a chromosome.
[0108] Copy refers to the number of copies of chromosomal segments.It can refer to the same copy of chromosomal segments, or the non-identical homologous copies of chromosomal segments, where different copies of chromosomal segments contain a substantially similar set of loci and differ in one or more alleles.Please note that in some cases of aneuploidy, such as M2 copy error, it is possible to have some copies of a given chromosomal segment that are identical and some copies of the same chromosomal segment that are not identical.
[0109] A haplotype typically refers to a combination of alleles at multiple loci that are inherited together on the same chromosome. A haplotype can refer to as few as two loci or an entire chromosome, depending on the number of recombination events that have occurred between a given set of loci. A haplotype can also refer to a set of single nucleotide polymorphisms (SNPs) on a single chromatid that are statistically associated.
[0110] Haplotype data, also called "phasing data" or "ordered gene data", refers to data from a single chromosome in a diploid or polyploid genome, i.e., either the separated maternal or paternal copies of a chromosome in a diploid genome.
[0111] Phasing refers to the act of determining haplotype genetic data for an individual given unordered diploid (or polyploid) genetic data. It may refer to the act of determining which of the two genes in the allele for a set of alleles found on one chromosome associates with each of the two homologous chromosomes in the individual.
[0112] Phasing data refers to genetic data for which one or more haplotypes have been determined.
[0113] A hypothesis refers to a set of possible ploidy states at a given set of chromosomes or possible allelic states at a given set of loci. A set of possibilities may contain one or more elements.
[0114] Copy number hypothesis, also called "ploidy state hypothesis", refers to a hypothesis about the number of copies of chromosomes in an individual. It can also refer to a hypothesis about the identity of each chromosome, including the parent of origin of each chromosome and which of the two chromosomes of the parent are present in the individual. It can also refer to a hypothesis about which chromosome or chromosome segment from related individuals in some cases corresponds genetically to a given chromosome from an individual.
[0115] Target individual refers to an individual whose genetic status is being determined. In some embodiments, only a limited amount of DNA is available from the target individual. In some embodiments, the target individual is a fetus. In some embodiments, there may be more than one target individual. In some embodiments, each fetus from a parent pair may be considered a target individual. In some embodiments, the genetic data determined is a set of allele calls. In some embodiments, the genetic data determined is a ploidy call.
[0116] A related individual refers to any individual who is genetically related to the target individual and therefore shares a haplotype block. In one context, a related individual may be a genetic parent of the target individual, or any genetic material derived from a parent, such as a sperm, polar body, embryo, fetus, or child. It may also refer to a sibling, parent, or grandparent.
[0117] Sibling refers to any individual whose genetic parent is the same as the individual in question. In some embodiments, it can refer to a live-born child, embryo or fetus, or one or more cells derived from an embryo or fetus at birth. Sibling can also refer to a monoploid individual derived from one parent, such as a sperm, a polar body, or any other set of haplotype genetic material. An individual can be considered to be a sibling of itself.
[0118] Fetal refers to "of the fetus" or "of a region of the placenta that is genetically similar to the fetus." In pregnant women, parts of the placenta are genetically similar to the fetus, and free floating fetal DNA found in maternal blood may come from parts of the placenta that have a genotype that matches the fetus. Note that the genetic information of half of the chromosomes in a fetus is inherited from the fetus's mother. In some embodiments, DNA from these maternally inherited chromosomes that originate from fetal cells is considered to be of "fetal origin" rather than "maternal origin."
[0119] DNA of fetal origin refers to DNA that was originally part of a cell whose genotype was essentially identical to that of the fetus.
[0120] DNA of maternal origin refers to DNA that was originally part of cells whose genotype was essentially identical to that of the mother.
[0121] A child may refer to an embryo, a blastocyst, or a fetus. It is noted that in embodiments of the present disclosure, the concepts described apply equally well to an individual that is a live-born baby, a fetus, an embryo, or a set of cells therefrom. Use of the term child may simply mean that the individual referred to as a child is the genetic offspring of the parents.
[0122] Parent refers to the genetic mother or father of an individual. An individual typically has two parents, a mother and a father, although this may not necessarily be the case, such as in genetic or chromosomal chimeras. A parent may be considered an individual.
[0123] Parental context refers to the genetic state of a given SNP in each of the two relevant chromosomes of one or both of the target's two parents.
[0124] Desired development (also known as "normal development") refers to the implantation of a viable embryo into the uterus resulting in a pregnancy, and / or the continuation of the pregnancy resulting in a live birth, and / or the absence of chromosomal abnormalities in the born child, and / or the absence of other undesirable genetic conditions in the born child, such as disease-related genes. The term "desired development" is meant to encompass anything that may be desired by a parent or health care facilitator. In some cases, "desired development" may refer to a non-viable or viable embryo that is useful for medical research or other purposes.
[0125] Uterine insertion refers to the process of implanting an embryo into the uterine cavity in the context of in vitro fertilization.
[0126] Maternal plasma refers to the plasma portion of the blood from a pregnant woman.
[0127] A clinical decision refers to a decision to take an action that has consequences that affect the health or survival of an individual. In the context of prenatal diagnosis, a clinical decision may refer to a decision to abort or not abort a fetus. A clinical decision may also refer to performing further testing, taking steps to mitigate an undesirable phenotype, or taking an action to prepare for the birth of a child with an abnormality.
[0128] The diagnostic box refers to one or a combination of machines designed to perform one or more aspects of the method disclosed herein. In one embodiment, the diagnostic box may be located at the point of care of the patient. In one embodiment, the diagnostic box may perform sequencing after targeted amplification. In one embodiment, the diagnostic box may function alone or with the assistance of a technician.
[0129] Informatics-based methods refer to methods that rely heavily on statistics to make sense of large amounts of data. In the context of prenatal diagnosis, this refers to methods designed to determine the ploidy state at one or more chromosomes or the allele state at one or more alleles by statistically inferring the most likely state, rather than by directly physically measuring the state, given a large amount of genetic data, for example from molecular arrays or sequencing. In one embodiment of the present disclosure, the informatics-based technology may be that disclosed in this patent. In one embodiment of the present disclosure, it may be PARENTAL SUPPORT™.
[0130] Primary genetic data refers to the analog intensity signal output by a genotyping platform. In the context of SNP arrays, primary genetic data refers to the intensity signal before any genotype calls are made. In the context of sequencing, primary genetic data refers to the analog measurements, similar to chromatograms, that come out of a sequencer before the identity of any base pairs is determined and before the sequence is mapped to a genome.
[0131] Secondary genetic data refers to the processed genetic data output by the genotyping platform. In the context of SNP arrays, secondary genetic data refers to the allele calls made by the software associated with the SNP array reader, where the software calls whether a given allele is present or absent in a sample. In the context of sequencing, secondary genetic data refers to the base pair identity of a sequence being determined, and in some cases, the sequence being mapped to a genome.
[0132] Non-invasive prenatal diagnosis (NPD), or "non-invasive prenatal screening" (NPS), refers to a method of determining the genetic status of a fetus carried by a mother using genetic material found in the mother's blood, which is obtained by drawing the mother's intravenous blood.
[0133] Preferential enrichment of DNA corresponding to a locus, or preferential enrichment of DNA at a locus, refers to any method that results in a higher percentage of molecules of DNA in a post-enrichment DNA mixture that corresponds to a locus than the percentage of molecules of DNA in the pre-enrichment DNA mixture that correspond to the locus. The method may include selective amplification of DNA molecules that correspond to the locus. The method may include removing DNA molecules that do not correspond to the locus. The method may include a combination of methods. Enrichment is defined as the percentage of DNA molecules in the post-enrichment mixture that correspond to the locus divided by the percentage of DNA molecules in the pre-enrichment mixture that correspond to the locus. Preferential enrichment may be performed at multiple loci. In some embodiments of the present disclosure, the enrichment is greater than 20. In some embodiments of the present disclosure, the enrichment is greater than 200. In some embodiments of the present disclosure, the enrichment is greater than 2,000. When preferential enrichment is performed at multiple loci, the enrichment may refer to the average enrichment of all loci in the set of loci.
[0134] Amplification refers to a method for increasing the number of copies of a molecule of DNA.
[0135] Selective amplification can refer to a method of increasing the copy number of a particular molecule of DNA or a molecule of DNA corresponding to a particular region of DNA. It can also refer to a method of increasing the copy number of a particular target molecule of DNA or a target region of DNA over non-target molecules or regions of DNA. Selective amplification can be a method of preferential enrichment.
[0136] Universal priming sequence refers to a DNA sequence that can be added to a population of target DNA molecules, for example, by ligation, PCR, or ligation-mediated PCR. Once added to a population of target molecules, a primer specific to the universal priming sequence can be used to amplify the target population using a single amplification primer pair. The universal priming sequence is typically not related to the target sequence.
[0137] A universal adaptor, or "ligation adaptor" or "library tag," is a DNA molecule that contains a universal priming sequence that can be covalently attached to the 5' and 3' ends of a population of target double-stranded DNA molecules. The addition of the adaptor provides universal priming sequences at the 5' and 3' ends of the target population from which PCR amplification can be performed, amplifying all molecules from the target population using a single amplification primer pair.
[0138] Targeting refers to methods used to selectively amplify or otherwise preferentially enrich molecules of DNA that correspond to a set of loci in a mixture of DNA.
[0139] A joint distribution model refers to a model that defines the probability of an event defined in terms of multiple random variables, given multiple random variables defined on the same probability space, where the probabilities of the variables are linked. In some embodiments, a degenerate case may be used where the probabilities of the variables are not linked.
[0140] hypothesis In the context of the present disclosure, a hypothesis refers to a possible genetic state. It may refer to a possible ploidy state. It may refer to a possible allelic state. A set of hypotheses may refer to a set of possible genetic states, a set of possible allelic states, a set of possible ploidy states, or a combination thereof. In some embodiments, a set of hypotheses may be designed such that one hypothesis from the set corresponds to the actual genetic state of any given individual. In some embodiments, a set of hypotheses may be designed such that all possible genetic states can be explained by at least one hypothesis from the set. In some embodiments of the present disclosure, one aspect of the method is to determine which hypothesis corresponds to the actual genetic state of the individual in question.
[0141] In another embodiment of the present disclosure, a step involves making a hypothesis. In some embodiments, it may be a copy number hypothesis. In some embodiments, it may involve a hypothesis about which segments of chromosomes from each related individual correspond genetically to which segments, if any, of other related individuals. Making a hypothesis may refer to the act of setting the limits of variables so that the entire set of possible genetic conditions under consideration is encompassed by those variables.
[0142] A "copy number hypothesis", also called a "ploidy hypothesis" or a "ploidy state hypothesis", may refer to a hypothesis regarding the possible ploidy state of a given chromosome copy, chromosome type, or chromosome segment in a target individual. It may also refer to two or more ploidy states of a chromosome type in an individual. A set of copy number hypotheses may refer to a set of hypotheses, each corresponding to a different possible ploidy state in an individual. A set of hypotheses may be for a set of possible ploidy states, a set of possible parental haplotype contributions, a set of possible percentages of fetal DNA in a mixed sample, or a combination thereof.
[0143] A normal individual contains one of each chromosome type from each parent. However, due to errors in meiosis and mitosis, it is possible for an individual to have zero, one, two, or more of a given chromosome type from each parent. In practice, it is rare to see more than one of a given chromosome from a parent. In this disclosure, some embodiments consider only the possible hypotheses that zero, one, or two copies of a given chromosome originate from a parent. It is a simple extension to consider more or less likely copies originating from a parent. In some embodiments, for a given chromosome, there are nine possible hypotheses: three possible hypotheses for zero, one, or two chromosomes of maternal origin multiplied by zero, one, or two possible hypotheses of paternal origin. (m,f) refers to the hypothesis that m is the number of a given chromosome inherited from the mother and f is the number of a given chromosome inherited from the father. Thus, the nine hypotheses are (0,0), (0,1), (0,2), (1,0), (1,1), (1,2), (2,0), (2,1) and (2,2). These are H 00 , H 01 , H 02 , H 10 , H 12 , H 20 , H 21 , and H 22 It may also be written as: 1,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,14,15,23,24,25,30,31,32,16,17,26,33,27,28,34,35,46,47,50,51,66,77,84,95,96,10,11,12,13,14,15,29,36,37,48,18,15,29,38,19,40,39,41,14,29,42,15,36,43,44,51,67,84,95,96,10,11,12,13,29,34,15,35,46,52,14,15,29,36,47,53,16,17,29,34,18,35,48,19,41,14,29,36,47,54,18,15,36,47,55,66,77,84,95,96,10,11,12,29,36,47,56,77,96,10,11,12,29,36,47,57,10,11,12,3,4,5,6,7,8,9,10,13,14,29,47,58,10,11,12,29,36,47,59,11,29,48,12,3,4,5,6,7,8,9,10,11,12,13,29,47,59,10
[0144] In some embodiments of the present disclosure, the ploidy hypothesis refers to the hypothesis that chromosomes from other related individuals correspond to chromosomes found in the genome of the target individual.In some embodiments, the key to the method is the fact that related individuals can be expected to share haplotype blocks, and using the measured genetic data from related individuals, together with the knowledge of which haplotype blocks match between the target individual and related individuals, it is possible to infer the correct genetic data for the target individual with higher reliability than using only the genetic measurements of the target individual.Thus, in some embodiments, the ploidy hypothesis can relate not only to the number of chromosomes, but also to which chromosomes in related individuals are identical or nearly identical to one or more chromosomes in the target individual.
[0145] Once a set of hypotheses is defined, as the algorithms operate on the input genetic data, they may output a determined statistical probability for each of the hypotheses under consideration. The probabilities of the various hypotheses may be determined by mathematically calculating, for each of the various hypotheses using the relevant genetic data as input, a value that has an equal probability, as described by one or more of the expert techniques, algorithms, and / or methods described elsewhere in this disclosure.
[0146] Once the probabilities of different hypotheses, as determined by multiple techniques, have been estimated, they can be combined. This may involve multiplying the probabilities determined by each technique for each hypothesis. The product of the hypothesis probabilities may be normalized. Note that one ploidy hypothesis refers to one possible ploidy state for a chromosome.
[0147] The process of "combining probabilities", also called "combining hypotheses", or combining the results of expert techniques, is a concept that should be familiar to those skilled in the art of linear algebra. One possible way of combining probabilities is as follows: When using expert techniques to evaluate a set of hypotheses given a set of genetic data, the output of the method is a set of probabilities associated in a one-to-one manner to each hypothesis in the set of hypotheses. When a set of probabilities determined by a first expert technique, each associated with one of the hypotheses in the set, is combined with a set of probabilities determined by a second expert technique, each associated with the same set of hypotheses, the two sets of probabilities are multiplied. This means that for each hypothesis in the set, the two probabilities associated with that hypothesis, as determined by the two expert methods, are multiplied together, and the corresponding product is the output probability. This process may be extended to any number of expert techniques. If only one expert technique is used, the output probability is the same as the input probability. If more than one expert technique is used, then the associated probabilities may be multiplied simultaneously. The products may be normalized so that the probability of a hypothesis in the set of hypotheses is 100%.
[0148] In some embodiments, if the combined probability for a given hypothesis is greater than the combined probability for any of the other hypotheses, that hypothesis may be considered to be determined to be most likely. In some embodiments, a hypothesis may be determined to be most likely, and if the normalized probability is greater than a threshold, a ploidy state, or other genetic state, may be called. In one embodiment, this may mean that the number and identity of chromosomes associated with that hypothesis may be referred to as a ploidy state. In one embodiment, this may mean that the identity of alleles associated with that hypothesis may be referred to as an allele state. In some embodiments, the threshold may be about 50% to about 80%. In some embodiments, the threshold may be about 80% to about 90%. In some embodiments, the threshold may be about 90% to about 95%. In some embodiments, the threshold may be about 95% to about 99%. In some embodiments, the threshold may be about 99% to about 99.9%. In some embodiments, the threshold may be greater than about 99.9%.
[0149] A ploidy hypothesis is created during an exemplary method of the present invention using a method, algorithm, technique, or subroutine that provides a likelihood. For example, in a specific illustrative example of an embodiment for determining the presence or absence of aneuploidy, a set of ploidy hypotheses is created for each sample in a set of samples, with each hypothesis associated with a particular copy number for a chromosome or chromosome segment of interest in the genome of the sample. For example, in an embodiment using quantitative non-allelic data such as QMM disclosed herein, the hypothesis can provide an estimate of a sample parameter, such as the variance of the starting amount of DNA in the sample due to pipetting variance or error or other measurement error, which can be used to normalize measurements (i.e., the measured genetic data) at some or all positions of the chromosome or chromosome segment of interest in that sample, and then a test statistic can be calculated as a variance-weighted average of these normalized measurements. Thus, in certain embodiments, the hypothesis provides a variance-weighted average test statistic for a given ploidy condition. The expectation and variance of the test statistic are calculated under each of the chromosome copy number hypotheses to form a Gaussian model of the maximum likelihood estimate. For example, a set of hypotheses in a NIPT analysis for non-allelic quantitative analysis can provide a variance-weighted average test statistic for disomy or trisomy in one or more of chromosomes 13, 18, and 21. In an exemplary embodiment of the invention in which a chromosome or chromosome segment of interest can be used to set sample parameters, the hypothesis can be a joint hypothesis for some or all copy numbers of chromosomes, e.g., chromosomes 13, 18, and 21. This is further described below with respect to quantitative methods that do not use non-target reference chromosomes.
[0150] In some embodiments of the present disclosure, the ploidy hypothesis may refer to the hypothesis that chromosomes from other related individuals correspond to the chromosomes found in the genome of the target individual. Some embodiments are the fact that related individuals may be expected to share haplotype blocks, and using measured genetic data from related individuals, together with knowledge of which haplotype blocks match between the target individual and related individuals, it is possible to infer correct genetic data for the target individual with higher confidence than using only the genetic measurements of the target individual. Thus, in some embodiments, the ploidy hypothesis may relate not only to the number of chromosomes, but also to which chromosomes in related individuals are identical or nearly identical to one or more chromosomes in the target individual.
[0151] An allele hypothesis, or "allele state hypothesis", may refer to a hypothesis regarding the possible allele states of a set of alleles. In some embodiments, the techniques, algorithms, or methods used take advantage of the fact that related individuals may share haplotype blocks, as described above, which may aid in the reconstruction of genetic data that was not completely measured. An allele hypothesis may also refer to a hypothesis regarding which chromosome, or chromosome segment, from related individuals, in some cases, corresponds genetically to a given chromosome from an individual. The theory of meiosis indicates that each chromosome of an individual is inherited from one of two parents, which is a nearly identical copy of the parent's chromosome. Thus, if the parental haplotypes, i.e., the phased genotypes of the parents, are known, the genotype of the child may also be inferred (the term "child" here is meant to include an individual formed from two gametes, one from the mother and one from the father). In one embodiment of the present disclosure, an allele hypothesis describes the possible allele states at a set of alleles that includes a haplotype at a chromosome or chromosome segment of interest, as well as which chromosomes from related individuals may match the chromosome(s) containing the set of alleles.
[0152] Once a set of hypotheses is defined, the algorithm operates on the input genetic data and outputs a determined statistical probability for each of the hypotheses under consideration. For example, in an embodiment of the invention, the method determines a probability value by comparing the genetic data with the predicted outcome for each hypothesis, the probability value indicating the likelihood that a sample has a particular number of copies of a chromosome or chromosome segment associated with the hypothesis.
[0153] The probabilities of the various hypotheses can be determined by mathematically calculating, for each of the various hypotheses, values that have equal probabilities, using the relevant genetic data as input, as described by one or more of the expert techniques, algorithms, and / or methods described elsewhere in this disclosure.
[0154] Once the probabilities of different hypotheses, as determined by multiple techniques, have been estimated, they can be combined. This may involve multiplying the probabilities determined by each technique for each hypothesis. The product of the hypothesis probabilities may be normalized. Note that one ploidy hypothesis refers to one possible ploidy state for a chromosome.
[0155] The process of "combining probabilities", also called "combining hypotheses", or combining the results of expert techniques, is a concept that should be familiar to those skilled in the art of linear algebra. In an exemplary method of the present invention, two methods are utilized to determine the presence or absence of aneuploidy or to determine the copy number of a chromosome, each of which provides a probability. In certain exemplary embodiments, the reliability of the determination is increased by combining the reliability selected for each method. For example, the reliability for a first method performing quantitative allelic analysis can be combined with the reliability from a second method performing quantitative non-allelic analysis.
[0156] In cases where the likelihood is determined by the first method in a way that is orthogonal or unrelated to the way in which it is determined for the second method, combining the likelihoods is straightforward and can be done by multiplication and normalization, or by using a formula such as R comb =R 1 R 2 / [R 1 R 2 +(1-R 1 )(1-R 2 )]
[0157] In the formula, R comb is the joint likelihood, R 1 and R 2 are the individual likelihoods. If the first and second methods are not orthogonal, i.e., there is correlation between the two methods, the likelihoods can still be combined, although the mathematics may be more complicated.
[0158] In some embodiments, the first probability and the second probability are weighted differently before combining the probabilities. In some embodiments, the first probability and the second probability are considered as independent events for the purpose of combining the two probability values. In some embodiments, the first probability and the second probability are considered as dependent events for the purpose of combining the two probability values. In some embodiments, the method further comprises obtaining a third probability value, the third probability value being indicative of the likelihood that the genome of the target has a copy number of a chromosome or chromosome segment associated with a particular hypothesis, the third probability value being derived from information that is a non-genetic clinical assay. Many non-genetic clinical assays have known probabilistic correlations with particular chromosome copy numbers or chromosome segment copy numbers. For each hypothesis, the combined first and second probability values may be combined with the third probability value to give a combined probability value indicative of the likelihood that the genome of the target cell has a copy number of the chromosome or chromosome segment of interest, which number is associated with a particular hypothesis. An example of such a non-genetic clinical assay includes cervical translucency measurement. In some embodiments, the non-genetic clinical assay is selected from the group consisting of measurement of beta-human chorionic gonadotropin, pregnancy associated plasma protein A, estriol, inhibin-A, and alpha-fetoprotein.
[0159] Without being limited by theory, the following disclosure further teaches how to combine probabilities. One possible way to combine probabilities is as follows: When using expert techniques to evaluate a set of hypotheses given a set of genetic data, the output of the method is a set of probabilities associated in a one-to-one manner to each hypothesis in the set of hypotheses. When a set of probabilities determined by a first expert technique, each associated with one of the hypotheses in the set, is combined with a set of probabilities determined by a second expert technique, each associated with the same set of hypotheses, the two sets of probabilities are multiplied. This means that for each hypothesis in the set, the two probabilities associated with that hypothesis, as determined by the two expert methods, are multiplied together, and the corresponding product is the output probability. This process may be extended to any number of expert techniques. If only one expert technique is used, the output probability is the same as the input probability. If more than one expert technique is used, then the associated probabilities may be multiplied simultaneously. The products may be normalized so that the probability of a hypothesis in the set of hypotheses is 100%.
[0160] In some embodiments, if the combined probability for a given hypothesis is greater than the combined probability for any of the other hypotheses, that hypothesis may be considered to be determined to be most likely. In some embodiments, a hypothesis may be determined to be most likely, and if the normalized probability is greater than a threshold, a ploidy state, or other genetic state, may be called. In one embodiment, this means that the number and identity of chromosomes associated with that hypothesis may be referred to as the ploidy state. In one embodiment, this means that the identity of alleles associated with that hypothesis is referred to as the allele state. In some embodiments, the threshold is about 50% to about 80%. In some embodiments, the threshold is about 80% to about 90%. In some embodiments, the threshold is about 90% to about 95%. In some embodiments, the threshold is about 95% to about 99%. In some embodiments, the threshold is about 99% to about 99.9%. In some embodiments, the threshold is greater than 99.9%. In other embodiments, a set of rules is used for the final risk call of the sample where a combined probability threshold is set, but different scenarios may be considered and may override the results of the probability threshold or may be used to enhance the calling power of the combined probability. For example, if there is a large difference in the probability of a given ploidy hypothesis, further analysis may be performed to determine, for example, whether there was an error in one of the methods.
[0161] Some embodiments of the invention employ a process of generating a subset of patients from a larger set of patients. The original set of patients is used as a source of target and non-target cells for analysis. In some embodiments of the invention, DNA samples obtained from patients are modified using standard molecular biology techniques in order to be sequenced on a DNA sequencer. In some embodiments, this technique involves forming a genetic library containing priming sites for the DNA sequencing procedure. In some embodiments, multiple loci may be targeted for site-specific amplification. In some embodiments, the target loci are polymorphic loci, e.g., single nucleotide polymorphisms. In embodiments involving the formation of a genetic library, the library can be coded using patient-specific DNA sequences, e.g., barcoding, thereby allowing multiple patients to be analyzed in a single flow cell (or flow cell equivalent) of a high-throughput DNA sequencer. Although the samples are mixed together in the DNA sequencer flow cell, determining the sequence of the barcodes allows for the identification of the patient source that contributed to the sequenced DNA.
[0162] One of skill in the art will appreciate that in those embodiments of the invention where the target DNA is not enriched for a particular locus, the entire genome may be sequenced, but assembly of the sequences into a complete genome is not required for use of the methods of the invention. Information regarding a particular locus can be readily determined from full genome sequencing.
[0163] In one embodiment of the present disclosure, the confidence may be calculated based on the accuracy of the determination of the ploidy status of the fetus. major The confidence level of the hypothesis is (1-H major / Σ(all H)). If the distributions of all of the hypotheses are known, it is possible to determine the confidence of the hypotheses. If the parent genotype information is known, it is possible to determine the distributions of all of the hypotheses. If knowledge of the expected distribution of the data for euploid fetuses and the expected distribution of the data for aneuploid fetuses is known, it is possible to calculate the confidence of the ploidy determination. If the parent genotype data is known, it is possible to calculate these expected distributions. In one embodiment, knowledge of the distribution of the test statistics around the normal hypothesis and around the abnormal hypothesis can be used to determine both the confidence of the call and refine the threshold to make a more confident call. This is particularly useful when the amount and / or percentage of fetal DNA in the mixture is low. It helps to avoid situations where a fetus that is actually aneuploid turns out to be euploid because the test statistic, such as the Z statistic, does not exceed the threshold created based on the optimized threshold when the percentage of fetal DNA is high.
[0164] Method for determining copy number of a chromosome or chromosome segment of interest by combining allelic and non-allelic data - Patent Application 20070229333 Another embodiment of the invention includes a method for determining the copy number of a chromosome or chromosome segment of interest in the genome of a target cell, such as a fetal cell or a tumor cell. Genetic data, e.g., DNA sequence data, can be obtained from a mixture of DNA, including DNA from one or more target cells and DNA from one or more non-target cells. The method can use a single patient or a set of patients. Genetic data is obtained from a patient. Genetic information is obtained at multiple loci. At least some of the loci, and potentially all of the loci, are polymorphic. The same loci are analyzed in both the target cell and the non-target cell. Several sequence reads are obtained for each locus. The number of sequence reads at each allele at a given locus is quantified. The quantitative data obtained can be obtained from a combination of the loci from the target cell and the non-target cell genome. The collected data is then tested against multiple copy number hypotheses, i.e., copy numbers of the chromosome or chromosome segment of interest. A first probability value is calculated for each hypothesis, i.e., the probability that the hypothesis is either true or false, given the measured genetic data. Thus, the likelihood that the genome of the target cell has the copy number of the chromosome or chromosome segment of interest specified by the hypothesis is determined. This first probability value is obtained using allelic data. A second probability value is calculated for each hypothesis, i.e., the probability that the hypothesis is either true or false, taking into account the measured genetic data. Thus, the likelihood that the genome of the target cell has the copy number of the chromosome or chromosome segment of interest specified by the hypothesis is determined. This second probability value is obtained using non-allelic data. For each hypothesis, the first probability value and the second probability value can be combined, for example by multiplication, to give a composite probability indicating the likelihood that the genome of the target cell has the copy number of the chromosome or chromosome segment associated with the hypothesis. The copy number of the chromosome or chromosome segment of interest in the genome of the target cell can be determined by selecting the copy number of the chromosome or chromosome segment associated with the hypothesis with the greatest binding probability, which is used to perform the determination of the copy number of the chromosome or chromosome segment in the sample of interest.In some embodiments where the genetic data is obtained from cell-free DNA obtained from the blood of a pregnant woman, hypotheses can include situations in which the mother is carrying multiple fetuses, for example twins.
[0165] Thus, in some embodiments, genetic data is obtained by simultaneously sequencing a mixture containing DNA from one or more target cells and from one or more non-target cells to provide genetic data at a set of loci from each member of a set of patients. In some embodiments, the target cells are fetal cells and the non-target cells are from the mother of the fetus. That is, in some embodiments directed to non-invasive prenatal diagnosis, the target cells may be fetal cells and the non-target cells may be mother cells. In some embodiments of the present invention, an example of a hypothesis that may be used to select a subset of patients may be the hypothesis that a particular chromosome or chromosome segment is diploid, i.e., exists in two copies. Examples of chromosomes for analysis include chromosomes 13, 18, 21, X and Y, which contain the segment. In some embodiments, the chromosomal segment analyzed for copy number is selected from the group consisting of chromosome 22q11.2, chromosome 1p36, chromosome 15q11-q13, chromosome 4p16.3, chromosome 5p15.2, chromosome 17p13.3, chromosome 22q13.3, chromosome 2q37, chromosome 3q29, chromosome 9q34, chromosome 17q21.31, and the ends of the chromosomes.
[0166] In some embodiments, the set of loci is present in a selected region of a chromosome. In some embodiments, the method is performed independently for different chromosomes or chromosome segments. The only upper limit imposed on the number of patients in the patient set is imposed by the DNA sequence generation capacity of the particular DNA sequencing technology selected (including patient multiplexing technology, e.g., barcoding technology, that is compatible with the sequencing technology). In an exemplary embodiment, there are at least 10 patients in the patient set. In some embodiments, there are at least 24 patients, in other embodiments, the patient set has at least 48 patients, and in other embodiments, the patient set has at least 96 patients in the patient set.
[0167] Methods for determining copy number of a chromosome or chromosome segment with hypotheses tested using a combination of allelic and non-allelic data An embodiment includes a method for determining the copy number of a chromosome or chromosomal segment of interest in the genome of a target cell, where genetic data is obtained from DNA derived from the target cell and from DNA derived from a non-target cell, the genetic data including (i) quantitative allelic data from a plurality of polymorphic loci, and (ii) quantitative non-allelic data from a plurality of polymorphic and / or non-polymorphic loci. The method includes creating a plurality of hypotheses, where each hypothesis is associated with a particular copy number of a chromosome or chromosomal segment in the genome of the target cell. A probability value is calculated for each hypothesis, where the probability value indicates the likelihood that the genome of the target cell has a copy number of the chromosome or chromosomal segment associated with the hypothesis, where a first probability value is derived from the allelic data and the non-allelic data obtained from at least one first locus. For example, the hypothesis may be tested using a model incorporating both the allelic data and the non-allelic data, thereby obtaining a probability value. Each calculated probability value may be combined to provide a composite probability indicating the likelihood that the genome of the target cell has a copy number of the chromosome or chromosomal segment associated with the hypothesis. The copy number of the chromosome or chromosome segment of interest in the genome of the target cell is determined by selecting the copy number of the chromosome or chromosome segment associated with the hypothesis with the greatest probability. In some embodiments where the genetic data is obtained from cell-free DNA obtained from the blood of a pregnant woman, the hypothesis can include a condition where the mother is carrying multiple fetuses, e.g., twins.
[0168] In some embodiments, a probability value for each hypothesis is obtained from allelic and non-allelic data obtained from a single locus. In some embodiments, the allelic data is tested on a model based on a distribution of possible allelic ratios associated with each hypothesis. In some embodiments, a probability value for each hypothesis is determined separately for genetic data from at least 1000 polymorphic loci. In some embodiments, calculating the probability value for each hypothesis includes (1) modeling predicted genetic data from DNA from the target cells for each hypothesis based on the obtained genetic data including DNA from non-target cells, (2) comparing the modeled genetic data from DNA from the target cells and the genetic data obtained from DNA from the target cells for each hypothesis, and (3) calculating a probability value for each hypothesis based on the difference between the modeled genetic data from DNA from the target cells and the genetic data obtained from DNA from the target cells. In some embodiments, the non-target cells are from the parents of the individual from whom the target cells are derived, and modeling the predicted genetic data further comprises using the rules of Mendelian inheritance to determine the predicted genetic data of the target cells as disclosed herein, and adjusting the predicted genetic data of the target cells to correct for biases in the system. Examples of such system biases include amplification bias, sequencing bias, processing bias, enrichment bias, and combinations thereof. The nature of such biases may vary according to the specific amplification technique, sequencing technique, processing, enrichment technique, etc. selected for the implementation of a particular embodiment. In some embodiments, the target cells are from a fetus, and the predicted genetic data includes genetic data from the parents of the fetus and genetic data from the fetus. In some embodiments, modeling the genetic data includes predicting, for each locus, the predicted distribution of allele measurements at that locus, and predicting, for each locus, the predicted relative amount of DNA at that locus (read depth). In some embodiments, predicting the predicted distribution of allele measurements can take into account linkage and crossover between different loci on the genome.In some embodiments, the predictive distribution is a binomial distribution.
[0169] Quantitative Non-allelic Maximum Likelihood ("QMM") Example Provided herein is an example of a quantitative method that can be used to determine the copy number of a chromosome of interest in a target individual. Note that this example involves normalization of target chromosome data using a reference chromosome found in other samples that are the same as the target chromosome (i.e., the chromosome of interest) but processed in a similar or identical manner. The method is described in the context of non-invasive prenatal aneuploidy testing, where the target individual is a fetus, and the DNA to be sequenced includes fetal DNA and possibly maternal DNA, for example, found in maternal plasma. Non-invasive prenatal aneuploidy testing seeks to determine fetal chromosome copy number based on free floating fetal DNA in maternal plasma. In quantitative methods, chromosome copy number classification is based on the number of sequence reads that map to each chromosome. Neither parental genotype nor allele information is used except to estimate the fetal fraction in the plasma. In this targeted sequencing approach, the number of sequence reads at each target SNP (single nucleotide polymorphism) is beneficial, in contrast to non-targeted sequencing approaches that tend to use sliding window average read depth, or similar average approaches. Based on the estimated fetal fraction, a maximum likelihood estimate is calculated based on a set of copy number hypotheses including monosomy, disomy, and trisomy. In this example, chromosomal segmental errors are not considered, and all positions on the same chromosome are assumed to have the same copy number. It should be clear to one skilled in the art how to apply this method to chromosomal segment copy number variants. Also, non-uniform fragmentation of the fetal or maternal genome may be incorporated, which is not done here.
[0170] Modeling of individual SNPs: The basic assumption of this method is that the number of sequence reads generated at a genomic location depends primarily on the number of genome copies at that location that enter the sequencing process. Targeted sequencing approaches are based on multiplex PCR, which means that the number of genome copies that enter sequencing is determined by both the number of chromosome copies in the original sample and the details of the PCR amplification process. Therefore, this method requires simplified models of both multiplex PCR and high-throughput sequencing.
[0171] In the original sample, one can assume that the amount of genome copies is the same at all positions, except for those due to chromosomal copy number variation. However, in the PCR process, each target position is amplified with a different efficiency. For each of the K PCR cycles, position i is amplified by a factor a i The number of observed reads at that position is x i This model can be written as Equation 1, where the sample factor c s is constant for each sample and represents sample parameters such as the initial amount of DNA and the total number of sequence reads. It can be considered as a sample-specific amplification factor. The chromosome copy number n i is the ploidy state or copy number of the chromosome in which position i is located.
number
[0172] However, small variations in experimental conditions mean that the amplification efficiency of the various PCR targets is not completely constant. This means that for the amplification efficiency of each target, the multiplicative noise term ε i Therefore, the model is expanded to Equation 2. x i =c s n i (a i ∈ i ) k (2)
[0173] Due to the multiplicative nature of the model, we work in log space and then logx i It is advantageous to consider the expectation and variance of . We can assume that the expectation of the log noise is zero. This is not exactly the same as assuming zero-mean noise, but it makes the math workable, as shown in Equation 3. E log x i =log n i +k log a i V log x i =k 2 V log ∈ i (3)
[0174] Sample normalization can be achieved by considering reads measured from positions located on chromosomes that are known, hypothesized, or assumed to have a copy number equal to 2. Other sample normalization methods exist, such as using other reference chromosomes, e.g., chromosomes 1 and 2, which are known to be disomic. Let D be the set of positions i located on chromosomes that are hypothesized to be disomic. The sample normalizer T s is defined as the average logarithmic count over location i in D, as detailed in Equation 4. This can be directly measured from each sample and is therefore considered a known quantity for further calculations. T s =E i ∈ D log x i =log c s +log2+kE i ∈ D Log A i (4)
[0175] Building a model from training data: A model for the efficiency of individual SNPs can be built from a set of training data with known chromosome copy numbers and fetal fractions. In the ideal case, plasma is collected from non-pregnant (euploid) women, so the fetal fraction is zero and there are no aneuploids. In this case, all samples contribute data to the model for all targets. In a more challenging case, pregnancy plasma with known chromosome copy numbers is used, and aneuploid samples are excluded from the data set. Thus, the model is still built from data where all chromosomes have the same copy number compared to disomy.
[0176] y i Let be the log-space normalized read depth at position i. One is y i (5) β as the average over the set of samples i The term β may be defined. i is the logarithmic spatial amplification model of position i that measures how its amplification efficiency compares to the average amplification efficiency of positions on the disomy. y i =log x i -T s = k log a i +k log ∈ i -k E i ∈ D Log A i β i =E s y i = k log a i -k E i ∈ D Log A i (5)
[0177] Similarly, σ i is y i is defined as the standard deviation over the sample of i and σ i , and form the amplification and dispersion models for set i of SNPs.
[0178] There are several subtleties to the calculation of the model, the most important being to note that the model does not remain constant for a fixed set of targets subjected to a fixed protocol.
[0179] Although the models are very similar, attempts to use a fixed model across multiple sequencing runs suffer from biases large enough to produce results with low fetal fractions that may be eliminated by training separately for separate experiments. As a result, in some embodiments, it is important to ensure that each sequencing run contains a sufficient number of samples for modeling.
[0180] Even within an experiment, there are typically some samples that do not fit the model. These are often, but not always, explained by locus dropout, which is discussed in more detail in a later section. Outlier samples are not well predicted by quality control metrics such as contamination level, spike ratio (a measure of starting amount of DNA), fetal fraction, or overall read depth. Samples are tested for fit by calculating the residuals z on each SNP with respect to the amplification and noise models. z i =(log x i -T s -β i ) / σ i (6)
[0181] Further log
number
[0182] Forming the test statistic and modeling SNP correlation: The test statistic for the chromosome copy number classification can be formed by averaging the normalized measurements at all positions on the chromosome. A variance-weighted average is chosen to minimize the variance of the test statistic. The normalized measurement y defined above i Consider the unknown copy number n i For a chromosomal location with y i has the properties described in Equation 7.
number
[0183] Let S be the set of positions on the current chromosome. A chromosome test statistic t is the set of y averaged over SNP i in S. i It is defined as the variance-weighted average of
number
[0184] Predicted values of t are calculated under each chromosome copy number hypothesis to form a Gaussian model of maximum likelihood estimates. The variance of the model for each hypothesis does not follow uniquely from the previously made assumptions that do not account for correlation between measurements. The simplest assumption of uncorrelated measurements was discarded because the variance observed on t was much higher than the model would suggest. Without suggesting any physical explanation for the correlation, y i andy jThe covariance with ρσ i σ j A single-parameter correlation model is proposed, which corresponds to a constant correlation coefficient between all positions i and j on the same chromosome. This model uses a single parameter to represent additional variance beyond that implied by the uncorrelated model. The variance of t using the constant correlation model is shown in Equation 9, which follows directly from the equation for the sum variance of a normal distribution with known correlation. (The Gaussian noise assumption is continued throughout.)
number
[0185] The maximum likelihood estimate of ρ for each chromosome is {β i} and {σ i} are calculated from the same modeling data following the estimation of
[0186] Chromosome copy number classification consists of the following steps utilizing the modeling developed in the above sections.
[0187] 1. Check the fit of the model. Calculate a set of residuals for the model provided using disomic chromosomes (1 and 2) and compare them to a standard normal distribution using the KS test. If the p-value obtained is too low, the sample is considered not to fit the model and cannot be classified.
[0188] 2. Copy number hypothesis generation. Using the provided fetal fraction, plasma copy numbers are calculated corresponding to each fetal copy number hypothesis. Fetal copy number hypothesis {h 1 ,h 2 ,h 3 For}={1,2,3}, the plasma copy number hypothesis is calculated using the fetal fraction according to Equation 10. The plasma copy number is a mixture of the fetal copy number, which depends on the hypothesis, and the maternal copy number, which is 2. n i =fh i +2(1-f)(10)
[0189] 3. Hypothesis modeling. The predicted value of the test statistic is the n corresponding to the ploidy hypothesis. i This is done according to Equation 7 and the definition of the test statistic. The variance model of the test statistic is hypothesis independent.
[0190] 4. Calculate the likelihood. The value of the test statistic is observed for the current chromosome. The data likelihood of each hypothesis is the likelihood of the test statistic under each of the corresponding normal distributions. The maximum likelihood estimate can then be reported or normalized using the prior distribution.
[0191] Copy number classification without non-targeted reference chromosomes (also called "QMM" method) As mentioned above, without using a reference chromosome or chromosome segment that is different from the target chromosome or chromosome segment, it is possible to assume that none of the chromosomes or chromosome segments have a known copy number. This is the sample normalization T, which is conditioned on the chromosome number hypothesis. s and the linear shift parameter α s We need an alternative way to estimate . Unlike approaches that use copy number hypotheses for each individual chromosome, this hypothesis space includes combination hypotheses for all training chromosomes.
[0192] In one embodiment, the following technique may be used to link the joint hypotheses to the individual hypotheses: For training chromosomes k ∈ {13, 18, 21}, p(D|h k ), h k Let ∈{1,2,3} be the pdf of the data conditioned on the individual copy number hypotheses for that chromosome. So, for example, for chromosome 13 we have:
number
[0193] The hypothesis probability, i.e., P(h k =1)=P(h k =2)=P(h kAssuming equal priors for P(D =3)=1 / 3, the above pdf is calculated. 13 |h 18 , h 21 , h 13 To calculate the hypothesis (h 18 ,h 21 ,h 13 ) corresponding to T s and α s Using the estimated values, the variance-weighted average test statistic is calculated. Similarly, for the other training chromosomes, p(D|h 18 ), p(D|h 21 ) is calculated. Since equal prior probabilities are assumed, the posterior probabilities are also calculated.
number
[0194] This represents a normalization step that provides confidence for each of the training chromosomes.
[0195] Next, the confidences of the remaining chromosomes are calculated, which gives estimates of the binding hypotheses of the training chromosomes.
number
[0196] Then, T corresponding to this hypothesis s and α s The estimate of can be used to calculate the variance-weighted average test statistic for each test chromosome.
number
[0197] In this method, a constant correlation coefficient model can be used to model the correlation between SNPs for a particular chromosome. For example, for a particular chromosome k, y i and y j The covariance of is, as mentioned above, ρ i σ i σ j Chromosome K is Nk With loci, the covariance matrix is given by:
number
[0198] This is the main diagonal
number
[0199] Example of Quantitative Allele Maximum Likelihood ("HET Rate") A method for determining ploidy status using allele maximum likelihood method is provided herein.This method is illustrated in the context of NIPT, but those skilled in the art will understand that it can be utilized for detecting circulating free tumor cells.In addition to the following description, detailed examples of how to implement the HET ratio method can be found in, among others, published US Patent Application No. 2012 / 0270212A1 and published US Patent Application No. 2011 / 0288780A1, all of which are incorporated herein by reference in their entirety.However, the HET ratio methods disclosed in these sources utilize data from separate reference chromosomes.
[0200] In the example of NIPT, the ploidy state of the fetus is given sequence data measured on free floating DNA isolated from maternal blood, the free floating DNA containing some DNA of maternal origin and some DNA of fetal / placental origin. In this example, the ploidy state of the fetus is determined using an allele maximum likelihood method and a calculated fraction of fetal DNA in the analyzed mixture. Also described are embodiments in which the fraction of fetal DNA or the percentage of fetal DNA in the mixture can be measured. In some embodiments, the fraction can be calculated using only genotyping measurements made on the maternal blood sample itself, which is a mixture of fetal and maternal DNA. In some embodiments, the fraction can also be calculated using the measured or otherwise known genotype of the mother and / or the measured or otherwise known genotype of the father.
[0201] For a particular chromosome, assume there are N SNPs, where: The parent genotypes from the ILLUMINA data are assumed to be correct: mother m = (m 1 ,...,m N ), father = (f 1 ,...,f N ), in the formula, m i , f i ∈(AA,AB,BB).
[0202] A set of NR array measurements S = (s 1 ,...,s nr ).
[0203] Derive the most likely copy number from the data For each copy number hypothesis H considered, derive the chromosome-wide data log likelihood LIK(H) and select the best hypothesis that maximizes LIK, i.e.,
number
[0204] The copy number hypotheses considered are as follows: Monosomy: ● Maternal H10 (one copy from the mother) Paternal H01 (one copy from the father) Disomy: H11 (one copy each from mother and father) Simple trisomy, no crossover consideration: Maternal lineage: H21_MATCHED (two identical copies from the mother, one copy from the father), H21_UNMATCHED (both copies from the mother, one copy from the father) Paternal lineage: H12_MATCHED (one copy from mother, two copies from father), H12_UNMATCHED (one copy from mother, both copies from father) Complex trisomies allowing crossover (using a joint distribution model): Maternal H21 (two copies from the mother, one copy from the father), Paternal H12 (one copy from the mother, two copies from the father)
[0205] In the absence of crossovers, each trisomy would be one of the matching or non-matching trisomy, regardless of whether the origin is mitosis, meiosis I, or meiosis II. Due to crossovers, the true trisomy is a combination of the two. First, a method is described for deriving the hypothesis likelihood for a simple hypothesis. Then, a method is described for combining the individual SNP likelihoods with crossovers to derive the hypothesis likelihood for a composite hypothesis. Initially, it is assumed that the true child fraction and other parameters such as the beta noise parameter (N) and possible error rate are known. A method for deriving the child fraction cf from the data is also described below.
[0206] Simple hypothesis LIK(D|H) For a simple hypothesis H, LIK(D|H), the log-likelihood of the data given hypothesis H across chromosomes is calculated as the sum of the log-likelihoods of the individual SNPs, i.e.:
number
[0207] This hypothesis does not assume linkage between SNPs and therefore does not utilize a joint distribution model.
[0208] Log likelihood per SNP On a particular SNP i, m i = true maternal genotype, f i = true paternal genotype, and cf = known or derived offspring fraction. i Let P(A|i,S) be the probability of having A on SNP i given sequence measurements S. Given a child hypothesis H, the log-likelihood of observed data D on SNP i is defined as P(D│m,f,c,H,cf,i)=P(SM|m,i)P(M│m,i)P(SF|f,i)P(F│f,i)P(S│m,c,H,cf,i), This results in the following: LIK(i,H) = log lik(x i │m i ,f i ,H,cf)=Σ c p(c│m i ,f i ,H)*loglik(x i |m i ,c,cf), ● where p(c|m,f,H) is the probability of obtaining the true child genotype=c given parents m,f and hypothesis H, which can be easily calculated. For example, p(c|m,f,H) for H11, H21 match and H21 no match is shown below: [Table 1]
[0209] P(D|m,f,c,H,i,cf) is the probability of the given data D for SNPi given the true maternal genotype m, the true paternal genotype f, the true child genotype c, the hypothesis H, and the child fraction cf. It can be separated into the probabilities of the maternal, paternal, and child data as follows: P(D│m,f,c,H,cf,i)=P(SM|m,i)P(M│m,i)P(SF|f,i)P(F│f,i)P(S│m,c,H,cf,i).
[0210] lik(xi|m,c,cf) is the probability that x i The pdf of the distribution that should follow x(x i ) Assuming a true mother m and a true child c, the probability of derivation x i Specifically, LIK(x i |m,c,cf)=pdfx(x i ).
[0211] In the simple case where Di of NR sequences in S is aligned up to SNPi, X~(1 / D i )Bin(p,D i ), where the probability of obtaining p=p(A|m,c,cf)=A is calculated for this mother-child mixture as follows:
number
[0212] c.f. correct is the corrected fraction of the child in the mixture:
number
[0213] A child with disomy cfcorrect =cf, however, the child's trisomy fraction in this chromosome mixture is actually a little higher:
number
[0214] In the more complicated case where exact alignment does not exist, X is the possible D i is a combination of binomial functions integrated over readings.
[0215] Using joint distribution models: LIK(H) for multiple hypotheses Since trisomies are usually not purely matched and unmatched due to crossover, in this section the results of the composite hypotheses H21 (maternal trisomy) and H12 (paternal trisomy) are derived to explain the possibility of crossover by combining matched and unmatched trisomies.
[0216] In the case of trisomy, if there was no crossover, the trisomy would simply be matched or unmatched trisomy. Matched trisomy is when a child inherits two copies of the same chromosomal segment from one parent. Unmatched trisomy is when a child inherits one copy of each homologous chromosomal segment from a parent. Due to crossover, some segments of a chromosome may have matching trisomy and other parts may have unmatched trisomy. In this section, we describe how to build a joint distribution model of the heterozygosity rate of a set of alleles.
[0217] Assume that on SNPi, LIK(i,Hm) is the fitness of the matching hypothesis H, LIK(i,Hu) is the fitness of the mismatching hypothesis H, and pc(i) = the probability of crossover between SNPi-1 and i. Then, we can calculate the total likelihood as follows: LIK(H)=Σ S、E LIK(S,E,1:N) where LIK(S, E, 1:N) is the likelihood starting from hypothesis S and ending with hypothesis E for SNP 1:N, where S=hypothesis of the first SNP and E=hypothesis of the last SNP, S, E∈(Hm,Hu). Recursively, the following can be calculated:
number
[0218] Next,
number
[0219] Deriving child fractions The above formula assumes a known offspring fraction, but this is not necessarily the case. In one embodiment, it is possible to find the most likely offspring fraction by maximizing the likelihood of disomy on a selected chromosome.
[0220] Specifically, we assume LIK(chr,H11,cf) = log-likelihood as above for the disomy hypothesis and for the offspring fraction of chromosome chr for the selected chromosomes in Cset (usually 1:16). Then the full likelihood is: ●Lik(cf)=Σ chr∈Cset LIk(chr,H11,cf), and
number
[0221] Any set of chromosomes can be used. It is also possible to derive offspring fractions without paternity data as follows:
[0222] Deriving copy number without paternity data Recall the formula for the simple hypothesis log-likelihood on SNPi.
number
[0223] Parent p(c│m i ,f i Determining the probability of a true offspring given a SNP (H,H) requires knowledge of the paternal genotype. If the paternal genotype is unknown but pAi, the population frequency of the A allele at this SNP, is known, it is possible to approximate the likelihood above by LIK(i,H) = log lik(x i │m i ,f i ,H,cf)=Σ c p(c│m i ,H)*loglik(x i │m i ,c,H,cf) During the ceremony,
number
[0224] in particular, p(AA│pA i )=(pA i ) 2 , p(AB│pA i )=2(pA i )*(1-pA i ), p(BB│pA i )=(1-pA i ) 2
[0225] Training methods without using control chromosomes or chromosome segments 3 data segments D 1 , D 2 , D 3 Suppose P(H) has a segment D 1 Assume the current priors above are: Let P be a parameter (e.g., the child fraction cf or the noise parameter np) with distribution P(p). Then the probability that a particular hypothesis H (with the prior P(H)) is true is:
number
number
number
number
[0226] therefore,
number
[0227] Significant processing advantages can be obtained when testing can be performed on only the chromosome(s) or chromosome segment(s) of interest, without the need for a control chromosome or chromosome segment. In one embodiment, the chromosome or chromosome segment of interest itself provides a baseline that can then be used to assess the accuracy of a given hypothesis. For example, using the formula:
number
number
[0228] In this formula, the probability P(H|D 1 ,p) is obtained for each grid point, and then P(p,D 1 ,D 2 ,D 3 ) is scaled by the best parameter distribution estimate given. Once the grid points are fixed, P(H|D 1 ,p) does not change. However, P(p,D 1 ,D 2 ,D 3 If there is no fixed hypothesis for P(H|D) (i.e., no control chromosome or chromosome segment is used), then P(H|D 1 ,D 2 ,D 3 ) can vary widely depending on what precedes each segment hypothesis.
[0229] In other words, the parameter distribution given all the data is a composite of the parameter distributions for each segment, so
number
[0230] To account for the lack of control, we use a uniform prior f prior (H) is obtained. For example, this may be obtained by estimating the child fraction using the allele ratio plot described above. Then, for each grid point p, the probability of the hypothesis is calculated ("grid-wise call"). P(H│D 1 ,p)~P(D 1 │H,p)P(H) where P(H) is the hypothesis used prior to the segment call. In one embodiment, this is done only once to provide an idea of the calls in the entire grid space.
[0231] For the first pass, f prior (H) is set to be P(H). The parameter distribution for each segment is then obtained using:
number
[0232] A composite parameter distribution is then obtained:
number
[0233] The (posterior) probability of each hypothesis is then obtained by combining the parameter scaling for each grid call:
number
[0234] This provides a new estimate of the distribution of hypotheses for each segment. prior (H) is the newly derived P(H│D 1 ,D 2 ,D 3 ) and the process (starting with computing the probability of the hypothesis for each grid point p) is repeated until convergence.
[0235] Once convergence is reached, the overall likelihood no longer changes to any appreciable extent. In one embodiment, this means that the function to be optimized is the data P(H│D) that is maximized by the best derived posterior P(H) and P(p) distributions. 1 ,D 2 ,D 3 ), i.e. the function to maximize is: L(D)=P(D 1 ,D 2 ,D 3 )~Σ H Σ p P(D│H,p)P(H)P(p).
[0236] The hypothesis with the final probability (i.e., call), child fraction, and noise parameters can then be output.
[0237] In certain embodiments of the present disclosure, the method of the present invention for determining aneuploidy can include a quantitative allele method, technique, or algorithm that can be used to determine the relative ratio of two or more different haplotypes that contain the same set of loci in a sample of DNA. The different haplotypes can represent two different homologous chromosomes from one individual, three different homologous chromosomes from a trisomic individual, three different homologous haplotypes from a mother and a fetus where one haplotype is shared between the mother and the fetus, three or four haplotypes from a mother and a fetus where one or two haplotypes are shared between the mother and the fetus, or other combinations. If one or more haplotypes are known, or if the diploid genotypes of one or more individuals are known, a set of alleles that are polymorphic between the haplotypes can be selected, and the average allele ratio can be determined based on the set of alleles that are uniquely derived from each of the haplotypes.
[0238] However, direct sequencing of such samples is extremely inefficient, as it results in many sequences for regions that are not polymorphic between different haplotypes in the sample, and therefore does not reveal information about the ratio of the two haplotypes. In order to increase the yield of allele information obtained by sequencing, methods are described herein that specifically target and enrich segments of DNA in a sample that are more likely to be polymorphic in the genome. It should be noted that for the allele ratios measured in the enriched sample to truly represent the actual haplotype ratios, it is important that there is little or no preferential enrichment of one allele compared to other alleles at a given locus in the target segment. Current methods known in the art that target polymorphic alleles are designed to ensure that at least some of any alleles present are detected. However, these methods are not designed for the purpose of measuring the allele ratios of polymorphic alleles present in the original mixture. It is not clear that any particular targeted enrichment method can produce an enriched sample in which the ratios of the various alleles in the enriched sample are approximately the same as the ratios of the alleles in the original non-amplified sample. Enrichment methods can theoretically be designed to achieve such a goal, but those skilled in the art recognize that current methods have numerous stochastic or deterministic biases. In an embodiment of the method described herein, multiple alleles found in a mixture of DNA corresponding to a given locus in a genome are amplified or preferentially enriched in such a way that the enrichment of each of the alleles is approximately the same. In other words, this method allows the relative amount of alleles present in the mixture as a whole to be increased, but the ratio between the alleles corresponding to each locus remains essentially the same as in the original mixture of DNA.For purposes of this disclosure, this means that the ratio of alleles in the original mixture divided by the ratio of alleles in the resulting mixture is between 0.5 and 1.5, 0.8 and 1.2, 0.9 and 1.1, 0.95 and 1.05, 0.98 and 1.02, 0.99 and 1.01, 0.995 and 1.005, 0.998 and 1.002, 0.999 and 1.001, or 0.9999 and 1.0001, in order for the ratio to remain essentially the same.
[0239] Allele distribution In certain embodiments, the objective of the method is to detect fetal copy number based on a maternal blood sample containing some free floating fetal DNA. In some embodiments, the fraction of fetal DNA compared to maternal DNA is unknown. The combination of a targeted method such as LIPS, followed by sequencing, results in a platform response consisting of a count of observed sequences associated with each allele at each SNP. A set of possible alleles, either A / T or C / G, is known at each SNP. Without loss of generality, the first allele is labeled A and the second allele is labeled B. Thus, the measurement at each SNP is the number of A sequences (N A ) and the number of B sequences (N B ) which are converted to the total sequence count (n) and the ratio of A alleles to the total (r) for future calculation purposes. The sequence count for a single SNP is called the read depth. The basic principle that allows copy number identification from this data is that the ratio of A and B sequences reflects the ratio of A and B alleles present in the DNA being measured. n=N A +N B r=N A / (N A +N B )
[0240] Measurements are first aggregated across SNPs from the same parental context based on unordered parent genotypes. Each context is defined by a maternal genotype and a paternal genotype, for a total of nine contexts. For example, all SNPs with maternal genotype AA and paternal genotype BB are members of the AA|BB context. The A allele is expressed as the proportion of maternal genotype r m and the paternal genotype ratio r f For example, the A allele is defined as present at a ratio r m = 1, and the proportion of fathers who are AB, r f = 0.5. Therefore, each context exists at r m and r f The genotype of the offspring is not necessarily predicted from the parent genotypes, but the allele ratio averaged over a number of SNPs can be predicted based on the assumption that the parent AB genotypes contribute equal proportions to A and B.
[0241] where n m is the number of maternal copies, and n r is the number of paternal copies of the chromosome (n m ,n f Consider a child copy number hypothesis of the form c The predicted value of (averaged over SNPs in a particular parental context) depends on the allele ratios of the parental context and the parental copy numbers.
number
[0242] In a mixture of maternal and fetal blood, copies of the allele are contributed both directly from the mother and from the child. Suppose the fraction of the child's DNA present in the mixture is δ. Then, in the mixture, the proportion r of A alleles in a given context is equal to the maternal proportion r. m and child ratio r c which can be reduced to a linear combination of the maternal and paternal proportions using Equation 1.
number
[0243] Equation 2 is the copy number hypothesis (n m ,n f ) predicts the expected proportion of A alleles for a SNP in a given context as a function of . Note that allele proportions on individual SNPs are not predicted by this formula as these depend on random assignment where at least one parent is heterozygous. Thus, the set of sequences from all SNPs in a particular context is combined. Suppose a context contains m SNPs, and recall that n sequences are generated from each SNP, then the data from that context consists of N=mn sequences. Each of the N sequences is considered as an independent random test where the theoretical rate of the A sequences is the allele proportion r. Thus, the measured rate of the A sequences is
number
[0244] The theoretical allelic ratio is calculated by the parental copy number (n m ,n f ), so each hypothesis h is a function of the SNP
number
number
number
[0245] The measurements in each of the nine contexts are assumed to be independent, taking into account the parent copy number, due to the general assumption of independent noise on each SNP. Thus, data from a particular chromosome consists of sequence measurements from context i, which ranges from 1 to 9. Thus, the entire chromosome
number
number
[0246] Parameter Estimation Equation 2 predicts allele ratios as a function of parental copy number hypotheses, but also includes the fraction of the child's DNA. Thus, the data likelihood for each chromosome is
number
number
[0247] Measure some chromosomes known to be disomic In this method, specific chromosomes that do not have copy number errors in the developmental state when the test is performed are measured. These chromosomes are called the training set T. The copy number hypothesis on these chromosomes is (1,1). Assuming that each chromosome is independent, the data likelihood of measurements from all chromosomes t in T is the product of the individual chromosome likelihoods. The child fraction δ can be chosen to maximize the data likelihood across the chromosomes of T conditioned on the disomy hypothesis. R t is the set of all contexts on chromosome t.
number
number
[0248] This optimization has only one degree of freedom, constrained between zero and one, and can therefore be easily solved using a variety of numerical methods. Then, to calculate the likelihood of each hypothesis on each chromosome, we use the solution δ * can be substituted into Eq. 2.
[0249] Measure only chromosomes that may have copy number errors If copy number errors are possible on all measured chromosomes, estimating fetal fractions in parallel with the copy number hypothesis will greatly increase the accuracy of the ploidy determination. Note that the same copy number errors present on all measured chromosomes would be very difficult to detect. For example, in each case, the contribution of maternal alleles compared to paternal alleles increases uniformly across all chromosomes and contexts, so maternal trisomy on all chromosomes at a given pediatric concentration will result in the same theoretical allele ratio as disomy on all chromosomes at a lower pediatric concentration.
[0250] A simple approach for classification of a limited set of chromosomes t is to consider a combined chromosome hypothesis H consisting of a combined set of hypotheses for all chromosomes to be tested. If the chromosome hypotheses consist of disomies, maternal trisomies, and paternal trisomies, the number of possible combined hypotheses is 3 T where T is the number of chromosomes examined. * (H) can be calculated conditionally on each joint hypothesis. Thus, the likelihood of a joint hypothesis is calculated as follows:
number
[0251] The joint hypothesis likelihood p(all data|H) can be calculated for each joint hypothesis H, and the maximum likelihood hypothesis is its corresponding estimate of the child fraction δ *( H) is selected.
[0252] Performance Specifications The ability to distinguish between parental copy number hypotheses is determined by the model described in the previous section. At the most general level, the difference in the predicted allele ratios under the different hypotheses must be large compared to the standard deviation of the measurements. Disomy and maternal trisomy, or hypothesis h 1 =(1,1) and h 2 Consider the example of distinguishing between the maternal allele ratio r = (2,1). Hypothesis 1 states that the maternal allele ratio r m and paternal allele ratio r f As a function of the allelic ratio r 1 and hypothesis 2 is the allele ratio r 2 Predict.
number
[0253] Allele ratio measurements
number
number
[0254] Substituting copy numbers for disomy (1,1) and maternal trisomy (2,1) in hypotheses 1 and 2, we obtain the following:
number
[0255] Overview of analytical methods used in the methods provided herein In a particular example of an embodiment of the present disclosure, using parental contexts and chromosomes known to be euploid, it is possible to estimate the percentage of DNA in maternal blood that is from the mother and the percentage of DNA in maternal blood that is from the fetus by a set of simultaneous equations. These simultaneous equations are made possible by knowledge of the alleles present in the father. In particular, alleles present in the father and absent in the mother provide a direct measure of fetal DNA. Then, looking at a particular chromosome of interest, such as chromosome 21, the measurements on this chromosome in the context of each parent are expressed as H mp where m represents the number of maternal chromosomes and p represents the number of paternal chromosomes, e.g., H 11 represents euploid and H21 and H 12 represent maternal and paternal trisomies, respectively.
[0256] This method differs from certain other methods for detecting chromosomal polyploidy in that it does not use a reference chromosome as a basis for comparing the allele ratios observed on the chromosome of interest to make the aneuploidy determination.
[0257] The present disclosure presents a method that may determine the ploidy state of a pregnant fetus in a non-invasive manner at one or more chromosomes using genetic information determined from fetal DNA found in maternal blood. The fetal DNA may be purified, partially purified, or unpurified, and genetic measurements may be performed on DNA from two or more individuals. Informatics-type methods may infer genetic information of a target individual, such as ploidy state, from bulk genotype measurements on a set of alleles. The set of alleles may include various subsets of alleles, one or more subsets may correspond to alleles found on the target individual but not on non-target individuals, and one or more other subsets may correspond to alleles found on non-target individuals and not on the target individual. The method may involve comparing the ratio of measured output intensities for various subsets of alleles given various potential ploidy states to the expected ratio. Platform responses may be determined, and corrections for system biases may be incorporated into the method.
[0258] The main assumptions of the method are: -The expected amount of genetic material in maternal blood from the mother is constant across all loci. -The expected amount of genetic material present in maternal blood from the fetus is constant across all loci, assuming the chromosomes are euploid. - The non-viable chromosomes (all except 13, 18, 21, X, Y) are all euploid in the fetus. In one embodiment, only some of the non-viable chromosomes need to be euploid on the fetus.
[0259] Common problem formulations: y ijk =g ijk (x ijk )+v ijk and X ijk is the amount of DNA on allele k=1 or 2 (1 represents allele A and 2 represents allele B), j=1...23 represents the chromosome number, i=1...N represents the locus number on the chromosome, gijk is the platform response for a particular locus and allele ijk, and v ijk is the independent noise in the measurement of that locus and allele. The amount of genetic material is x ijk =am ijk +Δc ijk where a is the amplification factor (or the net effect of leakage, diffusion, amplification, etc.) of genetic material present on each of the maternal chromosomes, and m ijk (0, 1, or 2) is the number of copies of a particular allele on the maternal chromosome, Δ is the amplification factor of genetic material present on each of the offspring chromosomes, and c ijk is the number of copies of a particular allele on the offspring chromosome (either 0, 1, 2, or 3). Note that for the first simplified description, a and Δ are assumed to be independent of the locus and allele, i.e., independent of i, j, and k. This gives: y ijk =g ijk (am ijk +Δc ijk )+v ijk
[0260] Approaches using an affine model that is uniform across all loci: One can model using an affine model; for simplicity, we assume that the model is the same for each locus and allele, but it will be understood after reading this disclosure how to modify the approach when the affine model depends on i, j, k. g ijk (x ijk )=b+amijk +Δc ijk Assuming that, The amplification factors a and Δ are used without loss of generality, with the addition of a y-axis intercept b, which defines the noise level in the absence of genetic material. The goal is to estimate a and Δ. Although it is possible to estimate b independently, we assume that the noise level is roughly constant across loci and only use a set of equations based on the parental context to estimate a and Δ. The measurements at each locus are y ijk =b+am ijk +Δc ijk +v ijk is given by:
[0261] Noise v ijk Assuming that A,B,C,D,E,i,i,d,i,i,i,j ...
number
[0262] During the ceremony
number
number
number
[0263] Now, we can cast all the predicted values y into a matrix as follows: T、k We can write a set of equations that describe: Y=B+A H P+v
[0264] During the ceremony,
number
number
number
number
number
number
[0265] To estimate A and Δ, or the matrix P, we aggregate the data over the set of chromosomes that may be assumed to be euploid on the offspring sample. This may include the chromosome under test, i.e., all chromosomes j=1...23 except j=13, 18, 21, X, and Y. (Note: Concordance tests can also be applied to the results on individual chromosomes to detect mosaic aneuploids on non-surviving chromosomes.) For clarity of notation, we define Y' as Y measured over all euploid sex chromosomes and Y'' as Y measured over a particular chromosome under test that may be aneuploid, e.g., chromosome 21. To estimate the parameters, we use the matrix
number
number
number
number
[0266] One simple approach to estimating the covariance matrix is to assume that all terms in v are independent (i.e., there are no off-diagonal terms) and that the variance of each term in v is 1 / N' T The trick is to be able to call the central limit theorem and find an 18x18 matrix so that it scales as
number
[0267] Once P' has been estimated, these parameters are used to determine the most likely hypothesis for the chromosome under study, such as chromosome 21. In other words, select the hypothesis:
number
[0268] Next, H * After finding * It is possible to estimate the confidence that one can have in determining H. For example, suppose there are two hypotheses under consideration: 11 (euploid) and H 21 (Maternal trisomy). H * =H 11 Assume that . Calculate the distance measure corresponding to each hypothesis.
number
[0269] The squares of these distance measures can be shown to be roughly distributed as chi-squared random variables with 18 degrees of freedom. Let Χ18 represent the corresponding probability density function of such a variable. We can then find the ratio of probabilities pH of each hypothesis according to:
number
[0270] Then, the equation
number
number
[0271] Variations of this method (1) For different biases b on each channel representing alleles A and B, the above approach may be modified. The bias matrix B is redefined as follows:
number
number
[0272] (2) In the general formulation, y ijk =g ijk (am ijk +Δc ijk )+v ijkFor all loci and alleles, the function g ijk can be directly measured or calibrated, so the function (which can be assumed to be monotonic in most genotyping platforms) can be inverted. Then, by inverting the function, the measurements regarding the amount of genetic material can be recast such that the system of equations is linear. That is, y’ ijk = g ijk -1(y ijk ) = am ijk + Δc ijk + v’ ijk . This approach is particularly good when g ijk is an affine function, as the inversion does not generate amplification or bias of the noise of v’ ijk .
[0273] (3) The modified noise term v’ ijk = g ijk -1(v ijk ) may be amplified or biased by the function inversion, so the above method may not be optimal from the perspective of noise. Another approach is to linearize the measurements around the operating point. That is, y ijk = g ijk (am ijk + Δc ijk ) + v ijk is recast as follows: [Number] Since it can be predicted that 30% or less of the cell-free floating DNA in maternal blood is derived from the child, Δ << a, and the expansion is a reasonable approximation. Alternatively, for platform responses such as the ILLUMINA bead array that increases monotonically and the second derivative is always negative, [Number] the linearization estimation can be improved according to. The obtained set of equations can be iteratively solved for a and Δ using methods such as Newton-Raphson optimization.
[0274] (4) Another common approach is to measure the total amount of DNA on the test chromosome (mother + fetus) and compare it to the amount of DNA on all other chromosomes, based on the assumption that the amount of DNA should be constant across all chromosomes. This is simpler, but has the drawback that confidence limits cannot be meaningfully estimated because we know how much the child contributes. However, we can look at the standard deviation across the other chromosome signals, which should be euploid, to estimate the signal variance and generate confidence limits. Since this method involves measurements of maternal DNA that are not in the child's DNA, these measurements do not contribute anything to the signal, but contribute directly to the noise. In addition, it is not possible to calibrate the amplification bias between different chromosomes. To address this last point, it is possible to combine the signals from all chromosomes by finding a regression function that links the average signal level of each chromosome to the average signal level of all other chromosomes, weighting it based on the variance of the regression fit, and examine whether the test chromosome of interest is within the tolerance range defined by the other chromosomes.
[0275] Incorporating Data Dropout Elsewhere in this disclosure, it is assumed that the probability of getting A is a direct function of the true maternal genotype, the true child genotype, the fraction of the child in the mixture, and the child copy number. It is also possible that maternal or child alleles may drop out, e.g., instead of having true child AB in the mixture, there is only A, in which case the probability of getting a combined sequence measurement of A is much higher. Assume that the dropout rate of the mother is MDO and the dropout rate of the child is CDO. In some embodiments, the dropout rate of the mother can be assumed to be zero, and the dropout rate of the child is relatively low, so the actual results are not severely affected by dropout. Nevertheless, they are incorporated into the algorithms here. Elsewhere, lik(x i │m i ,c,cf)=pdf x (x i ) is the True Mother m iGiven a sequence measurement S, assuming a true child c, x of A on SNP i i It is defined as the likelihood of obtaining the probability of obtaining the true mother (m i ) or child (c), but not mother (m) after possible dropout. d ) and children after possible dropout (c d ) The above equation can then be rewritten as:
number
number
[0276] For one set of experimental data, parent genotypes were measured, as well as true child genotypes where the child has maternal trisomies on chromosomes 14 and 21. Sequencing measurements were simulated for various values of child fraction, N different SNPs, and total number of reads, NR. From this data, it is possible to derive the most likely child fraction and derive copy number given known or derived child fractions.
[0277] In one embodiment, the method disclosed herein can be used to determine fetal aneuploidy by determining the copy number of maternal and fetal target chromosomes having target sequences in a mixture of maternal and fetal genetic material. The method can involve obtaining maternal tissue containing both maternal and fetal genetic material. In some embodiments, the maternal tissue can be maternal plasma or tissue isolated from maternal blood. The method can also involve obtaining a mixture of maternal and fetal genetic material from the maternal tissue by processing the maternal tissue. The method can involve distributing the obtained genetic material into a plurality of reaction samples, randomly providing individual reaction samples that include target sequences from the target chromosomes and individual reaction samples that do not include target sequences from the target chromosomes, and performing, for example, high throughput sequencing on the samples. The method can involve analyzing target sequences of genetic material present or absent in the individual reaction samples to provide a first number of binary results representing the presence or absence of a likely euploid fetal chromosome in the reaction sample and a second number of binary results representing the presence or absence of a likely aneuploid fetal chromosome in the reaction sample. Any of the numbers of binary outcomes may be calculated, for example, by information technology that counts sequence reads that map to a particular chromosome, a particular region of a chromosome, a particular locus, or a set of loci. The method may involve normalizing the number of binary events based on the length of the chromosome, the length of the region of the chromosome, or the number of loci in the set. The method may involve using the first number to calculate a predictive distribution of the number of binary outcomes of likely euploid fetal chromosomes in the reaction sample. The method may involve using the first number and the predicted fraction of fetal DNA found in the mixture to calculate a predictive distribution of the number of binary outcomes of likely aneuploid fetal chromosomes in the reaction sample, for example, by multiplying the predicted read count distribution of the number of binary outcomes of likely aneuploid fetal chromosomes by (1+n / 2), where n is the predicted fetal fraction. The fetal fraction may be estimated by multiple methods, some of which are described elsewhere in this disclosure.The method may involve using a maximum likelihood approach to determine whether the second number corresponds to whether the potentially aneuploid fetal chromosomes are euploid or aneuploid. The method may involve referring to the ploidy state of the fetus as the ploidy state that corresponds to the hypothesis that has the greatest likelihood of being correct given the measurement data.
[0278] A simple explanation of the allele ratio method of ploidy calling in NPD In one embodiment, the ploidy state of a pregnant fetus can be determined using a method that examines allele ratios. Some methods determine the ploidy state of a fetus by comparing the numerical sequencing output DNA count from a suspect chromosome to a reference euploid chromosome. In contrast to that concept, the allele ratio method determines the ploidy state of a fetus by looking at the allele ratios of different parental contexts on one chromosome. This method does not require the use of a reference chromosome. For example, imagine the following possible ploidy states and the allele ratios of the various parental contexts: (Note: The ratio "r" is defined as: 1 / r = fraction maternal DNA / fraction fetal DNA). [Table 3] *PU tri=paternal matching trisomy, PM tri=paternal matching trisomy, Note that this table represents only a subset of the parent contexts, the subset of possible ploidies that the method is designed to distinguish. In this case, from the set of parent contexts in a set of sequencing data, the A:B ratios of multiple alleles can be determined. Then, for each ploidy state and for each r value, several hypotheses can be stated, each hypothesis having a predicted pattern of A:B ratios for different parent contexts. Then, it can be determined which hypothesis best fits the experimental data. For example, using the above set of parent contexts, and a value of r=0.2, we can rewrite the chart as follows: (For example, we can calculate [#reads of allele A / # reads of allele B]. Therefore, 2+r:r becomes 2+0.2:0.2 → 2.2:0.2=11.) [Table 4]
[0279] Here you can see the ratio between the A:B ratios of different parent contexts. In this case, A:B AA|BB / A:B AA|AB can be predicted to be 11 / 21=0.524 on average for euploids, 5.5 / 12=0.458 on average for paternally unmatched trisomies, and 5.5 / 44=0.125 on average for paternally matched trisomies. The profiles of A:B ratios between different contexts are different in different ploidy states, and the profiles should be sufficiently distinct so that it may be possible to determine the ploidy state of the chromosomes with high accuracy. Note that the calculated value of r may be determined using different methods, or may be determined using a maximum likelihood approach to this method. In one embodiment, the method requires maternal genotype knowledge. In one embodiment, the method requires paternal genotype knowledge. In one embodiment, the method does not require knowledge of the paternal genotype. In one embodiment, the fetal fraction and the ratio of maternal to fetal DNA are essentially equivalent and can be used interchangeably after applying the appropriate linear algebraic transformations. In some embodiments, r=[fetal fraction] / [1% fetal fraction].
[0280] SNP classification using Phred scores The Phred score, q, is defined as: P(incorrect base call)=10^(-q / 10).
[0281] Let x = true genotype reference ratio = number of reference alleles / total number of alleles. For disomy, x in {0,0.5,1} corresponds to {MM,RM,RR}. Let z be the alleles observed in sequence z in {R,M}. Here, the likelihood of observing z=R is denoted and the true ratio of reference alleles in the genotype (i.e., P(z=R|x) P(z=R|x)=P(z=R|gc,x)P(gc)+P(z=R|bc,x)P(bc) where gc is the event of a correct call and bc is the event of a bad call.
[0282] P(gc) and P(bc) are calculated from the phred scores: P(z=R|gc,x)=x and P(z=R|bc,x)=1-x, assuming the probes are unbiased.
[0283] As a result, b=P(incorrect base call): P(z=R|x)=x(1-b)+(1-x)*b
[0284] Note that the probabilities of the reference allele measurements converge, as expected, to the reference allele proportions as the phred score improves.
[0285] Assuming that each sequence is generated independently and conditioned on the true genotype, the likelihood of a set of measurements at the same SNP is simply the product of the individual likelihoods. This method accounts for various phred scores. In another embodiment, it is possible to consider variations in the confidence of sequence mapping. Given a set of n sequences for a single SNP, the combination of likelihoods results in a polynomial of order n that can be evaluated at candidate allele ratios representing various hypotheses.
[0286] SNP classification using Phred thresholds When a large number of sequences are available for a single SNP, the polynomial likelihood function for the allele ratio becomes unwieldy. An alternative is to consider only base calls with high phred scores and then assume that they are correct. Each base read is an IID Bernoulli according to the true allele ratio, and the likelihood function is Gaussian. If r is the proportion of reference reads in the data, the likelihood function on x (true reference allele ratio) has mean = r and standard deviation = sqrt(r * (1-r) / n).
[0287] SNP bias correlation between samples Using the above two likelihood functions (polynomial, Gaussian), SNPs can be classified as RR, RM, or MM by considering the allele ratios {1, 0.5, 0}, or maximum likelihood estimates of the allele ratios can be calculated. If the same SNP is classified as RM in two different samples, the MLE estimates of the allele ratios can be compared to look for consistent "probe bias".
[0288] Using sequence length before determining DNA origin The length distribution of the sequences differs between maternal DNA and fetal DNA, and it has been reported that the fetus generally has shorter sequences. In one embodiment of the present disclosure, it is possible to construct prior distributions of the predicted lengths of both the mother (P(X|maternal)) and fetal DNA (P(X|fetal)) in the form of empirical data using prior knowledge. Considering a new unconfirmed DNA sequence of length x, it is possible to assign the probability that a given DNA sequence is either maternal DNA or fetal DNA based on the prior likelihood of x given either the mother or the fetus. In particular, when P(x|maternal) > P(x|fetal), the DNA sequence can be classified as maternal with P(x|maternal) = P(x|maternal) / [(P(x|maternal) + P(x|fetal))], and when p(x|maternal) < p(x|fetal), the DNA sequence can be classified as fetal with P(x|fetal) = P(x|fetal) / [(P(x|maternal) + P(x|fetal))]. In one embodiment of the present disclosure, by considering sequences that can be assigned as having a high probability to either the mother or the fetus, it is possible to determine the length distribution of the maternal and fetal sequences specific to that sample, and then use that sample-specific distribution as the expected size distribution of that sample.
[0289] Method for determining the average copy number in a set of target cells The method described above assumes that the DNA from the target cells is from one target cell or, if not, from target cells that are essentially genetically identical. This assumption may not hold, for example, in the case of a placental mosaic where the target is a fetus and the DNA from the fetus is from a plurality of cells where some of the placental cells are genetically different from other placental cells. For example, in many cases where the fetus is 47, XX+18 or 47, XY+18, the placenta is a mosaic that is a mixture of 46, XX and 47, XX+18 or 46, XY and 47, XY+18, respectively.
[0290] Another example involves the detection of cancer by copy number variants, where the target cells are from the tumor and the non-target cells are non-cancerous cells from the host. A hallmark of cancer is genomic instability, and many, if not all, tumors are genetically heterogeneous. Even small biopsies of tumor tissue show heterogeneity. The ways in which the genome of the cancer cells differs from the native host DNA are considered mutations. Some, but not necessarily all, of these mutations may drive the oncogenic properties of the cancer. In the case of liquid biopsies, i.e., detection of tumor DNA from cell-free DNA (cfDNA) in the bloodstream, the cell-free tumor DNA (ctDNA) is often heterogeneous and is believed to originate from apoptotic or necrotic cancer cells, representing some or all of the cells of the tumor. There are several types of mutations found in cancer, including but not limited to single nucleotide variants (SNVs), copy number variants (CNVs), point mutations, also called hypomethylation, hypermethylation, deletions, and duplications.
[0291] Considering the host's normal disomic genome as the baseline, the analysis of a mixture of normal and cancer cells gives the average difference between the baseline in the mixture and the DNA of the cell of origin of the ctDNA. For example, imagine if 10% of the DNA in the sample comes from cells with a deletion across the region of the chromosome targeted by the assay. The quantification method should show that the amount of reads corresponding to this region is predicted to be 95% of the amount predicted for a normal sample. This is because one of the two targeted chromosomal regions in each tumor cell with a deletion of the targeted region is missing, so the total amount of DNA mapping to this region is 90% (for normal cells) + 1 / 2 x 10% (for tumor cells) = 95%. Alternatively, the allele method should show that the ratio of alleles at heterozygous loci is 19:20 on average. Now imagine if 10% of the DNA in the sample comes from cells with a 5-fold focal amplification of the region of the chromosome targeted by the assay. The quantification method should show that the amount of reads corresponding to this region is predicted to be 125% of the amount predicted for a normal sample. This means that one of the two target chromosomal regions in each tumor cell with a 5-fold focal amplification is overcopied 5-fold across the target region, so the total amount of DNA mapping to this region is 90% (for normal cells) + (2 + 5) x 10% (for tumor cells) / 2 = 125%. Alternatively, the allele approach should show that the ratio of alleles at heterozygous loci is 25:20 on average. Note that using only the allele approach, a 5-fold focal amplification across a chromosomal region in a sample with 10% ctDNA may appear to be the same as a deletion across the same region in a sample with 40% ctDNA. In these two cases, the under-represented haplotype in the deletion case appears to be the haplotype without CNV in the case with the focal duplication, and the haplotype without CNV in the deletion case appears to be the haplotype that is over-represented in the case with the focal duplication. Combining the likelihoods produced by this allele approach with the likelihoods produced by the quantitative approach will distinguish between the two probabilities.
[0292] Parent Context Parental context refers to the genetic state of a given allele in each of the two related chromosomes of one or both of the two parents of the target. Note that in one embodiment, the parental context does not refer to the allelic state of the target, but rather to the allelic state of the parents. The parental context of a given SNP may consist of four base pairs, two paternal and two maternal, which may be the same or different from each other. Typically, "m 1 m 2 |f 1 f 2 " and m 1 and m 2 is the genetic state of a given SNP on the two maternal chromosomes, and f 1 and f 2 is the genetic state of a given SNP on the two paternal chromosomes. In some embodiments, the parental context is "f 1 f 2 |m 1 m 2 may also be written as 。 Note that the subscripts "1" and "2" refer to the genotypes of the first and second chromosomes at a given allele, and that the choice of which chromosome is labeled "1" and which is labeled "2" is arbitrary.
[0293] In this disclosure, A and B are often used to generally represent base pair identity, and it should be noted that A or B can equally well represent C (cytosine), G (guanine), A (adenine), or T (thymine). For example, for a given SNP-based allele, if the maternal genotype is T at that SNP on one chromosome and G at that SNP on the homologous chromosome, and the paternal genotype at that allele is G at that SNP on both homologous chromosomes, then the target individual's allele can be said to have a parent context of AB|BB, and the allele can also be said to have a parent context of AB|AA. It should be noted that theoretically, any of the four possible nucleotides can be present at a given allele, and thus, for example, it is possible for a mother to have a genotype of AT and a father to have a genotype of GC at a given allele. However, empirical data shows that in most cases, only two of the four possible base pairs are observed at a given allele. For example, when using a single tandem repeat, it is possible to have more than two parents, more than four, and even more than 10 contexts. In this disclosure, the discussion assumes that only two possible base pairs are observed at a given allele, although the embodiments disclosed herein may be modified to take into account cases where this assumption is not true.
[0294] "Parent context" may refer to a set or subset of target SNPs that have the same parent context. For example, if 1000 alleles are measured on a given chromosome on a target individual, the context AA|BB may refer to the set of all alleles in the 1000 allele group where the target maternal genotype is homozygous and the target paternal genotype is homozygous, but the maternal and paternal genotypes are different at that locus. If the parent data is unphased, and therefore AB=BA, there are nine possible parent contexts: AA|AA, AA|AB, AA|BB, AB|AA, AB|AB, AB|BB, BB|AA, BB|AB, and BB|BB. If the parent data is phased, and therefore AB≠BA, there are 16 different possible parent contexts. AA|AA, AA|AB, AA|BA, AA|BB, AB|AA, AB|AB, AB|BA, AB|BB, BA|AA, BA|AB, BA|BA, BA|BB, BB|AA, BB|AB, BB|BA, and BB|BB. With the exception of some SNPs on sex chromosomes, every SNP allele on a chromosome has one of these parental contexts. The set of SNPs whose parental context for one parent is heterozygous may be referred to as a heterozygous context.
[0295] Using Parent Context in NPD Non-invasive prenatal testing is an important technology that can be used to determine the genetic status of a fetus from genetic material obtained in a non-invasive manner, for example from a blood draw of the mother during pregnancy. The blood can be separated and the plasma isolated, followed by isolation of the plasma DNA. Size selection can be used to isolate DNA of the appropriate length. The DNA can be preferentially enriched at a set of loci. This DNA can then be measured by several means, for example by hybridization to a genotyping array and measuring fluorescence, or by sequencing on a high throughput sequencer.
[0296] When sequencing is used for fetal ploidy calling in the context of non-invasive prenatal diagnosis, there are several ways to use sequence data. The most common way sequence data can be used is to simply count the number of reads that map to a given chromosome. For example, suppose one is trying to determine the ploidy state of chromosome 21 of a fetus. Further imagine that the DNA in the sample consists of 10% DNA of fetal origin and 90% DNA of maternal origin. In this case, one can look at the average number of reads on a chromosome that is predicted to be disomic (e.g., chromosome 3) and compare it to the number of reads on chromosome 21, where the number of reads is adjusted for the number of base pairs on that chromosome that are part of the unique sequence. If the fetus is euploid, the amount of DNA per unit of the genome is predicted to be approximately equal at all positions (subject to stochastic variation). On the other hand, if the fetus was trisomic at chromosome 21, one would expect slightly more DNA per genetic unit from chromosome 21 than other positions on the genome. Specifically, one would expect approximately 5% more DNA from chromosome 21 in the mixture. When DNA is measured using sequencing, approximately 5% more uniquely mappable reads are expected from chromosome 21 than from other chromosomes for each unique segment. Observation of the amount of DNA from a particular chromosome, when adjusted for the number of uniquely mappable sequences to that chromosome, above a certain threshold, serves as the basis for aneuploidy diagnosis. Another method that can be used to detect aneuploidy is similar to the method above, except that it can take into account parental context.
[0297] When considering which alleles to target, one can consider that some parental contexts are more likely to be informative than others. For example, AA|BB and the symmetric context BB|AA are the most informative contexts because the fetus is known to have different alleles than the mother. For symmetry reasons, both the AA|BB and BB|AA contexts may be called AA|BB. Another set of informative parental contexts are AA|AB and BB|AB because in these cases the fetus has a 50% chance of having an allele that the mother does not have. For symmetry reasons, both the AA|AB and BB|AB contexts may be called AA|AB. A third set of informative parental contexts are AB|AA and AB|BB because in these cases the fetus has a known paternal allele that is also present in the maternal genome. For symmetry reasons, both the AB|AA and AB|BB contexts may be called AB|AA. A fourth parental context is AB|AB where the fetus has an unknown allele state and the mother has the same allele, whatever the allele state. The fifth parent context is AA|AA, and the mother and father are heterozygous.
[0298] Different implementations of the embodiment Disclosed herein is a method for determining the ploidy state of a target individual. The target individual may be a blast, an embryo, or a fetus. In some embodiments of the present disclosure, the method for determining the ploidy state of one or more chromosomes in a target individual may include any of the steps described herein, and combinations thereof.
[0299] In some embodiments, the source of genetic material used to determine the genetic status of the fetus may be fetal cells, such as nucleated fetal red blood cells isolated from maternal blood. The method may involve obtaining a blood sample from a pregnant mother. The method may involve using vision techniques to isolate fetal red blood cells based on the idea that a particular color combination is uniquely associated with nucleated red blood cells and that a similar color combination is not associated with any other present cells in the maternal blood. The color combination associated with nucleated red blood cells may include the red color of the hemoglobin around the nucleus (which color may be made more evident by staining), and the color of the nuclear material, which may be stained blue, for example. The nucleated red blood cells can be located by isolating the cells from the maternal blood, spreading them on a slide, and then identifying a point that sees both red (from the hemoglobin) and blue (from the nuclear material). The nucleated red blood cells may then be extracted using a micromanipulator, and genotyping and / or sequencing techniques may be used to measure the genotypic aspects of the genetic material in those cells.
[0300] In one embodiment, nucleated red blood cells can be stained with a dye that fluoresces only in the presence of fetal hemoglobin and not maternal hemoglobin, thus removing any ambiguity as to the maternal or fetal origin of the nucleated red blood cells. Some embodiments of the present disclosure may involve staining or otherwise marking the nuclear material. Some embodiments of the present disclosure may involve specifically marking fetal nuclear material using fetal cell specific antibodies.
[0301] There are many other methods for isolating fetal cells from maternal blood, or for isolating fetal DNA from maternal blood, or for enriching samples of fetal genetic material in the presence of maternal genetic material. Some of these methods are listed here, but this is not intended to be an exhaustive list. Some suitable techniques are listed here for convenience: fluorescently or otherwise tagged antibodies, size exclusion chromatography, magnetically or otherwise labeled affinity tags, epigenetic differences such as differential methylation between maternal and fetal cells at specific alleles, density gradient centrifugation followed by CD45 / 14 depletion and CD71 positive selection from CD45 / 14 negative cells, single or double Percoll gradients with different osmolarity, or using galactose specific lectin methods.
[0302] In one embodiment of the present disclosure, the target individual is a fetus, and different genotype measurements are performed on multiple DNA samples from the fetus. In some embodiments of the present disclosure, the fetal DNA sample is from isolated fetal cells, where fetal cells can be mixed with maternal cells. In some embodiments of the present disclosure, the fetal DNA sample is from free-floating fetuses, where fetal DNA can be mixed with free-floating maternal DNA. In some embodiments, the fetal DNA sample can be from maternal plasma or maternal blood, which contains a mixture of maternal DNA and fetal DNA. In some embodiments, fetal DNA may be mixed with maternal DNA in a maternal:fetal ratio ranging from 99.9:0.1% to 99:1%, 99:1% to 90:10%, 90:10% to 80:20%, 80:20% to 70:30%, 70:30% to 50:50%, 50:50% to 10:90%, or 10:90% to 1:99%, 1:99% to 0.1:99.9%.
[0303] In some embodiments, the genetic sample may be prepared and / or purified. There are several standard procedures known in the art to achieve such ends. In some embodiments, the sample may be centrifuged to separate the various layers. In some embodiments, the DNA may be isolated using filtration. In some embodiments, the preparation of DNA may involve amplification, separation, chromatographic purification, liquid-liquid separation, isolation, preferential enrichment, preferential amplification, targeted amplification, or any of several other techniques known in the art or described herein.
[0304] In some embodiments, the methods of the present disclosure may involve amplifying DNA. Amplifying DNA, the process of converting a small amount of genetic material into a larger amount of genetic material containing a similar genetic data set, can be done by a wide variety of methods, including but not limited to polymerase chain reaction (PCR). One method of amplifying DNA is whole genome amplification (WGA). There are several methods available for WGA, such as ligation-mediated PCR (LM-PCR), degenerate oligonucleotide primer PCR (DOP-PCR), and multiple displacement amplification (MDA). In LM-PCR, short DNA sequences, called adapters, are ligated to the blunt ends of the DNA. These adapters contain universal amplification sequences that are used to amplify DNA by PCR. In DOP-PCR, random primers that also contain universal amplification sequences are used in the first round of annealing and PCR. A second round of PCR is then used to further amplify the sequences using the universal primer sequences. MDA uses phi-29 polymerase, a highly processive non-specific enzyme that replicates DNA and has been used for single cell analysis. The main limitations to amplifying material from single cells are (1) the need to use very dilute DNA concentrations or very small reaction mixtures, and (2) the difficulty of reliably dissociating DNA from proteins throughout the genome. Nevertheless, single cell whole genome amplification has been used successfully for many years for a variety of applications. Other methods exist for amplifying DNA from DNA samples. DNA amplification converts the initial sample of DNA into a sample of DNA that is similar in sequence set but of much larger amounts. In some cases, amplification may not be required.
[0305] In some embodiments, DNA may be amplified using universal amplification, such as WGA or MDA. In some embodiments, DNA may be amplified by targeted PCR, or targeted amplification using circularization probes, for example. In some embodiments, DNA may be preferentially enriched using targeted amplification methods, or methods that result in complete or partial separation from undesired DNA, such as capture by hybridization approaches. In some embodiments, DNA may be amplified by using a combination of universal amplification and preferential enrichment methods. A more complete description of some of these methods can be found elsewhere in this document.
[0306] The genetic data of the target individual and / or related individuals can be converted from a molecular state to an electronic state by measuring the appropriate genetic material using tools and / or techniques from a group including, but not limited to, genotyping microarrays and high-throughput sequencing. Some high-throughput sequencing methods include Sanger DNA sequencing, pyrosequencing, ILLUMINA SOLEXA platform, ILLUMINA's genome analyzer, or APPLIED BIOSYSTEM's 454 sequencing platform, HELICOS' TRUE SINGLE MOLECULE SEQUENCING platform, HALCYON MOLECULAR's electron microscope sequencing, or any other sequencing method. All of these methods physically convert the genetic data stored in a sample of DNA into a set of genetic data that is typically stored in a memory device in the process of being processed.
[0307] Genetic data of a related individual may be measured by analyzing material obtained from groups including, but not limited to, bulk diploid tissue of the individual, one or more diploid cells from the individual, one or more haploid cells from the individual, one or more blast cells from the target individual, extracellular genetic material found on the individual, extracellular genetic material from the individual found in maternal blood, cells from the individual found in maternal blood, one or more embryos generated from (a) gamete(s) from the related individual, one or more blast cells obtained from such embryos, extracellular genetic material found on the related individual, genetic material known to be derived from the related individual, and combinations thereof.
[0308] In some embodiments, a set of at least one ploidy state hypothesis may be created for each of the chromosome types of interest of the target individual. Each ploidy state hypothesis may refer to one possible ploidy state of a chromosome or chromosome segment of the target individual. The set of hypotheses may include some or all of the possible ploidy states that a chromosome of a target individual may be expected to have. Some of the possible ploidy states may include nullosomy, monosomy, disomy, uniparental disomy, euploid, trisomy, matched trisomy, incompatible trisomy, maternal trisomy, paternal trisomy, tetrasomy, balanced (2:2) tetrasomy, unbalanced (3:1) tetrasomy, pentasomy, hexasomy, other aneuploids, and combinations thereof. Any of these aneuploid states may be mixed or partial aneuploids, such as unbalanced translocations, balanced translocations, Robertsonian translocations, recombinations, deletions, insertions, crossovers, and combinations thereof.
[0309] In some embodiments, knowledge of the determined ploidy state may be used to make a clinical decision. This knowledge is typically stored as a physical arrangement of material in a memory device and then converted into a report. The report may then be acted upon. For example, the clinical decision may be to terminate the pregnancy, or alternatively, the clinical decision may be to continue the pregnancy. In some embodiments, the clinical decision may involve a decision to take an intervention designed to reduce the severity of the phenotypic presentation of the genetic disorder, or related steps to prepare for a child with special needs.
[0310] In one embodiment of the present disclosure, any of the methods described herein may be modified to allow multiple targets to be derived from the same target individual, e.g., multiple blood draws from the same pregnant mother. This may improve the accuracy of the model, since multiple genetic measurements may provide more data from which target genotypes can be determined. In one embodiment, one set of target genetic data served as the reported primary data, and the other set served as data to double-check the primary target genetic data. In one embodiment, multiple sets of genetic data, each measured from genetic material collected from a target individual, are considered in parallel, and thus both sets of target genetic data help determine which sections of parental genetic data measured with high accuracy constitute the fetal genome.
[0311] In one embodiment, the method may be used for paternity testing purposes. For example, given SNP-based genotype information from a mother, and SNP-based genotype information from a man who may or may not be the genetic father, and measured genotype information from a mixed sample, it is possible to determine whether the man's genotype information actually represents the actual genetic father of a pregnant fetus. A simple way to do this is to simply look at possible contexts where the mother is AA and the father is AB or BB. In these cases, one may expect to see half (AA|AB) or all (AA|BB) of the paternal contribution, respectively. Given the expected ADO, it is easy to determine whether the observed fetal SNPs are correlated with the possible paternal SNPs.
[0312] One embodiment of the present disclosure may be as follows: A pregnant woman wants to know if her fetus is affected by Down's Syndrome and / or if her fetus is affected by Cystic Fibrosis, and the pregnant woman does not want to have a child affected by either of these conditions. A doctor draws the pregnant woman's blood and stains the hemoglobin with one marker to make it appear distinctly red and the nuclear material with another marker to make it appear distinctly blue. The doctor knows that the mother's red blood cells are typically anuclear, but because a high percentage of fetal cells contain nuclei, he is able to visually isolate some nucleated red blood cells by identifying cells that exhibit both red and blue colors. The doctor removes these cells from the slide with a micromanipulator and sends them to a lab where 10 individual cells are amplified and genotyped. Using genetic measurements, PARENTAL SUPPORT™ is able to determine that 6 out of 10 cells are maternal blood cells and 4 out of 10 cells are fetal cells. If a child is already born to a pregnant mother, PARENTAL SUPPORT™ can also be used to determine that the fetal cells are different from the cells of the birth child by calling high-confidence alleles on the fetal cells and showing that they are different from the cells of the birth child. Note that this method is similar in concept to the paternity testing embodiment of the present disclosure. The genetic data measured from the fetal cells can be of very low quality, including many allele dropouts, due to the difficulty of genotyping a single cell. Using the measured fetal DNA, along with high-confidence DNA measurements of the parents, the clinician can use PARENTAL SUPPORT™ to infer aspects of the fetal genome with high accuracy, thereby converting the genetic data contained in the genetic material from the fetus into a predicted genetic state of the fetus stored in a computer. The clinician can determine both the ploidy state of the fetus and the presence or absence of multiple disease-associated genes of interest. The fetus is found to be euploid and not a carrier for cystic fibrosis, and the mother decides to continue the pregnancy.
[0313] In one embodiment of the present disclosure, a pregnant mother wants to determine if her fetus is affected with any whole chromosomal abnormalities. The pregnant woman goes to the doctor and provides a sample of blood, and the pregnant woman and her husband provide a sample of their own DNA from a cheek swab. Laboratory researchers genotype the parental DNA using MDA protocols to amplify the parental DNA and use ILLUMINA INFINIUM arrays to measure the parental genetic data at a large number of SNPs. The researchers then spin the blood, collect the plasma, and use size-exclusion chromatography to isolate a sample of free-floating DNA. Alternatively, the researchers use one or more fluorescent antibodies, such as one specific for fetal hemoglobin, to isolate fetal red blood cells that have nuclei. The researchers then take the isolated or enriched fetal genetic material and amplify it using a library of appropriately designed 70-mer oligonucleotides, where the two ends of each oligonucleotide correspond to the flanking sequences on either side of the target allele. Upon addition of polymerase, ligase, and appropriate reagents, the oligonucleotides undergo gap-filling circularization to capture the desired allele. Exonuclease was added and heat inactivated, and the product was used directly as a template for PCR amplification. The PCR product was sequenced on an ILLUMINA genome analyzer. The sequence reads were then used as input for the PARENTAL SUPPORT™ method, which predicted the ploidy state of the fetus.
[0314] In another embodiment, a couple, a pregnant mother and an older mother, want to know if the pregnant fetus has Down's syndrome, Turner's syndrome, Prader-Willi syndrome, or some other whole chromosomal abnormality. An obstetrician draws blood from the mother and father. The blood is sent to a laboratory where a technician centrifuges the maternal sample to separate the plasma and buffy coat. The DNA in the buffy coat and paternal blood sample is transformed by amplification, and the genetic data encoded in the amplified genetic material is further transformed from molecularly stored genetic data to electronically stored genetic data by running the genetic material on a high-throughput sequencer to measure parental genotypes. The plasma sample is preferentially enriched at a set of loci using a 5,000-plex semi-nested targeted PCR method. The mixture of DNA fragments is prepared into a DNA library suitable for sequencing. The DNA is then sequenced using a high-throughput sequencing method, for example, an ILLUMINA GAIIx genome analyzer. Sequencing converts the information molecularly encoded in DNA into electronically encoded information in computer hardware. The ploidy state of the fetus may be determined using informatics-based techniques including the disclosed embodiments such as PARENTAL SUPPORT™. This may involve: calculating, in a computer, the probability of allele counts at a plurality of polymorphic loci from DNA measurements made on the prepared samples; generating, in a computer, a plurality of ploidy hypotheses each associated with a different possible ploidy state of the chromosome; constructing, in a computer, a joint distribution model of the expected allele counts at the plurality of polymorphic loci on the chromosome for each ploidy hypothesis; determining, in a computer, the relative probability of each of the plurality of ploidy hypotheses using the joint distribution model and the allele counts measured on the prepared samples; and calling the ploidy state of the fetus by selecting the ploidy state corresponding to the hypothesis with the greatest probability. The fetus is determined to have Down's syndrome. A report is printed or sent electronically to the pregnant woman's obstetrician and a diagnosis is sent to the woman. The woman, her husband and the doctor sit down to discuss their options.A couple decides to terminate a pregnancy based on the knowledge that the fetus is affected by a trisomic condition.
[0315] In one embodiment, a company may decide to offer a diagnostic technology designed to detect aneuploidy in a pregnant fetus from a maternal blood draw. Their product may include the mother drawing blood to an obstetrician. The obstetrician may also collect a genetic sample from the father of the fetus. The physician may isolate plasma from the maternal blood and purify DNA from the plasma. The physician may also isolate the buffy coat layer from the maternal blood and prepare DNA from the buffy coat. The physician may also prepare DNA from a paternal genetic sample. The clinician may add a universal amplification tag to DNA in the DNA derived from the plasma sample using the molecular biology techniques described in this disclosure. The clinician may amplify the universally tagged DNA. The clinician may preferentially enrich the DNA by several techniques including capture by hybridization and targeted PCR. The targeted PCR may include nesting, semi-nesting or semi-nesting, or any other approach that results in efficient enrichment of plasma-derived DNA. Targeted PCR may be massively multiplexed, for example with 10,000 primers in one reaction, targeting SNPs on chromosomes 13, 18, 21, X, and those loci common to both X and Y, and optionally other chromosomes. Selective enrichment and / or amplification may involve tagging each individual molecule with a different tag, molecular barcode, tag for amplification and / or tag for sequencing. The clinician may then sequence the plasma sample and, optionally, the prepared maternal and / or paternal DNA. The molecular biology process may be performed in whole or in part by the diagnostic box. The sequence data may be fed to a single computer or to another type of computing platform, such as may be found in the "cloud". The computing platform may calculate the allele count at the targeted polymorphic locus from the measurements made by the sequencer. The computing platform may generate multiple ploidy hypotheses relating to nullsomy, monosomy, disomy, matched trisomy, and unmatched trisomy for each of chromosomes 13, 18, 21, X, and Y.The computing platform may construct a joint distribution model of the predicted allele counts at the target locus on the chromosome for each ploidy hypothesis for each of the five chromosomes examined. The computing platform may use the joint distribution model and the allele counts measured for the preferentially enriched DNA from the plasma sample to determine the probability that each of the ploidy hypotheses is true. The computing platform may call the ploidy state of the fetus for each of chromosomes 13, 18, 21, X, and Y by selecting the ploidy state corresponding to the Germanic hypothesis with the greatest probability. A report including the called ploidy state may be generated, which may be sent electronically to the obstetrician, displayed on an output device, or a printed hard copy of the report may be delivered to the obstetrician. The obstetrician may inform the patient and optionally the father of the fetus, and may decide which clinical options are open to them and which option is most desirable.
[0316] In another embodiment, a pregnant woman (hereafter referred to as the "mother") may decide that she would like to know if her fetus(es) have any genetic abnormalities or other conditions. She may want to make sure there are no significant abnormalities before she is confident in continuing the pregnancy. The mother may go to her obstetrician and have a blood sample taken. The obstetrician may also take a genetic sample, such as a buccal swab from her cheek. She may also take a genetic sample from the father of the fetus, such as a buccal swab, sperm sample, or blood sample. The father may send the sample to a clinician. The clinician may enrich for a fraction of free floating fetal DNA in the maternal blood sample. The clinician may enrich for a fraction of enucleated fetal blood cells in the maternal blood sample. The clinician may use various aspects of the methods described herein to determine genetic data of the fetus. The genetic data may include the ploidy state of the fetus and / or the identity of one or several disease-associated alleles in the fetus. A report may be generated summarizing the results of the prenatal diagnosis. The report can be sent or mailed to a physician, who can inform the mother of the genetic status of the fetus. The mother can decide to terminate the pregnancy based on the fact that the fetus has one or more chromosomal or genetic abnormalities or undesirable conditions. The mother can also decide to continue the pregnancy based on the fact that the fetus does not have a gross chromosomal or genetic abnormality or genetic disease of interest.
[0317] Another example may involve a pregnant woman who has been artificially inseminated by a sperm donor and is pregnant. The mother wishes to minimize the risk of the fetus suffering from a genetic disease during her pregnancy. The mother has blood drawn at a blood collection clinic and using the techniques described in this disclosure, three nucleated fetal red blood cells are isolated and tissue samples are also collected from the mother and genetic father. The genetic material from the fetus and the mother and father are suitably amplified and genotyped using the ILLUMINA INFINIUM BEADARRAY and the methods described herein result in highly accurate parental and fetal genotype cleansing, phasing and fetal ploidy calls. The fetus is found to be euploid and phenotypic susceptibility is predicted from the reconstituted fetal genotype and a report is generated and sent to the mother's physician to determine which clinical decision is best.
[0318] In one embodiment, the raw genetic material of the mother and father is converted by amplification into amounts of DNA that are similar in sequence but are of higher abundance. The genotyping method then converts the genotype data encoded by the nucleic acids into genetic measurements that can be physically and / or electronically stored on a memory device as described above. The associated algorithms that make up the PARENTAL SUPPORT™ algorithm, discussed in detail herein, are converted into a computer program using a programming language. Then, by executing the computer program on computer hardware, instead of being physically encoded bits and bytes arranged in a pattern that represents the raw measurement data, they are converted into a pattern that represents a high confidence determination of the ploidy status of the fetus. The details of this conversion depend on the data itself, and the computer language and hardware system used to execute the method described herein. The data, now physically configured to represent a high quality ploidy determination of the fetus, is then converted into a report that can be sent to a medical professional. This conversion may be done using a printer or a computer display. The report may be a printed copy, paper or other suitable medium, or it may be electronic. In the case of an electronic report, it may be transmitted, physically stored in a memory device at a location on the computer accessible to the medical personnel, or displayed on a screen so that it may be read. In the case of a screen display, the data may be converted into a readable format by causing a physical transformation of pixels on the display device. The transformation may be accomplished by physically firing electrons at a phosphorescent screen, by changing an electrical charge that physically changes the transparency of a particular set of pixels on the screen, which may be in front of a substrate that emits or absorbs photons. This transformation may be accomplished by changing the nanoscale orientation of molecules in a liquid crystal, for example, from a nematic phase to a cholesteric or vicinal phase at a particular set of pixels. This transformation may be accomplished by an electrical current that causes a photon to be emitted from a particular set of pixels made from a plurality of light emitting diodes arranged in a meaningful pattern.This conversion may be accomplished by any other method used to display information, such as a computer screen, or some other output device or method of transmitting information. A medical professional may then take an action based on the report, such that the data in the report is converted into an action. The action may be to continue or terminate the pregnancy, in which case the pregnant fetus with the genetic abnormality is converted into a non-viable fetus. The conversions listed herein may be aggregated, for example, such that through several steps outlined in this disclosure, the genetic material of the pregnant mother and father can be converted into a medical decision consisting of terminating the fetus with the genetic abnormality or consisting of continuing the pregnancy. Alternatively, a set of genotype measurements can be converted into a report that helps a physician treat a pregnant patient.
[0319] In one embodiment of the present disclosure, the methods described herein can be used to determine the ploidy state of a fetus even if the host mother, i.e., the pregnant woman, is not the biological mother of the fetus she is carrying. In one embodiment of the present disclosure, the methods described herein can be used to determine the ploidy state of a fetus using only a maternal blood sample, without the need for a paternal genetic sample.
[0320] Some of the mathematics in the embodiments disclosed herein make assumptions regarding a limited number of aneuploidy states. In some cases, for example, only zero, one, or two chromosomes are predicted to originate from each parent. In some embodiments of the present disclosure, the mathematical derivation can be expanded to take into account other forms of aneuploidy, such as quadrosomy, pentasomy, hexasomy, where three chromosomes originate from one parent, without changing the basic concept of the present disclosure. At the same time, it is possible to focus only on a smaller number of ploidy states, for example, trisomy and disomy. It should be noted that ploidy determinations that indicate a non-complete number of chromosomes may indicate mosaicism in a sample of genetic material.
[0321] In some embodiments, the genetic abnormality is a type of aneuploidy, such as Down's syndrome (or trisomy 21), Edwards syndrome (trisomy 18), Patau syndrome (trisomy 13), Turner syndrome (45x), Klinefelter syndrome (males with two chromosomes), Prader-Willi syndrome, and DiGeorge syndrome (UPD15). Congenital disorders such as those listed in the preceding sentence are generally undesirable, and the knowledge that a fetus is affected by one or more phenotypic abnormalities may provide the basis for deciding to terminate the pregnancy, take the necessary precautions to prepare for the birth of a child with special needs, or take some therapeutic approach intended to reduce the severity of the chromosomal abnormality.
[0322] In some embodiments, the methods described herein can be used very early in gestation, for example, as early as 4 weeks, as early as 5 weeks, as early as 6 weeks, as early as 7 weeks, as early as 8 weeks, as early as 9 weeks, as early as 10 weeks, as early as 11 weeks, and as early as 12 weeks.
[0323] It is noted that it has been demonstrated that DNA from cancers parasitizing a host can be found in the host's blood. Just as genetic diagnoses can be made from measurements of mixed DNA found in maternal blood, so too can genetic diagnoses from measurements of mixed DNA found in host blood. Genetic diagnoses can include aneuploidy status, or genetic mutations. Any claim in this disclosure that reads about determining the ploidy status or genetic status of a fetus from measurements made in maternal blood can equally well be read about determining the ploidy status or genetic status of a cancer from measurements in host blood.
[0324] In some embodiments, the disclosed method allows for the determination of the ploidy state of the cancer, the method includes obtaining a mixed sample containing genetic material from the host and genetic material from the cancer, measuring DNA in the mixed sample, calculating the fraction of DNA from the cancer in the mixed sample, and using the measurements made on the mixed sample and the calculated fraction to determine the ploidy state of the cancer. In some embodiments, the method may further include administering a cancer therapeutic agent based on the determination of the ploidy state of the cancer. In some embodiments, the method may further include administering a cancer therapeutic agent based on the determination of the ploidy state of the cancer, the cancer therapeutic agent being taken from a group including pharmaceuticals, biological therapeutic agents, and antibody-based therapeutic agents, and combinations thereof.
[0325] In some embodiments, the methods disclosed herein are used in the context of preimplantation genetic diagnosis (PGD) for embryo selection during in vitro fertilization, where the target individual is an embryo and parental genotype data can be used to make ploidy determinations on the embryo from sequencing data from single or two cell biopsies from day 3 embryos or trophectoderm biopsies from day 5 or 6 embryos. In a PGD setting, only the child's DNA is measured and only a small number of cells are tested, typically 1-5, but as many as 10, 20 or 50 cells are tested. The total number of starting copies of A and B alleles (at SNPs) is then determined by the child's genotype and the number of cells. In NPD, the number of starting copies is very high, and therefore the allele ratios after PCR are expected to accurately reflect the starting ratios. However, the low number of starting copies in PGD means that contamination and imperfect PCR efficiency have a negligible effect on the allele ratios after PCR. This effect may be more important than read depth in predicting the variance of measured allele ratios after sequencing. The distribution of measured allele ratios given known offspring genotypes may be generated by Monte Carlo simulation of the PCR process based on PCR probe efficiency and the probability of contamination. Given the allele ratio distribution of each likely offspring genotype, the likelihood of various hypotheses can be calculated as described for NIPD.
[0326] Any of the embodiments disclosed herein may be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, or combinations thereof. The apparatus of the embodiments disclosed herein may be implemented in a computer program product embodied in a machine-readable storage device for execution by a programmable processor, and the method steps of the embodiments disclosed herein may be performed by the programmable processor executing a program of instructions to perform the functions of the embodiments disclosed herein by manipulating input data and generating output. The embodiments disclosed herein may be advantageously implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. Each computer program may be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language, if desired, and in any case the language may be a compiled or interpreted language. A computer program may be deployed in any form, such as a stand-alone program, or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed to be executed or interpreted on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communications network.
[0327] As used herein, a computer-readable storage medium refers to a physical or tangible storage device (as opposed to a signal), including, but not limited to, volatile and non-volatile, removable and non-removable media implemented in any method or technology for the tangible storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid-state memory technology, CD-ROM, DVD, or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other physical or material medium that can be used to tangibly store desired information or data or instructions and that can be accessed by a computer or processor.
[0328] Any of the methods described herein may include output of data in a physical format, such as, for example, on a computer screen or on printed paper. In the description of any embodiment elsewhere herein, it is understood that the described method may be combined with output of actionable data in a format that can be acted upon by a physician. In addition to this, the described method may be combined with the actual implementation of a clinical decision resulting in a clinical treatment, or with the implementation of a clinical decision to take no action. Some of the embodiments described herein for determining genetic data regarding a target individual may be combined with a decision to select one or more embryos for transfer in the context of IVF, and optionally with transfer of the embryos into the uterus of a prospective mother. Some of the embodiments described herein for determining genetic data regarding a target individual may be combined with notification of a potential chromosomal abnormality, or lack thereof, by a medical professional, and optionally with a decision to abort or not abort the fetus in the context of prenatal diagnosis. Some of the embodiments described herein may be combined with output of actionable data, making a clinical decision that results in clinical treatment, or making a clinical decision to take no action.
[0329] Targeted enrichment and sequencing The use of techniques to enrich samples of DNA at a set of target loci following sequencing as part of a method for non-invasive prenatal allele calling or ploidy calling may confer some unexpected advantages. In some embodiments of the present disclosure, the method involves measuring genetic data for use with informatics-based methods such as PARENTAL SUPPORT® (PS). The end result of some of the embodiments is actionable genetic data of the embryo or fetus. There are many methods that can be used to measure the genetic data of an individual and / or related individuals as part of the embodied methods. In one embodiment, a method for enriching the concentration of a set of target alleles is disclosed herein, the method comprising one or more of the following steps: targeted amplification of genetic material, addition of locus-specific oligonucleotide probes, ligation of specific DNA strands, isolation of a set of desired DNA, removal of undesired components of the reaction, detection of specific DNA sequences by hybridization, and detection of sequences of one or more DNA strands by DNA sequencing. In some cases, the DNA strands may refer to the target genetic material, in some cases they may refer to primers, in some cases they may refer to the synthesized sequences, or combinations thereof. These steps may be performed in several different orders. Given the highly diverse nature of molecular biology, it is generally not clear which methods and combinations of steps will perform poorly, well, or best in various situations.
[0330] For example, a universal amplification step of DNA prior to targeted amplification may confer several advantages, such as removing the risk of bottlenecks and reducing allelic bias. The DNA may be mixed with oligonucleotide probes that can hybridize with two adjacent regions of one target sequence on either side. After hybridization, the ends of the probes may be connected by adding a polymerase, a means for ligation, and any necessary reagents that allow circularization of the probes. After circularization, an exonuclease may be added to digest the non-circularized genetic material, followed by detection of the circularized probes. The DNA may be mixed with PCR primers that can hybridize with two adjacent regions of one target sequence on either side. After hybridization, the ends of the probes may be connected by adding a polymerase, a means for ligation, and any necessary reagents to complete the PCR amplification. The amplified or non-amplified DNA may be targeted by hybrid capture probes that target a set of loci, and after hybridization, the probes may be localized and separated from the mixture to provide a mixture of DNA enriched in the target sequences.
[0331] In some embodiments, detection of target genetic material may be performed in a multiplexed manner. The number of genetic target sequences that may be performed in parallel may range from 1-10, 10-100, 100-1000, 1000-10,000, 100,000-100,000, 100,000-1 million, or 1 million-10 million. It is noted that the prior art includes disclosures of successful multiplexed PCR reactions involving pools of up to about 50 or 100 primers, but no more. Previous attempts to multiplex more than 100 primers per pool have resulted in significant problems with undesirable side reactions such as primer dimer formation.
[0332] In some embodiments, the methods may be used to genotype a single cell, a small number of cells, 2-5 cells, 6-10 cells, 10-20 cells, 20-50 cells, 50-100 cells, 100-1000 cells, or small amounts of extracellular DNA, e.g., 1-10 picograms, 10-100 picograms, 100 picograms to 1 nanogram, 1-10 nanograms, 10-100 nanograms, or 100 nanograms to 1 microgram.
[0333] The use of methods to target certain loci followed by sequencing as part of a method for allele or ploidy calling may confer several unexpected advantages. Some ways in which DNA may be targeted or preferentially enriched include using circularization probes, ligated inversion probes (LIPs, MIPs), capture by hybridization methods such as SURESELECT, and targeted PCR or ligation-mediated PCR amplification strategies.
[0334] In some embodiments, the methods of the disclosure involve measuring genetic data for use with informatics-based methods such as PARENTAL SUPPORT® (PS). PARENTAL SUPPORT™ is an informatics-based approach to manipulate genetic data, aspects of which are described herein. The end result of some of the embodiments is actionable genetic data of the embryo or fetus, after which clinical decisions are made based on the actionable data. The algorithm behind the PS method can take measured genetic data of the target individual (often an embryo or fetus) and measured genetic data from related individuals to increase the accuracy with which the genetic status of the target individual is known. In one embodiment, the measured genetic data is used in the context of making ploidy determinations during prenatal genetic diagnosis. In one embodiment, the measured genetic data is used in the context of making ploidy determinations or allele calls on embryos during in vitro fertilization. In the aforementioned context, there are many methods that can be used to measure the genetic data of an individual and / or related individuals. The different methods include several steps, which often involve amplifying the genetic material, adding oligonucleotide probes, ligating specific DNA strands, isolating the set of desired DNA, removing undesired components of the reaction, detecting specific sequences of DNA by hybridization, detecting the sequence of one or more DNA strands by DNA sequencing methods. In some cases, the DNA strands may refer to the target genetic material, in some cases they may refer to primers, in some cases they may refer to synthesized sequences, or combinations thereof. These steps may be performed in several different orders. Given the highly diverse nature of molecular biology, it is generally not clear which methods and combinations of steps will perform poorly, well, or best in various situations.
[0335] It should be noted that theoretically, it is possible to target any number of loci in a genome, from one locus to well over a million loci. When a sample of DNA is targeted and then sequenced, the percentage of alleles read by the sequencer is enriched with respect to their natural abundance in the sample. Enrichment can be 1% (or less) to 10x, 100x, 1000x, or 10 millionx. There are approximately 3 billion base pairs in the human genome, and nucleotides containing about 75 million polymorphic loci. The more loci that are targeted, the lower the enrichment. The fewer the number of loci that are targeted, the greater the enrichment possible, and the greater the read depth can be achieved at those loci for a given number of sequence reads.
[0336] In one embodiment of the present disclosure, the targeting or preference may be entirely focused on SNPs. In one embodiment, the targeting or preference may be focused on any polymorphic site. Several commercial targeting products are available for enriching exons. Surprisingly, when using methods for NPD that rely on allele distribution, it is particularly advantageous to target exclusively SNPs or exclusively polymorphic loci. There are also published methods for NPD using sequencing, such as U.S. Patent No. 7,888,017, which involves read count analysis, where the read count focuses on counting the number of reads that map to a given chromosome, and the analyzed sequence reads do not focus on the regions of the genome that are polymorphic. These types of methodologies that do not focus on polymorphic alleles would not benefit much from targeting or preferential enrichment of a set of alleles.
[0337] In one embodiment of the present disclosure, it is possible to use targeted methods that focus on SNPs to enrich genetic samples in polymorphic regions of the genome. In one embodiment, it is possible to focus on a small number of SNPs, for example, 1 to 100 SNPs, or a larger number, for example, 100 to 1,000, 1,000 to 10,000, 10,000 to 100,000, or more than 100,000 SNPs. In one embodiment, it is possible to focus on one or a small number of chromosomes that are correlated with vivacious trisomy births, for example, chromosomes 13, 18, 21, X and Y, or some combination thereof. In one embodiment, it is possible to enrich the targeted SNPs by a small factor, for example, between 1.01 and 100 times, or by a larger factor, for example, between 100 and 1,000,000 times, or even more than 1,000,000 times. In one embodiment of the present disclosure, it is possible to use targeted methods to generate samples of DNA that are preferentially enriched in polymorphic regions of the genome. In one embodiment, this method can be used to create a mixture of DNA with any of these characteristics, where the mixture of DNA includes maternal DNA and also free floating fetal DNA. In one embodiment, this method can be used to create a mixture of DNA with any combination of these factors. For example, the method described herein can be used to generate a mixture of DNA that includes maternal DNA and fetal DNA and is preferentially enriched in DNA corresponding to 200 SNPs, all located on either chromosome 18 or 21, with an average enrichment of 1000-fold. In another example, this method can be used to create a mixture of DNA that is preferentially enriched in 10,000 SNPs, all or mostly located on chromosomes 13, 18, 21, X and Y, with an average enrichment per locus of more than 500-fold. Any of the targeting methods described herein can be used to create a mixture of DNA that is preferentially enriched at specific loci.
[0338] In some embodiments, the methods of the present disclosure further include measuring the DNA in the mixed fraction using a high-throughput DNA sequencer, wherein the DNA in the mixed fraction contains a disproportionate number of sequences from one or more chromosomes, the one or more chromosomes being taken from the group including chromosome 13, chromosome 18, chromosome 21, chromosome X, chromosome Y, and combinations thereof.
[0339] Three methods are described herein, multiplex PCR, targeted capture by hybridization, and ligated inversion probes (LIP), that may be used to obtain and analyze measurements from a sufficient number of polymorphic loci from maternal plasma samples to detect fetal aneuploidy. This is not meant to exclude other methods of selective enrichment of target loci. Other methods may be used equally well without changing the essence of the method. In each case, the assayed polymorphisms may include single nucleotide polymorphisms (SNPs), small indels, or STRs. The preferred method involves the use of SNPs. Each approach generates allele frequency data, and the allele frequency data for each target locus and / or the combined allele frequency distribution from these loci may be analyzed to determine fetal ploidy. Each approach has its own considerations due to the limited source material and the fact that maternal plasma is composed of a mixture of maternal and fetal DNA. This method may be combined with other approaches to provide a more accurate determination. In one embodiment, this method may be combined with a sequence counting approach as described in U.S. Pat. No. 7,888,017. The described approaches can also be used to non-invasively detect fetal paternity from maternal plasma samples. In addition, each approach can be applied to other mixtures of DNA or pure DNA samples to detect the presence or absence of aneuploid chromosomes, genotype multiple SNPs from degraded DNA samples, detect segmental copy number variations (CNVs), other genotypic states of interest, or some combination thereof.
[0340] Accurate measurement of allele distribution in a sample Current sequencing approaches can be used to estimate the distribution of alleles in a sample. One such method involves random sampling sequences from pooled DNA, called shotgun sequencing. The proportion of a particular allele in the sequencing data is typically very low and can be determined by simple statistics. The human genome contains about 3 billion base pairs. Thus, if the sequencing method used produces 100 bp reads, a particular allele will be measured approximately once in every 30 million sequence reads.
[0341] In one embodiment, the method of the present disclosure is used to determine the presence or absence of two or more different haplotypes containing the same set of loci in a sample of DNA from the measured allele distribution of loci from that chromosome. The different haplotypes can represent two different homologous chromosomes from one individual, three different homologous chromosomes from a trisomic individual, three different homologous haplotypes from a mother and fetus where one haplotype is shared between the mother and fetus, three or four haplotypes from a mother and fetus where one or two haplotypes are shared between the mother and fetus, or other combinations. Alleles that are polymorphic between haplotypes tend to be more informative, but any alleles where the mother and father are not homozygous for the same allele will yield useful information through the measured allele distribution beyond that available from simple read count analysis.
[0342] However, shotgun sequencing of such samples is highly inefficient, as it results in many sequences for chromosomes that are not polymorphic or of interest among different haplotypes in the sample, and therefore does not reveal information about the proportion of the target haplotype. Described herein is a method for specifically targeting and / or preferentially enriching segments of DNA in a sample that are more likely to be polymorphic in the genome, in order to increase the yield of allele information obtained by sequencing. It should be noted that for the measured allele distribution in the enriched sample to truly represent the actual amount present in the target individual, it is important that there is little or no preferential enrichment of one allele compared to other alleles at a given locus in the target segment. Current methods known in the art for targeting polymorphic alleles are designed to ensure that at least some of any alleles present are detected. However, these methods are not designed for the purpose of measuring the unbiased allele distribution of polymorphic alleles present in the original mixture. It is not clear that any particular method of targeted enrichment can produce an enriched sample whose measured allele distribution more accurately represents the allele distribution present in the original unamplified sample than any other method. While many enrichment methods can theoretically be expected to achieve such a goal, those skilled in the art are well aware that there are many stochastic or deterministic biases in current amplification, targeting, and other preferential enrichment methods. One embodiment of the method described herein allows multiple alleles found in a mixture of DNA corresponding to a given locus in a genome to be amplified or preferentially enriched in such a way that the enrichment of each of the alleles is approximately the same. In other words, the method allows the relative amount of alleles present in the mixture as a whole to be increased, but the ratio between the alleles corresponding to each locus remains essentially the same as in the original mixture of DNA. Prior art methods of preferential enrichment of loci can result in allele biases of more than 1%, more than 2%, more than 5%, or even more than 10%.This preferential enrichment may be due to capture bias when using a capture by hybridization approach, or amplification bias, which may be small for each cycle, but may be large when compounded over 20, 30, or 40 cycles. For purposes of this disclosure, the ratio remains essentially the same means that the ratio of alleles in the original mixture divided by the ratio of alleles in the resulting mixture is 0.95-1.05, 0.98-1.02, 0.99-1.01, 0.995-1.005, 0.998-1.002, 0.999-1.001, or 0.9999-1.0001. Note that the allele ratio calculations presented here may not be used to determine the ploidy state of the target individual, but may only be a metric used to measure allele bias.
[0343] In one embodiment, once a mixture is preferentially enriched at a set of target loci, it may be sequenced using any one of the previous, current, or next generation sequencing instruments that sequence clonal samples (samples generated from single molecules, e.g., ILLUMINA GAIIx, ILLUMINA HiSeq, Life Technologies SOLiD, 5500XL). This ratio may be assessed by sequencing through specific alleles within the target region. These sequencing reads may be analyzed and counted according to the allele type and the ratio of different alleles determined accordingly. For mutations that are one to a few bases in length, allele detection is done by sequencing, and it is essential that the sequencing reads span the allele in question to assess the allele composition of that captured molecule. The total number of captured molecules that are assayed for genotype can be increased by increasing the length of the sequencing read. Complete sequencing of all molecules ensures collection of the maximum amount of data available in the enrichment pool. However, sequencing is currently expensive, and a method that can measure allele distributions using a smaller number of sequence reads would be of great value. In addition, there are technical limitations on the maximum length of a read that can be performed, and there are limitations on accuracy as the read length increases. Maximum utility alleles are one to a few bases in length, but theoretically any allele shorter than the length of the sequencing read can be used. Allelic variation occurs in all types, but the examples provided herein focus on SNPs or variants that contain only a few adjacent base pairs. Larger variants, such as segmental copy number variants, can often be detected by a collection of these smaller variants, since the entire collection of SNPs within a segment is replicated. Variants larger than a few bases, such as STRs, require special consideration, and some targeting approaches work but others do not.
[0344] There are several targeting approaches that can be used to specifically isolate and enrich one or more variant positions in a genome. Typically, these rely on utilizing invariant sequences adjacent to the variant sequence. There is prior art on targeting in the context of sequencing where the substrate is maternal plasma (see, for example, Liao et al., Clin. Chem. 2011;57(1):pp.92-101). However, all of the approaches in the prior art use targeting probes that target exons and do not focus on targeting polymorphic regions of the genome. In one embodiment, the method of the present disclosure includes using a targeting probe that focuses exclusively or nearly exclusively on the polymorphic region. In one embodiment, the method of the present disclosure includes using a targeting probe that focuses exclusively or nearly exclusively on the SNP. In some embodiments of the present disclosure, the target polymorphic sites consist of at least 10% SNPs, at least 20% SNPs, at least 30% SNPs, at least 40% SNPs, at least 50% SNPs, at least 60% SNPs, at least 70% SNPs, at least 80% SNPs, at least 90% SNPs, at least 95% SNPs, at least 98% SNPs, at least 99% SNPs, at least 99.9% SNPs, or exclusively SNPs.
[0345] In one embodiment, the disclosed method can be used to determine genotypes (base composition of DNA at specific loci) and relative proportions of these genotypes from a mixture of DNA molecules, which may be derived from one or more genetically distinct individuals. In one embodiment, the disclosed method can be used to determine genotypes at a set of polymorphic loci and the relative ratios of the amounts of different alleles present at those loci. In one embodiment, the polymorphic loci can consist entirely of SNPs. In one embodiment, the polymorphic loci can include SNPs, single tandem repeats, and other polymorphisms. In one embodiment, the disclosed method can be used to determine the relative distribution of alleles at a set of polymorphic loci in a mixture of DNA, the mixture of DNA including DNA derived from the mother and DNA derived from the fetus. In one embodiment, the distribution of combined alleles can be determined on a mixture of DNA isolated from the blood of a pregnant woman. In one embodiment, the distribution of alleles at a set of loci can be used to determine the ploidy state of one or more chromosomes of a pregnant fetus.
[0346] In one embodiment, the mixture of DNA molecules may be derived from DNA extracted from multiple cells of one individual. In one embodiment, the original collection of cells from which the DNA is derived may contain a mixture of diploid or haploid cells of the same or different genotypes if the individual is mosaic (germline or somatic). In one embodiment, the mixture of DNA molecules may also be derived from DNA extracted from a single cell. In one embodiment, the mixture of DNA molecules may also be derived from DNA extracted from a mixture of two or more cells of the same or different individuals. In one embodiment, the mixture of DNA molecules may be derived from DNA isolated from biological material that is already free from cells, such as blood plasma, which is known to contain cell-free DNA. In one embodiment, this biological material may be a mixture of DNA from one or more individuals, as is the case during pregnancy, where fetal DNA has been shown to be present in the mixture. In one embodiment, the biological material may be from a mixture of cells found in maternal blood, some of the cells being of fetal origin. In one embodiment, the biological material may be cells from blood during pregnancy enriched for fetal cells.
[0347] Circularization probe Some embodiments of the present disclosure involve the use of "ligated inversion probes" (LIPs) previously described in the literature. LIP is a general term meant to encompass techniques involving the creation of circular molecules of DNA, where the probes are designed to hybridize to target regions of DNA on either side of a target allele such that the addition of an appropriate polymerase and / or ligase, and appropriate conditions, buffers, and other reagents, completes complementary inversion regions of DNA spanning the target allele to create a circular loop of DNA that captures the information found within the target allele. LIPs may also be referred to as pre-circularized probes, pre-circularized probes, or circularized probes. LIP probes may be linear DNA molecules between 50-500 nucleotides in length, and in one embodiment 70-100 nucleotides in length, and in some embodiments may be longer or shorter than described herein. Other embodiments of the present disclosure include different incarnations of the LIP technology, such as padlock probes and molecular inversion probes (MIPs).
[0348] One method of targeting specific locations for sequencing is to synthesize a probe in which the 3' and 5' ends of the probe anneal to the target DNA in an inverted manner at positions adjacent to and on either side of the target region, such that the addition of DNA polymerase and DNA ligase results in extension from the 3' end, adding bases to the single-stranded probe complementary to the target molecule (gap fill), followed by ligation of the new 3' end to the 5' end of the original probe, resulting in a circular DNA molecule that can then be isolated from background DNA. The probe ends are designed to flank the target region of interest. One aspect of this approach, commonly referred to as MIPS, has been used in conjunction with array technology to determine the nature of the sequence written. One drawback to the use of MIPs in the context of measuring allele ratios is that the hybridization, circularization, and amplification steps do not occur at equal rates for different alleles at the same locus. This results in measurements of allele ratios that are not representative of the actual allele ratios present in the original mixture.
[0349] In one embodiment, the circularization probe is constructed such that the region of the probe designed to hybridize upstream of the target polymorphic locus and the region of the probe designed to hybridize downstream of the target polymorphic locus are covalently connected via a non-nucleic acid backbone. This backbone can be any biocompatible molecule or combination of biocompatible molecules. Some examples of possible biocompatible molecules are poly(ethylene glycol), polycarbonate, polyurethane, polyethylene, polypropylene, sulfone polymers, silicone, cellulose, fluoropolymers, acrylic compounds, styrene block copolymers, and other block copolymers.
[0350] In one embodiment of the present disclosure, this approach is modified to be easily amenable to sequencing as a means of examining the filled sequence. In order to preserve the original allele ratio of the original sample, at least one important consideration must be taken into account. The variable position between different alleles in the gap-fill region should not be too close to the probe binding site, since there may be an initiation bias by DNA polymerase that results in variant differences. Another consideration is that there may be additional variations at the probe binding site that correlate with the variants in the gap-fill region, which may result in unequal amplification from different alleles. In one embodiment of the present disclosure, the 3' and 5' ends of the pre-circularized probe are designed to hybridize to bases that are one or several positions away from the variant position (polymorphic site) of the target allele. The number of bases between the polymorphic site (SNP or other) and the bases designed to hybridize to the 3' and 5' ends of the pre-circularized probe may be 1 base, 2 bases, 3 bases, 4 bases, 5 bases, 6 bases, 7-10 bases, 11-15 bases, or 16-20 bases, 20-30 base pairs, or 30-60 base pairs. The forward and reverse primers may be designed to hybridize different numbers of bases away from the polymorphic site. Circularized probes can be produced in large quantities with current DNA synthesis technology, allowing very large amounts of probes to be produced and potentially pooled, allowing many loci to be interrogated simultaneously. This has been reported to work with over 300,000 probes. Two articles that discuss methods involving circularized probes that can be used to measure genomic data of target individuals include: Porreca et al., Nature Methods, 2007 4(11), pp. 931-936., and Turner et al., Nature Methods, 2009, 6(5), pp. 315-316. The methods described in these articles may be used in combination with other methods described herein.Certain steps of the methods from these two articles may be used in combination with other steps from other methods described herein.
[0351] In some embodiments of the methods disclosed herein, the genetic material of the target individual is optionally amplified, followed by hybridization of the pre-circularized probe, performing gap-fill to fill the bases between the two ends of the hybridized probe, ligating the two ends to form a circularized probe, and amplifying the circularized probe, for example, using rolling circle amplification. Once the desired target allele information is captured by circularizing a properly designed oligonucleic acid probe, such as in a LIP system, the genetic sequence of the circularized probe may be measured to obtain the desired sequence data. In one embodiment, a properly designed oligonucleotide probe may be directly circularized on the non-amplified genetic material of the target individual and then amplified. It is noted that several amplification procedures may be used to amplify the original genetic material or the circularized LIP, including rolling circle amplification, MDA, or other amplification protocols. Different methods may be used to measure the genetic information on the target genome, for example, using high-throughput sequencing, Sanger sequencing, other sequencing methods, capture by hybridization, capture by circularization, multiplex PCR, other hybridization methods, and combinations thereof.
[0352] Once an individual's genetic material has been measured using one or a combination of the above methods, then an informatics-based method such as the PARENTAL SUPPORT™ method can be used with appropriate genetic measurements to determine the ploidy state of one or more chromosomes on the individual, and / or the genetic state of one or more alleles, specifically the alleles that correlate with the disease or genetic condition of interest. It should be noted that the use of LIP has been reported for multiplex capture of genetic sequences, followed by genotyping by sequencing. However, the use of sequencing data resulting from a LIP-based strategy for amplification of genetic material found in single cells, a small number of cells, or extracellular DNA has not been used to determine the ploidy state of a target individual.
[0353] The application of informatics-based methods to determine the ploidy status of an individual from genetic data measured by hybridization arrays such as ILLUMINA INFINIUM arrays or AFFYMETRIX gene chips has been described in reference documents elsewhere in this document. However, the methods described herein represent an improvement over methods previously described in the literature. For example, LIP-based approaches followed by high-throughput sequencing unexpectedly provide better genotype data due to the approach having better capacity for multiplexing, better capture specificity, better uniformity, and lower allele bias. Greater multiplexing allows more alleles to be targeted, giving more accurate results. With better uniformity, more target alleles are measured, giving more accurate results. A lower rate of allele bias results in a lower rate of miscalls, giving more accurate results. More accurate results lead to improved clinical outcomes and better medical care.
[0354] It is important to note that LIP can be used as a method to target specific loci in a sample of DNA for genotyping by methods other than sequencing. For example, LIP may be used to target DNA for genotyping using SNP arrays or other DNA or RNA based microarrays.
[0355] Ligation-mediated PCR Ligation-mediated PCR is a method of PCR used to preferentially enrich a sample of DNA by amplifying one or more loci in a mixture of DNA, the method comprising: obtaining a set of primer pairs, each primer of the pair containing a target specific sequence and a non-target sequence, the target specific sequence designed to anneal to a target region one upstream and one downstream of a polymorphic site, the set of primers being 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11-20, 21-30, 31-40, 32-46, 33-48, 34-50, 35-51, 36-52, 37-53, 38-54, 39-60, 40-46, 41-48, 42-46, 43-49, 44-55, 45-56, 46-57, 47-58, 48-59, 49-60, 50-61, 51-62, 52-63, 53-64, 54-65, 55-66, 56-70, 57-71, 58-72, 59-81, 62-82, 63-83, 64-85, 65-86, 66-87, 67-88, 68-90, 69-100, 69-101, 69-102, 69-103, 69-104, 69-105, 69-106, 69-107, 69-108, 69-109, 70-1109, 71-111, 72-112, 73-113, 74 obtaining a target molecule that can be separated by 41-50, 51-100, or more than 100; polymerizing DNA from the 3-prime end of the upstream primer to fill in the single-stranded region between it and the 5-prime end of the downstream primer with nucleotides complementary to the target molecule; ligating the last polymerized base of the upstream primer to the adjacent 5-prime base of the downstream primer; and amplifying only the polymerized and ligated molecule using non-target sequences contained in the 5-prime end of the upstream primer and the 3-prime end of the downstream primer. Pairs of primers for different targets may be mixed in the same reaction. The non-target sequences act as universal sequences so that all pairs of primers that are successfully polymerized and ligated can be amplified with a single amplification primer pair.
[0356] Capture by hybridization Preferential enrichment of a particular set of sequences in a target genome can be achieved in several ways. Elsewhere in this specification, there are descriptions of how LIP can be used to target a particular set of sequences, but in all of these applications, other targeting and / or preferential enrichment methods can be used equally well for the same purpose. One example of another targeting method is the capture by hybridization approach. Some examples of commercial capture by hybridization technology include AGILENT's SURE SELECT and ILLUMINA's TruSeq. In capture by hybridization, a set of oligonucleotides that are complementary or nearly complementary to the desired targeting sequence are hybridized to a mixture of DNA and then physically separated from the mixture. Once the desired sequence is hybridized to the targeting oligonucleotide, the effect of physically removing the targeting oligonucleotide is to also remove the target sequence. Once the hybridized oligos are removed, they can be heated above their melting temperature and amplified. Some methods of physically removing the targeting oligonucleotide are to covalently attach the targeting oligonucleotide to a solid support, such as a magnetic bead or a chip. Another way to physically remove targeting oligonucleotide is to covalently bind a molecular moiety that has strong affinity to another molecular moiety.One example of such a molecular pair is biotin and streptavidin, as used in SURE SELECT.Therefore, target sequence can be covalently bound to biotin molecule, and after hybridization, the biotinylated oligonucleotide that hybridizes to target sequence can be pulled down using a solid support that has streptavidin attached thereto.
[0357] Hybrid capture involves hybridization of a probe complementary to the target of interest to a target molecule. Hybrid capture probes were originally developed to target and enrich for large fractions of genomes with relative uniformity between targets. In that application, it was important that all targets were amplified with enough uniformity that all regions could be detected by sequencing, but no consideration was given to preserving the proportion of alleles in the original sample. After capture, the alleles present in the sample can be determined by direct sequencing of the captured molecules. These sequencing reads may be analyzed and counted according to allele type. However, using current technology, the measured allele distribution of the captured sequences typically does not represent the original allele distribution.
[0358] In one embodiment, the detection of alleles is performed by sequencing. To capture allele identity at a polymorphic site, it is essential that the sequencing read spans the allele in question to assess the allele composition of the captured molecule. The captured molecule is often of variable length when sequenced, so overlap with the variant position cannot be guaranteed unless the entire molecule is sequenced. However, cost considerations and technical limitations on the maximum possible length and accuracy of the sequencing read make sequencing of the entire molecule infeasible. In one embodiment, the read length can be increased from about 30 bases to about 50 bases or about 70 bases, greatly increasing the number of reads that overlap with variant positions in the target sequence.
[0359] Another way to increase the number of reads interrogating a position of interest is to reduce the length of the probe, as long as it does not introduce bias in the underlying enriched allele. The length of the synthesized probe should be long enough so that two probes designed to hybridize to two different alleles found at one locus will hybridize with approximately equal affinity to the various alleles in the original sample. Currently, methods known in the art typically describe probes longer than 120 bases. In the current embodiment, when an allele is one or more bases, the capture probe may be less than about 110 bases, less than about 100 bases, less than about 90 bases, less than about 80 bases, less than about 70 bases, less than about 60 bases, less than about 50 bases, less than about 40 bases, less than about 30 bases, and less than about 25 bases, which is sufficient to ensure equal enrichment from all alleles. When the mixture of DNA to be enriched using hybrid capture technology is a mixture containing free floating DNA isolated from blood, e.g., maternal blood, the average length of the DNA is very short, typically less than 200 bases. The use of shorter probes increases the likelihood that the hybrid capture probe will capture the desired DNA fragment. Larger variations may require longer probes. In one embodiment, the variation of interest is one (SNP) to several bases in length. In one embodiment, the target region in the genome can be preferentially enriched using hybrid capture probes, which are less than 90 bases in length, and may be less than 80 bases, less than 70 bases, less than 60 bases, less than 50 bases, less than 40 bases, less than 30 bases, or less than 25 bases. In one embodiment, to increase the likelihood that the desired allele is sequenced, the length of the probe designed to hybridize to the region adjacent to the polymorphic allele position can be reduced from more than 90 bases to about 80 bases, or about 70 bases, or about 60 bases, or about 50 bases, or about 40 bases, or about 30 bases, or about 25 bases.
[0360] To allow capture, there is a minimum overlap between the synthesized probe and the target molecule. This synthesized probe can be as short as possible while still being greater than this minimum required overlap. The effect of using a shorter probe length to target a polymorphic region is that there are more molecules that overlap with the target allele region. The fragmentation state of the original DNA molecule also influences the number of reads that overlap with the target allele. Some DNA samples, such as plasma samples, are already fragmented due to biological processes that take place in vivo. However, samples with longer fragments benefit from fragmentation prior to sequencing library preparation and enrichment. If both the probe and fragments are short (approximately 60-80 bp), maximum specificity can be achieved with a relatively small number of sequence reads that fail to overlap the critical region of interest.
[0361] In one embodiment, the hybridization conditions can be adjusted to maximize the uniformity in the capture of different alleles present in the original sample. In one embodiment, the hybridization temperature is reduced to minimize the difference in hybridization bias between alleles. Methods known in the art avoid using lower temperatures for hybridization because lowering the temperature has the effect of increasing hybridization of the probe to unintended targets. However, if the goal is to maintain the allele ratio with maximum fidelity, an approach using lower hybridization temperatures provides the most accurate allele ratio, despite the fact that current technology has moved away from this approach. Hybridization temperatures can also be increased to require a greater overlap between the target and the synthetic probe, so that only targets with substantial overlap of the target region are captured. In some embodiments of the present disclosure, the hybridization temperature is reduced from the normal hybridization temperature to about 40°C, about 45°C, about 50°C, about 55°C, about 60°C, about 65°C, or about 70°C.
[0362] In one embodiment, a hybrid capture probe can be designed such that the region of the capture probe that has DNA complementary to the DNA found in the region adjacent to the polymorphic allele is not directly adjacent to the polymorphic site. Instead, the capture probe can be designed such that the region of the capture probe that is designed to hybridize to the DNA adjacent to the polymorphic site of the target is separated from the portion of the capture probe that would be in van der Waals contact with the polymorphic site by a length equal to one or a small number of bases. In one embodiment, a hybrid capture probe is designed to hybridize to a region adjacent to but not across the polymorphic allele, which can be called an adjacent capture probe. The length of the adjacent capture probe can be less than about 120 bases, less than about 110 bases, less than about 100 bases, less than about 90 bases, less than about 80 bases, less than about 70 bases, less than about 60 bases, less than about 50 bases, less than about 40 bases, less than about 30 bases, and less than about 25 bases. The regions of the genome targeted by adjacent capture probes may be separated by polymorphic loci by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 to 20, or more than 20 base pairs.
[0363] Description of targeted capture-based disease screening test using targeted sequence capture. Custom targeted sequence capture such as currently offered by AGILENT (SURE SELECT), ROCHE-NIMBLEGEN, or ILLUMINA. Capture probes can be custom designed to ensure capture of various types of mutations. For point mutations, one or more probes overlapping the point mutation should be sufficient to capture and sequence the mutation.
[0364] For small insertions or deletions, one or more probes overlapping the mutation may be sufficient to capture and sequence the fragment containing the mutation. Hybridization may be less efficient, typically between probe-limited capture efficiencies designed against a reference genome sequence. To ensure capture of the fragment containing the mutation, two probes can be designed, one matching the normal allele and one matching the mutant allele. Longer probes may enhance hybridization. Multiple overlapping probes may enhance capture. Finally, by having probes that are not overlapping but immediately adjacent, the mutation may allow relatively similar capture efficiencies of normal and mutant alleles.
[0365] In the case of simple tandem repeats (STRs), probes that overlap these highly variable sites are unlikely to capture the fragments adequately. To enhance capture, probes can flank, but not overlap, the variable sites. The fragments can then be sequenced as usual to reveal the length and composition of the SRTs.
[0366] In the case of large deletions, a common approach currently used in exome capture systems may work, which is a series of overlapping probes. However, this approach may make it difficult to determine whether an individual is heterozygous. Targeting and evaluating SNPs within the captured region may potentially reveal loss of heterozygosity across the region, which indicates that the individual is a carrier. In one embodiment, it is possible to place non-overlapping or singleton probes across the potentially deleted region and use the number of fragments captured as a measure of heterozygosity. When an individual has a large deletion, half the number of fragments is expected to be available for capture compared to a non-deleted (diploid) reference locus. Thus, the number of reads obtained from the deleted region should be approximately half the number of reads obtained from a normal diploid locus. Aggregating and averaging the depth of sequencing reads from multiple singleton probes across a potentially deleted region can strengthen the signal and improve the reliability of the diagnosis. The two approaches of targeting SNPs to identify loss of heterozygosity and using multiple singleton probes to obtain a quantitative measure of the amount of underlying fragments from the locus can also be combined. Either or both of these strategies may be combined with other strategies to better achieve the same goal.
[0367] During testing for male fetal cfDNA detection, either an X-linked dominant mutation in which the mother and father are unaffected, or a dominant mutation in which the mother is unaffected, captured and sequenced in the same test, indicates high risk to the fetus, as indicated by the presence of a Y chromosome fragment. Detection of two mutant recessive alleles in the same gene in an unaffected mother means that the fetus has inherited a mutant allele from the father and potentially a second mutant allele from the mother. In all cases, follow-up testing with amniocentesis or chorionic villus sampling may be indicated.
[0368] The targeted capture-based disease screening test can be combined with a targeted capture-based non-invasive prenatal diagnostic test for aneuploidy.
[0369] There are several ways to reduce the variability of depth of read (DOR): for example, primer concentration can be increased, longer targeted amplification probes can be used, or more STA cycles can be performed (such as more than 25, more than 30, more than 35, or even more than 40).
[0370] Targeted PCR In some embodiments, PCR can be used to target specific locations in the genome. In plasma samples, the original DNA is highly fragmented (typically less than 500 bp, with an average length of less than 200 bp). In PCR, both forward and reverse primers must anneal to the same fragment to allow amplification. Therefore, if the fragment is short, the PCR assay must amplify a relatively short region as well. As with MIPS, bias in amplification from different alleles may occur if the polymorphic position is too close to the polymerase binding site. Currently, PCR primers targeting polymorphic regions (such as those containing SNPs) are typically designed such that the 3' end of the primer hybridizes to a base immediately adjacent to one or more polymorphic bases. In one embodiment of the present disclosure, the 3' end of both the forward and reverse PCR primers are designed to hybridize to a base that is one or several positions away from the variant position (polymorphic site) of the target allele. The number of bases between the polymorphic site (SNP or other) and the base designed to hybridize to the 3' end of the primer may be 1 base, 2 bases, 3 bases, 4 bases, 5 bases, 6 bases, 7 to 10 bases, 11 to 15 bases, or 16 to 20 bases. The forward and reverse primers may be designed to hybridize a different number of bases away from the polymorphic site.
[0371] PCR assays can be made in large quantities, but interactions between different PCR assays make it difficult to multiplex beyond about 100 assays. A variety of complex molecular techniques can be used to increase the level of multiplexing, but it will still be limited to less than 100, perhaps 200, or perhaps 500 assays per reaction. Samples containing large amounts of DNA can be split into multiple sub-reactions and then recombined before sequencing. For samples in which either the entire sample or some subset of DNA is limited, splitting the sample will introduce statistical noise. In one embodiment, a small or limited amount of DNA may refer to an amount of less than 10 pg, 10-100 pg, 100 pg-1 ng, 1-10 ng, or 10-100 ng. It should be noted that this method is particularly useful for small amounts of DNA, where other methods involving splitting into multiple pools can have significant problems related to the statistical noise introduced, but the method still offers the advantage of minimizing bias when performed on samples of any amount of DNA. In these situations, a universal pre-amplification step may be used to increase the overall sample amount. Ideally, this pre-amplification step should not significantly alter the allele distribution.
[0372] In one embodiment, the disclosed method can generate PCR products specific to a large number of target loci, specifically 1,000-5,000 loci, 5,000-10,000 loci, or more than 10,000 loci, for sequencing or genotyping by some other genotyping method from a limited sample, such as DNA from a single cell or body fluid. Currently, performing multiplex PCR reactions of more than 5-10 targets is a significant challenge and is often hindered by primer by-products (e.g., primer dimers) and other artifacts. When detecting target sequences using microarrays using hybridization probes, primer dimers and other artifacts may be ignored because they are not detected. However, when using sequencing as a detection method, the majority of the sequencing reads will sequence such artifacts and not sequence the desired target sequences in the sample. Methods described in the prior art that are used to multiplex and subsequently sequence more than 50 or 100 reactions in a single reaction typically yield more than 20%, often more than 50%, often more than 80%, and in some cases more than 90% of the sequence reads that are not targeted.
[0373] Typically, to perform targeted sequencing of multiple (n) (>50, >100, >500, or >1,000) targets of a sample, the sample can be split into several parallel reactions amplifying one individual target. This can be done in PCR multiwell plates or on commercially available platforms such as FLUIDIGM ACCESS ARRAY (48 reactions per sample in a microfluidic chip) or DROPLET PCR (100s to thousands of targets) from RAIN DANCE TECHNOLOGY. Unfortunately, these pooling and splitting methods are problematic because in samples containing limited amounts of DNA, there are often not enough copies of the genome to ensure that there is one copy of each region of the genome in each well. This is particularly serious when polymorphic loci are targeted and the relative proportions of alleles at the polymorphic loci are required, as the statistical noise introduced by pooling and splitting makes the measurement of the proportion of alleles present in the original sample of DNA very inaccurate. Described herein is a method for effectively and efficiently amplifying many PCR reactions that is applicable when only limited amounts of DNA are available. In one embodiment, the method will be applicable to the analysis of single cells, body fluids, mixtures of DNA (e.g., free floating DNA found in maternal plasma), biopsies, environmental and / or forensic samples.
[0374] In one embodiment, targeted sequencing may involve one, several, or all of the following steps: a) Generate and amplify libraries with adapter sequences at both ends of DNA fragments; b) After library amplification, split into multiple reactions; c) Generate and optionally amplify libraries with adapter sequences at both ends of DNA fragments; d) Perform 1000-10,000-plex amplification of selected targets using one target-specific "forward" primer and one tag-specific primer; e) Perform a second amplification using a "reverse" target-specific primer and one (or more) primers specific to the universal tag introduced as part of the target-specific forward primer in the first round; f) Perform 1000-plex preamplification of selected targets for a limited number of cycles; g) Split the product into multiple aliquots and amplify subpools of targets in individual reactions (e.g., 50-500-plex, which can be used up to singleplex); h) Pool products of parallel subpool reactions. i) During these amplifications, the primers can carry tags (partial or full length) compatible with sequencing so that the products can be sequenced.
[0375] Highly multiplexed PCR Disclosed herein is a method that allows targeted amplification across hundreds to thousands of target sequences (e.g., SNP loci) from a nucleic acid sample, such as genomic DNA obtained from plasma. The amplified sample is relatively free of primer-dimer products and may have low allele bias at the target loci. If during or after amplification, the products are tagged with adapters compatible with sequencing, analysis of these products can be performed by sequencing.
[0376] Highly multiplexed PCR amplification using methods known in the art produces primer-dimer products that are not suitable for sequencing in excess of the desired amplification products. These can be empirically reduced by excluding primers that form these products or by performing in silico selection of primers. However, the larger the number of assays, the more difficult this problem becomes.
[0377] One solution is to split the 5000-fold reaction into several smaller amplifications (e.g., 100 50-fold reactions or 50 100-fold reactions), or to use microfluidics, or to further split the sample into individual PCR reactions. However, when sample DNA is limited, such as non-invasive prenatal testing from pregnant plasma, splitting the sample into multiple reactions is hindered and should be avoided.
[0378] Described herein is a method for first globally amplifying the plasma DNA of a sample and then splitting the sample into multiple multiplexed target enrichment reactions with a more moderate number of target sequences per reaction. In one embodiment, the disclosed method can be used to preferentially enrich a DNA mixture at multiple loci, and includes one or more of the following steps: creating and amplifying a library from a mixture of DNA where the molecules in the library have adapter sequences ligated to both ends of the DNA fragments; splitting the amplified library into multiple reactions; and performing a first round of multiplex amplification of selected targets using one target-specific "forward" primer and one or more adapter-specific universal "reverse" primers. In one embodiment, the disclosed method further includes performing a second round of amplification using a "reverse" target-specific primer and one or more primers specific to the universal tag introduced as part of the target-specific forward primer in the first round. In one embodiment, the disclosed method may involve a fully nested, heminested, semi-nested, one-sided fully nested, one-sided heminested, or one-sided semi-nested PCR approach. In one embodiment, the disclosed method is used to preferentially enrich a DNA mixture at multiple loci, the method comprising performing multiplex preamplification of selected targets for a limited number of cycles, dividing the products into multiple aliquots, amplifying subpools of targets in individual reactions, and pooling the products of the parallel subpool reactions. It is noted that this approach can be used to perform targeted amplification for 50-500 loci, for 500-5,000 loci, for 5,000-50,000 loci, or even for 50,000-500,000 loci in a manner that produces low levels of allelic bias. In one embodiment, the primers have tags compatible with partial or full-length sequencing.
[0379] The workflow may include (1) extracting plasma DNA, (2) preparing a fragment library with universal adapters on both ends of the fragments, (3) amplifying the library with a universal primer specific to the adapters, (4) splitting the amplified sample "library" into multiple aliquots, (5) performing multiplex amplifications of the aliquots (e.g., about 100, 1,000, or 10,000 reactions with one target-specific primer and tag-specific primer per target), (6) pooling aliquots of one sample, (7) barcoding the samples, (8) mixing the samples and adjusting the concentration, and (9) sequencing the samples. The workflow may include multiple substeps including one of the steps listed (e.g., step (2), which is the step of preparing the library, may involve three enzymatic steps (blunting, dA tailing, and adapter ligation) and three purification steps). Workflow steps may be combined, split, or performed in a different order (e.g., barcoding and pooling of samples).
[0380] Amplification of the library may be performed in a manner biased to amplify short fragments more efficiently. In this manner, shorter sequences, e.g., mononucleosomal DNA fragments, may be preferentially amplified as cell-free fetal DNA (from the placenta) found in the circulation of pregnant women. It should be noted that the PCR assay may have a tag, e.g., a sequencing tag (usually a truncated form of 15-25 bases). After multiplexing, the PCR multiplexes of the samples are pooled, and then the tags are completed (including barcoding) by tag-specific PCR (which can also be done by ligation). Also, the entire sequencing tag may be added to the same reaction as the multiplexing. In the first cycle, the target may be amplified using a target-specific primer, and then the tag-specific primer may be taken over to complete the SQ adapter sequence. The PCR primer may not have a tag. The sequencing tag may be attached to the amplification product by ligation.
[0381] In one embodiment, highly multiplex PCR followed by evaluation of the amplified material by clonal sequencing may be used to detect fetal aneuploidy. Conventional multiplex PCR evaluates up to 50 loci simultaneously, whereas the techniques described herein can be used to simultaneously evaluate more than 50 loci, more than 100 loci, more than 500 loci, more than 1,000 loci, more than 5,000 loci, more than 10,000 loci, more than 50,000 loci, more than 100,000 loci. Experiments show that up to, including, more than 10,000 distinct loci can be simultaneously evaluated in a single reaction with sufficiently good efficiency and specificity to perform non-invasive prenatal aneuploidy diagnosis and / or highly accurate copy number calling. The assay may be combined in a single reaction using the entire cfDNA sample isolated from maternal plasma, a fraction thereof, or a further processed derivative of the cfDNA sample. The cfDNA or derivative may also be split into multiple parallel multiplex reactions. Optimal sample splitting and multiplexing is determined by trading off various performance specifications. Due to the limited amount of material, splitting the sample into multiple fractions may introduce sampling noise, handling time, and increase the chance of error. Conversely, a higher degree of multiplexing may result in more false amplifications and greater inequality in amplification, both of which may reduce test performance.
[0382] Two important relevant considerations in the application of the methods described herein are the limited amount of original plasma and the number of original molecules in this material from which allele frequencies or other measurements are obtained. If the number of original molecules is below a certain value, random sampling noise may become significant and affect the accuracy of the test. Typically, if measurements are performed on samples containing the equivalent of 500-1000 original molecules...
Claims
**Claim 1** A method for preparing a preparation of amplified DNA derived from a first blood sample or a fraction thereof of a pregnant woman useful for identifying a pregnancy having a high risk of preterm birth, preeclampsia, fetal growth retardation, spontaneous abortion, and / or non-viable birth, comprising: a) extracting cell-free DNA from the first blood sample or a fraction thereof to obtain a first extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; b) preparing a first preparation of amplified DNA by performing targeted multiplex amplification on the first extracted DNA to amplify 200 to 20,000 SNP loci in a single reaction volume to obtain amplified DNA, wherein the 200 to 20,000 SNP loci are located on one or more target chromosomes; c) analyzing the first preparation of the amplified DNA by performing high-throughput sequencing on the amplified DNA to obtain sequence reads and using the sequence reads to determine the ploidy status of the one or more target chromosomes, wherein a fetal fraction of less than 2.8% and / or no call of the ploidy status of the one or more target chromosomes indicates a pregnancy having a high risk of preterm birth, preeclampsia, fetal growth retardation, spontaneous abortion, and / or non-viable birth. **Claim 2** d) extracting cell-free DNA from a second blood sample or a fraction thereof of the pregnant woman collected over a long period to obtain a second extracted DNA comprising maternal cell-free DNA and fetal cell-free DNA; e) preparing a second preparation of amplified DNA by performing targeted multiplex amplification on the second extracted DNA to amplify the 200 to 20,000 SNP loci in a single reaction volume to obtain amplified DNA, wherein the 200 to 20,000 SNP loci are located on one or more target chromosomes; (f) subjecting the second preparation of the amplified DNA to high-throughput sequencing with respect to the amplified DNA to obtain sequence reads, and analyzing by using the sequence reads to determine the ploidy status of the one or more target chromosomes, wherein for each of the first and second blood samples, a fetal fraction of less than 2.8% and / or no call of the ploidy status of the one or more target chromosomes further indicates a pregnancy having a high risk of preterm birth, preeclampsia, fetal growth retardation, spontaneous abortion, and / or stillbirth, the method according to claim 1, further comprising analyzing. **Claim 3** The method according to claim 2, further comprising identifying, for each of the first and second blood samples, a pregnant woman with no call of the ploidy status of the one or more target chromosomes as having a risk of at least 40% of preterm birth before 37 weeks, preeclampsia, and / or fetal growth retardation. **Claim 4** The method according to claim 2, further comprising identifying, for each of the first and second blood samples, a pregnant woman with no call of the ploidy status of the one or more target chromosomes as having a risk of at least 50% of preterm birth before 37 weeks, preeclampsia, and / or fetal growth retardation. **Claim 5** The method according to claim 2, further comprising identifying, for each of the first and second blood samples, a pregnant woman with no call of the ploidy status of the one or more target chromosomes as having a risk of at least 15% of preeclampsia. **Claim 6** The method according to claim 2, further comprising identifying, for each of the first and second blood samples, a pregnant woman with no call of the ploidy status of the one or more target chromosomes as having a risk of at least 20% of preterm birth before 28 weeks. **Claim 7** The method according to claim 2, further comprising identifying, for each of the first and second blood samples, a pregnant woman with no call of the ploidy status of the one or more target chromosomes as having a risk of at least 25% of preterm birth before 34 weeks. **Claim 8** The method according to claim 2, further comprising identifying, for each of the first and second blood samples, a pregnant woman with no call of the ploidy status of the one or more target chromosomes as having a risk of at least 40% of preterm birth before 37 weeks. **Claim 9** The method according to claim 2, further comprising, for each of the first and second blood samples, identifying a pregnant woman without a call of the ploidy state of the one or more target chromosomes as having a risk of fetal growth retardation of at least 10%.
10. The method according to claim 2, further comprising, for each of the first and second blood samples, identifying a pregnant woman having a fetal fraction of less than 2.5%.
11. The method according to claim 2, further comprising repeating steps (d) to (f) for a third blood sample or a fraction thereof collected over a long period of time.
12. The method according to claim 1, wherein step (a) comprises extracting cell-free DNA from the plasma fraction of the blood sample.
13. The method according to claim 1, wherein step (b) comprises PCR amplification of 200 to 20,000 SNP loci using 200 to 20,000 pairs of target-specific PCR primers, or using universal primers and 200 to 20,000 target-specific primers.
14. The method according to claim 1, wherein step (b) comprises PCR amplification of 1,000 to 20,000 SNP loci using 1,000 to 20,000 pairs of target-specific PCR primers, or using universal primers and 1,000 to 20,000 target-specific primers.
15. The method according to claim 1, wherein step (b) comprises PCR amplification of 5,000 to 20,000 SNP loci using 5,000 to 20,000 pairs of target-specific PCR primers, or using universal primers and 5,000 to 20,000 target-specific primers.
16. The method according to claim 1, wherein the amplified DNA in step (b) each comprises 100 bp or less amplified from the extracted DNA.
17. The method according to claim 1, wherein the amplified DNA in step (b) each comprises 80 bp or less amplified from the extracted DNA.
18. The method according to claim 1, wherein the amplified DNA in step (b) each comprises 60 to 80 bp amplified from the extracted DNA.
19. The method according to claim 1, wherein step (b) further comprises barcode PCR after the targeted multiplex amplification.
20. The method according to claim 1, wherein the ploidy state of the one or more target chromosomes is determined by calculating the number of alleles at the SNP locus based on the sequence reads, creating a plurality of ploidy hypotheses each associated with a different possible ploidy state of the target chromosome, constructing a combined distribution model of the predicted number of alleles at the SNP locus on the target chromosome for each ploidy hypothesis, using the combined distribution model and the number of alleles to determine the relative probability of each of the ploidy hypotheses, and selecting the ploidy state corresponding to the hypothesis having the highest probability to call the ploidy state of the fetus.