Diagnosis of Fetal Chromosomal Aneuploidy Using Genomic Sequencing
By sequencing biological samples of pregnant women and determining fetal chromosome aneuploidy, the problems of insufficient accuracy and maternal background nucleic acid interference in the prior art are solved, and a highly accurate non-invasive fetal chromosome aneuploidy diagnosis is achieved.
Patent Information
- Application Number
- CN201710198531.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2007-07-23
- Filing Date
- 2008-07-23
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2028-10-02
AI Technical Summary
The existing non-invasive prenatal diagnosis methods detect fetal chromosomal aneuploidy, especially trisomy 21, have problems with insufficient accuracy and interference from maternal background nucleic acids, making it difficult to effectively distinguish between fetal and maternal nucleic acids, resulting in a high rate of false negative and false positive.
By sequencing nucleic acid molecules in biological samples of pregnant women, the amount of clinically relevant chromosomes and background chromosomes is determined, and the parameter ratio is used to compare with the cutoff value to achieve an accurate diagnosis of fetal chromosome aneuploidy.
It improves the accuracy of fetal chromosomal aneuploidy diagnosis, reduces the false negative and false positive rates, and achieves non-invasive detection with high sensitivity and high specificity.
Smart Images

Figure CN107083425B_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application for invention titled "Diagnosing Fetal Chromosomal Aneuploidies Using Genomic Sequencing" with the application date of July 23, 2008, application number 200880108377.1.
[0002] Priority Statement
[0003] This application claims the priority of U.S. Provisional Application No. 60 / 951438, filed on July 23, 2007, titled "DETERMINING A NUCLEIC ACID SEQUENCE IMBALANCE" (Attorney Docket No. 016285 - 005200US), and is a regular application thereof. The entire content of this provisional application is hereby incorporated by reference and used for all purposes.
[0004] Cross - Reference to Related Applications
[0005] This application also relates to a regular application titled "DETERMINING A NUCLEIC ACID SEQUENCE IMBALANCE" (Attorney Docket No. 016285 - 005210US) filed simultaneously herewith. The entire content of this application is hereby incorporated by reference and used for all purposes. Field of the Invention
[0006] The present invention generally relates to diagnosing and detecting fetal chromosomal aneuploidies by determining an imbalance between different nucleic acid sequences, and more particularly, to determining trisomy 21 (Down syndrome) and other chromosomal aneuploidies via the detection of a maternal sample (such as blood). Background of the Invention
[0007] Fetal chromosomal aneuploidies are caused by the presence of an abnormal dosage of chromosomes or chromosomal regions. The abnormal dosage can be abnormally high, such as the presence of an extra copy of chromosome 21 or a chromosomal region in trisomy 21; or abnormally low, such as the lack of a copy of the X chromosome in Turner syndrome.
[0008] Conventional prenatal diagnostic methods for fetal chromosomal aneuploidies such as trisomy 21 involve sampling fetal material by invasive methods such as amniocentesis or chorionic villus sampling, but this poses a limited risk of fetal loss. Non-invasive methods, such as screening by ultrasound scanning or biochemical markers, have been used to stratify pregnant women at risk prior to definitive invasive diagnostic methods. However, these screening methods typically measure secondary phenomena related to chromosomal aneuploidies such as trisomy 21 rather than the core chromosomal abnormalities, and thus the diagnostic accuracy is not optimal and has other drawbacks such as excessive influence of gestational age.
[0009] In 1997, cell-free fetal DNA was detected in maternal plasma, which provided new possibilities for non-invasive prenatal diagnosis (Lo, YMD and Chiu, RWK 2007 Nat Rev Genet 8, 71-77). Although this method is readily applicable to prenatal diagnosis of sex-linked disorders (Costa, JM et al. 2002 N Engl J Med 346, 1502) and certain single-gene disorders (Lo, YMDet al. 1998 N Engl J Med 339, 1734-1738), the application of this method for prenatal detection of fetal chromosomal aneuploidies still represents a considerable challenge (Lo, YMD and Chiu, RWK 2007, ibid.). First, fetal nucleic acids coexist with high-background nucleic acids of maternal origin in maternal plasma, and the high-background nucleic acids of maternal origin often interfere with the analysis of fetal nucleic acids (Lo, YMD et al. 1998 Am J Hum Genet 62, 768-775). Second, fetal nucleic acids circulate mainly in a cell-free form in maternal plasma, which makes it difficult to obtain gene or chromosomal dosage information of the fetal genome.
[0010] In recent years, significant developments have been made to overcome these challenges (Benachi, A & Costa, JM 2007 Lancet369, 440-442). One method is to detect fetal-specific nucleic acids in maternal plasma, thus overcoming the problem of maternal background interference (Lo, YMD and Chiu, RWK 2007, ibid.). The dosage of chromosome 21 is inferred from the ratio of polymorphic alleles in DNA / RNA molecules of placental origin. However, when the sample contains a low amount of target nucleic acid, the accuracy of this method is low and it is only applicable to fetuses that are heterozygous for the target polymorphism, and if one polymorphism is used, the target nucleic acid is only a subset of the population.
[0011] Dhallan et al. (Dhallan, R, et al. 2007, ibid., Dhallan, R, et al. 2007 Lancet 369, 474 - 481) described an alternative strategy to enrich the proportion of fetal DNA in maternal plasma. The proportion of chromosome 21 sequences contributed by the fetus in maternal plasma was determined by assessing the ratio of the fetal - specific alleles to the non - fetal - specific alleles of single - nucleotide polymorphisms (SNPs) on chromosome 21 that were paternally inherited. Similarly, the SNP ratio of a reference chromosome was calculated. Subsequently, an imbalance of fetal chromosome 21 was inferred by detecting a statistically significant difference between the SNP ratio of chromosome 21 and the SNP ratio of the reference chromosome, where a fixed p - value of less than or equal to 0.05 was used to define significance. To ensure a high degree of population coverage, more than 500 SNPs were targeted per chromosome. However, there is debate about the efficiency of formaldehyde in enriching fetal DNA to a high proportion (Chung, GTY, et al. 2005 Clin Chem 51, 655 - 658), and thus, the reproducibility of this method needs further evaluation. Additionally, since each fetus and mother provide many different SNPs for each chromosome, the power of the statistical test for SNP ratio comparison varies from case to case (Lo YMD & Chiu, RWK. 2007 Lancet 369, 1997). Furthermore, since these methods rely on the detection of genetic polymorphisms, they are limited to fetuses that are heterozygous for these polymorphisms.
[0012] Using polymerase chain reaction (PCR) and DNA quantification of chromosome 21 loci and reference loci in amniotic fluid cell cultures obtained from trisomy 21 and euploid fetuses, Zimmermann et al. (2002 Clin Chem 48, 362 - 363) were able to distinguish between these two groups of fetuses based on a 1.5 - fold increase in the DNA sequence of chromosome 21 in amniotic fluid cell cultures of trisomy 21 fetuses. Since a 2 - fold difference in DNA template concentration constitutes only a difference of one threshold cycle (Ct), the discrimination of a 1.5 - fold difference is the limit of conventional real - time PCR. To achieve a better degree of quantitative discrimination, alternative strategies are needed.
[0013] Digital PCR for detecting allelic ratio skewing in nucleic acid samples has been developed (Chang, HW et al. 2002 J Natl Cancer Inst 94, 1697 - 1703). Digital PCR is an amplification-based nucleic acid analysis technique that requires distributing a sample containing nucleic acids among a large number of discrete samples, where each sample on average contains no more than about 1 target sequence. By digital PCR, specific nucleic acid targets are amplified with sequence-specific primers to generate specific amplicons. Prior to nucleic acid analysis, the nucleic acid locus to be targeted and the type or set of sequence-specific primers to be included in the reaction are determined or selected.
[0014] Clinically, it has been demonstrated that digital PCR can be used to detect loss of heterozygosity (LOH) in tumor DNA samples (Zhou, W. et al. 2002 Lancet 359, 219 - 225). To analyze the results of digital PCR, previous studies have used sequential probability ratio testing (SPRT) to classify the experimental results as indicating the presence or absence of LOH in the sample (ElKaroui et al. 2006 Stat Med 25, 3124 - 3133).
[0015] In the method used in previous studies, the amount of data collected by digital PCR was quite low. Thus, the small number of data points and typical statistical fluctuations compromised the accuracy.
[0016] Therefore, non-invasive tests with high sensitivity and specificity are desired to minimize false negatives and false positives, respectively. However, fetal DNA exists at low absolute concentrations and represents a minor fraction of all DNA sequences in maternal plasma and serum. Thus, methods for non-invasive detection of fetal chromosomal aneuploidy by maximizing the amount of genetic information that can be inferred from the limited number of fetal nucleic acids present as a minor fraction in a biological sample containing maternal background nucleic acids are also desired. SUMMARY OF THE INVENTION
[0017] Embodiments of the present invention provide methods, systems, and devices for determining the presence of nucleic acid sequence imbalances (such as chromosomal imbalances) in biological samples obtained from pregnant women. Such determination can be made using parameters of the amount of a clinically relevant chromosomal region relative to other non-clinically relevant chromosomal regions (background regions) in the biological sample. In one aspect, the amount of chromosomes is determined by sequencing nucleic acid molecules in a maternal sample, such as urine, plasma, serum, and other suitable biological samples. The nucleic acid molecules in the biological sample are sequenced to sequence a genomic portion. To determine whether a change (i.e., an imbalance) relative to a reference amount exists, one or more cutoff values are selected, such as a ratio of the amounts of two chromosomal regions (or groups of chromosomal regions).
[0018] According to an exemplary embodiment, a biological sample received from a pregnant woman is analyzed for prenatal diagnosis of fetal chromosomal aneuploidy. The biological sample includes nucleic acid molecules. A portion of the nucleic acid molecules contained in the biological sample is sequenced. In one aspect, the amount of genetic information obtained is sufficient for the accuracy of the diagnosis, yet not excessive, in order to control costs and the input amount of the biological sample required.
[0019] Based on the sequencing, a first amount of a first chromosome is determined from sequences identified as originating from the first chromosome. A second amount of one or more second chromosomes is determined from sequences identified as originating from one of the second chromosomes. Subsequently, the parameters of the first amount and the second amount are compared with one or more cutoff values. Based on the comparison, a classification of fetal chromosomal aneuploidy for the first chromosome is determined. Sequencing helps to maximize the amount of genetic information that can be inferred from fetal nucleic acids present as a minor fraction in a biological sample containing maternal background nucleic acids.
[0020] According to an exemplary embodiment, a biological sample received from a pregnant woman is analyzed for prenatal diagnosis of fetal chromosomal aneuploidy. The biological sample includes nucleic acid molecules. The percentage of fetal DNA in the biological sample is determined. Based on this percentage and the desired accuracy, the number N of sequences to be analyzed is calculated. At least N nucleic acid molecules contained in the biological sample are randomly sequenced.
[0021] Based on the random sequencing, a first amount of a first chromosome is determined from sequences identified as originating from the first chromosome. A second amount of one or more second chromosomes is determined from sequences identified as originating from one of the second chromosomes. Subsequently, the parameters of the first amount and the second amount are compared with one or more cutoff values. Based on the comparison, a classification of fetal chromosomal aneuploidy for the first chromosome is determined. Random sequencing helps to maximize the amount of genetic information that can be inferred from fetal nucleic acids present as a minor fraction in a sample containing maternal background nucleic acids.
[0022] Other embodiments of the present invention relate to systems and computer-readable media related to the methods described herein.
[0023] A better understanding of the features and advantages of the present invention can be obtained by referring to the detailed description and the drawings below. Brief Description of the Drawings
[0024] Figure 1 is a flowchart of method 100 of an embodiment of the present invention, which is used for prenatal diagnosis of fetal chromosomal aneuploidy in a biological sample obtained from a pregnant individual.
[0025] Figure 2 is a flowchart of method 200 of an embodiment of the present invention, which is used for prenatal diagnosis of fetal chromosomal aneuploidy by random sequencing.
[0026] Figure 3A A graph showing the percentage representation of chromosome 21 sequences in a maternal plasma sample related to a trisomy 21 or euploid fetus, which is an embodiment of the present invention.
[0027] Figure 3B A graph showing the correlation between the fractional fetal DNA concentration determined by massively parallel sequencing and microfluidics digital PCR in an embodiment of the present invention.
[0028] Figure 4A A graph showing the percentage representation of aligned sequences for each chromosome, which is an embodiment of the present invention.
[0029] Figure 4B Indicates Figure 4A A graph showing the difference (%) in the percentage representation of each chromosome between the trisomy 21 and euploid cases shown.
[0030] Figure 5 A graph showing the correlation between the degree of over-representation of chromosome 21 sequences and the fractional fetal DNA concentration in maternal plasma related to a trisomy 21 fetus, which is an embodiment of the present invention.
[0031] Figure 6 A table showing a part of the human genome analyzed according to an embodiment of the present invention. T21 represents a sample obtained from a pregnancy related to a trisomy 21 fetus.
[0032] Figure 7Table showing the number of sequences required to distinguish euploid from trisomy 21 fetuses for embodiments of the present invention.
[0033] Figure 8A Table showing the first 10 starting positions of sequenced tags aligned with chromosome 21 for embodiments of the present invention.
[0034] Figure 8B Table showing the first 10 starting positions of sequenced tags aligned with chromosome 22 for embodiments of the present invention.
[0035] Figure 9 Block diagram showing an exemplary computer device that can be used with the systems and methods of embodiments of the present invention.
[0036] Definitions
[0037] As used herein, the term "biological sample" refers to any sample containing one or more nucleic acid molecules of interest collected from an individual (such as a human, such as a pregnant woman).
[0038] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) in single-stranded or double-stranded form and polymers thereof, and unless otherwise limited, the term includes nucleic acids containing known analogs of natural nucleotides, which have binding properties similar to those of reference nucleic acids and are metabolized in a manner similar to that of naturally occurring nucleotides. Unless otherwise specified, a particular nucleic acid sequence also implicitly includes its conservatively modified variants (such as degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences, as well as the explicitly recited sequences. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is replaced with a mixture of bases and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994)). The term nucleic acid is used interchangeably with gene, cDNA, mRNA, small non-coding RNA, microRNA (miRNA), Piwi-interacting RNA, and short hairpin RNA (shRNA) encoded by a gene or locus.
[0039] The term "gene" means a segment of DNA that is related to the production of a polypeptide chain. It can include regions before and after the coding region (leader region and non-transcribed tail region), as well as intervening sequences (introns) between individual coding segments (exons).
[0040] As used herein, the term "reaction" refers to any process related to a chemical, enzymatic, or physical action that indicates the presence or absence of a specific polynucleotide sequence of interest. Examples of "reactions" are amplification reactions such as polymerase chain reaction (PCR). Another example of a "reaction" is a sequencing reaction by synthesis or by ligation. An "information reaction" is a reaction that indicates the presence of one or more specific polynucleotide sequences of interest, and in one case, there is only one sequence of interest. As used herein, the term "well" refers to a reaction in a predetermined location and a limited structure, such as a well-shaped flask, a chamber, or a chamber in a PCR array.
[0041] As used herein, the term "clinically relevant nucleic acid sequence" may refer to a polynucleotide sequence corresponding to a larger genomic sequence fragment whose potential imbalance is being detected, or to the larger genomic sequence itself. One example is the sequence of chromosome 21. Other examples include chromosomes 18, 13, X, and Y. Other examples in addition to these include mutant gene sequences, genetic polymorphisms, or copy number variations inherited by a fetus from one or both of its parents. Other examples in addition to these include sequences that are mutated, deleted, or amplified in a malignant tumor, such as sequences that have undergone loss of heterozygosity or gene duplication. In certain embodiments, multiple clinically relevant nucleic acid sequences, or multiple markers equivalent to clinically relevant nucleic acid sequences, can be used to provide data for detecting imbalances. For example, data from 5 non - contiguous sequences from chromosome 21 can be used in an additive fashion to determine a possible chromosome 21 imbalance, thereby effectively reducing the required sample volume to 1 / 5.
[0042] As used herein, the term "background nucleic acid sequence" refers to a nucleic acid sequence whose normal ratio to a clinically relevant nucleic acid sequence is known, such as a ratio of 1:1. As an example, the background nucleic acid sequence and the clinically relevant nucleic acid sequence are two alleles from the same chromosome that are different due to heterozygosity. In another example, the background nucleic acid sequence is an allele that is heterozygous with another allele that is the clinically relevant nucleic acid sequence. Moreover, each of certain background nucleic acid sequences and clinically relevant nucleic acid sequences can be from different individuals.
[0043] As used herein, the term "reference nucleic acid sequence" refers to a nucleic acid sequence whose average concentration in each reaction is known or has been equivalently measured.
[0044] As used herein, the term "overrepresented nucleic acid sequence" refers to a nucleic acid sequence among two sequences of interest (such as a clinically relevant sequence and a background sequence), and this overrepresented nucleic acid sequence is more abundant than other sequences in a biological sample.
[0045] As used herein, the term "based on" means "at least partially based on" and refers to a value (or result) used to determine another value, such as a value that exists in the relationship between the input of a method and the output of that method. As used herein, the term "obtained" also refers to the relationship between the input of a method and the output of that method, such as the relationship that exists when the obtaining is a calculation of a formula.
[0046] As used herein, the term "quantitative data" means data obtained from one or more reactions and providing one or more numerical values. For example, the number of wells with a fluorescent label representing a particular sequence is quantitative data.
[0047] As used herein, the term "parameter" means a numerical value that characterizes a quantitative data set and / or the numerical relationship between quantitative data sets. For example, the ratio (or a function of the ratio) between the first amount of a first nucleic acid sequence and the second amount of a second nucleic acid sequence is a parameter.
[0048] As used herein, the term "cutoff value" means a numerical value whose value is used to arbitrate between two or more classification states (e.g., diseased and non-diseased) of a biological sample. For example, if a parameter is greater than the cutoff value, the quantitative data is classified into a first category (e.g., diseased state), or if the parameter is less than the cutoff value, the quantitative data is classified into another category (e.g., non-diseased state).
[0049] As used herein, the term "imbalance" means any significant deviation from a reference amount, which is defined by at least one cutoff value in the amount of a clinically relevant nucleic acid sequence. For example, the ratio of the reference amount is 3 / 5, so if the measured ratio is 1:1, there is an imbalance.
[0050] As used herein, the term "chromosomal aneuploidy" means a change in the quantitative number of chromosomes relative to the number of chromosomes in a diploid genome. Such a change can be an increase or a loss. The change can include all of a chromosome or a region of a chromosome.
[0051] As used herein, the term "random sequencing" means sequencing in which the nucleic acid fragments to be sequenced are not specifically identified or targeted prior to the sequencing procedure. Sequence-specific primers that target a specific gene locus are not required. The pool of nucleic acids to be sequenced varies with the sample and even with the analysis for the same sample. The characteristics of the nucleic acids being sequenced are revealed only by the resulting sequencing output. In certain embodiments of the present invention, a procedure for enriching a biological sample with a particular population of nucleic acid molecules sharing certain common characteristics can precede random sequencing. In one embodiment, each fragment in the biological sample has an equal probability of being sequenced.
[0052] As used herein, the term "fraction of the human genome" or "portion of the human genome" means a nucleotide sequence of less than 100% of the human genome, which consists of approximately three billion nucleotide base pairs. In the context of sequencing, the term refers to a nucleotide sequence of less than one-fold coverage of the human genome. The term can be expressed as a percentage or an absolute value of nucleotides / base pairs. As an example of use, the term can be used to represent the actual amount of sequencing performed. Embodiments can determine the minimum value required for the sequenced portion of the human genome to obtain an accurate diagnosis. As another example of use, the term refers to the amount of sequencing data used to obtain parameters or amounts for disease classification.
[0053] As used herein, the term "sequenced tag" means a sequenced string of nucleotides from any part or all of a nucleic acid molecule. For example, a sequenced tag can be a short string of sequenced nucleotides from a nucleic acid fragment, a short string of nucleotides at both ends of a nucleic acid fragment, or the sequencing of a complete nucleic acid fragment present in a biological sample. A nucleic acid fragment is any part of a larger nucleic acid molecule. A fragment (such as a gene) can exist separately (i.e., not linked) from other parts of the larger nucleic acid molecule. Detailed Description of the Invention
[0054] Embodiments of the present invention provide methods, systems, and devices for determining whether there is an increase or decrease (disease state) in the presence of clinically relevant chromosomes compared to a non-diseased state. Such determination can be made by utilizing a parameter of the amount of a clinically relevant chromosomal region in relation to other non-clinically relevant chromosomal regions (background regions) in a biological sample. Nucleic acid molecules of the biological sample are sequenced in order to sequence a genomic fraction, and the amount can be determined from the sequencing results. One or more cut-off values are selected for determining whether there is a change (i.e., imbalance) compared to a reference amount, for example, a ratio of the amounts of two chromosomal regions (or groups of chromosomal regions).
[0055] The change detected in the reference amount can be any deviation (upward or downward) in relation to a clinically relevant nucleic acid sequence compared to other non-clinically relevant sequences. Thus, the reference state can be any ratio or other amount (such as other than a 1-1 correspondence), and the measured state indicating a change, as determined by one or more cut-off values, can be any ratio or other amount different from the reference amount.
[0056] Clinically relevant chromosomal regions (also referred to as clinically relevant nucleic acid sequences) and background nucleic acid sequences can be from a first type of cell and one or more second types of cells. For example, fetal nucleic acid sequences from fetal / placental cells are present in a biological sample, such as maternal plasma that contains a background of maternal nucleic acid sequences from maternal cells. In one embodiment, a cut-off value is determined at least in part based on the percentage of the first type of cells in the biological sample. It should be noted that the percentage of fetal sequences in the sample can be determined by any fetal-derived locus and is not limited to measuring clinically relevant nucleic acid sequences. In another embodiment, a cut-off value is determined at least in part based on the percentage of tumor sequences in the biological sample, which biological sample, such as plasma, serum, saliva or urine, contains a background of nucleic acid sequences from non-malignant cells in the body.
[0057] I. General Methods
[0058] Figure 1 is a flow chart of method 100 of an embodiment of the present invention, which method 100 is for performing prenatal diagnosis of fetal chromosomal aneuploidy in a biological sample obtained from a pregnant individual.
[0059] In step 110, a biological sample from a pregnant woman is received. The biological sample can be plasma, urine, serum or any other suitable sample. The sample contains nucleic acid molecules of the fetus and the pregnant woman. For example, the nucleic acid molecules can be fragments of chromosomes.
[0060] In step 120, at least a portion of the plurality of nucleic acid molecules contained in the biological sample is sequenced. The portion that is sequenced represents a portion of the human genome. In one embodiment, the nucleic acid molecules are fragments of respective chromosomes. One end (e.g., 35 base pairs (bp)), both ends or the complete fragment can be sequenced. All nucleic acid molecules in the sample can be sequenced, or only a subset can be sequenced. As described in more detail below, the subset can be randomly selected.
[0061] In one embodiment, sequencing is performed using massively parallel sequencing. Massively parallel sequencing, such as can be achieved by the 454 platform (Roche) (Margulies, M. et al. 2005 Nature 437, 376 - 380), the Illumina Genome Analyzer (or Solexa platform), or the SOLiD System (Applied Biosystems), or the Helicos True Single Molecule DNA sequencing technology (Harris TD et al. 2008 Science, 320, 106 - 109), the Single Molecule Real Time (SMRT TM ) technology of Pacific Biosciences, and nanopore sequencing (Soni GV and Meller A. 2007 Clin Chem 53:1996 - 2001), allows sequencing of many nucleic acid molecules isolated from a sample in parallel, at high - order multiplexing (Dear Brief Funct Genomic Proteomic 2003; 1:397 - 416). Each of these platforms can sequence individual molecules of clonally amplified or even unamplified nucleic acid fragments.
[0062] Because in each run, hundreds of thousands to millions or even potentially hundreds of millions or billions of levels of sequencing reads are generated from each sample, the resulting sequencing reads form a representative signature of the mixture of nucleic acid species in the original sample. For example, the haplotype, transcriptome, and methylation signatures of the sequencing reads are similar to these representative signatures of the original sample (Brenner et al Nat Biotech 2000; 18:630 - 634; Taylor et al Cancer Res 2007; 67:8511 - 8518). Due to the large sampling of sequences from each sample, the number of identical sequences, such as the number of identical sequences generated by sequencing a nucleic acid pool at several - fold coverage or high redundancy, is also a good quantitative representation of the count of a particular nucleic acid species or locus in the original sample.
[0063] In step 130, based on sequencing (such as data from sequencing), a first quantity of a first chromosome (such as a clinically relevant chromosome) is determined. The first quantity is determined by sequences identified as being from the first chromosome. For example, each of these DNA sequences can subsequently be mapped to the human genome using a bioinformatics program. It is possible to discard a portion of such sequences from subsequent analysis because they are in repetitive regions of the human genome or in regions that have undergone inter-individual variation such as copy number variation. Thus, the quantity of the chromosome of interest or the quantity of one or more other chromosomes can be determined.
[0064] In step 140, based on sequencing, a second quantity of one or more second chromosomes is determined by sequences identified as being from one of the second chromosomes. In one embodiment, the second chromosomes are all other chromosomes except the first chromosome (i.e., the chromosome being tested). In another embodiment, the second chromosome is a single other chromosome.
[0065] There are many ways to determine the quantity of a chromosome, including but not limited to counting the number of sequenced tags, the number of sequenced nucleotides (base pairs), or the cumulative length of the sequenced nucleotides (base pairs) from a particular chromosome or chromosomal region.
[0066] In another embodiment, rules can be applied to the sequencing results to determine what is counted. In one aspect, a quantity can be obtained based on a portion of the sequencing output. For example, the sequencing output corresponding to nucleic acid fragments within a specified size range can be selected after bioinformatics analysis. Examples of size ranges are <300bp, <200bp, or <100bp.
[0067] In step 150, a parameter is determined from the first quantity and the second quantity. The parameter can be, for example, a simple ratio of the first quantity to the second quantity, or a ratio of the first quantity to the sum of the first quantity and the second quantity. In one aspect, each quantity can be an independent variable of a function or different functions, where the ratio of these different functions can subsequently be obtained. Those skilled in the art will understand the number of different suitable parameters.
[0068] In one embodiment, a parameter (such as a fractional expression) of a chromosome potentially related to chromosomal aneuploidy, such as aneuploidy of chromosome 21 or chromosome 18 or chromosome 13, can subsequently be calculated from the results of a bioinformatics program. The fractional expression can be obtained based on the quantity of all sequences (such as certain measurements of all chromosomes including the clinically relevant chromosome) or the quantity of a specific subset of chromosomes (such as a single other chromosome excluding the chromosome being tested).
[0069] In step 150, the parameter is compared to one or more cut-off values. The cut-off values can be determined in any number of suitable ways. Such ways include Bayesian-type likelihood methods, sequential probability ratio tests, false discovery, confidence intervals, receiver operating characteristic (ROC). Examples of the application of these methods and sample-specific methods are described in the concurrently filed application "DETERMINING A NUCLEIC ACID SEQUENCE IMBALANCE" (Attorney Docket No. 016285-005210US), which is incorporated herein by reference.
[0070] In one embodiment, the parameter (such as the fractional expression of a clinically relevant chromosome) is then compared to a reference range established in pregnancies involving normal (i.e., euploid) fetuses. It is possible that in certain variations of the procedure, the reference range (i.e., the cut-off value) can be adjusted according to the fractional concentration (f) of fetal DNA in a particular maternal plasma sample. If the fetus is male, the f value can be determined from the sequencing data set using, for example, sequences that can be mapped to the Y chromosome. The f value can also be determined, for example, using fetal epigenetic markers (Chan KCA et al 2006 Clin Chem 52, 2211-8), or by analysis of single nucleotide polymorphisms, in a separate assay.
[0071] In step 160, based on the comparison, it is determined whether there is a classification of fetal chromosomal aneuploidy for the first chromosome. In one embodiment, the classification is a definite presence (yes) or absence (no). In another embodiment, the classification can be unclassifiable or indeterminate. In yet another embodiment, the classification can be, for example, a score to be interpreted later by a physician.
[0072] II. Sequencing, Alignment, and Quantity Determination
[0073] As described above, only a portion of the genome is sequenced. On the one hand, even when the nucleic acid pool in a sample is sequenced at less than 100% genome coverage rather than at several-fold coverage, and in a portion of the captured nucleic acid molecules, most each nucleic acid species is sequenced only once. It is also possible to quantitatively determine the dosage imbalance of a particular chromosome or chromosomal region. In other words, the dosage imbalance of a chromosome or chromosomal region is inferred from the percentage expression of the locus among other locatable sequenced tags in the sample.
[0074] This is in contrast to the situation where nucleic acids in the same pool are sequenced multiple times in order to obtain redundancy or several-fold coverage, whereby each nucleic acid species is sequenced multiple times. In this case, the number of times a particular nucleic acid species has been sequenced relative to another nucleic acid species is related to their relative concentration in the original sample. As the multiple of coverage required to achieve an accurate representation of the nucleic acid species increases, the cost of sequencing increases.
[0075] In one example, a portion of such sequences can be from chromosomes associated with aneuploidy, such as chromosome 21 in this exemplary example. However, other sequences of such a sequencing exercise can be from other chromosomes. By considering the relative size of chromosome 21 compared to other chromosomes, a normalized frequency of the chromosome 21-specific sequences of such a sequencing exercise can be obtained within a reference range. If the fetus has trisomy 21, the normalized frequency of the sequences obtained from chromosome 21 of such a sequencing exercise will increase, thus allowing the detection of trisomy 21. The degree of change in the normalized frequency will depend on the fractional concentration of fetal nucleic acids in the analyzed sample.
[0076] In one embodiment, we perform single-end sequencing of human genomic DNA and human plasma DNA samples using an Illumina Genome Analyzer. The Illumina Genome Analyzer can sequence individual DNA molecules that are non-clonally amplified and captured on a solid surface called a flow cell. Each flow cell has eight lanes for sequencing eight separate samples or sample pools. Each lane can generate approximately 200 Mb of sequence, which is only a fraction of the three billion base pair sequence in the human genome. One lane of the flow cell is used to sequence each genomic DNA or plasma DNA sample. The short sequence tags generated are aligned with the human reference genome sequence and the chromosomal origin is noted. The total number of individually sequenced tags aligned to each chromosome is tabulated and compared to the relative size of each chromosome expected for the reference human genome or non-disease presenting samples. Then an increase or loss of a chromosome is determined.
[0077] The method described is merely an example of the gene / chromosome dosage strategy described currently. Optionally, paired-end sequencing can be performed. The number of sequenced tags that are aligned is counted and classified according to chromosomal location, rather than comparing the lengths of the sequenced fragments expected in the reference genome as described by Campbell et al (Nat Genet 2008; 40:722-729). The gain or loss of chromosomal regions or entire chromosomes is determined by comparing the tag counts with the expected chromosomal sizes in the reference genome or the expected chromosomal sizes of non-disease manifestation samples. Since paired-end sequencing allows inference of the size of the original nucleic acid fragment, an example is dedicated to counting the number of paired-sequenced tags corresponding to nucleic acid fragments of a specified size, such as <300 bp, <200 bp, or <100 bp.
[0078] In another embodiment, prior to sequencing, a sub-selection is also performed on a portion of the nucleic acid pool to be sequenced during the run. For example, hybridization-based techniques, such as oligonucleotide arrays, can be used to first perform a sub-selection on nucleic acid sequences from certain chromosomes, such as potential aneuploid chromosomes and other chromosomes not related to the detected aneuploidy. Another example is that, prior to sequencing, a sub-selection or enrichment is performed on certain sub-populations of the nucleic acid sequences of the sample pool. For example, as discussed above, it has been reported that fetal DNA molecules in maternal plasma consist of fragments shorter than maternal background DNA molecules (Chan et al Clin Chem 2004; 50:88-92). Thus, for example, by gel electrophoresis or size exclusion column or by a microfluidics-based approach, the nucleic acid sequences in the sample can be fractionated according to molecular size using one or more methods known to those skilled in the art. Additionally, optionally, in an example of analyzing cell-free fetal DNA in maternal plasma, the fetal nucleic acid fraction can be enriched by methods that suppress the maternal background, such as by adding formaldehyde (Dhallan et al JAMA 2004; 291:1114-9). In one embodiment, random sequencing is performed on a portion or sub-population of the preselected pool of nucleic acids.
[0079] Similarly, other single molecule sequencing strategies can also be used in this application, such as the Roche 454 platform, the Applied Biosystems SOLiD platform, the Helicos true single molecule DNA sequencing technology, the single molecule real-time technology (SMRTTM) of Pacific Biosciences, and nanopore sequencing.
[0080] III. Determination of Chromosome Quantity from Sequencing Output
[0081] After massively parallel sequencing, bioinformatics analysis is performed to localize the chromosomal origin of the sequenced tags. After this procedure, tags identified as being from potentially aneuploid chromosomes, i.e., chromosome 21 in this study, are quantitatively compared with all the sequenced tags or tags from one or more chromosomes not associated with aneuploidy. The correlation between the sequencing outputs of chromosome 21 and other non-chromosome 21 of the test sample is compared with the cut-off value obtained by the method described in the previous section to determine whether the sample is obtained from a pregnancy related to an euploid or trisomy 21 fetus.
[0082] Many different quantities, including but not limited to the following, can be obtained from the sequenced tags. For example, the number of sequenced tags aligned to a particular chromosome, i.e., the absolute count, can be compared with the absolute count of sequenced tags aligned to other chromosomes. Optionally, the fractional count of the quantity of sequenced tags of chromosome 21 can be compared with the fractional counts of other non-aneuploid chromosomes, with reference to all or some other sequenced tags. In this experiment, since 36 bp of each DNA fragment was sequenced, the number of nucleotides sequenced for a particular chromosome can be easily obtained by multiplying the count of sequenced tags by 36 bp.
[0083] In addition, since only one flow cell that can only sequence a part of the human genome is used to sequence each maternal plasma sample, statistically, most types of maternal plasma DNA fragments are only sequenced once, resulting in a count of sequenced tags. In other words, the nucleic acid fragments present in the maternal plasma sample are sequenced at a coverage less than 1-fold. Therefore, for any particular chromosome, the total number of nucleotides sequenced generally corresponds to the quantity, proportion or length of the part of the said chromosome that has been sequenced. Therefore, the quantitative determination of the expression level of a potentially aneuploid chromosome can be obtained from the partial number of nucleotides sequenced or the equivalent length of this potentially aneuploid chromosome, with reference to the similarly obtained quantities of other chromosomes.
[0084] IV. Enrichment of Nucleic Acid Pools for Sequencing
[0085] As mentioned above and established in the examples of the following section, only a portion of the human genome needs to be sequenced to distinguish trisomy 21 from euploid conditions. Thus, it may be possible and cost-effective to enrich a pool of nucleic acids to be sequenced before randomly sequencing a portion of the enriched pool. For example, fetal DNA molecules in maternal plasma consist of shorter fragments than maternal background DNA molecules (Chan et al Clin Chem 2004; 50:88-92). Thus, for example, by gel electrophoresis or size-exclusion columns or by microfluidics-based methods, nucleic acid sequences in a sample can be fractionated according to molecular size using one or more methods known to those skilled in the art.
[0086] In addition, optionally, in instances of analyzing cell-free fetal DNA in maternal plasma, the fetal nucleic acid fraction can be enriched by methods that suppress the maternal background such as addition of formaldehyde (Dhallan et al JAMA 2004; 291:1114-9). The proportion of sequences derived from the fetus will be enriched in a nucleic acid pool consisting of shorter fragments. Depending on Figure 7 , the number of tags to be sequenced required to distinguish euploid and trisomy 21 conditions will decrease as the fetal DNA fractional concentration increases.
[0087] Optionally, sequences from potentially aneuploid chromosomes and one or more chromosomes not associated with aneuploidy can be enriched by hybridization techniques such as oligonucleotide microarrays. The enriched pool of nucleic acids is then randomly sequenced. This will reduce the cost of sequencing.
[0088] V. Random Sequencing
[0089] Figure 2 is a flowchart of method 200 for prenatal diagnosis of fetal chromosomal aneuploidy using random sequencing according to an embodiment of the present invention. In one aspect of the massively parallel sequencing method, representative data for all chromosomes can be generated simultaneously. The source of specific fragments is not preselected. Sequencing is performed randomly and then followed by a database search to determine where a particular fragment originated. This is contrary to the case of amplifying a specific fragment of chromosome 21 and another specific fragment of chromosome 1.
[0090] In step 210, a biological sample is received from a pregnant woman. In step 220, for a desired accuracy, the number N of sequences to be analyzed is calculated. In one embodiment, the percentage of fetal DNA in the biological sample is first determined. This can be done by any suitable means known to those skilled in the art. The determination can be simply reading a value measured by another entity. In this embodiment, the calculation of the number N of sequences to be analyzed is based on the percentage. For example, when the percentage of fetal DNA decreases, the number of sequences to be analyzed will increase, while when the fetal DNA increases, the number of sequences to be analyzed can be reduced. The number N can be a fixed number or a relative number, such as a percentage. In another embodiment, a number N known to be sufficient for accurate disease diagnosis can be sequenced. Even in pregnancies with a fetal DNA concentration at the lower end of the normal range, the number N can be made sufficient.
[0091] In step 230, at least N of the plurality of nucleic acid molecules contained in the biological sample are randomly sequenced. The method is characterized in that, prior to sample analysis, i.e., sequencing, the nucleic acids to be sequenced are not specifically determined or targeted. The sequencing does not require sequence-specific primers targeting specific loci. The pool of nucleic acids being sequenced varies with the sample and even with different analyses of the same sample. Additionally, as described below ( Figure 6 ), the amount of sequencing output required for case diagnosis can vary between the samples being tested and a reference population. These aspects are significantly different from most molecular diagnostic methods, such as fluorescence-based methods in in situ hybridization, quantitative fluorescence PCR, quantitative real-time PCR, digital PCR, comparative genomic hybridization, microarray comparative genomic hybridization, etc., where the loci to be targeted need to be pre-determined and thus require the use of locus-specific primers or locus-specific probe pairs or panels.
[0092] In one embodiment, DNA fragments present in the plasma of a pregnant woman are randomly sequenced, and genomic sequences originally from the fetus or the mother are obtained. Random sequencing involves sampling (sequencing) a random portion of the nucleic acid molecules present in the biological sample. Since the sequencing is random, different subsets (portions) of the nucleic acid molecules (and thus the genome) can be sequenced in each analysis. This embodiment remains valid even when the subset varies with the sample or the analysis. Examples of the portion are about 0.1%, 0.5%, 1%, 5%, 10%, 20%, or 30% of the genome. In another embodiment, the portion is at least any one of these values.
[0093] The remaining steps 240 - 270 can be carried out in a manner similar to method 100.
[0094] VI. Post-Sequencing Selection of Sequenced Tag Pools
[0095] As described in Examples II and III below, a subset of the sequencing data is sufficient to distinguish trisomy 21 and aneuploidy cases. The subset of the sequencing data can be a certain proportion of the sequenced tags that convey certain property parameters. For example, in Example II, the sequenced tags uniquely aligned to the repeat-masked reference human genome are used. Optionally, a representative pool of nucleic acid fragments of all chromosomes can be sequenced, but efforts are made to compare the data on the potential aneuploid chromosomes and the data on many non-aneuploid chromosomes.
[0096] In addition, optionally, during the post-sequencing analysis, a secondary selection can be made on a subset of the sequencing output, and the subset includes the sequenced tags generated from the nucleic acid fragments corresponding to a specified size window in the original sample. For example, using an Illumina genome analyzer, paired-end sequencing involving sequencing of both ends of the nucleic acid fragment can be used. Subsequently, the sequencing data of each paired end is aligned with the reference human genome sequence. Subsequently, the distance or number of nucleotides spanning between the two ends can be deduced. The full length of the original nucleic acid fragment can also be deduced. Optionally, sequencing platforms such as the 454 platform, and possibly certain single molecule sequencing technologies, can sequence short nucleic acid fragments of full length, such as 20 bp. In this way, the actual length of the nucleic acid fragment can be directly known from the sequencing data.
[0097] Using other sequencing platforms, such as the Applied Biosystems SOLiD system (Applied Biosystems SOLiD system), such paired-end analysis is also possible. For the Roche 454 platform, because the read length of this 454 platform is increased compared with other massively parallel sequencing systems, it is also possible to determine the fragment length of the full sequence of the fragment.
[0098] Focusing the data analysis on a subset of the sequenced tags corresponding to the short nucleic acid fragments in the original maternal plasma sample has the advantage that the DNA sequences from the fetus are effectively enriched in the data set. This is because the fetal DNA molecules in maternal plasma consist of fragments shorter than the maternal background DNA molecules (Chan et al Clin Chem 2004; 50:88-92). According to Figure 7 , the number of the sequenced tags required to distinguish euploidy and trisomy 21 cases will decrease with the increase in the fetal DNA fractional concentration.
[0099] The selection after nucleic acid pool subpopulation sequencing is different from other nucleic acid enrichment strategies implemented prior to sample analysis, such as gel electrophoresis or size exclusion columns used to select nucleic acid molecules of a specific size, and such strategies require physical separation of the enriched pool from the nucleic acid background pool. Physical procedures can introduce more experimental steps and thus can incur problems such as contamination. Depending on the sensitivity and specificity required for disease determination, post-sequencing in silico selection of the sequencing output subpopulation can also allow for altered selection.
[0100] Bioinformatics, computational, and statistical methods for determining whether a maternal plasma sample is obtained from a pregnant woman carrying a fetus with trisomy 21 or aneuploidy can be compiled into a computer program product for determining parameters of the sequencing output. The running of the computer program includes determining the quantitative number of potential aneuploid chromosomes and the amounts of one or more other chromosomes. Parameters are determined and compared with appropriate cut-off values to determine whether there is fetal chromosomal aneuploidy for the potential aneuploid chromosomes. Examples
[0101] To illustrate, but not limit, the claimed invention, the following examples are provided.
[0102] I. Prenatal Diagnosis of Fetal Trisomy 21
[0103] Eight pregnant women were recruited for this study. All pregnant women were in the first or second trimester of pregnancy and had a singleton pregnancy. Four of them each carried a fetus with trisomy 21, and the other four each carried aneuploidy. Twenty milliliters of peripheral venous blood was collected from each individual. After centrifugation at 1600×g for 10 minutes, maternal plasma was harvested and further centrifuged at 16000×g for 10 minutes. Subsequently, DNA was extracted from 5 - 10 ml of each plasma sample. Maternal plasma DNA was used for massively parallel sequencing on an Illumina genome analyzer according to the manufacturer's instructions. During the sequencing and sequence data analysis process, the technician performing the sequencing was unaware of the fetal diagnosis.
[0104] Briefly, approximately 50 ng of maternal plasma DNA was used to prepare the DNA library. One could start with a lesser amount such as 15 ng or 10 ng of maternal plasma DNA. The maternal plasma DNA fragments were blunt-ended, ligated to Solexa adaptors, and fragments of 150 - 300 bp were selected by gel purification. Optionally, the blunt-ended and adaptor-ligated maternal plasma DNA fragments could be passed through a column (such as AMPure, Agencourt) to remove unligated adaptors without size selection prior to cluster generation. The adaptor-ligated DNA was hybridized to the surface of the flow cell, and DNA clusters were generated using an Illumina cluster station, followed by 36 cycles of sequencing on an Illumina Genome Analyzer. The DNA of each maternal plasma sample was sequenced on one flow cell. The sequencing reads were edited using the Solexa Analysis Pipeline. Subsequently, using the Eland application software, all reads were aligned to the repeat-masked reference human genome sequence, i.e., NCBI Build 36 (NCBI 36 assembly) (GenBank accession numbers: NC_000001 to NC_000024).
[0105] In this study, to reduce the complexity of data analysis, only sequences that had been mapped to unique positions in the repeat-masked human genome reference were further considered. Optionally, other subsets of the sequencing data or the entire set of sequencing data could be used. The total number of uniquely mappable sequences for each sample was counted. The number of sequences uniquely aligned to chromosome 21 was expressed as a proportion of the total count of sequences aligned to each sample. Since maternal plasma contains fetal DNA in the background DNA of maternal origin, trisomy 21 fetuses provide additional sequenced tags from chromosome 21 due to the presence of an extra copy of chromosome 21 in the fetal genome. Thus, in the maternal plasma from pregnancies with trisomy 21 fetuses, the percentage of chromosome 21 sequences was higher than that of chromosome 21 from pregnancies with euploid fetuses. The analysis did not require targeting fetal-specific sequences. The analysis also did not require prior physical separation of fetal nucleic acids from maternal nucleic acids. The analysis also did not require distinguishing or identifying fetal sequences from maternal sequences after sequencing.
[0106] Figure 3ARepresents the percentage of sequences mapped to chromosome 21 (chromosome 21 percentage representation) for each of 8 maternal plasma DNA samples. The chromosome 21 percentage representation in the maternal plasma of trisomy 21 pregnancies is significantly higher than that in euploid pregnancies. These data indicate that non-invasive prenatal diagnosis of fetal aneuploidy can be achieved by determining the percentage representation of aneuploid chromosomes compared to that of a reference population. Optionally, overrepresentation of chromosome 21 can be detected by comparing the experimentally obtained percentage representation of chromosome 21 with the percentage representation of chromosome 21 sequences expected for an euploid human genome. This can be done with or without masking repetitive regions in the human genome.
[0107] Five out of 8 pregnant women each carried a male fetus. Sequences mapped to the Y chromosome can be fetal specific. The percentage of sequences mapped to the Y chromosome is used to calculate the fetal DNA fractional concentration in the original maternal plasma sample. Moreover, the fetal DNA fractional concentration is also determined using microfluidic digital PCR, which involves zinc finger protein, X-linked (ZFX) and zinc finger protein, Y-linked (ZFY) paralogous genes.
[0108] Figure 3B Represents the correlation between the fetal DNA fractional concentration inferred from the percentage representation of the sequenced Y chromosome and the fetal DNA fractional concentration determined by ZFY / ZFX microfluidic digital PCR. There is a positive correlation between the fetal DNA fractional concentrations in maternal plasma determined by these two methods. The positive correlation coefficient (r) is 0.917 in Pearson correlation analysis.
[0109] For two representative cases, the percentages of maternal plasma DNA sequences aligned to each of 24 chromosomes (22 autosomes and X chromosome and Y chromosome) are shown in Figure 4A One pregnant woman carried a trisomy 21 fetus and the other pregnant women carried euploid fetuses. The percentage representation of sequences mapped to chromosome 21 is higher in the pregnant woman carrying a trisomy 21 fetus compared to those carrying normal fetuses.
[0110] The difference (%) in the percentage representation of each chromosome between the maternal plasma DNA samples of the above two cases is shown in Figure 4B The percentage difference for a specific chromosome is calculated using the following formula:
[0111] Percentage difference (%) = (P 21 - P E ) / P E × 100%, where
[0112] P 21 = the percentage of plasma DNA sequences aligned with a specific chromosome in pregnant women carrying fetuses with trisomy 21; and
[0113] P E = the percentage of plasma DNA sequences aligned with a specific chromosome in pregnant women carrying euploid fetuses.
[0114] As Figure 4B shown, compared with pregnant women carrying euploid fetuses, there is an overrepresentation of 11% of chromosome 21 sequences in the plasma of pregnant women carrying fetuses with trisomy 21. For sequences aligned with other chromosomes, the difference between the two cases is within 5%. Since the percentage representation of chromosome 21 increases in trisomy 21 compared with euploid maternal plasma samples, the difference (%) can optionally be referred to as the degree of overrepresentation of chromosome 21. In addition to the difference (%) and absolute difference in the percentage representation of chromosome 21, the ratio of the counts of the test sample and the reference sample can also be calculated, and this ratio represents the degree of overrepresentation of chromosome 21 in trisomy 21 compared with the euploid sample.
[0115] For 4 pregnant women each carrying an euploid fetus, 1.345% of their plasma DNA sequences on average were aligned with chromosome 21. Among 4 pregnant women carrying fetuses with trisomy 21, 3 of their fetuses were male. The percentage representation of chromosome 21 was calculated for each of these three cases. As described above, the difference (%) in the percentage representation of chromosome 21 in these three trisomy 21 cases was determined based on the average percentage representation of chromosome 21 obtained from the values of the 4 euploid cases. In other words, in this calculation, the average of the 4 cases of pregnant women carrying euploid fetuses was used as the reference. The fetal DNA fractional concentrations of these three male trisomy 21 cases were inferred from the percentage representation of their respective Y chromosome sequences.
[0116] The correlation between the degree of overrepresentation of chromosome 21 sequences and the fetal DNA fractional concentration is shown in Figure 5 . There is a significant positive correlation between the two parameters. The correlation coefficient (r) was 0.898 in the Pearson correlation analysis. These results indicate that the degree of overrepresentation of chromosome 21 sequences in maternal plasma is correlated with the fractional concentration of fetal DNA in the maternal plasma sample. Therefore, a cut-off value in the degree of overrepresentation of chromosome 21 sequences correlated with the fetal DNA fractional concentration can be determined to identify pregnancies associated with fetuses with trisomy 21.
[0117] The determination of the fractional concentration of fetal DNA in maternal plasma can also be performed independently of the sequencing run. For example, the concentration of Y-chromosome DNA can be pre-determined using real-time PCR, microfluidic PCR, or mass spectrometry. For example, we have shown in Figure 3B that there is a good correlation between the fetal DNA concentration estimated based on the Y-chromosome counts generated during the sequencing run and the ZFY / ZFX ratio generated outside the sequencing run. In fact, the fetal DNA concentration can be determined using loci other than the Y-chromosome and is applicable to female fetuses. For example, Chan et al. demonstrated that fetal-derived methylated RASSF1A sequences can be detected in pregnant women's plasma in the context of maternally-derived unmethylated RASSF1A sequences (Chan et al, Clin Chem 2006; 52:2211-8). Thus, the fractional concentration of fetal DNA can be determined by dividing the amount of methylated RASSF1A sequences by the amount of all RASSF1A (methylated and unmethylated) sequences.
[0118] For the implementation of our invention, maternal plasma is expected to be preferred over maternal serum because maternal blood cells release DNA during blood clotting. Thus, if serum is used, the fractional concentration of fetal DNA is expected to be lower in maternal plasma than in maternal serum. In other words, if maternal serum is used, more sequences are expected to be generated for the diagnosis of fetal chromosomal aneuploidy compared to plasma samples obtained simultaneously from the same pregnant woman.
[0119] In addition, another alternative way to determine the fractional concentration of fetal DNA is via quantification of polymorphic differences between the pregnant woman and the fetus (Dhallan R, et al. 2007 Lancet, 369, 474-481). An example of this method is to target polymorphic loci at which the pregnant woman is homozygous and the fetus is heterozygous. The amount of the fetal-specific allele is compared to the amount of the common allele in order to determine the fractional concentration of fetal DNA.
[0120] Contrary to the prior art for detecting chromosomal aberrations, which includes comparative genomic hybridization for detecting and quantifying one or more specific sequences, microarray comparative genomic hybridization, quantitative real-time polymerase chain reaction, massively parallel sequencing does not rely on the detection or analysis of a pre-determined or pre-defined set of DNA sequences. A randomly representative portion of the DNA molecules in the sample pool is sequenced. The number of different sequenced tags aligned to various chromosomal regions is compared between samples with or without the DNA species of interest. Chromosomal aberrations will be revealed by differences in the number (or percentage) of sequences aligned to any given chromosomal region in the sample.
[0121] In another embodiment, sequencing techniques for cell-free plasma DNA can be used to detect chromosomal aberrations in plasma DNA to detect specific cancers. Different cancers have a set of typical chromosomal aberrations. Changes (amplifications and deletions) in multiple chromosomal regions can be used. Thus, the proportion of sequences aligned to the amplified regions will increase, while the proportion of sequences aligned to the reduced regions will decrease. The percentage representation of each chromosome can be compared to the size of each corresponding chromosome in the reference genome, which is expressed as a percentage of the genomic representation of any given chromosome relative to the whole genome. Direct comparison or comparison with a reference chromosome can also be used.
[0122] II. Sequencing Only a Portion of the Human Genome
[0123] In the experiment described in Example I above, only one flow cell was used to sequence the maternal plasma DNA of each individual sample. After the sequencing run, the number of sequenced tags generated by each tested sample is shown in Figure 6 . T21 represents the samples obtained from pregnancies associated with trisomy 21 fetuses.
[0124] Since 36 bp of each sequenced maternal plasma DNA fragment was sequenced, the number of sequenced nucleotides / base pairs for each sample can be determined by multiplying the count of sequenced tags by 36 bp and is also shown in Figure 6 . Since there are approximately 3 billion base pairs in the human genome, the amount of sequencing data generated by each maternal plasma sample represents only about 10% to 13% of the portion.
[0125] Furthermore, in this study, as described in Example I above, only uniquely mappable sequenced tags, called U0 in the nomenclature of the Eland software, were used to demonstrate the overrepresentation of the amount of chromosome 21 sequences in each of the maternal plasma samples from pregnancies with trisomy 21 fetuses. As Figure 6 shown, the U0 sequences only represent a subset of all the sequenced tags generated by each sample and also represent an even smaller proportion, about 2%, of the human genome. These data indicate that sequencing only a portion of the human genome sequences present in the tested samples is sufficient to achieve the diagnosis of fetal aneuploidy.
[0126] III. Determination of the Quantity of Required Sequences
[0127] The sequencing results of plasma DNA from pregnant women carrying euploid male fetuses were used in this analysis. The number of sequenced tags that could be mapped without mismatch to the reference human genome sequence was 1,990,000. A subset of sequences was randomly selected from these 1,990,000 tags, and the percentage of sequences aligned to chromosome 21 was calculated in each subset. The number of sequences in the subset varied from 60,000 to 540,000 sequences. For each subset size, multiple subsets of the same number of sequenced tags were generated by randomly selecting the sequenced tags from the total pool until no other possible combinations remained. Subsequently, within each subset size, the average percentage of sequences aligned to chromosome 21 and its standard deviation (SD) were calculated from the multiple subsets. These data were compared across different subset sizes to determine the effect of subset size on the percentage distribution of sequences aligned to chromosome 21. Subsequently, the 5th and 95th percentiles of the percentage were calculated based on the mean and SD.
[0128] When a pregnant woman is carrying a fetus with trisomy 21, due to the extra dose of chromosome 21 from the fetus, the sequenced tags aligned to chromosome 21 should be overrepresented in the maternal plasma. The degree of overrepresentation depends on the percentage of fetal DNA in the maternal plasma DNA sample and is calculated using the following equation:
[0129] Per T21 =Per Eu ×(1 + f / 2), where,
[0130] Per T21 represents the percentage of sequences aligned to chromosome 21 in women carrying a fetus with trisomy 21; and
[0131] Per Eu represents the percentage of sequences aligned to chromosome 21 in women carrying an euploid fetus; and
[0132] f represents the percentage of fetal DNA in the maternal plasma DNA.
[0133] As Figure 7 shown, the SD of the percentage of sequences aligned to chromosome 21 decreases with an increase in the number of sequences in each subset. Therefore, as the number of sequences in each subset increases, the interval between the 5th and 95th percentiles decreases. When the 5%-95% intervals for the euploid and trisomy 21 cases do not overlap, it is possible to distinguish between the two groups of cases with an accuracy greater than 95%.
[0134] As Figure 7As shown, the minimum subgroup size for differentiating trisomy 21 cases from euploid cases depends on the percentage of fetal DNA. For fetal DNA percentages of 20%, 10%, and 5%, the minimum subgroup sizes for differentiating trisomy 21 and euploid cases are 120,000, 180,000, and 540,000 sequences, respectively. In other words, when the maternal plasma DNA sample contains 20% fetal DNA, the number of sequences that need to be analyzed to determine whether the fetus has trisomy 21 is 120,000. When the fetal DNA percentage is reduced to 5%, the number of sequences that need to be analyzed increases to 540,000.
[0135] Since data are generated using 36-base pair sequencing, 120,000, 180,000, and 540,000 sequences correspond to 0.14%, 0.22%, and 0.65% of the human genome, respectively. Since it has been reported that the lower range of fetal DNA concentration in maternal plasma obtained in early pregnancy is approximately 5% (Lo, YMD et al. 1998 Am J Hum Genet 62, 768 - 775), sequencing approximately 0.6% of the human genome can represent the minimum amount of sequencing required for a diagnosis with at least 95% accuracy in detecting fetal chromosomal aneuploidy in any pregnancy.
[0136] IV. Random Sequencing
[0137] To demonstrate that the DNA fragments being sequenced are randomly selected during the sequencing run, we obtained the sequenced tags generated from the 8 maternal plasma samples analyzed in Example I. For each maternal plasma sample, relative to the reference human genome sequence, i.e., NCBI Build 36, we determined the starting position of each 36bp sequenced tag that uniquely aligned to chromosome 21 without a mismatch. We then sorted the starting position numbers of the aligned sequenced tag pools from each sample in ascending order. We performed a similar analysis on chromosome 22. For illustrative purposes, the first 10 starting positions of chromosome 21 and chromosome 22 for each maternal plasma sample are shown in Figure 8A and Figure 8B respectively. From these tables, it can be seen that the sequenced pools of DNA fragments are different among the samples.
[0138] Using any suitable computer language, such as Java, C++, or Perl using, for example, conventional or object-oriented techniques, any software component or function described in this application can be executed as software code run by a processor. The software code can be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable media include random access memory (RAM), read-only memory (ROM), magnetic media such as hard disks or floppy disks, or optical media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, etc. The computer-readable medium can be any combination of such storage or transmission devices.
[0139] Such programs can also be encoded and transmitted using carrier signals suitable for propagation over wired, optical, and / or wireless networks that comply with various protocols including the Internet. Thus, the computer-readable medium of embodiments of the present invention can be generated using data signals encoded with such programs. The computer-readable medium encoded with program code can be assembled with a compatible device or provided independently by other devices (such as downloaded via the Internet). Any such computer-readable medium can be located on or within a computer program product (e.g., a hard disk or an entire computer system), and can exist on or within different computer program products in a system or network. The computer system can include a display screen, a printer, or other suitable displays that provide any of the results mentioned herein.
[0140] An example of a computer system is shown in Figure 9 In Figure 9 The subsystems shown in are interconnected via a system bus 975. Figure 9 Other subsystems are shown, such as printer 974, keyboard 978, hard disk 979, display screen 976 connected to a display adapter 982, etc. Peripheral devices and input / output (I / O) devices connected to an I / O controller 971 can be connected to the computer system in any number of ways known in the art, such as a serial port 977. For example, the serial port 977 or an external interface 981 can be used to connect a computer device to a wide area network such as the Internet, a mouse input device, or a scanner. Interconnecting via the system bus allows the central processing unit 973 to communicate with each subsystem and control the execution of instructions in the system memory 972 or the hard disk 979 and the exchange of information between subsystems. The system memory 972 and / or the hard disk 979 are specific manifestations of a computer-readable medium.
[0141] For purposes of illustration and description, the above presents a description of exemplary embodiments of the present invention. It is not intended to be comprehensive or to limit the invention to the exact forms described, and many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling those skilled in the art to best utilize the invention in various embodiments and with various modifications suited to the particular uses contemplated.
[0142] For all purposes, all publications, patents, and patent applications cited herein are incorporated by reference in their entireties.
Claims
1. A computer system for analyzing cell-free nucleic acids in a biological sample obtained from a pregnant individual, the computer system comprising: a device receiving sequenced tag data, the sequenced tags obtained from random sequencing of at least a portion of a plurality of cell-free nucleic acids from the genome of the pregnant individual and from the genome of at least one fetus contained in a biological sample of the pregnant individual, wherein the sequenced tags include sequenced tags corresponding to the cell-free nucleic acids from the genome of the pregnant individual and sequenced tags corresponding to the cell-free nucleic acids from the genome of at least one fetus, wherein the at least a portion represents a portion of the human genome and the portion of the human genome represents at least 0.5% of the human genome, wherein the biological sample is maternal plasma; means for determining the chromosomal origin of sequenced tags; means for determining a first quantity of a first chromosome based on the randomly sequenced data, the first quantity being determined by sequenced tags identified as being from the first chromosome; means for determining a second quantity of one or more second chromosomes based on the randomly sequenced data, the second quantity being determined by a sequenced tag identified as one of the second chromosomes; means for determining a parameter from the first quantity and the second quantity, wherein the parameter represents a ratio between the first quantity and the second quantity; as well as A device for determining a classification of whether a fetal chromosomal aneuploidy is present for the first chromosome, the device being used to compare the parameter with one or more cutoff values and, based on the comparison, determine a classification of whether a fetal chromosomal aneuploidy is present for the first chromosome, wherein the cutoff value is a reference value established in a normal biological sample.
2. The computer system of claim 1 , further comprising: Means for determining the percentage of fetal DNA in said biological sample.
3. The computer system of claim 1, wherein the first chromosome is chromosome 21, chromosome 18, chromosome 13, chromosome X, or chromosome Y.
4. The computer system of claim 1, wherein the first amount is obtained from a subpopulation of sequenced tags corresponding to nucleic acid fragments of a specified size.
5. The computer system of claim 1, wherein at least one of said cutoff values is related to the percentage of said fetal DNA in said biological sample.
6. The computer system of claim 5, wherein the percentage of fetal DNA in the biological sample is determined by any one or more of the ratio of Y chromosome sequences, fetal epigenetic markers, or using single nucleotide polymorphism analysis.
7. The computer system of claim 1, wherein said nucleic acid molecules of said biological sample have been enriched for sequences from at least one specific chromosome.
8. The computer system of claim 1, wherein the nucleic acid molecules of the biological sample have been enriched for sequences less than 300 bp.
9. The computer system of claim 1, wherein the nucleic acid molecules of the biological sample have been enriched for sequences less than 200 bp.
10. The computer system of claim 1, wherein the nucleic acid molecules of the biological sample have been amplified using polymerase chain reaction.
11. A computer system for analyzing cell-free DNA fragments in plasma obtained from a pregnant individual, the computer system comprising: a device receiving sequenced tag data, the sequenced tags obtained from random sequencing of at least a portion of a plurality of cell-free DNA fragments from the genome of the pregnant individual and from the genome of at least one fetus contained in plasma of the pregnant individual, wherein the sequenced tags include sequenced tags corresponding to the cell-free DNA fragments from the genome of the pregnant individual and sequenced tags corresponding to the cell-free DNA fragments from the genome of at least one fetus, wherein the at least a portion represents at least 0.5% of the human genome; means for determining the chromosomal origin of sequenced tags; means for determining a first amount of chromosome 21 based on the randomly sequenced data, wherein the first amount is determined by sequenced tags identified as being from chromosome 21; means for determining a second quantity of one or more second chromosomes based on the randomly sequenced data, the second quantity being determined by a sequenced tag identified as one of the second chromosomes; means for determining a parameter from the first quantity and the second quantity, wherein the parameter represents a ratio between the first quantity and the second quantity; as well as A device for determining a classification of whether a fetal chromosomal aneuploidy is present for the first chromosome, the device being used to compare the parameter with one or more cutoff values and, based on the comparison, determine a classification of whether a fetal chromosomal aneuploidy is present for the first chromosome, wherein the cutoff value is a reference value established in a normal biological sample.
12. The computer system of claim 11, further comprising: Means for determining the size of DNA fragments based on the random sequencing data, wherein the first amount is obtained from a subset of sequenced tags corresponding to nucleic acid molecules of a specified size.
13. The computer system of claim 12, wherein the random sequencing data comprises paired-end sequence data.
14. The computer system of any one of claims 11 to 13, further comprising: Means for determining the percentage of fetal DNA in said biological sample.
15. The computer system of any one of claims 11-14, wherein at least one of said cutoff values is related to the percentage of said fetal DNA in said biological sample.
16. The computer system of any one of claims 11-15, wherein the percentage of fetal DNA in the biological sample is determined by any one or more of the ratio of Y chromosome sequences, fetal epigenetic markers, or using single nucleotide polymorphism analysis.