Molecular analysis using cell-free fragments during pregnancy
By analyzing long cell-free DNA fragments using PacBio sequencing, the challenges of detecting and analyzing long DNA fragments in existing sequencing technologies are overcome, enabling accurate fetal tissue identification and disorder detection.
Patent Information
- Application Number
- JP2023216655
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-01-08
- Filing Date
- 2023-12-22
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-02-05
AI Technical Summary
Existing sequencing technologies, such as the Illumina platform, struggle with the detection and analysis of long cell-free DNA fragments greater than 600 bp due to amplification biases and secondary structure formation, leading to incomplete daughter strand synthesis and inaccurate representation of long DNA molecules in sequencing libraries.
Analyzing biological samples using long cell-free DNA fragments, particularly those over 200 bp, to identify multiple CpG sites, SNPs, and other genetic markers for fetal tissue origin and pregnancy-related disorders, utilizing methods like PacBio sequencing to capture and analyze these fragments effectively.
Enables more accurate and efficient analysis of fetal DNA by identifying multiple CpG sites and SNPs, determining fetal tissue origin, gestational age, and detecting genetic disorders through the analysis of long cell-free DNA fragments.
Smart Images

Figure 0007772389000011 
Figure 0007772389000012 
Figure 0007772389000013
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 970,634, filed February 5, 2020, and U.S. Provisional Patent Application No. 63 / 135,486, filed January 8, 2021, both of which are incorporated herein in their entirety for all purposes. [Background technology]
[0002] The modal size of circulating free DNA during pregnancy has been reported to be approximately 166 bp (Lo et al. Sci Transl Med. 2010;2:61ra91). There is little published data on fragments larger than 600 bp. One example is a study by Amicucci et al., who reported PCR amplification of an 8 kb fragment from the Y chromosome-derived basic protein Y2 gene (BPY2) from maternal plasma (Amicucci et al. Clin Chem 2000;40:301-2). It is unclear whether such data can be generalized across genomes. Indeed, there are many challenges in detecting long DNA fragments, such as those >600 bp, using massively parallel short-read sequencing technologies, e.g., the Illumina platform (Lo et al. Sci Transl Med. 2010;2:61ra91, Fan et al. Clin Chem. 2010;56:1278-86). These challenges include: (1) the recommended size range for Illumina sequencing platforms is typically 100–300 bp (De Maio et al. Micob Genom. 2019;5(9)). (2) DNA amplification is involved in sequencing library preparation (via PCR) or sequencing cluster generation via bridge amplification on a flow cell. Such amplification processes may favor amplifying shorter DNA fragments, in part due to the fact that long DNA templates (e.g., >600 bp) require a relatively long time to complete daughter strand synthesis compared to short DNA templates (e.g., <200 bp). Therefore, those long DNA molecules whose daughter strands are not fully generated during the PCR process within the fixed time frame for these PCR processes before or during sequencing on the Illumina platform are unavailable for downstream analysis. (3) Long DNA molecules are more likely to form secondary structures that hinder amplification.(4) Using Illumina sequencing technology, libraries are denatured, diluted, and spread on a two-dimensional surface, followed by bridge amplification, so that long DNA molecules are more likely to give rise to clusters containing two or more clonal DNA molecules than short DNA molecules (Head et al. Biotechniques. 2014;56:61-4). Summary of the Invention
[0003] The methods and systems described herein involve analyzing biological samples using long cell-free DNA fragments. These long cell-free DNA fragments enable analyses that are not possible or possible with shorter cell-free DNA fragments. Methylation CpG sites and single nucleotide polymorphism (SNP) status are often used to analyze DNA fragments in biological samples. CpG sites and SNPs are typically separated from the nearest CpG site or SNP by hundreds or thousands of base pairs. Most cell-free DNA fragments in biological samples are usually less than 200 bp in length. As a result, finding two or more consecutive CpG sites or SNPs on most cell-free DNA fragments is unlikely or impossible. Cell-free DNA fragments longer than 200 bp, including those longer than 600 bp or 1 kb, may contain multiple CpG sites and / or SNPs. The presence of multiple CpG sites and / or SNPs on long cell-free DNA fragments may enable more efficient and / or accurate analysis than with short cell-free DNA fragments alone. Long cell-free DNA fragments can be used to identify tissue of origin and / or provide information about the fetus in pregnant women. Furthermore, the accurate analysis of samples from pregnant women using long cell-free DNA fragments is surprising, since such long cell-free DNA fragments are expected to be primarily maternal in origin. Long cell-free DNA fragments of fetal origin are not expected to be present in sufficient quantities to provide information about the fetus.
[0004] Long cell-free DNA fragments containing SNPs can be used to determine the haplotype inherited by the fetus. Long cell-free DNA fragments may have multiple CpG sites, resulting in a methylation pattern that indicates the tissue of origin. In addition, trinucleotide repeats and other repetitive sequences may be present on long cell-free DNA fragments. These repeats can be used to determine the likelihood of a genetic disorder in the fetus or the father of the fetus. The amount of long cell-free DNA fragments can be used to determine gestational age. Similarly, the motifs at the ends of long cell-free DNA fragments can also be used to determine gestational age. Long cell-free DNA fragments (including, for example, the amount, length distribution, genomic location, methylation status, etc. of such fragments) can be used to determine pregnancy-related disorders.
[0005] These and other embodiments of the present disclosure are described in detail below. For example, other embodiments are directed to systems, devices, and computer-readable media associated with the methods described herein.
[0006] A better understanding of the nature and advantages of embodiments of the present disclosure may be obtained with reference to the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0007] [Figure 1A] 1 shows the size distribution of cell-free DNA determined according to an embodiment of the present invention: (A) 0-20 kb on a linear scale, (B) 0-20 kb on a logarithmic scale. [Figure 1B] 1 shows the size distribution of cell-free DNA determined according to an embodiment of the present invention: (A) 0-20 kb on a linear scale, (B) 0-20 kb on a logarithmic scale. [Figure 2A] 1 shows the size distribution of cell-free DNA determined according to an embodiment of the present invention: (A) 0-5 kb on a linear scale on the y-axis; (B) 0-5 kb on a logarithmic scale on the y-axis. [Figure 2B]1 shows the size distribution of cell-free DNA determined according to an embodiment of the present invention: (A) 0-5 kb on a linear scale on the y-axis; (B) 0-5 kb on a logarithmic scale on the y-axis. [Figure 3A] 1 shows the size distribution of cell-free DNA determined according to an embodiment of the present invention: (A) 0-400 bp on a linear scale on the y-axis; (B) 0-400 bp on a logarithmic scale on the y-axis. [Figure 3B] 1 shows the size distribution of cell-free DNA determined according to an embodiment of the present invention: (A) 0-400 bp on a linear scale on the y-axis; (B) 0-400 bp on a logarithmic scale on the y-axis. [Figure 4A] Figure 1 shows the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific) determined according to an embodiment of the present invention. (A) Linear scale of the y-axis, 0-20 kb. (B) Logarithmic scale of the y-axis, 0-20 kb. Blue lines indicate fragments carrying shared alleles (predominant of maternal origin), and red lines indicate fragments carrying fetal-specific alleles (of placental origin). [Figure 4B] Figure 1 shows the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific) determined according to an embodiment of the present invention. (A) Linear scale of the y-axis, 0-20 kb. (B) Logarithmic scale of the y-axis, 0-20 kb. Blue lines indicate fragments carrying shared alleles (predominant of maternal origin), and red lines indicate fragments carrying fetal-specific alleles (of placental origin). [Figure 5A] Figure 1 shows the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific) determined according to an embodiment of the present invention. (A) Linear scale of the y-axis, 0-5 kb. (B) Logarithmic scale of the y-axis, 0-5 kb. Blue lines indicate fragments carrying shared alleles (predominant of maternal origin), and red lines indicate fragments carrying fetal-specific alleles (of placental origin). [Figure 5B]Figure 1 shows the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific) determined according to an embodiment of the present invention. (A) Linear scale of the y-axis, 0-5 kb. (B) Logarithmic scale of the y-axis, 0-5 kb. Blue lines indicate fragments carrying shared alleles (predominant of maternal origin), and red lines indicate fragments carrying fetal-specific alleles (of placental origin). [Figure 6A] Figure 1 shows the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific) determined according to an embodiment of the present invention. (A) Linear scale of the y-axis, 0-1 kb. (B) Logarithmic scale of the y-axis, 0-1 kb. Blue lines indicate fragments carrying shared alleles (predominant of maternal origin), and red lines indicate fragments carrying fetal-specific alleles (of placental origin). [Figure 6B] Figure 1 shows the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific) determined according to an embodiment of the present invention. (A) Linear scale of the y-axis, 0-1 kb. (B) Logarithmic scale of the y-axis, 0-1 kb. Blue lines indicate fragments carrying shared alleles (predominant of maternal origin), and red lines indicate fragments carrying fetal-specific alleles (of placental origin). [Figure 7A] Figure 1 shows the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific) determined according to an embodiment of the present invention. (A) Linear scale of the y-axis, 0-400 bp. (B) Logarithmic scale of the y-axis, 0-400 bp. Blue lines indicate fragments carrying shared alleles (predominant of maternal origin), and red lines indicate fragments carrying fetal-specific alleles (of placental origin). [Figure 7B]Figure 1 shows the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific) determined according to an embodiment of the present invention. (A) Linear scale of the y-axis, 0-400 bp. (B) Logarithmic scale of the y-axis, 0-400 bp. Blue lines indicate fragments carrying shared alleles (predominant of maternal origin), and red lines indicate fragments carrying fetal-specific alleles (of placental origin). [Figure 8] 1 shows single molecule, double-stranded DNA methylation levels between fragments carrying maternal-specific alleles and fragments carrying fetal-specific alleles, according to an embodiment of the present invention. [Figure 9A] 1A-1C show (A) the matched distribution of single-molecule, double-stranded DNA methylation levels between fragments carrying maternal-specific alleles and fragments carrying fetal-specific alleles, and (B) receiver operating characteristic (ROC) analysis using single-molecule, double-stranded DNA methylation levels, according to an embodiment of the present invention. [Figure 9B] 1A-1C show (A) the matched distribution of single-molecule, double-stranded DNA methylation levels between fragments carrying maternal-specific alleles and fragments carrying fetal-specific alleles, and (B) receiver operating characteristic (ROC) analysis using single-molecule, double-stranded DNA methylation levels, according to an embodiment of the present invention. [Figure 10A] 1 shows the correlation between single molecule, double-stranded DNA methylation levels and fragment size of plasma DNA according to an embodiment of the present invention. (A) Size range of 0-20 kb. (B) Size range of 0-1 kb. [Figure 10B] 1 shows the correlation between single molecule, double-stranded DNA methylation levels and fragment size of plasma DNA according to an embodiment of the present invention. (A) Size range of 0-20 kb. (B) Size range of 0-1 kb. [Figure 11A]1 shows an example of a long fetal-specific DNA molecule identified in maternal plasma DNA of a pregnant woman, according to embodiments of the present invention. (A) The black bar indicates the long fetal-specific DNA molecule aligned to a region in chromosome 10 of the human reference genome. (B) A detailed view of the genes and epigenetics determined using PacBio sequencing according to the present disclosure. The bases highlighted in yellow (marked with an arrow) are likely due to sequence errors that can be corrected in some embodiments. [Figure 11B] 1 shows an example of a long fetal-specific DNA molecule identified in maternal plasma DNA of a pregnant woman, according to embodiments of the present invention. (A) The black bar indicates the long fetal-specific DNA molecule aligned to a region in chromosome 10 of the human reference genome. (B) A detailed view of the genes and epigenetics determined using PacBio sequencing according to the present disclosure. The bases highlighted in yellow (marked with an arrow) are likely due to sequence errors that can be corrected in some embodiments. [Figure 12A] 1 shows an example of a long maternal DNA molecule carrying shared alleles identified in maternal plasma DNA of a pregnant woman, according to an embodiment of the present invention. (A) The black bar indicates the long maternal-specific DNA molecule aligned to a region in chromosome 6 of the human reference. (B) A detailed diagram of the genetic and epigenetic information determined using PacBio sequencing according to an embodiment of the present invention. [Figure 12B] 1 shows an example of a long maternal DNA molecule carrying shared alleles identified in maternal plasma DNA of a pregnant woman, according to an embodiment of the present invention. (A) The black bar indicates the long maternal-specific DNA molecule aligned to a region in chromosome 6 of the human reference. (B) A detailed diagram of the genetic and epigenetic information determined using PacBio sequencing according to an embodiment of the present invention. [Figure 13] 1 shows frequency distributions for DNA from placenta (red) and maternal blood cells (blue) according to methylation levels at different resolutions from 1 kb to 20 kb, according to an embodiment of the present invention. [Figure 14A]1 shows frequency distributions for DNA from placenta (red) and maternal blood cells (blue) according to methylation levels within 16 kb and 24 kb windows, according to embodiments of the present invention. [Figure 14B] 1 shows frequency distributions for DNA from placenta (red) and maternal blood cells (blue) according to methylation levels within 16 kb and 24 kb windows, according to embodiments of the present invention. [Figure 15A] 1 shows an example of a long maternal-specific DNA molecule identified in maternal plasma DNA of a pregnant woman, according to an embodiment of the present invention. (A) The black bar indicates the long maternal-specific DNA molecule aligned to a region in chromosome 8 of a human reference. (B) A detailed diagram of the genes and epigenetics determined using PacBio sequencing according to an embodiment of the present invention. [Figure 15B] 1 shows an example of a long maternal-specific DNA molecule identified in maternal plasma DNA of a pregnant woman, according to an embodiment of the present invention. (A) The black bar indicates the long maternal-specific DNA molecule aligned to a region in chromosome 8 of a human reference. (B) A detailed diagram of the genes and epigenetics determined using PacBio sequencing according to an embodiment of the present invention. [Figure 16] 1 shows a diagram for estimating maternal inheritance of a fetus according to an embodiment of the present invention. [Figure 17] 1 illustrates the determination of genetic / epigenetic disorders in plasma DNA molecules using information of maternal and fetal origin according to an embodiment of the present invention. [Figure 18] 1 illustrates the identification of fetal abnormality fragments according to an embodiment of the present invention. [Figure 19A] 1 shows a diagram of error correction for cell-free DNA genotyping using PacBio sequencing, according to an embodiment of the invention. "." represents a base identical to the reference base in the Watson strand. "," represents a base identical to the reference base in the Crick strand. "Alphabetical letters" represent alternative alleles that differ from the reference allele. "*" represents an insertion. "^" represents a deletion. [Figure 19B]1 shows a diagram of error correction for cell-free DNA genotyping using PacBio sequencing, according to an embodiment of the invention. "." represents a base identical to the reference base in the Watson strand. "," represents a base identical to the reference base in the Crick strand. "Alphabetical letters" represent alternative alleles that differ from the reference allele. "*" represents an insertion. "^" represents a deletion. [Figure 19C] 1 shows a diagram of error correction for cell-free DNA genotyping using PacBio sequencing, according to an embodiment of the invention. "." represents a base identical to the reference base in the Watson strand. "," represents a base identical to the reference base in the Crick strand. "Alphabetical letters" represent alternative alleles that differ from the reference allele. "*" represents an insertion. "^" represents a deletion. [Figure 19D] 1 shows a diagram of error correction for cell-free DNA genotyping using PacBio sequencing, according to an embodiment of the invention. "." represents a base identical to the reference base in the Watson strand. "," represents a base identical to the reference base in the Crick strand. "Alphabetical letters" represent alternative alleles that differ from the reference allele. "*" represents an insertion. "^" represents a deletion. [Figure 19E] 1 shows a diagram of error correction for cell-free DNA genotyping using PacBio sequencing, according to an embodiment of the invention. "." represents a base identical to the reference base in the Watson strand. "," represents a base identical to the reference base in the Crick strand. "Alphabetical letters" represent alternative alleles that differ from the reference allele. "*" represents an insertion. "^" represents a deletion. [Figure 19F] 1 shows a diagram of error correction for cell-free DNA genotyping using PacBio sequencing, according to an embodiment of the invention. "." represents a base identical to the reference base in the Watson strand. "," represents a base identical to the reference base in the Crick strand. "Alphabetical letters" represent alternative alleles that differ from the reference allele. "*" represents an insertion. "^" represents a deletion. [Figure 19G]1 shows a diagram of error correction for cell-free DNA genotyping using PacBio sequencing, according to an embodiment of the invention. "." represents a base identical to the reference base in the Watson strand. "," represents a base identical to the reference base in the Crick strand. "Alphabetical letters" represent alternative alleles that differ from the reference allele. "*" represents an insertion. "^" represents a deletion. [Figure 20] 1 illustrates a method for analyzing a biological sample obtained from a woman carrying a fetus, according to an embodiment of the present invention. [Figure 21] 1 illustrates a method for analyzing a biological sample obtained from a woman carrying a fetus to determine haplotype inheritance, according to an embodiment of the present invention. [Figure 22] 1 shows methylation patterns for determining the tissue of origin of long DNA molecules in plasma according to an embodiment of the present invention. [Figure 23] 1 shows receiver operating characteristic (ROC) curves for determining fetal and maternal origin according to an embodiment of the present invention. [Figure 24] 1 shows pairwise methylation patterns according to an embodiment of the present invention. [Figure 25] 1 is a table of distribution of selected marker regions among different chromosomes according to an embodiment of the present invention. [Figure 26] 1 is a table of plasma DNA molecules based on single molecule methylation patterns using different percentages of buffy coat DNA molecules with a mismatch score greater than 0.3 as a selection criterion for marker regions, according to an embodiment of the present invention. [Figure 27] 1 illustrates a process flow for determining fetal inheritance in a non-invasive manner using placenta-specific methylation haplotypes, according to an embodiment of the present invention. [Figure 28] 1 illustrates the principle of non-invasive prenatal detection of fragile X syndrome using long cell-free DNA in maternal plasma, according to an embodiment of the present invention. [Figure 29] 1 illustrates maternal inheritance of a fetus based on methylation patterns, according to an embodiment of the present invention. [Figure 30]1 illustrates the qualitative analysis of maternal inheritance of a fetus using genetic and epigenetic information from plasma DNA molecules, according to an embodiment of the present invention. [Figure 31] 1 shows the detection rate of qualitative analysis of fetal maternal inheritance in a genome-wide method using genetic and epigenetic information from plasma DNA molecules compared to relative haplotype dosage (RHDO) analysis, according to an embodiment of the present invention. [Figure 32] 1 shows the relationship between the detection rate of paternal-specific variants in genome-wide methods and the number of sequenced plasma DNA molecules with different sizes used in the analysis, according to an embodiment of the present invention. [Figure 33] 1 illustrates a workflow for non-invasive detection of fragile X syndrome, according to an embodiment of the present invention. [Figure 34] 1 shows the methylation pattern of plasma DNA compared to the methylation profiles of placenta and buffy coat DNA, according to an embodiment of the present invention. [Figure 35] 1 is a table showing the distribution of CpG sites within 500 bp regions across the human genome, according to an embodiment of the present invention. [Figure 36] 1 is a table showing the distribution of CpG sites within 1 kb regions across the human genome, according to an embodiment of the present invention. [Figure 37] 1 is a table showing the distribution of CpG sites within 3 kb regions across the human genome, according to an embodiment of the present invention. [Figure 38] 1 is a table showing the proportional contribution of DNA molecules from different tissues in maternal plasma using methylation status matching analysis, according to an embodiment of the present invention. [Figure 39A] 1 shows the relationship between placental contribution and fetal DNA fraction estimated by the SNP approach, according to an embodiment of the present invention. [Figure 39B] 1 shows the relationship between placental contribution and fetal DNA fraction estimated by the SNP approach, according to an embodiment of the present invention. [Figure 40]1 illustrates a method for analyzing a biological sample obtained from a woman carrying a fetus to determine tissue of origin using methylation pattern analysis, according to an embodiment of the present invention. [Figure 41A] 1 shows the size distribution of cell-free DNA molecules from first, second, and third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 41B] 1 shows the size distribution of cell-free DNA molecules from first, second, and third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 42] 1 is a table showing the percentage of long plasma DNA molecules at different gestational stages, according to an embodiment of the invention. [Figure 43A] 1 shows the size distribution of DNA molecules covering fetal-specific alleles from maternal plasma at first, second, and third trimesters, according to an embodiment of the present invention. [Figure 43B] 1 shows the size distribution of DNA molecules covering fetal-specific alleles from maternal plasma at first, second, and third trimesters, according to an embodiment of the present invention. [Figure 44A] 1 shows the size distribution of DNA molecules covering maternal-specific alleles from maternal plasma at first, second, and third trimesters, according to an embodiment of the present invention. [Figure 44B] 1 shows the size distribution of DNA molecules covering maternal-specific alleles from maternal plasma at first, second, and third trimesters, according to an embodiment of the present invention. [Figure 45] 1 is a table of the proportion of long fetal and maternal plasma DNA molecules at different gestational stages, according to an embodiment of the invention. [Figure 46A] 1 shows a plot of the percentage of fetal-specific plasma DNA fragments in particular size ranges across different gestational stages, according to an embodiment of the present invention. [Figure 46B] 1 shows a plot of the percentage of fetal-specific plasma DNA fragments in particular size ranges across different gestational stages, according to an embodiment of the present invention. [Figure 46C] 1 shows a plot of the percentage of fetal-specific plasma DNA fragments in particular size ranges across different gestational stages, according to an embodiment of the present invention. [Figure 47A] 1 shows a graph of the percentage base content of the 5′ end of cell-free DNA molecules from first, second, and third trimester maternal plasma across a range of fragment sizes from 0 to 3 kb, according to an embodiment of the present invention. [Figure 47B] 1 shows a graph of the percentage base content of the 5′ end of cell-free DNA molecules from first, second, and third trimester maternal plasma across a range of fragment sizes from 0 to 3 kb, according to an embodiment of the present invention. [Figure 47C] 1 shows a graph of the percentage base content of the 5′ end of cell-free DNA molecules from first, second, and third trimester maternal plasma across a range of fragment sizes from 0 to 3 kb, according to an embodiment of the present invention. [Figure 48] 1 is a table of the percentage of terminal nucleotide bases among short and long cell-free DNA molecules from maternal plasma in first, second, and third trimesters, according to an embodiment of the present invention. [Figure 49] 1 is a table of the percentage of terminal nucleotide bases among short and long cell-free DNA molecules covering fetal-specific alleles from maternal plasma at first, second, and third trimesters, according to an embodiment of the present invention. [Figure 50] 1 is a table of the percentage of terminal nucleotide bases among short and long cell-free DNA molecules covering maternal-specific alleles from first-, second-, and third-trimester maternal plasma, according to an embodiment of the present invention. [Figure 51] 1 shows a hierarchical clustering analysis of short and long plasma cell-free DNA molecules using 256 terminal motifs, according to an embodiment of the present invention. [Figure 52A] 1 shows a principal component analysis of 4mer terminal motif profiles according to an embodiment of the present invention. [Figure 52B] 1 shows a principal component analysis of 4mer terminal motif profiles according to an embodiment of the present invention. [Figure 53] 1 is a table of the 25 most frequent terminal motifs among short plasma DNA molecules from maternal plasma during early pregnancy, according to an embodiment of the present invention. [Figure 54] 1 is a table of the 25 most frequent terminal motifs among short plasma DNA molecules from mid-trimester maternal plasma, according to an embodiment of the present invention. [Figure 55] 1 is a table of the 25 most frequent terminal motifs among short plasma DNA molecules from maternal plasma in late pregnancy, according to an embodiment of the present invention. [Figure 56] 1 is a table of the 25 most frequent terminal motifs among long plasma DNA molecules from maternal plasma during early pregnancy, according to an embodiment of the present invention. [Figure 57] 1 is a table of the 25 most frequent terminal motifs among long plasma DNA molecules from mid-trimester maternal plasma, according to an embodiment of the present invention. [Figure 58] 1 is a table of the 25 most frequent terminal motifs among long plasma DNA molecules from maternal plasma in late pregnancy, according to an embodiment of the present invention. [Figure 59A] 1A-C show scatter plots of motif frequencies of 16 NNXY motifs among short and long plasma DNA molecules in maternal plasma during (A) first trimester, (B) second trimester, and (C) third trimester, according to embodiments of the present invention. [Figure 59B] 1A-C show scatter plots of motif frequencies of 16 NNXY motifs among short and long plasma DNA molecules in maternal plasma during (A) first trimester, (B) second trimester, and (C) third trimester, according to embodiments of the present invention. [Figure 59C] 1A-C show scatter plots of motif frequencies of 16 NNXY motifs among short and long plasma DNA molecules in maternal plasma during (A) first trimester, (B) second trimester, and (C) third trimester, according to embodiments of the present invention. [Figure 60] 1 illustrates a method for analyzing a biological sample obtained from a woman carrying a fetus to determine gestational age, according to an embodiment of the present invention. [Figure 61] 1 illustrates a method for analyzing a biological sample obtained from a woman carrying a fetus to determine the likelihood of a pregnancy-related disorder, according to an embodiment of the present invention. [Figure 62] 1 is a table showing clinical information for four pre-eclampsia cases, according to an embodiment of the present invention. [Figure 63A]1 is a graph of the size distribution of cell-free DNA molecules from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 63B] 1 is a graph of the size distribution of cell-free DNA molecules from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 63C] 1 is a graph of the size distribution of cell-free DNA molecules from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 63D] 1 is a graph of the size distribution of cell-free DNA molecules from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 64A] 1 is a graph of the size distribution of cell-free DNA molecules from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 64B] 1 is a graph of the size distribution of cell-free DNA molecules from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 64C] 1 is a graph of the size distribution of cell-free DNA molecules from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 64D] 1 is a graph of the size distribution of cell-free DNA molecules from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 65A] 1 is a graph of the size distribution of DNA molecules covering fetal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 65B] 1 is a graph of the size distribution of DNA molecules covering fetal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 65C] 1 is a graph of the size distribution of DNA molecules covering fetal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 65D]1 is a graph of the size distribution of DNA molecules covering fetal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 66A] 1 is a graph of the size distribution of DNA molecules covering fetal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 66B] 1 is a graph of the size distribution of DNA molecules covering fetal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 66C] 1 is a graph of the size distribution of DNA molecules covering fetal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 66D] 1 is a graph of the size distribution of DNA molecules covering fetal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 67A] 1 is a graph of the size distribution of DNA molecules covering maternal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 67B] 1 is a graph of the size distribution of DNA molecules covering maternal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 67C] 1 is a graph of the size distribution of DNA molecules covering maternal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 67D] 1 is a graph of the size distribution of DNA molecules covering maternal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 68A] 1 is a graph of the size distribution of DNA molecules covering maternal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 68B]1 is a graph of the size distribution of DNA molecules covering maternal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 68C] 1 is a graph of the size distribution of DNA molecules covering maternal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 68D] 1 is a graph of the size distribution of DNA molecules covering maternal-specific alleles from pre-eclamptic and normotensive third trimester maternal plasma samples, according to an embodiment of the present invention. [Figure 69A] 1 is a graph of the percentage of short DNA molecules covering fetal-specific and maternal-specific alleles in pre-eclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing, according to an embodiment of the present invention. [Figure 69B] 1 is a graph of the percentage of short DNA molecules covering fetal-specific and maternal-specific alleles in pre-eclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing, according to an embodiment of the present invention. [Figure 70A] 1 is a graph of the percentage of short DNA molecules in pre-eclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing and Illumina sequencing, according to an embodiment of the present invention. [Figure 70B] 1 is a graph of the percentage of short DNA molecules in pre-eclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing and Illumina sequencing, according to an embodiment of the present invention. [Figure 71] 1 is a size ratio graph showing the relative proportions of short and long DNA molecules in pre-eclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing, according to an embodiment of the present invention. [Figure 72A] 1 shows the percentage of different ends of plasma DNA molecules in pre-eclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing, according to an embodiment of the present invention. [Figure 72B] 1 shows the percentage of different ends of plasma DNA molecules in pre-eclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing, according to an embodiment of the present invention. [Figure 72C] 1 shows the percentage of different ends of plasma DNA molecules in pre-eclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing, according to an embodiment of the present invention. [Figure 72D] 1 shows the percentage of different ends of plasma DNA molecules in pre-eclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing, according to an embodiment of the present invention. [Figure 73] 1 shows a hierarchical clustering analysis of pre-eclamptic and normotensive third trimester maternal plasma DNA samples using the frequency of plasma DNA molecules with each of the four types of fragment ends (the first nucleotide at the 5′ end of each strand): C-terminus, G-terminus, T-terminus, and A-terminus, according to an embodiment of the present invention. [Figure 74] 1 shows a hierarchical clustering analysis of pre-eclamptic and normotensive third trimester maternal plasma DNA samples using 16 di-nucleotide motifs XYNN (dinucleotide sequence of the first and second nucleotides from the 5′ end) according to an embodiment of the present invention. [Figure 75] 1 shows a hierarchical clustering analysis of pre-eclamptic and normotensive third trimester maternal plasma DNA samples using 16 dinucleotide motifs NNXY (dinucleotide sequence of the third and fourth nucleotides from the 5' end) according to an embodiment of the present invention. [Figure 76] 1 shows a hierarchical clustering analysis of pre-eclamptic and normotensive third trimester maternal plasma DNA samples using 256 4-nucleotide motifs (dinucleotide sequences of the first to fourth nucleotides from the 5' end) according to an embodiment of the present invention. [Figure 77A] 1 shows the contribution of T cells between four types of fragment ends in pre-eclamptic and normotensive maternal plasma DNA samples, according to an embodiment of the present invention. [Figure 77B]1 shows the contribution of T cells between four types of fragment ends in pre-eclamptic and normotensive maternal plasma DNA samples, according to an embodiment of the present invention. [Figure 77C] 1 shows the contribution of T cells between four types of fragment ends in pre-eclamptic and normotensive maternal plasma DNA samples, according to an embodiment of the present invention. [Figure 77D] 1 shows the contribution of T cells between four types of fragment ends in pre-eclamptic and normotensive maternal plasma DNA samples, according to an embodiment of the present invention. [Figure 78] 1 illustrates a method of analyzing a biological sample obtained from a woman carrying a fetus to determine the likelihood of a pregnancy-related disorder, according to an embodiment of the present invention. [Figure 79] FIG. 1 shows a diagram for estimating maternal inheritance of a fetus for a repeat-associated disease according to an embodiment of the present invention. [Figure 80] FIG. 1 shows a diagram for estimating fetal paternal inheritance for a repeat-associated disease according to an embodiment of the present invention. [Figure 81] 1 is a table showing examples of repeat expansion diseases. [Figure 82] 1 is a table showing examples of repeat expansion diseases. [Figure 83] 1 is a table showing examples of repeat expansion diseases. [Figure 84] 1 is a table showing examples of repeat expansion detection and repeat-associated methylation determination in a fetus according to an embodiment of the present invention. [Figure 85] 1 illustrates a method of analyzing a biological sample obtained from a woman carrying a fetus to determine the likelihood of a genetic disorder in the fetus, according to an embodiment of the present invention. [Figure 86] 1 illustrates a method for analyzing a biological sample obtained from a woman carrying a fetus to determine paternity, according to an embodiment of the present invention. [Figure 87] Methylation patterns for two representative plasma DNA molecules after size selection are shown. [Figure 88] 1 is a table of sequencing information for samples with and without size selection, according to an embodiment of the present invention. [Figure 89A]1 shows a graph of plasma DNA size profiles for samples with and without bead-based size selection, according to an embodiment of the present invention. [Figure 89B] 1 shows a graph of plasma DNA size profiles for samples with and without bead-based size selection, according to an embodiment of the present invention. [Figure 90A] 1 shows the size profile between fetal and maternal DNA molecules in a sample with size selection, according to an embodiment of the present invention. [Figure 90B] 1 shows the size profile between fetal and maternal DNA molecules in a sample with size selection, according to an embodiment of the present invention. [Figure 91] 1 is a statistical table of the number of plasma DNA molecules carrying informative SNPs between samples with and without size selection, according to an embodiment of the present invention. [Figure 92] 1 is a table of methylation levels in size-selected and non-size-selected plasma DNA samples according to an embodiment of the present invention. [Figure 93] 1 is a table of methylation levels of maternal or fetal-specific cell-free DNA molecules, according to embodiments of the present invention. [Figure 94] 1 is a table of the top 10 terminal motifs in samples with and without size selection, according to an embodiment of the present invention. [Figure 95] 1 is a receiver operating characteristic (ROC) graph showing that long plasma DNA molecules enhance the performance of tissue-of-origin analysis, according to an embodiment of the present invention. [Figure 96] 1 illustrates the principle of airport sequencing for plasma DNA molecules according to an embodiment of the present invention. [Figure 97] 1 is a table of the percentage of plasma DNA molecules within a particular size range and their corresponding methylation levels, according to an embodiment of the present invention. [Figure 98] 1 is a graph of size distribution and methylation patterns across different sizes, according to an embodiment of the present invention. [Figure 99]1 is a table of fetal DNA fractions determined using nanopore sequencing, according to an embodiment of the present invention. [Figure 100] 1 is a table of methylation levels between fetal-specific and maternal-specific DNA molecules according to an embodiment of the present invention. [Figure 101] 1 is a table of the percentage of plasma DNA molecules within specific size ranges for fetal and maternal DNA molecules and their corresponding methylation levels, according to an embodiment of the present invention. [Figure 102A] 1 is a graph of the size distribution of fetal and maternal DNA molecules determined by nanopore sequencing, according to an embodiment of the present invention. [Figure 102B] 1 is a graph of the size distribution of fetal and maternal DNA molecules determined by nanopore sequencing, according to an embodiment of the present invention. [Figure 103] 1 is a graph showing the difference in methylation levels between fetal and maternal DNA molecules based on a single informative SNP and two informative SNPs, according to an embodiment of the present invention. [Figure 104] 1 is a table of differences in methylation levels between fetal and maternal DNA molecules, according to an embodiment of the present invention. [Figure 105] 1 illustrates a measurement system according to an embodiment of the present invention. [Figure 106] 1 illustrates a computer system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0008] term A "tissue" corresponds to a group of cells that group together as a functional unit in a pregnant subject or its fetus. Two or more types of cells can be found within a single tissue. Different types of tissues can consist of different types of cells (e.g., liver cells, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (mother vs. fetus, tissues of a pregnant subject that has received a transplant, tissues of a pregnant organism or its fetus that has been infected with a microorganism or virus). A "reference tissue" can correspond to the tissue used to determine tissue-specific methylation levels. Multiple samples of the same tissue type from different pregnant individuals or their fetuses can be used to determine the tissue-specific methylation level of that tissue type.
[0009] A "biological sample" refers to any sample obtained from a pregnant subject (e.g., a human (or other animal), such as a pregnant woman, a pregnant person with or suspected of having a disorder, a pregnant organ transplant recipient, or a pregnant subject suspected of having a disease process involving an organ (e.g., the heart in myocardial infarction, the brain in stroke, or the hematopoietic system in anemia)) and containing one or more nucleic acid molecules of interest. A biological sample can be a body fluid, such as blood, plasma, serum, urine, vaginal fluid, vaginal wash, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirates from different parts of the body (e.g., thyroid, mammary gland), intraocular fluid (e.g., aqueous humor), etc. Stool samples can also be used. In various embodiments, the majority of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) may be cell-free, e.g., more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA may be cell-free. The centrifugation protocol may include, for example, obtaining a fluid portion at 3,000 g for 10 minutes and then recentrifuging at 30,000 g for an additional 10 minutes to remove residual cells. As part of the analysis of the biological sample, a statistically significant number of cell-free DNA molecules may be analyzed for the biological sample (e.g., to provide an accurate measurement). In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000, or more cell-free DNA molecules may be analyzed. At least the same number of sequence reads may be analyzed.
[0010] A "sequence read" refers to a chain of nucleotides sequenced from any portion or all of a nucleic acid molecule. For example, a sequence read can be a short chain of nucleotides (e.g., about 20-150 nucleotides) sequenced from a nucleic acid fragment, a short chain of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained, for example, using sequencing techniques or in a variety of ways using probes, for example, capture probes such as those used in hybridization arrays or microarrays, or amplification techniques such as polymerase chain reaction (PCR) or linear amplification using single primers or isothermal amplification. As part of the analysis of a biological sample, a statistically significant number of sequence reads can be analyzed, for example, at least 1,000 sequence reads can be analyzed. As other examples, at least 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000, or more sequence reads can be analyzed.
[0011] A "site" (also called a "genomic site") corresponds to a single site, which may be a single base position, or a group of correlated base positions, e.g., a CpG site, or a larger group of correlated base positions. A "locus" may correspond to a region that includes multiple sites. A locus may contain only one site, which would make the locus equivalent to the site in that context.
[0012] "Methylation state" refers to the state of methylation at a given site, e.g., a site may be either methylated, unmethylated, or possibly undetermined.
[0013] A "methylation index" for each genomic site (e.g., a CpG site) can refer to the proportion of DNA fragments (e.g., as determined from sequence reads or probes) that exhibit methylation at that site across the total number of reads covering that site. A "read" can correspond to information obtained from a DNA fragment (e.g., the methylation state of a site). A read can be obtained using a reagent (e.g., a primer or probe) that preferentially hybridizes to DNA fragments of a particular methylation state at one or more sites. Typically, such a reagent is applied after processing DNA molecules with a process that differentially modifies or recognizes them according to their methylation state, such as bisulfite conversion, or methylation-sensitive restriction enzymes, or methylation-binding proteins, or anti-methylcytosine antibodies, or single-molecule sequencing technologies that recognize methylcytosine and hydroxymethylcytosine (e.g., single-molecule real-time sequencing and nanopore sequencing (e.g., from Oxford Nanopore Technologies)).
[0014] The "methylation density" of a region may refer to the number of reads at sites within the region that show methylation divided by the total number of reads that cover sites in this region. This site may have specific characteristics, such as a CpG site. Thus, the "CpG methylation density" of a region refers to the number of reads that show CpG methylation divided by the total number of reads that cover CpG sites in this region (e.g., a specific CpG site, a CpG site within a CpG island, or a larger region). For example, the methylation density of each 100 kb bin in the human genome can be determined from the total number of unconverted cytosines (corresponding to methylated cytosines) after bisulfite treatment of CpG sites as a percentage of all CpG sites covered by sequence reads mapped to the 100 kb region. This analysis can also be performed for other bin sizes, such as 500 bp, 5 kb, 10 kb, 50 kb, or 1 Mb. A region may be the entire genome, or a chromosome, or a portion of a chromosome (e.g., a chromosome arm). The methylation index of a CpG site is the same as the methylation density of that region if the region contains only that CpG site. "Percentage of methylated cytosines" can refer to the number of cytosine sites "C" that are shown to be methylated (e.g., unconverted after bisulfite conversion) relative to the total number of cytosine residues analyzed in the region, i.e., including cytosines outside of CpG contexts. Examples of "methylation levels" include the methylation index, methylation density, the number of molecules methylated at one or more sites, and the percentage of molecules (e.g., cytosines) methylated at one or more sites.Aside from bisulfite conversion, other processes known to those skilled in the art can be used to examine the methylation state of DNA molecules, including, but not limited to, methylation-state-sensitive enzymes (e.g., methylation-sensitive restriction enzymes), methylation-binding proteins, single-molecule sequencing using methylation-state-sensitive platforms (e.g., nanopore sequencing (Schreiber et al. Proc Natl Acad Sci 2013;110:18910-18915), and single-molecule real-time sequencing (e.g., by Pacific Biosciences) (Flusberg et al. Nat Methods 2010;7:461-465)).
[0015] A "methylome" provides a measure of the amount of DNA methylation at multiple sites or loci in a genome. The methylome can correspond to the entire genome, a substantial portion of the genome, or a relatively small portion of the genome.
[0016] A "methylation profile" includes information related to DNA or RNA methylation at multiple sites or regions. Information related to DNA methylation may include, but is not limited to, the methylation index of CpG sites, the methylation density (abbreviated MD) of CpG sites in a region, the distribution of CpG sites across a contiguous region, the methylation pattern or level of each individual CpG site within a region containing two or more CpG sites, and non-CpG methylation. In one embodiment, a methylation profile may include the methylation or non-methylation patterns of two or more types of bases (e.g., cytosine or adenine). The methylation profile of a substantial portion of a genome can be considered equivalent to a methylome. "DNA methylation" in mammalian genomes typically refers to the addition of a methyl group to the 5' carbon of cytosine residues between CpG dinucleotides (i.e., 5-methylcytosine). DNA methylation can occur at cytosine in other contexts, such as CHG and CHH, where H is adenine, cytosine, or thymine. Cytosine methylation can also occur in the form of 5-hydroxymethylcytosine. 6Non-cytosine methylations such as -methyladenine have also been reported.
[0017] A "methylation pattern" refers to the order of methylated and unmethylated bases. For example, a methylation pattern can be the order of methylated bases on a single DNA strand, a single double-stranded DNA molecule, or another type of nucleic acid molecule. As an example, three consecutive CpG sites can have any of the following methylation patterns: UUU, MMM, UMM, UMU, UUM, MUM, MUU, or MMU, where "U" indicates an unmethylated site and "M" indicates a methylated site. Without being limited thereto, when extending this concept to base modifications including methylation, the term "modification pattern" will be used to refer to the order of modified and unmodified bases. For example, a modification pattern can be the order of modified bases on a single DNA strand, a single double-stranded DNA molecule, or another type of nucleic acid molecule. As an example, three consecutive potentially modifiable sites can have any of the following modification patterns: UUU, MMM, UMM, UMU, UUM, MUM, MUU, or MMU. Here, "U" indicates an unmodified site and "M" indicates a modified site. An example of a base modification that is not based on methylation is an oxidative change such as 8-oxoguanine.
[0018] The terms "hypermethylated" and "hypomethylated" can refer to the methylation density of a single DNA molecule as measured by the methylation level of that single molecule, e.g., the number of methylated bases or nucleotides in that molecule divided by the total number of methylable bases or nucleotides in that molecule. A hypermethylated molecule is a molecule whose methylation level of a single molecule is equal to or greater than a threshold, which can be defined for each application. This threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%. A hypomethylated molecule is a molecule whose methylation level of a single molecule is equal to or less than a threshold, which can be defined for each application and can vary for each application. This threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%.
[0019] The terms "hypermethylated" and "hypomethylated" may also refer to the methylation level of a population of DNA molecules as measured by the methylation level of a plurality of those molecules. A hypermethylated population of molecules is a population in which a plurality of the molecules have a methylation level above a threshold, which may be defined for each application and may vary from application to application. This threshold may be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%. A hypomethylated population of molecules is a population in which a plurality of the molecules have a methylation level below a threshold, which may be defined for each application and may vary from application to application. This threshold may be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%. In one embodiment, the population of molecules may be aligned to one or more selected genomic regions. In one embodiment, the selected genomic region may be associated with a disease such as a genetic disorder, an imprinting disorder, a metabolic disorder, or a neurological disorder. The selected genomic region may have a length of 50 nucleotides (nt), 100 nt, 200 nt, 300 nt, 500 nt, 1000 nt, 2 knt, 5 knt, 10 knt, 20 knt, 30 knt, 40 knt, 50 knt, 60 knt, 70 knt, 80 knt, 90 knt, 100 knt, 200 knt, 300 knt, 400 knt, 500 knt, or 1 Mnt.
[0020] The term "sequencing depth" refers to the number of times a locus is covered by sequence reads aligned to that locus. A locus can be as small as a nucleotide, or as large as a chromosome arm, or as large as the entire genome. Sequencing depth can be expressed as 50x, 100x, etc., where "x" refers to the number of times a locus is covered by sequence reads. Sequencing depth can also be applied to multiple loci or the entire genome, where x can refer to the average number of times a locus or haploid genome or the entire genome is sequenced, respectively. Ultra-deep sequencing can refer to a sequencing depth of at least 100x.
[0021] A "calibration sample" may correspond to a biological sample in which the fractional concentration of clinically relevant DNA (e.g., tissue-specific DNA fraction) is known or determined via a calibration method using tissue-specific alleles, such as in the case of transplantation in a pregnant subject, where alleles present in the donor's genome but absent in the recipient's genome can be used as markers for the transplanted organ. As another example, a calibration sample may correspond to a sample in which terminal motifs can be determined. A calibration sample may be used for both purposes.
[0022] A "calibration data point" includes a "calibration value" and a measured or known fractional concentration of clinically relevant DNA (e.g., DNA of a particular tissue type). The calibration value can be determined from relative frequencies (e.g., aggregate values) determined for calibration samples in which the fractional concentrations of clinically relevant DNA are known. Calibration data points can be defined in various ways, for example, as discrete points or as a calibration function (also called a calibration curve or calibration surface). A calibration function can be derived from additional mathematical transformations of the calibration data points.
[0023] A "separation value" is a difference or ratio that encompasses two values, such as two fractional contributions or two methylation levels. A separation value can be a simple difference or ratio. For example, the direct ratio of x / y is a separation value, as is x / (x+y). A separation value can include other factors, such as multiplicative factors. As another example, a difference or ratio of a function of values can be used, such as the difference or ratio of the natural logarithms (ln) of two values. A separation value can include a difference and a ratio.
[0024] "Separation values" and "aggregate values" (e.g., relative frequencies) are two examples of parameters (also called metrics) that provide a measure of a sample's variation between different classifications (states) and can therefore be used to determine various classifications. An aggregate value can be a separation value when the difference is taken between a set of relative frequencies of a sample and a reference set of relative frequencies, as is done, for example, in clustering.
[0025] As used herein, the term "classification" refers to any number or other characteristic associated with a particular property of a sample. For example, a "+" symbol (or the word "positive") can mean that the sample is classified as having a deletion or amplification. Classification can be binary (e.g., positive or negative) or can have more levels of classification (e.g., a scale of 1 to 10 or 0 to 1).
[0026] As used herein, the term "parameter" refers to a numerical value that characterizes a quantitative data set and / or a numerical relationship between quantitative data sets. For example, the ratio (or a function of the ratio) between a first amount of a first nucleic acid sequence and a second amount of a second nucleic acid sequence is a parameter.
[0027] The term "size profile" generally relates to the size of DNA fragments in a biological sample. A size profile can be a histogram that provides a distribution of a certain amount of DNA fragments of various sizes. Various statistical parameters (also called size parameters or simply parameters) can be used to distinguish one size profile from another. One parameter is the proportion of DNA fragments of a certain size or size range relative to all DNA fragments or DNA fragments of other sizes or ranges.
[0028] The terms "cutoff" and "threshold" refer to a predetermined number used in a certain operation. For example, a cutoff size may refer to the size above which fragments are excluded. A threshold may be a value above or below which a particular classification is applied. Either of these terms may be used in any of these contexts. A cutoff or threshold may be a "reference value" or may be derived from a reference value that represents a particular classification or distinguishes between two or more classifications. Such reference values may be determined in various ways, as will be understood by those skilled in the art. For example, a metric may be determined for two different cohorts of subjects with different known classifications, and a reference value may be selected as a representative of one classification (e.g., the average) or as a value between two clusters of the metric (e.g., selected to obtain a desired sensitivity and specificity). As another example, a reference value may be determined based on statistical analysis or sample simulation. Specific values such as cutoffs, thresholds, references, etc. may be determined based on the desired accuracy (e.g., sensitivity and specificity).
[0029] "Pregnancy-related disorder" includes any disorder characterized by abnormal relative expression levels of genes in maternal and / or fetal tissues and / or by abnormal clinical characteristics in the mother and / or fetus. These disorders include preeclampsia (Kaartokallio et al. Sci Rep. 2015;5:14107, Medina-Bastidas et al. Int J Mol Sci. 2020;21:3597), intrauterine growth restriction (Faxen et al. Am J Perinatol. 1998;15:9-13, Medina-Bastidas et al. Int J Mol Sci. 2020;21:3597), invasive placentation, preterm birth (Enquobahrie et al. BMC Pregnancy Childbirth. 2009;9:56), hemolytic disease of the newborn, placental insufficiency (Kelly et al. Endocrinology. 2017;158:743-755), and hydrops fetalis (Magor et al. al. Blood. 2015;125:2405-17), fetal malformations (Slonim et al. Proc Natl Acad Sci USA. 2009;106:9425-9), HELLP syndrome (Dijk et al. J Clin Invest. 2012;122:4003-4011), systemic lupus erythematosus (Hong et al. J Exp Med. 2019;216:1154-1169), and other maternal immune disorders.
[0030] The abbreviation "bp" refers to base pairs. In some cases, "bp" may be used to indicate the length of a DNA fragment even if the DNA fragment is single-stranded and does not contain base pairs. In the context of single-stranded DNA, "bp" may be interpreted as providing the length of the strand in nucleotides.
[0031] The abbreviation "nt" refers to a nucleotide. In some cases, "nt" may be used to indicate the length of a single-stranded DNA in bases. "nt" may also be used to indicate a relative position, such as upstream or downstream of the locus being analyzed. In the case of double-stranded DNA, "nt" may still refer to the length of a single strand rather than the total number of nucleotides in the two strands, unless the context clearly dictates otherwise. In some contexts regarding technical conceptualization, data display, processing, and analysis, "nt" and "bp" may be used interchangeably.
[0032] The term "machine learning model" may include models based on predicting test data using sample data (e.g., training data) and thus may include supervised learning. Machine learning models are often developed using a computer or processor. Machine learning models may include statistical models.
[0033] The term "data analytics framework" may include algorithms and / or models that can receive data as input and then output a predictive outcome. Examples of "data analytics frameworks" include statistical models, mathematical models, machine learning models, other artificial intelligence models, and combinations thereof.
[0034] The term "real-time sequencing" can refer to techniques that involve data collection or monitoring of the reactions involved in sequencing while they are progressing. For example, real-time sequencing can involve optical monitoring or imaging of DNA polymerase incorporating new bases.
[0035] The term "subsequence" may refer to a series of bases that is less than the complete sequence corresponding to a nucleic acid molecule. For example, if the complete sequence of a nucleic acid molecule contains five or more bases, a subsequence may contain one, two, three, or four bases. In some embodiments, a subsequence may refer to a series of bases that form a unit, where the unit is repeated multiple times in tandem. Examples include a 3-nt unit or subsequence repeated at a locus associated with a trinucleotide repeat disorder, a 1-nt to 6-nt unit or subsequence repeated 5-50 times as a microsatellite, or a 10-nt to 60-nt unit or subsequence repeated 5-50 times as a microsatellite or in other genetic elements such as Alu repeats.
[0036] The term "about" or "approximately" can mean within an acceptable error range of a particular value as determined by one of ordinary skill in the art, which depends in part on the method of measuring or determining the value, i.e., the limitations of the measurement system. For example, "about" can mean within 1 or more standard deviations, as is customary in the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term "about" or "approximately" can mean within an order of magnitude, within 5-fold, or more preferably within 2-fold of a value. When a specific value is described in this application and claims, unless otherwise specified, the term "about" should be assumed to be within an acceptable error range of the particular value. The term "about" can have the meaning commonly understood by one of ordinary skill in the art. The term "about" can refer to ±10%. The term "about" can refer to ±5%.
[0037] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intervening value between the upper and lower limits of that range, down to one-tenth of the lower limit, is also specifically disclosed. Each smaller range between any stated or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within the embodiments of the present disclosure. The upper and lower limits of these smaller ranges may be independently included or excluded in the range, and each range in which either one, both, or neither limit is included in the smaller range is also encompassed within the present disclosure, subject to any specifically excluded limit in the stated range. When a stated range includes one or both limits, ranges excluding either or both of those included limits are also included in the disclosure.
[0038] Standard abbreviations may be used, such as bp: base pairs, kb: kilobase, pi: picoliter, s or sec: seconds, min: minutes, h or hr: hours, aa: amino acids, nt: nucleotides, etc.
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present disclosure, some potential exemplary methods and materials can now be described.
[0040] Analysis of cell-free DNA molecules often involves primarily short cell-free DNA fragments as a result of limitations in analytical technology. The limited ability to obtain sequence information from long DNA molecules using Illumina sequencing technology was demonstrated by recent sequencing results of mouse cell-free DNA (Serpas et al., Proc Natl Acad Sci USA. 2019;116:641-649). When using Illumina sequencing in wild-type mice, only 0.02% of the sequenced DNA molecules were within the 600-2000 bp range. Even when using single-molecule real-time (SMRT) technology from Pacific Biosciences (i.e., PacBio SMRT sequencing) to sequence DNA libraries originally prepared for Illumina sequencing, only 0.33% of the sequenced DNA molecules were within the 600-2000 bp range. These reported data suggested that the sequencing step resulted in the loss of 93% of the long DNA molecules within the range of 600 bp to 2000 bp present in the original DNA library.
[0041] Due to the limitations of PCR in amplifying long DNA molecules, we speculate that a significant proportion of long cell-free DNA molecules are lost during the DNA library preparation step. Using gel electrophoresis, Jahr et al. reported the presence of many large fragments, e.g., approximately 10,000 kilobases (Jahr et al. Cancer Res. 2001;61:1659-65). However, the bands shown in gel electrophoresis images do not readily provide sequence information for these molecules in the gel, let alone epigenetic information.
[0042] Previously, we studied cell-free DNA extracted from maternal plasma using the Oxford Nanopore Technologies sequencing platform (Cheng et al. Clin Chem. 2015;61:1305-6). We observed a very small percentage of plasma DNA longer than 1 kb (0.06%-0.3%). We hypothesized that such a low percentage could be a result of the low sequencing accuracy of this platform.
[0043] In this field of cell-free DNA, most research has focused on short DNA molecules (e.g., less than 600 bp). The characteristics of long cell-free DNA molecules, including genetic and epigenetic information, have not been investigated. The present disclosure provides a systematic method for analyzing long cell-free DNA molecules, including their genetic and epigenetic information and their clinical utility in non-invasive prenatal testing, such as, but not limited to, non-invasive detection of single gene disorders, elucidation of fetal genome (e.g., non-invasive whole fetal genome sequencing), detection of de novo mutations at the genome-wide level, and detection / monitoring of pregnancy-related disorders, such as pre-eclampsia and preterm labor.
[0044] I. Cell-free DNA size analysis Cell-free DNA samples obtained from pregnant women were sequenced and found to contain a significant proportion of long DNA fragments. Accurate sequencing of long cell-free DNA fragments was demonstrated. The size profiles of these long cell-free DNA molecules were analyzed. The abundance of fetal and maternal long cell-free DNA molecules was compared. The long cell-free DNA molecules could be accurately aligned to a reference genome. The long cell-free DNA molecules could be used to determine the inheritance of haplotypes.
[0045] One plasma DNA sample from a pregnant woman in the third trimester was analyzed using PacBio SMRT sequencing. Double-stranded cell-free DNA molecules were ligated with hairpin adapters and subjected to single-molecule read-time sequencing utilizing a zero-mode waveguide and a single polymerase molecule (Eid et al. Science. 2009;323:133-8).
[0046] We sequenced 1.1 billion subreads, of which 659.3 million could be aligned to the human reference genome (hg19). Subreads were generated from 4.6 million PacBio Single Molecular Real-Time (SMRT) sequencing wells, each of which contained at least one subread that could be aligned to the human reference genome. On average, each molecule in an SMRT well was sequenced 143 times. In this example, there were 4.5 million circular consensus sequences (CCS), suggesting 4.5 million cfDNA molecules that could be used for downstream analysis. The size of each cfDNA was determined from the CCS by counting the number of identified bases.
[0047] Figures 1A and 1B show the size distribution of cell-free DNA from 0 to 20 kb. The y-axis indicates frequency. The x-axis indicates base pair size from 0 to 20 kb on a linear scale (Figure 1A) or logarithmic scale (Figure 1B). Because sequencing was performed across the entire length of the DNA molecule, the size of each DNA molecule can be determined directly by counting the number of nucleotides in the subreads or CCS. Measuring DNA fragment size can be achieved using any sequencing platform that can read the entire length of the DNA fragment and is not limited to the use of single-molecule sequencers. For example, Sanger sequencers can read up to 800 bp. Short-read sequencing, such as with Illumina platforms, can read up to 250 bp. Single-molecule sequencers, such as those from Pacific Biosciences and Oxford Nanopore, can read up to over 10,000 bp. DNA fragment size can also be determined after alignment to a reference genome, such as the human reference genome. DNA fragment size can be determined by paired-end sequencing followed by alignment to the reference genome. Figure 1B shows a long tail pattern. Among the 4.5 million CCS, 22.5% of cfDNA was larger than 200 bp, 19.0% was larger than 300 bp, 11.8% was larger than 400 bp, 10.6% was larger than 500 bp, 8.9% was larger than 600 bp, 6.4% was larger than 1 kb, 3.5% was larger than 2 kb, 1.9% was larger than 3 kb, 0.9% was larger than 4 kb, and 0.04% was larger than 10 kb. The longest observed length in the current PacBio SMRT results was 29,804 bp.
[0048] Plasma DNA from one pregnant subject was also sequenced on an Illumina sequencing platform using a PCR-based library preparation protocol (Lun et al. Clin Chem. 2013;59:1583-94). Among 18.2 million paired-end reads, 5.3% of cell-free DNA was larger than 200 bp, 2.0% was larger than 300 bp, 0.3% was larger than 400 bp, 0.2% was larger than 500 bp, and 0.2% was larger than 600 bp (Table 1). For comparison, size profiles were analyzed by aggregating single-molecule real-time sequencing data from five pregnant subjects (i.e., a total of 4.4 million CCSs). More plasma DNA molecules larger than 600 bp were observed (28.56%) compared with their counterparts obtained by the Illumina sequencing platform (0.2%). These results suggested that PacBio SMRT sequencing may enable the realization of 143-fold longer DNA molecules (longer than 600 bp). Although there were no readouts on the Illumina sequencing platform, 4.77% of plasma DNA molecules larger than 3 kb can be obtained using single-molecule real-time sequencing.
[0049] In contrast to previous reports (Cheng et al Clin Chem. 2015;61:1305-6) that showed a very small proportion of long plasma DNA molecules greater than 1 kb (0.06%-0.3%) using the Oxford Nanopore Technologies sequencing platform, we were able to obtain 21-fold more plasma DNA (6.4%) greater than 1 kb, demonstrating that PacBio SMRT sequencing was much more efficient at obtaining sequence information from long DNA populations.
[0050] Compared to paired-end short-read sequencing, such as the Illumina sequencing platform, long-read sequencing technologies, such as PacBio SMRT technology, have many advantages in determining the characteristics (e.g., length) of long DNA fragments. For example, long reads generally allow for more accurate alignment to the human reference genome (e.g., hg19). Long-read technology also allows for accurate determination of the length of plasma DNA molecules by directly counting the number of sequenced nucleotides. In contrast, paired-end short-read-based plasma DNA size estimation is an indirect method that estimates the size of plasma DNA molecules using the outermost coordinates of aligned paired-end reads. With such an indirect approach, alignment errors can lead to accurate size estimation. In this regard, the larger the size range between paired-end reads, the greater the likelihood of alignment errors. [Table 1]
[0051] Table 1. Comparison of size distribution between PacBio and Illumina sequencing of cell-free DNA.
[0052] Figures 2A and 2B show the size distribution of cfDNA from 0 to 5 kb. The y-axis indicates frequency. The x-axis indicates base pair sizes from 0 to 5 kb on a linear scale (Figure 2A) or logarithmic scale (Figure 2B). There was a series of major peaks occurring in a periodic pattern. Such periodic patterns even extended to molecules in the 1 kb to 2 kb range. The most frequent peak (2.6%) was 166 bp, consistent with previous findings using Illumina technology (Lo et al. Sci Transl Med. 2010;2:61ra91). The distance between adjacent major peaks in Figure 2B was approximately 200 bp, suggesting that long cfDNA production also involves nucleosomal structures.
[0053] Figures 3A and 3B show the size distribution of cell-free DNA from 0 to 400 bp. The y-axis indicates frequency. The x-axis indicates base pair size from 0 to 400 bp on a linear scale (Figure 3A) or logarithmic scale (Figure 3B). The previously reported characteristic features (Lo et al. Sci Transl Med. 2010;2:61ra91), with a most prominent peak at 166 bp and a 10 bp periodicity occurring in molecules less than 166 bp, were also reproducible using the new method disclosed herein. These results suggest that molecular size determination by counting the number of bases sequenced from a single molecule according to the present disclosure is reliable.
[0054] A. Size analysis of fetal and maternal DNA The sizes of maternal and fetal DNA fragments were analyzed and compared. As an example, buffy coat DNA and corresponding placental DNA from one pregnant woman were sequenced to obtain 59x and 58x haploid genome coverage, respectively. A total of 822,409 informative single nucleotide polymorphisms (SNPs) were identified for which the mother was homozygous and the fetus was heterozygous. Fetal-specific alleles are defined as alleles present in the fetal genome but absent in the maternal genome. Through PacBio sequencing, 2,652 fetal-specific fragments and 24,837 shared fragments (i.e., fragments carrying shared alleles, primarily of maternal origin) were identified in maternal plasma (M13160). The fetal DNA fraction was 21.8%.
[0055] Figures 4A and 4B show the size distribution of cell-free DNA among fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific). The x-axis indicates base pair sizes from 0 to 20 kb on a linear scale (Figure 4A) or logarithmic scale (Figure 4B). Both fragments carrying shared alleles (primarily maternal origin) and fetal-specific alleles (placental origin) exhibited long-tailed distributions, suggesting the presence of long DNA molecules derived from both fetal and maternal sources. For fragments of primarily maternal origin, 22.6% of plasma DNA molecules were larger than 2 kb in size, while for fragments of fetal origin, 8.5% of plasma DNA molecules were larger than 2 kb in size. These results suggested that fetal DNA molecules contained fewer long DNA molecules. The percentage of long DNA present in this SNP-based analysis of fetal and maternal origin of plasma DNA was, at first glance, much higher than that observed in the overall size analysis. Such discrepancies were likely due to the fact that long DNA molecules were more likely to cover one or more SNPs than short ones and thus favorably selected for SNP-based analysis. The relative proportion of long DNA molecules tagged with SNPs, which was skewed from the proportion of corresponding long DNA in the original pool, was governed by the size of those molecules. Among the fetal-specific DNA fragments, the longest was 16,186 bp, while among the fragments carrying shared alleles, the longest was 24,166 bp.
[0056] Figures 5A and 5B show the size distribution of cell-free DNA among fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific). The x-axis shows base pair sizes from 0 to 5 kb on a linear scale (Figure 5A) or logarithmic scale (Figure 5B). For both fetal-specific and shared DNA fragments, there was a series of major peaks occurring periodically for fragments less than 2 kb. The major peaks were likely consistent with nucleosome structures.
[0057] Figures 6A and 6B show the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific). The x-axis indicates base pair sizes from 0 to 1 kb on a linear scale (Figure 6A) or logarithmic scale (Figure 6B). For both fetal-specific and shared DNA fragments, there was a series of major peaks occurring periodically for fragments smaller than 1 kb. The major peaks were likely consistent with nucleosome structures. There appeared to be an observable shift in the fetal DNA size profile to the left of the size profile of the shared DNA fragments, suggesting that fetal DNA contains shorter DNA molecules than maternal DNA.
[0058] Figures 7A and 7B show the size distribution of cfDNA between fragments carrying shared alleles (shared) and fetal-specific alleles (fetal-specific). The x-axis indicates base pair sizes from 0 to 400 bp on a linear scale (Figure 7A) or logarithmic scale (Figure 7B). The previously reported characteristic features, with a most prominent peak at 166 bp and a 10 bp periodicity occurring in both fetal and maternal molecules below 166 bp, were also reproducible using the new method disclosed herein. These results suggest that molecular size determination by counting the number of sequenced bases from a single molecule according to the present disclosure is reliable.
[0059] B. Size and methylation analysis We analyzed the methylation levels of long cell-free maternal and fetal DNA molecules and found that the methylation levels of fetal DNA molecules were lower than those of maternal DNA molecules.
[0060] In PacBio SMRT sequencing, DNA polymerase mediates the incorporation of fluorescently labeled nucleotides into a complementary strand. The characteristics of the fluorescent pulses generated during DNA synthesis, including inter-pulse duration and pulse width, reflect polymerase kinetics, which can be used to determine nucleotide modifications, such as, but not limited to, 5-methylcytosine, using the approach described in our previous disclosure (U.S. Application No. 16 / 995,607, filed August 17, 2020, entitled "DETERMINATION OF BASE MODIFICATIONS OF NUCLEIC ACIDS") (the entire contents of which are incorporated herein by reference for all purposes).
[0061] In one embodiment, 95,210 fragments carrying maternal-specific alleles and 2,652 fragments carrying fetal-specific alleles were identified, respectively. Maternal-specific alleles are defined herein as alleles that are present in the maternal genome but absent in the fetal genome, and can be identified from SNPs for which the mother is heterozygous and the fetus is homozygous. In this example, a total of 677,375 such informative SNPs were identified. The size of each cell-free DNA molecule was determined. In one embodiment, because the methylation status in the genome is variable, for example, the methylation level of CpG islands is generally lower than that of regions without CpG islands, in order to minimize the variation introduced by genomic context, fragments larger than 1 kb, containing at least 5 CpG sites, and corresponding to a CpG density of less than 5% (i.e., the number of CpG sites in a molecule divided by the total length of the molecule is less than 0.05) can be selected in silico and used for downstream analysis.
[0062] Figure 8 shows the single-molecule, double-stranded DNA methylation levels between fragments carrying maternal-specific alleles and fragments carrying fetal-specific alleles. The y-axis shows the single-molecule, double-stranded DNA methylation levels in percent. The x-axis shows both fragments carrying maternal-specific alleles and fragments carrying fetal-specific alleles. The single-molecule, double-stranded DNA methylation levels of fragments carrying fetal-specific alleles (mean: 62.7%, interquartile range, IQR: 50.0%-77.2%) were lower than those of fragments carrying maternal-specific alleles (mean: 72.7%, IQR: 60.6%-83.3%) (P<0.0001).
[0063] Figure 9A shows the empirical distribution of single-molecule, double-stranded DNA methylation levels of fragments fitted by kernel density estimation implemented in the R package (r-project.org / ). Frequency is shown on the y-axis. The x-axis shows the single-molecule, double-stranded DNA methylation levels in percent. The distribution of fetal-specific long DNA fragments lies to the left of the distribution of maternal-specific fragments, suggesting lower single-molecule, double-stranded DNA methylation levels present in fetal DNA molecules.
[0064] Figure 9B shows a receiver operating characteristic (ROC) analysis using single-molecule, double-stranded DNA methylation levels. The y-axis indicates sensitivity. The x-axis indicates specificity. ROC analysis was performed using single-molecule, double-stranded DNA methylation levels to investigate the ability to distinguish between fetal and maternal DNA fragments using single-molecule, double-stranded DNA methylation levels. The area under the ROC curve (AUC) was found to be 0.62, which was greater than the random guess result of 0.5. In embodiments, spatial patterns of methylation states, such as sequences of methylation states in single molecules, or relative or absolute distances between modified bases and genomic coordinates, can be utilized to further improve the determination of fetal / maternal origin for fragments in plasma. In embodiments, methylation patterns can be combined with other fragmentation metrics (i.e., parameters related to DNA fragmentation), including, but not limited to, preferred ends (Chan et al. Proc Natl Acad Sci USA. 2016;113:E8159-8168), end motifs (Serpas et al. Proc Natl Acad Sci USA. 2019;116:641-649), size (Lo et al. Sci Transl Med. 2010;2:61ra), orientation recognition (i.e., orientation with respect to specific elements within the genome, e.g., open chromatin regions, fragmentation patterns (Sun et al. Genomes Res. 2019;29:418-427)), and topology type (e.g., linear vs. circular DNA molecules (Ma et al. Clin Chem. 2019;65:1161-1170)), to improve classification power for distinguishing fragments of placental (fetal) origin.
[0065] Figures 10A and 10B show that single-molecule, double-stranded DNA methylation levels of both fetal and maternal DNA fragments varied with fragment size. The y-axis indicates single-molecule, double-stranded DNA methylation levels as a percentage. The x-axis indicates sizes from 0 to >20 kb (Figure 10A) and 0 to >1 kb (Figure 10B). Meanwhile, single-molecule, double-stranded DNA methylation levels of fetal-specific DNA molecules were generally lower than those of maternal-specific DNA molecules, both in long (Figure 10A) and short (Figure 10B) ranges. This finding was consistent with current findings that, for short DNA molecules, fetal DNA methylation levels are lower than maternal DNA in the plasma of pregnant women (Lun et al. Clin Chem. 2013;59:1583-94).
[0066] In an embodiment, because the methylation level of fetal DNA molecules is relatively lower than that of maternal DNA molecules, molecules with single-molecule, double-stranded DNA methylation levels below a certain threshold, such as, but not limited to, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, and 5%, are selected to enrich cell-free DNA molecules of fetal origin in plasma DNA pools. For example, the fetal DNA fraction is 2.6% for fragments greater than 1 kb. If fragments (greater than 1 kb) with single-molecule, double-stranded methylation levels less than 50% are selected, the fetal DNA fraction of those further selected fragments greater than 1 kb increases to 5.6% (i.e., an increase of 115.4%). In another example, the fetal DNA fraction is 26.2% for fragments less than 200 bp. If we select fragments (less than 200 bp) with single-molecule, double-stranded DNA methylation levels below 50%, the fetal DNA fraction of those further selected fragments greater than 200 bp increases to 41.6% (i.e., 58.8%). Therefore, the use of thresholded single-molecule, double-stranded DNA methylation levels to enrich for fetal DNA may be more effective for longer DNA molecules under certain circumstances.
[0067] C. Haplotypes and methylation of long cell-free DNA In embodiments, the methods described herein can be used to obtain the base composition, size, and base modification of each single DNA molecule. SNP and methylation information from long cfDNA molecules can be used for haplotype determination. The use of long DNA molecules present in cfDNA pools as disclosed herein allows for phasing of variants in the genome by utilizing the haplotype information present in each consensus sequence, according to published methods (Edge et al. Genome Res. 2017;27:801-812, Wenger et al. Nat Biotechnol. 2019;37:1155-1162). This implementation of haplotype determination based on cfDNA sequence information differs from previous studies that relied on long DNA prepared from tissue DNA. Haplotypes within a genomic region are sometimes referred to as haplotype blocks. A haplotype block can be considered a set of alleles on a staged chromosome. In some embodiments, the haplotype block is extended as long as possible according to the set of sequence information supporting the two alleles physically linked on the chromosome, as well as the allelic overlap information between the different sequences.
[0068] Figures 11A and 11B show an example of a long fetal-specific DNA molecule identified in maternal plasma DNA from a pregnant woman. Among these fetal-specific DNA fragments, one 16,186-bp molecule was used in this embodiment. This fragment aligned to the region of chromosome 10 (chr10:56282981-56299166) of the human reference genome (Figure 11A) and carried seven fetal-specific alleles (Figure 11B). Six of the seven fetal-specific alleles matched the allele information inferred from deep sequencing of the maternal and fetal genomes (using the Illumina platform) (Figure 11B). Its methylation level was determined to be 27.1% according to the method described in this disclosure (Figure 11B), which was much lower than the average level of maternal-specific fragments (72.7%). These results suggest that single-molecule, double-stranded DNA methylation patterns can serve as markers for distinguishing cell-free DNA molecules of fetal and maternal origin.
[0069] Figures 12A and 12B show an example of a long maternal-specific DNA molecule carrying shared alleles identified in maternal plasma DNA from a pregnant woman. Among these fragments carrying shared alleles, the longest was 24,166 bp, which aligned to the human reference region of chromosome 6 (chr6:111074371-111098536) (Figure 12A) and carried 18 shared alleles (Figure 12B). All of these shared alleles were consistent with the allele information estimated from deep sequencing of the maternal and fetal genomes (using the Illumina platform) (Figure 12B). The methylation level was determined to be 66.9% according to the method described in this disclosure (Figure 12B). The genetic and epigenetic information of cell-free DNA molecules several kilobases in length could not be easily identified using short-read sequencing such as bisulfite sequencing (Illumina).
[0070] Here, we describe a method for determining the relative likelihood that a molecule originates from a pregnant woman or a fetus.In a pregnant woman, the DNA molecules that carry the fetal genotype actually originate from the placenta, while most of the DNA molecules that carry the maternal genotype originate from the maternal blood cells.In this method, first, construct a frequency distribution curve of DNA molecules according to the methylation level of both placental and maternal blood cells.To achieve this, the human genome is divided into bottles of different sizes.
[0071] Figure 13 shows frequency distributions for DNA from the placenta (red) and maternal blood cells (blue) according to methylation levels at different resolutions from 1 kb to 20 kb. Frequency is shown on the y-axis. Methylation levels are shown on the x-axis. Examples of bin sizes include, but are not limited to, 1 kb, 2 kb, 5 kb, 10 kb, 15 kb, and 20 kb. The methylation level of each bin was determined based on the number of methylated CpG sites divided by the total number of CpG sites. After determining the methylation levels of all bins, frequency distribution curves can be constructed for each of the placenta genome and maternal blood cell genome for different bin sizes.
[0072] Based on the methylation level of a long DNA molecule, the likelihood that it originates from the placenta or maternal blood cells can be determined by the relative abundance of the two types of DNA molecules at such methylation levels, as well as the fractional concentration of fetal DNA in the sample.
[0073] Let x and y be the frequencies of DNA molecules derived from placenta and maternal blood cells, respectively, with a particular methylation level, and f be the fractional concentration of fetal DNA in the sample.
[0074] The probability (P) that a DNA molecule originates from the fetus can be calculated as follows:
number
[0075] Figures 14A and 14B show frequency distributions for DNA from the placenta (red) and maternal blood cells (blue) according to methylation levels within a 16 kb (Figure 14A) and 24 kb (Figure 14B) window. Frequency is shown on the y-axis. Methylation levels are shown on the x-axis. Based on the frequency distribution plot for the 16 kb fragment (Figure 14A), the frequencies for DNA molecules derived from the placenta and maternal blood cells are 0.6% and 0.08%, respectively. Because the fetal DNA fraction is 21.8%, the probability that this DNA fragment originates from the placenta is 64%, suggesting a high probability of placental origin.
[0076] The probability that a DNA molecule originates from fetal tissue can also be calculated for a 24kb plasma DNA molecule and a methylation level of 66.9%.Based on the frequency distribution plot for the 24kb fragment, the frequencies for DNA molecules originating from placenta and maternal blood cells are 0.05% and 0.16%, respectively (Figure 14B).The probability that this DNA fragment originates from placenta is 0.8%, suggesting that it is very unlikely to be of placental origin.In other words, the likelihood that the molecule is of maternal origin is high.
[0077] This calculation can further consider the size of DNA molecules by referring to the size distribution curves for fetal and maternal DNA.Such analysis can be carried out using, for example but not limited to, Bayes' theorem, logistic regression, multiple regression and support vector machine, random forest analysis, classification and regression tree (CART), K-nearest neighbor algorithm.
[0078] Figures 15A and 15B show that the long DNA fragment in plasma was 18,896 bp in size, aligned to the human reference region of chromosome 8 (chr8:108694010-108712904) (Figure 15A), and carried seven maternal-specific alleles (Figure 15B). All of these maternal-specific alleles were consistent with the allelic information estimated from deep sequencing of the maternal and fetal genomes (Illumina technology) (Figure 15B). Its methylation level was determined to be 72.6% according to the method described in this disclosure (Figure 15B), which is comparable to the pooled methylation level of maternal-specific fragments (72.7%). Therefore, such molecules are more likely to be classified as fragments of maternal origin. The genetic and epigenetic information of cell-free DNA molecules that are several kilobases long could not be easily identified using short-read sequencing such as bisulfite sequencing (Illumina).
[0079] Using above method, the probability that this molecule originates from placenta can be calculated.Based on the frequency distribution plot of 19kb fragment, the frequency of the DNA molecule that originates from placenta and maternal blood cell is 0.65% and 0.23%, respectively.The probability that this DNA fragment originates from placenta is 43%, suggesting that it is highly likely to be of maternal origin.
[0080] D. Clinical Haplotyping Applications In embodiments, the ability to analyze both short and long DNA molecules in the plasma DNA of pregnant women allows for relative haplotype dosage (RHDO) analysis to be performed without the need for previous paternal, maternal, or fetal genotype information obtained from tissue (Lo et al. Sci Transl Med. 2010;2:61ra91, Hui et al. Clin Chem. 2017;63:513-524). This capability is more cost-effective and clinically applicable than previously possible.
[0081] Figure 16 illustrates this principle of how to perform RHDO analysis using cell-free DNA during pregnancy. Cell-free DNA is isolated from a pregnant woman and subjected to SMRT sequencing in step 1605. The size, allele information, and methylation status of each molecule, including long and short DNA molecules, can be determined according to the methods described in this disclosure. In step 1610, the sequenced molecules can be divided into two categories, namely, long DNA molecules and short DNA molecules, according to the size information. The cutoffs used to determine the long and short DNA categories include 150 bp, 180 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 550 bp, 600 bp, 650 bp, 700 bp, 750 bp, 800 bp, 850 bp, 900 bp, 950 bp, 1 kb, 1.1 kb, 1.2 kb, 1.3 kb, 1.4 kb, The length of the long DNA molecule may include, but is not limited to, 1.5 kb, 1.6 kb, 1.7 kb, 1.8 kb, 1.9 kb, 2 kb, 2.5 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 15 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 200 kb, 300 kb, 400 kb, 500 kb, or 1 Mb. In step 1615, in an embodiment, the allelic information present in the long DNA molecule can be used to construct maternal haplotypes, i.e., Hap I and Hap II. The short DNA molecules can be aligned to the maternal haplotypes according to the allelic information. Thus, the number of cell-free DNA molecules (e.g., short DNA) derived from maternal Hap I and Hap II can be determined.
[0082] In step 1620, haplotype imbalances can be analyzed. The imbalances can be in molecule number, molecule size, or molecule methylation status. In step 1625, the maternal inheritance of the fetus can be estimated. If the dosage of Hap I in maternal plasma DNA is overrepresented, the fetus is likely to inherit maternal Hap I. If not, the fetus is likely to inherit maternal Hap II. Different statistical approaches can be used to determine which maternal haplotypes are overrepresented, including, but not limited to, sequential probability ratio tests (SPRTs), binomial tests, chi-square tests, Student's t-tests, nonparametric tests (e.g., Wilcoxon tests), and hidden Markov models.
[0083] In addition to the counting analysis, in some embodiments, the methylation and size of short DNA molecules are also determined and assigned to maternal haplotypes. The methylation imbalance between two haplotypes (i.e., Hap I and Hap II) can be used to determine the maternal haplotype inherited by the fetus. If the fetus inherits Hap I, fragments carrying the Hap I allele will be more abundant in maternal plasma than those carrying the Hap II allele. Hypomethylation of DNA fragments derived from the fetus reduces the methylation level of Hap I compared to that of Hap II. In other words, if the methylation level of Hap I is lower than that of Hap II, the fetus is more likely to inherit maternal Hap I. Otherwise, the fetus is more likely to inherit maternal Hap II. In another embodiment, the probability that each fragment is derived from the fetus or the mother can be calculated as described above. For all fragments that align to Hap I, the aggregated probability that these fragments originate from the fetus can be determined based on Bayes' theorem. Similarly, the aggregated probability that these fragments originate from the fetus can be calculated for Hap II. The likelihood that Hap I or Hap II is inherited by the fetus can then be estimated based on the two aggregated probabilities.
[0084] In an embodiment, the size expansion or contraction between two haplotypes (i.e., Hap I and Hap II) can be used to determine the maternal haplotype inherited by the fetus. If the fetus inherits Hap I, fragments carrying the Hap I allele will be more abundant in maternal plasma compared to those carrying the Hap II allele. DNA fragments derived from the fetus will be relatively shorter than those derived from Hap II. In other words, if molecules derived from Hap I contain more DNA that is shorter than Hap II, the fetus is more likely to inherit maternal Hap I. Otherwise, the fetus is more likely to inherit maternal Hap II.
[0085] In some embodiments, a combined analysis of counts, size, and methylation between maternal Hap I and Hap II can be performed to estimate maternal inheritance of the fetus. For example, logistic regression can be used to combine three metrics including counts, size, and methylation status.
[0086] In clinical trials, haplotype-based analysis of count, size, and methylation status allows for determining whether a fetus has inherited a maternal haplotype associated with a genetic disorder, such as, but not limited to, Fragile X syndrome, muscular dystrophy, Huntington's disease, or beta-thalassemia. Detection of disorders involving repeats of DNA sequences in long cell-free reads is described separately in this disclosure.
[0087] E. Targeted Sequencing of Long Cell-Free DNA Molecules The method described in the present disclosure can be applied to analyze one or more selected long DNA fragments.In an embodiment, one or more long DNA fragments of interest can first be enriched by hybridization method, which allows the DNA molecules from the target region to hybridize with synthetic oligonucleotides with complementary sequences.In order to use the method described herein to decipher size, genetic information, and epigenetic information all together, it is preferable that the target DNA molecule is not amplified by PCR before being subjected to sequencing, because the base modification information of the original DNA molecule is not transferred to the PCR product.
[0088] To enrich these target regions without PCR amplification, several methods have been developed. In another embodiment, one or more target long DNA molecules can be enriched through the use of clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated protein 9 (Cas9) system (Stevens et al. PLOS One 2019;14(4):e0215441, Watson et al. Lab Invest 2020;100:135-146). Although such CRISPR-Cas9 cleavage alters the size of the original long DNA molecules, their genetic and epigenetic information is still preserved and can be retrieved using the methods described herein, including base content, haplotype (i.e., topological) information, de novo mutations, base modifications (e.g., 4mC (N4-methylcytosine), 5hmC (5-hydroxymethylcytosine), 5fC (5-formylcytosine), 5caC (5-carboxylcytosine), 1mA (N1-methyladenine), 3mA (N3-methyladenine), 7mA (N7-methyladenine), 3mC (N3-methylcytosine), 2mG (N2-methylguanine), 6mG (O6-methylguanine), 7mG (N7-methylguan ... Examples of base modifications include, but are not limited to, T (N3-methylthymine), 4mT (O4-methylthymine), and 8oxoG (8-oxo-guanine). In one embodiment, the ends of DNA molecules in a DNA sample are first dephosphorylated, making them less amenable to direct ligation to sequencing adaptors. The long DNA molecule of interest is then guided by a Cas9 protein accompanied by a guide RNA (crRNA) to create a double-stranded break. The long DNA molecule of interest, flanked on both sides by the double-stranded break, is then ligated to a sequencing adaptor designated by the selected sequencing platform. In another embodiment, DNA can be treated with an exonuclease to degrade DNA molecules not bound to the Cas9 protein (Stevens et al. PLOS One 2019;14(4):e0215441). Because these methods do not involve PCR amplification, the sequence of the original DNA molecule containing the base modification can be determined, and the base modification can be identified.
[0089] In embodiments, these methods can be used to target multiple long DNA molecules that share homologous sequences by designing guide RNAs based on a reference genome, such as the human reference genome (hg19), e.g., long interspersed nuclear element (LINE) repeats. In one example, such analysis can be used to analyze circulating cell-free DNA in maternal plasma for the detection of fetal aneuploidy (Kinde et al. PLOS One 2012;7(7):e41162). In embodiments, inactive or "dead" Cas9 (dCas9) and its associated single-stranded guide RNA (sgRNA) can be used to enrich for target long DNA molecules without cleaving double-stranded DNA molecules. For example, the 3' end of the sgRNA can be designed with an extra universal short sequence. Biotinylated single-stranded oligonucleotides complementary to the universal short sequence can be used to capture those target long DNA molecules bound by dCas9. In another embodiment, biotinylated dCas9 protein, sgRNA, or both can be used to facilitate enrichment.
[0090] In some embodiments, size selection can be performed to enrich long DNA fragments without limiting them to one or more specific genomic regions of interest, using approaches including, but not limited to, chemical, physical, enzymatic, gel-based, and magnetic bead-based methods, or a combination of two or more such approaches. In other embodiments, immunoprecipitation can be used to enrich DNA fragments with specific methylation profiles, as mediated by the use of anti-methylcytosine antibodies and methyl-binding proteins. The methylation profile of the bound or captured DNA can be determined using unmethylated sequencing.
[0091] F. General Concepts of Fetal Genetic Analysis Based on Long Plasma DNA Molecules Figure 17 shows the determination of genetic / epigenetic disorders in plasma DNA molecules using information on maternal and fetal origin. Long plasma DNA molecules can be determined to be of fetal or maternal origin in pregnant women according to the genetic and / or epigenetic profiles of the entire or partial CpG sites of the molecule [i.e., region (a)]. Genetic information can be, but is not limited to, sequence information, single nucleotide polymorphisms, insertions, deletions, tandem repeats, satellite DNA, microsatellites, minisatellites, inversions, etc. Epigenetic information can be the methylation status of one or more CpG sites in the plasma DNA molecule and their relative order. In other embodiments, the epigenetic information can be any of A, C, G, or T modifications. Long plasma DNA with tissue origin information can be used for non-invasive prenatal testing by determining the presence of genetic / epigenetic disorders in such long plasma DNA molecules [i.e., region (b)].
[0092] Figure 18 shows the identification of abnormal fetal fragments. As an example, according to the present disclosure, a long DNA fragment was identified as being of fetal origin based on the methylation pattern of region (a). Based on such molecules of fetal origin, the likelihood that the fetus is affected by a genetic or epigenetic disorder can be determined. Genetic disorders may include single-base variants, insertions, deletions, tandem repeats, satellite DNA, microsatellites, minisatellites, inversions, and the like. Examples of genetic disorders include, but are not limited to, beta-thalassemia, alpha-thalassemia, sickle cell disease, cystic fibrosis, sex-linked genetic disorders (e.g., hemophilia, Duchenne muscular dystrophy), spinal muscular atrophy, congenital adrenal hyperplasia, and the like. Epigenetic disorders may be abnormal levels of DNA methylation, such as increased (i.e., hypermethylation) or lost (hypomethylation) methylation. Examples of epigenetic disorders include, but are not limited to, Fragile X syndrome, Angelman syndrome, Prader-Willi syndrome, Facioscapulohumeral muscular dystrophy (FSHD), immunodeficiency, Centromeric Instability and Facial Anomalies (ICF) syndrome, etc. Genetic or epigenetic disorders may be found to reside within region (b).
[0093] G. Improving Sequencing Accuracy Sequencing accuracy can be improved by sequence reads of long cell-free DNA fragments. In Figure 11B, among the seven alleles in the long fetal-specific DNA molecule, there was one allele that appeared inconsistent between PacBio and Illumina sequencing.
[0094] Figures 19A-19G show an illustration of error correction for cell-free DNA genotyping using PacBio sequencing. The results of subread alignment for the seven sites in Figure 11B are visualized. The first line shows the genomic coordinates, and the second line shows the reference sequence. The third line and subsequent lines show the aligned subreads. For example, in Figure 19A, there are eight subreads across the region. A "." represents identity to the reference base in the Watson strand. A "," represents identity to the reference base in the Crick strand. An "alphabet letter" represents an alternative allele. An "*" represents an insertion or deletion. For the inconsistent site shown in Figure 19F, it can be seen that the major base was designated "T" in the consensus sequence. However, among the nine subreads for that site (Figure 19F), only five of the nine subreads (i.e., 56% major allele fraction (MAF)) were determined to be "T," while the others were determined to be "C." The major allele fraction at this site (Figure 19F) was lower than that at other sites (Figures 19A-E and Figure 19G) (MAF range: 67-89%). Therefore, if strict criteria were set for determining the base composition for each site in the consensus sequence, for example, using an MAF of at least 60%, this error site would be excluded from downstream interpretation. On the other hand, such an erroneous site could accidentally fall within a homopolymer (i.e., a series of consecutive identical bases, "TTTTTTT"). In embodiments, criteria can be set such that variants within homopolymers are flagged as QC failures and temporarily not used in downstream analysis. In embodiments, different mapping and base quality can be applied to correct or filter low-quality bases or subreads to improve base composition analysis.
[0095] As the sequencing accuracy of nanopore sequencing is further improved, embodiments of the present invention can be used in conjunction with such improved sequencing platforms, thereby providing improved accuracy.
[0096] H. Exemplary Methods Long cell-free DNA fragments can be sequenced from biological samples obtained from pregnant women who have cell-free DNA fragments. These long cell-free DNA fragments can be used to determine the inheritance of haplotypes by the fetus.
[0097] 1. Sequencing long cell-free DNA fragments 20 shows a method 2000 for analyzing a biological sample from a pregnant organism. The biological sample can contain a plurality of cell-free nucleic acid molecules. The biological sample can be any biological sample described herein. More than 20% of the cell-free nucleic acid molecules in the biological sample have a size greater than 200 nt (nucleotides).
[0098] In block 2010, the plurality of cell-free nucleic acid molecules are sequenced. The sequencing can be by single-molecule real-time techniques. In some embodiments, the sequencing can be by using a nanopore.
[0099] More than 20% of the sequenced cell-free nucleic acid molecules may have a length greater than 200 nt, hi some embodiments, 15-20%, 20-25%, 25-30%, 30-35%, or greater than 35% of the sequenced cell-free nucleic acid molecules may have a length greater than 200 nt.
[0100] In some embodiments, more than 11% of the sequenced plurality of cell-free nucleic acid molecules may have a length greater than 400 nt, hi embodiments, 5-10%, 10-15%, 15-20%, 20-25%, or greater than 25% of the sequenced plurality of cell-free nucleic acid molecules may have a length greater than 400 nt.
[0101] In some embodiments, more than 10% of the sequenced cell-free nucleic acid molecules may have a length greater than 500 nt, hi embodiments, 5-10%, 10-15%, 15-20%, 20-25%, or more than 25% of the sequenced cell-free nucleic acid molecules may have a length greater than 500 nt.
[0102] In embodiments, more than 8% of the sequenced cell-free nucleic acid molecules may have a length greater than 600 nt, hi embodiments, 5-10%, 10-15%, 15-20%, 20-25%, or greater than 25% of the sequenced cell-free nucleic acid molecules may have a length greater than 600 nt.
[0103] In some embodiments, more than 6% of the sequenced cell-free nucleic acid molecules may have a length greater than 1 knt. In some embodiments, 3-5%, 5-10%, 10-15%, 15-20%, 20-25%, or more than 25% of the sequenced cell-free nucleic acid molecules may have a length greater than 1 knt.
[0104] In embodiments, more than 3% of the sequenced cell-free nucleic acid molecules may have a length greater than 2 knt, 1-5%, 5-10%, 10-15%, 15-20%, 20-25%, or greater than 25% of the sequenced cell-free nucleic acid molecules may have a length greater than 2 knt.
[0105] In embodiments, more than 1% of the sequenced cell-free nucleic acid molecules may have a length greater than 3 knt, 1-5%, 5-10%, 10-15%, 15-20%, 20-25%, or greater than 25% of the sequenced cell-free nucleic acid molecules may have a length greater than 3 knt.
[0106] In some embodiments, at least 0.9% of the sequenced cell-free nucleic acid molecules may have a length greater than 4 knt. In some embodiments, 0.5-1%, 1-5%, 5-10%, 10-15%, 15-20%, or more than 20% of the sequenced cell-free nucleic acid molecules may have a length greater than 4 knt.
[0107] In some embodiments, at least 0.04% of the sequenced cell-free nucleic acid molecules may have a length greater than 10 knt. In some embodiments, 0.01-0.1%, 0.1%-0.5%, 0.5-1%, 1-5%, 5-10%, 10-15%, or more than 15% of the sequenced cell-free nucleic acid molecules may have a length greater than 4 knt.
[0108] The plurality of first nucleic acid molecules can comprise at least 10, 50, 100 or 200 cell-free nucleic acid molecules.The plurality of cell-free nucleic acid molecules can be from a plurality of different genome regions.For example, a plurality of chromosome arms or chromosomes can be covered by cell-free nucleic acid molecules.At least two of the plurality of cell-free nucleic acid molecules can correspond to non-overlapping regions.
[0109] The method of sequencing long cell-free DNA fragments can be used by any method described herein.The readings from sequencing can be used to determine fetal aneuploidy, abnormality (such as copy number abnormality), gene mutation or change, or inheritance of parental haplotype.The amount of sequence readings can represent the amount of cell-free DNA fragments.
[0110] 2. Haplotype inheritance 21 illustrates a method 2100 for analyzing a biological sample obtained from a woman carrying a fetus. The woman may have a first haplotype and a second haplotype within a first chromosomal region. The biological sample may include a plurality of cell-free DNA molecules from the fetus and the woman. The biological sample may be any biological sample described herein.
[0111] At block 2105, reads corresponding to a plurality of cell-free DNA molecules can be received. The reads can be sequence reads. In some embodiments, the method can include performing sequencing.
[0112] In block 2110, the size of a plurality of cell-free DNA molecules can be measured. The size can be measured by aligning one or more sequence reads corresponding to the ends of the DNA molecules to a reference genome. The size can be measured by full-length sequencing of the DNA molecules and counting the number of nucleotides in the full-length sequence. The genomic coordinates of the outermost nucleotides can be used to determine the length of the DNA molecules.
[0113] At block 2115, a first set of cell-free DNA molecules from the plurality of cell-free DNA molecules can be identified as having a size greater than or equal to a cutoff value. The cutoff value can be any cutoff associated with long DNA. For example, cutoffs may include 150bp, 180bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 1kb, 1.5kb, 2kb, 2.5kb, 3kb, 4kb, 5kb, 6kb, 7kb, 8kb, 9kb, 10kb, 15kb, 20kb, 30kb, 40kb, 50kb, 60kb, 70kb, 80kb, 90kb, 100kb, 200kb, 300kb, 400kb, 500kb, or 1Mb.
[0114] At block 2120, a sequence of a first haplotype and a sequence of a second haplotype from the reads corresponding to the first set of cell-free DNA molecules may be determined. Determining the sequence of the first haplotype and the sequence of the second haplotype may include aligning the reads corresponding to the first set of cell-free DNA molecules to the reads corresponding to a reference genome.
[0115] In some embodiments, determining the sequence of the first haplotype and the sequence of the second haplotype may not include a reference genome. Determining the sequence may include aligning a first subset of reads to a second subset of reads to identify different alleles at loci within the reads. The method may include determining that the first subset of reads has a first allele at the locus. The method may also include determining that the second subset of reads has a second allele at the locus. The method may further include determining that the first subset of reads corresponds to the first haplotype. Furthermore, the method may include determining that the second subset of reads corresponds to the second haplotype. The alignment may be similar to the alignment described in Figure 16.
[0116] In block 2125, a second set of cell-free DNA molecules from the plurality of cell-free DNA molecules can be aligned to the sequence of the first haplotype. The second set of cell-free DNA molecules can have a size smaller than a cutoff value. The second set of cell-free DNA molecules can be short DNA molecules of the first haplotype.
[0117] In block 2130, a third set of cell-free DNA molecules from the plurality of cell-free DNA molecules can be aligned to the sequence of the second haplotype. The third set of cell-free DNA molecules can have a size smaller than a cutoff value. The third set of cell-free DNA molecules can be short DNA molecules of the second haplotype.
[0118] In block 2135, a first value of a parameter may be measured using a second set of cell-free DNA molecules. The parameter may be a count of cell-free DNA molecules, a size profile of cell-free DNA molecules, or a methylation level of cell-free DNA molecules. The value may be a raw value or a statistical value (e.g., mean, median, mode, percentile, minimum, maximum). In some embodiments, the value may be normalized to the value of the parameter for a reference sample, another region, both haplotypes, or other size range.
[0119] At block 2140, a second value of the parameter may be measured using a third set of cell-free DNA molecules, the parameter being the same parameter as the second set of cell-free DNA molecules.
[0120] In block 2145, the first value may be compared to a second value. The comparison may use a separation value. The separation value may be calculated using the first value and the second value. The separation value may be compared to a cutoff value. The first separation value may be any separation value described herein. The cutoff value may be determined from a reference sample from a woman pregnant with a euploid fetus. In other embodiments, the cutoff value may be determined from a reference sample from a woman pregnant with an aneuploid fetus. In some embodiments, the cutoff value may be determined assuming an aneuploid fetus. For example, data from a reference sample from a woman pregnant with a euploid fetus may be adjusted to account for an increase or decrease in copy number of a chromosomal region for aneuploidy. The cutoff value may be determined from adjusting the data.
[0121] In 2150, the likelihood that the fetus will inherit the first haplotype can be determined based on a comparison between the first value and the second value. The likelihood can be determined based on a comparison between the separation value and the cutoff value. When the parameter is a size profile of the cell-free DNA molecules, the method can include determining that the fetus is more likely to inherit the first haplotype than the second haplotype if the first value is smaller than the second value, indicating that the second set of cell-free DNA molecules is characterized by a smaller size profile than the third set of cell-free DNA molecules. When the parameter is a methylation level of the cell-free DNA molecules, the method can include determining that the fetus is more likely to inherit the first haplotype than the second haplotype if the first value is smaller than the second value.
[0122] In some embodiments, the method may include identifying the number of repeats of a subsequence in one of the reads corresponding to the first set of cell-free DNA molecules. Determining the sequence of the first haplotype may include determining that the sequence includes the number of repeats of the subsequence. The first haplotype may include a repeat-associated disease, which may be any of those described herein. The likelihood that the fetus will inherit the repeat-associated disease may be determined. The likelihood that the fetus will inherit the repeat-associated disease may be equal to or similar to the likelihood that the fetus will inherit the first haplotype. Identifying the repeats of the sequence is described later in this disclosure, including in Figure 16.
[0123] II. Analysis of tissue of origin using methylation Long cell-free DNA molecules may have several methylation sites. As discussed in the present disclosure, the methylation level of long cell-free DNA molecules in pregnant women can be used to determine the tissue of origin. Furthermore, the methylation pattern present on long cell-free DNA molecules can be used to determine the tissue of origin.
[0124] Cells from placental tissues have unique methylome patterns compared with white blood cells and cells from tissues such as the liver, lung, esophagus, heart, pancreas, colon, small intestine, adipose tissue, adrenal gland, and brain (Sun et al., Proc Natl Acad Sci USA. 2015;112:E5503-12). The methylation profile of circulating fetal DNA in the maternal blood during pregnancy may be similar to that of the placenta, thus offering the potential for developing noninvasive fetal-specific biomarkers that are independent of fetal sex or genotype. However, bisulfite sequencing of maternal plasma DNA from pregnant women (e.g., using Illumina sequencing platforms) may lack the ability to distinguish between molecules of fetal and maternal origin due to a number of limitations: (1) plasma DNA can be degraded during bisulfite treatment, typically resulting in long DNA molecules being broken down into shorter molecules; and (2) DNA molecules larger than 500 bp may not be effectively sequenced on Illumina sequencing platforms for downstream analysis (Tan et al., Sci Rep. 2019;9:2856).
[0125] Methylation-based analyses of tissues of origin can focus on several differentially methylated regions (DMRs) and use aggregated methylation signals from multiple tissue-associated molecules instead of the methylation patterns of a single molecule (Sun et al., Proc Natl Acad Sci USA. 2015;112:E5503-12). Numerous studies have attempted to assess the contribution of placenta to the plasma DNA pool using methylation-sensitive restriction enzyme-based (Chan et al., Clin Chem. 2006;52:2211-8) or methylation-specific PCR-based approaches (Lo et al., Am J Hum Genet. 1998;62:768-75). However, these studies are only suitable for analyzing one or a few markers and can be difficult to use to analyze molecules on a genome-wide scale. However, these reads were estimated from amplified signals (i.e., PCR-based amplification during DNA library preparation and bridge amplification during sequencing cluster generation on a flow cell). Such an amplification step may create a bias in favor of short DNA molecules, resulting in the loss of information associated with long DNA molecules. Furthermore, Li et al. analyzed only reads associated with pre-mined DMRs (Li et al., Nuclei Acids Res. 2018;46:e89).
[0126] This disclosure describes a new approach to distinguish fetal and maternal DNA molecules in the plasma of pregnant women based on the methylation patterns of single DNA molecules without bisulfite treatment or DNA amplification. In embodiments, one or more long plasma DNA molecules are used for analysis (e.g., using bioinformatics and / or experimental assays for size selection). A long DNA molecule can be defined as a DNA molecule having a size of at least, but not limited to, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 100 kb, 200 kb, etc. Data on the presence and methylation status of longer cell-free DNA molecules in maternal plasma is lacking. For example, it is unclear whether the methylation status of such longer cell-free DNA molecules reflects the methylation status of cellular DNA in the tissue of origin, for example, because such longer fragments have more sites where the methylation status can change after fragmentation in the body, and such changes can occur while the fragments are circulating in plasma. For example, one study has shown that the methylation status of circulating DNA correlates with the size of the DNA fragment (Lun et al. Clin Chem. 2013;59:1583-94). Therefore, the feasibility of inferring the tissue of origin from such longer cell-free DNA molecules is unclear. Therefore, the approaches taken to identify tissue-associated methylation signatures, as well as the methodologies taken to determine and interpret the presence of such tissue-specific longer cell-free DNA molecules, are substantially different from those applied to short cell-free DNA analysis.
[0127] According to embodiments of the present disclosure, short and long DNA molecules can be identified and their biological characteristics determined, including, but not limited to, methylation patterns, fragment ends, size, and base composition. Short DNA molecules can be defined as DNA molecules with sizes, such as, but not limited to, less than 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 200 bp, and 300 bp. Short DNA molecules can also be DNA molecules that are not within the range considered long. A new approach for estimating the tissue of origin of circulating DNA molecules in the plasma of pregnant women is described. This new approach utilizes the methylation patterns of one or more long DNA molecules in the plasma. The longer the DNA molecule, the greater the number of CpG sites it is likely to contain. The presence of multiple CpG sites on plasma DNA molecules provides tissue of origin information even when the methylation status of any single CpG site is not informative for determining the tissue of origin. Such methylation patterns in long DNA molecules can include the methylation status for each CpG site, the order of the methylation status, and the distance between any two CpG sites. The methylation status between two CpG sites may depend on the distance between them. If CpG sites within a certain distance in a molecule (e.g., a CpG island) show tissue-specific patterns, statistical models can assign more weight to those signals during tissue-of-origin analysis.
[0128] Figure 22 illustrates this principle. Figure 22 shows methylation patterns for DNA molecules. Seven CpG sites for different tissues (placenta, liver, blood cells, and colon) and six plasma DNA fragments A-E are shown. Methylated CpG sites are shown in red, and unmethylated CpG sites are shown in green. As an example, consider seven CpG sites with varying methylation states across placenta, liver, blood cells, and colon tissues. Consider a scenario in which a single CpG site does not exhibit a specific methylation state in the placenta compared to other tissues. Therefore, the tissues of origin for plasma DNA molecules A, B, C, D, and E, which have different sizes, cannot be determined solely based on the methylation state at a single CpG site. In the case of plasma DNA molecules A and B, due to their relatively short size, they contain only three and four CpG sites, respectively. In an embodiment, the methylation patterns in DNA molecules containing two or more CpG sites can be defined as methylation haplotypes. As shown in Figure 22, plasma DNA molecules A and B could be contributed by either the placenta or the liver based on their methylation haplotypes because the placenta and liver shared the same methylation haplotypes at their genomic positions corresponding to molecules A (positions 1, 2, and 3) and B (positions 1, 2, 3, and 4). However, if long DNA molecules in plasma, such as molecules C, D, and E, could be obtained, those molecules C, D, and E could be unambiguously determined to be derived from the placenta based on the methylation haplotypes.
[0129] The reference pattern for a tissue may be based on a methylation pattern from a reference tissue. In some embodiments, the methylation pattern may be based on several reads and / or samples. The methylation level for each CpG site (also called a methylation index, MI, described below) may be used to determine whether the site is methylated.
[0130] A. Statistical model for methylation patterns In an embodiment, the likelihood that a plasma DNA molecule is derived from the placenta can be determined by comparing the methylation haplotype of a single DNA molecule with the methylation patterns in multiple reference tissues.Long plasma DNA molecules can be preferred for such analysis.A long DNA molecule can be defined as a DNA molecule having a size of at least 100bp, 200bp, 300bp, 400bp, 500bp, 600bp, 700bp, 800bp, 900bp, 1kb, 2kb, 3kb, 4kb, 5kb, 10kb, 20kb, 30kb, 40kb, 50kb, 100kb, 200kb, etc., but is not limited to.Reference tissues can include, but are not limited to, placenta, liver, lung, esophagus, heart, pancreas, colon, small intestine, adipose tissue, adrenal gland, brain, neutrophils, lymphocytes, basophils, eosinophils, etc. In one embodiment, the methylation haplotype of plasma DNA determined by single-molecule real-time sequencing and the methylome data based on whole-genome bisulfite sequencing of reference tissues can be analyzed synergistically to determine the likelihood that plasma DNA molecules are derived from placenta.As an example, whole-genome bisulfite sequencing was used to sequence placenta and buffy coat samples to an average genome coverage of 94x and 75x of the haploid genome, respectively.The methylation level (also called methylation index, MI) of each CpG site was calculated based on the number of sequenced cytosines (i.e., methylated, represented by C) and the number of sequenced thymines (i.e., unmethylated, represented by T) using the following formula:
number
[0131] CpG sites were stratified into three categories based on MI values estimated from placental DNA: 1. Category A CpG sites with MI values of 70 or greater. 2. Category B CpG sites with MI values between 30 and 70. 3. Category C CpG sites with MI values of 30 or less.
[0132] Similarly, the MI values of CpG sites estimated from buffy coat DNA were used to classify the CpG sites into three categories: 1. Category A CpG sites with MI values of 70 or greater. 2. Category B CpG sites with MI values between 30 and 70. 3. Category C CpG sites with MI values of 30 or less.
[0133] Categories used MI cutoffs of 30 and 70. Cutoffs may include other values, including 10, 20, 40, 50, 60, 80, or 90. In some embodiments, these categories may be used to determine a reference methylation pattern for a reference tissue (e.g., for use as illustrated in Figure 22). Category A sites may be considered methylated. Category C sites may be considered unmethylated. Category B sites may be considered uninformative and not included in the reference pattern.
[0134] For plasma DNA molecules with n CpG sites, the methylation status for each CpG site was determined by the approach described in our previous disclosure (U.S. Application No. 16 / 995,607). In some embodiments, the methylation status can be determined by bisulfite sequencing or nanopore sequencing. To determine the likelihood that a plasma DNA molecule originated from the placenta or maternal background, the methylation pattern of the molecule was analyzed in conjunction with previous methylation information from placenta and maternal buffy coat DNA. In embodiments, we utilized the principle that if a CpG site determined to be methylated (M) in a plasma DNA fragment coincides with a higher methylation index in the placenta, such an observation indicates that the molecule was more likely to originate from the placenta. If a CpG site determined to be methylated (M) in a plasma DNA molecule coincides with a lower methylation index in the placenta, such an observation indicates that the molecule was less likely to have originated from the placenta; if a CpG site determined to be unmethylated (U) in plasma DNA coincides with a lower methylation index in the placenta, such an observation indicates that the molecule was more likely to have originated from the placenta. If a CpG site determined to be unmethylated (U) in plasma DNA coincides with a higher methylation index in the placenta, such an observation indicates that the molecule was less likely to have originated from the placenta.
[0135] The following scoring scheme was implemented: An initial score (S), reflecting the likelihood of fetal origin for plasma DNA fragments, was set to 0. When the methylation status of plasma DNA molecules was compared with the previous methylation information of placental DNA, a. If a CpG site on a plasma DNA molecule is determined to be "M" and its counterpart in the placenta belongs to category A, a score of 1 is added to S (i.e., the score unit is increased by 1). b. If a CpG site on a plasma DNA molecule is determined to be "U" and its counterpart in the placenta belongs to category A, a score of 1 is subtracted from S (i.e., the score unit is reduced by 1). c. If a CpG site on a plasma DNA molecule is determined to be "M" and its counterpart in the placenta belongs to category B, a score of 0.5 is added to S. d. If a CpG site on a plasma DNA molecule is determined to be "U" and its counterpart in the placenta belongs to category B, a score of 0.5 is added to S. e. If a CpG site on a plasma DNA molecule is determined to be "M" and its counterpart in the placenta belongs to category C, a score of 1 is subtracted from S. f. If a CpG site on a plasma DNA molecule is determined to be "U" and its counterpart in the placenta belongs to category C, a score of 1 is added to S.
[0136] The above process is called "methylation status matching."
[0137] After all CpG sites in the plasma DNA molecule were processed, a final aggregate score S(placenta) was obtained for the plasma DNA molecule. In an embodiment, the number of CpG sites had to be at least 30, and the length of the plasma DNA molecule had to be at least 3 kb. Other numbers and lengths of CpG sites can be used, including but not limited to, any of those described herein.
[0138] A similar scoring scheme was applied when comparing the methylation status of plasma DNA molecules with the methylation level of corresponding sites in buffy coat DNA. After all CpG sites in a plasma DNA molecule were processed, a final aggregate score, S(buffy coat), was obtained for that plasma DNA molecule.
[0139] If S (placenta) was greater than S (buffy coat), the plasma DNA molecule was determined to be of fetal origin; otherwise, the plasma DNA molecule was determined to be of maternal origin.
[0140] The fetal-specific and maternal-specific DNA molecules used to evaluate the performance of inferring fetal-maternal origin for plasma DNA molecules were 17 and 405, respectively. Fetal-specific molecules were plasma DNA molecules carrying fetal-specific SNP alleles, while maternal-specific DNA molecules were those carrying maternal-specific SNP alleles.
[0141] Figure 23 shows the receiver operating characteristic curve (ROC) for determining fetal and maternal origin. The y-axis indicates sensitivity, and the x-axis indicates specificity. The red line represents the performance of distinguishing between molecules of fetal and maternal origin using the methylation status matching-based method presented in the present disclosure. The blue line represents the performance of distinguishing between molecules of fetal and maternal origin using the methylation level of a single molecule (i.e., the percentage of CpG sites determined to be methylated in a DNA molecule). Figure 23 shows that the area under the receiver operating characteristic curve (AUC) for the methylation status matching process (0.94) was significantly higher than that based on the methylation level of a single molecule (0.86) (P value < 0.0001, DeLong test). This suggests that analysis of the methylation patterns of long DNA molecules is useful for determining fetal / maternal origin.
[0142] In embodiments, the magnitude of the difference (ΔS) between S (placenta) and S (buffy coat) can be considered when determining whether plasma DNA is of fetal or maternal origin. The absolute value of ΔS may need to exceed a certain threshold, such as, but not limited to, 5, 10, 20, 30, 40, 50, etc. As an example, when a ΔS threshold of 10 was used, the positive predictive value (PPV) for detecting fetal DNA molecules improved from 14.95% to 91.67%.
[0143] In an embodiment, the methylation status of a CpG site will be affected by the methylation status of its neighboring CpG sites. The closer the nucleotide distance between any two CpG sites on a DNA molecule, the more likely the two CpG sites share the same methylation status. This phenomenon is called co-methylation. Numerous tissue-specific CpG island methylation events have been reported. Therefore, in some statistical models for tissue-of-origin analysis, more weight is assigned to dense clusters of CpG sites (e.g., CpG islands) that share the same methylation status. For scenarios "a" and "f," if the current CpG site under investigation is located within a genomic distance of 100 bp or less compared to the previous CpG site and the results of the methylation status matching process are identical for these two consecutive CpG sites, an additional point is added to the score S for the current CpG site. In the case of scenarios "b" and "e," if the current CpG site under investigation is located within a genomic distance of 100 bp or less compared to the previous CpG site and the results of the methylation status matching process are identical for these two consecutive CpG sites, an additional 1 point is subtracted from the score S for the current CpG site. However, if the current CpG site under investigation is located within a genomic distance of 100 bp or less compared to the previous CpG site and the results of the methylation status matching process for these two consecutive CpG sites are inconsistent, the above-mentioned default scoring scheme is used. On the other hand, if the current CpG site under investigation is located within a genomic distance of more than 100 bp compared to the previous CpG site, the above-mentioned scoring scheme with default parameters is used. Points other than 1 and distances other than 100 bp may be used, including any of those described herein.
[0144] In other embodiments, CpG sites are stratified into four or more categories based on the MI values estimated from placenta and buffy coat DNA. Previous methylation information of reference tissues can be estimated from single molecule real-time sequencing (i.e., nanopore sequencing and / or PacBio SMRT sequencing). The length of plasma DNA molecules may need to be at least 100bp, 200bp, 300bp, 400bp, 500bp, 600bp, 700bp, 800bp, 900bp, 1kb, 2kb, 3kb, 4kb, 5kb, 10kb, 20kb, 30kb, 40kb, 50kb, 100kb, 200kb, etc., but is not limited thereto. The number of CpG sites may need to be, but is not limited to, at least 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, etc.
[0145] In an embodiment, a probabilistic model can be used to characterize the methylation pattern of plasma DNA molecules. The methylation status of k CpG sites (k≧1) on a plasma DNA molecule is defined as M=(m1, m2, ..., m k ), where m i was 0 (for unmethylated state) or 1 (for methylated state) at CpG site i on the plasma DNA molecule. In an embodiment, the probability of M associated with a plasma DNA molecule derived from the placenta may depend on the reference methylation pattern in the placental tissue. The reference methylation patterns in the placental tissue for their corresponding CpG sites 1, 2, ..., k follow a beta distribution. The beta distribution is parameterized by two positive parameters α and β, denoted by Beta(α,β). Values derived from the beta distribution range from 0 to 1. Based on deep bisulfite sequencing data for the tissue of interest, the parameters α and β were determined by the number of cytosines (methylated) and thymines (unmethylated), respectively, sequenced at each CpG site for that particular tissue. In the case of placenta, such a beta distribution can be expressed as Beta(α P ,β p ) The probability of a plasma DNA molecule coming from the placenta, P(M|Placenta), is modeled by:
number
[0146] where "i" indicates the ith CpG site,
number
[0147] The probability of a plasma DNA molecule coming from the buffy coat (i.e., white blood cells), P(M|buffy coat), is modeled by:
number
[0148] where "i" indicates the ith CpG site,
number
[0149]
number
[0150] For a plasma DNA molecule, if P(M|placenta) is observed to be greater than P(M|buffy coat), then such a plasma DNA molecule is likely to originate from the placenta. Otherwise, it is likely to originate from the buffy coat. Using this model, an AUC of 0.79 was achieved.
[0151] B. Machine Learning Models In yet another embodiment, machine learning algorithms can be used to determine the fetal / maternal origin of specific plasma DNA molecules.To test the feasibility of using machine learning-based approaches to classify fetal and maternal DNA molecules in pregnant women, a graphical representation of methylation patterns for plasma DNA molecules was developed.
[0152] Figure 24 shows the definition for pairwise methylation patterns. Nine CpG sites are shown on a plasma DNA molecule. Methylated CpG sites are shown in red, and unmethylated CpG sites are shown in green. If the two CpG sites of a pair share the same methylation state (e.g., the first CpG and the fifth CpG), the pair is coded as 1, as shown at the position indicated by arrow "a." If the two CpG sites of a pair have different methylation states (e.g., the first CpG and the second CpG), the pair is coded as 0, as shown at the position indicated by arrow "b." The same coding rule was applied to all pairs of any two CpG sites on a DNA molecule.
[0153] As an example, a plasma DNA molecule containing nine CpG sites was used. The methylation pattern for this plasma DNA molecule, i.e., UMMMUUUMM (U and M represented unmethylated and methylated CpG, respectively), was determined using the approach described in our previous disclosure (U.S. Application No. 16 / 995,607). Pairwise comparison of the methylation status between any two CpG sites can be useful for machine learning or deep learning-based analysis. In this example, the same rule was applied to a total of 36 pairs. If there were a total of n CpG sites on the plasma DNA molecule, there would be n*(n-1) / 2 pairwise comparisons. A different number of CpG sites, such as 5, 6, 7, 8, 10, 11, 12, or 13, could be used. If the molecule contains a number of sites greater than the number used in the machine learning model, a sliding window can be used to divide the sites into an appropriate number of sites.
[0154] One or more molecules were obtained from each of the placenta and buffy coat DNA samples. The methylation patterns of these DNA molecules were determined by Pacific Bioscience (PacBio) Single-Molecule Real-Time (SMRT) sequencing according to the approach described in our previous publication (U.S. Application No. 16 / 995,607). The methylation patterns were converted into pairwise methylation patterns.
[0155] Pairwise methylation patterns associated with placental DNA and buffy coat DNA were used to train a convolutional neural network (CNN) to distinguish between molecules of potential fetal and maternal origin. Each target output (i.e., analogous to the dependent variable value) for DNA fragments from the placenta was assigned a value of "1," while each target output for DNA fragments from the buffy coat was assigned a value of "0." The pairwise methylation patterns were used to train the CNN model to determine its parameters (often referred to as weights). The optimal parameters for the CNN to distinguish between fetal and maternal origin of DNA fragments were obtained when the overall prediction error (binary value: 0 or 1) between the output score calculated by the sigmoid function and the desired target output reached a minimum by iteratively adjusting the model parameters. The overall prediction error was measured by the sigmoid cross-entropy loss function in a deep learning algorithm (https: / / keras.io / ). The model parameters learned from the training dataset were used to analyze DNA molecules (e.g., plasma DNA molecules) and output a probability score indicating the likelihood that the DNA molecule originated from the placenta or buffy coat. If the probability score of a plasma DNA fragment exceeded a certain threshold, the plasma DNA molecule was considered to be of fetal origin. Otherwise, it was considered to be of maternal origin. Thresholds include, but are not limited to, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.99, etc. In one example, using this CNN model, we achieved an AUC of 0.63 for determining whether a plasma DNA molecule was of fetal or maternal origin, demonstrating the feasibility of using a deep learning algorithm to estimate the tissue of origin of DNA molecules from maternal plasma. The performance of the deep learning algorithm will be further improved by acquiring more single-molecule real-time sequencing results.
[0156] In some other embodiments, statistical models may include, but are not limited to, linear regression, logistic regression, deep recurrent neural networks (e.g., long short-term memory, LSTM), Bayesian classifiers, hidden Markov models (HMM), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering for applications with noise (DBSCAN), random forest algorithms, support vector machines (SVM), etc. Different statistical distributions may be included, including, but not limited to, binomial distribution, Bernoulli distribution, gamma distribution, normal distribution, Poisson distribution, etc.
[0157] C. Placenta-specific methylation haplotypes The methylation status of each CpG site on a single DNA molecule can be determined using the approach described in our previous disclosure (U.S. Application No. 16 / 995,607) or any of the techniques described herein. In addition to single-molecule, double-stranded DNA methylation levels, the single-molecule methylation pattern of each DNA molecule can be determined, which can be the sequence of the methylation status of adjacent CpG sites along the single DNA molecule.
[0158] Different DNA methylation signatures can be found in different tissues and cell types. In embodiments, the tissue of origin of individual plasma DNA molecules can be inferred based on the methylation patterns of single molecules.
[0159] Genomic DNA from 10 buffy coat samples and 6 placental tissue samples was sequenced using SMRT sequencing (PacBio). By pooling together high-quality circular consensus sequencing (CCS) reads mapped from each sample type, we were able to achieve 58.7x and 28.7x coverage for buffy coat DNA and placental DNA, respectively.
[0160] Using a sliding window approach, the genome was divided into approximately 28.2 million overlapping windows of five CpG sites. In other embodiments, different window sizes may be used, such as, but not limited to, 2, 3, 4, 5, 6, 7, and 8 CpG sites. A non-overlapping window approach may also be used. Each window was considered a potential marker region. For each potential marker region, significant single-molecule methylation patterns were identified among all sequenced placental DNA molecules covering all five CpG sites within that marker region. A comparison was made between the CpG sites in plasma DNA molecules and the corresponding CpG sites in individual DNA molecules from the reference tissue. The single-molecule methylation pattern was then compared with the significant single-molecule methylation pattern in the placenta to calculate a mismatch score for each buffy coat DNA molecule covering all CpG sites within the same marker region.
number
[0161] A higher discrepancy score indicates that the methylation pattern of the buffy coat DNA molecules is more different from the significant single-molecule methylation pattern in the placenta. From the 28.2 million potential marker regions, we selected regions that showed substantial differences in single-molecule methylation patterns between pools of DNA molecules from placenta and buffy coat using the following criteria: a) more than 50% of the placenta DNA molecules had significant single-molecule methylation patterns, and b) more than 80% of the buffy coat DNA molecules had a discrepancy score greater than 0.3. Based on these criteria, 281,566 marker regions were selected for downstream analysis.
[0162] Figure 25 is a table showing the distribution of selected marker regions among different chromosomes. Column 1 shows the chromosome number. Column 2 shows the number of marker regions within the chromosome.
[0163] Here, we present the concept of tissue-of-origin classification for individual plasma DNA molecules based on single-molecule methylation patterns using plasma DNA molecules sequenced by SMRT sequencing covering either fetal-specific or maternal-specific alleles, as described previously in this disclosure. Any plasma DNA molecule covering a selected marker region with a methylation pattern identical to the methylation pattern of a prominent single molecule in the placenta is classified as a placenta-specific (i.e., fetal-specific) DNA molecule. In contrast, if the methylation pattern of a plasma DNA molecule is not identical to the methylation pattern of a prominent single molecule in the placenta, the molecule is classified as not placenta-specific. Correct classification in this analysis was defined by whether or not a placenta-specific methylation haplotype was present in the molecule, identifying fetal-specific DNA molecules as fetal-derived (i.e., placenta-specific) and maternal DNA molecules as non-fetal-derived (i.e., placenta-nonspecific). Previous methylation-based methods for tissue-of-origin analysis typically involved deconvolving the percentage or proportional contribution of a range of tissue contributors of cell-free DNA within a biological sample. An advantage of this method over previous methods is that evidence of a tissue's cell-free DNA contribution to a biological sample, such as placental-derived DNA in maternal plasma, can be determined regardless of the presence or absence of contributions from other tissues. Furthermore, the placental origin of any one cell-free DNA molecule can be determined by this method regardless of the fractional contribution of cell-free DNA molecules from that tissue.
[0164] Among the 28 DNA molecules covering fetal-specific alleles, 17 (61%) were classified as placenta-specific and 11 (39%) as not placenta-specific, whereas among the 467 DNA molecules covering maternal-specific alleles, 433 (93%) were classified as not placenta-specific and 34 (7%) as placenta-specific.
[0165] In embodiments, different percentages of buffy coat DNA molecules with a discrepancy score greater than 0.3 can be used as thresholds, including, but not limited to, greater than 60%, 70%, 75%, 80%, 85%, and 90%. By adjusting the criteria used in selecting marker regions, the overall classification accuracy of placental or non-placental origin of plasma DNA in pregnant subjects can be improved. This is particularly important in the setting of non-invasive prenatal testing, which seeks to determine whether disease-causing mutations or copy number abnormalities are present in the fetus.
[0166] Figure 26 is a table showing the classification of plasma DNA molecules based on single-molecule methylation patterns, using different percentages of buffy coat DNA molecules with a mismatch score greater than 0.3 as the selection criteria for marker regions. Column 1 shows the percentage of buffy coat DNA molecules with a mismatch score greater than 0.3%. Column 2 divides the DNA molecules into molecules covering fetal-specific alleles and molecules covering maternal-specific alleles. Columns 3 and 4 show the classification of DNA molecules as placenta-specific or not placenta-specific based on single-molecule methylation patterns. Column 5 shows the percentage of DNA molecules classified as the same as the specific allele in column 2.
[0167] Figure 27 shows a process flow for determining fetal inheritance non-invasively using placenta-specific methylation haplotypes. As shown in Figure 27, cell-free DNA from the plasma of a pregnant woman was extracted for single-molecule real-time sequencing. Long plasma DNA molecules were identified according to an embodiment of the present disclosure. The methylation status at each CpG site for each long plasma DNA molecule was determined according to an embodiment of the present disclosure. The methylation haplotype of each long plasma DNA molecule was determined according to an embodiment of the present disclosure. If a long plasma DNA molecule is identified as carrying a placenta-specific methylation haplotype, the genetic and epigenetic information associated with that molecule is considered to be inherited by the fetus. In an embodiment, if one or more long plasma DNA molecules containing a disease-causing mutation that is the same as a disease-causing mutation carried by the pregnant woman are determined to be of fetal origin based on methylation haplotype information according to an embodiment of the present disclosure, this suggests that the fetus inherited the mutation from the mother.
[0168] Embodiments include treatment of beta-thalassemia, sickle cell disease, alpha-thalassemia, cystic fibrosis, hemophilia A, hemophilia B, congenital adrenal hyperplasia, Duchenne muscular dystrophy, Becker muscular dystrophy, achondroplasia, thanatophoric dysplasia, von Willebrand disease, Noonan syndrome, hereditary hearing loss and deafness, various inborn errors of metabolism (e.g., citrullinemia type I, propionic acidemia, glycogen storage disease type Ia (von Gierke disease), glycogen storage disease type Ib / c (von Gierke disease), glycogen storage disease type II (Pompe disease), mucopolysaccharidosis (MPS) type I (Hurler / Hurler-Scheie / Scheie), MPS type II (Hunter syndrome), MPS type IIIA (Sanfilippo syndrome A), MPS type IIIB (Sanfilippo syndrome B), MPS type IIIC (Sanfilippo syndrome B), MPS type IIID (Sanfilippo syndrome C), MPS type IIIE (Sanfilippo syndrome D), MPS type IIIF (Sanfilippo syndrome E), MPS type IIIG (Sanfilippo syndrome F), MPS type IIIH (Sanfilippo syndrome F), MPS type IIII (Sanfilippo syndrome F), MPS type IIIIH ... The present invention may be applied to genetic disorders including, but not limited to, MPS IIIC (Sanfilippo syndrome C), MPS IIID (Sanfilippo syndrome D), MPS IVA (Morquio syndrome A), MPS IVB (Morquio syndrome B), MPS VI (Maroteaux-Lamy syndrome), MPS VII (Sly syndrome), mucolipidosis II (I-cell disease), metachromatic leukodystrophy, GM1 gangliosidosis, OTC deficiency (X-linked ornithine transcarbamylase deficiency), adrenoleukodystrophy (X-linked ALD), and Krabbe disease (globoid cell leukodystrophy).
[0169] In other embodiments, genetic disorders in fetuses may be associated with de novo DNA methylation in the fetal genome that was not present in the parental genome. One example is hypermethylation of the FMRP translational regulator 1 (FMR1) gene in fetuses with fragile X syndrome. Fragile X syndrome is caused by an expansion of a CGG trinucleotide repeat in the 5' untranslated region of the FMR1 gene. Normal alleles contain approximately 5-44 copies of the CGG repeat. Premutation alleles contain 55-200 copies of the CGG repeat. Full mutation alleles contain more than 200 copies of the CGG repeat.
[0170] Figure 28 illustrates the principle of noninvasive prenatal detection of fragile X syndrome in male fetuses from unaffected pregnant women carrying either normal or premutation alleles. In Figure 28, "n" represents the number of copies of CGG in the maternal genome, and "m" represents the number of copies of CGG in the fetal genome. The genome of an unaffected pregnant woman has 200 or fewer copies (i.e., n≦200) of CGG repeats and an unmethylated FMR1 gene. In contrast, the genome of a male fetus affected with fragile X syndrome has more than 200 copies (m>200) of CGG repeats and a methylated FMR1 gene. By performing single-molecule sequencing of maternal plasma DNA, multiple long DNA molecules can be identified from the genomic region of interest (e.g., the FMR1 gene) whose repeat number and methylation status can be simultaneously determined. Identifying one or more DNA molecules containing more than 200 copies of CGG repeats and covering a methylated FMR1 gene in the plasma of an unaffected woman indicates a high likelihood that the fetus has fragile X syndrome. In yet another embodiment, placenta-specific methylation haplotypes according to embodiments of the present disclosure can be used to further confirm the fetal origin of such plasma DNA molecules. If one or more molecules containing one or more regions within a molecule bearing a placenta-specific methylation haplotype are identified, and such molecules contain more than 200 copies of CGG repeats and cover a methylated FMR1 gene, it can be concluded with greater confidence that the fetus has fragile X syndrome. Conversely, identifying one or more molecules with a placenta-specific methylation haplotype, and such molecules contain less than 200 copies of CGG repeats and cover an unmethylated FMR1 gene, indicates a high likelihood that the fetus is not affected. In fragile X syndrome, a full mutation (more than 200 repeats) actually methylates the entire gene, turning off gene function. Therefore, particularly in the case of fragile X, detection of the methylated long allele (rather than showing a placental methylation profile) is highly suggestive of the fetus having the disease.
[0171] Detection of genetic disorders can be performed regardless of whether the mother's previous condition is known. Women with premutations may not have any symptoms, but may have mild symptoms, which are often only discovered later. If the mother's mutation status is unknown, one approach is to detect the long allele in plasma from women who do not appear to have the disease, or to analyze the mother's buffy coat and determine that it does not exhibit such a long allele. Another approach can be to combine the repeat length and the methylation status of cfDNA molecules. If the methylation status suggests a fetal pattern (methylation haplotype) and indicates a long allele, the fetus is likely to be affected. This approach can be applied to many trinucleotide disorders, such as Huntington's disease.
[0172] D. Non-invasive construction of the fetal genome using long plasma DNA molecules Methylation patterns can be used to determine the inheritance of haplotypes. Determining haplotype inheritance using a qualitative approach using methylation patterns can be more efficient than quantitative methods that characterize the amount of specific fragments. Methylation patterns can be used to determine maternal and paternal inheritance of haplotypes.
[0173] 1. Maternal inheritance of fetus Lo et al. demonstrated the feasibility of using parental haplotype information to construct a genome-wide genetic map and determine fetal mutation status from maternal plasma DNA sequences (Lo et al. Sci Transl Med. 2010;2:61ra91). This technique, called relative haplotype dosage (RHDO) analysis, is one approach to resolving maternal inheritance in fetuses. The principle was based on the fact that maternal haplotypes inherited by the fetus are relatively overrepresented in the plasma DNA of pregnant women compared with other maternal haplotypes not inherited by the fetus. Therefore, RHDO is a quantitative analysis method.
[0174] Embodiments present in this disclosure utilize methylation patterns in long plasma DNA molecules to determine the tissue of origin of those plasma DNA molecules. In one embodiment, the disclosure herein allows for the qualitative analysis of maternal inheritance of a fetus.
[0175] Figure 29 shows an example of determining maternal inheritance of a fetus. Genomic position P was heterozygous (A / G) in the maternal genome. Filled circles indicate methylated sites, and open circles indicate unmethylated sites. The methylation pattern in the placenta was "-MUMM-", where "M" represents a methylated cytosine at the CpG site and "U" represents an unmethylated cytosine. In one embodiment, the methylation patterns in the placenta and related reference tissues can be obtained from data previously generated from sequencing (e.g., single-molecule real-time sequencing and / or bisulfite sequencing). In plasma DNA, one non-paternal plasma DNA (indicated by Z) carrying an A allele at that particular genomic locus was found to exhibit a methylation pattern ("-MUMM-") that matched the methylation pattern in the placenta, in contrast to the methylation patterns in other tissues. No molecules carrying a G allele were found to exhibit a methylation pattern that matched the methylation pattern in the placenta. Thus, based on the presence of allele A and the "-MUMM-" methylation pattern, the fetus can be determined to have inherited maternal allele A.
[0176] Figure 30 shows qualitative analysis of fetal maternal inheritance using genetic and epigenetic information from plasma DNA molecules. As shown in the top branch of Figure 30, plasma DNA was extracted, followed by size selection of long DNA, according to an embodiment of the present disclosure. The size-selected plasma DNA molecules were subjected to single-molecule real-time sequencing (e.g., using a system manufactured by Pacific Biosciences). Genetic and epigenetic information was determined according to an embodiment of the present disclosure. For illustrative purposes, a molecule (X) was aligned to human chromosome 1, which contains a G allele at chromosomal position a (chr1:a) and an A allele at chromosomal position e (chr1:e). Molecule X has a C allele at chromosomal position d.
[0177] The CpG methylation status of molecule X is determined to be "-MUMM-", where "M" represents a methylated cytosine at the CpG site and "U" represents an unmethylated cytosine. A filled circle indicates a methylated site, and an open circle indicates an unmethylated site. As a result of analysis of a reference sample, placental DNA has been found to have a methylation pattern of "-MUMM-" in the region between positions a and e. Based on the methylation pattern of molecule X matching the methylation pattern of placental DNA, molecule X has been determined to be of placental origin in accordance with embodiments of the present disclosure.
[0178] As shown in the lower branch of Figure 30, DNA from maternal leukocytes was subjected to single-molecule real-time sequencing. Epigenetic and genetic information of maternal leukocytes was obtained according to an embodiment of the present disclosure. Gene alleles were staged into two haplotypes, namely, maternal haplotype I (Hap I) and maternal haplotype II (Hap II), using methods including, but not limited to, WhatsHap (Patterson et al. J Comput Biol. 2015;22:498-509), HapCUT (Bansal et al. Bioinformatics. 2008;24:i153-9), and HapCHAT (Beretta et al. BMC bioinformatics. 2018;19:252). Here, two haplotypes in the maternal genome, namely, "-ACGT-" (Hap I) and "-GTAC-" (Hap II), were obtained. Hap I was associated with wild-type variants, while Hap II was associated with disease-associated variants, which may include, but are not limited to, single nucleotide variants, insertions, deletions, translocations, inversions, repeat expansions, and / or other genetic structural changes.
[0179] For genomic location e, the maternal genotype was determined to be AA and the paternal genotype was determined to be GG. Due to the methylation pattern, plasma DNA molecule X was determined to be of placental origin. Because the maternal-specific allele A is present but the paternal-specific allele G is absent, molecule X was presumed to be inherited from one of the maternal haplotypes.
[0180] To further determine which maternal haplotype was inherited by the fetus, the allele information at genomic positions other than position chr1:e of this placenta-derived molecule X was compared with the maternal haplotype. As an example, molecule X has allele G at position a and allele C at position d. The presence of either of these alleles in molecule X indicates that molecule X should be assigned to maternal Hap II, which contains the same alleles.
[0181] Therefore, it can be concluded that the maternal haplotype II associated with the disease-associated variant was passed on to the fetus, and the unborn fetus was determined to be at risk for the disease.
[0182] Methylation pattern-based qualitative analysis of fetal maternal inheritance may require fewer plasma DNA molecules to conclude which maternal haplotypes were inherited by the fetus compared with RHDO, which was a quantitative analysis-based approach. A computer simulation analysis was performed to evaluate the detection rate of fetal maternal inheritance using genome-wide methods that used different numbers of plasma DNA molecules for analysis.
[0183] In the RHDO simulation analysis, N plasma DNA molecules were collectively aligned to M heterozygous SNPs within a haplotype block in the maternal genome. The fetal DNA fraction was f. The paternal genotypes of their corresponding SNPs were homozygous and identical to the maternal Hap I inherited by the fetus. Among the N plasma DNA molecules, the mean number of plasma DNA molecules aligned to maternal Hap I was N × (0.5 + f / 2), while the mean number of plasma DNA molecules aligned to maternal Hap II was N × (0.5 - f / 2). We assumed that the plasma DNA molecules sampled from the haplotypes followed a binomial distribution.
[0184] The number of plasma DNA molecules was assigned to Hap I (i.e., X) according to the following distribution: X~Bin(N,0.5+f / 2)(1), Here, "Bin" indicates the binomial distribution.
[0185] The number of plasma DNA molecules was assigned to Hap II (i.e., Y) according to the following distribution: Y~Bin(N,0.5-f / 2)(2).
[0186] Thus, plasma DNA molecules assigned to maternal Hap I are relatively over-represented in maternal plasma compared to maternal Hap II. To determine whether the over-representation was statistically significant, the difference in plasma DNA counts between the two maternal haplotypes was compared using the null hypothesis that the two haplotypes (denoted by X' and Y') were equally represented in plasma. X'~Bin(N,0.5)(3), Y'~Bin(N,0.5)(4).
[0187] The relative dosage difference between the two haplotypes was further defined as follows: D=(XY) / N(5), D' = (X' - Y') / N(6).
[0188] In one example, the statistic D reflecting relative haplotype dosage was compared to the mean (i.e., z-score) of D'(M) normalized by the standard deviation of D'(SD) as follows: z-score = (DM) / SD(7). A z-score greater than 3 indicated that Hap I was passed on to the fetus.
[0189] For RHDO analysis, 30,000 haplotype blocks were simulated across the entire genome, with Hap I inherited by the fetus, based on equations (1)-(7). The average length of a haplotype block was 100 kb. Each haplotype block contained an average of 100 SNPs, of which 10 SNPs were informative for contributing to haplotype imbalance. In one example, the fetal DNA fraction was 10%, and the median fragment size was 150 bp. By varying the number of plasma DNA molecules used in RHDO analysis from 1 million to 300 million, the percentage of haplotype blocks with a z-score greater than 3, referred to herein as the detection rate, was calculated. The number of plasma DNA molecules here was adjusted by the probability that plasma DNA covers an informative SNP site according to a Poisson distribution.
[0190] For computer simulations relating to methylation pattern-based qualitative analysis of fetal maternal inheritance, the following assumptions were made for illustrative purposes: 1) There were N plasma DNA molecules covering the haplotype blocks in the maternal genome used in the analysis. 2) The probability of a plasma DNA fragment used in tissue-of-origin analysis being at least 3 kb in length is indicated by a. 3) The probability of a plasma DNA molecule carrying more than 10 CpG sites was indicated by b. 4) The fetal DNA fraction of those fragments larger than 3 kb was indicated by f.
[0191] As shown in one embodiment of the present disclosure, accurate estimation of tissue of origin can be achieved for those plasma DNA molecules larger than 3 kb that have at least 10 CpG sites. The number of plasma DNA molecules that meet the above criteria (Z) was assumed to follow a Poisson distribution with a mean value of λ (i.e., N×a×b×f). Z~Poisson(λ)(8).
[0192] In one example, 30,000 haplotype blocks in which Hap I was inherited by the fetus were simulated based on Equation (8). The average length of each haplotype block was 100 kb. Each haplotype block contained an average of 100 SNPs, of which 20 heterozygous SNPs were phased into two maternal haplotypes. The fetal DNA fraction was 1%. After size selection, 40% of plasma DNA molecules were larger than 3 kb. 87.1% of plasma DNA molecules with at least 10 CpG sites were larger than 3 kb. The percentage of haplotype blocks with a Z value of 1 or greater indicated the detection rate. The computer simulation was repeated multiple times by varying the number of plasma DNA molecules (N) used for tissue-of-origin analysis by methylation pattern from 1 million to 300 million. The number of plasma DNA molecules used herein was further adjusted by the probability that plasma DNA would cover heterozygous SNPs according to the Poisson distribution.
[0193] Figure 31 shows the detection rate of qualitative analysis of fetal maternal inheritance in a genome-wide method using genetic and epigenetic information from plasma DNA molecules compared with relative haplotype dosage (RHDO) analysis. The number of molecules used in the analysis is shown on the x-axis. The detection rate of fetal maternal inheritance as a percentage is shown on the y-axis. The detection rate of fetal maternal inheritance was higher using the methylation pattern-based approach compared with RHDO. For example, using 100 million fragments, the detection rate based on methylation patterns was 100%, while the detection rate based on RHDO was only 55%. These results suggested that estimation of fetal maternal inheritance using the methylation pattern-based method was superior to that based on RHDO.
[0194] 2.Paternal inheritance of fetus The ability to obtain long plasma DNA molecules for analysis may help improve the detection rate of paternity-specific variants in the plasma DNA of pregnant women, as the use of long DNA molecules increases overall genome coverage compared to the use of an equal number of short DNA molecules. Further computer simulations were performed based on the following assumptions: 1) The fetal DNA fraction was f depending on the length L of the plasma DNA. This is f L where the subscript L indicated that plasma DNA molecules with a length of L bp were used in the analysis. 2) The number of paternal-specific variants that needed to be identified in maternal plasma DNA was V. 3) The number of plasma DNA molecules used in the analysis was N. 4) The number of plasma DNA molecules derived from a particular genomic locus or region followed a Poisson distribution.
[0195] In one example, the fetal DNA fraction of those plasma DNA molecules with sizes of 150 bp, 1 kb, and 3 kb was 10% (f 150bp =0.1), 2%(f 1kb =0.02), and 1%(f 3kb= 0.01). The number of paternal-specific variants in the genome was 250,000 (V = 250,000). The number of plasma DNA molecules (N) used in the analysis ranged from 50 million to 500 million.
[0196] Figure 32 shows the relationship between the detection rate of paternal-specific variants in genome-wide analysis and the number of sequenced plasma DNA molecules of different sizes used in the analysis. The number of sequenced molecules used in the analysis in millions is shown on the x-axis. The percentage of paternal-specific variants detected is shown on the y-axis. The different curves show the different sizes of DNA fragments used in the analysis: 3 kb on top, 1 kb in the middle, and 150 bp on the bottom. The longer the plasma DNA molecules used in the analysis, the higher the detection rate of paternal-specific variants can be achieved. For example, using 400 million plasma DNA molecules, the detection rates were 86%, 93%, and 98% when focusing on molecules with sizes of 150 bp, 1 kb, and 3 kb, respectively.
[0197] In other embodiments, other distributions may be used, including but not limited to Bernoulli distribution, Beta-Normal distribution, Normal distribution, Conway-Maxwell-Poisson distribution, Geometric distribution, etc. In some embodiments, Gibbs sampling and Bayes' theorem are used for maternal and paternal genetic analysis.
[0198] 3. Fragile X Genetic Analysis In an embodiment, methylation pattern-based determination of maternal inheritance in a fetus can facilitate non-invasive detection of fragile X syndrome using single-molecule real-time sequencing of maternal plasma DNA. Fragile X syndrome is a genetic disorder typically caused by an expansion of a CGG trinucleotide repeat within the FMR1 (Fragile X Mental Retardation 1) gene on the X chromosome. Fragile X syndrome and other disorders caused by repeat expansions are described elsewhere in this application. The method for detecting fragile X syndrome in a fetus can also be applied to any other expansion of the repeat disclosed herein.
[0199] Female subjects with a premutation, defined as having 55 to 200 copies of CGG repeats in the FMR1 gene, are at risk of giving birth to a child with fragile X syndrome. The likelihood of conceiving a fetus with fragile X syndrome depends on the number of CGG repeats present in the FMR1 gene. The higher the number of repeats in the mother, the higher the risk of the premutation expanding to a full mutation when passed on to the fetus. Maternal plasma samples were collected at 12 weeks gestation from women who were previously confirmed to carry a fragile X premutation allele of 115 ± 2 CGG repeats and had a son (proband) diagnosed with fragile X syndrome. Maternal plasma samples were then subjected to single-molecule real-time sequencing. In one example, single-molecule real-time sequencing was used to obtain 3.3 million circular consensus sequences (CCSs) aligned to the human reference genome, with a median subread depth of 75x per CCS (interquartile range: 14-237x). The genetic and epigenetic information for each sequenced plasma DNA can be determined according to the embodiment of the present disclosure.To obtain two maternal haplotypes of chromosome X, the Infinium Omni2.5Exome-8 Beadchip (Illumina) on iScan System, a microarray technology, was used to determine the genotypes of 2,000 SNPs on chromosome X for both the DNA extracted from maternal buffy coat and the proband's oral swab.Two maternal haplotypes, namely, Hap I and Hap II, can be estimated based on the genotype information of the maternal and proband genomes.
[0200] Figure 33 shows a workflow for noninvasive detection of fragile X syndrome. Across heterozygous SNP sites in maternal buffy coat DNA, alleles identical to the proband's genotype were used to define haplotypes associated with the premutation allele (i.e., Hap I), a potential precursor of the full mutation in the next generation. Conversely, alleles distinct from the proband's genotype were used to define haplotypes associated with the corresponding wild-type allele (Hap II). Maternal plasma DNA from the proband's mother, who was pregnant with a fetus, was subjected to single-molecule real-time sequencing. Sequencing reads were assigned to maternal Hap I and Hap II depending on whether the acquired genetic information was identical to Hap I or Hap II alleles across the genomic loci under investigation. According to an embodiment of the present disclosure, the methylation patterns of plasma DNA molecules were used to determine the tissue of origin of those plasma DNA molecules containing a specific number of CpG sites (i.e., DNA molecules identified as being of placental origin based on methylation pattern analysis were determined to be of fetal origin).
[0201] In scenario A, if fetal (i.e., placental) DNA molecules were detectable in those plasma DNA molecules assigned to maternal Hap I but not in those plasma DNA molecules assigned to maternal Hap II, Hap I would be determined to be passed on to the unborn fetus. The fetus would be determined to be at high risk for Fragile X syndrome. The placental origin of the plasma DNA molecules is based on the methylation status of the molecules, as discussed below.
[0202] In scenario B, if fetal DNA molecules are detectable in those plasma DNA molecules assigned to maternal Hap II but not in those plasma DNA molecules assigned to maternal Hap I, Hap II is determined to be passed on to the unborn fetus. The fetus is determined not to be affected with Fragile X Syndrome.
[0203] In embodiments, the definitions of "detectable" and "undetectable" for fetal DNA molecules may depend on the cutoff for the percentage of plasma DNA molecules identified as being of fetal (i.e., placental) origin. Cutoffs for "detectable" include, but are not limited to, greater than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, etc. Cutoffs for "undetectable" include, but are not limited to, less than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, etc. In some embodiments, the difference in the percentage of plasma DNA molecules determined to be of fetal origin between Hap I and Hap II may need to be, but is not limited to, greater than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, etc. In other embodiments, haplotype information can be obtained from long-read sequencing technologies (e.g., PacBio or nanopore sequencing) (Edge et al. Nat Commun. 2019;10:4660), synthetic long reads (e.g., using a platform from 10X Genomics) (Hui et al. Clin Chem. 2017;63:513-14), targeted locus amplification (TLA)-based phasing (Vermeulen et al. Am J Hum Genet. 2017;101:326-39), and statistical phasing (e.g., Shape-IT) (Delaneau et al. Nat Method. 2011;9:179-81).
[0204] In an embodiment, the maternal and fetal origin of plasma DNA molecules that were at least 200 bp long and contained at least five CpG sites (or any other cutoff for longer DNA molecules) could be determined using the methylation status matching approach described herein. We identified one plasma DNA molecule located at genomic position chrX:143,782,245-143,782,786 (3.2 Mb away from the FMR1 gene) with an allele (position: chrX:143782434, SNP accession number: rs6626483, allele genotype: C) identical to the corresponding allele on maternal Hap II but different from maternal Hap I.
[0205] Figure 34 shows the methylation pattern of plasma DNA compared with the methylation profiles of placenta and buffy coat DNA. The plasma DNA molecule contained five CpG sites. The methylation pattern was determined to be "MUUUU." This methylation pattern obtained from single-molecule real-time sequencing was compared with the reference methylation profiles of placenta tissue and buffy coat DNA samples obtained from bisulfite sequencing according to the methylation status matching approach described in this disclosure. The score for this molecule derived from the placenta [i.e., S(placenta)] was 2, which was greater than the score derived from the buffy coat [i.e., S(buffy coat)] of -3. Therefore, such plasma DNA molecule (chrX:143,782,245-143,782,786) was determined to be of fetal origin. However, no plasma DNA molecule carrying the maternal Hap I allele was observed, which was of fetal origin. Therefore, it was concluded that the fetus inherited maternal Hap II and was not affected by fragile X syndrome.
[0206] We hypothesized that the performance of the approach described herein may not be significantly affected by X-chromosome inactivation due to the following factors: 1) X-inactivation is not complete in humans. As many as one-third of genes on the X chromosome exhibit variable escape from X-inactivation (Cotton et al. Hum Mol Genet. 2015;25:1528-1539). CpG sites outside CpG islands (i.e., the majority of CpG sites) are similarly methylated in both sexes, suggesting that the methylation status of most CpG sites on the X chromosome may not be affected by X-inactivation (Yasukochi et al. Proc Natl Acad Sci USA. 2010;107:3704-9). 2) We used the methylation profile of sex-matched placental tissue for the unborn fetus. This strategy is useful for detecting maternal inheritance of the fetus using plasma DNA methylation patterns from women carrying male fetuses, because placental tissue containing male fetuses that should not be affected by X-inactivation has a unique methylation pattern that differs from other maternal tissues with more or less X-inactivation in specific regions.
[0207] DNA extracted from maternal buffy coat samples was further sequenced using single-molecule real-time sequencing. 2.3 million CCS were obtained with a median sub-read depth of 5x per CCS. Results confirmed that maternal Hap I carried a premutation allele with 124 CGG repeats, while maternal Hap II carried a wild-type allele with 43 CGG repeats. DNA extracted from fetal chorionic villus sampling was further sequenced using single-molecule real-time sequencing. 1.1 million CCS were obtained with a median sub-read depth of 4x per CCS. Results confirmed that the unborn fetus carried the wild-type allele.
[0208] E. Distribution of CpG sites in the human genome The longer the DNA fragment, the greater the probability that the fragment will have multiple CpG sites, which can be used for methylation pattern or other analyses.
[0209] Figure 35 shows the distribution of CpG sites within 500 bp regions across the human genome. Column 1 shows the number of CpG sites. Column 2 shows the number of 500 bp regions with that number of CpG sites. Column 3 shows the percentage of all regions represented by a region with a particular number of CpG sites. For example, 86.14% of 500 bp regions have at least one CpG site. Additionally, 11.08% of 500 bp regions have at least 10 CpG sites.
[0210] Figure 36 shows the distribution of CpG sites within 1 kb regions across the human genome. Column 1 shows the number of CpG sites. Column 2 shows the number of 1 kb regions with that number of CpG sites. Column 3 shows the percentage of all regions represented by a region with a particular number of CpG sites. For example, 91.67% of 500 bp regions have at least one CpG site. Also, 32.91% of 500 bp regions have at least 10 CpG sites.
[0211] Figure 37 shows the distribution of CpG sites within 3 kb regions across the human genome. Column 1 shows the number of CpG sites. Column 2 shows the number of 3 kb regions with that number of CpG sites. Column 3 shows the percentage of all regions represented by a region with a particular number of CpG sites. For example, 92.45% of 3 kb regions have at least one CpG site. Furthermore, 87.09% of 3 kb regions have at least 10 CpG sites.
[0212] In some embodiments, different numbers of CpG sites and different size cutoffs are used to maximize the sensitivity and specificity of placenta-specific marker identification and tissue of origin analysis.In general, CpG sites occur more frequently than SNPs.A DNA fragment of a given size is more likely to have more CpG sites than SNPs.The table shown above may show a lower percentage for regions with the same number of SNPs as CpG sites, because there are fewer SNPs than CpG sites within the same size region.As a result, using CpG sites allows more fragments to be used than using only SNPs, providing better statistics.
[0213] F. Examples of tissue-of-origin analysis In an embodiment, the tissue-of-origin analysis of maternal plasma can be expanded to two or more organs / tissues, including T cells, B cells, neutrophils, liver, and placenta. Nine maternal DNA samples were sequenced using single-molecule real-time sequencing. According to the methylation status matching approach described in this disclosure, the plasma DNA methylation patterns were used to estimate the placenta's contribution to maternal plasma DNA. For this methylation status matching analysis, in one embodiment, the methylation patterns of each DNA molecule in the maternal plasma DNA sample that was at least 500 bp in length and contained at least five CpG sites were compared with reference tissue methylation profiles obtained from bisulfite sequencing. Five tissues, including neutrophils, T cells, B cells, liver, and placenta, were used as reference tissues. Plasma DNA molecules were assigned to the tissue corresponding to the highest methylation status matching score for that plasma DNA molecule. The percentage of plasma DNA molecules assigned to a tissue compared to other tissues is considered to be the tissue's proportional contribution to the maternal plasma DNA of that sample. In embodiments, the sum of the proportional contributions of neutrophils, T cells, and B cells in maternal plasma provided a surrogate for the proportional contribution of hematopoietic cells.
[0214] Figure 38 shows the proportional contribution of DNA molecules from different tissues in maternal plasma using methylation status matching analysis. Column 1 shows sample identity. Column 2 shows hematopoietic cell contribution as a percentage. Column 3 shows liver contribution as a percentage. Column 4 shows placenta contribution as a percentage. Figure 38 shows that the major contributor of maternal plasma DNA is hematopoietic cells (median: 55.9%), which is consistent with previous reports (Sun et al. Proc Natl Acad Sci USA. 2015;112:E5503-12, Zheng et al. Clin Chem. 2012;58:549-58).
[0215] Figures 39A and 39B show the relationship between placental contribution and fetal DNA fraction estimated by the SNP approach. The x-axis shows the fetal fraction determined by the SNP approach. The y-axis shows the placental contribution as a percentage determined in maternal plasma using methylation status matching analysis. Figure 39A shows a good correlation between the placental contribution determined by methylation status matching analysis and the fetal DNA fraction estimated by SNP (Pearson's r = 0.95, P value < 0.0001). Tissue deconvolution analysis of maternal plasma DNA was further performed by comparing plasma DNA methylation densities determined by single-molecule real-time sequencing with various reference tissue methylation profiles obtained from bisulfite sequencing according to quadratic programming (Sun et al. Proc Natl Acad Sci USA. 2015;112:E5503-12). Figure 39B shows that using a methylation density-based approach, the correlation between placental contribution (Sun et al. Proc Natl Acad Sci USA. 2015;112:E5503-12) and fetal DNA fraction was reduced compared to using methylation status matching analysis (Pearson's r=0.65, P-value=0.059).
[0216] These data suggested that it was feasible to estimate the proportion of DNA molecules contributed by different tissues in maternal plasma DNA samples. In another embodiment, this method can also be used to measure DNA molecules from different cell types or tissues in samples obtained after invasive solid tissue biopsy or from solid tissues obtained after surgery. In some embodiments, using methylation patterns at the single DNA molecule level to estimate the proportional contribution of different tissues to maternal plasma DNA is superior to approaches based on aggregated methylation densities from all sequenced plasma DNA molecules across the genome.
[0217] G. Exemplary Methods 40 shows a method 4000 for analyzing a biological sample obtained from a woman carrying a fetus. The biological sample may include a plurality of cell-free DNA molecules from the fetus and the woman.
[0218] At block 4010, sequence reads corresponding to a plurality of cell-free DNA molecules may be received. In some embodiments, method 4000 may include performing sequencing of the cell-free DNA molecules.
[0219] In block 4020, the size of the plurality of cell-free DNA molecules may be measured. The measurement may include aligning the sequence reads to a reference genome. In some embodiments, the measurement may include full-length sequencing and counting the number of nucleotides in the full-length sequence. In some embodiments, the measurement may include physically separating the plurality of cell-free DNA molecules from the biological sample from other cell-free DNA molecules in the biological sample, where the other cell-free DNA molecules have a size smaller than a cutoff value. The physical separation may include any technique described herein, including the use of beads.
[0220] In block 4030, a set of cell-free DNA molecules from the plurality of cell-free DNA molecules can be identified as having a size equal to or greater than a cutoff value. The cutoff value can be 200 nt or greater. The cutoff value can be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cutoff value can be any cutoff value described herein for long cell-free DNA molecules. The size can be the number of CpG sites rather than the length of the molecule. For example, the cutoff value can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more CpG sites.
[0221] In block 4040, the methylation status of each of the plurality of sites can be determined for one cell-free DNA molecule of the set of cell-free DNA molecules. The plurality of sites can include at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more CpG sites. At least one of the plurality of sites can be methylated. Two sites of the plurality of sites can be separated by at least 160 nt, 170 nt, 180 nt, 190 nt, 200 nt, 250 nt, or 500 nt. The method can include sequencing the plurality of cell-free DNA molecules to obtain sequence reads, and determining the methylation status of the site by measuring characteristics corresponding to the nucleotide of the site and the nucleotide adjacent to the site. For example, methylation can be determined as described in U.S. Application No. 16 / 995,607.
[0222] At block 4050, a methylation pattern may be determined. The methylation pattern may indicate the methylation status at each site of the plurality of sites.
[0223] At block 4060, the methylation pattern may be compared to one or more reference patterns. Each of the one or more reference patterns may be determined for a particular tissue type. In some embodiments, the comparison may include determining the number of sites that match the reference pattern.
[0224] The reference pattern of one or more reference patterns can be determined by measuring the methylation density of each reference site of a plurality of reference sites using DNA molecules from a reference tissue. The methylation density of each reference site of the plurality of reference sites can be compared to one or more threshold methylation densities. Each reference site of the plurality of reference sites can be identified as methylated, unmethylated, or uninformative based on comparing its methylation density to one or more threshold methylation densities, and the plurality of sites is a plurality of reference sites identified as methylated or unmethylated. Uninformative sites can include those whose methylation densities are between two threshold methylation densities. For example, the methylation index of an uninformative site can be 30 to 70 or any other range, as described herein.
[0225] In step 4070, the tissue of origin of the cell-free DNA molecule can be determined using the methylation pattern. The tissue of origin can be placenta. The tissue of origin can be fetal or maternal. The method can further include determining that the tissue of origin is the reference tissue if the methylation pattern matches the reference pattern, as described with reference to FIG. 22 . A match can refer to a perfect match. In some embodiments, the tissue of origin can be determined to be the reference tissue if the methylation pattern matches a certain percentage of the sites in the reference pattern. For example, the methylation pattern can match at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, or more of the sites in the reference pattern.
[0226] The method may include determining the tissue of origin by comparing the methylation pattern with a first reference methylation pattern from a first reference tissue of the plurality of reference tissues to determine a similarity score. The similarity score may be calculated using a methylation state matching process or a beta distribution probability model described herein. The similarity score may be compared to a threshold. If the similarity score exceeds the threshold, the tissue of origin may be determined to be the first reference tissue. The similarity score may be a first similarity score. The method may further include calculating a threshold by determining a second similarity score by comparing the methylation pattern with a second reference methylation pattern from a second reference tissue of the plurality of reference tissues. The first reference tissue and the second reference tissue may be different tissues. The threshold may be a second similarity score. The first reference tissue may have the highest similarity score compared to all other reference tissues.
[0227] The first reference methylation pattern may include a first subset of sites having at least a first probability of being methylated for the first reference tissue. For example, the first subset of sites may be sites that are methylated or that are considered normally methylated. The first reference methylation pattern may include a second subset of sites having at most a second probability of being methylated for the first reference tissue. For example, the second subset of sites may be sites that are unmethylated or that are considered normally unmethylated. Determining the similarity score may include increasing the similarity score if one site of the plurality of sites is methylated and the site of the plurality of sites is within the first subset of sites, and decreasing the similarity score if one site of the plurality of sites is methylated and the site of the plurality of sites is within the second subset of sites. The similarity score may be determined in a manner similar to the methylation state matching approach described herein.
[0228] The first reference methylation pattern comprises a plurality of sites, each of which is characterized by a probability of being methylated and a probability of being unmethylated in the first reference tissue. A similarity score can be determined by determining the probability of the methylation state of each of the plurality of sites in the reference tissue corresponding to the methylation state of the site in the cell-free DNA molecule. The similarity score can be determined by calculating the product of the plurality of probabilities. The product can be the similarity score. The probability can be determined by a beta distribution, similar to the approach described herein.
[0229] Method 4000 may further include determining a tissue of origin for each cell-free DNA molecule of the set of cell-free DNA molecules. This determination may include determining the methylation state of each of the plurality of respective sites, each of the plurality of respective sites corresponding to a cell-free DNA molecule. Determining the tissue of origin may further include determining a methylation pattern. Furthermore, determining the tissue of origin may also include comparing the methylation pattern to at least one reference pattern of one or more reference patterns. In some embodiments, the comparison of methylation patterns may be similar to that shown in FIG. 22 and the accompanying description. In FIG. 22, placenta, liver, blood cells, and colon are examples of reference tissues with the reference patterns shown. FIG. 38 shows hematopoietic cells as another example of a reference tissue.
[0230] In some embodiments, the amount of cell-free DNA molecules corresponding to each tissue of origin can be determined. Each tissue of origin can include each reference tissue of multiple reference tissues. The fractional contribution of the tissue of origin can be determined using the amount of cell-free DNA molecules corresponding to each tissue of origin. For example, the tissue of origin can be the placenta. Other tissues of origin can include hematopoietic cells and the liver. For example, the fractional contribution of the placenta can be determined by dividing the amount of cell-free DNA molecules by the total amount of cell-free DNA molecules corresponding to all tissues of origin. In some embodiments, the fraction calculated by dividing the amount of cell-free DNA molecules by the total amount of cell-free DNA molecules can be related to the fractional contribution via a function or a set of calibration data points. Both the function and the set of calibration data points can be determined from multiple calibration samples with known fractional contributions of the tissue of origin. Each calibration data point can specify a fractional contribution corresponding to a calibrated value of the fraction. The function can represent a linear or nonlinear fit of the calibration data points and can relate the fractional contribution to the fraction of the tissue of origin or other parameters, including the tissue of origin. An embodiment for determining fractional contribution can be similar to that described in Figures 39A and 39B.
[0231] A machine learning model may be used to determine the tissue of origin. The model may be trained by receiving multiple training methylation patterns, each training methylation pattern having a methylation state at one or more sites of the multiple sites, and each training methylation pattern determined from DNA molecules from a known tissue. Each molecule from the known tissue may be cellular DNA. The training may include storing multiple training samples, each training sample including one of the multiple training methylation patterns and a label indicating the known tissue corresponding to the training methylation pattern. The training may include using the multiple training samples to optimize parameters of the model based on the output of the model, which matches or does not match the corresponding label when the multiple training methylation patterns are input to the model. The parameters may include a first parameter indicating whether one site of the multiple sites has the same methylation state as another site of the multiple sites. For example, the model may be similar to the paired comparison in FIG. 24. The parameters may include a second parameter indicating the distance between the sites of the multiple sites. In some embodiments, the machine learning model may not require alignment of the methylation sites to a reference genome. The output of the model may specify the tissue corresponding to the input methylation pattern.
[0232] The machine learning model may be a convolutional neural network (CNN) or any model described herein. The model may include, but is not limited to, linear regression, logistic regression, deep recurrent neural network (e.g., long short-term memory, LSTM), Bayesian classifier, hidden Markov model (HMM), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering for applications with noise (DBSCAN), random forest algorithm, and support vector machine (SVM).
[0233] Paternity can be determined by method 4000. The tissue of origin can be a fetus. The method can further include aligning one sequence read of the plurality of sequence reads to a first region of a reference genome, the first region including a plurality of sites corresponding to alleles, the plurality of sites including a threshold number of sites; determining a first haplotype using each allele present at each site of the plurality of sites; comparing the first haplotype with a second haplotype corresponding to a male subject; and using the comparison to determine a classification of the likelihood that the male subject is the father of the fetus. The male subject can be considered to have a high probability of being the father if the haplotypes match, or a low probability of being the father if the haplotypes do not match. In some embodiments, the first haplotype can be compared to both haplotypes of the male subject.
[0234] In an embodiment, paternity can be tested when the tissue of origin is a fetus by aligning one sequence read from the plurality of sequence reads to a first region of a reference genome. The first region can include a first plurality of sites corresponding to alleles. The plurality of sites can include a threshold number of sites. The threshold number of sites can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more sites. The alleles at each site in the plurality of sites can be compared with the alleles at the corresponding site in the genome of the male subject. A classification of the likelihood that the male subject is the father of the fetus can be determined using the comparison. The male subject can be considered likely to be the father if a certain number or percentage of alleles match, or unlikely to be the father if less than that number or percentage match. The cutoff percentage can be 100%, 90%, 80%, or 70%.
[0235] In some embodiments, haplotypes can be determined. The method can include aligning sequence reads corresponding to each cell-free DNA molecule of a set of cell-free DNA molecules to a reference genome. The sequence reads can be identified as corresponding to haplotypes present in a woman. The haplotypes present in a woman can be known from genotyping the woman. In some embodiments, the woman's haplotype can be known by analyzing the concentration of DNA fragments of the haplotype in a biological sample from the woman. The tissue of origin can be determined as a fetus using methylation patterns. The haplotype can be determined to be a fetal haplotype of maternal inheritance.
[0236] Inheritance of haplotypes can be determined using methylation of a reference tissue rather than using known methylation profiles such as those associated with imprinted loci. A match or similarity score between a methylation pattern and a reference pattern can preclude knowledge of whether a given allele or site is methylated based on the parent from which it was inherited.
[0237] A haplotype can be identified as carrying a disease-causing genetic mutation or change. Identifying a haplotype as carrying a disease-causing genetic mutation can include identifying a genetic mutation or change in a first sequence read. The genetic mutation can include a single-base difference, a deletion, or an insertion. A first methylation level can be measured in a second sequence read corresponding to a first genomic location within a first distance of the first sequence read. A second methylation level can also be measured in a third sequence read corresponding to a second genomic location within a second distance of the first sequence read. The first distance can be 100 nt, 200 nt, 300 nt, 400 nt, 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 2 knt, 5 knt, or 10 knt. The second sequence read and the third sequence read can be on the same chromosome arm as the first sequence read. The first methylation level and the second methylation level can be associated with a genetic mutation or alteration. The first methylation level and the second methylation level can be greater than one or two threshold levels associated with the genetic mutation or alteration. The threshold levels can be determined using subjects known to have or not have a genetic mutation or alteration. The method can include classifying the fetus as having a high probability of having a disease caused by the genetic mutation or alteration.
[0238] A fetal-specific methylation pattern can be determined. The method can include, for each cell-free DNA molecule of the set of cell-free DNA molecules, aligning sequence reads corresponding to the cell-free DNA molecule to a reference genome. The method can include identifying sequence reads as corresponding to a region. The region can be determined by receiving a plurality of fetal sequence reads corresponding to a plurality of fetal DNA molecules from fetal tissue. The method can include receiving a plurality of maternal sequence reads corresponding to a plurality of maternal DNA molecules. The method can include, for each fetal sequence read of the plurality of fetal sequence reads, determining a fetal methylation state of each methylation site of a plurality of methylation sites within the region. The method can include, for each maternal sequence read of the plurality of maternal sequence reads, determining a maternal methylation state of each methylation site of the plurality of methylation sites.
[0239] A method for determining a fetal-specific methylation pattern may include determining a value of a parameter that characterizes the amount of sites where the fetal methylation state differs from the maternal methylation state. The method may include comparing the value of the parameter to a threshold. The parameter may be the percentage of sites that differ between fetal and maternal DNA molecules. The percentage may be a mismatch score, as described herein. The threshold may indicate a minimum level of mismatch score and may be 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or higher. In some embodiments, the threshold may represent the average mismatch score of maternal or fetal DNA molecules. The method may include determining that the value of the parameter exceeds the threshold. In some embodiments, a certain percentage of maternal or fetal DNA molecules may be required to have a value of the parameter that exceeds the threshold. For example, the percentage may be 50%, 60%, 70%, 80%, 90%, or higher. In some embodiments, a certain percentage of fetal DNA molecules corresponding to the region may be required to have the fetal-specific methylation pattern. For example, the percentage may be 40%, 50%, 60%, 70%, 80%, or more. This method may be similar to the method described in FIG.
[0240] The method may include enriching a biological sample for cell-free DNA molecules from the tissue of origin. Enriching the biological sample may include selecting and amplifying a set of cell-free DNA molecules. As described herein, enrichment may include size-based selection. In some embodiments, enrichment may include methylation pattern-based selection. For example, methyl-CpG binding domain (MBD)-based capture and sequencing may be used. Cell-free DNA may be incubated with a tagged MBD protein capable of binding to methylated cytosines. The protein-DNA complex may then be precipitated with antibody-conjugated magnetic beads. DNA molecules with more methylated CpG sites may be preferentially enriched for downstream analysis.
[0241] III. Changes in long cell-free DNA fragments with gestational age The amount of long cell-free DNA fragments can vary with gestational age.Long cell-free DNA fragments can be used to determine gestational age.In addition, long cell-free DNA fragments can be more abundant in specific terminal motifs than short cell-free DNA fragments, and the relative amount of specific terminal motifs can vary with gestational age.The amount of terminal motifs can also be used to determine gestational age.The deviation between the gestational age determined using long cell-free DNA fragments and the gestational age determined by other clinical techniques can indicate pregnancy-related disorders.In some embodiments, long cell-free DNA fragments can be used to determine the likelihood of pregnancy-related disorders without necessarily determining gestational age.
[0242] A. Size analysis of fetal and maternal DNA Plasma DNA from two pregnant women in the first trimester (gestational age: 13 weeks), two in the second trimester (gestational age: 21-22 weeks), and five in the third trimester (gestational age: 38 weeks) was sequenced using single-molecule real-time (SMRT) sequencing (PacBio). For each case, a median of 176 million subreads (range: 49-685 million) were obtained, of which 128 million subreads (range: 35-507 million) could be aligned to the human reference genome (hg19). Each molecule in the SMRT well was sequenced an average of 107 times. A median of 965,308 (range: 251,686-2,871,525) high-quality circulating consensus sequence (CCS) reads, defined as CCS reads with at least three subreads, could be used for downstream analysis.
[0243] All sequenced molecules from samples obtained from each trimester were pooled together for size analysis. There were a total of 1.94 million, 5.09 million, and 4.45 million cell-free DNA molecules for the first-, second-, and third-trimester maternal plasma samples, respectively.
[0244] Figures 41A and 41B show the size distribution of cell-free DNA molecules from first-, second-, and third-trimester maternal plasma samples within the size range of 0 to 5 kb. The x-axis shows size. The y-axis shows frequency. Size distributions are plotted on a linear scale on the y-axis ranging from 0 to 5 kb for Figure 41A and on a logarithmic scale on the y-axis ranging from 0 to 5 kb for Figure 41B. Plasma DNA from all three trimesters showed the expected major peak at 166 bp, as shown in Figure 41A, and a series of major peaks occurring in a periodic pattern spanning molecules in the 1 kb to 2 kb range, as shown in Figure 41B.
[0245] Figure 42 is a table showing the percentage of long plasma DNA molecules at different gestational ages. Column 1 shows the gestational age associated with the plasma sample. Column 2 shows the percentage of DNA molecules longer than 500 bp. Column 3 shows the percentage of DNA molecules longer than 1 kb. Compared with first and second trimesters, third trimesters had an increased frequency of plasma DNA molecules ≥ 500 bp. The percentages of plasma DNA molecules longer than 500 bp were 15.8%, 16.1%, and 32.3% for first, second, and third trimesters, respectively. The percentages of plasma DNA molecules longer than 1 kb were 11.3%, 10.6%, and 21.4% for first, second, and third trimesters, respectively. While first- and second-trimester maternal plasma showed similar percentages of long cell-free DNA molecules, third-trimester maternal plasma had approximately twice the percentage of long DNA molecules.
[0246] For all maternal plasma DNA samples analyzed for this disclosure, the genotypes of DNA extracted from their paired maternal buffy coat and fetal samples were determined using the Infinium Omni2.5Exome-8 Beadchip (Illumina) on the iScan System, an array hybridization-based genotyping method. Fetal samples were obtained by chorionic villus sampling, amniocentesis, or placental sampling, depending on whether the case was in the first, second, or third trimester. A median of 203,647 informative single nucleotide polymorphisms (SNPs) for which the mother was homozygous and the fetus was heterozygous were identified for each case. When the sequenced DNA molecules for all cases from each trimester were pooled together, a total of 1,362, 2,984, and 6,082 DNA molecules covering fetal-specific alleles were identified for the first, second, and third trimesters, respectively. A median of 210,820 informative SNPs for which the mother was heterozygous and the fetus was homozygous were identified for each case. A total of 30,574, 65,258, and 78,346 DNA molecules covering maternal-specific alleles were identified for the first, second, and third trimesters, respectively. Among all maternal plasma samples, the median fetal DNA fraction determined from sequencing data for DNA molecules 600 bp or smaller was 15.6% (range, 7.6-26.7%).
[0247] Figures 43A and 43B show size distributions of DNA molecules covering fetal-specific alleles from maternal plasma at first, second, and third trimesters. The x-axis shows size. The y-axis shows frequency. Size distributions are plotted on a linear scale on the y-axis ranging from 0 to 3 kb for Figure 43A and on a logarithmic scale on the y-axis ranging from 0 to 3 kb for Figure 43B.
[0248] Figures 44A and 44B show size distributions of DNA molecules covering maternal-specific alleles from maternal plasma at first, second, and third trimesters. The x-axis shows size. The y-axis shows frequency. Size distributions are plotted on a linear scale on the y-axis ranging from 0 to 3 kb for Figure 44A and on a logarithmic scale on the y-axis ranging from 0 to 3 kb for Figure 44B.
[0249] As shown in Figures 43A-44B, plasma DNA molecules covering fetal and maternal-specific alleles from all three trimesters exhibit a long-tailed distribution, suggesting the presence of long DNA molecules derived from both fetal and maternal sources in all three trimesters.
[0250] Figure 45 is a table of the percentage of long fetal and maternal plasma DNA molecules at different gestational stages. Column 1 shows the gestational age associated with the plasma sample. Column 2 shows the percentage of fetal DNA molecules longer than 500 bp. Column 3 shows the percentage of maternal DNA molecules longer than 500 bp. Column 4 shows the percentage of fetal DNA molecules longer than 1 kb. Column 5 shows the percentage of maternal DNA molecules longer than 1 kb. Among the pool of DNA molecules in maternal plasma, those covering fetal-specific alleles (placental origin) had a smaller percentage of long DNA molecules compared to those covering maternal-specific alleles. The percentage of long plasma DNA molecules covering fetal-specific alleles with a size greater than 500 bp was 19.8%, 23.2%, and 31.7% for first, second, and third trimesters, respectively. The percentage of long plasma DNA molecules covering fetal-specific alleles with a size greater than 1 kb was 15.2%, 16.5%, and 19.9% for first, second, and third trimesters, respectively.
[0251] Despite the fact that maternal plasma in the first and second trimesters contained a smaller proportion of long DNA molecules, and fetal DNA molecules contained fewer long DNA molecules in all three trimesters compared to the third trimester, our previous disclosure and the method described herein allow for the analysis of a significant proportion of long DNA molecules in plasma, which was previously impossible with short-read sequencing techniques. Furthermore, different size selection strategies, including but not limited to electrophoresis, chromatography, and bead-based methods, can be used to enrich long DNA fragments in plasma samples.
[0252] Figures 46A, 46B, and 46C show plots of the percentage of fetal-specific plasma DNA fragments in specific size ranges across different gestational stages. The gestational ages of the evaluated pregnancies were verified by ultrasound to confirm age. Figure 46A shows the results for DNA fragments 150 bp or shorter. Figure 46B shows the results for DNA fragments between 150 and 600 bp. Figure 46C shows the results for DNA fragments 600 bp or longer. The graphs have the percentage of fetal-specific fragments on the y-axis and the gestational age on the x-axis. As shown in the graphs, the percentage of fetal-specific fragments shorter than 150 bp (Figure 46A) and longer than 600 bp (Figure 46C) both achieve a certain discriminatory power in distinguishing between late-trimester samples and first- and mid-trimester samples compared with the percentage of fetal-specific fragments in the 150-600 bp range (Figure 46B). The percentage of fetal-specific fragments longer than 600 bp may provide the best discriminatory power. This conclusion was supported by the fact that the absolute minimum distance between the third trimester group and the combined first and second trimester group was 0.38 when the proportion of fetal-specific fragments shorter than 150 bp was used, whereas the corresponding value was 3.76 when the proportion of fetal-specific fragments greater than 600 bp was used. These results suggested that the use of long DNA molecules to reflect pathophysiological conditions is superior to the use of short DNA molecules.
[0253] B. Plasma DNA end analysis In addition to size, the first nucleotide of the 5' end of both the Watson and Crick strands of each sequenced DNA molecule was determined separately. This analysis consisted of four types of termini: A-, C-, G-, and T-termini. The percentage of plasma DNA molecules with specific termini from maternal plasma samples obtained from each trimester was calculated. The percentages of A-, C-, G-, and T-termini at each fragment size were further analyzed.
[0254] Figures 47A, 47B, and 47C show graphs of the percentage base content of the 5' end of cell-free DNA molecules from first-, second-, and third-trimester maternal plasma across the fragment size range of 0 to 3 kb. Figure 47A shows first-trimester maternal plasma. Figure 47B shows second-trimester maternal plasma. Figure 47C shows third-trimester maternal plasma. Percentage base content is shown on the y-axis. Fragment size in base pairs is shown on the x-axis. As can be seen in the graphs, C-termini were overrepresented in many size ranges (mostly less than 1 kb) and varied across different size ranges for first-, second-, and third-trimester samples. The plasma DNA end patterns for third-trimester samples appeared different from those for first- and second-trimester samples. For example, the T- and G-end curves were mixed across the 105- to 172-bp size range but diverged in the first- and second-trimester samples. For longer fragments (e.g., greater than about 1 kb), the C-terminal fragment is not the most abundant fragment: the G-terminal fragment overtakes the C-terminal fragment at about 1 kb, and then the A-terminal fragment becomes more abundant than the G-terminal fragment at about 2 kb.
[0255] Figure 48 is a table of the percentage of terminal nucleotide bases among short and long cell-free DNA molecules from first-, second-, and third-trimester maternal plasma. Column 1 shows the terminal base of the molecule. Column 2 shows the expected percentage points and species. Column 3 shows the percentage of terminal species among fragments 500 bp or less for first-trimester maternal plasma. Column 4 shows the percentage of terminal species among fragments greater than 500 bp for first-trimester maternal plasma. Columns 5 and 6 are the same as columns 3 and 4, respectively, except for second-trimester maternal plasma and instead of first-trimester maternal plasma. Columns 7 and 8 are the same as columns 3 and 4, respectively, except for third-trimester maternal plasma and instead of first-trimester maternal plasma.
[0256] If cell-free DNA fragmentation were completely random, the proportions of terminal nucleotide bases should reflect the composition of the human genome, which is 29.5% A, 29.5% T, 20.5% C, and 20.5% G, as shown in the second column of Figure 48. In contrast to random fragmentation, the 5' ends of short cell-free DNA molecules ≤500 bp showed substantial overrepresentation of C-termini (30.4%, 30.4%, and 31.3% for maternal plasma in first, second, and third trimester, respectively), slight overrepresentation of G-termini (27.4%, 26.9%, and 25.3% for first, second, and third trimester, respectively), and underrepresentation of A-termini (19.8%, 19.4%, and 19.3% for first, second, and third trimester, respectively), and underrepresentation of T-termini (22.4%, 23.3%, and 24.1% for first, second, and third trimester, respectively).
[0257] However, compared with short cell-free DNA molecules, cell-free DNA molecules longer than 500 bp showed a significant increase in the proportion of A-termini (29.6%, 26.0%, and 26.7% for maternal plasma in the first, second, and third trimesters, respectively), a slight increase in the proportion of G-termini (31.0%, 29.5%, and 29.9% for the first, second, and third trimesters, respectively), a significant decrease in the proportion of T-termini (13.9%, 16.9%, and 16.4% for the first, second, and third trimesters, respectively), and a slight decrease in the proportion of C-termini (25.5%, 27.5%, and 27.1% for the first, second, and third trimesters, respectively).
[0258] Figure 49 is a table showing the percentage of terminal nucleotide bases among short and long cell-free DNA molecules covering fetal-specific alleles from maternal plasma at first, second, and third trimesters. Figure 50 is a table showing the percentage of terminal nucleotide bases among short and long cell-free DNA molecules covering maternal-specific alleles from maternal plasma at first, second, and third trimesters. Column 1 shows the terminal base of the molecule. Column 2 shows the expected percentage of points and species. Column 3 shows the percentage of terminal species between fragments of 500 bp or less for maternal plasma at first trimesters. Column 4 shows the percentage of terminal species between fragments greater than 500 bp for maternal plasma at first trimesters. Columns 5 and 6 are the same as columns 3 and 4, respectively, except for maternal plasma at second trimesters and instead of maternal plasma at first trimesters. Columns 7 and 8 are the same as columns 3 and 4, respectively, except for maternal plasma at third trimesters and instead of maternal plasma at first trimesters. Figures 49 and 50 show that such differences in the proportion of terminal nucleotide bases between short and long cell-free DNA molecules remained unchanged even when DNA molecules covering fetal and maternal-specific alleles were examined separately.
[0259] Figure 51 shows hierarchical clustering analysis of short and long plasma cell-free DNA molecules using 256 4-mer terminal motifs. Each column indicates the sample used to analyze the frequency of terminal motifs based on short fragments (indicated by cyan in the first row) and long fragments (indicated by yellow in the first row), respectively. Starting from the second row, each row indicates the type of terminal motif. The frequency of terminal motifs is shown as a series of color gradients according to the row's normalized frequency (z-score) (i.e., the number of standard deviations below or above the mean frequency across samples). Redder colors indicate a higher frequency of the terminal motif, while bluer colors indicate a lower frequency of the terminal motif.
[0260] In Figure 51, short and long cell-free DNA molecules were characterized by analyzing 4-mer terminal motif profiles. For each sequenced DNA molecule, the first 4-nucleotide sequence (4-mer motif) at the 5' end of both the Watson and Crick strands was determined separately. For each maternal plasma sample, the frequency of each plasma DNA terminal motif was calculated separately for short (≤500 bp) and long (>500 bp) plasma DNA molecules. Hierarchical clustering analysis based on the frequency of 256 4-mer terminal motifs showed that the terminal motif profiles of long DNA molecules across different maternal plasma samples formed distinct clusters from short DNA molecules. These results suggested that long and short DNA molecules had different fragmentation properties. In embodiments, the relative perturbation of these terminal motifs between long and short DNA molecules is used to indicate the contribution of cell-free DNA from cell death pathways, such as, but not limited to, apoptosis and necrosis. Increased activity from these cell death pathways may be associated with pregnancy-related and other disorders.
[0261] Figures 52A and 52B show principal component analysis (PCA) using 4-mer terminal motif profiles for classification analysis. Figure 52A shows short cell-free DNA molecules (≤500 bp) from different trimesters. Figure 52B shows long cell-free DNA molecules (>500 bp) from maternal plasma samples from different trimesters. The percentages in parentheses on the x- and y-axes represent the amount of variability explained by the corresponding component. Each blue dot represents a maternal plasma sample from the first trimester. Each yellow dot represents a maternal plasma sample from the second trimester. Each red dot represents a maternal plasma sample from the third trimester. The ellipses represent the 95% confidence level for grouping data points from a particular trimester. Compared to short cell-free DNA molecules (Figure 52A) (also described in U.S. Application No. 15 / 787,050), the 4-mer terminal motif profile of long cell-free DNA molecules (Figure 52B) provided a clearer separation between maternal plasma samples from the first, second, and third trimesters. In embodiments, terminal motif profiles of long plasma DNA molecules can be utilized for molecular gestational age assessment, either alone or in combination with other maternal plasma DNA characteristics, including but not limited to methylation level and size.
[0262] For example, a neural network was used to train a model to predict gestational age based on 256 terminal motifs, global methylation levels, and the proportion of fragments ≥ 600 bp in size. The output variables were 1, 2, and 3, representing first, second, and third trimesters. The input variables included 256 terminal motifs, global methylation levels, and the proportion of fragments ≥ 600 bp in size. A leave-one-out approach was used to evaluate the performance of predicting gestational age. For a dataset containing nine samples, the leave-one-out approach was performed by selecting one sample as a test sample and using the remaining eight samples to train a neural network-based model. Based on the established model, such test samples were determined to be ≥ 1, ≥ 2, or ≥ 3. This process was then repeated for the remaining untested samples. This training and testing process was repeated a total of nine times. By comparing the test results with clinical information regarding gestational age, eight of the nine samples (89%) were correctly predicted for gestational age. In another embodiment, such analysis may be performed using, for example but not limited to, Bayes' theorem, logistic regression, multiple regression and support vector machines, random forest analysis, classification and regression trees (CART), K-nearest neighbor algorithms.
[0263] All sequenced molecules from each trimester were then pooled together for downstream terminal motif analysis. The 256 terminal motifs were ranked according to their frequency among short and long plasma DNA molecules.
[0264] Figures 53-58 are tables of the 25 terminal motifs with the highest frequency for DNA fragments of specific lengths (shorter or longer than 500 bp) and for different gestational stages. Figures 53, 54, and 55 are tables containing terminal motifs sorted by their rank for short fragments (less than 500 bp). In Figures 53-55, column 1 shows the terminal motif. Column 2 shows the frequency rank of the motif for short fragments. Column 3 shows the frequency rank of the motif for long fragments. Column 4 shows the frequency of the motif for short fragments. Column 5 shows the frequency of the motif for long fragments. Column 6 shows the fold change (the frequency of the motif for short fragments divided by the frequency of the motif for long fragments).
[0265] Figures 56, 57, and 58 are tables containing terminal motifs sorted by their rank in long fragments (>500 bp). In Figures 56-58, column 1 shows the terminal motif. Column 2 shows the frequency rank of the motif in long fragments. Column 3 shows the frequency rank of the motif in short fragments. Column 4 shows the frequency of the motif in long fragments. Column 5 shows the frequency of the motif in short fragments. Column 6 shows the fold change (frequency of the motif in long fragments divided by the frequency of the motif in short fragments).
[0266] Figures 53 and 56 are from first trimester samples, Figures 54 and 57 are from second trimester samples, and Figures 55 and 58 are from third trimester samples.
[0267] Among the top 25 terminal motifs with the highest frequency among short plasma DNA molecules, 11 of them began with a CC dinucleotide. Overall, CC-initiated terminal motifs accounted for 14.66%, 14.66%, and 15.13% of short plasma DNA terminal motifs in maternal plasma during the first, second, and third trimesters, respectively. Among the top 25 terminal motifs with the highest frequency among long plasma DNA molecules, 4-mer motifs ending with a TT dinucleotide accounted for 9 of them in maternal plasma during the second and third trimesters and 10 of them in maternal plasma during the first trimester.
[0268] For each sequenced DNA molecule, the dinucleotide sequences of the third nucleotide (X) and the fourth nucleotide (Y) from the 5' end of both the Watson and Crick strands were determined separately. X and Y are one of the four nucleotide bases in DNA. There were 16 possible NNXY motifs: NNAA, NNAT, NNAG, NNAC, NNTA, NNTT, NNTG, NNTC, NNGA, NNGT, NNGG, NNGC, NNCA, NNCT, NNCG, and NNCC.
[0269] Figures 59A, 59B, and 59C show scatter plots of motif frequencies of 16 NNXY motifs between short and long plasma DNA molecules. Figure 59A shows results for the first trimester. Figure 59B shows results for the second trimester. Figure 59C shows results for the third trimester. Motif frequencies of long fragments are shown on the y-axis. Motif frequencies of short fragments are shown on the x-axis. Each circle represents one of the 16 NNXY motifs. The pairs of dotted lines in each scatter plot indicate a 1.5-fold increase (upper line) and decrease (lower line) in motif frequency of long plasma DNA molecules (>500 bp) compared to short plasma DNA molecules (≤500 bp). Circles located outside the shaded region represent motifs with a fold change of more than 1.5.
[0270] While the ends of short plasma DNA molecules showed a high frequency of 4-mer motifs beginning with a CC dinucleotide (CCNN) (Jiang et al. Cancer Discov 2020;10(5):664-673, Chan et al. Am J Hum Genet 2020;107(5):882-894), the ends of long plasma DNA molecules showed a greater than 1.5-fold increase in the frequency of 4-mer motifs ending in TT (NNTT) across all three trimesters (Figure 11). NNTT motifs accounted for 18.94%, 15.22%, and 15.30% of long plasma DNA terminal motifs in maternal plasma during the first, second, and third trimesters, respectively. In contrast, NNTT motifs accounted for only 9.53%, 9.29%, and 8.91% of short plasma DNA terminal motifs in maternal plasma during the first, second, and third trimesters, respectively.
[0271] As previously reported by Han et al., cell-free DNA newly released into plasma from dying cells was enriched in A-terminal fragments longer than 150 bp. DNA fragmentation factor beta (DFFB), the major intracellular nuclease involved in DNA fragmentation during apoptosis, was found to be involved in the generation of such fragments (Han et al. Am J Hum Genet 2020;106:202-214). In this study, we demonstrate that cell-free DNA molecules longer than 500 bp were also enriched in A-terminal fragments, suggesting that DFFB may also be involved in the generation of these fragments. In normal pregnancy, trophoblast apoptosis increases with advancing gestation (Sharp et al. Am J Reprod Immuno 2010;64(3):159-69). Indeed, our finding that the proportion of long DNA molecules covering fetal-specific alleles increases with advancing gestational age may reflect increased trophoblast apoptosis with advancing gestational age.
[0272] In embodiments, the method described herein can be used to analyze the long cell-free DNA molecules in maternal plasma for predicting, screening, and monitoring the progress of placenta-related pregnancy complications, including but not limited to preeclampsia, intrauterine growth restriction (IUGR), preterm labor, and gestational trophoblastic disease.In placenta-related pregnancy complications, such as preeclampsia (Leung et al. Am J Obstet Gynecol 2001;184:1249-1250), IUGR (Smith et al. Am J Obstet Gynecol 1997;177:1395-1401, Levy et al. Am J Obstet Gynecol 2002;186:1056-1061), and gestational trophoblastic disease, the level of trophoblast apoptosis has been reported to be elevated. Furthermore, elevated fetal DNA levels in maternal plasma have been reported in preeclampsia (Lo et al. Clin Chem 1999;45(2):184-8, Smid et al. Ann NY Acad Sci 2001;945:132-7), IUGR (Sekizawa et al. Am J Obstet Gynecol 2003;188:480-4), and preterm labor (Leung et al. Lancet 1998;352(9144):1904-5). We hypothesized that increased placental apoptosis in placenta-related pregnancy complications would increase the proportion of long cell-free DNA molecules of placental origin in maternal plasma samples. Therefore, long cell-free DNA molecules of placental origin themselves, as well as long DNA signatures including but not limited to A-terminal fragments and NNTT motifs, may serve as biomarkers of placental apoptosis.
[0273] Although 1-nucleotide and 4-nucleotide motifs are used in the above analysis, in other embodiments, motifs of other lengths may be used, for example, 2, 3, 5, 6, 7, 8, 9, 10, or more.
[0274] C. Exemplary Methods Long cell-free DNA fragments can be used to determine the gestational age of women carrying fetuses.The amount of long cell-free DNA fragments changes with gestational age and can be used to determine gestational age.The terminal motifs of cell-free DNA fragments also change with gestational age and can be used to determine gestational age.If the gestational age determined using long cell-free DNA fragments significantly deviates from the gestational age determined by other clinical techniques, the pregnant woman and / or fetus may be considered to have a pregnancy-related disorder.In some embodiments, it may not be necessary to determine gestational age to determine the likelihood of pregnancy-related disorder.
[0275] 1. Gestation period Figure 60 shows a method 6000 for analyzing a biological sample obtained from a woman carrying a fetus. Gestational age can be determined and used to classify the likelihood of a pregnancy-related disorder. The biological sample can include a plurality of cell-free DNA molecules from the fetus and the woman.
[0276] Sequence reads corresponding to a plurality of cell-free DNA molecules can be received. In some embodiments, sequencing can be performed to obtain the sequence reads.
[0277] At block 6020, the size of the plurality of cell-free DNA molecules may be measured. The size may be measured in a manner similar to that described in Figure 21. The size may be measured using sequence reads.
[0278] At block 6030, a first quantity of cell-free DNA molecules having a size greater than a cutoff value can be measured. The quantity can be the number, total length, or mass of the cell-free DNA molecules.
[0279] At block 6040, a value of a normalization parameter using the first amount can be generated. The value of the normalization parameter can be the first amount normalized by the total number of cell-free DNA molecules, the number of cell-free DNA molecules from the fetus or the mother, or the number of DNA molecules from a specific region. For example, the normalization parameter can be the proportion of fetal-specific fragments, as illustrated in Figures 46A-C.
[0280] In block 6050, the value of the normalization parameter may be compared to one or more calibration data points. Each calibration data point may specify a gestational age corresponding to a calibration value of the normalization parameter. For example, a gestational age of a particular gestational age or a particular number of weeks may correspond to a calibration value of the normalization parameter. The one or more calibration data points may be determined from a plurality of calibration samples having known gestational ages and including cell-free DNA molecules having a size greater than a cutoff value. In some embodiments, the calibration data points are determined from a function correlating gestational age with the value of the normalization parameter.
[0281] At block 6060, a gestational age may be determined using the comparison. The gestational age may be considered the period corresponding to the calibration value that is closest to the value of the normalized parameter. In some embodiments, the gestational age may be considered the most advanced period for which the value of the normalized parameter exceeds the calibration value.
[0282] The method may further include determining a reference gestational age of the fetus using ultrasound or the date of the woman's last menstrual period. The method may also include comparing the gestational age to the reference gestational age. The method may also further include using the comparison of the gestational age to the reference gestational age to determine a classification of the likelihood of a pregnancy-related disorder. For example, a difference between the gestational age and the reference gestational age may indicate a pregnancy-related disorder. The difference may be a different gestational age or a difference in gestational age by a minimum number of weeks (e.g., 1, 2, 3, 4, 5, 6, 7, or more weeks).
[0283] The method may further include using terminal motifs. For example, the method may further include determining a first subsequence corresponding to at least one end of cell-free DNA molecules having a size greater than a cutoff value. The first amount may be of cell-free DNA molecules having a size greater than the cutoff value and having the first subsequence at one or more ends of each cell-free DNA molecule. The first subsequence may be or may include 1, 2, 3, 4, 5, or 6 nucleotides. As illustrated in Figures 52A and 52B, terminal motifs may be used to determine gestational age through PCA analysis. Calibration samples may be used with different terminal motifs and known gestational ages and subjected to PCA analysis. Other classification and regression algorithms, such as linear discriminant analysis, logistic regression, support vector machines, linear regression, and nonlinear regression, may be used with terminal motifs. Classification and regression algorithms may associate gestational ages with specific terminal motifs and / or fragments of specific sizes.
[0284] The terminal motif can be any motif discussed in Figures 47-59 or 94. The rank or frequency of the terminal motif can be compared to the rank or frequency of the terminal motif in a calibration sample from a subject of known gestational age. The rank or frequency of the terminal motif can then be used to determine the gestational age. A terminal motif present at a rank or frequency that deviates from the rank or frequency determined from a reference sample of the same gestational age can indicate a pregnancy-related disorder.
[0285] Generating a value for the normalization parameter can include (a) normalizing the first amount by the total amount of cell-free DNA molecules having a size greater than the cutoff value; (b) normalizing the first amount by a second amount of cell-free DNA molecules having a size greater than the cutoff value and ending in a second subsequence, where the second subsequence is different from the first subsequence; or (c) normalizing the first amount by a third amount of cell-free DNA molecules having a size less than the cutoff value.
[0286] 2. Pregnancy-related disorders 61 illustrates a method 6100 for analyzing a biological sample obtained from a woman carrying a fetus. Embodiments may include classifying the likelihood of a pregnancy-related disorder without necessarily determining gestational age. The biological sample may include a plurality of cell-free DNA molecules from the fetus and the woman.
[0287] Sequence reads corresponding to a plurality of cell-free DNA molecules can be received. In some embodiments, sequencing can be performed to obtain the sequence reads.
[0288] At block 6120, the sizes of the plurality of cell-free DNA molecules may be measured. The sizes may be obtained in a manner similar to that described in Figure 21. Measuring the sizes may use the received sequence reads.
[0289] In block 6130, a first amount of cell-free DNA molecules having a size greater than a cutoff value can be measured. The cutoff value can be 200 nt or greater. The cutoff value can be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cutoff value can be any cutoff value described herein for long cell-free DNA molecules. The first amount can be a numerical value or a frequency.
[0290] In block 6140, a first value of a normalization parameter using the first amount may be generated. Generating the value of the normalization parameter may include measuring a second amount of cell-free DNA molecules having a size smaller than a cutoff value and calculating a ratio of the first amount and the second amount. The cutoff value may be the first cutoff value. The second cutoff value may be smaller than the first cutoff value. The second amount may include cell-free DNA molecules having a size smaller than the second cutoff value, or the second amount may include all cell-free DNA molecules in the plurality of cell-free DNA molecules. The normalization parameter may be a measure of the frequency of long cell-free DNA molecules.
[0291] At block 6150, a second value corresponding to an expected value of the normalized parameter for a healthy pregnancy may be obtained. The second value may depend on the gestational age of the fetus. The second value may be an expected value. In some embodiments, the second value may be a cutoff value that distinguishes from abnormal values.
[0292] Obtaining the second value may include obtaining the second value from a calibration table relating measurements of the pregnant woman to calibration values of the normalized parameter. The calibration table may be generated by obtaining a first table relating gestational ages to measurements of the pregnant female subject. A second table relating gestational ages to calibration values of the normalized parameter may be obtained. The data in the first and second tables may be from the same subject or different subjects. A calibration table relating measurements to calibration values may be created from the first and second tables. The calibration table may include a function relating calibration values to measurements.
[0293] The measurement of the pregnant female subject can be the time since the last menstrual period or a characteristic of an image (e.g., an ultrasound) of the pregnant female subject. The measurement of the pregnant female subject can be a characteristic of an image of the pregnant female subject. For example, the characteristic of the image can include the length, size, appearance, or anatomical structure of the female subject's fetus. The characteristic can include a biometric measurement, such as crown-rump length or femur length. The appearance of specific organs can be used, including the appearance of a four-chambered heart or the vertebrae of the spine. Gestational age can be determined by a physician from an ultrasound image (e.g., Committee on Obstetric Practice et al., "Methods for estimating the due date," Committee Opinion, No. 700, May 2017).
[0294] In some embodiments, the machine learning model may associate one or more calibration data points with characteristics of the image. The model may be trained by receiving a plurality of training images. Each training image may be from a female subject known to be free of a pregnancy-related disorder or known not to have a pregnancy-related disorder. The female subjects may have a variety of gestational ages. Training may include storing a plurality of training samples from the female subjects. Each training sample may include a known value of a normalization parameter associated with the training image. The model may be trained using the plurality of training samples by optimizing parameters of the model based on an output of the model that matches or does not match the image with the known value of the normalization parameter. The output of the model may specify a value of the normalization parameter corresponding to the image. A second value of the normalization parameter may be generated by inputting an image of a woman into the machine learning model.
[0295] At block 6160, a deviation between the first value of the normalization parameter and the second value of the normalization parameter may be determined. The deviation may be a separation value.
[0296] At block 6170, a classification of the likelihood of a pregnancy-related disorder may be determined using the deviation. If the deviation exceeds a threshold, a pregnancy-related disorder may occur. The threshold may indicate a statistically significant difference. The threshold may indicate a difference of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100%.
[0297] Pregnancy-related disorders may include preeclampsia, intrauterine growth restriction, invasive placentation, preterm birth, hemolytic disease of the newborn, placental insufficiency, hydrops fetalis, fetal malformations, hemolysis, elevated liver enzymes, and low platelet count (HELLP) syndrome, or systemic lupus erythematosus.
[0298] IV. Size and Termination Analysis for Pregnancy-Related Disorders Size and / or terminal analysis of long DNA molecules was used to determine the likelihood of preeclampsia. Such methods may also be applied to other pregnancy-related disorders. DNA extracted from maternal plasma samples of four pregnant women diagnosed with preeclampsia was subjected to single-molecule real-time (SMRT) sequencing (PacBio).
[0299] Figure 62 is a table showing clinical information for four pre-eclampsia cases. Column 1 shows the case number. Column 2 shows the gestational age in weeks at the time of blood collection. Column 3 shows the sex of the fetus. Column 4 shows clinical information regarding pre-eclampsia (PET).
[0300] M12804 was a case of severe preeclampsia (PET) and pre-existing IgA nephropathy. M12873 was a case of chronic hypertension with mixed mild PET. M12876 was a case of severe late-onset PET. M12903 was a case of severe late-onset PET with intrauterine growth restriction (IUGR). Five normotensive third-trimester maternal plasma samples were used as controls for subsequent analyses in this disclosure.
[0301] For the four preeclamptic and five normotensive third-trimester maternal plasma DNA samples analyzed for this disclosure, DNA extracted from their paired maternal buffy coat and placenta samples was genotyped with an Infinium Omni2.5Exome-8 Beadchip (Illumina) on an iScan System.
[0302] Plasma DNA concentrations in each sample were quantified using the Qubit dsDNA High Sensitivity Assay with a Qubit Fluorometer (ThermoFisher Scientific). The mean plasma DNA concentrations for preeclampsia and third-trimester cases were 95.4 ng / mL (range, 52.1–153.8 ng / mL) and 10.7 ng / mL (range, 6.4–19.1 ng / mL), respectively. The mean plasma DNA concentrations for preeclampsia cases were approximately 9-fold higher than those for third-trimester cases.
[0303] The mean fetal DNA fractions determined from sequencing data of DNA molecules ≤600 bp covering informative single nucleotide polymorphisms (SNPs) for which the mother was homozygous and the fetus was heterozygous were 22.6% (range, 16.6–25.7%) and 20.0% (range, 15.6–26.7%) for preeclamptic and normotensive third-trimester maternal plasma samples, respectively.
[0304] A. Size Analysis According to an embodiment of the present disclosure, size analysis was performed on maternal plasma samples from preeclamptic and normotensive third-trimester pregnancy. Figures 63A-63D and 64A-64D show the size distribution of plasma DNA molecules from preeclamptic and normotensive third-trimester pregnancy cases. The x-axis represents size, and the y-axis represents frequency. Size distributions are plotted on a linear x-axis scale ranging from 0 to 1 kb for Figures 63A-63D and on a logarithmic x-axis scale ranging from 0 to 5 kb for Figures 64A-64D. Figures 63A and 64A show sample M12804. Figures 63B and 64B show sample M12873. Figures 63C and 64C show sample M12876. Figures 63D and 64D show sample M12903.
[0305] The blue line represents the size distribution of all sequenced plasma DNA molecules pooled from five normotensive, third-trimester cases. The red line represents the size distribution of sequenced plasma DNA molecules from individual preeclampsia cases. In Figures 63A-63D, the blue lines represent the shorter peaks below 200 bp and the higher peaks between 300 and 400 bp. In Figures 64A-64D, the blue line corresponds to the higher peak at 1 kb.
[0306] Overall, plasma DNA size profiles in preeclamptic patients were shorter than those in normotensive, third-trimester pregnant women, with an increased height of the 166 bp peak and an increased proportion of DNA molecules shorter than 166 bp (Figures 63A-63D). These changes were more pronounced in two severe preeclamptic cases, M12876 and M12903. In M12903, a case of preeclampsia with intrauterine growth restriction (IUGR), the changes were even more dramatic.
[0307] Three of the four preeclamptic plasma samples showed a reduced proportion of long plasma DNA molecules with sizes between 200 and 5,000 bp (Figures 64B-64D). The proportions of long plasma DNA molecules greater than 500 bp in M12873, M12876, and M12903 were 11.7%, 8.9%, and 4.5%, respectively, whereas the proportion of long plasma DNA molecules in pooled sequencing data from five normotensive, third-trimester cases was 32.3%. A plasma sample from a case of severe preeclampsia with preexisting IgA nephropathy (PET) (M12804) showed a reduced proportion of shorter DNA molecules less than 2,000 bp but an increased proportion of longer DNA molecules greater than 2,000 bp compared to pooled sequencing data from five normotensive, third-trimester cases (Figure 2A). The proportion of long plasma DNA molecules in M12804 was 34.9%.
[0308] Figures 65A-65D and Figures 66A-66D show the size distribution of DNA molecules covering fetal-specific alleles from pre-eclamptic and normotensive third-trimester maternal plasma samples. Each of Figures A-D shows a different pre-eclamptic sample. The x-axis shows size. The y-axis shows frequency in Figures 65A-65D and cumulative frequency in Figures 66A-66D. In Figures 66A-66D, sizes range from 0 to 35 kb.
[0309] The blue line in each graph represents the size distribution of all sequenced plasma DNA molecules covering the fetal-specific alleles pooled from five normotensive, third-trimester cases. The red line in each graph represents the size distribution of plasma DNA molecules covering the fetal-specific alleles sequenced from individual preeclampsia cases. In Figures 65A-65D, the blue lines represent the shorter peaks below 200 bp and the higher peaks between 300 and 400 bp. In Figures 66A-66D, the blue lines represent the lower peaks between 100 and 1000 bp.
[0310] Figures 67A-67D and Figures 68A-68D show the size distribution of DNA molecules covering fetal-specific alleles from pre-eclamptic and normotensive third-trimester maternal plasma samples. Each of Figures A-D shows a different pre-eclamptic sample. The x-axis shows size. The y-axis shows frequency in Figures 67A-67D and cumulative frequency in Figures 68A-68D. In Figures 68A-68D, sizes range from 0 to 35 kb.
[0311] The blue line in each graph represents the size distribution of all sequenced plasma DNA molecules covering maternal-specific alleles pooled from five normotensive, third-trimester cases. The red line in each graph represents the size distribution of plasma DNA molecules covering maternal-specific alleles from individual preeclampsia cases. In Figure 67A, the blue lines represent the higher peaks below 200 bp and the higher peaks between 300 and 400 bp. In Figures 67B-67D, the blue lines represent the shorter peaks below 200 bp. In Figure 68A, the blue line corresponds to the higher line between 1000 and 10,000 bp. In Figures 68B-68D, the blue line corresponds to the lower line between 100 and 1000 bp.
[0312] The phenomenon of plasma DNA shortening was observed in both DNA molecules covering fetal-specific alleles (Figures 65B-65D and Figures 66B-66D) and maternal-specific alleles (Figures 67B-67D and Figures 68B-68D) in three of four preeclamptic plasma samples compared with normotensive late-pregnancy maternal plasma samples. The exception was case M12804, a patient with severe PET and pre-existing IgA nephropathy, which showed an increased proportion of shorter DNA molecules less than 1 kb and a decreased proportion of longer DNA molecules greater than 1 kb among their plasma DNA molecules covering fetal-specific alleles (Figures 65A and 66A). Indeed, the plasma DNA molecules covering maternal-specific alleles in case M12804 showed a lengthened size profile (Figures 67A and 68A).
[0313] Figures 69A and 69B are graphs of the percentage of short DNA molecules covering (A) fetal-specific alleles and (B) maternal-specific alleles in preeclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing. The y-axis shows the percentage of short DNA fragments less than 150 bp. The x-axis shows normal and PET samples.
[0314] In this embodiment, the proportion of short DNA molecules is defined as the percentage of maternal plasma DNA molecules with a size of less than 150bp.M12804 had pre-existing IgA nephropathy, but other samples did not, so this case was excluded from this analysis.The pre-eclampsia plasma sample group showed a significant increase in the proportion of short DNA molecules covering fetal-specific alleles (P=0.036, Wilcoxon rank sum test) and maternal-specific alleles (P=0.036, Wilcoxon rank sum test) compared with the normal blood pressure control plasma sample group.
[0315] Figures 70A and 70B are graphs of the percentage of short DNA molecules in preeclamptic and normotensive maternal plasma samples sequenced by (A) PacBio SMRT sequencing and (B) Illumina sequencing. The y-axis shows the percentage of short DNA fragments less than 150 bp.
[0316] In this study, the proportion of short DNA molecules was defined as the percentage of maternal plasma DNA molecules with a size less than 150 bp. Case M12804 was excluded from this analysis because it exhibited a different size profile compared to other preeclamptic cases in this cohort, likely due to preexisting IgA nephropathy. The preeclamptic plasma samples showed a significantly increased proportion of short DNA molecules (median: 28.0%, range: 25.8-35.1%) compared to normotensive control plasma samples (median: 12.1%, range: 8.5-15.8%) (P = 0.036, Wilcoxon rank-sum test). In contrast, in a previous cohort of four preeclamptic and four gestational age-matched normotensive maternal plasma DNA samples subjected to bisulfite conversion and Illumina sequencing, there was no significant difference in the proportion of short DNA molecules in preeclamptic and control plasma samples (P = 0.340, Wilcoxon rank sum test) (Figure 70B).
[0317] In some embodiments, a cutoff of 20% can be used for the percentage of short DNA molecules in maternal plasma samples sequenced by PacBio SMRT sequencing to determine whether a pregnancy is at high or low risk of developing preeclampsia. Maternal plasma samples with a percentage of short DNA molecules greater than 20% are determined to be at high risk of developing preeclampsia, while maternal plasma samples with a percentage of short DNA molecules less than 20% are determined to be at low risk of developing preeclampsia. Using this cutoff, both sensitivity and specificity were 100%. In some other embodiments, the cutoff for the percentage of short DNA molecules used can include, but is not limited to, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, etc. In another embodiment, the percentage of short DNA molecules in maternal plasma samples is used to monitor and evaluate the severity of preeclampsia during pregnancy.
[0318] In an embodiment, a size ratio indicating the relative proportion of short and long DNA molecules was calculated for each sample using the following equation:
number
[0319] Figure 71 is a size ratio graph showing the relative proportions of short and long DNA molecules in preeclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing. The y-axis shows the size ratio. The x-axis shows normal and PET samples. The preeclamptic plasma sample group showed a significantly higher size ratio compared to the normotensive control plasma sample group (P=0.016, Wilcoxon rank sum test).
[0320] In some embodiments, the size profile generated by long-read sequencing platforms, including but not limited to PacBio SMRT sequencing and Oxford Nanopore sequencing, can be used to predict the occurrence and severity of preeclampsia during pregnancy.In some embodiments, the size profile of plasma DNA molecules can be analyzed to monitor the progression of preeclampsia and the development of severe preeclampsia characteristics, including but not limited to liver damage and kidney damage.In some embodiments, the size parameters used in analysis can include, but are not limited to, the proportion of short or long DNA molecules, and the size ratio that indicates the relative proportion of short DNA molecules and long DNA molecules. The cutoffs used to determine short and long DNA categories may include, but are not limited to, 150 bp, 180 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 550 bp, 600 bp, 650 bp, 700 bp, 750 bp, 800 bp, 850 bp, 900 bp, 950 bp, 1 kb, etc. The size ranges used in determining the size ratio of short and long molecules may include, but are not limited to, 50-150 bp, 50-166 bp, 50-200 bp, 200-400 bp, 200-1000 bp, 200-5000 bp, or other combinations.
[0321] Size end analysis may include using the method described in method 6100 of FIG.
[0322] B. Fragment end analysis According to an embodiment of the present disclosure, fragment end analysis was performed on maternal plasma samples from preeclamptic and normotensive women in late pregnancy. For each sequenced plasma DNA molecule, the first nucleotide of the 5' end of both the Watson and Crick strands was determined. The proportions of T-, C-, A-, and G-terminal fragments were determined for each plasma DNA sample.
[0323] Figures 72A-72D show the percentage of different ends of plasma DNA molecules in preeclamptic and normotensive maternal plasma samples sequenced using PacBio SMRT sequencing. The x-axis shows normal third-trimester samples and PET samples. The y-axis shows the percentage of a given end. Figure 72A shows the percentage of T-terminus, Figure 72B shows the percentage of C-terminus, Figure 72C shows the percentage of A-terminus, and Figure 72D shows the percentage of G-terminus. The preeclamptic plasma sample group showed a significantly increased percentage of T-terminus plasma DNA molecules (P=0.016, Wilcoxon rank sum test) and a significantly decreased percentage of G-terminus plasma DNA molecules (P=0.016, Wilcoxon rank sum test) compared to the normotensive control plasma sample group.
[0324] Figure 73 shows a hierarchical clustering analysis of pre-eclampsia and normotensive third trimester maternal plasma DNA samples using four types of fragment termini (the first nucleotide at the 5' end of each strand): C-terminus, G-terminus, T-terminus, and A-terminus. Each column represents a plasma DNA sample. The first row indicates which group each sample belongs to: cyan represents normotensive third trimester maternal plasma DNA samples, and orange represents pre-eclampsia plasma DNA samples. Cyan covers the first five columns. Orange covers the last four columns.
[0325] Starting from the second row, each row indicates a type of fragment end. The frequency of the terminal motif is shown as a series of color gradients according to the row's normalized frequency (z-score) (i.e., the number of standard deviations below or above the mean frequency across samples). Redder colors indicate a higher frequency of the terminal motif, and bluer colors indicate a lower frequency of the terminal motif. Hierarchical clustering analysis based on the frequencies of the four types of fragment ends showed that the fragment end profiles of preeclamptic plasma DNA samples formed a different cluster than normotensive third-trimester plasma DNA samples.
[0326] In embodiments, for each sequenced DNA molecule, the dinucleotide sequence of the first nucleotide (X) and the second nucleotide (Y) can be determined separately from the 5'-end of both the Watson strand and the Crick strand. X and Y are one of the four nucleotide bases in DNA. There are 16 possible two-nucleotide terminal motifs, XYNN, namely, AANN, ATNN, AGNN, ACNN, TANN, TTNN, TGNN, TCNN, GANN, GTNN, GGNN, GCNN, CANN, CTNN, CGNN, and CCNN. According to embodiments of the present disclosure, for each sequenced DNA molecule, the dinucleotide sequence of the third nucleotide (X) and the fourth nucleotide (Y) can be determined separately from the 5'-end of both the Watson strand and the Crick strand. There are 16 possible two-nucleotide NNXY motifs. The first four-nucleotide sequence (4mer motif) of the 5'-end of both the Watson strand and the Crick strand can also be determined separately for each sequenced DNA molecule.
[0327] Figure 74 shows a hierarchical clustering analysis of maternal plasma DNA samples from pre-eclampsia and normotensive third trimesters using 16 dinucleotide motifs, XYNN (dinucleotide sequences of the first and second nucleotides from the 5' end). Figure 75 shows a hierarchical clustering analysis of maternal plasma DNA samples from pre-eclampsia and normotensive third trimesters using 16 dinucleotide motifs, NNXY (dinucleotide sequences of the third and fourth nucleotides from the 5' end). Figure 76 shows a hierarchical clustering analysis of maternal plasma DNA samples from pre-eclampsia and normotensive third trimesters using 256 tetranucleotide motifs (dinucleotide sequences of the first through fourth nucleotides from the 5' end).
[0328] In Figures 74-76, the first row indicates which group each sample belongs to, with cyan representing normotensive, third-trimester maternal plasma DNA samples and orange representing preeclamptic plasma DNA samples. Cyan covers the first five columns. Orange covers the last four columns. Starting with the second row, each row indicates the type of fragment end. The frequency of the terminal motif is shown as a series of color gradients according to the row's normalized frequency (z-score) (i.e., the number of standard deviations below or above the mean frequency across samples). Redder colors indicate a higher frequency of the terminal motif, while bluer colors indicate a lower frequency of the terminal motif.
[0329] These results suggest that the plasma DNA in preeclampsia samples and non-preeclampsia samples have different fragmentation characteristics. In one embodiment, to predict the onset of preeclampsia during pregnancy, terminal motif profiles generated from long-read sequencing platforms, including but not limited to PacBio SMRT sequencing and Oxford Nanopore sequencing, can be used. Although the above analysis used 1-nucleotide, 2-nucleotide, and 4-nucleotide motifs, in other embodiments, motifs of other lengths, such as 3, 5, 6, 7, 8, 9, 10, or more, can be used.
[0330] In some embodiments, fragment end analysis and tissue-of-origin analysis can be combined to improve the performance of predicting, detecting, and monitoring pregnancy-related conditions, including but not limited to preeclampsia. First, fragment end analysis can be performed on each maternal plasma sample to separate the plasma DNA molecules into four fragment end categories: T-, C-, A-, and G-end fragments. Then, using methylation status matching analysis according to an embodiment of the present disclosure, tissue-of-origin analysis can be performed separately using plasma DNA molecules from each fragment end category for each maternal plasma DNA sample. The proportional contribution of different tissues among one of the fragment end categories was defined as the percentage of plasma DNA molecules in the corresponding fragment end category assigned to the corresponding tissue compared to other tissues.
[0331] Three and five plasma DNA samples from pregnant women with and without preeclampsia were analyzed using single-molecule real-time sequencing. Median values of 658,722, 889,900, 851,501, and 607,554 were obtained for plasma fragments with A-, C-, G-, and T-termini, respectively. For A-termini fragments, the methylation patterns of any fragments with at least 10 CpG sites were compared with reference methylation profiles of neutrophils, T cells, B cells, liver, and placenta, according to the methylation state matching approach described in this disclosure. Plasma DNA fragments were assigned to the tissue corresponding to the highest score of matching methylation states between those tissues. Using this method, a median of 2.43% (range: 0.73–5.50%) of A-termini fragments were assigned to T cells (i.e., T cell contribution) across all analyzed samples. The C-, G-, and T-termini fragments were further analyzed using a similar method. A median T cell contribution of 3.20% (range: 1.55–5.19%), 3.52% (range: 1.53–6.27%), and 2.22% (0–7.79%) was observed for those fragments with C-, G-, and T-termini, respectively.
[0332] Figures 77A-77D show T cell contributions among DNA molecules belonging to different fragment end categories, i.e., (A) T-end, (B) C-end, (C) A-end, and (D) G-end, in preeclamptic and normotensive maternal plasma DNA samples. The x-axis represents normal third-trimester and PET samples. The y-axis represents T cell contribution as a percentage. The results showed that among G-terminal fragments, T cell contribution was significantly reduced in preeclamptic plasma samples compared with normotensive third-trimester plasma samples (P=0.036, Wilcoxon rank-sum test). In an embodiment, a 3% cutoff for T cell contributions among all G-terminal fragments in maternal plasma DNA samples can be used to determine whether a pregnancy is at high or low risk of developing preeclampsia.
[0333] C. Exemplary Methods
[00103] Figure 78 shows a method 7800 for analyzing a biological sample obtained from a woman carrying a fetus. The biological sample may include a plurality of cell-free DNA molecules from the fetus and the woman. The method may generate a likelihood classification of a pregnancy-related disorder. The pregnancy-related disorder may be pre-eclampsia or any pregnancy-related disorder described herein.
[0334] Sequence reads corresponding to a plurality of cell-free DNA molecules can be received.
[0335] At block 7810, the size of the plurality of cell-free DNA molecules can be measured. The size can be measured by alignment or nucleotide counting, or any technique described herein, including FIG.
[0336] In block 7820, a set of cell-free DNA molecules having a size greater than a cutoff value can be identified. The cutoff value can be any cutoff value for long cell-free DNA fragments, including 500nt, 600nt, 700nt, 800nt, 900nt, 1knt, 1.1knt, 1.2knt, 1.3knt, 1.4knt, 1.5knt, 1.6knt, 1.7knt, 1.8knt, 1.9knt, or 2knt. The cutoff value can be any cutoff value described herein for long cell-free DNA molecules.
[0337] At block 7830, a value for a terminal motif parameter using the first amount may be generated. A first amount of cell-free DNA molecules in the set having a first subsequence at one or more ends of the cell-free DNA molecules in the set may be measured. In some embodiments, the terminal motif parameter may be the first amount normalized by the total amount of all subsequences at the end. In some embodiments, the end may be the 3' end. In some embodiments, the end may be the 5' end.
[0338] The first subsequence can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides in length. The first subsequence can include the final nucleotide at the end of each cell-free DNA molecule. For example, the first subsequence can be the XYNN pattern shown in Figure 74. In some embodiments, the first subsequence may not include the final nucleotide(s) at the end of each cell-free DNA molecule. For example, the first subsequence can include the NNXY pattern of Figure 75.
[0339] The second amount of cell-free DNA molecules can be measured, which have a subsequence different from the first subsequence at one or more ends of the cell-free DNA molecules.The value of terminal motif parameter can be generated by using the ratio of the second amount and the third amount.For example, the second amount can be divided by the third amount, or the third amount can be divided by the second amount.
[0340] At block 7840, the value of the terminal motif parameter may be compared to a reference value. The threshold value may be a value that represents a statistically significant difference from the value of the associated parameter for subjects without a pregnancy-related disorder. The threshold value may be determined from one or more reference subjects with normal pregnancies or one or more reference subjects with a pregnancy-related disorder.
[0341] In some embodiments, the value of the terminal motif parameter can be compared to a threshold value, and the value of the second terminal motif parameter can be compared to a second threshold value. A second amount of cell-free DNA molecules having a second subsequence different from the first subsequence at one or more ends of the cell-free DNA molecules can be measured. Thus, the amount of the different terminal motif can be determined. A value of a second terminal motif parameter can be generated using the second amount. The value of the second terminal motif parameter can be compared to a second threshold value. The second threshold value can be the same as or different from the first threshold value. Additional subsequences can be used in the same manner as the first and second subsequences. In some embodiments, all possible subsequences can be used in the comparison to the threshold value.
[0342] At block 7850, a classification of the likelihood of a pregnancy-related disorder may be determined using the comparison. If the value of the size parameter or the value of the terminal motif parameter exceeds a threshold, a pregnancy-related disorder may occur.
[0343] In some embodiments, determining the classification of the likelihood of a pregnancy-related disorder may use a comparison of the value of the second terminal motif parameter to a second cutoff value. If the value of the first terminal motif parameter exceeds a first threshold and the value of the second terminal motif parameter exceeds a second threshold, a pregnancy-related disorder may occur.
[0344] The method may include using a size parameter in addition to the terminal motif parameter. A second set of cell-free DNA molecules having a size within a first size range may be identified. The first size range may include sizes greater than a cutoff value. The first size range may include sizes greater than the cutoff value. The first size range may be less than 550nt, 600nt, 650nt, 700nt, 750nt, 800nt, 850nt, 900nt, 950nt, 1nt, 1.5knt, 2knt, 3knt, 5knt, or greater. A value for the size parameter may be generated using a second amount of cell-free DNA molecules in the second set. The value of the size parameter may be compared with a second threshold. Determining a classification of the likelihood of a pregnancy-related disorder may involve comparing the value of the size parameter with a second threshold. If one or both of the first and second thresholds are exceeded, the classification may indicate a high likelihood of having a pregnancy-related disorder.
[0345] The size parameter can be a normalization parameter. For example, a third amount of cell-free DNA molecules in a second size range can be measured. The second size range can include sizes below the first cutoff value. The second size range can include all sizes. The second size range can include 50-150 nt, 50-166 nt, 50-200 nt, and 200-400 nt. The second size range can include any size of the short cell-free DNA fragments described herein. The second size range can exclude sizes in the first size range. The value of the size parameter can be generated by determining the ratio of the second amount to the third amount. For example, the second amount can be divided by the third amount, or the third amount can be divided by the second amount.
[0346] Either of the amounts of cell-free DNA molecules can be cell-free DNA molecules from a particular tissue of origin. For example, the tissue of origin can be T cells or another tissue of origin described herein. The second amount can be similar to the T cell contribution illustrated in Figures 77A-77D. The contribution from the tissue of origin can be determined using methylation status or patterns as described in this disclosure.
[0347] V. Repeat Expansion-Associated Disorders Long cell-free DNA fragments obtained from pregnant women can be used to identify repeat expansions in genes. Repeat expansions in genes can result in neuromuscular diseases. Tandem repeat expansions have been associated with human diseases, including, but not limited to, neurodegenerative disorders such as Fragile X syndrome, Huntington's disease, and spinocerebellar ataxia. These tandem repeat expansions can occur in protein-coding regions of genes (Machado-Joseph disease, Holliver syndrome, Huntington's disease) or non-coding regions (Friedrich ataxia, myotonic dystrophy, and some forms of Fragile X syndrome). Expansions containing minisatellites, pentanucleotides, tetranucleotides, and multiple trinucleotide repeats are associated with fragile sites. These disease-associated expansions can be caused by replication slippage, asymmetric recombination, or epigenetic abnormalities. The number of repeats in a sequence refers to the total number of times a subsequence appears. For example, "CAGCAG" contains two repeats. The number of repeats cannot be 1, since a repeat contains at least two instances of a subsequence. A subsequence may be understood to be a repeat unit.
[0348] In an embodiment, analysis of long cell-free DNA in pregnant women can facilitate the detection of repeat-associated diseases. For example, trinucleotide repeats represent repeated stretches of 3-bp motifs in DNA sequences. One example is the sequence "CAGCAGCAG" containing three 3-bp "CAG" motifs. Microsatellite expansions, typically trinucleotide repeat expansions, have been reported to play an important role in neurological disorders (Kovtun et al. Cell Res. 2008;18:198-213, McMurray et al. Nat Rev Genet. 2010;11:786-99). One example is that more than 55 CAG repeats (total 165 bp) in the ATXN3 gene are pathogenic, resulting in spinocerebellar ataxia type 3 (SCA3), a disease characterized by progressive motor problems. This condition is inherited in an autosomal dominant pattern. Therefore, one copy of the altered gene is sufficient to cause the disorder. To determine the repeat number of microsatellites, polymerase chain reaction (PCR) is typically used to amplify the genomic region of interest, followed by subjecting the PCR products to numerous different techniques, such as capillary electrophoresis (Lyon et al. J Mol Diagn. 2010;12:505-11), Southern blot analysis (Hsiao et al. J Clin Lab Anal. 1999;13:188-93), melting curve analysis (Lim et al. J Mol Diagn. 2014;17:302-14), and mass spectrometry (Zhang et al. Anal Methods. 2016;8:5039-44). However, these methods are labor-intensive and time-consuming, making them difficult to apply to high-throughput screening in actual clinical practice, such as prenatal testing. Sanger sequencing, however, presents significant challenges in inferring long repeats from complex sequence traces through manual inspection.Illumina sequencing technologies and Ion Torrent are notoriously difficult to sequence in GC-rich (or GC-poor) regions containing repeats (Ashely et al. 2016;17:507-22), and the length of DNA, including extended DNA, easily exceeds the length of the sequence read (Loomis et al. Genome Res. 2013;23:121-8).
[0349] Another example is myotonic dystrophy, an autosomal dominant disorder caused by a CTG repeat expansion ranging from 50 to 4000 CTG repeats near the DMPK gene. Molecular diagnosis of DM is routinely performed prenatally by invasively analyzing the number of CTG repeats on fetal genomic DNA.
[0350] In contrast to short-read sequencing (hundreds of bases), the method described herein can obtain long DNA molecules (several kilobases) from maternal plasma DNA.The method described herein can be used to non-invasively determine whether a fetus inherits the disease from an affected mother.
[0351] Figure 79 shows a diagram for estimating maternal inheritance of a fetus for a repeat-associated disease. In step 7905, cell-free DNA from a pregnancy was subjected to single-molecule real-time (e.g., PacBio SMRT) sequencing. In step 7910, the sequencing results were divided into long DNA and short DNA categories according to the present disclosure. In step 7915, the allele information present in the long DNA molecules can be used to construct maternal haplotypes, i.e., Hap I and Hap II. Hap I and Hap II can each contain an expanded repeat of a trinucleotide subsequence (e.g., CTG). In step 7920, haplotype imbalances can be analyzed, similar to that described in Figure 16. In step 7925, maternal inheritance of a fetus can be estimated. The methods described herein allow for the use of sequence information from long DNA molecules according to the present disclosure to determine not only haplotypes (e.g., Hap I and Hap II), but also to determine the haplotype with the disorder-causing expanded repeat (e.g., affected Hap I). According to the methods described herein, counts, sizes, or methylation states from short DNA molecules distributed across maternal Hap I and Hap II can be used to determine whether the fetus inherited maternal Hap I (affected) or Hap II (unaffected) in this example.
[0352] Figure 80 shows a diagram for estimating the paternal inheritance of a fetus for a repeat-associated disease. Cell-free DNA from pregnancy can be used to determine whether the fetus inherited the affected paternal haplotype. As shown in Figure 80, cell-free DNA from pregnancy of an unaffected woman whose husband is affected with a repeat expansion disease (e.g., 70 CTG repeats) (e.g., 5 CTG repeats for Hap I and 6 CTG repeats for Hap II) was subjected to PacBio SMRT sequencing. Long DNA molecules were identified and used to determine the haplotype and repeat number. If a haplotype with a long stretch of CTG repeats (e.g., 70 CTG repeats in this example) is present in the maternal plasma of an unaffected pregnant woman, it suggests that the fetus inherited the affected paternal haplotype. In some embodiments, the DNA containing the expanded repeat also carries one or more additional paternal-specific alleles not present in the maternal genome. This situation is useful for confirming paternal inheritance.
[0353] In another embodiment, cell-free DNA from pregnancy can be used to determine whether a fetus inherits an affected paternal haplotype. As shown in Figure 80, cell-free DNA from a pregnant woman whose husband is affected with a repeat expansion disease (e.g., 70 CTG repeats) (e.g., 5 CTG repeats for Hap I and 6 CTG repeats for Hap II) was subjected to PacBio SMRT sequencing, and long sequenced DNA molecules were identified and used to determine the haplotype and repeat number. If a haplotype with a long stretch of CTG repeats (e.g., 70 CTG repeats in this example) is present in the maternal plasma of an unaffected pregnant woman, it suggests that the fetus has inherited the affected paternal haplotype. In some embodiments, the DNA containing the expanded repeat also carries one or more additional paternal-specific alleles that are not present in the maternal genome. This situation is useful for confirming paternal inheritance.
[0354] Figures 81, 82, and 83 are tables showing examples of repeat expansion diseases. Column 1 shows the repeat expansion associated disease. Column 2 shows the repeat subsequence. Column 3 shows the repeat number in normal subjects. Column 4 shows the repeat number in affected subjects. Column 5 shows the genetic location associated with the repeat. Column 6 lists the gene name. Column 7 lists the pattern of inheritance. Tables are taken from omicslab.genetics.ac.cn / dred / index.php.
[0355] A. Example of repeat expansion detection It has been reported that paternally inherited expanded CAG repeats can be detected in maternal plasma using a direct PCR approach followed by fragment analysis on a 3130XL Genetic Analyzer (Oever et al. Prenat Diagn. 2015;35:945-9). Because the size of the expanded alleles begins only with repeats greater than 35 trinucleotides [i.e., DNA regions spanning the repeats 105 bp (35 × 3) or greater in length], noninvasive prenatal testing for Huntington's disease was achievable by PCR. Many expanded repeats, particularly most trinucleotide repeat disorders (Orr et al. Annu. Rev. Neurosci. 2007;30:575-621), do not involve repeats greater than 300 bp in length, exceeding the size of a short fetal DNA molecule, as previously documented. DNA with large extended repeats makes PCR difficult (Orr et al. Annu. Rev. Neurosci. 2007;30:575-621). As suggested by the study of Oever et al., the signal intensity of long CAG repeats is often much lower than that of smaller repeats, and this phenomenon is observed in both genomic DNA and plasma DNA, making the sensitivity for detecting these long CAG repeats lower (Oever et al. Prenat Diagn. 2015;35:945-9). Another limitation of PCR is the inability to preserve methylation signals during amplification. In one embodiment, single-molecule real-time sequencing of long DNA molecules allows for the determination of tandem repeat polymorphisms across one or more regions and their associated methylation levels.
[0356] Figure 84 is a table showing examples of repeat expansion detection and repeat-associated methylation determination in a fetus. Column 1 indicates the repeat type by number of base pairs. Column 2 indicates the repeat unit. Column 3 indicates the genomic location. Column 4 indicates the reference base, the sequence present in the human reference genome. Column 5 indicates the paternal genotype. Column 6 indicates the maternal genotype. Column 7 indicates the fetal genotype. Column 8 indicates the fetal DNA methylation level associated with the paternal allele. Column 9 indicates the fetal DNA methylation level associated with the maternal allele.
[0357] Figure 84 shows numerous examples of 1-bp, 2-bp, 3-bp, and 4-bp tandem repeats. For example, at the genomic location chr3:192384705-192384706, a "GATA" tandem repeat was identified. The paternal genotype at this locus was T(GATA)3 / T(GATA)5, with allele 1 having three repeat units and allele 2 having five repeat units. Compared to the reference allele T(GATA)3, the paternal allele 2 suggested a genetic event involving a repeat expansion. The maternal genotype at this locus was T / T, indicating a genetic event involving a repeat contraction. The fetal genotype at this locus was T(GATA)5 / T, suggesting that the fetus inherited the paternal allele 2 (i.e., T(GATA)5) and the maternal allele T. The methylation levels associated with the paternal and maternal alleles were 50.98 and 62.8, respectively. These results suggested that the use of tandem repeat polymorphisms allows for the determination of maternal and paternal inheritance of a fetus. This technique allows for the identification of distinct methylation patterns associated with the two alleles. Another example shows that a fetus inherited a repeat expansion [(TAAA)3] from its mother at genomic location chr4:73237157-73237158. Fetal molecules containing the maternally inherited repeat expansion showed a higher methylation level (95.65%) compared to fetal molecules containing the paternal allele (62.84%). These data suggested that repeats, repeat structures, and associated methylation changes can be detected. In one embodiment, a specific cutoff can be used to determine whether the methylation difference between maternal and paternal inheritance was significant. The cutoff is an absolute difference in methylation levels such as, but not limited to, greater than 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or 90%. Determining maternal inheritance can be similar to the method described in method 2100 of Figure 21.
[0358] B. Exemplary Methods Partial sequence repeats can be used to determine fetal information. For example, the presence of partial sequence repeats can be used to determine that a molecule is of fetal origin. Furthermore, partial sequence repeats can indicate the likelihood of genetic disorders. Partial sequence repeats can be used to determine the inheritance of maternal and / or paternal haplotypes. Furthermore, the paternity of a fetus can be determined using partial sequence repeats.
[0359] 1. Fetal Origin Analysis Using Partial Sequence Repeats Figure 85 shows a method 8500 of analyzing a biological sample obtained from a woman carrying a fetus, the biological sample including cell-free DNA molecules from the fetus and the woman. The likelihood of a genetic disorder in the fetus can be determined.
[0360] In block 8510, a first sequence read corresponding to one cell-free DNA molecule of the cell-free DNA molecules can be received. The cell-free DNA molecule can have a length greater than a cutoff value. The cutoff value can be 200 nt or more. The cutoff value can be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cutoff value can be any cutoff value described herein for long cell-free DNA molecules.
[0361] In step 8520, the first read may be aligned to a region of a reference genome. The region may be known to potentially contain repeats of a subsequence. The region may correspond to any of the positions or genes in Figures 81-83. The subsequence may be a trinucleotide sequence, including any of those described herein.
[0362] At block 8530, the number of repeats of the subsequence in the first sequence read corresponding to the cell-free DNA molecule can be determined.
[0363] In block 8540, the number of repeats of the subsequence may be compared to a threshold number. The threshold number may be 55, 60, 75, 100, 150, or more. The threshold number may be different for different genetic disorders. For example, the threshold may reflect the minimum number of repeats in affected subjects, the maximum number of repeats in normal subjects, or a number between these two numbers (see Figures 81-83).
[0364] In block 8550, a classification of the likelihood that the fetus has a genetic disorder may be determined using a comparison of the number of iterations to a threshold number. If the number of iterations exceeds the threshold, it may be determined that the fetus has a high likelihood of having a genetic disorder. The genetic disorder may be Fragile X Syndrome or any of the disorders listed in Figures 81-83.
[0365] In some embodiments, the method may include repeating the classification for several different target loci, each of which is known to potentially have subsequence repeats. A plurality of sequence reads corresponding to a cell-free DNA molecule may be received. The plurality of sequence reads may be aligned to a plurality of regions of a reference genome. The plurality of regions may be known to potentially contain subsequence repeats. The plurality of regions may be non-overlapping regions. Each region of the plurality of regions may have a different SNP. The plurality of regions may be derived from different chromosome arms or chromosomes. The plurality of regions may cover at least 0.01%, 0.1%, or 1% of the reference genome. The number of subsequence repeats may be identified in the plurality of sequence reads. The number of subsequence repeats may be compared to multiple threshold numbers. Each threshold number may indicate the presence or likelihood of a different genetic disorder. For each of the multiple genetic disorders, a classification of the likelihood that the fetus has the respective genetic disorder may be determined using a comparison of the plurality of threshold numbers to one of the multiple threshold numbers.
[0366] The cell-free DNA molecule can be determined to be of fetal origin.Determining the fetal origin can include receiving a second sequence read corresponding to the cell-free DNA molecule of maternal origin obtained from buffy coat or a sample of a woman before pregnancy.The second sequence read can be aligned to a region of a reference genome.A second repeat number of the partial sequence can be identified in the second sequence read.The second repeat number can be determined to be less than the first repeat number.
[0367] Determining fetal origin can include determining the methylation level of the cell-free DNA molecule using methylated and unmethylated sites of the cell-free DNA molecule. The methylation level can be compared with a reference level. The method can include determining that the methylation level exceeds the reference level. The methylation level can be the number or percentage of methylated sites.
[0368] Determining fetal origin can include determining the methylation pattern of multiple sites of the cell-free molecule. A similarity score can be determined by comparing the methylation pattern with a reference pattern from maternal or fetal tissue. The similarity score can be compared to one or more thresholds. The similarity score can be any similarity score described herein, including, for example, the one described in method 4000.
[0369] 2. Paternity analysis using partial sequence repeats 86 shows a method 8600 of analyzing a biological sample obtained from a woman carrying a fetus, the biological sample including cell-free DNA molecules from the fetus and the woman. The biological sample can be analyzed to determine the paternity of the fetus.
[0370] In block 8610, a first sequence read corresponding to one cell-free DNA molecule of the cell-free DNA molecules may be received. The method may include determining that the cell-free DNA molecule is of fetal origin. The cell-free DNA molecule may be determined to be of fetal origin by any method described herein, including, for example, that described in method 8500. The cell-free DNA molecule may have a size greater than a cutoff value. The cutoff value may be 200 nt or greater. The cutoff value may be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cutoff value may be any cutoff value described herein for long cell-free DNA molecules.
[0371] At block 8620, the first read may be aligned to a first region of a reference genome. The first region may be known to have subsequence repeats.
[0372] At block 8630, a first number of repeats of a first subsequence in a first sequence read corresponding to the cell-free DNA molecule can be identified. The first subsequence can include an allele.
[0373] At block 8640, the sequence data obtained from the male subject may be analyzed to determine whether a second number of repeats of the first subsequence is present within the first region, the second number of repeats including at least two instances of the first subsequence. The sequence data may be obtained by extracting a biological sample from the male subject and performing sequencing on DNA in the biological sample.
[0374] At block 8650, a classification of the likelihood that the male subject is the father of the fetus may be determined using a determination of whether the second number of repeats of the first subsequence is present. The classification may be that the male subject is likely to be the father if it is determined that the second number of repeats of the first subsequence is present. The classification may be that the male subject is likely not the father if it is determined that the second number of repeats of the first subsequence is not present.
[0375] The method may include comparing the first number of replicates to a second number of replicates. Determining a classification of the likelihood that the male subject is the father may include using a comparison of the first number of replicates to the second number of replicates. The classification may be that the male subject is likely to be the father if the first number of replicates is within a threshold of the second number of replicates. The threshold may be within 10%, 20%, 30%, or 40% of the second number of replicates.
[0376] The method may include using multiple regions of repeats. For example, the cell-free DNA molecule is a first cell-free DNA molecule. The method may include receiving a second sequence read corresponding to a second cell-free DNA molecule of the cell-free DNA molecule. The method may also include aligning the second sequence read to a second region of a reference genome. The method may further include identifying a first number of repeat...
Claims
1. 1. A method of analyzing a biological sample obtained from a woman carrying a fetus, the biological sample comprising a plurality of cell-free DNA molecules from the fetus and the woman, the method comprising: receiving sequence reads corresponding to the plurality of cell-free DNA molecules; measuring the size of the plurality of cell-free DNA molecules; identifying a set of cell-free DNA molecules from the plurality of cell-free DNA molecules as having a size greater than or equal to a cutoff value, wherein the cutoff value is at least 500 n; For one cell-free DNA molecule of said set of cell-free DNA molecules: determining the methylation status of each site of the plurality of sites; determining a methylation pattern, determining the methylation pattern, wherein the methylation pattern indicates the methylation state of each site of the plurality of sites using one or more sequence reads corresponding to the cell-free DNA molecule; comparing the methylation pattern to one or more reference patterns, each of the one or more reference patterns determined for a particular tissue type; and using the methylation pattern to determine the tissue of origin of the cell-free DNA molecule.
2. The method of claim 1, wherein the cutoff value is 600 nt.
3. The method of claim 1, wherein the cutoff value is 1 knt.
4. the tissue of origin for each cell-free DNA molecule of the set of cell-free DNA molecules; determining the methylation state of each site of a plurality of respective sites, wherein each site of the plurality corresponds to the cell-free DNA molecule; determining the methylation pattern; and 10. The method of claim 1, further comprising determining the methylation pattern by comparing the methylation pattern to at least one reference pattern of the one or more reference patterns.
5. determining the amount of cell-free DNA molecules corresponding to each tissue of origin; 5. The method of claim 4, further comprising using the amount of cell-free DNA molecules corresponding to each tissue of origin to determine the fractional contribution of the tissue of origin in the biological sample.
6. measuring the size of the plurality of cell-free DNA molecules; 2. The method of claim 1, comprising aligning the sequence reads to a reference genome.
7. measuring the size of the plurality of cell-free DNA molecules; determining the full length sequence of the plurality of cell-free DNA molecules; and and counting the number of nucleotides in each cell-free DNA molecule of the plurality of cell-free DNA molecules.
8. measuring the size of the plurality of cell-free DNA molecules; 10. The method of claim 1, comprising physically separating the plurality of cell-free DNA molecules from the biological sample from other cell-free DNA molecules in the biological sample, wherein the other cell-free DNA molecules have a size smaller than the cutoff value.
9. One reference pattern of the one or more reference patterns is measuring the methylation density of each reference site of the plurality of reference sites using DNA molecules from the reference tissue; comparing the methylation density of each reference site of the plurality of reference sites to one or more threshold methylation densities; and 2. The method of claim 1, wherein the methylation density is determined by identifying each reference site of the plurality of reference sites as methylated, unmethylated, or uninformative based on comparing the methylation density to the one or more threshold methylation densities, wherein the plurality of sites are the plurality of reference sites identified as methylated or unmethylated.
10. The method of claim 1 , wherein the tissue of origin is placenta.
11. The method of claim 1 , wherein the tissue of origin is fetal or maternal.
12. the tissue of origin is of fetal origin; The method comprises: aligning one of the sequence reads to a first region of a reference genome, the first region including a plurality of sites corresponding to alleles, the plurality of sites including a threshold number of sites; determining a first haplotype using each allele present at each site of the plurality of sites; comparing the first haplotype to a second haplotype corresponding to a male subject; 12. The method of claim 11, further comprising: using the comparison to determine a classification of the likelihood that the male subject is the father of the fetus.
13. the tissue of origin is of fetal origin; The method comprises: aligning one of the sequence reads to a first region of a reference genome, the first region comprising a first plurality of sites corresponding to alleles, the plurality of sites comprising a threshold number of sites; comparing an allele at each site of the plurality of sites to an allele at the corresponding site in the genome of a male subject; 12. The method of claim 11, further comprising: using the comparison to determine a classification of the likelihood that the male subject is the father of the fetus.
14. For each cell-free DNA molecule of said set of cell-free DNA molecules, aligning the sequence reads corresponding to the cell-free DNA molecules to a reference genome; identifying the sequence reads as corresponding to a haplotype present in the female; determining the tissue of origin as fetal using the methylation pattern; 12. The method of claim 11, further comprising determining the haplotype to be a maternally inherited fetal haplotype.
15. Identifying the haplotype as carrying a disease-causing genetic mutation or alteration; and assisting in classifying the fetus as likely to have the disease caused by the genetic mutation or alteration.
16. identifying the haplotype as carrying the disease-causing genetic mutation; identifying the genetic mutation or alteration in the first sequence read; measuring a first methylation level in a second sequence read corresponding to a first genomic location within a first distance of the first sequence read; measuring a second methylation level in a third sequence read corresponding to a second genomic location within a second distance of the first sequence read; 16. The method of claim 15, wherein the first methylation level and the second methylation level are associated with the genetic variation.
17. For each cell-free DNA molecule of said set of cell-free DNA molecules, aligning the sequence reads corresponding to the cell-free DNA molecules to a reference genome; identifying the sequence reads as corresponding to a region, the region comprising: receiving a plurality of fetal sequence reads corresponding to a plurality of fetal DNA molecules from fetal tissue; receiving a plurality of maternal sequence reads corresponding to a plurality of maternal DNA molecules; determining, for each fetal sequence read of the plurality of fetal sequence reads, a fetal methylation state of each methylation site of a plurality of methylation sites within the region; determining, for each maternal sequence read of the plurality of maternal sequence reads, a maternal methylation state of each methylation site of the plurality of methylation sites; determining the value of a parameter characterizing the amount of sites where the fetal methylation state differs from the maternal methylation state; comparing the value of the parameter to a threshold; and The method of claim 11 , further comprising: determining that the value of the parameter exceeds the threshold.
18. Determining the tissue of origin of the cell-free DNA molecules comprises inputting the methylation pattern into a machine learning model, the model comprising: receiving a plurality of training methylation patterns, each training methylation pattern having a methylation state at one or more sites of said plurality of sites, each training methylation pattern determined from DNA molecules from a known tissue; storing a plurality of training samples, each training sample including one of the plurality of training methylation patterns and a label indicating the known tissue corresponding to the training methylation pattern; 2. The method of claim 1, wherein the model is trained by using the plurality of training samples to optimize parameters of the model based on the output of the model that matches or does not match the label that corresponds to the plurality of training methylation patterns when the training methylation patterns are input to the model, wherein the output of the model specifies the tissue that corresponds to the input methylation patterns.
19. 20. The method of claim 18, wherein the machine learning model comprises a convolutional neural network (CNN), linear regression, logistic regression, a deep recurrent neural network, a Bayesian classifier, a hidden Markov model (HMM), a linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering for applications with noise (DBSCAN), a random forest algorithm, or a support vector machine (SVM).
20. 19. The method of claim 18, wherein each DNA molecule from the known tissue is cellular DNA.
21. 20. The method of claim 18, wherein the parameters of the model include a first parameter that indicates whether one site of the plurality of sites has the same methylation state as another site of the plurality of sites.
22. The method of claim 18 , wherein the parameters of the model include a second parameter indicative of a distance between sites of the plurality of sites.
23. one reference pattern of the one or more reference patterns corresponds to a reference tissue; 10. The method of claim 1, wherein the method further comprises determining that the tissue of origin is the reference tissue if the methylation pattern matches the reference pattern.
24. The method of claim 1 , wherein the plurality of sites comprises at least five CpG sites.
25. determining the tissue of origin using the methylation pattern, determining a similarity score by comparing the methylation pattern to a first reference methylation pattern from a first reference tissue of a plurality of reference tissues; comparing the similarity score to a threshold; and determining the tissue of origin to be the first reference tissue if the similarity score exceeds the threshold.
26. the similarity score is a first similarity score; The method comprises: The threshold value is 26. The method of claim 25, further comprising calculating by determining a second similarity score by comparing the methylation pattern to a second reference methylation pattern from a second reference tissue of the plurality of reference tissues, wherein the first reference tissue and the second reference tissue are different tissues, and the threshold is the second similarity score.
27. the first reference methylation pattern comprises a first subset of sites having at least a first probability of being methylated for the first reference tissue; the first reference methylation pattern comprises a second subset of sites having at most a second probability of being methylated for the first reference tissue, wherein the first probability is greater than the second probability; Determining the similarity score comprises: increasing the similarity score if a site of the plurality of sites is methylated and the site of the plurality of sites is within the first subset of sites; and decreasing the similarity score if a site of the plurality of sites is methylated and the site of the plurality of sites is within the second subset of sites.
28. the first reference methylation pattern comprises the plurality of sites, each site of the plurality of sites characterized by a probability of being methylated and a probability of being unmethylated for the first reference tissue; The similarity score is For each of the plurality of sites, determining the probability in the reference tissue corresponding to the methylation state of the site in the cell-free DNA molecule; 26. The method of claim 25, wherein the similarity score is determined by calculating a product of the plurality of probabilities, the product being the similarity score.
29. 29. The method of claim 28, wherein the probability is determined using a beta distribution.
30. sequencing the plurality of cell-free DNA molecules to obtain the sequence reads; 2. The method of claim 1, further comprising determining the methylation status of the site by measuring properties corresponding to the nucleotide at the site and the nucleotides adjacent to the site.
31. 2. The method of claim 1, wherein each cell-free DNA molecule of the set of cell-free DNA molecules comprises at least a certain number of CpG sites.
32. 10. The method of claim 1, wherein at least one site of the plurality of sites is methylated.
33. 2. The method of claim 1, wherein two sites of said plurality of sites are separated by at least 160 nt.
34. A computer program comprising instructions for controlling a computer device to carry out the method of any one of claims 1 to 33.
35. 35. A computer readable medium comprising the computer program of claim 34.
36. A computer device comprising the computer program of claim 34.
Citation Information
Patent Citations
Non-invasive determination of methylome of tumor from plasma
US20150011403A1
Methylation pattern analysis of tissues in a DNA mixture
US20160017419A1
Methylation pattern analysis of haplotypes in tissues in a DNA mixture
US20170029900A1
Using cell-free DNA fragment size to determine copy number variations
US20170220735A1
Gestational age assessment by methylation and size profiling of maternal plasma DNA
US20180105807A1