Molecular analysis using long free fragments from pregnant women
By using PacBio sequencing technology to analyze cell-free DNA fragments longer than 200 bp, the problem of accurately obtaining fetal genetic information in existing technologies has been solved, enabling efficient and accurate fetal genetic analysis and gestational age determination.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-05
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies struggle to effectively analyze cell-free DNA fragments longer than 200 bp, especially on the Illumina sequencing platform, making it difficult to accurately obtain fetal genetic information.
PacBio sequencing technology was used to analyze cell-free DNA fragments longer than 200 bp. By identifying multiple CpG sites and SNPs, combined with methylation patterns and repetitive sequences, fetal genetic information was determined.
It enables efficient and accurate analysis of fetal genetic information, identifies fetal genetic diseases and gestational age, and provides analysis of long cell-free DNA fragments of maternal origin.
Smart Images

Figure CN115066504B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Application No. 62 / 970,634, filed February 5, 2020, and U.S. Provisional Application No. 63 / 135,486, filed January 8, 2021, the entire contents of which are incorporated herein by reference for all purposes. Background Technology
[0003] The modal size of circulating cell-free DNA in pregnant women has been reported to be approximately 166 bp (Lo et al., *Sci Transl Med.* 2010; 2:61ra91). Very little data exists regarding fragments larger than 600 bp. One example is the work of Amicucci et al., who reported an 8 kb amplification of the Y chromosome basic protein Y2 gene (BPY2) from maternal plasma using PCR (Amicucci et al., *Clin Chem.* 2000; 40:301-2). It is unknown whether this data is applicable across the entire genome. In fact, detecting such long DNA fragments, such as those larger than 600 bp, using massively parallel short-read sequencing technologies, such as the Illumina platform, presents many challenges (Lo et al., *Sci Transl Med.* 2010; 2:61ra91; Fan et al., *Clinical Chem.* 2010; 56:1278-86). These challenges include: (1) the recommended size range for Illumina sequencing platforms typically spans 100–300 bp (De Maio et al., Microbial Genomics, 2019; 5(9)); and (2) DNA amplification should involve sequencing library preparation on flow channels (via PCR) or sequencing cluster generation (via bridge amplification). Such amplification processes can facilitate the amplification of shorter DNA fragments, partly due to the fact that long DNA templates (e.g., >600 bp) require a relatively longer time to complete substrand synthesis compared to short DNA templates (e.g., <200 bp). Therefore, within a fixed timeframe of these PCR processes before or during sequencing on the Illumina platform, those long DNA molecules that fail to be fully generated during the PCR process will be unusable in downstream analysis; (3) long DNA molecules will have a greater probability of forming secondary structures that hinder amplification; (4) using Illumina sequencing technology, long DNA molecules will be more likely to generate clusters containing more than one clone DNA molecule than short DNA molecules because the library is denatured, diluted and diffused on a two-dimensional surface, followed by bridging amplification (Head et al. Biotechniques. 2014; 56:61-4). Summary of the Invention
[0004] The methods and systems described herein relate to the analysis of biological samples using long cell-free DNA fragments. Using these long cell-free DNA fragments allows for analyses that were not previously considered or would be impossible with shorter cell-free DNA fragments. The state of methylated CpG sites and single nucleotide polymorphisms (SNPs) is often used to analyze DNA fragments in biological samples. CpG sites and SNPs are typically separated from the nearest CpG site or SNP by hundreds or thousands of base pairs. Most cell-free DNA fragments in biological samples are typically less than 200 bp in length. Therefore, finding two or more consecutive CpG sites or SNPs on most cell-free DNA fragments is unlikely or impossible. Cell-free DNA fragments longer than 200 bp, including those longer than 600 bp or 1 kb, may contain multiple CpG sites and / or SNPs. The presence of multiple CpG sites and / or SNPs on long cell-free DNA fragments allows for more efficient and / or more accurate analyses compared to short cell-free DNA fragments alone. Long cell-free DNA fragments can be used to identify tissues of origin and / or to provide information about the fetus in a pregnant woman's body. Furthermore, the accurate analysis of samples from pregnant women using long cell-free DNA fragments was unexpected, as we expected these fragments to be primarily of maternal origin. We did not expect long cell-free DNA fragments of fetal origin to exist in an amount sufficient to provide information about the fetus.
[0005] Long cell-free DNA fragments containing SNPs can be used to determine fetal genetic haplotypes. These fragments, by having multiple CpG sites, can exhibit methylation patterns indicative of the tissue of origin. Additionally, trinucleotide repeats and other repeat sequences can be present on these long cell-free DNA fragments. These repeat sequences can be used to determine the likelihood of a fetal genetic condition or fetal kinship. The amount of the long cell-free DNA fragment can be used to determine gestational age. Similarly, motifs at the ends of long cell-free DNA fragments can also be used to determine gestational age. Long cell-free DNA fragments (including, for example, the amount, length distribution, genomic location, methylation status, etc.) can be used to identify pregnancy-related conditions.
[0006] These and other embodiments of this disclosure are described in detail below. For example, other embodiments relate to systems, apparatus, and computer-readable media in relation to the methods described herein.
[0007] A better understanding of the nature and advantages of the embodiments of this disclosure can be obtained by referring to the following detailed description and accompanying drawings. Attached Figure Description
[0008] Figure 1A and Figure 1BThe size distribution of free DNA as determined according to embodiments of the present invention is shown. (A) 0-20 kb on a linear scale, (B) 0-20 kb on a logarithmic scale.
[0009] Figure 2A and Figure 2B The diagram shows the determined size distribution of free DNA according to an embodiment of the present invention. (A) 0-5 kb on a linear scale along the y-axis. (B) 0-5 kb on a logarithmic scale along the y-axis.
[0010] Figure 3A and Figure 3B The diagram shows the determined size distribution of free DNA according to an embodiment of the present invention. (A) 0-400 bp on a linear scale along the y-axis. (B) 0-400 bp on a logarithmic scale along the y-axis.
[0011] Figure 4A and Figure 4B This diagram shows the size distribution of cell-free DNA, as determined according to embodiments of the present invention, between segments carrying shared alleles (shared) and segments carrying fetal-specific alleles (fetal-specific). (A) 0-20 kb bp on a linear scale along the y-axis. (B) 0-20 kb on a logarithmic scale along the y-axis. The blue line indicates segments carrying shared alleles (primarily maternally originating), and the red line indicates segments carrying fetal-specific alleles (placental originating).
[0012] Figure 5A and Figure 5B This diagram shows the size distribution of cell-free DNA, as determined according to embodiments of the present invention, between segments carrying shared alleles (shared) and segments carrying fetal-specific alleles (fetal-specific). (A) 0-5 kb bp on a linear scale along the y-axis. (B) 0-5 kb on a logarithmic scale along the y-axis. The blue line indicates segments carrying shared alleles (primarily maternally originating), and the red line indicates segments carrying fetal-specific alleles (placental originating).
[0013] Figure 6A and Figure 6B This diagram shows the size distribution of cell-free DNA, as determined according to embodiments of the present invention, between segments carrying shared alleles (shared) and segments carrying fetal-specific alleles (fetal-specific). (A) 0-1 kb on a linear scale along the y-axis. (B) 0-1 kb on a logarithmic scale along the y-axis. The blue line indicates segments carrying shared alleles (primarily maternally originating), and the red line indicates segments carrying fetal-specific alleles (placental originating).
[0014] Figure 7A and Figure 7BThis diagram shows the size distribution of cell-free DNA, as determined according to embodiments of the present invention, between segments carrying shared alleles (shared) and segments carrying fetal-specific alleles (fetal-specific). (A) 0-400 bp on a linear scale along the y-axis. (B) 0-400 bp on a logarithmic scale along the y-axis. The blue line indicates segments carrying shared alleles (primarily maternally derived), and the red line indicates segments carrying fetal-specific alleles (placental-derived).
[0015] Figure 8 This illustrates the degree of single-molecule double-stranded DNA methylation between a segment carrying a maternally specific allele and a segment carrying a fetal specific allele, according to embodiments of the present invention.
[0016] Figure 9A and Figure 9B This invention presents (A) a fitted distribution of single-molecule double-stranded DNA methylation levels between fragments carrying maternally specific alleles and fragments carrying fetal specific alleles, and (B) a receiver operating characteristic (ROC) analysis using single-molecule double-stranded DNA methylation levels.
[0017] Figure 10A and Figure 10B This illustrates the correlation between the degree of single-molecule double-stranded DNA methylation and the fragment size of plasma DNA according to embodiments of the present invention. (A) Size range of 0-20 kb. (B) Size range of 0-1 kb.
[0018] Figure 11A and Figure 11B Examples of identified long fetal-specific DNA molecules in maternal plasma DNA from pregnant women according to embodiments of the present invention are shown. (A) Black bars indicate long fetal-specific DNA molecules aligned to regions in chromosome 10 of the human reference genome. (B) Detailed illustration of the genetic and epigenetic information determined using PacBio sequencing in this disclosure. Bases highlighted in yellow (marked by arrows) may be attributable to sequence errors, which may be corrected in some embodiments.
[0019] Figure 12A and Figure 12B Examples of identified long maternal DNA molecules carrying shared alleles in maternal plasma DNA from pregnant women according to embodiments of the present invention are shown. (A) Black bars indicate long maternally specific DNA molecules aligned to a region in chromosome 6 of a human reference. (B) Detailed illustration of genetic and epigenetic information determined using PacBio sequencing according to embodiments of the present invention.
[0020] Figure 13This diagram shows the frequency distribution of DNA from the placenta (red) and DNA from maternal blood cells (blue) at different resolutions from 1kb to 20kb according to the degree of methylation, based on embodiments of the present invention.
[0021] Figure 14A and Figure 14B This diagram shows the frequency distribution of DNA from the placenta (red) and DNA from maternal blood cells (blue) within 16-kb and 24-kb windows according to the degree of methylation, based on embodiments of the present invention.
[0022] Figure 15A and Figure 15B Examples of identified long maternal-specific DNA molecules in maternal plasma DNA of pregnant women according to embodiments of the present invention are shown. (A) Black bars indicate long maternal-specific DNA molecules aligned to regions in chromosome 8 of a human reference. (B) Detailed illustration of genetic and epigenetic information determined using PacBio sequencing according to embodiments of the present invention.
[0023] Figure 16 This diagram illustrates the maternal genetic information of the fetus according to an embodiment of the present invention.
[0024] Figure 17 The illustration depicts the identification of genetic / epigenetic disorders in plasma DNA molecules using maternal and fetal origin information according to an embodiment of the present invention.
[0025] Figure 18 The illustration depicts the identification of abnormal fetal segments according to an embodiment of the present invention.
[0026] Figure 19A-19G This diagram illustrates error correction for cell-free DNA genotyping using PacBio sequencing according to an embodiment of the present invention. '.' indicates a base identical to the reference base in the Watson strand. ',' indicates a base identical to the reference base in the Crick strand. 'Letter' indicates a substitute allele that differs from the reference allele. '*' indicates an insertion. '^' indicates a deletion.
[0027] Figure 20 This invention illustrates a method for analyzing biological samples obtained from pregnant women according to an embodiment of the present invention.
[0028] Figure 21 This invention illustrates a method for analyzing biological samples obtained from pregnant women to determine haplotype inheritance according to an embodiment of the present invention.
[0029] Figure 22 This illustrates the methylation pattern of the tissue of origin for determining the origin of long DNA molecules in plasma according to an embodiment of the present invention.
[0030] Figure 23 The present invention displays receiver operating characteristic (ROC) curves for determining fetal and maternal origins according to embodiments of the present invention.
[0031] Figure 24 This illustrates the paired methylation pattern according to an embodiment of the present invention.
[0032] Figure 25 This is a distribution table of selected marker regions in different chromosomes according to embodiments of the present invention.
[0033] Figure 26 A classification table of plasma DNA molecules based on the single-molecule methylation pattern of plasma DNA molecules according to embodiments of the present invention is constructed using different percentages of ESR (erythrocyte sedimentation rate) brown-yellow layer DNA molecules with mismatch fractions greater than 0.3 as the selection criteria for marker regions.
[0034] Figure 27 This illustrates a method flow for determining fetal genetics in a non-invasive manner using placental-specific methylation haplotypes, according to an embodiment of the present invention.
[0035] Figure 28 The illustration depicts the principle of a non-invasive prenatal test for fragile X syndrome using long cell-free DNA from maternal plasma according to an embodiment of the present invention.
[0036] Figure 29 The illustration depicts maternal genetics of a fetus based on methylation patterns according to an embodiment of the present invention.
[0037] Figure 30 This illustration depicts a qualitative analysis of maternal genetics in a fetus using genetic and epigenetic information from plasma DNA molecules, according to an embodiment of the present invention.
[0038] Figure 31 The illustration shows the detection rate of qualitative analysis of maternal genetics in the fetus performed in a genome-wide manner using genetic and epigenetic information of plasma DNA molecules, compared to relative haplotype dose (RHDO) analysis, according to embodiments of the present invention.
[0039] Figure 32 This illustrates the relationship between the detection rate of paternal-specific variants in a genome-wide manner according to embodiments of the present invention and the number of sequenced plasma DNA molecules of different sizes used for analysis.
[0040] Figure 33 This illustrates a workflow for non-invasive detection of Fragile X syndrome according to an embodiment of the present invention.
[0041] Figure 34This illustrates the methylation pattern of plasma DNA according to embodiments of the present invention, relative to the methylation profile of placental and erythrocyte sedimentation rate (ESR) brown-yellow layer DNA.
[0042] Figure 35 This is a table showing the distribution of CpG sites in a 500-bp region throughout the human genome, according to an embodiment of the present invention.
[0043] Figure 36 This is a table showing the distribution of CpG sites in the 1-kb region throughout the human genome, according to an embodiment of the present invention.
[0044] Figure 37 This is a table showing the distribution of CpG sites in the 3-kb region throughout the human genome, according to an embodiment of the present invention.
[0045] Figure 38 This table illustrates the contribution ratio of different tissues in maternal plasma to DNA molecules using methylation state matching analysis, according to embodiments of the present invention.
[0046] Figure 39A and Figure 39B This illustrates the relationship between placental contribution and fetal DNA fraction inferred via the SNP pathway according to embodiments of the present invention.
[0047] Figure 40 This invention illustrates a method for analyzing biological samples obtained from pregnant women using methylation pattern analysis in order to determine the tissue of origin, according to an embodiment of the present invention.
[0048] Figure 41A and Figure 41B This illustrates the size distribution of cell-free DNA molecules from maternal plasma samples taken during early, mid, and late pregnancy, according to embodiments of the present invention.
[0049] Figure 42 A table showing the proportions of long plasma DNA molecules at different stages of pregnancy, according to embodiments of the present invention.
[0050] Figure 43A and Figure 43B This illustrates the size distribution of DNA molecules encompassing fetal-specific alleles from maternal plasma in the early, middle, and late stages of pregnancy, according to embodiments of the present invention.
[0051] Figure 44A and Figure 44B This illustrates the size distribution of DNA molecules encompassing maternal-specific alleles from maternal plasma in early, mid, and late pregnancy, according to embodiments of the present invention.
[0052] Figure 45 This is a table showing the proportions of fetal and maternal plasma DNA molecules at different gestational stages according to embodiments of the present invention.
[0053] Figure 46A , Figure 46B and Figure 46C This diagram shows the proportions of fetal-specific plasma DNA fragments across a specific size range at different gestational stages, according to embodiments of the present invention.
[0054] Figure 47A , Figure 47B and Figure 47C This diagram shows the base content ratio at the 5' end of free DNA molecules from maternal plasma in early, mid, and late pregnancy, spanning a fragment size range of 0kb to 3kb, according to embodiments of the present invention.
[0055] Figure 48 This is a table showing the ratio of terminal nucleotide bases in short and long cell-free DNA molecules from maternal plasma in the early, middle, and late stages of pregnancy, according to embodiments of the present invention.
[0056] Figure 49 This is a table showing the ratio of terminal nucleotide bases in short and long cell-free DNA molecules containing fetal-specific alleles from maternal plasma in the early, middle, and late stages of pregnancy, according to embodiments of the present invention.
[0057] Figure 50 This is a table showing the ratio of terminal nucleotide bases in short and long cell-free DNA molecules containing maternally specific alleles from maternal plasma in early, mid, and late pregnancy, according to embodiments of the present invention.
[0058] Figure 51 The illustration depicts hierarchical cluster analysis of short and long plasma-free DNA molecules using 256 terminal motifs according to an embodiment of the present invention.
[0059] Figure 52A and Figure 52B Principal component analysis showing the terminal motif profile of the tetramer according to an embodiment of the present invention.
[0060] Figure 53 This is a table of the 25 terminal motifs with the highest frequency in short plasma DNA molecules derived from maternal plasma in early pregnancy, according to an embodiment of the present invention.
[0061] Figure 54 This is a table of the 25 terminal motifs with the highest frequency in short plasma DNA molecules derived from maternal plasma during mid-pregnancy, according to an embodiment of the present invention.
[0062] Figure 55 This is a table of the 25 terminal motifs with the highest frequency in short plasma DNA molecules from maternal plasma in late pregnancy, according to an embodiment of the present invention.
[0063] Figure 56This is a table of the 25 terminal motifs with the highest frequency in long plasma DNA molecules derived from maternal plasma in early pregnancy, according to an embodiment of the present invention.
[0064] Figure 57 This is a table of the 25 terminal motifs with the highest frequency in long plasma DNA molecules derived from maternal plasma during mid-pregnancy, according to an embodiment of the present invention.
[0065] Figure 58 This is a table of the 25 terminal motifs with the highest frequency in long plasma DNA molecules derived from maternal plasma in late pregnancy, according to an embodiment of the present invention.
[0066] Figure 59A , Figure 59B and Figure 59C This invention presents a motif frequency scatter plot of 16 NNXY motifs in short and long plasma DNA molecules in maternal plasma during (A) early pregnancy, (B) mid pregnancy, and (C) late pregnancy, according to an embodiment of the present invention.
[0067] Figure 60 This invention illustrates a method for analyzing biological samples obtained from pregnant women to determine gestational age, according to an embodiment of the present invention.
[0068] Figure 61 This invention illustrates a method for analyzing biological samples obtained from pregnant women to classify pregnancy-related conditions, according to embodiments of the present invention.
[0069] Figure 62 This is a table showing clinical information for four cases of preeclampsia according to an embodiment of the present invention.
[0070] Figures 63A-63D This is a size distribution diagram of cell-free DNA molecules from maternal plasma samples of preeclampsia and normal-blood-pressure women in late pregnancy, according to an embodiment of the present invention.
[0071] Figures 64A-64D This is a size distribution diagram of cell-free DNA molecules from maternal plasma samples of preeclampsia and normal-blood-pressure women in late pregnancy, according to an embodiment of the present invention.
[0072] Figure 65A-65D This is a size distribution map of DNA molecules encompassing fetal-specific alleles from maternal plasma samples of preeclampsia and normal-blood-tension women in late pregnancy, according to embodiments of the present invention.
[0073] Figures 66A-66D This is a size distribution map of DNA molecules encompassing fetal-specific alleles from maternal plasma samples of preeclampsia and normal-blood-tension women in late pregnancy, according to embodiments of the present invention.
[0074] Figures 67A-67DThis is a size distribution diagram of DNA molecules encompassing maternal-specific alleles from plasma samples of mothers in late pregnancy with preeclampsia and normal blood pressure, according to embodiments of the present invention.
[0075] Figures 68A-68D This is a size distribution diagram of DNA molecules encompassing maternal-specific alleles from plasma samples of mothers in late pregnancy with preeclampsia and normal blood pressure, according to embodiments of the present invention.
[0076] Figure 69A and Figure 69B This is a diagram showing the proportion of short DNA molecules containing fetal-specific and maternal-specific alleles in plasma samples of preeclampsia and normal-blood pressure mothers sequenced by PacBio SMRT according to an embodiment of the present invention.
[0077] Figure 70A and Figure 70B This is a graph showing the proportion of short DNA molecules in plasma samples from mothers with preeclampsia and normal blood pressure, sequenced by PacBio SMRT and Illumina sequencing according to embodiments of the present invention.
[0078] Figure 71 This is a size ratio diagram showing the relative proportions of short and long DNA molecules in plasma samples from preeclampsia and normal-blood pressure mothers sequenced using PacBio SMRT according to embodiments of the present invention.
[0079] Figures 72A-72D This shows the proportion of different ends of plasma DNA molecules in plasma samples from preeclampsia and normal-blood pressure mothers sequenced by PacBio SMRT according to embodiments of the present invention.
[0080] Figure 73 This illustrates a hierarchical cluster analysis of maternal plasma DNA samples from preeclampsia and normal-blood pressure late pregnancy, based on embodiments of the present invention, using the frequencies of plasma DNA molecules having four types of fragment ends (the first nucleotide at the 5' end of each strand), namely the C-terminus, G-terminus, T-terminus, and A-terminus.
[0081] Figure 74 This illustrates a hierarchical cluster analysis of maternal plasma DNA samples from preeclampsia and normal-blood-tension late pregnancy using 16 dinucleotide motifs XYNN (dinucleotide sequences derived from the first and second nucleotides at the 5' end) according to embodiments of the present invention.
[0082] Figure 75 This illustrates a hierarchical cluster analysis of maternal plasma DNA samples from preeclampsia and normal-blood-tension late pregnancy using 16 dinucleotide motifs NNXY (dinucleotide sequences derived from the third and fourth nucleotides at the 5' end) according to embodiments of the present invention.
[0083] Figure 76 This illustrates a hierarchical cluster analysis of maternal plasma DNA samples from preeclampsia and normal-blood-tension late pregnancy using 256 tetranucleotide motifs (dinucleotide sequences from the first to the fourth nucleotide at the 5' end) according to embodiments of the present invention.
[0084] Figures 77A-77D This invention illustrates T cell contributions at the ends of four types of fragments in plasma DNA samples from mothers with preeclampsia and normal blood pressure, according to embodiments of the present invention.
[0085] Figure 78 This invention illustrates a method for analyzing biological samples obtained from pregnant women to determine the likelihood of pregnancy-related conditions, according to embodiments of the present invention.
[0086] Figure 79 A diagram illustrating maternal inheritance of repetitive sequence-related diseases in the fetus according to an embodiment of the present invention.
[0087] Figure 80 A diagram illustrating paternal inheritance of repetitive sequence-related diseases in the fetus according to an embodiment of the present invention.
[0088] Figure 81 , Figure 82 and Figure 83 A table to display instances of diseases with extended repetitive sequences.
[0089] Figure 84 A table showing examples of repetitive sequence expansion detection and repetitive sequence-related methylation determination in the fetus according to embodiments of the present invention.
[0090] Figure 85 This invention illustrates a method for analyzing biological samples obtained from pregnant women to determine the likelihood of genetic diseases in the fetus, according to embodiments of the present invention.
[0091] Figure 86 This invention illustrates a method for analyzing biological samples obtained from pregnant women to determine kinship, according to an embodiment of the present invention.
[0092] Figure 87 This shows the methylation patterns of two representative plasma DNA molecules after size selection.
[0093] Figure 88 This is a sequencing information table for size-selected and non-size-selected samples according to an embodiment of the present invention.
[0094] Figure 89A and Figure 89BA diagram showing the plasma DNA size profiles of samples selected based on bead size and samples not selected based on bead size, according to embodiments of the present invention.
[0095] Figure 90A and Figure 90B This shows the size profile between fetal DNA molecules and maternal DNA molecules in a size-selected sample according to an embodiment of the present invention.
[0096] Figure 91 This is a statistical table showing the number of plasma DNA molecules carrying informative SNPs between size-selected samples and non-size-selected samples according to embodiments of the present invention.
[0097] Figure 92 A table showing the degree of methylation in size-selected and non-size-selected plasma DNA samples according to embodiments of the present invention.
[0098] Figure 93 This is a table showing the degree of methylation in maternal or fetal specific cell-free DNA molecules according to embodiments of the present invention.
[0099] Figure 94 This is a table of the first 10 terminal motifs in size-selected and non-size-selected samples according to an embodiment of the present invention.
[0100] Figure 95 This is a receiver operating characteristic (ROC) plot illustrating the enhanced tissue efficacy analysis of long plasma DNA molecules according to an embodiment of the present invention.
[0101] Figure 96 The principle of airport sequencing for plasma DNA molecules according to an embodiment of the present invention is illustrated.
[0102] Figure 97 This is a table showing the percentage of plasma DNA molecules and their corresponding methylation levels within a specific size range according to embodiments of the present invention.
[0103] Figure 98 This is a diagram illustrating the size distribution and methylation patterns across different sizes according to an embodiment of the present invention.
[0104] Figure 99 This is a table showing the fetal DNA fraction determined using nanopore sequencing according to an embodiment of the present invention.
[0105] Figure 100 This is a table showing the degree of methylation between fetal-specific DNA molecules and maternal-specific DNA molecules according to embodiments of the present invention.
[0106] Figure 101This is a table showing the percentage of plasma DNA molecules within a specific size range and the corresponding methylation levels of fetal and maternal DNA molecules according to embodiments of the present invention.
[0107] Figure 102A and Figure 102B This is a size distribution map of fetal DNA molecules and maternal DNA molecules determined by nanopore sequencing according to an embodiment of the present invention.
[0108] Figure 103 A diagram illustrating the differences in methylation levels between fetal DNA molecules and maternal DNA molecules based on a single informative SNP and two informative SNPs, according to embodiments of the present invention.
[0109] Figure 104 This is a table showing the difference in methylation levels between fetal DNA molecules and maternal DNA molecules according to embodiments of the present invention.
[0110] Figure 105 A measurement system according to an embodiment of the present invention is illustrated.
[0111] Figure 106 This illustrates a computer system according to an embodiment of the present invention.
[0112] the term
[0113] A “tissue” corresponds to a group of cells grouped into functional units within a pregnant individual or their fetus. More than one type of cell can be found in a single tissue. Different types of tissues may be composed of different types of cells (e.g., hepatocytes, alveolar cells, or blood cells), but may also correspond to tissues from different organisms (mother and fetus; tissues from a pregnant individual who has received a transplant; tissues from a pregnant organism infected with a microorganism or virus or its fetus). A “reference tissue” may correspond to the tissue used to determine the degree of tissue-specific methylation. Multiple samples of the same tissue type from different pregnant individuals or their fetuses can be used to determine the degree of tissue-specific methylation for that tissue type.
[0114] "Biosample" refers to any sample taken from a pregnant individual (e.g., a human (or other animal), such as a pregnant woman, a person with a disease or a suspected pregnant person with a disease, a pregnant organ transplant recipient, or a pregnant individual suspected of having a disease involving an organ (e.g., the heart in a myocardial infarction or the brain in a stroke or the hematopoietic system in anemia) and containing one or more nucleic acid molecules of interest. Biosamples can be bodily fluids, such as blood, plasma, serum, urine, vaginal fluid, vaginal douche fluid, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirated fluid from different parts of the body (e.g., the thyroid gland, the breast), intraocular fluid (e.g., aqueous humor), etc. Fecal samples may also be used. In various embodiments, a majority of the DNA in a biosample enriched with cell-free DNA (e.g., a plasma sample obtained via centrifugation) may be cell-free, for example, more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA may be cell-free. The centrifugation protocol may include, for example, centrifuging at 3,000 g for 10 minutes to obtain a fluid fraction, and then centrifuging again at, for example, 30,000 g for 10 minutes to remove residual cells. As part of the biological sample analysis, a statistically significant number of cell-free DNA molecules in the biological sample may be analyzed (e.g., to provide accurate measurement results). In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more cell-free DNA molecules may be analyzed. At least the same number of sequence reads may be analyzed.
[0115] A “sequence read” refers to a string of nucleotides sequenced from any part or all of a nucleic acid molecule. For example, a sequence read can be a short string of nucleotides (e.g., 20-150 nucleotides) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained in various ways, such as using sequencing technologies or using probes, for example, in hybridization arrays or capture probes that can be used in microarrays; or amplification technologies, such as polymerase chain reaction (PCR) or linear or isothermal amplification using a single primer. As part of a biological sample analysis, a statistically significant number of sequence reads may be analyzed, for example, at least 1,000 sequence reads may be analyzed. As other examples, at least 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 or more sequence reads may be analyzed.
[0116] A "site" (also called a "genomic site") corresponds to a single location, which can be a single base position or a group of related base positions, such as a CpG site or a larger group of related base positions. A "locus" can correspond to a region containing multiple sites. A locus can contain only one site, in which case the locus is equivalent to a single site.
[0117] "Methylation state" refers to the methylation state at a given location. For example, the site can be methylated, unmethylated, or in some cases, undetermined.
[0118] The “methylation index” of each genomic site (e.g., a CpG site) refers to the proportion of DNA fragments exhibiting methylation at that site (e.g., determined by sequence reads or probes) relative to the total number of reads covering that site. A “read” may correspond to information obtained from a DNA fragment (e.g., the methylation state at the site). Reads can be obtained using reagents (e.g., primers or probes) that preferentially hybridize to DNA fragments having a specific methylation state at one or more sites. Typically, these reagents are applied after treatment with methods that differentially modify or identify DNA molecules based on their methylation state, such as bisulfite conversion, methylation-sensitive restriction enzymes, methylation-binding proteins, anti-methylcytosine antibodies, or single-molecule sequencing technologies that identify methylcytosine and hydroxymethylcytosine (e.g., single-molecule real-time sequencing and nanopore sequencing (e.g., from Oxford Nanopore Technologies)).
[0119] The “methylation density” of a region can refer to the number of reads at sites within a methylated region divided by the total number of reads covering those sites within the region. Sites can have specific characteristics, such as CpG sites. Therefore, the “CpG methylation density” of a region can refer to the number of reads showing CpG methylation divided by the total number of reads covering CpG sites within the region (e.g., specific CpG sites, CpG sites within CpG islands, or larger regions). For example, the methylation density of each 100kb site in the human genome can be determined by measuring the total number of unconverted cytosines (corresponding to methylated cytosines) at CpG sites after bisulfite treatment as the proportion of all CpG sites covered by a sequence read located in a 100kb region. This analysis can also be performed for other site sizes, such as 500bp, 5kb, 10kb, 50kb, or 1Mb. A region can be the entire genome, a chromosome, or a portion of a chromosome (e.g., a chromosome arm). The methylation index of a CpG site is the same as the methylation density of a region when it contains only the CpG site. The "ratio of methylated cytosine" can refer to the number of methylated cytosine sites "C's" relative to the total number of cytosine residues analyzed, i.e., the total number of cytosines in the region excluding the CpG case). Methylation index, methylation density, the count of molecules methylated at one or more sites, and the ratio of molecules (e.g., cytosines) methylated at one or more sites are examples of "degree of methylation". In addition to bisulfite conversion, other methods known to those skilled in the art can be used to query the methylation status of DNA molecules, including but not limited to enzymes sensitive to methylation status (e.g., methylation-sensitive restriction enzymes), methylation-binding proteins, single-molecule sequencing using platforms sensitive to methylation status (e.g., nanopore sequencing (Schreiber et al., Proceedings of the National Academy of Sciences, 2013; 110:18910-18915) and single-molecule real-time sequencing (e.g., single-molecule real-time sequencing from Pacific Biosciences (Flusberg et al., Nature Methods, 2010; 7:461-465)).
[0120] The methylome provides a measure of the amount of DNA methylation at multiple sites or loci in the genome. The methylome can correspond to the entire genome, a significant portion of the genome, or one or more relatively small parts of the genome.
[0121] A “methylation profile” contains information related to DNA or RNA methylation at multiple sites or regions. Information related to DNA methylation may include, but is not limited to, the methylation index of CpG sites, the methylation density (MD) of CpG sites in a region, the distribution of CpG sites in adjacent regions, the methylation pattern or degree of each individual CpG site within a region containing more than one CpG site, and non-CpG methylation. In one embodiment, a methylation profile may include methylation or non-methylation patterns of more than one type of base (e.g., cytosine or adenine). A significant portion of the methylation profile of the genome can be considered equivalent to a methylome. In mammalian genomes, “DNA methylation” typically refers to the addition of a methyl group to the 5' carbon of a cytosine residue in a CpG dinucleotide (i.e., 5-methylcytosine). DNA methylation can occur in cytosine in other cases, such as CHG and CHH, where H is adenine, cytosine, or thymine. Cytosine methylation can also occur in the form of 5-hydroxymethylcytosine. Non-cytosine methylation, such as N, has also been reported. 6 -Methyladenine.
[0122] A “methylation pattern” refers to the sequence of methylated and unmethylated bases. For example, a methylation pattern can be the sequence of methylated bases on a single DNA strand, a single double-stranded DNA molecule, or another type of nucleic acid molecule. As an example, three consecutive CpG sites can have any of the following methylation patterns: UUU, MMM, UMM, UMU, UUM, MUM, MUU, or MMU, where “U” indicates an unmethylated site and “M” indicates a methylated site. When we extend this concept to base modifications that include, but are not limited to, methylation, we will use the term “modification pattern,” which refers to the sequence of modified and unmodified bases. For example, a modification pattern can be the sequence of modified bases on a single DNA strand, a single double-stranded DNA molecule, or another type of nucleic acid molecule. As an example, three consecutive potentially modifiable sites can have any of the following modification patterns: UUU, MMM, UMM, UMU, UUM, MUM, MUU, or MMU, where “U” indicates an unmodified site and “M” indicates a modified site. An example of non-methylated base modification is the oxidation change in 8-oxo-guanine.
[0123] The terms "hypermethylation" and "hypomethylation" can refer to the methylation density of a single DNA molecule, as measured by the degree of methylation of its individual molecules, for example, the number of methylated bases or nucleotides within the molecule divided by the total number of methylatable bases or nucleotides within the molecule. A hypermethylated molecule is a molecule in which the degree of methylation of its individual molecules is equal to or higher than a threshold, which can be defined according to different applications. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%. A hypomethylated molecule is a molecule in which the degree of methylation of its individual molecules is equal to or lower than a threshold, which can be defined and varied according to different applications. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%.
[0124] The terms "hypermethylation" and "hypomethylation" can also refer to the degree of methylation of a population of DNA molecules, as measured by the degree of multimolecular methylation of these molecules. A population of highly methylated molecules is one in which the degree of multimolecular methylation is equal to or higher than a threshold, which can be defined and varied depending on the application. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%. A population of hypomethylated molecules is one in which the degree of multimolecular methylation is equal to or lower than a threshold, which can be defined depending on the application. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, and 95%. In one embodiment, the population of molecules can be compared to one or more selected genomic regions. In one embodiment, one or more selected genomic regions may be associated with diseases such as genetic disorders, memory disorders, metabolic disorders, or neurological disorders. One or more selected genomic regions may be 50 nucleotides (nt), 100 nt, 200 nt, 300 nt, 500 nt, 1000 nt, 2 knt, 5 knt, 10 knt, 20 knt, 30 knt, 40 knt, 50 knt, 60 knt, 70 knt, 80 knt, 90 knt, 100 knt, 200 knt, 300 knt, 400 knt, 500 knt, or 1 Mnt in length.
[0125] The term "sequencing depth" refers to the number of times a locus is covered by sequence reads aligned to that locus. A locus may be as small as a nucleotide, as large as a chromosome arm, or as large as the entire genome. Sequencing depth can be expressed as 50×, 100×, etc., where "×" indicates the number of times the locus is covered by sequence reads. Sequencing depth can also be applied to multiple loci or the entire genome; in this case, × can refer to the average number of times the locus, haploid genome, or entire genome is sequenced. Ultra-deep sequencing can refer to a sequencing depth of at least 100×.
[0126] A “calibration sample” may correspond to a biological sample whose clinically relevant DNA fraction concentration (e.g., tissue-specific DNA fraction) is known or determined via a calibration method, such as using tissue-specific alleles, as in transplantation in pregnant individuals, which are present in the donor genome but not in the recipient genome and can be used as markers for the transplanted organ. As another example, a calibration sample may correspond to a sample from which terminal motifs can be determined. Calibration samples can be used for two purposes.
[0127] A “calibration data point” comprises a “calibration value” and a measured or known fractional concentration of clinically relevant DNA (e.g., DNA from a specific tissue type). The calibration value can be determined by a relative frequency (e.g., ensemble value) measured for a calibration sample in which the fractional concentration of clinically relevant DNA is known. Calibration data points can be defined in various ways, such as as discrete points or as a calibration function (also known as a calibration curve or calibration surface). The calibration function can be derived from additional mathematical transformations of the calibration data points.
[0128] A “separation value” corresponds to the difference or ratio involving two values, such as two contribution fractions or two degrees of methylation. A separation value can be a simple difference or ratio. As an example, the proportions of x / y and x / (x+y) are separation values. A separation value can include other factors, such as multiplicative factors. As another example, the difference or ratio can be a function of the values, such as the difference or ratio of the natural logarithms (ln) of the two values. A separation value can include both differences and ratios.
[0129] "Separation values" and "aggregate values" (e.g., belonging to relative frequencies) are two instances of parameters (also called measures) that provide a measure of sample variation between different categories (states) and can therefore be used to determine different categories. Aggregate values can be separation values, for example, when taking the difference between a set of relative frequencies of a sample and a set of reference relative frequencies, as is generally done in a cluster.
[0130] As used herein, the term "classification" refers to any one or more numbers or other characters associated with a specific characteristic of a sample. For example, the symbol "+" (or the word "positive") can indicate that a sample is classified as having a missing or augmented feature. Classifications can be binary (e.g., positive or negative) or have more levels of classification (e.g., levels from 1 to 10 or from 0 to 1).
[0131] As used herein, the term "parameter" refers to a numerical value that characterizes a set of quantitative data and / or the numerical relationship between sets of quantitative data. For example, the ratio (or a function of the ratio) between a first quantity of a first nucleic acid sequence and a second quantity of a second nucleic acid sequence is a parameter.
[0132] The term "size profile" broadly refers to the size of DNA fragments in a biological sample. A size profile can be a histogram providing the distribution of a given number of DNA fragments of various sizes. Various statistical parameters (also called size parameters or simply parameters) can be used to distinguish one size profile from another. A parameter is the percentage of DNA fragments of a specific size or size range relative to all DNA fragments or relative to another size or size range.
[0133] The terms “cutoff value” and “threshold” refer to predetermined numerical values used in operation. For example, a cutoff size may refer to the size at which segments larger than it are excluded. A threshold may be a value above or below which a particular classification applies. Either of these terms may be used in either of these contexts. A cutoff value or threshold may be a “reference value” or derived from a value that represents a particular classification or distinguishes between two or more classifications. As those skilled in the art will understand, such reference values can be determined in various ways. For example, a metric may be determined for individuals in two different groups with different known classifications, and a reference value may be selected to represent the value between a metric of one classification (e.g., the mean) or two clusters (e.g., selected to obtain desired sensitivity and specificity). As another example, a reference value may be determined based on statistical analysis or simulation of a sample. Specific cutoff values, thresholds, reference values, etc., may be determined based on desired accuracy (e.g., sensitivity and specificity).
[0134] "Pregnancy-related conditions" encompass any condition characterized by abnormal relative expression levels of genes in maternal and / or fetal tissues or by abnormal clinical features in the mother and / or fetus. These conditions include, but are not limited to, maternal preeclampsia (Kaartokallio et al., *Scientific Reports*, 2015; 5:14107; Medina-Bastidas et al., *International Journal of Molecular Sciences*, 2020; 21:3597), intrauterine growth restriction (Faxén et al., *American Journal of Perinatology*, 1998; 15:9-13; Medina-Bastidas et al., *International Journal of Molecular Sciences*, 2020; 21:3597), invasive placental formation, and preterm birth (Enquobahrie et al., *BMC Pregnancy and Childbirth*, BMC...). (PregnancyChildbirth.) 2009; 9:56), neonatal hemolytic disease, placental insufficiency (Kelly et al. Endocrinology. 2017; 158:743-755), fetal hydrops (Magor et al. Blood. 2015; 125:2405-17), fetal malformations (Slonim et al. Proceedings of the National Academy of Sciences 2009; 106:9425-9), HELLP syndrome (Dijk et al. Journal of Clinical Research 2012; 122:4003-4011), systemic lupus erythematosus (Hong et al. Journal of Experimental Medicine 2019; 216:1154-1169), and other immune diseases.
[0135] The abbreviation "bp" stands for base pair. In some cases, "bp" can be used to indicate the length of a DNA fragment, even if the DNA fragment is single-stranded and does not contain base pairs. In the case of single-stranded DNA, "bp" can be interpreted as providing length in units of nucleotides.
[0136] The abbreviation "nt" stands for nucleotide. In some cases, "nt" can be used to indicate the length of a single strand of DNA, measured in bases. Additionally, "nt" can be used to indicate relative position, such as upstream or downstream of the analyzed locus. For double-stranded DNA, unless the context clearly indicates otherwise, "nt" may still refer to the length of a single strand rather than the total number of nucleotides in both strands. In some situations concerning technical conceptualization, data display, processing, and analysis, "nt" and "bp" are used interchangeably.
[0137] The term "machine learning model" can refer to a model that makes predictions about test data using sample data (such as training data), and therefore can include supervised learning. Machine learning models are often developed using computers or processors. Machine learning models can include statistical models.
[0138] The term "data analytics framework" can encompass algorithms and / or models that take data as input and subsequently output predicted results. Examples of "data analytics frameworks" include statistical models, mathematical models, machine learning models, other artificial intelligence models, and combinations thereof.
[0139] The term "real-time sequencing" can refer to techniques that involve collecting or monitoring data during the progress of the reaction involved in sequencing. For example, real-time sequencing may involve optical monitoring or imaging of DNA polymerase and the presence of new bases.
[0140] The term "subsequence" can refer to a string of bases smaller than the complete sequence of a nucleic acid molecule. For example, when the complete sequence of a nucleic acid molecule contains 5 or more bases, the subsequence can contain 1, 2, 3, or 4 bases. In some embodiments, a subsequence can refer to a string of bases forming a unit, wherein the unit is repeated multiple times in a tandem sequence. Examples include 3nt units or subsequences repeated at loci associated with trinucleotide repeat sequence disorders, 1nt to 6nt units or subsequences repeated 5 to 50 times as microsatellites, 10nt to 60nt units or subsequences repeated 5 to 50 times as microsatellites, or included in other genetic elements such as Alu repeat sequences.
[0141] The term "about / approximately" may mean within an acceptable margin of error for a particular value as determined by a person skilled in the art, depending in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, "about" may mean within 1 or greater than 1 standard deviation. Alternatively, "about" may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or methods, the term "about" may mean within a certain order of magnitude of the value, within 5 times, and more preferably within 2 times. When a particular value is described in this application and claims, unless otherwise stated, the term "about" should be assumed to mean within an acceptable margin of error for the particular value. The term "about" may have the meaning as commonly understood by a person skilled in the art. The term "about" may mean ±10%. The term "about" may mean ±5%.
[0142] Where a range of values is provided, it should be understood that, unless the context clearly indicates otherwise, each interpolated value between the upper and lower limits of the range is also specifically disclosed, accurate to the tenths of the lower limit unit. Each smaller range between any stated or interpolated value within the stated range, and any other stated or interpolated value within the stated range, is covered within embodiments of this disclosure. The upper and lower limits of these smaller ranges may be independently included in or excluded from the range, and each range containing any limit, infinity, or both limits is also covered within this disclosure, subject to any specifically excluded limit within the stated range. If the stated range contains one or both of the limits, the range excluding any or both of those included limits is also included within this disclosure.
[0143] Standard abbreviations can be used, such as bp (base pair); kb (kilobase); pi (picoliter); s or sec (second); min (minute); h or hr (hour); aa (amino acid); nt (nucleotide); etc.
[0144] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. While any methods and materials similar to or equivalent to those described herein may be used in the practice or testing of embodiments of this disclosure, some potential and exemplary methods and materials are hereby described. Detailed Implementation
[0145] The analysis of cell-free DNA molecules primarily involves short cell-free DNA fragments, often attributed to limitations in analytical techniques. The limited ability of Illumina sequencing to obtain sequence information from long DNA molecules is evidenced in recent sequencing results of cell-free DNA from mice (Serpas et al., *Proceedings of the National Academy of Sciences*, 2019; 116:641-649). Only 0.02% of the sequenced DNA molecules in wild-type mice using Illumina sequencing were in the 600bp and 2000bp range. Even when sequencing DNA libraries originally prepared for Illumina sequencing using single-molecule real-time (SMRT) technology from Pacific Biosciences (i.e., PacBio SMRT sequencing), only 0.33% of the sequenced DNA molecules in the 600bp and 2000bp range remained. These reported data indicate that the sequencing step results in the loss of 93% of the long DNA molecules in the 600bp and 2000bp range present in the original DNA library.
[0146] We hypothesize that, due to the limitations of PCR in amplifying the long DNA molecules described above, the DNA library preparation step also results in the loss of a significant proportion of long, free DNA molecules. Jahr et al. reported the presence of large fragments, such as ~10,000 kilobases, using gel electrophoresis (Jahr et al., *Cancer Res.*, 2001; 61:1659-65). However, the bands shown in the gel electrophoresis images will not readily provide sequence information for these molecules in the gel, let alone epigenetic information.
[0147] We have previously used the Oxford Nanopore Technologies sequencing platform to study cell-free DNA extracted from maternal plasma (Cheng et al., Clin Chem. 2015; 61:1305-6). We observed a very small percentage (0.06% to 0.3%) of long plasma DNA greater than 1 kb. We hypothesize that this low percentage may be a result of the platform's low sequencing accuracy.
[0148] In the field of cell-free DNA, most research focuses on short DNA molecules (e.g., <600 bp). The characteristics of the genetic and epigenetic information contained in long cell-free DNA molecules are unexplored. This disclosure provides a systematic approach to analyzing long cell-free DNA molecules, including decoding their genetic and epigenetic information and their clinical utility in non-invasive prenatal testing, such as, but not limited to, non-invasive detection of monogenic disorders, elucidation of the fetal genome (e.g., non-invasive whole fetal genome sequencing), detection of re-mutations at the genome-wide level, and detection / monitoring of pregnancy-related conditions such as preeclampsia and preterm birth.
[0149] I. Free DNA Size Analysis
[0150] Cell-free DNA samples obtained from pregnant women were sequenced, revealing that a significant portion of the DNA fragments were long. Accurate sequencing of these long cell-free DNA fragments was demonstrated. The size profile of these long cell-free DNA molecules was analyzed. The amount of fetal and maternal long cell-free DNA molecules was compared. Long cell-free DNA molecules can be more accurately compared with a reference genome. Long cell-free DNA molecules can be used to determine haplotype inheritance.
[0151] A plasma DNA sample from a pregnant woman in late pregnancy was analyzed using PacBio SMRT sequencing. Double-stranded free DNA molecules were bound to hairpin adaptors and subjected to single-molecule real-time sequencing using zero-mode waveguides and single polymerase molecules (Eid et al., Science, 2009; 323:133-8).
[0152] We sequenced 1.1 billion subreads, of which 659.3 million were aligned to the human reference genome (hg19). The subreads were generated from 4.6 million PacBio single-molecule real-time (SMRT) sequencing wells, each containing at least one alignment-alignable subread to the human reference genome. On average, each molecule in the SMRT well was sequenced an average of 143 times. In this example, 4.5 million circular common sequences (CCS) were found, indicating 4.5 million cell-free DNA molecules available for downstream analysis. The size of each cell-free DNA molecule was determined by counting the number of identified bases using the CCS.
[0153] Figure 1A and Figure 1B Displays the size distribution of free DNA from 0kb to 20kb. The y-axis shows frequency. The x-axis shows linear scaling. Figure 1A ) or logarithmic scale ( Figure 1B DNA fragment size measures from 0 kb to 20 kb in base pairs. Because sequencing is performed on full-length DNA molecules, the size of each DNA molecule can be determined directly by counting the number of nucleotides in subreads or CCS. DNA fragment size measurements can be achieved using any sequencing platform capable of reading full-length DNA fragments, and are not limited to single-molecule sequencers. For example, the Sanger sequencer can read at 800 bp. Short-read sequencing, such as that performed on the Illumina platform, can read at 250 bp. Single-molecule sequencers, such as those from Pacific Biosciences and Oxford Nanopore, can read at over 10,000 bp. DNA fragment size can also be determined after alignment with a reference genome, such as the human reference genome. DNA fragment size can be determined by bilateral sequencing followed by alignment with a reference genome. Figure 1B A long-tail pattern was observed. Among the 4.5 million CCSs, 22.5% contained cell-free DNA longer than 200 bp, 19.0% longer than 300 bp, 11.8% longer than 400 bp, 10.6% longer than 500 bp, 8.9% longer than 600 bp, 6.4% longer than 1 kb, 3.5% longer than 2 kb, 1.9% longer than 3 kb, 0.9% longer than 4 kb, and 0.04% longer than 10 kb. The longest cell-free DNA observed in the current PacBio SMRT results was 29,804 bp.
[0154] A plasma DNA sample from a pregnant individual was also sequenced on the Illumina sequencing platform using a PCR-based library preparation protocol (Lun et al., Clinical Chemistry, 2013; 59:1583-94). Of the 18.2 million bilateral reads, 5.3% were cell-free DNA larger than 200 bp, 2.0% were larger than 300 bp, 0.3% were larger than 400 bp, 0.2% were larger than 500 bp, and 0.2% were larger than 600 bp (Table 1). For comparison, we analyzed size profiles by aggregating single-molecule real-time sequencing data from five pregnant individuals (i.e., a total of 4.4 million CCS). We observed a greater proportion of plasma DNA molecules larger than 600 bp (28.56%) compared to the corresponding 0.2% larger than 600 bp plasma DNA molecules obtained via the Illumina sequencing platform. These results demonstrate that PacBio SMRT sequencing enables us to obtain DNA molecules longer than 143 times (over 600 bp). We can obtain 4.77% of plasma DNA molecules larger than 3 kb using single-molecule real-time sequencing, which is not available on the Illumina sequencing platform.
[0155] In contrast to previous reports (Cheng et al., Clinical Chemistry 2015; 61:1305-6) showing a very small percentage (0.06% to 0.3%) of long plasma DNA molecules larger than 1 kb using the Oxford Nanopore Technologies sequencing platform, we were able to obtain more than 21-fold (6.4%) of plasma DNA larger than 1 kb, demonstrating that PacBio SMRT sequencing is far more efficient in obtaining sequence information from long DNA populations.
[0156] Compared to bilateral short-read sequencing platforms like Illumina, long-read sequencing technologies such as PacBio SMRT offer several advantages in determining the characteristics (e.g., length) of long DNA fragments. For example, long reads generally allow for more accurate alignment with a human reference genome (e.g., hg19). Long-read technologies also allow for accurate determination of plasma DNA molecule length by directly counting the sequenced nucleotides. In contrast, plasma DNA size assessment based on bilateral short reads is an indirect method of inferring the size of plasma DNA molecules using the outermost coordinates of aligned bilateral reads. For such indirect approaches, alignment errors will affect accurate size inference. In this regard, increasing the size span between bilateral reads will increase the probability of alignment errors.
[0157]
[0158] Table 1. Comparison of size distribution between PacBio and Illumina sequencing of cell-free DNA.
[0159] Figure 2A and Figure 2B Displays the size distribution of free DNA from 0kb to 5kb. The y-axis shows frequency. The x-axis shows linear scaling. Figure 2A ) or logarithmic scale ( Figure 2B The size, measured in base pairs, ranges from 0 kb to 5 kb. A series of main peaks appear with a periodic pattern. This periodic pattern even extends to molecules in the 1 kb and 2 kb range. The peak with the highest frequency (2.6%) is at 166 bp, which is consistent with previous findings using Illumina technology (Lo et al., Science Translational Medicine, 2010; 2:61ra91). Figure 2B The distance between adjacent main peaks in the data is about 200 bp, indicating that the generation of long free DNA will also involve nucleosome structure.
[0160] Figure 3A and Figure 3B Displays the size distribution of free DNA from 0 bp to 400 bp. The y-axis shows frequency. The x-axis shows linear scale. Figure 3A ) or logarithmic scale ( Figure 3B Sizes ranging from 0 bp to 400 bp in base pairs. The characteristic features previously reported (Lo et al., *Science Translational Medicine*, 2010; 2:61ra91), with a dominant peak at 166 bp and a periodic 10 bp occurrence in molecules smaller than 166 bp, can also be reproduced using the novel method of this disclosure. These results demonstrate that determining molecular size by counting the number of bases sequenced from single molecules according to this disclosure is reliable.
[0161] A. Size analysis of fetal and maternal DNA
[0162] The sizes of maternal and fetal DNA fragments were analyzed and compared. As an example, erythrocyte sedimentation rate (ESR) amber layer DNA and matched placental DNA from a pregnant woman were sequenced to obtain 59× and 58× haploidentical genome coverage, respectively. We identified a total of 822,409 informative single nucleotide polymorphisms (SNPs), with the mother being homozygous and the fetus heterozygous. Fetal-specific alleles were defined as those alleles present in the fetal genome but not in the maternal genome. We identified 2,652 fetal-specific fragments and 24,837 shared fragments (i.e., fragments carrying shared alleles; primarily of maternal origin) in maternal plasma (M13160) via PacBio sequencing. The fetal DNA fraction was 21.8%.
[0163] Figure 4A and Figure 4BThis displays the size distribution of cell-free DNA between segments carrying shared alleles (shared) and segments carrying fetal-specific alleles (fetal-specific). The x-axis shows the linear scale (…). Figure 4A ) or logarithmic scale ( Figure 4B Sizes ranging from 0 kb to 20 kb in base pairs. Both fragments carrying shared alleles (primarily maternally derived) and fragments carrying fetal-specific alleles (placental-derived) exhibit long-tailed distributions, indicating the presence of long DNA molecules from both fetal and maternal origins. For the predominantly maternally derived fragments, 22.6% of plasma DNA molecules were larger than 2 kb, compared to 8.5% for fetal-derived fragments. These results suggest that fetal DNA molecules contain fewer long DNA molecules. The percentage of long DNA present in this SNP-based analysis of fetal and maternal origins in plasma DNA appears to be significantly higher than the percentage observed in the overall size analysis. This difference may be attributed to the fact that long DNA molecules are more likely to cover one or more SNPs than short DNA molecules, and therefore long DNA will be advantageously selected for SNP-based analysis. The relative proportion of SNP-tagged long DNA molecules deviating from the proportion of long DNA in the corresponding original pool will be controlled by the size of those molecules. Among those fetal-specific DNA fragments, the longest DNA fragment is 16,186 bp, while among those fragments carrying shared alleles, the longest fragment is 24,166 bp.
[0164] Figure 5A and Figure 5B This displays the size distribution of cell-free DNA between segments carrying shared alleles (shared) and segments carrying fetal-specific alleles (fetal-specific). The x-axis shows the linear scale (…). Figure 5A ) or logarithmic scale ( Figure 5B The size of DNA fragments ranging from 0kb to 5kb, measured in base pairs, is considered. For both fetal-specific and shared DNA fragments, fragments smaller than 2kb exhibit a series of dominant peaks appearing periodically. These dominant peaks may be correlated with nucleosome structures.
[0165] Figure 6A and Figure 6B This displays the size distribution of cell-free DNA between segments carrying shared alleles (shared) and segments carrying fetal-specific alleles (fetal-specific). The x-axis shows the linear scale (…). Figure 6A ) or logarithmic scale ( Figure 6BThe size of fetal DNA ranges from 0 kb to 1 kb, measured in base pairs. For those smaller than 1 kb segments, both fetal-specific and shared DNA segments, a series of dominant peaks appear in a periodic manner. These dominant peaks may correspond to nucleosome structures. There appears to be an observable shift in the fetal DNA size profile toward the left of the shared DNA segment size profile, suggesting that fetal DNA will include more short DNA molecules than maternal DNA.
[0166] Figure 7A and Figure 7B This displays the size distribution of cell-free DNA between segments carrying shared alleles (shared) and segments carrying fetal-specific alleles (fetal-specific). The x-axis shows the linear scale (…). Figure 7A ) or logarithmic scale ( Figure 7B Sizes ranging from 0 bp to 400 bp in base pairs. The previously reported (Lo et al., Science Translational Medicine 2010; 2:61ra91) characteristic features with a dominant peak at 166 bp and a 10 bp periodicity in both fetal and maternal molecules smaller than 166 bp can also be reproduced using the novel method of this disclosure. These results demonstrate that determining molecular size by counting the number of bases sequenced from single molecules according to this disclosure is reliable.
[0167] B. Size and methylation analysis
[0168] Analysis of the methylation levels of long-chain maternal and fetal DNA molecules revealed that the methylation level of fetal DNA molecules was lower than that of maternal DNA molecules.
[0169] In PacBio SMRT sequencing, DNA polymerase mediates the incorporation of fluorescently labeled nucleotides into complementary strands. The characteristics of the fluorescence pulses generated during DNA synthesis, including the inter-pulse duration and pulse width, reflect polymerase kinetics that can be used to determine, for example, but not limited to, nucleotide modifications of 5-methylcytosine using the pathways described in our previously published application (U.S. Application No. 16 / 995,607, filed August 17, 2020, entitled "Determination of Base Modifications of Nucleic Acids"), the entire contents of which are incorporated herein by reference for all purposes.
[0170] In this embodiment, we identified 95,210 fragments carrying maternally specific alleles and 2,652 fragments carrying fetal specific alleles. Maternally specific alleles are defined herein as those alleles present in the maternal genome but not in the fetal genome that can be identified from SNPs in which the mother is heterozygous and the fetus is homozygous. In this example, we identified a total of 677,375 informative SNPs. We determined the size of each cell-free DNA molecule. In one embodiment, when the methylation state in the genome is variable, for example, the methylation level of CpG islands is generally lower than that of regions without CpG islands, to minimize variability introduced by genomic conditions, we may use computer simulations to select fragments larger than 1 kb, containing at least 5 CpG sites, and corresponding to a CpG density of less than 5% (i.e., the number of CpG sites in the molecule divided by the total length of the molecule < 0.05) for downstream analysis.
[0171] Figure 8 This chart shows the degree of single-molecule double-stranded DNA methylation between segments carrying maternally specific alleles and segments carrying fetal specific alleles. The y-axis shows the degree of single-molecule double-stranded DNA methylation as a percentage. The x-axis shows both segments carrying maternally specific alleles and segments carrying fetal specific alleles. The degree of single-molecule double-stranded DNA methylation in segments carrying fetal specific alleles (mean: 62.7%; IQR: 50.0%–77.2%) was lower than the corresponding degree of single-molecule double-stranded DNA methylation in segments carrying maternally specific alleles (mean: 72.7%; IQR: 60.6%–83.3%) (P<0.0001).
[0172] Figure 9A This displays the empirical distribution of single-molecule double-stranded DNA methylation levels in fragments fitted using kernel density estimation performed in the R suite (r-project.org / ). Frequency is shown on the y-axis. The x-axis shows single-molecule double-stranded DNA methylation levels as a percentage. The distribution of fetal-specific long DNA fragments is to the left of the distribution of maternal-specific fragments, indicating the presence of lower single-molecule double-stranded DNA methylation levels in fetal DNA molecules.
[0173] Figure 9BThe image shows a receiver operating characteristic (ROC) analysis performed using the degree of methylation of single-molecule double-stranded DNA. The y-axis shows sensitivity, and the x-axis shows specificity. ROC analysis using the degree of methylation of single-molecule double-stranded DNA to study the dynamics of distinguishing fetal DNA fragments from maternal DNA fragments revealed an area under the ROC curve (AUC) of 0.62, greater than 0.5 for random guessing. In the examples, we can further refine the determination of fetal / maternal origin of fragments in plasma by utilizing spatial patterns of methylation states (e.g., the sequence of methylated states) and the relative or absolute distances between modified bases and genomic coordinates within a single molecule. In embodiments, we can combine methylation patterns with other fragmentomic metrics (i.e., parameters relating to DNA fragmentation), including but not limited to preferred ends (Chan et al., Proceedings of the National Academy of Sciences 2016; 113:E8159-8168), terminal motifs (Serpas et al., Proceedings of the National Academy of Sciences 2019; 116:641-649), size (Lo et al., Science Translational Medicine 2010; 2:61ra), orientation perception (i.e., regarding specific elements within the genome, such as open chromatin regions, orientation of fragmentation patterns (Sun et al., Genome Research 2019; 29:418-427)), and topological forms (e.g., linear pairs of circular DNA molecules (Ma et al., Clinical Chemistry 2019; 65:1161-1170)), thereby improving the classification dynamics for distinguishing fragments of placental origin (fetal origin).
[0174] Figure 10A and Figure 10B This displays the degree of single-molecule double-stranded DNA methylation in both fetal and maternal DNA fragments, varying according to fragment size. The y-axis shows the degree of single-molecule double-stranded DNA methylation as a percentage. The x-axis shows sizes from 0 kb to greater than 20 kb. Figure 10A ) and sizes from 0kb to greater than 1kb ( Figure 10B On the other hand, in a long range ( Figure 10A ) and short range ( Figure 10B Within both categories, the degree of single- and double-stranded DNA methylation in fetal-specific DNA molecules is generally lower than that in maternal-specific DNA molecules. For short DNA molecules, this finding is consistent with the current understanding that the degree of methylation of fetal DNA in maternal plasma is lower than that of maternal DNA (Lun et al., Clinical Chemistry 2013; 59:1583-94).
[0175] In this embodiment, because the methylation level of fetal DNA molecules is relatively lower than that of maternal DNA molecules, we will select molecules with single-molecule double-stranded DNA methylation levels less than, but not limited to, specific thresholds of 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, and 5%, to enrich fetal-origin cell-free DNA molecules in the plasma DNA pool. For example, for fragments >1kb, the fetal DNA fraction is 2.6%. If we select fragments with single-molecule double-stranded methylation levels <50% (>1kb), the fetal DNA fraction of those further selected >1kb fragments will increase to 5.6% (i.e., an increase of 115.4%). In another example, for fragments <200bp, the fetal DNA fraction is 26.2%. If we select fragments with a single-molecule double-stranded DNA methylation level <50% (<200bp), the fetal DNA fraction of those fragments >200bp, after further selection, will increase to 41.6% (i.e., 58.8%). Therefore, in some cases, the use of threshold single-molecule double-stranded DNA methylation level for enriching fetal DNA will be more effective for long DNA molecules.
[0176] C. Haplotype and long-term free DNA methylation
[0177] In the embodiments, we can use the methods described in this disclosure to obtain the base composition, size, and base modifications of each individual DNA molecule. SNP and methylation information of long cell-free DNA molecules can be used for haplotype analysis. The use of long DNA molecules present in the cell-free DNA pool disclosed in this disclosure will allow for the phasing of variants in the genome using haplotype information present in each common sequence, according to, but not limited to, the disclosed methods (Edge et al., *Genome Research* 2017; 27:801-812; Wenger et al., *Nature Biotechnology* 2019; 37:1155-1162). The implementation of haplotype determination based on cell-free DNA sequence information differs from previous studies that had to rely on long DNA prepared from tissue DNA. Haplotypes within a genomic region are sometimes referred to as haplotype blocks. Haplotype blocks can be considered as a group of alleles that have been phased on a chromosome. In some embodiments, haplotype blocks are extended as long as possible based on a set of sequence information supporting two alleles physically linked on the chromosome and allele overlap information between different sequences.
[0178] Figure 11A and Figure 11BExamples of identified long fetal-specific DNA molecules in maternal plasma DNA are shown. Among those fetal-specific DNA fragments, we specifically illustrate that embodiments of our invention use an alignment with a region in chromosome 10 (chr10:56282981-56299166) of the human reference genome. Figure 11A And carries 7 fetal-specific alleles ( Figure 11B The 16,186 bp molecule. Six of the seven fetal-specific alleles were consistent with alleles inferred from deep sequencing of the maternal and fetal genomes (using the Illumina platform). Figure 11B According to the method described in this disclosure, the degree of methylation was determined to be 27.1%. Figure 11B The methylation rate was significantly lower than the average level of maternally specific fragments (72.7%). These results suggest that single-molecule double-stranded DNA methylation patterns could serve as markers for distinguishing fetal from maternally originating cell-free DNA molecules.
[0179] Figure 12A and Figure 12B Examples of identified long maternal DNA molecules carrying shared alleles in maternal plasma DNA from pregnant women are shown. Among those segments carrying shared alleles, alignments are made with regions on chromosome 6 (chr6:111074371-111098536) of the human reference. Figure 12A And carries 18 shared alleles ( Figure 12B The longest fragment is 24,166 bp. All those shared alleles are consistent with allele information inferred from deep sequencing of the maternal and fetal genomes (using the Illumina platform). Figure 12B According to the method described in this disclosure, the degree of methylation was determined to be 66.9% (). Figure 12B Genetic and epigenetic information in free DNA molecules, which are approximately 1,000 base pairs in length, cannot be easily identified using short-read sequencing methods such as bisulfite sequencing (Illumina).
[0180] Here, we describe a method for determining the relative probability that a molecule originated from the pregnant woman or the fetus. In pregnant women, DNA molecules carrying the fetal genotype actually originate from the placenta, while most DNA molecules carrying the maternal genotype originate from maternal blood cells. In this method, we first construct frequency distribution curves of DNA molecules based on the degree of methylation of DNA molecules from both the placenta and maternal blood cells. To achieve this, we assign the human genome to nucleotides of different sizes.
[0181] Figure 13The frequency distribution of placental DNA (red) and maternal blood cell DNA (blue) at different resolutions from 1kb to 20kb is displayed according to the degree of methylation. Frequency is shown on the y-axis. Methylation degree is shown on the x-axis. Examples of locus sizes include, but are not limited to, 1kb, 2kb, 5kb, 10kb, 15kb, and 20kb. The methylation degree of each locus is determined based on the number of methylated CpG sites divided by the total number of CpG sites. After determining the methylation degree of all loci, frequency distribution curves for each locus in the placental genome and maternal blood cell genome can be constructed for different locus sizes.
[0182] Based on the degree of methylation of long DNA molecules, their likelihood of originating from the placenta or maternal blood cells can be determined by the relative abundance of the two types of DNA molecules at this level of methylation and the fractional concentration of fetal DNA in the sample.
[0183] Let x and y be the frequencies of DNA molecules derived from the placenta and maternal blood cells at a specific degree of methylation, respectively, and f be the fractional concentration of fetal DNA in the sample.
[0184] The probability (P) that a DNA molecule originates from a fetus can be calculated as follows:
[0185]
[0186] Based on the previous example, consider a plasma DNA molecule of 16kb with a methylation level of 27.1%.
[0187] Figure 14A and Figure 14B The results show that based on the degree of methylation, at 16kb ( Figure 14A ) and 24kb ( Figure 14B Frequency distribution of DNA from the placenta (red) and DNA from maternal blood cells (blue) within the window. Frequency is shown on the y-axis. Methylation level is shown on the x-axis. Frequency distribution based on 16kb fragments ( Figure 14A The frequencies of DNA molecules originating from the placenta and maternal blood cells were 0.6% and 0.08%, respectively. When the fetal DNA fraction was 21.8%, the probability that this DNA fragment originated from the placenta was 64%, indicating an increased likelihood of placental origin.
[0188] The probability that plasma DNA molecules with 24kb and 66.9% methylation levels originated from fetal tissue can also be calculated. Based on the frequency distribution of the 24kb fragment, the frequencies of DNA molecules originating from the placenta and maternal blood cells were 0.05% and 0.16%, respectively. Figure 14BThe probability that this DNA fragment originated from the placenta is 0.8%, indicating that it is extremely unlikely to be of placental origin. In other words, there is a high probability that the molecule is of maternal origin.
[0189] This calculation can be further considered by taking into account the size of DNA molecules by referring to the size distribution curves of fetal DNA and maternal DNA. The analysis can be performed, for example but not limited to, using Bayes' theorem, logistic regression, multiple regression and support vector machines, random forest analysis, classification and regression trees (CART), and the K-nearest neighbor algorithm.
[0190] Figure 15A and Figure 15B The region shown is on chromosome 8 (chr8:108694010-108712904) of the human reference. Figure 15A ) compared and carried 7 maternally specific alleles ( Figure 15B The length of the long DNA fragment in the plasma was 18,896 bp. All of those maternally specific alleles were consistent with the allele information inferred from deep sequencing (Illumina technology) of the maternal and fetal genomes. Figure 15B According to the method described in this disclosure, the degree of methylation was determined to be 72.6% (). Figure 15B The overall methylation level (72.7%) was comparable to that of maternally specific fragments. Therefore, such molecules are more likely to be classified as maternally derived fragments. The genetic and epigenetic information of free DNA molecules approximately 1000 kilobases in length cannot be readily identified using short-read sequencing methods such as bisulfite sequencing (Illumina).
[0191] The probability that this molecule originated from the placenta can be calculated using the method described above. Based on the frequency distribution of the 19kb fragment, the frequencies of DNA molecules originating from the placenta and maternal blood cells are 0.65% and 0.23%, respectively. The probability that this DNA fragment originated from the placenta is 43%, indicating an increased likelihood of maternal origin.
[0192] D. Applications of clinical haplotype analysis
[0193] In this embodiment, the ability to analyze both short and long DNA molecules in the plasma of pregnant women will allow us to perform relative haplotype dose (RHDO) analysis (Lo et al., Science Translational Medicine 2010; 2:61ra91; Hui et al., Clinical Chemistry 2017; 63:513-524) without requiring prior paternal, maternal, or fetal genotype information obtained from the tissue. This capability will be more cost-effective and clinically applicable than previously possible.
[0194] Figure 16This illustration depicts the principle of how we can perform RHDO analysis using cell-free DNA from pregnant women. Cell-free DNA is isolated from the pregnant woman and subjected to SMRT sequencing at stage 1605. The size, allele information, and methylation status of each molecule, comprising long and short DNA molecules, can be determined according to the methods described in this disclosure. At stage 1610, based on the size information, the sequenced molecules can be classified into two categories: long DNA molecules and short DNA molecules. Cutoff values used to determine the long and short DNA categories may include, but are not limited to, 150bp, 180bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 1kb, 1.1kb, 1.2kb, 1.3kb, etc. 1.4kb, 1.5kb, 1.6kb, 1.7kb, 1.8kb, 1.9kb, 2kb, 2.5kb, 3kb, 4kb, 5kb, 6kb, 7kb, 8kb, 9kb, 10kb, 15kb, 20kb, 30kb, 40kb, 50kb, 60kb, 70kb, 80kb, 90kb, 100kb, 200kb, 300kb, 400kb, 500kb, or 1Mb. In an embodiment, at stage 1615, allele information present in long DNA molecules can be used to construct maternal haplotypes, namely Hap I and Hap II. Short DNA molecules can be compared with maternal haplotypes based on allele information. Therefore, the number of cell-free DNA molecules (e.g., short DNA) derived from maternal Hap I and Hap II can be determined.
[0195] At stage 1620, haplotype imbalances can be analyzed. Imbalances can be molecular counts, molecular sizes, or molecular methylation states. At stage 1625, maternal inheritance in the fetus can be inferred. If the dose of Hap I in maternal plasma DNA is excessively present, the fetus is likely to inherit maternal Hap I. Otherwise, the fetus is likely to inherit maternal Hap II. Various statistical approaches, including but not limited to the successive probability ratio test (SPRT), binomial test, chi-squared test, Student's t-test, nonparametric tests (e.g., the Wilcoxon test), and hidden Markov models, will be used to determine which maternal haplotype is excessively present.
[0196] In this embodiment, in addition to counting analysis, the methylation and size of short DNA molecules were determined and assigned to maternal haplotypes. The methylation imbalance between the two haplotypes (i.e., Hap I and Hap II) can be used to determine the maternal haplotype inherited by the fetus. If the fetus has inherited Hap I, there are more segments carrying the Hap I allele in the maternal plasma than segments carrying the Hap II allele. Hypomethylation of fetal DNA segments will result in lower methylation levels of Hap I than Hap II. In other words, if Hap I methylation shows lower methylation levels than Hap II, the fetus is more likely to inherit maternal Hap I. Otherwise, the fetus is more likely to inherit maternal Hap II. In another embodiment, the probability of an individual segment originating from the fetus or mother can be calculated as described above. For all segments compared to Hap I, the aggregate probability of these segments originating from the fetus can be determined based on Bayes' theorem. Similarly, the aggregate probability of these segments originating from the fetus can be calculated for Hap II. Subsequently, the fetus's likelihood of inheriting Hap I or Hap II can be inferred based on the probabilities of the two sets of probabilities.
[0197] In this embodiment, the lengthening or shortening of the size between the two haplotypes (i.e., Hap I and Hap II) can be used to determine the maternal haplotype inherited by the fetus. If the fetus has inherited Hap I, there will be more segments carrying the Hap I allele in the maternal plasma compared to segments carrying the Hap II allele. The DNA segment derived from the fetus will be relatively shorter than the DNA segment derived from Hap II. In other words, if the molecule derived from Hap I contains more short DNA than Hap II, the fetus is more likely to inherit maternal Hap I. Otherwise, the fetus is more likely to inherit maternal Hap II.
[0198] In some embodiments, we can perform a combined analysis of counts, sizes, and methylation between maternal Hap I and Hap II to infer maternal genetics in the fetus. For example, we can use logistic regression to combine those three measures that include count, size, and methylation status.
[0199] In clinical practice, haplotype-based analyses of counts, sizes, and methylation status can allow for the determination of whether an unborn fetus has inherited a maternal haplotype associated with a genetic condition, such as, but not limited to, single-gene disorders including Fragile X syndrome, muscular dystrophy, Huntington's disease, or β-thalassemia. This disclosure separately describes the detection of conditions involving repetitive sequences in long cell-free reads of DNA.
[0200] E. Targeted sequencing of long free DNA molecules
[0201] The methods described in this disclosure can also be applied to the analysis of one or more selected long DNA fragments. In embodiments, the one or more long DNA fragments of interest may first be enriched by a hybridization method that allows DNA molecules from one or more regions of interest to hybridize with synthetic oligonucleotides having complementary sequences. In order to decode all size, genetic, and epigenetic information into one using the methods described in this disclosure, the target DNA molecule is preferably not subjected to PCR amplification prior to sequencing, because base modification information in the original DNA molecule will not be transferred to the PCR product.
[0202] Several methods have been developed for enriching these target regions without performing PCR amplification. In another embodiment, one or more target long DNA molecules can be enriched using the clustered regularly spaced short palindromic repeat (CRISPR)-CRISPR-associated protein 9 (Cas9) system (Stevens et al. PLOS One 2019; 14(4):e0215441; Watson et al. Lab Invest 2020; 100:135-146). Even though the CRISPR-Cas9-mediated cleavage alters the size of the original long DNA molecule, its genetic and epigenetic information is preserved and can be obtained using the methods described in this disclosure, including but not limited to base content, haplotype (i.e., phase) information, remutation, and base modifications (e.g., 4mC (N4-methylcytosine), 5hmC (5-hydroxymethylcytosine), 5fC (5-formylcytosine), 5caC (5-carboxycytosine), 1mA (N1-methyladenine), 3mA (N3-methyladenine), 7mA (N7-methyladenine), 3mC (N3-methylcytosine), 2mG (N2-methylguanine), 6mG (O6-methylguanine), 7mG (N7-methylguanine), 3mT (N3-methylthymidine), 4mT (O4-methylthymidine). Pyrimidine) and 8oxoG (8-oxoguanine). In this example, the ends of DNA molecules in the DNA sample are first dephosphorylated to make them less likely to bind directly to the sequencing adaptor. Subsequently, the Cas9 protein and guide RNA (crRNA) guide the long DNA molecule of interest to produce a double-stranded cut. Then, the long DNA molecule of interest, which has been double-stranded on both sides, is bound to the sequencing adaptor specified by the selected sequencing platform. In another example, DNA can be treated with an exonuclease to degrade DNA molecules that are not bound by the Cas9 protein (Stevens et al., PLOS ONE 2019; 14(4):e0215441). Since these methods do not involve PCR amplification, the raw DNA molecules with base modifications can be sequenced and the base modifications can be identified.
[0203] In embodiments, these methods can be used to design guide RNAs, such as long scattered nuclear element (LINE) repeat sequences, to target long DNA molecules sharing a large number of homologous sequences by referencing a reference genome such as the human reference genome (hg19). In one instance, such analysis can be used to analyze circulating cell-free DNA in maternal plasma to detect fetal aneuploidy (Kinde et al., PLOS ONE 2012; 7(7):e41162). In embodiments, deactivated or 'dead' Cas9 (dCas9) and its associated single guide RNA (sgRNA) can be used to enrich targeted long DNA molecules without cleaving double-stranded DNA molecules. For example, the 3' end of the sgRNA can be designed to carry an additional universal short sequence. We can use biotin-labeled single-stranded oligonucleotides complementary to said universal short sequence to capture those targeted long DNA molecules bound by dCas9. In another embodiment, we can use biotin-labeled dCas9 protein or sgRNA or both to facilitate enrichment.
[0204] In embodiments, we may perform size selection to enrich long DNA fragments using pathways including, but not limited to, chemical, physical, enzymatic, gel-based, and magnetic bead-based methods, or a combination of many other approaches, without limitation on one or more specific genomic regions of interest. In other embodiments, immunoprecipitation may be used to enrich DNA fragments with a specific methylation profile, such as by using anti-methylcytosine antibodies and methyl-binding proteins. The methylation profile of the bound or captured DNA can be determined using unmethylated sensing sequencing.
[0205] F. General Concepts of Fetal Genetic Analysis Based on Long Plasma DNA Molecules
[0206] Figure 17 The identification of genetic / epigenetic disorders in plasma DNA molecules containing information of maternal and fetal origin. Long plasma DNA molecules in a pregnant woman can be identified as of fetal or maternal origin based on the genetic and / or epigenetic profile of CpG sites in the whole or a portion of the molecule [i.e., region (a)]. Genetic information can be, but is not limited to, sequence information, single nucleotide polymorphisms, insertions, deletions, tandem repeats, satellite DNA, microsatellites, small satellites, inversions, etc. Epigenetic information can be the methylation status of one or more CpG sites and their relative order in the plasma DNA molecule. In another embodiment, epigenetic information can be modifications of any of A, C, G, or T. Long plasma DNA with tissue origin information can be used for non-invasive prenatal testing by determining the presence of genetic and / or epigenetic disorders in such long plasma DNA molecules [i.e., region (b)].
[0207] Figure 18The identification of abnormal fetal fragments is illustrated. As an example, the methylation pattern of region (a) of this disclosure identifies long DNA fragments of fetal origin. We can determine the likelihood of a fetus being affected by a genetic or epigenetic condition based on such fetal-origin molecules. Genetic conditions may involve single nucleotide variants, insertions, deletions, tandem repeats, satellite DNA, microsatellites, small satellites, inversions, etc. Examples of genetic conditions include, but are not limited to: β-thalassemia, α-thalassemia, sickle cell anemia, cystic fibrosis, X-linked genetic conditions (e.g., hemophilia, Duchenne muscular dystrophy), spinal muscular atrophy, congenital adrenal hyperplasia, etc. Epigenetic conditions may be, for example, abnormal DNA methylation levels of increased (i.e., hypermethylation) or decreased (hypomethylation). Examples of epigenetic disorders include, but are not limited to, fragile X syndrome, Angelman's syndrome, Prader-Willi syndrome, facial shoulder and arm muscular atrophy (FSHD), immunodeficiency, centromere instability and facial abnormalities (ICF) syndrome, etc. Genetic or epigenetic disorders can be found in region (b).
[0208] G. Improve sequencing accuracy
[0209] Sequencing accuracy can be improved using sequence reads of long, cell-free DNA fragments. Figure 11B Among the seven alleles in the long fetal-specific DNA molecule, one appears to be inconsistent between PacBio and Illumina sequencing.
[0210] Figure 19A-19G This diagram illustrates the error correction for cell-free DNA genotyping performed using PacBio sequencing. We visually observed... Figure 11B The alignment results for those 7 sites are shown. Column 1 indicates genomic coordinates; column 2 is the reference sequence. Column 3 and subsequent columns indicate the aligned subreads. For example, in... Figure 19A There are 8 subreads that cross the region. '.' indicates a match to the reference base in the Watson sequence. ',' indicates a match to the reference base in the Crick sequence. 'Letter' indicates a substitute allele. '*' indicates an insertion and / or deletion. We can see that... Figure 19F The major base at the inconsistency site shown is called 'T' in the common sequence. However, among the nine subreads at said site ( Figure 19F Of the nine subreads, only five (56% major allele fraction (MAF)) were identified as 'T', while the other subreads were identified as 'C'. The major allele fraction at this locus ( Figure 19FThe major allele fraction was lower than that at other loci. Figure 19A -E and Figure 19G (MAF range: 67%-89%). Therefore, if we set a strict criterion for determining the base composition of each site in the common sequence, for example, using at least 60% MAF, this error site will be excluded from downstream interpretation. On the other hand, such error sites happen to fall within homopolymers (i.e., a series of consecutive identical bases 'TTTTTTT'). In an example, we can set a criterion that causes variants within homopolymers to be flagged as QC failures and temporarily excluded from downstream analysis. In an example, we can apply different positioning quality and base quality to correct or filter low-quality bases or subreads to improve base composition analysis.
[0211] With further improvements in sequencing accuracy of nanopore sequencing, embodiments of the present invention can also be used with such improved sequencing platforms and thus produce improved accuracy.
[0212] H. Instance Methods
[0213] Long cell-free DNA fragments from biological samples containing cell-free DNA obtained from pregnant women can be sequenced. These long cell-free DNA fragments can be used to determine the haplotype inheritance in the fetus.
[0214] 1. Sequencing long cell-free DNA fragments
[0215] Figure 20 Methods for analyzing biological samples of pregnant organisms (2000). Biological samples may contain multiple cell-free nucleic acid molecules. Biological samples may be any biological sample described herein. More than 20% of the cell-free nucleic acid molecules in the biological sample are larger than 200 nt (nucleotides).
[0216] At block 2010, multiple free nucleic acid molecules are sequenced. Sequencing can be performed using single-molecule real-time technology. In some embodiments, sequencing can be performed using nanopores.
[0217] More than 20% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 200 nt. In some embodiments, 15%-20%, 20%-25%, 25%-30%, 30%-35%, or more than 35% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 200 nt.
[0218] In some embodiments, more than 11% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 400 nt. In embodiments, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 400 nt.
[0219] In some embodiments, more than 10% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 500 nt. In embodiments, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 500 nt.
[0220] In this embodiment, more than 8% of the sequenced cell-free nucleic acid molecules may have a length greater than 600 nt. In this embodiment, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the sequenced cell-free nucleic acid molecules may have a length greater than 600 nt.
[0221] In some embodiments, more than 6% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 1 knt. In embodiments, 3%-5%, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 1 knt.
[0222] In the examples, more than 3% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 2 knt. In the examples, 1%-5%, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 2 knt.
[0223] In the examples, more than 1% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 3 knt. In the examples, 1%-5%, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 3 knt.
[0224] In some embodiments, at least 0.9% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 4 knt. In embodiments, 0.5%-1%, 1%-5%, 5%-10%, 10%-15%, 15%-20%, or more than 20% of the sequenced multiple cell-free nucleic acid molecules may have a length greater than 4 knt.
[0225] In some embodiments, at least 0.04% of the sequenced plurality of cell-free nucleic acid molecules may have a length greater than 10 knt. In embodiments, 0.01% to 0.1%, 0.1% to 0.5%, 0.5%-1%, 1%-5%, 5%-10%, 10%-15%, or more than 15% of the sequenced plurality of cell-free nucleic acid molecules may have a length greater than 4 knt.
[0226] Multiple cell-free nucleic acid molecules may comprise at least 10, 50, 100, 150, or 200 cell-free nucleic acid molecules. These multiple cell-free nucleic acid molecules may originate from multiple different genomic regions. For example, multiple chromosome arms or chromosomes may be covered by cell-free nucleic acid molecules. At least two of the multiple cell-free nucleic acid molecules may correspond to non-overlapping regions.
[0227] Sequencing methods for long cell-free DNA fragments can be used by any of the methods described herein. Reads from sequencing can be used to determine fetal aneuploidy, aberrations (e.g., duplicate number aberrations), genetic mutations or variations, or parental haplotype inheritance. The number of sequence reads can represent the amount of cell-free DNA fragment.
[0228] 2. Haplotype inheritance
[0229] Figure 21 Method 2100 shows the analysis of biological samples obtained from women carrying fetuses. The women may have a first haplotype and a second haplotype in the first chromosomal region. The biological sample may contain multiple cell-free DNA molecules from the fetus and the woman. The biological sample may be any biological sample described herein.
[0230] At block 2105, reads corresponding to multiple cell-free DNA molecules can be received. These reads may be sequence reads. In some embodiments, the method may include performing sequencing.
[0231] At block 2110, the size of multiple free DNA molecules can be measured. Size can be measured by aligning one or more sequence reads corresponding to the ends of the DNA molecule with a reference genome. Size can also be measured by performing full-length sequencing on the DNA molecule and subsequently counting the number of nucleotides in the full-length sequence. The genomic coordinates at the outermost nucleotides can be used to determine the length of the DNA molecule.
[0232] At block 2115, the first group of free DNA molecules from multiple free DNA molecules can be identified as having a size greater than or equal to a cutoff value. The cutoff value can be any cutoff value associated with long DNA. For example, cutoff values can include 150bp, 180bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 1kb, 1.5kb, 2kb, 2.5kb, 3kb, 4kb, 5kb, 6kb, 7kb, 8kb, 9kb, 10kb, 15kb, 20kb, 30kb, 40kb, 50kb, 60kb, 70kb, 80kb, 90kb, 100kb, 200kb, 300kb, 400kb, 500kb, or 1Mb.
[0233] At block 2120, the sequences of the first haplotype and the second haplotype can be determined from the reads corresponding to the first group of cell-free DNA molecules. Determining the sequences of the first haplotype and the second haplotype may involve aligning the reads corresponding to the first group of cell-free DNA molecules with a reference genome.
[0234] In some embodiments, the sequences for determining the first haplotype and the second haplotype may not include a reference genome. Sequence determination may include aligning a first subgroup of the read with a second subgroup of the read to identify different alleles at loci within the read. The method may include determining that the first subgroup of the read has a first allele at a locus. The method may also include determining that the second subgroup of the read has a second allele at a locus. The method may further include determining that the first subgroup of the read corresponds to a first haplotype. Additionally, the method may include determining that the second subgroup of the read corresponds to a second haplotype. The alignment may be used with... Figure 16 The descriptions are similar.
[0235] At block 2125, a second group of cell-free DNA molecules from multiple cell-free DNA molecules can be aligned with the sequence of the first haplotype. The second group of cell-free DNA molecules may have a size smaller than the cutoff value. The second group of cell-free DNA molecules may be short DNA molecules of the first haplotype.
[0236] At block 2130, a third group of cell-free DNA molecules from multiple cell-free DNA molecules can be aligned with the sequence of the second haplotype. The third group of cell-free DNA molecules can have a size smaller than the cutoff value. The third group of cell-free DNA molecules can be short DNA molecules of the second haplotype.
[0237] At block 2135, the first value of the second set of cell-free DNA molecule measurement parameters can be used. The parameters can be the count of cell-free DNA molecules, the size profile of cell-free DNA molecules, or the degree of methylation of cell-free DNA molecules. Values can be raw values or statistical values (e.g., mean, median, mode, percentile, minimum, maximum). In some embodiments, the values can be normalized to the values of parameters from a reference sample, another region, two haplotypes, or other size ranges.
[0238] At block 2140, the second value of the measurement parameters for the third group of free DNA molecules can be used. These parameters are the same as those for the second group of free DNA molecules.
[0239] At block 2145, a first value can be compared with a second value. The comparison can use a separation value. A separation value can be calculated using the first and second values. The separation value can be compared with a cutoff value. The separation value can be any separation value described herein. The cutoff value can be determined from a reference sample of pregnant women carrying euploid fetuses. In other embodiments, the cutoff value can be determined from a reference sample of pregnant women carrying aneuploid fetuses. In some embodiments, assuming an aneuploid fetus, a cutoff value can be determined. For example, data from a reference sample of pregnant women carrying euploid fetuses can be adjusted to account for increases or decreases in the number of copies of aneuploid chromosomal regions. The cutoff value can be determined from the adjusted data.
[0240] At 2150, the probability of the fetus inheriting a first haplotype can be determined based on a comparison of a first value and a second value. This probability can be determined based on a comparison of a segregation value and a cutoff value. When the parameter is the size profile of cell-free DNA molecules, the method may include determining, when the first value is less than the second value, the probability that the fetus has a higher probability of inheriting a first haplotype than a second haplotype, indicating that the second group of cell-free DNA molecules is characterized by a smaller size profile than the third group of cell-free DNA molecules. When the parameter is the degree of methylation of cell-free DNA molecules, the method may include determining, when the first value is less than the second value, the probability that the fetus has a higher probability of inheriting a first haplotype than a second haplotype.
[0241] In some embodiments, the method may include identifying the number of repetitive sequences in a subsequence of a read corresponding to a first set of cell-free DNA molecules. Determining the sequence of a first haplotype may include determining the number of repetitive sequences in the subsequence of the determined sequence. The first haplotype may include repetitive sequence-associated diseases, which may be any repetitive sequence-associated disease described herein. The likelihood of a fetus inheriting a repetitive sequence-associated disease may be determined. The likelihood of a fetus inheriting a repetitive sequence-associated disease may be comparable to or similar to the likelihood of a fetus inheriting the first haplotype. The repetitive sequences for identifying the sequence are subsequently described in this disclosure, including using Figure 16 Describe it.
[0242] II. Analysis of the tissue of origin using methylation
[0243] Long cell-free DNA molecules can have several methylation sites. As discussed in this disclosure, the degree of methylation of long cell-free DNA molecules in a pregnant woman can be used to determine the tissue of origin. Furthermore, the methylation patterns present on long cell-free DNA molecules can be used to determine the tissue of origin.
[0244] Cells derived from placental tissue possess a unique methylomic pattern compared to leukocytes and cells from tissues such as, but not limited to, the liver, lungs, esophagus, heart, pancreas, colon, small intestine, adipose tissue, adrenal glands, and brain (Sun et al., Proceedings of the National Academy of Sciences, 2015; 112:E5503-12). The methylation profile of circulating fetal DNA in maternal blood can be similar to that in the placenta, thus offering the potential to explore methods for generating non-invasive, fetal-specific biomarkers independent of fetal sex or genotype. However, bisulfite sequencing of maternal plasma DNA from pregnant women (e.g., using the Illumina sequencing platform) may lack the ability to distinguish between fetal and maternal molecules due to several limitations: (1) plasma DNA can be degraded during bisulfite treatment, and long DNA molecules will typically break into shorter molecules; (2) DNA molecules larger than 500 bp may not be efficiently sequenced by the Illumina sequencing platform for downstream analysis (Tan et al., Sci Rep., 2019; 9:2856).
[0245] For methylation-based analyses of tissues of origin, we can focus on several differentially methylated regions (DMRs) and use aggregated methylation signals from multiple molecules associated with DMRs (Sun et al., *Proceedings of the National Academy of Sciences*, 2015; 112:E5503-12) instead of single-molecule methylation models. Several studies have attempted to assess the placenta's contribution to the plasma DNA pool using methylation-sensitive restriction enzyme-based approaches (Chan et al., *Clinical Chemistry*, 2006; 52:2211-8) or methylation-specific PCR-based approaches (Lo et al., *American Journal of Human Genetics*, 1998; 62:768-75). However, those studies are only suitable for analyzing one or a few markers and may be challenging for analyzing molecules at the genome-wide scale. These reads, however, are inferred from amplified signals (i.e., PCR-based amplification in the flow cell during DNA library preparation and bridging amplification during sequencing cluster generation). The amplification step could potentially induce a preference for short DNA molecules, leading to information loss associated with long DNA molecules. Furthermore, Li et al. only analyzed those reads associated with previously developed DMRs (Li et al., Nucleic Acids Res. 2018; 46:e89).
[0246] In this disclosure, we describe a novel approach for distinguishing fetal DNA molecules from maternal DNA molecules in maternal plasma based on the methylation patterns of individual DNA molecules without bisulfite treatment and DNA amplification. In embodiments, one or more long plasma DNA molecules will be used for analysis (e.g., using bioinformatics and / or size-selective experimental analysis). Long DNA molecules can be defined as DNA molecules with sizes of at least, but not limited to, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 100 kb, 200 kb, etc. Limited data exist regarding the presence and methylation status of longer free DNA molecules in maternal plasma. For example, it is unknown whether the methylation state of the longer cell-free DNA molecules will reflect the methylation state of the cellular DNA of the tissue of origin, for instance, because longer fragments have more sites where the methylation state may change after fragmentation in the body; such changes may occur when the fragment circulates in the plasma. For example, studies have shown that the methylation state of circulating DNA is related to the size of the DNA fragment (Lun et al., Clinical Chemistry, 2013; 59:1583-94). Therefore, the feasibility of inferring the tissue of origin from the longer cell-free DNA molecules is unknown. Consequently, the approaches taken to identify tissue-associated methylation markers and the methods used to determine and interpret the presence of the tissue-specific longer cell-free DNA molecules are substantially different from those used for the analysis of short cell-free DNA.
[0247] According to embodiments of this disclosure, we can identify short and long DNA molecules and determine their biological characteristics, including but not limited to methylation patterns, fragment ends, size, and base composition. Short DNA molecules can be defined as DNA molecules with sizes smaller than, but not limited to, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 200 bp, 300 bp, etc. Short DNA molecules can also be DNA molecules outside the range considered long. We describe a novel approach for inferring the tissue of origin of circulating DNA molecules in pregnant women's plasma. This novel approach utilizes methylation patterns on one or more long DNA molecules in plasma. The longer the DNA molecule, the more CpG sites it may contain. The presence of multiple CpG sites on plasma DNA molecules will provide information about the tissue of origin, even if the methylation state of any single CpG site may not provide information for determining the tissue of origin. The methylation pattern in the long DNA molecule may include the methylation state of each CpG site, the order of the methylation states, and the distance between any two CpG sites. The methylation state between two CpG sites may depend on the distance between the two CpG sites. When CpG sites (e.g., CpG islands) within a specific distance in a molecule exhibit tissue-specific patterns, statistical models can assign more weight to those signals during tissue-of-origin analysis.
[0248] Figure 22 This principle is illustrated schematically. Figure 22 The methylation patterns of DNA molecules are shown. Seven CpG sites and six plasma DNA fragments AE representing different tissues (placenta, liver, hematopoietic cells, colon) are shown. Methylated CpG sites are shown in red, and unmethylated CpG sites are shown in green. As an example, we consider seven CpG sites with various methylation states across placental, liver, hematopoietic, and colon tissues. We consider the scenario where a single CpG site does not exhibit a placenta-specific methylation state relative to other tissues. Therefore, the tissue of origin for those plasma DNA molecules A, B, C, D, and E with variable sizes can be determined not only based on the methylation state at a single CpG site. For plasma DNA molecules A and B, because the two molecules are relatively short, they contain only 3 and 4 CpG sites, respectively. In the examples, the methylation pattern in DNA molecules containing more than one CpG site can be defined as a methylation haplotype. Figure 22As shown, plasma DNA molecules A and B can be attributed to either the placenta or the liver based on their methylation haplotypes because the placenta and liver share the same methylation haplotypes at those genomic locations corresponding to molecule A (locations 1, 2, and 3) and those corresponding to molecule B (locations 1, 2, 3, and 4). However, when we can obtain long DNA molecules such as molecules C, D, and E from plasma, we can definitively determine which molecules C, D, and E originate from the placenta based on their methylation haplotypes.
[0249] The reference pattern of an organization may be based on the methylation pattern of a reference organization. In some embodiments, the methylation pattern may be based on several reads and / or samples. The degree of methylation at each CpG site (also known as the methylation index MI and described below) can be used to determine whether a site is methylated.
[0250] A. Statistical models for methylation patterns
[0251] In this embodiment, the likelihood that the plasma DNA molecule is derived from the placenta can be determined by comparing the methylation haplotype of a single DNA molecule with methylation patterns in multiple reference tissues. Long plasma DNA molecules may be advantageous for the analysis. Long DNA molecules can be defined as DNA molecules with sizes of at least, but not limited to, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 100 kb, 200 kb, etc. Reference tissues may include, but are not limited to, placenta, liver, lung, esophagus, heart, pancreas, colon, small intestine, adipose tissue, adrenal gland, brain, neutrophils, lymphocytes, basophils, eosinophils, etc. In this embodiment, we can determine the likelihood that plasma DNA molecules originated from the placenta by using synergistic analysis of plasma DNA methylation haplotypes determined by single-molecule real-time sequencing and methylome data from whole-genome bisulfite sequencing of reference tissues. As an example, placental and erythrocyte sedimentation rate (ESR) amber samples were sequenced to an average 94-fold and 75-fold genome coverage, respectively, using whole-genome bisulfite sequencing. The degree of methylation (also known as the methylation index MI) at each CpG site was calculated using the following formula based on the number of sequenced cytosine (i.e., methylated, denoted by C) and the number of sequenced thymine (i.e., unmethylated, denoted by T):
[0252]
[0253] CpG sites are classified into three categories based on MI values inferred from placental DNA:
[0254] 1. Category A: CpG sites with an MI value ≥ 70.
[0255] 2. Category B CpG sites with MI values between 30 and 70.
[0256] 3. CpG sites with a category CMI value ≤30.
[0257] Similarly, CpG sites were classified into three categories using the MI values at CpG sites inferred from the erythrocyte sedimentation rate (ESR) amber layer DNA:
[0258] 1. Category A: CpG sites with an MI value ≥ 70.
[0259] 2. Category B CpG sites with MI values between 30 and 70.
[0260] 3. CpG sites with a category CMI value ≤30.
[0261] The categories use MI cutoff values of 30 and 70. Cutoff values may include other values, including 10, 20, 40, 50, 60, 80, or 90. In some embodiments, these categories can be used to determine the reference methylation pattern of a reference tissue (e.g., using...). Figure 22 (Described for use). Category A sites can be considered methylated. Category C sites can be considered unmethylated. Category B sites can be considered non-informative and not included in the reference pattern.
[0262] For a plasma DNA molecule with n CpG sites, the methylation status of each CpG site is determined using the method described in our previous disclosure (U.S. Application No. 16 / 995,607). In some embodiments, the methylation status can be determined by bisulfite sequencing or nanopore sequencing. To determine the likelihood that a plasma DNA molecule is of placental or maternal origin, the methylation pattern of the molecule is analyzed in conjunction with prior methylation information from placental and maternal erythrocyte sedimentation rate (ESR) amber DNA. In embodiments, we utilize the following principle: if a CpG site in a plasma DNA fragment that is identified as methylated (M) is consistent with a higher methylation index in the placenta, such an observation indicates that the molecule is more likely to be of placental origin. If a CpG site in a plasma DNA molecule that is identified as methylated (M) is consistent with a lower methylation index in the placenta, such an observation indicates that the molecule is less likely to be of placental origin; if a CpG site in a plasma DNA molecule that is identified as unmethylated (U) is consistent with a lower methylation index in the placenta, such an observation indicates that the molecule is more likely to be of placental origin. If unmethylated (U) CpG sites in plasma DNA are consistent with a higher methylation index in the placenta, such observations would indicate that the molecule is unlikely to have originated from the placenta.
[0263] We implemented the following scoring procedure. An initial score (S) reflecting the probability of fetal origin of the plasma DNA fragment was set to 0. When comparing the methylation status of plasma DNA molecules with the methylation information of prior placental DNA,
[0264] a. If the CpG site on the plasma DNA molecule is identified as 'M' and its counterpart in the placenta belongs to category A, then add 1 point to S (i.e., add 1 to the score unit).
[0265] b. If the CpG site on the plasma DNA molecule is identified as 'U' and its counterpart in the placenta belongs to category A, then deduct 1 point from S (i.e., reduce the score unit by 1).
[0266] c. If the CpG site on the plasma DNA molecule is identified as 'M' and its counterpart in the placenta belongs to category B, then add 0.5 points to S.
[0267] d. If the CpG site on the plasma DNA molecule is identified as 'U' and its counterpart in the placenta belongs to category B, then add 0.5 points to S.
[0268] e. If the CpG site on the plasma DNA molecule is identified as 'M' and its counterpart in the placenta belongs to category C, then deduct 1 point from S.
[0269] f. If the CpG site on the plasma DNA molecule is identified as 'U' and its counterpart in the placenta belongs to category C, then add 1 point to S.
[0270] We call the above process 'methylation state matching'.
[0271] After all CpG sites in the plasma DNA molecule have been processed, the final collection fraction S (placenta) of the plasma DNA molecule is obtained. In the examples, the number of CpG sites is required to be at least 30 and the length of the plasma DNA molecule is required to be at least 3 kb. Other numbers and lengths of CpG sites may be used, including but not limited to any number and length described herein.
[0272] A similar scoring procedure is applied when comparing the methylation status of plasma DNA molecules with the methylation level of erythrocyte sedimentation rate (ESR) brown-yellow layer DNA at the corresponding sites. After all CpG sites in the plasma DNA molecules have been treated, the final aggregate score S (ESR brown-yellow layer) of the plasma DNA molecules is obtained.
[0273] If S(placental) > S(erythrocyte sedimentation rate, brownish-yellow layer), then the plasma DNA molecules are determined to be of fetal origin; otherwise, the plasma DNA molecules are determined to be of maternal origin.
[0274] Seventeen fetal-specific DNA molecules and 405 maternal-specific DNA molecules were identified as potential candidates for evaluating the efficacy of plasma DNA molecules in inferring fetal-maternal origin. Fetal-specific molecules are plasma DNA molecules carrying fetal-specific SNP alleles, while maternal-specific DNA molecules are plasma DNA molecules carrying maternal-specific SNP alleles.
[0275] Figure 23 The receiver operating characteristic (ROC) curves used to determine fetal and maternal origin are displayed. The y-axis shows sensitivity, and the x-axis shows specificity. The red line represents the effectiveness of distinguishing fetal and maternal origin molecules using the methylation state matching method present in this disclosure. The blue line represents the effectiveness of distinguishing fetal and maternal origin molecules using the degree of single-molecule methylation (i.e., the proportion of CpG sites in a DNA molecule that have been identified as methylated). Figure 23 The area under the receiver operating characteristic (AUC) for the methylation state matching process (0.94) was significantly higher than the AUC based on single-molecule methylation level (0.86) (P < 0.0001; DeLong test). This suggests that methylation pattern analysis of long DNA molecules could be used to determine fetal / maternal origin.
[0276] In this embodiment, when determining whether plasma DNA is of fetal or maternal origin, the difference (ΔS) between S (placental) and S (erythrocyte sedimentation rate, brown layer) can be considered. The absolute value of ΔS can be required to exceed a specific threshold, such as, but not limited to, 5, 10, 20, 30, 40, 50, etc. For example, when we use 10 as the threshold for ΔS, the positive predictive value (PPV) in fetal DNA molecular detection increases from 14.95% to 91.67%.
[0277] In this embodiment, the methylation state of a CpG site is influenced by the methylation state of its neighboring CpG sites. The closer the nucleotide distance between any two CpG sites on a DNA molecule, the more likely the two CpG sites will share the same methylation state. This phenomenon is known as comethylation. Multiple tissue-specific CpG island methylations have been reported; therefore, in some statistical models used for tissue-of-origin analysis, more weight is assigned to dense clusters of CpG sites (e.g., CpG islands) that share the same methylation state. For scenarios 'a' and 'f', if the currently explored CpG site is located within a genomic distance of no more than 100 bp relative to the previous CpG site and the methylation state matching process for these two consecutive CpG sites yields the same result, an additional 1 point is added to the score S of the currently explored CpG site. For scenarios 'b' and 'e', if the currently explored CpG site is located within a genomic distance of no more than 100 bp relative to the previous CpG site and the methylation status matching results of these two consecutive CpG sites are the same, then an additional 1 point will be deducted from the score S of the currently explored CpG site. However, if the currently explored CpG site is located within a genomic distance of no more than 100 bp relative to the previous CpG site, but the methylation status matching results of these two consecutive CpG sites are inconsistent, then the aforementioned preset scoring procedure will be used. On the other hand, if the currently explored CpG site is located within a genomic distance greater than 100 bp relative to the previous CpG site, then the aforementioned scoring procedure with preset parameters will be used. Scores other than 1 and distances other than 100 bp can be used, including any scores and distances described herein.
[0278] In other embodiments, CpG sites are classified into more than three categories based on MI values inferred from placental and erythrocyte sedimentation rate (ESR) amber layer DNA. Methylation information from previous reference tissues can be inferred from single-molecule real-time sequencing (i.e., nanopore sequencing and / or PacBioSMRT sequencing). Plasma DNA molecule lengths may be required to be at least, but not limited to, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 100 kb, 200 kb, etc. The number of CpG sites may be required to be at least, but not limited to, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, etc.
[0279] In this embodiment, we can use a probabilistic model to characterize the methylation patterns of plasma DNA molecules. The methylation status of k CpG sites (k≥1) on a plasma DNA molecule is represented as M=(m1, m2, ..., m k), where m at the CpG site i on the plasma DNA molecule i The value is 0 (for the unmethylated state) or 1 (for the methylated state). In the examples, the M probability associated with placental plasma DNA molecules is determined by a reference methylation pattern in the placental tissue. The reference methylation pattern in the placental tissue for those corresponding to 1, 2, ..., k CpG sites will follow a beta (β) distribution. The beta (β) distribution is parameterized by two positive parameters α and β, denoted as beta (α, β). The values derived from the beta (β) distribution will range from 0 to 1. Based on deep bisulfite sequencing data of the tissue of interest, parameters α and β are determined by the number of sequenced cytosine (methylated) and thymine (unmethylated) at each CpG site in the specific tissue, respectively. For the placenta, such a beta (β) distribution is denoted as beta (α... P β p The probability P(M|placental) of plasma DNA molecules derived from the placenta will be modeled as follows:
[0280]
[0281] Where 'i' represents the i-th CpG site; beta The beta (β) distribution indicates the methylation pattern associated with the i-th CpG site in the placenta; P is the probability of the observed plasma DNA molecules having a given methylation pattern at k CpG sites.
[0282] The probability P(M|erythrocyte sedimentation rate brown layer) of plasma DNA molecules originating from the erythrocyte sedimentation rate brown layer (i.e., white blood cells) will be modeled as follows:
[0283]
[0284] Where 'i' represents the i-th CpG site; beta The beta (β) distribution indicates the methylation pattern associated with the i-th CpG site in the erythrocyte sedimentation rate (ESR) brown-yellow layer DNA. P is the probability of co-occurrence of observed plasma DNA molecules with a given methylation pattern at k CpG sites.
[0285] Beta and Beta It can be determined from whole-genome bisulfite sequencing results of placental and erythrocyte sedimentation rate (ESR) brown-yellow layer DNA, respectively.
[0286] For plasma DNA molecules, if we observe P(M|placental) > P(M|erythrocyte sedimentation rate brown layer), then such plasma DNA molecules are likely derived from the placenta; otherwise, they are likely derived from the erythrocyte sedimentation rate brown layer. We used this model to achieve an AUC of 0.79.
[0287] B. Machine Learning Models
[0288] In other embodiments, we can use machine learning algorithms to determine the fetal / maternal origin of specific plasma DNA molecules. To test the feasibility of using a machine learning-based approach to classify fetal and maternal DNA molecules in pregnant women, we developed a methylation pattern mapping method for plasma DNA molecules.
[0289] Figure 24 The diagram illustrates the definition of paired methylation patterns. Nine CpG sites are displayed on a plasma DNA molecule. Methylated CpG sites are shown in red, and unmethylated CpG sites are shown in green. When two CpG sites in a pair share the same methylation state (e.g., CpG 1 and CpG 5), the pair is encoded as 1, as shown by the arrow at position 'a'. When two CpG sites in a pair have different methylation states (e.g., CpG 1 and CpG 2), the pair is encoded as 0, as shown by the arrow at position 'b'. The encoding rule also applies to all pairs on the DNA molecule that have any two CpG sites.
[0290] We use a plasma DNA molecule containing 9 CpG sites as an example. The methylation pattern of this plasma DNA molecule is determined by the pathway described in our previous publication (U.S. Application No. 16 / 995,607), namely UMMMUUUMM (where U and M represent unmethylated CpG and methylated CpG, respectively). Pairwise comparisons of the methylation state between any two CpG sites are applicable to machine learning or deep learning-based analyses. In this example, the rule also applies to a total of 36 pairs. If there are a total of n CpG sites on the plasma DNA molecule, there will be n×(n-1) / 2 comparison pairs. Different numbers of CpG sites can be used, including 5, 6, 7, 8, 10, 11, 12, 13, etc. If the molecule contains more sites than the number used in the machine learning model, a sliding window can be used to divide the sites into an appropriate number of sites.
[0291] We obtained one or more molecules from placental and erythrocyte sedimentation rate (ESR) amber layer DNA samples, respectively. The methylation patterns of those DNA molecules were determined by Pacific Bioscience (PacBio) single-molecule real-time (SMRT) sequencing according to the pathway described in our previous publication (U.S. Application No. 16 / 995,607). Those methylation patterns were then translated into paired methylation patterns.
[0292] Paired methylation patterns associated with placental DNA and those associated with ESR (erythrocyte sedimentation rate) tannin DNA were used to train a convolutional neural network (CNN) to distinguish molecules potentially of fetal or maternal origin. Target outputs (i.e., values similar to dependent variables) for DNA fragments from the placenta were assigned '1', while those for DNA fragments from the ESR tannin were assigned '0'. The paired methylation patterns were used for training to determine the parameters (often called weights) for the CNN model. The optimal parameters for the CNN distinguishing fetal-maternal origin of DNA fragments were obtained when the total prediction error between the output scores calculated using the sigmoid function and the desired target output (binary values: 0 or 1) was minimized through iterative tuning of the model parameters. The total prediction error was measured using the sigmoid cross-entropy loss function in deep learning algorithms (https: / / keras.io / ). The model parameters learned from the training dataset were used to analyze DNA molecules (such as plasma DNA molecules) to output probabilistic scores indicating the likelihood that a DNA molecule originated from the placenta or ESR tannin. If the probabilistic score of a plasma DNA fragment exceeds a certain threshold, then such plasma DNA molecules are considered to be of fetal origin. Otherwise, they are considered to be of maternal origin. The thresholds will include, but are not limited to, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.99, etc. In one instance, we used this CNN model to achieve an AUC of 0.63 to determine whether a plasma DNA molecule is of fetal or maternal origin, indicating that it is possible to use a deep learning algorithm to infer the tissue of origin of DNA molecules from maternal plasma. The performance of the deep learning algorithm will be further improved by obtaining more single-molecule real-time sequencing results.
[0293] In some other embodiments, the statistical model may include, but is not limited to, linear regression, logistic regression, deep recurrent neural networks (e.g., Long Short-Term Memory (LSTM)), Bayes' classifier, Hidden Markov Model (HMM), Linear Discriminant Analysis (LDA), k-means clustering, density-based spatial clustering (DBSCAN) for noisy applications, random forest algorithms, and support vector machines (SVM). Different statistical distributions will be involved, including but not limited to binomial distribution, Bernoulli distribution, gamma distribution, normal distribution, Poisson distribution, etc.
[0294] C. Placental-specific methylated haplotypes
[0295] The methylation state of each CpG site on a single DNA molecule can be determined using the pathways described in our previous disclosure (U.S. Application No. 16 / 995,607) or any of the techniques described herein. In addition to the degree of single-molecule double-stranded DNA methylation, we can determine the single-molecule methylation pattern of each DNA molecule, which can be a sequence of methylation states along adjacent CpG sites of a single DNA molecule.
[0296] Different DNA methylation markers can be found in different tissues and cell types. In the examples, we can infer the tissue of origin based on the single-molecule methylation pattern of individual plasma DNA molecules.
[0297] Genomic DNA from ten erythrocyte sedimentation rate (ESR) tannin samples and six placental tissue samples was sequenced using SMRT sequencing (PacBio). We were able to achieve 58.7-fold and 28.7-fold coverage of ESR tannin DNA and placental DNA, respectively, by combining localized high-quality circular co-sequencing (CCS) reads from each sample type.
[0298] The genome was divided into approximately 28.2 million overlapping windows of 5 CpG sites using a sliding window approach. In other embodiments, different window sizes of 2, 3, 4, 5, 6, 7, and 8 CpG sites can be used, but are not limited to. A non-overlapping window approach can also be used. Each window is considered a potential marker region. For each potential marker region, we identified the major single-molecule methylation pattern in all sequenced placental DNA molecules covering all 5 CpG sites within the marker region. A comparison was made between the CpG sites of plasma DNA molecules and the corresponding CpG sites of individual DNA molecules from a reference tissue. Subsequently, we calculated the mismatch fraction of each erythrocyte sedimentation rate (ESR) amber layer DNA molecule covering all CpG sites within the same marker region by comparing its single-molecule methylation pattern with the major single-molecule methylation pattern in the placenta.
[0299]
[0300] The number of mismatched CpG sites refers to the number of CpG sites in the ESR (erythrocyte sedimentation rate) brown-yellow layer DNA molecules that show a different methylation state compared to the main monomolecular methylation pattern in the placenta.
[0301] A higher mismatch score indicates that the methylation pattern of DNA molecules in the ESR (erythrocyte sedimentation rate) tandem layer is more distinct from the predominant monomolecular methylation pattern in the placenta. We selected marker regions from 28.2 million potential marker regions that demonstrated substantial differences in monomolecular methylation patterns between the pool of DNA molecules from the placenta and the ESR tandem layer using the following criteria: a) more than 50% of placental DNA molecules had a predominant monomolecular methylation pattern; and b) more than 80% of ESR tandem layer DNA molecules had a mismatch score greater than 0.3. Based on these criteria, we selected 281,566 marker regions for downstream analysis.
[0302] Figure 25 This is a table showing the distribution of selected marker regions on different chromosomes. The first column displays the number of chromosomes. The second column displays the number of marker regions on each chromosome.
[0303] We hereby clarify our concept of tissue-of-origin classification for individual plasma DNA molecules based on single-molecule methylation patterns, using SMRT sequencing to cover fetal-specific or maternal-specific alleles as previously described in this disclosure. Any plasma DNA molecule covering a selected marker region with a methylation pattern identical to the predominant single-molecule methylation pattern in the placenta will be classified as placenta-specific (i.e., fetal-specific) DNA molecule. Conversely, if the single-molecule methylation pattern of a plasma DNA molecule differs from the predominant single-molecule methylation pattern in the placenta, we classify the molecule as non-placenta-specific. Correct classification in this analysis is defined by the presence of a placenta-specific methylation haplotype in the molecule, in a manner that fetal-specific DNA molecules are identified as fetal-origin (i.e., placenta-specific) and maternal DNA molecules are identified as non-fetal-origin (i.e., placenta-non-specific). Previous methylation-based methods for tissue-of-origin analysis typically involve deconvolution of a range of tissue contributors to cell-free DNA within a biological sample. The advantage of the method of the present invention over previous methods is that evidence of tissue contribution to cell-free DNA (e.g., placental-derived DNA in maternal plasma) in a biological sample can be determined without considering the presence or absence of contributions from other tissues. Furthermore, the placental origin of any cell-free DNA molecule can be determined using the method of the present invention without considering the contribution fraction of said tissue to the cell-free DNA molecule.
[0304] Of the 28 DNA molecules containing fetal-specific alleles, 17 (61%) were classified as placenta-specific, and 11 (39%) as non-placenta-specific. On the other hand, of the 467 DNA molecules containing maternal-specific alleles, 433 (93%) were classified as non-placenta-specific, and 34 (7%) as placenta-specific.
[0305] In embodiments, we can use different percentages of ESR (erythrocyte sedimentation rate) amber layer DNA molecules with mismatch scores greater than 0.3 as thresholds, including but not limited to greater than 60%, 70%, 75%, 80%, 85%, and 90%. We can improve the overall classification accuracy of placental or non-placental origin of plasma DNA in pregnant individuals by adjusting the criteria used in marker region selection. This is particularly crucial in non-invasive prenatal testing settings when attempting to determine the presence of pathogenic mutations or duplicate number aberrations in the fetus.
[0306] Figure 26 A classification table of plasma DNA molecules based on single-molecule methylation patterns, using the percentage of ESR-affected brown-yellow layer DNA molecules with mismatch fractions greater than 0.3 as the selection criterion for marker regions. The first column shows the percentage of ESR-affected brown-yellow layer DNA molecules with mismatch fractions greater than 0.3%. The second column divides DNA molecules into those encompassing fetal-specific alleles and those encompassing maternal-specific alleles. The third and fourth columns show the classification of DNA molecules based on single-molecule methylation patterns as placental-specific or non-placental-specific. The fifth column shows the percentage of DNA molecules classified with the same specific alleles as those in the second column.
[0307] Figure 27 This demonstrates a methodological workflow for determining fetal genetics using placental-specific methylation haplotypes in a non-invasive manner. For example... Figure 27 As shown, cell-free DNA from the pregnant woman's plasma is extracted for single-molecule real-time sequencing. Long plasma DNA molecules are identified according to embodiments of this disclosure. The methylation status at each CpG site of each long plasma DNA molecule is determined according to embodiments of this disclosure. The methylation haplotype of each long plasma DNA molecule is determined according to embodiments of this disclosure. If a long plasma DNA molecule is identified as carrying a placenta-specific methylation haplotype, the genetic and epigenetic information associated with said molecule is considered to be inherited by the fetus. In embodiments, if, based on methylation haplotype information, one or more long plasma DNA molecules containing the same pathogenic mutation as the pathogenic mutation carried by the pregnant woman are determined to be of fetal origin according to embodiments of this disclosure, it indicates that the fetus has inherited the mutation from the mother.
[0308] The examples are applicable to genetic diseases including, but not limited to, the following: β-thalassemia, sickle cell anemia, α-thalassemia, cystic fibrosis, hemophilia A, hemophilia B, congenital adrenal hyperplasia, Duchenne muscular dystrophy, Becker muscular dystrophy, achondroplasia, lethal dysplasia, von Willebrand disease, Noonan syndrome, hereditary hearing loss and deafness, various congenital metabolic defects (e.g., type I citrullinemia, propionic acidemia, type Ia glycogen storage disease (von Gierke disease), type Ib / c glycogen storage disease (von Gierke disease), type II glycogen storage disease (Pompe disease)). Mucopolysaccharide storage disease (MPS) type I (Hurler / Hurler-Scheie / Scheie), type II (Hunter syndrome), type IIIA (Sanfilippo syndrome A), type IIIB (Sanfilippo syndrome B), type IIIC (Sanfilippo syndrome C), type IIID (Sanfilippo syndrome D), type IVA (Morquio syndrome A), type IVB (Morquio syndrome B), type VI (Maroteaux-Lamy syndrome), type VII (Sly syndrome). Syndrome, mucolipidemia II (I-cell disease), metachromatic leukotrophic disorder, GM1 ganglioside storage disease, OTC deficiency (X-linked ornithine carbamoyltransferase deficiency), adrenoleukotrophic disorder (X-linked ALD), Krabbe disease (globose cell leukotrophic disorder), etc.
[0309] In other embodiments, fetal genetic disorders may be associated with remethylation of DNA in the fetal genome that is not present in the parental genome. An example would be hypermethylation of the FMRP translation regulator 1 (FMR1) gene in a fetus with fragile X syndrome. Fragile X syndrome is caused by the expansion of the CGG trinucleotide repeat sequence in the 5' untranslated region of the FMR1 gene. A normal allele would contain approximately 5 to 44 copies of the CGG repeat sequence. A quasi-mutant allele would contain 55 to 200 copies of the CGG repeat sequence. A fully mutant allele would contain more than 200 copies of the CGG repeat sequence.
[0310] Figure 28This diagram illustrates the principle of non-invasive prenatal testing for fragile X chromosome in male fetuses of unaffected pregnant women carrying normal or quasi-mutant alleles. Figure 28 In this context, 'n' represents the number of CGG copies in the maternal genome; 'm' represents the number of CGG copies in the fetal genome. The genome of an unaffected pregnant woman will contain the FMR1 gene, which has no more than 200 copies of the CGG repeat sequence (i.e., n ≤ 200) and is unmethylated. In contrast, the genome of a male fetus affected by fragile X will contain a methylated FMR1 gene with more than 200 copies of the CGG repeat sequence (m > 200). We can identify multiple long DNA molecules from the genomic region of interest (e.g., the FMR1 gene) from which the number of repeat sequences and methylation status can be simultaneously determined by performing single-molecule sequencing of maternal plasma DNA. If we identify one or more methylated DNA molecules in the plasma of an unaffected woman that encompass the FMR1 gene, contain more than 200 copies of the CGG repeat sequence, this indicates that the fetus may have fragile X. In yet another embodiment, we can further determine the fetal origin of the plasma DNA molecule using placental-specific methylation haplotypes according to embodiments of this disclosure. If we identify one or more molecules containing one or more regions of a haplotype carrying a placenta-specific methylation, and these molecules encompass the FMR1 gene, contain more than 200 CGG repeat sequences, and are methylated, we can more confidently conclude that the fetus has fragile X syndrome. Conversely, if we identify one or more molecules with a placenta-specific methylation haplotype, and these molecules encompass the FMR1 gene, contain fewer than 200 CGG repeat sequences, and are not methylated, it indicates that the fetus is likely unaffected. In the case of fragile X syndrome, a complete mutation (>200 repeat sequences) effectively methylates the entire gene and shuts it down. Therefore, especially for fragile X syndrome, the detection of long alleles (rather than those showing a profile of placental methylation) will be highly indicative of the fetus having the disease.
[0311] Detection of genetic disorders can be performed with or without prior knowledge of the mother's condition. Women with quasi-mutations may not have any symptoms, but some may have mild symptoms, often only becoming aware of them afterward. If the mother's mutation status is unknown, one approach is to detect long alleles in the plasma of women who do not appear to have the disease, or to analyze the maternal erythrocyte sedimentation rate (ESR) tannin layer and determine if it does not show such long alleles. Another approach is to combine the repetitive sequence length of the cfDNA molecule with its methylation status. If the methylation status indicates a fetal pattern (methylated haplotype) and shows long alleles, the fetus may be affected. This approach is applicable to many trinucleotide disorders, such as Huntington's disease.
[0312] D. Non-invasive construction of the fetal genome using long plasma DNA molecules.
[0313] Methylation patterns can be used to determine haplotype inheritance. Qualitative approaches using methylation patterns to determine haplotype inheritance are more efficient than quantitative methods that characterize the amount of a specific fragment. Methylation patterns can be used to determine maternal and paternal inheritance of haplotypes.
[0314] 1. Maternal inheritance in the fetus
[0315] Lo et al. demonstrated the feasibility of constructing whole-genome genetic maps and determining fetal mutation status from maternal plasma DNA sequences using parental haplotype information (Lo et al., *Science Translational Medicine*, 2010; 2:61ra91). This technique, known as relative haplotype dose (RHDO) analysis, is a pathway for analyzing maternal inheritance in the fetus. The principle is based on the fact that the maternal haplotype inherited by the fetus will be relatively over-presented in the maternal plasma DNA when compared to another maternal haplotype that was not transmitted to the fetus. Therefore, RHDO is a quantitative analytical method.
[0316] Embodiments present in this disclosure utilize methylation patterns in long plasma DNA molecules to determine the tissue of origin of said plasma DNA molecules. In one embodiment, the disclosure herein would allow for qualitative analysis of maternal genetics in a fetus.
[0317] Figure 29 This illustrates an example of determining maternal inheritance in the fetus. Genomic location P in the maternal genome (A / G) is heterozygous. Filled circles indicate methylated sites, and hollow circles indicate unmethylated sites. The methylation pattern in the placenta is “-MUMM-”, where “M” represents methylated cytosine at a CpG site and “U” represents unmethylated cytosine at a CpG site. In one embodiment, methylation patterns in the placenta and relevant reference tissues may be obtained from data previously generated by sequencing (e.g., single-molecule real-time sequencing and / or bisulfite sequencing). In plasma DNA, a non-paternal plasma DNA (represented by Z) carrying allele A at the specific genomic locus was found to exhibit a methylation pattern (“-MUMM-”) compatible with the methylation pattern in the placenta compared to other tissues. No molecule carrying allele G exhibiting a methylation pattern compatible with the methylation pattern in the placenta was found. Therefore, the presence of allele A and the “-MUMM-” methylation pattern allows for the determination of maternal allele A in the fetus.
[0318] Figure 30 This demonstrates a qualitative analysis of maternal inheritance in the fetus using genetic and epigenetic information from plasma DNA molecules. For example... Figure 30 As shown in the top branch, according to an embodiment of this disclosure, plasma DNA is extracted and then size-selected for long DNA. The size-selected plasma DNA molecules undergo single-molecule real-time sequencing (e.g., using a system manufactured by Pacific Biosciences). Genetic and epigenetic information is determined according to an embodiment of this disclosure. For illustrative purposes, molecule (X) is compared to human chromosome 1, which contains allele G at chromosome position a (chr1:a) and allele A at chromosome position e (chr1:e). Molecule X has allele C at chromosome position d.
[0319] The CpG methylation state of molecule X was determined to be "-MUMM-", where "M" represents methylated cytosine at the CpG site and "U" represents unmethylated cytosine at the CpG site. Filled circles indicate methylated sites, and hollow circles indicate unmethylated sites. As a result of analysis of a reference sample, placental DNA is known to have a "-MUMM-" methylation pattern in the region between positions a and e. According to embodiments of this disclosure, based on the methylation pattern of molecule X that matches the methylation pattern of placental DNA, molecule X was determined to be of placental origin.
[0320] like Figure 30 As shown in the lower branch, DNA from maternal leukemia cells undergoes single-molecule real-time sequencing. Epigenetic and genetic information of the maternal leukemia cells is obtained according to embodiments of this disclosure. Genetic alleles are phased into two haplotypes, namely maternal haplotype I (Hap I) and maternal haplotype II (Hap II), using methods including but not limited to WhatsHap (Patterson et al., *J. Comput. Biol.* 2015; 22:498-509), HapCUT (Bansal et al., *Bioinformatics.* 2008; 24:153-9), and HapCHAT (Beretta et al., *BMC. Bioinformatics.* 2018; 19:252). Here, we obtain two haplotypes from the maternal genome, namely "-ACGT-" (Hap I) and "-GTAC-" (Hap II). Hap I is associated with one or more wild-type variants, while Hap II is associated with one or more disease-associated variants. One or more disease-associated variants may include, but are not limited to, single nucleotide variants, insertions, deletions, translocations, inversions, repetitive sequence extensions, and / or other genetic structural variations.
[0321] For genomic location e, the maternal genotype was determined to be AA and the paternal genotype to be GG. Due to the methylation pattern, plasma DNA molecule X was determined to be of placental origin. Because of the presence of the maternal-specific allele A but the absence of the paternal-specific allele G, molecule X was inferred to have been inherited from one of the maternal haplotypes.
[0322] To further determine which maternal haplotype was transmitted to the fetus, we compared the allelic information of this placental molecule X at genomic positions other than chr1:e with the maternal haplotype. As an example, molecule X has allele G at position a and allele C at position d. The presence of either of these alleles in molecule X indicates that molecule X should be assigned to maternal Hap II, which contains the same alleles.
[0323] Therefore, we can conclude that maternal haplotype II associated with one or more disease-related variants is transmitted to the fetus. The unborn fetus is identified as being at risk of being affected by the disease.
[0324] Compared to RHDO, a quantitative analysis-based approach, qualitative analysis of maternal inheritance in the fetus based on methylation patterns may require fewer plasma DNA molecules to determine which maternal haplotype the fetus inherits. We performed computer simulation analyses on a genome-wide basis with varying numbers of plasma DNA molecules for analysis to assess the detection rate of maternal inheritance in the fetus.
[0325] For the RHDO simulation analysis, N plasma DNA molecules from the haplotype blocks of the maternal genome were collectively aligned with M heterozygous SNPs. The fetal DNA fraction was f. Those corresponding to SNPs had paternal homozygous genotypes identical to the maternal Hap I genotype transmitted to the fetus. Among the N plasma DNA molecules, the mean value of plasma DNA molecules aligned with maternal Hap I was N×(0.5+f / 2), while the mean value of plasma DNA molecules aligned with maternal Hap II was N×(0.5-f / 2). We assumed that the plasma DNA molecules sampled from the haplotypes followed a binomial distribution.
[0326] The number of plasma DNA molecules are assigned to Hap I (i.e., X) and follow the following distribution:
[0327] X~Bin(N,0.5+f / 2) (1),
[0328] "Bin" represents a binomial distribution.
[0329] The number of plasma DNA molecules are assigned to Hap II (i.e., Y) and follow the following distribution:
[0330] Y~Bin(N,0.5-f / 2) (2).
[0331] Therefore, plasma DNA molecules assigned to maternal Hap I will be relatively overpresented in maternal plasma compared to maternal Hap II. To determine whether overpresentation was statistically significant, we compared the difference in plasma DNA counts between two maternal haplotypes under the null hypothesis that both haplotypes (denoted by X' and Y') were also overpresented in plasma.
[0332] X'~Bin(N,0.5) (3),
[0333] Y'~Bin(N,0.5) (4).
[0334] We further define the relative dose difference between two haplotypes as follows:
[0335] D = (XY) / N (5),
[0336] D'=(X'-Y') / N (6).
[0337] In one instance, the statistical value D, reflecting the relative haplotype dose, is compared with the mean (M) of D', and standardized to the following (i.e., z-score) using the standard deviation (SD) of D':
[0338] z-score = (D–M) / SD (7).
[0339] A z-score >3 indicates that Hap I is transmitted to the fetus.
[0340] For RHDO analysis, based on formulas (1) through (7), we simulated 30,000 haplotype blocks throughout the entire genome of the fetus via Hap I. The average length of the haplotype blocks was 100 kb. Each haplotype block contained an average of 100 SNPs, of which 10 SNPs would provide information in contributing to haplotype imbalance. In one instance, the fetal DNA fraction was 10% and the median fragment size was 150 bp. We calculated the percentage of haplotype blocks using a z-score >3 by varying the number of plasma DNA molecules used for RHDO analysis in the range of 1 million to 300 million, which is referred to herein as the detection rate. The number of plasma DNA molecules in this paper was adjusted for the probability of plasma DNA covering informative SNP sites according to the Passon distribution.
[0341] For computer simulations relating to qualitative analysis of maternal genetics in the fetus based on methylation patterns, we make the following assumptions for illustrative purposes:
[0342] 1) There are N plasma DNA molecules covering haplotype blocks in the maternal genome used for analysis.
[0343] 2) The probability of a plasma DNA fragment of at least 3 kb in length being used for tissue of origin analysis is denoted by a.
[0344] 3) The probability of a plasma DNA molecule carrying more than 10 CpG sites is represented by b.
[0345] 4) The fetal DNA fraction of those fragments >3kb is represented by f.
[0346] As illustrated in one embodiment of this disclosure, we can achieve accurate inference of the tissue of origin for plasma DNA molecules larger than 3 kb that have at least 10 CpG sites. It is assumed that the number (Z) of plasma DNA molecules satisfying the above criteria follows a Passon distribution, with a mean value of λ (i.e., N × a × b × f).
[0347] Z~Pason(λ) (8).
[0348] In one instance, based on Equation (8), we simulated 30,000 haplotype blocks in which Hap I was transferred to the fetus. The average length of each haplotype block was 100 kb. Each haplotype block contained an average of 100 SNPs, of which 20 heterozygous SNPs would be phased into two maternal haplotypes. The fetal DNA fraction was 1%. After size selection, 40% of plasma DNA molecules were >3 kb in size. 87.1% of plasma DNA molecules with at least 10 CpG sites were >3 kb in size. The percentage of haplotype blocks with a Z-value ≥1 indicates the detection rate. We repeated the computer simulation multiple times by varying the number (N) of plasma DNA molecules (N) for tissue-of-origin analysis in the range of 1 million to 300 million using methylation patterns. The number of plasma DNA molecules in this paper was further adjusted according to the Passon distribution by the probability of plasma DNA covering heterozygous SNPs.
[0349] Figure 31 This shows the detection rate of qualitative analysis of fetal maternal inheritance using genetic and epigenetic information from plasma DNA molecules in a genome-wide manner, compared to relative haplotype dose (RHDO) analysis. The number of molecules used in the analysis is shown on the x-axis. The detection rate of fetal maternal inheritance as a percentage is shown on the y-axis. The detection rate of fetal maternal inheritance using the methylation pattern-based approach is higher than that using RHDO. For example, using 100 million fragments, the detection rate based on the methylation pattern is 100%, while the detection rate based on RHDO is only 55%. These results indicate that inference of fetal maternal inheritance using the methylation pattern-based approach is superior to that using the RHDO-based approach.
[0350] 2. Paternal inheritance in the fetus
[0351] The ability to obtain long plasma DNA molecules for analysis can be applied to improve the detection rate of paternal-specific variants in maternal plasma DNA because using the same number of long DNA molecules will increase total genome coverage compared to using short DNA molecules. We further performed computer simulations based on the following assumptions:
[0352] 1) The fetal DNA fraction is f, depending on the length L of the plasma DNA. It is rewritten as f. L The subscript L indicates that a plasma DNA molecule of length L (in bp) is used for analysis.
[0353] 2) The number of father-specific variants identified in maternal plasma DNA needs to be V.
[0354] 3) The number of plasma DNA molecules used for analysis is N.
[0355] 4) The number of plasma DNA molecules originating from a specific genomic locus or region follows the Parsons distribution.
[0356] In one instance, the fetal DNA fractions of plasma DNA molecules with sizes of 150 bp, 1 kb, and 3 kb were 10% (f 150bp =0.1), 2% (f 1kb =0.02) and 1% (f 3kb =0.01). The number of father-specific variants in the genome was 250,000 (V = 250,000). The number of plasma DNA molecules (N) used for analysis ranged from 50 million to 500 million.
[0357] Figure 32 This diagram shows the relationship between the detection rate of father-specific variants analyzed in a genome-wide manner and the number of sequenced plasma DNA molecules of different sizes used for analysis. The number of sequenced molecules used for analysis, in millions, is shown on the x-axis. The percentage of father-specific variants detected is shown on the y-axis. Different curves show the different sizes of DNA fragments used for analysis, with a top of 3kb, a middle of 1kb, and a bottom of 150bp. The longer the plasma DNA molecules used for analysis, the higher the achievable detection rate of father-specific variants. For example, using 400 million plasma DNA molecules, the detection rates were 86%, 93%, and 98% when focusing on molecules of 150bp, 1kb, and 3kb, respectively.
[0358] In other embodiments, other distributions, including but not limited to Bernoulli distribution, Beta (β) normal distribution, normal distribution, Conway-Maxwell-Poisson distribution, geometric distribution, etc., may be used. In some embodiments, Gibbs sampling and Beth's theorem will be used for maternal and paternal genetic analysis.
[0359] 3. X-chromosome fragility genetic analysis
[0360] In this embodiment, the determination of maternally inherited methylation patterns in the fetus facilitates non-invasive detection of fragile X syndrome using single-molecule real-time sequencing of maternal plasma DNA. Fragile X syndrome is a genetic disorder typically caused by an expansion of the CGG trinucleotide repeat sequence within the FMR1 (Fragile X Disorder 1) gene on the X chromosome. Fragile X syndrome and other conditions caused by repeat sequence expansions are described elsewhere in this application. Methods for detecting fragile X syndrome in the fetus are also applicable to any other repeat sequence expansions disclosed herein.
[0361] Women with a quasi-mutation, defined as having 55 to 200 copies of the CGG repeat sequence in the FMR1 gene, are at risk of having a child with fragile X syndrome. The likelihood of having a fetus with fragile X syndrome depends on the number of CGG repeat sequences present in the FMR1 gene. The greater the number of repeat sequences in the mother, the higher the risk of the quasi-mutation expanding to a full mutation when passed to the fetus. Maternal plasma samples are collected at 12 weeks of gestation from women who have been previously identified as carrying an X-chromosome fragile quasi-mutation allele with 115 ± 2 CGG repeat sequences and who have a son diagnosed with fragile X syndrome (the primary disease patient). The maternal plasma is then subjected to single-molecule real-time sequencing. In one example, we used single-molecule real-time sequencing to obtain 3.3 million circular common sequences (CCS) aligned to a human reference genome, with a median subread depth of 75-fold / CCS (interquartile range: 14-237-fold). The genetic and epigenetic information of each sequenced plasma DNA can be determined according to embodiments of this disclosure. To obtain the two maternal haplotypes on chromosome X, we used the Infinium Omni2.5Exome-8 Beadchip microarray technology on the iScan system (Illumina) to genotype 2,000 SNPs on chromosome X using two DNA samples extracted from the maternal erythrocyte sedimentation rate (ESR) amber layer and the buccal swab from the primary disease patient. The two maternal haplotypes, Hap I and Hap II, were inferred based on genotypic information from the maternal and primary disease patient genomes.
[0362] Figure 33This describes a workflow for non-invasive detection of fragile X chromosome. In the heterozygous conjugation SNP sites of maternal erythrocyte sedimentation rate (ESR) amber layer DNA, alleles identical to the primary disease genotype are used to define haplotypes associated with quasi-mutant alleles (i.e., Hap I), which are potential precursors to complete mutations in offspring. Conversely, alleles different from the primary disease genotype are used to define haplotypes associated with corresponding wild-type alleles (Hap II). Maternal plasma DNA from mothers carrying fetuses of the primary disease undergoes single-molecule real-time sequencing. Sequencing reads are assigned to maternal Hap I and Hap II, depending on whether the obtained genetic information is identical to the Hap I or Hap II alleles at the studied genomic loci. According to embodiments of this disclosure, the methylation pattern of plasma DNA molecules is used to determine the tissue of origin for plasma DNA molecules containing a specific number of CpG sites (i.e., DNA molecules identified as placental-origin based on methylation pattern analysis are determined to be fetal-origin).
[0363] In scenario A, if fetal (i.e., placental) DNA molecules are detectable in plasma DNA molecules assigned to maternal Hap I but not in those assigned to maternal Hap II, then Hap I will be identified as transmitted to the unborn fetus. The fetus will be identified as being at high risk of being affected by fragile X chromosome. The placental origin of plasma DNA molecules will be based on the methylation status of the molecules, as discussed below.
[0364] In scenario B, if fetal DNA molecules are detectable in plasma DNA molecules assigned to maternal Hap II but not in plasma DNA molecules assigned to maternal Hap I, then Hap II will be determined to have been transmitted to the unborn fetus. The fetus will be determined to be unaffected by fragile X chromosome.
[0365] In embodiments, the definitions of "detectable" and "undetectable" fetal DNA molecules can be determined by a cutoff value for the percentage of plasma DNA molecules identified as originating from the fetus (i.e., the placenta). Cutoff values for "detectable" may include, but are not limited to, greater than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, etc. Cutoff values for "undetectable" may include, but are not limited to, less than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, etc. In some embodiments, the difference in the percentage of plasma DNA molecules identified as originating from the fetus between Hap I and Hap II may be required to be greater than, but not limited to, 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, etc. In some other embodiments, haplotype information may be obtained from long-read sequencing technologies (e.g., PacBio or nanopore sequencing) (Edge et al., Nature Communications, 2019; 10:4660), synthetic long reads (e.g., using a platform from 10X Genomics) (Hui et al., Clinical Chemistry, 2017; 63:513-14), phasing based on targeted locus amplification (TLA) (Vermeulen et al., American Journal of Human Genetics, 2017; 101:326-39), and statistical phasing (e.g., Shape-IT) (Delaneau et al., Nature Methods, 2011; 9:179-81).
[0366] In the embodiments, we can determine the maternal and fetal origin of plasma DNA molecules containing at least 200 bp and at least 5 CpG sites (or any other cutoff value for long DNA molecules) according to the methylation state matching pathway disclosed in this application. We identified a plasma DNA molecule located at genomic position chrX:143,782,245-143,782,786 (3.2 Mb from the FMR1 gene), in which the allele (position: chrX:143782434; SNP register number: rs6626483; allele genotype: C) is identical to the corresponding allele on maternal Hap II, but different from the corresponding allele on maternal Hap I.
[0367] Figure 34This shows the methylation pattern of plasma DNA compared to the methylation profiles of placental and erythrocyte sedimentation rate (ESR) tannin DNA. Plasma DNA molecules contain 5 CpG sites. The methylation pattern was identified as “MUUUU”. Following the methylation status matching pathway described in this disclosure, this methylation pattern obtained from single-molecule real-time sequencing was compared to the reference methylation profiles of placental tissue and ESR tannin DNA samples obtained from bisulfite sequencing. The fraction of this molecule originating from the placenta [i.e., S(placental)] was 2, which is greater than the fraction of molecules originating from the ESR tannin [i.e., S(ESR tannin)] - 3. Therefore, these plasma DNA molecules (chrX: 143, 782, 245-143, 782, 786) were identified as fetal in origin. However, we did not observe any plasma DNA molecules carrying alleles of maternal Hap I from fetal origin. Therefore, we conclude that the fetus inherited maternal Hap II and was not affected by fragile X chromosome disease.
[0368] We hypothesize that the efficacy of the pathway described in this paper is not significantly affected by X chromosome inactivation due to the following factors:
[0369] 1) X inactivation is incomplete in humans. Up to one-third of the genes on the X chromosome show variable evasion from X inactivation (Cotton et al., Human Molecular Genetics, 2015; 25:1528-1539). CpG sites outside CpG islands (i.e., most CpG sites) are methylated to similar degrees in both sexes, indicating that the methylation status of most CpG sites on the X chromosome is unaffected by X inactivation (Yasukochi et al., Proceedings of the National Academy of Sciences, 2010; 107:3704-9).
[0370] 2) We use methylation profiles of placental tissue matched to the sex of unborn fetuses. This strategy can be used to detect maternal genetics of the fetus using plasma DNA methylation patterns from women carrying male fetuses, because placental tissue involving male fetuses, presumably unaffected by X-ray inactivation, will have a unique methylation pattern that differs from other maternal tissues that are more or less affected by X-ray inactivation in specific regions.
[0371] We further sequenced DNA extracted from maternal erythrocyte sedimentation rate (ESR) amber layer samples using single-molecule real-time sequencing. We obtained 2.3 million CCSs with a median read depth of 5-fold / CCS. The results confirmed that maternal Hap I carried a quasi-mutant allele with 124 CGG repeat sequences and maternal Hap II carried a wild-type allele with 43 CGG repeat sequences. Additionally, we further sequenced DNA extracted from chorionic villus sampling of unborn fetuses using single-molecule real-time sequencing. We obtained 1.1 million CCSs with a median read depth of 4-fold / CCS. The results confirmed that the unborn fetus carried a wild-type allele.
[0372] E. Distribution of CpG sites in the human genome
[0373] Longer DNA fragments increase the likelihood of having multiple CpG sites. These multiple CpG sites can be used for methylation patterns or other analyses.
[0374] Figure 35 This displays the distribution of CpG sites in 500 bp regions throughout the human genome. The first column shows the number of CpG sites. The second column shows the number of 500 bp regions and the number of CpG sites. The third column shows the proportion of all regions represented by regions with a specific number of CpG sites. For example, 86.14% of 500 bp regions will have at least one CpG site. Additionally, 11.08% of 500 bp regions will have at least 10 CpG sites.
[0375] Figure 36 This displays the distribution of CpG sites in 1kb regions throughout the human genome. The first column shows the number of CpG sites. The second column shows the number of 1kb regions and the number of CpG sites. The third column shows the proportion of all regions represented by regions with a specific number of CpG sites. For example, 91.67% of 500bp regions will have at least one CpG site. Furthermore, 32.91% of 500bp regions will have at least 10 CpG sites.
[0376] Figure 37 This displays the distribution of CpG sites in 3kb regions throughout the human genome. The first column shows the number of CpG sites. The second column shows the number of 3kb regions and the number of CpG sites. The third column shows the proportion of all regions represented by regions with a specific number of CpG sites. For example, 92.45% of 3kb regions will have at least one CpG site. Additionally, 87.09% of 3kb regions will have at least 10 CpG sites.
[0377] In some embodiments, different numbers and size cutoffs of CpG sites are used to maximize the sensitivity and specificity of placental-specific marker identification and tissue of origin analysis. Generally, CpG sites occur more frequently than SNPs. DNA fragments of a given size may have more CpG sites than SNPs. The table shown above illustrates a lower proportion of regions with the same number of SNPs as CpG sites because regions of the same size contain fewer SNPs than CpG sites. Therefore, using CpG sites allows for the use of more fragments and provides better statistics than using SNPs alone.
[0378] F. Examples of Origin Tissue Analysis
[0379] In this embodiment, the analysis of the originating tissue in maternal plasma can be extended to more than two organs / tissues containing T cells, B cells, neutrophils, liver, and placenta. We used single-molecule real-time sequencing to sequence nine maternal DNA samples. We inferred the placental contribution to maternal plasma DNA using plasma DNA methylation patterns according to the methylation state matching pathway described in this disclosure. In one embodiment, for this methylation state matching analysis, the methylation patterns of each DNA molecule in the maternal plasma DNA sample that is at least 500 bp in length and contains at least five CpG sites are compared to a reference tissue methylation profile obtained from bisulfite sequencing. Five tissues containing neutrophils, T cells, B cells, liver, and placenta are used as reference tissues. Plasma DNA molecules are assigned to the tissue corresponding to the highest methylation state matching score of said plasma DNA molecule. The percentage of plasma DNA molecules assigned to a tissue relative to other tissues is considered the proportion of that tissue's contribution to the maternal plasma DNA of the sample. In this embodiment, the sum of the contribution proportions of neutrophils, T cells, and B cells in maternal plasma provides an alternative representation of the contribution proportions of hematopoietic cells.
[0380] Figure 38 This displays the percentage contribution of DNA molecules from different tissues in maternal plasma using methylation status matching analysis. The first column shows sample identification. The second column shows hematopoietic cell contribution as a percentage. The third column shows liver contribution as a percentage. The fourth column shows placental contribution as a percentage. Figure 38 The results showed that hematopoietic cells were the main contributors to maternal plasma DNA (median: 55.9%), which is consistent with previous reports (Sun et al. Proceedings of the National Academy of Sciences 2015; 112:E5503-12; Zheng et al. Clinical Chemistry 2012; 58:549-58).
[0381] Figure 39A and Figure 39BThis diagram illustrates the relationship between placental contribution and fetal DNA fraction inferred via the SNP pathway. The X-axis shows the fetal fraction determined via the SNP pathway. The Y-axis shows the placental contribution in maternal plasma as a percentage, determined using methylation status matching analysis. Figure 39A The study showed a strong correlation between placental contribution determined by methylation status matching analysis and fetal DNA fraction inferred from SNPs (Pearson's correlation coefficient = 0.95; P < 0.0001). We further performed tissue deconvolution analysis of maternal plasma DNA based on a secondary program by comparing plasma DNA methylation density determined by single-molecule real-time sequencing with methylation profiles from various reference tissues obtained from bisulfite sequencing (Sun et al., Proceedings of the National Academy of Sciences, 2015; 112: E5503-12). Figure 39B The study showed that using a methylation density-based approach reduced the correlation between placental contribution (Sun et al., Proceedings of the National Academy of Sciences, 2015; 112:E5503-12) and fetal DNA fraction compared to using methylation state matching analysis (Pearson correlation coefficient = 0.65; P = 0.059).
[0382] These data demonstrate that it is feasible to infer the proportion of DNA molecules contributed by different tissues in maternal plasma DNA samples. In another embodiment, this method can also be used to measure DNA molecules from different cell types or tissues in samples obtained after invasive solid tissue biopsy or from solid tissues obtained after surgery. In some embodiments, the use of methylation patterns at the level of individual DNA molecules to infer the proportion of contribution of different tissues to maternal plasma DNA is superior to the use of a pathway based on the aggregate methylation density of all sequenced plasma DNA molecules from the entire genome.
[0383] G. Instance Methods
[0384] Figure 40 Method 4000 shows the analysis of biological samples obtained from women carrying fetuses. Biological samples may contain multiple cell-free DNA molecules from the fetus and the woman.
[0385] At block 4010, sequence reads corresponding to a plurality of cell-free DNA molecules may be received. In some embodiments, method 4000 may include performing sequencing of cell-free DNA molecules.
[0386] At block 4020, the size of multiple cell-free DNA molecules can be measured. The measurement may include aligning a sequence read to a reference genome. In some embodiments, the measurement may include full-length sequencing of nucleotides in a full-length sequence and counting their number. In some embodiments, the measurement may include physically separating multiple cell-free DNA molecules from other cell-free DNA molecules in the biological sample, wherein the other cell-free DNA molecules have a size smaller than a cutoff value. Physical separation may include any of the techniques described herein, including the use of beads.
[0387] At block 4030, a group of cell-free DNA molecules from multiple cell-free DNA molecules can be identified as having a size greater than or equal to a cutoff value. The cutoff value can be greater than or equal to 200 nt. The cutoff value can be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cutoff value can be any cutoff value described herein for long cell-free DNA molecules. The size can be the number of CpG sites rather than the molecule length. For example, the cutoff value can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more CpG sites.
[0388] At block 4040, for one cell-free DNA molecule in the group of cell-free DNA molecules, the methylation status at each of a plurality of sites can be determined. The plurality of sites may contain at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more CpG sites. At least one of the plurality of sites may be methylated. Two sites of the plurality of sites may be spaced at least 160 nt, 170 nt, 180 nt, 190 nt, 200 nt, 250 nt, or 500 nt apart. The method may include sequencing the plurality of cell-free DNA molecules to obtain sequence reads and determining the methylation status of the site by measuring the characteristics of the nucleotides corresponding to the site and the nucleotides of neighboring sites. For example, methylation may be generally determined as in U.S. Application No. 16 / 995,607.
[0389] At block 4050, the methylation pattern can be determined. The methylation pattern can indicate the methylation status at each of multiple sites.
[0390] At block 4060, a methylation pattern can be compared with one or more reference patterns. Each of the one or more reference patterns can be determined for a specific tissue type. In some embodiments, the comparison may include determining the number of sites that match the reference patterns.
[0391] The reference patterns in one or more reference patterns can be determined by measuring the methylation density at each of the multiple reference sites using DNA molecules from a reference tissue. The methylation density at each of the multiple reference sites can be compared to one or more threshold methylation densities. Each of the multiple reference sites can be identified as methylated, unmethylated, or non-informative based on the comparison of methylation density to one or more threshold methylation densities, wherein the multiple sites are multiple reference sites identified as methylated or unmethylated. Non-informative sites may include sites having a methylation density between two threshold methylation densities. For example, the methylation index of a non-informative site may be between 30 and 70 or in any other range as described herein.
[0392] At step 4070, the methylation pattern can be used to determine the tissue of origin of the free DNA molecule. The tissue of origin can be the placenta. The tissue of origin can be the fetus or the mother. The method may include determining the tissue of origin as a reference tissue when the methylation pattern matches a reference pattern, and using... Figure 22 The description is similar. Matching can refer to an exact match. In some embodiments, identifying the origin tissue as the reference tissue can occur when the methylation pattern matches a reference pattern at a specific percentage of sites. For example, the methylation pattern may match a reference pattern at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, or more sites.
[0393] The method may include determining the origin organization by comparing a methylation pattern with a first reference methylation pattern from a plurality of reference organizations to determine a similarity score. The similarity score can be calculated using the methylation state matching method described herein or a beta (β) distribution probability model. The similarity score can be compared to a threshold. When the similarity score exceeds the threshold, the origin organization is determined to be the first reference organization. The similarity score may be the first similarity score. The method may further include calculating the threshold by comparing a methylation pattern with a second reference methylation pattern from a plurality of reference organizations to determine a second similarity score. The first and second reference organizations may be different organizations. The threshold may be the second similarity score. The first reference organization may have the highest similarity score compared to all other reference organizations.
[0394] The first reference methylation pattern may include a first subgroup of sites having at least a first methylation probability for a first reference tissue. For example, the first subgroup of sites may be sites considered methylated or generally considered methylated. The first reference methylation pattern may include a second subgroup of sites having at most a second methylation probability for a first reference tissue. For example, the second subgroup of sites may be sites considered unmethylated or generally considered unmethylated. Determining a similarity score may involve increasing the similarity score when one site among a plurality of sites is methylated and said site is in the first subgroup of sites, and decreasing the similarity score when one site among a plurality of sites is methylated and said site is in the second subgroup of sites. The similarity score may be determined to be similar to a methylation state matching pathway described herein.
[0395] The first reference methylation pattern comprises multiple sites, each characterized by a methylation probability and an unmethylation probability relative to the first reference tissue. For each site, the similarity score can be determined by identifying the probability in the reference tissue corresponding to the methylation state of the site in the free DNA molecule. The similarity score can be determined by calculating the product of the multiple probabilities. The product can be the similarity score. The probabilities can be determined using a beta (β) distribution, similar to the approach described herein.
[0396] Method 4000 may further include determining the origin organization of each free DNA molecule in the group of free DNA molecules. This determination may include determining the methylation state at each of a plurality of corresponding sites, wherein the plurality of corresponding sites correspond to free DNA molecules. Determining the origin organization may further include determining a methylation pattern. Additionally, determining the origin organization may also include comparing the methylation pattern with at least one of one or more reference patterns. In some embodiments, the comparison of methylation patterns may be compared with... Figure 22 Similar to the accompanying description. In Figure 22 In the illustration, the placenta, liver, blood cells, and colon are examples of reference tissues with the illustrated reference pattern. Figure 38 Hematopoietic cells are shown in another instance of the reference tissue.
[0397] In some embodiments, the amount of cell-free DNA molecules corresponding to each origin tissue can be determined. Each origin tissue may include reference tissues among a plurality of reference tissues. The contribution fraction of the origin tissue can be determined using measurements of cell-free DNA molecules corresponding to each origin tissue. For example, the origin tissue may be the placenta. Other origin tissues may include hematopoietic cells and the liver. For example, the contribution fraction of the placenta can be determined by dividing the amount of cell-free DNA molecules by the total number of cell-free DNA molecules corresponding to all origin tissues. In some embodiments, the fraction calculated by dividing the amount of cell-free DNA molecules by the total number of cell-free DNA molecules can be correlated with the contribution fraction via a function or a set of calibration data points. Both the function and the set of calibration data points can be determined from multiple calibration samples having known contribution fractions of origin tissues. Each calibration data point can specify a contribution fraction corresponding to a calibration value of the fraction. The function can represent a linear or nonlinear fit of the calibration data points and can correlate the contribution fraction with the fraction of the origin tissue or other parameters involving the origin tissue. Examples of determining the contribution fraction can be compared with those already used Figure 39A and Figure 39B The described embodiments are similar.
[0398] Machine learning models can be used to determine the tissue of origin. The model can be trained by receiving multiple training methylation patterns, each having a methylation state at one or more sites among multiple sites, determined by DNA molecules from known tissues. These molecules from known tissues can be cellular DNA. Training may involve storing multiple training samples, each containing one of the multiple training methylation patterns and a label indicating the known tissue corresponding to that training methylation pattern. Training may also involve optimizing the model's parameters based on the model's output when the multiple training methylation patterns are input into the model, according to whether the labels match or do not match. The parameters may include a first parameter indicating whether one site among the multiple sites has the same methylation state as another site among the multiple sites. For example, the model can be compared with... Figure 24 The pairwise comparison is similar. The parameters may include a second parameter indicating the distance between each point across multiple sites. In some embodiments, the machine learning model may not require alignment of methylation sites with a reference genome. The model's output may specify the tissue corresponding to the input methylation pattern.
[0399] The machine learning model can be a convolutional neural network (CNN) or any model described herein. The model may include, but is not limited to, linear regression, logistic regression, deep recurrent neural networks (e.g., Long Short-Term Memory, LSTM), Bayesian classifiers, Hidden Markov Models (HMMs), Linear Discriminant Analysis (LDA), k-means clustering, density-based spatial clustering for noisy applications (DBSCAN), random forest algorithms, and support vector machines (SVMs).
[0400] Kinship can be determined by method 4000. The tissue of origin may be the fetus. The method may further include aligning one sequence read from the sequence read with a first region of a reference genome, the first region comprising multiple loci corresponding to alleles, the multiple loci comprising a threshold number of loci, determining a first haplotype using the corresponding alleles present at each of the multiple loci, comparing the first haplotype with a second haplotype corresponding to a male individual, and using the comparison to determine a classification of the probability that the male individual is the father of the fetus. If the haplotypes match, the male individual may be considered the father, or if the haplotypes do not match, the male individual may not be considered the father. In some embodiments, the first haplotype may be compared with two haplotypes of the male individual.
[0401] In an embodiment, when the tissue of origin is a fetus, kinship can be tested by aligning one sequence read from the sequence read with a first region of a reference genome. The first region may contain a first plurality of loci corresponding to alleles. The plurality of loci may contain a threshold number of loci. The threshold number of loci may be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more loci. Alleles at each of the plurality of loci can be compared with alleles at corresponding loci in the male individual's genome. The comparison can be used to determine a classification of the likelihood that the male individual is the father of the fetus. If a specific number or percentage of alleles match, the male individual may be considered the father, and if less than the stated number or percentage of alleles match, the male individual may not be considered the father. The cutoff percentage may be 100%, 90%, 80%, or 70%.
[0402] In some embodiments, haplotypes can be determined. The method may include, for each cell-free DNA molecule in the group, comparing a sequence read corresponding to the cell-free DNA molecule with a reference genome. The sequence read can be identified as corresponding to a haplotype present in a woman. The haplotype present in a woman can be determined from genotyping. In some embodiments, a woman's haplotype can be determined by analyzing the concentration of DNA fragments of the haplotype in a biological sample from the woman. Methylation patterns can be used to determine if the tissue of origin is fetal. The haplotype can be determined to be a maternally inherited fetal haplotype.
[0403] Haplotype inheritance can be determined using methylation profiles of a reference tissue rather than using known methylation profiles such as those associated with memorized loci. Matching or similarity scores between methylation patterns and reference patterns can exclude knowledge based on whether the parents, given alleles, or loci are methylated.
[0404] Haplotypes can be identified as carrying pathogenic genetic mutations or variations. Identifying a haplotype as carrying a pathogenic genetic mutation may be included in the identification of the genetic mutation or variation within the first sequence read. Genetic variations may include single nucleotide differences, deletions, or insertions. The degree of first methylation in a second sequence read corresponding to a first genomic position within a first distance of the first sequence read can be measured. The degree of second methylation in a third sequence read corresponding to a second genomic position within a second distance of the first sequence read can also be measured. The first distance can be 100 nt, 200 nt, 300 nt, 400 nt, 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 2 knt, 5 knt, or 10 knt. The second and third sequence reads may be located on the same chromosome arm as the first sequence read. The degree of first and second methylation may be associated with the genetic mutation or variation. The degree of first and second methylation may be greater than one or two threshold levels associated with the genetic mutation or variation. Threshold levels can be determined using individuals known to have or not have the genetic mutation or variation. The method may include classifying the fetus as potentially having a disease caused by a genetic mutation or variation.
[0405] Fetal-specific methylation patterns can be determined. The method may include, for each cell-free DNA molecule in the group, aligning a sequence read corresponding to the cell-free DNA molecule with a reference genome. The method may include identifying a sequence read corresponding to a region. The region can be determined by receiving multiple fetal sequence reads corresponding to multiple fetal DNA molecules from fetal tissue. The method may include receiving multiple maternal sequence reads corresponding to multiple maternal DNA molecules. The method may include determining the fetal methylation status at each of a plurality of methylation sites within the region for each fetal sequence read in the plurality of fetal sequence reads. The method may include determining the maternal methylation status at each of a plurality of methylation sites for each maternal sequence read in the plurality of maternal sequence reads.
[0406] A method for determining a fetal-specific methylation pattern may include determining the value of a parameter characterizing the amount of sites in which the fetal methylation state differs from the maternal methylation state. The method may include comparing the value of the parameter to a threshold. The parameter may be the proportion of different sites between fetal and maternal DNA molecules. The proportion may be the mismatch fraction described herein. The threshold may indicate a minimum level of mismatch fraction and may be 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or greater. In some embodiments, the threshold may represent the average mismatch fraction of maternal or fetal DNA molecules. The method may include determining that the value of the parameter exceeds the threshold. In some embodiments, a specific percentage of the maternal or fetal DNA molecules may be required to have a parameter value exceeding the threshold. For example, the percentage may be 50%, 60%, 70%, 80%, 90%, or greater. In some embodiments, a specific percentage of fetal DNA molecules corresponding to the region may be required to have a fetal-specific methylation pattern. For example, the percentage may be 40%, 50%, 60%, 70%, 80%, or greater. This method can be used with… Figure 25 The methods described are similar.
[0407] The method may include enriching a biological sample from the tissue of origin to obtain cell-free DNA molecules. Enriching the biological sample may involve selecting and amplifying the group of cell-free DNA molecules. Enrichment may include size-based selection as described herein. In some embodiments, enrichment may include selection based on methylation patterns. For example, capture and sequencing based on the methyl-CpG binding domain (MBD) may be used. The cell-free DNA may be incubated with a tagged MBD protein that can bind methylated cytosine. Subsequently, the protein-DNA complex may be precipitated using antibody-bound magnetic beads. DNA molecules with more methylated CpG sites may be preferentially enriched for downstream analysis.
[0408] III. Changes in long cell-free DNA fragments with gestational age
[0409] The amount of long cell-free DNA fragments can vary with gestational age. Long cell-free DNA fragments can be used to determine gestational age. Additionally, compared to shorter cell-free DNA fragments, long cell-free DNA fragments can be more abundant in specific terminal motifs, and the relative amount of these motifs can vary with gestational age. The amount of these terminal motifs can also be used to determine gestational age. The deviation between gestational age determined using long cell-free DNA fragments and gestational age determined by other clinical techniques can indicate pregnancy-related conditions. In some embodiments, long cell-free DNA fragments can be used to determine the likelihood of a pregnancy-related condition without having to determine gestational age.
[0410] A. Size analysis of fetal and maternal DNA
[0411] Plasma DNA from two pregnant women in early pregnancy (gestational age: 13 weeks), two pregnant women in mid-pregnancy (gestational age: 21-22 weeks), and five pregnant women in late pregnancy (gestational age: 38 weeks) was sequenced using single-molecule real-time (SMRT) sequencing (PacBio). For each case, a median of 176 million (range: 49 million-685 million) subreads were obtained, of which 128 million (range: 35 million-507 million) subreads were comparable to the human reference genome (hg19). Each molecule in the SMRT well was sequenced an average of 107 times. A median of 965,308 (range: 251,686-2,871,525) high-quality circular co-sequencing (CCS) reads, defined as CCS reads with at least three subreads, were available for downstream analysis.
[0412] All sequenced molecules from samples obtained at each stage of pregnancy were synthesized for size analysis. Maternal plasma samples from early, mid, and late pregnancy contained a total of 1.94 million, 5.09 million, and 4.45 million cell-free DNA molecules, respectively.
[0413] Figure 41A and Figure 41B This displays the size distribution of cell-free DNA molecules ranging from 0kb to 5kb in maternal plasma samples from early, mid, and late pregnancy. The x-axis shows size. The y-axis shows frequency. Figure 41A Plot the size distribution of the y-axis on a linear scale in the range of 0kb to 5kb, and for Figure 41B Plot the size distribution of the y-axis on a logarithmic scale, ranging from 0 kb to 5 kb. Plasma DNA from all three gestational periods is shown below. Figure 41A The expected main peak shown is at 166bp, and as shown... Figure 41B The diagram shows a series of main peaks that appear in a periodic pattern extending to molecules in the 1kb and 2kb range.
[0414] Figure 42A table showing the proportion of long plasma DNA molecules at different stages of pregnancy. The first column shows the gestational age associated with the plasma sample. The second column shows the proportion of DNA molecules longer than 500 bp. The third column shows the proportion of DNA molecules longer than 1 kb. Compared to early and mid-pregnancy, the frequency of plasma DNA molecules of 500 bp or longer increases in late pregnancy. The proportions of long plasma DNA molecules longer than 500 bp in early, mid, and late pregnancy are 15.8%, 16.1%, and 32.3%, respectively. The proportions of long plasma DNA molecules longer than 1 kb in early, mid, and late pregnancy are 11.3%, 10.6%, and 21.4%, respectively. While similar proportions of long cell-free DNA molecules in maternal plasma are observed in early and mid-pregnancy, the proportion of long DNA molecules in maternal plasma in late pregnancy is approximately twice that of the aforementioned long DNA molecules.
[0415] For all maternal plasma DNA samples analyzed in this disclosure, DNA extracted from paired maternal erythrocyte sedimentation rate (ESR) amber layer and fetal samples was genotyped using the Infinium Omni 2.5 Exome-8 Beadchip on the iScan system (Illumina) as an array hybridization-based genotyping method. Fetal samples were obtained via chorionic villus sampling, amniocentesis, or placental sampling, depending on whether the case was from early, mid, or late pregnancy. For each case, a median of 203,647 informative single nucleotide polymorphisms (SNPs) were identified, with the mother being homozygous and the fetus heterozygous. When the sequenced DNA molecules from each gestational period for all cases were aggregated, a total of 1,362, 2,984, and 6,082 DNA molecules covering fetal-specific alleles were identified in early, mid, and late pregnancy, respectively. On the other hand, for each case, a median of 210,820 informative SNPs were identified, where the mother was heterozygous and the fetus was homozygous. We identified a total of 30,574, 65,258, and 78,346 DNA molecules covering maternally specific alleles in early, mid, and late pregnancy, respectively. The median fetal DNA fraction determined from sequencing data of DNA molecules ≤600 bp was 15.6% (range 7.6%–26.7%) across all maternal plasma samples.
[0416] Figure 43A and Figure 43B This displays the size distribution of DNA molecules encompassing fetal-specific alleles from maternal plasma in early, mid, and late pregnancy. The x-axis shows size. The y-axis shows frequency. For Figure 43A Plot the size distribution of the y-axis on a linear scale in the range of 0kb to 3kb, and for Figure 43BPlot the size distribution of the y-axis on a logarithmic scale in the range of 0kb to 3kb.
[0417] Figure 44A and Figure 44B This displays the size distribution of DNA molecules encompassing maternal-specific alleles from maternal plasma in early, mid, and late pregnancy. The x-axis shows size. The y-axis shows frequency. For Figure 44A Plot the size distribution of the y-axis on a linear scale in the range of 0kb to 3kb, and for Figure 44B Plot the size distribution of the y-axis on a logarithmic scale in the range of 0kb to 3kb.
[0418] like Figures 43A to 44B As shown, plasma DNA molecules containing fetal-specific alleles and plasma DNA molecules containing maternal-specific alleles from all three gestational periods exhibit a long-tailed distribution, indicating that long DNA molecules from both fetal and maternal sources are present in all three gestational periods.
[0419] Figure 45 This table shows the proportions of long fetal and maternal plasma DNA molecules at different gestational stages. The first column shows the gestational age associated with the plasma sample. The second column shows the proportion of fetal DNA molecules longer than 500 bp. The third column shows the proportion of maternal DNA molecules longer than 500 bp. The fourth column shows the proportion of fetal DNA molecules longer than 1 kb. The fifth column shows the proportion of maternal DNA molecules longer than 1 kb. Within the pool of DNA molecules in maternal plasma, DNA molecules containing fetal-specific alleles (of placental origin) have a smaller proportion of long DNA molecules compared to those containing maternally specific alleles. The proportions of long plasma DNA molecules containing fetal-specific alleles and larger than 500 bp in early, mid, and late pregnancy are 19.8%, 23.2%, and 31.7%, respectively. The proportions of long plasma DNA molecules containing fetal-specific alleles and larger than 1 kb in early, mid, and late pregnancy are 15.2%, 16.5%, and 19.9%, respectively.
[0420] Regardless of the fact that a smaller proportion of long plasma DNA molecules are present in maternal plasma during early and mid-pregnancy compared to late pregnancy, and that fetal DNA molecules contain fewer long DNA molecules across all three gestational periods, the methods described in our previous and present disclosures allow us to analyze a considerable proportion of long plasma DNA molecules that were previously impossible to analyze using short-read sequencing technologies. Furthermore, we can use different size selection strategies, including but not limited to electrophoresis, chromatography, and bead-based methods, to enrich long DNA fragments in plasma samples.
[0421] Figure 46A , Figure 46B and Figure 46C This plot shows the proportions of fetal-specific plasma DNA fragments within a specific size range at different stages of pregnancy. The gestational age of the pregnancies being evaluated was verified using dating ultrasound. Figure 46A Results show DNA fragments less than or equal to 150 bp. Figure 46B Results showing DNA fragments ranging from 150 to 600 bp. Figure 46C The results show DNA fragments greater than or equal to 600 bp. The graph displays the proportion of fetal-specific fragments on the y-axis and gestational age on the x-axis. As shown in the graph, the proportion of fetal-specific fragments in the range of 150 bp to 600 bp ( Figure 46B Compared to ), the proportion of fetal-specific fragments shorter than 150bp ( Figure 46A The proportion of fetal-specific fragments longer than 600 bp () Figure 46C Both approaches will achieve a specific distinguishing dynamic between samples from late pregnancy and those from early and mid-pregnancy. The proportion of fetal-specific fragments longer than 600 bp provides the best distinguishing dynamic. This conclusion is demonstrated by the fact that when using a proportion of fetal-specific fragments shorter than 150 bp, the absolute minimum distance between the combined group of late pregnancy and early / mid-pregnancy is 0.38, while when using a proportion of fetal-specific fragments longer than 600 bp, the corresponding value is 3.76. These results indicate that the use of longer DNA molecules reflects pathophysiological states and is superior to the use of shorter DNA molecules.
[0422] B. Plasma DNA End Analysis
[0423] In addition to size, we determined the first nucleotide at the 5' end of both the Watson and Crick strands for each sequenced DNA molecule. This analysis consisted of four types of ends: A-terminus, C-terminus, G-terminus, and T-terminus. The percentage of plasma DNA molecules with specific ends from maternal plasma samples obtained at each stage of pregnancy was calculated. The percentages of A-terminus, C-terminus, G-terminus, and T-terminus at each fragment size were further analyzed.
[0424] Figure 47A , Figure 47B and Figure 47C This graph shows the proportion of bases at the 5' end of cell-free DNA molecules from maternal plasma during early, mid, and late pregnancy, spanning a fragment size range from 0kb to 3kb. Figure 47A Displays maternal plasma from early pregnancy. Figure 47B Displays maternal plasma from mid-pregnancy. Figure 47CThis diagram shows maternal plasma in late pregnancy. Base content is shown as a percentage on the y-axis. Fragment size, in base pairs, is shown on the x-axis. As seen in the diagram, C-terminal abundance is presented across many size ranges (mostly less than 1 kb) and varies depending on the size range used for samples from early, mid, and late pregnancy. The plasma DNA end pattern in late pregnancy samples appears to differ from that in early and mid pregnancy samples. For example, T-terminal and G-terminal curves are mixed at sizes ranging from 105 bp to 172 bp, but diverge in early and mid pregnancy samples. For longer fragments (e.g., greater than approximately 1 kb), the C-terminal fragment is not the most abundant. At approximately 1 kb, the G-terminal fragment surpasses the C-terminal fragment, and subsequently, at approximately 2 kb, the A-terminal fragment becomes more abundant than the G-terminal fragment.
[0425] Figure 48 This table shows the proportions of terminal nucleotide bases in short and long cell-free DNA molecules from maternal plasma in early, mid, and late pregnancy. The first column displays the bases at the molecular ends. The second column shows the expected proportions and species. The third column shows the proportion of terminal species in fragments of 500 bp or less from early pregnancy maternal plasma. The fourth column shows the proportion of terminal species in fragments greater than 500 bp from early pregnancy maternal plasma. Columns five and six are similar to columns three and four, respectively, except that maternal plasma from mid-pregnancy is used instead of early pregnancy maternal plasma. Columns seven and eight are similar to columns three and four, respectively, except that maternal plasma from late pregnancy is used instead of early pregnancy maternal plasma.
[0426] If the cell-free DNA fragmentation is completely random, the terminal nucleotide base ratios should reflect the composition of the human genome, which is 29.5% A, 29.5% T, 20.5% C, and 20.5% G. Figure 48 As shown in the second column. In contrast to random fragmentation, the 5' ends of short cell-free DNA molecules ≤500 bp showed considerable overpresentation at the C-terminus (30.4%, 30.4%, and 31.3% for maternal plasma in early, mid, and late pregnancy, respectively), slight overpresentation at the G-terminus (27.4%, 26.9%, and 25.3% for early, mid, and late pregnancy, respectively), and very low presentation at the A-terminus (19.8%, 19.4%, and 19.3% for early, mid, and late pregnancy, respectively), as well as very low presentation at the T-terminus (22.4%, 23.3%, and 24.1% for early, mid, and late pregnancy, respectively).
[0427] However, when compared with short cell-free DNA molecules, long cell-free DNA molecules (>500 bp) showed a considerably increased proportion of A-terminus (29.6%, 26.0%, and 26.7% in maternal plasma during early, mid, and late pregnancy, respectively), a slightly increased proportion of G-terminus (31.0%, 29.5%, and 29.9% in early, mid, and late pregnancy, respectively), a considerably decreased proportion of T-terminus (13.9%, 16.9%, and 16.4% in early, mid, and late pregnancy, respectively), and a slightly decreased proportion of C-terminus (25.5%, 27.5%, and 27.1% in early, mid, and late pregnancy, respectively).
[0428] Figure 49 This table shows the ratio of terminal nucleotide bases in short and long cell-free DNA molecules containing fetal-specific alleles, derived from maternal plasma during early, mid, and late pregnancy. Figure 50 This table presents the proportions of terminal nucleotide bases in short and long cell-free DNA molecules containing maternally specific alleles from maternal plasma in early, mid, and late pregnancy. The first column shows the bases at the molecular ends. The second column shows the expected proportion points and species. The third column shows the proportion of terminal species in fragments of 500 bp or less from early pregnancy maternal plasma. The fourth column shows the proportion of terminal species in fragments greater than 500 bp from early pregnancy maternal plasma. Columns five and six are similar to columns three and four, respectively, except that maternal plasma from mid-pregnancy is used instead of early pregnancy maternal plasma. Columns seven and eight are similar to columns three and four, respectively, except that maternal plasma from late pregnancy is used instead of early pregnancy maternal plasma. Figure 49 and Figure 50 The difference in the ratio of terminal nucleotide bases in short and long cell-free DNA molecules remains unchanged, even when we examine DNA molecules containing fetal-specific alleles and DNA molecules containing maternal-specific alleles separately.
[0429] Figure 51 This diagram illustrates hierarchical cluster analysis of short and long cell-free DNA molecules using 256 tetrameric terminal motifs. Each column indicates the sample used for analyzing terminal motif frequencies based on short fragments (indicated in cyan in the first column) and long fragments (indicated in yellow in the first column), respectively. Starting from the second column, each column indicates the type of terminal motif. Terminal motif frequencies are presented as a series of color gradients based on column-normalized frequencies (z-scores) (i.e., the standard deviation of frequencies below or above the mean frequency across the entire sample). Redder colors indicate higher terminal motif frequencies, while bluer colors indicate lower terminal motif frequencies.
[0430] exist Figure 51 In this study, we characterized short and long cell-free DNA molecules by analyzing their tetramer terminal motif profiles. We determined the first 4-nucleotide sequence (tetramer motif) at the 5' end of both Watson and Crick genes for each sequenced DNA molecule. For each maternal plasma sample, the frequencies of each plasma DNA terminal motif were calculated for short plasma DNA molecules (≤500 bp) and long plasma DNA molecules (>500 bp). Hierarchical cluster analysis based on 256 tetramer terminal motif frequencies showed that the terminal motif profiles of long DNA molecules formed different clusters than those of short DNA molecules throughout the various maternal plasma samples. These results indicate that long and short DNA molecules have different fragmentation characteristics. In this example, we will use the relative perturbation of these terminal motifs between long and short DNA molecules to indicate the contribution of cell-free DNA originating from cell death pathways such as, but not limited to, apoptosis and necrosis. Enhanced activity from these cell death pathways may be associated with pregnancy-related conditions and other conditions.
[0431] Figure 52A and Figure 52B This shows the principal component analysis (PCA) performed using the tetramer terminal motif profile for classification analysis. Figure 52A Showing short cell-free DNA molecules (≤500bp) from different stages of pregnancy. Figure 52B Displays long cell-free DNA molecules (>500bp) from maternal plasma samples at different gestational stages. Percentages in parentheses on the X and y axes represent the amount of variability expressed by the corresponding component. Blue dots represent maternal plasma samples from early pregnancy. Yellow dots represent maternal plasma samples from mid-pregnancy. Red dots represent maternal plasma samples from late pregnancy. Ellipses indicate the 95% confidence level used to group data points from a specific gestational stage. (Compared to short cell-free DNA molecules...) Figure 52A (Also described in U.S. Application No. 15 / 787,050) compared to long free DNA molecules ( Figure 52B The tertiary motif profile of the tetramer produces a clearer interval between maternal plasma samples from early, mid, and late pregnancy. In examples, we can utilize the tertiary motif profile of individual long plasma DNA molecules or in combination with other maternal plasma DNA features, including but not limited to methylation degree and size, for molecular gestational age assessment.
[0432] For example, we used a neural network to train a model for predicting gestational age based on 256 terminal motifs, total methylation level, and the proportion of fragments ≥600 bp in size. The output variables were 1, 2, and 3, representing early, mid, and late pregnancy. The input variables included 256 terminal motifs, total methylation level, and the proportion of fragments ≥600 bp in size. We used a leave-one-out method to evaluate the effectiveness of the gestational age prediction. For a dataset consisting of 9 samples, a leave-one-out method was used, with one sample selected as the test sample and the remaining 8 samples used to train the model based on the neural network. Based on the established model, this test sample was determined to be 1, 2, or 3. We then repeated this method on the other untested samples. We repeated this training and testing process a total of 9 times. By comparing those test results with clinical information about gestational age, 8 out of 9 samples (89%) were appropriately predicted in terms of gestational age. In another embodiment, the analysis may be performed, for example but not limited to, using Bayes' theorem, logistic regression, multiple regression and support vector machines, random forest analysis, classification and regression trees (CART), and the K nearest neighbor algorithm.
[0433] Next, all sequenced molecules from samples obtained at each stage of pregnancy were synthesized for downstream terminal motif analysis. These motifs were then categorized based on the frequency of 256 terminal motifs in both short and long plasma DNA molecules.
[0434] Figures 53 to 58 A table of the 25 most frequent terminal motifs for DNA fragments of specific lengths (shorter or longer than 500 bp) and different gestational stages. Figure 53 , Figure 54 and Figure 55 This is a list of terminal motifs sorted by terminal motif ranking in short fragments (<500bp). Figures 53 to 55 In the format, the first column displays the terminal sequence number. The second column displays the sequence frequency class in the short segment. The third column displays the sequence frequency class in the long segment. The fourth column displays the sequence frequency in the short segment. The fifth column displays the sequence frequency in the long segment. The sixth column displays the change factor (sequence frequency in the short segment divided by sequence frequency in the long segment).
[0435] Figure 56 , Figure 57 and Figure 58 This is a list of terminal motifs sorted by terminal motif ranking in long fragments (>500bp). Figures 56 to 58 In the format, the first column displays the terminal sequence number. The second column displays the sequence frequency class in the long segment. The third column displays the sequence frequency class in the short segment. The fourth column displays the sequence frequency in the long segment. The fifth column displays the sequence frequency in the short segment. The sixth column displays the change factor (sequence frequency in the long segment divided by sequence frequency in the short segment).
[0436] Figure 53 and Figure 56 Sample taken from early pregnancy. Figure 54 and Figure 57 Sample from mid-pregnancy. Figure 55 and Figure 58 Sample from late pregnancy.
[0437] Of the top 25 most frequent terminal motifs in short plasma DNA molecules, 11 begin with the C(C) dinucleotide. In maternal plasma during early, mid, and late pregnancy, C(C)-initialized terminal motifs accounted for 14.66%, 14.66%, and 15.13% of short plasma DNA terminal motifs, respectively. Of the top 25 most frequent terminal motifs in long plasma DNA molecules, TT-terminated tetrameric motifs accounted for 9 in mid and late pregnancy maternal plasma and 10 in early pregnancy maternal plasma.
[0438] We determined the dinucleotide sequences of the third nucleotide (X) and fourth nucleotide (Y) at the 5' end of both the Watson and Crick genes for each sequenced DNA molecule. X and Y can be one of the four nucleotide bases in DNA. There are 16 possible NNXY motifs: NNAA, NNAT, NNAG, NNAC, NNTA, NNTT, NNTG, NNTC, NNGA, NNGT, NNGG, NNGC, NNCA, NNCT, NNCG, and NNCC.
[0439] Figure 59A , Figure 59B and Figure 59C This shows the motif frequency scatter plot of 16 NNXY motifs in short and long plasma DNA molecules. Figure 59A This shows the results of early pregnancy. Figure 59B This shows the results from the second trimester. Figure 59C This displays results for late pregnancy. Motif frequencies of long fragments are shown on the y-axis. Motif frequencies of short fragments are shown on the x-axis. Each circle represents one of the 16 NNXY motifs. The dotted lines in each scatter plot represent a 1.5-fold increase (upper line) and a 1.5-fold decrease (lower line) in motif frequencies in long plasma DNA molecules (>500 bp) compared to short plasma DNA molecules (≤500 bp). Circles outside the shaded area represent motifs with a fold change >1.5.
[0440] In all three gestational periods (Figure 11), when short plasma DNA molecules exhibited a high frequency of tetrameric motifs starting with CC dinucleotides (CCNN) (Jiang et al., Cancer Discov, 2020; 10(5):664-673; Chan et al., American Journal of Human Genetics, 2020; 107(5):882-894), the frequency of tetrameric motifs ending with TT (NNTT) at the ends of long plasma DNA molecules increased by >1.5-fold. In maternal plasma during early, mid, and late pregnancy, NNTT motifs accounted for 18.94%, 15.22%, and 15.30% of long plasma DNA terminal motifs, respectively. Conversely, in maternal plasma during early, mid, and late pregnancy, NNTT motifs accounted for only 9.53%, 9.29%, and 8.91% of short plasma DNA terminal motifs, respectively.
[0441] As previously reported by Han et al., cell-free DNA newly released into plasma from dead cells is enriched for A-terminal fragments >150 bp. DNA fragmentation factor β (DFFB), a major intracellular nuclease involved in DNA fragmentation during apoptosis, was found to be responsible for generating these fragments (Han et al., *American Journal of Human Genetics*, 2020; 106:202-214). In this disclosure, we have shown that long cell-free DNA molecules >500 bp are also enriched for A-terminal fragments, suggesting that DFFB may also be responsible for generating these fragments. In normal pregnancy, trophoblastic cell apoptosis increases with gestational progression (Sharp et al., *American Journal of Reproductive Immunology*, 2010; 64(3):159-69). Indeed, our finding that the proportion of long DNA molecules covering fetal-specific alleles increases with gestational progression may reflect an increase in trophoblastic cell apoptosis with gestational progression.
[0442] In this embodiment, we can use the methods described herein to analyze long cell-free DNA molecules in maternal plasma for the prediction, screening, and development monitoring of placental-related pregnancy complications, including but not limited to preeclampsia, intrauterine growth restriction (IUGR), preterm birth, and trophoblastic disorders of pregnancy. Increased levels of trophoblastic cell apoptosis have been reported in placental-related pregnancy complications such as preeclampsia (Leung et al., *American Journal of Obstetrics and Gynecology*, 2001; 184:1249-1250), IUGR (Smith et al., *American Journal of Obstetrics and Gynecology*, 1997; 177:1395-1401; Levy et al., *American Journal of Obstetrics and Gynecology*, 2002; 186:1056-1061), and trophoblastic disorders of pregnancy. In addition, elevated levels of fetal DNA in maternal plasma have been reported in the following conditions: preeclampsia (Lo et al., Clinical Chemistry, 1999; 45(2):184-8; Smid et al., Annular Journal of the New York Academy of Sciences, 2001; 945:132-7), IUGR (Sekizawa et al., American Journal of Obstetrics and Gynecology, 2003; 188:480-4), and preterm birth (Leung et al., The Lancet, 1998; 352(9144):1904-5). We hypothesize that in placental-related pregnancy complications, the proportion of placental-originating long cell-free DNA molecules in maternal plasma samples increases due to increased placental cell apoptosis. Therefore, placental-originating long cell-free DNA molecules themselves, as well as long DNA markers containing, but not limited to, A-terminal fragments and NNTT motifs, may serve as biomarkers for placental cell apoptosis.
[0443] Although single nucleotide motifs and four nucleotide motifs were used in the analysis above, motifs of other lengths, such as 2, 3, 5, 6, 7, 8, 9, 10 or longer, may be used in other embodiments.
[0444] C. Instance Methods
[0445] Long cell-free DNA fragments can be used to determine the gestational age of a pregnant woman. The amount of long cell-free DNA fragments varies with gestational age and can be used to determine gestational age. The terminal motifs of cell-free DNA fragments also vary with gestational age and can be used to determine gestational age. When the gestational age determined using long cell-free DNA fragments deviates significantly from that determined by other clinical techniques, the pregnant woman and / or fetus may be considered to have a pregnancy-related condition. In some embodiments, determining gestational age may not be necessary to determine the likelihood of a pregnancy-related condition.
[0446] 1. Gestational age
[0447] Figure 60This demonstrates 6000 methods for analyzing biological samples obtained from women carrying a fetus. Gestational age can be determined and used to classify the likelihood of pregnancy-related conditions. The biological sample may contain multiple cell-free DNA molecules from the fetus and the woman.
[0448] Sequence reads corresponding to multiple free DNA molecules are acceptable. In some embodiments, sequencing for obtaining the sequence reads may be performed.
[0449] At block 6020, the size of multiple free DNA molecules can be measured. The size can be compared with... Figure 21 The measurement is done in a similar manner to the description. Dimensions can be measured using sequential reads.
[0450] At block 6030, the first quantity of free DNA molecules with a size greater than the cutoff value can be measured. The quantity can be the number of free DNA molecules, the total length, or the mass.
[0451] At block 6040, the first quantity can be used to generate the value of the normalized parameter. The value of the normalized parameter can be the first quantity normalized by the total number of cell-free DNA molecules, by the number of cell-free DNA molecules from the fetus or mother, or by the number of DNA molecules from a specific region. For example, if using... Figure 46A-C The standardized parameter described can be the proportion of fetal-specific fragments.
[0452] At block 6050, the value of the normalization parameter can be compared to one or more calibration data points. Each calibration data point can specify gestational age corresponding to a calibration value of the normalization parameter. For example, gestational age at a specific gestational stage or week can correspond to a calibration value of the normalization parameter. One or more calibration data points can be determined from multiple calibration samples having a known gestational age and containing free DNA molecules with a size greater than a cutoff value. In some embodiments, calibration data points are determined from functionally relevant gestational age having values of the normalization parameter.
[0453] At block 6060, gestational age can be determined using comparison. Gestational age can be considered as the age corresponding to the calibration value closest to the value of the normalized parameter. In some embodiments, gestational age can be considered as the highest age corresponding to the calibration value below the value of the normalized parameter.
[0454] The method may further include using ultrasound or the date of the woman's last menstrual period to determine a reference gestational age of the fetus. The method may also include comparing gestational age with a reference gestational age. The method may further include using the comparison of gestational age with a reference gestational age to determine a classification of the likelihood of pregnancy-related conditions. For example, a difference between gestational age and a reference gestational age can indicate a pregnancy-related condition. The difference may be a difference in gestational age at different gestational stages or at a minimum gestational week (e.g., 1, 2, 3, 4, 5, 6, 7 weeks or more).
[0455] The method may further include the use of terminal motifs. For example, the method may include determining a first subsequence corresponding to at least one end of a free DNA molecule having a size greater than a cutoff value. The first subsequence may belong to a free DNA molecule having a size greater than the cutoff value and having the first subsequence at one or more ends of the corresponding free DNA molecule. The first subsequence may be or contain 1, 2, 3, 4, 5, or 6 nucleotides. (As used...) Figure 52A and Figure 52B As described, terminal motifs can be used to determine gestational age via PCA analysis. Calibrated samples with different terminal motifs and known gestational ages that have undergone PCA analysis can be used. Other classification and regression algorithms, such as linear discriminant analysis, logistic regression, support vector machines, linear regression, and nonlinear regression, can be used on the terminal motifs. Classification and regression algorithms can correlate gestational age with specific terminal motifs and / or specific size segments.
[0456] The terminal motif can be represented by Figure 47-59 or Figure 94 Any motif discussed. The rank or frequency of a terminal motif can be compared to the rank or frequency of a terminal motif in a calibrated sample from individuals with known gestational age. The rank or frequency of the terminal motif can then be used to determine gestational age. Terminal motifs present at a rank or frequency that deviates from that determined from a reference sample with the same gestational age may indicate pregnancy-related conditions.
[0457] The value of the generated normalization parameter may include (a) normalizing the first amount by the total amount of free DNA molecules with a size greater than the cutoff value; (b) normalizing the first amount by the second amount of free DNA molecules with a size greater than the cutoff value and ending with a second subsequence, the second subsequence being different from the first subsequence; or (c) normalizing the first amount by the third amount of free DNA molecules with a size less than the cutoff value.
[0458] 2. Pregnancy-related conditions
[0459] Figure 61 Method 6100 illustrates the analysis of biological samples obtained from women carrying a fetus. Examples may include the possibility of classifying pregnancy-related conditions without determining gestational age. The biological sample may contain multiple cell-free DNA molecules from the fetus and the woman.
[0460] Sequence reads corresponding to multiple free DNA molecules are acceptable. In some embodiments, sequencing for obtaining the sequence reads may be performed.
[0461] At block 6120, the size of multiple free DNA molecules can be measured. The size can be compared with... Figure 21 The method described is similar to that used for measurement. Dimensions can be measured using the accepted sequence of reads.
[0462] At block 6130, a first quantity of cell-free DNA molecules having a size greater than a cutoff value can be measured. The cutoff value can be greater than or equal to 200 nt. The cutoff value can be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cutoff value can be any cutoff value described herein for long cell-free DNA molecules. The first quantity can be a number or a frequency.
[0463] At block 6140, a first value for a normalized parameter can be generated using a first quantity. Generating the normalized parameter value may involve measuring a second quantity containing free DNA molecules of a size smaller than a cutoff value; and calculating the ratio of the first quantity to the second quantity. The cutoff value may be a first cutoff value. The second cutoff value may be less than the first cutoff value. The second quantity may contain free DNA molecules of a size smaller than the second cutoff value, or the second quantity may contain all free DNA molecules among a plurality of free DNA molecules. The normalized parameter may be a measure of the frequency of long free DNA molecules.
[0464] At block 6150, a second value corresponding to the expected value of standardized parameters for a healthy pregnancy can be obtained. This second value may be determined based on the gestational age of the fetus. The second value may be the expected value. In some embodiments, the second value may be a cutoff value that distinguishes it from outliers.
[0465] Obtaining a second value may involve obtaining a calibration table of measurements of pregnant women with standardized parameters. The calibration table can be generated from a first table of measurements of individual pregnant women with gestational age. A second table of gestational age with standardized parameters and calibration values can be obtained. Data in the first and second tables may come from the same or different individuals. A calibration table of measurements with calibration values can be generated from the first and second tables. The calibration table may contain functions relating to the calibration values of the measurements.
[0466] Measurements of an individual pregnant woman can be taken as time since her last menstrual period or as image (e.g., ultrasound) features. For example, image features may include the length, size, appearance, or anatomy of the fetus. Features may include biometric measurements such as crown-rump length or femur length. The appearance of specific organs may be used, including the four-chambered heart or the vertebrae on the spinal cord. Gestational age can be determined by a practicing physician from ultrasound images (e.g., Committee on Obstetric Practice et al., "Methods for estimating the duedate," Committee Opinion, Vol. 700, May 2017).
[0467] In some embodiments, the machine learning model may associate one or more calibration data points with image features. The model can be trained by receiving multiple training images. Each training image may come from a female individual known to be free of pregnancy-related conditions or known to not have pregnancy-related conditions. The female individual may have a range of gestational ages. Training may involve storing multiple training samples from the female individuals. Each training sample may contain known values of normalized parameters associated with the training images. The model can be trained by optimizing the model's parameters using the output of the model based on matching or non-matching images with known values of normalized parameters from the multiple training samples. The model's output may specify values of normalized parameters corresponding to the images. A second value of the normalized parameters can be generated by inputting images of the female individuals into the machine learning model.
[0468] At block 6160, the deviation between the first and second values of the normalized parameter can be determined. The deviation can be the separation value.
[0469] At block 6170, a bias can be used to classify the likelihood of a pregnancy-related condition. A pregnancy-related condition is considered likely when the bias exceeds a threshold. The threshold can indicate a statistically significant difference. The threshold can indicate a difference of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100%.
[0470] Pregnancy-related conditions may include / include preeclampsia, intrauterine growth restriction, invasive placenta formation, preterm birth, neonatal hemolytic disease, placental insufficiency, fetal hydrops, fetal malformations, hemolysis, elevated liver enzymes and low platelet count (HELLP) syndrome or systemic lupus erythematosus.
[0471] IV. Size and distal analysis of pregnancy-related conditions
[0472] Size and / or end analysis of long DNA molecules was used to determine the likelihood of preeclampsia. This method can also be applied to other pregnancy-related conditions. DNA extracted from maternal plasma samples of four pregnant women diagnosed with preeclampsia underwent single-molecule real-time (SMRT) sequencing (PacBio).
[0473] Figure 62 This table displays clinical information for four cases of preeclampsia. The first column shows the case number. The second column shows the gestational age in weeks at the time blood sampling was stopped. The third column shows the sex of the fetus. The fourth column shows clinical information regarding preeclampsia (PET).
[0474] M12804 is a case of severe preeclampsia (PET) with pre-existing IgA nephropathy. M12873 is a case of chronic hypertension with superimposed mild PET. M12876 is a case of severe late-onset PET. M12903 is a case of severe late-onset PET with intrauterine growth restriction (IUGR). Five maternal plasma samples from late-pregnancy women with normal blood pressure were used as controls for subsequent analysis in this disclosure.
[0475] For the four preeclampsia cases and the five normal-blood-tension maternal plasma DNA samples from late pregnancy, DNA extracted from their paired maternal erythrocyte sedimentation rate (ESR) amber layer and placental samples was genotyped using the Infinium Omni2.5Exome-8 Beadchip system on the iScan system (Illumina).
[0476] Plasma DNA concentrations in each sample were quantified using high-sensitivity Qubit dsDNA analysis performed with a Qubit fluorometer (Thermo Fisher Scientific). The mean plasma DNA concentrations in preeclampsia cases and late pregnancy cases were 95.4 ng / mL (range 52.1–153.8 ng / mL) and 10.7 ng / mL (range 6.4–19.1 ng / mL), respectively. The mean plasma DNA concentration in preeclampsia cases was approximately nine times higher than that in late pregnancy cases.
[0477] For maternal plasma samples in late pregnancy with preeclampsia and normal blood pressure, the mean fetal DNA fractions determined from sequencing data of DNA molecules ≤600 bp containing informative single nucleotide polymorphisms (SNPs) in which the mother was homozygous and the fetus was heterozygous were 22.6% (range 16.6%–25.7%) and 20.0% (range 15.6%–26.7%), respectively.
[0478] A. Dimensional Analysis
[0479] Size analysis was performed on maternal plasma samples from preeclampsia and normal-blood-tension late pregnancy, according to embodiments of this disclosure. Figures 63A-63D and Figures 64A-64D This displays the size distribution of plasma DNA molecules from late-pregnancy cases with preeclampsia and normal blood pressure. The x-axis shows size. The y-axis shows frequency. For Figures 63A-63D Plot the size distribution of the x-axis on a linear scale in the range of 0kb to 1kb, and for Figures 64A-64D Plot the size distribution of the x-axis on a logarithmic scale in the range of 0kb to 5kb. Figure 63A and Figure 64A Sample M12804 is shown. Figure 63B and Figure 64B Sample M12873 is shown. Figure 63C and Figure 64C Sample M12876 is shown. Figure 63D and Figure 64D Sample M12903 is shown.
[0480] The blue line represents the size distribution of all sequenced plasma DNA molecules from five normal-blood-tension late-pregnancy cases. The red line represents the size distribution of sequenced plasma DNA molecules from a single case of preeclampsia. Figures 63A-63D In the diagram, the blue line represents the shorter peak at 200bp and the higher peak between 300bp and 400bp. Figures 64A-64D In the middle, the blue line corresponds to the line with the higher peak at 1kb.
[0481] Generally, the plasma DNA size profile of patients with preeclampsia is shorter than that of pregnant women in late pregnancy with normal blood pressure, with an increased height of the 166-bp peak and an increased proportion of DNA molecules shorter than 166bp. Figures 63A-63D These changes were more pronounced in the two severe preeclampsia cases, M12876 and M12903. The changes were even more dramatic in the preeclampsia case M12903, which also had intrauterine growth restriction (IUGR).
[0482] Three out of four preeclampsia plasma samples showed a proportion of reduced-size long plasma DNA molecules ranging from 200 to 5000 bp. Figure 64B-64DThe proportions of long plasma DNA molecules >500 bp in M12873, M12876, and M12903 were 11.7%, 8.9%, and 4.5%, respectively, while the proportion of long plasma DNA molecules in the composite sequencing data from five normotensive late-pregnancy cases was 32.3%. Compared with the composite sequencing data from five normotensive late-pregnancy cases, plasma samples from a case of severe preeclampsia (PET) with pre-existing IgA nephropathy (M12804) showed a reduced proportion of shorter DNA molecules <2000 bp, but an increased proportion of longer DNA molecules >2000 bp. Figure 2A The proportion of long plasma DNA molecules in M12804 was 34.9%.
[0483] Figure 65A-65D and Figures 66A-66D This diagram shows the size distribution of DNA molecules encompassing fetal-specific alleles from maternal plasma samples in late preeclampsia and normal-blood pressure pregnancies. Figures A through D represent different preeclampsia samples. The x-axis indicates size. Figure 65A-65D The y-axis displays frequency and Figures 66A-66D The y-axis in the graph displays the cumulative frequency. Figures 66A-66D The size ranges from 0kb to 35kb.
[0484] The blue lines in each figure represent the size distribution of all sequenced plasma DNA molecules encompassing fetal-specific alleles, synthesized from five cases of late pregnancy with normal blood pressure. The red lines in each figure represent the size distribution of sequenced plasma DNA molecules encompassing fetal-specific alleles, from individual cases of preeclampsia. Figure 65A-65D In the diagram, the blue line represents the shorter peak at 200bp and the higher peak between 300bp and 400bp. Figures 66A-66D In the diagram, the blue line corresponds to the lower peak between 100bp and 1000bp.
[0485] Figures 67A-67D and Figures 68A-68D This diagram shows the size distribution of DNA molecules encompassing fetal-specific alleles from maternal plasma samples in late preeclampsia and normal-blood pressure pregnancies. Figures A through D represent different preeclampsia samples. The x-axis indicates size. Figures 67A-67D The y-axis displays frequency and Figures 68A-68D The y-axis in the graph displays the cumulative frequency. Figures 68A-68D The size ranges from 0kb to 35kb.
[0486] The blue lines in each figure represent the size distribution of sequenced plasma DNA molecules encompassing maternal-specific alleles, synthesized from five normotensive late-pregnancy cases. The red lines in each figure represent the size distribution of sequenced plasma DNA molecules encompassing maternal-specific alleles, from individual preeclampsia cases. Figure 67A In the diagram, the blue line represents the peak at 200bp and the peak between 300bp and 400bp. Figures 67B-67D In the diagram, the blue line represents the shorter peak at 200bp. Figure 68A In the diagram, the blue line corresponds to the line with a higher peak between 1000bp and 10000bp. Figures 68B-68D In the diagram, the blue line corresponds to the lower peak between 100bp and 1000bp.
[0487] When compared with maternal plasma samples from late pregnancy with normal blood pressure, plasma DNA shortening was observed in three out of four preeclampsia plasma samples, covering fetal-specific alleles in DNA molecules. Figure 65B-65D and Figures 66B-66D ) and DNA molecules that encompass maternally specific alleles ( Figures 67B-67D and Figures 68B-68D Both were observed. An exception is case M12804, a severe PET study with pre-existing IgA nephropathy, which showed an increased proportion of shorter DNA molecules (less than 1 kb) and a decreased proportion of longer DNA molecules (greater than 1 kb) in plasma DNA molecules containing fetal-specific alleles. Figure 65A and Figure 66A In fact, plasma DNA molecules containing maternally specific alleles in case M12804 showed an elongated size profile. Figure 67A and Figure 68A ).
[0488] Figure 69A and Figure 69B This graph shows the proportion of short DNA molecules containing (A) fetal-specific alleles and (B) maternal-specific alleles, sequenced by PacBio SMRT sequencing in plasma samples from mothers with preeclampsia and normal blood pressure. The y-axis shows the proportion of short DNA fragments <150 bp. The x-axis shows normal and PET samples.
[0489] In this embodiment, the proportion of short DNA molecules was defined as the percentage of maternal plasma DNA molecules with a size less than 150 bp. M12804 was excluded from this analysis because this case had pre-existing IgA nephropathy, but other samples did not. When compared with the normotensive control plasma sample group, the preeclampsia plasma sample group showed a significantly increased proportion of short DNA molecules covering fetal-specific alleles (P = 0.036, Wilcoxon rank sum test) and a significantly increased proportion of short DNA molecules covering maternal-specific alleles (P = 0.036, Wilcoxon rank sum test).
[0490] Figure 70A and Figure 70B This graph shows the proportion of short DNA molecules sequenced by (A) PacBio SMRT sequencing and (B) Illumina sequencing in plasma samples from mothers with preeclampsia and normal blood pressure. The y-axis shows the proportion of short DNA fragments <150 bp.
[0491] In this embodiment, the proportion of short DNA molecules was defined as the percentage of maternal plasma DNA molecules smaller than 150 bp. M12804 was excluded from this analysis because this case may have shown a different size profile compared to other preeclampsia cases in this cohort due to the pre-existing IgA nephropathy present in this case. The preeclampsia plasma cohort showed a significantly increased proportion of short DNA molecules (median: 28.0%; range: 25.8%–35.1%) compared to the normotensive control plasma cohort (median: 12.1%; range: 8.5%–15.8%) (P = 0.036, Wilcarson rank-sum test). Conversely, in the previous cohorts of four preeclampsia and four gestationally age-matched normotensive maternal plasma DNA samples that underwent bisulfite conversion and Illumina sequencing, the proportion of short DNA molecules in preeclampsia plasma and control plasma samples was not significantly different (P = 0.340, Wilcarson rank-sum test). Figure 70B ).
[0492] In some embodiments, a 20% cutoff value for the proportion of short DNA molecules in a maternal plasma sample sequenced by PacBio SMRT sequencing can be used to determine whether a pregnant woman is at high or low risk of developing preeclampsia. Maternal plasma samples with a proportion of short DNA molecules higher than 20% will be identified as being at high risk of preeclampsia, while those with a proportion of short DNA molecules lower than 20% will be identified as being at low risk. With this cutoff value, both sensitivity and specificity are 100%. In some other embodiments, the cutoff value for the proportion of short DNA molecules used may include, but is not limited to, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, etc. In another embodiment, the proportion of short DNA molecules in the maternal plasma sample will be used to monitor and assess the severity of preeclampsia during pregnancy.
[0493] In the embodiments, the size ratio of the relative proportion of indicator short DNA molecules to long DNA molecules in each sample is calculated using the following equation.
[0494]
[0495] P(50-150) indicates the proportion of sequenced plasma DNA molecules with a size in the range of 50bp to 150bp; and P(200-1000) indicates the proportion of sequenced plasma DNA molecules with a size in the range of 200bp to 1000bp.
[0496] Figure 71 A plot showing the size ratio of short to long DNA molecules in plasma samples from preeclampsia and normotensive mothers sequenced by PacBio SMRT sequencing. The y-axis shows the size ratio. The x-axis shows normal and PET samples. The preeclampsia plasma sample group showed a significantly higher size ratio compared to the normotensive control plasma sample group (P = 0.016, Wilcarson rank-sum test).
[0497] In embodiments, we can utilize size profiles generated by long-read sequencing platforms including, but not limited to, PacBio SMRT sequencing and Oxford Nanopore sequencing to predict the development and severity of preeclampsia in pregnant women. In some embodiments, we can monitor the development of preeclampsia and the development of severe preeclampsia characteristics, including but not limited to liver and kidney damage, by analyzing the size profile of plasma DNA molecules. In some embodiments, the size parameters used in the analysis may include, but are not limited to, the proportion of short or long DNA molecules and size ratios indicating the relative proportion of short to long DNA molecules. Cutoff values used to determine short and long DNA categories may include, but are not limited to, 150bp, 180bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 1kb, etc. The size range used to determine the size ratio of short molecules to long molecules may include, but is not limited to, 50-150bp, 50-166bp, 50-200bp, 200-400bp, 200-1000bp, 200-5000bp, or other combinations.
[0498] Size end analysis may include usage Figure 61 The method described in method 6100.
[0499] B. Fragment End Analysis
[0500] Fragment end analysis was performed on maternal plasma samples from preeclampsia and normal-blood-tension late pregnancy, according to embodiments of this disclosure. The first nucleotide at the 5' end of both Watson and Crick fragments was determined for each sequenced plasma DNA molecule. The proportions of T-terminal, C-terminal, A-terminal, and G-terminal fragments were determined for each plasma DNA sample.
[0501] Figures 72A-72D This displays the proportions of different ends of plasma DNA molecules in plasma samples from mothers with preeclampsia and normal blood pressure, sequenced using PacBio SMRT sequencing. The x-axis shows samples from normal late pregnancy and PET samples. The y-axis shows predetermined end proportions. Figure 72A Display the T-end ratio. Figure 72B Displays the C-end ratio. Figure 72C Display the proportion of end A. Figure 72D The proportion of G-terminal DNA molecules was shown. Compared with the normotensive control plasma sample group, the preeclampsia plasma sample group showed a significantly increased proportion of T-terminal plasma DNA molecules (P = 0.016, Wilcarson rank-sum test) and a significantly decreased proportion of G-terminal plasma DNA molecules (P = 0.016, Wilcarson rank-sum test).
[0502] Figure 73 This displays a hierarchical cluster analysis of maternal plasma DNA samples from preeclampsia and normal-blood pressure late pregnancy using four types of fragment ends (the first nucleotide at the 5' end of each strand): C-terminus, G-terminus, T-terminus, and A-terminus. Each column indicates the plasma DNA sample. The first column indicates the group to which each sample belongs, with cyan indicating maternal plasma DNA samples from normal-blood pressure late pregnancy and orange indicating preeclampsia plasma DNA samples. Cyan covers the first five columns. Orange covers the last four columns.
[0503] Starting from the second column, each column indicates the type of fragment ends. Based on column-normalized frequencies (z-scores) (i.e., the standard deviation of frequencies below or above the mean frequency across the entire sample), the terminal motif frequencies are presented as a series of color gradients. The redder the color, the higher the terminal motif frequency, and the bluer the color, the lower the terminal motif frequency. Hierarchical cluster analysis based on the four types of fragment end frequencies showed that the fragment end profile of preeclampsia plasma DNA samples formed distinct clusters compared to plasma DNA samples from late-pregnancy women with normal blood pressure.
[0504] In the embodiments, we can determine the dinucleotide sequences of the first nucleotide (X) and the second nucleotide (Y) at the 5' end of both the Watson and Crick stocks for each sequenced DNA molecule. X and Y can be one of the four nucleotide bases in DNA. There are 16 possible dinucleotide terminal motifs XYNN, namely AANN, ATNN, AGNN, ACNN, TANN, TTNN, TGNN, TCNN, GANN, GTNN, GGNN, GCNN, CANN, CTNN, CGNN, and CCNN. We can determine the dinucleotide sequences of the third nucleotide (X) and the fourth nucleotide (Y) at the 5' end of both the Watson and Crick stocks for each sequenced DNA molecule according to the embodiments of this disclosure. There are 16 possible dinucleotide motifs NNXY. We can also determine the first tetranucleotide sequence (tetrameric motif) at the 5' end of both the Watson and Crick stocks for each sequenced DNA molecule.
[0505] Figure 74 Hierarchical cluster analysis of maternal plasma DNA samples from preeclampsia and normal-blood-tension late pregnancy using 16 dinucleotide motifs XYNN (dinucleotide sequences of the first and second nucleotides at the 5' end). Figure 75 Hierarchical cluster analysis of maternal plasma DNA samples from preeclampsia and normal-blood-tension late pregnancy using 16 dinucleotide motifs NNXY (dinucleotide sequences of the third and fourth nucleotides at the 5' end). Figure 76This shows a hierarchical cluster analysis of maternal plasma DNA samples from preeclampsia and normal-blood-tension late pregnancy using 256 tetranucleotide motifs (dinucleotide sequences from the first to the fourth nucleotide at the 5' end).
[0506] exist Figures 74-76 In the diagram, the first column indicates which group each sample belongs to, with cyan indicating maternal plasma DNA samples from late pregnancy with normal blood pressure and orange indicating plasma DNA samples from preeclampsia. Cyan covers the first five columns, and orange covers the last four. Starting from the second column, each column indicates the type of fragment ends. Based on the column's standardized frequency (z-score) (i.e., the standard deviation of the frequency below or above the mean frequency in the entire sample), the terminal motif frequencies are presented as a series of color gradients. The redder the color, the higher the terminal motif frequency, and the bluer the color, the lower the terminal motif frequency.
[0507] These results indicate that plasma DNA in preeclampsia and non-preeclampsia samples exhibits different fragmentation characteristics. In one embodiment, we can utilize terminal motif profiles generated by long-read sequencing platforms including, but not limited to, PacBio SMRT sequencing and Oxford Nanopore sequencing to predict the development of preeclampsia in pregnant women. Although single nucleotide motifs, dinucleotide motifs, and tetranucleotide motifs were used in the analysis above, motifs of other lengths, such as 3, 5, 6, 7, 8, 9, 10, or longer, may be used in other embodiments.
[0508] In some embodiments, fragment end analysis and origin tissue analysis can be combined to improve the efficacy of predicting, detecting, and monitoring pregnancy-related conditions, including but not limited to preeclampsia. First, fragment end analysis can be performed on each maternal plasma sample to separate plasma DNA molecules into four fragment end categories: T-terminal, C-terminal, A-terminal, and G-terminal fragments. Subsequently, methylation status matching analysis can be used, according to embodiments of this disclosure, to perform origin tissue analysis on each maternal plasma DNA sample individually using plasma DNA molecules from each of the fragment end categories. The contribution proportion of a different tissue within one fragment end category is defined as the percentage of plasma DNA molecules in the corresponding fragment end category assigned to the corresponding tissue relative to other tissues.
[0509] We used single-molecule real-time sequencing to analyze plasma DNA samples from three and five pregnant women with and without preeclampsia. We obtained median 658,722, 889,900, 851,501, and 607,554 plasma fragments with A-terminals, C-terminals, G-terminals, and T-terminals, respectively. For fragments with A-terminals, we compared the methylation patterns of any fragment with at least 10 CpG sites with reference methylation profiles of neutrophils, T cells, B cells, liver, and placenta, according to the methylation status matching pathway described in this disclosure. Plasma DNA fragments were assigned to the tissues corresponding to the highest fraction of methylation status matches among those tissues. Using this method, a median of 2.43% (range: 0.73%–5.50%) of A-terminal fragments were assigned to T cells (i.e., T-cell contribution) in all samples analyzed. We further analyzed fragments with C-terminals, G-terminals, and T-terminals in a similar manner. For fragments with C-terminus, G-terminus, and T-terminus, median T-cell contributions were observed at 3.20% (range: 1.55%–5.19%), 3.52% (range: 1.53%–6.27%), and 2.22% (0%–7.79%), respectively.
[0510] Figures 77A-77D This diagram shows the T-cell contribution in DNA molecules belonging to different segment end categories—(A)T-terminus, (B)C-terminus, (C)A-terminus, and (D)G-terminus—in maternal plasma DNA samples from women with preeclampsia and normal blood pressure. The x-axis shows samples from normal late pregnancy and PET samples. The y-axis shows T-cell contribution as a percentage. Results showed that, within the G-terminal segments, the T-cell contribution in preeclampsia plasma samples was significantly reduced compared to plasma samples from normal late pregnancy (P = 0.036, Wilcoxon rank-sum test). In this example, we can use a 3% cutoff value for T-cell contribution in all G-terminal segments of the maternal plasma DNA sample to determine whether a pregnant woman is at high or low risk of developing preeclampsia.
[0511] C. Instance Methods
[0512] Figure 78 Method 7800 shows the analysis of biological samples obtained from women carrying a fetus. The biological sample may contain multiple cell-free DNA molecules from the fetus and the woman. The method may generate a classification of possible pregnancy-related conditions. Pregnancy-related conditions may include preeclampsia or any pregnancy-related condition described herein.
[0513] Sequence reads corresponding to multiple free DNA molecules are acceptable.
[0514] At block 7810, the size of multiple free DNA molecules can be measured. Size can be determined by aligning nucleotides or counting the number of nucleotides or by inclusion. Figure 21Any technique described herein shall be used for measurement.
[0515] At block 7820, a group of cell-free DNA molecules with sizes larger than the cutoff value can be identified. The cutoff value can be any cutoff value used for long cell-free DNA fragments, including 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cutoff value can be any cutoff value described herein for long cell-free DNA molecules.
[0516] At block 7830, a first amount can be used to generate a value for the terminal motif parameter. A first amount of free DNA molecules in the group having a first subsequence at one or more ends can be measured. In some embodiments, the terminal motif parameter may be a first amount normalized by the total amount of all subsequences at an end. In some embodiments, the end may be a 3' end. In some embodiments, the end may be a 5' end.
[0517] The length of the first subsequence can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. The first subsequence can be included in the last nucleotide at the end of the corresponding free DNA molecule. For example, the first subsequence could be... Figure 74 The XYNN pattern is shown. In some embodiments, the first subsequence may not include the last nucleotide or nucleotides at the end of the corresponding free DNA molecule. For example, the first subsequence may include... Figure 75 The NNXY pattern.
[0518] A second quantity can be measured for a cell-free DNA molecule that has a subsequence at one or more ends that differs from the first subsequence. The value of the terminal motif parameter can be generated using the ratio of the second quantity to the third quantity. For example, the second quantity can be divided by the third quantity, or the third quantity can be divided by the second quantity.
[0519] At block 7840, the value of the terminal motif parameter can be compared to a threshold. The threshold can be a value representing a statistically significant difference relative to the value of the relevant parameter in individuals without pregnancy-related conditions. The threshold can be determined from one or more reference individuals with normal pregnancies or one or more reference individuals with pregnancy-related conditions.
[0520] In some embodiments, the value of the terminal motif parameter can be compared to a threshold, and the value of the second terminal motif parameter can be compared to a second threshold. A second quantity of a free DNA molecule having a second subsequence different from the first subsequence at one or more ends can be measured. Thus, the quantity of different terminal motifs can be determined. The value of the second terminal motif parameter can be generated using the second quantity. The value of the second terminal motif parameter can be compared to a second threshold. The second threshold can be the same as or different from the first threshold. Additional subsequences can be used in the same manner as the first and second subsequences. In some embodiments, all possible subsequences can be used for comparison with the threshold.
[0521] At block 7850, a classification of the likelihood of a pregnancy-related condition can be determined using comparison. A pregnancy-related condition is considered likely when the value of the size parameter or the terminal motif parameter exceeds a threshold.
[0522] In some embodiments, the classification of the likelihood of a pregnancy-related condition can be performed by comparing the value of a second terminal motif parameter with a second cutoff value. A pregnancy-related condition may be considered likely when the value of a first terminal motif parameter exceeds a first threshold and the value of a second terminal motif parameter exceeds a second threshold.
[0523] The method may include the use of size parameters other than terminal motif parameters. A second group of cell-free DNA molecules having sizes within a first size range can be identified. The first size range may include sizes larger than a cutoff value. The first size range may be smaller than 550 nt, 600 nt, 650 nt, 700 nt, 750 nt, 800 nt, 850 nt, 900 nt, 950 nt, 1 nt, 1.5 knt, 2 knt, 3 knt, 5 knt, or larger. The value of the size parameter can be generated using a second amount of cell-free DNA molecules from the second group. The value of the size parameter can be compared to a second threshold. The classification of the likelihood of a pregnancy-related condition can be determined using the comparison of the size parameter value to the second threshold. When one or both of the first and second thresholds are exceeded, the classification may indicate a pregnancy-related condition.
[0524] Size parameters can be standardized parameters. For example, a third quantity of cell-free DNA molecules within a second size range can be measured. The second size range can include sizes smaller than a first cutoff value. The second size range can include all sizes. The second size range can include 50-150 nt, 50-166 nt, 50-200 nt, and 200-400 nt. The second size range can include any size described herein for short cell-free DNA fragments. The second size range can exclude sizes within the first size range. The value of the size parameter can be generated by determining the ratio of the second quantity to the third quantity. For example, the second quantity can be divided by the third quantity, or the third quantity can be divided by the second quantity.
[0525] Any of the stated amounts of free DNA molecules may be free DNA molecules derived from a specific tissue of origin. For example, the tissue of origin may be a T cell or another tissue of origin described herein. The second amount may be used with... Figures 77A-77D The T-cell contribution described is similar. The contribution of the tissue of origin can be determined using the methylation status or pattern as described in this disclosure.
[0526] V. Diseases related to repeat sequence extension
[0527] Long cell-free DNA fragments obtained from pregnant women can be used to identify repetitive sequence extensions in genes. Repetitive sequence extensions in genes can lead to neuromuscular diseases. Tandem repeat extensions have been associated with human diseases including, but not limited to, neurodegenerative disorders such as Fragile X syndrome, Huntington's disease, and spinocerebellar ataxia. These tandem repeat extensions can occur in protein-coding regions of genes (Machado-Joseph disease, Haw River syndrome, Huntington's disease) or non-coding regions (Friedrichataxia, myotonic dystrophy, some forms of Fragile X syndrome). Extensions involve microsatellites, pentanucleotides, tetranucleotides, and many trinucleotide repeats have been associated with fragility sites. Extensions associated with these diseases may be caused by replication slippage, asymmetric recombination, or epigenetic aberrations. The number of repeat sequences in a sequence refers to the total number of occurrences of the subsequence. For example, “CAGCAG” contains two repeat sequences. Because a repeat sequence contains at least two instances of a subsequence, the number of repeat sequences cannot be one. A subsequence can be understood as a repeating unit.
[0528] In this embodiment, analysis of long cell-free DNA in pregnant women can facilitate the detection of diseases associated with repetitive sequences. For example, a trinucleotide repeat sequence represents a repetitive extension of a 3 bp motif in a DNA sequence. One example is the sequence 'CAGCAGCAG', which comprises three 3 bp 'CAG' motifs. Microsatellite extensions, typically trinucleotide repeat sequence extensions, have been reported to play a key role in neurological disorders (Kovtun et al., *Cell Research*, 2008; 18:198-213; McMurray et al., *Nature Reviews Genetics*, 2010; 11:786-99). One example is the pathogenicity of more than 55 CAG repeat sequences (totaling 165 bp) in the ATXN3 gene, leading to spinocerebellar ataxia type 3 (SCA3) characterized by progressive motor problems. This condition is inherited in an autosomal dominant pattern. Therefore, a single copy of the altered gene is sufficient to cause the aforementioned condition. To determine the number of repetitive sequences in microsatellites, polymerase chain reaction (PCR) is typically used to amplify the genomic region of interest, and the PCR products are then subjected to several different techniques, including capillary electrophoresis (Lyon et al., *Journal of Molecular Diagnostics*, 2010; 12:505-11), southern ink dot analysis (Hsiao et al., *Journal of Clinical Laboratory Analysis*, 1999; 13:188-93), melting curve analysis (Lim et al., *Journal of Molecular Diagnostics*, 2014; 17:302-14), and mass spectrometry (Zhang et al., *Analytical Methods*, 2016; 8:5039-44). However, these methods are labor-intensive, time-consuming, and difficult to apply to high-throughput screening in real-world clinical practice such as prenatal testing. Sanger sequencing presents considerable challenges in inferring long repetitive sequences from complex sequence traces through manual examination. It is well known that Illumina sequencing technology and Ion Torrent have considerable difficulty in sequencing GC-rich (or GC-poor) regions with those repetitive sequences (Ashely et al. 2016; 17:507-22), and the length of DNA including expanded repetitive sequences is prone to exceed the read length (Loomis et al. Genome Research 2013; 23:121-8).
[0529] Another example is myotonic dystrophy and autosomal dominant disorders caused by the expansion of CTG repeat sequences in the range of 50 to 4000 CTG repeat sequences adjacent to the DMPK gene. Molecular diagnosis of DM is routinely performed in prenatal diagnosis by invasively analyzing the number of CTGs on the fetal genomic DNA.
[0530] In contrast to short-read sequencing (hundreds of bases), the method described in this disclosure is able to obtain long DNA molecules (multiple kilobases) from maternal plasma DNA. We can use the method described in this disclosure to non-invasively determine whether an unborn fetus has inherited the disease from an affected mother.
[0531] Figure 79 A diagram illustrating the inference of maternal inheritance in the fetus for diseases related to repetitive sequences. In stage 7905, cell-free DNA from the pregnant woman undergoes single-molecule real-time sequencing (e.g., PacBio SMRT). In stage 7910, the sequencing results are classified into long DNA and short DNA categories according to this disclosure. In stage 7915, allele information present in long DNA molecules can be used to construct maternal haplotypes, namely Hap I and Hap II. Hap I and Hap II may each contain expanded repetitive sequences of trinucleotide subsequences (e.g., CTG). In stage 7920, haplotype imbalances can be analyzed, and compared with... Figure 16 Similarly, in stage 7925, maternal inheritance in the fetus can be inferred. According to this disclosure, the method described herein allows us not only to identify haplotypes (e.g., Hap I and Hap II), but also to use sequence information of long DNA molecules to determine which haplotype has the disease-causing expanded repetitive sequence (e.g., affected Hap I). In this example, we can determine whether the fetus has inherited maternal Hap I (affected) or Hap II (unaffected) using the count, size, or methylation status of short DNA molecules distributed throughout maternal Hap I and Hap II, according to the method described herein.
[0532] Figure 80 This diagram illustrates the inference of paternal inheritance in fetuses for diseases related to repetitive sequences. We can use cell-free DNA from the pregnant woman to determine whether the fetus has inherited an affected paternal haplotype. Figure 80 As shown, cell-free DNA (e.g., 5 CTG repeats for Hap I and 6 CTG repeats for Hap II) from unaffected pregnant women whose husbands are affected by a disease involving extended repeat sequences (e.g., 70 CTG repeats) is subjected to PacBio SMRT sequencing. The sequenced long DNA molecules are identified and used to determine haplotypes and the number of repeat sequences. If a haplotype with a long extension of CTG repeat sequences (e.g., 70 CTG repeats in this example) is present in the maternal plasma of the unaffected pregnant woman, it indicates that the fetus has inherited the affected paternal haplotype. In some embodiments, the DNA containing the extended repeat sequences also carries one or more other paternal-specific alleles not present in the maternal genome. This would be useful for confirming paternal inheritance.
[0533] In another embodiment, we can use cell-free DNA from the pregnant woman's body to determine whether the fetus has inherited an affected paternal haplotype. For example... Figure 80 As shown, cell-free DNA (e.g., 5 CTG repeats for Hap I and 6 CTG repeats for Hap II) from unaffected pregnant women whose husbands are affected by a disease involving extended repeat sequences (e.g., 70 CTG repeats) is subjected to PacBio SMRT sequencing. The sequenced long DNA molecules are identified and used to determine haplotypes and the number of repeat sequences. If a haplotype with a long extension of CTG repeat sequences (e.g., 70 CTG repeats in this example) is present in the maternal plasma of the unaffected pregnant woman, it indicates that the fetus has inherited the affected paternal haplotype. In some embodiments, the DNA containing the extended repeat sequences also carries one or more other paternal-specific alleles not present in the maternal genome. This would be useful for confirming paternal inheritance.
[0534] Figure 81 , Figure 82 and Figure 83 This table displays instances of diseases caused by repetitive sequence expansion. The first column shows diseases associated with repetitive sequence expansion. The second column shows the repetitive subsequences. The third column shows the number of repetitive sequences in normal individuals. The fourth column shows the number of repetitive sequences in affected individuals. The fifth column shows the genetic location associated with the repetitive sequence. The sixth column lists the gene names. The seventh column lists the inheritance patterns. The table is sourced from omicslab.genetics.ac.cn / dred / index.php.
[0535] A. An example of repeat sequence expansion detection
[0536] According to reports, paternally inherited expanded CAG repeat sequences can be detected directly via PCR in maternal plasma and subsequently analyzed on a 3130XL genetic analyzer (Oever et al., Prenatal Diagn. 2015; 35:945-9). Non-invasive prenatal testing for Huntington's disease can be achieved via PCR because the size of expanded alleles only starts from >35 trinucleotide repeat sequences [i.e., DNA regions spanning repeats of 105 bp (35 × 3) or longer]. Many expanded repeat sequences, especially those for most trinucleotide repeat disorders (Orr et al., Annu. Rev. Neurosci. 2007; 30:575-621), involve repeat sequences of 300 bp or longer than the size of short fetal DNA molecules previously reported. DNA with large expanded repeat sequences will pose challenges for PCR (Orr et al., Annu. Rev. Neurosci. 2007; 30:575-621). As demonstrated by Once et al., long CAG repeat sequences often exhibit significantly lower signal strengths compared to smaller repeat sequences, a phenomenon observed in both genomic and plasma DNA, resulting in lower sensitivity for detecting these long CAG repeat sequences (Oever et al., Prenatal Diagnosis 2015; 35:945-9). Another limitation of PCR is the inability to preserve methylation signals during amplification. In one embodiment, single-molecule real-time sequencing of long DNA molecules would allow for the determination of tandem repeat polymorphisms and their associated methylation levels in one or more regions.
[0537] Figure 84 This table displays examples of repetitive sequence expansion detection and repetitive sequence-associated methylation determination in the fetus. The first column shows the type of repetitive sequence in multiple base pairs. The second column shows the repetitive unit. The third column shows the genomic location. The fourth column shows the reference base, i.e., the sequence present in the human reference genome. The fifth column shows the paternal genotype. The sixth column shows the maternal genotype. The seventh column shows the fetal genotype. The eighth column shows the degree of fetal DNA methylation associated with the paternal allele. The ninth column shows the degree of fetal DNA methylation associated with the maternal allele.
[0538] Figure 84Multiple instances of 1bp, 2bp, 3bp, and 4bp tandem repeat sequences are shown. For example, a “GATA” tandem repeat sequence is identified at the genomic location chr3:192384705-192384706. The father at this locus has a genotype of T(GATA)3 / T(GATA)5, where allele 1 has 3 repeat units and allele 2 has 5 repeat units. The paternal allele 2 indicates a genetic event involving repeat sequence expansion compared to the reference allele T(GATA)3. The mother at this locus has a genotype of T / T, showing a genetic event involving repeat sequence contraction. The fetal genotype at this locus is T(GATA)5 / T, indicating that the fetus inherits paternal allele 2 (i.e., T(GATA)5) and maternal allele T. The methylation levels associated with the paternal and maternal alleles are 50.98% and 62.8%, respectively. These results indicate that the use of tandem repeat polymorphisms will allow for the determination of maternal and paternal inheritance in the fetus. This technique will allow for the identification of different methylation patterns associated with two alleles. Another example shows that at the genomic location chr4:73237157-73237158, the fetus has inherited a repeat sequence extension from the mother [(TAAA)3]. Fetal molecules containing repeat sequence extensions inherited from the mother showed a higher level of methylation (95.65%) compared to fetal molecules containing paternal alleles (62.84%). These data demonstrate that we can detect repeat sequences, repeat sequence structures, and associated methylation changes. In one embodiment, we can use a specific cutoff value to determine whether the methylation difference between maternal and paternal inheritance is significant. The cutoff value will be the absolute difference among methylation levels greater than, but not limited to, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or 90%. Maternal inheritance can be determined using... Figure 21 The method described in method 2100 is similar.
[0539] B. Instance Methods
[0540] Subsequence repeat sequences can be used to determine fetal information. For example, the presence of subsequence repeat sequences can help determine that a molecule is of fetal origin. Additionally, subsequence repeat sequences can indicate the likelihood of genetic disorders. Subsequence repeat sequences can be used to determine the inheritance of maternal and / or paternal haplotypes. Furthermore, fetal kinship can be determined using subsequence repeat sequences.
[0541] 1. Fetal origin analysis using subsequence repeat sequences
[0542] Figure 85Method 8500 shows the analysis of biological samples obtained from women carrying fetuses, the biological samples containing cell-free DNA molecules from both the fetus and the woman. The possibility of genetic diseases in the fetus can be determined.
[0543] At block 8510, a first sequence read corresponding to one of the cell-free DNA molecules can be received. The cell-free DNA molecule may have a length greater than a cutoff value. The cutoff value may be greater than or equal to 200 nt. The cutoff value may be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cutoff value may be any cutoff value described herein for long cell-free DNA molecules.
[0544] At step 8520, the first sequence read can be aligned with a region of the reference genome. This region is known to potentially contain repetitive sequences of the subsequence. The region may correspond to... Figures 81-83 The subsequence can be a trinucleotide sequence, including any trinucleotide sequence described herein.
[0545] At block 8530, the number of repeating sequences corresponding to the subsequence in the first read of the free DNA molecule can be identified.
[0546] At block 8540, the number of repeat sequences in the subsequence can be compared to a threshold number. The threshold number can be 55, 60, 75, 100, 150, or more. The threshold number can vary depending on the genetic condition. For example, the threshold could reflect the minimum number of repeat sequences in an affected individual, the maximum number of repeat sequences in a normal individual, or a number between these two (see [link to relevant documentation]). Figures 81-83 ).
[0547] At block 8550, a comparison of the number of repeat sequences to a threshold number can be used to determine the classification of the fetus's likelihood of having a genetic condition. When the number of repeat sequences exceeds the threshold number, the fetus can be identified as potentially having a genetic condition. The genetic condition could be fragile X syndrome or... Figures 81-83 Any of the conditions listed herein.
[0548] In some embodiments, the method may include classifying multiple different target loci, each known to potentially contain repetitive sequences of subsequences. Multiple sequence reads corresponding to cell-free DNA molecules may be received. The multiple sequence reads may be aligned to multiple regions of a reference genome. The multiple regions are known to potentially contain repetitive sequences of subsequences. The multiple regions may be non-overlapping regions. Each region among the multiple regions may have different SNPs. The multiple regions may originate from different chromosome arms or chromosomes. The multiple regions may cover at least 0.01%, 0.1%, or 1% of the reference genome. The number of repetitive sequences of subsequences in the multiple sequence reads may be identified. The number of repetitive sequences of subsequences may be compared to multiple threshold numbers. Each threshold number may indicate the presence or likelihood of different genetic conditions. For each of multiple genetic conditions, a classification of the likelihood of a fetus having the corresponding genetic condition may be determined using a comparison with one of the multiple threshold numbers.
[0549] Cell-free DNA molecules can be determined to be of fetal origin. Determining fetal origin may involve receiving a second sequence read corresponding to a maternally derived cell-free DNA molecule obtained from a pre-pregnancy erythrocyte sedimentation rate (ESR) amber layer or sample. The second sequence read can be aligned to the region of a reference genome. A second number of repetitive sequences within the second sequence read can be identified. It can be determined that the second number of repetitive sequences is less than the first number of repetitive sequences.
[0550] Determining fetal origin may involve determining the degree of methylation of a cell-free DNA molecule using methylated and unmethylated sites. The degree of methylation may be compared to a reference level. The method may include determining if the degree of methylation exceeds a reference level. The degree of methylation may be the number or proportion of methylated sites.
[0551] Determining fetal origin may involve identifying methylation patterns at multiple sites of the free molecule. A similarity score can be determined by comparing the methylation patterns to reference patterns from maternal or fetal tissue. The similarity score can be compared to one or more thresholds. The similarity score can be any similarity score described herein, including, for example, the similarity score described using method 4000.
[0552] 2. Kinship analysis using subsequence repeat sequences
[0553] Figure 86 Method 8600 shows the analysis of biological samples obtained from women carrying fetuses, the biological samples containing cell-free DNA molecules from both the fetus and the woman. The biological samples can be analyzed to determine the father of the fetus.
[0554] At block 8610, a first sequence read corresponding to one of the cell-free DNA molecules can be received. The method may include determining that the cell-free DNA molecule is of fetal origin. The cell-free DNA molecule can be determined to be of fetal origin by any of the methods described herein, including, for example, those described with method 8500. The cell-free DNA molecule may have a size greater than a cutoff value. The cutoff value may be greater than or equal to 200 nt. The cutoff value may be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cutoff value may be any cutoff value described herein for long cell-free DNA molecules.
[0555] At block 8620, the first sequence read can be aligned with the first region of the reference genome. It is known that the first region may contain repetitive sequences of subsequence.
[0556] At block 8630, the first number of repeating sequences corresponding to the first subsequence in the first read of the free DNA molecule can be identified. The first subsequence may contain alleles.
[0557] At block 8640, sequence data obtained from a male individual can be analyzed to determine whether a second number of repetitive sequences of the first subsequence exist in the first region. The second number of repetitive sequences contain at least two instances of the first subsequence. The sequence data can be obtained by extracting a biological sample from a male individual and performing sequencing on the DNA in the biological sample.
[0558] At block 8650, the possibility of a male individual being the father of the fetus can be classified by determining whether a second number of repeating sequences of the first subsequence exist. This classification could be that the male individual is likely the father when the second number of repeating sequences of the first subsequence are present, or that the male individual is unlikely to be the father when the second number of repeating sequences of the first subsequence are absent.
[0559] The method may include comparing a first number of repeating sequences with a second number of repeating sequences. Classifying the probability that a male individual is the father may include using the comparison of the first number of repeating sequences with the second number of repeating sequences. The classification may be that the male individual is likely the father when the first number of repeating sequences falls within a threshold of the second number of repeating sequences. The threshold may be within 10%, 20%, 30%, or 40% of the second number of repeating sequences.
[0560] The method may include using multiple regions of repetitive sequences. For example, a cell-free DNA molecule is a first cell-free DNA molecule. The method may include receiving a second sequence read corresponding to a second cell-free DNA molecule in the said cell-free DNA molecule. The method may also include aligning the second sequence read to a second region of a reference genome. The method may further include identifying a first number of repetitive sequences corresponding to a second subsequence in the second sequence read of the second cell-free DNA molecule. The method may include analyzing sequence data obtained from a male individual to determine whether a second number of repetitive sequences of the second subsequence are present in the second region. Classifying the probability that the male individual is the father of the fetus may further include using the determination of whether a second number of repetitive sequences of the second subsequence are present in the second region. The probability classification may be that the male individual is more likely to be the father of the fetus when the repetitive sequences are present in both the first and second regions in the male individual's sequence data.
[0561] VI. Size selection for enriching long plasma DNA molecules
[0562] In our embodiments, we can physically select DNA molecules having one or more desired size ranges prior to analysis (e.g., single-molecule real-time sequencing). As an example, size selection can be performed using a solid-phase reversible immobilization technique. In other embodiments, size selection can be performed using electrophoresis (e.g., using a Coastal Genomic system or a Pippin size selection system). Our approach differs from previous practices that primarily focused on shorter DNA (Li et al., JAMA 2005; 293:843-9), where fetal DNA is known in the art to be shorter than maternal DNA (Chan et al., Clinical Chemistry 2004; 50:88-92).
[0563] Size selection techniques can be applied to any of the methods described herein and to any size described herein. For example, free DNA molecules can be enriched by electrophoresis, magnetic beads, hybridization, immunoprecipitation, amplification, or CRISPR. The resulting enriched sample may have a higher concentration or a higher proportion of specific size fragments than the original biological sample.
[0564] A. Size selection using electrophoresis
[0565] In this embodiment, taking advantage of the fact that the electrophoretic mobility of DNA depends on the DNA size, we can use a gel electrophoresis-based approach to select target DNA molecules having a desired size range, such as, but not limited to, ≥100bp, ≥200bp, ≥300bp, ≥400bp, ≥500bp, ≥600bp, ≥700bp, ≥800bp, ≥900bp, ≥1kb, ≥2kb, ≥3kb, ≥4kb, ≥5kb, ≥6kb, ≥7kb, ≥8kb, ≥9kb, ≥10kb, ≥20kb, ≥30kb, ≥40kb, ≥50kb, ≥60kb, ≥70kb, ≥80kb, ≥90kb, ≥100kb, ≥200kb; or other size ranges, including sizes larger than any of the cutoff values described herein. For example, the LightBench (Coastal Genomics) automated gel electrophoresis system was used for DNA size selection. In principle, shorter DNA molecules will move faster than longer DNA molecules during gel electrophoresis. We applied this size selection technique to a plasma DNA sample (M13190) to select DNA molecules larger than 500 bp. We used a 3% size selection cartridge with an 'in-channel filter' (ICF) collection device and a loading buffer with internal size markers for size selection. The DNA library was loaded into the gel and electrophoresis was started. When the target size was reached, the first segment <500 bp was retrieved from the ICF. Operation was resumed and allowed to complete electrophoresis to obtain the second segment ≥500 bp. We used single-molecule real-time sequencing (PacBio) to sequence the second segment with a molecular size ≥500 bp. We obtained 1,434 high-quality circular common sequences (CCS) (i.e., 1,434 molecules). Of these, 97.9% of the sequenced molecules were larger than 500 bp. The proportion of such DNA molecules larger than 500 bp was much higher than the proportion of corresponding DNA molecules without size selection (10.6%). The total methylation of those molecules was determined to be 75.5%.
[0566] Figure 87The methylation patterns of two representative plasma DNA molecules are shown after size selection in molecule (I) and molecule (II). Molecule I (chr21:40,881,731-40,882,812) is 1.1 kb long and has 25 CpG sites. The monomolecular methylation level (i.e., the number of methylated sites divided by the total number of sites) of molecule I was determined to be 72.0% using the pathway described in our previous publication (U.S. Application No. 16 / 995,607). Molecule II (chr12:63,108,065-63,111,674) is 3.6 kb long and has 34 CpG sites. The monomolecular methylation level of molecule II was determined to be 94.1%. This demonstrates that size-selective methylation analysis allows us to efficiently analyze the methylation of long DNA molecules and compare the methylation status between two or more molecules.
[0567] B. Size selection using beads
[0568] Solid-phase reversible immobilization technology uses paramagnetic beads to selectively bind nucleic acids, depending on the size of the DNA molecule. These beads comprise a polystyrene core, magnetite, and a carboxylate-modified polymer coating. DNA molecules selectively bind to the beads in the presence of polyethylene glycol (PEG) and salt, depending on the concentrations of PEG and salt in the reaction. PEG causes negatively charged DNA to bind to carboxyl groups on the bead surface, which is then collected in the presence of a magnetic field. Molecules of the desired size are eluted from the magnetic beads using an elution buffer such as 10 mM Tris-HCl, pH 8 buffer, or water. The volume ratio of PEG to DNA determines the size of the DNA molecules we can obtain. A lower PEG:DNA ratio results in more long molecules being retained on the beads.
[0569] 1. Sample processing
[0570] Peripheral blood samples from two pregnant women in late pregnancy were collected in EDTA blood tubes. Peripheral blood samples were collected and centrifuged at 1,600×g for 10 minutes at 4°C. The plasma fraction was further centrifuged at 16,000×g for 10 minutes at 4°C to remove residual cells and debris. The erythrocyte sedimentation rate (ESR) amber fraction was centrifuged at 5,000×g for 5 minutes at room temperature to remove residual plasma. Placental tissue was collected immediately after delivery. Plasma DNA extraction was performed using the QIAamp circulating nucleic acid kit (Qiagen). ESR amber and placental tissue DNA extraction was performed using the QIAamp DNA mini kit (Qiagen).
[0571] 2. Plasma DNA Size Selection
[0572] The extracted plasma DNA samples were aliquoted into two equal portions. One portion from each patient underwent size selection using AMPure XP SPRI beads (Beckman Coulter, Inc.). 50 μL of each extracted plasma DNA sample was thoroughly mixed with 25 μL of AMPureXP solution and incubated at room temperature for 5 minutes. The beads were separated from the solution using a magnet and washed with 180 μL of 80% ethanol. Subsequently, the beads were resuspended in 50 μL of water and vortexed for 1 minute to elute the size-selected DNA from the beads. The beads were then removed to obtain the size-selected DNA solution.
[0573] 3. Single nucleotide polymorphism recognition
[0574] Genotyping of fetal and maternal genomic DNA samples was performed using the iScan system (Illumina). Single nucleotide polymorphisms (SNPs) were identified. The placental genotype was compared with the mother's genotype to identify fetal-specific and maternal-specific alleles. Fetal-specific alleles were defined as alleles present in the fetal genome but not in the maternal genome. In one embodiment, fetal-specific alleles were determined by analyzing SNP sites where the mother was homozygous and the fetus was heterozygous. Maternal-specific alleles were defined by alleles present in the maternal genome but not in the fetal genome. In one embodiment, fetal-specific alleles were determined by analyzing SNP sites where the mother was heterozygous and the fetus was homozygous.
[0575] 4. Single-molecule real-time sequencing
[0576] SMRTbell template preparation kit 1.0-SPv3 (Pacific Biosciences) was used to construct single-molecule real-time (SMRT) sequencing templates using two size-selected samples and their corresponding unselected samples. DNA was purified using 1.8×AMPure PB beads, and library sizes were estimated using a TapeStation instrument (Agilent). Sequencing primer bonding and polymerase binding conditions were calculated using SMRT Link version 5.1.0 software (Pacific Biosciences). In short, sequencing primers v3 were bonded to the sequencing template, and polymerase binding to the template was subsequently performed using Sequel binding and internal control kit 2.1 (Pacific Biosciences). Sequencing was performed on Sequel SMRT Cell 1M v2. Sequencing films were collected for 20 hours on the Sequel system using Sequel sequencing kit 2.1 (Pacific Biosciences).
[0577] 5. Dimensional Analysis
[0578] Figure 88 This table displays sequencing information for size-selected and non-size-selected samples. The first column is the sample identifier. The second column lists the sample groups—regardless of whether size selection was used. The third column lists the number of sequenced molecules. ...
Claims
1. A method for analyzing a biological sample obtained from a woman carrying a fetus for non-diagnostic purposes, said biological sample comprising a plurality of cell-free DNA molecules from said fetus and said woman, said method comprising: Receive sequence reads corresponding to the plurality of free DNA molecules; Measure the size of the plurality of free DNA molecules; A group of free DNA molecules from the plurality of free DNA molecules is identified as having a size greater than or equal to a cutoff value, wherein the cutoff value is at least 500 nt; and For one free DNA molecule in the group of free DNA molecules: Determine the methylation status at each of the multiple sites. Determine the methylation pattern, where: The methylation pattern uses one or more sequence reads corresponding to the free DNA molecule to indicate the methylation state at each of the plurality of sites. The methylation pattern is compared with one or more reference patterns, each of which is determined for a specific tissue type; and The methylation pattern is used to determine the origin organization of the free DNA molecule.
2. The method according to claim 1, wherein the cutoff value is 600 nt.
3. The method according to claim 1, wherein the cutoff value is 1 knt.
4. The method according to any one of claims 1 to 3, further comprising determining the origin of each free DNA molecule in the group of free DNA molecules by: The methylation state at each of a plurality of corresponding sites is determined, wherein the plurality of corresponding sites correspond to the free DNA molecule. Determine the methylation mode, and The methylation mode is compared with at least one of the one or more reference modes.
5. The method of claim 4, further comprising: The amount of free DNA molecules corresponding to each tissue of origin was measured, and The contribution percentage of the originating tissue in the biological sample is determined using the measurement corresponding to the free DNA molecules of each originating tissue.
6. The method of claim 1, wherein measuring the size of the plurality of free DNA molecules comprises: The sequence reads were compared with the reference genome.
7. The method of claim 1, wherein measuring the size of the plurality of free DNA molecules comprises: Full-length sequencing was performed on the aforementioned multiple free DNA molecules, and Count the number of nucleotides in each of the plurality of free DNA molecules.
8. The method of claim 1, wherein measuring the size of the plurality of free DNA molecules comprises: The plurality of cell-free DNA molecules from the biological sample are physically separated from other cell-free DNA molecules in the biological sample, wherein the other cell-free DNA molecules have a size smaller than the cutoff value.
9. The method of claim 1, wherein one of the one or more reference modes is determined by: The methylation density at each reference site in multiple reference sites was measured using DNA molecules from a reference tissue. The methylation density at each of the plurality of reference sites is compared with one or more threshold methylation densities, and Each of the plurality of reference sites is identified as methylated, unmethylated, or non-informative by comparing the methylation density with one or more threshold methylation densities, wherein the plurality of sites are the plurality of reference sites identified as methylated or unmethylated.
10. The method of claim 1, wherein the tissue of origin is the placenta.
11. The method of claim 1, wherein the tissue of origin is a fetus or a mother.
12. The method according to claim 11, wherein: The tissue of origin is fetal. The method further includes: One sequence read from the sequence read is compared with a first region of a reference genome. The first region includes multiple sites corresponding to alleles, and the multiple sites contain a threshold number of sites. The first haplotype is determined using the corresponding alleles present at each of the multiple loci. Compare the first haplotype with the second haplotype corresponding to male individuals, and The comparison is used to classify the likelihood that the male individual is the father of the fetus.
13. The method according to claim 11, wherein: The tissue of origin is fetal. The method further includes: One sequence read from the given sequence is aligned with a first region of a reference genome. This first region includes a first plurality of loci corresponding to alleles, and the plurality of loci comprises a threshold number of loci. Compare the alleles at each of the multiple loci with the corresponding alleles at the same locus in the genome of a male individual, and... The comparison is used to classify the likelihood that the male individual is the father of the fetus.
14. The method of claim 11, further comprising: For each free DNA molecule in the aforementioned group of free DNA molecules: The sequence reads corresponding to the free DNA molecule were compared with a reference genome. The sequence read was identified as corresponding to a haplotype present in the female. The methylation pattern was used to determine that the tissue of origin was fetal, and The haplotype was determined to be a maternally inherited fetal haplotype.
15. The method of claim 14, further comprising: The haplotype was identified as carrying a pathogenic genetic mutation.
16. The method of claim 15, wherein identifying the haplotype as carrying the pathogenic genetic mutation comprises: Identify the genetic mutation in the first sequence read. The degree of first methylation in a second sequence read corresponding to a first genomic position within a first distance of the first sequence read is measured, and The degree of second methylation in a third sequence read corresponding to a second genomic position within a second distance of the first sequence read is measured, wherein: The first degree of methylation and the second degree of methylation are related to the genetic mutation.
17. The method of claim 11, further comprising: For each free DNA molecule in the aforementioned group of free DNA molecules: The sequence reads corresponding to the free DNA molecule were compared with a reference genome. The sequence read segment is identified as corresponding to a region, wherein the region is determined by the following: Receive multiple fetal sequence reads corresponding to multiple fetal DNA molecules from fetal tissue. Receive multiple parent sequence reads corresponding to multiple parent DNA molecules. For each of the multiple fetal sequence reads, determine the fetal methylation status at each methylation site among the multiple methylation sites within the region. For each of the multiple parent sequence reads, determine the parent methylation state at each of the multiple methylation sites. The value of a parameter is determined to characterize the amount of sites where the fetal methylation state differs from the maternal methylation state. Compare the value of the parameter with a threshold, and The value of the parameter is determined to exceed the threshold.
18. The method of claim 1, wherein determining the tissue of origin of the free DNA molecule comprises inputting the methylation pattern into a machine learning model, the model being trained by: Multiple training methylation patterns are received, each training methylation pattern having a methylation state at one or more sites among the multiple sites, and each training methylation pattern is determined by DNA molecules from known tissues. Storing multiple training samples, each training sample containing one of the multiple training methylation patterns and a label indicating the known tissue corresponding to the training methylation pattern, and When the plurality of training methylation patterns are input into the model, the parameters of the model are optimized using the plurality of training samples based on the plurality of outputs of the model that match or do not match the corresponding labels, wherein one output of the model indicates the tissue corresponding to the input methylation pattern.
19. The method of claim 18, wherein the machine learning model comprises a convolutional neural network (CNN), linear regression, logistic regression, deep recurrent neural network, Bayes' classifier, hidden Markov model (HMM), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering with noisy applications (DBSCAN), random forest algorithm, or support vector machine (SVM).
20. The method of claim 18, wherein each DNA molecule from the known tissue is cellular DNA.
21. The method of claim 18, wherein the parameters of the model include a first parameter indicating whether one of the plurality of sites has the same methylation state as another of the plurality of sites.
22. The method of claim 18, wherein the parameters of the model include a second parameter indicating the distance between each of the plurality of sites.
23. The method according to any one of claims 1 and 18, wherein one of the one or more reference modes corresponds to a reference organization. The method further includes determining the origin tissue as the reference tissue when the methylation pattern matches the reference pattern.
24. The method according to any one of claims 1 and 18, wherein the plurality of sites comprises at least 5 CpG sites.
25. The method according to any one of claims 1 and 18, wherein determining the tissue of origin using the methylation pattern comprises: A similarity score is determined by comparing the methylation pattern with a first reference methylation pattern from a first reference tissue among a plurality of reference tissues. Compare the similarity score with the threshold; and When the similarity score exceeds the threshold, the originating organization is determined to be the first reference organization.
26. The method of claim 25, wherein: The similarity score is the first similarity score. The method further includes: The threshold is calculated as follows: A second similarity score is determined by comparing the methylation pattern with a second reference methylation pattern from a second reference organization among the plurality of reference organizations, wherein the first reference organization and the second reference organization are different organizations, and the threshold is the second similarity score.
27. The method of claim 25, wherein: The first reference methylation pattern includes a first subgroup of sites having at least a first methylation probability for the first reference tissue. The first reference methylation pattern includes a second subgroup site having at most a second methylation probability for the first reference tissue, and Determining the similarity score includes: The similarity score is increased when one of the plurality of sites is methylated and that site is located in the first subgroup of sites. The similarity score is reduced when one of the plurality of sites is methylated and the site is located in the second subgroup of sites.
28. The method of claim 25, wherein: The first reference methylation pattern includes the plurality of sites, wherein each site is characterized by a methylation probability and an unmethylation probability relative to the first reference tissue. The similarity score is determined by the following: For each of the plurality of sites: Determine the probability in the reference tissue of the methylation state corresponding to the site in the free DNA molecule. Calculate the product of multiple probabilities, where the product is the similarity score.
29. The method of claim 28, wherein the probability is determined using a beta (β) distribution.
30. The method according to any one of claims 1 and 18, further comprising: Sequencing of the multiple free DNA molecules to obtain sequence reads, and The methylation state of a site is determined by measuring the characteristics of the nucleotide corresponding to the site and the nucleotides adjacent to the site.
31. The method according to any one of claims 1 and 18, wherein the size of the plurality of free DNA molecules includes the number of CpG sites.
32. The method according to any one of claims 1 and 18, wherein at least one of the plurality of sites is methylated.
33. The method according to any one of claims 1 and 18, wherein two sites of the plurality of sites are spaced at least 160 nt apart.
34. A computer-readable medium storing instructions that, when executed, control a computing system to perform a method for analyzing a biological sample obtained from a woman pregnant with a fetus, the biological sample comprising a plurality of cell-free DNA molecules from the fetus and the woman, the method comprising: Receive sequence reads corresponding to the plurality of free DNA molecules; Measure the size of the plurality of free DNA molecules; A group of free DNA molecules from the plurality of free DNA molecules is identified as having a size greater than or equal to a cutoff value, wherein the cutoff value is at least 500 nt; and For one free DNA molecule in the group of free DNA molecules: Determine the methylation status at each of the multiple sites. Determine the methylation pattern, where: The methylation pattern uses one or more sequence reads corresponding to the free DNA molecule to indicate the methylation state at each of the plurality of sites. The methylation pattern is compared with one or more reference patterns, each of which is determined for a specific tissue type; and The methylation pattern is used to determine the origin organization of the free DNA molecule.
35. The computer-readable medium of claim 34, wherein the cutoff value is 600 nt.
36. The computer-readable medium of claim 34, wherein the cutoff value is 1 knt.
37. The computer-readable medium of claim 34, wherein the method further comprises determining the origin organization of each free DNA molecule in the group of free DNA molecules by: The methylation state at each of a plurality of corresponding sites is determined, wherein the plurality of corresponding sites correspond to the free DNA molecule. Determine the methylation mode, and The methylation mode is compared with at least one of the one or more reference modes.
38. The computer-readable medium of claim 37, wherein the method further comprises: The amount of free DNA molecules corresponding to each tissue of origin was measured, and The contribution percentage of the originating tissue in the biological sample is determined using the measurement corresponding to the free DNA molecules of each originating tissue.
39. The computer-readable medium of claim 34, wherein measuring the size of the plurality of free DNA molecules comprises: The sequence reads were compared with the reference genome.
40. The computer-readable medium of claim 34, wherein measuring the size of the plurality of free DNA molecules comprises: Full-length sequencing was performed on the aforementioned multiple free DNA molecules, and Count the number of nucleotides in each of the plurality of free DNA molecules.
41. The computer-readable medium of claim 34, wherein measuring the size of the plurality of free DNA molecules comprises: The plurality of cell-free DNA molecules from the biological sample are physically separated from other cell-free DNA molecules in the biological sample, wherein the other cell-free DNA molecules have a size smaller than the cutoff value.
42. The computer-readable medium of claim 34, wherein one of the one or more reference modes is determined by: The methylation density at each reference site in multiple reference sites was measured using DNA molecules from a reference tissue. The methylation density at each of the plurality of reference sites is compared with one or more threshold methylation densities, and Each of the plurality of reference sites is identified as methylated, unmethylated, or non-informative by comparing the methylation density with one or more threshold methylation densities, wherein the plurality of sites are the plurality of reference sites identified as methylated or unmethylated.
43. The computer-readable medium of claim 34, wherein the tissue of origin is the placenta.
44. The computer-readable medium of claim 34, wherein the tissue of origin is a fetus or a mother.
45. The computer-readable medium of claim 44, wherein: The tissue of origin is fetal. The method further includes: One sequence read from the sequence read is compared with a first region of a reference genome. The first region includes multiple sites corresponding to alleles, and the multiple sites contain a threshold number of sites. The first haplotype is determined using the corresponding alleles present at each of the multiple loci. Compare the first haplotype with the second haplotype corresponding to male individuals, and The comparison is used to classify the likelihood that the male individual is the father of the fetus.
46. The computer-readable medium of claim 44, wherein: The tissue of origin is fetal. The method further includes: One sequence read from the given sequence is aligned with a first region of a reference genome. This first region includes a first plurality of loci corresponding to alleles, and the plurality of loci comprises a threshold number of loci. Compare the alleles at each of the multiple loci with the corresponding alleles at the same locus in the genome of a male individual, and... The comparison is used to classify the likelihood that the male individual is the father of the fetus.
47. The computer-readable medium of claim 44, wherein the method further comprises: For each free DNA molecule in the aforementioned group of free DNA molecules: The sequence reads corresponding to the free DNA molecule were compared with a reference genome. The sequence read was identified as corresponding to a haplotype present in the female. The methylation pattern was used to determine that the tissue of origin was fetal, and The haplotype was determined to be a maternally inherited fetal haplotype.
48. The computer-readable medium of claim 47, wherein the method further comprises: The haplotype was identified as carrying a pathogenic genetic mutation.
49. The computer-readable medium of claim 48, wherein identifying the haplotype as carrying the pathogenic genetic mutation comprises: Identify the genetic mutation in the first sequence read. The degree of first methylation in a second sequence read corresponding to a first genomic position within a first distance of the first sequence read is measured, and The degree of second methylation in a third sequence read corresponding to a second genomic position within a second distance of the first sequence read is measured, wherein: The first degree of methylation and the second degree of methylation are related to the genetic mutation.
50. The computer-readable medium of claim 44, wherein the method further comprises: For each free DNA molecule in the aforementioned group of free DNA molecules: The sequence reads corresponding to the free DNA molecule were compared with a reference genome. The sequence read segment is identified as corresponding to a region, wherein the region is determined by the following: Receive multiple fetal sequence reads corresponding to multiple fetal DNA molecules from fetal tissue. Receive multiple parent sequence reads corresponding to multiple parent DNA molecules. For each of the multiple fetal sequence reads, determine the fetal methylation status at each methylation site among the multiple methylation sites within the region. For each of the multiple parent sequence reads, determine the parent methylation state at each of the multiple methylation sites. The value of a parameter is determined to characterize the amount of sites where the fetal methylation state differs from the maternal methylation state. Compare the value of the parameter with a threshold, and The value of the parameter is determined to exceed the threshold.
51. The computer-readable medium of claim 34, wherein determining the origin of the free DNA molecule comprises inputting the methylation pattern into a machine learning model, the model being trained by: Multiple training methylation patterns are received, each training methylation pattern having a methylation state at one or more sites among the multiple sites, and each training methylation pattern is determined by DNA molecules from known tissues. Storing multiple training samples, each training sample containing one of the multiple training methylation patterns and a label indicating the known tissue corresponding to the training methylation pattern, and When the plurality of training methylation patterns are input into the model, the parameters of the model are optimized using the plurality of training samples based on the plurality of outputs of the model that match or do not match the corresponding labels, wherein one output of the model indicates the tissue corresponding to the input methylation pattern.
52. The computer-readable medium of claim 51, wherein the machine learning model comprises a convolutional neural network (CNN), linear regression, logistic regression, deep recurrent neural network, Bayes' classifier, hidden Markov model (HMM), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering with noisy applications (DBSCAN), random forest algorithm, or support vector machine (SVM).
53. The computer-readable medium of claim 51, wherein each DNA molecule from the known tissue is cellular DNA.
54. The computer-readable medium of claim 51, wherein the parameters of the model include a first parameter indicating whether one of the plurality of sites has the same methylation state as another of the plurality of sites.
55. The computer-readable medium of claim 51, wherein the parameters of the model include a second parameter indicating the distance between each point in the plurality of sites.
56. The computer-readable medium of claim 34, wherein one of the one or more reference modes corresponds to a reference organization. The method further includes determining the origin tissue as the reference tissue when the methylation pattern matches the reference pattern.
57. The computer-readable medium of claim 34, wherein the plurality of sites comprises at least 5 CpG sites.
58. The computer-readable medium of claim 34, wherein determining the origin tissue using the methylation pattern comprises: A similarity score is determined by comparing the methylation pattern with a first reference methylation pattern from a first reference tissue among a plurality of reference tissues. Compare the similarity score with the threshold; and When the similarity score exceeds the threshold, the originating organization is determined to be the first reference organization.
59. The computer-readable medium according to claim 58, wherein: The similarity score is the first similarity score. The method further includes: The threshold is calculated as follows: A second similarity score is determined by comparing the methylation pattern with a second reference methylation pattern from a second reference organization among the plurality of reference organizations, wherein the first reference organization and the second reference organization are different organizations, and the threshold is the second similarity score.
60. The computer-readable medium of claim 58, wherein: The first reference methylation pattern includes a first subgroup of sites having at least a first methylation probability for the first reference tissue. The first reference methylation pattern includes a second subgroup site having at most a second methylation probability for the first reference tissue, and Determining the similarity score includes: The similarity score is increased when one of the plurality of sites is methylated and that site is located in the first subgroup of sites. The similarity score is reduced when one of the plurality of sites is methylated and the site is located in the second subgroup of sites.
61. The computer-readable medium according to claim 58, wherein: The first reference methylation pattern includes the plurality of sites, wherein each site is characterized by a methylation probability and an unmethylation probability relative to the first reference tissue. The similarity score is determined by the following: For each of the plurality of sites: Determine the probability in the reference tissue of the methylation state corresponding to the site in the free DNA molecule. Calculate the product of multiple probabilities, where the product is the similarity score.
62. The computer-readable medium of claim 61, wherein the probability is determined using a beta (β) distribution.
63. The computer-readable medium of claim 34, wherein the method further comprises: Sequencing of the multiple free DNA molecules to obtain sequence reads, and The methylation state of a site is determined by measuring the characteristics of the nucleotide corresponding to the site and the nucleotides adjacent to the site.
64. The computer-readable medium of claim 34, wherein the size of the plurality of free DNA molecules includes the number of CpG sites.
65. The computer-readable medium of claim 34, wherein at least one of the plurality of sites is methylated.
66. The computer-readable medium of claim 34, wherein two of the plurality of sites are spaced at least 160 nt apart.
67. A computer program product comprising instructions that, when executed, control a computing system to perform the method as claimed in any one of claims 1-33 or the method performed by a computer-readable medium as claimed in any one of claims 34-66.
68. A computer-readable storage medium comprising the computer program product of claim 67.
69. A computing system comprising the computer program product according to claim 67.
Citation Information
Patent Citations
Gestational age assessment by methylation and size profiling of maternal plasma DNA
US20180105807A1
Determination of base modifications of nucleic acids
US20210047679A1
Methylation pattern analysis of haplotypes in tissues in DNA mixture
CN108138233A
Universal haplotype-based noninvasive prenatal testing for single gene diseases
CN109996894A