Molecular analysis using long free fragments in pregnant women

By analyzing free DNA fragments longer than 200 bp using PacBio SMRT sequencing technology, the problem of loss of long DNA molecules in the current technology during PCR amplification is solved, and the accurate identification and analysis of fetal and maternal genetic information is achieved, and important data on fetal health and gestational age are provided.

CN120048339APending Publication Date: 2025-05-27THE CHINESE UNIVERSITY OF HONG KONG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510195201.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2021-01-08
Filing Date
2021-02-05
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

It is difficult to effectively analyze and sequence free DNA fragments longer than 600 bp, especially when using the Illumina sequencing platform, long DNA molecules are easily lost during PCR amplification, resulting in incomplete sequencing results.

Method used

By using PacBio SMRT sequencing technology and the properties of long free DNA fragments, free DNA fragments longer than 200 bp can be accurately sequenced and analyzed, including identification of fetal specific alleles and maternal specific alleles, thereby determining fetal haplotypes and maternal genetic information.

Benefits of technology

Efficient sequencing and analysis of long free DNA fragments is achieved, which can accurately identify the genetic information of the fetus and maternal body, provide important data on fetal health and gestational age, and help detect pregnancy-related diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048339A_ABST
    Figure CN120048339A_ABST
Patent Text Reader

Abstract

The methods and systems described herein relate to analyzing a biological sample from a pregnant individual using long free DNA fragments. Deoxyribonucleic acid (DNA) fragments of biological samples are often analyzed using methylated CpG sites and the status of single nucleotide polymorphisms (SNPs). The CpG site and SNP are generally spaced from the nearest CpG site or SNP by hundreds or thousands of base pairs. It is little or impossible to find two or more contiguous CpG sites or SNPs on the majority of free DNA fragments. A free DNA fragment longer than 600 bp may comprise a plurality of CpG sites and / or SNPs. The presence of multiple CpG sites and / or SNPs on a long free DNA fragment as compared to a single short free DNA fragment may allow for analysis. Long free DNA fragments may be used to identify origin tissue and / or to provide information about a fetus in a pregnant female.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application is a divisional application of Chinese Patent Application No. 202180013180.5, and claims the priority of U.S. Provisional Application No. 62 / 970,634 filed on February 5, 2020 and U.S. Provisional Application No. 63 / 135,486 filed on January 8, 2021. The entire contents of the two applications are incorporated herein by reference for all purposes. Background Art

[0003] The modal size of circulating free DNA in pregnant women has been reported to be approximately 166 bp (Lo et al., Sci Transl Med. 2010; 2: 61ra91). There is very little published data on fragments larger than 600 bp. One example is the work of Amicucci et al. (Amicucci et al., Clin Chem 2000; 40: 301-2) reporting the amplification of an 8 kb fragment of the basic protein Y2 gene (BPY2) of the Y chromosome from maternal plasma using PCR. It is not known whether such data is generalizable across the genome. In fact, there are many challenges in detecting such long DNA fragments, such as those larger than 600 bp, using massively parallel short-read sequencing technologies, such as those using the Illumina platform (Lo et al., Sci Transl Med 2010; 2: 61ra91; Fan et al., Clin Chem 2010; 56: 1278-86). These challenges include: (1) the recommended size range of the Illumina sequencing platform typically spans 100-300 bp (De Maio et al., MicobGenom. 2019; 5(9)); (2) DNA amplification should be involved in the preparation of the sequencing library (via PCR) or the generation of sequencing clusters (via bridge amplification) on the flow cell. Such amplification processes may favor the amplification of shorter DNA fragments, in part due to the fact that long DNA templates (e.g., >600 bp) should require a relatively long time compared to short DNA templates (e.g., <200 bp) to complete daughter strand synthesis. Thus, within the fixed time frame of these PCR processes before or during sequencing on the Illumina platform, those long DNA molecules whose daughter strands are not fully generated during the PCR process will be unavailable for downstream analysis; (3) long DNA molecules will have a higher probability of forming secondary structures that impede amplification; (4) using Illumina sequencing technology, long DNA molecules are more likely to produce clusters containing more than one clonal DNA molecule compared to short DNA molecules, because the library is denatured, diluted, and spread on a two-dimensional surface, followed by bridge amplification (Head et al., Biotechniques. 2014; 56: 61-4). Summary of the Invention

[0004] The methods and systems described herein relate to the analysis of biological samples using long cell-free DNA fragments. Using these long cell-free DNA fragments allows for analyses that were not previously considered or that were not possible with shorter cell-free DNA fragments. The DNA fragments of a biological sample are often analyzed for the status of methylated CpG sites and single nucleotide polymorphisms (SNPs). CpG sites and SNPs are typically separated from the nearest CpG site or SNP by hundreds or thousands of base pairs. Most cell-free DNA fragments in a biological sample are typically less than 200 bp in length. Thus, it is unlikely or impossible to find two or more consecutive CpG sites or SNPs on most cell-free DNA fragments. Cell-free DNA fragments longer than 200 bp, including those longer than 600 bp or 1 kb, can contain multiple CpG sites and / or SNPs. The presence of multiple CpG sites and / or SNPs on long cell-free DNA fragments can allow for more efficient and / or accurate analyses compared to short cell-free DNA fragments alone. Long cell-free DNA fragments can be used to identify the tissue of origin and / or to provide information about a fetus in a pregnant female. Additionally, it was unexpected that samples from pregnant women could be accurately analyzed using long cell-free DNA fragments because we expected that the long cell-free DNA fragments would be primarily of maternal origin. We did not expect that long cell-free DNA fragments of fetal origin would be present in amounts sufficient to provide information about the fetus.

[0005] Long cell-free DNA fragments that have SNPs can be used to determine the haplotype inherited by the fetus. Long cell-free DNA fragments can have a methylation pattern indicative of the tissue of origin by having multiple CpG sites. Additionally, trinucleotide repeats and other repeat sequences can be present on long cell-free DNA fragments. These repeat sequences can be used to determine the likelihood of a genetic disorder in the fetus or fetal relatedness. The amount of long cell-free DNA fragments can be used to determine gestational age. Similarly, the motif at the end of a long cell-free DNA fragment can also be used to determine gestational age. Long cell-free DNA fragments (including, for example, the amount, length distribution, genomic location, methylation status, etc. of the fragment) can be used to determine pregnancy-related disorders.

[0006] These and other embodiments of the disclosure are described in detail below. For example, other embodiments relate to systems, devices, and computer-readable media related to the methods described herein.

[0007] A better understanding of the nature and advantages of the embodiments of the disclosure can be obtained by reference to the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1A and Figure 1B Shows the size distribution of determined cell-free DNA according to an embodiment of the invention. (A) 0 - 20 kb on a linear scale, (B) 0 - 20 kb on a logarithmic scale.

[0009] Figure 2A and Figure 2B Show the determined size distribution of cell-free DNA according to an embodiment of the present invention. (A) 0 - 5 kb on the linear scale of the y-axis. (B) 0 - 5 kb on the logarithmic scale of the y-axis.

[0010] Figure 3A and Figure 3B Show the determined size distribution of cell-free DNA according to an embodiment of the present invention. (A) 0 - 400 bp on the linear scale of the y-axis. (B) 0 - 400 bp on the logarithmic scale of the y-axis.

[0011] Figure 4A and Figure 4B Show the determined size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fragments carrying fetus-specific alleles (fetus-specific) according to an embodiment of the present invention. (A) 0 - 20 kb bp on the linear scale of the y-axis. (B) 0 - 20 kb on the logarithmic scale of the y-axis. The blue line indicates fragments carrying shared alleles (mainly of maternal origin) and the red line indicates fragments carrying fetus-specific alleles (of placental origin).

[0012] Figure 5A and Figure 5B Show the determined size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fragments carrying fetus-specific alleles (fetus-specific) according to an embodiment of the present invention. (A) 0 - 5 kb bp on the linear scale of the y-axis. (B) 0 - 5 kb on the logarithmic scale of the y-axis. The blue line indicates fragments carrying shared alleles (mainly of maternal origin) and the red line indicates fragments carrying fetus-specific alleles (of placental origin).

[0013] Figure 6A and Figure 6B Show the determined size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fragments carrying fetus-specific alleles (fetus-specific) according to an embodiment of the present invention. (A) 0 - 1 kb on the linear scale of the y-axis. (B) 0 - 1 kb on the logarithmic scale of the y-axis. The blue line indicates fragments carrying shared alleles (mainly of maternal origin) and the red line indicates fragments carrying fetus-specific alleles (of placental origin).

[0014] Figure 7A and Figure 7BShows the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fragments carrying fetal-specific alleles (fetal-specific) determined according to an embodiment of the present invention. (A) 0 - 400 bp on a linear scale of the y-axis. (B) 0 - 400 bp on a logarithmic scale of the y-axis. The blue line indicates fragments carrying shared alleles (primarily of maternal origin) and the red line indicates fragments carrying fetal-specific alleles (of placental origin).

[0015] Figure 8 Shows the degree of single-molecule double-stranded DNA methylation between fragments carrying maternal-specific alleles and fragments carrying fetal-specific alleles according to an embodiment of the present invention.

[0016] Figure 9A and Figure 9B Shows (A) the fitted distribution of the degree of single-molecule double-stranded DNA methylation between fragments carrying maternal-specific alleles and fragments carrying fetal-specific alleles and (B) the receiver operating characteristic (ROC) analysis using the degree of single-molecule double-stranded DNA methylation according to an embodiment of the present invention.

[0017] Figure 10A and Figure 10B Shows the correlation between the degree of single-molecule double-stranded DNA methylation and the fragment size of plasma DNA according to an embodiment of the present invention. (A) Size range of 0 - 20 kb. (B) Size range of 0 - 1 kb.

[0018] Figure 11A and Figure 11B Shows examples of identified long fetal-specific DNA molecules in maternal plasma DNA of pregnant women according to an embodiment of the present invention. (A) The black bars indicate long fetal-specific DNA molecules aligned to regions in chromosome 10 of the human reference genome. (B) A detailed illustration of the genetic and epigenetic information determined using PacBio sequencing in the present disclosure. The bases highlighted in yellow (marked by arrows) may be attributed to sequence errors, which may be corrected in some embodiments.

[0019] Figure 12A and Figure 12B Shows examples of identified long maternal DNA molecules carrying shared alleles in maternal plasma DNA of pregnant women according to an embodiment of the present invention. (A) The black bars indicate long maternal-specific DNA molecules aligned to regions in chromosome 6 of the human reference. (B) A detailed illustration of the genetic and epigenetic information determined using PacBio sequencing according to an embodiment of the present invention.

[0020] Figure 13Shows the frequency distribution of DNA from the placenta (red) and DNA from maternal blood cells (blue) at different resolutions from 1 kb to 20 kb according to the degree of methylation, according to an embodiment of the present invention.

[0021] Figure 14A and Figure 14B Shows the frequency distribution of DNA from the placenta (red) and DNA from maternal blood cells (blue) within 16-kb and 24-kb windows according to the degree of methylation, according to an embodiment of the present invention.

[0022] Figure 15A and Figure 15B Shows examples of identified long maternal-specific DNA molecules in maternal plasma DNA according to an embodiment of the present invention. (A) The black bars indicate long maternal-specific DNA molecules aligned to regions in chromosome 8 of the human reference. (B) A detailed illustration of the genetics and epigenetics determined using PacBio sequencing according to an embodiment of the present invention.

[0023] Figure 16 Shows an inferred maternal genetic diagram of the fetus according to an embodiment of the present invention.

[0024] Figure 17 Illustrates the determination of genetic / epigenetic disorders in plasma DNA molecules using information on maternal origin and fetal origin according to an embodiment of the present invention.

[0025] Figure 18 Illustrates the identification of abnormal fetal fragments according to an embodiment of the present invention.

[0026] Figure 19A - 19G Shows an error correction illustration of cell-free DNA genotyping using PacBio sequencing according to an embodiment of the present invention. '.' represents a base identical to the reference base in the Watson strand. ',' represents a base identical to the reference base in the Crick strand. 'Letter' represents an alternative allele different from the reference allele. '*' represents an insertion. '^' represents a deletion.

[0027] Figure 20 Shows a method for analyzing a biological sample obtained from a female carrying a fetus according to an embodiment of the present invention.

[0028] Figure 21 Shows a method for analyzing a biological sample obtained from a female carrying a fetus to determine haplotypic inheritance according to an embodiment of the present invention.

[0029] Figure 22 Shows a methylation pattern for determining the origin tissue of long DNA molecules in plasma according to an embodiment of the present invention.

[0030] Figure 23 Displays the receiver operating characteristic (ROC) curve for determining fetal origin and maternal origin according to an embodiment of the present invention.

[0031] Figure 24 Displays paired methylation patterns according to an embodiment of the present invention.

[0032] Figure 25 Is a distribution table of selected marker regions among different chromosomes according to an embodiment of the present invention.

[0033] Figure 26 Is a classification table of plasma DNA molecules based on the single-molecule methylation pattern of plasma DNA molecules according to an embodiment of the present invention, using different percentages of buffy coat DNA molecules with a mismatch score greater than 0.3 as the selection criterion for marker regions.

[0034] Figure 27 Displays the method flow for non-invasively determining fetal inheritance using placenta-specific methylation haplotypes according to an embodiment of the present invention.

[0035] Figure 28 Illustrates the principle of non-invasive prenatal detection of fragile X syndrome using long free DNA in maternal plasma according to an embodiment of the present invention.

[0036] Figure 29 Illustrates maternal inheritance of a fetus based on methylation patterns according to an embodiment of the present invention.

[0037] Figure 30 Illustrates qualitative analysis of maternal inheritance of a fetus using genetic and epigenetic information of plasma DNA molecules according to an embodiment of the present invention.

[0038] Figure 31 Illustrates the detection rate of qualitative analysis of maternal inheritance of a fetus in a genome-wide manner using genetic and epigenetic information of plasma DNA molecules compared to relative haplotype dosage (RHDO) analysis according to an embodiment of the present invention.

[0039] Figure 32 Displays the relationship between the detection rate of paternal-specific variants in a genome-wide manner and the number of sequenced plasma DNA molecules of different sizes used for analysis according to an embodiment of the present invention.

[0040] Figure 33 Displays the workflow of non-invasive detection of fragile X syndrome according to an embodiment of the present invention.

[0041] Figure 34Displays the methylation pattern of plasma DNA compared to the methylation profiles of placental and buffy coat DNA according to an embodiment of the present invention.

[0042] Figure 35 Table showing the CpG site distribution in 500-bp regions across the entire human genome according to an embodiment of the present invention.

[0043] Figure 36 Table showing the CpG site distribution in 1-kb regions across the entire human genome according to an embodiment of the present invention.

[0044] Figure 37 Table showing the CpG site distribution in 3-kb regions across the entire human genome according to an embodiment of the present invention.

[0045] Figure 38 Table showing the contribution ratios of different tissues to DNA molecules in maternal plasma using methylation status matching analysis according to an embodiment of the present invention.

[0046] Figure 39A and Figure 39B Displays the relationship between the placental contribution inferred through the SNP approach and the fetal DNA fraction according to an embodiment of the present invention.

[0047] Figure 40 Displays a method for analyzing a biological sample obtained from a female carrying a fetus to determine the origin tissue using methylation pattern analysis according to an embodiment of the present invention.

[0048] Figure 41A and Figure 41B Displays the size distribution of cell-free DNA molecules from maternal plasma samples in the first, second, and third trimesters of pregnancy according to an embodiment of the present invention.

[0049] Figure 42 Table showing the proportion of long plasma DNA molecules in different trimesters of pregnancy according to an embodiment of the present invention.

[0050] Figure 43A and Figure 43B Displays the size distribution of DNA molecules covering fetal-specific alleles from maternal plasma in the first, second, and third trimesters of pregnancy according to an embodiment of the present invention.

[0051] Figure 44A and Figure 44B Displays the size distribution of DNA molecules covering maternal-specific alleles from maternal plasma in the first, second, and third trimesters of pregnancy according to an embodiment of the present invention.

[0052] Figure 45 Table showing the proportions of long fetal and maternal plasma DNA molecules in different trimesters of pregnancy according to an embodiment of the present invention.

[0053] Figure 46A 、 Figure 46B and Figure 46C display a proportional graph of fetal-specific plasma DNA fragments within a specific size range across different gestational periods according to an embodiment of the present invention.

[0054] Figure 47A 、 Figure 47B and Figure 47C display a proportional graph of the base content at the 5'-end of cell-free DNA molecules within a fragment size range from 0 kb to 3 kb from maternal plasma in the first, second, and third trimesters of pregnancy according to an embodiment of the present invention.

[0055] Figure 48 is a table of the terminal nucleotide base ratios among short and long cell-free DNA molecules from maternal plasma in the first, second, and third trimesters of pregnancy according to an embodiment of the present invention.

[0056] Figure 49 is a table of the terminal nucleotide base ratios among short and long cell-free DNA molecules covering fetal-specific alleles from maternal plasma in the first, second, and third trimesters of pregnancy according to an embodiment of the present invention.

[0057] Figure 50 is a table of the terminal nucleotide base ratios among short and long cell-free DNA molecules covering maternal-specific alleles from maternal plasma in the first, second, and third trimesters of pregnancy according to an embodiment of the present invention.

[0058] Figure 51 illustrates a hierarchical clustering analysis of short and long plasma cell-free DNA molecules using 256 terminal motifs according to an embodiment of the present invention.

[0059] Figure 52A and Figure 52B display a principal component analysis of the 4-mer terminal motif profile according to an embodiment of the present invention.

[0060] Figure 53 is a table of the 25 terminal motifs with the highest frequencies among short plasma DNA molecules from maternal plasma in the first trimester of pregnancy according to an embodiment of the present invention.

[0061] Figure 54 is a table of the 25 terminal motifs with the highest frequencies among short plasma DNA molecules from maternal plasma in the second trimester of pregnancy according to an embodiment of the present invention.

[0062] Figure 55 is a table of the 25 terminal motifs with the highest frequencies among short plasma DNA molecules from maternal plasma in the third trimester of pregnancy according to an embodiment of the present invention.

[0063] Figure 56Table of the 25 terminal motifs with the highest frequencies among the long plasma DNA molecules from maternal plasma in the first trimester according to an embodiment of the present invention.

[0064] Figure 57 Table of the 25 terminal motifs with the highest frequencies among the long plasma DNA molecules from maternal plasma in the second trimester according to an embodiment of the present invention.

[0065] Figure 58 Table of the 25 terminal motifs with the highest frequencies among the long plasma DNA molecules from maternal plasma in the third trimester according to an embodiment of the present invention.

[0066] Figure 59A 、 Figure 59B and Figure 59C Show the motif frequency scatter plots of 16 NNXY motifs among short and long plasma DNA molecules in maternal plasma in (A) the first trimester, (B) the second trimester, and (C) the third trimester according to an embodiment of the present invention.

[0067] Figure 60 Show a method for analyzing a biological sample obtained from a female carrying a fetus to determine the gestational age according to an embodiment of the present invention.

[0068] Figure 61 Show a method for analyzing a biological sample obtained from a female carrying a fetus to classify the likelihood of pregnancy-related disorders according to an embodiment of the present invention.

[0069] Figure 62 Table showing the clinical information of four preeclampsia cases according to an embodiment of the present invention.

[0070] Figure 63A - 63D Size distribution plots of free DNA molecules from maternal plasma samples in the third trimester of preeclampsia and normotensive pregnancies according to an embodiment of the present invention.

[0071] Figure 64A - 64D Size distribution plots of free DNA molecules from maternal plasma samples in the third trimester of preeclampsia and normotensive pregnancies according to an embodiment of the present invention.

[0072] Figure 65A - 65D Size distribution plots of DNA molecules covering fetal-specific alleles from maternal plasma samples in the third trimester of preeclampsia and normotensive pregnancies according to an embodiment of the present invention.

[0073] Figure 66A - 66D Size distribution plots of DNA molecules covering fetal-specific alleles from maternal plasma samples in the third trimester of preeclampsia and normotensive pregnancies according to an embodiment of the present invention.

[0074] Figure 67A - 67DSize distribution plot of DNA molecules covering maternal-specific alleles from preeclamptic and normotensive late pregnancy maternal plasma samples according to an embodiment of the present invention.

[0075] Figure 68A - 68D Size distribution plot of DNA molecules covering maternal-specific alleles from preeclamptic and normotensive late pregnancy maternal plasma samples according to an embodiment of the present invention.

[0076] Figure 69A and Figure 69B Proportion plot of short DNA molecules covering fetal-specific alleles and maternal-specific alleles in preeclamptic and normotensive maternal plasma samples sequenced by PacBio SMRT sequencing according to an embodiment of the present invention.

[0077] Figure 70A and Figure 70B Proportion plot of short DNA molecules in preeclamptic and normotensive maternal plasma samples sequenced by PacBio SMRT sequencing and Illumina sequencing according to an embodiment of the present invention.

[0078] Figure 71 Size ratio plot indicating the relative proportions of short and long DNA molecules in preeclamptic and normotensive maternal plasma samples sequenced by PacBio SMRT sequencing according to an embodiment of the present invention.

[0079] Figure 72A - 72D Shows the proportions of different ends of plasma DNA molecules in preeclamptic and normotensive maternal plasma samples sequenced by PacBio SMRT sequencing according to an embodiment of the present invention.

[0080] Figure 73 Shows hierarchical cluster analysis of preeclamptic and normotensive late pregnancy maternal plasma DNA samples using the frequencies of plasma DNA molecules having four types of fragment ends (the first nucleotide at the 5'-end of each strand), namely C-end, G-end, T-end, and A-end, respectively, according to an embodiment of the present invention.

[0081] Figure 74 Shows hierarchical cluster analysis of preeclamptic and normotensive late pregnancy maternal plasma DNA samples using 16 dinucleotide motifs XYNN (dinucleotide sequences of the first and second nucleotides from the 5'-end) according to an embodiment of the present invention.

[0082] Figure 75 Shows hierarchical cluster analysis of preeclamptic and normotensive late pregnancy maternal plasma DNA samples using 16 dinucleotide motifs NNXY (dinucleotide sequences of the third and fourth nucleotides from the 5'-end) according to an embodiment of the present invention.

[0083] Figure 76 Hierarchical cluster analysis of maternal plasma DNA samples in late pregnancy with preeclampsia and normotensive pregnancy using 256 tetranucleotide motifs (a dinucleotide sequence from the first nucleotide to the fourth nucleotide at the 5′ end) according to an embodiment of the present invention is shown.

[0084] Figure 77A - 77D The contribution of T cells among four types of fragment ends in maternal plasma DNA samples of preeclampsia and normotensive mothers according to an embodiment of the present invention is shown.

[0085] Figure 78 A method of analyzing a biological sample obtained from a female pregnant with a fetus to determine the likelihood of a pregnancy-related disorder is shown according to an embodiment of the present invention.

[0086] Figure 79 A graphical representation showing inferred maternal inheritance of a repeat sequence-associated disease in a fetus according to an embodiment of the present invention.

[0087] Figure 80 A diagram showing inferred paternal inheritance of a repeat sequence-associated disease in a fetus according to an embodiment of the present invention.

[0088] Figure 81 , Figure 82 and Figure 83 A table showing examples of repeat expansion diseases.

[0089] Figure 84 A table showing examples of repeat expansion detection and repeat-associated methylation determination in a fetus according to an embodiment of the present invention.

[0090] Figure 85 A method of analyzing a biological sample obtained from a female pregnant with a fetus in order to determine the likelihood of a genetic disorder in the fetus is shown according to an embodiment of the present invention.

[0091] Figure 86 A method of analyzing a biological sample obtained from a female pregnant with a fetus to determine kinship according to an embodiment of the present invention is shown.

[0092] Figure 87 Shown are the methylation patterns of two representative plasma DNA molecules after size selection.

[0093] Figure 88 This is a sequencing information table of size-selected samples and non-size-selected samples according to an embodiment of the present invention.

[0094] Figure 89A and Figure 89BDiagram showing the plasma DNA size profiles of bead-based size-selected samples and non-bead-based size-selected samples according to an embodiment of the present invention.

[0095] Figure 90A and Figure 90B Diagram showing the size profile between fetal DNA molecules and maternal DNA molecules in size-selected samples according to an embodiment of the present invention.

[0096] Figure 91 Statistical table of the number of plasma DNA molecules carrying informative SNPs between size-selected samples and non-size-selected samples according to an embodiment of the present invention.

[0097] Figure 92 Table of methylation levels in size-selected plasma DNA samples and non-size-selected plasma DNA samples according to an embodiment of the present invention.

[0098] Figure 93 Table of methylation levels in maternal or fetal-specific free DNA molecules according to an embodiment of the present invention.

[0099] Figure 94 Table of the top 10 end motifs in size-selected samples and non-size-selected samples according to an embodiment of the present invention.

[0100] Figure 95 Receiver operating characteristic (ROC) graph showing enhanced origin tissue efficacy analysis of long plasma DNA molecules according to an embodiment of the present invention.

[0101] Figure 96 Illustrating the principle of airport sequencing for plasma DNA molecules according to an embodiment of the present invention.

[0102] Figure 97 Table of the percentage of plasma DNA molecules within a specific size range and their corresponding methylation levels according to an embodiment of the present invention.

[0103] Figure 98 Diagram of size distribution and methylation patterns across different sizes according to an embodiment of the present invention.

[0104] Figure 99 Table of fetal DNA fractions determined using nanopore sequencing according to an embodiment of the present invention.

[0105] Figure 100 Table of the methylation levels between fetal-specific DNA molecules and maternal-specific DNA molecules according to an embodiment of the present invention.

[0106] Figure 101A table of the percentage of plasma DNA molecules within a specific size range and their corresponding degrees of methylation of fetal DNA molecules and maternal DNA molecules according to an embodiment of the present invention.

[0107] Figure 102A and Figure 102B A size distribution diagram of fetal DNA molecules and maternal DNA molecules determined by nanopore sequencing according to an embodiment of the present invention.

[0108] Figure 103 A diagram showing the difference in the degree of methylation between fetal DNA molecules and maternal DNA molecules based on a single informative SNP and two informative SNPs according to an embodiment of the present invention.

[0109] Figure 104 A table of the difference in the degree of methylation between fetal DNA molecules and maternal DNA molecules according to an embodiment of the present invention.

[0110] Figure 105 Illustrates a measurement system according to an embodiment of the present invention.

[0111] Figure 106 Displays a computer system according to an embodiment of the present invention.

[0112] The term

[0113] "tissue" corresponds to a group of cells grouped as a functional unit in a pregnant individual or their fetus. More than one type of cell may be found in a single tissue. Different types of tissues may be composed of different types of cells (e.g., hepatocytes, alveolar cells, or blood cells), but may also correspond to tissues from different organisms (mother and fetus; tissues in a pregnant individual who has received a transplant; tissues of a pregnant organism or its fetus infected with a microorganism or virus). "Reference tissue" may correspond to the tissue used to determine the tissue-specific degree of methylation. Multiple samples of the same tissue type from different pregnant individuals or their fetuses may be used to determine the tissue-specific degree of methylation of the tissue type.

[0114] "Biological sample" refers to any sample taken from a pregnant individual (such as a human (or other animal), such as a pregnant woman, a diseased person or a pregnant person suspected of having a disease, a pregnant organ transplant recipient or a pregnant individual suspected of having a disease course involving an organ (such as the heart in myocardial infarction or the brain in stroke or the hematopoietic system in anemia)) and containing one or more nucleic acid molecules of interest. The biological sample can be a body fluid, such as blood, plasma, serum, urine, vaginal fluid, vaginal lavage fluid, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirate from different parts of the body (such as the thyroid, breast), intraocular fluid (such as aqueous humor), etc. Stool samples can also be used. In various embodiments, most of the DNA in a biological sample that has been enriched for cell-free DNA (such as a plasma sample obtained via a centrifugation protocol) can be cell-free, for example, more than 50%, 60%, 70%, 80%, 90%, 95% or 99% of the DNA can be cell-free. The centrifugation protocol can include, for example, obtaining a fluid fraction by centrifugation at 3,000 g for 10 minutes and then centrifuging again at, for example, 30,000 g for 10 minutes to remove residual cells. As part of the analysis of the biological sample, a statistically significant number of cell-free DNA molecules in the biological sample can be analyzed (such as to provide an accurate measurement result). In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 or more cell-free DNA molecules can be analyzed. At least the same number of sequence reads can be analyzed.

[0115] "Sequence read" refers to a string of nucleotides sequenced from any part or all of a nucleic acid molecule. For example, a sequence read can be a short string of nucleotides (such as 20 - 150 nucleotides) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained in various ways, such as using sequencing technologies or using probes, such as in hybridization arrays or capture probes as can be used in microarrays; or amplification technologies, such as polymerase chain reaction (PCR) or linear amplification or isothermal amplification using a single primer. As part of the analysis of the biological sample, a statistically significant number of sequence reads can be analyzed, for example, at least 1,000 sequence reads can be analyzed. As other examples, at least 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 or more sequence reads can be analyzed.

[0116] "Locus" (also referred to as "genomic locus") corresponds to a single locus, which can be a single base position or a set of related base positions, such as a CpG locus or a larger set of related base positions. "Locus" can correspond to a region containing multiple loci. A locus can contain only one locus, in which case the locus is equivalent to a locus in that case.

[0117] "Methylation state" refers to the methylation state at a given locus. For example, a locus can be methylated, unmethylated, or in some cases undetermined.

[0118] The "methylation index" of each genomic locus (e.g., CpG locus) can refer to the proportion of DNA fragments (e.g., as determined by sequence reads or probes) that show methylation at the locus relative to the total number of reads covering the locus. "Read" can correspond to information obtained from a DNA fragment (e.g., the methylation state at a locus). Reads can be obtained using reagents (e.g., primers or probes) that preferentially hybridize to DNA fragments having a specific methylation state at one or more loci. Typically, the reagent is applied after treatment with a method that differentially modifies or differentially identifies DNA molecules depending on the methylation state of the DNA molecule, such as bisulfite conversion, or methylation-sensitive restriction enzymes, or methylation-binding proteins, or anti-methylcytosine antibodies, or single-molecule sequencing techniques that identify methylcytosine and hydroxymethylcytosine (e.g., single-molecule real-time sequencing and nanopore sequencing (e.g., from Oxford Nanopore Technologies)).

[0119] The "methylation density" of a region can refer to the number of reads at sites within the region showing methylation divided by the total number of reads covering the sites in the region. The sites can have specific characteristics, such as being CpG sites. Thus, the "CpG methylation density" of a region can refer to the number of reads showing CpG methylation divided by the total number of reads covering CpG sites (such as specific CpG sites, CpG sites within CpG islands, or larger regions) in the region. For example, the methylation density of each 100 kb bin in the human genome can be determined by the total number of cytosines that are not converted at CpG sites (which correspond to methylated cytosines) after bisulfite treatment as a proportion of all CpG sites covered by sequence reads mapped to the 100 kb region. This analysis can also be performed for other bin sizes, such as 500 bp, 5 kb, 10 kb, 50 kb, or 1 Mb, etc. The region can be the entire genome or a chromosome or a part of a chromosome (such as a chromosome arm). The methylation index of a CpG site is the same as the methylation density of the region when the region contains only that CpG site. The "proportion of methylated cytosines" can refer to the number of cytosine sites "C's" that are shown to be methylated (such as not converted after bisulfite conversion) relative to the total number of cytosine residues analyzed, i.e., the total number of cytosines in the region excluding the CpG cases. Methylation index, methylation density, molecular count methylated at one or more sites, and proportion of molecules (such as cytosines) methylated at one or more sites are examples of "degree of methylation". In addition to bisulfite conversion, other methods known to those skilled in the art can be used to interrogate the methylation status of DNA molecules, including but not limited to enzymes sensitive to methylation status (such as methylation-sensitive restriction enzymes), methyl-binding proteins, single molecule sequencing using platforms sensitive to methylation status (such as nanopore sequencing (Schreiber et al., Proc Natl Acad Sci 2013; 110: 18910 - 18915) and single molecule real-time sequencing (such as single molecule real-time sequencing from Pacific Biosciences) (Flusberg et al., Nat Methods 2010; 7: 461 - 465)).

[0120] A "methylome" provides a measure of the amount of DNA methylation at multiple sites or loci in the genome. The methylome can correspond to all of the genome, a substantial portion of the genome, or one or more relatively small parts of the genome.

[0121] "Methylation profile" contains information related to DNA or RNA methylation at multiple sites or regions. Information related to DNA methylation may include, but is not limited to, methylation index of CpG sites, methylation density (abbreviated as MD) of CpG sites in a region, distribution of CpG sites on contiguous regions, methylation pattern or degree of each individual CpG site within a region containing more than one CpG site, and non-CpG methylation. In one embodiment, the methylation profile may include methylation or non-methylation patterns of more than one type of base (e.g., cytosine or adenine). A substantial portion of the methylation profile of a genome may be considered equivalent to the methylome. "DNA methylation" in mammalian genomes generally refers to the addition of a methyl group to the 5'-carbon of the cytosine residue in CpG dinucleotides (i.e., 5-methylcytosine). DNA methylation can also occur in cytosine in other contexts such as CHG and CHH, where H is adenine, cytosine, or thymine. Cytosine methylation can also be in the form of 5-hydroxymethylcytosine. Non-cytosine methylation has also been reported, such as N 6 -methyladenine.

[0122] "Methylation pattern" refers to the order of methylated and non-methylated bases. For example, a methylation pattern can be the order of methylated bases on a single DNA strand, a single double-stranded DNA molecule, or another type of nucleic acid molecule. As an example, three consecutive CpG sites can have any of the following methylation patterns: UUU, MMM, UMM, UMU, UUM, MUM, MUU, or MMU, where "U" indicates an unmethylated site and "M" indicates a methylated site. When we extend this concept to include base modifications not limited to methylation, we will use the term "modification pattern", which refers to the order of modified and unmodified bases. For example, a modification pattern can be the order of modified bases on a single DNA strand, a single double-stranded DNA molecule, or another type of nucleic acid molecule. As an example, three consecutive potentially modifiable sites can have any of the following modification patterns: UUU, MMM, UMM, UMU, UUM, MUM, MUU, or MMU, where "U" indicates an unmodified site and "M" indicates a modified site. An example of a base modification not based on methylation is an oxidative change such as in 8-oxo-guanine.

[0123] The terms "hypermethylation" and "hypomethylation" can refer to the methylation density of a single DNA molecule, as measured by its single-molecule methylation level, e.g., the number of methylated bases or nucleotides within the molecule divided by the total number of methylatable bases or nucleotides within the molecule. A hypermethylated molecule is one in which the single-molecule methylation level is equal to or higher than a threshold, which can be defined according to different applications. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%. A hypomethylated molecule is one in which the single-molecule methylation level is equal to or lower than a threshold, which can be defined and can vary according to different applications. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%.

[0124] The terms "hypermethylation" and "hypomethylation" can also refer to the methylation level of a population of DNA molecules, as measured by their multi-molecule methylation level. A hypermethylated population of molecules is one in which the multi-molecule methylation level is equal to or higher than a threshold, which can be defined and can vary according to different applications. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%. A hypomethylated population of molecules is one in which the multi-molecule methylation level is equal to or lower than a threshold, which can be defined according to different applications. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, and 95%. In one embodiment, a population of molecules can be aligned with one or more selected genomic regions. In one embodiment, one or more selected genomic regions can be disease-related, such as a genetic disorder, a memory disorder, a metabolic disorder, or a neurological disorder. The length of one or more selected genomic regions can be 50 nucleotides (nt), 100 nt, 200 nt, 300 nt, 500 nt, 1000 nt, 2 knt, 5 knt, 10 knt, 20 knt, 30 knt, 40 knt, 50 knt, 60 knt, 70 knt, 80 knt, 90 knt, 100 knt, 200 knt, 300 knt, 400 knt, 500 knt, or 1 Mnt.

[0125] The term "sequencing depth" refers to the number of times a locus is covered by sequence reads aligned to the locus. The locus can be as small as a nucleotide, or as large as a chromosomal arm, or as large as an entire genome. The sequencing depth can be expressed as 50×, 100×, etc., where "×" refers to the number of times the locus is covered by sequence reads. The sequencing depth can also be applied to multiple loci or the entire genome, in which case, "×" can refer to the average number of times the locus or the haploid genome or the entire genome is sequenced, respectively. Ultra-deep sequencing can refer to a sequencing depth of at least 100×.

[0126] A "calibration sample" can correspond to the following biological samples: the fractional concentration of clinically relevant DNA (e.g., tissue-specific DNA fraction) of which is known or determined via a calibration method, such as using alleles specific to the tissue, e.g., in transplantation in a pregnant individual, where these alleles are present in the donor genome but not in the recipient genome which can be used as a marker for the transplanted organ. As another example, a calibration sample can correspond to a sample from which the end motif can be determined. The calibration sample can be used for two purposes.

[0127] A "calibration data point" contains a "calibration value" and the measured or known fractional concentration of clinically relevant DNA (e.g., DNA of a specific tissue type). The calibration value can be determined by, e.g., the relative frequency (e.g., aggregate value) measured for a calibration sample, the fractional concentration of clinically relevant DNA of which is known. The calibration data point can be defined in various ways, e.g., as discrete points or as a calibration function (also called a calibration curve or calibration surface). The calibration function can be derived from additional mathematical transformation of the calibration data points.

[0128] A "separation value" corresponds to a difference or ratio involving two values, e.g., two contributing fractions or two degrees of methylation. The separation value can be a simple difference or ratio. As an example, the direct ratio of x / y and x / (x + y) is a separation value. The separation value can contain other factors such as multiplicative factors. As other examples, the difference or ratio of functions of the values can be used, e.g., the difference or ratio of the natural logarithm (ln) of two values. The separation value can contain both differences and ratios.

[0129] "Separation values" and "aggregate values" (e.g., belonging to relative frequencies) are two examples of parameters (also called metrics) that provide a measure of the sample that varies between different classifications (states), and can thus be used to determine different classifications. An aggregate value can be a separation value, e.g., when taking the difference between a set of relative frequencies of a sample and a set of reference relative frequencies, as is commonly done in clustering.

[0130] As used herein, the term "classification" refers to any one or more numbers or one or more other characters associated with a particular characteristic of a sample. For example, the symbol "+" (or the word "positive") can indicate that the sample is classified as having a deletion or an amplification. A classification can be binary (e.g., positive or negative) or have more classification levels (e.g., a scale from 1 to 10 or from 0 to 1).

[0131] As used herein, the term "parameter" means a numerical value that characterizes a quantitative data set and / or a numerical relationship between quantitative data sets. For example, the ratio (or a function of the ratio) between a first amount of a first nucleic acid sequence and a second amount of a second nucleic acid sequence is a parameter.

[0132] The term "size profile" generally relates to the sizes of DNA fragments in a biological sample. A size profile can be a histogram that provides a distribution of a certain amount of DNA fragments of various sizes. Various statistical parameters (also referred to as size parameters or simply parameters) can be used to distinguish one size profile from another. One parameter is the percentage of DNA fragments of a particular size or size range relative to all DNA fragments or relative to DNA fragments of another size or range.

[0133] The terms "cutoff value" and "threshold value" refer to a predetermined numerical value used in an operation. For example, a cutoff size can refer to the size above which fragments are excluded. A threshold value can be a value above or below which a particular classification applies. Either of these terms can be used in either of these situations. A cutoff value or a threshold value can be a "reference value" that indicates a particular classification or discriminates between two or more classifications or is derived from such a reference value. As those skilled in the art will appreciate, such reference values can be determined in various ways. For example, a metric can be determined for two different groups of individuals with different known classifications, and a reference value can be selected to represent a classification (e.g., an average) or a value between the metrics of the two groups (e.g., selected to obtain a desired sensitivity and specificity). As another example, a reference value can be determined based on a statistical analysis or simulation of a sample. A particular cutoff value, threshold value, reference value, etc. can be determined based on a desired accuracy (e.g., sensitivity and specificity).

[0134] "Pregnancy-related disorders" include any disorder characterized by an abnormal relative expression level of a gene in maternal and / or fetal tissue or an abnormal clinical feature in the mother and / or fetus. These disorders include but are not limited to pre-eclampsia in the mother (Kaartokallio et al., Sci Rep. 2015;5:14107; Medina-Bastidas et al., Int J Mol Sci. 2020;21:3597), intrauterine growth restriction (Faxén et al., Am J Perinatol. 1998;15:9-13; Medina-Bastidas et al., Int J Mol Sci. 2020;21:3597), invasive placentation, preterm birth (Enquobahrie et al., BMC Pregnancy Childbirth. 2009;9:56), haemolytic disease of the newborn, placental insufficiency (Kelly et al., Endocrinology. 2017;158:743-755), hydrops fetalis (Magor et al., Blood. 2015;125:2405-17), fetal malformations (Slonim et al., Proc Natl Acad Sci USA. 2009;106:9425-9), HELLP syndrome (Dijk et al., J Clin Invest. 2012;122:4003-4011), systemic lupus erythematosus (Hong et al., J Exp Med. 2019;216:1154-1169) and other immune diseases.

[0135] The abbreviation "bp" refers to base pair. In some cases, "bp" may be used to denote the length of a DNA fragment, even if the DNA fragment may be single-stranded and does not contain base pairs. In the case of single-stranded DNA, "bp" may be interpreted as providing the length in nucleotides.

[0136] The abbreviation "nt" refers to nucleotide. In some cases, "nt" may be used to denote the length of single-stranded DNA in bases. In addition, "nt" may be used to denote a relative position, such as upstream or downstream of a locus being analysed. For double-stranded DNA, unless the context clearly indicates otherwise, "nt" may still refer to the single-stranded length rather than the total number of nucleotides in both strands. In some cases regarding technical conceptualization, data display, processing and analysis, "nt" and "bp" may be used interchangeably.

[0137] The term "machine learning model" can include a model that makes predictions on test data based on the use of sample data (e.g., training data), and can thus include supervised learning. Machine learning models are often developed using a computer or a processor. A machine learning model can include a statistical model.

[0138] The term "data analysis framework" can include algorithms and / or models that can take data as input and then output the predicted results. Examples of "data analysis frameworks" include statistical models, mathematical models, machine learning models, other artificial intelligence models, and combinations thereof.

[0139] The term "real-time sequencing" can refer to techniques involving data collection or monitoring during the progress of the reactions involved in sequencing. For example, real-time sequencing can involve optically monitoring or imaging DNA polymerase as new bases are added.

[0140] The term "subsequence" can refer to a string of bases that is less than the complete sequence corresponding to a nucleic acid molecule. For example, when the complete sequence of a nucleic acid molecule contains 5 or more bases, a subsequence can contain 1, 2, 3, or 4 bases. In some embodiments, a subsequence can refer to a string of bases that forms a unit, where the unit is repeated multiple times in a tandem continuous manner. Examples include 3nt units or subsequences that are repeated at loci associated with trinucleotide repeat disorders, 1nt to 6nt units or subsequences that are repeated 5 to 50 times as microsatellites, 10nt to 60nt units or subsequences that are repeated 5 to 50 times as minisatellites, or include repeats such as Alu repeats in other genetic elements.

[0141] The term "about / approximately" can mean within an acceptable error range of a particular value as determined by a person of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, in accordance with practice in the art, "about" can mean within 1 or more standard deviations. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or methods, the term "about" can mean within an order of magnitude of the value, within 5-fold, and more preferably within 2-fold. When a particular value is recited in this application and the claims, unless otherwise stated, it should be assumed that the term "about" means within an acceptable error range of the particular value. The term "about" can have the meaning as commonly understood by a person of ordinary skill in the art. The term "about" can refer to ±10%. The term "about" can refer to ±5%.

[0142] In the case of providing a range of values, it should be understood that, unless the context clearly indicates otherwise, each intervening value between the upper and lower limits of the range is also specifically disclosed, to the tenth of the unit of the lower limit. Each smaller range between any of the stated values or intervening values within the stated range and any other stated value or intervening value within the stated range is encompassed within the embodiments of the present disclosure. The upper and lower limits of these smaller ranges may independently be included within or excluded from the range, and each range where either limit, no limit, or both limits are included within the smaller range is also encompassed within the present disclosure, subject to any specifically excluded limits within the stated range. If the stated range includes one or both of the limits, ranges excluding either or both of the included limits are also included within the present disclosure.

[0143] Standard abbreviations may be used, such as bp, base pair; kb, kilobase; pi, picoliter; s or sec, second; min, minute; h or hr, hour; aa, amino acid; nt, nucleotide; and the like.

[0144] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Although any methods and materials similar or equivalent to those described herein may be used in the practice or testing of the embodiments of the present disclosure, some potential and exemplary methods and materials are now described. Detailed Description

[0145] Analysis of cell-free DNA molecules mainly involves short cell-free DNA fragments, which is often attributed to the limitations of analysis techniques. The limited ability to obtain sequence information from long DNA molecules using Illumina sequencing technology was demonstrated in recent sequencing results of mouse cell-free DNA (Serpas et al., Proceedings of the National Academy of Sciences of the United States of America 2019; 116:641 - 649). Only 0.02% of the sequenced DNA molecules in wild-type mice using Illumina sequencing were in the range of 600 bp and 2000 bp. Even when sequencing a DNA library originally prepared for Illumina sequencing using the single molecule real-time (SMRT) technology from Pacific Biosciences (i.e., PacBio SMRT sequencing), there were still only 0.33% of the sequenced DNA molecules in the range of 600 bp and 2000 bp. These reported data indicate that the sequencing step will lose 93% of the long DNA molecules in the range of 600 bp and 2000 bp present in the original DNA library.

[0146] We speculate that due to the limitations of PCR in amplifying the long DNA molecules described above, the DNA library preparation step will also lose a significant proportion of the long cfDNA molecules. Jahr et al. reported the presence of large-sized fragments of many kilobases, such as ~10,000 kilobases, using gel electrophoresis (Jahr et al., Cancer Res. 2001; 61:1659-65). However, the bands shown in the gel electrophoresis images will not be readily amenable to providing sequence information of these molecules in the gel, let alone epigenetic information.

[0147] We have previously used the Oxford Nanopore Technologies sequencing platform to study cfDNA extracted from maternal plasma (Cheng et al., Clin Chem. 2015; 61:1305-6). We observed a very low proportion of long plasma DNA greater than 1 kb (0.06% to 0.3%). We hypothesized that such low percentages might be the result of the low sequencing accuracy of this platform.

[0148] In the field of cfDNA, most studies have focused on short DNA molecules (e.g., <600 bp). The properties of the genetic and epigenetic information contained in long cfDNA molecules are unexplored. The present disclosure provides a systematic way to analyze long cfDNA molecules, the systematic way including decoding their genetic and epigenetic information and their clinical utility in non-invasive prenatal testing, such as but not limited to non-invasive detection of single-gene disorders, elucidation of the fetal genome (e.g., non-invasive whole fetal genome sequencing), detection of de novo mutations at the whole genome level, and detection / monitoring of pregnancy-related disorders such as preeclampsia and preterm labor.

[0149] I. cfDNA Size Analysis

[0150] cfDNA samples obtained from pregnant women were sequenced, and a significant portion of the DNA fragments were found to be long. Accurate sequencing of the long cfDNA fragments was demonstrated. The size profiles of these long cfDNA molecules were analyzed. The amounts of fetal long cfDNA molecules and maternal long cfDNA molecules were compared. The long cfDNA molecules could be more accurately aligned with a reference genome. The long cfDNA molecules could be used to determine haplotypic inheritance.

[0151] A plasma DNA sample from a pregnant woman in the third trimester of pregnancy was analyzed using PacBio SMRT sequencing. Double-stranded cfDNA molecules were ligated to hairpin adapters and subjected to single-molecule real-time sequencing using zero-mode waveguides and a single polymerase molecule (Eid et al., Science. 2009; 323:133-8).

[0152] We sequenced 1.1 billion subreads, of which 659.3 million subreads could be aligned to the human reference genome (hg19). The subreads were generated from 4.6 million PacBio single molecule real-time (SMRT) sequencing pores, and each pore contained at least one subread that could be aligned to the human reference genome. On average, each molecule in the SMRT pores was sequenced 143 times. In this example, there were 4.5 million circular consensus sequences (CCS), indicating 4.5 million free DNA molecules available for downstream analysis. The size of each free DNA was determined by the CCS by counting the number of bases that had been identified.

[0153] Figure 1A and Figure 1B shows the size distribution of free DNA from 0 kb to 20 kb. The y-axis shows the frequency. The x-axis shows the size in base pairs from 0 kb to 20 kb on a linear scale ( Figure 1A ) or a logarithmic scale ( Figure 1B ). Since the sequencing is performed on full-length DNA molecules, the size of each DNA molecule can be directly determined by counting the number of nucleotides in the subreads or CCS. The DNA fragment size measurement can be achieved using any sequencing platform that can read through the full-length DNA fragment and is not limited to the use of single molecule sequencers. For example, a Sanger sequencer can read through 800 bp. Short read sequencing such as by the Illumina platform can read through 250 bp. Single molecule sequencers such as Pacific Biosciences and Oxford Nanopore can read through more than 10,000 bp. The size of the DNA fragment can also be determined after alignment to a reference genome such as the human reference genome. The size of the DNA fragment can be determined by paired-end sequencing followed by alignment to the reference genome. Figure 1B shows a long-tailed pattern. Among the 4.5 million CCS, there were 22.5% of free DNA larger than 200 bp, 19.0% larger than 300 bp, 11.8% larger than 400 bp, 10.6% larger than 500 bp, 8.9% larger than 600 bp, 6.4% larger than 1 kb, 3.5% larger than 2 kb, 1.9% larger than 3 kb, 0.9% larger than 4 kb, and 0.04% larger than 10 kb. The longest free DNA observed in the current PacBio SMRT results was 29,804 bp.

[0154] Plasma DNA from a pregnant individual was also sequenced on an Illumina sequencing platform using a PCR-based library preparation protocol (Lun et al. Clin Chem 2013; 59: 1583-94). Among 18.2 million paired-end reads, there was 5.3% cell-free DNA greater than 200 bp, 2.0% greater than 300 bp, 0.3% greater than 400 bp, 0.2% greater than 500 bp, 0.2% greater than 600 bp (Table 1). As a comparison, we analyzed the size profile by pooling single molecule real-time sequencing data from 5 pregnant individuals (i.e., a total of 4.4 million CCS). Compared to the corresponding plasma DNA molecules greater than 600 bp (0.2%) obtained by the Illumina sequencing platform, we observed more plasma DNA molecules greater than 600 bp (28.56%). These results indicate that PacBio SMRT sequencing enables us to achieve more than 143-fold longer DNA molecules (longer than 600 bp). We could obtain 4.77% plasma DNA molecules greater than 3 kb using single molecule real-time sequencing, but there were no reads in the Illumina sequencing platform.

[0155] In contrast to previous reports (Cheng et al. Clin Chem 2015; 61: 1305-6) showing a very small proportion of long plasma DNA molecules greater than 1 kb (0.06% to 0.3%) using an Oxford Nanopore Technologies sequencing platform, we could obtain more than 21-fold greater plasma DNA greater than 1 kb (6.4%), indicating that PacBio SMRT sequencing is much more efficient in obtaining sequence information from long DNA populations.

[0156] Compared to paired-end short-read sequencing such as that of the Illumina sequencing platform, long-read sequencing technologies such as PacBio SMRT technology have several advantages in characterizing long DNA fragments (such as length). For example, long reads generally allow us to align more accurately to the human reference genome (such as hg19). Long-read technologies will also allow us to accurately determine the length of plasma DNA molecules by directly counting the number of nucleotides sequenced. In contrast, plasma DNA size assessment based on paired-end short reads is an indirect method of inferring the size of plasma DNA molecules using the outermost coordinates of aligned paired-end reads. For such indirect approaches, errors in alignment will cause inaccurate size inference. In this regard, an increase in the size span between paired-end reads will increase the probability of errors in alignment.

[0157]

[0158] Table 1. Comparison of size distributions between PacBio and Illumina sequencing of cell-free DNA.

[0159] Figure 2A and Figure 2B show the size distribution of cell-free DNA from 0 kb to 5 kb. The y-axis shows the frequency. The x-axis shows the size in base pairs from 0 kb to 5 kb on a linear scale ( Figure 2A ) or a logarithmic scale ( Figure 2B ). There is a series of main peaks that appear with a periodic pattern. The periodic pattern even extends to molecules in the range of 1 kb and 2 kb. The peak with the highest frequency (2.6%) is below 166 bp, which is consistent with previous findings using Illumina technology (Lo et al., Science Translational Medicine 2010; 2:61ra91). Figure 2B The distance between adjacent main peaks in

[0160] Figure 3A and Figure 3B show the size distribution of cell-free DNA from 0 bp to 400 bp. The y-axis shows the frequency. The x-axis shows the size in base pairs from 0 bp to 400 bp on a linear scale ( Figure 3A ) or a logarithmic scale ( Figure 3B ). The characteristic features previously reported (Lo et al., Science Translational Medicine 2010; 2:61ra91) of having the most main peak below 166 bp and 10-bp periodicity in molecules smaller than 166 bp can also be reproduced using the novel methods of the present disclosure. These results indicate that determining the molecule size by counting the number of bases from single-molecule sequencing of the present disclosure is reliable.

[0161] A. Size analysis for fetal DNA and maternal DNA

[0162] The sizes of maternal and fetal DNA fragments were analyzed and compared. As an example, buffy coat DNA from a pregnant woman and matched placental DNA were sequenced to obtain 59× and 58× haploid genome coverage, respectively. We identified a total of 822,409 informative single nucleotide polymorphisms (SNPs) where the mother was homozygous and the fetus was heterozygous. Fetal-specific alleles were defined as those present in the fetal genome but not in the maternal genome. We identified 2,652 fetal-specific fragments and 24,837 shared fragments (i.e., fragments carrying shared alleles; mainly of maternal origin) in maternal plasma (M13160) via PacBio sequencing. The fetal DNA fraction was 21.8%.

[0163] Figure 4A and Figure 4BShow the size distribution of cell-free DNA between the fragments carrying shared alleles (shared) and the fragments carrying fetal-specific alleles (fetal-specific). The x-axis shows the size in base pairs from 0 kb to 20 kb on a linear scale ( Figure 4A ) or a logarithmic scale ( Figure 4B ). Both the fragments carrying shared alleles (mostly of maternal origin) and the fragments carrying fetal-specific alleles (of placental origin) exhibit a long-tailed distribution, indicating the presence of long DNA molecules from both fetal and maternal sources. For the fragments of predominantly maternal origin, 22.6% of the plasma DNA molecules have a size greater than 2 kb, while for the fragments of fetal origin, 8.5% of the plasma DNA molecules have a size greater than 2 kb. These results suggest that fetal DNA molecules contain fewer long DNA molecules. The percentage of long DNA in this SNP-based analysis of the fetal and maternal origins of plasma DNA seems to be much higher than the percentage of long DNA observed in the total size analysis. The difference may be attributed to the fact that the probability of a long DNA molecule covering one or more SNPs is higher than that of a short DNA molecule covering one or more SNPs, and thus long DNA will be preferentially selected for SNP-based analysis. The relative proportion of long DNA molecules with added SNP tags deviating from the proportion of long DNA in the corresponding original pool will be controlled by the size of those molecules. Among those fetal-specific DNA fragments, the longest DNA fragment is 16,186 bp, while among those fragments carrying shared alleles, the longest fragment is 24,166 bp.

[0164] Figure 5A and Figure 5B Show the size distribution of cell-free DNA between the fragments carrying shared alleles (shared) and the fragments carrying fetal-specific alleles (fetal-specific). The x-axis shows the size in base pairs from 0 kb to 5 kb on a linear scale ( Figure 5A ) or a logarithmic scale ( Figure 5B ). For those fragments less than 2 kb for both fetal-specific DNA fragments and shared DNA fragments, there is a series of main peaks that appear in a periodic manner. The main peaks may be aligned with the nucleosome structure.

[0165] Figure 6A and Figure 6B Show the size distribution of cell-free DNA between the fragments carrying shared alleles (shared) and the fragments carrying fetal-specific alleles (fetal-specific). The x-axis shows the size in base pairs from 0 kb to 5 kb on a linear scale ( Figure 6A ) or a logarithmic scale ( Figure 6B) Sizes in base pairs from 0 kb to 1 kb. For those fragments less than 1 kb for both fetal-specific DNA fragments and shared DNA fragments, there is a series of main peaks that appear in a periodic manner. The main peaks may align with the nucleosome structure. There appears to be an observable shift of the fetal DNA size profile towards the left of the shared DNA fragment size profile, indicating that fetal DNA will include more short DNA molecules than maternal DNA.

[0166] Figure 7A and Figure 7B show the size distribution of cell-free DNA between fragments carrying shared alleles (shared) and fragments carrying fetal-specific alleles (fetal-specific). The x-axis shows the size in base pairs from 0 bp to 400 bp on a linear scale ( Figure 7A ) or a logarithmic scale ( Figure 7B ). The characteristic features previously reported (Lo et al., Science Translational Medicine 2010; 2:61ra91) of having the most main peak at 166 bp and a 10-bp periodicity in both fetal and maternal molecules less than 166 bp can also be reproduced using the novel methods of the present disclosure. These results indicate that it is reliable to determine the molecule size by counting the number of bases from single-molecule sequencing of the present disclosure.

[0167] B. Size and methylation analysis

[0168] The methylation levels of long cell-free maternal and fetal DNA molecules were analyzed. It was found that the methylation level of fetal DNA molecules was lower than that of maternal DNA molecules.

[0169] In PacBio SMRT sequencing, DNA polymerase mediates the incorporation of fluorescently labeled nucleotides into the complementary strand. The characteristics of the fluorescence pulses generated during DNA synthesis, including the inter-pulse duration and pulse width, will reflect the polymerase kinetics, which can be used to determine nucleotide modifications such as, but not limited to, 5-methylcytosine using the approach described in our previously published case (U.S. Application No. 16 / 995,607, filed on August 17, 2020, entitled "DETERMINATION OF BASE MODIFICATIONS OF NUCLEIC ACIDS"), the entire content of which is incorporated herein by reference for all purposes.

[0170] In an embodiment, we separately identified 95,210 fragments carrying maternal-specific alleles and 2,652 fragments carrying fetal-specific alleles. Maternal-specific alleles are herein defined as those alleles that are present in the maternal genome but not in the fetal genome and that can be identified from SNPs in which the mother is heterozygous and the fetus is homozygous. In this example, we identified a total of 677,375 of such informative SNPs. We determined the size of each cell-free DNA molecule. In one embodiment, when the methylation status in the genome is variable, for example, the degree of methylation of CpG islands is generally lower than that of regions without CpG islands, to minimize the variability introduced by the genomic situation, we can computationally select fragments greater than 1 kb, containing at least 5 CpG sites and corresponding to a CpG density of less than 5% (i.e., the number of CpG sites in the molecule divided by the total length of the molecule < 0.05) for downstream analysis.

[0171] Figure 8 Shows the single-molecule double-stranded DNA methylation levels between fragments carrying maternal-specific alleles and fragments carrying fetal-specific alleles. The y-axis shows the single-molecule double-stranded DNA methylation level in percentage. The x-axis shows both fragments carrying maternal-specific alleles and fragments carrying fetal-specific alleles. The single-molecule double-stranded DNA methylation level of fragments carrying fetal-specific alleles (mean: 62.7%; interquartile range IQR: 50.0% - 77.2%) is lower than the corresponding single-molecule double-stranded DNA methylation level of fragments carrying maternal-specific alleles (mean: 72.7%; IQR: 60.6% - 83.3%) (P < 0.0001).

[0172] Figure 9A Shows the empirical distribution of the single-molecule double-stranded DNA methylation levels of fragments fitted by kernel density estimation implemented in the R package (r-project.org / ). The frequency is shown on the y-axis. The x-axis shows the single-molecule double-stranded DNA methylation level in percentage. The distribution of fetal-specific long DNA fragments is to the left of the distribution of maternal-specific fragments, indicating that lower single-molecule double-stranded DNA methylation levels are present in fetal DNA molecules.

[0173] Figure 9BShows receiver operating characteristic (ROC) analysis using the degree of single-molecule double-stranded DNA methylation. The y-axis shows sensitivity. The x-axis shows specificity. ROC analysis was performed using the degree of single-molecule double-stranded DNA methylation to study the power of differentiating fetal DNA fragments from maternal DNA fragments using the degree of single-molecule double-stranded DNA methylation, and the area under the ROC curve (AUC) was found to be 0.62, greater than the random guessing result of 0.5. In embodiments, we can utilize the spatial pattern of the methylation state (such as the sequence of the methylation state), the relative or absolute distance between modified bases and genomic coordinates in a single molecule to further improve the determination of the fetal / maternal origin of fragments in plasma. In embodiments, we can combine methylation patterns with other fragmentomic metrics (i.e., parameters regarding DNA fragmentation), which include but are not limited to preferred ends (Chan et al., Proceedings of the National Academy of Sciences of the United States of America 2016; 113:E8159-8168), end motifs (Serpas et al., Proceedings of the National Academy of Sciences of the United States of America 2019; 116:641-649), size (Lo et al., Science Translational Medicine 2010; 2:61ra), orientation awareness (i.e., regarding specific elements within the genome, such as the orientation of open chromatin regions, fragmentation patterns (Sun et al., Genomes Res. 2019; 29:418-427)), topological form (e.g., linear vs. circular DNA molecules (Ma et al., Clinical Chemistry 2019; 65:1161-1170)), thereby improving the classification power of differentiating fragments of placental origin (fetal origin).

[0174] Figure 10A and Figure 10B Shows the degree of single-molecule double-stranded DNA methylation of both fetal DNA fragments and maternal DNA fragments according to the change in fragment size. The y-axis shows the degree of single-molecule double-stranded DNA methylation in percentage. The x-axis shows sizes from 0 kb to greater than 20 kb ( Figure 10A ) and sizes from 0 kb to greater than 1 kb ( Figure 10B ). On the other hand, within both long range ( Figure 10A ) and short range ( Figure 10B ), the degree of single-molecule double-stranded DNA methylation of fetal-specific DNA molecules is generally lower than that of maternal-specific DNA molecules. For short DNA molecules, this finding is consistent with the current understanding that the methylation degree of fetal DNA in pregnant women's plasma is lower than that of maternal DNA (Lun et al., Clinical Chemistry 2013; 59:1583-94).

[0175] In an embodiment, since the methylation level of fetal DNA molecules is relatively lower than that of maternal DNA molecules, we will select molecules with a single-molecule double-stranded DNA methylation level less than a specific threshold such as but not limited to 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, and 5% to enrich cell-free DNA molecules of fetal origin in the plasma DNA pool. For example, for fragments >1 kb, the fetal DNA fraction is 2.6%. If we select fragments >1 kb with a single-molecule double-stranded methylation level <50%, the fetal DNA fraction of those further selected fragments >1 kb will increase to 5.6% (i.e., a 115.4% increase). In another example, for fragments <200 bp, the fetal DNA fraction is 26.2%. If we select fragments <200 bp with a single-molecule double-stranded methylation level <50%, the fetal DNA fraction of those further selected fragments >200 bp will increase to 41.6% (i.e., 58.8%). Thus, in some cases, the use of a threshold single-molecule double-stranded DNA methylation level for enriching fetal DNA will be more effective for long DNA molecules.

[0176] C. Haplotypes and Long Cell-Free DNA Methylation

[0177] In an embodiment, we can use the methods described in the present disclosure to obtain the base composition, size, and base modifications of each individual DNA molecule. The SNP and methylation information of long cell-free DNA molecules can be used for haplotype analysis. The use of long DNA molecules present in the cell-free DNA pool disclosed in the present disclosure will allow phasing of variants in the genome by using haplotype information present in each common sequence according to but not limited to published methods (Edge et al., Genome Res. 2017; 27:801 - 812; Wenger et al., Nat Biotechnol. 2019; 37:1155 - 1162). Embodiments for determining haplotypes based on the sequence information of cell-free DNA are different from previous studies that had to rely on long DNA prepared from tissue DNA. Haplotypes within a genomic region are sometimes referred to as haplotype blocks. A haplotype block can be regarded as a set of alleles that have been phased on a chromosome. In some embodiments, the haplotype block will extend as long as possible based on a set of sequence information that supports two alleles physically linked on a chromosome and allele overlap information between different sequences.

[0178] Figure 11A and Figure 11BExamples of identified long fetal-specific DNA molecules in maternal plasma DNA of pregnant women are shown. Among those fetal-specific DNA fragments, we hereby illustrate an example of our invention using a 16,186 bp molecule that aligns with a region in chromosome 10 (chr10:56282981-56299166) of the human reference genome ( Figure 11A ) and carries 7 fetal-specific alleles ( Figure 11B ). Six out of the 7 fetal-specific alleles are fetal-specific alleles that are consistent with the allele information inferred from deep sequencing of the maternal and fetal genomes (using the Illumina platform) ( Figure 11B ). According to the method described in this disclosure, its methylation level was determined to be 27.1% ( Figure 11B ), much lower than the average level (72.7%) of maternal-specific fragments. These results indicate that the single-molecule double-stranded DNA methylation pattern will serve as a marker to distinguish cell-free DNA molecules of fetal and maternal origin.

[0179] Figure 12A and Figure 12B Examples of identified long maternal DNA molecules carrying shared alleles in maternal plasma DNA of pregnant women are shown. Among those fragments carrying shared alleles, the longest fragment that aligns with a region in chromosome 6 (chr6:111074371-111098536) of the human reference ( Figure 12A ) and carries 18 shared alleles ( Figure 12B ) is 24,166 bp. All those shared alleles are consistent with the allele information inferred from deep sequencing of the maternal and fetal genomes (using the Illumina platform) ( Figure 12B ). According to the method described in this disclosure, its methylation level was determined to be 66.9% ( Figure 12B ). The genetic and epigenetic information of cell-free DNA molecules of about kilobase length cannot be easily identified by using short-read sequencing such as bisulfite sequencing (Illumina).

[0180] Here, we describe a method for determining the relative likelihood that a molecule is derived from a pregnant woman or a fetus. In pregnant women, DNA molecules carrying the fetal genotype actually originate from the placenta, while most DNA molecules carrying the maternal genotype originate from maternal blood cells. In this method, we first construct a frequency distribution curve of DNA molecules based on the methylation levels of DNA molecules for both the placenta and maternal blood cells. To achieve this, we divide the human genome into different sized bins.

[0181] Figure 13Show the frequency distributions of DNA from the placenta (red) and DNA from maternal blood cells (blue) at different resolutions from 1 kb to 20 kb according to the degree of methylation. The frequencies are shown on the y-axis. The degree of methylation is shown on the x-axis. Examples of bin sizes include, but are not limited to, 1 kb, 2 kb, 5 kb, 10 kb, 15 kb, and 20 kb. The degree of methylation of each bin is determined based on the number of methylated CpG sites divided by the total number of CpG sites. After determining the degree of methylation of all bins, frequency distribution curves for each of the placental genome and the maternal blood cell genome can be constructed for different bin sizes.

[0182] Based on the degree of methylation of long DNA molecules, the likelihood that they are from the placenta or maternal blood cells can be determined by the relative abundances of the two types of DNA molecules at such a degree of methylation and the fractional concentration of fetal DNA in the sample.

[0183] Let x and y be the frequencies of DNA molecules from the placenta and maternal blood cells, respectively, at a specific degree of methylation, and let f be the fractional concentration of fetal DNA in the sample.

[0184] The probability (P) that a DNA molecule is from the fetus can be calculated as:

[0185]

[0186] According to the previous example, consider plasma DNA molecules with a 16 kb and 27.1% degree of methylation.

[0187] Figure 14A and Figure 14B Show the frequency distributions of DNA from the placenta (red) and DNA from maternal blood cells (blue) within windows of 16 kb ( Figure 14A ) and 24 kb ( Figure 14B ) according to the degree of methylation. The frequencies are shown on the y-axis. The degree of methylation is shown on the x-axis. Based on the frequency distribution plot for the 16 kb fragment ( Figure 14A ), the frequencies of DNA molecules from the placenta and maternal blood cells are 0.6% and 0.08%, respectively. When the fetal DNA fraction is 21.8%, the probability that this DNA fragment is from the placenta is 64%, indicating an increased likelihood of placental origin.

[0188] The probability that a DNA molecule of a plasma DNA molecule with a 24 kb and 66.9% degree of methylation is from fetal tissue can also be calculated. Based on the frequency distribution plot for the 24 kb fragment, the frequencies of DNA molecules from the placenta and maternal blood cells are 0.05% and 0.16%, respectively ( Figure 14B)。The probability that this DNA fragment is derived from the placenta is 0.8%, indicating that it is highly unlikely to be of placental origin. In other words, there is a high likelihood that the molecule is of maternal origin.

[0189] This calculation can further consider the size of the DNA molecule by referring to the size distribution curves of fetal DNA and maternal DNA. The analysis can be performed, for example but not limited to, using Bayes's theorem, logistic regression, multiple regression, and support vector machines, random forest analysis, classification and regression trees (CART), and the K-nearest neighbor algorithm.

[0190] Figure 15A and Figure 15B shows alignment with the region in chromosome 8 (chr8: 108694010-108712904) of the human reference ( Figure 15A ) and the size of the long DNA fragment in plasma carrying 7 maternal-specific alleles ( Figure 15B ) is 18,896 bp. All those maternal-specific alleles are consistent with the allele information inferred from deep sequencing (Illumina technology) of the maternal and fetal genomes ( Figure 15B ). According to the method described in the present disclosure, its methylation level was determined to be 72.6% ( Figure 15B ), showing equivalence to the combined methylation level (72.7%) of the maternal-specific fragment. Therefore, such molecules are more likely to be classified as fragments of maternal origin. The genetic and epigenetic information of cell-free DNA molecules of about a kilobase in length cannot be easily identified by using short-read sequencing such as bisulfite sequencing (Illumina).

[0191] The probability that this molecule is derived from the placenta can be calculated using the method described above. Based on the frequency distribution plot of 19 kb fragments, the frequencies of DNA molecules derived from the placenta and maternal blood cells are 0.65% and 0.23%, respectively. The probability that this DNA fragment is derived from the placenta is 43%, indicating an increased likelihood that it is of maternal origin.

[0192] D. Clinical haplotype analysis application

[0193] In the examples, the ability to analyze both short DNA molecules and long DNA molecules in pregnant women's plasma DNA will allow us to perform relative haplotype dosage (RHDO) analysis (Lo et al., Science Translational Medicine 2010; 2:61ra91; Hui et al., Clinical Chemistry 2017; 63:513-524) without requiring prior paternal or maternal or fetal genotype information obtained from tissue. This ability will be more cost-effective and clinically applicable than previously possible.

[0194] Figure 16Illustrate the principle of how we can use cell-free DNA in pregnant women for RHDO analysis. Cell-free DNA is isolated from a pregnant woman and subjected to SMRT sequencing at stage 1605. The size, allele information, and methylation status of each molecule, including long DNA molecules and short DNA molecules, can be determined according to the methods described in the present disclosure. At stage 1610, based on the size information, we can divide the sequenced molecules into two categories, namely long DNA molecules and short DNA molecules. The cut-off values for determining the long DNA category and the short DNA category can include but are not limited to 150bp, 180bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 1kb, 1.1kb, 1.2kb, 1.3kb, 1.4kb, 1.5kb, 1.6kb, 1.7kb, 1.8kb, 1.9kb, 2kb, 2.5kb, 3kb, 4kb, 5kb, 6kb, 7kb, 8kb, 9kb, 10kb, 15kb, 20kb, 30kb, 40kb, 50kb, 60kb, 70kb, 80kb, 90kb, 100kb, 200kb, 300kb, 400kb, 500kb, or 1Mb. In an embodiment, at stage 1615, the allele information present in the long DNA molecules can be used to construct maternal haplotypes, namely Hap I and Hap II. The short DNA molecules can be aligned with the maternal haplotypes according to the allele information. Therefore, the number of cell-free DNA molecules (such as short DNA) derived from maternal Hap I and Hap II can be determined.

[0195] At stage 1620, the imbalance of the haplotypes can be analyzed. The imbalance can be the molecule count, the molecule size, or the molecule methylation status. At stage 1625, the maternal inheritance of the fetus can be inferred. If the dosage of Hap I in the maternal plasma DNA is overrepresented, the fetus will likely inherit maternal Hap I. Otherwise, the fetus will likely inherit maternal Hap II. Different statistical approaches, including but not limited to the sequential probability ratio test (SPRT), the binomial test, the Chi-squared test, Student's t-test, non-parametric tests (such as the Wilcoxon test), and the hidden Markov model, will be used to determine which maternal haplotype is overrepresented.

[0196] In an embodiment, in addition to count analysis, the methylation and size of short DNA molecules are also determined and assigned to the maternal haplotypes. The methylation imbalance between two haplotypes, i.e., Hap I and Hap II, can be used to determine the maternal haplotype inherited by the fetus. If the fetus has inherited Hap I, there are more fragments carrying the allele of Hap I in maternal plasma compared to the fragments carrying the allele of Hap II. The hypomethylation of DNA fragments derived from the fetus results in a lower methylation level of Hap I than that of Hap II. In other words, if the methylation of Hap I is shown to be lower than that of Hap II, the fetus is more likely to inherit maternal Hap I. Otherwise, the fetus is more likely to inherit maternal Hap II. In another embodiment, the probability that a single fragment is derived from the fetus or the mother can be calculated as described above. For all fragments aligned to HapI, the combined probability that these fragments are derived from the fetus can be determined based on Bayes' theorem. Similarly, the combined probability that these fragments are derived from the fetus can be calculated for Hap II. Subsequently, the likelihood that the fetus inherits Hap I or Hap II can be inferred based on the two combined probabilities.

[0197] In an embodiment, the lengthening or shortening of size between two haplotypes, i.e., Hap I and Hap II, can be used to determine the maternal haplotype inherited by the fetus. If the fetus has inherited Hap I, there are more fragments carrying the allele of Hap I in maternal plasma compared to the fragments carrying the allele of Hap II. The DNA fragments derived from the fetus will be relatively shorter than the DNA fragments derived from Hap II. In other words, if the molecules derived from Hap I contain more short DNA than Hap II, the fetus is more likely to inherit maternal Hap I. Otherwise, the fetus is more likely to inherit maternal Hap II.

[0198] In some embodiments, we can perform a combined analysis of count, size, and methylation between maternal Hap I and Hap II to infer the maternal inheritance of the fetus. For example, we can use logistic regression to combine those three metrics including count, size, and methylation status.

[0199] In clinical practice, the haplotype-based analysis of count, size, and methylation status will allow determination of whether the unborn fetus has inherited a maternal haplotype associated with a genetic disorder, such as but not limited to monogenic disorders including fragile X syndrome, muscular dystrophy, Huntington disease, or β-thalassemia. Disorders related to repetitive sequences in DNA sequences in long free reads are separately described in this disclosure.

[0200] E. Targeted Sequencing of Long Free DNA Molecules

[0201] The method described in the present disclosure can also be applied to analyze one or more selected long DNA fragments. In an embodiment, the one or more long DNA fragments of interest can be first enriched by a hybridization method that allows DNA molecules from one or more regions of interest to hybridize with synthetic oligonucleotides having complementary sequences. In order to decode all size, genetic and epigenetic information into one using the method described in the present disclosure, the target DNA molecule is preferably not amplified by PCR before undergoing sequencing, because the base modification information in the original DNA molecule will not be transferred to the PCR product.

[0202] Several methods have been developed to enrich these target regions without performing PCR amplification. In another embodiment, one or more target long DNA molecules can be enriched via the use of the clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR associated protein 9 (Cas9) system (Stevens et al., PLoS One 2019; 14(4):e0215441; Watson et al., Lab Invest 2020; 100:135-146). Even though the CRISPR-Cas9-mediated cleavage will change the size of the original long DNA molecule, its genetic and epigenetic information is still preserved and can be obtained using the methods described in the present disclosure, including but not limited to base content, haplotype (i.e., phase) information, re-mutation, base modification (e.g., 4mC (N4-methylcytosine), 5hmC (5-hydroxymethylcytosine), 5fC (5-formylcytosine), 5caC (5-carboxylcytosine), 1mA (N1-methyladenine), 3mA (N3-methyladenine), 7mA (N7-methyladenine), 3mC (N3-methylcytosine), 2mG (N2-methylguanine), 6mG (O6-methylguanine), 7mG (N7-methylguanine), 3mT (N3-methylthymine), 4mT (O4-methylthymine), and 8oxoG (8-oxoguanine). In an embodiment, the ends of the DNA molecules in the DNA sample are first dephosphorylated to make them less likely to directly ligate to the sequencing adaptor. Subsequently, the Cas9 protein is guided by the guide RNA (crRNA) to the long DNA molecule of interest to produce a double-strand cleavage. Then, the long DNA molecule of interest flanked by double-strand cleavage on both sides is ligated to the sequencing adaptor specified by the selected sequencing platform. In another embodiment, the DNA can be treated with an exonuclease to degrade the DNA molecules not bound by the Cas9 protein (Stevens et al., PLoS One 2019; 14(4):e0215441). Since these methods do not involve PCR amplification, the original DNA molecules with base modifications can be sequenced, and the base modifications can be determined.

[0203] In embodiments, these methods can be used to design guide RNAs by reference to a reference genome such as the human reference genome (hg19), for example, long interspersed nuclear element (LINE) repeats to target long DNA molecules that share substantial homologous sequences. In one example, such assays can be used to analyze circulating cell-free DNA in maternal plasma to detect fetal aneuploidy (Kinde et al., PLoS ONE 2012;7(7):e41162). In embodiments, deactivated or 'dead' Cas9 (dCas9) and its associated single guide RNA (sgRNA) can be used to enrich targeted long DNA without cleaving double-stranded DNA molecules. For example, the 3' end of the sgRNA can be designed to carry an additional universal short sequence. We can use a biotinylated single-stranded oligonucleotide complementary to the universal short sequence to capture those targeted long DNA molecules bound by dCas9. In another embodiment, we can use a biotinylated dCas9 protein or sgRNA or both to facilitate enrichment.

[0204] In embodiments, we can perform size selection to enrich long DNA fragments using approaches that include, but are not limited to, chemical, physical, enzymatic, gel-based, and bead-based methods or methods that incorporate far more than these approaches, without limitation to one or more specific genomic regions of interest. In other embodiments, immunoprecipitation can be used to enrich DNA fragments having a specific methylation profile, such as mediated by using anti-methylcytosine antibodies and methyl-binding proteins. The methylation profile of the bound or captured DNA can be determined using methylation-insensitive sequencing.

[0205] F. General concepts of fetal genetic analysis based on long plasma DNA molecules

[0206] Figure 17 Depiction of the determination of genetic / epigenetic disorders in plasma DNA molecules with information of maternal and fetal origin. Long plasma DNA molecules in a pregnant woman can be determined to be of fetal or maternal origin based on the genetic and / or epigenetic profile of CpG sites in the whole or part of the molecule [i.e., region (a)]. The genetic information can be, but is not limited to, sequence information, single nucleotide polymorphisms, insertions, deletions, tandem repeats, satellite DNA, microsatellites, minisatellites, inversions, etc. The epigenetic information can be the methylation status of one or more CpG sites and their relative order in the plasma DNA molecule. In another embodiment, the epigenetic information can be a modification of any of A, C, G, or T. Long plasma DNA with tissue origin information can be used for non-invasive prenatal testing by determining the presence of genetic and / or epigenetic disorders in such long plasma DNA molecules [i.e., region (b)].

[0207] Figure 18Illustrate the identification of fetal abnormal segments. As an example, based on the methylation patterns in region (a) of the present disclosure, long DNA segments are identified as being of fetal origin. We can determine the likelihood that a fetus is affected by a genetic or epigenetic disorder based on such fetal-origin molecules. Genetic disorders can involve single nucleotide variants, insertions, deletions, tandem repeats, satellite DNA, microsatellites, minisatellites, inversions, etc. Examples of genetic disorders include, but are not limited to: β-thalassemia, α-thalassemia, sickle cell anemia, cystic fibrosis, X-linked disorders (such as hemophilia, Duchenne muscular dystrophy), spinal muscular atrophy, congenital adrenal hyperplasia, etc. Epigenetic disorders can be, for example, abnormal DNA methylation levels such as increased methylation (i.e., hypermethylation) or loss (hypomethylation). Examples of epigenetic disorders include, but are not limited to: fragile X syndrome, Angelman's syndrome, Prader-Willi syndrome, facioscapulohumeral muscular dystrophy (FSHD), immunodeficiency, instability of centromeres and facial anomalies (ICF) syndrome, etc. Genetic or epigenetic disorders can be found in region (b).

[0208] G. Improve sequencing accuracy

[0209] Sequencing accuracy can be improved using sequence reads of long cell-free DNA fragments. In Figure 11B , among 7 alleles in a long fetal-specific DNA molecule, 1 allele seems to be inconsistent between PacBio and Illumina sequencing.

[0210] Figure 19A - 19G Illustration showing error correction of cell-free DNA genotyping using PacBio sequencing. We visually inspected the Figure 11B subread alignment results for those 7 loci. The first column indicates the genomic coordinates; the second column is the reference sequence. The third column and subsequent columns indicate the aligned subreads. For example, in Figure 19A , there are 8 subreads passing through the region. '.' indicates the same as the reference base in the Watson strand. ',' indicates the same as the reference base in the Crick strand. 'Letter' indicates an alternative allele. '*' indicates an insertion and / or deletion. We can see that Figure 19F the major base at the inconsistent locus shown in Figure 19F is called 'T' in the consensus sequence. However, among 9 subreads at the locus ( Figure 19FMinor allele fraction lower than other loci ( Figure 19A -E and Figure 19G )(MAF range: 67% - 89%). Thus, if we use, for example, a MAF setting of at least 60% to determine strict criteria for the base composition of each locus in the consensus sequence, this error locus will be excluded from downstream interpretation. On the other hand, such error loci happen to fall within homopolymers (i.e., a series of consecutive identical bases 'TTTTTTT'). In an embodiment, we can set criteria such that variants within homopolymers are flagged as QC failures and not used for downstream analysis for the time being. In an embodiment, we can apply different locus quality and base quality to correct or filter low-quality bases or subreads to improve base composition analysis.

[0211] In the case where the sequencing accuracy of nanopore sequencing is further improved, embodiments of the present invention can also be used with such improved sequencing platforms and thus produce improved accuracy.

[0212] H. Example methods

[0213] Long cell-free DNA fragments from a biological sample having cell-free DNA fragments obtained from a pregnant woman can be sequenced. These long cell-free DNA fragments can be used to determine the haplotypic inheritance of the fetus.

[0214] 1. Sequencing long cell-free DNA fragments

[0215] Figure 20 A method 2000 for analyzing a biological sample of a pregnant organism is shown. The biological sample can contain a plurality of cell-free nucleic acid molecules. The biological sample can be any biological sample described herein. More than 20% of the cell-free nucleic acid molecules in the biological sample have a size greater than 200 nt (nucleotides).

[0216] At block 2010, a plurality of cell-free nucleic acid molecules are sequenced. The sequencing can be performed by single molecule real-time technology. In some embodiments, the sequencing can be performed by using a nanopore.

[0217] More than 20% of the plurality of sequenced cell-free nucleic acid molecules can have a length greater than 200 nt. In some embodiments, 15% - 20%, 20% - 25%, 25% - 30%, 30% - 35% or more than 35% of the plurality of sequenced cell-free nucleic acid molecules can have a length greater than 200 nt.

[0218] In some embodiments, more than 11% of the plurality of sequenced cell-free nucleic acid molecules can have a length greater than 400 nt. In an embodiment, 5% - 10%, 10% - 15%, 15% - 20%, 20% - 25% or more than 25% of the plurality of sequenced cell-free nucleic acid molecules can have a length greater than 400 nt.

[0219] In some embodiments, more than 10% of the plurality of sequenced free nucleic acid molecules may have a length greater than 500 nt. In embodiments, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the plurality of sequenced free nucleic acid molecules may have a length greater than 500 nt.

[0220] In embodiments, more than 8% of the plurality of sequenced free nucleic acid molecules may have a length greater than 600 nt. In embodiments, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the plurality of sequenced free nucleic acid molecules may have a length greater than 600 nt.

[0221] In some embodiments, more than 6% of the plurality of sequenced free nucleic acid molecules may have a length greater than 1 knt. In embodiments, 3%-5%, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the plurality of sequenced free nucleic acid molecules may have a length greater than 1 knt.

[0222] In embodiments, more than 3% of the plurality of sequenced free nucleic acid molecules may have a length greater than 2 knt. In embodiments, 1%-5%, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the plurality of sequenced free nucleic acid molecules may have a length greater than 2 knt.

[0223] In embodiments, more than 1% of the plurality of sequenced free nucleic acid molecules may have a length greater than 3 knt. In embodiments, 1%-5%, 5%-10%, 10%-15%, 15%-20%, 20%-25%, or more than 25% of the plurality of sequenced free nucleic acid molecules may have a length greater than 3 knt.

[0224] In some embodiments, at least 0.9% of the plurality of sequenced free nucleic acid molecules may have a length greater than 4 knt. In embodiments, 0.5%-1%, 1%-5%, 5%-10%, 10%-15%, 15%-20%, or more than 20% of the plurality of sequenced free nucleic acid molecules may have a length greater than 4 knt.

[0225] In some embodiments, at least 0.04% of the plurality of sequenced free nucleic acid molecules may have a length greater than 10 knt. In embodiments, 0.01% to 0.1%, 0.1% to 0.5%, 0.5%-1%, 1%-5%, 5%-10%, 10%-15%, or more than 15% of the plurality of sequenced free nucleic acid molecules may have a length greater than 4 knt.

[0226] The plurality of free nucleic acid molecules may comprise at least 10, 50, 100, 150, or 200 free nucleic acid molecules. The plurality of free nucleic acid molecules may be from multiple different genomic regions. For example, multiple chromosome arms or chromosomes may be covered by the free nucleic acid molecules. At least two of the plurality of free nucleic acid molecules may correspond to non-overlapping regions.

[0227] Methods for sequencing long free DNA fragments can be used by any of the methods described herein. Reads from the sequencing can be used to determine fetal aneuploidy, aberrations (e.g., copy number aberrations), genetic mutations or variations, or parental haplotype inheritance. The amount of sequence reads can represent the amount of free DNA fragments.

[0228] 2. Haplotype inheritance

[0229] Figure 21 Method 2100 for displaying an analysis of a biological sample obtained from a female carrying a fetus is shown. The female may have a first haplotype and a second haplotype in a first chromosomal region. The biological sample may comprise a plurality of free DNA molecules from the fetus and the female. The biological sample may be any biological sample described herein.

[0230] At block 2105, reads corresponding to the plurality of free DNA molecules may be received. The reads may be sequence reads. In some embodiments, the method may comprise performing sequencing.

[0231] At block 2110, the sizes of the plurality of free DNA molecules may be measured. The size may be measured by aligning one or more sequence reads corresponding to the ends of the DNA molecules with a reference genome. The size may be measured by performing full-length sequencing of the DNA molecules and then counting the number of nucleotides in the full-length sequence. The genomic coordinates at the outermost nucleotides may be used to determine the length of the DNA molecule.

[0232] At block 2115, a first set of free DNA molecules from the plurality of free DNA molecules may be identified as having a size greater than or equal to a cut-off value. The cut-off value may be any cut-off value associated with long DNA. For example, the cut-off value may include 150 bp, 180 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 550 bp, 600 bp, 650 bp, 700 bp, 750 bp, 800 bp, 850 bp, 900 bp, 950 bp, 1 kb, 1.5 kb, 2 kb, 2.5 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 15 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 200 kb, 300 kb, 400 kb, 500 kb, or 1 Mb.

[0233] At block 2120, the sequences of the first haplotype and the second haplotype can be determined from reads corresponding to the first set of cell-free DNA molecules. Determining the sequences of the first haplotype and the second haplotype can include aligning the reads corresponding to the first set of cell-free DNA molecules to a reference genome.

[0234] In some embodiments, determining the sequences of the first haplotype and the second haplotype may not include a reference genome. Determining the sequences can include aligning a first subgroup of reads to a second subgroup of reads to identify different alleles at loci in the reads. The method can include determining that the first subgroup of reads has a first allele at a locus. The method can also include determining that the second subgroup of reads has a second allele at a locus. The method can further include determining that the first subgroup of reads corresponds to the first haplotype. Additionally, the method can include determining that the second subgroup of reads corresponds to the second haplotype. The alignment can be similar to the alignment described by Figure 16 described.

[0235] At block 2125, a second set of cell-free DNA molecules from a plurality of cell-free DNA molecules can be aligned to the sequence of the first haplotype. The second set of cell-free DNA molecules can have a size less than a cutoff value. The second set of cell-free DNA molecules can be short DNA molecules of the first haplotype.

[0236] At block 2130, a third set of cell-free DNA molecules from a plurality of cell-free DNA molecules can be aligned to the sequence of the second haplotype. The third set of cell-free DNA molecules can have a size less than a cutoff value. The third set of cell-free DNA molecules can be short DNA molecules of the second haplotype.

[0237] At block 2135, a first value of a parameter can be measured using the second set of cell-free DNA molecules. The parameter can be a count of cell-free DNA molecules, a size profile of cell-free DNA molecules, or a methylation level of cell-free DNA molecules. The value can be an original value or a statistical value (e.g., mean, median, mode, percentile, minimum, maximum). In some embodiments, the value can be normalized to a value of the parameter of a reference sample, another region, both haplotypes, or other size ranges.

[0238] At block 2140, a second value of the parameter can be measured using the third set of cell-free DNA molecules. The parameter is the same parameter as that of the second set of cell-free DNA molecules.

[0239] At block 2145, a first value may be compared with a second value. The comparison may use a separation value. The separation value may be calculated using the first value and the second value. The separation value may be compared with a cut-off value. The separation value may be any separation value described herein. The cut-off value may be determined from a reference sample from a pregnant female carrying an euploid fetus. In other embodiments, the cut-off value may be determined from a reference sample from a pregnant female carrying an aneuploid fetus. In some embodiments, assuming an aneuploid fetus, the cut-off value may be determined. For example, data from a reference sample from a pregnant female carrying an euploid fetus may be adjusted to account for an increased or decreased copy number of chromosomal regions of the aneuploid. The cut-off value may be determined from the adjusted data.

[0240] At 2150, the likelihood of the fetus inheriting a first genetic haplotype may be determined based on the comparison of the first value and the second value. The likelihood may be determined based on the comparison of the separation value and the cut-off value. When the parameter is the size profile of the cell-free DNA molecules, the method may include determining that the fetus has a higher likelihood of inheriting the first genetic haplotype than the second genetic haplotype when the first value is less than the second value, indicating that a second set of cell-free DNA molecules is characterized by a smaller size profile than a third set of cell-free DNA molecules. When the parameter is the degree of methylation of the cell-free DNA molecules, the method may include determining that the fetus has a higher likelihood of inheriting the first genetic haplotype than the second genetic haplotype when the first value is less than the second value.

[0241] In some embodiments, the method may include identifying the number of repeat sequences of a subsequence in a read corresponding to the first set of cell-free DNA molecules. Determining the sequence of the first haplotype may include determining that the sequenced sequence includes the number of repeat sequences of the subsequence. The first haplotype may include a repeat sequence-related disease, and the repeat sequence-related disease may be any repeat sequence-related disease described herein. The likelihood that the fetus inherits the repeat sequence-related disease may be determined. The likelihood that the fetus inherits the repeat sequence-related disease may be comparable or similar to the likelihood that the fetus inherits the first genetic haplotype. The identification of the repeat sequences of the sequence is described later in this disclosure, including using Figure 16 described by.

[0242] II. Using Methylation Analysis to Determine the Tissue of Origin

[0243] The long cell-free DNA molecules may have several methylation sites. As discussed in this disclosure, the degree of methylation of the long cell-free DNA molecules in a pregnant female may be used to determine the tissue of origin. Additionally, the methylation pattern present on the long cell-free DNA molecules may be used to determine the tissue of origin.

[0244] Compared to white blood cells and cells from tissues such as, but not limited to, liver, lung, esophagus, heart, pancreas, colon, small intestine, adipose tissue, adrenal gland, brain, etc., cells from placental tissue have unique methylomic patterns (Sun et al., Proceedings of the National Academy of Sciences of the United States of America 2015; 112: E5503-12). The methylation profile of circulating fetal DNA in the blood of pregnant mothers may be similar to that of circulating fetal DNA in the placenta, thus providing the possibility to explore means for generating non-invasive fetal-specific biomarkers regardless of fetal sex or genotype. However, bisulfite sequencing of maternal plasma DNA in pregnant women (e.g., using the Illumina sequencing platform) may lack the ability to distinguish molecules of fetal origin from those of maternal origin due to several limitations: (1) Plasma DNA may be degraded during bisulfite treatment, and long DNA molecules will generally break into shorter molecules; (2) DNA molecules larger than 500 bp may not be effectively sequenced by the Illumina sequencing platform for downstream analysis (Tan et al., Scientific Reports 2019; 9: 2856).

[0245] For methylation-based analysis of the tissue of origin, we can focus on several differentially methylated regions (DMRs) and use the collective methylation signals from multiple molecules associated with the DMRs (Sun et al., Proceedings of the National Academy of Sciences of the United States of America 2015; 112: E5503-12) instead of single-molecule methylation patterns. Multiple studies have attempted to use methylation-sensitive restriction enzyme-based approaches (Chan et al., Clinical Chemistry 2006; 52: 2211-8) or methylation-specific PCR-based approaches (Lo et al., American Journal of Human Genetics 1998; 62: 768-75) to assess the contribution of the placenta to the plasma DNA pool. However, those studies are only suitable for analyzing one or a few markers and may be challenging for analyzing molecules on a genome-wide scale. However, those reads are inferred from amplified signals (i.e., PCR-based amplification during DNA library preparation in the flow cell and bridge amplification during sequencing cluster generation). The amplification step may potentially cause a preference for short DNA molecules, resulting in loss of information associated with long DNA molecules. In addition, Li et al. only analyzed those reads associated with previously developed DMRs (Li et al., Nucleic Acids Research 2018; 46: e89).

[0246] In the present disclosure, we describe novel approaches for differentiating fetal DNA molecules from maternal DNA molecules in maternal plasma without bisulfite treatment and DNA amplification based on the methylation patterns of individual DNA molecules. In embodiments, one or more long plasma DNA molecules will be used for analysis (e.g., using bioinformatics and / or experimental analysis for size selection). Long DNA molecules can be defined as DNA molecules having a size of at least but not limited to 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 100 kb, 200 kb, etc. There is little data on the presence and methylation status of longer free DNA molecules in maternal plasma. For example, it is not known whether the methylation status of the longer free DNA molecules will reflect the methylation status of the cellular DNA of the originating tissue, e.g., since long fragments have more sites where the methylation status may change after fragmentation occurs in the body; such changes may occur when the fragments circulate in plasma. For example, studies have shown that the methylation status of circulating DNA is related to the size of the DNA fragments (Lun et al., Clinical Chemistry 2013; 59:1583-94). Thus, the feasibility of inferring the originating tissue from the longer free DNA molecules is unknown. Therefore, the approaches taken to identify tissue-related methylation signatures and the methods taken to determine and interpret the presence of the tissue-specific longer free DNA molecules are substantially different from the approaches and methods applied to short free DNA analysis.

[0247] According to embodiments of the present disclosure, we can identify short DNA molecules and long DNA molecules and determine their biological characteristics including but not limited to methylation patterns, fragment ends, sizes, and base compositions. Short DNA molecules can be defined as DNA molecules having a size less than but not limited to 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 200 bp, 300 bp, etc. Short DNA molecules can be DNA molecules not within the range considered long. We describe a new approach for inferring the origin tissue of circulating DNA molecules in maternal plasma. This new approach utilizes methylation patterns on one or more long DNA molecules in plasma. The longer the DNA molecule, the more CpG sites it may contain. The presence of multiple CpG sites on a plasma DNA molecule will provide origin tissue information even if the methylation status of any single CpG site may not provide information for determining the origin tissue. The methylation pattern in a long DNA molecule can include the methylation status of each CpG site, the order of the methylation statuses, and the distance between any two CpG sites. The methylation status between two CpG sites can depend on the distance between the two CpG sites. When CpG sites (such as CpG islands) within a specific distance in a molecule exhibit tissue-specific patterns, a statistical model can assign more weight to those signals during origin tissue analysis.

[0248] Figure 22 This principle is schematically illustrated. Figure 22 The methylation pattern of a DNA molecule is shown. Seven CpG sites and six plasma DNA fragments A - E for different tissues (placenta, liver, blood cells, colon) are shown. Methylated CpG sites are shown in red, and unmethylated CpG sites are shown in green. As an example, we consider 7 CpG sites with various methylation statuses across placenta, liver, blood cells, and colon tissues. We consider the following scenario: A single CpG site does not exhibit a methylation status specific to the placenta compared to other tissues. Thus, the origin tissue of those plasma DNA molecules A, B, C, D, and E with variable sizes can be determined not only based on the methylation status at a single CpG site. For plasma DNA molecules A and B, since the two molecules are relatively short in size, they contain only 3 and 4 CpG sites, respectively. In an embodiment, the methylation pattern in a DNA molecule containing more than one CpG site can be defined as a methylation haplotype. As Figure 22As shown, plasma DNA molecules A and B can be contributed by the placenta or the liver based on their methylation haplotypes, because the placenta and the liver share the same methylation haplotypes at those genomic positions corresponding to molecule A (positions 1, 2, and 3) and those genomic positions corresponding to molecule B (positions 1, 2, 3, and 4). However, when we can obtain long DNA molecules such as molecules C, D, and E in plasma, it is possible to unambiguously determine that those molecules C, D, and E are derived from the placenta based on the methylation haplotypes.

[0249] The reference pattern of a tissue can be based on the methylation pattern of a reference tissue. In some embodiments, the methylation pattern can be based on several reads and / or samples. The degree of methylation of each CpG site (also referred to as the methylation index MI and described below) can be used to determine whether the site is methylated.

[0250] A. Statistical Model for Methylation Patterns

[0251] In an embodiment, the likelihood that a plasma DNA molecule is derived from the placenta can be determined by comparing the methylation haplotype of a single DNA molecule with the methylation patterns in multiple reference tissues. Long plasma DNA molecules can be advantageous for such analysis. Long DNA molecules can be defined as DNA molecules having a size of at least but not limited to 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 100 kb, 200 kb, etc. Reference tissues can include but are not limited to the placenta, liver, lung, esophagus, heart, pancreas, colon, small intestine, adipose tissue, adrenal gland, brain, neutrophil, lymphocyte, basophil, eosinophil, etc. In an embodiment, we can determine the likelihood that a plasma DNA molecule is derived from the placenta by co-analyzing the methylation haplotype of plasma DNA determined by single molecule real-time sequencing and the methylome data based on whole genome bisulfite sequencing of reference tissues. As an example, the placenta and buffy coat samples were sequenced to an average of 94-fold and 75-fold genome coverage of the haploid genome respectively using whole genome bisulfite sequencing. The degree of methylation (also referred to as the methylation index MI) of each CpG site was calculated using the following formula based on the number of sequenced cytosines (i.e., methylated, represented by C) and the number of sequenced thymines (i.e., unmethylated, represented by T):

[0252]

[0253] CpG sites were classified into three categories based on the MI values inferred from placental DNA:

[0254] 1. CpG sites with a category A MI value ≥ 70.

[0255] 2. CpG sites with a class B MI value between 30 and 70.

[0256] 3. CpG sites with a class C MI value of ≤ 30.

[0257] Similarly, CpG sites are classified into three classes using the MI values at the CpG sites inferred from buffy coat DNA:

[0258] 1. CpG sites with a class A MI value ≥ 70.

[0259] 2. CpG sites with a class B MI value between 30 and 70.

[0260] 3. CpG sites with a class C MI value of ≤ 30.

[0261] The classes use MI cut-off values of 30 and 70. The cut-off values can include other numerical values, including 10, 20, 40, 50, 60, 80, or 90. In some embodiments, these classes can be used to determine a reference methylation pattern for a reference tissue (e.g., as described for use). Class A sites can be considered methylated. Class C sites can be considered unmethylated. Class B sites can be considered non-informative and are not included in the reference pattern. Figure 22 For a plasma DNA molecule with n CpG sites, the methylation status of each CpG site is determined by the approach described in our previous publication (U.S. Application No. 16 / 995,607). In some embodiments, the methylation status can be determined by bisulfite sequencing or by nanopore sequencing. To determine the likelihood that a plasma DNA molecule is derived from a placental or maternal background, the methylation pattern of the molecule is analyzed in combination with the methylation information in the previous placental and maternal buffy coat DNA. In an embodiment, we utilize the following principle: If the CpG sites determined to be methylated (M) in a plasma DNA fragment are consistent with a higher methylation index in the placenta, then such an observation will indicate that this molecule is more likely to be derived from the placenta. If the CpG sites determined to be methylated (M) in a plasma DNA molecule are consistent with a lower methylation index in the placenta, then such an observation will indicate that this molecule is less likely to be derived from the placenta; if the CpG sites determined to be unmethylated (U) in the plasma DNA are consistent with a lower methylation index in the placenta. Then such an observation will indicate that this molecule is more likely to be derived from the placenta. If the CpG sites determined to be unmethylated (U) in the plasma DNA are consistent with a higher methylation index in the placenta, then such an observation will indicate that this molecule is less likely to be derived from the placenta.

[0262]

[0263] ​We implement the following scoring process. An initial score (S) reflecting the likelihood of fetal origin of plasma DNA fragments is set to 0. When comparing the methylation status of plasma DNA molecules with the methylation information of the previous placental DNA,

[0264] a. If a CpG site on a plasma DNA molecule is determined to be 'M' and its counterpart in the placenta belongs to category A, add 1 point to S (i.e., increase the score unit by 1).

[0265] b. If a CpG site on a plasma DNA molecule is determined to be 'U' and its counterpart in the placenta belongs to category A, subtract 1 point from S (i.e., decrease the score unit by 1).

[0266] c. If a CpG site on a plasma DNA molecule is determined to be 'M' and its counterpart in the placenta belongs to category B, add 0.5 points to S.

[0267] d. If a CpG site on a plasma DNA molecule is determined to be 'U' and its counterpart in the placenta belongs to category B, add 0.5 points to S.

[0268] e. If a CpG site on a plasma DNA molecule is determined to be 'M' and its counterpart in the placenta belongs to category C, subtract 1 point from S.

[0269] f. If a CpG site on a plasma DNA molecule is determined to be 'U' and its counterpart in the placenta belongs to category C, add 1 point to S.

[0270] We refer to the above process as'methylation status matching'.

[0271] After all CpG sites in the plasma DNA molecule have been processed, the final aggregate score S(placenta) of the plasma DNA molecule is obtained. In the examples, the number of CpG sites is required to be at least 30 and the length of the plasma DNA molecule is required to be at least 3 kb. Other numbers and lengths of CpG sites can be used, including but not limited to any numbers and lengths described herein.

[0272] When comparing the methylation status of plasma DNA molecules with the methylation degree of buffy coat DNA at corresponding sites, a similar scoring process will be applied. After all CpG sites in the plasma DNA molecule have been processed, the final aggregate score S(buffy coat) of the plasma DNA molecule is obtained.

[0273] If S(placenta) > S(buffy coat), it is determined that the plasma DNA molecule is of fetal origin; otherwise, it is determined that the plasma DNA molecule is of maternal origin.

[0274] There are 17 fetal-specific DNA molecules and 405 maternal-specific DNA molecules for evaluating the efficacy of inferring the fetal-maternal origin of plasma DNA molecules. The fetal-specific molecules are plasma DNA molecules carrying fetal-specific SNP alleles, while the maternal-specific DNA molecules are plasma DNA molecules carrying maternal-specific SNP alleles.

[0275] Figure 23 Receiver operating characteristic curves (ROCs) are shown for determining fetal origin and maternal origin. The y-axis shows sensitivity, and the x-axis shows specificity. The red line represents the efficacy of molecules that distinguish fetal origin from maternal origin using the methylation-state matching method present in this disclosure. The blue line represents the efficacy of molecules that distinguish fetal origin from maternal origin using the single-molecule methylation level (i.e., the proportion of CpG sites determined to be methylated in a DNA molecule). Figure 23 The area under the receiver operating characteristic curve (AUC) for the methylation-state matching process (0.94) is shown to be significantly higher than the AUC based on the single-molecule methylation level (0.86) (P value < 0.0001; DeLong test). This indicates that the analysis of the methylation pattern of long DNA molecules can be used to determine fetal / maternal origin.

[0276] In an embodiment, when determining whether plasma DNA is of fetal origin or maternal origin, the difference (ΔS) between S (placenta) and S (buffy coat) can be considered. The absolute value of ΔS can be required to exceed a specific threshold such as, but not limited to, 5, 10, 20, 30, 40, 50, etc. As an illustration, when we use 10 as the threshold for ΔS, the positive predictive value (PPV) in the detection of fetal DNA molecules increases from 14.95% to 91.67%.

[0277] In an embodiment, the methylation status of a CpG site will be affected by the methylation status of its neighboring CpG sites. The closer the nucleotide distance between any two CpG sites on a DNA molecule, the more likely the two CpG sites will share the same methylation status. This phenomenon has been termed co-methylation. Multiple tissue-specific CpG island methylations have been reported; thus, in some statistical models for origin tissue analysis, more weight will be assigned to CpG sites (such as CpG islands) in dense clusters that share the same methylation status. For scenarios 'a' and 'f', if the currently explored CpG site is located within a genomic distance of no more than 100 bp relative to the previous CpG site and the results of the methylation status matching process for these two consecutive CpG sites are the same, 1 point will be additionally added to the score S of the currently explored CpG site. For scenarios 'b' and 'e', if the currently explored CpG site is located within a genomic distance of no more than 100 bp relative to the previous CpG site and the results of the methylation status matching process for these two consecutive CpG sites are the same, 1 point will be additionally deducted from the score S of the currently explored CpG site. However, if the currently explored CpG site is located within a genomic distance of no more than 100 bp relative to the previous CpG site, but the results of the methylation status matching process for these two consecutive CpG sites are inconsistent, the aforementioned preset scoring process will be used. On the other hand, if the currently explored CpG site is located within a genomic distance greater than 100 bp relative to the previous CpG site, the scoring process with preset parameters described above will be used. Scores other than 1 and distances other than 100 bp can be used, including any scores and distances described herein.

[0278] In other embodiments, CpG sites are classified into more than three categories based on the MI value inferred from placental and buffy coat DNA. The methylation information of the previous reference tissue can be inferred from single molecule real-time sequencing (i.e., nanopore sequencing and / or PacBio SMRT sequencing). The length of the plasma DNA molecule can be required to be at least but not limited to 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 100 kb, 200 kb, etc. The number of CpG sites can be required to be at least but not limited to 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, etc.

[0279] In an embodiment, we can use a probability model to characterize the methylation pattern of plasma DNA molecules. The methylation status of k CpG sites (k≥1) on a plasma DNA molecule is denoted as M=(m 1 、m 2, …, m k ), where m at the CpG site i on the plasma DNA molecule i is 0 (for the unmethylated state) or 1 (for the methylated state). In the examples, the probability M associated with the plasma DNA molecules derived from the placenta can depend on the reference methylation pattern in the placental tissue. The reference methylation pattern in the placental tissue for those corresponding 1, 2, …, k CpG sites will follow a beta distribution. The beta distribution is parameterized by two positive parameters α and β denoted as beta(α, β). The range of values derived from the beta distribution will be from 0 to 1. Based on the deep bisulfite sequencing data of the tissue of interest, the parameters α and β are determined respectively by the number of sequenced cytosines (methylated) and thymines (unmethylated) at each CpG site of the specific tissue. For the placenta, such a beta distribution is denoted as beta(α P , β p ). The probability P(M|placenta) of the plasma DNA molecules derived from the placenta will be modeled as follows:

[0280]

[0281] where ‘i’ represents the i-th CpG site; beta denotes the beta(β) distribution associated with the methylation pattern at the i-th CpG site in the placenta; P is the joint probability of the observed plasma DNA molecules having a given methylation pattern among the k CpG sites.

[0282] The probability P(M|buffy coat) of the plasma DNA molecules derived from the buffy coat (i.e., white blood cells) will be modeled as follows:

[0283]

[0284] where ‘i’ represents the i-th CpG site; beta denotes the beta(β) distribution associated with the methylation pattern at the i-th CpG site in the buffy coat DNA. P is the joint probability of the observed plasma DNA molecules having a given methylation pattern among the k CpG sites.

[0285] Beta and beta can be determined respectively from the whole-genome bisulfite sequencing results of the placenta and buffy coat DNA.

[0286] For plasma DNA molecules, if we observe that P(M|placenta) > P(M|buffy coat), then such plasma DNA molecules are likely to be derived from the placenta; otherwise, they are likely to be derived from the buffy coat. We achieved an AUC of 0.79 using this model.

[0287] B. Machine learning model

[0288] In yet other embodiments, we can use machine learning algorithms for determining the fetal / maternal origin of specific plasma DNA molecules. To test the feasibility of using a machine learning-based approach to classify fetal DNA molecules and maternal DNA molecules in pregnant women, we developed a methylation pattern mapping method for plasma DNA molecules.

[0289] Figure 24 Shows the definition of paired methylation patterns. Nine CpG sites are shown on the plasma DNA molecule. Methylated CpG sites are shown in red, and unmethylated CpG sites are shown in green. When the two CpG sites in a pair share the same methylation status (e.g., the 1st CpG and the 5th CpG), the pair will be encoded as 1, as shown at position 'a' indicated by the arrow. When the two CpG sites in a pair have different methylation statuses (e.g., the 1st CpG and the 2nd CpG), the pair will be encoded as 0, as shown at position 'b' indicated by the arrow. The encoding rule applies equally to all pairs of any 2 CpG sites on the DNA molecule.

[0290] We use a plasma DNA molecule containing 9 CpG sites as an example. The methylation pattern of this plasma DNA molecule is determined by the method described in our previous publication (U.S. Application No. 16 / 995,607), i.e., U-M-M-M-U-U-U-M-M (U and M represent unmethylated CpG and methylated CpG, respectively). Pairwise comparisons of the methylation status between any two CpG sites are applicable to machine learning or deep learning-based analysis. In this example, the rule applies equally to a total of 36 pairs. If there are a total of n CpG sites on the plasma DNA molecule, then there will be n×(n - 1) / 2 comparison pairs. Different numbers can be used, including 5, 6, 7, 8, 10, 11, 12, 13 CpG sites, etc. If the molecule contains more sites than the number used in the machine learning model, a sliding window can be used to divide the sites into an appropriate number of sites.

[0291] We obtain one or more molecules from DNA samples from the placenta and buffy coat, respectively. The methylation patterns of those DNA molecules are determined by Pacific Bioscience (PacBio) single molecule real-time (SMRT) sequencing according to the procedures described in our prior publication (U.S. application Ser. No. 16 / 995,607). Those methylation patterns are translated into paired methylation patterns.

[0292] The paired methylation patterns associated with placental DNA and the paired methylation patterns associated with buffy coat DNA are used to train a convolutional neural network (CNN) to distinguish molecules of potential fetal origin from maternal origin. Each target output (i.e., analogous to a dependent variable value) of a DNA fragment from the placenta is assigned a '1', while each target output of a DNA fragment from the buffy coat is assigned a '0'. The paired methylation patterns are used for training to determine the parameters (often called weights) for the CNN model. The optimal parameters for the CNN for distinguishing the fetal-maternal origin of DNA fragments are obtained when the total prediction error between the output scores calculated by a sigmoid function and the desired target outputs (binary values: 0 or 1) reaches a minimum by repeatedly adjusting the model parameters. The total prediction error is measured by the S cross-entropy loss function in deep learning algorithms (https: / / keras.io / ). The model parameters learned from the training dataset are used to analyze DNA molecules (such as plasma DNA molecules) to output a probabilistic score indicating the likelihood that the DNA molecule is from the placenta or buffy coat. If the probabilistic score of a plasma DNA fragment exceeds a specific threshold, such plasma DNA molecules are considered of fetal origin. Otherwise, it will be considered of maternal origin. The threshold will include but not be limited to 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.99, etc. In one example, we used this CNN model to achieve an AUC of 0.63 to determine whether plasma DNA molecules are of fetal origin or maternal origin, indicating that it is possible to use deep learning algorithms to infer the origin tissue of DNA molecules from maternal plasma. The performance of the deep learning algorithm is further improved by obtaining more single molecule real-time sequencing results.

[0293] In some other embodiments, the statistical model may include, but is not limited to, linear regression, logistic regression, deep recurrent neural networks (such as long short-term memory LSTM), Bayes's classifier, hidden Markov model (HMM), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering of applications with noise (DBSCAN), random forest algorithm, and support vector machine (SVM), etc. Different statistical distributions will be involved, including but not limited to binomial distribution, Bernoulli distribution, gamma distribution, normal distribution, Poisson distribution, etc.

[0294] C. Methylation haplotypes specific to the placenta

[0295] The methylation status of each CpG site on a single DNA molecule can be determined using the approach described in our previous publication (U.S. Application No. 16 / 995,607) or any of the techniques described herein. In addition to the degree of single-molecule double-stranded DNA methylation, we can determine the single-molecule methylation pattern of each DNA molecule, which can be the sequence of the methylation status of adjacent CpG sites along a single DNA molecule.

[0296] Different DNA methylation signatures can be found in different tissues and cell types. In the examples, we can infer the origin tissue of individual plasma DNA molecules based on their single-molecule methylation patterns.

[0297] Genomic DNA from ten buffy coat samples and six placental tissue samples was sequenced using SMRT sequencing (PacBio). We were able to achieve 58.7-fold and 28.7-fold coverage of buffy coat DNA and placental DNA, respectively, by aggregating the mapped high-quality circular consensus sequencing (CCS) reads from each sample type.

[0298] The genome is partitioned into approximately 28.2 million overlapping windows of 5 CpG sites using a sliding window approach. In other embodiments, different window sizes such as, but not limited to, 2, 3, 4, 5, 6, 7, and 8 CpG sites may be used. We may also use a non-overlapping window approach. Each window is considered a potential marker region. For each potential marker region, we identify the major single-molecule methylation pattern among all the sequenced placental DNA molecules that cover all 5 CpG sites within the marker region. A comparison is made between the CpG sites in the plasma DNA molecules and the corresponding CpG sites in the individual DNA molecules of the reference tissue. Subsequently, we calculate the mismatch score for each buffy coat DNA molecule that covers all CpG sites within the same marker region, which is done by comparing its single-molecule methylation pattern to the major single-molecule methylation pattern in the placenta.

[0299]

[0300] Where the number of mismatched CpG sites refers to the number of CpG sites that show a methylation status in the buffy coat DNA molecules that is different compared to the major single-molecule methylation pattern in the placenta.

[0301] A higher mismatch score indicates that the methylation pattern of the buffy coat DNA molecule is more different from the major single-molecule methylation pattern in the placenta. We use the following criteria to select marker regions that show substantial single-molecule methylation pattern differences between the DNA molecule pools from the placenta and the buffy coat: a) more than 50% of the placental DNA molecules have the major single-molecule methylation pattern; and b) more than 80% of the buffy coat DNA molecules have a mismatch score greater than 0.3. We select 281,566 marker regions based on these criteria for downstream analysis.

[0302] Figure 25 Table of the distribution of the selected marker regions among different chromosomes. The first column shows the chromosome number. The second column shows the number of marker regions in the chromosome.

[0303] We hereby describe our concept of classifying the origin tissues of individual plasma DNA molecules based on single molecule methylation patterns using plasma DNA molecules sequenced by SMRT sequencing that cover fetal-specific alleles or maternal-specific alleles as previously described in this disclosure. Any plasma DNA molecule that covers a selected marker region and has the same methylation pattern as the major single molecule methylation pattern in the placenta will be classified as a placenta-specific (i.e., fetal-specific) DNA molecule. Conversely, if the single molecule methylation pattern of a plasma DNA molecule is different from the major single molecule methylation pattern in the placenta, we classify this molecule as non-placenta-specific. The correct classification in this analysis is defined by whether a placenta-specific methylation haplotype is present in the molecule in such a way that fetal-specific DNA molecules are identified as being of fetal origin (i.e., specific to the placenta) and maternal DNA molecules are identified as being non-fetal origin (i.e., not specific to the placenta). Previous methylation-based methods for origin tissue analysis typically involved deconvoluting the percentage or contribution ratio of a range of tissue contributors of free DNA within a biological sample. The advantage of the method of the present invention over previous methods is that evidence of tissue contribution of free DNA (e.g., placenta-derived DNA in maternal plasma) can be determined without considering the presence or absence of contributions from other tissues. In addition, the placental origin of any free DNA molecule can be determined by the method of the present invention without considering the fractional contribution of the tissue to the free DNA molecule.

[0304] Among 28 DNA molecules that cover fetal-specific alleles, 17 (61%) were classified as placenta-specific, and 11 (39%) were classified as non-placenta-specific. On the other hand, among 467 DNA molecules that cover maternal-specific alleles, 433 (93%) were classified as non-placenta-specific, and 34 (7%) were classified as placenta-specific.

[0305] In an embodiment, we can use different percentages of buffy coat DNA molecules with a mismatch score greater than 0.3 as thresholds, including but not limited to greater than 60%, 70%, 75%, 80%, 85%, and 90%, etc. We can improve the overall classification accuracy of the placental or non-placental origin of plasma DNA in pregnant individuals by adjusting the criteria used in marker region selection. This is particularly crucial in a non-invasive prenatal testing environment when we attempt to determine whether a pathogenic mutation or copy number aberration is present in the fetus.

[0306] Figure 26A classification table of plasma DNA molecules based on the use of different percentages of buffy coat DNA molecules with a mismatch score greater than 0.3 as a selection criterion for marker regions for single molecule methylation patterns of plasma DNA molecules. The first column shows the percentage of buffy coat DNA molecules with a mismatch score greater than 0.3%. The second column divides the DNA molecules into DNA molecules covering fetal-specific alleles and DNA molecules covering maternal-specific alleles. The third and fourth columns show the classification of DNA molecules as placenta-specific or non-placenta-specific based on single molecule methylation patterns. The fifth column shows the percentage of DNA molecules classified identically to the specific alleles in the second column.

[0307] Figure 27 Shows a method flow for determining fetal inheritance using placenta-specific methylation haplotypes in a non-invasive manner. As Figure 27 shown, cell-free DNA from pregnant women's plasma is extracted for single molecule real-time sequencing. Long plasma DNA molecules are identified according to the examples in the present disclosure. The methylation status at each CpG site of each long plasma DNA molecule is determined according to the examples in the present disclosure. The methylation haplotype of each long plasma DNA molecule is determined according to the examples in the present disclosure. If a long plasma DNA molecule is identified as carrying a placenta-specific methylation haplotype, the genetic and epigenetic information associated with the molecule is considered to be inherited from the fetus. In an example, if one or more long plasma DNA molecules containing the same pathogenic mutation as the pathogenic mutation carried by the pregnant woman are determined to be of fetal origin based on methylation haplotype information according to the examples in the present disclosure, it indicates that the fetus has inherited the mutation from the mother.

[0308] The embodiments can be applicable to genetic diseases including but not limited to the following: β-thalassemia, sickle cell anemia, α-thalassemia, cystic fibrosis, hemophilia A, hemophilia B, congenital adrenal hyperplasia, Duchenne muscular dystrophy, Becker muscular dystrophy, achondroplasia, lethal dysplasia, von Willebrand disease, Noonan syndrome, hereditary hearing loss and deafness, various inborn errors of metabolism (such as type I citrullinemia, propionic academia, type Ia glycogen storage disease (von Gierke disease), type Ib / c glycogen storage disease (von Gierke disease), type II glycogen storage disease (Pompe disease), type I mucopolysaccharidosis (MPS) (Hurler / Hurler-Scheie / Scheie), type II MPS (Hunter syndrome), type IIIA MPS (Sanfilippo syndrome A), type IIIB MPS (Sanfilippo syndrome B), type IIIC MPS (Sanfilippo syndrome C), type IIID MPS (Sanfilippo syndrome D), type IVA MPS (Morquio syndrome A), type IVB MPS (Morquio syndrome B), type VI MPS (Maroteaux-Lamy syndrome), type VII MPS (Sly syndrome), mucolipidosis II (I-cell disease), metachromatic leukodystrophy, GM1 gangliosidosis, OTC deficiency (X-linked ornithine transcarbamylase deficiency), adrenoleukodystrophy (X-linked ALD), Krabbe disease (globoid cell leukodystrophy), etc.).

[0309] In other embodiments, the genetic disease of the fetus may be related to de novo DNA methylation in the fetal genome that does not exist in the parental genome. An example would be the hypermethylation of the FMRP translational regulator 1 (FMR1) gene in a fetus with fragile X syndrome. Fragile X syndrome is caused by an expansion of the CGG trinucleotide repeat sequence in the 5' untranslated region of the FMR1 gene. A normal allele will contain approximately 5 to 44 copies of the CGG repeat sequence. A premutation allele will contain 55 to 200 copies of the CGG repeat sequence. A full mutation allele will contain more than 200 copies of the CGG repeat sequence.

[0310] Figure 28Illustrates the principle of non-invasive prenatal testing for Fragile X syndrome in male fetuses of unaffected pregnant women carrying normal or premutation alleles. In Figure 28 where 'n' represents the number of CGG repeats in the maternal genome; 'm' represents the number of CGG repeats in the fetal genome. The genome of an unaffected pregnant woman will have the FMR1 gene, which has a CGG repeat sequence with no more than 200 repeats (i.e., n ≤ 200) and is unmethylated. In contrast, the genome of a male fetus affected by Fragile X syndrome will have an FMR1 gene with more than 200 CGG repeat copies (m > 200) and is methylated. We can identify multiple long DNA molecules from the genomic regions of interest (such as the FMR1 gene) where the repeat number and methylation status can be determined simultaneously by performing single-molecule sequencing of maternal plasma DNA. If we identify one or more DNA molecules in the plasma of an unaffected woman that cover the FMR1 gene, contain more than 200 CGG repeat copies, and are methylated, it will indicate that the fetus is likely to have Fragile X syndrome. In yet another embodiment, we can further determine the fetal origin of the plasma DNA molecules according to the embodiments in the present disclosure using placenta-specific methylation haplotypes. If we identify one or more molecules that contain one or more regions within a molecule carrying a placenta-specific methylation haplotype and the molecule covers the FMR1 gene, contains more than 200 CGG repeat copies, and is methylated, we can conclude with more certainty that the fetus has Fragile X syndrome. Conversely, if we identify one or more molecules with a placenta-specific methylation haplotype and the molecule covers the FMR1 gene, contains fewer than 200 CGG repeat copies, and is not methylated, it will indicate that the fetus is likely to be unaffected. In the case of Fragile X syndrome, a full mutation (>200 repeats) actually methylates the entire gene and shuts off gene function. Therefore, especially for Fragile X, the detection of long alleles that are methylated (rather than showing a placental methylation profile) will highly indicate that the fetus has the disease.

[0311] Testing for genetic disorders can be performed with or without knowledge of the mother's previous status. Women with a premutation may not have any symptoms, but some women with a premutation may have mild symptoms and often only find out later. If we do not know the maternal mutation status, one approach is to test for long alleles in the plasma of women who do not seem to have the disease or analyze the maternal buffy coat and determine that it does not show such long alleles. As another approach, we can combine the repeat length of cfDNA molecules with the methylation status. If the methylation status indicates a fetal pattern (methylation haplotype) and shows a long allele, the fetus may be affected. This approach is applicable to many trinucleotide disorders such as Huntington's disease.

[0312] D. Non-invasive Construction of Fetal Genome with Long Plasma DNA Molecules

[0313] Methylation patterns can be used to determine haplotype inheritance. Using a qualitative approach with methylation patterns to determine haplotype inheritance can be more efficient than a quantitative method that characterizes the amount of a specific fragment. Methylation patterns can be used to determine maternal and paternal inheritance of haplotypes.

[0314] 1. Maternal Inheritance of the Fetus

[0315] Lo et al. demonstrated the feasibility of constructing a genome-wide genetic map using parental haplotype information and determining the mutational status of the fetus from maternal plasma DNA sequences (Lo et al., Science Translational Medicine 2010; 2:61ra91). This technique has been termed relative haplotype dosage (RHDO) analysis and is a means of resolving the maternal inheritance of the fetus. The principle is based on the fact that the maternal haplotype inherited by the fetus will be relatively overrepresented in the pregnant woman's plasma DNA when compared to another maternal haplotype that was not transmitted to the fetus. Thus, RHDO is a quantitative analysis method.

[0316] Examples present in the present disclosure utilize methylation patterns in long plasma DNA molecules to determine the origin tissue of the plasma DNA molecules. In one embodiment, the disclosure herein will allow for a qualitative analysis of the maternal inheritance of the fetus.

[0317] Figure 29 An example of determining the maternal inheritance of the fetus is shown. The genomic position P in the maternal genome (A / G) is heterozygous. The filled circles indicate methylation sites, and the open circles indicate unmethylated sites. The methylation pattern in the placenta is "-M-U-M-M-", where "M" represents methylated cytosine at a CpG site and "U" represents unmethylated cytosine at a CpG site. In one embodiment, the methylation patterns in the placenta and associated reference tissues can be obtained from data previously generated by sequencing (e.g., single molecule real-time sequencing and / or bisulfite sequencing). In the plasma DNA, a non-paternal plasma DNA (designated Z) carrying allele A at the specific genomic locus was found to exhibit a methylation pattern ("-M-U-M-M-") that is compatible with the methylation pattern in the placenta relative to the methylation patterns of other tissues. No molecules carrying allele G that exhibited a methylation pattern compatible with the methylation pattern in the placenta were found. Thus, based on the presence of allele A and the "-M-U-M-M-" methylation pattern, it can be determined that the fetus inherited the maternal allele A.

[0318] Figure 30 A qualitative analysis of the maternal inheritance of the fetus using genetic and epigenetic information of plasma DNA molecules is shown. AsFigure 30 As shown in the top branch of, according to an embodiment in the present disclosure, plasma DNA is extracted and then size selection for long DNA is performed. The size-selected plasma DNA molecules are subjected to single molecule real-time sequencing (e.g., using a system manufactured by Pacific Biosciences). Genetic and epigenetic information is determined according to an embodiment in the present disclosure. For illustrative purposes, molecule (X) is aligned with human chromosome 1 that has allele G at chromosomal position a (chr1:a) and allele A at chromosomal position e (chr1:e). Molecule X has allele C at chromosomal position d.

[0319] The CpG methylation status of this molecule X is determined to be "-M-U-M-M-", where "M" represents methylated cytosine at the CpG site and "U" represents unmethylated cytosine at the CpG site. The filled circles indicate methylation sites, and the hollow circles indicate unmethylated sites. As a result of the analysis of the reference sample, it is known that placental DNA has a methylation pattern of "-M-U-M-M-" in the region between positions a and e. According to an embodiment in the present disclosure, based on the methylation pattern of molecule X that matches the methylation pattern of placental DNA, molecule X is determined to be of placental origin.

[0320] As Figure 30 As shown in the lower branch of, DNA from maternal white blood cells is subjected to single molecule real-time sequencing. Epigenetic and genetic information of maternal white blood cells is obtained according to an embodiment in the present disclosure. The genetic alleles are phased into two haplotypes, namely maternal haplotype I (Hap I) and maternal haplotype II (Hap II), using methods including but not limited to WhatsHap (Patterson et al., Journal of Computational Biology 2015; 22:498-509), HapCUT (Bansal et al., Bioinformatics 2008; 24:i153-9), HapCHAT (Beretta et al., BMC Bioinformatics 2018; 19:252), etc. Here, we obtain two haplotypes in the maternal genome, namely "-A-C-G-T-" (Hap I) and "-G-T-A-C-" (Hap II). Hap I is associated with one or more wild-type variants, while Hap II is associated with one or more disease-related variants. One or more disease-related variants may include but are not limited to single nucleotide variants, insertions, deletions, translocations, inversions, repeat sequence expansions, and / or other genetic structural variations.

[0321] For genomic location e, the maternal genotype was determined to be AA and the paternal genotype was determined to be GG. Due to the methylation pattern, plasma DNA molecule X was determined to be of placental origin. Since the maternal-specific allele A was present but the paternal-specific allele G was absent, molecule X was inferred to be inherited from one of the maternal haplotypes.

[0322] To further determine which maternal haplotype was transmitted to the fetus, we compared the allelic information at genomic locations other than position chr1:e of this placental-derived molecule X with the maternal haplotypes. As an example, molecule X had allele G at position a and allele C at position d. The presence of either of these alleles in molecule X indicated that molecule X should be assigned to maternal Hap II, which contained the same alleles.

[0323] Therefore, we can conclude that maternal haplotype II, which is associated with one or more disease-related variants, was transmitted to the fetus. The unborn fetus was determined to be at risk of being affected by the disease.

[0324] Compared with RHDO, which is a quantitative analysis-based approach, the methylation pattern-based qualitative analysis of maternal inheritance for the fetus may require fewer plasma DNA molecules to draw conclusions about which maternal haplotype the fetus inherits. We performed in silico analysis genome-wide with different numbers of plasma DNA molecules for analysis to evaluate the detection rate of maternal inheritance for the fetus.

[0325] For the RHDO simulation analysis, N plasma DNA molecules were collectively aligned with M heterozygous SNPs in the haplotype blocks of the maternal genome. The fetal DNA fraction was f. The paternal genotypes of those corresponding SNPs were homozygous and the same as maternal Hap I transmitted to the fetus. Among the N plasma DNA molecules, the average number of plasma DNA molecules aligned with maternal Hap I was N×(0.5 + f / 2), while the average number of plasma DNA molecules aligned with maternal Hap II was N×(0.5 - f / 2). We assumed that the plasma DNA molecules sampled from the haplotypes followed a binomial distribution.

[0326] The number of plasma DNA molecules assigned to Hap I (i.e., X) followed the following distribution:

[0327] X ∼ Bin(N,0.5 + f / 2)(1),

[0328] where "Bin" represents the binomial distribution.

[0329] The number of plasma DNA molecules assigned to Hap II (i.e., Y) followed the following distribution:

[0330] Y ~ Bin(N, 0.5 - f / 2) (2).

[0331] Thus, compared to the maternal Hap II, plasma DNA molecules assigned to the maternal Hap I will be relatively overrepresented in maternal plasma. To determine whether the overrepresentation is statistically significant, we compare the difference in plasma DNA counts between the two maternal haplotypes under the null hypothesis that two of the haplotypes (represented by X' and Y') are also represented in plasma.

[0332] X' ~ Bin(N, 0.5) (3),

[0333] Y' ~ Bin(N, 0.5) (4).

[0334] We further define the relative dosage difference between two haplotypes as follows:

[0335] D = (X - Y) / N (5),

[0336] D' = (X' - Y') / N (6).

[0337] In one example, the statistic D reflecting the relative haplotype dosage is compared to the mean (M) of D', and is normalized by the standard deviation (SD) of D' to the following (i.e., z-score):

[0338] z - score = (D – M) / SD (7).

[0339] A z - score > 3 indicates that Hap I is transmitted to the fetus.

[0340] For RHDO analysis, based on formulas (1) to (7), we simulate 30,000 haplotype blocks across the entire genome where Hap I is transmitted to the fetus. The average length of a haplotype block is 100 kb. Each haplotype block contains on average 100 SNPs, of which 10 SNPs will provide information in contributing to haplotype disequilibrium. In one example, the fetal DNA fraction is 10% and the median fragment size is 150 bp. We calculate the percentage of haplotype blocks with a z - score > 3 by varying the number of plasma DNA molecules used for RHDO analysis in the range of 1 million to 300 million, which percentage is referred to herein as the detection rate. The number of plasma DNA molecules herein is adjusted according to the Poisson distribution by the probability of plasma DNA covering informative SNP sites.

[0341] For computer simulations related to the qualitative analysis of maternal inheritance for the fetus based on methylation patterns, we make the following assumptions for illustrative purposes:

[0342] 1) There are N plasma DNA molecules in the maternal genome for analysis that cover haplotype blocks.

[0343] 2) The probability of plasma DNA fragments with a length of at least 3 kb for origin tissue analysis is represented by a.

[0344] 3) The probability of plasma DNA molecules carrying more than 10 CpG sites is represented by b.

[0345] 4) The fetal DNA fraction of those fragments > 3 kb is represented by f.

[0346] As depicted in an embodiment of the present disclosure, we can achieve an accurate inference of the origin tissue of plasma DNA molecules greater than 3 kb that have at least 10 CpG sites. Assuming that the number of plasma DNA molecules (Z) meeting the above criteria follows a Poisson distribution, where the mean is λ (i.e., N × a × b × f).

[0347] Z ~ Poisson(λ) (8).

[0348] In one example, based on formula (8), we simulated 30,000 haplotype blocks in which Hap I was transmitted to the fetus. The average length of each haplotype block was 100 kb. Each haplotype block contained an average of 100 SNPs, of which 20 heterozygous SNPs would be phased into two maternal haplotypes. The fetal DNA fraction was 1%. After size selection, there were 40% plasma DNA molecules with a size > 3 kb. There were 87.1% plasma DNA molecules with a size > 3 kb that had at least 10 CpG sites. The percentage of haplotype blocks with a Z value ≥ 1 indicates the detection rate. We repeated the computer simulation operation multiple times by varying the number of plasma DNA molecules (N) for origin tissue analysis in the range of 1 million to 300 million using methylation pattern changes. The number of plasma DNA molecules herein was further adjusted according to the Poisson distribution by the probability of plasma DNA covering heterozygous SNPs.

[0349] Figure 31 Shows the detection rate of the qualitative analysis of maternal inheritance of the fetus using the genetic and epigenetic information of plasma DNA molecules in a genome-wide manner compared to relative haplotype dosage (RHDO) analysis. The number of molecules for analysis is shown on the x-axis. The detection rate of maternal inheritance of the fetus in percentage form is shown on the y-axis. Compared with RHDO, using the methylation pattern-based approach, the detection rate of maternal inheritance of the fetus is higher. For example, using 100 million fragments, the detection rate based on the methylation pattern is 100%, while the detection rate based on RHDO is only 55%. These results indicate that the inference of maternal inheritance of the fetus using the methylation pattern-based method will be superior to the inference of maternal inheritance of the fetus using the RHDO-based method.

[0350] 2. Paternal inheritance of the fetus

[0351] The ability to obtain long plasma DNA molecules for analysis can be applied to improve the detection rate of paternal-specific variants in maternal plasma DNA because the use of the same number of long DNA molecules will increase the total genomic coverage compared to the use of short DNA molecules. We further perform computer simulations based on the following assumptions:

[0352] 1) The fetal DNA fraction is f depending on the plasma DNA length L. It is rewritten as f L , where the subscript L indicates that plasma DNA molecules with a length of L bp are used for analysis.

[0353] 2) The number of paternal-specific variants to be identified in maternal plasma DNA is V.

[0354] 3) The number of plasma DNA molecules used for analysis is N.

[0355] 4) The number of plasma DNA molecules derived from a specific genomic locus or region follows a Poisson distribution.

[0356] In one example, the fetal DNA fractions of plasma DNA molecules with sizes of 150 bp, 1 kb, and 3 kb are 10% (f 150bp = 0.1), 2% (f 1kb = 0.02), and 1% (f 3kb = 0.01), respectively. The number of paternal-specific variants in the genome is 250,000 (V = 250,000). The number of plasma DNA molecules used for analysis (N) ranges from 50 million to 500 million.

[0357] Figure 32 Shows the relationship between the detection rate of paternal-specific variants performed in a genome-wide manner and the number of sequenced plasma DNA molecules of different sizes used for analysis. The number of sequenced molecules used for analysis in millions is shown on the x-axis. The percentage of detected paternal-specific variants is shown on the y-axis. Different curves show different sizes of DNA fragments used for analysis, with 3 kb at the top, 1 kb in the middle, and 150 bp at the bottom. The longer the plasma DNA molecules used for analysis, the higher the detection rate of paternal-specific variants that can be achieved. For example, using 400 million plasma DNA molecules, when focusing on molecules with sizes of 150 bp, 1 kb, and 3 kb, the detection rates are 86%, 93%, and 98%, respectively.

[0358] In other embodiments, other distributions including but not limited to Bernoulli distribution, beta (β) normal distribution, normal distribution, Conway-Maxwell-Poisson distribution, geometric distribution, etc. may be used. In some embodiments, Gibbs sampling and Bayes' theorem will be used for maternal and parental genetic analysis.

[0359] 3. Fragile X chromosome genetic analysis

[0360] In an embodiment, the determination of fetal maternal inheritance based on methylation patterns can facilitate non-invasive detection of fragile X syndrome using single molecule real-time sequencing of maternal plasma DNA. Fragile X syndrome is a genetic disorder typically caused by an expansion of a CGG trinucleotide repeat sequence within the FMR1 (fragile X mental retardation 1) gene on the X chromosome. Fragile X syndrome and other disorders caused by repeat sequence expansions are described elsewhere in this application. The method for detecting fragile X syndrome in a fetus can also be applied to any other repeat sequence expansion disclosed herein.

[0361] Female individuals with a premutation are at risk of having a child with fragile X syndrome, where the premutation is defined as having 55 to 200 CGG repeat sequence copies in the FMR1 gene. The likelihood of having a fetus with fragile X syndrome depends on the number of CGG repeat sequences present in the FMR1 gene. The greater the number of repeat sequences in the mother, the higher the risk of expansion from a premutation to a full mutation when passed on to the fetus. Maternal plasma samples were collected at 12 weeks of gestation from a woman who had previously been confirmed to carry a fragile X premutation allele with 115 ± 2 CGG repeat sequences and had a son (the proband) diagnosed with fragile X syndrome. Subsequently, the maternal plasma was subjected to single molecule real-time sequencing. In one example, we obtained 3.3 million circular consensus sequences (CCS) aligned to the human reference genome using single molecule real-time sequencing, with a median subread depth of 75x / CCS (interquartile range: 14 - 237x). The genetic and epigenetic information of each sequenced plasma DNA can be determined according to the embodiments of this disclosure. To obtain the two maternal haplotypes of chromosome X, we used the Infinium Omni2.5Exome-8 Beadchip, a microarray technology on the iScan system (Illumina), to genotype 2,000 SNPs on chromosome X for two DNAs extracted from the maternal buffy coat and the proband's buccal swab. The two maternal haplotypes, namely Hap I and Hap II, can be inferred based on the genotype information of the maternal and proband genomes.

[0362] Figure 33Disclosed is a workflow for non-invasive detection of Fragile X syndrome. Among the heterozygous SNP sites of maternal buffy coat DNA, the allele identical to the genotype of the proband is used to define a haplotype (i.e., Hap I) related to the premutation allele, which is a potential precursor of a full mutation in the offspring. On the other hand, the allele different from the genotype of the proband is used to define a haplotype (Hap II) related to the corresponding wild-type allele. Maternal plasma DNA from the mother of a pregnant fetus of the proband is subjected to single molecule real-time sequencing. Depending on whether the genetic information obtained is identical to the alleles of Hap I or Hap II in those genomic loci under study, the sequencing reads are assigned to maternal Hap I and Hap II. According to an embodiment in the present disclosure, the methylation pattern of plasma DNA molecules is used to determine the origin tissue of those plasma DNA molecules containing a specific number of CpG sites (i.e., DNA molecules identified as being of placental origin based on methylation pattern analysis will be determined to be of fetal origin).

[0363] In scenario A, if fetal (i.e., placental) DNA molecules can be detected in those plasma DNA molecules assigned to maternal Hap I but not in those plasma DNA molecules assigned to maternal Hap II, then Hap I will be determined to be transmitted to the unborn fetus. The fetus will be determined to be at high risk of being affected by Fragile X syndrome. The placental origin of plasma DNA molecules will be based on the methylation status of the molecules as discussed below.

[0364] In scenario B, if fetal DNA molecules can be detected in those plasma DNA molecules assigned to maternal Hap II but not in those plasma DNA molecules assigned to maternal Hap I, then Hap II will be determined to be transmitted to the unborn fetus. The fetus will be determined to be unaffected by Fragile X syndrome.

[0365] In embodiments, the definitions of "detectable" and "undetectable" fetal DNA molecules may depend on a cut-off value of the percentage of plasma DNA molecules identified as being of fetal (i.e., placental) origin. Cut-off values for "detectable" may include, but are not limited to, greater than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, etc. Cut-off values for "undetectable" may include, but are not limited to, less than 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, etc. In some embodiments, a difference in the percentage of plasma DNA molecules determined to be of fetal origin between Hap I and Hap II may be required to be greater than, but not limited to, 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, etc. In some other embodiments, haplotype information may be obtained from long-read sequencing techniques (such as PacBio or nanopore sequencing) (Edge et al., Nat Commun. 2019; 10:4660), synthetic long reads (such as using a platform from 10X Genomics) (Hui et al., Clin Chem. 2017; 63:513-14), phasing based on targeted locus amplification (TLA) (Vermeulen et al., Am J Hum Genet. 2017; 101:326-39), and statistical phasing (such as Shape-IT) (Delaneau et al., Nat Methods. 2011; 9:179-81).

[0366] In embodiments, we may determine the maternal and fetal origin of those plasma DNA molecules that are at least 200 bp and contain at least 5 CpG sites (or any other cut-off value for long DNA molecules) according to the methylation state matching pathway disclosed in the present application. We identified a plasma DNA molecule located at genomic position chrX:143,782,245-143,782,786 (3.2 Mb away from the FMR1 gene), where the allele (position: chrX:143782434; SNP accession number: rs6626483; allele genotype: C) is the same as the corresponding allele on maternal Hap II, but different from the corresponding allele on maternal Hap I.

[0367] Figure 34Show the methylation pattern of plasma DNA compared to the methylation profiles of placental and buffy coat DNA. Plasma DNA molecules contain five CpG sites. The methylation pattern was determined to be "M-U-U-U-U". According to the methylation status matching approach described in this disclosure, this methylation pattern obtained from single molecule real-time sequencing was compared with the reference methylation profiles of placental tissue and buffy coat DNA samples obtained from bisulfite sequencing. The fraction of this molecule derived from the placenta [i.e., S(placenta)] was 2, which is greater than the fraction of the molecule derived from the buffy coat [i.e., S(buffy coat)] -3. Therefore, such plasma DNA molecules (chrX:143,782,245-143,782,786) were determined to be of fetal origin. However, we did not observe any plasma DNA molecules carrying the allele of maternal Hap I of fetal origin. Therefore, we conclude that the fetus inherits maternal Hap II and is not affected by fragile X syndrome.

[0368] We envision that the efficacy of the approach described herein may not be significantly affected by X chromosome inactivation due to the following factors:

[0369] 1) X inactivation is not complete in humans. Up to one-third of the genes on the X chromosome show variable escape from X inactivation (Cotton et al., Hum Mol Genet. 2015; 25:1528-1539). CpG sites outside CpG islands (i.e., most CpG sites) are methylated to a similar extent in both sexes, indicating that the methylation status of most CpG sites in the X chromosome may not be affected by X inactivation (Yasukochi et al., Proc Natl Acad Sci USA 2010; 107:3704-9).

[0370] 2) We used the methylation profiles of sex-matched placental tissue of the unborn fetus. This strategy can be used to detect the maternal inheritance of the fetus using the plasma DNA methylation pattern of women carrying male fetuses because the placental tissue involving male fetuses, which is presumably not affected by X inactivation, will have a unique methylation pattern different from other maternal tissues that are more or less involved in X inactivation of specific regions.

[0371] We further sequenced the DNA extracted from the maternal buffy coat sample using single molecule real-time sequencing. We obtained 2.3 million CCSs, with a median subread depth of 5X / CCS. The results confirmed that maternal Hap I carried a premutation allele with 124 CGG repeats and maternal Hap II carried a wild-type allele with 43 CGG repeats. In addition, we further sequenced the DNA extracted from chorionic sampling of the unborn fetus using single molecule real-time sequencing. We obtained 1.1 million CCSs, with a median subread depth of 4X / CCS. The results confirmed that the unborn fetus carried a wild-type allele.

[0372] E. Distribution of CpG sites in the human genome

[0373] Longer DNA fragments result in a greater chance of larger fragments having multiple CpG sites. These multiple CpG sites can be used for methylation pattern or other analyses.

[0374] Figure 35 Shows the distribution of CpG sites in 500bp regions across the human genome. The first column shows the number of CpG sites. The second column shows the number of 500bp regions and the number of CpG sites. The third column shows the proportion of all regions represented by regions with a specific number of CpG sites. For example, 86.14% of 500bp regions will have at least 1 CpG site. Additionally, 11.08% of 500bp regions will have at least 10 CpG sites.

[0375] Figure 36 Shows the distribution of CpG sites in 1kb regions across the human genome. The first column shows the number of CpG sites. The second column shows the number of 1kb regions and the number of CpG sites. The third column shows the proportion of all regions represented by regions with a specific number of CpG sites. For example, 91.67% of 500bp regions will have at least 1 CpG site. In addition, 32.91% of 500bp regions will have at least 10 CpG sites.

[0376] Figure 37 Shows the distribution of CpG sites in 3kb regions across the human genome. The first column shows the number of CpG sites. The second column shows the number of 3kb regions and the number of CpG sites. The third column shows the proportion of all regions represented by regions with a specific number of CpG sites. For example, 92.45% of 3kb regions will have at least 1 CpG site. Additionally, 87.09% of 3kb regions will have at least 10 CpG sites.

[0377] In some embodiments, different numbers and different size cutoffs of CpG sites will be used to maximize the sensitivity and specificity of placental specific marker identification and origin tissue analysis. Generally, the occurrence frequency of CpG sites is higher than that of SNPs. A DNA fragment of a given size may have more CpG sites than SNPs. The table shown above may show a lower proportion of regions having the same number of SNPs as CpG sites because there are fewer SNPs in the same size region than CpG sites. Thus, using CpG sites allows the use of more fragments and provides better statistics than using only SNPs.

[0378] F. Examples of Origin Tissue Analysis

[0379] In an embodiment, we can extend the origin tissue analysis in maternal plasma to more than two organs / tissues including T cells, B cells, neutrophils, liver, and placenta. We used single molecule real-time sequencing to sequence 9 maternal DNA samples. We inferred the contribution of the placenta to maternal plasma DNA using the plasma DNA methylation pattern according to the methylation state matching approach described in the present disclosure. In one embodiment, for this methylation state matching analysis, the methylation pattern of each DNA molecule in the maternal plasma DNA sample that is at least 500 bp in length and contains at least 5 CpG sites is compared with the reference tissue methylation profile obtained from bisulfite sequencing. Five tissues including neutrophils, T cells, B cells, liver, and placenta are used as reference tissues. The plasma DNA molecules will be assigned to the tissue corresponding to the maximum methylation state matching score of the plasma DNA molecule. The percentage of plasma DNA molecules assigned to a tissue relative to other tissues will be regarded as the contribution ratio of the tissue to the maternal plasma DNA of the sample. In an embodiment, the sum of the contribution ratios of neutrophils, T cells, and B cells in maternal plasma provides an alternative representation of the contribution ratio of hematopoietic cells.

[0380] Figure 38 Shows the contribution ratios of DNA molecules from different tissues in maternal plasma using methylation state matching analysis. The first column shows sample identification. The second column shows the hematopoietic cell contribution in percentage. The third column shows the liver contribution in percentage. The fourth column shows the placenta contribution in percentage. Figure 38 Shows that the major contributor to maternal plasma DNA is hematopoietic cells (median: 55.9%), which is consistent with previous reports (Sun et al. Proc Natl Acad Sci USA 2015; 112:E5503-12; Zheng et al. Clin Chem 2012; 58:549-58).

[0381] Figure 39A and Figure 39BDisplays the relationship between placental contribution and the fraction of fetal DNA inferred through the SNP approach. The X-axis shows the fraction of the fetus determined through the SNP approach. The Y-axis shows the placental contribution in maternal plasma in percentage form determined by using methylation status matching analysis. Figure 39A Displays a good correlation between the placental contribution determined by methylation status matching analysis and the fraction of fetal DNA inferred through SNPs (Pearson's r = 0.95; P value < 0.0001). We further performed tissue deconvolution analysis of maternal plasma DNA by quadratic programming by comparing the plasma DNA methylation density determined by single molecule real-time sequencing with various reference tissue methylation profiles obtained from bisulfite sequencing (Sun et al. Proceedings of the National Academy of Sciences of the United States of America 2015; 112: E5503-12). Figure 39B Displays a reduced correlation between placental contribution (Sun et al. Proceedings of the National Academy of Sciences of the United States of America 2015; 112: E5503-12) and the fraction of fetal DNA when using a methylation density-based approach compared to using methylation status matching analysis (Pearson's correlation coefficient = 0.65; P value = 0.059).

[0382] These data indicate that it is feasible to infer the proportion of DNA molecules contributed by different tissues in a maternal plasma DNA sample. In another embodiment, this method can also be used to measure DNA molecules from different cell types or tissues in a sample obtained after an invasive solid tissue biopsy or from solid tissue obtained after surgery. In some embodiments, the use of methylation patterns at the single DNA molecule level to infer the proportion of contributions from different tissues to maternal plasma DNA will be superior to the use of an approach based on the collective methylation density of all sequenced plasma DNA molecules across the genome.

[0383] G. Example Method

[0384] Figure 40 Displays method 4000 for analyzing a biological sample obtained from a female pregnant with a fetus. The biological sample can contain multiple cell-free DNA molecules from the fetus and the female.

[0385] At block 4010, sequence reads corresponding to multiple cell-free DNA molecules can be received. In some embodiments, method 4000 can include performing sequencing of the cell-free DNA molecules.

[0386] At block 4020, the sizes of multiple cell-free DNA molecules can be measured. The measurement can include aligning sequence reads to a reference genome. In some embodiments, the measurement can include full-length sequencing of nucleotides in a full-length sequence and counting their number. In some embodiments, the measurement can include physically separating multiple cell-free DNA molecules from a biological sample from other cell-free DNA molecules in the biological sample, where the other cell-free DNA molecules have a size less than a cut-off value. The physical separation can include any technique described herein, including using beads.

[0387] At block 4030, a set of cell-free DNA molecules from multiple cell-free DNA molecules can be identified as having a size greater than or equal to a cut-off value. The cut-off value can be greater than or equal to 200 nt. The cut-off value can be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cut-off value can be any cut-off value described herein for long cell-free DNA molecules. The size can be the number of CpG sites rather than the molecule length. For example, the cut-off value can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more CpG sites.

[0388] At block 4040, for a cell-free DNA molecule in the set of cell-free DNA molecules, the methylation status at each of multiple sites can be determined. The multiple sites can include at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more CpG sites. At least one of the multiple sites can be methylated. Two sites among the multiple sites can be spaced at least 160 nt, 170 nt, 180 nt, 190 nt, 200 nt, 250 nt, or 500 nt. The method can include sequencing multiple cell-free DNA molecules to obtain sequence reads and determining the methylation status of a site by measuring the characteristics of the nucleotide corresponding to the site and the nucleotides of neighboring sites. For example, methylation can be determined generally as in U.S. Application No. 16 / 995,607.

[0389] At block 4050, a methylation pattern can be determined. The methylation pattern can indicate the methylation status at each of multiple sites.

[0390] At block 4060, the methylation pattern can be compared to one or more reference patterns. Each of the one or more reference patterns can be determined for a specific tissue type. In some embodiments, the comparison can include determining the number of sites that match the reference pattern.

[0391] The reference patterns in one or more reference patterns can be determined by measuring the methylation density at each of a plurality of reference sites using DNA molecules from a reference tissue. The methylation density at each of the plurality of reference sites can be compared with one or more threshold methylation densities. Each of the plurality of reference sites can be identified as methylated, unmethylated, or non-informative based on comparing the methylation density with one or more threshold methylation densities, where the plurality of sites are the plurality of reference sites identified as methylated or unmethylated. Non-informative sites can include sites having a methylation density between two threshold methylation densities. For example, the methylation index of a non-informative site can be between 30 and 70 or within any other range described herein.

[0392] At step 4070, the origin tissue of the cell-free DNA molecules can be determined using the methylation pattern. The origin tissue can be the placenta. The origin tissue can be fetal or maternal. The method can include determining the origin tissue as the reference tissue when the methylation pattern matches the reference pattern, similar to that described with Figure 22 The match can refer to an exact match. In some embodiments, determining the origin tissue as the reference tissue can occur when the methylation pattern matches the reference pattern at a particular percentage of sites. For example, the methylation pattern can match the reference pattern at at least 60%, 70%, 80%, 85%, 90%, 95%, 97% or more of the sites.

[0393] The method can include determining the origin tissue by determining a similarity score by comparing the methylation pattern with a first reference methylation pattern from a first reference tissue among a plurality of reference tissues. The similarity score can be calculated using the methylation state matching method or the beta (β) distribution probability model described herein. The similarity score can be compared with a threshold. When the similarity score exceeds the threshold, the origin tissue can be determined as the first reference tissue. The similarity score can be a first similarity score. The method can further include calculating the threshold by determining a second similarity score by comparing the methylation pattern with a second reference methylation pattern from a second reference tissue among a plurality of reference tissues. The first reference tissue and the second reference tissue can be different tissues. The threshold can be the second similarity score. The first reference tissue can have the highest similarity score compared to all other reference tissues.

[0394] The first reference methylation pattern may include a first subgroup of sites having at least a first methylation probability for the first reference tissue. For example, the first subgroup of sites may be sites that are considered methylated or are generally considered methylated. The first reference methylation pattern may include a second subgroup of sites having at most a second methylation probability for the first reference tissue. For example, the second subgroup of sites may be sites that are considered unmethylated or are generally considered unmethylated. Determining the similarity score may include increasing the similarity score when a site among the plurality of sites is methylated and the site among the plurality of sites is in the first subgroup of sites, and decreasing the similarity score when a site among the plurality of sites is methylated and the site among the plurality of sites is in the second subgroup of sites. The similarity score may be determined to be similar to the methylation status matching approach described herein.

[0395] The first reference methylation pattern includes a plurality of sites, wherein each site among the plurality of sites is characterized by a methylation probability and an unmethylation probability for the first reference tissue. For each site among the plurality of sites, the similarity score may be determined by determining the probability in the reference tissue corresponding to the methylation status of the site in the cell-free DNA molecule. The similarity score may be determined by calculating the product of the plurality of probabilities. The product may be the similarity score. The probability may be determined by a beta (β) distribution, similar to the approach described herein.

[0396] Method 4000 may further include determining the origin tissue of each cell-free DNA molecule in the set of cell-free DNA molecules. This determination may include determining the methylation status at each site among a plurality of corresponding sites, where the plurality of corresponding sites correspond to the cell-free DNA molecule. Determining the origin tissue may further include determining a methylation pattern. Additionally, determining the origin tissue may also include comparing the methylation pattern with at least one reference pattern among one or more reference patterns. In some embodiments, the comparison of the methylation patterns may be similar to Figure 22 and the accompanying description. In Figure 22 , placenta, liver, blood cells, and colon are examples of reference tissues having the illustrated reference patterns. Figure 38 Hematopoietic cells are shown as another example of a reference tissue.

[0397] In some embodiments, the amount of cell-free DNA molecules corresponding to each originating tissue can be measured. Each originating tissue can include each of a plurality of reference tissues. The contribution fraction of the originating tissue can be determined using the measurement of the cell-free DNA molecules corresponding to each originating tissue. For example, the originating tissue can be the placenta. Other originating tissues can include hematopoietic cells and the liver. For example, the contribution fraction of the placenta can be determined by dividing the amount of cell-free DNA molecules by the total cell-free DNA molecules corresponding to all originating tissues. In some embodiments, the fraction calculated by dividing the amount of cell-free DNA molecules by the total cell-free DNA molecules can be related to the contribution fraction via a function or a set of calibration data points. Both the function and the set of calibration data points can be determined from a plurality of calibration samples having known contribution fractions of the originating tissue. Each calibration data point can specify the contribution fraction corresponding to the calibration value of the fraction. The function can represent a linear or non-linear fit of the calibration data points and can relate the contribution fraction to the fraction of the originating tissue or other parameters related to the originating tissue. Embodiments for determining the contribution fraction can be similar to the embodiments already described using Figure 39A and Figure 39B are described.

[0398] A machine learning model can be used to determine the originating tissue. The model can be trained by receiving a plurality of training methylation patterns, each training methylation pattern having a methylation state at one or more of a plurality of sites, each training methylation pattern being determined from a DNA molecule from a known tissue. Each molecule from a known tissue can be cellular DNA. Training can include storing a plurality of training samples, each training sample including one of the plurality of training methylation patterns and a label indicating the known tissue corresponding to the training methylation pattern. Training can include using the plurality of training samples to optimize the parameters of the model based on the output of the model when the plurality of training methylation patterns are input into the model and matching or not matching the corresponding labels. The parameters can include a first parameter indicating whether one site among the plurality of sites has the same methylation state as another site among the plurality of sites. For example, the model can be similar to the Figure 24 pairwise comparison. The parameters can include a second parameter indicating the distance between each site among the plurality of sites. In some embodiments, the machine learning model may not require alignment of the methylation sites with a reference genome. The output of the model can specify the tissue corresponding to the input methylation pattern.

[0399] The machine learning model can be a convolutional neural network (CNN) or any model described herein. The model can include, but is not limited to, linear regression, logistic regression, deep recurrent neural networks (e.g., long short-term memory, LSTM), Bayesian classifiers, hidden Markov models (HMM), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering of applications with noise (DBSCAN), random forest algorithms, and support vector machines (SVM).

[0400] Relatedness can be determined by method 4000. The originating tissue can be a fetus. The method can further include aligning one of the sequence reads to a first region of a reference genome, the first region including a plurality of sites corresponding to alleles, the plurality of sites including a threshold number of sites, determining a first haplotype using the corresponding alleles present at each of the sites in the plurality of sites, comparing the first haplotype to a second haplotype corresponding to a male individual, and using the comparison to determine a classification of the likelihood that the male individual is the father of the fetus. If the haplotypes match, the male individual may be considered the father, or if the haplotypes do not match, the male individual may not be considered the father. In some embodiments, the first haplotype can be compared to two haplotypes of the male individual.

[0401] In an embodiment, when the originating tissue is a fetus, relatedness can be tested by aligning one of the sequence reads to a first region of a reference genome. The first region can include a first plurality of sites corresponding to alleles. The plurality of sites can include a threshold number of sites. The threshold number of sites can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more sites. The alleles at each of the sites in the plurality of sites can be compared to the alleles at the corresponding sites in the genome of the male individual. The comparison can be used to determine a classification of the likelihood that the male individual is the father of the fetus. If a particular number or percentage of alleles match, the male individual may be considered the father, and if fewer than the number or percentage of alleles match, the male individual may not be considered the father. The cutoff percentage can be 100%, 90%, 80% or 70%.

[0402] In some embodiments, haplotypes can be determined. The method can include, for each of the cell-free DNA molecules in the population of cell-free DNA molecules, aligning the sequence read corresponding to the cell-free DNA molecule to the reference genome. The sequence read can be identified as corresponding to a haplotype present in a female. The haplotypes present in the female can be known from genotyping the female. In some embodiments, the haplotype of the female can be known by analyzing the concentration of DNA fragments of the haplotype in a biological sample from the female. The originating tissue being a fetus can be determined using a methylation pattern. The haplotype can be determined to be a maternally inherited fetal haplotype.

[0403] Inheritance of a haplotype can be determined using methylation of a reference tissue rather than using a known methylation profile such as a methylation profile associated with imprinted loci. Matching or similarity scores of methylation patterns to reference patterns can exclude parents based on inheritance thereof, regardless of whether a given allele or locus is methylated.

[0404] A haplotype can be identified as carrying a pathogenic genetic mutation or variant. Identifying a haplotype as carrying a pathogenic genetic mutation can include identifying a genetic mutation or variant in a first sequence read. The genetic variant can include a single nucleotide difference, a deletion, or an insertion. A first methylation level can be measured in a second sequence read corresponding to a first genomic position within a first distance of the first sequence read. A second methylation level can also be measured in a third sequence read corresponding to a second genomic position within a second distance of the first sequence read. The first distance can be 100 nt, 200 nt, 300 nt, 400 nt, 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 2 knt, 5 knt, or 10 knt. The second sequence read and the third sequence read can be on the same chromosome arm as the first sequence read. The first methylation level and the second methylation level may be related to the genetic mutation or variant. The first methylation level and the second methylation level can be greater than one or two threshold levels related to the genetic mutation or variant. The threshold levels can be determined using individuals known to have or not have the genetic mutation or variant. The method can include classifying a fetus as potentially having a disease caused by a genetic mutation or variant.

[0405] A fetal-specific methylation pattern can be determined. The method can include, for each cell-free DNA molecule in the population of cell-free DNA molecules, aligning a sequence read corresponding to the cell-free DNA molecule to a reference genome. The method can include identifying the sequence read as corresponding to a region. The region can be determined by receiving a plurality of fetal sequence reads corresponding to a plurality of fetal DNA molecules from fetal tissue. The method can include receiving a plurality of maternal sequence reads corresponding to a plurality of maternal DNA molecules. The method can include determining a fetal methylation state at each methylation site among a plurality of methylation sites within the region for each fetal sequence read among the plurality of fetal sequence reads. The method can include determining a maternal methylation state at each methylation site among the plurality of methylation sites for each maternal sequence read among the plurality of maternal sequence reads.

[0406] A method for determining fetal-specific methylation patterns can include determining a value of a parameter that characterizes the amount of sites where the fetal methylation status is different from the maternal methylation status. The method can include comparing the value of the parameter with a threshold. The parameter can be the proportion of sites that are different between fetal DNA molecules and maternal DNA molecules. The proportion can be the mismatch fraction described herein. The threshold can indicate the lowest level of the mismatch fraction and can be 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater. In some embodiments, the threshold can represent the average mismatch fraction of maternal or fetal DNA molecules. The method can include determining that the value of the parameter exceeds the threshold. In some embodiments, a specific percentage of the values of the parameter for maternal or fetal DNA molecules can be required to exceed the threshold. For example, the percentage can be 50%, 60%, 70%, 80%, 90% or greater. In some embodiments, a specific percentage of the fetal DNA molecules corresponding to the region can be required to have a fetal-specific methylation pattern. For example, the percentage can be 40%, 50%, 60%, 70%, 80% or greater. This method can be used with Figure 25 the method described.

[0407] The method can include enriching a biological sample from the source tissue to obtain cell-free DNA molecules. Enriching the biological sample can include selecting and amplifying the set of cell-free DNA molecules. Enrichment can include size-based selection as described herein. In some embodiments, enrichment can include selection based on methylation patterns. For example, methylation-CpG binding domain (MBD)-based capture and sequencing can be used. The cell-free DNA can be incubated with a tagged MBD protein that can bind methylated cytosine. Subsequently, the protein-DNA complex can be precipitated using antibody-conjugated magnetic beads. DNA molecules with more methylated CpG sites can be preferentially enriched for downstream analysis.

[0408] III. Variation of long cell-free DNA fragments with gestational age

[0409] The amount of long cell-free DNA fragments can vary with gestational age. Long cell-free DNA fragments can be used to determine gestational age. Additionally, long cell-free DNA fragments can be more abundant in specific end motifs compared to shorter cell-free DNA fragments, and the relative amount of the specific end motif can vary with gestational age. The amount of the end motif can also be used to determine gestational age. The deviation of the gestational age determined using long cell-free DNA fragments from the gestational age determined via other clinical techniques can indicate pregnancy-related disorders. In some embodiments, long cell-free DNA fragments can be used to determine the likelihood of pregnancy-related disorders without having to determine gestational age.

[0410] A. Size analysis for fetal DNA and maternal DNA

[0411] Plasma DNA from two pregnant women in the first trimester (gestational age: 13 weeks), plasma DNA from two pregnant women in the second trimester (gestational age: 21 - 22 weeks), and plasma DNA from five pregnant women in the third trimester (gestational age: 38 weeks) were sequenced using single molecule real-time (SMRT) sequencing (PacBio). For each case, a median of 176 million (range: 49 million - 685 million) subreads were obtained, of which 128 million (range: 35 million - 507 million) subreads could be aligned to the human reference genome (hg19). Each molecule in the SMRT cell was sequenced an average of 107 times. A median of 965,308 (range: 251,686 - 2,871,525) high-quality circular consensus sequencing (CCS) reads, defined as CCS reads with at least 3 subreads, were available for downstream analysis.

[0412] All sequenced molecules from samples obtained from each trimester were combined for size analysis. There were a total of 1.94 million, 5.09 million, and 4.45 million cell-free DNA molecules in the maternal plasma samples from the first trimester, second trimester, and third trimester, respectively.

[0413] Figure 41A and Figure 41B show the size distribution of cell-free DNA molecules in the 0 kb to 5 kb size range in maternal plasma samples from the first trimester, second trimester, and third trimester. The x-axis shows the size. The y-axis shows the frequency. For Figure 41A , the size distribution on the y-axis in the 0 kb to 5 kb range is plotted on a linear scale, and for Figure 41B , the size distribution on the y-axis in the 0 kb to 5 kb range is plotted on a logarithmic scale. Plasma DNA from all three trimesters shows the expected main peak below 166 bp as shown in Figure 41A , and shows a series of main peaks in a periodic pattern extending to molecules in the 1 kb and 2 kb ranges as shown in Figure 41B .

[0414] Figure 42Table showing the proportion of long plasma DNA molecules in different trimesters of pregnancy. The first column shows the gestational age associated with the plasma sample. The second column shows the proportion of DNA molecules longer than 500 bp. The third column shows the proportion of DNA molecules longer than 1 kb. The frequency of plasma DNA molecules of 500 bp or greater than 500 bp increases in the third trimester of pregnancy compared to the first and second trimesters. The proportions of long plasma DNA molecules greater than 500 bp in the first, second, and third trimesters of pregnancy are 15.8%, 16.1%, and 32.3%, respectively. The proportions of long plasma DNA molecules greater than 1 kb in the first, second, and third trimesters of pregnancy are 11.3%, 10.6%, and 21.4%, respectively. When the long free DNA molecules in maternal plasma in the first and second trimesters of pregnancy show similar proportions, the proportion of long DNA molecules in maternal plasma in the third trimester of pregnancy is approximately twice that of the aforementioned long DNA molecules.

[0415] For all maternal plasma DNA samples analyzed in this disclosure, the DNA extracted from their paired maternal buffy coats and fetal samples was genotyped on an Infinium Omni2.5Exome-8 Beadchip on an iScan system (Illumina) as an array-based hybridization genotyping method. Fetal samples were obtained by chorionic villus sampling, amniocentesis, or placental sampling, depending on whether the case was from the first, second, or third trimester of pregnancy, respectively. For each case, a median of 203,647 informative single nucleotide polymorphisms (SNPs) were identified where the mother was homozygous and the fetus was heterozygous. When the sequenced DNA molecules from each trimester for all cases were combined, we identified a total of 1,362, 2,984, and 6,082 DNA molecules covering fetal-specific alleles in the first, second, and third trimesters of pregnancy, respectively. On the other hand, for each case, a median of 210,820 informative SNPs were identified where the mother was heterozygous and the fetus was homozygous. We identified a total of 30,574, 65,258, and 78,346 DNA molecules covering maternal-specific alleles in the first, second, and third trimesters of pregnancy, respectively. Among all maternal plasma samples, the median fetal DNA fraction determined from the sequencing data of DNA molecules ≤600 bp was 15.6% (range 7.6% - 26.7%).

[0416] Figure 43A and Figure 43B show the size distribution of DNA molecules covering fetal-specific alleles in maternal plasma from the first, second, and third trimesters of pregnancy. The x-axis shows the size. The y-axis shows the frequency. For Figure 43A , the size distribution on the y-axis is plotted on a linear scale in the range from 0 kb to 3 kb, and for Figure 43B, plot the size distribution of the y-axis on a logarithmic scale in the range of 0 kb to 3 kb.

[0417] Figure 44A and Figure 44B Show the size distribution of DNA molecules covering maternal-specific alleles from maternal plasma in the first trimester, second trimester, and third trimester of pregnancy. The x-axis shows the size. The y-axis shows the frequency. For Figure 44A , plot the size distribution of the y-axis on a linear scale in the range of 0 kb to 3 kb, and for [[ID=26 , plot the size distribution of the y-axis on a logarithmic scale in the range of 0 kb to 3 kb.

[0418] As ​ shown, plasma DNA molecules covering fetal-specific alleles and plasma DNA molecules covering maternal-specific alleles from all three trimesters exhibit a long-tailed distribution, indicating the presence of long DNA molecules from both fetal and maternal sources in all three trimesters.

[0419] ​ is a table of the proportions of long fetal and maternal plasma DNA molecules for different trimesters of pregnancy. The first column shows the gestational age associated with the plasma sample. The second column shows the proportion of fetal DNA molecules longer than 500 bp. The third column shows the proportion of maternal DNA molecules longer than 500 bp. The fourth column shows the proportion of fetal DNA molecules longer than 1 kb. The fifth column shows the proportion of maternal DNA molecules longer than 1 kb. Among the pool of DNA molecules in maternal plasma, compared to DNA molecules covering maternal-specific alleles, DNA molecules covering fetal-specific alleles (of placental origin) have a smaller proportion of long DNA molecules. The proportions of long plasma DNA molecules covering fetal-specific alleles with a size greater than 500 bp in the first trimester, second trimester, and third trimester of pregnancy are 19.8%, 23.2%, and 31.7%, respectively. The proportions of long plasma DNA molecules covering fetal-specific alleles with a size greater than 1 kb in the first trimester, second trimester, and third trimester of pregnancy are 15.2%, 16.5%, and 19.9%, respectively.

[0420] Regardless of the fact that there is a smaller proportion of long plasma DNA molecules in maternal plasma in the first trimester and second trimester compared to the third trimester of pregnancy and that fetal DNA molecules contain fewer long DNA molecules in all three trimesters, the methods described in our previous publications and this present disclosure allow us to analyze a significant proportion of long plasma DNA molecules that were previously impossible to analyze with short-read sequencing technologies. Additionally, we can use different size selection strategies including but not limited to electrophoresis-based, chromatography-based, and bead-based methods to enrich long DNA fragments in plasma samples.

[0421] ​ ,​ and ​ show the proportional graphs of fetal-specific plasma DNA fragments in specific size ranges during different pregnancy trimesters. The gestational age of the evaluated pregnancy cases was verified by dating ultrasound for determining the pregnancy trimester. ​ Show the results of DNA fragments less than or equal to 150 bp. Figure 46B Show the results of DNA fragments from 150 to 600 bp. Figure 46C Show the results of DNA fragments greater than or equal to 600 bp. The graphs have the proportion of fetal-specific fragments on the y-axis and gestational age on the x-axis. As shown in the graphs, compared with the proportion of fetal-specific fragments in the range of 150 bp to 600 bp ( Figure 46B ), both the proportion of fetal-specific fragments shorter than 150 bp ( Figure 46A ) and the proportion of fetal-specific fragments longer than 600 bp ( Figure 46C ) will achieve specific discriminatory power to distinguish late pregnancy samples from early and mid-pregnancy samples. The proportion of fetal-specific fragments longer than 600 bp can provide the best discriminatory power. This conclusion is demonstrated by the fact that when using the proportion of fetal-specific fragments shorter than 150 bp, the absolute minimum distance between the combined group of late pregnancy and early and mid-pregnancy is 0.38, while when using the proportion of fetal-specific fragments greater than 600 bp, the corresponding value is 3.76. These results indicate that the use of long DNA molecules reflects the pathophysiological state, which is superior to the use of short DNA molecules.

[0422] B. Plasma DNA End Analysis

[0423] In addition to size, for each sequenced DNA molecule, we determined the first nucleotide at the 5'-end of both the Watson strand and the Crick strand. This analysis consists of 4 types of ends, namely the A-end, C-end, G-end, and T-end. The percentage of plasma DNA molecules with specific ends obtained from maternal plasma samples from each pregnancy trimester was calculated. The percentages of the A-end, C-end, G-end, and T-end for each fragment size were further analyzed.

[0424] Figure 47A 、 Figure 47B and Figure 47C Show the proportional graphs of the base content at the 5'-end of free DNA molecules from maternal plasma in early, mid-, and late pregnancy across the fragment size range from 0 kb to 3 kb. Figure 47A Show maternal plasma in early pregnancy. Figure 47B Show maternal plasma in mid-pregnancy. Figure 47CMaternal plasma in the third trimester of pregnancy is shown. The base content in percentage is shown on the y-axis. The fragment size in base pairs is shown on the x-axis. As seen in the schema, the C-terminus is overrepresented across many size ranges (mostly less than 1 kb) and varies according to different size ranges for first trimester, second trimester, and third trimester samples. The plasma DNA end pattern in the third trimester samples seems to be different from that in the first trimester and second trimester samples. For example, the T- and G-end curves are mixed at sizes in the range of 105 bp to 172 bp, but they diverge in the first trimester and second trimester samples. For longer fragments (e.g., greater than about 1 kb), the C-terminal fragments are not the most abundant. The G-terminal fragments exceed the C-terminal fragments at about 1 kb, and then the A-terminal fragments become more abundant than the G-terminal fragments at about 2 kb.

[0425] Figure 48 Table of terminal nucleotide base ratios among short and long free DNA molecules from maternal plasma in the first trimester, second trimester, and third trimester of pregnancy. The first column shows the base at the end of the molecule. The second column shows the expected ratio points and species. The third column shows the ratio of terminal species among fragments ≤500 bp in maternal plasma in the first trimester of pregnancy. The fourth column shows the ratio of terminal species among fragments >500 bp in maternal plasma in the first trimester of pregnancy. The fifth and sixth columns are similar to the third and fourth columns respectively, except that maternal plasma in the second trimester replaces maternal plasma in the first trimester. The seventh and eighth columns are similar to the third and fourth columns respectively, except that maternal plasma in the third trimester replaces maternal plasma in the first trimester.

[0426] If the fragmentation of free DNA were completely random, the terminal nucleotide base ratios should reflect the composition of the human genome, which is 29.5% A, 29.5% T, 20.5% C, and 20.5% G, as Figure 48 shown in the second column of. In contrast to random fragmentation, the 5'-ends of short free DNA molecules ≤500 bp show a substantial overrepresentation of the C-terminus (30.4%, 30.4%, and 31.3% for maternal plasma in the first trimester, second trimester, and third trimester respectively), a slight overrepresentation of the G-terminus (27.4%, 26.9%, and 25.3% for the first trimester, second trimester, and third trimester respectively), an underrepresentation of the A-terminus (19.8%, 19.4%, and 19.3% for the first trimester, second trimester, and third trimester respectively), and an underrepresentation of the T-terminus (22.4%, 23.3%, and 24.1% for the first trimester, second trimester, and third trimester respectively).

[0427] However, when compared to short cell-free DNA molecules, long cell-free DNA molecules >500bp show a considerable increase in the proportion of A ends (29.6%, 26.0%, and 26.7% for first trimester, second trimester, and third trimester maternal plasma, respectively), a slight increase in the proportion of G ends (31.0%, 29.5%, and 29.9% for first trimester, second trimester, and third trimester, respectively), a considerable decrease in the proportion of T ends (13.9%, 16.9%, and 16.4% for first trimester, second trimester, and third trimester, respectively), and a slight decrease in the proportion of C ends (25.5%, 27.5%, and 27.1% for first trimester, second trimester, and third trimester, respectively).

[0428] Figure 49 Table of terminal nucleotide base proportions among short and long cell-free DNA molecules covering fetal-specific alleles from first trimester, second trimester, and third trimester maternal plasma. Figure 50 Table of terminal nucleotide base proportions among short and long cell-free DNA molecules covering maternal-specific alleles from first trimester, second trimester, and third trimester maternal plasma. The first column shows the base at the end of the molecule. The second column shows the expected proportion point and species. The third column shows the proportion of terminal species among fragments ≤500bp in first trimester maternal plasma. The fourth column shows the proportion of terminal species among fragments >500bp in first trimester maternal plasma. The fifth and sixth columns are similar to the third and fourth columns, respectively, except that second trimester maternal plasma replaces first trimester maternal plasma. The seventh and eighth columns are similar to the third and fourth columns, respectively, except that third trimester maternal plasma replaces first trimester maternal plasma. Figure 49 and Figure 50 show that the said differences in terminal nucleotide base proportions among short and long cell-free DNA molecules remain the same even when we examine DNA molecules covering fetal-specific alleles and DNA molecules covering maternal-specific alleles separately.

[0429] Figure 51 Illustrates hierarchical clustering analysis of short and long plasma cell-free DNA molecules using 256 4-mer end motifs. Each column indicates the samples used to analyze end motif frequencies based on short fragments (represented by cyan in the first column) and long fragments (represented by yellow in the first column), respectively. Starting from the second column, each column indicates the type of end motif. Based on the column-normalized frequency (z-score) (i.e., the number of standard deviations above or below the average frequency in the entire sample), the end motif frequencies are presented in a series of color gradients. The redder the color, the higher the end motif frequency, and the bluer the color, the lower the end motif frequency.

[0430] InFigure 51 In this, we characterized short and long cell-free DNA molecules by analyzing the tetramer end-motif profiles thereof. We determined, for each sequenced DNA molecule, the first 4-nucleotide sequence (tetramer motif) at the 5'-end of both the Watson strand and the Crick strand. For each maternal plasma sample, the plasma DNA end-motif frequencies of short plasma DNA molecules (≤500 bp) and long plasma DNA molecules (>500 bp) were calculated separately. Hierarchical cluster analysis based on 256 tetramer end-motif frequencies showed that the end-motif profiles of long DNA molecules in different maternal plasma samples formed clusters distinct from those of short DNA molecules. These results indicate that long DNA and short DNA have different fragmentation characteristics. In the examples, we will use the relative perturbation of these end-motifs between long DNA molecules and short DNA molecules to indicate the contribution of cell-free DNA from cell death pathways such as, but not limited to, apoptosis and necrosis. Increased activity from these cell death pathways may be associated with pregnancy-related disorders and other disorders.

[0431] Figure 52A and Figure 52B shows principal component analysis (PCA) using the tetramer end-motif profiles for classification analysis. Figure 52A shows short cell-free DNA molecules (≤500 bp) from different trimesters of pregnancy. Figure 52B shows long cell-free DNA molecules (>500 bp) from maternal plasma samples from different trimesters of pregnancy. The percentages in parentheses on the X-axis and Y-axis represent the amount of variability accounted for by the corresponding components. Each blue dot represents a maternal plasma sample in the first trimester of pregnancy. Each yellow dot represents a maternal plasma sample in the second trimester of pregnancy. Each red dot represents a maternal plasma sample in the third trimester of pregnancy. The ellipses represent the 95% confidence level for grouping the data points from a particular trimester. Compared to short cell-free DNA molecules ( Figure 52A )(also described in U.S. Application No. 15 / 787,050), the tetramer end-motif profiles of long cell-free DNA molecules ( Figure 52B ) show a clearer separation between maternal plasma samples in the first trimester, second trimester, and third trimester of pregnancy. In the examples, we can utilize the end-motif profiles of individual long plasma DNA molecules or in combination with other maternal plasma DNA characteristics including, but not limited to, methylation level and size for molecular gestational age assessment.

[0432] For example, we use a neural network to train a model for predicting gestational age based on 256 end motifs, the total degree of methylation, and the proportion of fragments with a size ≥ 600 bp. The output variables are 1, 2, and 3, representing the first trimester, second trimester, and third trimester of pregnancy. The input variables include 256 end motifs, the total degree of methylation, and the proportion of fragments with a size ≥ 600 bp. We use the leave-one-out method to evaluate the efficacy of predicting gestational age. For a dataset including 9 samples, the leave-one-out method is performed in such a way that one sample is selected as the test sample and the remaining 8 samples are used to train the model based on the neural network. Based on the established model, such a test sample is determined to be 1, 2, or 3. Subsequently, we repeat this method for the other samples that have not been tested. For such training and testing processes, we repeat a total of 9 times. By comparing those test results with the clinical information regarding gestational age, 8 out of 9 samples (89%) are properly predicted in terms of gestational age. In another embodiment, the analysis can be performed using, for example but not limited to, Bayes' theorem, logistic regression, multiple regression, and support vector machines, random forest analysis, classification and regression trees (CART), and the K-nearest neighbor algorithm.

[0433] Next, all the sequenced molecules from the samples obtained from each trimester are combined together for downstream end motif analysis. They are ranked according to the frequencies of 256 end motifs among short plasma DNA molecules and long plasma DNA molecules.

[0434] Figures 53 to 58 Table of the 25 end motifs with the highest frequencies for DNA fragments of a specific length (shorter or longer than 500 bp) and different trimesters of pregnancy. Figure 53 、 Figure 54 and Figure 55 Table of end motifs sorted by end motif rank in short fragments (< 500 bp). In Figures 53 to 55 , the first column shows the end motif. The second column shows the motif frequency rank in short fragments. The third column shows the motif frequency rank in long fragments. The fourth column shows the motif frequency in short fragments. The fifth column shows the motif frequency in long fragments. The sixth column shows the fold change (motif frequency in short fragments divided by motif frequency in long fragments).

[0435] Figure 56 、 Figure 57 and Figure 58 Table of end motifs sorted by end motif rank in long fragments (> 500 bp). In Figures 56 to 58 , the first column shows the end motif. The second column shows the motif frequency rank in long fragments. The third column shows the motif frequency rank in short fragments. The fourth column shows the motif frequency in long fragments. The fifth column shows the motif frequency in short fragments. The sixth column shows the fold change (motif frequency in long fragments divided by motif frequency in short fragments).

[0436] Figure 53 and Figure 56 from early pregnancy samples. Figure 54 and Figure 57 from mid-pregnancy samples. Figure 55 and Figure 58 from late pregnancy samples.

[0437] Among the top 25 most frequent terminal motifs in short plasma DNA molecules, 11 of them start with the CC dinucleotide. In maternal plasma in early pregnancy, mid-pregnancy, and late pregnancy, the terminal motifs starting with CC together account for 14.66%, 14.66%, and 15.13% of the short plasma DNA terminal motifs, respectively. Among the top 25 most frequent terminal motifs in long plasma DNA molecules, the tetramer motif ending with the TT dinucleotide accounts for 9 of them in maternal plasma in mid-pregnancy and late pregnancy and 10 of them in maternal plasma in early pregnancy.

[0438] We determined the dinucleotide sequences of the third nucleotide (X) and the fourth nucleotide (Y) at the 5' end of both the Watson strand and the Crick strand for each sequenced DNA molecule. X and Y can be one of the four nucleotide bases in DNA. There are 16 possible NNXY motifs, namely NNAA, NNAT, NNAG, NNAC, NNTA, NNTT, NNTG, NNTC, NNGA, NNGT, NNGG, NNGC, NNCA, NNCT, NNCG, and NNCC.

[0439] Figure 59A 、 Figure 59B and Figure 59C Show the motif frequency scatter plots of the 16 NNXY motifs in short plasma DNA molecules and long plasma DNA molecules. Figure 59A Show the results of early pregnancy. Figure 59B Show the results of mid-pregnancy. Figure 59C Show the results of late pregnancy. The motif frequencies of the long fragments are shown on the y-axis. The motif frequencies of the short fragments are shown on the x-axis. Each circle represents one of the 16 NNXY motifs. The dotted line pairs in each scatter plot refer to a 1.5-fold increase (upper line) and decrease (lower line) in the motif frequency in long plasma DNA molecules (>500 bp) compared to short plasma DNA molecules (≤500 bp). Circles located outside the shaded area represent motifs with a change multiple >1.5.

[0440] In all three trimesters (Figure 11), when the ends of short plasma DNA molecules showed a high frequency of a 4-mer motif starting with the CC dinucleotide (CCNN) (Jiang et al., Cancer Discov 2020;10(5):664-673; Chan et al., Am J Hum Genet 2020;107(5):882-894), the ends of long plasma DNA molecules showed a >1.5-fold increase in the frequency of a 4-mer motif ending with TT (NNTT). In maternal plasma in the first, second, and third trimesters of pregnancy, the NNTT motif accounted for 18.94%, 15.22%, and 15.30% of the long plasma DNA end motifs, respectively. Conversely, in maternal plasma in the first, second, and third trimesters of pregnancy, the NNTT motif accounted for only 9.53%, 9.29%, and 8.91% of the short plasma DNA end motifs, respectively.

[0441] As previously reported by Han et al., enrichment of the A-ended fragments >150 bp for free DNA newly released from dying cells into plasma was performed. DNA fragmentation factor β (DFFB), which is the major intracellular nuclease involved in DNA fragmentation during apoptosis, was found to be responsible for generating the fragments (Han et al., Am J Hum Genet 2020;106:202-214). In the present disclosure, we have shown that long free DNA molecules >500 bp were also enriched for A-ended fragments, indicating that DFFB may also be responsible for generating these fragments. In normal pregnancy, trophoblast apoptosis increases with gestational development (Sharp et al., Am J Reprod Immuno 2010;64(3):159-69). Indeed, our finding that the proportion of long DNA molecules covering fetal-specific alleles increases with gestational development may reflect the increase in trophoblast apoptosis with gestational development.

[0442] In embodiments, we can use the methods described herein to analyze long cell-free DNA molecules in maternal plasma for the prediction, screening, and monitoring of the development of placenta-related pregnancy complications, including but not limited to preeclampsia, intrauterine growth restriction (IUGR), preterm birth, and gestational trophoblastic disease. Increased levels of trophoblast apoptosis have been reported in placenta-related pregnancy complications such as preeclampsia (Leung et al., Am J Obstet Gynecol 2001; 184: 1249-1250), IUGR (Smith et al., Am J Obstet Gynecol 1997; 177: 1395-1401; Levy et al., Am J Obstet Gynecol 2002; 186: 1056-1061), and gestational trophoblastic disease. In addition, increased fetal DNA content in maternal plasma has been reported in the following: preeclampsia (Lo et al., Clin Chem 1999; 45(2): 184-8; Smid et al., Ann N Y Acad Sci 2001; 945: 132-7), IUGR (Sekizawa et al., Am J Obstet Gynecol 2003; 188: 480-4), and preterm birth (Leung et al., Lancet 1998; 352(9144): 1904-5). We hypothesize that in placenta-related pregnancy complications, due to increased placental apoptosis, the proportion of long cell-free DNA molecules of placental origin in maternal plasma samples increases. Thus, long cell-free DNA molecules of placental origin themselves, as well as long DNA signatures including but not limited to A-end fragments and NNTT motifs, may serve as biomarkers of placental apoptosis.

[0443] Although single nucleotide motifs and 4-nucleotide motifs were used in the above analysis, in other embodiments, motifs of other lengths such as 2, 3, 5, 6, 7, 8, 9, 10, or longer can be used.

[0444] C. Example Methods

[0445] Long cell-free DNA fragments can be used to determine the gestational age of a female carrying a fetus. The amount of long cell-free DNA fragments varies with gestational age and can be used to determine gestational age. The end motifs of the cell-free DNA fragments also vary with gestational age and can be used to determine gestational age. When the gestational age determined using long cell-free DNA fragments significantly deviates from the gestational age determined via other clinical techniques, the pregnant female and / or the fetus can be considered to have a pregnancy-related disorder. In some embodiments, it may not be necessary to determine gestational age to determine the likelihood of a pregnancy-related disorder.

[0446] 1. Gestational Age

[0447] Figure 60A method 6000 for analyzing a biological sample obtained from a female carrying a fetus is shown. The gestational age can be determined and used to classify the likelihood of pregnancy-related disorders. The biological sample can contain multiple free DNA molecules from the fetus and the female.

[0448] Sequence reads corresponding to the multiple free DNA molecules can be received. In some embodiments, sequencing for obtaining the sequence reads can be performed.

[0449] At block 6020, the sizes of the multiple free DNA molecules can be measured. The sizes can be measured in a manner similar to the manner described by Figure 21 The sizes can be measured using the sequence reads.

[0450] At block 6030, a first quantity of free DNA molecules having a size greater than a cutoff value can be measured. The quantity can be the number, total length, or mass of the free DNA molecules.

[0451] At block 6040, a value of a normalization parameter can be generated using the first quantity. The value of the normalization parameter can be the first quantity normalized by the total number of free DNA molecules, by the number of free DNA molecules from the fetus or the mother, or by the number of DNA molecules from a specific region. For example, as described by Figures 46A - C the normalization parameter can be the fetal-specific fragment ratio.

[0452] At block 6050, the value of the normalization parameter can be compared with one or more calibration data points. Each calibration data point can specify a gestational age corresponding to a calibration value of the normalization parameter. For example, a gestational age at a specific trimester or a specific number of weeks can correspond to a calibration value of the normalization parameter. The one or more calibration data points can be determined from a plurality of calibration samples having a known gestational age and containing free DNA molecules having a size greater than the cutoff value. In some embodiments, the calibration data points are determined from a functionally related gestational age having a value of the normalization parameter.

[0453] At block 6060, the gestational age can be determined using the comparison. The gestational age can be considered the age corresponding to the calibration value closest to the value of the normalization parameter. In some embodiments, the gestational age can be considered the highest age corresponding to a calibration value below the value of the normalization parameter.

[0454] The method can further include determining a reference gestational age of the fetus using ultrasound or the date of the female's last menstrual period. The method can also include comparing the gestational age with the reference gestational age. The method can further include classifying the likelihood of pregnancy-related disorders using the comparison of the gestational age and the reference gestational age. For example, the difference between the gestational age and the reference gestational age can indicate a pregnancy-related disorder. The difference can be a difference in gestational age at different trimesters or a difference in the smallest number of weeks (e.g., 1, 2, 3, 4, 5, 6, 7 weeks or more).

[0455] The method may further comprise using end motifs. For example, the method may comprise determining a first subsequence corresponding to at least one end of a free DNA molecule having a size greater than a cutoff value. A first quantity may pertain to free DNA molecules having a size greater than the cutoff value and having the first subsequence at one or more ends of the corresponding free DNA molecule. The first subsequence may be or comprise 1, 2, 3, 4, 5, or 6 nucleotides. As described Figure 52A and Figure 52B , end motifs may be used to determine gestational age via PCA analysis. Calibration samples having different end motifs and known gestational ages and subjected to PCA analysis may be used. Other classification and regression algorithms such as linear discriminant analysis, logistic regression, support vector machines, linear regression, non-linear regression, etc. may be used on the end motifs. The classification and regression algorithms may correlate gestational age with specific end motifs and / or specific sized fragments.

[0456] The end motif may be any motif discussed in FIGS. 47 - 59 or Figure 94 . The end motif rank or frequency may be compared to the end motif rank or frequency in a calibration sample from an individual having a known gestational age. Subsequently, the end motif rank or frequency may be used to determine gestational age. An end motif present at a rank or frequency deviating from the rank or frequency determined from a reference sample having the same gestational age may indicate a pregnancy-related disorder.

[0457] Generating a value of the normalization parameter may comprise (a) normalizing the first quantity by the total amount of free DNA molecules having a size greater than the cutoff value; (b) normalizing the first quantity by a second quantity of free DNA molecules having a size greater than the cutoff value and ending with a second subsequence different from the first subsequence; or (c) normalizing the first quantity by a third quantity of free DNA molecules having a size less than the cutoff value.

[0458] 2. Pregnancy-Related Disorders

[0459] Figure 61 Disclosed is a method 6100 of analyzing a biological sample obtained from a female carrying a fetus. Embodiments may include classifying the likelihood of a pregnancy-related disorder without having to determine gestational age. The biological sample may comprise a plurality of free DNA molecules from the fetus and the female.

[0460] Sequence reads corresponding to the plurality of free DNA molecules may be received. In some embodiments, sequencing for obtaining the sequence reads may be performed.

[0461] At block 6120, the sizes of the plurality of free DNA molecules may be measured. The sizes may be obtained in a manner similar to that described Figure 21 . Measuring the sizes may use the received sequence reads.

[0462] At block 6130, a first quantity of cell-free DNA molecules having a size greater than a cut-off value can be measured. The cut-off value can be greater than or equal to 200 nt. The cut-off value can be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cut-off value can be any cut-off value described herein for long cell-free DNA molecules. The first quantity can be a number or a frequency.

[0463] At block 6140, a first value of a normalization parameter can be generated using the first quantity. Generating the value of the normalization parameter can include measuring a second quantity of cell-free DNA molecules having a size less than the cut-off value; and calculating a ratio of the first quantity to the second quantity. The cut-off value can be a first cut-off value. A second cut-off value can be less than the first cut-off value. The second quantity can include cell-free DNA molecules having a size less than the second cut-off value or the second quantity can include all cell-free DNA molecules among a plurality of cell-free DNA molecules. The normalization parameter can be a measure of the frequency of long cell-free DNA molecules.

[0464] At block 6150, a second value corresponding to an expected value of the normalization parameter for a healthy pregnancy can be obtained. The second value can depend on the gestational age of the fetus. The second value can be an expected value. In some embodiments, the second value can be a cut-off value that is distinct from an outlier value.

[0465] Obtaining the second value can include obtaining the second value from a calibration table of measurement results for pregnant women having calibration values of the normalization parameter. The calibration table can be generated by obtaining a first table of gestational ages having individual measurement results for pregnant women. A second table of gestational ages having calibration values of the normalization parameter can be obtained. The data in the first table and the second table can be from the same individual or different individuals. A calibration table of measurement results having calibration values can be generated from the first table and the second table. The calibration table can include a function of the calibration values of the measurement results.

[0466] Measurement results for a pregnant female individual can be the time since the last menstrual period or image (e.g., ultrasound) characteristics of the pregnant female individual. The measurement result for the pregnant female individual can be an image characteristic of the pregnant female individual. For example, the image characteristic can include the length, size, appearance, or anatomy of the fetus of the female individual. The characteristic can include biometric measurements such as crown-rump length or femur length. The appearance of a particular organ can be used, including the appearance of a four-chamber heart or vertebrae on the spinal cord. The gestational age can be determined by a practitioner from an ultrasound image (e.g., Committee on Obstetric Practice et al., "Methods for estimating the due date", Committee Opinion, No. 700, May 2017).

[0467] In some embodiments, the machine learning model can correlate one or more calibration data points with the image characteristics. The model can be trained by receiving a plurality of training images. Each training image can be from a female individual known to be free of pregnancy-related disorders or known not to have pregnancy-related disorders. The female individual can have a range of gestational ages. The training can include storing a plurality of training samples from the female individual. Each training sample can include a known value of a standardized parameter associated with the training image. The model can be trained by optimizing the parameters of the model using the plurality of training samples based on the output of the model for matching or non-matching images with known values of the standardized parameter. The output of the model can specify a value of the standardized parameter corresponding to the image. A second value of the standardized parameter can be generated by inputting the female image into the machine learning model.

[0468] At block 6160, a deviation between the first value of the standardized parameter and the second value of the standardized parameter can be determined. The deviation can be a separation value.

[0469] At block 6170, the deviation can be used to determine a classification of the likelihood of a pregnancy-related disorder. The pregnancy-related disorder can be likely when the deviation exceeds a threshold. The threshold can indicate a statistically significant difference. The threshold can indicate a difference of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100%.

[0470] Pregnancy-related disorders can include preeclampsia, intrauterine growth restriction, invasive placentation, preterm birth, neonatal hemolytic disease, placental insufficiency, fetal hydrops, fetal anomalies, hemolysis, elevated liver enzymes and low platelet count (HELLP) syndrome, or systemic lupus erythematosus.

[0471] IV. Dimension and Endpoint Analysis for Pregnancy-Related Disorders

[0472] The sizing and / or end analysis of long DNA molecules is used to determine the likelihood of preeclampsia. The method can also be applied to other pregnancy-related disorders. DNA extracted from maternal plasma samples of four pregnant women diagnosed with preeclampsia was subjected to single molecule real-time (SMRT) sequencing (PacBio).

[0473] Figure 62 Table showing the clinical information of four preeclampsia cases. The first column shows the case number. The second column shows the gestational age in weeks at the time of blood sampling cessation. The third column shows the fetal sex. The fourth column shows the clinical information regarding preeclampsia (PET).

[0474] M12804 is a case of severe preeclampsia (PET) with pre-existing IgA nephropathy. M12873 is a case of chronic hypertension with superimposed mild PET. M12876 is a case of severe late-onset PET. M12903 is a case of severe late-onset PET with intrauterine growth restriction (IUGR). Five maternal plasma samples from normotensive late pregnancy were used as controls for subsequent analysis in this disclosure.

[0475] For the four preeclampsia and the five normotensive late pregnancy maternal plasma DNA samples analyzed in this disclosure, DNA extracted from their paired maternal buffy coats and placental samples was genotyped on an Infinium Omni2.5Exome-8 Beadchip on an iScan system (Illumina).

[0476] The plasma DNA concentration of each sample was quantified by Qubit dsDNA high-sensitivity assay performed with a Qubit fluorometer (ThermoFisher Scientific). The mean plasma DNA concentrations of the preeclampsia cases and the late pregnancy cases were 95.4 ng / mL (range 52.1 - 153.8 ng / mL) plasma and 10.7 ng / mL (range 6.4 - 19.1 ng / mL) plasma, respectively. The mean plasma DNA concentration of the preeclampsia cases was approximately 9-fold higher than that of the late pregnancy cases.

[0477] For the preeclampsia and normotensive late pregnancy maternal plasma samples, the mean fetal DNA fractions determined from sequencing data of ≤600 bp DNA molecules covering informative single nucleotide polymorphisms (SNPs) where the mother is homozygous and the fetus is heterozygous were 22.6% (range 16.6% - 25.7%) and 20.0% (range 15.6% - 26.7%), respectively.

[0478] A. Sizing analysis

[0479] Size analysis was performed on maternal plasma samples in late pregnancy with preeclampsia and normotensive pregnancy according to embodiments in the present disclosure. Figures 63A - 63D and Figures 64A - 64D show the size distributions of plasma DNA molecules from cases of preeclampsia and normotensive late pregnancy. The x-axis shows size. The y-axis shows frequency. For Figures 63A - 63D , the size distribution with the x-axis on a linear scale in the range of 0 kb to 1 kb is plotted, and for Figures 64A - 64D , the size distribution with the x-axis on a logarithmic scale in the range of 0 kb to 5 kb is plotted. Figure 63A and Figure 64A show sample M12804. Figure 63B and Figure 64B show sample M12873. Figure 63C and Figure 64C show sample M12876. Figure 63D and Figure 64D show sample M12903.

[0480] The blue line represents the size distribution of all sequenced plasma DNA molecules combined from five cases of normotensive late pregnancy. The red line represents the size distribution of sequenced plasma DNA molecules from a single preeclampsia case. In Figures 63A - 63D , the blue line is the line of the shorter peak below 200 bp and the line of the higher peak between 300 bp and 400 bp. In Figures 64A - 64D , the blue line corresponds to the line of the higher peak at 1 kb.

[0481] Generally, the plasma DNA size profile of preeclampsia patients is shorter than that of normotensive late pregnancy pregnant women, with an increase in the height of the 166-bp peak and an increase in the proportion of DNA molecules shorter than 166 bp ( Figures 63A - 63D ). These changes are more obvious in the two severe preeclampsia cases M12876 and M12903. The changes are even more drastic in the preeclampsia case M12903 accompanied by intrauterine growth restriction (IUGR).

[0482] Three of the four preeclampsia plasma samples showed a decreased proportion of long plasma DNA molecules with sizes of 200 - 5000 bp ( Figures 64B - 64D)。The proportions of long plasma DNA molecules >500 bp in M12873, M12876, and M12903 were 11.7%, 8.9%, and 4.5%, respectively, while the proportion of long plasma DNA molecules in the pooled sequencing data from five normotensive late pregnancy cases was 32.3%. Compared with the pooled sequencing data from five normotensive late pregnancy cases, plasma samples from cases of preeclampsia (PET) (M12804) associated with pre-existing IgA nephropathy showed a reduced proportion of shorter DNA molecules <2000 bp but an increased proportion of longer DNA molecules >2000 bp ( Figure 2A )。The proportion of long plasma DNA molecules in M12804 was 34.9%.

[0483] Figures 65A - 65D and Figures 66A - 66D show the size distributions of DNA molecules covering fetal-specific alleles in preeclampsia and normotensive late pregnancy maternal plasma samples. Each of Panels A to D shows different preeclampsia samples. The x-axis shows size. Figures 65A - 65D the y-axis in Figures 66A - 66D shows frequency and Figures 66A - 66D the y-axis in

[0484] shows cumulative frequency. In Figures 65A - 65D , the size ranges from 0 kb to 35 kb. Figures 66A - 66D In

[0485] Figures 67A - 67D and Figures 68A - 68D show the size distributions of DNA molecules covering fetal-specific alleles in preeclampsia and normotensive late pregnancy maternal plasma samples. Each of Panels A to D shows different preeclampsia samples. The x-axis shows size. Figures 67A - 67D the y-axis in Figures 68A - 68D shows frequency and Figures 68A - 68D the y-axis in

[0486] The blue lines in each figure represent the size distribution of all sequenced plasma DNA molecules covering maternal-specific alleles from five cases of normal blood pressure in late pregnancy. The red lines in each figure represent the size distribution of sequenced plasma DNA molecules covering maternal-specific alleles from a single case of preeclampsia. In Figure 67A the blue line is the line of the higher peak below 200 bp and the line of the higher peak between 300 bp and 400 bp. In Figures 67B - 67D the blue line is the line of the shorter peak below 200 bp. In Figure 68A the blue line corresponds to the line of the higher peak between 1000 bp and 10000 bp. In Figures 68B - 68D the blue line corresponds to the line of the lower peak between 100 bp and 1000 bp.

[0487] When compared with plasma samples from normal blood pressure in late pregnancy, plasma DNA shortening was observed in both DNA molecules covering fetal-specific alleles ( Figures 65B - 65D and Figures 66B - 66D ) and DNA molecules covering maternal-specific alleles ( Figures 67B - 67D and Figures 68B - 68D ) in three of the four preeclampsia plasma samples. The exception was case M12804 of severe PET with pre-existing IgA nephropathy, which showed an increased proportion of shorter DNA molecules less than 1 kb and a decreased proportion of longer DNA molecules greater than 1 kb among those plasma DNA molecules covering fetal-specific alleles ( Figure 65A and Figure 66A ). Indeed, the plasma DNA molecules covering maternal-specific alleles in case M12804 showed an elongated size profile ( Figure 67A and Figure 68A ).

[0488] Figure 69A and Figure 69B are graphs of the proportions of short DNA molecules covering (A) fetal-specific alleles and (B) maternal-specific alleles sequenced by PacBio SMRT sequencing in preeclampsia and normal blood pressure maternal plasma samples. The y-axis shows the proportion of short DNA fragments <150 bp. The x-axis shows normal samples and PET samples.

[0489] In the examples, the proportion of short DNA molecules was defined as the percentage of maternal plasma DNA molecules smaller than 150 bp. M12804 was excluded from this analysis because this case had pre-existing IgA nephropathy, whereas the other samples did not have pre-existing IgA nephropathy. When compared to the normotensive control plasma sample group, the preeclamptic plasma sample group showed significantly increased proportions of short DNA molecules covering fetal-specific alleles (P = 0.036, Wilcoxon rank sum test) and short DNA molecules covering maternal-specific alleles (P = 0.036, Wilcoxon rank sum test).

[0490] Figure 70A and Figure 70B Graphs of the proportions of short DNA molecules sequenced by (A) PacBio SMRT sequencing and (B) Illumina sequencing in preeclamptic and normotensive maternal plasma samples. The y-axis shows the proportion of short DNA fragments <150 bp.

[0491] In the examples, the proportion of short DNA molecules was defined as the percentage of maternal plasma DNA molecules smaller than 150 bp. M12804 was removed from this analysis because this case may have shown a different size profile compared to the other preeclamptic cases in this group due to the pre-existing IgA nephropathy in this case. When compared to the normotensive control plasma sample group (median: 12.1%; range: 8.5% - 15.8%), the preeclamptic plasma sample group showed a significantly increased proportion of short DNA molecules (median: 28.0%; range: 25.8% - 35.1%) (P = 0.036, Wilcoxon rank sum test). In contrast, in a previous group of four preeclamptic and four gestation-matched normotensive maternal plasma DNA samples that had undergone bisulfite conversion and Illumina sequencing, the proportions of short DNA molecules in preeclamptic plasma and control plasma samples were not significantly different (P = 0.340, Wilcoxon rank sum test)( Figure 70B ).

[0492] In some embodiments, we may use a 20% cutoff value for the proportion of short DNA molecules in maternal plasma samples to be sequenced by PacBio SMRT sequencing to determine whether a pregnant woman is at high risk or low risk of developing preeclampsia. A maternal plasma sample with a proportion of short DNA molecules higher than 20% will be determined to be at high risk of developing preeclampsia, while a maternal plasma sample with a proportion of short DNA molecules lower than 20% will be determined to be at low risk of developing preeclampsia. In the case of using this cutoff value, both the sensitivity and the specificity are 100%. In some other embodiments, the cutoff value for the proportion of short DNA molecules used may include but is not limited to 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, etc. In another embodiment, the proportion of short DNA molecules in maternal plasma samples will be used to monitor and evaluate the severity of preeclampsia during pregnancy.

[0493] In an embodiment, the size ratio indicating the relative proportion of short DNA molecules to long DNA molecules in each sample is calculated using the following equation.

[0494]

[0495] where P(50 - 150) indicates the proportion of sequenced plasma DNA molecules with sizes in the range of 50 bp to 150 bp; and P(200 - 1000) indicates the proportion of sequenced plasma DNA molecules with sizes in the range of 200 bp to 1000 bp.

[0496] Figure 71 A size ratio plot indicating the relative proportion of short DNA molecules to long DNA molecules in preeclampsia and normotensive maternal plasma samples sequenced by PacBio SMRT sequencing. The y-axis shows the size ratio. The x-axis shows normal samples and PET samples. The preeclampsia plasma sample group shows a significantly higher size ratio when compared to the normotensive control plasma sample group (P = 0.016, Wilcoxon rank sum test).

[0497] In embodiments, we can utilize size profiles generated by long-read sequencing platforms including but not limited to PacBio SMRT sequencing and Oxford Nanopore sequencing to predict the development and severity of preeclampsia in pregnant women. In some embodiments, we can monitor the development of preeclampsia and the development of severe preeclampsia features including but not limited to liver and kidney damage by analyzing the size profiles of plasma DNA molecules. In some embodiments, the size parameters used in the analysis can include but are not limited to the proportion of short DNA molecules or long DNA molecules and the size ratio indicating the relative proportion of short DNA molecules to long DNA molecules. The cut-off values for determining short DNA and long DNA categories can include but are not limited to 150bp, 180bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 1kb, etc. The size ranges for determining the size ratio of short molecules to long molecules can include but are not limited to 50-150bp, 50-166bp, 50-200bp, 200-400bp, 200-1000bp, 200-5000bp, or other combinations.

[0498] Size end analysis can include using the method described in method 6100 in Figure 61 .

[0499] B. Fragment end analysis

[0500] Fragment end analysis is performed on late pregnancy maternal plasma samples of preeclampsia and normotensive pregnancies according to embodiments in the present disclosure. The first nucleotide at the 5' end of both the Watson strand and the Crick strand is determined for each sequenced plasma DNA molecule. The proportions of T-end, C-end, A-end, and G-end fragments are determined for each plasma DNA sample.

[0501] Figures 72A - 72D Shows the different end proportions of plasma DNA molecules in preeclampsia and normotensive maternal plasma samples sequenced by PacBio SMRT sequencing. The x-axis shows late pregnancy normal samples and PET samples. The y-axis shows the established end proportions. Figure 72A Shows the T-end proportion. Figure 72B Shows the C-end proportion. Figure 72C Shows the A-end proportion. Figure 72D Shows the G-end proportion. When compared with the group of normotensive control plasma samples, the preeclampsia plasma sample group shows a significantly increased proportion of T-end plasma DNA molecules (P = 0.016, Wilcoxon rank sum test) and a significantly decreased proportion of G-end plasma DNA molecules (P = 0.016, Wilcoxon rank sum test).

[0502] Figure 73 Shown is a hierarchical cluster analysis of maternal plasma DNA samples in late pregnancy with preeclampsia and normotensive pregnancy using four types of fragment ends (the first nucleotide at the 5'-end of each strand), namely C-end, G-end, T-end, and A-end. Each column indicates a plasma DNA sample. The first column indicates which group each sample belongs to, where cyan indicates maternal plasma DNA samples in late pregnancy with normotensive pregnancy and orange indicates preeclampsia plasma DNA samples. Cyan covers the first five columns. Orange covers the last four columns.

[0503] Starting from the second column, each column indicates the type of fragment end. Based on the column-normalized frequency (z-score), which is the number of standard deviations above or below the average frequency in the entire sample, the end motif frequencies are presented in a series of color gradients. The redder the color, the higher the end motif frequency, and the bluer the color, the lower the end motif frequency. The hierarchical cluster analysis based on the frequencies of the four types of fragment ends shows that the fragment end profiles of preeclampsia plasma DNA samples form a cluster different from that of plasma DNA samples in late pregnancy with normotensive pregnancy.

[0504] In an embodiment, we can determine the dinucleotide sequence of the first nucleotide (X) and the second nucleotide (Y) at the 5'-ends of both the Watson strand and the Crick strand for each sequenced DNA molecule separately. X and Y can be one of the four nucleotide bases in DNA. There are 16 possible dinucleotide end motifs XYNN, namely AANN, ATNN, AGNN, ACNN, TANN, TTNN, TGNN, TCNN, GANN, GTNN, GGNN, GCNN, CANN, CTNN, CGNN, and CCNN. We can determine the dinucleotide sequence of the third nucleotide (X) and the fourth nucleotide (Y) at the 5'-ends of both the Watson strand and the Crick strand for each sequenced DNA molecule separately according to the embodiments in the present disclosure. There are 16 possible dinucleotide NNXY motifs. We can also determine the first tetranucleotide sequence (4-mer motif) at the 5'-ends of both the Watson strand and the Crick strand for each sequenced DNA molecule separately.

[0505] Figure 74 Shown is a hierarchical cluster analysis of maternal plasma DNA samples in late pregnancy with preeclampsia and normotensive pregnancy using 16 dinucleotide motifs XYNN (the dinucleotide sequence of the first nucleotide and the second nucleotide at the 5'-end). Figure 75 Shown is a hierarchical cluster analysis of maternal plasma DNA samples in late pregnancy with preeclampsia and normotensive pregnancy using 16 dinucleotide motifs NNXY (the dinucleotide sequence of the third nucleotide and the fourth nucleotide at the 5'-end). Figure 76Shows hierarchical clustering analysis of maternal plasma DNA samples in late pregnancy with preeclampsia and normotensive pregnancies using 256 tetranucleotide motifs (dinucleotide sequences of the first to fourth nucleotides at the 5' end).

[0506] In Figures 74 - 76 , the first column indicates which group each sample belongs to, where cyan indicates maternal plasma DNA samples in late pregnancy with normal blood pressure and orange indicates preeclampsia plasma DNA samples. Cyan covers the first five columns. Orange covers the last four columns. Starting from the second column, each column indicates the type of fragment end. Based on the column-normalized frequency (z-score) (i.e., the number of standard deviations above or below the average frequency in the entire sample), the end motif frequencies are presented in a series of color gradients. The redder the color, the higher the end motif frequency, and the bluer the color, the lower the end motif frequency.

[0507] These results indicate that plasma DNA in preeclampsia samples and non-preeclampsia samples has different fragmentation characteristics. In one embodiment, we can utilize the end motif profiles generated by long-read sequencing platforms including but not limited to PacBio SMRT sequencing and Oxford Nanopore sequencing to predict the development of preeclampsia in pregnant women. Although single nucleotide motifs, dinucleotide motifs, and tetranucleotide motifs were used in the above analysis, in other embodiments, motifs of other lengths such as 3, 5, 6, 7, 8, 9, 10, or longer can be used.

[0508] In some embodiments, we can combine fragment end analysis and origin tissue analysis to improve the efficacy of prediction, detection, and monitoring of pregnancy-related conditions including but not limited to preeclampsia. First, we can perform fragment end analysis on each maternal plasma sample to separate plasma DNA molecules into four fragment end categories, namely T-end, C-end, A-end, and G-end fragments. Subsequently, we can perform origin tissue analysis on each maternal plasma DNA sample using methylation status matching analysis, using plasma DNA molecules from each of the fragment end categories separately. The contribution ratio of different tissues in one of the fragment end categories is defined as the percentage of plasma DNA molecules in the corresponding fragment end category assigned to the corresponding tissue relative to other tissues.

[0509] We analyzed plasma DNA samples from three and five pregnant women with and without preeclampsia using single molecule real-time sequencing. We obtained median plasma fragment counts of 658,722, 889,900, 851,501, and 607,554 with A, C, G, and T ends, respectively. For fragments with an A end, we compared the methylation pattern of any fragment with at least 10 CpG sites to the reference methylation profiles of neutrophils, T cells, B cells, liver, and placenta according to the methylation state matching pathway described in the present disclosure. Plasma DNA fragments were assigned to the tissue corresponding to the highest score of methylation state matching among those tissues. Using this method, a median of 2.43% (range: 0.73% - 5.50%) of the A-end fragments were assigned to T cells (i.e., T cell contribution) among all samples analyzed. We further analyzed those fragments with C, G, and T ends in a similar manner. For those fragments with C, G, and T ends, median T cell contributions of 3.20% (range: 1.55% - 5.19%), 3.52% (range: 1.53% - 6.27%), and 2.22% (0% - 7.79%) were observed, respectively.

[0510] Figures 77A - 77D Shows T cell contributions among DNA molecules belonging to different fragment end classes, namely (A) T end, (B) C end, (C) A end, and (D) G end, in preeclampsia and normotensive maternal plasma DNA samples. The x-axis shows late normal pregnancy samples and PET samples. The y-axis shows T cell contribution as a percentage. The results show that among G-end fragments, the T cell contribution in preeclampsia plasma samples was significantly reduced compared to late normotensive pregnancy plasma samples (P = 0.036, Wilcoxon rank sum test). In an example, we can use a 3% cutoff of the T cell contribution among all G-end fragments in the maternal plasma DNA sample to determine whether a pregnant woman is at high or low risk of developing preeclampsia.

[0511] C. Example Method

[0512] Figure 78 Shows method 7800 for analyzing a biological sample obtained from a female carrying a fetus. The biological sample can include multiple cell-free DNA molecules from the fetus and the female. The method can generate a classification of the likelihood of a pregnancy-related disorder. The pregnancy-related disorder can be preeclampsia or any pregnancy-related disorder described herein.

[0513] Sequence reads corresponding to multiple cell-free DNA molecules can be received.

[0514] At block 7810, the size of multiple cell-free DNA molecules can be measured. The size can be determined by aligning nucleotides or counting the number of nucleotides or by including Figure 21measured by any of the techniques described herein.

[0515] At block 7820, a set of cell-free DNA molecules having a size greater than a cut-off value can be identified. The cut-off value can be any cut-off value for long cell-free DNA fragments, including 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cut-off value can be any cut-off value for long cell-free DNA molecules described herein.

[0516] At block 7830, a value of an end motif parameter can be generated using a first quantity. A first quantity of the cell-free DNA molecules in the set having a first subsequence at one or more ends of the cell-free DNA molecules in the set can be measured. In some embodiments, the end motif parameter can be the first quantity normalized by the total quantity of all subsequences at an end. In some embodiments, the end can be the 3' end. In some embodiments, the end can be the 5' end.

[0517] The length of the first subsequence can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. The first subsequence can include the last nucleotide at the end of the corresponding cell-free DNA molecule. For example, the first subsequence can be the XYNN pattern as shown in Figure 74 In some embodiments, the first subsequence can not include the last nucleotide or nucleotides at the end of the corresponding cell-free DNA molecule. For example, the first subsequence can include the NNXY pattern of Figure 75 .

[0518] A second quantity of cell-free DNA molecules having a subsequence different from the first subsequence at one or more ends of the cell-free DNA molecules can be measured. The value of the end motif parameter can be generated using a ratio of the second quantity to a third quantity. For example, the second quantity can be divided by the third quantity or the third quantity can be divided by the second quantity.

[0519] At block 7840, the value of the end motif parameter can be compared to a threshold. The threshold can be a value representing a statistically significant difference relative to the value of the relevant parameter for an individual not suffering from a pregnancy-related disorder. The threshold can be determined from one or more normal pregnancy reference individuals or one or more reference individuals suffering from a pregnancy-related disorder.

[0520] In some embodiments, the value of an end motif parameter can be compared to a threshold, and the value of a second end motif parameter can be compared to a second threshold. A second quantity of free DNA molecules having a second subsequence different from the first subsequence at one or more ends of the free DNA molecules can be measured. Thus, the quantity of different end motifs can be determined. The value of the second end motif parameter can be generated using the second quantity. The value of the second end motif parameter can be compared to the second threshold. The second threshold can be the same as or different from the first threshold. Additional subsequences can be used in the same manner as the first and second subsequences. In some embodiments, all possible subsequences can be used for comparison to the threshold.

[0521] At block 7850, a classification of the likelihood of a pregnancy-related disorder can be determined using a comparison. When the value of a size parameter or the value of an end motif parameter exceeds a threshold, a pregnancy-related disorder can be likely.

[0522] In some embodiments, a classification of the likelihood of a pregnancy-related disorder can be determined using a comparison of the value of a second end motif parameter to a second cutoff value. When the value of a first end motif parameter exceeds a first threshold and the value of a second end motif parameter exceeds a second threshold, a pregnancy-related disorder can be likely.

[0523] The method can include using a size parameter other than the end motif parameter. A second set of free DNA molecules having a size within a first size range can be identified. The first size range can include sizes greater than a cutoff value. The first size range can include sizes greater than a cutoff value. The first size range can be less than 550 nt, 600 nt, 650 nt, 700 nt, 750 nt, 800 nt, 850 nt, 900 nt, 950 nt, 1 nt, 1.5 knt, 2 knt, 3 knt, 5 knt, or greater. The value of the size parameter can be generated using the second quantity of free DNA molecules in the second set. The value of the size parameter can be compared to a second threshold. A classification of the likelihood of a pregnancy-related disorder can be determined using the comparison of the value of the size parameter to the second threshold. When one or both of a first threshold and a second threshold are exceeded, the classification can be likely to have a pregnancy-related disorder.

[0524] The size parameter can be a standardized parameter. For example, a third amount of free DNA molecules within a second size range can be measured. The second size range can include sizes less than a first cutoff value. The second size range can include all sizes. The second size range can include 50 - 150 nt, 50 - 166 nt, 50 - 200 nt, 200 - 400 nt. The second size range can include any size for short free DNA fragments described herein. The second size range can exclude sizes within a first size range. The value of the size parameter can be generated by determining the ratio of a second amount to the third amount. For example, the second amount can be divided by the third amount or the third amount can be divided by the second amount.

[0525] Any one of the amounts of free DNA molecules can be a free DNA molecule from a specific origin tissue. For example, the origin tissue can be a T cell or another origin tissue described herein. The second amount can be similar to the contribution of T cells described with Figures 77A - 77D The contribution of the origin tissue can be determined using methylation status or patterns as described in the present disclosure.

[0526] V. Repeat Expansion-Related Diseases

[0527] Long free DNA fragments obtained from pregnant women can be used to identify repeat expansions in genes. Repeat expansions in genes can lead to neuromuscular diseases. Tandem repeat expansions have been associated with human diseases including but not limited to neurodegenerative disorders such as Fragile X syndrome, Huntington's disease, and spinocerebellar ataxia. These tandem repeat expansions can occur in the protein-coding region of genes (Machado-Joseph disease, Haw River syndrome, Huntington's disease) or non-coding regions (Friedreich ataxia, myotonic dystrophy, some forms of Fragile X syndrome). The expansions involve minisatellites, pentanucleotides, tetranucleotides, and many trinucleotide repeats have been associated with fragile sites. The expansions associated with these diseases can be caused by replication slippage or unequal recombination or epigenetic aberrations. The number of repeat sequences in a sequence refers to the total number of occurrences of the subsequence. For example, "CAGCAG" contains two repeat sequences. Since a repeat sequence contains at least two instances of the subsequence, the number of repeat sequences cannot be 1. The subsequence can be understood as the repeating unit.

[0528] In embodiments, the analysis of long cell-free DNA in pregnant women can facilitate the detection of repeat sequence-related diseases. For example, trinucleotide repeats represent repeated stretches of 3 bp motifs in a DNA sequence. An example is the sequence 'CAGCAGCAG' which includes three 3 bp 'CAG' motifs. Microsatellite expansions, which are typically trinucleotide repeat expansions, have been reported to play a key role in neurological disorders (Kovtun et al., Cell Res. 2008; 18:198-213; McMurray et al., Nat Rev Genet. 2010; 11:786-99). An example is that more than 55 CAG repeats (total 165 bp) in the ATXN3 gene are pathogenic, leading to Spinocerebellar Ataxia type 3 (SCA3) disease characterized by progressive motor problems. This condition is inherited in an autosomal dominant pattern. Thus, one copy of the altered gene is sufficient to cause the disorder. To determine the number of repeats of a microsatellite, polymerase chain reaction (PCR) is typically used to amplify the genomic region of interest and the PCR products are then subjected to a number of different techniques such as: capillary electrophoresis (Lyon et al., J Mol Diagn. 2010; 12:505-11), Southern blotting (Hsiao et al., J Clin Lab Anal. 1999; 13:188-93), melting curve analysis (Lim et al., J Mol Diagn. 2014; 17:302-14) and mass spectrometry (Zhang et al., Anal Methods. 2016; 8:5039-44). However, these methods are labor-intensive and time-consuming and are difficult to apply to high-throughput screening in real clinical practice such as prenatal testing. Sanger sequencing has considerable difficulty in inferring long repeats from complex sequence traces via manual inspection. It is well known that Illumina sequencing technology and Ion Torrent have considerable difficulty in sequencing GC-rich (or GC-poor) regions with those repeats (Ashely et al., 2016; 17:507-22) and the DNA length including the expanded repeats is prone to exceed the sequence read length (Loomis et al., Genome Res. 2013; 23:121-8).

[0529] Another example is myotonic dystrophy, an autosomal dominant disorder caused by CTG repeat expansions in the range of 50 to 4000 CTG repeats adjacent to the DMPK gene. Molecular diagnosis of DM is routinely performed in prenatal diagnosis by invasively analyzing the number of CTGs on fetal genomic DNA.

[0530] In contrast to short-read sequencing (hundreds of bases), the methods described in this disclosure are capable of obtaining long DNA molecules (multiple kilobases) from maternal plasma DNA. We can use the methods described in this disclosure to non-invasively determine whether an unborn fetus has inherited this disease from an affected mother.

[0531] Figure 79 Illustration showing maternal inheritance of an inferred fetus for a repeat sequence-related disease. At stage 7905, the cell-free DNA in a pregnant woman is subjected to single-molecule real-time (e.g., PacBio SMRT) sequencing. At stage 7910, the sequenced results are classified into long DNA and short DNA categories according to this disclosure. At stage 7915, the allelic information present in the long DNA molecules can be used to construct maternal haplotypes, namely Hap I and Hap II. Hap I and Hap II can each contain an expanded repeat sequence of a trinucleotide subsequence (e.g., CTG). At stage 7920, the imbalance of the haplotypes can be analyzed, similar to that described as Figure 16 In accordance with this disclosure, the methods described herein allow us to not only determine haplotypes (e.g., Hap I and Hap II), but also use the sequence information of the long DNA molecules to determine which haplotype has the expanded repeat sequence causing the disorder (e.g., affected Hap I). In this example, we can use the count, size, or methylation status of the short DNA molecules distributed throughout the maternal Hap I and Hap II to determine whether the fetus has inherited maternal Hap I (affected) or Hap II (unaffected) according to the methods described herein.

[0532] Figure 80 Illustration showing paternal inheritance of an inferred fetus for a repeat sequence-related disease. We can use the cell-free DNA in a pregnant woman to determine whether the fetus has inherited an affected paternal haplotype. As Figure 80 shown, the cell-free DNA (e.g., 5 CTG repeats for Hap I and 6 CTG repeats for Hap II) in an unaffected pregnant woman whose husband is affected by a repeat sequence expansion disease (e.g., 70 CTG repeats) is subjected to PacBio SMRT sequencing, and the sequenced long DNA molecules are identified and used to determine the haplotypes and the number of repeat sequences. If a haplotype with a long stretch of CTG repeats (e.g., 70 CTG repeats in this example) is present in the maternal plasma of the unaffected pregnant woman, it indicates that the fetus has inherited the affected paternal haplotype. In some embodiments, the DNA containing the expanded repeat sequence also carries one or more other paternal-specific alleles that are not present in the maternal genome. This situation will apply to confirm paternal inheritance.

[0533] In another embodiment, we can use the cell-free DNA in a pregnant woman to determine whether the fetus has inherited an affected paternal haplotype. As Figure 80 shown, the cell-free DNA in an unaffected pregnant woman whose husband is affected by a repeat expansion disease (e.g., 70 CTG repeats) (e.g., 5 CTG repeats for Hap I and 6 CTG repeats for Hap II) is subjected to PacBio SMRT sequencing. The sequenced long DNA molecules are identified and used to determine the haplotype and the number of repeat sequences. If the haplotype with a long stretch of CTG repeats (e.g., 70 CTG repeats in this example) is present in the maternal plasma of the unaffected pregnant woman, it indicates that the fetus has inherited the affected paternal haplotype. In some embodiments, the DNA containing the expanded repeat also carries one or more other paternal-specific alleles that are not present in the maternal genome. This situation will apply to confirm paternal inheritance.

[0534] Figure 81 , Figure 82 and Figure 83 are tables showing examples of repeat expansion diseases. The first column shows the repeat expansion-related diseases. The second column shows the repeat subsequences. The third column shows the number of repeat sequences in normal individuals. The fourth column shows the number of repeat sequences in affected individuals. The fifth column shows the genetic location related to the repeat sequences. The sixth column lists the gene names. The seventh column lists the inheritance patterns. The tables are from omicslab.genetics.ac.cn / dred / index.php.

[0535] A. Examples of Repeat Expansion Detection

[0536] It has been reported that paternally inherited expanded CAG repeats can be detected in maternal plasma using a direct approach by PCR and subsequent fragment analysis on a 3130XL genetic analyzer (Oever et al., Prenat Diagn. 2015; 35: 945-9). Non-invasive prenatal testing for Huntington's can be achieved by PCR because the size of the expanded allele only starts from >35 trinucleotide repeats [i.e., a DNA region with a length of 105 bp (35×3) or higher across the repeats]. Many expanded repeats for most trinucleotide repeat disorders (Orr et al., Annu. Rev. Neurosci. 2007; 30: 575-621) will involve repeats with a length of 300 bp or higher, which are beyond the size of the short fetal DNA molecules recorded in previous reports. DNA with large expanded repeats will pose difficulties for PCR (Orr et al., Annu. Rev. Neurosci. 2007; 30: 575-621). As shown in the study by Oever et al., the signal intensity of long CAG repeats is often much lower compared to that of smaller repeats, and this phenomenon is observed in both genomic DNA and plasma DNA, resulting in lower sensitivity for detecting those long CAG repeats (Oever et al., Prenat Diagn. 2015; 35: 945-9). Another limitation of PCR is the inability to preserve methylation signals during amplification. In one embodiment, single molecule real-time sequencing of long DNA molecules will allow determination of tandem repeat polymorphisms and their associated methylation levels in one or more regions.

[0537] Figure 84 Table showing examples of detection of repeat expansion and determination of repeat-associated methylation in a fetus. The first column shows the repeat type in multiple base pairs. The second column shows the repeat unit. The third column shows the genomic location. The fourth column shows the reference base, i.e., the sequence present in the human reference genome. The fifth column shows the paternal genotype. The sixth column shows the maternal genotype. The seventh column shows the fetal genotype. The eighth column shows the degree of fetal DNA methylation related to the paternal allele. The ninth column shows the degree of fetal DNA methylation related to the maternal allele.

[0538] Figure 84 Showing multiple examples of 1bp, 2bp, 3bp, and 4bp tandem repeats. For example, the "GATA" tandem repeat is identified at the genomic location of chr3: 192384705-192384706. The genotype of the father at this locus is T(GATA) 3 / T(GATA) 5, where allele 1 has 3 repeat units and allele 2 has 5 repeat units. Compared with the reference allele T(GATA) 3 , the paternal allele 2 indicates a genetic event involving repeat sequence expansion. The mother's genotype at this locus is T / T, showing a genetic event involving repeat sequence contraction. The fetal genotype at this locus is T(GATA) 5 / T, indicating that the fetus has inherited the paternal allele 2 (i.e., T(GATA) 5 ), and the maternal allele T. The methylation levels associated with the paternal allele and the maternal allele are 50.98% and 62.8%, respectively. These results indicate that the use of tandem repeat sequence polymorphisms will allow the determination of the maternal and paternal inheritance of the fetus. This technique will allow the identification of different methylation patterns associated with the two alleles. Another example shows that at the genomic position of chr4:73237157 - 73237158, the fetus has inherited a repeat sequence expansion [(TAAA) 3 from the mother. Compared with the methylation level (62.84%) of the fetal molecule containing the paternal allele, the fetal molecule containing the repeat sequence expansion inherited from the mother shows a higher methylation level (95.65%). These data indicate that we can detect repeat sequences, repeat sequence structures, and related methylation changes. In one embodiment, we can use a specific cutoff value to determine whether the methylation difference between maternal inheritance and paternal inheritance is significant. The cutoff value will be the absolute difference in methylation levels greater than but not limited to 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or 90%, etc. The determination of maternal inheritance can be similar to the method described by method 2100 using Figure 21 .

[0539] B. Example Methods

[0540] Subsequence repeat sequences can be used to determine fetal information. For example, the presence of subsequence repeat sequences can be used to determine that a molecule is of fetal origin. Additionally, subsequence repeat sequences can indicate the likelihood of a genetic disorder. Subsequence repeat sequences can be used to determine the inheritance of maternal and / or paternal haplotypes. Additionally, the relatedness of a fetus can be determined using subsequence repeat sequences.

[0541] 1. Fetal Origin Analysis Using Subsequence Repeat Sequences

[0542] Figure 85 Method 8500 for analyzing a biological sample obtained from a female carrying a fetus is shown, where the biological sample contains cell-free DNA molecules from the fetus and the female. The likelihood of a genetic disorder in the fetus can be determined.

[0543] At block 8510, a first sequence read corresponding to one of the cell-free DNA molecules among the cell-free DNA molecules may be received. The cell-free DNA molecule may have a length greater than a cut-off value. The cut-off value may be greater than or equal to 200 nt. The cut-off value may be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cut-off value may be any cut-off value for long cell-free DNA molecules described herein.

[0544] At step 8520, the first sequence read may be aligned with a region of a reference genome. The region may be known to potentially contain repetitive sequences of subsequences. The region may correspond to Figures 81 - 83 either a position in or a gene in. The subsequence may be a trinucleotide sequence, including any trinucleotide sequence described herein.

[0545] At block 8530, the number of repetitive sequences corresponding to a subsequence in the first sequence read of the cell-free DNA molecule may be identified.

[0546] At block 8540, the number of repetitive sequences of the subsequence may be compared with a threshold number. The threshold number may be 55, 60, 75, 100, 150, or more. For different genetic disorders, the threshold number may be different. For example, the threshold may reflect the minimum number of repetitive sequences in an affected individual, the maximum number of repetitive sequences in a normal individual, or a number between these two numbers (see Figures 81 - 83 ).

[0547] At block 8550, the comparison of the number of repetitive sequences with the threshold number may be used to determine a classification of the likelihood that the fetus has a genetic disorder. When the number of repetitive sequences exceeds the threshold number, the fetus may be determined to possibly have a genetic disorder. The genetic disorder may be Fragile X syndrome or Figures 81 - 83 any disorder listed in.

[0548] In some embodiments, the method may include classifying multiple different target loci, each of which is known to potentially have a repetitive sequence of a subsequence. Multiple sequence reads corresponding to cell-free DNA molecules may be received. The multiple sequence reads may be aligned with multiple regions of a reference genome. The multiple regions are known to potentially contain repetitive sequences of subsequences. The multiple regions may be non-overlapping regions. Each of the multiple regions may have different SNPs. The multiple regions may be from different chromosomal arms or chromosomes. The multiple regions may cover at least 0.01%, 0.1%, or 1% of the reference genome. The number of repetitive sequences of subsequences in the multiple sequence reads may be identified. The number of repetitive sequences of subsequences may be compared with multiple threshold numbers. Each threshold number may indicate the presence or likelihood of a different genetic disorder. For each of the multiple genetic disorders, a comparison with one of the multiple threshold numbers may be used to determine a classification of the likelihood that the fetus has the corresponding genetic disorder.

[0549] The cell-free DNA molecules may be determined to be of fetal origin. Determination of fetal origin may include receiving second sequence reads corresponding to cell-free DNA molecules of maternal origin obtained from the buffy coat or sample of the female prior to pregnancy. The second sequence reads may be aligned with the regions of the reference genome. The second number of repetitive sequences of subsequences in the second sequence reads may be identified. It may be determined that the second number of repetitive sequences is less than the first number of repetitive sequences.

[0550] Determination of fetal origin may include determining the methylation level of the cell-free DNA molecules using methylated and unmethylated sites of the cell-free DNA molecules. The methylation level may be compared with a reference level. The method may include determining that the methylation level exceeds the reference level. The methylation level may be the number or proportion of methylated sites.

[0551] Determination of fetal origin may include determining the methylation pattern of multiple sites of the cell-free molecules. A similarity score may be determined by comparing the methylation pattern with a reference pattern from maternal or fetal tissue. The similarity score may be compared with one or more thresholds. The similarity score may be any similarity score described herein, including, for example, the similarity score described using method 4000.

[0552] 2. Kinship analysis using subsequence repetitive sequences

[0553] Figure 86 Method 8600 for analyzing a biological sample obtained from a female carrying a fetus, the biological sample including cell-free DNA molecules from the fetus and the female. The biological sample may be analyzed to determine the father of the fetus.

[0554] At block 8610, a first sequence read corresponding to one of the cell-free DNA molecules in the cell-free DNA molecules may be received. The method may include determining that the cell-free DNA molecule is of fetal origin. The cell-free DNA molecule may be determined to be of fetal origin by any method described herein, any method including, for example, the method described with method 8500. The cell-free DNA molecule may have a size greater than a cut-off value. The cut-off value may be greater than or equal to 200 nt. The cut-off value may be at least 500 nt, including 600 nt, 700 nt, 800 nt, 900 nt, 1 knt, 1.1 knt, 1.2 knt, 1.3 knt, 1.4 knt, 1.5 knt, 1.6 knt, 1.7 knt, 1.8 knt, 1.9 knt, or 2 knt. The cut-off value may be any cut-off value described herein for long cell-free DNA molecules.

[0555] At block 8620, the first sequence read may be aligned with a first region of a reference genome. The first region is known to have a repetitive sequence of subsequences.

[0556] At block 8630, a first number of repetitive sequences corresponding to a first subsequence in the first sequence read corresponding to the cell-free DNA molecule may be identified. The first subsequence may include an allele.

[0557] At block 8640, sequence data obtained from a male individual may be analyzed to determine whether a second number of repetitive sequences of the first subsequence are present in the first region. The second number of repetitive sequences includes at least two instances of the first subsequence. The sequence data may be obtained by extracting a biological sample from the male individual and performing sequencing on the DNA in the biological sample.

[0558] At block 8650, the determination of whether the second number of repetitive sequences of the first subsequence are present may be used to classify the likelihood that the male individual is the father of the fetus. The classification may be that the male individual is likely to be the father when it is determined that the second number of repetitive sequences of the first subsequence are present. The classification may be that the male individual is likely not to be the father when it is determined that the second number of repetitive sequences of the first subsequence are not present.

[0559] The method may include comparing the first number of repetitive sequences with the second number of repetitive sequences. Classifying the likelihood that the male individual is the father may include using the comparison of the first number of repetitive sequences with the second number of repetitive sequences. The classification may be that the male individual is likely to be the father when the first number of repetitive sequences is within a threshold of the second number of repetitive sequences. The threshold may be within 10%, 20%, 30%, or 40% of the second number of repetitive sequences.

[0560] The method can include using multiple regions of a repetitive sequence. For example, a cell-free DNA molecule is a first cell-free DNA molecule. The method can include receiving a second sequence read corresponding to a second cell-free DNA molecule in the cell-free DNA molecules. The method can also include aligning the second sequence read with a second region of a reference genome. The method can further include identifying a first number of repetitive sequences corresponding to a second subsequence in the second sequence read corresponding to the second cell-free DNA molecule. The method can include analyzing sequence data obtained from a male individual to determine whether a second number of repetitive sequences of the second subsequence are present in the second region. The classification of determining the likelihood that the male individual is the father of the fetus can further include using the determination of whether the second number of repetitive sequences of the second subsequence are present in the second region. The classification of the likelihood can be that the likelihood that the male individual is the father of the fetus is higher when the repetitive sequences are present in both the first region and the second region in the sequence data of the male individual.

[0561] VI. Size selection for enriching long plasma DNA molecules

[0562] In embodiments, we can physically select DNA molecules having one or more desired size ranges prior to analysis (e.g., single molecule real-time sequencing). As an example, size selection can be performed using solid-phase reversible immobilization techniques. In other embodiments, size selection can be performed using electrophoresis (e.g., using a Coastal Genomic system or a Pippin size selection system). Our approach is different from previous operations that mainly focused on shorter DNA (Li et al., Journal of the American Medical Association (JAMA) 2005; 293:843-9), and it is well known in the art that fetal DNA is shorter than maternal DNA (Chan et al., Clinical Chemistry 2004; 50:88-92).

[0563] Size selection techniques can be applied to any of the methods described herein and for any size described herein. For example, cell-free DNA molecules can be enriched by electrophoresis, magnetic beads, hybridization, immunoprecipitation, amplification, or CRISPR. The resulting enriched sample can have a greater concentration or a higher proportion of specific size fragments than the biological sample before enrichment.

[0564] A. Size selection by electrophoresis

[0565] In an embodiment, taking advantage of the fact that the electrophoretic mobility of DNA depends on the DNA size, we can use a gel electrophoresis-based approach to select target DNA molecules having a desired size range, such as but not limited to ≥100 bp, ≥200 bp, ≥300 bp, ≥400 bp, ≥500 bp, ≥600 bp, ≥700 bp, ≥800 bp, ≥900 bp, ≥1 kb, ≥2 kb, ≥3 kb, ≥4 kb, ≥5 kb, ≥6 kb, ≥7 kb, ≥8 kb, ≥9 kb, ≥10 kb, ≥20 kb, ≥30 kb, ≥40 kb, ≥50 kb, ≥60 kb, ≥70 kb, ≥80 kb, ≥90 kb, ≥100 kb, ≥200 kb; or other size ranges, including sizes greater than any cut-off value described herein. For example, the LightBench (Coastal Genomics) automated gel electrophoresis system is used for DNA size selection. In principle, during gel electrophoresis, shorter DNA will move faster than longer DNA. We applied this size selection technique to a plasma DNA sample (M13190) aiming to select DNA molecules greater than 500 bp. We used a 3% size selection cartridge with an 'in-channel filter' (ICF) collection device and a loading buffer with internal size markers for size selection. The DNA library was loaded into the gel and electrophoresis was started. When the target size was reached, the first fraction <500 bp was retrieved from the ICF. The operation was resumed and electrophoresis was allowed to complete to obtain the second fraction ≥500 bp. We used single molecule real-time sequencing (PacBio) to sequence the second fraction of molecules with a size ≥500 bp. We obtained 1,434 high-quality circular consensus sequences (CCS) (i.e., 1,434 molecules). Among them, 97.9% of the sequenced molecules were greater than 500 bp. Such a proportion of DNA molecules greater than 500 bp is much higher than the proportion (10.6%) of the corresponding DNA molecules without size selection. The total methylation of those molecules was determined to be 75.5%.

[0566] Figure 87Shows the methylation patterns of two representative plasma DNA molecules after size selection in (I) molecule I and (II) molecule II. Molecule I (chr21:40,881,731-40,882,812) is 1.1 kb long and has 25 CpG sites. The single-molecule methylation degree of molecule I (i.e., the number of methylated sites divided by the total number of sites) was determined to be 72.0% using the approach described in our previous publication (U.S. application No. 16 / 995,607). Molecule II (chr12:63,108,065-63,111,674) is 3.6 kb long and has 34 CpG sites. The single-molecule methylation degree of molecule II was determined to be 94.1%. This indicates that size-selection-based methylation analysis allows us to effectively analyze the methylation of long DNA molecules and compare the methylation status between two or more molecules.

[0567] B. Size Selection with Beads

[0568] Solid-phase reversible immobilization technology uses paramagnetic beads to selectively bind nucleic acids depending on the DNA molecule size. Such beads contain a polystyrene core, magnetite, and a carboxylate-modified polymer coating. DNA molecules will selectively bind to the beads depending on the PEG and salt concentrations in the reaction in the presence of polyethylene glycol (PEG) and salt. PEG enables negatively charged DNA to bind to the carboxyl groups on the bead surface, and the negatively charged DNA is collected in the presence of a magnetic field. Molecules of the desired size are eluted from the magnetic beads using an elution buffer such as 10 mM Tris-HCl, pH 8 buffer or water. The volume ratio of PEG to DNA will determine the size of the DNA molecules we can obtain. The lower the PEG:DNA ratio, the more long molecules will be retained on the beads.

[0569] 1. Sample Processing

[0570] Peripheral blood samples from two pregnant women in the third trimester were collected in EDTA blood tubes. The peripheral blood samples were collected and centrifuged at 1,600 × g for 10 minutes at 4°C. The plasma fraction was further centrifuged at 16,000 × g for 10 minutes at 4°C to remove residual cells and debris. The buffy coat fraction was centrifuged at 5,000 × g for 5 minutes at room temperature to remove residual plasma. Placental tissue was collected immediately after delivery. Plasma DNA extraction was performed using the QIAamp Circulating Nucleic Acid Kit (Qiagen). Buffy coat and placental tissue DNA extraction were performed using the QIAamp DNA Mini Kit (Qiagen).

[0571] 2. Plasma DNA Size Selection

[0572] The extracted plasma DNA samples are divided into two aliquots. One aliquot from each patient undergoes size selection using AMPure XP SPRI beads (Beckman Coulter, Inc.). 50 μL of each extracted plasma DNA sample is thoroughly mixed with 25 μL of AMPure XP solution and incubated for 5 minutes at room temperature. The beads are separated from the solution using a magnet and washed with 180 μL of 80% ethanol. Subsequently, the beads are resuspended in 50 μL of water and vortexed for 1 minute to elute the size-selected DNA from the beads. Subsequently, the beads are removed to obtain the size-selected DNA solution.

[0573] 3. Single nucleotide polymorphism identification

[0574] The fetal genomic DNA samples and maternal genomic DNA samples are genotyped using an iScan system (Illumina). Single nucleotide polymorphisms (SNPs) are identified. The genotype of the placenta is compared with the genotype of the mother to identify fetal-specific alleles and maternal-specific alleles. Fetal-specific alleles are defined as alleles that are present in the fetal genome but not in the maternal genome. In one embodiment, those fetal-specific alleles can be determined by analyzing those SNP loci in which the mother is homozygous and the fetus is heterozygous. Maternal-specific alleles are defined by alleles that are present in the maternal genome but not in the fetal genome. In one embodiment, those fetal-specific alleles can be determined by analyzing those SNP loci in which the mother is heterozygous and the fetus is homozygous.

[0575] 4. Single molecule real-time sequencing

[0576] Two size-selected samples and their corresponding unselected samples are subjected to single molecule real-time (SMRT) sequencing template construction using the SMRTbell Template Prep Kit 1.0-SPv3 (Pacific Biosciences). The DNA is purified with 1.8×AMPure PB beads, and the library size is estimated using a TapeStation instrument (Agilent). The sequencing primer annealing and polymerase binding conditions are calculated using SMRT Link version 5.1.0 software (Pacific Biosciences). Briefly, sequencing primer v3 is annealed to the sequencing template, and then the polymerase is bound to the template using the Sequel Binding and Internal Control Kit 2.1 (Pacific Biosciences). Sequencing is performed on a Sequel SMRT Cell 1M v2. Sequencing movies are collected for 20 hours on a Sequel system using the Sequel Sequencing Kit 2.1 (Pa...

Claims

1. A computer program product comprising instructions that, when executed, control a computer system to perform a method of analyzing a biological sample obtained from a pregnant female, the biological sample comprising a plurality of cell-free DNA molecules from the fetus and the female, the method comprising: measuring the sizes of the plurality of cell-free DNA molecules; identifying a set of cell-free DNA molecules having sizes greater than a cut-off value, wherein the cut-off value is at least 500 nt; generating a value of a terminal motif parameter using a first quantity, wherein generating the value of the terminal motif parameter comprises: measuring the first quantity of the cell-free DNA molecules in the set that have a first subsequence at one or more ends of the cell-free DNA molecules in the set; comparing the value of the terminal motif parameter with a threshold; and using the comparison to determine a classification of the likelihood of a pregnancy-related disorder.

2. The computer program product according to claim 1, wherein the method further comprises: measuring a second quantity of cell-free DNA molecules having a subsequence different from the first subsequence at one or more ends of the cell-free DNA molecules, and wherein: generating the value of the terminal motif parameter comprises using a ratio of the first quantity to the second quantity.

3. The computer program product according to claim 1, wherein the length of the first subsequence is 1, 2, 3, or 4 nucleotides.

4. The computer program product according to claim 3, wherein the first subsequence includes the last nucleotide at the end of the corresponding cell-free DNA molecule.

5. The computer program product according to claim 1, wherein: the threshold is a first threshold, and the terminal motif parameter is a first terminal motif parameter, the method further comprises: measuring a second quantity of cell-free DNA molecules having a second subsequence different from the first subsequence at one or more ends of the cell-free DNA molecules, generating a value of a second terminal motif parameter using the second quantity, and comparing the value of the second terminal motif parameter with a second threshold; wherein: determining the classification of the likelihood of the pregnancy-related disorder uses the comparison of the value of the second terminal motif parameter with the second threshold, wherein the pregnancy-related disorder is likely when the value of the first terminal motif parameter exceeds the first threshold and the value of the second terminal motif parameter exceeds the second threshold.

6. The computer program product according to claim 1, wherein the first quantity of cell-free DNA molecules comprises cell-free DNA molecules determined to be from the originating tissue.

7. The computer program product according to claim 1, wherein: the threshold is a first threshold, and the set of cell-free DNA molecules is a first set of cell-free DNA molecules, the method further comprises: identifying a second set of cell-free DNA molecules having sizes within a first size range, the first size range including sizes greater than the cut-off value, generating a value of a size parameter using a second quantity of the cell-free DNA molecules in the second set, and comparing the value of the size parameter with a second threshold, The classification that determines the likelihood of the pregnancy-related disorder includes comparing the value of the size parameter with the second threshold.

8. The computer program product according to claim 1, wherein the pregnancy-related disorder includes preeclampsia, intrauterine growth restriction, invasive placental formation, preterm birth, neonatal hemolytic disease, placental insufficiency, fetal hydrops, fetal malformation, hemolysis, elevated liver enzymes and low platelet count (HELLP) syndrome, or systemic lupus erythematosus.

9. The computer program product according to any one of claims 1 to 8, wherein the cut-off value is 600 nt.

10. The computer program product according to any one of claims 1 to 8, wherein the cut-off value is 1,000 nt.

11. The computer program product according to any one of claims 1 to 10, wherein: enriching the plurality of free DNA molecules relative to the biological sample to obtain a size greater than or equal to the cut-off value, wherein more than 20% of the free DNA molecules in the biological sample have a size greater than 200 nt.

12. The computer program product according to claim 11, wherein the method further comprises: enriching the plurality of free DNA molecules using electrophoresis.

13. The computer program product according to claim 11, wherein the method further comprises: enriching the plurality of free DNA molecules using magnetic beads for size-selective binding of free DNA molecules.

14. The computer program product according to claim 11, wherein the method further comprises: enriching the plurality of free DNA molecules using hybridization, immunoprecipitation, amplification, or CRISPR.

15. The computer program product according to any one of claims 12 to 14, wherein the enrichment is for obtaining a size greater than 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, or 1 knt.

16. The computer program product according to any one of claims 1 to 10, wherein the plurality of free DNA molecules are enriched relative to the biological sample to obtain a methylation profile, the method further comprises: enriching the plurality of free DNA molecules using immunoprecipitation.

17. The computer program product according to claim 1, wherein the plurality of free DNA molecules include at least 10,000 molecules.

18. The computer program product according to claim 1, wherein the method further comprises receiving readings corresponding to the plurality of free DNA molecules, wherein the readings are obtained by single molecule sequencing.

19. The computer program product according to claim 18, wherein the single molecule sequencing includes optical monitoring of a DNA polymerase that incorporates new bases into a complementary strand of the plurality of free DNA molecules.

20. A computer-readable storage medium comprising the computer program product according to any one of claims 1 - 19.

21. A computing system comprising one or more processors and the computer program product according to any one of claims 1 - 19.

Citation Information

Patent Citations

  • Gestational age assessment by methylation and size profiling of maternal plasma DNA

    US20180105807A1

  • Determination of base modifications of nucleic acids

    US20210047679A1