Analysis of Nucleic Acids Associated with Extracellular Vesicles
By concentrating and analyzing long DNA fragments from extracellular particles, the method enhances the fetal DNA fraction and sensitivity of NIPT, addressing challenges of low fetal DNA and improving diagnostic accuracy.
Patent Information
- Application Number
- JP2024566282
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-10
- Filing Date
- 2023-05-09
- Publication Date
- 2025-06-10
AI Technical Summary
Non-invasive prenatal testing (NIPT) faces challenges with low fetal DNA fraction, leading to test failures, low sensitivity in detecting fetal aneuploidy, and high false negative rates, especially in samples with fetal DNA fractions below 4%.
The method involves purifying and concentrating extracellular particles from blood samples to enrich fetal DNA, selecting long DNA fragments greater than 200 bp, and using sequencing techniques to analyze these fragments, thereby increasing the fetal DNA fraction and genetic/epigenetic informativeness.
This approach significantly increases the fetal DNA fraction, improves the sensitivity of detecting fetal aneuploidy, reduces false negative rates, and enables more accurate genetic and epigenetic analysis, facilitating earlier and more reliable NIPT.
Smart Images

Figure 2025517662000005 
Figure 2025517662000006 
Figure 2025517662000007
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application claims priority from U.S. Provisional Patent Application No. 63 / 340,316, filed May 10, 2022, entitled "Analysis Of Nucleic Acids Associated With Extracellular Vesicles", and is a PCT application, the entire content of which is incorporated herein by reference for all purposes.
Background Art
[0002] The discovery of cell - free fetal DNA in maternal plasma has opened up a series of new means for non - invasive prenatal testing (NIPT), including the detection of chromosomal aneuploidies and the diagnosis of single - gene disorders. The accuracy of NIPT is usually affected by the fractional concentration of fetal DNA in the maternal plasma sample, which is commonly referred to as the fetal DNA fraction (Chiu et al. BMJ 2011;342:c7401; Canick et al. Prenat. Diagn. 2013;33:667 - 674). For the following reasons, when analyzing samples with a relatively low fetal DNA fraction, there is a need to improve the performance of NIPT.
[0003] For example, when the fetal DNA fraction is low, test failure or no result is likely to occur (Porreco et al. Am. J. Obstet. Gynecol. 2014;211:365.e1-365.e12). Second, the sensitivity of detecting fetal aneuploidy in pregnancies with low fetal DNA fraction is low (Chiu et al. BMJ 2011;342:c7401; Jiang et al. Bioinformatics 2012;28:2883-2890, npj Genomic Med. 2016;1:16013; Hui et al. Prenatal Diagnosis 2020;40:155-163), which results in a high false negative rate. Simply repeating the assay is likely to lead to misdecision or no decision. Previously, Canick et al. reported four false negatives among 212 cases with Down syndrome, all of which had a relatively low fetal DNA fraction of 4% - 7% (Canick et al. Prenat. Diagn 2013;33:667-674). Finally, the prevalence of fetal aneuploidy appears to vary in pregnancies with different fetal DNA fractions. For example, in patients with a fetal DNA fraction below 4%, the prevalence of aneuploidy was 4.7%, which was demonstrated to be significantly higher compared to the prevalence of 0.4% in the overall cohort (Norton et al. N. Engl. J. Med. 2015;372:1589-1597). Therefore, patients with low fetal DNA fraction cannot be ignored. Therefore, it is desirable to provide an improved technique to address such problems.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
[0005] In various examples, cell-free DNA from extracellular particles (EP) is analyzed. The sample can be purified for extracellular particles. By way of example, purification can include centrifugation, washing, and nuclease treatment. To increase the fetal fraction, purification can concentrate the sample for certain types of EP (e.g., long EP). In this way, the desired particle population can be selected for analysis of their nucleic acids. As part of the analysis of DNA molecules (fragments) from the concentrated sample, DNA molecules larger than a certain size can be selected, which can increase genetic and / or epigenetic informativeness without side effects (e.g., reduction of the fetal DNA fraction). Long DNA fragments can be analyzed in various ways, including using short-read sequencing techniques that perform fragmentation prior to sequencing and using long-read sequencing techniques.
[0006] In one example, the method includes receiving a blood sample from a pregnant female carrying a fetus. One or more purification steps can concentrate extracellular particles to produce a concentrated sample. The extracellular particles can contain cell-free nucleic acids (e.g., DNA and / or RNA) inside the membrane. The membrane of the extracellular particles can be disrupted to expose the cell-free nucleic acid molecules from the extracellular particles. An assay can be applied to the cell-free nucleic acid molecules to obtain sequence reads. The cell-free nucleic acid molecules from inside the EP and / or bound to the surface of the EP can be assayed. The size of the cell-free nucleic acid molecules can be determined. By way of example, the sequence reads can be used to determine the size of the cell-free nucleic acid molecules, or physical techniques such as electrophoresis or PCR using amplicons of different sizes can be used. For example, if the size threshold is 200 bp or more, a set of cell-free nucleic acid molecules larger than the size threshold can be identified. The sequence reads can be analyzed to determine genomic characteristics of the fetus.
[0007] In another example, a blood sample from a pregnant female carrying a fetus can contain extracellular particles and particle-free nucleic acids. The extracellular particles can contain cell-free nucleic acids inside the membrane. Physical separation techniques can preferentially select at least some of the extracellular particles, thereby obtaining a particle concentrated sample, which can be processed using processing techniques that remove excess particle-free nucleic acids, thereby obtaining a processed particle concentrated sample. The processing techniques can include washing the particle concentrated sample with an ionic solution and applying nuclease to the particle concentrated sample. The processing techniques can increase the fractional concentration of fetal nucleic acids in the processed particle concentrated sample compared to the particle concentrated sample. The membrane of the extracellular particles can be disrupted to expose the cell-free nucleic acids from the extracellular particles. An assay can be applied to the cell-free nucleic acid molecules to obtain sequence reads. The cell-free nucleic acid molecules from inside the EP and / or bound to the surface of the EP can be assayed. The sequence reads can be analyzed to determine genomic characteristics of the fetus or the pregnancy of the female.
[0008] In another example, a blood sample of a female who is pregnant with a fetus can contain extracellular particles and particle-free nucleic acid molecules. The extracellular particles can contain cell-free nucleic acid molecules inside the membrane. One or more purification steps can concentrate the extracellular particles to produce a concentrated sample. The membrane of the extracellular particles can be disrupted to expose the cell-free nucleic acid molecules from the extracellular particles. Sequencing techniques can be applied to the cell-free nucleic acid molecules to obtain sequence reads. The cell-free nucleic acid molecules from inside and / or bound to the surface of the EP can be sequenced. At least some of the sequence reads can exceed 600 bp. The sequence reads can be analyzed to determine genomic characteristics of the fetus or the pregnancy of the female.
[0009] In another example, a blood sample of a female who is pregnant with a fetus can contain extracellular particles and particle-free nucleic acid molecules. The extracellular particles can contain cell-free nucleic acid molecules inside the membrane. One or more purification steps can concentrate the extracellular particles to produce a concentrated sample. The membrane of the extracellular particles can be disrupted to expose the cell-free nucleic acid molecules from the extracellular particles. At least some of the cell-free nucleic acid molecules from the extracellular particles are at least 600 bp. Fragmentation techniques can be applied to the cell-free nucleic acid molecules. After applying the fragmentation techniques, sequencing techniques can be applied to the cell-free nucleic acid molecules to obtain sequence reads. The cell-free nucleic acid molecules from inside and / or bound to the surface of the EP can be sequenced. The sequence reads can be analyzed to determine genomic characteristics of the fetus or the pregnancy of the female.
[0010] These and other embodiments of the present disclosure are described in detail below. For example, other embodiments are directed to systems, devices, and computer-readable media related to the methods described herein.
[0011] A better understanding of the nature and advantages of the embodiments of the present disclosure can be obtained by reference to the following detailed description and the accompanying drawings.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 6A
Figure 6B
Figure 7
Figure 8A
Figure 8B
Figure 8C
Figure 10A
Figure 10B
Figure 9
Figure 11
Figure 12A
Figure 12B
Figure 12C
Figure 13
Figure 14
Figure 15
Figure 16A
Figure 16B
Figure 17A
Figure 17B
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
[0013] TERMINOLOGY The term "tissue" corresponds to a group of cells grouped together as a functional unit. More than one type of cell can be found in a single tissue. Different types of tissues can consist of different types of cells (e.g., hepatocytes, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (mother vs. fetus). A "reference tissue" can correspond to a tissue used to determine tissue-specific methylation patterns. Multiple samples of the same tissue type from different individuals can be used to determine tissue-specific methylation patterns for that tissue type (e.g., fetal tissue).
[0014] A "biological sample" refers to any sample collected from a pregnant woman and containing one or more nucleic acid molecules of interest (e.g., DNA and / or RNA). Biological samples can be body fluids such as blood, plasma, serum, urine, vaginal fluid, fluid from a hydatid cyst (e.g., of the testis), vaginal lavage fluid, pleural effusion, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, excreted fluid from the nipple, aspirated fluid from different parts of the body (e.g., thyroid, breast), intraocular fluid (e.g., aqueous humor), etc. Stool samples can also be used. In various embodiments, most of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) can be cell-free, e.g., more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free. The centrifugation protocol can include, for example, 3,000 g × 10 minutes, obtaining the fluid portion, and re-centrifuging at, for example, 30,000 g for an additional 10 minutes to remove residual cells. Other centrifugation protocols can be used at various forces (rotation speeds) such as at least 1,600 g, 5,000 g, 10,000 g, 16,000 g, 20,000 g, 30,000 g, 40,000 g, 50,000 g, 60,000 g, 70,000 g, 80,000 g, 90,000 g, 100,000 g, and 110,000 g, and for various times, e.g., at least 5 minutes, 10 minutes, 15 minutes, 20 minutes, 30 minutes, 40 minutes, 1 hour, or 2 hours, and this can be repeated. Other centrifugation protocols are described herein. As part of the analysis of a biological sample, a statistically significant number of cell-free DNA molecules can be analyzed for the biological sample (e.g., to provide accurate measurements). In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 10,000, or 50,000, or 100,000, or 500,000, or 1,000,000, or 5,000,000 or more cell-free DNA molecules can be analyzed. At least the same number of sequence reads can be analyzed.
[0015] "Extracellular vesicles" (EVs) are also referred to as "extracellular particles" (EPs) and refer to small localized particles with specific physical and / or chemical characteristics such as volume, density, mass, electronegativity, and permeability, which can be released from cells and can occur from living cells or during cell death such as apoptosis or necrosis. EPs may or may not have a membrane in which genomic material is present (e.g., non-membrane-bound protein-nucleic acid complexes). Such particles can contain proteins, nucleic acids (DNA and / or RNA), lipids, metabolites, and organelles from the parent cell. EVs can be divided according to size and synthetic pathway and can be referred to as exosomes, microvesicles, and apoptotic bodies. In certain contexts, exosomes are membrane-bound EVs generated in the endosomal compartment of most eukaryotic cells. Microvesicles (also called ectosomes or microparticles) are a type of extracellular vesicle (EV) released from the cell membrane. EVs can be referred to as small (SEV) or large (LEV) depending on size. By way of example, EVs can have diameters ranging from a few nanometers to a few micrometers. EVs can play a role in cell-to-cell communication and can transport molecules such as mRNA, miRNA, and proteins between cells. Any of the above terms are interchangeable and refer to EVs or EPs. Exemplary numbers of particles that can be analyzed include at least 100, 500, 1,000, 5,000, 10,000, 50,000, and 100,000 particles.
[0016] As used herein, the term "fragment" (e.g., a DNA or RNA fragment) can refer to a portion of a polynucleotide or polypeptide sequence that contains at least three contiguous nucleotides. A nucleic acid fragment can retain the biological activity and / or some of the properties of the parent polynucleotide. A nucleic acid fragment can be double-stranded or single-stranded, methylated or non-methylated, intact or nicked, complexed or non-complexed with other macromolecules, e.g., lipid particles or proteins. A nucleic acid fragment can be a linear or circular fragment.
[0017] "Cell-free DNA" (cfDNA) can include DNA from extracellular particles and DNA that is not from extracellular particles. "Extracellular particle DNA", "EP DNA", and "EV DNA" (such terms may use cfDNA in place of DNA) refer to cell-free DNA from extracellular particles. Such EP DNA can include DNA within the membrane of the particle and DNA bound to the surface of the EP. EP-associated DNA can also refer to such EP DNA from within the EP and / or bound to the surface of the EP. "Particle-free DNA", "EP-free DNA", and "EV-free DNA" refer to cell-free DNA that is not from extracellular particles. Such terms can also be used more generally for RNA or nucleic acids.
[0018] "Clinically relevant DNA" can refer to DNA from a specific tissue source that is measured, for example, to determine the fractional concentration of such DNA or to classify the phenotype of a sample (e.g., plasma). An example of clinically relevant DNA is fetal DNA in maternal plasma.
[0019] The term "assay" generally refers to a technique for determining the characteristics of a nucleic acid or a sample of nucleic acids (e.g., a statistically significant number of nucleic acids), and the characteristics of the subject from which the sample was obtained. An assay (e.g., a first assay or a second assay) generally refers to a technique for determining the amount of nucleic acid in a sample, the genomic identity of the nucleic acid in the sample, the copy number variation of the nucleic acid in the sample, the methylation state of the nucleic acid in the sample, the fragment size distribution of the nucleic acid in the sample, the mutation state of the nucleic acid in the sample, or the fragmentation pattern of the nucleic acid in the sample. Any assay known to those of ordinary skill in the art can be used to detect any of the nucleic acid characteristics referred to herein. Nucleic acid characteristics include sequence, amount, genomic identity, copy number, methylation state at one or more nucleotide positions, nucleic acid size, mutation of the nucleic acid at one or more nucleotide positions, and the pattern of fragmentation of the nucleic acid (e.g., the nucleotide positions at which the nucleic acid fragments). The term "assay" can be used interchangeably with the term "method". An assay or method can have a particular sensitivity and / or specificity (e.g., based on the selection of one or more cut-off values), and their relative usefulness as diagnostic tools can be measured using receiver operating characteristic (ROC) curve area under the curve (AUC) statistics.
[0020] "Array read" refers to a strand of nucleotides obtained from any part or all of a nucleic acid molecule. For example, an array read can be a short strand of nucleotides (e.g., 20 to 150 nucleotides) sequenced from a nucleic acid fragment, a short strand of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. An array read can be a long nucleotide strand (e.g., hundreds or thousands of nucleotides) sequenced from a nucleic acid fragment. Array reads can be obtained in various ways, for example, using sequencing techniques or using probes, such as hybridization arrays or capture probes that can be used in microarrays, or single primers or amplification techniques such as polymerase chain reaction (PCR) or linear amplification using isothermal amplification. Exemplary sequencing techniques include massively parallel sequencing, targeted sequencing, Sanger sequencing, sequencing by ligation, ion semiconductor sequencing, and single molecule sequencing (e.g., using nanopores or single molecule real-time sequencing (e.g., from Pacific Biosciences)). Such sequencing can be random sequencing or targeted sequencing (e.g., by using capture probes that hybridize to a specific region or by amplifying a specific region, both of which enrich such regions). Exemplary PCR techniques include real-time PCR and digital PCR (e.g., droplet digital PCR). As part of the analysis of a biological sample, a statistically significant number of array reads can be analyzed, for example, at least 1,000 array reads can be analyzed. Other examples include analyzing at least 5,000, 10,000, or 50,000, or 100,000, or 500,000, or 1,000,000, or 5,000,000 or more array reads.
[0021] "Single molecule sequencing" refers to sequencing of a single template DNA molecule to obtain sequence reads without the need to interpret the base sequence information from the cloned copy of the template DNA molecule. Single molecule sequencing can sequence the entire molecule or only a part of the DNA molecule. Most of the DNA molecule can be sequenced, for example, more than 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%. The sequence reads (or reads from both ends) can be aligned to a reference genome. When both ends are aligned (e.g., as part of the read of the entire fragment or in the case of paired ends), greater accuracy can be achieved in the alignment and the length of the fragment can be obtained.
[0022] The term "allele" refers to alternative DNA sequences at the same physical genomic locus, which may or may not result in different phenotypic traits. In any particular diploid organism that has two copies of each chromosome (except for the sex chromosomes in a male human subject), the genotype of each gene contains a pair of alleles present at that locus, which are the same in homozygotes and different in heterozygotes. A population or species of organisms typically contains multiple alleles at each locus among the various individuals. A genomic locus at which more than one allele is found in the population is called a polymorphic site. Allelic variation at a locus can be measured as the number of alleles present (i.e., the degree of polymorphism) or the proportion of heterozygotes in the population (i.e., the heterozygosity rate). As used herein, the term "polymorphism" refers to any variation between individuals in the human genome, regardless of its frequency. Examples of such variations include, but are not limited to, single nucleotide polymorphisms, simple tandem repeat polymorphisms, insertion-deletion polymorphisms, mutations (which can cause diseases), and copy number variations. The term "haplotype" can refer to a combination of alleles or epigenetic markers (e.g., methylation) at multiple loci that are transmitted together on the same chromosome or chromosomal region. A haplotype can refer to as few as a pair of loci, or a chromosomal region, or an entire chromosome or chromosomal arm.
[0023] As used herein, the term "locus" or its plural form "loci" is any length location or address of nucleotides (or base pairs). Loci can vary across the genome.
[0024] The term "fractionated fetal DNA concentration" is used interchangeably with the terms "proportion of fetal DNA" and "fraction of fetal DNA", and refers to the proportion of fetal DNA molecules present in a biological sample derived from a fetus (e.g., a maternal plasma or serum sample) (Lo et al, Am J Hum Genet. 1998, 62:768-775, Lun et al., Clin Chem. 2008:54:1664-1672).
[0025] The terms "size profile" and "size distribution" generally relate to the sizes of DNA fragments in a biological sample. A size profile can be a histogram providing the distribution of amounts of DNA fragments of various sizes. Various statistical parameters (also referred to as size parameters or simply parameters) can distinguish one size profile from another. One parameter is the percentage of DNA fragments of a particular size or size range relative to all DNA fragments, or relative to DNA fragments of another size or range of sizes.
[0026] A "calibration sample" can correspond to a biological sample, and the fractional concentration or other measurable value of its clinically relevant DNA (e.g., fetal-specific DNA fraction) is known or determined via a calibration method, e.g., using tissue-specific alleles in pregnancy, such that alleles present in the fetal genome but not in the maternal genome can be used as markers for the fetus. As another example, a calibration sample can correspond to a sample for which calibration values for other characteristics are determined, and such other characteristics can be used to estimate the fractional concentration (or other measurable value).
[0027] "Calibration data points" include "calibration values" and measured or known fractional concentrations of clinically relevant DNA (e.g., DNA of a specific tissue type). Calibration values can be determined from the relative frequencies (e.g., aggregated values) determined for calibration samples where the fractional concentration of clinically relevant DNA is known. Calibration data points can be defined in various ways, for example, as discrete points or as a calibration function (also called a calibration curve or calibration surface). The calibration function can be derived from additional mathematical transformations of the calibration data points.
[0028] As used herein, the term "parameter" means a numerical value that characterizes a quantitative data set and / or a numerical relationship between quantitative data sets. For example, the ratio (or a function of the ratio) between a first amount of a first nucleic acid sequence and a second amount of a second nucleic acid sequence is a parameter.
[0029] "Separation value" corresponds to a difference or ratio that includes two values, for example, two fractional contributions, two size values / parameters, two methylation levels, or two numbers. A separation value is an example of a parameter. A separation value can be a simple difference or ratio. As an example, the direct ratio of x / y is a separation value, as is x / (x + y). A separation value can include other factors, such as multiplicative factors. As another example, a difference or ratio of functions of values, such as the difference or ratio of the natural logarithm (ln) of two values, can be used. A separation value can include differences and ratios. By comparing a separation value to a threshold, it can be determined whether the separation between two values is statistically significant.
[0030] "DNA methylation" in a mammalian genome typically refers to the addition of a methyl group to the 5' carbon of a cytosine residue in a CpG dinucleotide (i.e., 5-methylcytosine). DNA methylation can occur at cytosine in other contexts, such as CHG and CHH, where H is adenine, cytosine, or thymine. Cytosine methylation can also be in the form of 5-hydroxymethylcytosine. Non-cytosine methylation, such as N6-methyladenine, has also been reported.
[0031] The "methylation level" is, for example, an example of the relative abundance of methylated DNA molecules (e.g., at specific sites) and other DNA molecules (e.g., all other DNA molecules at specific sites or only unmethylated DNA molecules). The amount of the other DNA molecules can function as a normalization factor. As another example, the intensity of methylated DNA molecules (e.g., fluorescence or electric field strength) relative to the intensity of all or unmethylated DNA molecules can be determined. The relative abundance can also include the intensity per volume. The methylation level can be determined using methylation recognition assays such as methylation recognition sequence determination or PCR. Exemplary methylation recognition sequence determination can include, for example, bisulfite sequencing or single molecule techniques using nanopores or single molecule real-time sequencing as described in U.S. Publication No. 2021 / 0047679-A1.
[0032] The "methylation pattern" refers to a series of methylation states at multiple sites of a fragment, genome, or sample (e.g., including a specific tissue type). The methylation state at a site can be unmethylated (U) or methylated (M). In the case of a sample or genome, the methylation state can be a percentage. A reference methylation pattern can be designated as methylated when the methylation level at a site exceeds a specified threshold (e.g., 70%, 75%, 80%, 85%, 90%, 95%, or 99%). A reference methylation pattern can be designated as unmethylated when the methylation level at a site is less than a specified threshold (e.g., 30%, 25%, 20%, 15%, 10%, 5%, or 1%). Thus, the methylation pattern of a fragment (a series of M and U at sites) can be compared and matched with the reference methylation pattern of fetal tissue. Optionally, the reference methylation patterns of various tissues can be obtained from single molecule sequencing that expresses as a methylation pattern across individual molecules, and the methylation state can be binary (0 or 1 representing the unmethylated state and the methylated state, respectively).
[0033] As used herein, the term "classification" refers to any number or other character related to a particular characteristic of a sample. For example, a "+" symbol (or the word "positive") may indicate that the sample is classified as having a deletion or amplification. The classification can be binary (e.g., positive or negative) or can have more levels of classification (e.g., a scale of 1-10 or 0-1).
[0034] The terms "cut-off" and "threshold" refer to a predetermined number used in an operation. For example, the cut-off size may refer to the size above which fragments are excluded. The threshold may be a value above or below which a particular classification applies. Either of these terms can be used in either of these situations. The cut-off or threshold can be a "reference value", or can represent a particular classification, or can be derived from a reference value that distinguishes two or more classifications. The cut-off can be determined in advance, with or without reference to the characteristics of the sample or subject. For example, the cut-off can be selected based on the age or gender of the subject being examined. The cut-off can be selected after and based on the output of the test data. For example, a particular cut-off can be used when the sequencing of a sample reaches a certain depth. As another example, a reference subject having known classifications of one or more pathologies and measured characteristic values (e.g., methylation level, statistical size value, or number) can be used to determine a reference level to distinguish different pathologies and / or classifications of pathologies (e.g., whether the subject has a pathology). The reference value can be selected as representing a value that lies between one classification (e.g., an average value) or two clusters of measurement criteria (e.g., selected to achieve a desired sensitivity and specificity). As another example, the reference value can be determined based on a statistical simulation of the sample. Either of these terms can be used in either of these situations. Such reference values can be determined in various ways, as will be understood by those skilled in the art. For example, the measurement criteria can be determined for two different cohorts of subjects having different known classifications, and the reference value can be selected as representing a value that lies between one classification (e.g., an average value) or two clusters of measurement criteria (e.g., selected to achieve a desired sensitivity and specificity). As another example, the reference value can be determined based on a statistical simulation of the sample. Particular values such as cut-offs, thresholds, references, etc. can be determined based on the desired accuracy (e.g., sensitivity and specificity).
[0035] As used herein, the terms "array imbalance" or "abnormality" mean any significant deviation defined by at least one cut-off value in the amount of a clinically relevant chromosomal region from a reference amount in a pregnant woman's maternal plasma DNA. Array imbalances can include chromosomal dosage imbalances, allelic imbalances, variant dosage imbalances, copy number imbalances, haplotype dosage imbalances, and other similar imbalances.
[0036] "Fetal genomic characteristics" can refer to characteristics of fetal DNA, such as characteristics of fetal DNA fragments and / or the fetal genome. The genomic characteristics can be genetic and / or epigenetic. By way of example, genomic characteristics can include array imbalances, genotypes (e.g., genetic alleles), haplotypes (e.g., genetic haplotypes), variants (e.g., variant alleles), and methylation levels (e.g., at specific sites that can be inferred based on gene imprinting). Such characteristics can be determined by analyzing DNA in a biological sample from a pregnant woman.
[0037] "Genomic characteristics of pregnancy" can be pregnancy-related disorders. "Pregnancy-related disorders" include any disorder characterized by abnormal relative expression levels of genes in maternal and / or fetal tissues and / or by abnormal clinical characteristics in the mother and / or fetus. These disorders include, but are not limited to, hypertension, gestational diabetes, infections, preterm labor, pregnancy loss / miscarriage, fetal growth restriction (FGR), preeclampsia (Kaartokallio et al. Sci Rep. 2015;5:14107; Medina-Bastidas et al. Int J Mol Sci. 2020;21:3597), intrauterine fetal growth retardation (Faxen et al. Am J Perinatol. 1998;15:9-13; Medina-Bastidas et al. Int J Mol Sci. 2020;21:3597), invasive placentation, preterm birth (Enquobahrie et al. BMC Pregnancy Childbirth. 2009;9:56), neonatal hemolytic disease, placental insufficiency (Kelly et al. Endocrinology. 2017;158:743-755), fetal hydrops (Magor et al. Blood. 2015;125:2405-17), fetal malformations (Slonim et al. Proc Natl Acad Sci USA. 2009;106:9425-9), HELLP syndrome (Dijk et al. J Clin Invest. 2012;122:4003-4011), systemic lupus erythematosus (Hong et al. JExpMed. 2019;216:1154-1169), and other maternal immune disorders.
[0038] The term "machine learning model" may include a model based on predicting test data using sample data (e.g., training data), and thus may include supervised learning. Machine learning models are often developed using a computer or a processor. A machine learning model may include a statistical model.
[0039] The terms "about" or "approximately" can mean within an acceptable error range for a particular value determined by one of ordinary skill in the art, which depends in part on the limitations of the measuring system, i.e., how the value is measured or determined. For example, "about" can mean within one standard deviation, or more than one standard deviation, in accordance with the conventions within the relevant art. Alternatively, "about" can mean within a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, especially with respect to living systems or biological processes, the terms "about" or "approximately" can mean within one order of magnitude, within fivefold, more preferably within twofold of a value. When a particular value is recited in the present application and claims, unless otherwise stated, the term "about" meaning within the acceptable error range of the particular value should be assumed. The term "about" can have the meaning generally understood by one of ordinary skill in the art. The term "about" can refer to ±10%. The term "about" can refer to ±5%.
[0040] When a range of values is provided, unless the context clearly indicates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit between the upper and lower limits of that range, is also specifically disclosed. Each smaller range between any of the recited values or intervening values within the recited range and any other recited value or intervening value within the recited range is included within the embodiments of the present disclosure. The upper and lower limits of these smaller ranges can be independently included within or excluded from the range, and each range in which either limit, neither limit, or both limits of the smaller range are included is also included within the present disclosure subject to any specifically excluded limits within the recited range. When the recited range includes one or both of the limits, ranges excluding either or both of these included limits are also included within the present disclosure.
[0041] Standard abbreviations can be used, such as bp: base pair, kb: kilobase, pi: picoliter, s or sec: second, min: minute, h or hr: hour, aa: amino acid, nt: nucleotide, etc.
[0042] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present disclosure, some potential and exemplary methods and materials are described herein.
[0043] Detailed Description As described above, NIPT may suffer from problems due to low fetal fraction in some samples. Apart from the low fetal DNA fraction, the relatively fragmented nature of cell-free DNA would be another potential limitation of NIPT in certain situations. For example, short DNA molecules make it technically difficult to directly construct fetal genetic / epigenetic haplotypes from maternal plasma. The length of cell-free DNA has been shown to be mostly below 200 bp by next-generation short-read sequencing (Illumina) (Lo et al. Sci Transl Med. 2010;2:61ra91). Sequencing short plasma cell-free DNA would not be efficient for analyzing genetics and / or epigenetics in a haplotype pattern. This is because single nucleotide polymorphisms (SNPs) or CpG sites are typically separated from the nearest SNP or CpG site by only hundreds or thousands of base pairs. Thus, NIPT also suffers from problems due to the short size of the typically used cell-free DNA fragments. Typical NIPT analysis is of cell-free DNA. An improved approach would be to simultaneously obtain long DNA molecules and enable enhancement of the fetal signal in NIPT.
[0044] Instead of using free-floating self-free DNA molecules that do not contain any particles (vesicles), embodiments can use self-free DNA molecules within particles. The use of such a specific type of self-free DNA in particles (also referred to as particle cfDNA) enables the ability to capture and use long fetal DNA fragments and potentially the ability to concentrate samples for long fetal DNA fragments, i.e., the ability to increase the percentage of long DNA in the sample. The selection of particle DNA fragments having a size larger than the size threshold can increase the fetal DNA fraction.
[0045] Some embodiments can perform certain purification steps for concentrating particle DNA, such as physical separation (such as filtration or centrifugation), washing with an ionic solution (such as saline), and / or nuclease treatment. Alone or in combination with the selection of long DNA (i.e., larger than the size threshold), the purification can result in an increase in the fetal DNA fraction, thereby enabling a higher accuracy and / or more efficient assay (e.g., the same accuracy can be achieved using a smaller sample). This is possible because any statistical analysis involving fetal DNA (e.g., changes from expected / normal values) can be more easily detected because the changes are more prominent with a higher fetal fraction.
[0046] The analysis of long DNA fragments can be enhanced or enabled by using long-read sequencing techniques such as nanopore sequencing (e.g., Oxford Nanopore Technologies) and single-molecule real-time sequencing (e.g., Pacific Biosciences), synthetic long-read sequencing (Illumina), and linked-read technologies (10X genomics, Tell-seq), the latter two of which involve ligating a set of short DNA fragments as if they were derived from longer fragments. Additionally or alternatively, long DNA molecules can be analyzed by fragmenting them and then using short-read sequencing techniques.
[0047] I. Extracellular Particles The presence of membrane-bound or non-membrane-bound extracellular particles (EPs) in body fluids (e.g., plasma) has been previously reported (Malkin et al. Cell Death Dis. 2020;11:584). Cells can release such extracellular particles in various ways. For example, during apoptosis, cells release apoptotic bodies, which are a type of large extracellular vesicle. Some active release processes such as secretion generate microvesicles. Exosomes, which are a major cause of small EPs, have different ways of forming membrane vesicles using intracellular membranes instead of the cell membrane. Because of the different ways of forming EVs, their sizes vary greatly.
[0048] Most studies on EPs have focused on mRNA and miRNA (Zhou et al. Sig Transduct Target Ther. 2020;5:144). A practically meaningful approach based on EP-related DNA in the clinical context of NIPT is still not available. The sizes of EPs vary greatly, with diameters ranging from a few nanometers to a few micrometers. These particles can be broadly classified into nanoparticles (e.g., exosomes), microparticles (microvesicles), and apoptotic bodies according to their diameter sizes. Nanoparticles are typically referred to as EPs smaller than 100 nm, microparticles are usually those in the range of 100 nm - 1 μm, and apoptotic bodies are usually those with a diameter size of 1 μm - 5 μm. In a less accurate way, EPs can be roughly separated into two classes, namely, large EPs (>=200 nm) (LEPs) and small EPs (<200 nm) (SEPs). The intracellular origins of LEPs and SEPs are different (e.g., LEPs are formed by using the cell membrane, while SEPs are formed by intracellular membranes or proteins). Therefore, the genetic information associated with them can be handled differently.
[0049] Several groups have attempted to test the potential use of LEP-related DNA in the plasma of pregnant women. In an initial study, Bischoff et al. reported that the fetal DNA fraction showed an increase by analyzing DNA from the nucleic acid-positive acellular particle fraction sorted by flow cytometry (Bischoff et al. Hum Reprod Update 2005;11:59-67). This study quantified DYS1 (ChrY) and GAPDH sequences using real-time PCR to measure the fetal fraction in male pregnancies, but did not perform any analysis using the positive acellular particle fraction.
[0050] Orozsco et al. (Orozco et al. Placenta. 2009;30:10.; Goswami et al. Placenta. 2006;27:1.) demonstrated that placenta-derived DNA-related LEP (leukocyte antigen G positive (HLA-G+) or placental alkaline phosphatase positive (PLAP+)) was significantly increased in the maternal plasma of pregnant subjects compared to plasma from non-pregnant controls. Orozsco et al. detected placental LEP using antibodies and PicoGreen (a double-stranded DNA fluorescent dye), but were unable to elucidate the genetic and epigenetic information of fetal DNA molecules. Furthermore, both studies were based on flow cytometry sorting, which is only suitable for the analysis of LEP with a diameter size >1μm, thus resulting in low-resolution LEP separation.
[0051] However, in recent reports using array determination to analyze aneuploidy, the results of EP-related DNA were inferior. In this other report, ultra-parallel sequencing of EP-related DNA in maternal plasma was used to attempt to detect fetal chromosomal aneuploidy and single gene disorders (Zhang et al. BMC Med Genomic. 2019;12:151). However, the analysis of EP-related DNA was shown to be inferior to the analysis of normal cell-free DNA (i.e., particle-free cfDNA). The fetal DNA fraction in EP-related DNA was 2-fold lower than the fetal DNA fraction in plasma cell-free DNA (Zhang et al. BMC Med Genomic. 2019;12:151). Furthermore, the length of EP-related DNA was shorter than that of cfDNA (median size: 152.4 bp vs. 168.5 bp), and thus each EP-related DNA fragment would provide even less information than cfDNA. Such results suggested that EP is not beneficial for the use of long DNA fragments, provides a lower fetal DNA fraction, and thus is not beneficial for performing NIPT.
[0052] More recently, Lucas Brandon Edelman disclosed a patent application (WO2020 / 0002862A1) regarding a method for analyzing circulating microparticles that briefly discussed potential applications in NIPT without actual examples and disclosed implementation steps. The technology presented in Lucas Brandon Edelman's disclosure focused on barcoding DNA molecules inside the microparticles, enabling tracking of whether two or more DNA molecules are derived from the same microparticle. The concept of the technology is similar to the "linked-read technology" developed by 10x Genomics (Hui et al. Clin Chem. 2017;63:513-524). However, the disclosure by Lucas Brandon Edelman did not select a specific subpopulation of microparticles based on the physical and / or biological properties of the microparticles to improve the performance of NIPT or for the selection of a subpopulation of nucleic acid molecules.
[0053] Overall, there is still a lack of a practically meaningful approach. The present disclosure reports, for example, a new method that can selectively analyze a subset of extracellular particles that simultaneously enrich a DNA molecule of interest (e.g., a fetal DNA molecule) and a long DNA molecule by selecting long DNA molecules in which fetal DNA is enriched. Surprisingly, a fetal fraction higher than 50% can be achieved according to the techniques disclosed herein. These methods included sequencing DNA molecules associated with extracellular particles and analyzing genetic and / or epigenetic information that can substantially enhance the diagnostic power of NIPT. The present disclosure is beneficial to groups at risk of low fetal DNA fraction, which may be caused, for example, by a high body mass index (Hui et al. Prenatal Diagnosis 2020;40:155-163). The disclosed techniques may also enable NIPT to be performed earlier than customarily recommended by many authorities, for example, at 10 weeks.
[0054] II. Workflow for EP isolation The present disclosure provides various techniques for obtaining EP DNA (e.g., DNA contained in the EP as opposed to DNA bound to the outside of the EP) using one or more purification steps that can provide particles of desired size and content. The results in later sections show that certain purification and / or in silico techniques consistently increase the fetal fraction above 40% and provide surprising results in the ability to obtain long DNA fragments, which can enable new functionality, for example, to determine haplotypes in a more efficient and accurate manner. Various experimental procedures can be used to obtain extracellular particles (EPs) of potentially specific sizes.
[0055] Figure 1 shows a first exemplary workflow 100 for EP separation and analysis. As shown, the blood sample 102 in the sample holder is centrifuged at 1600 g for 10 minutes, which is performed twice. This first centrifugation step generates a pellet at the bottom of the vial, where the pellet contains live and dead cells. After removing the cell pellet, an optional filtration step 106 can filter the residual material (supernatant) (e.g., using a 5 μm filter) to ensure that cells do not proceed to the next step. This intermediate supernatant (plasma) after filtration contains LEV DNA but is significantly diluted with cell-free DNA and SEV DNA. A typical NIPT test is based on the liquid fraction, i.e., the supernatant from 1600 g × 2 (twice) for 10 minutes each, or 1600 g for 10 minutes, + 16,000 g for 10 minutes. When collecting plasma at 1600 g for 10 minutes (e.g., to remove cells) + 16,000 g for 10 minutes, the LEV portion is mostly removed and the remaining plasma can be considered LEV-free DNA. Other centrifugation protocols with different forces (rotation speeds), times, and numbers of centrifugation steps can vary.
[0056] In centrifugation step 108, the filtered supernatant can be centrifuged at 20,000 g for 40 minutes to collect a pellet enriched with LEV. The LEV pellet can be collected directly and contains a certain plasma carry-over labeled as LEV without further processing corresponding to sample 110. The remaining supernatant contains SEV and particle-free DNA. As another example, ion washing (e.g., using phosphate-buffered saline, PBS) can be used to provide LEV with washing corresponding to sample 120. The washing can remove some particle-free DNA. After ion washing, the sample can be further centrifuged (e.g., at 20,000 g for 40 minutes) to further separate and extract the LEV.
[0057] As a further process, after performing ion washing, nuclease treatment (e.g., with DNase I) can be applied. The nuclease treatment can further degrade nucleic acids that are not within the membrane of the LEV, thereby enabling the removal of such particle-free DNA and resulting in sample 130. Such DNA bound to the outside of the EV may be EV-associated DNA, but the purpose of purification may be to remove such EV-associated DNA and obtain a highly concentrated sample for the DNA within the EV membrane. Thus, with more processing, the outside DNA can be increasingly removed. The DNA in any of the samples can be isolated for sequencing.
[0058] Typically, DNA in plasma does not undergo physical fragmentation because the DNA is naturally fragmented. However, it has been recognized that long DNA can occur within vesicles. For sequencing such long DNA (e.g., greater than 600 bp) on certain platforms, such as Illumina or other short-read sequencing platforms, some embodiments can perform a physical fragmentation process so that such DNA can be sequenced. Exemplary fragmentation techniques can include mechanical shearing, enzymatic fragmentation such as Tn5 transposase-based tagging, DNASE1, DNASE1L3, and / or DFFB treatment, light, sonication, or chemical DNA fragmentation that uses a combination of heat and a divalent metal cation such as magnesium or zinc to break nucleic acids. In some embodiments, bisulfite treatment can be used to fragment DNA molecules. The level of fragmentation can be specified to shorten the average fragment length to less than a specified size (e.g., 600 bp), such as down to 200 bp. In one embodiment, long-read sequencing techniques such as single-molecule sequencing (e.g., using nanopores or single-molecule real-time sequencing (e.g., from Pacific Biosciences)) can be used. In addition to or instead of sequencing, probe-based techniques such as PCR can be used.
[0059] Bioinformatics analysis can be of various types and can include multiple steps. The analysis can be genetic and / or epigenetic. For example, sequencing can provide sequence reads that are aligned to a reference genome to determine the genomic location of the reads. Such sequence reads can be analyzed for various characteristics at certain positions, sites, or regions, such as the number, size of the DNA fragment, methylation level, terminal position within the genome, amount of overhang (jaggedness) at the ends of the fragment, and motifs at the ends of the fragment, such as 3-mers or 4-mers at the ends of DNA fragments. Such fragment end analysis can preferably be used when no separate physical fragmentation is performed. Using such characteristics, various abnormalities, pathologies, or disorders can be detected, including copy number aberrations, and sequence variations (including variations that can be single nucleotides or more), and haplotype inheritance.
[0060] Figure 2 shows a second exemplary workflow 200 for EP separation and analysis. Workflow 200 is similar to workflow 100. The exemplary method includes, but is not limited to, the following two aspects: (1) selecting a desired subset of EPs that enrich DNA molecules of fetal origin, and (2) performing genetic and / or epigenetic analysis of those selected DNA molecules. In the first aspect, the selection of EPs can be performed based on selecting EPs having their diameter sizes, for example, diameters of 200 nm to 5 μm (LEP) and <200 nm (SEP). As an example, such selection of EPs can be performed based on centrifugation and ultracentrifugation.
[0061] As shown, the procedure for obtaining LEP with washing and / or nuclease treatment is the same as the procedure for sample 120 and sample 130 for workflow 100. For the supernatant 208 containing SEP and particle-free DNA, filtration is performed (e.g., using a 0.22 micrometer filter), followed by centrifugation at 110,000 g for 4 hours to obtain sample 212. The liquid fraction of sample 212 can be used as the final supernatant (FSN) that mostly contains particle-free DNA. The pellet from sample 212 can be further processed (e.g., by ion washing and / or nuclease treatment) to obtain sample 214, and sample 214 can be centrifuged again at 110,000 g for 4 hours. The remaining pellet can concentrate SEP that can be extracted and analyzed, for example, as described below.
[0062] A. Purification of Samples for EP In various embodiments, EPs can be separated into populations of different sizes based on fractionation centrifugation or other physical separation techniques such as filtration or flow cytometry. Such physical separation can be carried out by any of the methods described herein. In one example, the collected blood can be subjected to two runs of centrifugation at 1,600 g for 10 minutes each to remove cells. The obtained supernatant can be filtered through a filter (e.g., a 5 μm mesh polycarbonate filter) to minimize cell contamination. The filtered supernatant can then be centrifuged at 20,000 g for 40 minutes to collect LEP. LEP can be treated, for example, with DNase I, either before or after ionic washing (e.g., PBS washing), thus removing DNA molecules outside the particles. The treatment can be only ionic washing. The DNase I and PBS-treated materials can be further centrifuged at 20,000 g for 40 minutes. The remaining plasma can be filtered, for example, using one or more 0.22 μm mesh polycarbonate filters and centrifuged at 110,000 g for 4 hours to collect SEP. SEP can be further washed with an ionic solution such as PBS (regardless of the presence or absence of DNase I treatment) and recentrifuged at 110,000 g for 4 hours to purify SEP. DNA from both LEP and SEP, and particle-free cfDNA from FSN can be subjected to DNA extraction and sequencing.
[0063] 1. Size Separation The diameter size selection of EPs can be performed in various ways, for example, using a plurality of techniques including, but not limited to, density gradient centrifugation, size exclusion chromatography, polymer-based precipitation (e.g., using ExoQuick), filtration (including wash filters to obtain EPs captured by the filter), ultrafiltration, tangential flow filtration, asymmetric flow field-flow fractionation, and affinity-based methods.
[0087] Since the sedimentation rate of particles depends on the particle size at a specific centrifugal force and liquid viscosity, the EPs collected at a specific centrifugal force and liquid viscosity will reflect the particle size. As shown in Figure 2, after removing cells by centrifugation at 1,600 g and a 5-μm filter, the EPs can be collected by centrifugation at 20,000 g, followed by DNase I treatment and washing with phosphate-buffered saline (PBS). The DNase I- and PBS-treated materials can be further centrifuged at 20,000 g to collect the LEP. The remaining plasma can be filtered through a 0.22-μm filter (e.g., a mesh polycarbonate filter) and centrifuged at 110,000 g to collect the supernatant (e.g., the final supernatant (FSN)) in which particle-free cfDNA molecules are concentrated.
[0064] The particles from the previous centrifugation at 110,000 g can be washed with an anionic solution such as PBS (regardless of the presence or absence of DNase I treatment) and recentrifuged at 110,000 g to collect the SEP. Thus, as separate fractions from the above procedure, the LEP, SEP, and FSN can be obtained. The corresponding DNA molecules, namely, LEP-related DNA, SEP-related DNA, and particle-free cfDNA, can be extracted by a DNA extraction kit (e.g., the QIAamp Circulating Nucleic Acid Kit (QIAGEN)).
[0065] As various examples, the target diameter size of EP can include, but is not limited to, 30 nm to 100 nm, 30 nm to 150 nm, 30 nm to 200 nm, 100 nm to 1 μm, 100 nm to 3 μm, 100 nm to 5 μm, 1 μm to 3 μm, 1 μm to 5 μm, or other combinations of diameters. Different centrifugal forces can be used according to the target diameter size of EP, for example, but not limited to, 100 g, 200 g, 300 g, 400 g, 500 g, 600 g, 700 g, 800 g, 900 g, 1,000 g, 1,100 g, 1,200 g, 1,300 g, 1,400 g, 1,500 g, 2,000 g, 3,000 g, 4,000 g, 5,000 g, 10,000 g, 20,000 g, 40,000 g, 50,000 g, 100,000 g, 200,000 g, 300,000 g, 400,000 g, 500,000 g, etc., or different combinations. Different durations of centrifugation can be used, for example, but not limited to, 1 second, 5 seconds, 10 seconds, 20 seconds, 30 seconds, 40 seconds, 50 seconds, 1 minute, 5 minutes, 10 minutes, 20 minutes, 30 minutes, 40 minutes, 50 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 10 hours, 20 hours, 1 day, 2 days, etc. Such exemplary values can be used in any of the exemplary techniques described herein.
[0066] As described above, filtration can also be used. Exemplary filter sizes are 2 μm, 3 μm, 4 μm, 5 μm, 6 μm, 7 μm, 8 μm, 9 μm, 10 μm, etc., corresponding to different filtration intensities. In some embodiments, the target LEV is less than 1 μm and potentially greater than 200 nm. Such exemplary values can be used in any of the exemplary techniques described herein.
[0067] Centrifugal force and filter size are two important parameters for obtaining a desired population of vesicles such as LEVs. In various embodiments, the centrifugal force for a second centrifugation (e.g., centrifugation step 108) can be 10,000 g, 11,000 g, 12,000 g, 13,000 g, 14,000 g, 15,000 g, 16,000 g, 17,000 g, 18,000 g, 19,000 g, 20,000 g, etc. to pellet and concentrate the LEVs after a first centrifugation at a centrifugal force such as 500 g, 600 g, 700 g, 800 g, 900 g, 1,000 g, 1,100 g, 1,200 g, 1,300 g, 1,400 g, 1,500 g, 1,600 g, 1,700 g, 1,800 g, 1,900 g, 2,000 g, 5,000 g, 10,000 g, etc., but is not limited thereto. Without limitation, a first filter step can be added to remove unwanted particles between any two centrifugations at a size such as 1 um, 2 um, 3 um, 4 um, 5 um, 6 um, 7 um, 8 um, 9 um, 10 um, etc. Without limitation, a second filter step can be added to further concentrate the desired particles between any two centrifugations at a size such as 0.1 um, 0.2 um, 0.3 um, 0.4 um, 0.5 um, 0.6 um, 0.7 um, 0.8 um, 0.9 um, 1 um, etc. The duration of centrifugation can be unlimited, such as 1 second, 5 seconds, 10 seconds, 20 seconds, 30 seconds, 40 seconds, 50 seconds, 1 minute, 5 minutes, 10 minutes, 20 minutes, 30 minutes, 40 minutes, 50 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 10 hours, 20 hours, 1 day, 2 days, etc. The order of centrifugation and filtration can be variable. The purity of the DNA associated with the LEVs can be further improved using ionic buffer washes (PBS washes) and / or enzymatic digestion (e.g., DNASE1).
[0068] 2. Concentration In certain embodiments, the desired population of EPs can be further enriched before, after, or without centrifugation, and in combination with centrifugation. Enrichment can be for DNA from a particular type of cell. For example, without limitation, immunoprecipitation-based, immunoaffinity-based, aptamer affinity-based, flow cytometry-based methods (e.g., fluorescence-activated cell sorting (FACS)), or microfluidics-based techniques can be used to select EPs from fetal tissues such as syncytiotrophoblasts using protein markers (e.g., syncytin-1 and placental alkaline phosphatase (PLAP)). When performing NIPT, enrichment for syncytiotrophoblast cells may be desirable because such cells are specific to the placenta and carry a surface protein marker (e.g., PLAP) that facilitates selection. For example, a fluorophore (e.g., PerCP) can be used to stain PLAP via its specific antibody.
[0069] Such identification of fetal-derived particles can be used to enrich samples for fetal DNA. Further, DNA from a given particle can be identified (e.g., barcoded) such that after fragmentation, small fragments from the same particle can be reassembled together to create a single long read. For example, sequence reads can be aligned to a reference genome, and if two reads are adjacent to each other (e.g., within 1, 2, 3, 4, or 5 bases) and from the same particle, they can be assumed to be from the same long fragment, thereby providing sequence reads greater than 600 bp. Such techniques can be referred to as linked-read sequencing.
[0070] 3. Processes for Further Purification Various processes can be performed at various times, e.g., before or after physical separation techniques such as centrifugation. Such processes can be performed individually or together, e.g., continuously, and potentially applied more than once, with other processes or separation steps in between.
[0071] One treatment is ion washing. The washing buffer (e.g., phosphate buffered saline) can have an osmotic pressure, ionic strength, and / or pH similar to that of plasma. Such a treatment can remove particle-free nucleic acids in the sample and / or bound to the outside of the EVs. In other embodiments, other solutions including, but not limited to, saline, HEPES (4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid), MOPS (3-(N-morpholino)propanesulfonic acid), and TBS (Tris buffered saline) can be used to wash the sample (e.g., LEP or SEP).
[0072] Another exemplary treatment is nuclease treatment that can degrade particle-free nucleic acids in the sample and / or bound to the outside of the EVs. When such nucleic acids are degraded and removed from the surface of the EVs, they can be removed, for example, by washing or by a size selection process such as centrifugation.
[0073] As an example of nuclease treatment, DNase I treatment can be applied during the isolation of EPs to remove DNA outside the EPs. Other DNA nucleases include, but are not limited to, TREX1 (3 prime repair exonuclease 1), AEN (apoptosis-enhancing nuclease), EXO1 (exonuclease 1), DNASE2 (deoxyribonuclease 2), ENDOG (endonuclease G), APEX1 (apurinic / apyrimidinic endodeoxyribonuclease 1), FEN1 (flap structure-specific endonuclease 1), DNASE1L1 (deoxyribonuclease 1-like 1), DNASE1L2 (deoxyribonuclease 1-like 2), and EXOG (exo / endonuclease G).
[0074] B. Analysis For a second aspect of this exemplary workflow of FIGS. 1 and 2, DNA isolated from different EP sources can then be analyzed to reveal inner genetic and / or epigenetic information, for example, using PCR (including real-time PCR or digital PCR) or a sequencing platform. After the procedure of isolating LEP or SEP, the membrane on the particle can be disrupted, thereby exposing the DNA fragments. The fragmented DNA can then be analyzed. Such analysis can utilize the enrichment of long DNA fragments and / or the increase in the fetal DNA fraction.
[0075] In the present disclosure, EP can be used to enrich long DNA molecules as it is envisioned that the protective environment of EP prevents nuclease degradation of the long DNA molecules associated therewith (e.g., reduces the accessibility of DNA nucleases). Long DNA molecules can be defined as having a size greater than a size threshold, such as, but not limited to, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 100 kb, 500 kb, 1 Mb, etc.
[0076] DNA fragments of a desired size range can be selected physically (e.g., using electrophoresis) or in silico (e.g., by determining the length of the DNA fragments and selecting the fragments within the size range). Electrophoresis can be performed before genomic analysis, for example, before sequencing or PCR analysis.
[0077] To analyze long DNA on a short read platform, DNA shearing (e.g., physically, enzymatically, or chemically) can be performed on the DNA collected from the selected EP so that the long DNA molecules present in the EP can be sequenced by short read sequencing techniques (e.g., Illumina). Alternatively, long read sequencing techniques, including but not limited to nanopore sequencing (e.g., Oxford Nanopore Technologies) and single molecule real-time sequencing (e.g., Pacific Biosciences), can be performed on the DNA collected from the selected EP. Analysis of the DNA molecules can include, but is not limited to, counting, size profiling, fragment end analysis, nucleotide variant analysis, and epigenetic analysis, or other techniques described herein.
[0078] As shown and described herein, some of the techniques for analyzing EPs can not only enrich DNA molecules of fetal origin but also enable long DNA molecules, thus facilitating genetic and / or epigenetic analysis. Previous reports were unable to achieve these objectives, for example, for the following reasons: the separation of the desired EPs to enrich tissue-specific DNA molecules was not established, and the long DNA molecules inside the EPs were not effectively analyzed. The techniques described herein can use long read sequencing techniques to evaluate the DNA inside the EP or the artificially fragmented long DNA molecules inside the EP so that short read sequencing may be suitable for evaluating EP-related long DNA molecules.
[0079] Additionally, certain methods do not effectively remove contaminating DNA outside the EP. Certain embodiments can remove DNA outside the EP by combining DNase I treatment and PBS washing, followed by recentrifugation. The improved efficiency may be due to the fact that DNase I digestion is more efficient with naked DNA than with DNA protected by histones, and further saline (e.g., PBS) washing can remove remaining nucleosomes after DNase I treatment.
[0080] Significantly and surprisingly high concentrations can provide higher statistical accuracy, for example, because the background noise of maternal DNA is reduced to a minority in some cases. Additionally or alternatively, any assay can be more efficient (e.g., use less sample or use less reagent) because fewer DNA fragments are required to be analyzed for the same level of accuracy. For example, with an increase in the fetal DNA fraction, most of the DNA fragments are from fetal tissues with imbalances, so sequence imbalances (or other genomic characteristics) can be detected earlier.
[0081] Furthermore, the analysis can benefit from the use of long DNA fragments. Such long DNA fragments can be useful for haplotyping because heterozygous loci from multiple fragments overlap. The fetal genome can be reconstructed in this way. Additionally, long DNA molecules carry more CpG sites, facilitating the determination of plasma DNA molecules of placental origin based on their respective methylation patterns. Thus, the fetal methylome can be reconstructed using the methylation patterns along each long DNA molecule.
[0082] Additionally, since various sub-samples (e.g., LEP, SEP, and FSN) can be generated from a single blood sample, measurements that use all of the samples can be combined or compared. For example, measurements of genomic and / or epigenetic characteristics can be performed on each sample, and a majority or all consensus in the determination can be used to determine the classification. In this way, sensitivity and specificity can be improved.
[0083] C. General Results Table 1 provides a summary of differently prepared LEP samples (different sample types) from blood samples of the same patient in the third trimester with a male fetus. As described above, three different LEP DNA samples were collected. In Table 1, the DNA concentration refers to the initial DNA concentration after isolation, i.e., how much of a particular type of DNA is present per 1 ml of plasma (generated by centrifuging at 1600 g 2x). The input refers to how much DNA was used in the library preparation. The total mapped reads refer to how much the DNA fragments were mapped to the human genome after sequencing.
Table 1
[0084] The DNA concentration decreases with further processing of PBS washing and further reduction in DNase I treatment. This indicates a decrease in DNA outside the LEV. Surprisingly, the number of mapped reads does not decrease as much as the decrease in DNA concentration and input DNA for preparing the library. Therefore, there is a higher proportion of usable DNA in the final sample after washing and nuclease treatment.
[0085] According to some reports, mitochondrial DNA is highly concentrated in LEP. For three different types of samples, the contributions of mitochondrial DNA and nuclear DNA in LEP DNA were analyzed. The main contributing factor is still nuclear DNA across all three samples, but its amount decreases with processing (e.g., PBS washing treatment and DNase I treatment). In the case of LEP samples without further processing, nuclear DNA is about 98% while mitochondrial DNA is about 2%. In the case of LEP samples with PBS washing, nuclear DNA is about 92% while mitochondrial DNA is about 8%. In the case of LEP samples with PBS washing and DNase I treatment, nuclear DNA is about 87% while mitochondrial DNA is about 13%. These results indicate the enrichment of DNA from LEP by further processing.
[0086] Table 2 shows the number of DNA fragments having alleles specific to the fetus and the number of DNA fragments having shared alleles (i.e., those shared between the mother and the fetus). The total number of fragments is across all loci and not only those having fetus-specific alleles. Table 2 shows an increase in the fetal fraction. The fetal fraction increased from 18% to 44% and further to 75% for the three samples. As shown by the differently processed samples, more non-LEV-related DNA was removed and more fetal DNA was obtained, indicating that the DNA within LEP is mainly fetal.
Table 2
[0087] III. Exemplary Techniques for Determining Fetal Fraction The data presented in this specification shows an increase in the fetal DNA fraction using samples purified for LEP, and an increase in the fraction for long DNA fragments. Various techniques can be used to determine the fetal DNA fraction. For example, fetal-specific markers can be used. Examples of fetal-specific markers include alleles or epigenetic markers such as methylation levels. Another example for measuring the fetal DNA fraction is to use size, as described, for example, in US Patent Publication No. 2013 / 0237431.
[0088] In the present disclosure, for illustrative purposes, DNA molecules obtained from LEP without PBS washing and DNase I treatment, LEP with PBS washing, LEP with PBS washing and DNase I treatment, SEP with PBS washing and DNase I treatment, and the final supernatant (FSN) were sequenced. Blood was collected from four pregnant women (third trimester: n = 3, first trimester: n = 1). According to an embodiment of the present disclosure, cell-free DNA molecules from FSN, DNA molecules from LEP and SEP were subjected to short-read sequencing (75bp x 2 paired-end mode, Illumina), and the median was 17.41 million paired-end sequencing reads (range: 6.01 million to 48.67 million).
[0089] The maternal buffy coat and placental tissue genotypes were obtained using a microarray-based genotyping technology (HumanOmni2.5 Genotyping Array, Illumina), and informative SNPs were identified (i.e., the mother was homozygous (represented as the AA genotype) and the fetus was heterozygous (represented as the AB genotype)). Fetal-specific DNA fragments were identified as DNA fragments carrying the fetal-specific allele at the informative SNP site. In this scenario, the B allele was fetal-specific, and the DNA fragment carrying the B allele was inferred to be derived from fetal tissue. The number of fetal-specific molecules (p) carrying the fetal-specific allele (B) was determined. The number of molecules (q) carrying the shared allele (A) was determined. In third-trimester cases, the fetal DNA fraction across cell-free DNA molecules from FSN, and DNA molecules from LEP and SEP was calculated as 2p / (p + q). * Calculated by 100%.
[0090] In cases where placental tissue genotype information was not available, the non-maternal DNA fraction was used to infer the fetal DNA fraction according to previously published methods (Jiang et al. NPJ Genom Med. 2016;1:16013 and US Publication No. 2017 / 0081720). The non-maternal DNA fraction was defined as the fraction of DNA molecules carrying alleles different from those of the mother.
[0091] Figure 3 shows the correlation between the fetal DNA fraction and the non-maternal DNA fraction. Such a correlation can be used to determine the fetal DNA fraction when placental tissue genotype information is not available, and the non-maternal DNA fraction can be determined using the percentage of non-maternal alleles detected. For example, the homozygous locus (AA) can be determined from maternal genotyping, and those that are not A are the non-maternal fraction. The fetal DNA fraction was determined using fetal-specific markers.
[0092] To convert the non-maternal DNA fraction to the fetal DNA fraction, a calibration curve 308 between the fetal DNA fraction and the non-maternal DNA fraction was determined using 12 third-trimester samples. The formula for converting the non-maternal DNA fraction to the fetal DNA fraction is shown below: F = X * 4.4484 - 2.4558. In the formula, F represents the fetal DNA fraction and X represents the non-maternal DNA fraction.
[0093] When determining the fetal DNA fraction, various techniques can use one or more first calibration data points, which can be obtained from one or more calibration samples having a known / measured fetal DNA fraction and a determined calibration value (e.g., size, methylation level, non-maternal fraction, etc.). In one embodiment, one or more calibration data points can be obtained. Each calibration data point specifies a clinically relevant DNA fraction concentration corresponding to the calibration value of a parameter (e.g., regarding size, methylation level, non-maternal fraction, etc.). The calibration function (curve) can be fitted to the plurality of calibration data points (e.g., by minimizing the least squares error) such that a newly measured calibration value can be input into the calibration function that outputs the estimated fetal DNA fraction.
[0094] IV. Fetal Fraction Analysis of Extracellular Vesicles For example, as described above, various procedures can be performed to purify samples for LEP and SEP and obtain different sample types (fractions) from a blood sample. The fetal fractions of different sample types were determined. Fragment size analysis was also performed. Surprisingly, long DNA was observed, contrary to what was seen in previous studies. Furthermore, an increase in the fetal fraction was seen among the long DNA, which was also surprising. Various NIPT techniques can advantageously use the increased fetal DNA fraction and long DNA fragments, as described herein.
[0095] A. Enriched Fetal Fraction in LEP The fetal fractions of different sample types were compared, and the fetal fractions of different treatments for the LEP sample type were compared.
[0096] 1. Fetal fraction in different parts of EP-related DNA samples Figure 4 shows the fetal DNA fractions in different EP-related DNA samples. This plot illustrates the fetal DNA contributions in different EP-related DNA samples. As shown here, LEP-related DNA shows substantial enrichment in DNA molecules of fetal origin, while the fetal DNA fraction of SEP-related DNA is slightly lower than that of FSN. Both the SEP sample and the LEP sample were treated with PBS washing and DNase I treatment.
[0097] As shown in Figure 4, the fetal DNA fraction was 77.00% in the DNA molecules obtained from LEP (i.e., LEP with PBS washing and DNase I treatment), which showed a 5.50-fold enrichment compared to SEP (i.e., SEP with PBS washing and DNase I treatment, fetal DNA fraction: 14.01%). These data suggested that extracellular particles separated by a certain centrifugation setting could concentrate fetal DNA molecules. In this case, the DNA molecules obtained from LEP had a higher fetal DNA fraction (fetal DNA fraction: 17.98%) than the cell-free DNA obtained from FSN. FSN was generally considered to be similar to plasma DNA prepared for NIPT. Such a high increase indicates that the embodiments of the present disclosure can simultaneously analyze a series of diameter size ranges of extracellular particles (e.g., LEP and SEP), and thus determine the optimal diameter size range of particles for concentrating target molecules (e.g., fetal DNA).
[0098] 2. Increase in fetal fraction in LEP with washing and DNase treatment Furthermore, it was demonstrated that removing the DNA outside the LEP helps to concentrate fetal DNA. Such DNA outside the LEP can be removed using physiological saline washing and / or nuclease treatment.
[0099] Figures 5A-5B show the enrichment of fetal DNA in LEP-related DNA in third-trimester pregnant women. Figure 5A shows the total fetal DNA fractions of (1) LEP without PBS washing and DNase I treatment, (2) LEP with PBS washing, (3) LEP with PBS washing and DNase I treatment, and (4) cell-free DNA obtained from FSN in Case 1 and Case 2. Figure 5B shows the fetal DNA fractions across different chromosomes of (1) LEP without PBS washing and DNase I treatment, (2) LEP with PBS washing, (3) LEP with PBS washing and DNase I treatment, and (4) cell-free DNA obtained from FSN in Case 1 and Case 2.
[0100] As shown in Figure 5A, plasma samples from two third-trimester pregnant women were analyzed. The fetal DNA fractions in LEP with PBS washing and DNase I treatment samples were 77.00% and 41.83% for Case 1 and Case 2, respectively. These fetal DNA fractions were 4.65-fold and 2.86-fold higher than those in LEP without PBS washing and DNase I treatment (Case 1: 16.57%, Case 2: 14.64%) and 4.28-fold and 2.33-fold higher than those in cell-free DNA obtained from FSN (Case 1: 17.98%, Case 2: 17.92%). These results indicated that LEP with PBS washing and DNase I treatment was beneficial for enriching fetal DNA from maternal plasma samples. In addition, such enrichment could be observed across the whole genome (Figure 5B). As shown in Figure 5B, the substantial increase in LEP samples with washing and nuclease treatment persisted for each chromosome.
[0101] The enrichment of fetal DNA fractions in LEP was also extended to first-trimester pregnant women.
[0102] Figures 6A-6B show the enrichment of fetal DNA in LEP-related DNA in first-trimester pregnant women. Figure 6A shows the overall fetal DNA fraction among (1) LEP without PBS washing and DNase I treatment, (2) LEP with PBS washing, (3) LEP with PBS washing and DNase I treatment, and (4) cell-free DNA obtained from FSN in one first-trimester case (Case 3). Figure 6B shows the fetal DNA fraction across different chromosomes among (1) LEP without PBS washing and DNase I treatment, (2) LEP with PBS washing and DNase I treatment, (3) LEP with PBS washing and DNase I treatment, and (4) cell-free DNA obtained from FSN in one first-trimester case (Case 3).
[0103] Analysis of first-trimester cases (12 weeks) revealed that the fetal DNA fraction of LEP with PBS washing and DNase I treatment samples was 38.25%, which was 4.18-fold higher than that of LEP without PBS washing and DNase I treatment (9.15%) and 3.89-fold higher than that of cell-free DNA obtained from FSN (9.83%) (Figure 6A). These results indicated that LEP with PBS washing and DNase I treatment was useful for enriching fetal DNA from maternal plasma samples even in the early pregnancy. In addition, such enrichment could be observed across the entire genome even in first-trimester cases (Figure 6B). This indicates that the enrichment is not biased towards one or two chromosomes but instead is applied across the entire genome.
[0104] B. Enrichment of Long DNA Fragments Furthermore, it is demonstrated that long fetal DNA fragments are present in LEP and corresponding reads can be obtained from LEP. Contrary to previous studies, it shows that long DNA actually exists in LEP. Short-read and long-read sequencing technologies are used to show the presence of long DNA fragments in LEP. Also, it shows that the enrichment of long DNA fragments from the fetus, for example, when only long DNA is analyzed, the fetal fraction increases. Also, the effect of the treatment on the fetal fraction is analyzed.
[0105] 1. Presence of Long DNA in LEP To determine whether long DNA fragments are present in LEP, electrophoretic measurements were carried out. Such measurements were carried out with and without fragmentation of DNA from the LEP sample. The samples were not washed or treated with nuclease. Fragmentation serves to show whether, using short-read sequencing, the resulting smaller fragments can be analyzed from the fragmentation of long DNA fragments. TapeStation High Sensitivity D1000 results (TapeStation, Agilent Technologies) were used for the electrophoretic measurements.
[0106] Figure 7 shows the presence of long DNA in LEP as revealed by mechanical shearing. The plots illustrate the TapeStation results of LEP-related DNA with and without mechanical shearing. The DNA concentration (represented by rectangles) of 50 - 600 bp was quantified and shown at the top of each lane. The reference scale of different sizes is shown on the left.
[0107] The amount (0.1 ng) of <600 bp DNA molecules obtained from LEP without mechanical shearing was much less than that with mechanical shearing (Covaris; 1.2 ng). This result indicates that long DNA is present inside the LEP and was fragmented by Covaris into a size range measurable by TapeStation HS D1000. The size range of the box (about 50 - 600 bp) corresponds to the size range that can be sequenced using a short-read platform. Fragmentation and increase of DNA fragments within this range indicate that these unexpectedly long fragments can be sequenced using the fragmentation step, thereby increasing the amount of DNA that can be analyzed and, in some cases, increasing the fetal fraction, as will be shown later.
[0108] 2. Enrichment of Long DNA in LEP Considering the presence of long DNA fragments in LEP, for example, to examine the increase in long DNA in LEP, the amount of long DNA fragments between sample types was compared. To study long DNA, fragmentation was used along with short-read sequencing. Even after fragmentation was performed, the resulting DNA was approximately 200 - 400 bp, as shown in Figure 7.
[0109] Figures 8A - 8C show the enrichment of long DNA in LEP-related DNA. In Figures 8A - 8B, the frequency refers to the percentage of DNA fragments in the sample that are below or above 200 bp. The two samples are FSN and LEP with PBS washing and DNase I treatment. The horizontal axis divides the data into two groups based on fragment size, namely, groups above and below 200 bp.
[0110] The plots show the enrichment of long DNA in LEP-related DNA. For LEP with PBS washing and DNase I digestion, 44.9% of the DNA molecules were longer than 200 bp (Figure 8A), while for FSN, only 4.4% of the DNA molecules were longer than 200 bp (Figure 8B). The percentage of long DNA molecules increased by approximately 10-fold. Furthermore, the total long reads increased: 2,697,610 (LEP with washing and nuclease digestion) versus 410,646 (FSN). Thus, even when mechanically sheared to a target size of 200 bp, a significant proportion of long molecules (i.e., 44.9% of DNA molecules > 200 bp) was still observed in the DNA obtained from LEP (Case 1), which was much higher than the cell-free DNA obtained from FSN (i.e., 4.4% of DNA molecules > 200 bp).
[0111] Figure 8C shows the size distribution of cell-free DNA from LEP-related DNA and FSN. The plot shows the percentage of DNA fragments in a sample that are of a particular size. The X-axis is the fragment length measured in bp. The size distribution of DNA in LEP was substantially longer than that in FSN. As can be seen, FSN has a sharp peak around 166 bp, which is typical of plasma. However, the processed LEP sample has a long tail with significant DNA fragments up to 400 bp, and this is the case even after the fragmentation step. Thus, the overall size profile of LEP-related DNA is shifted towards larger sizes compared to the cell-free DNA of FSN.
[0112] Considering that fragmentation was performed, these results suggested that these DNA molecules with lengths exceeding 200 bp were derived from even longer DNA molecules (e.g., several kilobases).
[0113] Figure 10A shows the size profiles of all DNA in various sample types corresponding to different treatments. The DNA was still fragmented, for example, using mechanical shearing, light, or sonication. As shown, the size distribution 1001 of DNA from LEV without further treatment remained similar to the typical distribution of plasma DNA. However, the size distribution 1002 (profile) of LEV with PBS washing, and the size distribution 1003 of LEV with PBS washing and DNase I treatment showed that the DNA inside LEP had an average longer length than the untreated sample. Since the DNA was fragmented to 200 bp, this size profile does not provide the original size and could be even longer. If the DNA fragment is shorter than 200 bp, it is not fragmented.
[0114] Figure 10B shows the size profiles of fetal DNA in various sample types corresponding to different treatments. Similar to the previous whole nuclear DNA size distribution, the size distribution of fetal DNA showed the same trend. Without further processing, the DNA size distribution 1011 remained similar to a typical plasma DNA distribution. The size distribution 1012 (profile) of LEV with PBS washing, and the size distribution 1013 of LEV with PBS washing and DNase I treatment showed that their inner DNA could have an average longer length than the untreated samples.
[0115] For both Figures 10A and 10B, many DNA fragments from the treated samples have longer sizes exceeding 200 BP. In contrast, the FSN in Figure 8C has few DNA fragments exceeding 200 BP. The distributions of LEP (especially the treated) and cell-free DNA are very different because the treated LEV samples do not peak at 166 bp.
[0116] Instead of using fragmentation and short-read sequencing platforms, LEP-related DNA can be sequenced with long-read sequencing technologies such as single molecule real-time sequencing (pool of two samples in the third trimester of pregnancy).
[0117] Figure 9 illustrates how single molecule real-time sequencing reveals the enrichment of long DNA in LEP-related DNA. The vertical axis shows the percentage of DNA fragments in the sample exceeding a certain size threshold. Three size thresholds: 200 bp, 600 bp, and 1000 bp are used. For each size threshold, two sample types: FSN and LEP with washing nuclease treatment were tested. LEP-related DNA showed a substantial increase in DNA molecules of sizes longer than 200 bp (87.67%), 600 bp (72.60%), and 1000 bp (49.32%) compared to cell-free DNA from FSN (>200 bp cell-free DNA ratio: 36.05%, >600 bp: 11.93%, >1000 bp: 6.87%).
[0118] As shown in Figure 9, the number of DNA molecules having a length >600 bp was substantially higher in LEP-related DNA (72.60%) compared to cell-free DNA from FSN (11.93%). This result confirmed the previous finding that a substantial proportion of the DNA present in LEP cannot be directly sequenced on the Illumina short-read sequencing platform. In addition, the LEP samples contained a much higher amount of DNA molecules having a length >1 kb (49.32%) compared to cell-free DNA from FSN (6.87%). These data further suggested that analyzing LEP-related DNA according to embodiments of the present disclosure can enrich fetal-origin molecules, obtain more long DNA molecules, and thus facilitate the improvement of NIPT.
[0119] The ability to obtain such a high percentage of long DNA fragments can provide various advantages. For example, the use of methylation information in CpG sites and / or variants in long DNA molecules will facilitate the determination of fetal maternal inheritance. Determine whether the observed DNA fragments from LEP are of fetal origin (e.g., using fetal-specific markers), and thereby determine whether any genetic / epigenetic changes linked to such DNA fragments are transmitted to the fetus, if present. In this way, the analysis of gene imprinting can be enabled using such long DNA fragments. Any use of the fetal-specific markers described herein can be carried out in various ways, such as genetic markers (e.g., sequence alleles) or epigenetic markers (e.g., methylation markers or fragmentation patterns such as terminal motifs or terminal positions).
[0120] Paramagnetic beads provide another way to analyze the length of DNA fragments. Based on solid-phase reversible immobilization technology, paramagnetic beads can be used to selectively concentrate nucleic acids based on DNA molecule size. Such beads include a polystyrene core, magnetite, and a carboxylate-modified polymer coating. DNA molecules selectively bind to the beads in the presence of PEG and salts depending on the concentration of polyethylene glycol (PEG) and salts during the reaction. PEG causes negatively charged DNA to bind to carboxyl groups on the bead surface, which are collected in the presence of a magnetic field. Molecules of the desired size are eluted from the magnetic beads using an elution buffer, such as 10 mM Tris-HCl, pH 8 buffer or water. The bead-to-sample volume ratio determines the size of the DNA molecules that can be obtained. The lower the bead-to-sample ratio, the longer the molecules retained on the beads.
[0121] Figure 11 shows that long LEP-related DNA can be concentrated with paramagnetic beads. The vertical axis shows the percentage of DNA fragments of a specific size using two different protocols, 0.8x and 1.2x. The horizontal axis divides the DNA fragments into two size ranges (greater than 200 bp and less than 200 bp). The left plot is for all DNA in the sample, while the right plot is for fetal DNA only. The fetal DNA used in the right plot was identified using a fetal-specific marker. The LEP samples were washed and treated with nuclease.
[0122] As shown in Figure 11, when using a bead-to-sample ratio of 0.8X, long DNA molecules in both the maternal and fetal DNA populations are enriched (DNA molecules with size >200 bp: 91.2%, fetal DNA molecules with size >200 bp: 87.9%) compared to using a ratio of 1.2X (DNA molecules with size >200 bp: 44.9%, fetal DNA molecules with size >200 bp: 46.6%). Therefore, this plot illustrates the enrichment of long DNA (e.g., >200 bp) in LEP-related total DNA (left panel) and LEP-related fetal DNA (right panel) by using paramagnetic beads with a bead-to-sample ratio of 0.8X.
[0123] 3. Enrichment of the long fetal fraction in LEP As shown with the paramagnetic bead data, a similar enrichment of long DNA can only be found when analyzing fetal DNA. The DNA was fragmented and subjected to short-read sequencing. Fetal DNA was identified using fetal-specific alleles.
[0124] Figures 12A - 12C show the enrichment of long fetal DNA in LEP-related DNA. In Figures 12A - 12B, the frequency refers to the percentage of DNA fragments in the sample that are below or above 200 bp. The two samples are FSN and LEP with PBS washing and DNase I treatment. The horizontal axis divides the data into two groups based on fragment size, namely, groups above and below 200 bp. Fetal DNA was identified using fetal-specific markers.
[0125] Similar to the plot when analyzing all DNA, the plot illustrates the enrichment of long fetal DNA in LEP-related DNA. Such long DNA enrichment after DNA shearing can also be observed in the fetal DNA population (i.e., DNA molecules with size >200 bp: 46.6% in LEP vs. 4.3% in cell-free DNA). Also, the percentage of long fetal DNA molecules increases by approximately 10-fold. Additionally, the total number of long reads increases from 72,152 to 2,155,708.
[0126] Figure 12C shows the size distribution of LEP-related fetal DNA and cell-free fetal DNA from FSN. The size profile shows behavior similar to other previous size profiles shown herein, with LEP DNA being longer. The overall size profile of LEP-related fetal DNA is relatively shifted towards larger sizes compared to fetal cell-free DNA of FSN.)
[0127] 4. Effect of treatment on the fetal fraction of long DNA fragments The effect of LEP purification and treatment steps on the fetal fraction of DNA fragments of different sizes, including exceeding a size threshold (e.g., 200 bp, 600 bp, or 1000 bp), was also analyzed. The fetal fraction remains stable for LEP-treated samples and significantly increases for LEP samples that are treated and washed. Thus, long DNA fragments can be obtained without the corresponding decrease in the fetal fraction, as observed in standard plasma samples.)
[0128] Figure 13 shows the fetal fraction in LEP with various treatments compared to FSN. The results correspond to Case 1 of Figure 5A. The vertical axis is the fetal fraction determined using fetal-specific markers. The plot shows the fetal DNA fraction for those DNA molecules greater than 200 bp among LEP without PBS wash and DNase I treatment, LEP with PBS wash, LEP with PBS wash and DNase treatment, and cell-free DNA obtained from FSN.)
[0129] As shown in Figure 13, the fetal DNA fraction in >200 bp DNA molecules obtained from LEP with only wash and wash / treatment was higher than that in cell-free DNA obtained from FSN. For LEP samples with wash and treatment, the fetal fraction is nearly 80%. Thus, LEP-based analysis will promote the enrichment of those long DNA molecules of fetal origin.)
[0130] Figure 14 shows the fetal fraction versus fragment size for various sample types. The analysis used a pool of six cases in the third trimester of pregnancy. After fragmentation, sequencing was performed on a short-read sequencing platform. The vertical axis is the fetal fraction, and the horizontal axis is the fragment size. The fetal fraction was determined using one or more fetal-specific markers in a set of one or more loci. Fragments were used to determine whether the fragment covered one of the loci corresponding to the fetal-specific marker. The fetal fraction was determined using the ratio of the number of fragments with fetal-specific markers to the total number of fragments covering any one of the loci.
[0131] As shown, the fetal DNA fraction in the DNA pool from the washed LEV sample 1408 appeared to be relatively stable as the size of the DNA fragments increased. In contrast, the fetal DNA fraction in the DNA pool from the FSN sample 1410 (considered equivalent to plasma) decreased dramatically as the size of the DNA fragments increased. These results indicate that the embodiments can obtain more long fetal DNA with DNA from LEV with DNase treatment.
[0132] The complex ability to have a high fetal fraction among long DNA fragments provides various advantages, such as enabling more efficient techniques for determining fetal genomic characteristics. For example, when the fetal fraction is close to 50%, fetal-specific alleles will be included in a significant proportion of the DNA fragments. For example, since sequencing errors can be easily removed, it may not be necessary to genotype the fetus. The sequencing errors will be far fewer than the actual fetal-specific alleles. Therefore, when the number of rDNA fragments at a locus is at least 10-15% of the fragments at a given locus, its allele (different from the maternal allele) can be identified as a fetal-specific allele. Also, when long DNA fragments are available, such fetal-specific fragments are likely to cover CpG sites, thereby enabling the detection of fetal epigenetic characteristics. Additionally, such long DNA fragments are likely to contain multiple fetal-specific alleles, thereby enabling the determination of fetal haplotypes by joining together fragments with fetal-specific alleles. Similarly, for long fetal DNA fragments, multiple fetal-specific epigenetic markers are likely to be present in the same fragment, thereby enabling the fetal DNA to be identified and joined together to identify both haplotypes.
[0133] C.SEP SEP samples prepared by the methods described in FIGS. 1 and 2 were also analyzed. SEP has a size of less than approximately 200 nm. The analysis examined the proportion of DNA fragments of different sizes in plasma and SEP samples, and the fetal fraction of these two sample types. Analysis of SEP-related DNA indicates that long fetal DNA molecules can be obtained.
[0134] To analyze long-sized DNA from the SEP source, single molecule real-time sequencing (Pacific Biosciences) was performed on a pool of 5 SEP-related DNA samples and paired untreated plasma DNA samples from third-trimester pregnant women, generating 870,000 and 980,000 circular consensus sequences (CCS), respectively. The lengths of fetal DNA molecules (SEP-DNA) obtained from SEP ranged from 50 bp to 23,026 bp.
[0135] Figure 15 shows the size distributions of SEP-related DNA and paired plasma DNA. The vertical axis is the percentage of DNA fragments occurring within a given size range for each of the two samples (SEP and plasma).
[0136] Figure 15 shows an increase in long DNA fragments in the SEP sample. The size distribution of SEP-related fetal DNA molecules is shifted towards larger sizes, suggesting that SEP-related fetal DNA enriched long fetal DNA molecules. For example, DNA molecules >200 bp account for 86.9% and 56.3% of SEP-related fetal DNA and plasma fetal DNA, respectively. The percentage (13.0%) of DNA fragments within the size range of 2,000 - 3,500 bp in SEP-related fetal DNA was 4.6-fold higher than that (2.8%) of plasma fetal DNA.
[0137] Compared to plasma, the peak of the size distribution is switched from the main peak of about 150 - 600 bp to the size range of 600 - 2,000 bp. This indicates that long fragments are also enriched in the SEP sample compared to plasma. Importantly, the single molecule sequencing technology was able to detect these long fragments that had been overlooked in previous studies.
[0138] Figures 16A-16B show the analysis of fetal DNA molecules in SEP-related DNA using different size ranges. Figure 16A shows the fetal DNA fractions across different DNA size ranges in plasma and SEP samples. In Figure 16A, the vertical axis is the fetal DNA fraction measured using fetal-specific markers. The horizontal axis shows three size ranges, each of which represents the fetal fraction in plasma and SEP samples.
[0139] It was assumed that the fetal DNA fraction would vary depending on the different sizes of DNA molecules obtained from SEP. Indeed, in the smaller size ranges (50-600 bp and 600-3000 bp), the fetal fraction in the SEP sample was lower than that in plasma. However, for DNA in the range of 3000-5000, the fetal fraction was higher in SEP compared to plasma. Thus, in the case of very long DNA, the decrease in the fetal fraction in plasma DNA is much more dramatic than in SEP. Therefore, in the case of long DNA, SEP can provide more fetal DNA and longer fetal DNA than plasma.
[0140] More specifically, in the fragment size range of 3,000-5,000 bp, the fetal DNA fraction was higher in SEP-related DNA than in plasma DNA (1.9% vs. 1.2%). In contrast, the fetal DNA fraction was lower in SEP-related DNA than in plasma DNA for both fragment size ranges of 50-600 bp (19.1% vs. 22.9%) and 600-3,000 bp (6.4% vs. 7.8%).
[0141] Figure 16B shows the amount of fetal DNA fragments having a size of >5 kb per million total CCS from SEP-related DNA and plasma DNA. CCS can be regarded as equivalent to DNA fragments. Such enrichment seen for fragments of 3000 - 5000 bp can be extended to DNA fragments of >5 kb size, and the number of fetal DNA fragments of >5 kb size was surprisingly 5-fold higher in SEP-related DNA compared to paired plasma DNA. In plasma, there are less than 5 reads, while SEP has about 25 reads, which is at least 5-fold or more. Thus, long fetal DNA molecules were enriched in SEP-related DNA compared to plasma. This analysis of SEP, being different from previous studies by Zhang et al. which used short-read sequencing, was able to detect only DNA molecules of less than 600 bp (Zhang et al. BMC Med Genomics. 2019;12:151).
[0142] These data suggested that in some embodiments, long fetal DNA could also be enriched using SEP-related DNA with fragment size selection. Fragment size selection can be performed in silico or physically (e.g., gel-based or bead-based DNA size selection).
[0143] V. Fetal Analysis For example, after purification of LEP or SEP, the membrane can then be disrupted to expose the DNA fragments, and various types of analysis can be performed using the DNA extracted from the particles. The DNA fragments can be analyzed using various assays such as the various types of sequencing and PCR described herein. Such assays can provide information regarding the DNA fragments such as the sequence (including terminal motifs), location within the reference genome (e.g., including the genomic location after alignment and at the ends of the DNA fragment), methylation status at various sites (e.g., CpG sites), and size (e.g., determined from the length of the entire sequence or from the alignment of the terminal sequences as can be done from paired-end reads). Such information can provide characteristics at a particular position, site, or region such as the number, size of the DNA fragment, methylation level, terminal position within the genome, amount of overhang (jaggedness) at the ends of the fragment, and motifs at the ends of the fragment, e.g., 3-mer or 4-mer at the ends of the DNA fragment.
[0144] Various examples of bioinformatics analysis have already been discussed. For example, copy number aberrations (or other sequence imbalances) can be detected based on the number of DNA fragments in one region, or haplotypes can be compared to reference values such as the number of DNA fragments in different regions or other haplotypes. Also, differences between methylation levels or sizes, and regions / haplotypes can be used. Additional examples are provided below.
[0145] A. Maternal inheritance of the fetus A higher fetal DNA fraction present in LEP-related DNA would improve the resolution and accuracy of maternal genetic analysis of the fetus. For example, using sequencing results from LEP-related DNA, relative haplotype dosage (RHDO) analysis based on the sequential probability ratio test (SPRT) (Lo et al. Sci Transl Med. 2010;2:61ra91 and US Publication No. 2011 / 0105353) can be used to infer the maternal genetics of the fetus. As described in US Publication No. 2017 / 0029900, methylation haplotypes can also be used. As other examples besides SPRT, without limitation, RHDO analysis based on binomial distribution, Poisson distribution, gamma distribution, beta distribution, hidden Markov model, etc. can be used.
[0146] The RHDO method can use the difference in the number of alleles of heterozygous loci (e.g., SNPs) between maternal haplotypes in the sample, i.e., Hap I and Hap II, respectively. When maternal Hap I is inherited by the fetus, the number of plasma DNA molecules derived from maternal Hap I will be relatively more compared to maternal Hap II. Otherwise, maternal Hap II will be relatively overly abundant. NhapI and NhapII are the measured allele numbers of Hap I and Hap II, respectively, which can be assumed to follow a Poisson distribution.
Number
[0147] The difference in the number of alleles between maternal haplotypes, N hapI -N hapII is, N * the mean of f and
Number
[0148] The fetus can inherit either haplotype I or II from the mother. Thus, if Z is < 3 but > -3, it means that there is insufficient statistical evidence to determine fetal inheritance in that region. The RHDO process can start from any genomic location and progressively accumulate sequenced reads mapped to SNPs present with maternal Hap I and Hap II, respectively. When maternal inheritance classification is performed during the accumulation of sequenced reads for RHDO analysis, the RHDO process can resume at the following heterozygous loci.
[0149] RHDO analysis was applied to three samples across the whole genome (i.e., DNA from FSN, LEP with PBS wash, and LEP with PBS wash and DNase I treatment). For each sample, the median of 129,199 SNPs with heterozygous maternal genotypes was analyzed using 29 million sequenced results (range: 107,550 - 136,642).
[0150] For the sequencing results obtained from LEP with FSN, PBS washing, and LEP with PBS washing and DNase I treatment, respectively, RHDO classifications of 678, 1033, and 1727 were obtained. There are more classifications for LEP with PBS washing and DNase I treatment, and thus higher resolution.
[0151] Figures 17A - 17B show the analysis of LEP - related DNA enabling higher resolution of maternal inheritance determination. Figure 17A shows the haplotype block size distribution determined to be inherited by the fetus from the analysis of cell - free DNA (FSN), DNA from LEP with PBS washing, and DNA from LEP with PBS washing and DNase I treatment, respectively. The vertical axis is the size of the haplotype block size, and the width of the line indicates more blocks of that size. Figure 17B shows exemplary genomic regions with maternal inheritance patterns from the analysis of cell - free DNA (FSN), DNA from LEP with PBS washing, and DNA from LEP with PBS washing and DNase I treatment, respectively.
[0152] As shown in the violin plot of Figure 17A, the median of the maternal haplotype block sizes determined to be inherited by the fetus is significantly smaller in LEP with PBS washing and DNase I treatment (1.24 Mb) compared to FSN (3.03 Mb) and LEP with PBS washing (1.70 Mb). This result suggests that LEP - related DNA enabled achieving higher resolution in determining the maternal inheritance of the fetus. The same conclusion was reached using the N50 statistic (i.e., FSN: 5.26 Mb, LEP with PBS washing: 3.78 Mb, LEP with PBS washing and DNase I treatment: 1.73 Mb). N50 is defined as the length corresponding to the haplotype block at which the cumulative length of the haplotype blocks reaches 50% of the total length of all blocks after ranking all haplotype blocks in descending order by length.
[0153] Figure 17B shows an exemplary genomic region (chr1: 174,000,000 - 200,000,000) that shows several maternally inherited haplotype blocks determined to be inherited by the fetus by analyzing DNA sequencing data from LEP with FSN, LEP with PBS washing, and LEP with PBS washing and DNase I treatment, respectively, according to an embodiment of the present disclosure. It can be observed that fetal maternal inheritance can be achieved with higher resolution in LEP-related DNA.
[0154] Therefore, these results suggest that the analysis of LEP-related DNA enables better performance in the detection of single-gene disorders by non-invasive methods. For example, the high resolution of the RHDO analysis in Figure 17B can enable accurate indication of fetal recombination when fetal recombination is present. Recombination present in the fetus would confound the RHDO analysis in low-resolution RHDO analysis (i.e., using FSN). For example, a 100-Mb region is more likely to contain recombination than a 1-Mb region. Therefore, a 100-Mb resolution RHDO analysis concludes that maternal haplotype I of size 100 Mb was inherited by the fetus. However, in reality, there is recombination within 90 Mb - 100 Mb that carries the disease-causing gene. Therefore, an incorrect interpretation that the fetus is affected by this disease would occur.
[0155] On the other hand, in a 1-Mb resolution RHDO analysis (e.g., LEP with washing and nuclease treatment), it can be seen that before the 90-Mb location, there are many blocks classified as Hap I inherited by the fetus, and the pattern continues where many blocks after the 90-Mb location are classified as Hap II inherited by the fetus. In this way, a correct interpretation regarding whether the fetus is affected could be achieved. The use of LEP DNA enables high resolution of the RHDO analysis and thus improves the performance of single-gene disorder detection.
[0156] Haplotypic genes and single gene disorders are examples of fetal genomic characteristics. Other fetal genomic characteristics can be determined, such as sequence imbalances, genotypes (e.g., genetic alleles), haplotypes (e.g., genetic haplotypes), mutations (e.g., mutant alleles), and methylation levels.
[0157] B. Pregnancy Analysis In addition to fetal genomic characteristics, genomic characteristics of pregnancy can be determined. Diagnostic values of particle-associated DNA can be extended to pregnancy complications (e.g., preeclampsia). Increased plasma EP has been reported in preeclamptic patients (Orozco et al. Placenta. 2009;30:10.; Goswami et al. Placenta. 2006;27:1.), indicating that EP-related DNA levels may be promising biomarkers for those diseases. Thus, DNA molecules obtained from LEP, SEP, and FSN can be used to inform pregnancy complications including, but not limited to, hypertension, gestational diabetes, infections, preeclampsia, preterm labor, pregnancy loss / miscarriage, and fetal growth restriction (FGR). Subjects with preeclampsia can have less amounts of long cfDNA.
[0158] Additionally, the method can distinguish RNA molecules contributed by the mother and fetus in an EP sample. Thus, the method can identify changes in the contribution from other individuals (i.e., mother or fetus) to a mixture at a particular locus or for a particular gene, even when the contribution from one individual changes or moves in the opposite direction. Such changes cannot be easily detected when measuring the overall expression level of a gene, regardless of the tissue or individual of origin.
[0159] C. Advantages of High Fetal Fraction The ability to have a high fetal fraction among DNA fragments provides various advantages, for example, enabling more efficient techniques for determining fetal genomic characteristics. For example, when the fetal fraction is close to 50%, fetal-specific alleles will comprise a significant proportion (e.g., at least 10%, 15%, or 20%) of the DNA fragments. For example, when the fetal fraction is 50%, the fetal-specific alleles at a fetal heterozygous locus will comprise approximately 25% of the DNA fragments.
[0160] As a result of the high fetal proportion, for example, sequencing errors will occur at a much lower rate and can be easily removed, so it may not be necessary to genotype fetal cells. The sequencing errors will be far fewer than the actual fetal-specific alleles. Thus, when the number of DNA fragments at a locus at least exceeds the threshold of fragments at a given locus (e.g., 10%, 15%, or 20%), the allele (different from the maternal allele) can be identified as a fetal-specific allele. Such genotyping of the fetus using a purified blood sample from the mother can provide information regarding fetal mutations, including de novo mutations, since a significant portion of the fragments with mutations will be present.
[0161] In addition to the improved functionality, increased accuracy and efficiency (e.g., using smaller samples or fewer assay reactions and / or reagents) can be achieved. For example, since there are more fetal DNA fragments in the sample (e.g., after purification of LEP and / or fragment size selection of LEP or SEP), there will be a greater separation between two classifications of fetal genomic characteristics. For example, since the overexpression or underexpression is greater, there will be a greater separation (e.g., indicating copy number abnormalities) between two classifications of sequence imbalance.
[0162] Because the overexpression or underexpression is greater, the threshold (cut-off) for classification will be reached earlier, i.e., with fewer DNA fragments. Thus, smaller samples and / or fewer assay reactions (e.g., less sequencing or digital PCR) can be performed. Thus, the higher concentration of DNA molecules derived from the placenta can provide a higher sensitivity approach in the detection of fetal abnormalities, including but not limited to the detection of chromosomal aneuploidies (e.g., trisomies 21, 18 or 13), and single gene disorders (e.g., cystic fibrosis, hemochromatosis, Tay-Sachs, beta / alpha thalassemia, and sickle cell disease).
[0163] D. Advantages of Using Long DNA Surprisingly, the data herein show an increase in long DNA fragments. This is in contrast to previous studies by Zhang et al. that found shorter and fewer DNA fragments. The techniques described herein provide preferential enrichment of LEP, for example, by using pellets of large particles obtained after centrifugation at least 10 minutes at greater than 10,000g. Further, the use of fragmentation by long read sequencing techniques (such as nanopore sequencing (e.g., Oxford Nanopore Technologies) and single molecule real time sequencing (e.g., Pacific Biosciences)) or short read sequencing techniques can provide sequence reads of long DNA fragments.
[0164] 1. Haplotype and Mutation Analysis (For example, after purification of LEP) There are advantages in having a greater amount (e.g., raw amount or percentage) of long DNA fragments in the sample. In addition to the advantage of obtaining a higher fetal DNA fraction in LEV for RHDO analysis, genetic and epigenetic analyses of long DNA fragments in LEV and SEV can be performed. For example, the use of methylation information in CpG sites and / or variants in long DNA molecules will facilitate the determination of fetal maternal inheritance. Determine whether the observed DNA fragments from LEP are of fetal origin (e.g., using fetal-specific alleles that can be identified via the above techniques), and thereby determine whether any genetic / epigenetic changes linked to such DNA fragments, if present, are transmitted to the fetus.
[0165] In long DNA fragments (e.g., 500bp, 600bp, 700bp, 800bp, 900bp, 1000bp, 5000bp, or more), it is more likely to have both SNP sites and CpG sites, or multiple SNP sites, or multiple CpG sites on the same fragment. The allelic state and / or methylation state at such positions can provide increased ability and accuracy for determining haplotypes. Multiple values (alleles or methylation states) on the same fragment can be compared to parental haplotypes or other reference haplotypes (e.g., from a particular population). In this way, haplotypes can be identified.
[0166] Longer DNA fragments can also help with de novo assembly, for example, to determine haplotypes and / or de novo mutations. When there is a higher likelihood of multiple heterozygous loci (for alleles with the same sequence or for methylation status), the variation of such fragments overlapping on one heterozygous locus increases. Thus, such fragments can extend the haplotype, for example, by identifying the same allele on the fragment, but the fragment can also extend to another heterozygous locus where another overlapping fragment can be identified, etc. Additionally, fetal vesicles can be identified (e.g., using fetal-specific proteins outside the vesicles), and any of such DNA fragments (short or long) can be ligated together or can fill gaps from haplotyping focused on long DNA fragments.
[0167] Also, there are advantages by having a greater amount (e.g., raw amount or percentage) of long DNA fragments in a fetal sample (e.g., after fragment size selection of LEP or SEP). For example, for each region, three haplotypes can be essentially determined, namely, two from the mother (one of which is shared with the fetus) and one that is paternal. And if there are de novo mutations, all four haplotypes can be determined. When the fetal DNA fraction is high, each branch (haplotyped) has a sufficient number of DNA fragments to support (determine) each haplotype. Or, when the fetal DNA fraction is very high (e.g., above 70%), the two fetal haplotypes can be determined with confidence by determining the two most common haplotypes. Further, the haplotype can have higher resolution the higher the fetal DNA fraction, as shown in FIG. 17B.
[0168] 2. Origin tissue analysis As a further illustration of the advantages of obtaining and analyzing long DNA fragments, the longer a DNA molecule, the greater the number of CpG sites it is likely to contain. Different cell types carry different methylation patterns across CpG sites. For example, cells from placental tissue have a unique methylome pattern compared to cells from tissues such as white blood cells, and liver, lung, esophagus, heart, pancreas, colon, small intestine, adipose tissue, adrenal gland, and brain.
[0169] The methylation pattern can function as a "molecular barcode" to track the cell identity of DNA molecules derived from the LEP of a pregnant woman. For example, the methylation pattern can be represented as "---M----U-------M------", where "M" represents a methylated CpG site, "U" represents an unmethylated CpG site, and the dashed lines represent arbitrary nucleotide distances between any two CpG sites or surrounding different CpG sites. Longer DNA molecules carrying more CpG sites increase the complexity of the "molecular barcode" and enable higher specificity in tissue origin analysis for DNA molecules derived from the LEP of pregnancy compared to shorter DNA molecules.
[0170] For example, because many tissues share the same methylation state, it is not possible to determine which organ contributes to a DNA molecule containing one CpG site in a DNA mixture from a pregnant subject (e.g., the mixture contains DNA from placenta, liver, intestine, lung, heart, brain, T cells, B cells, neutrophils, megakaryocytes, and red blood cells). In contrast, based on the single molecule methylation pattern across a series of CpG sites, there can be a higher likelihood (specificity) of accurately determining which organs contribute to a DNA molecule containing a sufficient number of CpG sites (e.g., >30 CpG sites). Determining the origin tissue of LEP DNA molecules in a pregnant woman can be implemented by comparing the methylation pattern of LEP DNA of a certain size (e.g., >1000 bp) with the reference methylation patterns of various tissues including but not limited to placenta, liver, intestine, lung, heart, brain, T cells, B cells, neutrophils, megakaryocytes, and red blood cells.
[0171] Comparing the LEV DNA methylation with a reference methylation pattern can include, but is not limited to, artificial intelligence-based algorithms such as edit distance calculation (e.g., the minimum edit distance referring to the tissue contributing to such molecules being analyzed), bitwise operations, naive Bayes classifiers, random forest trees, support vector machines, gradient boosting, hidden Markov models, convolutional neural networks, and deep recurrent neural networks.
[0172] Figure 18 shows an example of using EV DNA molecules for non-invasive prenatal testing. EV DNA molecules determined to be of placental origin based on the methylation pattern according to embodiments of the present disclosure can be used for non-invasive prenatal testing (NIPT) of pregnant women. Examples of such NIPT include detection of fetal chromosomal aneuploidy, detection of single gene diseases, detection of fetal copy number abnormalities, and the like.
[0173] As depicted in Figure 18, the biological sample 1810 shows EVs in the plasma of a pregnant woman. The biological sample 1810 also includes cell-free DNA not shown.
[0174] In step 1815, the desired EVs (e.g., small or large) are selected using physical, chemical, and / or biological properties (e.g., size) as described herein. The concentrated sample 1820 shows EVs within the desired size range.
[0175] In step 1825, DNA is extracted from the EVs in the concentrated sample 1820, for example, by disrupting the membranes of the EVs. The extracted DNA 1830 includes long DNA with a high fetal fraction as shown herein.
[0176] After extraction, at step 1835, the DNA fragments can be analyzed. For example, methylation recognition sequencing such as bisulfite treatment, single molecule sequencing, enzyme methyl-seq (EM-seq). Sequence reads 1840 indicate methylated CpG sites (M) and unmethylated CpG sites (U).
[0177] At step 1850, analyze the sequence reads to obtain one or more characteristics such as the amount of DNA (potentially determined by alignment to a reference genome at a particular location or region), fragment size (e.g., by determining the length of long reads of the entire DNA molecule or by aligning paired-end reads), fragmentation pattern (such as end motifs or terminal positions in the reference genome), and methylation pattern.
[0178] At step 1860, DNA fragments (especially long DNA fragments, e.g., at least 600 bp or other lengths described herein) are identified as corresponding to a particular reference tissue. Different tissues have different methylation patterns. Such reference methylation patterns can be determined by analyzing cells of a particular reference tissue. The reference methylation pattern can be designated as methylated when the methylation level at a site exceeds a specified threshold (e.g., 70%, 75%, 80%, 85%, 90%, 95%, or 99%). The reference methylation pattern can be designated as unmethylated when the methylation level at a site is less than a specified threshold (e.g., 30%, 25%, 20%, 15%, 10%, 5%, or 1%). A particular location within the genome can have a pattern rather than being unique to a particular tissue. Optionally, the reference methylation patterns of various tissues can be obtained from single molecule sequencing that expresses the methylation pattern across individual molecules, and the methylation state can be binary (0 or 1 representing the unmethylated state and methylated state, respectively).
[0179] When a long DNA fragment has multiple CpG sites (e.g., as shown in FIG. 18), the pattern at the aligned location in the reference genome can be compared to the reference pattern of one or more reference tissues at the aligned location. Whether the methylation pattern (U and M at specific positions) is the same at each of the positions can be used to determine the closest matching reference tissue, or potentially, to provide a match only if the methylation patterns are exactly the same. Such identification can be accurate for long DNA fragments covering multiple CpG sites, e.g., CpG sites greater than 4, 5, 6, 7, 8, 9, or 10. In some embodiments, only reference fetal (placental) tissue is used to identify fetal DNA fragments.
[0180] After fetal DNA fragments are identified using the methylation pattern, the fetal DNA can be analyzed to perform NIPT. Such identified DNA fragments are likely to be fetal DNA fragments, so the fetal fraction will be very high (e.g., +90%). Thus, the sequences of such identified fetal fragments can be used to indicate a disease such as a single-gene disorder or to identify the presence of one or more sequences (e.g., alleles and / or mutations) associated with two or more diseases. The partial or entire genome of the fetus can be determined, for example, using assembly techniques with the identified fetal DNA.
[0181] Thus, the sequencing techniques of the methods described herein can include methylation-aware sequencing. Then, for each of the plurality of sequence reads, the methylation pattern at the CpG sites of the sequence read can be determined. The sequence reads can also be aligned to genomic locations within the reference genome. The methylation pattern can be compared to the reference methylation pattern of fetal tissue at the genomic location. In this way, the sequence reads can be identified as corresponding to fetal DNA molecules based on the comparison.
[0182] Once fetal DNA fragments are identified, the fetal DNA can be analyzed. For example, using sequence reads identified as corresponding to fetal DNA molecules based on methylation patterns, it is possible to determine whether the fetus has genomic abnormalities (such as copy number, mutations, epigenetic disorders, etc.). Such determinations can use various characteristics of fetal DNA fragments across the genome or in specific regions, such as number, size, and fragmentation.
[0183] 3. Combined Sequence and Methylation Analysis The methylation pattern can be used to determine the origin tissue (placental origin) of LEP-related DNA molecules. It is possible to determine whether a single nucleotide variant (SNV) linked to the methylation pattern is inherited by the fetus. Also, it is possible to determine whether de novo mutations present in maternal plasma DNA are of fetal origin according to their linked methylation patterns. As an inference, inheritance can be determined based on SNVs, and this can be used to determine whether the observed abnormal methylation pattern is inherited by the fetus. Thus, fetal genetic and epigenetic gene analysis can be synergistic with each other.
[0184] The fetal genotype and / or haplotype can be determined by analyzing fetal DNA. For example, one or more haplotypes of the fetus can be determined using sequence reads identified as corresponding to fetal DNA molecules based on the fetal methylation pattern. Determining one or more haplotypes of the fetus can include determining the first maternal haplotype as being inherited by the fetus. Determining one or more haplotypes of the fetus can include determining the first paternal haplotype as being inherited by the fetus.
[0185] In addition to using fetal-specific methylation patterns to identify fetal DNA, fetal-specific alleles can also be used. Thus, the method can identify sequence reads as having fetal-specific alleles and can determine the methylation pattern at the CpG sites of the sequence reads. The methylation pattern can then be used to determine whether the fetus has an epigenetic abnormality. For example, the methylation pattern of a particular fetal DNA molecule can have a pattern that matches a pattern known to correspond to an epigenetic abnormality. Such epigenetic abnormalities can include fragile X syndrome.
[0186] 4. Increase in length while maintaining the fetal fraction The approach disclosed herein enables the selective analysis of long DNA molecules from LEP without a decrease in the fetal DNA fraction. As noted above, the higher the concentration of DNA molecules of placental origin, the higher the accuracy (e.g., sensitivity and / or specificity) that can be achieved. In contrast, for particle-free cfDNA, which is much more fragmented compared to LEP DNA molecules, the selective analysis of long DNA molecules often comes at the significant expense of a decrease in the fetal DNA fraction. Thus, the use of LEP DNA molecules according to embodiments of the present disclosure can result in higher performance for NIPT.
[0187] VI. Methods Various methods are described above and are explained in this section. Purification and / or processing of a blood sample for extracellular vesicles (e.g., LEP and SEP) can be performed, resulting in an increase in the fetal fraction and / or long DNA fragments. Nucleic acid fragments (DNA and / or RNA) of a certain length, e.g., larger than a size threshold, can be selected, which can result in an increase in the fetal fraction. The analysis can involve different types of assays, including sequencing and probe-based techniques such as digital PCR. When performing sequencing, long nucleic acid fragments can be analyzed by using long read techniques or by further fragmenting the nucleic acid fragments and then using short read techniques.
[0188] A. Concentration for NIPT analysis using washing and / or DNase treatment A blood sample can be purified for EP using physical separation techniques such as centrifugation and / or filtration, for example. The sample can then be processed, for example, by ionic washing and / or nuclease treatment. In this way, the sample can be concentrated for vesicles (particles) and thus for fetal nucleic acids.
[0189] FIG. 19 is a flowchart illustrating a method 1900 for purifying and processing a blood sample of a pregnant female with a fetus. The female may be pregnant with more than one fetus, which also applies to other techniques described herein. Method 1900 and other methods described herein can be performed using a computer system in part or in whole, e.g., with a computer system that controls physical processes.
[0190] At block 1910, a blood sample from a female carrying a fetus is received. The blood sample contains extracellular particles and particle-free nucleic acids. The blood sample can be a plasma sample or can contain other components, such as blood cells. The extracellular particles contain cell-free nucleic acids inside the membrane. For example, each extracellular particle can contain cell-free nucleic acids inside its respective membrane. The blood sample can be received by a measurement system that can perform physical and in-silico steps.
[0191] At block 1920, a physical separation technique preferentially selects at least a portion of the extracellular particles, thereby obtaining a particle-enriched sample. For example, the physical separation technique can preferentially select particles that are below an upper threshold value and / or above a lower threshold value. Examples of such thresholds are provided in Section II.A.1. For example, the upper threshold value can be 10 microns, 9 microns, 8 microns, 7 microns, 6 microns, 5 microns, or 4 microns, 3 microns, 2 microns. The lower threshold value can be 200 nm, 300 nm, 400 nm, 500 nm, 600 nm, 700 nm, 800 nm, 900 nm, or 100 nm. As used herein, the term "preferentially" refers to a technique that increases the percentage of extracellular particles having a desired characteristic (e.g., a specified size), thereby obtaining an enriched sample with a higher percentage of extracellular particles having the desired characteristic than the original sample.
[0192] Examples of physical separation are provided herein, for example, in Sections II.A and Figures 1 and 2. For example, one or more stages of centrifugation can be performed. The pellet after centrifugation can be extracted and processed later. Centrifugation parameters (e.g., force and time) can be selected to obtain particles of a desired size, e.g., large or small. Examples of force and time, as well as some centrifugation stages, are provided in other sections of this document. Another example of physical separation is filtration, which is described in other sections.
[0193] One or more initial stages of centrifugation can be used to remove cells, for example, by centrifuging at 500 g or more for at least 10 minutes. One or more subsequent centrifugation stages can be at 10,000 g or more for at least 10 minutes, resulting in a pellet of LEP, which can be removed. Thus, one or more subsequent centrifugation stages can be used to remove LEP. Further centrifugation can preferentially select SEP from the supernatant.
[0194] In block 1930, the particle-concentrated sample is processed using a processing technique to remove excess particle-free nucleic acid, thereby obtaining a processed particle-concentrated sample. As an example, the processing technique can include ion washing of the particle-concentrated sample in an ionic solution (e.g., phosphate-buffered saline (PBS) or other saline), and / or applying a nuclease to the particle-concentrated sample. Either one of these two processes can be performed multiple times and can be performed alternately. For example, washing can be performed first, then nuclease treatment can be performed, and subsequently another washing can be performed. The centrifugation step can also be performed during any of the processing steps.
[0195] The processing technique can increase the fractional concentration of fetal nucleic acid in the processed particle-concentrated sample compared to the particle-concentrated sample. Such an increase is shown in various figures such as Figures 4, 5A - 5B, and 6A - 6B. Examples of such washing and nuclease treatment are described herein. Washing can remove nucleic acid floating in the sample, and nuclease treatment can remove nucleic acid bound to the membrane of the vesicles (particles).
[0196] In block 1940, cell-free nucleic acid molecules from extracellular particles are exposed by disrupting (e.g., lysing) the membrane of the extracellular particles. Such disruption can be carried out in various ways, for example, by mechanical disruption, sonication, enzymatic hydrolysis (e.g., proteinase K), detergents (e.g., ionic detergents such as sodium dodecyl sulfate (SDS) or non-ionic detergents such as Triton® X-100), osmotic shock, and freeze-thaw methods. As a result of the disruption, the membrane is disrupted, thereby releasing the inner nucleic acid molecules (fragments). These cell-free nucleic acid molecules (fragments) can be DNA and / or RNA.
[0197] In block 1950, the cell-free nucleic acid molecules are assayed to obtain sequence reads. Different types of assays can be used, including probe-based techniques such as sequencing and digital PCR. As described herein, various forms of sequencing can be performed by using long-read technology or by further fragmenting the nucleic acid fragments and then using short-read technology. The cell-free nucleic acid molecules bound to the inside of and / or on the surface of the EP can be assayed.
[0198] In block 1960, sequence reads are analyzed to determine genomic characteristics of a fetus or pregnancy. Examples of such characteristics are described above. For example, such sequence reads can be analyzed for various characteristics at a particular position, site, or region, such as number, size of nucleic acid fragments, methylation level, terminal position within the genome, amount of overhang (jaggedness) at the ends of the fragments, and motifs at the ends of the fragments, such as 3-mers or 4-mers at the ends of nucleic acid fragments. Further details of such techniques are described in U.S. Publication Nos. 2011 / 0105353, 2014 / 0019064, 2013 / 0237431, 2014 / 0195164, 2014 / 0315200, 2016 / 0017419, 2016 / 0217251, 2017 / 0073774, 2017 / 0024513, 2018 / 0105807, 2018 / 0142300, 2019 / 0341127, 2020 / 0056245, and 2020 / 0199656.
[0199] Examples of such analyses related to pregnancy can include techniques from U.S. Publication Nos. 2014 / 0243212 (RNA signatures specific to preeclampsia) and 2018 / 0372726 (e.g., using reference expression regions of one or more expression markers). Genomic characteristics of pregnancy can be related to one or more complications that reduce the number of women who carry a fetus to term.
[0200] In one example, analyzing sequence reads can be used to determine a genotype. Determining the genotype of a fetus at a locus can include aligning the sequence reads to a reference genome and can include determining that the locus contains a first allele when at least a specified percentage (e.g., 10%, 15%, 20%, etc.) of the sequence reads contain the first allele at the locus. The genotype can indicate a mutation. Other examples can include determining a genetic haplotype, determining the tissue of origin of a nucleic acid molecule, and determining the fetal DNA percentage.
[0201] B. Concentration of Fetal Fraction by Size Selection of Nucleic Acid Fragments Blood samples can be purified for EP, for example, using centrifugation and / or filtration. After assaying cell-free nucleic acid molecules (e.g., DNA and / or RNA), the size of the cell-free nucleic acid molecules can be determined and only certain nucleic acid fragments (molecules) can be selected. In this way, the sample can be enriched for fetal nucleic acids.
[0202] FIG. 20 is a flowchart illustrating a method 2000 for analyzing a blood sample of a pregnant woman containing a fetus, including selecting nucleic acid molecules based on size.
[0203] In block 2010, a blood sample of a pregnant woman carrying a fetus is received. The blood sample contains extracellular particles and particle-free nucleic acids. As with other methods, the blood sample can be a plasma sample or can contain other components, such as blood cells. The extracellular particles contain cell-free nucleic acids inside the membrane, as can occur with other methods described herein. The blood sample can be received by a measurement system capable of performing physical and in-silico steps.
[0204] In block 2020, one or more purification steps are performed to concentrate the extracellular particles, thereby producing a concentrated sample. The one or more purification steps can include one or more physical separation techniques and / or treatment techniques. The physical separation technique can preferentially select at least a portion of the extracellular particles, thereby obtaining a particle-concentrated sample. The physical separation technique can be performed in the same manner as block 1920 of method 1900. The treatment technique can be performed in the same manner as block 1930 of method 1900.
[0205] As an example, one or more purification steps can include filtration using one or more filters or flow cytometry. As another example, one or more purification steps can include centrifugation. One or more purification steps can preferentially select extracellular particles that are larger than a specified size.
[0206] As another example, one or more purification steps can perform a physical separation technique that preferentially selects at least a portion of the extracellular particles, thereby obtaining a particle-concentrated sample, and process the particle-concentrated sample using a processing technique that removes excess particle-free nucleic acid molecules, thereby obtaining a processed particle-concentrated sample. As an example, the physical separation technique can include at least one stage of centrifugation, for example, centrifugation at 16,000 g or more for at least 10 minutes. The processing technique can include washing the particle-concentrated sample with an ionic solution and / or applying a nuclease to the particle-concentrated sample. The first processing technique can increase the fractional concentration of fetal DNA in the processed particle-concentrated sample compared to the particle-concentrated sample.
[0207] In block 2030, cell-free nucleic acid molecules from extracellular particles are exposed by disrupting the membrane of the extracellular particles. Block 2030 can be implemented in the same manner as block 1940 of method 1900.
[0208] In block 2040, the cell-free nucleic acid molecules are assayed to obtain sequence reads. As an example, assaying can include sequencing or digital PCR. Block 2040 can be implemented in the same manner as block 1950 of method 1900. Cell-free nucleic acid molecules from inside and / or bound to the surface of the EP can be assayed.
[0209] In block 2050, the size of the cell-free nucleic acid molecule is determined. The size can be determined in various ways, for example, using physical techniques such as sequence reads or electrophoresis techniques or differential amplification. The size can correspond to the length, mass, or weight of the nucleic acid molecule. The size can be a size range. The size can be determined in various ways. For example, the length of the entire sequence (such as can be determined using long-read sequencing such as single-molecule sequencing) can be used as the size. Thus, assaying can include sequencing the entire entirety of each cell-free nucleic acid molecule, thereby generating one sequence read for each of the cell-free nucleic acid molecules, and determining the size of the cell-free nucleic acid molecule can include counting the nucleotides in the sequence read of the cell-free nucleic acid molecule.
[0210] As another example, the size can be determined by aligning the end sequences of the fragments such that it can be done using paired-end reads so that it is not necessary to sequence the entire fragment. Thus, determining the size of the cell-free nucleic acid molecule can include aligning one or more sequence reads to a reference genome for each of the cell-free nucleic acid molecules.
[0211] In some embodiments, the size of the nucleic acid molecule can be determined using physical techniques such as electrophoresis. In such embodiments, the physical size measurement can be performed before assaying the nucleic acid molecule. Thus, sequence reads may not be used to determine the size in such embodiments.
[0212] In yet another example, determining the size of the cell-free nucleic acid molecule can include performing digital PCR with different amplicon sizes. For example, different primers can amplify molecules of different lengths, resulting in amplicons of different lengths over the digital reaction. Also, different probes can detect the presence of amplicons of various sizes.
[0213] In block 2060, a set of cell-free nucleic acid molecules larger than a size threshold is identified. The size threshold can be 200 bp or more. As described herein, other exemplary size thresholds are 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 100 kb, 500 kb, and 1 Mb.
[0214] If size is determined using physical separation techniques, the set of cell-free nucleic acid molecules can be identified before the assay is performed. For example, cell-free nucleic acid molecules within a certain size range can be captured and then their nucleic acids can be assayed. If size is determined using sequence reads, cell-free nucleic acids that are within the desired range can be identified and their sequence information can be used.
[0215] In block 2070, the sequence reads of the set of cell-free nucleic acid molecules are analyzed to determine the genomic characteristics of the fetus. Block 2070 can be performed in the same manner as block 1960 of method 1900.
[0216] C. Sequencing for Long Reads A blood sample can be purified for EP using, for example, centrifugation and / or filtration. Long-read sequencing techniques can be performed to capture the surprisingly abundant long nucleic acid fragments in EP. In this way, the sample can be enriched for fetal nucleic acids and long cell-free fetal nucleic acid molecules (e.g., DNA and / or RNA) can be sequenced. Since cell-free nucleic acid molecules in plasma are known to be short (because they are naturally fragmented), performing long-read sequencing of cell-free nucleic acid molecules will be different from the prior art.
[0217] FIG. 21 is a flowchart illustrating a method 2100 for analyzing a blood sample of a pregnant female including performing long-read sequencing.
[0218] In block 2110, a blood sample of a pregnant female is received. The blood sample includes extracellular particles and particle-free nucleic acid molecules. The extracellular particles include cell-free nucleic acid molecules inside the membrane.
[0219] In block 2120, one or more purification steps are performed to concentrate the extracellular particles, thereby generating a concentrated sample. Block 2120 can be performed in the same manner as block 1920 of method 1900.
[0220] In block 2130, cell-free nucleic acid molecules from the extracellular particles are exposed by disrupting the membrane of the extracellular particles. Block 2130 can be performed in the same manner as block 1930 of method 1900.
[0221] In block 2140, sequencing technology is used to sequence the cell-free nucleic acid molecules to obtain sequence reads. The sequencing technology is such that at least a portion of the sequence reads exceeds a size threshold, for example, 600 bp. Other such size thresholds can be 700 bp, 800 bp, 900 bp, or 1000 bp, or other size thresholds described herein. As an example, the sequencing technology can include single molecule sequencing such as nanopore sequencing (e.g., Oxford Nanopore Technologies) and single molecule real-time sequencing (e.g., Pacific Biosciences). The sequencing technology can sequence short reads and long reads. Cell-free nucleic acid molecules from inside and / or bound to the surface of the EP can be sequenced.
[0222] Other examples of long-read sequencing techniques include synthetic long-read sequencing (Illumina) and linked-read technologies (10X genomics, Tell-seq). In such embodiments, long nucleic acid molecules are fragmented into compartments, and their partial sequences are tagged with the same barcode sequence (i.e., molecular barcode). Different long nucleic acid molecules are assigned to different compartments and tagged with different molecular barcodes. Thus, fragments derived from a long nucleic acid molecule can be reassembled into the original long nucleic acid molecule based on the same molecular barcode. As an example, the compartments can be implemented using droplets, beads, serial dilutions, or wells.
[0223] At block 2150, sequence reads are analyzed to determine genomic characteristics of the fetus. All of the nucleic acid fragments can be sequenced, and thus, the analyzed sequence reads, including sequence reads from long DNA fragments, can be of various lengths. The sequence reads can be of the entire nucleic acid fragment or only of the ends. Block 2150 can be implemented in the same manner as block 1970 of method 1900.
[0224] As an example, analyzing the sequence reads can include determining the fetus's haplotype by aligning sequence reads longer than 600 bp to each other, for example, as part of a de novo assembly. At least some of the aligned sequence reads can include multiple heterozygous loci. The aligned sequence reads can share heterozygous loci with the same allele, thereby allowing different sequence reads to have overlapping alignments in different amounts and at different loci.
[0225] D. Implementation of Fragmentation of Short-Read Platforms A blood sample can be purified for EP, for example, using centrifugation and / or filtration. To capture the surprisingly long nucleic acid fragments present in EP, the cell-free nucleic acid fragments (e.g., DNA and / or RNA) extracted from EP can be further fragmented and sequenced using a short-read sequencing platform. In this way, the sample can be enriched for fetal nucleic acids, and long cell-free fetal nucleic acid molecules can be sequenced. Since cell-free nucleic acid molecules in plasma are known to be short (because they are naturally fragmented), performing a fragmentation step will be different from the prior art.
[0226] Figure 22 is a flowchart illustrating a method 2200 for analyzing a blood sample of a pregnant female with a fetus, including performing fragmentation and short-read sequencing.
[0227] At block 2210, a blood sample of a pregnant female with a fetus is received. The blood sample includes extracellular particles and particle-free nucleic acid molecules. The extracellular particles include cell-free nucleic acid molecules inside the membrane.
[0228] At block 2220, one or more purification steps are performed to concentrate the extracellular particles, thereby generating a concentrated sample. Block 2220 can be performed in the same manner as block 1920 of method 1900.
[0229] At block 2230, the cell-free nucleic acid molecules from the extracellular particles are exposed by disrupting the membrane of the extracellular particles. At least a portion of the cell-free nucleic acid molecules from the extracellular particles are at least 600 bp. Block 2230 can be performed in the same manner as block 1930 of method 1900.
[0230] In block 2240, fragmentation techniques are applied to cell-free nucleic acid molecules. Fragmentation can shorten the length of long nucleic acid fragments, such that they can be sequenced using short-read sequencing platforms such as Illumina. Mechanical shearing, enzymatic fragmentation such as Tn5 transposase-based tagging, DNASE1, DNASE1L3, and / or DFFB treatment, chemical DNA fragmentation using light, sonication, or a combination of heat with a divalent metal cation such as magnesium or zinc to break down the nucleic acid. In some embodiments, bisulfite treatment can be used to fragment nucleic acid molecules.
[0231] In block 2250, after applying the fragmentation technique, the cell-free nucleic acid molecules are sequenced to obtain sequence reads. Since at least some of the long nucleic acid molecules are fragmented, the resulting fragments can be sequenced on a short-read sequencing platform. Cell-free nucleic acid molecules bound to the inside of the EP and / or to the surface of the EP can be sequenced.
[0232] In block 2260, the sequence reads are analyzed to determine genomic characteristics of the fetus or the pregnancy of the female. Block 2260 can be carried out in the same manner as block 1970 of method 1900.
[0233] For any of the methods described herein, the analysis can, for example, determine the genetic haplotype from the mother. As an example, analyzing the sequence reads can include using the sequence reads to determine the difference in the number of alleles at heterozygous loci of two maternal haplotypes, and using the difference in the number of alleles to determine the genetic haplotype for each of a plurality of regions. As shown in FIG. 17B, the average haplotype block size can be less than 2 Mb or 1.5 Mb.
[0234] VII. Exemplary System FIG. 23 illustrates a measurement system 2300 according to an embodiment of the present disclosure. The system shown includes a sample 2305, such as cell-free nucleic acid molecules (e.g., DNA and / or RNA) within an assay device 2310, and an assay 2308 can be performed on the sample 2305. For example, the sample 2305 can be contacted with the reagents of the assay 2308 to provide a signal of physical properties 2315 (e.g., sequence information of cell-free nucleic acid molecules). Examples of assay devices can be flow cells that include assay probes and / or primers, or tubes through which droplets (along with the droplets containing the assay) move. The physical properties 2315 (e.g., fluorescence intensity, voltage, or current) from the sample are detected by a detector 2320. The detector 2320 can perform measurements at intervals (e.g., periodic intervals) to obtain data points that constitute a data signal. In one embodiment, an analog-to-digital converter converts the analog signal from the detector into digital format at multiple times. The assay device 2310 and the detector 2320 can form an assay system, such as a sequencing system that performs sequencing according to the embodiments described herein. A data signal 2325 is transmitted from the detector 2320 to a logic system 2330. As an example, the data signal 2325 can be used to determine the sequence and / or location in a reference genome of nucleic acid molecules (e.g., DNA and / or RNA). The data signal 2325 can include various measurements performed simultaneously, such as different colors of fluorescent dyes or different electrical signals for different molecules of the sample 2305, and thus the data signal 2325 can correspond to multiple signals. The data signal 2325 can be stored in a local memory 2335, an external memory 2340, or a storage device 2345.
[0235] The logic system 2330 can be, or can include, a computer system, an ASIC, a microprocessor, a graphics processing unit (GPU), etc. It can also include, or be coupled with, a display (e.g., a monitor, an LED display, etc.) and a user input device (e.g., a mouse, a keyboard, buttons, etc.). The logic system 2330 and other components can be part of a stand-alone or network-connected computer system, or can be directly attached to, or incorporated into, a device (e.g., an array determination device) that includes the detector 2320 and / or the assay device 2310. The logic system 2330 can also include software executed by the processor 2350. The logic system 2330 can include a computer-readable medium that stores instructions for controlling the measurement system 2300 to perform any of the methods described herein. For example, the logic system 2330 can provide commands to a system that includes the assay device 2310 so that an array determination or other physical operation is performed. Such physical operations can be performed in a specific order, e.g., reagents are added and removed in a specific order. Such physical operations can be performed by a robotic system, e.g., including a robotic arm, so that a sample can be obtained and used to perform an assay.
[0236] The measurement system 2300 can also include a treatment device 2360 that can provide treatment to a subject. The treatment device 2360 can be used to determine and / or perform treatment. Examples of such treatment can include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormonal therapy, and stem cell transplantation. The logic system 2330 can be connected to the treatment device 2360, for example, to provide the results of the methods described herein. The treatment device can receive inputs from other devices, such as an imaging device, and user input (e.g., to control the treatment, such as control of a robotic system).
[0237] Any of the computer systems referred to in this specification may utilize any suitable number of subsystems. An example of such a subsystem is shown in FIG. 24 in computer system 10. In some embodiments, the computer system includes a single computer device, and the subsystems can be components of the computer device. In other embodiments, the computer system can include a plurality of devices, each a subsystem, along with internal components. The computer system can include desktop and laptop computers, tablets, mobile phones, and other mobile devices.
[0238] The subsystems shown in FIG. 24 are interconnected via a system bus 75. A printer 74, a keyboard 78, a storage device 79, a monitor 76 coupled to a display adapter 82 (e.g., a display screen such as an LED), and other additional subsystems are shown. Peripheral devices and I / O devices coupled to an input / output (I / O) controller 71 can be connected to the computer system by any number of means known in the art such as an input / output (I / O) port 77 (e.g., USB, FireWire®). For example, the computer system 10 can be connected to a wide area network such as the Internet, a mouse input device, or a scanner using the I / O port 77 or an external interface 81 (e.g., Ethernet®, Wi-Fi, etc.). The interconnection via the system bus 75 enables the central processor 73 to communicate with each subsystem and to control the execution of a plurality of instructions from the system memory 72 or the storage device 79 (e.g., a fixed disk such as a hard drive or an optical disk) and the exchange of information between subsystems. The system memory 72 and / or the storage device 79 can embody a computer-readable medium. Another subsystem is a data collection device 85 such as a camera, a microphone, an accelerometer, etc. Any of the data referred to in this specification can be output from one component to another and can be output to the user.
[0239] A computer system can include a plurality of the same components or subsystems connected together, for example, by an external interface 81, by an internal interface, or via a removable storage device that can be connected and removed from one component to another. In some embodiments, a computer system, subsystem, or device can communicate via a network. In such cases, one computer can be regarded as a client and another computer can be regarded as a server, and each can be part of the same computer system. The client and the server can each include a plurality of systems, subsystems, or components.
[0240] Aspects of embodiments can be implemented in the form of control logic using a hardware circuit (e.g., an application-specific integrated circuit or a field-programmable gate array) and / or computer software stored in a memory with a modular or integrated form generally programmable processor, and thus, the processor can include a memory storing software instructions configuring the hardware circuit, and an FPGA or ASIC with configuration instructions. As used herein, a processor can include a single-core processor, a multi-core processor on the same integrated chip, or a plurality of processing units on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, those skilled in the art will appreciate and understand other means and / or methods of implementing the embodiments of the present disclosure using hardware and combinations of hardware and software.
[0241] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using, for example, conventional techniques or object-oriented techniques, using any suitable computer language such as Java®, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive or floppy® disk, or optical media such as a compact disk (CD) or digital versatile disk (DVD) or Blu-ray disk, flash memory, and the like. The computer-readable medium may be any combination of such devices. Additionally, the order of operations may be rearranged. A process may end when its operations are complete, but may have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, and the like. When a process corresponds to a function, its end may correspond to the calling of the function or returning the function to the main function.
[0242] Such programs can also be encoded and transmitted using carrier signals adapted for transmission via various protocols compliant wired, optical, and / or wireless networks, including the Internet. Accordingly, a computer-readable medium can be created using a data signal encoded with such a program. A computer-readable medium encoded with program code can be packaged (as firmware) with a compatible device or provided separately from other devices (e.g., via an Internet download). Any such computer-readable medium can reside on or within a single computer product (e.g., a hard drive, CD, or an entire computer system) and can exist on or within different computer products in a system or network. A computer system can include a monitor, printer, or other suitable display for providing a user with any of the results referred to herein.
[0243] Any of the methods described herein may be implemented, in whole or in part, using a computer system that includes one or more processors configured to perform the steps. Any operations performed by the processor (e.g., alignment, determination, comparison, calculation, computation) may be performed in real time. The term "real time" may refer to a computing operation or process that is completed within a particular time constraint. The time constraint may be one minute, one hour, one day, or seven days. Accordingly, embodiments may be directed to a computer system configured to perform any of the steps of the methods described herein, and potentially, different components perform respective steps or groups of respective steps. Although presented as numbered steps, the steps of the methods herein may be performed simultaneously, at different times, or in a different order. Additionally, portions of these steps may be used in conjunction with portions of other steps from other methods. Also, all or portions of the steps may be optional. Additionally, any of the steps of any of the methods may be performed using a module, unit, circuit, or other means of the system for performing these steps.
[0244] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the present disclosure. However, other embodiments of the present disclosure may be directed to specific embodiments related to each individual aspect, or specific combinations of these individual aspects.
[0245] The foregoing description of the exemplary embodiments of the present disclosure has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the present disclosure to the exact forms described, and many modifications and variations are possible in light of the above teachings.
[0246] The recitation of "a", "an", or "the" is intended to mean "one or more" unless specifically indicated to the contrary. The use of "or" is intended to mean "and / or" unless specifically indicated to the contrary and is not intended to mean "exclusive or". References to a "first" component do not necessarily require the provision of a second component. Further, references to a "first" or "second" component do not limit the referenced component to a particular location unless expressly recited. The term "based on" is intended to mean "at least in part based on".
[0247] Claims may be drafted to exclude any element that may be optional. Accordingly, this description is intended to serve as a precedent for the use of exclusive terms such as "solely", "only", etc., or the use of "negative" limitations in connection with the recitation of elements of the claims.
[0248] All patents, patent applications, publications, and descriptions described herein are hereby incorporated by reference in their entirety for all purposes. None are admitted to be prior art. In the event of any conflict between this application and the references provided herein, this application shall control.
Claims
**Claim 1** A method comprising: receiving a blood sample from a female who is pregnant with a fetus; performing one or more purification steps to concentrate extracellular particles, thereby producing a concentrated sample, wherein the extracellular particles contain cell-free nucleic acids inside a membrane; exposing cell-free nucleic acid molecules from the extracellular particles by disrupting the membranes of the extracellular particles; assaying the cell-free nucleic acid molecules to obtain sequence reads; determining the size of the cell-free nucleic acid molecules; identifying a set of cell-free nucleic acid molecules that are larger than a size threshold, wherein the size threshold is 200 bp or more; analyzing the sequence reads of the set of cell-free nucleic acid molecules to determine genomic characteristics of the fetus; and a method comprising the steps above. **Claim 2** The method according to claim 1, wherein the size of the cell-free nucleic acid molecules is determined using the sequence reads. **Claim 3** The method according to claim 1, wherein the size threshold is 600 bp or more. **Claim 4** The method according to claim 1, wherein the size threshold is 1,000 bp or more. **Claim 5** Determining the size of the cell-free nucleic acid molecules comprises: aligning one or more sequence reads for each of the cell-free nucleic acid molecules to a reference genome, the method according to claim 1. **Claim 6** Determining the size of the cell-free nucleic acid molecules comprises: performing digital PCR with different amplicon sizes, the method according to claim 1. **Claim 7** Determining the size of the cell-free nucleic acid molecules comprises performing electrophoresis techniques, the method according to claim 1. **Claim 8** The method according to claim 7, wherein the assaying is performed after identifying a set of cell-free nucleic acid molecules using the electrophoresis techniques. **Claim 9** The assaying comprises sequencing the entirety of each of the cell-free nucleic acid molecules, thereby generating one sequence read for each of the cell-free nucleic acid molecules, and determining the size of the cell-free nucleic acid molecules comprises counting nucleotides in the sequence reads of the cell-free nucleic acid molecules, the method according to claim 1. **Claim 10** A method comprising: Receiving a blood sample from a female who is pregnant with a fetus, wherein the blood sample contains extracellular particles and particle-free nucleic acids, and the extracellular particles contain cell-free nucleic acids inside the membrane, Performing a physical separation technique that preferentially selects at least a portion of the extracellular particles, thereby obtaining a particle-concentrated sample, Processing the particle-concentrated sample using a processing technique for removing excess particle-free nucleic acids, thereby obtaining a processed particle-concentrated sample, wherein the processing technique includes washing the particle-concentrated sample with an ionic solution and applying a nuclease to the particle-concentrated sample, and the processing technique increases the fractional concentration of fetal nucleic acids in the processed particle-concentrated sample compared to the particle-concentrated sample, Exposing cell-free nucleic acid molecules from the extracellular particles by disrupting the membrane of the extracellular particles, Assaying the cell-free nucleic acid molecules to obtain sequence reads, Analyzing the sequence reads to determine genomic characteristics of the fetus or the pregnancy of the female, A method comprising.
11. The method according to claim 1 or claim 10, wherein the assaying includes sequencing using a sequencing technique or digital PCR.
12. The method according to claim 10, wherein the ionic solution is phosphate-buffered saline (PBS).
13. A method comprising: Receiving a blood sample from a female who is pregnant with a fetus, wherein the blood sample contains extracellular particles and particle-free nucleic acids, and the extracellular particles contain cell-free nucleic acids inside the membrane, Performing one or more purification steps to concentrate the extracellular particles, thereby generating a concentrated sample, Exposing cell-free nucleic acid molecules from the extracellular particles by disrupting the membrane of the extracellular particles, Sequencing the cell-free nucleic acid molecules using a sequencing technique to obtain sequence reads, wherein at least a portion of the sequence reads exceeds 600 bp, Analyzing the sequence reads to determine genomic characteristics of the fetus or the pregnancy of the female, A method comprising.
14. The method according to claim 13, wherein at least a portion of the sequence reads exceeds 1000 bp.
15. A method comprising: Receiving a blood sample from a woman who is pregnant with a fetus, wherein the blood sample contains extracellular particles and particle-free nucleic acids, and the extracellular particles contain cell-free nucleic acids inside the membrane, Performing one or more purification steps to concentrate the extracellular particles, thereby generating a concentrated sample; Exposing cell-free nucleic acid molecules from the extracellular particles by disrupting the membrane of the extracellular particles, wherein at least a portion of the cell-free nucleic acid molecules from the extracellular particles are at least 600 bp; Applying a fragmentation technique to the cell-free nucleic acid molecules; After applying the fragmentation technique, sequencing the cell-free nucleic acid molecules using a sequencing technique to obtain sequence reads; Analyzing the sequence reads to determine genomic characteristics of the fetus or the pregnancy of the woman; A method comprising.
16. The method according to claim 15, wherein the fragmentation technique comprises one or more selected from the group consisting of mechanical shearing, enzymatic fragmentation, nuclease treatment, light, sonication, chemical DNA fragmentation, or bisulfite treatment.
17. The method according to any one of claims 10, 13, and 15, wherein the genomic characteristics are those of the pregnancy, and the genomic characteristics of the pregnancy are related to one or more complications that reduce the woman carrying the fetus to full term.
18. The one or more purification steps are: Performing a physical separation technique that preferentially selects at least a portion of the extracellular particles, thereby obtaining a particle concentrated sample; Processing the particle concentrated sample using a processing technique to remove excess particle-free nucleic acids, thereby obtaining a processed particle concentrated sample, wherein the processing technique includes washing the particle concentrated sample with an ionic solution or applying a nuclease to the particle concentrated sample;
19. The method according to claim 10 or claim 18, wherein the processing technique includes washing the particle concentrated sample with the ionic solution and applying the nuclease to the particle concentrated sample.
20. The method according to claim 10 or claim 18, wherein the processing technique includes washing with the ionic solution, and the ionic solution is PBS.
21. The method according to claim 18 or claim 20, wherein the processing technique comprises applying the nuclease to the particle-concentrated sample.
22. The method according to claim 10 or claim 20, wherein the nuclease is selected from the group consisting of DNase I, TREX1 (3'-repair exonuclease 1), AEN (apoptosis-enhancing nuclease), EXO1 (exonuclease 1), DNASE2 (deoxyribonuclease 2), ENDOG (endonuclease G), APEX1 (apurinic / apyrimidinic endodeoxyribonuclease 1), FEN1 (flap structure-specific endonuclease 1), DNASE1L1 (deoxyribonuclease 1-like 1), DNASE1L2 (deoxyribonuclease 1-like 2), and EXOG (exo / endonuclease G).
23. The method according to claim 10 or claim 18, wherein the physical separation technique preferentially selects particles that are below an upper threshold value and above a lower threshold value.
24. The method according to claim 10 or claim 18, wherein the physical separation technique comprises at least one stage of centrifugation.
25. The method according to claim 24, wherein the centrifugation is performed at 16,000 g or more for at least 10 minutes.
26. The method according to any one of claims 1, 10, 13, and 15, wherein filtration or flow cytometry using one or more filters is used to concentrate or select the extracellular particles of a specified size.
27. The method according to any one of claims 1, 13, and 15, wherein the one or more purification steps comprise centrifugation.
28. The method according to any one of claims 1, 10, 13, and 15, wherein the extracellular particles larger than the specified size are preferentially selected or concentrated, and the specified size is at least 200 nm.
29. Analyzing the sequence reads comprises using the sequence reads to determine the difference in the number of alleles at heterozygous loci of two maternal haplotypes; and using the difference in the number of alleles to determine the genetic haplotypes for each of a plurality of regions, wherein the average haplotype block size is less than 2 Mb. The method according to claims 1, 10, 13, and 15, comprising.
30. Analyzing the sequence reads comprises determining the genotype of the fetus at a locus by aligning the sequence reads to a reference genome and when at least 15% of the sequence reads contain a first allele at the locus, determining that the locus contains the first allele, the method according to any one of claims 1, 10, 13, and 15. **Claim 31** The method according to claim 30, wherein the genomic characteristic is that of the fetus and the genotype exhibits a mutation. **Claim 32** analyzing the sequence reads is determining the haplotype of the fetus by aligning sequence reads longer than 600 bp with each other, wherein the haplotype is the genomic characteristic of the fetus, the aligned sequence reads share a heterozygous locus with the same allele, and at least a part of the aligned sequence reads contains a plurality of heterozygous loci, the method according to any one of claims 1, 10, 13, and 15. **Claim 33** The method according to any one of claims 1, 10, 13, and 15, wherein the genomic characteristic of the fetus is a sequence imbalance at a locus or region of the fetal genome of the fetus. **Claim 34** Disrupting the membrane of the extracellular particles includes mechanical disruption, sonication, enzymatic hydrolysis, detergents, osmotic shock, or freeze-thaw, the method according to any one of claims 1, 10, 13, and 15. **Claim 35** The method according to any one of claims 1, 10, 13, and 15, wherein the sequence reads are also obtained from self-free nucleic acid molecules bound to the membrane of the extracellular particles. **Claim 36** The method according to any one of claims 11, 13, and 15, wherein the sequencing technique includes single molecule sequencing. **Claim 37** The method according to claim 36, wherein the sequencing technique uses nanopores. **Claim 38** The method according to any one of claims 11, 13, and 15, wherein the sequencing technique includes linked-read sequencing. **Claim 39** The method according to any one of claims 11, 13, and 15, wherein the sequencing technique includes methylation-aware sequencing. **Claim 40** said analyzing is for each of a plurality of sequence reads, determining the methylation pattern at the CpG sites of the sequence reads, thereby determining the methylation pattern. aligning the array reads to genomic locations within a reference genome; comparing the methylation pattern to a reference methylation pattern of fetal tissue at the genomic location; identifying, based on the comparing, the array reads as corresponding to fetal nucleic acid molecules; The method according to claim 39, comprising: **Claim 41** wherein the analyzing further comprises: determining whether the fetus has a genomic abnormality using the array reads identified as corresponding to fetal nucleic acid molecules based on the methylation pattern, the method according to claim 40; **Claim 42** wherein the analyzing further comprises: determining one or more haplotypes of the fetus using the array reads identified as corresponding to fetal nucleic acid molecules based on the methylation pattern, the method according to claim 40; **Claim 43** The method according to claim 42, wherein determining the one or more haplotypes of the fetus comprises determining a first maternal haplotype as being inherited by the fetus. **Claim 44** The method according to claim 42, wherein determining the one or more haplotypes of the fetus comprises determining a first paternal haplotype as being inherited by the fetus. **Claim 45** identifying array reads as having fetal-specific alleles; determining a methylation pattern at CpG sites of the array reads; The method according to claim 39, further comprising determining whether the fetus has an epigenetic abnormality using the methylation pattern. **Claim 46** The method according to claim 45, wherein the epigenetic abnormality is fragile X syndrome. **Claim 47** A computer product comprising a non-transitory computer-readable medium storing a plurality of instructions that, when executed, cause a computer system to perform the method according to any one of the preceding claims. **Claim 48** A system comprising: the computer product according to claim 47; and one or more processors for executing instructions stored on the computer-readable medium. **Claim 49** A system comprising means for performing any of the above methods. **Claim 50** A system comprising one or more processors configured to perform any of the above methods. **Claim 51** A system comprising modules that respectively perform the steps described in any of the above methods.