Methods for identification of cell-free DNA fragments of maternal origin in plasma
By employing genomic site analysis and machine learning to identify maternal-specific alleles, the method addresses the challenge of distinguishing maternal cfDNA in NIPT, improving the accuracy and applicability of non-invasive prenatal testing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- IDENTIFAI GENETICS LTD
- Filing Date
- 2025-10-30
- Publication Date
- 2026-05-07
AI Technical Summary
Existing non-invasive prenatal testing (NIPT) methods struggle to effectively distinguish and isolate cell-free DNA (cfDNA) fragments of maternal origin from those of fetal origin in maternal plasma, as maternal alleles are often shared with the fetus, making differentiation challenging.
A method involving genomic site analysis, machine learning, and deep learning to identify maternal-specific alleles and extract fragmentomic features, using techniques such as Bayesian statistics and machine learning models to classify cfDNA reads as maternal or fetal, without requiring fetal genotyping.
Enables accurate identification and classification of maternal cfDNA fragments, enhancing the precision of NIPT by improving genetic predictions and enabling targeted prenatal interventions.
Smart Images

Figure IL2025050957_07052026_PF_FP_ABST
Abstract
Description
[0001] METHODS FOR IDENTIFICATION OF CELL-FREE DNA FRAGMENTS OF MATERNAL ORIGIN IN PLASMA TECHNOLOGICAL FIELD
[0002] The present disclosure relates to the field of genetic analysis.
[0003] REFERENCES:
[0004] Rabinowitz T, et al. Bayesian-based noninvasive prenatal diagnosis of single-gene disorders. Genome Res. 2019. https: / / doi.org / 10.1101 / gr.235796.118.
[0005] Rabinowitz T, Deri-Rozov S, Shomron N. et al., Improved noninvasive fetal variant calling using standardized benchmarking approaches. Computational and Structural Biotechnology Journal. 2021; 19:509–17.
[0006] BACKGROUND
[0007] Cell-free DNA (cfDNA) present in maternal blood during pregnancy is typically used for analyzing the fetal genome.
[0008] Rabinowitz et al (2019) describes a genome wide non-invasive prenatal test (NIPT) approach, termed noninvasive prenatal variant calling. Using Hoobari, the first noninvasive fetal variant caller, they were able to genotype all fetal positions, including biparental loci and indels (US2021 / 0340601). Hoobari, is a tool which employs a Bayesian algorithm to predict the inheritance of monogenic diseases, irrespective of their mode of inheritance or parental origin (Rabinowitz et al. 2019; US2021 / 0340601; Rabinowitz et al. 2021). This tool enables the estimation of the likelihood that each cfDNA fragment is of fetal origin by considering its length. Hoobari can detect mutations caused by single nucleotide polymorphisms (SNPs), or small indels - insertions and deletions of bases in the genome.
[0009] GENERAL DESCRIPTION
[0010] In one aspect the present invention provides a method for identifying fragmentomic features typical to cell free DNA sequence reads of maternal origin in a biological sample comprising: a. receiving reads of sequencing data of (i) maternal cell-free DNA, and optionally (ii) maternal genomic DNA (gDNA) (iii) paternal gDNA from a pair parenting the fetus, and / or (iv) fetal gDNA;
[0011] b. identifying genomic sites in which the maternal genotype is heterozygous (referred to as “AB”), and the fetal genotype is homozygous (referred to as “AA” or “BB”), thereby each of said genomic sites is represented by a maternal-fetal allele combination of AB-AA or AB-BB,
[0012] c. identifying the maternal-specific allele in each of said maternal-fetal allele combinations, wherein in the AB-AA combination the maternal- specific allele is B, and wherein in the AB-BB combination the maternal- specific allele is A,
[0013] d. identifying cfDNA reads comprising the maternal-specific allele in each of said maternal-fetal allele combinations, wherein in the AB-AA combination the maternal-specific allele is B, and wherein in the AB-BB combination the maternal-specific allele is A;
[0014] e. obtaining information on one or more fragmentomic features for each cfDNA read comprising a maternal-specific allele (referred to as maternal fragment); and optionally
[0015] f. introducing the obtained information on the fragmentomic features of the maternal fragment to a machine learning model;
[0016] thereby identifying fragmentomic features typical to sequence reads of maternal origin.
[0017] In one embodiment, the information on the one or more fragmentomic features is obtained using a method selected from a group consisting of descriptive statistics, manual filtration threshold definition, inferential statistics, regression analysis, clustering, graphbased approaches, Bayesian statistics, Markovian models, statistical learning models, machine learning models, and deep learning models.
[0018] In one embodiment, identifying genomic sites in which the mother is heterozygous, and the fetus is homozygous is performed by genotyping maternal gDNA and fetal gDNA.
[0019] In one embodiment, identifying genomic sites in which the mother is heterozygous, and the fetus is homozygous is performed by genotyping the maternal cfDNA and optionally the maternal gDNA and predicting fetal homozygosity by calculating total fetal fraction and the VAF of the fetal fraction in the cfDNA.
[0020] In one embodiment, identifying genomic sites in which the mother is heterozygous, and the fetus is homozygous is performed by genotyping the maternal gDNA and sequencing the maternal cfDNA and predicting fetal homozygosity by noninvasive fetal variant calling.
[0021] In one embodiment, said variant calling is performed by applying a Bayesian procedure.
[0022] In one embodiment, said Bayesian procedure comprises prior probabilities calculated using sequencing data of at least one of said parents.
[0023] In one embodiment, said predicting fetal homozygosity is performed using the Hoobari algorithm.
[0024] In one embodiment, the method further comprises extracting a phased fetal VCF. In one embodiment, identifying genomic sites in which the mother is heterozygous, and the fetus is homozygous is performed by extracting Y chromosome features.
[0025] In one embodiment, the one or both gDNA sequencing data and the cfDNA sequencing data is obtained by deep whole genome sequencing (WGS), whole exome sequencing (WES), targeted sequencing, panel sequencing, gene sequencing, long-read genome sequencing, paired-end sequencing, single read sequencing, or amplicon sequencing.
[0026] In one embodiment, said WGS or WES data is obtained by deep sequencing. In one embodiment, said fragmentomic features comprise one or more of read quality mapping, read base qualities, fragment length, short / long read ratio, end motifs, cleavage patterns around methylation sites, read endpoint preferred end, DNA / accessibility / nucleosome, distance to nearest nucleosome, transcription factor binding sites, regional fetal fraction, regional sequence composition, read sequence composition, and number of sequence errors in the read.
[0027] In one embodiment, said machine learning model is a read classifier machine learning model.
[0028] In one embodiment, said machine learning model is selected from a group consisting of clustering, association rule algorithms, feature evaluation algorithms, subset selection algorithms, support vector machines, classification rules, cost - sensitive classifiers, vote algorithms, stacking algorithms, Bayesian networks, decision trees, random forest algorithms, neural networks, convolutional neural networks, instance -based algorithms, linear modeling algorithms, k - nearest neighbors ( KNN ) analysis, ensemble learning algorithms, boosting algorithms probabilistic models, graphical models, logistic regression methods (including multinomial logistic regression methods), gradient ascent methods, dimensionality reduction methods, singular value decomposition methods, principal component analysis, and a combination thereof.
[0029] In one embodiment, said reads of sequencing data are short reads.
[0030] In one embodiment, said reads of sequencing data are long reads.
[0031] In one embodiment, said reads of sequencing data comprise short reads and long reads.
[0032] In another aspect, the present invention provides a method for classifying sequence reads as being of maternal origin in cfDNA obtained from a biological sample obtained from a pregnant woman, comprising:
[0033] a. receiving reads of sequencing data of said cell-free DNA (cfDNA) obtained from a biological sample obtained from a pregnant woman; b. extracting fragmentomic features for each of said reads; and c. analyzing the fragmentomic features to identify features typical to cfDNA reads of maternal -origin;
[0034] wherein a sequence read presenting features typical to cfDNA reads of maternal origin is classified as being of maternal origin.
[0035] In one embodiment, the features typical to cfDNA sequence reads of maternal origin are identified using the methods of the invention as described above.
[0036] In another aspect, the present invention provides a method of determining whether a fetus possesses a genetic disease, comprising:
[0037] a. Obtaining cfDNA from a biological sample of a pregnant woman; b. obtaining sequencing data of said cfDNA;
[0038] c. classifying sequence reads as being of maternal origin, whereby the remaining reads are defined as being of fetal origin;
[0039] d. analyzing said reads of fetal origin to determine the presence or absence of a mutation associated with a genetic disease; and e. Only if said analyzing identifies that the fetus possesses a mutation associated with a genetic disease:
[0040] (i) administering a prenatal or a post-natal treatment for said genetic disease in an amount effective to prevent or treat said disease, wherein said treatment comprises pharmaceutical based intervention, surgery, genetic therapy, nutritional therapy, or combinations thereof; or
[0041] (ii) performing a pregnancy termination.
[0042] In one embodiment, said classifying sequence reads as being of maternal origin is performed using the methods of the invention.
[0043] In another aspect, the present invention provides a computer software product, comprising a computer-readable medium in which program instructions are stored, which instructions, when read by a data processor, configure the data processor to (1) receive reads of sequencing data of (i) maternal cell-free DNA, and optionally (ii) maternal genomic DNA (gDNA) (iii) paternal gDNA from a pair parenting the fetus, and / or (iv) fetal gDNA, and to (2) execute the method according to the invention.
[0044] In another aspect, the present invention provides a system for identifying fragmentomic features typical to sequence reads of maternal origin and / or for classifying sequence reads as being of maternal origin in cfDNA obtained from a pregnant woman, comprising: an input utility for receiving reads of sequencing data of (i) maternal cell-free DNA, and optionally (ii) maternal genomic DNA (gDNA) (iii) paternal gDNA from a pair parenting the fetus, and / or (iv) fetal gDNA; and a data processor configured for analyzing said data for executing the method according to the invention.
[0045] BRIEF DESCRIPTION OF THE DRAWINGS
[0046] For better understanding the subject matter that is disclosed herein and to exemplify how it may be carried out in practice, embodiments will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which:
[0047] Figure 1 is an outline exemplifying a method suitable for identifying sequence reads of maternal origin, according to various exemplary embodiments of the present invention. Figure 2 is a graph showing the frequency of various fragment lengths (no. of nucleotides per fragment) of maternal and fetal fragments.
[0048] DETAILED DESCRIPTION OF EMBODIMENTS
[0049] Biological samples obtained from a pregnant woman (e.g., blood plasma samples) comprise cfDNA of both the mother and the fetus. The present invention provides methods and systems for identifying cfDNA of maternal origin in such biological samples.
[0050] Most applications of cfDNA for non-invasive prenatal testing (NIPT) focus on extracting fetal fragments. This is possible thanks to alleles that are not found in the mother but are found in the fetus, typically alleles that the fetus inherits from the father. For instance, if the father carries two copies of an allele A (denoted A / A) and the mother carries two copies of allele G (denoted G / G), the fetal genotype is expected to be A / G, inheriting A from the mother and G from the father. In such cases, sequenced DNA fragments (i.e., reads) corresponding to the G allele are necessarily fetal-specific.
[0051] Identifying maternal cfDNA-derived fragments is more challenging, since one of her alleles is always shared with the fetus by inheritance. In the above example, allele G is shared between the mother and fetus and is not maternal-exclusive. This is termed ‘shared allele’ rather than maternal-specific. Moreover, while fetal DNA is often unique and can be distinguished from the mother, the maternal-derived cfDNA is always present in the maternal plasma, thus making its differentiation from fetal DNA and its isolation during pregnancy a more challenging task.
[0052] Therefore extraction / identification of maternal-specific reads (i.e. non-fetal reads) requires a unique approach, as provided in the present invention.
[0053] The identification of maternal-specific reads can be useful in various manners, for example for improving genetic predictions in non-invasive prenatal tests (NIPT), or for finding variants that can determine the best way to give birth. For example, in various factor deficiencies there is a preference for c-section delivery and avoidance of vacuum assistance or forceps to reduce the risk of intracranial hemorrhage. Examples of factor deficiencies include, but are not limited to FI (fibrinogen), FII (Prothrombin), FV, FVII, FVIII (Hemophilia A), FIX (Hemophilia B), FX, FXI (Hemophilia C) FXII, and FXIII. Figure 1 shows an exemplary outline of an embodiment of the method of the invention exemplifying a workflow for identifying sequence reads of maternal origin in cfDNA obtained from plasma of a pregnant woman. This embodiment comprises an invasive step of obtaining fetal genomic DNA information. As a result of the analysis as will be described below, features of maternal-specific reads are learnt and may then be implemented in further methods of NIPT which do not require fetal genotyping for identifying reads of maternal origin.
[0054] Thus, in accordance with this first embodiment, maternal (and optionally paternal) as well as fetal genomic DNA are obtained from biological samples and sequenced.
[0055] The maternal (and optionally paternal) genomic DNA may be obtained from any cellular or tissue source as known in the art, for example but not limited to, from a blood sample, or saliva e.g., from peripheral blood mononuclear cells (PBMC).
[0056] Fetal genomic DNA may be obtained from any cellular or tissue source as known in the art, for example but not limited to, from amniotic fluid, chorionic villus sampling (CVS), cord blood, or from any fetal tissue taken post-natal.
[0057] In addition, cell-free DNA from a maternal biological sample (e.g., maternal plasma) which comprises both cfDNA of maternal origin and cfDNA of fetal origin is obtained and sequenced.
[0058] The maternal and fetal genotype is predicted by alignment and mapping to a human reference genome and variant calling to identify genetic variants.
[0059] Upon analysis of the maternal and fetal genotype, a set of variants is defined in which the maternal genotype is heterozygous (namely having different alleles for the same genetic locus), and the fetal genotype is homozygous (namely having the same allele at the examined genetic locus). For demonstration purposes, these alleles are defined as alleles A and B, i.e., denoted AB for the heterozygous locus in the maternal genome, and AA or BB for the homozygous locus in the fetal genome. Therefore, two possible maternal-fetal genotype combinations comprise the set of variants: ABAA and ABBB depending on the type of homozygosity in the fetus.
[0060] Next, the maternal-specific allele in each position in the set of positions described above is defined. In ABAA positions, the maternal-specific allele is B; in ABBB positions, the maternal-specific allele is A.
[0061] cfDNA sequence fragments carrying the maternal specific allele are thus defined as maternal cfDNA fragments or reads. Features and characteristics of the maternal fragments are learned using statistical, artificial intelligence (AI) including machine learning (ML) and deep learning (DL), or by any method known in the art.
[0062] Information learned using the maternal cfDNA fragments is used for classification of new, unseen cfDNA fragments, to detect maternal fragments without using the supported allele information, or the maternal or fetal genomic information.
[0063] In one embodiment, termed herein “a cross-sample method”, information learned from one or more family cases is applied over another, different family case, which was not used for learning.
[0064] In one embodiment, termed herein “an in-sample method”, information learned from the first variant set (i.e., the ABAA and ABBB set) is utilized on other sets of variants within the same family.
[0065] In one embodiment, a combined method integrates both the cross-sample and the in-sample methods.
[0066] In one embodiment, the method of the invention is performed using variant allele fraction (VAF)-based boosting without the use of fetal genomic DNA information.
[0067] As used herein, the term VAF refers to the ratio of reads showing the alternate allele divided by the number of all reads covering a variant.
[0068] The method comprises the following steps:
[0069] (i) Maternal cfDNA and optionally maternal genomic DNA are sequenced and mapped to the human reference genome
[0070] (ii) In an embodiment wherein, maternal genomic DNA was sequenced, a high-confidence maternal heterozygous positions list is created, namely, based on the mother’s genomic DNA analysis one or more locations are selected in which the mother is heterozygous. Maternal heterozygous loci undergo stringent filtering based on multiple criteria, for instance, the depth of coverage, to keep only high confidence predictions. In an embodiment, only positions where the VAF is ~0.5 are kept.
[0071] In an embodiment wherein maternal genomic DNA was not sequenced, and only cfDNA sequence data is provided, maternal-heterogenous-fetal-homogenous positions are determined with mitigated confidence.
[0072] (iii) A maternal heterozygous and fetal homozygous positions list is created, namely, the plasma cfDNA is examined in positions corresponding to the maternal heterozygous positions list. A VAF equal to 0.5 ± (0.5 * fetal fraction) suggests fetal homozygosity. Stringent filtering is done based on VAF and other parameters.
[0073] For example, if the fetal fraction is 10%, in a total of 100 fragments, 90 fragments originate from the mother and 10 from the fetus. The mother is heterozygous; therefore 45 fragments comprise allele A and 45 fragments comprise allele B. The fetus is homozygous and hence all 10 fetal fragments comprise allele A.
[0074] This calculation results in a total number of 55 A fragments and 45 B fragments: 55 / 100 = 50 / 100 + 5 / 100
[0075] 5 / 100 = 5% = 1 / 2 of the fetal fraction.
[0076] VAF when the mother is heterozygous and fetus is homozygous = 0.5*(maternal fraction) + l*(fetal fraction) = 0.5*(1-fetal fraction) + fetal_fraction = 0.5 - 0.5*ff + ff = 0.5 + 0.5*ff.
[0077] (iv) Maternal-specific reads are extracted, namely, in positions where the mother is heterozygous and the fetus is homozygous, reads supporting a non-fetal allele are assumed maternal. If, for instance, the maternal genotype is AB and the fetal genotype BB, then reads showing allele A are assumed maternal.
[0078] In another embodiment, the method of the invention comprises genotype-based boosting. Namely, maternal (and optionally paternal) genomic DNA and maternal cfDNA are sequenced, and after alignment and mapping to a human reference genome, variant calling is performed, for example using Hoobari (Rabinowitz et al 2019) to identify variants. Then, positions where the maternal genotype is 0 / 1, and the fetal predicted genotype is either 0 / 0 or 1 / 1 are extracted. Only high confidence variant calls are kept. Maternal fragments are extracted, Hoobari is run once again while utilizing the maternal reads (to estimate the maternal noise or further train the read-classifier).
[0079] In accordance with this embodiment Hoobari serves as a boosting algorithm, as high confidence results are used to further improve the analysis and re-run the algorithm.
[0080] In another embodiment, the method of the invention comprises haplotype-based Boosting. In accordance with this embodiment, a phased fetal VCF is extracted using Hoobari. This allows to extract reads corresponding to the maternal haplotype. Accordingly, greater confidence is achieved in the first group of positions extracted using Hoobari. Even in positions where the fetus is 0 / 1, it would be possible to identify the maternal read, since the fetus would be either 0|1 or 1|0, namely, it becomes possible to tell which allele was inherited from which parent; in 0|l heterozygosity, allele 0 was inherited from the mother, and in 1|0 heterozygosity it was inherited from the father. It would therefore be possible to make a deductive reasoning of the linkage to the first group’s position.
[0081] In another embodiment, the method of the invention further comprises a Male-X Solution. Namely, in male fetuses, it is possible to extract fetal fragments using the Y chromosome. There are regions in chromosome X that are not shared with the Y chromosome. These regions are termed non-pseudo autosomal regions (non-PAR), and they appear only once in the fetus. In such reads or regions, the expected minor allele frequency, which is based on the fetal fraction, genotype, and haplotype information, is used to extract maternal-specific reads.
[0082] Accordingly, in a first aspect, the present invention provides a method for identifying fragmentomic features typical to cell free DNA sequence reads of maternal origin in a biological sample comprising:
[0083] a. receiving reads of sequencing data of (i) maternal genomic DNA (gDNA) and optionally paternal gDNA from a pair parenting the fetus, (ii) maternal cell-free DNA (cfDNA), and optionally (iii) fetal gDNA. b. identifying genomic sites in which the maternal genotype is heterozygous (referred to as “AB”), and the fetal genotype is homozygous (referred to as “AA” or “BB”), thereby each of said genomic sites is represented by a maternal-fetal allele combination of AB-AA or AB-BB,
[0084] c. identifying the maternal-specific allele in each of said maternal-fetal allele combinations, wherein in the AB-AA combination the maternal- specific allele is B, and wherein in the AB-BB combination the maternal- specific allele is A,
[0085] d. identifying cfDNA reads comprising the maternal-specific allele in each of said maternal-fetal allele combinations, wherein in the AB-AA combination the maternal-specific allele is B, and wherein in the AB-BB combination the maternal-specific allele is A;
[0086] e. obtaining information on one or more fragmentomic features for each cfDNA read comprising a maternal-specific allele (referred to as maternal fragment); and optionally f. introducing the obtained information on the fragmentomic features of the maternal fragment to a machine learning model;
[0087] thereby identifying fragmentomic features typical to sequence reads of maternal origin.
[0088] The present invention also provides a method for classifying sequence reads as being of maternal origin in cfDNA obtained from a pregnant woman, comprising:
[0089] a. receiving reads of sequencing data of said cell-free DNA (cfDNA) obtained from a pregnant woman;
[0090] b. extracting fragmentomic features for each of said reads; and
[0091] c. analyzing the fragmentomic features to identify features typical to cfDNA reads of maternal-origin;
[0092] wherein a sequence read presenting features typical to cfDNA reads of maternal origin is classified as being of maternal origin.
[0093] Cell-free DNA (cfDNA) also referred to as “circulating free DNA” are DNA fragments existing outside of cells in vi vo circulating in body fluids such as blood plasma. The fragments of cfDNA typically have lengths ranging from about 150 to 200 base pairs (bp), and averaging about 170 bp, which presumably relates to the length of a DNA stretch wrapped around a nucleosome. During pregnancy, cell-free fetal DNA can be found circulating in maternal plasma. Thus, the cfDNA in maternal plasma is a mixture of both maternal and fetal DNA, both the total amount of cfDNA, and the fraction of fetal DNA within it, increases throughout pregnancy.
[0094] The term cfDNA also refers to fragments of DNA that have been obtained from the in vivo extracellular sources and separated, isolated, or otherwise manipulated in vitro. cfDNA can be obtained by extracting DNA from blood plasma after removal of intact cells. Methods for extracting cfDNA are well known in the art, for example, as shown in the Examples below.
[0095] The term “genomic DNA” or “gDNA” herein refers to DNA existing in a cell in vivo and containing a complete genome of the cell or organism. The term also refers to DNA that has been obtained from the in vivo cell and separated, isolated, or otherwise manipulated in vitro. Typically, the cell is isolated prior to being subjected to lysis to produce in vitro cellular DNA. The term gDNA used herein does not include cfDNA. Genomic DNA can be obtained for example, but not limited to, from blood cells. Fetal genomic DNA can be obtained, for example, but not limited to, from cells in the placenta or amniotic fluid.
[0096] To obtain DNA sequencing data, the DNA containing samples are subjected to DNA sequencing methods including, but not limited to, deep whole genome sequencing (WGS), whole exome sequencing (WES), next generation sequencing (NGS), targeted sequencing, panel sequencing, gene sequencing, long-read genome sequencing, paired-end sequencing, single end sequencing, and amplicon sequencing.
[0097] Paired-end sequencing obtains one read from each end of a nucleic acid fragment resulting in “paired-end reads”.
[0098] The term “Next Generation Sequencing” (NGS) herein refers to sequencing methods that allow for massively parallel sequencing of clonally amplified molecules and of single nucleic acid molecules. Non-limiting examples of NGS include sequencing-by-synthesis using reversible dye terminators, and sequencing-by-ligation.
[0099] Deep sequencing refers to sequencing a genomic region multiple times, sometimes hundreds or even thousands of times. Deep sequencing of the genome allows researchers to detect rare genetic variants.
[0100] As used herein the term “deep whole genome sequencing” refers to deep sequencing of the entire genome.
[0101] The sequencing is repeated multiple times, for example, but not limited to between 10 times (10X) and 1000 times (1000X), e.g., 10 times (10X), 20 times (20X), 30 times (30X), 50 times (50X), 100 times (100X), 150 times (X150), 200 times (200X), 300 times (300X), 500 times (500X), or 1000 times (1000X).
[0102] Plasma cfDNA can be subjected to varying sequencing depths.
[0103] In one non-limiting example, the cfDNA in plasma is sequenced 300 times (300X); in other embodiments, the cfDNA in maternal plasma is sequenced 50 times (50X), 100 times (100X), or 150 times (X150).
[0104] In addition, genomic DNA is also subjected to whole genome sequencing. Such genomic DNA may be obtained from any cell type, for example from blood cells, e.g., leukocytes. In an embodiment, whole genome sequencing of the paternal and maternal genomic DNA is performed to a targeted depth of about 20X to 40X, for example 30X.
[0105] Whole genome sequencing may be performed using any method known in the art for short or long reads, for example, but not limited to, sequencing by synthesis using Illumina’s sequencing series (e.g., the HiSeq X Ten System, HiSeq 4000, NovaSeq 6000, and Novaseq X and X-plus), WGS by Ultima Genomics sequencing machines, pore-based sequencing, e.g., nanopore WGS sequencing using MinION device (Oxford Nanopore Technologies),, as well as sequencing systems by PacBio technologies.
[0106] The sequencing generates “reads” which are sequences of DNA fragments of varying lengths. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A T C G). It may be stored in a memory device and processed as appropriate. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information.
[0107] The sequencing input may be of long reads (e.g., from about 1 KBP (kilobase pairs) to about 100KBP, or more) or short reads (e.g., from between about 50 base pairs and 400 base pairs), or a combination of long reads and short reads.
[0108] After sequencing, the reads are aligned to a human reference genome based on sequence similarities.
[0109] The identification of the genetic variants (i.e., variant sites or mutations) can be performed using a variant calling approach, which is generally based on alignment of the DNA sequencing data and the application of a commercially available variant caller.
[0110] As used herein, the terms “aligned”, “alignment”, or “aligning” refer to the process of comparing a read to a reference sequence and thereby determining whether the read is contained in the reference sequence. If the reference sequence contains the read, the read may be mapped to a particular location in the reference sequence. In some cases, alignment simply tells whether the read is present or absent in the reference sequence.
[0111] Sequence alignment techniques that can be used according to some embodiments of the present invention include, without limitation, Burrows Wheeler Aligner (BWA), ABA, ALE, AMAP, anon, BAli-Phy, Base-By-Base, BHAOS / DIALIGN, Bowtie, Bowtie 2, ClustalW, CodonCode Aligner, Comass, DECIPHER, DIALIGN-TX, DIALIGN-T, DNA Alignment, DNA Baser Sequence Assembler, EDNA, FSA, Geneious, Kalign, MAFFT, MARNA, MA VID, MSA, MSAProbs, MULTALIN, Multi-LAGEN, MUSCLE, Opal, Pecan, Phylo, Praline, PicXAA, POA, Probalign, ProbCons, PROMALS3D, PRRN / PRRP, PSAlign, RevTrans, SAGA, SAM, Se-AI, STAR, STAR-Fusion, StatAlign, Stemloc, T-Coffee, UGENE, VectorFriends, NovoAlign, and GLProbs. Exemplary variant callers suitable for the present embodiments include, without limitation, Genome Analysis Toolkit (GATK) and Freebayes. For example, Freebayes can comprise an alignment based on literal sequences of reads aligned to a particular target, not their precise alignment. GATK can comprise: (i) pre-processing; (ii) variant discovery; and (iii) callset refinement. Pre-processing can comprise starting from raw sequence data, e.g., in FASTQ or uBAM format, and producing analysis-ready BAM files; processing can include alignment to a reference genome as well as data cleanup operations to correct for technical biases and make the data suitable for analysis; variant discovery can comprise starting from analysis-ready BAM files and producing a callset in VCF format; processing can involve identifying sites where one or more individuals display possible genomic variation, and applying filtering methods appropriate to the experimental design; callset refinement can comprise starting and ending with a VCF callset; processing can involve using metadata to assess and improve genotyping accuracy, attach additional information and evaluate the overall quality of the callset.
[0112] Also contemplated are variant callers such as, but not limited to, Platypus, VarScan, Bowtie analysis, MuTect, Google DeepVariant, and / or SAMtools. For example, Bowtie analysis can comprise implementing the Burrows-Wheeler transform for aligning. MuTect can comprise: (i) pre-processing; (ii) statistical analysis; and (iii) postprocessing. Pre-processing can comprise an initial alignment of sequencing reads; statistical analysis can comprise using two Bayesian classifiers, one classifier can detect whether a SNP is non-reference at a given site and, for those sites that are found as nonreference, the other classifier can make sure that the normal does not carry the SNP; postprocessing can comprise removal of artifacts of sequencing, short read alignments, and hybrid capture. SAMtools can comprise storing, manipulating, and aligning sequencing reads stored as SAM files.
[0113] In an embodiment, the step of determining a probability that a sequence read in the maternal plasma cfDNA is of fetal origin is performed for example by an algorithm, e.g. the algorithm Hoobari as described in Rabinowitz et al., 2019 and US2021 / 0340601, to calculate the fetal fraction (FF), i.e., the percent of fetal derived cfDNA within the maternal plasma cfDNA, and to calculate the fragment length distributions. This step may also comprise extracting various fragment-level characteristics. As used herein the term “fetal fraction” or “FF” refers to the portion of fetal cfDNA, within the total amount of cfDNA in the maternal blood. The portion of fetal cfDNA within maternal blood (the fetal fraction) varies throughout the pregnancy, and between individuals, hence this is not regarded as a constant but as a variable. Low levels of fetal cfDNA are referred to as a low fetal fraction.
[0114] In an embodiment, said determining the probabilities comprises applying a Bayesian procedure. Optionally, said Bayesian procedure comprises prior probabilities calculated using sequencing data of at least one of said parents.
[0115] In an embodiment, this procedure further comprises recalibration of the output of said Bayesian procedure using machine learning.
[0116] In a specific embodiment the determination of the probability, for each DNA fragment (or read), to be of fetal origin is performed using variant calling, for example using the Hoobari algorithm as described in Rabinowitz et al., 2019 and US2021 / 0340601.
[0117] The term “fraginentomic features” refers to molecular characteristics of DNA reads, as well as to genomic, epigenetic and alignment features of the DNA read. Fragmentomic features include, but are not limited to, Read quality mapping, Read base qualities, Fragment length, short / long read ratio, DNA fragment end motifs, Cleavage patterns around methylation sites, Read endpoint preferred end, DNA accessibility / nucleosome positioning inference, Distance to nearest nucleosome, Transcription factor binding sites, Regional fetal fraction, Regional sequence composition, Read sequence composition, and Number of sequence errors in the read.
[0118] The terms “fragment length” and “fragment size” are used interchangeably herein to refer to a parameter that relates to the length or size of a nucleic acid fragment, e.g., cfDNA fragments obtained from a bodily fluid
[0119] As used herein the term “phasing” or “haplotype phasing” refers to haplotype deduction.
[0120] A method of determining whether a fetus possesses a genetic disease, comprising:
[0121] a. Obtaining cfDNA from a biological sample of a pregnant woman; b. obtaining sequencing data of said cfDNA;
[0122] c. classifying sequence reads as being of maternal origin, whereby the remaining reads are defined as being of fetal origin; d. analyzing said reads of fetal origin to determine the presence or absence of a mutation associated with a genetic disease; and
[0123] e. Only if said analyzing identifies that the fetus possesses a mutation associated with a genetic disease:
[0124] (i) administering a prenatal or a post-natal treatment for said genetic disease in an amount effective to prevent or treat said disease, wherein said treatment comprises pharmaceutical based intervention, surgery, genetic therapy, nutritional therapy, or combinations thereof; or
[0125] (ii) performing a pregnancy termination
[0126] The fetal genotype, namely the presence or absence of a mutation associated with a genetic disease, can be predicted using one or more of the following: descriptive statistics, manual filtration threshold definition, inferential statistics, regression analysis, clustering, graph-based approaches, Bayesian statistics, Markovian models, statistical learning models, machine learning models, or deep learning models.
[0127] Analysis of the sequencing data and the identification of fragmentomic features typical to sequence reads of maternal origin or the classification of sequence reads as being of maternal origin, as well as the diagnosis / genotyping derived therefrom are typically performed using various computer executed algorithms and programs. Therefore, certain embodiments employ processes involving data stored in or transferred through one or more computer systems or other processing systems. Embodiments disclosed herein also relate to apparatus for performing these operations.
[0128] Thus, in an embodiment, the methods of the invention are implemented using a computer comprising one or more processors and system memory.
[0129] EXAMPLES
[0130] Materials and Methods
[0131] Sample collection and DNA extraction
[0132] DNA samples were collected during week 11 of the pregnancy with informed consent. DNA from chorionic villus sampling (CVS) was extracted using the magLEAD 12gC, MagDEA Dx kit (ExScale, Chiba, Japan). Peripheral maternal blood was collected using 2-4 Ethylene-diamine-tetra-acetic acid (EDTA) tubes. Within one hour of collection, plasma was separated from blood by centrifugation at room temperature for 10 minutes at 1600 x g. The plasma was then centrifuged again at 16,000 x g for 10 minutes at room temperature to remove any residual cells. Extraction of cfDNA was performed using the MagMAX™ Cell-Free DNA Isolation Kit (Thermo Fisher). Parental (maternal and paternal) genomic DNA was extracted from peripheral blood mononuclear cells (PBMC) using a standard protocol that includes (i) huffy coat separation and (ii) DNA purification using Mag-Bind® Blood & Tissue DNA kit (Omega Bio-tek Inc) according to the manufacturer's instructions.
[0133] Library preparation and sequencing
[0134] Library preparation was performed using the TruSeq DNA PCR-Free Library Prep Kit (Illumina), for genomic DNA, the Accel-NGS 2S PCR-free Library Prep Kit for cfDNA samples that underwent WGS, and the SureSelect XT HS2-V8 for cfDNA samples that underwent WES, according to the manufacturer's instructions. This was followed by sequencing using the NovaSeq platform (Illumina) targeting 150-bp paired-end reads across every DNA sample from each family unit.
[0135] Paternal and maternal genomic DNA were each sequenced using PCR-free 3 Ox WGS.
[0136] Fetal genomic DNA obtained using Chorionic Villus Sampling (CVS) was sequenced using PCR-free 3 Ox WGS.
[0137] Parental, fetal and cfDNA were analyzed as described in Rabinowitz et al, 2019.
[0138] Alignment to the genome
[0139] Reads were aligned to the Genome Reference Consortium Human Build 38 (GRCh38 / hg38) using Burrows-Wheeler v0.7.834 with default parameters. Duplicate reads and reads mapping to multiple locations were excluded from downstream analysis.
[0140] Variant calling
[0141] Single-nucleotide substitutions and small insertions and deletions were identified using the Genome Analysis Toolkit (GATK) HaplotypeCaller software v4.2.4.0 applying default parameters and Hoobari. Sequence alignment, removal of duplicate readalignments, parental variant identification, and non-invasive fetal variant calling were executed as outlined in Rabinowitz et al, 2019. Noninvasive fetal variant calling
[0142] Hoobari was run using the parental variants and the cfDNA pre-processing results database as input. The output was a standard variant call format (VCF) file. The analysis of the results was held using several software dedicated to VCF manipulation, such as vcflib and vcftools.
[0143] Bayesian noninvasive genotyping
[0144] At each site of interest, a Bayesian calculation was applied. For each possible fetal genotype:
[0145] P(data|G)P(G)
[0146] P(G|data) = ∑i=1nP(data|Gi)P(Gi)
[0147]
[0148] where G is the fetal genotype and Gi is the ith possible fetal genotype out of n possibilities. For bi-allelic variants, it would be either homozygous for the reference allele (AA), heterozygous (Aa), or homozygous for the alternate allele (aa). P(G) is the prior probability for each genotype and was calculated by Mendelian laws. The data variable denotes the reads that cover a site and P(data | G) denotes the likelihood function, which is defined in this Example as a product of the likelihood of each read:
[0149] P(data|G) = ∏j=1mP(rj|G,GM,f) = ∏j=1m(P(rj|fet)P(fet) + P(rj|mat)P(mat)),
[0150]
[0151] The likelihood of a read rj depends on the fetal genotype and is calculated using the maternal genotype and the fetal fraction. P(rj|fet) and P(rj|mat) are the probabilities of a read-observation that supports a certain allele, given that the read is fetal or maternal, respectively. This depends on the tested fetal genotype Gi, the maternal genotype GM and the observed allele. P(fet) and P(mat) are the probabilities of observing a fetal or maternal read based only on the fetal fraction, and regardless of the allele that it supports. To utilize the size differences between fetal and maternal fragments, the fetal fraction used for each read was calculated only from reads with the same fragment size. For reads that are not properly paired or have a fragment size of >500, the total fetal fraction is used. Example 1:
[0152] Fig. 2 shows maternal and fetal fragment length distribution, calculated using fragments obtained from ABAA and ABBB positions
Claims
CLAIMS:
1. A method for identifying fragmentomic features typical to cell-free DNA sequence reads of maternal origin in a biological sample comprising: a. receiving reads of sequencing data of (i) maternal cell-free DNA, and optionally (ii) maternal genomic DNA (gDNA) (iii) paternal gDNA from a pair parenting the fetus, and / or (iv) fetal gDNA;b. identifying genomic sites in which the maternal genotype is heterozygous (referred to as “AB”), and the fetal genotype is homozygous (referred to as “AA” or “BB”), thereby each of said genomic sites is represented by a maternal-fetal allele combination of AB-AA or AB-BB,c. identifying the maternal-specific allele in each of said maternal-fetal allele combinations, wherein in the AB-AA combination the maternal- specific allele is B, and wherein in the AB-BB combination the maternal- specific allele is A,d. identifying cfDNA reads comprising the maternal-specific allele in each of said maternal-fetal allele combinations, wherein in the AB-AA combination the maternal-specific allele is B, and wherein in the AB-BB combination the maternal-specific allele is A;e. obtaining information on one or more fragmentomic features for each cfDNA read comprising a maternal-specific allele (referred to as maternal fragment); and optionallyf. introducing the obtained information on the fragmentomic features of the maternal fragment to a machine learning model;thereby identifying fragmentomic features typical to sequence reads of maternal origin.
2. The method of claim 1 wherein the information on the one or more fragmentomic features is obtained using a method selected from a group consisting of descriptive statistics, manual filtration threshold definition, inferential statistics, regression analysis, clustering, graph-based approaches, Bayesian statistics, Markovian models, statistical learning models, machine learning models, and deep learning models.
3. The method of claim 1 wherein identifying genomic sites in which the mother is heterozygous, and the fetus is homozygous is performed by genotyping maternal gDNA and fetal gDNA.
4. The method of claim 1 wherein identifying genomic sites in which the mother is heterozygous, and the fetus is homozygous is performed by genotyping the maternal cfDNA and optionally the maternal gDNA and predicting fetal homozygosity by calculating total fetal fraction and the VAF of the fetal fraction in the cfDNA.
5. The method of claim 1 wherein identifying genomic sites in which the mother is heterozygous, and the fetus is homozygous is performed by genotyping the maternal gDNA and sequencing the maternal cfDNA and predicting fetal homozygosity by noninvasive fetal variant calling.
6. The method of claim 5 wherein said variant calling is performed by applying a Bayesian procedure.
7. The method of claim 6 wherein said Bayesian procedure comprises prior probabilities calculated using sequencing data of at least one of said parents.
8. The method of any one of claims 5 to 7 wherein said predicting fetal homozygosity is performed using the Hoobari algorithm.
9. The method of claim 5, wherein the method further comprises extracting a phased fetal VCF.
10. The method of claim 1 wherein identifying genomic sites in which the mother is heterozygous, and the fetus is homozygous is performed by extracting Y chromosome features.
11. A method for classifying sequence reads as being of maternal origin in cfDNA obtained from a biological sample of a pregnant woman, comprising:a. receiving reads of sequencing data of said cell-free DNA (cfDNA) obtained from a biological sample of a pregnant woman;b. extracting fragmentomic features for each of said reads; and c. analyzing the fragmentomic features to identify features typical to cfDNA reads of maternal -origin;wherein a sequence read presenting features typical to cfDNA reads of maternal origin is classified as being of maternal origin.
12. The method of claim 11, wherein the features typical to cfDNA sequence reads of maternal origin are identified using the method of any one of claims 1 to 10.
13. A method of determining whether a fetus possesses a genetic disease, comprising:a. Obtaining cfDNA from a biological sample of a pregnant woman; b. obtaining sequencing data of said cfDNA;c. classifying sequence reads as being of maternal origin, whereby the remaining reads are defined as being of fetal origin;d. analyzing said reads of fetal origin to determine the presence or absence of a mutation associated with a genetic disease; ande. Only if said analyzing identifies that the fetus possesses a mutation associated with a genetic disease:(i) administering a prenatal or a post-natal treatment for said genetic disease in an amount effective to prevent or treat said disease, wherein said treatment comprises pharmaceutical based intervention, surgery, genetic therapy, nutritional therapy, or combinations thereof; or(ii) performing a pregnancy termination.
14. The method of claim 13 wherein said classifying sequence reads as being of maternal origin is performed using the method of any one of claims 11 or 12.
15. The method of any one of the preceding claims, wherein one or both gDNA sequencing data and the cfDNA sequencing data is obtained by deep whole genome sequencing (WGS), whole exome sequencing (WES), targeted sequencing, panel sequencing, gene sequencing, long-read genome sequencing, paired-end sequencing, single read sequencing, or amplicon sequencing.
16. The method of claim 14 wherein said WGS or WES data is obtained by deep sequencing.
17. The method of any one of the preceding claims, wherein said fragmentomic features comprise one or more of read quality mapping, read base qualities, fragment length, short / long read ratio, end motifs, cleavage patterns around methylation sites, read endpoint preferred end, DNA / accessibility / nucleosome, distance to nearest nucleosome, transcriptionfactor binding sites, regional fetal fraction, regional sequence composition, read sequence composition, and number of sequence errors in the read.
18. The method of any one of the preceding claims wherein said machine learning model is a read classifier machine learning model.
19. The method of any one of the preceding claims wherein said machine learning model is selected from a group consisting of clustering, association rule algorithms, feature evaluation algorithms, subset selection algorithms, support vector machines, classification rules, cost - sensitive classifiers, vote algorithms, stacking algorithms, Bayesian networks, decision trees, random forest algorithms, neural networks, convolutional neural networks, instance - based algorithms, linear modeling algorithms, k - nearest neighbors ( KNN ) analysis, ensemble learning algorithms, boosting algorithms probabilistic models, graphical models, logistic regression methods (including multinomial logistic regression methods), gradient ascent methods, dimensionality reduction methods, singular value decomposition methods, principal component analysis, and a combination thereof.
20. The method of any one of claims 1-17 wherein said reads of sequencing data are short reads.
21. The method of any one of claims 1-17 wherein said reads of sequencing data are long reads.
22. The method of any one of claims 1-17 wherein said reads of sequencing data comprise short reads and long reads.
23. A computer software product, comprising a computer-readable medium in which program instructions are stored, which instructions, when read by a data processor, configure the data processor to (1) receive reads of sequencing data of (i) maternal cell-free DNA, and optionally (ii) maternal genomic DNA (gDNA) (iii) paternal gDNA from a pair parenting the fetus, and / or (iv) fetal gDNA, and to (2) execute the method according to any one of claims 1-22.
24. A system for identifying fragmentomic features typical to sequence reads of maternal origin and / or for classifying sequence reads as being of maternal origin in cfDNA obtained from a pregnant woman, comprising: an input utility for receiving reads of sequencing data of (i) maternal cell-free DNA, and optionally (ii) maternal genomic DNA (gDNA) (iii) paternal gDNA from a pairparenting the fetus, and / or optionally (iv) fetal gDNA; and a data processor configured for analyzing said data for executing the method according to any one of claims 1-20.