Noninvasive fetal genetic variant identification using distribution bias detection
By analyzing distribution biases in segment features of maternal and fetal DNA, the method addresses the inefficiencies of NIPT in detecting CNVs, achieving accurate fetal genotyping with zero false positives.
Patent Information
- Application Number
- PCT/IL2025/050639
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-25
- Filing Date
- 2025-07-23
- Publication Date
- 2026-01-29
AI Technical Summary
Current non-invasive prenatal testing (NIPT) methods struggle to accurately detect smaller copy number variants (CNVs) in fetal DNA due to high false positive rates and inefficiencies, particularly when analyzing cell-free fetal DNA (cffDNA), which are responsible for up to 15% of developmental disorders in live-born babies.
A method that analyzes the distribution bias of segment features, such as length, methylation patterns, and enriched end motifs, in maternal and fetal DNA to identify potential CNVs by comparing distributions with reference sets, using deep whole genome sequencing and algorithms to filter out false positives.
The method effectively identifies deletions exceeding 100 kilobases with zero false positives, confirmed by amniotic fluid and paternal DNA analysis, demonstrating high precision and reliability in fetal genotyping.
Smart Images

Figure 00000027_0000 
Figure 00000028_0000 
Figure 00000029_0000
Abstract
Description
[0001] NONINVASIVE FETAL GENETIC VAIU’ANT IDENTIFICATION USING DISTRIBUTION BIAS DETECTION
[0002] TECHNOLOGICAL FIELD
[0003] The present disclosure relates to the field of prenatal genetic analysis.
[0004] REFERENCES
[0005] Chiu and Lo, (2021 ) Prenatal Diagnosis 41.10: 1193-1201.
[0006] Hochstenbach et al (2009) Eur. J Med Genet 52, 161-169.
[0007] Layer et al (2014) Genome Biol 15, 1-19.
[0008] Rabinowitz et al (2019) Genome Res. 29 (3), pp. 428-438
[0009] Raman et al (2019) Nucleic Acids Res 47, 1605-1614.
[0010] Shaikh (2017) Curr. Genet. Med. Rep. 14;5(4): 183-190
[0011] Ye et al (2009) Bioinformatics 25, 2865-2871.
[0012] BACKGROUND
[0013] Non-invasive prenatal testing (NIPT) is the process of assessing the health of an unborn fetus by determining the risk that the fetus will be bom with genetic abnormalities. In the last decade, NIPT has emerged as a risk-free alternative to amniocentesis, enabling detection of genetic abnormalities during pregnancy. This method relies on the existence of cell-free fetal DNA (cffDNA) as a fraction of total cell-free DNA (cfDNA) circulating in maternal plasma from the early weeks of gestation through to birth.
[0014] In an NIPT procedure, blood is drawn from the mother, cfDNA is extracted and sequenced and is then used to gain genetic information about the fetus. Current NIPT tests are offered in clinics worldwide and can detect large genetic aberrations on a wholechromosome scale, e.g., aneuploidy and large microdeletions or microduplications. Smaller CNVs that are currently missed by NIPT are responsible for up to 15% of live bom babies with developmental disorders (Hochstenbach et al., 2009). The majority of pathologic CNVs are de-novo. While various tools exist for CNV identification, they are primarily designed for genomic DNA, leading to a high rate of false positives when applied to cffDNA (Layer et al., 2014; Ye et al., 2009). Others, tailored for cffDNA, effectively identify only larger microdeletions above 1 Mb (Raman et al., 2019). Rabinowitz et al (2019) describe a different approach for genome wide NIPT of monogenic disorders, defining them as a unique case of variant calling, termed noninvasive prenatal vanant calling. Accordingly, a Bayesian genotyping algorithm utilizes the information of each read, covering each candidate variant, and a machine learning-based fine-tuning step subsequently incorporates information from previously verified results. By accounting for each read, the authors were able to utilize characteristics that separate fetal and maternal DNA, such as fragment length. The algorithm was implemented as Hoobari, the first nonin vasive fetal variant caller, that was able to genotype all fetal positions, including biparental loci and indels. However, performance in biparental loci and indels was lower than in positions in which only one parent is heterozygous (US2021 / 0340601).
[0015] GENERAL DESCRIPTION
[0016] In a first of its aspects, the present invention provides a method for genotyping a fetus, comprising: i. receiving reads of sequencing data of (i) maternal plasma cell-free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting the fetus; ii. determining the distribution of a segment feature for each of the reads of sequencing data of the cfDNA; and iii. identifying potential genomic sites at which the fetus may have a copy number variant (CNV) by comparing the segment feature distribution of the cfDNA with a segment feature distribution reference set; wherein a deviation from the segment feature distribution of the reference set indicates the presence of a fetal CNV; thereby genotyping said fetus.
[0017] In one embodiment, the segment feature is segment length, hence in accordance with this embodiment, the present invention provides a method for genotyping a fetus, comprising: i. receiving reads of sequencing data of (i) maternal plasma cell-free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting the fetus; ii. determining the read length distribution for the reads of sequencing data of the cfDNA; and iii. identifying potential genomic sites at which the fetus may have a copy number variant (CNV) by comparing the read length distribution of the cfDNA with a read length distribution reference set; wherein a deviation from the read length distribution of the reference set indicates the presence of a fetal CNV; thereby genotyping said fetus.
[0018] In one embodiment, said reads of sequencing data are short reads.
[0019] In one embodiment, said reads of sequencing data are long reads.
[0020] In one embodiment, said reads of sequencing data comprise short reads and long reads.
[0021] In one embodiment, one or both cfDNA sequencing data and the gDNA sequencing data is obtained by a method selected from a group consisting of whole genome sequencing (WGS), whole exome sequencing (WES), next generation sequencing (NGS), targeted sequencing, panel sequencing, gene sequencing, long-read genome sequencing, paired-end sequencing, single end sequencing, and amplicon sequencing.
[0022] In one embodiment, WGS or WES data is obtained by deep sequencing.
[0023] In one embodiment, the read length distribution for the cfDNA sequence reads is determined by: i. Mapping the cfDNA sequence reads into bins, wherein the bins are obtained by dividing each chromosome into non-overlapping segments; and ii. Measuring the length of each read mapped into each bin, thereby determining read length distribution in each bin and in the entire chromosome.
[0024] In one embodiment, the read length distribution reference set is generated by: i. receiving reads of sequencing data of maternal plasma cell -free DNA (cfDNA) from a plurality of women carrying a fetus; ii. Mapping the cfDNA sequence reads into bins, wherein the bins are obtained by dividing each chromosome into non-overlapping segments; and iii. Measuring the length of each read mapped into each bin, thereby generating a read length distribution reference set. In one embodiment, for each bin, cumulative difference values are calculated for the plurality of women carrying a fetus and the values are converted into a Z-score representing the length distribution reference.
[0025] In one embodiment, the Z-score is transformed into a rolling average.
[0026] In one embodiment, the window size for the moving average is between about 5 bins and 100 bins, e.g., 20 bins.
[0027] In one embodiment, the size of the bin is between about 1,000 bases and 50,000 bases, e.g., 5,000 bases.
[0028] In one embodiment, peak points that fall within the top l%to 5%, e.g., 1%, of values indicate the center of a potential CNV.
[0029] In one embodiment, the method further comprises a step of normalized coverage calculation.
[0030] In one embodiment, the method further comprises a step of fdtering out false positive CNV determinations.
[0031] In one embodiment, after said step (i) in claim 1 the reads of sequencing data are aligned to a reference genome, thereby producing a BAM file.
[0032] In one embodiment, the method further comprises a step of determining the probability that the fetus has a CNV using fetal variant calling.
[0033] In one embodiment, the method further comprises a step of determining the probability that the fetus has a CNV using the Hoobari algorithm.
[0034] In one embodiment, the method further comprises a step of determining the probability that the fetus has a CNV using haplotype-based analysis.
[0035] In one embodiment, the method further comprises a step of determining the probability that the fetus has a CNV using fragmentomics-based analysis.
[0036] In one embodiment, said copy number variant is a deletion variant or a duplication variant.
[0037] In another aspect, the present invention provides a computer software product, comprising a computer-readable medium in which program instructions are stored, which instructions, when read by a data processor, configure the data processor to (1) receive reads of sequencing data of (i) maternal cell-free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting a fetus, and to (2) execute the method according to the present invention. In another aspect, the present invention provides a system for genotyping a fetus, comprising: an input utility for receiving reads of sequencing data of (i) maternal cell-free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting a fetus; and a data processor configured for analyzing said data for executing the method according to the present invention.
[0038] BRIEF DESCRIPTION OF THE DRAWINGS
[0039] For better understanding the subject matter that is disclosed herein and to exemplify how it may be carried out in practice, embodiments will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which:
[0040] Figure 1 is a flowchart diagram of a method suitable for fetal genotyping, according to various exemplary embodiments of the present invention.
[0041] Figure 2A is a graph showing segment length distribution (density) inside a fetal only deletion and in the entire chromosome, dashed lines are the limits of each peak.
[0042] Figure 2B is a graph showing a q-q plot of the segment length distribution inside the deletion as compared with the entire chromosome.
[0043] Figure 2C is a graph showing the Z score of the sum of difference in quantiles between the bins inside the deletion and the entire chromosome. Each line is the value for a single plasma sample. The light-colored line represents the sample harboring the deletion.
[0044] Figure 3A is a graph showing a rolling average of the Z score data shown in Figure 3B, dark-colored dots mark the distribution peaks.
[0045] Figure 3B is a graph showing the length Z-score as a function of the chromosomal location in chromosome 4, in a sample that has a Denovo deletion in the location indicated by the dashed line.
[0046] Figure 4A is a graph showing the rolling average (mean) of the coverage ratio in a representative deletion as a function of the chromosomal location. The location of the length score peak is represented by a dashed line marked by an arrow. The deletion borders are defined as the data points on both sides that exceed 0 for the first time (marked by the dashed lines).
[0047] Figure 4B is a graph showing the rolling average (mean) of the length Z-score for the same deletion shown in Fig. 4A. The arrow points to the local maximum length and the limits are defined as the first data points on both sides of the deletion that are above 0 (marked by the dashed lines).
[0048] DETAILED DESCRIPTION OF EMBODIMENTS
[0049] The present invention is based on the finding that analysis of the distribution of various parameters typical to cfDNA segments, such as their average length, can be used to identify copy number variations (CNVs) in the fetal genome.
[0050] The present disclosure leverages the distribution bias typical to various features between cfDNA segments of maternal origin and cfDNA segments of fetal origin, e.g., the difference in the length distribution between fetal and maternal cfDNA (Rabinowitz et al., 2019), different methylation patterns, and different enriched end motifs (Chiu and L.O, 2021). Consequently, in regions with a deletion that exists only in the fetus, a relative increase in the average length of DNA segments is expected, compared to the rest of the genome. Conversely, areas with a de novo or paternal duplication would exhibit a decreased average segment length.
[0051] The CNV may be de-novo variations or inherited variations, i.e., both paternal and maternal variations. The method and system of the present disclosure employ deep whole genome sequencing (WGS) of cfDNA of the maternal plasma during pregnancy combined with WGS of the maternal blood cells.
[0052] The present disclosure thus provides a method for genome-wide noninvasive prenatal genotyping that comprises algorithms that use a segment feature, e.g., a segment length variance, a methylation pattern, or enriched end motifs, to detect potential fetal CNVs. The potential fetal CNVs then undergo a filtering process, which incorporates various parameters such as coverage depth and their relation to complex genomic regions, to confirm them as true CNVs. As shown in the examples below, in a study involving 10 families, the method of the present disclosure effectively identified deletions exceeding 100 kilobases, resulting in two verified positive cases without any false positives. Detailed analyses, including the examination of amniotic fluid and paternal DNA, confirmed the authenticity of these findings, highlighting the absence of additional de novo or paternal deletions in the embryos' sequencing data. This outcome demonstrates the precision and reliability of the method of the present disclosure.
[0053] To identify regions of variance, e.g., where the read length distribution, the methylation pattern, or the enriched end motifs, differ from the expected one, each chromosome is segmented into discrete, non -overlapping bins. For each bin, the segment features of individual segments mapped to the bin, are recorded, and compared to the distribution of the segment feature of all segments mapped to the chromosome. The method of the invention was exemplified using segment length distribution. As shown in the Examples below, comparison of segment lengths within a fetal only deletion versus the entire chromosome revealed that segments in the deletion area are notably longer than those in the rest of the chromosome.
[0054] Accordingly, in an aspect, the present invention provides a method for genotyping a fetus, comprising: i. receiving reads of sequencing data of (i) maternal plasma cell-free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting the fetus; ii. determining the distribution of a segment feature for each of the reads of sequencing data of the cfDNA; iii. identifying potential genomic sites at which the fetus may have a copy number variant (CNV) by comparing the segment feature distribution of the cfDNA with a segment feature distribution reference set; wherein a deviation from the segment feature distribution of the reference set indicates the presence of a fetal CNV; thereby genotyping said fetus.
[0055] A flowchart describing this process is illustrated in Figure 1.
[0056] As used herein the term “segment feature” encompasses any feature which may be differentially distributed between maternal cfDNA segments and fetal cfDNA fragments. For example, but not limited to the extent of methylation in the region, the level of enrichment for specific end motifs and the segment length.
[0057] Accordingly, in an aspect, the present invention provides a method for genotyping a fetus, comprising: i. receiving reads of sequencing data of (i) maternal plasma cell-free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting the fetus; ii. determining the read length distribution for the reads of sequencing data of the cfDNA; iii. identifying potential genomic sites at which the fetus may have a copy number variant (CNV) by comparing the read length distribution of the cfDNA with a read length distribution reference set; wherein a deviation from the read length distribution of the reference set indicates the presence of a fetal CNV; thereby genotyping said fetus.
[0058] In one embodiment, the method of the present disclosure comprises the following steps:
[0059] 1. Sample sequencing - Deep whole genome sequencing is performed on cfDNA from the maternal plasma of pregnant women. This is complemented by whole genome sequencing of the maternal and optionally paternal blood cells.
[0060] 2. Alignment- The sequenced data is aligned to a reference genome, resulting in the generation of BAM files.
[0061] 3. Normalized coverage calculation- a suitable software tool (algorithm) for normalizing the data, together with a reference panel, is applied to the BAM files, lliis algorithm outputs a bin-wise coverage file that normalizes the coverage of each genomic bin to account for GC content and other sample-specific effects. This normalized coverage is important for filtering potential deletions later in the process.
[0062] 4. Creation of a segment feature Reference: a. Each chromosome is divided into non-overlapping bins, bins that are within genomic sites that are not sufficiently mapped or in centromeres are removed. b. The segment feature distribution is calculated for each bin and compared to the entire chromosome, with quintiles being used for detailed comparison. c. The difference between the quintiles of each bin and the entire chromosome is calculated and summed. d. The sum of differences for each bin across reference panel of several families (for example, of at least 5 families) is then converted to a Z-score to quantify deviation from the norm.
[0063] 5. Identification of candidate CNVs: a. The segment feature distribution Z-score is smoothed using a moving average, and local extremes within the top or bottom 1% of data points are marked as potential CNV centers. This is repeated with varying window sizes to detect CNVs of different lengths. b. CNV boundaries are determined based on the moving average of bin-wise coverage obtained, with boundaries defined by the points where the moving average changes sign around a length score peak. ration of candidate CNVs: a. Candidate CNVs are filtered based on a variety of features to distinguish true CNVs from artifacts, maternal inheritances, or statistical noise. These features include but are not limited to: i. Mean Coverage (Mean_cov): The average Z-score for coverage within the CNV region. The Z-score is derived from the moving average of bin-wise coverage and indicates the deviation in coverage from other genomic regions and samples, with adjustments made for GC content. A lower value indicates a deletion, and a higher value indicates a duplication. ii. Coverage difference from other families (Cov diff) Measures the extent to which the coverage within a CNV region in the sample deviates from the coverage observed in the same region across other samples. The Cov_diff value is then determined by finding the closest Z-score value among other samples to the mean coverage Z-score (Mean_cov) of the sample with the potential deletion. iii. Coverage P-value (p_val_cov): A statistical metric used to evaluate the rarity and significance of the coverage variation observed within a CNV region, relative to other regions of the same chromosome and sample. It specifically assesses how often one encounters regions of identical size displaying the same or lower mean coverage (Mean_cov) levels as those found within the CNV area. The calculation is predicated on the hypothesis that de novo and large paternal CNVs are rare events. For duplication the feature value is 1- p_val_cov. iv. Mean segment feature (mean feature): The average segment feature Z-score within the CNVs. This value reflects how extreme is the feature distribution inside the CNV compared to the same region in other samples and other regions in the sample. v. Segment feature significance (p feature): A statistical metric used to evaluate the rarity and significance of the segment feature variation observed within a CNV region, relative to other regions of the same chromosome and sample. It specifically assesses how often one encounters regions of identical size displaying the same or higher mean feature (mean_feature) levels as those found within the CNV area. vi. Post-Deletion Mean feature (mean after): The average length Z-score calculated for a segment that begins at the CNVs endpoint and extends to half of the CNV size. This feature is used to test if the segment feature difference is specific for the CNV region or a property of the genomic region. vii. Pre-Deletion Mean feature (mean before): This represents the average feature Z-score for a segment starting from a position. The average feature Z-score calculated for a segment that begins at the CNVs endpoint and extends to half of the CNV size. This feature is used to test if the segment feature difference is specific for the CNV region or a property of the genomic region. viii. Mean Coverage in Maternal Sample (mean cov M): The average Z-score of coverage within the CNVs in the maternal sample. This feature is used to validate that the coverage difference is indeed specific to the plasma and not due to an inherited variant from the mother. ix. Surrounding Region Size in Maternal Sample (len_M): Indicates the extent of the region around the lowest Z-score point inside the deletion in the maternal sample, where all values are below 0. For duplication this is the region around the highest Z- score where all values are above zero. If len_M is similar or larger than the CNV size it implies the CNV is an artifact or inherited variant from the mother. x. Frequency of 'NA' Values in Maternal CNV (freq na M): The occurrence of NA values within the CNVs region in the maternal sample. NAs occur when the region has poor mappability. High proportion of NAs implies the CNV is an artifact. xi. Overlap with problematic genomic regions (prob regions): The proportion of the overlap with complex genomic regions like blacklist regions and regions that have alternate haplotype. A significant overlap implies the CNV is an artifact. xii. Maximum Overlap with other CNV detection algorithms in maternal Sample (per M overlap): The maximal overlap between the potential CNV and CNVs detected by another CNV detection algorithm in the maternal sample. A significant overlap implies the CNV is inherited from the mother or is an artifact. xiii. Coverage-Length Overlap (overlap_ feature _cov): Describes the overlap between the CNV boundaries as determined by coverage and those determined by segment length. CNV limits for the feature are determined in a similar way to limits by coverage as the points on either side of the length peak where the length z- score reverse sign (for example as shown in Figure 4B). b. CNVs that were not identified by other algorithms are filtered with a stringent filter set that comprises one or more of the following features: p_ feature, mean_cov_M _cov, freq_na_M, prob_regions, per_M_overlap, mean_after, mean_before, len_M, overlap_ feature _cov, mean_cov. c. CNVs that are identified by other algorithms are filtered with a lenient filter set that comprises one or more of the following features: p_ feature, mean_cov_M, freq_na_M, prob_regions, per_M_overlap, overlap_ feature _cov.
[0064] In an embodiment, the method further comprises performing CNV prediction using additional CNV detection algorithms, e.g., Pindel (Ye et al., 2009) or Lumpy (Layer et al., 2014). The algorithms are run on both plasma and maternal BAM files to identify potential CNVs. The CNV prediction step using additional CNV detection algorithms is optional and may be used to increase the algorithm power and accuracy but is not mandatory.
[0065] "Blood sample” herein refers to a whole blood sample that has not been fractionated or separated into its component parts as well as to a fractionated blood sample. Whole blood is often combined with an anticoagulant such as EDTA or ACD during the collection process but is generally otherwise unprocessed.
[0066] “Blood fractionation” is the process of fractionating whole blood or separating it into its component parts. This is typically done by centrifuging the blood. The resulting components are: (a) a clear solution of blood plasma in the upper phase (which can be separated into its own fractions), (b) a buffy coat, which is a thin layer of leukocytes (white blood cells) mixed with platelets in the middle, and (c) erythrocytes (red blood cells) at the bottom of the centrifuge tube in the hematocrit fraction.
[0067] “Blood plasma” or “plasma” is the liquid component of blood (the blood component excluding cells). It makes up about 55% of total blood by volume. It is mostly water (93% by volume), and contains dissolved proteins including albumins, immunoglobulins, and fibrinogen, as well as glucose, clotting factors, electrolytes, hormones, carbon dioxide, and cell free DNA.
[0068] Blood plasma can be prepared by centrifuging a tube of whole blood in the presence of an anti-coagulant until the blood cells are separated and pulled down to the bottom of the tube. The blood plasma is then poured or drawn off.
[0069] “Cell-free DNA” (cfDNA) also referred to as “circulating free DNA” are DNA fragments existing outside of cells in vivo circulating in body fluids such as blood plasma. The fragments of cfDNA typically have lengths ranging from about 150 to 200 base pairs (bp), and averaging about 170 bp, which presumably relates to the length of a DNA stretch wrapped around a nucleosome During pregnancy, cell-free fetal DNA can be found circulating in maternal plasma. Thus, the cfDNA in maternal plasma is a mixture of both maternal and fetal DNA; both the total amount of cfDNA, and the fraction of fetal DNA within it. increases throughout pregnancy.
[0070] The term cfDNA also refers to fragments of DNA that have been obtained from the in vivo extracellular sources and separated, isolated, or otherwise manipulated in vitro. cfDNA can be obtained by extracting DNA from blood plasma after removal of intact cells. Methods for extracting cfDNA are well known in the art, for example, as shown in the Examples below'. In addition to the sequencing of cfDNA, paternal and maternal genomic DNA data may also be obtained by sequencing DNA derived from a cell-containing sample using a WGS approach, to assign prior probabilities to plasma sequencing reads as to their origins (fetal / matemal).
[0071] The term “genomic DNA ” or “gDNA ” herein refers to DNA existing in a cell in vivo and containing a complete genome of the cell or organism. The term also refers to DNA that has been obtained from the in vivo cell and separated, isolated, or otherwise manipulated in vitro. Typically, the cell is isolated prior to being subjected to lysis to produce in vitro cellular DNA. The term gDNA as used herein does not include cfDNA.
[0072] The term “sample” herein refers to a sample typically derived from a biological fluid, cell, tissue, organ, or organism comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence that is to be analyzed for the presence of a genetic variant. Such samples include but are not limited to blood or a blood fraction (for example, peripheral blood mononuclear cells) obtained from whole blood samples.
[0073] The sample is preferably obtained from a human subject.
[0074] The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. The sample may be used fresh or thawed after being frozen.
[0075] The DNA is extracted using standard protocols, e.g., as described in the examples below.
[0076] The term “genomic site” refers to a unique position in the genome. It may be represented by a chromosome ID, chromosome position and orientation on a reference genome.
[0077] The maternal genomic DNA (gDNA) data, maternal cell-free DNA (cfDNA) data, and optionally, the paternal gDNA data are obtained by a sequencing method including, but not limited to, deep whole genome sequencing (WGS), whole exome sequencing (WES), next generation sequencing (NGS), targeted sequencing, panel sequencing, gene sequencing, long-read genome sequencing, paired-end sequencing, single end sequencing, and amplicon sequencing.
[0078] The term “Next Generation Sequencing” (NGS) herein refers to sequencing methods that allow for massively parallel sequencing of clonally amplified molecules and of single nucleic acid molecules. Non-limiting examples of NGS include sequencing -bysynthesis using reversible dye terminators, and sequencing-by-ligation. Deep sequencing refers to sequencing a genomic region multiple times, sometimes hundreds or even thousands of times. Deep sequencing of the genome allows researchers to detect rare genetic variants.
[0079] As used herein the term “deep whole genome sequencing” refers to deep sequencing of the entire genome. In the context of the present invention cell-free DNA extracted from maternal blood plasma during pregnancy is subjected to deep whole genome sequencing. The maternal blood plasma samples may be obtained at any stage of the pregnancy, preferably between weeks 7-38 of the pregnancy.
[0080] The sequencing is repeated multiple times, for example, but not limited to between 10 times (10X) and 1000 times (1000X), e.g., 10 times ( 10X), 20 times (20X), 30 times (30X), 50 times (50X), 100 times (100X), 150 times (X150), 200 times (200X), 300 times (300X), 500 times (500X), or 1000 times (1000X).
[0081] In one non-limiting example, the cfDNA in maternal plasma is sequenced 300 times (300X).
[0082] In addition, genomic maternal and, optionally, paternal DNA is also subjected to whole genome sequencing. Such genomic DNA may be obtained from any cell type, for example from blood cells, e.g., leukocytes. In an embodiment, whole genome sequencing of paternal and maternal genomic DNA is performed to a targeted depth of between about 20X and 40X, for example 30X.
[0083] Whole genome sequencing may be performed using any method known in the art, for example, the HiSeq X Ten System (Illumina) or HiSeq 4000 (Illumina).
[0084] The sequencing generates “reads” which are sequences of DNA fragments of varying lengths. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A T C G). It may be stored in a memory device and processed as appropriate. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information.
[0085] The sequencing input may be of long reads (e.g., from about 1 KBP (kilogram base pairs) to about 100KBP, or more) or short reads (e.g., from between about 50 base pairs and 400 base pairs), or a combination of long reads and short reads.
[0086] After sequencing, the reads are aligned to a human reference genome based on sequence similarities. As used herein, the terms “aligned”, “alignment”, or “aligning” refer to the process of comparing a read to a reference sequence and thereby determining whether the read is contained in the reference sequence. If the reference sequence contains the read, the read may be mapped to a particular location in the reference sequence. In some cases, alignment simply tells whether the read is present or absent in the reference sequence.
[0087] Optionally, additional information (also referred to herein as “metadata”) pertaining to one or both the parents is also received. The received metadata optionally and preferably includes at least one, more preferably more than one, of the following features: mutation carrier status of the parents, ethnicity of the parents, body mass index (BMI), and week of pregnancy.
[0088] The method of the present disclosure is suitable for the detection of copy number variation (CNV).
[0089] The term “copy number variation ” (CNV) herein refers to any structural genome variant in which the amount of a certain genomic sequence is altered either increased or decreased. As such CNV encompasses deletions, insertions, duplications, multiplications, and translocations. CNV also encompasses chromosomal aneuploidies and partial aneuploidies.
[0090] CNV may cause various pathologies such as developmental disorders, e.g., neurodevelopmental disorders (see for example Shaikh, 2017).
[0091] Accordingly, in an embodiment, and only if the genotyping identifies that the fetus possesses a copy number variant associated with a disease or disorder:
[0092] (i) Performing an invasive prenatal test (e.g., by amniocentesis);
[0093] (ii) administering a prenatal or a post-natal treatment for said disease or disorder in an amount effective to prevent or treat said disease or disorder, wherein said treatment comprises pharmaceutical based intervention, surgery, genetic therapy, nutritional therapy, or combinations thereof; or
[0094] (iii) performing a pregnancy termination.
[0095] The term “aneuploidy” herein refers to an imbalance of genetic material caused by a loss or gain of a whole chromosome, or part of a chromosome. The identification of the maternal and paternal variants (i.e., variant sites or mutations) can be performed using a variant calling approach, which is generally based on alignment of the DNA sequencing data and the application of a variant caller.
[0096] Sequence alignment techniques that can be used according to some embodiments of the present invention include, without limitation, Burrows Wheeler Aligner (BWA), ABA, ALE, AMAP, anon, BAli-Phy, Base-By-Base, BHAOS / DI ALIGN, Bowtie, Bowtie 2, ClustalW, CodonCode Aligner, Comass, DECIPHER, DIALIGN-TX, DIALIGN-T, DNA Alignment, DNA Baser Sequence Assembler, EDNA, FSA, Geneious, Kalign, MAFFT, MARNA, MAVID, MSA, MSAProbs, MULTALIN, Multi- LAGEN, MUSCLE, Opal, Pecan, Phylo, Praline, PicXAA, POA, Probalign, ProbCons, PR0MALS3D, PRRN / PRRP, PSAlign, RevTrans, SAGA, SAM, Se-AI, STAR, STAR- Fusion, StatAlign, Stemloc, T-Coffee, UGENE, VectorFriends, and GLProbs.
[0097] Exemplary variant callers suitable for the present embodiments include, without limitation, Genome Analysis Toolkit (GATK) and Freebayes. For example, Freebayes can comprise an alignment based on literal sequences of reads aligned to a particular target, not their precise alignment. GATK can comprise: (i) pre-processing; (ii) variant discovery; and (iii) callset refinement. Pre-processing can comprise starting from raw sequence data, e.g., in FASTQ or uBAM format, and producing analysis-ready BAM files; processing can include alignment to a reference genome as well as data cleanup operations to correct for technical biases and make the data suitable for analysis; variant discovery can comprise starting from analysis-ready BAM files and producing a callset in VCF format; processing can involve identifying sites where one or more individuals display possible genomic variation, and applying filtering methods appropriate to the experimental design; callset refinement can comprise starting and ending with a VCF callset; processing can involve using metadata to assess and improve genotyping accuracy, attach additional information and evaluate the overall quality of the callset.
[0098] Also contemplated are variant callers such as, but not limited to, Platypus, VarScan, Bowtie analysis, MuTect and / or SAMtools. For example, Bowtie analysis can comprise implementing the Burrows-Wheeler transform for aligning. MuTect can comprise: (i) pre-processing; (ii) statistical analysis; and (iii) post-processing. Preprocessing can comprise an initial alignment of sequencing reads; statistical analysis can comprise using two Bayesian classifiers, one classifier can detect whether a SNP is nonreference at a given site and, for those sites that are found as non-reference, the other classifier can make sure that the normal does not carry the SNP; post-processing can comprise removal of artifacts of sequencing, short read alignments and hybrid capture. SAMtools can comprise storing, manipulating, and aligning sequencing reads stored as SAM files.
[0099] In various exemplary embodiments of the invention the method further comprises the determination of the probability, for each variant site, to be of fetal origin, also referred to as site level prediction.
[0100] In an embodiment, the determination of the probability of the variant to be of fetal origin comprises constructing a fetal size distribution and a maternal size distribution, binning said fetal size distribution and calculating a fetal fraction for each fragment size bin, and calculating, for at least one size and at least one fragment at said at least one site, a probability that said fragment is fetal, based on a fetal fraction of a respective fragment size bin to which said fragment belongs.
[0101] As used herein the term “fetal fraction ” or “FF” refers to the portion of fetal cfDNA, within the total amount of cfDNA in the maternal blood. The portion of fetal cfDNA within maternal blood (the fetal fraction) varies throughout the pregnancy, and between individuals, hence this is not regarded as a constant but as a variable. Low levels of fetal cfDNA are referred to as a low fetal fraction.
[0102] In an embodiment, said determining the probabilities comprises applying a Bayesian procedure. Optionally, said Bayesian procedure comprises prior probabilities calculated using sequencing data of at least one of said parents.
[0103] In an embodiment, this procedure further comprises recalibration of the output of said Bayesian procedure using machine learning.
[0104] In a specific embodiment the determination of the probability, for each variant site, to be of fetal origin is performed using variant calling, for example using the Hoobari algorithm as described in Rabinowitz et al., 2019 and WO2021 / 0340601 .
[0105] The method of the invention can also be combined with analysis of fragmentomic features of DNA reads.
[0106] The term “fragmentomic features” refers to molecular characteristics of DNA reads, as well as to genomic, epigenetic and alignment features of the DNA read. Fragmentomic features include, but are not limited to, Read quality mapping, Read base qualities, Fragment length, short / long read ratio, DNA fragment end motifs, Cleavage paterns around methylation sites, Read endpoint preferred end, DNA accessibility / nucleosome positioning inference, Distance to nearest nucleosome, Transcription factor binding sites, Regional fetal fraction, Regional sequence composition, Read sequence composition, and Number of sequence errors in the read.
[0107] In certain embodiments, the variant calling probabilities and the fragmentomics- based probabilities are multiplied to calculate joint probabilities, for which a maximumlikelihood approach is applied to predict the genotype at each site.
[0108] As used herein the term “heterozygous” refers to different versions (alleles) of a genomic locus. The term “homozygous” refers to the presence of the same versions (alleles) of the genomic locus.
[0109] The term “locus” is used to refer to the specific location of a nucleic acid sequence or variant on a reference chromosome.
[0110] The term ’’about” as used herein indicates values that may deviate up to 1%, more specifically 5%, more specifically 10%, more specifically 15%, and in some cases up to 20% higher or lower than the value referred to, the deviation range including integer values, and, if applicable, non-integer values as well, constituting a continuous range.
[0111] It must be noted that, as used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise.
[0112] Throughout this specification and the Examples and claims which follow, unless the context requires otherwise, the word “comprise”, and variations such as “comprises” and “comprising” . will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.
[0113] Disclosed and described, it is to be understood that this invention is not limited to the specific examples, methods’ steps, and compositions disclosed herein as such methods’ steps and compositions may vary somewhat. It is also to be understood that the terminology applied herein is used for the purpose of describing specific embodiments only and not intended to be limiting since the scope of the present invention will be limited only by the appended claims and equivalents thereof. EXAMPLES
[0114] Materials and Methods
[0115] Sample collection and DNA extraction.
[0116] Samples from each family were collected during week 11-33 of the pregnancy with informed consent. DNA from chorionic villus sampling (CVS) or amniocentesis specimen was extracted using the magLEAD 12gC, MagDEA Dx kit (ExScale, Chiba, Japan). Peripheral maternal blood was collected using 2-4 Ethylene-diamine-tetra-acetic acid (EDTA) tubes. Within one hour of collection, plasma was separated from blood by centrifugation at room temperature for 10 minutes at 1600 x g. The plasma was then centrifuged again at 16,000 x g for 10 minutes at room temperature to remove any residual cells. Extraction of cfDNA was performed using the QIAamp Circulating Nucleic Acid Kit (Qiagen). Removal of excess salts resulting from cfDNA purification was conducted using Agencourt AMPure XP beads (Beckman Coulter, Inc.) at a 2X ratio to cfDNA volume. Parental (maternal and paternal) genomic DNA was extracted from peripheral blood mononuclear cells (PBMC) using a standard protocol that includes (i) buffy coat separation and (ii) DNA purification using the magLEAD 12gC, MagDEA Dx kit (ExScale, Chiba, Japan) according to the manufacturer's instructions.
[0117] Library preparation and sequencing
[0118] Library preparation was performed using the TruSeq DNA PCR-Free Library Prep Kit (Illumina), for genomic DNA, and the Accel-NGS 2S PCR-free Library Prep Kit for cfDNA samples, according to the manufacturer's instructions. This was followed by sequencing using the NovaSeq platform (Illumina) targeting 150-bp paired-end reads across every DNA sample from each family unit.
[0119] Both parental and fetal genomic DNA libraries were sequenced aiming for a conventional depth of 30 x .
[0120] Cell-free DNA samples were not fragmented during library preparation and were sequenced to a requested coverage of 300x, using the NovaSeq platform (Illumina) with 150-bp paired-end reads. Alignment to the genome
[0121] Reads were aligned to the Genome Reference Consortium Human Build 38 (GRCh38 / hg38) using Burrows-Wheeler vO.7.834 with default parameters. Duplicate reads, resulting from PCR clonality or optical duplicates, and reads mapping to multiple locations were excluded from downstream analysis.
[0122] Example 1: Analysis of segment lengths distribution
[0123] To identify regions where the read length distribution differs from the expected one, each chromosome is segmented into discrete, non-overlapping bins. For each bin, the lengths of individual segments mapped to the bin are recorded and compared to the length distribution of all segments mapped to the chromosome. As illustrated in Figure 2A, the comparison of segment lengths within a fetal only deletion versus the entire chromosome reveals that segments in the deletion area are notably longer than those in the rest of the chromosome. To quantify the disparity in length distributions between each bin and the whole chromosome, the cumulative difference was calculated across the 100 quantiles of both distributions, as depicted in the q-q plot shown in Figure 2B. A q-q plot is a plot of the quantiles of one data set (i.e., the deletion quantiles) against the quantiles of the second data set (i.e., the entire chromosome quantiles). A quantile refers to the fraction, or percentage of points below a given value. This approach provides a single metric that encapsulates the variance in segment lengths. Subsequently, to ensure that the observed differences in segment length were not merely artifacts of region-specific characteristics, a comparative analysis is employed across multiple samples. Specifically, for each bin the cumulative quantiles difference is calculated in multiple families and the values are converted into a Z-score. This method effectively normalizes the length distribution within a specific bin against all bins in the same chromosome and the same bin in other samples. Figure 2C highlights this analysis for three bins situated within a deletion that exists in a single sample. Notably, the Z-score for the sample that harbors the deletion consistently ranks as the highest or second highest, indicating significant deviation in segment length distribution within these bins.
[0124] Example 2: principles of CNV detection
[0125] Next, following the establishment of the length distribution reference, CNV detection was performed. In this phase, regions across each chromosome that show significant deviations in their length distributions were identified for each cfDNA sample. Given the highly variable nature of the length Z-score, as illustrated in Figure 3A, which complicates the straightforward identification of deletions, a strategy of transforming it into a rolling average was adopted, as depicted in Figure 3B. This approach smooths out the data, facilitating the recognition of significant patterns. Hie selection of the window size for the moving average is crucial, as it directly influences the minimum size of the detectable CNVs. For instance, choosing a window size of 20 bins, each encompassing 5 kb, sets the threshold for CNV detection at sizes exceeding 100 kb. Utilizing the smoothed data from the rolling average, peak points that fall within the top 1% of values are pinpointed, as shown in Figure 3B. Those peaks are then considered the center of a potential CNV.
[0126] To determine the CNV boundaries, sequencing coverage data is employed, utilizing the principle that a CNV present in the fetus but absent in the mother should manifest as a discernible difference in coverage, typically averaging half of the fetal fraction in the sample. However, coverage is susceptible to influences beyond mere CNV presence, including GC content and sample -specific biases. A software tool is applied to mitigate these factors. In the present Example WisecondorX (Raman et al., 2019) was applied, but any suitable algorithm that can perform this function may be used. As part of its analysis pipeline WisecondorX segments the chromosome into bins and for each bin, it calculates a value that reflects the coverage ratio in relation to all other bins, adjusted for regional and sample-specific variances. To reduce noise, the coverage data was transformed into a rolling mean. The CNV boundaries were then determined by identifying points adjacent to the length peak where the coverage's rolling mean changes sign, as illustrated in Figure 4A.
[0127] To refine the list of CNV candidates, several additional features were incorporated into the analysis framework. These features serve to filter out candidates that may represent false positives, such as artifacts, maternal inheritances, or anomalies attributable to statistical noise rather than genuine biological variance. Key features for flagging potential artifacts include the degree of CNV overlap with known problematic genomic regions, the prevalence of NAs within the CNV's coverage data, and the anomaly scores of length measurements in adjacent regions. To identify CNVs potentially inherited from the mother, the coverage levels within the maternal sample and the degree of overlap between the candidate CNV and any detected within the maternal genome were considered. Discerning authentic CNVs from statistical aberrations involves evaluating the significance and magnitude of variations in both coverage and length distributions, alongside the alignment of CNV boundaries as determined by coverage data and those determined by length metrics. The determination of CNV boundaries based on length employs a methodology akin to that used for coverage, pinpointing the transition points adjacent to the peak where the length z-score changes direction, as illustrated in Figure 4B. To bolster the reliability of this approach, two tailored filtering protocols were devised: a general one applicable to CNVs uniquely detected by the algorithm of the present disclosure and a more lenient one for CNVs also identified by other analytical tools.
[0128] Example 3: CNV detection in a cohort of 10 families
[0129] To evaluate the effectiveness of the method of the present disclosure, a trial was conducted using a cohort consisting of ten families selected from a dataset. The focus was on identifying de novo or paternal deletions exceeding 100 kilobases. 1,076 potential deletion events were detected using the method of the disclosue. Among these, 25 deletions overlaped deletions identified by at least one other CNV detection algoritem.
[0130] Delitions that did not overlap with other algorithms were filtered with the folowing filters:
[0131] • p_ feature <0.05.
[0132] • mean_ feature > mean_after.
[0133] • mean_ feature > mean_before.
[0134] • mean_after>0.5.
[0135] • mean_before>0.5.
[0136] • mean_cov< mean_cov_M.
[0137] • freq_na_M<0.1.
[0138] • prob_regions<0.2.
[0139] • per_M_overlap<0.5.
[0140] • overlap_ feature _cov>0.7.
[0141] • cov_diff / mean_cov >0.5.
[0142] • mean_cov< -0.04. Delitions that overlapped with other algorithms were filtered using the folowing filters:
[0143] • p_ feature <0.05
[0144] • mean_cov< mean_cov_M
[0145] • freq_na_M<0.1
[0146] • prob_regions<0.2
[0147] • per_M_overlap<0.5
[0148] • overlap_ feature _cov>0.7.
[0149] The feature that was used as a filter in this example was segment length.
[0150] After applying the filtration criteria, the list was narrowed down to three promising candidate deletions. Among these, was a deletion approximately 880 kilobases in length located on chromosome 4. This deletion was not flagged by any other CNV detection tool. Subsequent confirmatory tests using genomic sequencing data from the embryo's amniotic fluid and the father's DNA established this deletion as a genuine de novo event, affirming the efficacy of the method of the present disclosure in uncovering novel deletions.
[0151] The remaining two candidate deletions had also been flagged by other CNV detection algorithms. Upon closer manual inspection, it was observed that for one of these deletions, the boundaries identified by the other algorithms did not align well with the areas showing reduced coverage, casting doubt on its validity. Conversely, the second deletion, a 164-kilobase deletion on the X chromosome, was verified as a legitimate paternal deletion, further demonstrating the utility of the method of the present disclosure in detecting meaningful genetic variations.
Claims
CLAIMS:
1. A method for genotyping a fetus, comprising: i. receiving reads of sequencing data of (i) maternal plasma cell -free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting the fetus; ii. determining the distribution of a segment feature for each of the reads of sequencing data of the cfDNA; and iii. identifying potential genomic sites at which the fetus may have a copy number variant (CNV) by comparing the segment feature distribution of the cfDNA with a segment feature distribution reference set; wherein a deviation from the segment feature distribution of the reference set indicates the presence of a fetal CNV; thereby genotyping said fetus.
2. A method for genotyping a fetus, comprising: i. receiving reads of sequencing data of (i) maternal plasma cell-free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting the fetus; ii. determining the read length distribution for the reads of sequencing data of the cfDNA; and iii. identifying potential genomic sites at which the fetus may have a copy number variant (CNV) by comparing the read length distribution of the cfDNA with a read length distribution reference set; wherein a deviation from the read length distribution of the reference set indicates the presence of a fetal CNV; thereby genotyping said fetus.
3. The method of any one of claims 1 or 2 wherein said reads of sequencing data are short reads.
4. The method of any one of claims 1 or 2 wherein said reads of sequencing data are long reads.
5. The method of any one of claims 1 or 2 wherein said reads of sequencing data comprise short reads and long reads.
6. The method of any one of claims 1 to 5, wherein one or both cfDNA sequencing data and the gDNA sequencing data is obtained by a method selected from a group consisting of whole genome sequencing (W GS), whole exome sequencing (WES), next generation sequencing (NGS), targeted sequencing, panel sequencing, gene sequencing,long-read genome sequencing, paired-end sequencing, single end sequencing, and amplicon sequencing.
7. The method of claim 6 wherein said WGS or WES data is obtained by deep sequencing.
8. The method of any one of claims 2 to 7, wherein the read length distribution for the cfDNA sequence reads is determined by: i. Mapping the cfDNA sequence reads into bins, wherein the bins are obtained by dividing each chromosome into non-overlapping segments; and ii. Measuring the length of each read mapped into each bin, thereby determining read length distribution in each bin and in the entire chromosome.
9. The method of any one of claims 2 to 8, wherein the read length distribution reference set is generated by: i. receiving reads of sequencing data of maternal plasma cell -free DNA (cfDNA) from a plurality of women carrying a fetus; ii. Mapping the cfDNA sequence reads into bins, wherein the bins are obtained by dividing each chromosome into non-overlapping segments; and iii. Measuring the length of each read mapped into each bin, thereby generating a read length distribution reference set.
10. The method of claim 9, wherein for each bin cumulative difference values are calculated for the plurality of women carrying a fetus and the values are converted into a Z-score representing the length distribution reference.
11. The method of claim 10, wherein the Z-score is transformed into a rolling average.
12. The method of claim 11 , wherein the window size for the moving average is between about 5 bins and 100 bins, e.g., 20 bins.
13. The method of any one of claims 8 to 12, wherein the size of the bin is between about 1,000 bases and 50,000 bases, e.g., 5,000 bases.
14. The method of any one of claims 11 to 13 wherein peak points that fall within the top 1% to 5%, e.g., 1%, of values indicate tire center of a potential CNV.
15. The method of any one of claims 2 to 14 further comprising a step of normalized coverage calculation.
16. The method of any one of claims 2 to 15 further comprising a step of filtering out false positive CNV determinations.
17. The method of any one of the preceding claims, wherein after said step (i) in claim 1 the reads of sequencing data are aligned to a reference genome, thereby producing a BAM file.
18. The method of any one of the preceding claims, further comprising a step of determining the probability that the fetus has a CNV using fetal variant calling.
19. The method of any one of the preceding claims, further comprising a step of determining the probability that the fetus has a CNV using the Hoobari algorithm.
20. The method of any one of the preceding claims further comprising a step of determining the probability that the fetus has a CNV using haplotype-based analysis.
21. The method of any one of the preceding claims further comprising a step of determining the probability that the fetus has a CNV using fragmentomics-based analysis.
22. The method of any one of the preceding claims wherein said copy number variant is a deletion variant or a duplication variant.
23. A computer software product, comprising a computer-readable medium in which program instructions are stored, which instructions, when read by a data processor, configure the data processor to (1) receive reads of sequencing data of (i) maternal cell- free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting a fetus, and to (2) execute the method according to any one of claims 1- 22.
24. A system for genotyping a fetus, comprising: an input utility for receiving reads of sequencing data of (i) maternal cell-free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting a fetus; and a data processor configured for analyzing said data for executing the method according to any one of claims
Citation Information
Patent Citations
Method and system for identifying gene disorder in maternal blood
US20210340601A1
Combined size-and count-based analysis of maternal plasma for detection of fetal subchromosomal aberrations
WO2016116033A1
Noninvasive fetal variant identification using haplotype analysis
WO2024105671A1