Methods for noise reduced and bias reduced genetic analysis
Combining different DNA sequencing methods with computational analysis to assess and correct noise and bias patterns addresses the limitations of existing genome-wide genotyping, improving accuracy and reducing costs for genetic analysis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-03-26
AI Technical Summary
Current genome-wide genotyping and mutation detection methods face challenges such as noise and bias, particularly in targeted sequencing approaches like WES and WGS, which affect the detection of structural variations and mutations in non-targeted regions, and are costly for applications like cell-free DNA sequencing in liquid biopsies and noninvasive prenatal testing.
A method combining different DNA sequencing techniques, such as WES and WGS, with computational analysis to assess and correct noise and bias patterns, enabling accurate genetic predictions by comparing noise and bias profiles across multiple sequencing data sets.
This approach improves the accuracy and reduces costs of genome-wide mutation detection by balancing the advantages of various sequencing technologies, enhancing sensitivity and specificity in genetic analysis, especially for noninvasive prenatal testing.
Smart Images

Figure IL2025050833_26032026_PF_FP_ABST
Abstract
Description
[0001] METHODS FOR NOISE REDUCED AND BIAS REDUCED GENETIC ANALYSIS
[0002] TECHNOLOGICAL FIELD
[0003] The present disclosure relates to the field of genetic analysis.
[0004] REFERENCES:
[0005] Goodwin S, et al., Coming of age: ten years of next-generation sequencing technologies. Nature Reviews Genetics 17, 333-351 (2016)
[0006] Kivioja T, et al., Counting absolute numbers of molecules using unique molecular identifiers. Nature Methods 9, 72-74 (2012)
[0007] Meynert A.M., et al., Variant detection sensitivity and biases in whole genome and exome sequencing. BMC Bioinformatics 15, Article no. 247 (2014)
[0008] Rabinowitz T, et al. Bayesian-based noninvasive prenatal diagnosis of single-gene disorders. Genome Res. 2019. https: / / doi.org / 10.1101 / gr.235796.118.
[0009] Rabinowitz T, Deri-Rozov S, Shomron N. et al., Improved noninvasive fetal variant calling using standardized benchmarking approaches. Computational and Structural Biotechnology Journal . 2021; 19:509-17.
[0010] Shestak A.G., et al., Allelic drop out is a common phenomenon that reduces the diagnostic yield of PCR-based sequencing of targeted gene panels. Front Genet. Volume 12 (2021)
[0011] Tsao D.S., et al., Nature Scientific Reports 9, Article no. 14382 (2019)
[0012] US2021 / 0340601.
[0013] BACKGROUND
[0014] Genome -wide approaches for genotyping and mutation detection rely on deoxyribonuclease (DNA) sequencing followed by computational analysis. DNA sequencing is typically performed using next-generation sequencing (NGS), an error prone technology which requires reading each genomic locus several times, i.e., with deeper coverage. Errors in DNA sequencing are typically caused by noise and bias (Goodwin et al, 2016); while noise can be mitigated with more data, i.e., deeper coverage, or using outlier filtration, bias persists despite these techniques.
[0015] NGS can be used to sequence the entire genome; this is called whole genome sequencing (WGS). Historically, factors such as sequencing cost, initial sample requirements, speed, and throughput have driven the development of targeted sequencing approaches, like whole-exome sequencing (WES) and other smaller NGS gene panels. These approaches focus on specific regions of the genome but have inherent limitations.
[0016] One limitation is bias and noise: Different genomic regions may be under or over- represented or even overlooked, leading to genotyping errors (Meynert et al, 2014). For example, false negative results in mutation detection can be due to an overlooked allele; this is termed "drop-out alleles" (Shestak et al, 2021). Dedicated methods to capture specific regions result in their own patterns of sequencing bias and noise. Another limitation is complexity, time and cost: Targeted sequencing requires specific kits, trained personnel, and meticulous primer and probe design. Low input protocols increase noise and bias, necessitating DNA amplification via PCR, which introduces additional bias and noise.
[0017] To address these issues, technological and computational solutions were devised, such as improved sequencing machines and unique molecular identifiers (UMIs). Computational methods include alignment algorithms, PCR-duplicate removal, basequality correction, and variant detection.
[0018] Despite advancements, targeted sequencing remains limited in detecting structural variations (SVs) and mutations in non-targeted regions. WGS covers the entire genome uniformly without the need for amplification, thus alleviating many limitations of targeted approaches. However, WGS has its own challenges, such as the high cost of deep sequencing required for certain applications like cell-free DNA (cfDNA) sequencing in liquid biopsies and noninvasive prenatal testing (NIPT).
[0019] NIPT relies on detecting slight imbalances in genetic material corresponding to specific genomic loci within maternal plasma cfDNA, correlated with fetal DNA amounts (fetal fraction). Larger chromosomal abnormalities provide sufficient genetic material for confident predictions, but smaller mutations are harder to detect, especially with low fetal fractions, as seen in early pregnancy stages or due to factors like BMI and ethnicity.
[0020] NGS technology, while powerful, is prone to sequencing errors that require deeper coverage, thus increasing costs. Various approaches have been employed to address these challenges, such as Digital PCR which provides accurate molecule counts for detecting imbalances but is not high throughput and is limited to a few mutations per test; WES and NGS Panels which allow ultra-deep sequencing at feasible costs but suffer from noise, bias, limited genomic coverage, and require careful design; Noise and Bias Reduction Methods which use unique molecular identifiers (UMIs) (Kivioja et al, 2012) to reduce noise or quantitative counting templates (QCTs™) to reduce bias (Tsao et al, 2019), though these methods are limited in throughput and not suitable for WGS; PCR-free WGS which eliminates PCR-related bias, enabling better detection of structural variations and overall superior performance compared to WES and NGS panels. However, the high cost of achieving deep sequencing coverage remains a limitation.
[0021] Rabinowitz et al (2019) describe a genome wide NIPT approach, termed noninvasive prenatal variant calling. Using Hoobari, the first noninvasive fetal vanant caller, they were able to genotype ail fetal positions, including biparental loci and indels (US2021 / 0340601). Hoobari. is atool which employs a Bayesian algorithm to predict the inheritance of monogenic diseases, irrespective of their mode of inheritance or parental origin (Rabinowitz et al. 2019; US2021 / 0340601; Rabinowitz et al. 2021 ) . This tool enables the estimation of the likelihood that each cfDNA fragment is of fetal origin by considering its length. Hoobari can detect mutations caused by single nucleotide polymorphisms (SNPs), or small indels - insertions and deletions of bases in the genome.
[0022] GENERAL DESCRIPTION
[0023] In a first of its aspects, the present invention provides a method of genetic analysis of a subject, comprising: a. Obtaining one or more samples comprising DNA; b. Sequencing the DNA in said one or more samples, wherein said sequencing is performed using at least a first and a second DNA sequencing method, thereby obtaining two or more sets of DNA sequencing data, for each of said DNA samples; wherein the first and the second sequencing methods are different sequencing methods; c. Constructing a profile of noise and / or bias and / or error for each of said two or more sets of DNA sequencing data, or for a subset thereof, wherein said profile comprises one or more features for each analyzed genomic locus; d. Assessing systemic noise and bias by comparing the noise and / or bias and / or error profiles of the two or more sets of DNA sequencing data (reads), or a subset of reads thereof, thereby obtaining one of more feature indicators describing the relationship between the two or more sets of DNA sequencing data; e. Performing a genetic analysis of the DNA in said one or more samples to provide genetic predictions; and f. Applying the noise and bias indicators obtained in (d) to correct the outcome of the genetic predictions obtained in (e).
[0024] In one embodiment, the genetic analysis comprises one or more of copy number prediction, genotyping, mutation prediction, DNA fragment classification and DNA fragment origin de-convolution.
[0025] In one embodiment, said DNA is cell-free DNA (cfDNA).
[0026] In one embodiment, said sample comprising DNA is a maternal plasma sample comprising cfDNA, and wherein said genetic analysis of a subject comprises genotyping a fetus.
[0027] In one embodiment, said sample comprising DNA is (i) a maternal plasma sample comprising cfDNA and (ii) maternal and / or paternal genomic DNA.
[0028] In one embodiment, said sample comprising DNA is a subject’s body fluid (e.g., plasma) sample comprising cfDNA, and wherein said genetic analysis comprises one or more of cancer detection, assessment of cancer progression, or assessment of response to therapy in the subject.
[0029] In one embodiment, said sample comprising DNA is (i) a subject’s body fluid (e.g., plasma) sample comprising cfDNA and (ii) genomic DNA of said subject.
[0030] In one embodiment, said genetic analysis comprises variant calling. In one embodiment, said DNA sequencing methods are selected from a group consisting of whole genome sequencing (WGS), whole exome sequencing (WES), next generation sequencing (NGS), targeted sequencing, panel sequencing, gene sequencing, long-read genome sequencing, paired-end sequencing, single end sequencing, and amplicon sequencing.
[0031] In one embodiment, said first and second DNA sequencing methods comprise a combination of NGS and WGS.
[0032] In one embodiment, said first and second DNA sequencing methods comprise WES and WGS.
[0033] In one embodiment, said assessment of systemic noise and bias and said genetic predictions are performed on reads of the same DNA sample or a different sample from the same individual.
[0034] In one embodiment, said assessment of systemic noise and bias is based on cumulative data obtained from one or more samples or individuals and said genetic predictions are performed on another DNA sample.
[0035] In one embodiment, the sequencing data obtained by a first of the sequencing methods is used for boosting or filtration of the sequencing data obtained by a second of the sequencing methods.
[0036] In one embodiment, said boosting or filtration is performed by uniting or intersecting of concordant or discordant genotyped genomic loci or sequenced DNA reads.
[0037] In one embodiment, the sequencing data obtained by a first of the sequencing methods is used for phasing of genomic loci within of the sequencing data obtained by a second of the sequencing methods.
[0038] In one embodiment, said assessing of noise and bias is performed by comparing the two or more sets of DNA sequencing data to obtain information on one or more features of a genomic locus.
[0039] In one embodiment, the information on one or more of the features is obtained using a method selected from a group consisting of descriptive statistics, manual filtration threshold definition, inferential statistics, regression analysis, clustering, graph-based approaches, Bayesian statistics, Markovian models, statistical learning models, machine learning models, and deep learning models.
[0040] In one embodiment, said reads of sequencing data are short reads.
[0041] In one embodiment, said reads of sequencing data are long reads. In one embodiment, said reads of sequencing data comprise short reads and long reads.
[0042] In one embodiment, providing genetic predictions is based on at least one Sequence Alignment Map (SAM) parameter.
[0043] In one embodiment, providing genetic predictions is based at least on an observed template length.
[0044] In one embodiment, providing genetic predictions, or determining the probability that a variant is of fetal origin comprises calculating a total fetal fraction.
[0045] In one embodiment, the method further comprises constructing a fetal size distribution and a maternal size distribution, wherein said providing genetic predictions of step e in claim 1 comprises binning said fetal size distribution and calculating a fetal fraction for each fragment size bin, and calculating, for at least one size and at least one fragment at said at least one site, a probability that said fragment is fetal, based on a fetal fraction of a respective fragment size bin to which said fragment belongs.
[0046] In one embodiment, providing genetic predictions comprises applying a Bayesian procedure.
[0047] In one embodiment, said Bayesian procedure comprises prior probabilities calculated using sequencing data of at least one of said parents.
[0048] In one embodiment, the method further comprises recalibration output of said Bayesian procedure using machine learning.
[0049] In one embodiment, said providing genetic predictions is performed using fetal variant calling.
[0050] In one embodiment, said providing genetic predictions is performed using the Hoobari algorithm.
[0051] In one embodiment, said providing genetic predictions comprises fragmentomics- based probability that the sequence read is of fetal origin.
[0052] In one embodiment, the method comprises extracting fragmentomic features for each cfDNA read identified as overlapping a potential genomic site where the fetus may have a variant.
[0053] In one embodiment, said fragmentomic features comprise one or more of read quality mapping, read base qualities, fragment length, short / long read ratio, end motifs, cleavage patterns around methylation sites, read endpoint preferred end, DNA / accessibility / nucleosome, distance to nearest nucleosome, transcription factor binding sites, regional fetal fraction, regional sequence composition, read sequence composition, and number of sequence errors in the read.
[0054] In one embodiment, said providing genetic predictions comprises multiplying the variant calling-based probability with the fragmentomics-based probability to obtain a calculated joint probability.
[0055] In another aspect, the invention provides a method for genotyping a fetus, comprising: a. Obtaining samples comprising DNA wherein the samples are of (i) maternal plasma cell-free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting the fetus; b. Sequencing the DNA in said samples, wherein said sequencing is performed using at least a first and a second DNA sequencing method, thereby obtaining two or more sets of DNA sequencing data, for each of said DNA samples; wherein the first and the second sequencing methods are different sequencing methods; c. Constructing a profile of noise and / or bias and / or error for each of said two or more sets of DNA sequencing data, or for a subset thereof, wherein said profile comprises one or more features for each analyzed genomic locus d. Assessing systemic noise and bias by comparing the noise and / or bias and / or error profiles of the two or more sets of DNA sequencing data or a subset of reads thereof, thereby obtaining one of more feature indicators describing the relationship between the two or more sets of DNA sequencing data for each of said samples; e. identifying potential genomic sites at which the fetus may have a genetic variant; and f. Applying the noise and bias feature indicators obtained in (d) to correct the identification of potential genomic sites at which the fetus may have a genetic variant; thereby genotyping said fetus. In another aspect, the invention provides a computer software product, comprising a computer-readable medium in which program instructions are stored, which instructions, when read by a data processor, configure the data processor to (1) receive reads of two or more sets of DNA sequencing data, wherein said sequencing is performed using at least a first and a second DNA sequencing method; wherein the first and the second sequencing methods are different sequencing methods, and to (2) execute the method according to the invention, as described above.
[0056] In another aspect, the invention provides a system for performing genetic analysis, comprising: an input utility for receiving reads of two or more sets of DNA sequencing data of the same DNA sample, wherein said sequencing is performed using at least a first and a second DNA sequencing method; wherein the first and the second sequencing methods are different sequencing methods; and a data processor configured for analyzing said data for executing the method according to the invention, as described above.
[0057] BRIEF DESCRIPTION OF THE DRAWINGS
[0058] For better understanding the subject matter that is disclosed herein and to exemplify how it may be carried out in practice, embodiments will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which:
[0059] Figure 1 is an outline exemplifying a method suitable for genomic analysis, according to various exemplary embodiments of the present invention.
[0060] Figure 2A-2D are graphs comparing values of sensitivity (SEN), specificity (SPE), and rate of filtered out results (NO_CALL) as tested in coding regions (i.e., exome regions) in four inheritance categories: 2A - Biparental, 2B - Compound, 2C - Maternal, 2D - Paternal, and type of sequencing run: WES100X, WGS300X, and WGS45X+WES100.
[0061] DETAILED DESCRIPTION OF EMBODIMENTS
[0062] The present invention provides methods and systems that combine performing DNA sequencing of a sample using different sequencing techniques (also referred to as sequencing assays) followed by computational methods to mitigate the challenges of DNA analysis, for example using cfDNA for genetic analysis of liquid biopsies and for NIPT. This approach aims to balance the advantages of various technologies while addressing their limitations, providing a more accurate and cost-effective solution for genome-wide mutation detection and genotyping.
[0063] Specifically, the methods and systems of the invention comprise combining different sequencing assays, for example, WES or other NGS panels with WGS, and using these assays together and separately to verify, improve or modify the predicted genetic results (e.g., the genotype at a given genomic locus).
[0064] Figure 1 shows an exemplary outline of an embodiment of the method of the invention showing the combination of multiple sequencing methods for noise and bias learning and correction.
[0065] First, a DNA sample is obtained. Next, the DNA in the DNA sample is sequenced using two or more DNA sequencing assays, for example using WES and WGS. Any number of different sequencing assays may be employed (represented by n in Figure 1). Each of the different assays is performed on the same DNA sample. Next, the bias and noise (also referred to herein as “systemic noise”) patterns for each type of assay are learned.
[0066] Accordingly, in a first of its aspects, the present invention provides a method of genetic analysis of a subject, comprising: a. Obtaining one or more samples comprising DNA; b. Sequencing the DNA in said one or more samples, wherein said sequencing is performed using at least a first and a second DNA sequencing method, thereby obtaining two or more sets of DNA sequencing data, for each of said DNA samples; wherein the first and the second sequencing methods are different sequencing methods; c. Constructing a profile of noise and / or bias and / or error for each of said two or more sets of DNA sequencing data, or for a subset thereof, wherein said profile comprises one or more features for each analyzed genomic locus; d. Assessing systemic noise and bias by comparing the noise and / or bias and / or error profiles of the two or more sets of DNA sequencing data (reads), or a subset of reads thereof, thereby obtaining one of more feature indicators describing the relationship between the two or more sets of DNA sequencing data; e. Performing a genetic analysis of the DNA in said one or more samples to provide genetic predictions; and f. Applying the noise and bias feature indicators obtained in (d) to correct the outcome of the genetic predictions obtained in (e).
[0067] In the context of the present invention, genetic analysis comprises one or more of copy number prediction, genotyping, mutation prediction, DNA fragment classification and DNA fragment origin de -convolution.
[0068] In an embodiment, each of the copy number prediction, genotyping, mutation prediction, DNA fragment classification and DNA fragment origin de-convolution are associated with a disease or disorder, and therefore the genetic analysis encompasses determining whether the subject carries a genetic disease.
[0069] As used herein the term “noise and / or bias and / or error profile” refers to the assessment of noise and / or bias and / or error patterns in a set of DNA sequencing data or a subset thereof. The profile comprises a set of features (may also be referred to as statistics) for each of the analyzed genomic loci, wherein said features are features of the DNA sequence.
[0070] As used herein the term “assessing systemic noise and bias” refers to assessment of the connection, relationship, or similarity between specific noise and / or bias and / or error profiles (i.e., patterns) in a first set of DNA sequencing data, or a subset thereof, and these patterns as they are observed in another set of DNA sequencing data, or a subset thereof, with respect to the same genomic locus. The comparison of the profiles results in feature indicators describing the relationship between a given feature in one set or subset of DNA sequencing data and at least one other set or subset of DNA sequencing data.
[0071] Systemic noise and bias can be learned by comparing data obtained by the different sequencing methods on various features of a genomic locus. Genetic predictions can then be supported, verified or changed according to the assessment of noise and bias. The features that can be used for assessing noise and bias include but are not limited to:
[0072] Set A:
[0073] Predicted fetal genotype.
[0074] - Alternate alleles in the genomic locus.
[0075] Total count of alternate alleles in genomic locus.
[0076] Probabilities of each possible genotype.
[0077] Count of reads covering a genomic locus (i.e., read depth).
[0078] Counts of reads supporting the human reference allele.
[0079] Counts of reads supporting an alternate allele, i.e., a mutation.
[0080] Genotype quality, for example the probability difference between most probable and second most probable genotype.
[0081] Predicted variant type: for example, single nucleotide variant (SNV), small insertion or deletion, etc.
[0082] SNV type, i.e. allele transition or transversion.
[0083] Features describing the counts, distributions and ratios of the position of an alternate allele within the read alignment.
[0084] Features describing the counts, distributions and ratios of reads arising from the forward and reverse strand.
[0085] Features describing the counts, distributions and ratios of DNA fragment lengths of reads covering the genomic locus.
[0086] Features describing the genomic context: proximity to genes or location in gene (such as 3’UTR, 5’UTR, intron, exon), proximity to telomere or centromere, overlap with repetitive elements.
[0087] Features describing the counts, distributions and ratios of read mates (reads arising from the same DNA fragment) which overlap each other, including subsets of such reads, e.g., overlapping read mates which support the same allele (concordant reads) or a different allele (discordant).
[0088] Features describing the counts, distributions and ratios of sequencing base quality or read mapping quality for reads covering the genomic locus, or for any subset of reads (for instance, the read mapping quality of concordant reads covering a genomic locus. Features describing the read-based phasing (haplotype prediction) of a genomic locus: o Whether phasing was successful, i.e. some reads or read mates cover at least one other variant adjacent to the assessed genomic locus. o Predicted haplotype probability, e.g., the probability difference between the most and second-most probable haplotypes. o Whether the genotype at a genomic locus was changed due to haplotype information. o Count and ratio of genomic loci at a region with changed genotyped.
[0089] Features describing the population-based phasing of a genomic locus o Whether phasing was successful, i.e., whether the variant could be found in population reference sets. o Predicted haplotype probability, e.g., the probability difference between the most and second-most probable haplotypes. o Whether the genotype at a genomic locus was changed due to haplotype information. o Count and ratio of genomic loci at a region with changed genotyped.
[0090] Statistics about the genomic regions encapsulating the genomic locus or variant in question: o GC content o End-motif content o End-coordinate content o Fragment length content, e.g. long / short ratio
[0091] Population based features: o Allele frequency for each alternate allele o Inbreeding coefficient as estimated from the genotype probabilities persample when compared against the Hardy-Weinberg expectation
[0092] Features related to structural variant detection: o Number of split reads o Alignment position of split reads o Orientation of split reads o Discordant read pair insert size o Discordant read orientation o Discordant read mapping quality o Average read depth in genomic region of interest o Variability in read depth in genomic region of interest o Length of read coverage gap o GC content surrounding potential structural variant o Homopolymer lengths surrounding potential structural variant o Sequence motif at breakpoint junction o Average and variability of the mapping quality of reads supporting the structural variant
[0093] Features related to structural variant detection using long-read sequencing o Number of phased variants supporting the structural variants o Number of full or almost full-length reads supporting structural variant
[0094] Set B:
[0095] Features describing the relationship (e.g., difference or distance) between a feature in one of the samples (parental, assay 1 or assay 2) and other samples within the same family.
[0096] Set C:
[0097] Features describing the relationship between any described feature (including Set A and Set B features) and the same features within a group of verified or assumed verified negative and positive variant or genotype results.
[0098] Any combination of the above features.
[0099] As used herein the term “phasing” or “haplotype phasing” refers to haplotype deduction.
[0100] Based on the set of values (i.e., the statistics) and relationship indicators mentioned above, the genetic predictions are corrected by applying a suitable mathematical model, e.g., a trained ML / AI model or a statistical correction model.
[0101] As used herein to “provide genetic predictions” means applying scores concerning the probability of having a specific genotype (i.e., homozygous, heterozygous, etc.) with respect to a genetic variant, also referred to as “probability scores”. In one embodiment, the method of the invention is used for NIPT. In such case, the DNA sample is a parental (maternal and optionally paternal) DNA sample and maternal cfDNA.
[0102] In an embodiment, the method of genetic analysis in accordance with the present invention refers to determining whether a fetus possesses a genetic disease.
[0103] If said method of genetic analysis identifies that the fetus possesses a mutation associated with a genetic disease, the method further comprises one of the following steps:
[0104] (i) Performing an invasive prenatal test (for example, amniocentesis, or chorionic villus sampling) to further confirm the predictions of the NIPT,
[0105] (ii) administering a prenatal or a post-natal treatment for said genetic disease in an amount effective to prevent or treat said disease, wherein said treatment comprises pharmaceutical based intervention, surgery, genetic therapy, nutritional therapy, or combinations thereof; or
[0106] (iii) performing a pregnancy termination.
[0107] In another embodiment, the method of the invention is used for diagnosis, for example of malignancy or immune status (e.g., a pathologic immune status such as autoimmunity or inflammation) in a tested subject. In such a case, the DNA sample is a fluid biopsy taken from the subject.
[0108] If said method of genetic analysis identifies that the subject possesses a mutation associated with malignancy or a pathologic immune status, the method further comprises one of the following steps:
[0109] (i) Performing a further test (for example, taking a biopsy) to further confirm the predictions of the method of the invention, or
[0110] (ii) administering treatment for said malignancy or pathologic immune status in an amount effective to prevent or treat the disease, wherein said treatment comprises pharmaceutical based intervention, surgery, genetic therapy, nutritional therapy, or combinations thereof. Cell-free DNA (cfDNA) also referred to as “circulating free DNA” are DNA fragments existing outside of ceils in vivo circulating in body fluids such as blood plasma. The fragments of cfDNA typically have lengths ranging from about 150 to 200 base pairs (bp), and averaging about 170 bp, which presumably relates to the length of a DNA stretch wrapped around a nucleosome. During pregnancy, cell-free fetal DNA can be found circulating in maternal plasma. Thus, the cfDNA in maternal plasma is a mixture of both maternal and fetal DNA; both the total amount of cfDNA, and the fraction of fetal DNA within it, increases throughout pregnancy.
[0111] The term cfDNA also refers to fragments of DNA that have been obtained from the in vivo extracellular sources and separated, isolated, or otherwise manipulated in vitro. cfDNA can be obtained by extracting DNA from blood plasma after removal of intact cells. Methods for extracting cfDNA are well known in the art, for exampie, as shown in the Examples below.
[0112] The term “genomic DNA ” or “gDNA ” herein refers to DNA existing in a cell in vivo and containing a complete genome of the cell or organism. The term also refers to DNA that has been obtained from the in vivo cell and separated, isolated, or otherwise manipulated in vitro. Typically, the cell is isolated prior to being subjected to lysis to produce in vitro cellular DNA. The term gDNA as used herein does not include cfDNA.
[0113] To obtain DNA sequencing data, the DNA containing samples are subjected to at least two different DNA sequencing methods including, but not limited to, deep whole genome sequencing (WGS), whole exome sequencing (WES), next generation sequencing (NGS), targeted sequencing, panel sequencing, gene sequencing, long-read genome sequencing, paired-end sequencing, single end sequencing, and amplicon sequencing.
[0114] The term “Next Generation Sequencing” (NGS) herein refers to sequencing methods that allow for massively parallel sequencing of clonally amplified molecules and of single nucleic acid molecules. Non-limiting examples of NGS include sequencing -bysynthesis using reversible dye terminators, and sequencing-by-ligation.
[0115] Deep sequencing refers to sequencing a genomic region multiple times, sometimes hundreds or even thousands of times. Deep sequencing of the genome allows researchers to detect rare genetic variants.
[0116] As used herein the term “deep whole genome sequencing” refers to deep sequencing of the entire genome. The sequencing is repeated multiple times, for example, but not limited to between 10 times (10X) and 1000 times (1000X), e.g., 10 times (10X), 20 times (20X), 30 times (30X), 50 times (50X), 100 times (100X), 150 times (X150), 200 times (200X), 300 times (300X), 500 times (500X), or 1000 times (1000X).
[0117] Plasma cfDNA can be subjected to varying sequencing depths.
[0118] In one non-limiting example, the cfDNA in plasma is sequenced 300 times (300X); in other embodiments, the cfDNA in maternal plasma is sequenced 50 times (50X), 100 times (100X), or 150 times (X150).
[0119] In addition, genomic DNA is also subjected to whole genome sequencing. Such genomic DNA may be obtained from any cell type, for example from blood cells, e.g., leukocytes. In an embodiment, whole genome sequencing of paternal and maternal genomic DNA is performed to a targeted depth of between about 20X and 40X, for example 3 OX.
[0120] Whole genome sequencing may be performed using any method known in the art for short or long reads, for example, but not limited to, sequencing by synthesis using Illumina’s sequencing series (e.g., the HiSeq X Ten System, HiSeq 4000, NovaSeq 6000, and Novaseq X and X-plus), W GS by Ultima Genomics sequencing machines, pore-based sequencing, e.g., nanopore WGS sequencing using MinlON device (Oxford Nanopore Technologies), , as well as sequencing systems by PacBio technologies.
[0121] The sequencing generates “reads” which are sequences of DNA fragments of varying lengths. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A T C G). It may be stored in a memory device and processed as appropriate. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information.
[0122] The sequencing input may be of long reads (e.g., from about 1 KBP (kilobase pairs) to about 100KBP, or more) or short reads (e.g., from between about 50 base pairs and 400 base pairs), or a combination of long reads and short reads.
[0123] After sequencing, the reads are aligned to a human reference genome based on sequence similarities. The identification of the genetic variants (i.e., variant sites or mutations) can be performed using a variant calling approach, which is generally based on alignment of the DNA sequencing data and the application of a commercially available variant caller.
[0124] As used herein, the terms “aligned”, “alignment”, or “aligning” refer to the process of comparing a read to a reference sequence and thereby determining whether the read is contained in the reference sequence. If the reference sequence contains the read, the read may be mapped to a particular location in the reference sequence. In some cases, alignment simply tells whether the read is present or absent in the reference sequence.
[0125] Sequence alignment techniques that can be used according to some embodiments of the present invention include, without limitation, Burrows Wheeler Aligner (BWA), ABA, ALE, AMAP, anon, BAli-Phy, Base-By-Base, BHAOS / DIALIGN, Bowtie, Bowtie 2, ClustalW, CodonCode Aligner, Comass, DECIPHER, DIALIGN-TX, DIALIGN-T, DNA Alignment, DNA Baser Sequence Assembler, EDNA, FSA, Geneious, Kalign, MAFFT, MARNA, MA VID, MSA, MSAProbs, MULTALIN, Multi- LAGEN, MUSCLE, Opal, Pecan, Phylo, Praline, PicXAA, POA, Probalign, ProbCons, PR0MALS3D, PRRN / PRRP, PSAlign, RevTrans, SAGA, SAM, Se-AI, STAR, STAR- Fusion, StatAlign, Stemloc, T-Coffee, UGENE, VectorFriends, NovoAlign, and GLProbs.
[0126] Exemplary variant callers suitable for the present embodiments include, without limitation, Genome Analysis Toolkit (GATK) and Freebayes. For example, Freebayes can comprise an alignment based on literal sequences of reads aligned to a particular target, not their precise alignment. GATK can comprise: (i) pre-processing; (ii) variant discovery; and (iii) callset refinement. Pre-processing can comprise starting from raw sequence data, e.g., in FASTQ or uBAM format, and producing analysis-ready BAM files; processing can include alignment to a reference genome as well as data cleanup operations to correct for technical biases and make the data suitable for analysis; variant discovery can comprise starting from analysis-ready BAM files and producing a callset in VCF format; processing can involve identifying sites where one or more individuals display possible genomic variation, and applying filtering methods appropriate to the experimental design; callset refinement can comprise starting and ending with a VCF callset; processing can involve using metadata to assess and improve genotyping accuracy, attach additional information and evaluate the overall quality of the callset. Also contemplated are variant callers such as, but not limited to, Platypus, VarScan, Bowtie analysis, MuTect, Google DeepVariant, and / or SAMtools. For example, Bowtie analysis can comprise implementing the Burrows-Wheeler transform for aligning. MuTect can comprise: (i) pre-processing; (ii) statistical analysis; and (iii) postprocessing. Pre-processing can comprise an initial alignment of sequencing reads; statistical analysis can comprise using two Bayesian classifiers, one classifier can detect whether a SNP is non-reference at a given site and, for those sites that are found as nonreference, the other classifier can make sure that the normal does not carry the SNP; postprocessing can comprise removal of artifacts of sequencing, short read alignments, and hybrid capture. SAMtools can comprise storing, manipulating, and aligning sequencing reads stored as SAM files.
[0127] In an embodiment, the step of determining a probability that a sequence read in the maternal plasma cfDNA is of fetal origin is performed for example by an algorithm, e.g. the algorithm Hoobari as described in Rabinowitz et al., 2019 and US2021 / 0340601, to calculate the fetal fraction (FF), i.e., the percent of fetal derived cfDNA within the maternal plasma cfDNA, and to calculate the fragment length distributions. This step may also comprise extracting various fragment-level characteristics.
[0128] In various exemplary embodiments of the invention the method comprises the determination of the probability, of each DNA fragment (or read), to be of fetal origin. Using these read-level probabilities, a prediction ofthe genotype is made, i.e., a prediction whether the fetus carries a normal allele, a pre-mutation, or a full-mutation.
[0129] In an embodiment, the determination of the probability of the DNA fragment (or read) to be of fetal origin comprises constructing a fetal size distribution and a maternal size distribution, binning said fetal size distribution and calculating a fetal fraction for each fragment size bin, and calculating, for at least one size and at least one fragment at said at least one site, a probability that said fragment is fetal, based on a fetal fraction of a respective fragment size bin to which said fragment belongs.
[0130] As used herein the term “fetal fraction ” or “FF” refers to the portion of fetal cfDNA, within the total amount of cfDNA in the maternal blood. The portion of fetal cfDNA within maternal blood (the fetal fraction) varies throughout the pregnancy, and between individuals, hence this is not regarded as a constant but as a variable. Low levels of fetal cfDNA are referred to as a low fetal fraction. In an embodiment, said determining the probabilities comprises applying a Bayesian procedure. Optionally, said Bayesian procedure comprises prior probabilities calculated using sequencing data of at least one of said parents.
[0131] In an embodiment, this procedure further comprises recalibration of the output of said Bayesian procedure using machine learning.
[0132] In a specific embodiment the determination of the probability, for each DNA fragment (or read), to be of fetal origin is performed using variant calling, for example using the Hoobari algorithm as described in Rabinowitz et al., 2019 and US2021 / 0340601.
[0133] The variant calling prediction can optionally be combined with analysis of fragmentomic features of the DNA reads.
[0134] The term “fragmentomic features” refers to molecular characteristics of DNA reads, as well as to genomic, epigenetic and alignment features of the DNA read. Fragmentomic features include, but are not limited to, Read quality mapping, Read base qualities, Fragment length, short / long read ratio, DNA fragment end motifs, Cleavage patterns around methylation sites, Read endpoint preferred end, DNA accessibility / nucleosome positioning inference, Distance to nearest nucleosome, Transcription factor binding sites, Regional fetal fraction, Regional sequence composition, Read sequence composition, and Number of sequence errors in the read.
[0135] In certain embodiments, the variant calling probabilities, and the fragmentomics- based probabilities are multiplied to calculate joint probabilities, for which a maximumlikelihood approach is applied to predict the genotype at each site.
[0136] EXAMPLES
[0137] Materials and Methods
[0138] Sample collection and DNA extraction
[0139] DNA samples from one family (mother + father) were collected during week 11 of the pregnancy with informed consent. DNA from chorionic villus sampling (CVS) was extracted using the magLEAD 12gC, MagDEA Dx kit (ExScale, Chiba, Japan). Peripheral maternal blood was collected using 2-4 Ethylene-diamine-tetra-acetic acid (EDTA) tubes. Within one hour of collection, plasma was separated from blood by centrifugation at room temperature for 10 minutes at 1600 x g. The plasma was then centrifuged again at 16,000 x g for 10 minutes at room temperature to remove any residual cells. Extraction of cfDNA was performed using the MagMAX™ Cell-Free DNA Isolation Kit (Thermo Fisher). Parental (maternal and paternal) genomic DNA was extracted from peripheral blood mononuclear cells (PBMC) using a standard protocol that includes (i) buffy coat separation and (ii) DNA purification using Mag-Bind® Blood & Tissue DNA kit (Omega Bio-tek Inc) according to the manufacturer's instructions.
[0140] Library preparation and sequencing
[0141] Library preparation was performed using the TruSeq DNA PCR-Free Library Prep Kit (Illumina), for genomic DNA, the Accel-NGS 2S PCR-free Library Prep Kit for cfDNA samples that underwent WGS, and the Sure Select XT HS2-V8 for cfDNA samples that underwent WES, according to the manufacturer's instructions. This was followed by sequencing using the NovaSeq platform (Illumina) targeting 150-bp paired-end reads across every DNA sample from each family unit.
[0142] Paternal and maternal genomic DNA were each sequenced using PCR-free 3 Ox WGS.
[0143] Fetal genomic DNA obtained using Chorionic Villus Sampling (CVS) was sequenced using PCR-free 30x WGS.
[0144] Parental, fetal and cfDNA were analyzed as described in Rabinowitz et al, 2019.
[0145] Maternal cfDNA (containing fetal cfDNA) was analyzed using 3 sequencing approaches:
[0146] - WES-only: WES lOOx
[0147] - WGS-only: WGS 300x
[0148] Combined: WES lOOx + WGS 45x
[0149] Parental and fetal genomic DNA were analyzed in the same way for all three cfDNA sequencing approaches.
[0150] Alignment to the genome
[0151] Reads were aligned to the Genome Reference Consortium Human Build 38 (GRCh38 / hg38) using Burrows-Wheeler vO.7.834 with default parameters. Duplicate reads and reads mapping to multiple locations were excluded from downstream analysis. Variant calling
[0152] Single-nucleotide substitutions and small insertions and deletions were identified using the Genome Analysis Toolkit (GATK) Haplotype Caller software v4.2.4.0 applying default parameters and Hoobari. Sequence alignment, removal of duplicate readalignments, parental variant identification, and non-invasive fetal variant calling were executed as outlined in Rabinowitz et al, 2019.
[0153] Noninvasive fetal variant calling
[0154] Hoobari was run using the parental variants and the cfDNA pre-processing results database as input. The output was a standard variant call format (VCF) file. The analysis of the results was held using several software dedicated for VCF manipulation, such as vcflib and vcftools.
[0155] Bayesian noninvasive genotyping
[0156] At each site of interest, a Bayesian calculation was applied. For each possible fetal genotype: where G is the fetal genotype and Gi is the ith possible fetal genotype out of n possibilities. For bi-allelic variants, it would be either homozygous for the reference allele (AA), heterozygous (Aa), or homozygous for the alternate allele (aa). P(G) is the prior probability for each genotype and was calculated by Mendelian laws. The data variable denotes the reads that cover a site and Pldaia | G) denotes the likelihood function, which is defined in this Example as a product of the likelihood of each read:
[0157] The likelihood of a read depends on the fetal genotype and is calculated using the maternal genotype and the fetal fraction. are the probabilities of a read-observation that supports a certain allele, given that the read is fetal or maternal, respectively. This depends on the tested fetal genotype G,. the maternal genotype GM and the observed allele. P(fet) and P(mat) are the probabilities of observing a fetal or maternal read based only on the fetal fraction, and regardless of the allele that it supports. In order to utilize the size differences between fetal and maternal fragments, the fetal fraction used for each read was calculated only from reads with the same fragment size. For reads that are not properly paired or have a fragment size of >500, the total fetal fraction is used.
[0158] Example 1: Combined sequencing improves sensitivity and specificity
[0159] The fetal fraction was calculated using the algorithm Hoobari as described in Rabinowitz et al., 2019 . The calculated fetal fraction was 10.87%.
[0160] A comparison of the sensitivity, specificity and no-call rate (percentage of results that are filtered out due to low confidence) was done between 4 inheritance categories, across the 3 approaches (Figure 2):
[0161] Maternal - genomic loci in which only the mother carries a variant, i.e., an allele different from the allele that appears in the human reference genome.
[0162] Paternal - genomic loci in which only the father carries a variant. Biparental - genomic loci in which both parents carry a variant.
[0163] Compound - i.e., compound heterozygosity, a situation in which both parents are carriers of different variants in the same gene.
[0164] The combination of sequencing methods resulted in the following:
[0165] Improved specificity over the WGS-only approach in all categories, except for biparental.
[0166] Improved sensitivity over the WES-only approach, but lower than the WGS-only approach.
[0167] Lower no-call rates than in both the WES-only and the WGS-only approaches.
[0168] Overall, these results show that a method combining two sequencing approaches performed for the same sample can overcome the limitations of both WES and WGS when performed as mono-sequencing methods.
Claims
CLAIMS:
1. A method of genetic analysis of a subject, comprising: a. Obtaining one or more samples comprising DNA; b. Sequencing the DNA in said one or more samples, wherein said sequencing is performed using at least a first and a second DNA sequencing method, thereby obtaining two or more sets of DNA sequencing data, for each of said DNA samples; wherein the first and the second sequencing methods are different sequencing methods; c. Constructing a profile of noise and / or bias and / or error for each of said two or more sets of DNA sequencing data, or for a subset thereof, wherein said profile comprises one or more features for each analyzed genomic locus; d. Assessing systemic noise and bias by comparing the noise and / or bias and / or error profiles of the two or more sets of DNA sequencing data (reads), or a subset of reads thereof, thereby obtaining one of more feature indicators describing the relationship between the two or more sets of DNA sequencing data; e. Performing a genetic analysis of the DNA in said one or more samples to provide genetic predictions; and f. Applying the noise and bias indicators obtained in (d) to correct the outcome of the genetic predictions obtained in (e).
2. The method of claim 1 wherein the genetic analysis comprises one or more of copy number prediction, genotyping, mutation prediction, DNA fragment classification and DNA fragment origin de-convolution.
3. The method of any one of claims 1 or 2, wherein said DNA is cell -free DNA (cfDNA).
4. The method of any one of the preceding claims wherein said sample comprising DNA is maternal plasma sample comprising cfDNA, and wherein said genetic analysis of a subject comprises genotyping a fetus.
5. The method of claim 4 wherein said sample comprising DNA is (i) a maternal plasma sample comprising cfDNA and (ii) maternal and / or paternal genomic DNA.
6. The method of any one of claims 1 to 3, wherein said sample comprising DNA is a subject’s body fluid sample (e.g., plasma) comprising cfDNA, and wherein said genetic analysis comprises one or more of cancer detection, assessment of cancer progression, or assessment of response to therapy in the subject.
7. The method of claim 6 wherein said sample comprising DNA is (i) a subject’s body fluid sample (e.g., plasma) comprising cfDNA and (ii) genomic DNA of said subject.
8. The method of any one of the preceding claims wherein said genetic analysis comprises variant calling.
9. The method of any one of the preceding claims, wherein said DNA sequencing methods are selected from a group consisting of whole genome sequencing (WGS), whole exome sequencing (WES), next generation sequencing (NGS), targeted sequencing, panel sequencing, gene sequencing, long-read genome sequencing, paired-end sequencing, single end sequencing, and amplicon sequencing.
10. The method of claim 9 wherein said first and second DNA sequencing methods comprise a combination of NGS and WGS.
11. The method of claim 9 wherein said first and second DNA sequencing methods comprise WES and WGS.
12. The method of any one of the preceding claims wherein said assessment of systemic noise and bias and said genetic predictions are performed on reads of the same DNA sample or a different sample from the same individual.
13. The method of any one of claims 1 to 11 wherein said assessment of systemic noise and bias is based on cumulative data obtained from one or more samples or individuals and said genetic predictions are performed on another DNA sample.
14. The method of any one of claims 1 to 13 wherein the sequencing data obtained by a first of the sequencing methods is used for boosting or filtration of the sequencing data obtained by a second of the sequencing methods.
15. The method of claim 14, wherein said boosting or filtration is performed by uniting or intersecting of concordant or discordant genotyped genomic loci or sequenced DNA reads.
16. The method of any one of claims 1 to 15 wherein the sequencing data obtained by a first of the sequencing methods is used for phasing of genomic loci within of the sequencing data obtained by a second of the sequencing methods.
17. The method of any one of claims 1 to 16 wherein said assessing of noise and bias is performed by comparing the two or more sets of DNA sequencing data to obtain information on one or more features of a genomic locus.
18. The method of claim 17 wherein the information on one or more of the features is obtained using a method selected from a group consisting of descriptive statistics, manual filtration threshold definition, inferential statistics, regression analysis, clustering, graphbased approaches, Bayesian statistics, Markovian models, statistical learning models, machine learning models, and deep learning models.
19. The method of any one of claims 1-18 wherein said reads of sequencing data are short reads.
20. The method of any one of claims 1-18 wherein said reads of sequencing data are long reads.
21. The method of any one of claims 1-18 wherein said reads of sequencing data comprise short reads and long reads.
22. The method of any one of the preceding claims, wherein providing genetic predictions is based on at least one Sequence Alignment Map (SAM) parameter.
23. The method of any one of the preceding claims, wherein providing genetic predictions is based at least on an observed template length.
24. The method of claim 4, wherein providing genetic predictions comprises calculating a total fetal fraction.
25. The method of claim 24, further comprising constructing a fetal size distribution and a maternal size distribution, wherein said providing genetic predictions of step e in claim 1 comprises binning said fetal size distribution and calculating a fetal fraction for each fragment size bin, and calculating, for at least one size and at least one fragment at said at least one site, a probability that said fragment is fetal, based on a fetal fraction of a respective fragment size bin to which said fragment belongs.
26. The method of any one of claims 24 or 25 wherein providing genetic predictions comprises applying a Bayesian procedure.
27. The method of claim 26 wherein said Bayesian procedure comprises prior probabilities calculated using sequencing data of at least one of said parents.
28. The method of claim 26 or 27, further comprising recalibration output of said Bayesian procedure using machine learning.
29. The method of any one of claims 24 to 28 wherein said providing genetic predictions is performed using fetal variant calling.
30. The method of any one of the preceding claims wherein said providing genetic predictions is performed using the Hoobari algorithm.
31. The method of any one of claims 24 to 30 wherein said providing genetic predictions comprises fragmentomics-based probability that the sequence read is of fetal origin.
32. The method of claim 31 wherein the method comprises extracting fragmentomic features for each cfDNA read identified as overlapping a potential genomic site where the fetus may have a variant.
33. The method of claim 32, wherein said fragmentomic features comprise one or more of read quality mapping, read base qualities, fragment length, short / long read ratio, end motifs, cleavage patterns around methylation sites, read endpoint preferred end, DNA / accessibility / nucleosome, distance to nearest nucleosome, transcription factor binding sites, regional fetal fraction, regional sequence composition, read sequence composition, and number of sequence errors in the read.
34. The method of any one of claims 32 to 33 wherein said providing genetic predictions comprises multiplying the variant calling -based probability with the fragmentomics-based probability to obtain a calculated joint probability.
35. A method for genotyping a fetus, comprising: a. Obtaining samples comprising DNA wherein the samples are of (i) maternal plasma cell-free DNA (cfDNA), and (ii) maternal and optionally paternal genomic DNA (gDNA) from a pair parenting the fetus; b. Sequencing the DNA in said samples, wherein said sequencing is performed using at least a first and a second DNA sequencing method, thereby obtaining two or more sets of DNA sequencing data, for each of said DNA samples; wherein the first and the second sequencing methods are different sequencing methods; c. Constructing a profile of noise and / or bias and / or error for each of said two or more sets of DNA sequencing data, or for a subset thereof,wherein said profile comprises one or more features for each analyzed genomic locus; d. Assessing systemic noise and bias by comparing the noise and / or bias and / or error profiles of the two or more sets of DNA sequencing data or a subset of reads thereof, thereby obtaining one of more feature indicators describing the relationship between the two or more sets of DNA sequencing data for each of said samples; e . identifying potential genomic sites at which the fetus may have a genetic variant; and f. Applying the noise and bias feature indicators obtained in (d) to correct the identification of potential genomic sites at which the fetus may have a genetic variant; thereby genotyping said fetus.
36. A computer software product, comprising a computer-readable medium in which program instructions are stored, which instructions, when read by a data processor, configure the data processor to (1) receive reads of two or more sets of DNA sequencing data, wherein said sequencing is performed using at least a first and a second DNA sequencing method; wherein the first and the second sequencing methods are different sequencing methods, and to (2) execute the method according to any one of claims 1-35.
37. A system for performing genetic analysis, comprising: an input utility for receiving reads of two or more sets of DNA sequencing data of the same DNA sample, wherein said sequencing is performed using at least a first and a second DNA sequencing method; wherein the first and the second sequencing methods are different sequencing methods; and a data processor configured for analyzing said data for executing the method according to any one of claims 1-35.
Citation Information
Patent Citations
Method and device for extracting somatic mutations from single-cell transcriptome sequencing data
US20240120026A1
Noninvasive fetal variant identification using haplotype analysis
WO2024105671A1