High-resolution and non-invasive fetal sequencing
The NIFS method uses a Bayesian Gaussian mixture model to estimate fetal fractions and assign genetic variant origins in maternal plasma, enhancing diagnostic yields for fetal genetic screening and reducing the need for invasive procedures.
Patent Information
- Application Number
- JP2025512747
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-30
- Filing Date
- 2023-08-30
- Publication Date
- 2025-09-04
AI Technical Summary
Current non-invasive prenatal screening methods have low diagnostic yields for pathogenic and likely pathogenic variants, requiring invasive procedures like amniocentesis that pose risks to the mother and fetus, and lack comprehensive genetic screening of the fetal coding sequence or entire genome.
A high-resolution non-invasive fetal sequencing (NIFS) method using cell-free DNA from maternal plasma, employing a Bayesian Gaussian mixture model to estimate fetal fractions and assign genetic variant origins based on fragment size and sequencing characteristics, enabling detection of single nucleotide variants, indels, and copy number variations without paternal or maternal samples.
The method achieves 99.7% detection of SNV sites comparable to invasive exome sequencing, with 93.4% accuracy, allowing for early detection of fetal variants and maternal carrier screening, reducing the need for invasive procedures and improving diagnostic yields.
Smart Images

Figure 2025529155000017 
Figure 2025529155000018 
Figure 2025529155000019
Abstract
Description
Detailed Description of the Invention
[0001] Priority claim This application claims the benefit of U.S. Provisional Application No. 63 / 402,379, filed August 30, 2022, the entire contents of which are incorporated herein by reference.
[0002] Federally funded research and development This invention was made with government support under Grant No. HD081256 awarded by the National Institutes of Health. The government has certain rights in this invention.
[0003] Technical Field Provided herein are methods (e.g., computer-implemented methods) for assigning maternal or fetal origin to one or more genetic variants in cell-free DNA (cfDNA) from a sample from a pregnant mammal, preferably a pregnant human, using a probabilistic model for assigning maternal or fetal origin to genetic variants in DNA from a sample obtained from the pregnant mammal, where the model assigns maternal or fetal origin based on a combination of fetal fraction and DNA fragment size.
[0004] background Early fetal genetic diagnosis can guide clinical management, predict outcome, and provide the basis for precision medicine. 1-3
[0005] overview Provided herein is a computer-implemented method that can be used to assign maternal or fetal origin to one or more genetic variants in cell-free DNA (cfDNA) from a sample from a pregnant mammal, preferably a pregnant human.The method includes: (a) accessing from a memory a probabilistic model for assigning maternal or fetal origin to genetic variants in DNA from a sample obtained from a pregnant mammal, where the model assigns maternal or fetal origin based on fetal fraction and / or DNA fragment size in combination with other sequencing characteristics; (b) inputting into the model a set of values that represent one or more genetic variants detected in cfDNA from a peripheral blood sample from the pregnant mammal, where the values include empirically determined sequence information for each genetic variant, such as the ratio of different bases in the reads and DNA fragment size information, such as rank sum statistics; and (c) using the model to assign maternal or fetal origin to one or more genetic variants.
[0006] In some embodiments, the genetic variant comprises a single nucleotide variant (SNV), an indel, and / or a copy number variation (CNV).
[0007] In some embodiments, an initial set of values representing one or more genetic variants is obtained by a method comprising: aligning raw sequencing reads derived from cfDNA to a reference genome sequence; converting the raw sequencing reads to consensus reads; realigning the consensus reads to the reference genome sequence, thereby generating a set of aligned consensus reads; identifying consensus reads that differ from the reference genome; assigning the consensus reads that differ from the reference genome as alternative alleles and assigning the consensus reads that match the reference genome as reference alleles; and determining a fragment size rank-sum statistic that represents an estimated fragment size distribution of reads that support the reference allele compared to the fragment size distribution of reads that support the alternative allele, thereby obtaining an initial set of values representing sequence identity and DNA fragment size rank-sum statistic for the one or more genetic variants.
[0008] In some embodiments, each of the raw sequencing reads includes a unique molecular identifier (UMI), and the method includes converting the raw sequencing reads into a single consensus read for each UMI.
[0009] In some embodiments, the method further includes, prior to step (b), selecting the set of candidate variants by a method that includes: accessing from memory a machine learning classifier, optionally a random forest-based model, where the machine learning classifier is trained using a predetermined set of filter criteria and a subset of sites present in the sample or reference sample to identify potential false positive (FP) sites; inputting the initial set of variants into the machine learning classifier; and using the trained machine learning classifier to filter out the set of variants enriched in false positive (FP) sites, thereby selecting a set of candidate variants from the initial set.
[0010] In some embodiments, the probability model uses k-arithmetic means or a Bayesian mixture model that simultaneously estimates fetal fractions and assigns fetal or maternal origin for each variant site in the set. In some embodiments, the Bayesian mixture model is a Bayesian Gaussian mixture model constrained on variant allele fractions and fragment size, e.g., fragment size rank sum statistics.
[0011] In some embodiments, the fetal fraction of a sample is modeled as a latent variable (f), and the arithmetic mean of the variant allele fraction distribution for each component is set based on f.
[0012] In some embodiments, the fetal fraction is estimated based on a reference fetal fraction determined based on clusters derived from VAF across sites.
[0013] In some embodiments, the method further includes outputting a list of one or more genetic variants identified as having fetal origin and / or one or more genetic variants identified as having maternal origin.
[0014] In some embodiments, the method includes comparing the genetic variant to a database containing a list of genetic variants and information about variants with potential medical relevance to the fetus or the mother, identifying potentially medically relevant variants present in the fetus or the mother, and outputting a list of one or more genetic variants identified as having fetal origin and / or one or more genetic variants identified as having potentially medically relevant maternal origin.
[0015] In some embodiments, the methods further include methods that may further include recommending further testing based on the presence of a potentially medically relevant variant.
[0016] In some embodiments, further testing includes amniocentesis or chorionic villus sampling (CVS); further monitoring of the fetus by ultrasound; or genetic testing of the mother.
[0017] In some embodiments, the method further comprises using high-throughput sequencing on the cfDNA extracted from a single sample of peripheral blood from the mother, and optionally exome capture is performed before sequencing.The method of the present invention does not need (and typically does not) use paternal blood samples or sequences (for example, for benchmarking or any other purposes), and optionally does not use a separate maternal-only sample (for example, for benchmarking or any other purposes).The method may, but does not necessarily, comprise determining maternal genotype from leukocytes as described herein, and in some embodiments, the method is carried out only using cfDNA from a single sample of plasma from the mother.Therefore, the method of the present invention can be carried out using a single sample without requiring separate samples from maternal and paternal genomes, or for normalization to a reference panel.
[0018] In some embodiments, adapters with common PCR primer sequences and unique molecular identifiers (UMIs) are attached to the cfDNA, and PCR amplification is performed prior to sequencing.
[0019] In some embodiments, the method further includes enriching the sample for fetal DNA, optionally including fetal protein-coding genes or other regions of the fetal genome that may be relevant to clinical interpretation or variant identification, by contacting the cfDNA with a plurality of oligonucleotides that bind to portions of the fetal genome.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs. Methods and materials for use in the present invention are described herein. Other suitable methods and materials known in the art can also be used. Materials, methods, and examples are illustrative only and are not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.
[0021] Other features and advantages of the invention will become apparent from the following detailed description and drawings, and from the claims. DESCRIPTION OF THE DRAWINGS [Brief explanation of the drawings]
[0022] [Figure 1A] Workflow for noninvasive fetal exome screening using NIFS. Figure 1A shows the process of cell-free DNA (cfDNA) extraction from maternal plasma followed by exome capture. We generated 51 libraries spanning gestational ages and calculated the fetal fraction (fetal cfDNA / total cfDNA) in each sample. [Figure 1B] Workflow for noninvasive fetal exome screening using NIFS. Figure 1B highlights the novel variant detection method developed to account for the fetal fraction and corresponding unique allele fraction at each site depending on the maternal and fetal genotype combination present in the cfDNA. Each cluster represents a unique maternal / fetal genotype combination, and clusters are colored by the genotype generated from direct exome sequencing (ES) of maternal and fetal DNA. These data highlight the merging of clusters at maternal heterozygous sites at lower coverage and fetal fractions, as well as novel and paternally derived variants that were apparently resolved regardless of fetal fraction (fetal 0 / 1; maternal 0 / 0). [Figure 1C] Workflow for noninvasive fetal exome screening using NIFS. Figure 1C shows the application of NIFS to 14 cases referred for invasive testing and representative variants of clinical interest in Table 4, including a pathogenic splice variant in COL2A1 (NC_000012.12:g.47982610C>T) in a fetus with micrognathia consistent with Stickler syndrome, a 4MB pathogenic deletion on chromosome 7 in a fetus with multiple congenital anomalies (NC_000007.14:g.155368937-159327017del), and maternal carrier screening to detect a PAH risk variant associated with phenylketonuria (PKU) (NC_000012.12:g.102866632C>T). All variants were orthogonally validated (Table 2). [Figure 2] An example workflow divides data processing into three stages. In the alignment and preprocessing stage, raw sequencing reads derived from exome-based sequencing (ES) of cfDNA are aligned to the reference genome, grouped by unique molecular identifiers (UMIs), and converted to a single consensus read for each UMI. The consensus reads are then realigned to the reference and base quality scores are recalibrated, generating a set of aligned consensus reads ready for downstream variant calling and analysis. Next, in the variant filtering and genotyping stage of the workflow, Mutect2 is used to identify candidate variant sites. Variants are filtered using a set of hard filters and a random forest-based model trained on a subset of sites present in the sample. A Bayesian mixture model is used to simultaneously estimate fetal fractions and assign fetal and maternal genotypes to each site. Finally, in variant interpretation, all passing variants are annotated and evaluated to create a list of clinically relevant variants for interpretation. [Figure 3A]The ratings of "Unfiltered variant detection," "Filtered variant detection," "Overall genotyping performance," "Expected paternal or novel variant detection," and "Genotyping accuracy for maternal heterozygous variants" are plotted against fetal fraction. Theoretical sensitivity and detection of non-maternal variants is strong across fetal fractions, but genotyping accuracy, especially for variants heterozygous in the mother, is poorer at lower fetal fractions. Sensitivity: TP / (TP+FN); PPV: TP / (TP+FP); Genotyping accuracy: Percentage of maternal heterozygous variants assigned the correct fetal genotype. [Figure 3B] The ratings of "Unfiltered variant detection," "Filtered variant detection," "Overall genotyping performance," "Expected paternal or novel variant detection," and "Genotyping accuracy for maternal heterozygous variants" are plotted against fetal fraction. Theoretical sensitivity and detection of non-maternal variants is strong across fetal fractions, but genotyping accuracy, especially for variants heterozygous in the mother, is poorer at lower fetal fractions. Sensitivity: TP / (TP+FN); PPV: TP / (TP+FP); Genotyping accuracy: Percentage of maternal heterozygous variants assigned the correct fetal genotype. [Figure 3C] The ratings of "Unfiltered variant detection," "Filtered variant detection," "Overall genotyping performance," "Expected paternal or novel variant detection," and "Genotyping accuracy for maternal heterozygous variants" are plotted against fetal fraction. Theoretical sensitivity and detection of non-maternal variants is strong across fetal fractions, but genotyping accuracy, especially for variants heterozygous in the mother, is poorer at lower fetal fractions. Sensitivity: TP / (TP+FN); PPV: TP / (TP+FP); Genotyping accuracy: Percentage of maternal heterozygous variants assigned the correct fetal genotype. [Figure 4]We were able to separate male and female cases through assessment of chrY sequencing coverage. Examination of the number of intervals on chrY using mapped sequencing reads allowed us to detect a confirmed male vanishing twin with a female fetus, which had read coverage across a much greater proportion of chrY than other samples from pregnancies with female fetuses (Table 13). While examining predicted chrY copy status was less accurate, we found an extreme case in which the mother had received a stem cell transplant from a male donor and therefore had six-fold higher coverage on chrY than expected. Furthermore, lower normalized chrY depth distinguished twin pregnancies of discordant gender. We also highlight a single sample with a lower-than-expected fetal fraction that showed low coverage on chrY, and a pregnancy with two XX fetuses that clustered with other samples from pregnancies with a single XX fetus. [Figure 5] Variant allele fraction graph. Histogram of observed variant allele fractions (proportion of reads supporting the alternative allele) plotted for all autosomal sites observed in a sample with 38% fetal fraction at 268x coverage. Distribution peaks are shown along with assignment to maternal or fetal genome genotypes based on the arithmetic mean and variance by fetal fraction and coverage. [Figure 6] Overview of Non-Invasive Fetal Sequencing (NIFS). Plasma (a mixture of fetal and maternal DNA; Step 1) is collected, and then DNA is extracted using a UMI (Step 2). UMI is deposited and exome sequences are captured (Step 3). Bioinformatics analysis is then performed to generate exome-wide, accurate fetal and maternal genotypes (Step 4). Interpretation is performed using best clinical practice (Step 5). [Figure 7A]A) The process for extracting cell-free DNA (cfDNA) from maternal plasma followed by exome capture is shown. We were able to extract both plasma, which consists of fetal and maternal DNA, and DNA from white blood cells, which contains only maternal DNA. The unique maternal DNA from white blood cells can be used for independent variant validation and maternal carrier screening. B) UMIs are attached, exomes are captured, sequencing libraries are constructed, and sequencing is performed on an Illumina machine. [Figure 7B] A) The process for extracting cell-free DNA (cfDNA) from maternal plasma followed by exome capture is shown. We were able to extract both plasma, which consists of fetal and maternal DNA, and DNA from white blood cells, which contains only maternal DNA. The unique maternal DNA from white blood cells can be used for independent variant validation and maternal carrier screening. B) UMIs are attached, exomes are captured, sequencing libraries are constructed, and sequencing is performed on an Illumina machine. [Figure 8] Fragment Size as a Genotyping Feature: Each plot shows the distribution of fragment sizes for reads that can be definitively assigned fetal or maternal origin. Based on genotypes derived from maternal and umbilical cord blood or amniocentesis WES, reads were assigned origin if they overlapped with a variant site that allowed one allele to be unambiguously associated with either the fetus or the mother. For example, if the maternal genome is homozygous and the fetal genome is heterozygous, all reads supporting the alternative allele can be assigned fetal origin. In each sample, fetal-derived fragments are enriched for small fragments and depleted for large fragments compared to maternal-derived fragments. [Figure 9]Filtering and genotyping performance. Cell-free fetal DNA is enriched for short fragments compared to maternal DNA, and we devised a rank-sum test to indicate these deviations. A lower rank-sum statistic indicates an increased number of short fragments, indicating that the variant is more likely to be of fetal origin. This information correlates well with VAF predictions, and our genotyping method uses both of these metrics. [Figure 10] Variant Calling Workflow: Overview of the variant calling process, including initial variant detection with mutect, which can optionally be filtered by maternal genotype when generated from leukocytes as described in Figure 7. Variants are first filtered with machine learning techniques to remove false positives and then genotyped with Bayesian Gaussian mixture models as described below. [Figure 11] Diagram of the graphical model used for genotype assignment. The model is a Bayesian-Gaussian mixture model defined over the variant allele fractions and fragment size statistics (calculated by InsertSizeRankSumTest) at each site, where the arithmetic means of the variant allele fraction components are constrained by a latent variable estimating the fetal fraction. Model information is shown in the table below. [Figure 12] An example of a data processing workflow. [Figure 13] FIG. 1 is a schematic diagram of an exemplary computer system.
[0023] Detailed Description Non-invasive prenatal screening (NIPS) is transforming for the detection of aneuploidies. However, recent systematic benchmarking studies have shown that this current state-of-the-art, low-resolution approach captures only a low fraction of pathogenic and likely pathogenic variants in fetuses with ultrasound-detected structural abnormalities (approximately 26% diagnostic yield from aneuploidy screening alone), whereas much higher diagnostic yields (42–48%) can be achieved by sequencing and analyzing nucleotide changes and copy number variants (CNVs) that alter protein-coding sequences in the human genome (Lowther et al. 2020, BioRxiv; forthcoming AJHG). Several recent studies have demonstrated that targeted approaches to a small panel of genes or CNVs known to be relevant for prenatal diagnosis using cfDNA. 4-15 Although recent advances have shown that genetic screening can be targeted in the fetal coding sequence or the entire genome during pregnancy, at present, comprehensive genetic screening of the fetal coding sequence or the entire genome still requires invasive procedures, such as amniocentesis, which carry inherent risks to the mother and fetus.
[0024] Demonstrated herein is an integrated molecular and computational process to facilitate scalable, non-invasive fetal sequencing (NIFS) for discovering and annotating individual nucleotide sequence variations and CNVs from circulating cell-free DNA (cfDNA) extracted from maternal plasma. This high-resolution NIFS approach encompasses virtually all interpretable pathogenic variants in fetal diagnostic testing (see Lowther et al.'s benchmark from invasive studies of fetal structural anomaly (FSA) cases), or can be applied to whole genome sequencing (NIFS-G) to interrogate all 22,000+ protein-coding genes in the fetal exome (NIFS-E). Here, we focus on NIFS-E as a novel approach to simultaneously provide non-invasive interrogation of the complete fetal exome, as well as routine maternal carrier screening during pregnancy without the need for paternal samples (Figure 1A-C). The success of this method has implications for replacing the current standard of care microarrays and exome sequencing from invasive procedures for prenatal genetic diagnosis, as well as the enterprise of neonatal sequencing, newborn screening, and maternal carrier testing.
[0025] Consistent with current applications of NIPS, we sequenced samples from 51 pregnancies, including 37 samples from the third trimester for method development and 14 samples collected during the first (n = 5) or second (n = 9) trimesters, and observed fetal fractions ranging from 6% to 51%. Applying NIFS, we captured and sequenced 22,995 genes (18,049 protein-coding genes) from the fetal genome. We detected and genotyped single nucleotide variants (SNVs) and indels using a custom pipeline that applied a Bayesian Gaussian mixture model to account for fetal fraction, incorporating approaches such as cfDNA fragment length and sequencing characteristics into our model. We further leveraged these characteristics and read depth ratios for CNV discovery from cfDNA samples. At a sequencing cost much lower than clinical whole-genome sequencing, the NIFS method generated a median exome-wide read coverage of 203x. We further benchmarked NIFS on 11 samples using germline exome sequencing from umbilical cord blood or amniocentesis and found that it captured 99.7% of all 298,576 SNV sites detectable from standard exome sequencing, while maintaining a median sensitivity of 93.0% and accuracy of 93.4% after accurate genotyping of all variants. For all samples, fetal sex was accurately inferred. Importantly, fetal fractions from our method had minimal impact on novel or paternally inherited variants, suggesting the ability to screen for de novo mutations in the fetus very early in pregnancy.
[0026] As a validation experiment, we assessed the clinical utility of NIFS across 14 pregnancies referred for routine genetic testing and detected 100% of variants of interest identified from current standard-of-care clinical trials, including a pathogenic novel CNV in a fetus with multiple congenital anomalies, a potentially pathogenic splice variant in COL2A1 in a fetus with micrognathia, and a homozygous pathogenic indel in CFTR that may cause cystic fibrosis. This NIFS approach accessed both maternal and fetal cfDNA and provided highly sensitive (98.3% sensitivity vs. standard exome sequencing) detection for maternal SNVs and carrier screening, which yielded at least one reportable variant in 57.1% of mothers evaluated, compared with previous estimates. 16,17 .
[0027] The potential utility of NIFS in prenatal screening is broad. This method offers nucleotide-resolution screening to replace the current low-resolution NIPS approach. It could also provide a rapid reflex test for pregnancies with ultrasound abnormalities before invasive procedures are required. Where exome sequencing is currently warranted, variant discovery could be utilized and reinterpreted in the neonatal period. 18,19 This could dramatically shorten the time to diagnosis for many conditions. We also identified variants that may modulate the risk of later-onset related conditions (i.e., BRCA2 variants associated with breast cancer risk), and we believe that such variants may be associated with the current standard of care. 20,21 Based on reportable secondary findings using existing guidelines, interpretation of parental risk can be based on criteria from the American College of Medical Genetics. Indeed, the NIFS approach provides genetic data comparable to current invasive prenatal exome sequencing and requires the same detailed guidelines and procedures for interpretation and appropriate return of results. 22These analyses demonstrate that the complete fetal exome is accessible using novel molecular and analytical approaches such as NIFS from the same maternal plasma samples already routinely collected for low-resolution fetal screening.
[0028] Non-invasive fetal sequencing (NIFS) method Because genetic variants present in cfDNA are a mixture of fetal and maternal fragments, variant allele fractions can be used to inform predictions regarding minor variant genotypes. Variant-supporting reads depend on maternal and fetal genotypes and fetal fraction. These patterns help predict both maternal and fetal genotypes. Figure 5 shows an example graph of VAF plotted against frequency annotated to show the components assigned using the Bayesian-Gaussian mixture model described herein. As fetal fraction decreases, cluster means shift. Sequencing depth can also affect results, as lower coverage results in greater within-cluster variance. Low coverage and low fetal fraction can validate the ability to distinguish fetal genotypes based solely on VAF at sites where the mother is heterozygous.
[0029] In this method, fetal variants are uniquely detectable with high sensitivity and specificity using the NIFS analysis pipeline described herein, which also takes fragment size into account.
[0030] Figure 6 provides a schematic diagram of an exemplary NIFS workflow. An exemplary workflow is shown in Figure 2. As shown, the method is performed on a sample collected from a pregnant woman (step 1, although the method can be performed on a previously collected sample and does not require a sample collection step). In step 2, cfDNA (and optionally maternal DNA, e.g., obtained from white blood cells) is extracted from the sample. Exome capture is then optionally performed, and the cfDNA (and optionally maternal DNA) is sequenced in step 3. Bioinformatics analysis of the sample is performed in step 4, and variant interpretation is performed in step 5.
[0031] The sample (i.e., peripheral blood sample) can be collected using methods known in the art. In some embodiments, 5 to 40 ml, e.g., 20 ml, is collected by blood draw in a pregnant subject. The method can be used in mammals, e.g., human or non-human veterinary subjects. The method of the present invention does not require (and typically does not) use a paternal blood sample or sequence (e.g., for benchmarking or any other purposes), and optionally does not use a separate maternal-only sample (e.g., for benchmarking or any other purposes). The method may, but does not necessarily, include determining maternal genotypes from white blood cells as described herein; in some embodiments, the method is performed using only cfDNA from a single sample of plasma from the mother. Thus, the method of the present invention can be performed using a single sample without the need for separate samples from maternal and paternal genomes or for normalization to a reference panel.
[0032] After the sample is subjected to plasma separation, cfDNA is extracted from the plasma (representing a mixture of fetal and maternal cfDNA). DNA can also be optionally extracted from white blood cells (only maternal DNA can be used for validation). DNA extraction can be performed using methods known in the art, for example, as shown in Figure 7A. Exemplary methods are described in the Library Creation Methods section below. Briefly, plasma is mixed with magnetic beads that bind cfDNA, and then a magnetic field is applied to concentrate the beads, which are then washed, separated, and eluted. Other methods for isolating cfDNA are known in the art, and can be used, for example, spin column-based, manual magnetic bead-based, automated magnetic bead-based methods, etc. (e.g., Polatoglou et al., Diagnostics (Basel). 2022 Oct;12(10):2550; Bronkhorst et al., Tumour Biol. 2020 Apr;42(4):1010428320916314; Bronkhorst et al., Tumour Biol. 2019 Aug;41(8):1010428319866369; Michelson et al., Mitochondrion. 2023 Jul;71:26-39; Streleckiene et al., Biopreserv Biobank. 2019 Dec;17(6):553-561; Solassol et al., Clin Chem Lab Med. 2018 Aug 28;56(9):e243-e246).Numerous kits are available for isolation, including the QIAamp Circulating Nucleic Acid Kit (QiaM, 55114, Qiagen GmbH, Hilden, Germany), NucleoSpin Plasma XS (Macherey-Nagel 740900.50, High Sensitivity Protocol - MNaS, Macherey-Nagel GmbH, Düren, Germany), QIAmp MinElute ccfDNA Mini Kit (QiaS, 55204, Qiagen GmbH, Hilden, Germany), cfPure Cell-Free DNA Extraction Kit (BChM, K5011610-BC, BioChain Inc., Newark, CA, USA), and MagMAX Cell-Free DNA Isolation Kit (TFiM, A29319, Thermo Fisher Scientific, Waltham, MA, USA). Automated methods include the MagNA Pure 24 Total NA Isolation Kit (RocA, 07658036001, Roche Diagnostics GmbH, Penzberg, Germany), the NextPrep-Mag™ cfDNA Automated Isolation Kit (PerkinElmer), and the cfNA2000 protocol of the MagNA Pure 24 System (Roche Diagnostics).
[0033] Preferably, the adapter with common primer sequence and unique molecular identifier (UMI) is attached to DNA to maximize sequence coverage, and PCR is used to amplify the library.Then, this method can include the optional step of enriching the sample for fetal DNA, for example, by contacting cfDNA with a plurality of oligonucleotides that bind to a part of fetal protein-coding genes, for example, TWIST target panel (Alliance Clinical Research Exome), thereby targeting all 22,995 genes from fetal genome (or 18,049 protein-coding genes, or a subset thereof) or a subset thereof, or from other regions of fetal genome that may be relevant for clinical interpretation or variant identification; or this method can include sequencing all nucleotides in genome without exome capture (this method is referred to as NIFS genome; NIFS-G). The UMI-tagged DNA (either from the total cfDNA population, e.g., genomic DNA or exome-enriched DNA) is then sequenced using high-throughput / next-generation sequencing methods, preferably to an average sequencing depth of about 100X, 150X, or 200X. Filtered sequencing depths (i.e., after UMIs are used to filter out related reads) of at least 200, 250, or 300X are preferred for the first and second trimesters, and at least 100X (more preferably, 200, 250, or 300X) for the third trimester. See Figure 7B.
[0034] Sequencing can be automated Sanger sequencing (e.g., using an ABI 3730x1 Genome Analyzer), pyrosequencing on solid support (e.g., using 454 sequencing, Roche), sequencing-by-synthesis with reversible termination (e.g., using an ILLUMINA® Genome Analyzer), sequencing-by-ligation (ABI SOLiD®) or sequencing-by-synthesis with virtual terminators (HELISCOPE®); Molecule sequencing (see Voskoboynik et al. eLife 2013 2:e00569 and U.S. patent application Ser. No. 13 / 608,778, filed Sep. 10, 2012); DNA nanoball sequencing; single molecule real time (SMRT) sequencing; Nanopore It can be carried out by using methods known in the art, including DNA sequencing; hybridization sequencing; mass spectrometry sequencing; and microfluidic Sanger sequencing.Exemplary next-generation sequencing methods known to those skilled in the art include massively parallel signature sequencing (MPSS), polony sequencing, pyrosequencing (454), Illumina (Solexa) sequencing synthesis, SOLiD sequencing by ligation, ion semiconductor sequencing (Ion Torrent sequencing), DNA nanoball sequencing, chain termination sequencing (Sanger sequencing), heliscope single molecule sequencing, single molecule real time (SMRT) sequencing (Pacific Biosciences); flow-based sequencing (for example, Ultima sequencing) and nanopore sequencing.
[0035] Then, novel bioinformatics analysis methods are used to detect and identify variants from sequencing data, discovering short variants (e.g., single nucleotide variants (SNVs) and indels) and CNVs. Generally, data processing methods can be divided into three stages: alignment and preprocessing; variant filtering and genotyping; and variant interpretation. See, for example, Figure 9.
[0036] In the alignment and preprocessing stage, raw sequencing reads derived from cfDNA are aligned to a reference genome, grouped by UMI, and converted into a single consensus read for each UMI. The consensus reads are then realigned to the reference, and base quality scores are optionally recalibrated to improve read quality, resulting in a set of aligned consensus reads ready for downstream variant calling and analysis. Maternal genotyping can optionally be performed, and then germline variants can be filtered using the maternal genome or database. Genotyping data is optionally in Variant Call Format (VCF), a file format used to encode genetic variant sites and genotypes.
[0037] Next, in the variant filtering and genotyping phase of the workflow, candidate variant sites are first identified by comparison with a reference genome (e.g., GRCh38 using Mutect2). Candidate variants can be filtered to remove potential false positive (FP) sites, for example, using a set of hard filters and a machine learning classifier, such as a random forest-based model, a support vector machine (SVM), or a neural net trained on a subset of sites present in the sample. A probability model is then applied to estimate fetal fractions and assign fetal and / or maternal genotypes to all observed variant sites in the cfDNA sequencing data, for example, using k-arithmetic means or Bayesian mixture models, to simultaneously estimate fetal fractions and assign fetal and maternal genotypes to each site.
[0038] In some embodiments, the probabilistic model simultaneously estimates fetal fraction and assigns fetal and maternal genotypes to all variant sites observed in the cfDNA sequencing data using a constrained 2D Bayesian Gaussian mixture model with five components, where each component represents a different combination of maternal and fetal genotypes for autosomal variants. Combinations are defined across two dimensions: variant allele fraction (VAF) and a fragment size rank-sum statistic summing the difference between fragment sizes of reads supporting the reference and alternative alleles, as described herein (e.g., in the section on variant detection in cfDNA with Mutect2). See, e.g., Figure 8. The center of the cfDNA VAF cluster is determined by the fetal fraction (FF). The model (shown in Figure 11) can be fitted using stochastic variational inference (e.g., using Pyro).
[0039] Table A shows the five components used in an exemplary Bayesian Gaussian mixture model. [Table A]
[0040] As shown in Figure 9, incorporating fragment size and variant allele fraction (VAF) into the probabilistic model allows for accurate assignment of origin (maternal or fetal or both) based on, for example, assignment to one of the five components shown above.
[0041] Finally, in variant interpretation, all passing variants are annotated and evaluated to generate a list of clinically relevant variants for interpretation. Annotation can be performed by referencing one or more databases; for example, variants are identified based on the gene and functional consequences (e.g., based on RefSeq). 4 ), allele frequencies (e.g., based on gnomAD v2.1.1 and gnomAD v3.0), and Rare Exome Variant Ensemble Learner (REVEL), which predicts the deleteriousness of each nucleotide change in the genome. 16 Score, ClinVar 17 annotation (updated 2023-04-30), and the Online Mendelian Inheritance in Man database (OMIM, version 2022-07-08 18 ) based genotype (e.g., recessive, phenotype, etc.). Variants may be further filtered, e.g., if they had an allele frequency less than 5, or if they were filtered out in gnomAD v2.1.1 and gnomAD v3.0. 6 Variants can be included if they were not reported in ClinVar, or can be excluded if they were determined to be likely_benign or benign / likely_benign, or synonymous variants in ClinVar. See, e.g., Figure 12 for an exemplary data processing workflow.
[0042] The results can then be used to output a list from each sample for further consideration, preferably including all ClinVar annotated pathogenic / likely pathogenic variants, all frameshift / stopgain variants, splice AI scores, and 19 The list includes all predicted splice variants with a REVEL score >0.95, all non-frameshift variants >15 amino acids, and all non-synonymous variants with a REVEL score >0.7. The list can be shared with, for example, a healthcare provider or the mother.
[0043] Based on the presence of a variant in the fetus that may adversely affect the health of the fetus or the mother (e.g., a pathogenic or likely pathogenic variant, including one associated with poor health or outcome), the method can further include recommending further testing, e.g., invasive testing such as amniocentesis or chorionic villus sampling (CVS), and / or further monitoring by ultrasound.
[0044] Based on the presence of a variant in the fetus that may be harmful, the method may further include recommending further testing, e.g., genetic testing, to confirm the variant.
[0045] Standard computing devices and systems can be used and implemented to perform the methods described herein. Computing devices include various forms of digital computers, such as laptops, desktops, mobile devices, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. In some embodiments, the computing device is a mobile device, such as a personal digital assistant, a mobile phone, a smartphone, a tablet, or other similar computing device. The components, their associations and relationships, and their functions described herein are exemplary only and do not limit the practice of the invention(s) described and / or claimed herein.
[0046] A computing device typically includes one or more of a processor, memory, storage device, high-speed interfaces connecting to the memory and high-speed expansion ports, and low-speed interfaces connecting to low-speed buses and storage devices. The components are interconnected using various buses and may be mounted on a common motherboard or otherwise as needed. The processor may process instructions for execution within the computing device, including instructions stored in memory or storage device for displaying graphical information for a GUI on an external input / output device, such as a display coupled to the high-speed interface. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and memory types, as needed. Multiple computing devices may also be connected, each providing a portion of the operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0047] 13 shows an exemplary computer system 500 including a processor 510, a memory 520, a storage device 530, and an input / output device 540. Each of the components 510, 520, 530, and 540 may be interconnected, for example, by a system bus 550. The processor 510 may process instructions for execution within the system 500. In some implementations, the processor 510 is a single-threaded processor, a multi-threaded processor, or another type of processor. The processor 510 may process instructions stored in the memory 520 or the storage device 530. The memory 520 and the storage device 530 may store information within the system 500.
[0048] The input / output device 540 provides input and output operations to the system 500. In some implementations, the input / output device 540 can include one or more of a network interface device, such as an Ethernet card, a serial communication device, such as an RS-232 port, or a wireless interface device, such as an 802.11 card, a 3G wireless modem, a 4G wireless modem, or a 5G wireless modem, or both. In some implementations, the input / output device can include a driver device configured to receive input data and send output data to other input / output devices, such as a keyboard, a printer, and a display device 560. In some implementations, mobile computing devices, mobile communication devices, and other devices can be used. In some embodiments, the method is performed using an apparatus including a sequencing machine, such as an Illumina sequencer. [Example]
[0049] Example The present invention is further described in the following examples, which do not limit the scope of the invention described in the claims.
[0050] method In the examples below, the following methods were used: Sampling Samples were collected from the Brigham and Women's Hospital (BWH) LIFECODES longitudinal biorepository 23The samples were collected from the MAPing pregnancy study pregnancies based at the Center for Fetal Medicine at BWH (Tables 1-2). This resource was designed to facilitate research related to prenatal screening and the diagnosis and understanding of the genetic basis of fetal structural abnormalities. Although we collected samples from any gestational age, with initial technology development focused on the third trimester, the 14 samples with fetal abnormalities were primarily collected from the first and second trimesters (Tables 3 and 12). Women were enrolled at a prenatal visit any time during pregnancy, and peripheral blood samples were collected in two Streck collection tubes (Streck, La Vista, NE), providing a maximum of 20 mL of maternal blood.
[0051] A total of 51 samples were collected from 49 singleton, 1 dizygotic, and 1 monozygotic twin pregnancies. Samples were collected across all three trimesters, and 11 samples had matched confirmatory samples to generate benchmark data using fetal DNA obtained from umbilical cord blood or amniocentesis.
[0052] How to create a library After sample collection, we isolated cell-free DNA (cfDNA) from maternal serum using Streck's recommended "Double Spin Protocol 2." We extracted maternal genomic DNA from all samples using sedimented maternal leukocytes. To perform a broad benchmark of maternal variant discovery, we collected maternal germline DNA (gDNA) from 28 mothers and performed standard exome sequencing (ES) on the Broad Institute Genomics Platform (Table 3). We extracted DNA from the cfDNA portion of the isolated serum using the NextPrep Mag cfDNA Isolation Kit (catalog no. NOVA-3825-03). We then determined the size and concentration of cfDNA fragments via Tapestation (cfDNA Tape, Agilent Technologies) and QuBit (Broad Range DNA, Agilent Technologies), respectively. To convert cfDNA into sequenceable fragments, we used the NEBNext® Ultra™ II DNA Library Prep Kit for Illumina® from New England Biolabs (NEB) according to the manufacturer's protocol with the following modifications: 1) the NEB adapter and USER enzyme steps were replaced with direct ligation of xGen Stubby unique molecular identifier (UMI) adapters ordered from Integrated DNA Technology (IDT). 2) The NEB primers were replaced with xGen Dual Index Primers (IDT). After adapter ligation, PCR was performed for 12 cycles. Following this initial PCR amplification, libraries were multiplexed into batches of up to 16 (up to 8 μg total material) and exome capture was performed using Alpha Broad Exome baits from TWIST Bioscience, targeting 194,202 exon regions, under the IDT xGen hybridization protocol.Briefly, multiplexed libraries were combined with human Cot DNA and xGen blocking oligos and dehydrated before resuspension in hybridization buffer and bait. After a 4-hour incubation, the bait-hybridized libraries were combined with buffer-resuspended streptavidin beads and washed several times to remove any non-hybridized libraries, followed by 15 rounds of on-bead, post-capture PCR. PCR-amplified libraries were purified using SPRI bead cleanup, and exome libraries were analyzed with Tapestation (D1000, Agilent Technologies) and QuBit (Broad-Range DNA, Agilent Technologies) before multiplexing and sequencing on an Illumina NovaSeq. We were able to obtain additional material for benchmark analysis in a subset of participants in this study, including fetal umbilical cord blood collected at delivery (n = 7), paternal DNA (n = 7), and in four cases, DNA was extracted from cultured cells derived from amniocentesis (Table 3).
[0053] Array Creation We sequenced cfDNA libraries on the Broad Institute Genomics Platform in pooled multiplex sequencing runs on an Illumina Novaseq S4 flow cell. Our multiplexing strategy attempted to generate as many unique sequencing reads as possible while keeping raw sequence duplication (without considering UMIs) below 75%. We note that at a depth of 200x, assuming a 25% fetal fraction (the median fetal fraction observed across our samples), each target was expected to have an arithmetic mean coverage of approximately 50 reads of fetal origin. Eight samples (MGB038, MGB039, MGB40, MGB016, MGB043, MGB046, MGB047, and MGB048) were sequenced across two S4 flow cells, and the raw sequencing reads were pooled and processed together.
[0054] Data processing workflow overview An overview of an exemplary data processing workflow is shown in Figure 2. The workflow is divided into three major sections, which are described at a high level in this section and in more detail below. The first stage preprocessed the raw sequencing reads, constructed consensus reads for each UMI found in the sequencing dataset, and aligned these consensus reads to the reference genome. For a more detailed description of the methods and tools, see "Aligning and Preprocessing cfDNA Sequence Data." The next stage of processing used the somatic variant calling tool Mutect2 to generate candidate variant call sites from the aligned consensus reads (see the section "Variant Detection in cfDNA"). A machine learning-based approach was then used to train a classifier on each sample's dataset and filter variant sites that were likely artifacts (see the section "cfDNA Variant Filtering"). A Bayesian mixture model was then used to simultaneously estimate fetal fractions and assign fetal and maternal genotypes to each variant site ("cfDNA Genotyping"). Finally, a set of protocols was developed to annotate, interpret, and manage variant sites to generate a list of potentially clinically relevant variants for each sample ("Variant Determination").
[0055] Alignment and preprocessing of cfDNA sequence data To generate high-quality GRCh38 aligned clam files for variant calling, we used the following pipeline. First, UMIs were extracted from each read using the open-source fgbio ExtractUmisFromBam (github.com / fulcrumgenomics / fgbio) from Fulcrum Genomics. Several steps were then performed using the open-source Picard tools from the Broad Institute of MIT and Harvard (broadinstitute.github.io / picard / ), including sorting the data by query name using Picard SortSam. Illumina adapters were identified and marked with Picard's MarkIlluminaAdapters. Reads were then converted to FASTQ using Picard's SamToFastq and aligned with the open-source BWA-MEM aligner. 24 We aligned the sequences to the GRCh38 reference genome using the GRCh38 library and merged them back into BAM files using Picard's MergeBamAlignment. We then used the open-source GATK library from the Broad Institute (gatk.broadinstitute.org). 25The PrintReads tool in the framework was used to remove a small number of degenerately mapped fragments with a mapped fragment length of less than 19 bp. Reads were then grouped by UMI using the fgbio tool GroupReadsByUmi. We generated consensus duplex reads using fgbio CallDuplexConsensusReads with the parameters --error-rate-pre-umi=45, --error-rate-post-umi=30, --min-input-base-quality=10, and --min-reads=0. Consensus reads were filtered with fgbio FilterConsensusReads using the parameters --min-reads 0 0 0, --max-read-error-rate 0.35, --max-base-error-rate 0.3, --min-base-quality 40, and --max-no-call-fraction 0.25, and then clipped with fgbio ClipBam using the parameters --clipping-mode=Hard, and --clip-overlapping-reads=true. Match information was fixed and Picard's FixMateInformation was added to the matching CIGAR tag. Consensus reads were sorted by coordinate using Picard's SortSam. Finally, base quality scores were recalibrated using the GATK BaseRecalibrator and ApplyBQSR tools. Metric collection, as well as variant calling, filtering, and genotyping, were then applied to the covered target intervals of the Twist Broad Custom Exome Kit (Twist Alliance Clinical Research Exome). A publicly available version of the covered target data is available at twistbioscience.com / resources / data-files / twist-alliance-clinical-research-exome-349-mb-bed-files.
[0056] This target interval list was downloaded from the UCSC Genome Browser (NCBI annotation release 110) to the GRCh38 refSeq database. 26 The number of genes covered was calculated by crossing
[0057] Coverage Analysis We applied the Picard tool CollectHsSequencingMetrics to collect coverage statistics across all exome targets based on aligned consensus reads. In Table 3, we report for each sample the arithmetic mean coverage across all target intervals and the fraction of target bases with at least 50x coverage by cfDNA sequencing reads (of mixed maternal and fetal origin). We also multiplied each sample's per-target arithmetic mean coverage metric by that sample's estimated fetal fraction to generate the proportion of exome target intervals with at least 8x and 10x arithmetic mean estimated fetal read coverage; both values for each sample are reported in Table 4. Finally, for samples with matched paternal gDNA ES, we report the median and interquartile range of the number of reads supporting only the paternal allele in Table 7, "Genotyping Performance and Coverage Based on Paternal gDNA ES." These values can be used to infer the distribution of coverage of these sites by sequencing reads originating from the fetal genome, half of which are expected to support the paternal allele.
[0058] Detection of variant sites in cfDNA We used the open source software tool Mutect2 from the Broad Institute of MIT and Harvard with the parameter --max-mnp-distance 0 to split MNP variants into separate records and the following parameters to generate the annotations used for filtering: -G StandardAnnotation -G StandardHCAnnotation -A MappingQualityZero -A TandemRepeat -A CountNs27 After identifying candidate variant sites using the cfDNA genotyping engine, we generated additional annotations for use in genotyping by modifying GATK to add an InsertSizeRankSum annotation to each variant based on the fragment sizes of reads supporting the reference allele and alternative allele at each site (see the section "cfDNA Genotyping"). To generate this annotation, we used a Mann-Whitney U test (implemented by GATK's RankSumTest) to compare the distribution of estimated fragment sizes of reads supporting the reference allele with the distribution of fragment sizes of reads supporting the alternative allele. The annotation value is the Z-score of the U statistic. Fragment sizes were estimated for each read determined to be informative at each site by the Mutect2 / GATK assembly-based calling engine based on the mapped insert size reported by BWA in the BAM file for each read pair, and were adjusted to account for insertions and deletions reported in the read's CIGAR and mate CIGAR.
[0059] cfDNA variant filtering To remove potential false positive (FP) sites due to sequencing, library preparation, or alignment errors, a variant site filter was developed that included hard filtering rules and a random forest-based classifier that assigned each variant site a score reflecting the likelihood that the site was a true positive (TP) variant. The filtering rules were as follows:
[0060] 1. Any site in a cfDNA sample was filtered if Mutect2 identified two or more alternative alleles and at least one of the alternative alleles was an indel.
[0061] 2. A machine learning classifier (described in detail below) was applied to score the variants, filtering out any variants with a score below a cutoff determined by assessing sensitivity against a gold standard set of common variants.
[0062] 3. Indels that passed our random forest filter but were likely to be recurrent sequencing errors were hard-filtered based on a list of recurrent artifactual indel calls observed in our dataset. To construct this list, we identified all indel sites with an allele count of at least 5 in the subset of cfDNA samples from this study that did not have a matched cord blood or amniocentesis sample and were not among the samples referred for genetic testing. From this resulting list, we used gnomAD v3.1.2 28 We removed any sites present at any allele frequency in the database. Using the remaining sites, we created a catalog of 969 indels that represent recurrent artifacts in our data. We applied a filter to remove indel sites with positions and alternate alleles that matched one of the sites in this catalog.
[0063] 4. Any site confidently called by Mutect2 that is in phase with an SNV site that does not pass one or more of the above filters was also filtered. Mutect2 calls a specific set of sites that are in phase with each other based on the number of reads that span two or more sites in the set and support the same combination of alleles. This information is recorded in the phase set ID (PID) annotation for the variant. This filter captures clustered sets of sites that represent mapping errors when reads originate from other paralogous sequences in the genome that contain multiple paralog-specific variants.
[0064] In addition to the filters above, we applied two filters after genotyping based on identifying variant sites with unexpectedly low numbers of reads supporting the alternative allele (see section cfDNA genotyping).
[0065] The machine learning classifier described in step 2 above is based on the principle of label-free learning, where only the positive training labels are known with certainty in the training dataset. 29 A new instance of this classifier was trained for each sample using only the data from that sample. We used gnomAD v3 with a maximum subpopulation frequency of at least 0.1 (given by the AF_popmax annotation in the gnomAD data) because variant sites common to the population are likely to be real. 28 We assigned an initial positive label to sites present in the . All other sites were initially assigned a negative training label. Next, we trained a random forest using 800 estimators implemented by the scikitlearn package. After training the classifiers, we scored each variant with the predicted probability that the site was true positive according to the classifier. We then identified a cutoff for this score for PASS filter status using the FilterVariantTranches tool in GATK. This tool finds the optimal cutoff that yields the required estimated sensitivity based on a set of common SNPs and indels provided as a resource with best-practice pipelines. In our pipeline, we requested a sensitivity of 99.5 for the --snp-tranche and 95.0 for the --indel-tranche. The following features, chosen to be independent of allele fraction and added based on the genomic context of the site coordinates or generated by Mutect2, were selected for assessment in the random forest:
[0066] Indel: A binary characteristic indicating whether a variant is a SNP or an indel ●SOR: Strand bias test statistic ●MQ: Multiplicative root mean square mapping quality ●MQRankSum: Mapping quality biased rank sum test ●ReadPosRankSum: Read position bias rank sum test BaseQRankSum: Testing base quality score bias for reference and alternative alleles ●MPOS: Median distance of the site from the end of the lead ECNT: Number of events in the aggregated haplotype containing the variant NCount: The number of reads in the pileup with N base calls (created by forming double consensus reads) at the variant site. ●DP: Depth of variant site SEGDUP: A binary characteristic indicating whether the site is within a segment duplication. ●LCR: Li et al. 30 A binary characteristic indicating whether the site is within a low-complexity region, as defined by the LCR-hs38 resource provided by SIMPLEREP: A binary feature indicating whether the site is within an annotated single repeat. • STR: A binary trait indicating whether GATK / Mutect classifies the site as falling within a short tandem repeat sequence.
[0067] Note that the fact that variants were observed in repetitive genomic regions (annotated by SEGDUP, LCR, SIMPLEREP, and STR annotations) was used as a feature for training the classifier, not as a hard filter, with the goal of allowing the classifier to make confident calls in those regions of the genome.
[0068] cfDNA genotyping and fetal fraction estimation We developed a machine learning-based model to simultaneously estimate fetal fraction and assign fetal and maternal genotypes to all variant sites observed in cfDNA sequencing data. Our model consists of a constrained Bayesian Gaussian mixture model with five components, each representing a different combination of maternal and fetal genotypes for autosomal variants. We defined the mixture across two dimensions: variant allele fraction and the fragment size rank-sum statistic, which summarizes the difference between fragment sizes of reads supporting the reference allele and the alternative allele, as described in the section on variant site detection in cfDNA. The fetal fraction of the sample was modeled as a latent variable (f), and the arithmetic mean of the variant allele fraction distribution for each component was set based on this: 0 / 0 represents the homozygous reference genotype, 0 / 1 represents the heterozygous genotype, and 1 / 1 represents the homozygous alternative genotype. The components and their arithmetic means were defined as follows: ("Cluster 0": fetal 0 / 1, maternal 0 / 0) f / 2; ("Cluster 1": fetal 0 / 0, maternal 0 / 1) (1 - f) / 2; ("Cluster 2": fetal 0 / 1, maternal 0 / 1) 0.5; ("Cluster 3": fetal 1 / 1, maternal 0 / 1) f + (1 - f) / 2; ("Cluster 4": fetal 0 / 1, maternal 1 / 1) 1 - (f / 2). Each data dimension was modeled independently; i.e., the covariance matrix for each component was diagonal. Before performing inference on the model parameters, we removed a subset of sites that appeared to be outliers, including sites with no passing filter status (as set by the filtering procedure described in cfDNA variant filtering above), sites with a cfDNA VAF less than 0.025 or greater than 0.975, and sites with fragment size statistics that were missing, less than -4, or greater than 4. To further clean the data, we removed sites that did not pass outlier tests for cfDNA VAF and fragment size statistics. The outlier tests were performed by fitting the IsolationForest outlier classifier from the sklearn.ensemble package to the data with a contamination parameter of 0.05. We used Pyro 31We defined a genotyping mixture model and fitted it to the data for each sample using stochastic variational inference. We used Pyro's AutoDelta guide function to find the maximum posterior value for each parameter. To initialize the model, we first created an initial estimate of the fetal fraction. We did this by identifying the location of a cluster of sites in the VAF distribution that represent sites that are homozygous in the maternal lineage and heterozygous in the fetus ("cluster 4"). We calculated a Gaussian kernel density estimate for all sites with a VAF less than 0.975 and used the scipy.signal.argrelextrema function to initialize the fetal fraction by identifying the peak in the density with a maximum value corresponding to cluster 4. To estimate the initial value of the arithmetic mean of the fragment size statistic distribution, we found 500 sites with cfDNA VAFs closest to the expected VAF of cluster 4 based on the estimated fetal fraction and used the median of their fragment size statistic values. Once the arithmetic mean of the fragment size statistic distribution for maternal homozygous variants / fetal heterozygous sites was estimated, we initialized the arithmetic means of the other fragment size component distributions by multiplying this value by the vector [-1.0, 0.5, 0.0, -0.5, 1.0] to match the expected relative contribution of maternal versus fetal reads observed for sites in each cluster.
[0069] After fitting the model parameters using stochastic variational inference, we re-added all sites filtered from the above model and solved for optimal cluster assignment parameters for all autosomal sites to the dataset by fully enumerating all latent variables using Pyro's enumeration strategy for discrete latent variables, with a guide function that fixed the learned model parameters but varied the assignment probabilities. We then estimated the likelihood of each possible fetal genotype by summing the cluster component assignment probabilities: the likelihood of a 0 / 0 (ref / ref) fetal genotype at that site is the probability of that site's assignment to cluster 1; the likelihood of a 0 / 1 (ref / alt) fetal genotype is the sum of the assignment probabilities for clusters 0, 2, and 4; and the likelihood of a 1 / 1 (alt / alt) fetal genotype is the probability of assignment to cluster 3. We automatically assigned homozygous alternating genotypes to sites likely to be homozygous alternating genotypes in the cfDNA samples (i.e., VAF greater than 0.975). Similarly, maternal genotype likelihoods were set as follows: the likelihood of a maternal 0 / 0 genotype was set to the probability of assignment to cluster 0; the likelihood of a maternal 0 / 1 genotype was set to the sum of the assignment probabilities of clusters 1, 2, and 3; and the likelihood of a maternal 1 / 1 genotype was set to assignment to cluster 4.
[0070] We applied this model to all autosomal variants in all samples and to variants on the X chromosome in samples with a predicted fetal sex chromosome ploidy of XX. For samples with a predicted fetal sex chromosome ploidy of XY, we used a model that genotyped variants only within the pseudoautosomal region (PAR) of the X chromosome. For chromosome X variants outside the PAR in XY samples, we defined three Gaussian components for each possible pair of maternal and fetal genotypes (excluding variants that were homozygous in both the mother and the fetus). We defined these components based on the parameters learned in training the autosomal model as follows: for clusters representing maternal heterozygous variants in which the fetus carries the variant, we set the VAF arithmetic mean to 1 / (2-f); for clusters representing maternal heterozygous variants in which the fetus does not carry the variant, we set the VAF arithmetic mean to (1-f) / (2-f); and a third cluster representing variants in the fetus (i.e., de novo mutations), with a variant that is the homozygous reference and a VAF arithmetic mean of f / (2-f). The fragment size arithmetic means of these clusters were set to the arithmetic means learned in the autosomal model for clusters 1, 3, and 0, respectively, with variances equal to five times the fragment size variance from autosomal cluster 0 (to account for additional mutations observed at these sites). We assigned genotypes to these variants by calculating the likelihood that each variant was generated by each of these Gaussian components and assigning the variant to that cluster's genotype set accordingly.
[0071] After genotyping, we applied two additional filters to the resulting variant calls. First, we excluded calls with a variant allele fraction too low to be generated by a cluster representing variants in cluster 0, where the fetus is heterozygous and the mother is the homozygous reference. To do this, we performed a lower binomial distribution test on the number of observed reads supporting the alternative allele out of the total depth at that site, using the expected VAF for that cluster with a binomial probability of f / 2, and excluded any sites for which the p-value for this test was less than 1e-5. Second, we excluded indel calls in which the alternative allele was supported by three or fewer reads, as we found high error rates for these variants.
[0072] True data processing: Standard ES from gDNA variant calling in maternal, paternal, fetal, fetal cord blood, and amniocentesis samples gDNA libraries were prepared from maternal, paternal, and fetal cord blood and amniocentesis samples according to standard ES protocols at the Broad Institute Genomics Platform (Cambridge, MA). After Illumina sequencing, reads were aligned and variants were identified according to GATK best practice guidelines. 25 Briefly, following adapter sequence marking and clipping, the preprocessed reads were run using BWA-MEM with default parameters. 24The sequences were aligned to a human reference using GATK. Duplicate reads were marked using Picard MarkDuplicates and excluded from downstream analysis. Base recalibration was performed using GATK BaseRecalibrator and ApplyBQSR (using sites of known mutations from the GATK Reference Bundle). Germline single-nucleotide variants (SNVs) and indels were called for each sample using GATK HaplotypeCaller in GVCF mode, followed by joint genotyping across all maternal and fetal DNA-derived samples and variant filtering with GATK VQSR. To ensure a high-quality set of genotypes for use in benchmarking, we used a large-scale family sequencing project. 32 We further applied a stringent set of variant filters previously used in
[10] . Briefly, variant sites were removed if they overlapped with low-complexity regions of the genome. Variant genotypes meeting any of the following criteria were filtered: depth less than 10; allelic balance <0.25 or >0.75; probability of allelic balance (based on a binomial distribution with an arithmetic mean of 0.5) being less than 1e-9; or fewer than 90 reads being informative for the genotype. The amniocentesis sample from study participant MGB043 was sequenced at Boston Children's Hospital (Boston, MA) using a protocol from GeneDx (Stamford, CT). Sequencing data from this sample was realigned to hg38 and then reprocessed according to the informatics steps described above. For this sample only, benchmarking was limited to the intersection of the exome target region of the Broad Custom Exome kit used for the remaining samples with the GeneDx kit.
[0073] Benchmarking and Evaluation Variants were compared to "true" genotype data derived from gDNA ES from either matched cord blood, amniocentesis, maternal DNA collected from leukocytes, or paternal samples (see section gDNA ES Variant Calling in Maternal, Paternal, Fetal Cord, and Amniocentesis Samples). For comparison of cfDNA variants with cord blood or amniocentesis, five sets of assessments were performed, reported in Tables 5, 6, 8, and 10:
[0074] Site-level comparison of variants not removed by our filtering method (see the "cfDNA Variant Filtering" section), which did not consider fetal genotype at the site (Table 10, "After Filter Variant Detection"). This evaluation provides an assessment of the limits of sensitivity of cfDNA sequencing at the depth used in this study after attempting to remove sequencing artifacts and other errors from the sequencing data. As with the unfiltered variant detection evaluation below, we excluded maternal variants that were not transmitted to the fetus from this evaluation so that the PPV metric demonstrates the method's ability to distinguish errors from true biological mutations.
[0075] Site-level comparison of all variant sites detected in the cfDNA sequencing data (and therefore of either maternal or fetal origin, or both), regardless of the final filter status or fetal genotype assigned to that site by our bioinformatics pipeline (also referred to as "Unfiltered Variant Detection" in Table 10). This evaluation provides an assessment of the theoretical limits to the sensitivity of cfDNA sequencing at the depth used in this study, without attempting to remove sequencing artifacts and other errors. We excluded sites that were present in the mother (according to maternal and umbilical cord blood or amniocentesis gDNA ES data) but not transmitted to the fetus; therefore, FPs in this evaluation are expected to represent true sequencing or mapping errors, as opposed to fetal / maternal genotyping failures.
[0076] Comparison of all fetal genotypes assigned by our model with all genotypes called in cord blood or amniocentesis ES data (Table 5, "Overall Genotyping Performance"). This assessment assesses the accuracy of our genotyping model, which attempts to assign fetal and maternal genotypes to all sites detected in cfDNA (which is a mixture of cfDNA fragments with maternal and fetal origins). See Supplementary, Methods section "cfDNA Genotyping" in the description of the genotyping model. In contrast to the "Unfiltered Mutation Detection" and "Filtered Variant Detection" assessments, untransmitted maternal variants are included. The results of this assessment represent the full capability of our informatics method to determine fetal genotypes at all sites in the exome given only the cfDNA sequencing sample.
[0077] Comparison of all variants assigned a fetal heterozygous genotype to a maternal homozygous reference genotype (Table 6, "Expected paternal or novel variant detection") at sites that were present in the cord blood or amniocentesis gDNA ES data but not in the participant's maternal gDNA ES data. This evaluation excluded variant sites detected in the maternal gDNA ES data and assessed only variants called in the NIFS data that the genotyping pipeline assigned to the "fetal 0 / 1; maternal 0 / 0" cluster. This evaluation characterizes the method's accuracy in detecting paternally inherited variants as well as novel mutations.
[0078] Assessment of our method's ability to accurately genotype variants that are heterozygous in the mother (Table 8, "NIFS Genotype Accuracy for Maternal Variant Heterozygosity"). This evaluation focuses only on sites where maternal gDNA ES data indicate the mother is heterozygous for variants that pass the filter context. These sites are important for recessive disease diagnosis, but are more difficult to genotype due to the low fetal fraction. For this evaluation, we report a single accuracy metric, which is the proportion of truly maternal heterozygous sites that were assigned a passing filter context and the correct fetal genotype by NIFS.
[0079] All of the above assessments, except for "NIFS genotype accuracy in maternal variant heterozygosity," were performed by Real Time Genomics 33,34We used the vcfeval tool in the Realtime Genomics (RTG; realtimegenomics.com / products / rtg-tools) platform, which performs haplotype-based analysis to match variants across samples and is a widely accepted standard for evaluating genomic variant calling. All benchmark analyses were restricted to intervals targeted by exome capture panels on autosomes. Evaluations of "unfiltered variant detection" and "post-filtered variant detection" in comparisons with cord blood and amniocentesis samples were performed by matching sites regardless of the called genotype. In these two evaluations, the presence of the same variant site with a matched genomic location and alternate allele in both the cfDNA sample and the validation data was counted as a true positive (this was achieved using the vcfeval parameter - squared ploidy). Meanwhile, comparisons of "overall genotyping performance" required each called fetal allele in our pipeline's output to match an allele present in the called genotype in the validation data. The benchmark evaluates true positives (TP), FP, true negatives (TN) and false negatives (FN). Sensitivity and PPV were calculated by RTG vcfeval as follows: ●PPV=TP / (TP+FP) ●Sensitivity=TP / (TP+FN)
[0080] For all of the above assessments except for "Genotyping Accuracy for Maternal Heterozygous Variants," we excluded from the assessment all regions with coverage of fewer than 10 reads in ES data from cord blood or amniocentesis samples—in other words, all regions for which the variant site calls in these regions were not counted as TP, FP, or FN. This assessment was performed using the same analysis script as "Genotyping Accuracy by Maternal and Fetal Genotypes," listed in Table 9 below. We note that the metrics presented in this assessment can be calculated by summarizing the results for each of the genotype clusters corresponding to maternal heterozygous variants in Table 9. See also Figures 3A-C.
[0081] In addition to the above analyses, we also used matched cord blood or amniocentesis gDNA ES data for a more detailed breakdown of NIFS sensitivity and genotype accuracy for all confirmed variants in the fetal and maternal exomes (Table 9, "Genotype Accuracy by Maternal and Fetal Genotypes"). For this assessment, we compared all NIFS calls made from cfDNA to the combined set of all variants called in either the maternal gDNA ES or the cord blood / amniocentesis gDNA ES. For each combination of maternal and fetal genotype present in this comparison set of maternal and fetal variants, we calculated the percentage of sites with concordant positions and alternative alleles present in the raw Mutect2 cfDNA VCF (reported in Table 10), as well as the percentage of sites that were unfiltered and assigned the correct fetal genotype by the cfDNA variant calling pipeline. These assessments were performed using custom analysis scripts that matched variant calls in maternal gDNA ES, cord blood or amniocentesis gDNA ES, and cfDNA sequencing data by genomic location and alternative alleles (as opposed to the haplotype-based method implemented in RTG vcfeval).
[0082] The second set of evaluations compared maternal genotypes predicted by our model with variants detected in ES sequencing of maternal gDNA extracted from sedimented maternal leukocytes. The results of this evaluation are reported in two parts, "Maternal Variant Detection" and "Maternal Genotyping Performance," in Table 11. We matched any called variant site for the "Maternal Variant Detection" comparison, regardless of the maternal genotype assigned by NIFS or gDNA variant calling. We required perfect genotype agreement between gDNA and ES calls for the "Maternal Genotyping Performance" evaluation. For these maternal evaluations, we excluded sites for which maternal gDNA ES data had less than 10x read coverage. These evaluations were performed using the RTG vcfeval tool.
[0083] Finally, for participants with matched gDNA ES data derived from paternal blood samples, we assessed the proportion of sites in the cfDNA data that were assigned non-reference fetal genotypes, excluding sites present in the maternal gDNA ES data. The results of this assessment are reported in Table 7, "Genotyping Performance and Coverage Based on Paternal gDNA ES." This assessment is an alternative method for calculating the PPV of NIFS calls predicted to be either paternally inherited or novel mutations, in addition to the "Predicted Paternal or Novel Variant Detection" results reported in Table 6. For this analysis, we used RTG vcfeval to calculate the PPV of all calls assigned to cluster 0 (the cluster representing fetal heterozygous and maternal homozygous reference variants) against the set of paternal ES variants, and restricted the assessment to sites that did not match any variant sites in the maternal ES data. We excluded from this assessment regions where the paternal ES data provided less than 10x coverage. We also report the number of reads supporting the alternative allele for each of these confirmed paternal variants detected by NIFS.
[0084] For sample MGB043, the amniocentesis sample was sequenced at Boston Children's Hospital using a different exome capture kit provided by GeneDx, and therefore we limited all evaluations to the set of exome-targeted intervals covered by both the Twist Custom Exome list and the GeneDx exome targets used for the NIFS sample (n=194, 202 intervals).
[0085] Family Relationship Inference Predicted genetic relationships (between cfDNA, parental and umbilical cord blood and amniocentesis samples) were identified after variant calling. 35 To confirm the suspected familial relationship in our cohort, we filtered the cfDNA variants to include only those with a gnomAD allele frequency (AF_popmax) greater than 0.05 and a fetal genotype inference quality score greater than 10. Processing the resulting predicted genotyped clusters with KING (parameters - association - degree 2) verified the predicted relationships with an estimated proportion of identical genomes by offspring (KING metric PropIBD) of at least 0.4 (with one exception: a paternal-fetal pair with 0.32 propIBD, which we manually confirmed).
[0086] Copy number variant (CNV) detection We developed a sliding-window binning method to investigate significant copy state deviations using coverage collected from GATK CollectReadCounts with GC correction. Copy state was normalized against a subset of control NIFS libraries (no fetal abnormality cases, Table 12) using GATK CreateReadCountPanelOfNormals and DenoiseReadCounts. We excluded highly variable capture intervals in control cfDNA samples with a median absolute deviation (MAD) greater than the third quartile + 1.5 * interquartile range (IQR). We then calculated the median copy ratio for each sample within bins representing a genome-spanning sliding window of 3 MB size with a 100 kb offset. A final filtering step was applied to remove 1 MB bins with more than 10 control samples classified as outliers based on bin-by-bin IQR analysis. Because only one validated CNV event was observed in our study cohort, we were unable to perform extensive benchmarking or sensitivity analysis of CNVs. We followed the methodology of Fu et al. 32 We note that our previous gDNA ES study reported by demonstrated accurate CNV discovery beyond the resolution of individual genes, up to the routine detection of events spanning more than two exons, highlighting the feasibility of discovering CNVs at single-exon resolution. Detection of these events in cfDNA will be challenging due to the mixing of maternal and fetal DNA, but more data will enable the development and thorough benchmarking of improved methods.
[0087] Gender determination We investigated the ability of NIFS to determine fetal sex given the robust coverage of chrY and chrX. We first focused on chrY for gender delineation, considering that any reads on chrY above some artifact should indicate a male fetus. Indeed, the presence of any coverage (from GATK CollectReadCounts) in binned intervals on chrY was highly discriminatory for sex determination (Figure 4). However, accurate prediction of chrY copy state, determined by dividing the median coverage across all intervals chrY by the fetal fraction, remained challenging due to the relatively low and variable coverage on chrY compared to the rest of the genome.
[0088] Variant Classification We analyzed each sample for potentially pathogenic variants in the fetus and mother using genotypes derived from the cfDNA results. 36 Merge was applied to generate multi-sample VCFs of all samples with cffDNA sequencing. 37 and bcftools, this merged VCF was used to generate genetic and functional results (RefSeq 26 ), allele frequency (gnomAD v2.1.1 and gnomAD v3.0), REVEL 38 Score, ClinVar 39 Annotation (updated 2023-04-30), and Online Mendelian Inheritance in Man (OMIM, version 2022-07-08) 40 ) were annotated with gene-specific disease information such as genotype (e.g., recessive) from gnomAD v2.1.1 and gnomAD v3.0. We analyzed alleles with allele frequencies less than 5 or with genotypes less than 5, and then analyzed the data for alleles less than 5. We also ... 28We included variants and excluded synonymous variants if they were not reported in the 2016 WHO Guidelines for Genetic Algorithms (2016). We then generated a list from each sample for further consideration, which included all ClinVar-annotated pathogenic / likely pathogenic variants, all frameshift / stop-gain variants, and splice AI scores. 41 This included all predicted splice variants with a REVEL score >0.95, all non-frameshift variants >15 amino acids, and all non-synonymous variants with a REVEL score >0.7. Variants that did not pass the filter were removed, except for ClinVar P / LP variants. From this set, we filtered variants with fewer than four alternate reads and variants determined as likely_benign or benign / likely_benign in ClinVar.
[0089] We confirmed fetal genotypes using the method described above, with the caveat that for a small subset of indels staged with high-quality SNVs, we used SNV genotypes given the higher SNV genotyping accuracy. We selected disease gene variants from OMIM for further analysis. We manually reviewed each of the remaining variants using the Integrated Genomics Viewer (IGV) and removed variants that appeared to be of low quality or were present in multiple NIFS samples (indicating they were likely technical artifacts). ACMG criteria were used. 42-44 The pathogenicity of variants was examined and clinical relevance was assessed based on the CNVs described by Riggs et al. 45The assessment was performed according to the Clingen and ACMG guidelines provided by [the authors]. We assessed potential carrier variants for 28 samples with concordant maternal germline exome sequencing data (Table 3) and further filtered these based on their ClinVar pathogenicity status. Variants were considered if they were listed as pathogenic or likely pathogenic in ClinVar with clinical significance corresponding to two or more gold stars (i.e., criteria provided by practice guidelines, expert panel review, or multiple submitters with no discrepancies). Variants with genotypes corresponding to maternal carrier status were selected. As before, variants were reviewed for potential clinical relevance (Table 14). All identified variants were confirmed by maternal germline exome sequencing.
[0090] [Table 1]
[0091] [Table 2]
[0092] [Table 3]
[0093] [Table 4]
[0094] [Table 5]
[0095] [Table 6]
[0096] [Table 7]
[0097]
Table 8
[0098]
Table 9-1
Table 9-2
[0099]
Table 10
[0100]
Table 11
[0101]
Table 12
[0102]
Table 13
[0103]
Table 14
[0104] References: 1.Lowther C,Valkanas E,Giordano JL,ら Systematic evaluation of genome sequencing for the assessment of fetal structural anomalies[Donald].bioRxiv.2020;biorxiv.org / content / 10.1101 / 2020.08.12.248526.abstract 2.Talkowski ME,Ordulu Z,Pillalamarri V,ら Clinical diagnosis by whole-genome sequencing of a prenatal sample.N Engl J Med 2012;367(23):2226-32. 3.Tolusso LK,Hazelton P,Wong B,Swarr DT.Beyond diagnostic yield:prenatal exome sequencing results in maternal,neonatal,and family clinical management changes.Genet Med 2021;23(5):909-17. 4.Gregg AR,Skotko BG,Benkendorf JL,ら Noninvasive prenatal screening for fetal aneuploidy,2016 update:a position statement.Genet Med 2016;18(10):1056-65. 5.American College of Obstetricians and Gynecologists’ Committee on Practice Bulletins-Obstetrics,Committee on Genetics,Society for Maternal-Fetal Medicine.Screening for Fetal Chromosomal Abnormalities:ACOG Practice Bulletin,Number 226.Obstet Gynecol 2020;136(4):e48-69.
[0105] 6.Bianchi DW,Parker RL,Wentworth J,ら DNA sequencing versus standard prenatal aneuploidy screening.N Engl J Med 2014;370(9):799-808. 7.Norton ME,Jacobsson B,Swamy GK,ら Cell-free DNA analysis for noninvasive examination of trisomy.N Engl J Med 2015;372(17):1589-97. 8.Yatsenko SA,Peters DG,Saller DN,Chu T,Clemens M,Rajkovic A.Maternal cell-free DNA-based screening for fetal microdeletion and the importance of careful diagnostic follow-up.Genet Med 2015;17(10):836-8. 9.Zhang J,Li J,Saucier JB,ら Non-invasive prenatal sequencing for multiple Mendelian monogenic disorders using circulating cell-free fetal DNA.Nat Med 2019;25(3):439-47. 10.Breveglieri G,D’Aversa E,Finotti A,Borgatti M.Non-invasive Prenatal Testing Using Fetal DNA.Mol Diagn Ther 2019;23(2):291-9.
[0106] 11.Dungan JS,Klugman S,Darilek S,ら Noninvasive prenatal screening(NIPS)for fetal chromosome abnormalities in a general-risk population:An evidence-based clinical guideline of the American College of Medical Genetics and Genomics(ACMG).Genet Med 2023;25(2):100336. 12.Rose NC,Barrie ES,Malinowski J,ら Systematic evidence-based review:The application of noninvasive prenatal screening using cell-free DNA in general-risk pregnancies.Genet Med 2022;24(7):1379-91. 13.Fan HC,Gu W,Wang J,Blumenfeld YJ,El-Sayed YY,Quake SR.Non-invasive prenatal measurement of the fetal genome.Nature 2012;487(7407):320-4. 14.Provenzano A,Farina A,Seidenari A,ら Prenatal Noninvasive Trio-WES in a Case of Pregnancy-Related Liver Disorder.Diagnostics(Basel)[Basel]2021;11(10).Available from:dx.doi.org / 10.3390 / diagnostics11101904 15.Filer DL,Mieczkowski PA,Brandt A,ら The diagnostic utility of noninvasive prenatal exome sequencing is limited by sequencing depth and fetal fraction.Prenat [ PubMed ] 2021;dx.doi.org / 10.1002 / pd.6009
[0107] 16.Guo MH,Gregg AR.Estimating yields of prenatal carrier screening and implications for design of expanded carrier screening panels.Genet Med 2019;21(9):1940-7. 17.Ben-Shachar R,Svenson A,Goldberg JD,Muzzey DA Data-driven evaluation of the size and content of expanded carrier screening panels.Genet Med 2019;21(9):1931-9. 18.Saunders CJ,Miller NA,Soden SE,らRapid whole-genome sequencing for genetic disease diagnosis in neonatal intensive care units.Sci Transl Med 2012;4(154):154ra135. 19.Balloux F,Bronstad Brynildsrud O,van Dorp L,ら From Theory to Practice:Translating Whole-Genome Sequencing(WGS)into the Clinic.Trends Microbiol 2018;26(12):1035-48. 20.Miller DT,Lee K,Abul-Husn NS,ら ACMG SF v3.1 list for reporting of secondary findings in clinical exome and genome sequencing:A policy statement of the American College of Medical Genetics and Genomics(ACMG).Genet Med 2022;24(7):1407-14.
[0108] 21.Monaghan KG,Leach NT,Pekarek D,Prasad P,Rose NC,ACMG Professional Practice and Guidelines Committee.The use of fetal exome sequencing in prenatal diagnosis:a points to consider document of the American College of Medical Genetics and Genomics(ACMG).Genet Med 2020;22(4):675-80. 22. Van den Veyver IB, Chandler N, Wilkins-Haug LE, Wapner RJ, Chitty LS, ISPD Board of Directors. International Society for Prenatal Diagnosis Updated Position Statement on the use of genome-wide sequencing for prenatal diagnosis. Prenat Diagn 2022;42(6):796-803. 23. McElrath TF, Lim KH, Pare E, et al. Longitudinal evaluation of predictive value for preeclampsia of circulating angiogenic factors through pregnancy. Am J Obstet Gynecol 2012;207(5):407.e1-7. 24. Li H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM [Internet]. arXiv [q-bio.GN]. 2013; available at arxiv.org / abs / 1303.3997
[0109] 25. Poplin R, Ruano-Rubio V, DePristo MA, et al. Scaling accurate genetic variant discovery to tens of thousands of samples [Internet]. bioRxiv.2018 [cited 2019 Nov 21]; available at: 201178.biorxiv.org / content / 10.1101 / 201178v3.abstract 26. O’Leary NA, Wright MW, Brister JR, et al. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Res 2016;44(D1):D733-45. 27. Benjamin D, Sato T, Cibulskis K, Getz G, Stewart C, Lichtenstein L. Calling Somatic SNVs and Indels with Mutect2 [Internet]. bioRxiv. 2019 [cited 2022 Apr 12];861054. Available from biorxiv.org / content / 10.1101 / 861054v1 28. Karczewski KJ, Francioli LC, Tiao G, et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 2020;581(7809):434-43. 29. Bekker J, Davis J. Learning from positive and unlabeled data: a survey. Mach Learn 2020;109(4):719-60. 30. Li H. Toward better understanding of artifacts in variant calling from high-coverage samples. Bioinformatics 2014;30(20):2843-51.
[0110] 31. Bingham E, Chen JP, Jankowiak M, et al. Pyro: Deep Universal Probabilistic Programming. J Mach Learn Res 2019;20(28):1-6. 32. Fu JM, Satterstrom FK, Peng M, et al. Rare coding variation provides insight into the genetic architecture and phenotypic context of autism. Nat Genet 2022;54(9):1320-31. 33.Cleary JG, Braithwaite R, Gaastra K, et al. Joint variant and de novo mutation identification on pedigrees from high-throughput sequencing data.J Comput Biol 2014;21(6):405-19. 34. Cleary JG, Braithwaite R, Gaastra K, et al. Comparing Variant Call Files for Performance Benchmarking of Next-Generation Sequencing Variant Calling Pipelines [Internet].bioRxiv.2015 [cited 2023 Jun 15]; available at: 023754.biorxiv.org / content / 10.1101 / 023754v2 35. Manichaikul A, Mychaleckyj JC, Rich SS, Daly K, Sale M, Chen WM. Robust relationship inference in genome-wide association studies. Bioinformatics 2010;26(22):2867-73.
[0111] 36. Danecek P, Bonfield JK, Liddle J, et al. Twelve years of SAMtools and BCFtools. Gigascience [Internet] 2021;10(2). Available at: dx.doi.org / 10.1093 / gigascience / giab008 37. Wang K, Li M, Hakonarson H. ANNOVAR: functional annotation of genetic variants from high - throughput sequencing data. Nucleic Acids Res 2010;38(16):e164. 38. Ioannidis NM, Rothstein JH, Pejaver V, et al. REVEL: An Ensemble Method for Predicting the Pathogenicity of Rare Missense Variants. Am J Hum Genet 2016;99(4):877 - 85. 39. Landrum MJ, Lee JM, Benson M, et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res 2018;46(D1):D1062 - 7. 40. Amberger JS, Bocchini CA, Schiettecatte F, Scott AF, Hamosh A. Omim.org: Online Mendelian Inheritance in Man (OMIM®), an online catalog of human genes and genetic disorders. Nucleic Acids Res 2015;43(Database issue):D789 - 98.
[0112] 41. Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, et al. Predicting Splicing from Primary Sequence with Deep Learning. Cell 2019;176(3):535 - 548.e24. 42.Gregg AR,Aarabi M,Klugman S,ら Screening for autosomal recessive and X-linked conditions during pregnancy and preconception:a practice resource of the American College of Medical Genetics and Genomics(ACMG).Genet Med 2021;23(10):1793-806. 43.Harrison SM,Biesecker LG,Rehm HL.Overview of Specifications to the ACMG / AMP Variant Interpretation Guidelines.Curr Protoc Hum Genet 2019;103(1):e93. 44.Richards S,Aziz N,Bale S,ら Standards and guidelines for the interpretation of sequence variants:a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology.Genet Med 2015;17(5):405-24. 45.Riggs ER,Andersen EF,Cherry AM,ら Technical standards for the interpretation and reporting of constitutional copy-number variants:a joint consensus recommendation of the American College of Medical Genetics and Genomics(ACMG)and the Clinical Genome Resource(ClinGen).Genet Med 2020;22(2):245-57.
[0113] Other embodiments While the present invention has been described in conjunction with its detailed description, it should be understood that the foregoing description is intended to be illustrative, and not limiting, of the scope of the invention, which is defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
Claims
1. 1. A computer-implemented method for assigning maternal or fetal origin to one or more genetic variants in cell-free DNA (cfDNA) from a sample derived from a pregnant mammal, preferably a pregnant human, said method comprising: (a) accessing, from memory, a probabilistic model for assigning maternal or fetal origin to genetic variants in DNA from a sample obtained from a pregnant mammal, wherein the model assigns maternal or fetal origin based on fetal fraction and / or DNA fragment size in combination with other sequencing characteristics; (b) inputting into the model a set of values representing one or more genetic variants detected in the cfDNA from a peripheral blood sample from a pregnant mammal, the values including, for each genetic variant, empirically determined sequence information, e.g., the ratio of different bases in reads, and DNA fragment size information, e.g., a rank sum statistic; (c) using the model to assign a maternal or fetal origin of the one or more genetic variants; and 20. A computer-implemented method comprising:
2. 2. The method of claim 1, wherein the genetic variant comprises a single nucleotide variant (SNV), an indel, and / or a copy number variation (CNV).
3. an initial set of values representing the one or more genetic variants, Aligning raw sequencing reads derived from the cfDNA to a reference genome sequence; converting the raw sequencing reads into consensus reads; realigning the consensus reads to the reference genome sequence, thereby generating a set of aligned consensus reads; identifying consensus reads that differ from the reference genome; assigning consensus reads that differ from the reference genome as alternative alleles and consensus reads that match the reference genome as reference alleles; determining a fragment size rank sum statistic representing the distribution of estimated fragment sizes of reads supporting the reference allele compared to the distribution of fragment sizes of reads supporting the alternative allele; obtained by a process comprising thereby obtaining an initial set of values representing sequence identity and DNA fragment size rank sum statistics for one or more genetic variants; The method of claim 1.
4. 4. The method of claim 3, wherein each of the raw sequencing reads comprises a unique molecular identifier (UMI), and the method comprises converting the raw sequencing reads into a single consensus read for each UMI.
5. Before step (b), accessing from memory a machine learning classifier, optionally a random forest-based model, that is trained using a set of predetermined filter criteria and a subset of sites present in the sample or reference sample to identify potential false positive (FP) sites; inputting the initial set of variants into the machine learning classifier; using the trained machine learning classifier to filter out a set of variants enriched in false positive (FP) sites, thereby selecting a set of candidate variants from the initial set; 2. The method of claim 1, further comprising selecting a set of candidate variants by a method comprising:
6. 2. The method of claim 1, wherein the probabilistic model is a Bayesian mixture model that simultaneously estimates fetal fraction and assigns fetal or maternal origin for each variant site in the set.
7. 7. The method of claim 6, wherein the Bayesian mixture model is a Bayesian Gaussian mixture model constrained on variant allele fractions and fragment size rank sum statistics.
8. 7. The method of claim 6, wherein the fetal fraction of the sample is modeled as a latent variable (f) and the arithmetic mean of the variant allele fraction distribution is set for each component based on f.
9. 7. The method of claim 6, wherein the fetal fraction is estimated based on a reference fetal fraction determined based on clusters derived from VAF across sites.
10. 10. The method of claims 1-9, further comprising outputting a list of one or more genetic variants identified as having fetal origin and / or one or more genetic variants identified as having maternal origin.
11. comparing the genetic variant to a database containing a list of genetic variants and information about variants of potential medical relevance to the fetus or the mother; identifying potentially medically relevant variants present in the fetus or the mother; outputting a list of the one or more genetic variants identified as having fetal origin and / or one or more genetic variants identified as having potentially medically relevant maternal origin; The method of claims 1 to 9, further comprising:
12. 12. The method of claim 11, further comprising: the method further comprising recommending further testing based on the presence of a potentially medically relevant variant.
13. 13. The method of claim 12, wherein the further testing comprises amniocentesis or chorionic villus sampling (CVS), further monitoring of the fetus by ultrasound, or genetic testing of the mother.
14. 14. The method of claims 1-13, further comprising using high-throughput sequencing on cfDNA extracted from a single sample of peripheral blood from the mother, optionally wherein exome capture is performed prior to said sequencing.
15. 15. The method of claim 14, wherein an adapter having a common PCR primer sequence and a unique molecular identifier (UMI) is attached to the cfDNA, and PCR amplification is performed prior to the sequencing.
16. 15. The method of claim 14, further comprising enriching the sample for fetal DNA, optionally including fetal protein-coding genes or other regions of the fetal genome that may be relevant to clinical interpretation or variant identification, by contacting the cfDNA with a plurality of oligonucleotides that bind to portions of the fetal genome.