Methods and compositions for single-gene non-invasive prenatal testings
The method enhances sgNIPT by extracting and enriching cell-free DNA from maternal plasma to detect pathogenic variants in target genes, addressing the limitations of current tests and achieving high sensitivity and specificity in identifying fetal genotypes for autosomal recessive disorders.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NATERA INC
- Filing Date
- 2025-10-14
- Publication Date
- 2026-04-23
AI Technical Summary
Current single-gene non-invasive prenatal tests (sgNIPT) lack sufficient sensitivity and specificity to detect pathogenic variants associated with autosomal recessive disorders, particularly compound and homozygous recessive mutants, without requiring a paternal sample.
A method involving the extraction of cell-free DNA from maternal plasma, targeted enrichment of target variant loci, high-throughput sequencing, and determination of fetal genotypes using hybrid capture probes and sequencing to identify pathogenic variants in target genes associated with autosomal recessive disorders.
Enables the detection of most pathogenic variants associated with autosomal recessive disorders with high sensitivity and specificity, including compound and homozygous recessive mutants, from fetal cell-free DNA in maternal plasma, without needing a paternal sample.
Smart Images

Figure IMGF000041_0001 
Figure IMGF000042_0001 
Figure IMGF000043_0001
Abstract
Description
Atty. Dkt. No.: N.059.W0.01METHODS AND COMPOSITIONS FOR SINGLE-GENE NON-INVASIVE PRENATAL TESTINGSCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefits of U.S. Provisional Application No. 63 / 707,085, filed October 14, 2024, and U.S. Provisional Application No. 63 / 855,743, filed August 1, 2025, the contents of which are hereby incorporated by reference in their entirety.BACKGROUND
[0002] Non-invasive prenatal testing (NIPT) is being widely used to screen for chromosome abnormalities in a fetus, such as Down syndrome (trisomy 21), Patau syndrome (trisomy 13), and Edwards syndrome (trisomy 18). As a blood test, NIPT is generally less risky than invasive procedures such as amniocentesis (e.g., a biopsy), and therefore generally more preferable for pregnant people and clinicians.
[0003] The American College of Obstetricians and Gynecologists (ACOG) recommends carrier screening for certain autosomal recessive diseases, such as cystic fibrosis (CF), spinal muscular atrophy (SMA), alpha thalassemia, beta thalassemia, and sickle cell disease. Although single-gene NIPT (sgNIPT) tests have been considered for supplementation of carrier screening and confirmation of fetal genotypes, currently available sgNIPT tests are not able to detect pathogenic variants associated with autosomal recessive disorders with sufficient coverage, sensitivity and specificity.SUMMARY OF THE DISCLOSURE
[0004] Provided here are methods, compositions, and systems for non-invasive prenatal testing (sgNIPT) that can detect most if not all pathogenic variants associated with common autosomal recessive disorders, including not only compound heterozygous mutants but also homozygous recessive mutants, from fetal cell-free DNA (cfDNA) present in the maternal plasma, with high sensitivity and high specificity, and without having to acquire and test a paternal sample.Atty. Dkt. No.: N.059.W0.01
[0005] In one aspect, the disclosure herein relates to a method for preparing a non- naturally occurring composition, comprising: extracting cell-free DNA from a plasma fraction of a blood sample of a pregnant person, wherein the pregnant person is a heterozygous carrier of at least one pathogenic variant in at least one target gene associated with autosomal recessive disorders, wherein the extracted cell-free DNA comprises a mixture of maternal cell-free DNA and fetal cell-free DNA; performing targeted enrichment on the extracted cell-free DNA or DNA derived therefrom to enrich a plurality of target variant loci in a plurality of target genes and generating enriched DNA, wherein the target variant loci comprise the at least one pathogenic variant; performing high-throughput sequencing on the enriched DNA or DNA derived thereof and generating sequence reads, and determining fetal genotypes of one or more of the target genes from the sequence reads.
[0006] In another aspect, the disclosure herein relates to a method for preparing a non- naturally occurring composition, comprising: genotyping cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof to identify the pregnant person as a heterozygous carrier of at least one pathogenic variant in at least one target gene associated with autosomal recessive disorders; extracting cell-free DNA from a plasma fraction of a blood sample of the pregnant person, wherein the extracted cell-free DNA comprises a mixture of maternal cell-free DNA and fetal cell-free DNA; performing targeted enrichment on the extracted cell-free DNA or DNA derived therefrom to enrich a plurality of target variant loci in a plurality of target genes and generating enriched DNA, wherein the target variant loci comprise the at least one pathogenic variant; and performing high-throughput sequencing on the enriched DNA or DNA derived thereof and generating sequence reads, and determining fetal genotypes of one or more of the target genes from the sequence reads.
[0007] In some embodiments, the at least one pathogenic variant comprises a single nucleotide variant (SNV), an indel, a copy number variation (CNV), a gene fusion, a chromosomal rearrangement (e.g„ deletion, duplication, inversion, or translocation), or a combination thereof.
[0008] In some embodiments, the target genes comprise one or more of CFTR, HBA1, HBA2, HBB, or SMN1. In some embodiments, the autosomal recessive disorders comprise oneAtty. Dkt. No.: N.059.W0.01 or more of cystic fibrosis, alpha thalassemia, beta thalassemia, sickle cell disease, or spinal muscular atrophy.
[0009] In some embodiments, the target genes comprise one or more of MEFV, GBA, PAH, HEXA, ATP7B, DHCR7, IKBKAP, CPT2, ASPA, GALC, AC ADM, GAA. PKHD1, or GALT. In some embodiments, the autosomal recessive disorders comprise one or more of familial mediterranean fever, Gaucher disease, phenylalanine hydroxylase deficiency, hexosaminidase A deficiency (Tay-Sachs disease), Wilson disease, Smith-Lemil-Optiz syndrome, familial dysautonomia, carnitine palmitoyltransferase II deficiency, Canavan disease, Krabbe disease, medium chain acyl-CoA dehydrogenase deficiency, Pompe disease, autosomal recessive polycystic kidney disease, or galactosemia.
[0010] In some embodiments, the method further comprises performing long-read sequencing on cellular DNA extracted from a blood sample of the pregnant person or a buffy- coat fraction thereof and generating phased haplotypes of the pregnant person at one or more of the target genes.
[0011] In some embodiments, the targeted enrichment further enriches a plurality of phasing SNP loci within and / or flanking one or more of the target genes, wherein the plurality of phasing SNP loci each has a minor allele frequency (MAF) of at least 1 %, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, or at least 10%.
[0012] In some embodiments, the plurality of phasing SNP loci are located up to 20,000 to 500,000 bases from at least one of the target variant loci.
[0013] In some embodiments, the plurality of phasing SNP loci comprise at least 100, at least 200, or at least 400 phasing SNP loci per target gene. At an average MAF of 5%, this translates to at least 5, at least 10, or at least 20 informative heterozygous SNPs per haplotype.
[0014] In some embodiments, the method comprises using (i) the phased haplotypes of the pregnant person at one or more of the target genes (e.g., phased haplotypes generated from long-read sequencing of cellular DNA or derivative thereof, with the cellular DNA being extracted from a buffy-coat fraction of a blood sample of the pregnant person), and (ii) theAtty. Dkt. No.: N.059.W0.01 sequence reads at phasing SNP loci that are heterozygous in the pregnant person (e.g., sequence reads generated from sequencing of cell-free DNA or derivative thereof, with the cell-free DNA being extracted from a plasma fraction of the blood sample), to determine the most likely fetal genotype and / or the most likely haplotype the fetus inherited from the pregnant person.
[0015] In some embodiments, the method further comprises identifying deletion in one or more of the target genes using sequence reads at the plurality of phasing SNP loci.
[0016] In some embodiments, the method further comprises quantifying fetal fraction in the plasma fraction using the sequence reads.
[0017] In some embodiments, the method further comprises appending an adapter comprising a universal primer binding site to the extracted cell-free DNA or DNA derived therefrom and generating adapted DNA prior to the targeted enrichment.
[0018] In some embodiments, the adapter further comprises a molecular index / barcode sequence, and wherein sequence reads generated from the high-throughput sequencing are grouped together using the molecular index / barcode sequence.
[0019] In some embodiments, the adapter does not comprise a molecular index / barcode sequence, and wherein sequence reads generated from the high-throughput sequencing are grouped together using the fragment-end sequences of the extracted cell-free DNA or DNA derived therefrom.
[0020] In some embodiments, the method further comprises amplifying the adapted DNA using a primer that binds to the universal primer binding site and generating adapted- amplified DNA prior to the targeted enrichment.
[0021] In some embodiments, the adapted- amplified DNA further comprises a sample index / barcode sequence and / or a sequencing primer binding site.
[0022] In some embodiments, the targeted enrichment comprises preforming targeted multiplex amplification or linked target capture to enrich the target loci. In some embodiments, the targeted enrichment comprises preforming targeted probe capture to enrich the target loci.Atty. Dkt. No.: N.059.W0.01
[0023] In some embodiments, the targeted probe capture is performed using probes covering the entire exons of one or more of the target genes, and / or wherein the targeted probe capture is performed using probes covering the entire exons and introns of one or more of the target genes. In some embodiments, the targeted probe capture is performed using probes covering the entire exons of one or more of the target genes that carry a pathogenic variant, and / or wherein the targeted probe capture is performed using probes covering the entire exons and introns of one or more of the target genes that carry a pathogenic variant.
[0024] In some embodiments, the targeted probe capture is performed using probes covering the entire exons of CFTR. In some embodiments, the targeted probe capture is performed using probes covering the entire exons of CFTR that carry a pathogenic variant.
[0025] In some embodiments, the targeted probe capture is performed using tiling probes providing at least 2X, at least 3X, or at least 4X tiled coverage of the entire exons of one or more of the target genes with or without a pathogenic variant, and / or wherein the targeted probe capture is performed using tiling probes providing at least 2X, at least 3X, or at least 4X tiled coverage of the entire exons and introns of one or more of the target genes with or without a pathogenic variant. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 2X, at least 3X, or at least 4X tiled coverage of the entire exons of one or more of the target genes that carry a pathogenic variant, and / or wherein the targeted probe capture is performed using tiling probes providing at least 2X, at least 3X, or at least 4X tiled coverage of the entire exons and introns of one or more of the target genes that carry a pathogenic variant.
[0026] In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons of CFTR. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons of CFTR that carry a pathogenic variant.
[0027] In some embodiments, the targeted probe capture is performed using hybrid capture probes of 80-200 nucleotides in length, 100-150 nucleotides in length, or 110-130 nucleotides in length.Atty. Dkt. No.: N.059.W0.01
[0028] In some embodiments, the high-throughput sequencing is performed with a median depth of read of at least 30,000, at least 40,000, at least 50,000, at least 60,000, at least 70,000, 30.000-150,000, or 40,000-100,000 per target variant locus.
[0029] In some embodiments, the method comprises determining fetal genotypes of one or more target variant loci at which the pregnant person is a heterozygous carrier of a pathogenic variant.
[0030] In some embodiments, the method comprises determining fetal genotypes of one or more target variant loci at which the pregnant person is a homozygous wildtype, wherein the pregnant person is a heterozygous earner of a pathogenic variant at a different locus of the same target gene.
[0031] In some embodiments, the method does not comprise adding a known quantity of an internal control nucleic acid to the extracted cell-free DNA or DNA derived thereof prior to performing high-throughput sequencing.
[0032] In a further aspect, the disclosure herein relates to a composition comprising a panel of hybrid capture probes that (i) cover the entire exons of one or more target genes associated with autosomal recessive disorders (e.g., CFTR, SMN1), (ii) cover the entire exons and introns of one or more target genes associated with autosomal recessive disorders (e.g., HBB, HBA1, HBA2), and / or (iii) cover a plurality of common SNP loci (e.g., MAF>=5%) within or flanking one or more of the target genes.
[0033] In a further aspect, the disclosure herein relates to a composition comprising a panel of hybrid capture probes that (i) cover the entire exons of one or more target genes associated with autosomal recessive disorders (e.g.. CFTR, SMN1) that carry a pathogenic variant, (ii) cover the entire exons and introns of one or more target genes associated with autosomal recessive disorders (e.g., HBB, HBA1, HBA2) that carry a pathogenic variant, and / or (iii) cover a plurality of common SNP loci (e.g., MAF>=5%) within or flanking one or more of the target genes.Atty. Dkt. No.: N.059.W0.01
[0034] In an additional aspect, the disclosure herein relates to a system comprising (a) a panel of hybrid capture probes for performing target enrichment on extracted cell-free DNA or DNA derived therefrom, wherein the extracted cell-free DNA is extracted from a plasma fraction of a blood sample of a pregnant person and comprises a mixture of maternal cell-free DNA and fetal cell-free DNA, wherein the pregnant person is a heterozygous carrier of at least one pathogenic variant in at least one target gene associated with autosomal recessive disorders, wherein the panel of hybrid capture probes (i) cover the entire exons of one or more target genes associated with autosomal recessive disorders (e.g., CFTR, SMN1), (ii) cover the entire exons and introns of one or more target genes associated with autosomal recessive disorders (e.g., HBB, HBA1, HBA2), and / or (iii) cover a plurality of common SNP loci (e.g., MAF>=5%) within or flanking one or more of the target genes; (b) a sequencer for performing high- throughput sequencing on the enriched DNA or DNA derived thereof and generating sequence reads; and (c) a computer processor or a cloud-based analysis pipeline for determining fetal genotypes of one or more of the target genes from the sequence reads.
[0035] In an additional aspect, the disclosure herein relates to a system comprising (a) a panel of hybrid capture probes for performing target enrichment on extracted cell-free DNA or DNA derived therefrom, wherein the extracted cell-free DNA is extracted from a plasma fraction of a blood sample of a pregnant person and comprises a mixture of maternal cell-free DNA and fetal cell-free DNA, wherein the pregnant person is a heterozygous carrier of at least one pathogenic variant in at least one target gene associated with autosomal recessive disorders, wherein the panel of hybrid capture probes (i) cover the entire exons of one or more target genes associated with autosomal recessive disorders (e.g., CFTR, SMN1) that carry a pathogenic variant, (ii) cover the entire exons and introns of one or more target genes associated with autosomal recessive disorders (e.g., HBB, HBA1, HBA2) that cany a pathogenic variant, and / or (iii) cover a plurality of common SNP loci (e.g., MAF>=5%) within or flanking one or more of the target genes; (b) a sequencer for performing high-throughput sequencing on the enriched DNA or DNA derived thereof and generating sequence reads; and (c) a computer processor or a cloud-based analysis pipeline for determining fetal genotypes of one or more of the target genes from the sequence reads.Atty. Dkt. No.: N.059.W0.01BRIEF DESCRIPTION OF THE DRAWINGS
[0036] FIG. 1A-1F show exemplary embodiments of the methods described herein. Fig. 1A shows one embodiment of the sgNIPT assay workflow described herein, which uses hybrid capture to enrich target loci from cfDNA library. Fig. IB shows another embodiment of the sgNIPT assay workflow described herein, which uses multiplex PCR amplification to enrich target loci from a cfDNA library. Fig. 1C shows a further embodiment of the sgNIPT assay workflow described herein, which uses Linked Target Capture (LTC) to enrich target loci from a cfDNA library. Fig. ID shows an additional embodiment of the sgNIPT assay workflow described herein, which using maternal phasing (e.g., by long-read sequencing of cellular DNA extracted from a buffy-coat fraction of a maternal blood sample) for improving fetal genotype calls (e.g., using phased maternal haplotypes in combination with sequence reads of target variant loci and common / phasing SNP loci).
[0037] FIG. 2 shows an exemplary assay workflow for validation of the sgNIPT methods described herein. Briefly, cell line samples mimicking cfDNA extracted from maternal plasma are subject to library preparation to generate a DNA library, which include the steps of adapter ligation (adding universal priming site and / or molecular index / barcode sequence) followed by library amplification (adding sequencing adapter and / or sample index / barcode sequence). The DNA library is then subject to target enrichment by the 19-gene panel of hybrid capture probes described in Example 1 to generate a targeted sequencing library, wherein the panel of hybrid capture probes allows 4X tiling of variant target loci and includes phasing SNPs for all genes. The targeted sequencing library is subject to high-throughput sequencing on NovaSeq S4 with >60,000 base coverage and ~40 samples per run.
[0038] FIG. 3A-3B show exemplary embodiments of the methods described herein. Fig. 3A shows one embodiment of the sgNIPT calling workflow described herein, which involves (a) calling of fetal genotypes at one or more maternal homozygous wildtype location(s) of a target gene, wherein the pregnant person is a heterozygous carrier of a pathogenic variant at one or more different location(s) of said target gene, as well as (b) calling of fetal genotypes at one or more maternal carrier / heterozygous location(s), wherein the pregnant person is a heterozygous carrier of a pathogenic variant at said heterozygous location(s). Fig. 3B shows oneAtty. Dkt. No.: N.059.W0.01 embodiment of the sgNTPT calling workflow described herein, which involves (a) calling of fetal genotypes at one or more maternal homozygous wildtype location(s) of a target gene, wherein the pregnant person is a heterozygous carrier of a pathogenic variant at one or more different location(s) of said target gene, as well as (b) calling of fetal genotypes at one or more maternal carrier / heterozygous location(s) at high proportion pathogenic loci, wherein the pregnant person is a heterozygous earner of a pathogenic variant at said heterozygous location(s).
[0039] FIG. 4 shows exemplary embodiments of the sgNIPT assay workflow described herein, which uses Linked Target Capture (LTC) and Probe-Dependent Primers (PDPs) to enrich target variant loci from cfDNA library.
[0040] FIG. 5 shows exemplary embodiments of remapping and normalization for highly accurate calling of SMN1 deletions.
[0041] FIG. 6 shows exemplary embodiments of the sgNIPT assay workflow described herein, which uses Linked Target Capture (LTC) and Probe-Dependent Primers (PDPs) to enrich target variant loci from cfDNA library.
[0042] FIG. 7 shows assay results of F508del mixtures with earner mother, including (A) LTC data and (B) hybrid capture data. The allele fraction for mixtures using a parent heterozygous for F508del (GM07553 and GM07826). Cell line GM07826 is homozygous for F508del, GM08334 and GM07552 are heterozygous at the site, and GM07827 is homozygous wild type. Significant separation of the three categories (affected, het, and WT) can be seen starting at 4%.
[0043] FIG. 8 shows assay results of F508del mixtures with non-carrier mother, including (A) LTC data and (B) hybrid capture data. The allele fraction for mixtures using a parent homozygous WT at the F508del locus (GM07461). Cell line GM07469 is heterozygous at the site, and GM07827 is homozygous wild type. Significant separation of the categories (het, and WT) can be seen starting at 4%.
[0044] FIG. 9 shows assay results of R553x mixtures with carrier mother, including (A) LTC data and (B) hybrid capture data. The allele fraction for mixtures using a parentAtty. Dkt. No.: N.059.W0.01 heterozygous at the R553X locus (GM07461). Cell line GM07469 is heterozygous at the site, and GM07458 is homozygous wild type. Significant separation of the two categories (het, and WT, no cell lines homozygous for the mutation were available) can be seen at 4%.
[0045] FIG. 10 shows assay results of R553x mixtures with non-carrier mother, including (A) LTC data and (B) hybrid capture data. The allele fraction for mixtures using a parent homozygous WT at the R553x locus (GM07553). Cell line GM07552 is heterozygous at the site. Significant separation of the categories (het, and WT) can be seen at 4%.DETAILED DESCRIPTION
[0046] Provided here are methods, compositions, and systems for non-invasive prenatal testing (sgNIPT) that can detect most if not all pathogenic variants associated with common autosomal recessive disorders, including not only compound heterozygous mutants but also homozygous recessive mutants, from fetal cell-free DNA (cfDNA) present in the maternal plasma, with high sensitivity and high specificity, and without having to acquire and test a paternal sample.
[0047] Provided herein in one aspect is a method for preparing a non-naturally occurring composition, comprising: extracting cell-free DNA from a plasma fraction of a blood sample of a pregnant person, wherein the pregnant person is a heterozygous carrier of at least one pathogenic variant in at least one target gene associated with autosomal recessive disorders, wherein the extracted cell-free DNA comprises a mixture of maternal cell-free DNA and fetal cell-free DNA; performing targeted enrichment on the extracted cell-free DNA or DNA derived therefrom to enrich a plurality of target variant loci in a plurality of target genes and generating enriched DNA, wherein the target variant loci comprise the at least one pathogenic variant; performing high-throughput sequencing on the enriched DNA or DNA derived thereof and generating sequence reads, and determining fetal genotypes of one or more of the target genes from the sequence reads.
[0048] Provided herein in another aspect is a method for preparing a non-naturally occurring composition, comprising: genotyping cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof to identify the pregnant person as aAtty. Dkt. No.: N.059.W0.01 heterozygous carrier of at least one pathogenic variant in at least one target gene associated with autosomal recessive disorders; extracting cell-free DNA from a plasma fraction of a blood sample of the pregnant person, wherein the extracted cell-free DNA comprises a mixture of maternal cell-free DNA and fetal cell-free DNA; performing targeted enrichment on the extracted cell-free DNA or DNA derived therefrom to enrich a plurality of target variant loci in a plurality of target genes and generating enriched DNA, wherein the target variant loci comprise the at least one pathogenic variant; and performing high-throughput sequencing on the enriched DNA or DNA derived thereof and generating sequence reads, and determining fetal genotypes of one or more of the target genes from the sequence reads.
[0049] Provided herein in a further aspect is a composition comprising a panel of hybrid capture probes that (i) cover the entire exons of one or more target genes associated with autosomal recessive disorders (e.g., CFTR, SMN1), (ii) cover the entire exons and introns of one or more target genes associated with autosomal recessive disorders (e.g., HBB, HBA1, HBA2), and / or (iii) cover a plurality of common SNP loci (e.g., MAF>=5%) within or flanking one or more of the target genes (e.g., distance to closest target variant loci up to 20kbp-500kbp).
[0050] Provided herein in an additional aspect is a system comprising (a) a panel of hybrid capture probes for performing target enrichment on extracted cell-free DNA or DNA derived therefrom, wherein the extracted cell-free DNA is extracted from a plasma fraction of a blood sample of a pregnant person and comprises a mixture of maternal cell-free DNA and fetal cell-free DNA, wherein the pregnant person is a heterozygous carrier of at least one pathogenic variant in at least one target gene associated with autosomal recessive disorders, wherein the panel of hybrid capture probes (i) cover the entire exons of one or more target genes associated with autosomal recessive disorders (e.g.. CFTR, SMN1). (ii) cover the entire exons and introns of one or more target genes associated with autosomal recessive disorders (e.g., HBB, HBA1, HBA2), and / or (iii) cover a plurality of common SNP loci (e.g., MAF>=5%) within or flanking one or more of the target genes (e.g., distance to closest target variant loci up to 20kbp-500kbp); (b) a sequencer for performing high-throughput sequencing on the enriched DNA or DNA derived thereof and generating sequence reads; and (c) a computer processor for determining fetal genotypes of one or more of the target genes from the sequence reads.Atty. Dkt. No.: N.059.W0.01
[0051] Carrier Screening
[0052] In some embodiments, carrier screening is performed on a pregnant person to detect the presence or absence or at least one pathogenic variant in a target gene associated with an autosomal recessive disorder. In some embodiments, the carrier screen comprises Horizon™ genetic carrier screening test.
[0053] In some embodiments, carrier screening is performed on a pregnant person to detect the presence or absence or at least one pathogenic variant in a target gene associated with one or more of cystic fibrosis, alpha thalassemia, beta thalassemia, sickle cell disease, or spinal muscular atrophy. In some embodiments, carrier screening is performed on a pregnant person to detect the presence or absence or at least one pathogenic variant in one or more of CFTR, HBA1, HBA2, HBB, or SMNl.
[0054] In some embodiments, carrier screening is performed on a pregnant person to detect the presence or absence or at least one pathogenic variant in a target gene associated with one or more of familial mediterranean fever, Gaucher disease, phenylalanine hydroxylase deficiency, hexosaminidase A deficiency (Tay-Sachs disease), Wilson disease, Smith-Lemil- Optiz syndrome, familial dysautonomia, carnitine palmitoyltransferase II deficiency, Canavan disease, Krabbe disease, medium chain acyl-CoA dehydrogenase deficiency, Pompe disease, autosomal recessive polycystic kidney disease, or galactosemia. In some embodiments, carrier screening is performed on a pregnant person to detect the presence or absence or at least one pathogenic variant in one or more of MEFV, GBA, PAH, HEXA, ATP7B, DHCR7, IKBKAP, CPT2, ASPA, GALC, ACADM, GAA, PKHD1, or GALT.
[0055] In some embodiments, the method described herein comprises genotyping cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof to identify the pregnant person as a heterozygous earner of at least one pathogenic variant in a target gene associated with an autosomal recessive disorder.
[0056] In some embodiments, the method described herein comprises genotyping cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof to identify the pregnant person as a heterozygous earner of at least one pathogenic variant in aAtty. Dkt. No.: N.059.W0.01 target gene associated with one or more of cystic fibrosis, alpha thalassemia, beta thalassemia, sickle cell disease, or spinal muscular atrophy. In some embodiments, the method described herein comprises genotyping cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof to identify the pregnant person as a heterozygous carrier of at least one pathogenic variant in one or more of CFTR, HBA1, HBA2, HBB, or SMN1.
[0057] In some embodiments, the method described herein comprises genotyping cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof to identify the pregnant person as a heterozygous earner of at least one pathogenic variant in a target gene associated with one or more of familial mediterranean fever, Gaucher disease, phenylalanine hydroxylase deficiency, hexosaminidase A deficiency (Tay-Sachs disease). Wilson disease, Smith-Lemil-Optiz syndrome, familial dysautonomia, carnitine palmitoyltransferase II deficiency, Canavan disease, Krabbe disease, medium chain acyl-CoA dehydrogenase deficiency, Pompe disease, autosomal recessive polycystic kidney disease, or galactosemia. In some embodiments, the method described herein comprises genotyping cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof to identify the pregnant person as a heterozygous earner of at least one pathogenic variant in one or more of MEFV, GBA, PAH, HEXA, ATP7B, DHCR7, IKBKAP, CPT2, ASPA, GALC, AC ADM, GAA, PKHD1, or GALT.
[0058] Phasing of Haplotypes of Pregnant Person
[0059] In some embodiments, the method described herein comprises phasing of the haplotypes of the pregnant person at a target gene associated with an autosomal recessive disorder. In some embodiments, the method described herein comprises performing long-read sequencing on cellular DNA extracted from a blood sample of the pregnant person or a buffy- coat fraction thereof and generating phased haplotypes of the pregnant person at a target genes associated with an autosomal recessive disorder. Examples of long-read sequencing include PacBio High Fidelity (HiFi) sequencing (https: / / www.pacb.com / technology / hifi-sequencing / ), Oxford Nanopore long read sequencing (https: / / nanoporetech.com / platform / technology), and Ultima Genomics UG-100 (https: / / www.ultimagenomics.com / ug-100).Atty. Dkt. No.: N.059.W0.01
[0060] In some embodiments, the method described herein comprises phasing of the haplotypes of the pregnant person at one or more target genes associated with cystic fibrosis, alpha thalassemia, beta thalassemia, sickle cell disease, or spinal muscular atrophy. In some embodiments, the method described herein comprises phasing of the haplotypes of the pregnant person at one or more of CFTR, HBA1, HBA2, HBB, or SMN1.
[0061] In some embodiments, the method described herein comprises phasing of the haplotypes of the pregnant person at one or more target genes associated with familial mediterranean fever, Gaucher disease, phenylalanine hydroxylase deficiency, hexosaminidase A deficiency (Tay-Sachs disease), Wilson disease, Smith-Lemil-Optiz syndrome, familial dysautonomia, carnitine palmitoyltransferase II deficiency, Canavan disease, Krabbe disease, medium chain acyl-CoA dehydrogenase deficiency, Pompe disease, autosomal recessive polycystic kidney disease, or galactosemia. In some embodiments, the method described herein comprises phasing of the haplotypes of the pregnant person at one or more of MEFV, GB A, PAH, HEXA, ATP7B, DHCR7, IKBKAP, CPT2, ASPA, GALC, ACADM, GAA, PKHD1, or GALT.
[0062] In some embodiments, the method described herein comprises performing long- read sequencing on cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof and generating phased haplotypes of the pregnant person at one or more target genes associated with cystic fibrosis, alpha thalassemia, beta thalassemia, sickle cell disease, or spinal muscular atrophy. In some embodiments, the method described herein comprises performing long-read sequencing on cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof and generating phased haplotypes of the pregnant person at one or more of CFTR, HBA1, HBA2. HBB, or SMN1.
[0063] In some embodiments, In some embodiments, the method described herein comprises performing long-read sequencing on cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof and generating phased haplotypes of the pregnant person at one or more target genes associated with familial mediterranean fever, Gaucher disease, phenylalanine hydroxylase deficiency, hexosaminidase A deficiency (Tay- Sachs disease), Wilson disease, Smith-Lemil-Optiz syndrome, familial dysautonomia, carnitineAtty. Dkt. No.: N.059.W0.01 palmitoyltransferase TI deficiency, Canavan disease, Krabbe disease, medium chain acyl-CoA dehydrogenase deficiency, Pompe disease, autosomal recessive polycystic kidney disease, or galactosemia. In some embodiments, the method described herein comprises performing long- read sequencing on cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof and generating phased haplotypes of the pregnant person at one or more of MEFV, GBA, PAH, HEXA. ATP7B. DHCR7, IKBKAP. CPT2. ASPA. GALC, AC ADM, GAA, PKHD1, or GALT.
[0064] In some embodiments, the method described herein comprising extracting cellular DNA from a buffy-coat fraction of a blood sample of the pregnant person and extracting cell-free DNA from a plasma fraction of the same blood sample; performing long-read sequencing on the extracted cellular DNA or derivative thereof and generating phased haplotypes of the pregnant person at one or more target genes; and performing sequencing on the extracted cell-free DNA or derivative thereof and determining the most likely fetal genotype and / or the most likely haplotype the fetus inherited from the pregnant person, such as described in the following sections. One exemplary embodiment is shown in Fig. ID.
[0065] Extraction of cfDNA
[0066] After a pregnant person has been identified as a heterozygous carrier of at least one pathogenic variant in a gene associated with an autosomal recessive disorder, single-gene noninvasive prenatal testing (sgNIPT) can be performed as a reflex assay to determine the genotype of the fetus. In some embodiments, the sgNIPT comprises extracting cell-free DNA from a plasma fraction of a blood sample of a pregnant person, wherein the pregnant person is a heterozygous carrier of at least one pathogenic variant in at least one target gene associated with autosomal recessive disorders, wherein the extracted cell-free DNA comprises a mixture of maternal cell-free DNA and fetal cell-free DNA.
[0067] Methods that are particularly useful in exemplary embodiments include methods for isolating circulating free DNA (cfDNA) from a liquid sample, and in illustrative embodiments from a blood, serum, or plasma sample.Atty. Dkt. No.: N.059.W0.01
[0068] In certain illustrative embodiments, isolation of cfDNA from a liquid (e.g., blood or blood derivative sample such as a serum or plasma sample) can involve binding DNA molecules from a sample to a matrix and isolating the DNA molecules in the presence of a solvent. In some embodiments, the method further comprises incubating the biological sample comprising DNA molecules with a protease, prior to contacting the DNA molecules to the matrix. In some embodiments, the method can further include the steps of washing the matrix with a wash buffer to remove impurities, and optionally, drying the matrix. Enriched nucleic acid samples can be eluted from the matrix with an elution buffer.
[0069] Other methods for nucleic acid isolation, for example cfDNA isolation, and optional enrichment of certain cfDNA can include ion exchange columns, or microfluidic devices, such as solid phase isolation, based on DNA capture by immobilized beads or functionalized surface. Additional methods include liquid phase isolation, utilizing an electric field, or chemical reagents, instead of a functionalized surface. In illustrative embodiments herein, isolation of cfDNA from a patient sample is performed using a DNA isolation kit (e.g., QIAamp Circulating Nucleic Acid kit (Qiagen)).
[0070] In some embodiments, cfDNA or their derivatives of certain sizes can be enriched before or after subjecting the cfDNA to methods herein. In some embodiments, size selection can be performed before the sequencing library preparation. In some embodiments, size selection can be performed after the sequencing library preparation and before sequencing. In some embodiments, size selection is performed on a sequencing-ready pool. Enriched cfDNA molecules can be, for example 50 to 1200 base pairs in length, 70 to 500 base pairs in length, 100 to 200 base pairs in length, 80 to 180 base pairs in length, or between 80 and 140 base pairs in length. In some embodiments, the enriched cfDNA molecules are between 60 and 200 bp in length, between 60 and 150 bp in length, between 80 and 180 bp in length, or between 80 and 140 bp in length, before the enriched cfDNA molecules, or derivatives thereof, are ligated to adapters in methods herein. Such enrichment methods can be performed for example using the methods described in WO2018 / 156418 and WO2019161244, each of which is incorporated herein by reference in its entirety.Atty. Dkt. No.: N.059.W0.01
[0071] In some embodiments, the sample is enriched for fetal DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is fetal cell-free DNA or derivatives thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of fetal cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170 bp. from 160 to 190 bp, or from 170 bp to 220 bp.
[0072] Library Preparation
[0073] In some embodiments, the method further comprises appending an adapter to the extracted DNA or derivative thereof and generating adapted DNA before performing targeted enrichment (e.g., with a panel of oligonucleotide probes). In some embodiments, the adapter is a Y-adapter. In some embodiments, the adapter comprises a universal priming site.
[0074] In some embodiments, the adapter comprises a molecular barcode or index sequence, wherein sequence reads generated from the high-throughput sequencing can be grouped together using the molecular barcode or index sequence. In some embodiments, the adapters does not comprise a molecular barcode or index sequence, wherein sequence reads generated from the high-throughput sequencing can be grouped together using the fragment-end sequence of the extracted cell-free DNA or DNA derived therefrom. In some embodiments, the sequence reads that are grouped together using the molecular barcode or index sequence and / or the fragment-end sequence can be subject to error correction to correct sequencing errors and generate a consensus sequence.
[0075] In some embodiments, the method further comprises amplifying the adapted DNA using a primer that binds to the universal primer binding site and generating adapted- amplified DNA before performing targeted enrichment (e.g., with a panel of oligonucleotide probes). In some embodiments, the adapted-amplified DNA further comprises a sequencing adapter sequence or sequencing primer binding site for high-throughput sequence. In some embodiments, the adapted-amplified DNA further comprises a sample barcode or index sequence, which allows multiplexed sequencing of pooled sequencing libraries (e.g., multiplexed sequencing of sequencing libraries generated from multiple samples).Atty. Dkt. No.: N.059.W0.01
[0076] Typically, methods herein include a step of appending nucleic acid adapters to extracted DNA molecules or to nucleic acid derivatives generated therefrom. For example, adapters may be appended on to the DNA molecules by ligation. The extracted DNA molecules in illustrative embodiments are extracted from a sample of a subject. In some embodiments, appending nucleic acid adapters is performed after the extracted DNA molecules are fragmented to form fragmented DNA molecules. Typically, methods include exposing the extracted DNA molecules to one or more polymerases or kinases, such as Klenow Large Fragment Polymerase and T4 polynucleotide kinase (PNK), as well as a ligase, such as T4 ligase. In some embodiments, the extracted DNA molecules or the fragmented DNA molecules are exposed to one or more polymerases and / or kinases to generate the nucleic acid derivatives. In some embodiments, the method further comprises appending adapters to the nucleic acid derivatives generated therefrom. In some embodiments, the extracted DNA molecules are not fragmented prior to appending nucleic acid adapters thereto. In some embodiments, the extracted DNA molecules are cfDNA molecules.
[0077] In some embodiments, adapters are ligated to the extracted DNA molecules. In some embodiments, before such ligation, the extracted DNA molecules can be modified to form sample nucleic acid derivatives, for example to make them more amenable to adapter ligation. For example, extracted DNA molecules can be blunt ended, nucleotides can be added to the extracted DNA molecules or blunted-ended derivative therefrom, and / or phosphate moieties can be added or removed from the ends of sample DNA molecules or derivatives thereof. In some embodiments, prior to ligation, the extracted DNA molecules may be blunt ended, and then a single adenosine base can be added to the 3’ end. In some embodiments, prior to ligation the DNA may be cleaved using a restriction enzyme or some other cleavage method. In some embodiments, during ligation the 3’ adenosine of the DNA fragments and the complementary 3’ thymidine overhang of an adapter can enhance ligation efficiency. In some embodiments, adapter ligation is performed using a T4 ligase.
[0078] In some embodiments, adapters containing one or more universal priming sequences are utilized in methods herein. In some embodiments, the adapters are Y adapters, for example in methods involving NGS sequencing. In some embodiments, the adapters eachAtty. Dkt. No.: N.059.W0.01 comprises a universal priming site. Tn some embodiments, the adapter may comprise a barcode or index sequence. Thus, multiple samples can be analyzed in the same sequencing reaction. The sample barcode or index sequence can be used to process data according to the sample from which the data was generated.
[0079] In some embodiments, the adapter may comprise a molecular barcode or index sequence. In some embodiments, the number of adapters having different molecular barcode or index sequences is between 10 to 1,000, and wherein the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different molecular barcode or index sequences in the ligation reaction is at least 1,000: 1. The number of different molecular barcode or index sequences in the ligation reaction, in certain embodiments, ranges fromlO to 50, 10 to 100, 50 to 200, 100 to 300, 200 to 500, 300 to 600, 500 to 700, 600 to 800 or 700 to 1,000. In some embodiments, there are at least 1, 10, 20, 30, 40, 50, or at least 100; 200, 300, 400, 500, 600, 700. 800, 900, or 1000 different molecular barcode or index sequences in the ligation reaction. In some embodiments, the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different molecular barcode or index sequences in the ligation reaction is at least 10,000: 1. In some embodiments, the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different molecular barcode or index sequences in the ligation reaction ranges from 50.000: 1 to 50:1. from 25,000: 1 to 100:1, from 10,000: 1 to 100: 1 , from 10:000: 1 to 8,000: 1 to 500: 1 , from 5,000: 1 to 200: 1 , from 10,000: 1 to 50:1. In some embodiments, the methods disclosed herein result in at least 100; 200; 500; 750; 1,000; 2,000; 5,000; 7,500; 10,000: 20,000; 25,000; 30,000; 40,000; 50,000 different molecular barcode or index sequences to each one template nucleic acid or cfDNA molecules.
[0080] An exemplary library preparation protocol is provided in Example 2.
[0081] Targeted Enrichment
[0082] In some embodiments, the method described herein comprises performing targeted enrichment on the extracted cell-free DNA or DNA derived therefrom to enrich a plurality of target variant loci in a plurality of target genes associated with autosomal recessiveAtty. Dkt. No.: N.059.W0.01 disorders and generating enriched DNA, wherein the target variant loci comprise at least one pathogenic variant carried by the pregnant person.
[0083] In some embodiments, the at least one pathogenic variant comprises a single nucleotide variant (SNV), an indel, a copy number variation (CNV), a gene fusion, a chromosomal rearrangement (e.g„ deletion, duplication, inversion, or translocation), or a combination thereof.
[0084] In some embodiments, the target genes are associated with one or more of cystic fibrosis, alpha thalassemia, beta thalassemia, sickle cell disease, or spinal muscular atrophy. In some embodiments, the target genes comprise one or more of CFTR, HBA1, HBA2, HBB, or SMN1.
[0085] In some embodiments, the target genes are associated with one or more of familial mediterranean fever, Gaucher disease, phenylalanine hydroxylase deficiency, hexosaminidase A deficiency (Tay-Sachs disease), Wilson disease, Smith-Lemil-Optiz syndrome, familial dysautonomia, carnitine palmitoyltransferase II deficiency, Canavan disease, Krabbe disease, medium chain acyl-CoA dehydrogenase deficiency, Pompe disease, autosomal recessive polycystic kidney disease, or galactosemia. In some embodiments, the target genes comprise one or more of MEFV, GBA, PAH, HEXA, ATP7B, DHCR7, IKBKAP, CPT2, ASPA, GALC, AC ADM, GAA, PKHD1, or GALT.
[0086] In some embodiments, the pregnant woman is a heterozygous carrier of a pathogenic variant in the CFTR gene, and the target variant loci encompass said pathogenic variant loci in the CFTR gene. In some embodiments, the target variant loci encompass all pathogenic SNVs / indels in the CFTR gene that are known to be associated with autosomal recessive disorders.
[0087] In some embodiments, the pregnant woman is a heterozygous carrier of a pathogenic variant in the HBA1 gene, and the target variant loci encompass said pathogenic variant loci in the HBA1 gene. In some embodiments, the target variant loci encompass all pathogenic SNVs / indels in the HBA1 gene that are known to be associated with autosomal recessive disorders. In some embodiments, the pathogenic variant is a chromosomalAtty. Dkt. No.: N.059.W0.01 rearrangement involving HBA1 / HBA2. Tn some embodiments, the pathogenic variant is a large deletion of HBA1 or a portion thereof, such as from one of the chromosomes. In some embodiments, the pathogenic variant is a large deletion of both HBA 1 and HBA2, such as from one of the chromosomes. Having less than 2 total copies of HBA is known to be associated with autosomal recessive disorders.
[0088] In some embodiments, the pregnant woman is a heterozygous carrier of a pathogenic variant in the HBA2 gene, and the target variant loci encompass said pathogenic variant loci in the HBA2 gene. In some embodiments, the target variant loci encompass all pathogenic SNVs / indels in the HBA2 gene that are known to be associated with autosomal recessive disorders. In some embodiments, the pathogenic variant is a chromosomal rearrangement involving HBA1 / HBA2. In some embodiments, the pathogenic variant is a large deletion of HBA2 or a portion thereof, such as from one of the chromosomes. In some embodiments, the pathogenic variant is a large deletion of both HBA 1 and HB A2, such as from one of the chromosomes. Having less than 2 total copies of HBA is known to be associated with autosomal recessive disorders.
[0089] In some embodiments, the pregnant woman is a heterozygous carrier of a pathogenic variant in the HBB gene, and the target variant loci encompass said pathogenic variant loci in the HBB gene. Tn some embodiments, the target variant loci encompass all pathogenic SNVs / indels in the HBB gene that are known to be associated with autosomal recessive disorders.
[0090] In some embodiments, the pregnant woman is a heterozygous carrier of a pathogenic variant in the SMN 1 gene, and the target variant loci encompass said pathogenic variant loci in the SMN1 gene. In some embodiments, the target variant loci encompass all pathogenic SNVs / indels in the SMN1 gene that are known to be associated with autosomal recessive disorders. In some embodiments, the pathogenic variant is a chromosomal rearrangement involving SMN1 / 2. In some embodiments, the pathogenic variant is a large deletion of SMN1 or a portion thereof (e.g., exon 7), such as from one of the chromosomes. In some embodiments, the pathogenic variant is a conversion of SMN 1 and SMN2, such as through a point mutation or SNV.Atty. Dkt. No.: N.059.W0.01
[0091] In some embodiments, the targeted enrichment further enriches a plurality of phasing SNP loci within and / or flanking one or more of the target genes, wherein the plurality of phasing SNP loci each has a minor allele frequency (MAF) of at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, or at least 10%. In some embodiments, MAF refers to the frequency at which the second most common allele occurs in a human population. In some embodiments, SNPs with a MAF of >5% are considered common variants, and SNPs with a MAF of <5% are considered rare variants.
[0092] In some embodiments, the plurality of phasing SNP loci are located up to 20kbp- 500kbp from at least one of the target variant loci. In some embodiments, the plurality of phasing SNP loci are located up to 20kbp-500kbp from the closest target variant loci. In some embodiments, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the phasing SNP loci are located no more than 100,000 bp, 50,000 bp, 20,000 bp, 10,000 bp, 5.000 bp, 2,000 bp, 1.000 bp, 500 bp, 200 bp. or 100 bp from the closest target variant loci.
[0093] In some embodiments, the plurality of phasing SNP loci comprise at least 100, at least 200, or at least 400 phasing SNP loci per target gene. At an average MAF of 5%, this translate to at least 5, at least 10, or at least 20 informative heterozygous SNPs per haplotype. In some embodiments, the phasing SNPs are heterozygous in the pregnant person.
[0094] In some embodiments, the phased haplotypes of the pregnant individual covering one or more of the target genes can be generated using sequence reads at the plurality of phasing SNP loci (e.g., sequence reads of informative heterozygous SNPs). Long range PCR amplicons can be designed to span multiple heterozygous loci. Long reads generated by PacBio can be used to assemble haplotypes that span an entire a target gene and nearby common SNPs or parts of the gene that include the target variant. In some embodiments the method comprises using (i) the phased haplotypes of the pregnant person at one or more of the target genes (e.g., phased haplotypes generated from long-read sequencing of cellular DNA or derivative thereof, with the cellular DNA being extracted from a buffy-coat fraction of a blood sample of the pregnant person), and (ii) the sequence reads at phasing SNP loci that are heterozygous in the pregnant person (e.g., sequence reads generated from sequencing of cell-free DNA or derivative thereof,Atty. Dkt. No.: N.059.W0.01 with the cell-free DNA being extracted from a plasma fraction of the blood sample), to determine the most likely fetal genotype and / or the most likely haplotype the fetus inherited from the pregnant person.
[0095] In some embodiments, fetal fraction of cfDNA in the plasma fraction of the maternal blood sample can be estimated using sequence reads at the plurality of phasing SNP loci or other SNP loci specifically included in target enrichment to enable estimation of fetal fraction. In some embodiments, deletion in one or more of the target genes can be identified using sequence reads at the plurality of phasing SNP loci.
[0096] In some embodiments, the targeted enrichment comprises preforming targeted probe capture to enrich the target loci. In some embodiments, the targeted enrichment comprises preforming linked target capture to enrich the target loci. In some embodiments, the targeted enrichment comprises preforming targeted multiplex amplification to enrich the target loci.
[0097] Hybrid Capture
[0098] In some embodiments, the targeted enrichment technique can involve fragment capture by hybridization (i.e., hybrid capture). Although any hybrid capture method can be used to perform methods herein that include a targeted enrichment step, in some embodiments, a method of the present disclosure may involve using any of the hybrid capture methods disclosed herein to enrich cfDNA of one or more target genes associated with autosomal recessive disorders. In some embodiments described herein, the targeted enrichment steps can be performed after cfDNA molecules are extracted. In some embodiments described herein, the targeted enrichment steps can be performed after appending adapters to the extracted cfDNA molecules. In some embodiments described herein, the targeted enrichment steps can be performed after PCR amplification of the adapted DNA.
[0099] In capture by hybridization, hybrid capture oligonucleotide probes complementary to one or both strands of specific target DNA sequences, or DNA derived therefrom in a sample, are utilized, i.e., the probes may be strand specific. The specific target DNA sequence in illustrative embodiments overlaps with or is found within a target region of a sample DNA molecule such as a cfDNA. Thus, hybrid capture probes when used in methodsAtty. Dkt. No.: N.059.W0.01 herein can be designed to bind to a DNA molecule that contains at least one target region or a portion thereof. In some embodiments, the hybrid capture probes can be designed to bind to a target DNA sequence within or overlapping a target region. In other examples, the hybrid capture probes can be designed to bind to a common region that is flanking but not overlapping the target region and that can be a common region that was added to some, most, almost all or all of the DNA in a sample, or added to all amplicons using a common sequence on at least one primer of a primer pair. In illustrative embodiments, a hybrid capture probe or set thereof, are designed to bind to a target DNA sequence within target region, or set of target regions, respectively.
[0100] In some embodiments, the hybrid capture probes can be designed to bind to the canonical sequence of the target, or to bind to the pathogenic variant sequence separately or in addition to the canonical sequence.
[0101] Hybrid capture probes may be added to a prepared sample and hybridized through a denature-reannealing process to form duplexes of exogenous-endogenous fragments (e.g., hybrid capture probes bound to sample DNA molecules, or DNA derived therefrom). These duplexes may then be physically separated from the sample by various means. In some embodiments, once the hybrid capture probes are removed, the sample DNA molecules, or DNA derived therefrom can be amplified. Some ways to physically remove the hybrid capture probes are by covalently bonding the hybrid capture probes to a solid support, for example a magnetic bead, or a chip. Another way to physically remove the hybrid capture probes is by covalently bonding them to a molecular moiety with a strong affinity for another molecular moiety. An example of such a molecular pair is biotin and streptavidin, such as is used in xGen™ NGS Hybridization Capture (IDT) or in SURE SELECT (Agilent). Thus, hybrid capture probes, for example that bind to a target DNA sequence within or overlapping a target region of a DNA molecule obtained or derived from a sample, can be covalently attached to a biotin molecule, and after hybridization with sample DNA or DNA derived therefrom, a solid support with streptavidin affixed can be used to pull down the biotinylated hybrid capture probes, which are hybridized to DNA molecules obtained or derived from a sample that include a target region that includes the target DNA sequence recognized by the hybrid capture probes. Thus, in someAtty. Dkt. No.: N.059.W0.01 embodiments, the hybrid capture probes are immobilized, directly or indirectly to a solid support. In some embodiments, the hybrid capture probes include a binding partner, for example biotin.
[0102] In some embodiments of any of the aspects herein, the hybrid capture probes can be a part of a set of at least two hybrid capture probes. In some embodiments, the set includes at least one hybrid capture probe for each target region. In some embodiments, the set includes two or more hybrid capture probes for each target region. In some embodiments, the set includes three or more hybrid capture probes for each target region. In some embodiments, the set includes four or more hybrid capture probes for each target region.
[0103] In some embodiments of any of the aspects herein, the hybrid capture probes can have a length in the range of 30 bases to 170 bases, 30 bases to 160 bases, 30 bases to 150 bases, 30 bases to 140 bases, 30 bases to 130 bases, 30 bases to 120 bases, 30 bases to 110 bases, 30 bases to 100 bases, 30 bases to 90 bases. 30 bases to 80 bases, 30 bases to 70 bases, 30 bases to 60 bases, 30 bases to 50 bases, 40 bases to 160 bases, 40 bases to 150 bases, 40 bases to 140 bases, 40 bases to 130 bases, 40 bases to 120 bases, 40 bases to 110 bases, 40 bases to 100 bases, 40 bases to 90 bases, 40 bases to 80 bases, 40 bases to 70 bases, 40 bases to 60 bases, 50 bases to 150 bases, 50 bases to 140 bases, 50 bases to 130 bases, 50 bases to 120 bases, 50 bases to 110 bases. 50 bases to 100 bases, 50 bases to 90 bases, 50 bases to 80 bases. 50 bases to 70 bases. 60 bases to 140 bases, 60 bases to 130 bases, 60 bases to 120 bases, 60 bases to 1 10 bases, 60 bases to 100 bases, 60 bases to 90 bases, 60 bases to 80 bases, 70 bases to 130 bases, 70 bases to 120 bases. 70 bases to 110 bases, 70 bases to 100 bases. 70 bases to 90 bases, 80 bases to 120 bases, 80 bases to 110 bases, 80 bases to 100 bases, 90 bases to 120 bases, 90 bases to 110 bases, 100 bases to 165 bases, 100 bases to 150 bases, 100 bases to 140 bases, 100 bases to 130 bases, 100 bases to 120 bases, 110 bases to 150 bases, 110 bases to 140 bases, 110 bases to 130 bases, 120 bases to 150 bases, or 130 bases to 160 bases.
[0104] In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the entire exons of one or more of the target genes. In some embodiments, the targeted probe capture is performed using probes covering the entire exons and introns of one or more of the target genes. In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the entire exons of one or more of the targetAtty. Dkt. No.: N.059.W0.01 genes that carry a pathogenic variant. In some embodiments, the targeted probe capture is performed using probes covering the entire exons and introns of one or more of the target genes that carry a pathogenic variant.
[0105] In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the entire exons of CFTR. In some embodiments, the targeted probe capture is performed using probes covering the entire exons of CFTR that carry a pathogenic variant. In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the entire exons of SMN1. In some embodiments, the targeted probe capture is performed using probes covering the entire exons of SMN 1 that carry a pathogenic variant. In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the entire exons of HBB. In some embodiments, the targeted probe capture is performed using probes covering the entire exons of HBB that carry a pathogenic variant. In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the entire exons of HBA1. In some embodiments, the targeted probe capture is performed using probes covering the entire exons of HBA1 that carry a pathogenic variant. In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the entire exons of HBA2. In some embodiments, the targeted probe capture is performed using probes covering the entire exons of HBA2 that carry a pathogenic variant.
[0106] In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the entire exons and introns of HBB. In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the entire exons and introns of HBA1. In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the entire exons and introns of HBA2. In some embodiments, the targeted probe capture is performed using hybrid capture probes covering the intergenic region between HBA1 and HBA2.
[0107] In some embodiments, the targeted probe capture is performed using tiling probes providing at least 2X, at least 3X, or at least 4X tiled coverage of the entire exons of one or more of the target genes with or without a pathogenic variant. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 2X, at least 3X, or at least 4XAtty. Dkt. No.: N.059.W0.01 tiled coverage of the entire exons and introns of one or more of the target genes with or without a pathogenic variant. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 2X, at least 3X, or at least 4X tiled coverage of the entire exons of one or more of the target genes that carry a pathogenic variant. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 2X, at least 3X, or at least 4X tiled coverage of the entire exons and introns of one or more of the target genes that carry a pathogenic variant.
[0108] In some embodiments, the hybrid capture probes are 80-200 nucleotides in length, 100-150 nucleotides in length, or 110-130 nucleotides in length. In some embodiments, the hybrid capture panel was designed as a 4X tiling probe set (e.g., 4 baits per base with ~90 bp overlap for ~120bp probes). In some embodiments, the hybrid capture panel was designed as a 2X tiling probe set (2 baits per base with ~60 bp overlap for ~120bp probes).
[0109] In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons of CFTR. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons of SMN1. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons of HBB. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons of HBA1. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons of HB A2.
[0110] In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons and introns of HBB. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons and introns of HBA1. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons and introns of HB A2. In some embodiments, the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the intergenic region between HBA1 and HBA2.Atty. Dkt. No.: N.059.W0.01
[0111] Linked target capture using probe-dependent primers
[0112] In some embodiments, the targeted enrichment technique can involve probedependent primers. Probe-dependent primers (PDPs) have been disclosed (Pel, et al. “Rapid and highly-specific generation of targeted DNA sequencing libraries enabled by linking capture probes with universal primers” PLoS ONE 13(12):e0208283 (2018); WO 2017168332A1 “Linked duplex target capture”, which are hereby incorporated by reference in their entirety). Such embodiments can be considered linked target capture (LTC) methods. Briefly, in an LTC method, a target-specific probe is linked to a universal primer. The target specific probe is designed to hybridize to a target of interest such as one of the target variant loci. The universal primer linked to the probe is designed to hybridize to a universal priming site in the adaptors that have been ligated to the cell-free DNA. The binding of the probe to the target brings the linked universal primer into proximity with the universal priming site and in fact, the ability of the universal primer to bind to the universal priming site and be extended depends on the probe binding to its target. Results to-date have shown that the universal primers do not hybridize to adaptors attached to fragments that do not include the target. The linked target capture is highly target specific and primer extension depends on successful probe binding. For that reason, the universal primers linked to the probes are "probe-dependent primers". LTC may be performed using only one PDP, e.g., for a linear amplification, but preferably uses paired forward and reverse PDPs (as shown in Fig. 1 (b) of Pel 2018 as target-capture PCR 1 ) to amplify the fragment exponentially. The bound probe does not interfere with primer extension to copy the entire fragment (including the distal adaptor) when a strand-displacing polymerase is used. It is noted that the cell-free DNA fragment is copied by primers extended from within the ligated adaptors, so the entirety of the fragment is copied into amplicons. Due to the probe, LTC gives the target specificity of conventional PCR with gene- specific primers while due to the priming sites in the adaptors, the entire fragment is amplified. The resulting amplification products will include a copy of the entire target-containing fragment with adaptors at both ends. Preferably, PDPs are designed to incorporate non-extendable capture probes linked 5’ to 5’ with a primer. Multiple linker types are possible as discussed below. Typically, probes of PDPs can be between 30 to 70 nucleotides in length, and include or comprise a 3’ inverted dT base or 3’ C3 spacer to inhibit polymerase extension. In some embodiments, probes are designed to cover the desired regionAtty. Dkt. No.: N.059.W0.01 with overlap between forward and reverse probes. Tn some embodiments, the probes are between 20 and 100 nucleotides in length. In some embodiments, the size of the probe can be between 20 and 40 nucleotides, between 30 and 50 nucleotides, between 40 and 60. between 50 and 70 between 60 and 80, between 70 and 90, 80 and 100, 90 and 110, 100 and 120 nucleotides in length. In some embodiments, at least one of the probes of a PDP pair comprises a sample index.
[0113] In PDPs, forward and reverse probes can be designed to bind to nucleic acid sequences within or near a target variant loci on a sample DNA molecule to enrich nucleic acid molecules comprising the target variant loci of interest or copies thereof. Typically, the primer portion of a PDP is a universal primer designed to bind to a universal primer site on the appended adapter. In some embodiments, the PDP is designed with a sequencer binding sequence, such as an Illumina flow cell binding sequence, incorporated therein. In some embodiments, the sequencer flow cell binding sequence is between the probe and universal primer, and adjacent to the primer. Linked primers of the invention may also include sequencing tags to ensure that all cluster reads originate from the same linked template molecule. The lengths of the primers can be extended or shortened at the 5' end or the 3' end to produce primers with desired melting temperatures. Also, the annealing position of each primer pair can be designed such that the sequence and length of the primer pairs yield the desired melting temperature. In illustrative embodiments, the primer is a low melting temperature universal primer complementary to a portion of the ligated adapter.
[0114] The primer can be tailed or untailed depending on the specific requirements. In some embodiments, the universal primer comprises an A tail. The length of the primers of the PDP can range from 5 to 40 nucleotides in length. In certain embodiments, the PDP primers are between 10 and 25 nucleotides long. In embodiments, the primers of the PDP can range from 5 to 15 nucleotides, from 10 to 25 nucleotides, from 15 to 35 nucleotides, or from 25 to 40 nucleotides in length.
[0115] Typically, probe dependent primers comprise a linker between the probe and the primer. Probe and primer portions of the PDP are typically linked by a polyethylene glycol derivative, an oligosaccharide, a lipid, a hydrocarbon, a polymer, or a protein. In some embodiments, the linker is a PEG molecule, or derivative thereof. In some embodiments, theAtty. Dkt. No.: N.059.W0.01 linker is an oligosaccharide. Tn some embodiments, the linker is a lipid. Tn some embodiments, the linker is a hydrocarbon. In some embodiments, the linker is a polymer. In some embodiments, the linker is a protein, or portion thereof. Linkers based on click chemistry is described in WO2017 / 168332A1, which is incorporated herein by reference in its entirety.
[0116] Targeted amplification
[0117] In some embodiments, the targeted enrichment technique can involve targeted multiplex amplification (e.g., PCR or isothermal amplification). Methods in some aspects herein include performing one or, in some embodiments, two or more amplifications. Such amplifications in certain illustrative embodiments include at least one targeted amplification wherein at least one primer and in certain embodiments both primers of a primer pair, one or more primer pairs, or a set of primer pairs used for the amplification are each designed to bind to a specific nucleic acid sequence at or near, typically within, a genomic region of interest comprising a target variant loci (i.e. are target- specific primers) to generate target region amplicons. In some embodiments, methods herein include one or more universal amplifications.
[0118] A number of amplification technologies can be used with methods herein. For example, such amplification can be an isothermal amplification (e.g., recombinase polymerase amplification (RPA) (Kersting et al. 2014 Microchim Acta 181 (13-14), 1715-1723, (incorporated by reference in its entirety)), a ligase -based amplification, PCR, or a combination thereof (e.g., ligation-mediated PCR). In some illustrative embodiments, the targeted amplification is a targeted PCR(s) that is performed using a PCR reaction mixture that includes one primer pair, or in illustrative embodiments a set of primer pairs, and at least a portion of the library of DNA molecules comprising the extracted cfDNA or DNA derived therefrom (e.g., adapted DNA, adapted- amplified DNA).
[0119] Typically, at least one primer of a primer pair used for targeted amplification herein, is a target-specific primer designed to bind to a specific nucleic acid sequence in or near, typically within, a genomic region of interest comprising the target variant loci, which in illustrative examples can be genomic regions where pathogenic variants (e.g., SNVs., indels) are associated with autosomal recessive disorders. A target- specific primer can be designed to bindAtty. Dkt. No.: N.059.W0.01 to any sequence within or near a target region for amplification of the target region or a portion of the target region. One of the advantages of the methods described herein is increased flexibility in primer / probe design for targeted amplification or enrichment. In some embodiments, one primer of the one or more primer pairs or the set of primer pairs in the reaction mixture used for a targeted amplification is a universal primer and binds to a primer binding site on an adapter. Thus, for example, in such embodiments a universal primer that binds an adapter primer binding site can be used for an amplification reaction along with a targetspecific primer that binds a primer binding site on a sample DNA region.
[0120] Target-specific primers typically define the ends of target region amplicons, which typically encompass at least a portion of the target region. In some embodiments, a PCR can be performed using two target-specific primers. The target region amplicon in such embodiment would extend from the sample DNA region bound by target-specific primer on a 5’ end to the sample DNA region bound by primer on the 3’ end. In some embodiments, a PCR can be performed using a universal primer and a target-specific primer. The target region amplicon in such embodiment would extend from the sample DNA region bound by target- specific primer on a 3’ end of one strand to the end of the sample DNA fragment on the 5’ end of that strand.
[0121] In some methods herein, a universal amplification of the library of DNA molecules comprising the extracted cfDNA or DNA derived therefrom can be performed before the targeted amplification. Such universal amplification can be performed for example using a universal primer pair that binds primer binding sites in the adapter. Thus, in some embodiments, the methods herein include performing a universal PCR using the adapted DNA molecules, and a universal PCR primer pair comprising primers designed to bind universal primer binding sequences on the adapters, before performing one or more targeted PCRs.
[0122] The one or more primer pairs in illustrative embodiments is a set of primer pairs. In some embodiments, the set of primer pairs is a set of between 2 and 1,000, 2 and 500, 2 and 250, 2 and 200, 2and 150, 2 and 100, 2 and 50 or 2 and 10 primer pairs, or between 5 and 1,000, 5 and 500, 5 and 250, 5 and 200, 5 and 150, 5 and 100, 5 and 50 or 5 and 10 primer pairs, or between 50 and 250 or between 100 and 200 primer pars.Atty. Dkt. No.: N.059.W0.01
[0123] In some embodiments, at least one of the primer pairs comprises a universal primer and a target- specific primer. In some embodiments, at least one of the primer pairs comprises two target-specific primers. In some embodiments, at least one of the primers comprises a sequencing tag. In some embodiments, at least one of the primers comprises a sample index. In some embodiments, at least one of the primers comprises biotin modification. In some embodiments, performing a PCR further comprises using primers comprising a sequencing tag. In some embodiments, performing a PCR further comprises using primers comprising a sample index. In some embodiments, the primers of the primer pairs are probedependent primers and the amplification is a target capture polymerase chain reaction.
[0124] In some embodiments, one or both of the primer binding sites of a primer pair can include at least a portion of one of the adapter sequences. In some embodiments, one of the primer binding sites of a primer pair can include at least a portion of the adapter sequences. In some embodiments, both of the primer binding sites of a primer pair can include at least a portion of the adapter sequences. In some embodiments, neither of the primer binding sites of a primer pair include any of the adapter sequences.
[0125] Methods as described herein, in some embodiments, can include multiple amplification cycles (e.g., multiple PCR temperature cycles), and in some embodiments can include several sequential PCR reactions performed during the same set of temperature cycles. In some embodiments, amplification cycles can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 cycles. In some embodiments, amplification cycles can include at least 7, 8, 9, or 10 cycles. In illustrative embodiments, amplification cycles can include at least 11, 12, 13, 14, 15, 16, or 17 cycles.
[0126] Typically, in embodiments described herein, PCR amplification is performed by adding a PCR reaction mixture to the DNA template followed by addition of a polymerase enzyme, and then amplified through multiple amplification cycles. In some embodiments, the PCR reaction mixture contains one or more primer pairs, deoxynucleotides (dNTPs), PCR reaction buffer, and deionized water. In some embodiments, the dNTPs can comprise a mixture of dATP, dCTP, dGTP and dTTP. In some embodiments, the final concentration of each dNTP in the reaction mixture can range from 0.05 mM to 0.5 mM dNTPs, for example, from 0.05 mMAtty. Dkt. No.: N.059.W0.01 to 0.5 mM, 0.05 mM to 0.1 M, 0.05 mM to 0.15 mM, 0.05 to 0.2 mM, 0.05 mM to 0.25 mM, 0.05 mM to 0.3 mM, 0.05 mM to 0.35 mM, 0.05 to 0.4 mM 0.05 mM to 0.45 mM, from .1 mM to 0.5 mM, 0.15 mM to 0.5 mM, or 0.2 mM to 0.5 mM, 0.25 to 0.5 mM, 0.3 mM to 0.5 mM, 0.35 mM to 0.5 mM, 0.4 mM 0.5 mM, or from 0.45 mM to 0.5 mM. In illustrative embodiments, the final concentration of each dNTP in the reaction mixture is between 0.15 mM and 0.25 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is 0.2 mM.
[0127] PCR buffer solution creates a suitable environment for the polymerase chain reaction and can contain many different components, including magnesium chloride (MgC12), potassium chloride (KC1), dimethyl sulfoxide (DMSO), and glycerin or bovine serum albumin (BSA). In some embodiments, the concentration of KC1 can be between 25 and 50 mM, between 25 and 75 mM, between 25 and 100 mM, between 30 and 100 mM, between 50 and 100 mM, or between 70 and 100 mM. In some embodiments, the MgC12 concentration can be in the range of 0.5 mM to 5 mM, between 0.5 mM to 4.5 mM, 0.5 to 4.0 mM, 0.5 to 3.5 mM, 0.5 to 3.0 mM, 0.5 to 2.5 mM, 0.5 to 2.0 mM, 0.5 to 1.5 mM, 1.0 to 5 mM, 1.5 to 5mM, 2.0 to 5 mM, 2.5 to 5 mM, 3.0 to 5 mM, 3.5 to 5 mM, 4.0 to 5 mM, or 4.5 to 5 mM. In some embodiments, the concentration of MgC12 is 2.0 mM.
[0128] In illustrative embodiments, the buffer solution is a Q5® Reaction Buffer (B9027S, New England Biolabs, Inc.). In some embodiments, the reaction buffer is Standard Taq Reaction Buffer (B9014S, New England Biolabs, Inc). In some embodiments, the reaction buffer is a Standard Taq (Mg-free) Reaction Buffer (B9015S, New England Biolabs, Inc.).
[0129] In some embodiments, a DNA polymerase is used to produce DNA amplicons using DNA as a template. In some embodiments, the polymerase is a Q5® DNA Polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). Q5® High- Fidelity DNA polymerase is a high-fidelity, thermostable, DNA polymerase with 3'— ► 5' exonuclease activity, fused to a processivity-enhancing Sso7d domain. Q5® High-Fidelity DNA polymerase lacks 5 '—> 3 "exonuclease activity and strand displacement activity.Atty. Dkt. No.: N.059.W0.01
[0130] In some embodiments, the polymerase is a T4 DNA polymerase (M0203S, New England BioLabs, Inc.). T4 DNA Polymerase catalyzes the synthesis of DNA in the 5' > 3' direction and requires the presence of template and primer. This enzyme has a 3'^ 5 ' exonuclease activity which is much more active than that found in DNA Polymerase I. T4 DNA polymerase lacks 5'—> 3' exonuclease activity and strand displacement activity.
[0131] In some embodiments of any of the aspects herein, the length of the primers can be between 10 to 100 nucleotides, such as between 10 to 75 nucleotides, 10 to 40 nucleotides, 10 to 35 nucleotides, 10 to 30 nucleotides, 10 to 20 nucleotides, 15 to 100 nucleotides, 20 to 100 nucleotides, from 25 to 100 nucleotides, from 30 to 100 nucleotides from 35 to 100 nucleotides, from 40 to 100 nucleotides, from 45 to 100 nucleotides, from 50 to 100 nucleotides, from 55 to 100 nucleotides, from 60 to 100 nucleotides, from 65 to 100 nucleotides, from 70 to 100 nucleotides, or from 75 to 100 nucleotides. In some embodiments, the range of the length of the primers is between 5 to 50 nucleotides, such as 5 to 40 nucleotides, 5 to 20 nucleotides, or 5 to 10 nucleotides. In some embodiments, the primers are between 5 and 50 bp in length, between 10 and 40 bp in length, between 15 and 30 bp in length, between 15 and 25 bp in length, between 20 and 40 bp in length, between 25 and 50 bp in length, or between 30 and 50 bp in length. In some embodiments, the primers are between 25 and 100 bp in length, between 35 and 100 bp in length, between 45 and 100 bp in length, between 55 and 100 bp in length, between 65 and 100 bp in length, or between 75 and 100 bp in length.
[0132] In some embodiments of any of the aspects or embodiments herein, the number of primer pairs can range from 1 to 100,000 primer pairs that each bind to one or more primer binding sequences. In some embodiments, the primer pairs are a part of a set of primer pairs. In some embodiments, the set of primers range from 2 to 100,000, from 2 to 10,000, from 2 to 1,000, from 2 to 100, from 2 to 50, from 10 to 100, from 50 to 100, from 100 to 200, from 100 to 500, from 100 to 1,000, from 100 to 10,000, from 100 to 100,000, from 1,000 to 100,00, or from 10,000 to 100,000 primer pairs. In some embodiments, the number of primer pairs can range from 10 to 10,000, 10 to 1,000, 10 to 100, 10 to 50, 10 to 40, 10 to 30, 15 to 30, or 15 to 25 primer pairs.Atty. Dkt. No.: N.059.W0.01
[0133] Tn some embodiments, PCR is used to generate very short amplicons. cfDNA (such as fetal cfDNA in maternal serum) is highly fragmented. For fetal cfDNA, the fragment sizes are distributed in approximately a Gaussian fashion with a mean of 160 bp. a standard deviation of 15 bp, a minimum size of about 100 bp, and a maximum size of about 220 bp. Because cfDNA fragments are short, the likelihood of both primer sites being present the likelihood of a fragment of length L comprising both the forward and reverse primers sites is the ratio of the length of the amplicon to the length of the fragment. Under ideal conditions, assays in which the amplicon is 45, 50, 55, 60, 65, or 70 bp will successfully amplify from 72%, 69%, 66%, 63%, 59%, or 56%, respectively, of available template fragment molecules. Thus, in some embodiments target amplicons generated in method herein are between 40 and 100, 40 and 75, or 45 and 70 bp in length. The amplicon length is the distance between the 5-prime ends of the forward and reverse priming sites. In an embodiment, a substantial fraction of the amplicons are between 25 on the low end of the range, and 100 bp, 90 bp, 80 bp, 70 bp, 65 bp, 60 bp, 55 bp, 50 bp, or 45 bp on the high end of the range.
[0134] Sequencing to Generate Sequence Reads
[0135] In some embodiments, the method described herein comprises performing sequencing on the enriched DNA or DNA derived thereof and generating sequence reads. In some embodiments, the sequencing is next-generation sequencing or high-throughput sequencing.
[0136] DNA sequencing techniques, particularly high throughput next-generation sequencing techniques (often referred to as massively parallel sequencing techniques) such as those employed NOVASEQ (ILLUMINA), MISEQ (ILLUMINA), HISEQ (ILLUMINA), ION TORRENT (LIFE TECHNOLOGIES), GENOME ANALYZER ILX (ILLUMINA), GS FLEX+(ROCHE 454) etc., can be used for determining the sequences of the enriched DNA or DNA derived thereof to elucidate the sequence of the original cfDNA. High throughput genetic sequencers are amenable to the use of barcoding (i.e., sample tagging with distinctive nucleic acid sequences) so as to identify specific samples from individuals thereby permitting the simultaneous analysis of multiple samples in a single run of the DNA sequencer. Methods as described herein that utilize NGS detection, in some embodiments can have an average depth ofAtty. Dkt. No.: N.059.W0.01 read of at least 0.1 , 0.5, 1 , 10, 50, 100, 200, 500, 1000, 2000, 2900, 3000, 3500, 4000, 5000, 10,000, 50,000, 75,000, 100,000, 130,000, 150,000, 175,000, or 200,000.
[0137] Methods herein can include analyzing data obtained from next- generation sequencing techniques. In some embodiments of methods herein, the enriched DNA or DNA derived thereof can be subjected to sequencing using next-generation sequencing techniques. For a skilled artisan, algorithm design tools are available that can be used and / or adapted to analyze the sequencing data. In addition, those skilled in the art can determine appropriate parameters for measuring alignment to a consensus sequence and / or to a known target region sequence, including any algorithms needed to achieve maximal alignment over the length of the sequences being compared.
[0138] Sequencing reads can be demultiplexed using an in-house tool and mapped using the Burrows- Wheeler alignment software, Bwa mem function (BWA, Burrows-Wheeler Alignment Software (see Li H. and Durbin R. (2010) Fast and accurate long-read alignment with Burrows-Wheeler Transform. Bioinformatics.) in single end or paired end mode to a version of reference genome. The reference genome used can be hgl9 or hg38. Amplification statistics QC can be performed by analyzing one or more of, but not limiting to, total reads, number of mapped reads, number of mapped reads on target, and number of reads counted.
[0139] Methods herein can include a background error model that can be constructed using normal, healthy, or non-diseased liquid samples, in illustrative embodiments, normal, healthy, or non-diseased plasma samples, which are sequenced on the same sequencing run to account for run-specific artifacts. In some embodiments, 5. 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250. or more than 250 normal, healthy, or non-diseased liquid samples, in illustrative embodiments, plasma samples can be analyzed on the same sequencing run. The number of samples that can be sequenced on the same sequencing run can be in the range of 5 to 500, 5 to 400, 5 to 300, 5 to 250, 20 to 250, 30 to 250, 50 to 250, 75 to 250, 100 to 250, 50 to 500, or 100 to 500. Sample barcodes are used in illustrative embodiments. In some illustrative embodiments, 20, 25, 40, or 50 normal samples (e.g., plasma samples) can be analyzed on the same sequencing run. Outlier samples can be iteratively removed from the model to account for noise and contamination. In some embodiments, samples with a Z score of greater than 5, 6, 7, 8, 9, or 10Atty. Dkt. No.: N.059.W0.01 are removed from the data analysis. For each base substitution of every genomic loci, the DOR weighted mean and standard deviation of the error can be calculated.
[0140] Methods herein can include calculating percent identity that can be calculated by determining the number of matched positions in aligned DNA sequences, dividing the number of matched positions by the total number of aligned DNA sequences, and multiplying by 100. A matched position refers to a position in which identical nucleotides occur at the same position in aligned DNA sequences. The percent identity over a particular length can be determined by counting the number of matched positions over that length and dividing that number by the length followed by multiplying the resulting value by 100. A non-limiting example for calculating the percent identity, can be, if (i) a 500-nucleotide DNA target sequence is compared to a subject DNA sequence, (ii) an alignment program presents 200 nucleotides from the target DNA sequence aligned with a region of the subject DNA sequence where the first and last nucleotides of that 200-nucleotide region are matches, and (iii) the number of matches over those 200 aligned nucleotides is 180, then the 500-nucleotide nucleic acid target sequence contains a length of 200 and a sequence identity over that length of 90 percent (i.e., 180, 200x100=90).
[0141] In some embodiments, the uniformity in DOR can be measured using standard methods such as. but not limiting to, DOR slope, normalized median depth of read (nmDOR), or breadth of read (BOR). DOR slope represents the slope of the line in the linear portion of a list of loci sorted in descending DOR order. Closer to zero is better, as it represents a flat line. In some embodiments, the uniformity in DOR can be measured using the percent of reads in the 90th- 95th percentile. For this measurement, the loci are sorted in descending DOR order. In illustrative embodiments, a DOR distribution using the 90th-95th percentile contains 5 percent of reads. The reads of all loci between the 90th percentile and 95th percentile can be counted and divided by the total reads for all loci.
[0142] In some embodiments, the magnitude of the DOR slope can be less than 0.005, 0.001, 0.0005, 0.0001, 0.00005, 0.00001, 0.000005, or 0.000001. The magnitude of the DOR slope can be between 0 and 0.005, such as 0.000001 to 0.005, such as between 0.000005 to 0.00001, 0.00001 to 0.00005, 0.00005 to 0.0001, 0.0001 to 0.0005, 0.0005 to 0.001, or 0.001 to 0.005. The percent of reads in the 90th-95th percentile can be between 0.2 and 9 percent, such asAtty. Dkt. No.: N.059.W0.01 between 0.2 to 8 percent, 0.2 to 7 percent, 0.2 to 6 percent, 0.4 to 9 percent, 0.4 to 8 percent, 0.4 to 7 percent, 0.4 to 6 percent, 1 to 9 percent, 1 to 8 percent, 1 to 7 percent, 1 to 6 percent, 2 to 9 percent, 2 to 8 percent, 2 to 7 percent, 2 to 6 percent, 3 to 9 percent, 3 to 8 percent. 3 to 7 percent, 3 to 6 percent, 0.2 to 1.0 percent, 1 to 2 percent, 2 to 3 percent, 2 to 4 percent, 3 to 4 percent, 4 to 5 percent, 5 to 6 percent, or 6 to 8 percent, or 7 to 9 percent. In some embodiments of methods herein, the method can produce a composition comprising at least 100 different amplicons (e.g., at least 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000. 50,000, 75,000. or 100,000 non-identical amplicons) with the magnitude of the DOR slope in any of the ranges herein, or with a percent of reads in the 90th-95th percentile in any of the ranges herein. In some embodiments, different amplicons can range in between 100 to 500,000, 100 to 400,000, 100 to 300,000, 100 to 200,000, 100 to 100,000, 100 to 75,000, 100 to 50,000, 100 to 40,000, 100 to 30,000, 100 to 25,000, 100 to 20,000, or 100 to 15,000 non-identical amplicons.
[0143] In some embodiments, the high-throughput sequencing is performed with a median depth of read of at least 100, at least 200, at least 500, at least 1,000, at least 2,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 60,000, at least 70,000, 1,000-200,000, 2,000-200,000, 5,000-200,000, 10,000-150,000, 20,000-150,000, 30,000-150,000, or 40,000-100,000 per target variant locus.
[0144] Determining Fetal Genotypes from Sequence Reads
[0145] In some embodiments, the method described herein further comprises determining fetal genotypes of one or more of the target genes from the sequence reads. In some embodiments, the method comprises determining fetal genotypes of one or more target variant loci at which the pregnant person is a heterozygous carrier of a pathogenic variant. The observed read counts or molecule counts supporting the pathogenic allele and the reference allele at the pathogenic variant can be compared to the expected allele counts for different possible fetal genotypes. The fetal genotype that is most likely to produce the observed allele counts can be computed given fetal fraction computed using common SNP loci. In some embodiments a bias model generated for a specific pathogenic allele can be used to improve the accuracy of calling the fetal genotypes.Atty. Dkt. No.: N.059.W0.01
[0146] In some embodiments, the method comprises determining fetal genotypes of one or more target variant loci at which the pregnant person is a homozygous wildtype, wherein the pregnant person is a heterozygous earner of a pathogenic variant at a different locus of the same target gene. Any reads or molecules with the pathogenic allele at these loci can be indicative of a fetal haplotype inherited from the father, and indicative of fetus inheriting a defective copy of the gene from the father.
[0147] In some embodiment, the method described herein comprises calling a fetus having homozygous recessive mutations in a target gene associated with an autosomal recessive disease. In some embodiment, the method described herein comprises calling a fetus as having compound heterozygous mutations in a target gene associated with an autosomal recessive disease. Evidence of fetus inheriting a pathogenic allele at a locus at which the pregnant person is homozygous wildtype indicates the possibility of the fetus being compound heterozygous for the recessive disorder. Given that the paternal copy of the gene is affected, and the pregnant person is a carrier for the recessive disorder, the probability that the fetus inherited two defective copies is 50%. In some embodiments, the fetus can be called high risk for the disorder just based on the paternally inherited pathogenic allele. In some embodiments, the fetal genotype call at the maternal heterozygous call can be used to adjust the risk status of the fetus. A heterozygous or homozygous mutant fetal genotype call at the maternal heterozygous locus can result in an increased risk classification for a compound heterozygous call, whereas a homozygous wildtype fetal genotype call can result in a reduced fetal risk call. In the absence of any paternally inherited pathogenic alleles at maternal homozygous loci, the risk calculation can be based on the fetal genotype call at the maternal heterozygous locus.
[0148] In some embodiments, the maternal phased haplotypes can be used to improve the accuracy and confidence of the fetal genotype calls at the pathogenic maternal heterozygous locus. Given the maternal haplotypes, the common SNP alleles on the same haplotype as the pathogenic allele can be determined at each common SNP locus that is heterozygous in the pregnant person. The allele counts at each one of these positions can be used to compute the most likely fetal genotype, and hence the most likely haplotype the fetus inherited from the pregnant person. In some embodiments, the allele counts at all of the heterozygous loci withinAtty. Dkt. No.: N.059.W0.01 the phased haplotype, including one or more pathogenic loci, can be used to estimate the likelihood of the fetus inheriting either of the two maternal haplotypes. The distance of the pathogenic locus to the common SNP loci and the chance of recombination between the two loci can be factored in to these calculations to give higher weightage to allele counts that are closer to the pathogenic locus.
[0149] In some embodiments, the method described herein does not comprise adding a known quantity of an internal control / reference nucleic acid to the extracted cell-free DNA or DNA derived thereof prior to performing high-throughput sequencing. In some embodiments, the method described herein does not comprise adding a known quantity of an internal control / reference nucleic acid to the extracted cell-free DNA or DNA derived thereof and coamplifying the internal control nucleic acid with the extracted cell-free DNA or DNA derived thereof. In some embodiments, the internal control / reference nucleic acid is a synthetic molecule.
[0150] Relative Variant Dosage Analysis, Relative Haplotype Dosage Analysis
[0151] In some embodiments, consider a single SNP where the expected allele fraction present in the plasma is r (based on the maternal and fetal genotypes). The reference allele can be denoted as A and the variant or alternate allele can be denoted as B. The expected allele fraction can be defined as the expected fraction of B alleles in the combined maternal and fetal DNA. For maternal genotype gmand child genotype gc. the expected allele fraction can be given by equation 1, assuming that the genotypes are represented as allele fractions as well (i.e., a homozygous AA genotype is represented as an allele fraction of 0, a heterozygous or AB genotype is represented as an allele fraction of 0.5, and a homozygous BB genotype is represented as an allele fraction of 1).
[0152] Determining fetal fraction
[0153] In equation 1, f represents the fetal fraction, which is defined as the fraction of fetal DNA in the mixture of fetal and maternal DNA. In some embodiments, the fetal fractionAtty. Dkt. No.: N.059.W0.01 may be determined as follows. A maximum likelihood estimate of the fetal fraction f for a prenatal test may be derived without the use of paternal information. The fetal fraction is estimated from the set of SNPs where the maternal genotype is 0 or 1, resulting in a set of only two possible fetal genotypes. Define So as the set of SNPs with maternal genotype 0 and Si as the set of SNPs with maternal genotype 1. The possible fetal genotypes on So are 0 and 0.5, resulting in a set of possible allele fractions Ro(f)={ O.f / 2 } . Similarly, Ri(f)={ l-f / 2, 1 }. This method can be trivially extended to include SNPs where maternal genotype is 0.5, but these SNPs will be less informative due to the larger set of possible allele fractions. The observation at a SNP consists of the number of mapped reads with each allele present, naand nb, which sum to the depth of read d. In some embodiments, the depth of molecule is used to calculate the allele fraction instead of the depth of read. The simplest model for the observation likelihood is a binomial distribution which assumes that each of the d reads is drawn independently from a large pool that has allele fraction r. Define Nao and Nbo as the vectors formed by nasand Ubs for SNPs s in So, and Naiand Nbi similarly for Si. The probability for observing Nao and Nbo reads for alleles A and B across SNPs in So is given by the following.
[0154] The probability P( i ,M>i I r) for observing Naiand Nbi reads for alleles A and B across SNPs in Si is defined similarly. The maximum likelihood estimate fAof f is defined by equation 2. r = arg maxi P(Nao , / Vho I / )P( i , i I / ) (2)
[0155] The argmax in equation 2 can be calculated over a range of values of f. In some embodiments, the range is a grid of equally spaced values in the interval [0, 0.5].
[0156] The binomial model can be extended in a number of ways. When the maternal and fetal genotypes are either all A or all B, the expected allele fraction in plasma will be 0 or 1, and the binomial probability will not be well-defined. In practice, unexpected alleles are sometimes observed. Thus, one embodiment of the method described here comprises using training data to model the rate of the unexpected allele appearing on each SNP, and using thisAtty. Dkt. No.: N.059.W0.01 model to correct the expected allele fraction. When the expected allele fraction is not 0 or 1 , the observed allele fraction may not converge with a sufficiently high depth of read to the expected allele fraction due to amplification bias or other phenomena. The allele fraction can then be modeled as a beta distribution centered at the expected allele fraction, leading to a beta-binomial distribution for P(na, nblr) which has higher variance than the binomial.
[0157] Determining the fetal genotype at loci of interest
[0158] There are seven (7) possible maternal-fetal genotype combinations at a given locus of interest, listed as follows:Mother: AA, Fetus: AA (gm= 0, gc= 0)Mother: AA, Fetus: AB (gm= 0, gc= 0.5)Mother: AB, Fetus: AA (gm= 0.5, gc= 0)Mother: AB, Fetus: AB (gm= 0.5, gc= 0.5)Mother: AB, Fetus: BB (gm= 0.5, gc= 1)Mother: BB, Fetus: AB (gm= 1, gc = 0.5)Mother: BB, Fetus: BB (gm= 1, gc= 1)
[0159] For each matemal-fetal genotype combination above, the expected allele fraction AF can be determined using equation 1 (assuming the fetal fraction is known or has been estimated as described above). In some embodiments, a modified version of equation 1, that takes into account error rates for specific alleles, is used. The modified version of equation 1 may take the following form:
[0160] Here g’cand g’mare adjusted versions of gcandm, defined as follows: g’c= [ 1 -> (1 - EB^A) when the genotype is BB,0.5 -> 0.5 - EB^A + EA- B when the genotype is AB, and0 -> EA-^B when the genotype is AA]
[0161] Here EA^B is the error rate corresponding to a true allele A being read as B, and EB^A is the error rate corresponding to a true allele B being read as A. g’m is adjusted in an identical fashion.Atty. Dkt. No.: N.059.W0.01
[0162] Fetal genotype can be determined by using a joint distribution model to create a set of expected allele fractions for the possible maternal-fetal genotype combinations above, comparing the expected allele fractions to the actual allele fractions measured and observed in the mixed sample, and choosing the maternal-fetal genotype combination whose expected allele fraction pattern most closely matches the observed allele fraction. The model for the fetal genotype call at a single locus of interest is defined as F(a. b, g’c, g’m. f) (3), or the probability of observing na=a and nn=b given the maternal and fetal genotypes, which also depends on the fetal fraction through equation 1 or its adjusted form. The functional form of F may be a binomial distribution, beta-binomial distribution, or similar functions as discussed above.F(<x&,g’c ,#’m ,f) = P(«a=a nb =b I AF’(g’c,g’m, / )) (3)
[0163] In some embodiments, the variance in the observed allele fractions at common SNPs is estimated and incorporated into (3) as a measure of how well the model fits the observed data:F(a,b,g'c,g’m,f NP) = P(na=a, nb =b I AF’(g’c,g'm,f), NP)
[0164] Here NP represents the additional variance in the data which is not captured by the model. NP may be estimated by fitting a beta-binomial distribution to the allele fractions of common SNPs, performing a constrained one-dimensional search over a range of values, and selecting the value that provides the maximum likelihood fit to the allele fractions of the common SNPs.
[0165] In some embodiments, the observed allele fraction at the locus of interest is adjusted to remove bias due to assay artifacts prior to applying the model on the observed allele fraction. Determination of bias at the locus of interest may have been previously performed using training data by quantifying the deviation of the observed allele fraction from the expected allele fraction in the training data.
[0166] In some embodiments, F(rzAg’c ,g’m ,f NP) is calculated for each of the seven possible maternal-fetal genotype combinations and then weighted by multiplying by a prior probability prior(g\ ,#’m) for the corresponding maternal-fetal genotype combination. prioFg'cAtty. Dkt. No.: N.059.W0.01, ’m) may either be uniform across all maternal -fetal genotype combinations or based on population frequencies. A normalized probability for each maternal-fetal genotype combination is calculated by summing the probabilities and dividing each probability by the sum, as follows: Normalized probability = F(a,b,g’c,g’m,f, NP) * prior(g:,g’m) / F(a,b,g’c,g’m,f, NP) * prior(g’c, ’m)
[0167] The fetal genotype can be determined to be the fetal genotype in the maternal- fetal genotype combination with the highest normalized probability. In some embodiments, a threshold is applied to the normalized probability and a call is made only if at least one maternal- fetal genotype has a normalized probability above the threshold.
[0168] In some embodiments, the maternal haplotype is separately constructed from maternal genomic DNA (e.g., buffy coat, buccal swab, saliva) and the allele fractions at SNPs in the plasma sample are used to determine the maternal haplotype inherited by the fetus. This additional information is incorporated into the probability model in (3) to improve the determination of the fetal genotype.
[0169] In one specific example, assume the total molecular depth at a particular locus is 3000, and the fetal fraction is 10%. Based on the fetal fraction, approximately 2700 molecules came from the mother and approximately 300 came from the fetus. If the mother is known as heterozygous at said locus, there are approximately 1350 molecules from each of the two alleles that are from maternal origin (e.g., assume mutation is A>C variant). Based on the fetal genotype, there are the following possibilities: (a) wildtype fetus (A / A) - all 300 reads from the fetus will carry the A allele, so overall split will be A: 1650 and C:1350; (b) heterozygous fetus (A / C) - approximately 150 reads from each of the fetal alleles, so overall split will be A: 1500 and C:1500; (c) homozygous mutant fetus (C / C) - all 300 reads from the fetus will carry the C allele, so overall split will be A: 1350 and C:1650. Given these expected numbers, a call can be made based on which one of these better explains the observed allele counts, after taking into the bias at that position.
[0170] Remapping and Normalization for SMNAtty. Dkt. No.: N.059.W0.01
[0171] In some embodiments, a remapping workflow is applied to allow for highly accurate calling of SMN 1 deletions. A custom genome reference is built by identification and masking of regions on chromosome 5 that are homologous to SMN 1 gene and surrounding common single nucleotide polymorphism (SNP) targets. The read from the initial alignment to a standard reference aligning to any of these regions or SMN 1 itself are extracted and remapped to the custom reference. In some embodiments, these reads are further processed to group by either fragment ends and / or molecular identifiers to obtain an accurate fragment count at each locus in SMN 1 and other regions in chromosome 5 with high identity to that locus. The aligned reads are then grouped to correctly identify reads coming from the same initial fragment, and processed to generate accurate molecular counts. In some embodiments, the coverage at positions in SMN1 is further normalized by using coverage at a set of preselected control loci in other chromosomes. The target loci are split into two categories: (a) positions that are completely identical in both SMN1 and SMN2, and (b) positions that have >=1 base position different between SMN1 and SMN2. The normalized coverage and allele counts in category (b) are jointly used to infer the copy numbers of both SMN1 and SMN2 in pregnant person. These copy number estimates are applied to allele counts and coverage at loci from both categories described above to estimate the copy number of SMN1 in the fetus. An exemplary process is shown in Fig. 5.
[0172] Accordingly, another aspect of the instant disclosure is directed to a method of performing an initial mapping to the standard reference genome, extracting the reads mapping to SMN1, SMN2, and other smaller regions in chromosome 5 that have high homology with parts of SMN1, and remapping these reads to a custom reference in which all regions in chromosome 5 that are homologous to SMN 1 are masked. This method enables accurate quantification of reads coming from both SMN genes and other regions enriched by the target enrichment process.
[0173] A further aspect of the instant disclosure is directed to a method of normalizing coverage at target loci SMN 1 using a set of control loci in other chromosomes that are not usually involved in copy number variations. This method improves the ability to call SMN deletions.Atty. Dkt. No.: N.059.W0.01
[0174] An additional aspect of the instant disclosure is directed to a method of using allele fractions and coverage at differential loci in SMN1 to call both SMN1 and SMN2 copy numbers in the pregnant individual simultaneously.WORKING EXAMPLES.
[0175] Example 1 - Hybrid Capture Panel for sgNIPT
[0176] A hybrid capture panel was designed for sgNIPT covering 19 genes causing autosomal recessive disorders (CFTR, HBA1, HBA2, HBB, SMN1, MEFV, GBA, PAH, HEXA, ATP7B, DHCR7, IKBKAP, CPT2. ASPA, GALC, ACADM, GAA, PKHD1, GALT). This panel include tiling probes covering (i) all pathogenic SNVs / indels of SMN1 / 2, HBA1 / 2, HBB, and CFTR, (ii) complete gene (intron and exon) involved in gene deletion for HBA1, HBA2, and HBB and any pathogenic SNVs outside of HBA1 / 2 and HBB, (iii) variants that discriminate SMN1 / 2 and HBA1 / 2, (iv) intergenic region between HBA1 and HBA2 covering common deletion breakpoints, (v) all exons for SMN1 and CFTR and any pathogenic SNVs outside the exons of CFTR. Additionally, the panel may include probes targeting common SNPs in and around each target gene for phased haplotype analysis, fetal fraction estimation, and deletion calling (e.g., based on lack of heterozygosity since SMN1 / 2 and HBA1 / 2 have no known breakpoint), probes targeting SRY regions of interest for fetal sex determination, probes targeting ancestry informative SNPs, probes targeting additional SNPs for fetal fraction estimation, and / or baits for both the reference and variant sequence of F508del in CFTR. The hybrid capture panel was designed as a 4X tiling probe set (e.g., 4 baits per base with ~90 bp overlap for ~120bp probes) or as a 2X tiling probe set (2 baits per base with ~60 bp overlap for ~120bp probes). The hybrid capture panel was used for maternal allele quantification at heterozygous loci (in mother), as well as paternal allele calling at both heterozygous and homozygous loci.
[0177] Common SNPs for each gene: From lOOOGenome project data, 5, 10, and 20 informative heterozygous SNPs per haplotype require about 100, 200, and 400 SNPs, respectively, at an average of 5% MAF. Although the expectation was to include 200 SNPs per gene for phasing in order to obtain 10 informative heterozygous SNPs at an average 5% MAF,Atty. Dkt. No.: N.059.W0.01 experimental results show that 100 SNPs can provide 25-49 SNPs per haplotype due to high MAF.
[0178] Large deletions and insertions: The current Horizon workflow uses fusion / hybrid sequences to detect large deletions, followed by looking at ratios of reads that map to the reference genome versus the fusion / hybrid sequences, which can be confirmed with Sanger sequencing. Similar design strategy was used here for large deletions with known breakpoints to cover the two wildtype coordinates of the junction plus the hybrid junction.
[0179] Example 2 - sgNIPT Library Preparation
[0180] End-Repair: (i) mixing purified cfDNA with End-Repair Master Mix (final concentrations: lOOuM dNTP; 0.3 U Klenow; 7.5 U T4-PNK, lx NEB buffer); (ii) incubating at 25C for 30 minutes followed by 20 minutes at 75C; (iii) A-tailing by adding A-tailing Master Mix (final concentrations: 1.2uM dATP; 3.75 U Klenow); and (iv) incubating at 25C for 30 minutes followed by 15 minutes at 75C.
[0181] Ligation: (i) adding 4ul barcoded forward and reverse adaptors (final concentration 1.45 pM); (ii) adding Ligation Master Mix (final concentrations: 13.6 U T4 PNK; 2,780 U T4 DNA Ligase; lx ligation buffer with 6% PEG); (iii) incubating at 20C for 4 hours; (iv) stopping the reaction by adding 5ul 0.5M EDTA; (v) storing at -20C or proceeding to DNA purification and elute in 35ul buffer or water.
[0182] Library amplification: (i) adding purified ligation product to PCR Master Mix (final concentrations: 300uM dNTP; 2 U Kapa HiFi Polymerase; 2.5uM forward and reverse primer; lx Kapa HiFi buffer); (ii) amplifying in thermocycler (3 minutes at 95C, 12 cycles of 98C, 55C, 68C for 20 seconds each, 68C for 5 minutes, hold at 4C); (iii) DNA purification and elute in 35ul buffer or water.
[0183] Example 3 - HC sgNIPT Workflow
[0184] A proof-of-concept assay workflow for validation of the sgNIPT methods described herein is shown in Fig. 2. Briefly cfDNA extracted from maternal plasma are subjectAtty. Dkt. No.: N.059.W0.01 to library preparation to generate a DNA library, which include the steps of adapter ligation (adding universal priming site and / or molecular index / barcode sequence) followed by library amplification (adding sequencing adapter and / or sample index / barcode sequence), as described in Example 2. The libraries from anywhere between 6-48 samples are pooled together. The pooled libraries are split into 6-12 individual pools and subject to target enrichment by the 19- gene panel of hybrid capture probes described in Example 1 to generate a targeted sequencing library, wherein the panel of hybrid capture probes allows 4X tiling of variant target loci and includes phasing SNPs for all genes. The targeted sequencing libraries from the 6-12 individual pools are combined into single pool and sequencing primers are added to generate a single sequencing pool. The sequencing pool is subject to high-throughput sequencing on NovaSeq S4 with >60.000 base coverage and ~40 samples per run. The experimental readout of this proof- of-concept assay workflow is described in Example 4.
[0185] Example 4 - HC sgNIPT Results
[0186] Gene-level calling strategy A (putative positive): Fig. 3A shows an exemplary embodiment of the sgNIPT calling workflow, which involves (a) calling of fetal genotypes at one or more maternal homozygous wildtype location(s) of a target gene, wherein the pregnant person is a heterozygous carrier of a pathogenic variant at one or more different location(s) of said target gene, as well as (b) calling of fetal genotypes at one or more maternal carrier / heterozygous location(s), wherein the pregnant person is a heterozygous carrier of a pathogenic variant at said heterozygous location(s).
[0187] Performance estimates were generated for two scenarios. (A) High no call rate scenario - fix no call rate without phasing at 8% - 10% for all genes and evaluate improvements in sensitivity and specificity from phasing compared to no phasing; (B) Low no call rate scenario - minimize no call rate for FF > 2.8% (NC rate is -3%) and evaluate improvements in sensitivity and specificity from phasing compared to no phasing.
[0188] In scenario A, phasing reduces NC rate from 8%- 10% to 4%-6%, and provides a boost in specificity of ~1% across genes. For CFTR and HBB, phasing additionally leads to -50% relative boost in PPV compared to no phasing. In scenario B, phasing provides a boost inAtty. Dkt. No.: N.059.W0.01 specificity of l %-4% across genes. For CFTR and HBB, phasing additionally leads to -10% absolute increase in PPV compared to no phasing.
[0189] Gene-level calling strategy B (putative negative): Strategy B limits the heterozygous call to pathogenic loci that have high carrier prevalence (in this case >20%) for the gene. A low prevalence mutation is unlikely to be inherited from both parents, so this strategy improves specificity while not reducing sensitivity by much. Fig. 3B shows an exemplary embodiment of the sgNIPT calling workflow, which involves (a) calling of fetal genotypes at one or more maternal homozygous wildtype location(s) of a target gene, wherein the pregnant person is a heterozygous carrier of a pathogenic variant at one or more different location(s) of said target gene, as well as (b) calling of fetal genotypes at one or more maternal carrier / heterozygous location(s) at high proportion pathogenic loci, wherein the pregnant person is a heterozygous carrier of a pathogenic variant at said heterozygous location(s).
[0190] Performance estimates were generated for two scenarios. (A) High no call rate scenario - fix no call rate without phasing at 8% - 10% for all genes and evaluate improvements in sensitivity and specificity from phasing compared to no phasing; (B) Low no call rate scenario - minimize no call rate for FF > 2.8% (NC rate is -3%) and evaluate improvements in sensitivity and specificity from phasing compared to no phasing.
[0191] In scenario A, phasing reduces NC rate from 6%-8% to 4%-5%, and provides a boost in specificity of 0.5%- 1.5% across genes. For CFTR and HBB, phasing additionally leads to -50% relative boost in PPV compared to no phasing. In scenario B, phasing slightly lower sensitivity compared to Strategy A due to not making homozygous recessive calls for low frequency variants phasing, and provides a boost in specificity of 0.5%- 1.5% across genes due to fewer false positives from low frequency variants.
[0192] Example 5 - LTC Panel for sgNIPT
[0193] An exemplary proof-of-concept Linked Target Capture (LTC) panel was designed for sgNIPT. The LTC targets for detection of maternal and paternal alleles are a subset of targets of the hybrid capture panel. The LTC targets span a wide range of GC content and cover different types of variants. To test for copy number changes / complex rearrangements / repeats,Atty. Dkt. No.: N.059.W0.01 the LTC targets tile across almost all of HBA1 / 2, including introns (~5kb total), except: (i) 700bp in the intergenic region around a SINE repeat with low mappability which has the additional benefit of splitting the region into <3kbp segments, and (ii) 2x200bp gaps to split the remaining two ~3kb targets in half for faster runtime of the design pipeline. Further, the LTC targets include pathogenic / likely pathogenic SNVs / indels in CFTR around and including F508del, HBA and HBB. In addition, the LTC targets include common SNPs on either side of SNVs / indels of CFTR and HBB to evaluate the recovery of germline variants with LTC and to generate phased haplotypes for calling fetal genotypes (87 total). Finally, the LTC targets include variants that are present in cell lines for assay validation (CFTR: PHE508DEL, ARG553TER, GLY551ASP; HBA: -SEA, --FIL), as well as 2 SRY regions of interest (30bp) for fetal sex determination. In total, the POC LTC sgNIPT panel covers 254 targets (4152bp in total target size), contains 960 probes, partially covers pathogenic / likely pathogenic and common variants in CFTR, HBA1 / 2, and HBB including the most common clinically relevant variants (e.g. CFTR F508del), and covers common target types including SNPs. SNVs. indels. repeats, copy number changes and complex rearrangements.
[0194] Example 6 - LTC sgNIPT Workflow and Results
[0195] FIG. 4 shows exemplary embodiments of the sgNIPT assay workflow described herein, which uses Linked Target Capture (LTC) and Probe-Dependent Primers (PDPs) to enrich target variant loci from cfDNA library.
[0196] Probes were conjugated with primers to form PDPs in two separate equally-sized sub-pools to allow differential probe boosting of each sub-pool. Conjugated PDPs were tested in LTC. Libraries were prepared from 33 ng monoDNA derived from cell lines from 3 CFTR families, where the children were homozygous mutant, homozygous reference or heterozygous for CFTR F508del and R553X mutations, in homozygous or heterozygous maternal backgrounds. Libraries were also prepared from 33 ng of single-donor healthy cfDNA.
[0197] LTC was performed using conjugated PDPs, with four replicates for each of the family combination (96 total samples). Sequencing was performed on a NovaSeq SI 300 cycle flowcell (1200 million clusters, 12 million per sample). Data were analyzed for Standard LTCAtty. Dkt. No.: N.059.W0.01QC Metrics, including on-target rate, molecular recovery, VAF accuracy, and VAF consistency between replicates. The results indicate that the probe-dependent primers function as intended and with no off-target amplification; the universal primers do not anneal and extend except when the target for the linked probe is in the fragment and the probe anneals to the target.
[0198] Mononucleosomal DNA produced from cell lines and single-donor healthy cfDNA were used as input to library preparation. The sgNIPT PDP panel was conjugated following the standard conjugation procedure. After library preparation and conjugation, LTC was run for each sample in quadruplicate (2 replicates from each of 2 library preparations). LTC consisted of three steps. The first was a 12-cycle target capture PCR applied to the libraries using PDPs to amplify all target regions, followed by a cleanup. Cleaned up material was then added to a second target capture PCR, where four cycles of PCR with PDPs were run to add the full Illumina read 1 and read 2 primer sequences to the library, followed by a second cleanup.Finally, a 12-cycle barcoding PCR using barcoded primers added sample-specific barcodes and the full Illumina library sequence. After a final cleanup, the LTC libraries were quantified, pooled and sequenced on a NovaSeq SP flowcell.
[0199] With respect to sequencing, a minimum depth of 1 million clusters per kilobase of LTC coverage per sample is generally required. The bait sizes for the sgNIPT LTC panel covers ~13k bp, meaning -13 million cluster per sample are needed. For 96 samples, a NovaSeq S I flowcell provides a minimum of 12 million clusters per sample. PE150 sequencing (150 cycles should cover ~160bp length of cell-free DNA). The 150bp reads provide the best chance for unambiguous mapping.
[0200] The POC sgNIPT LTC panel was functional and gave signal from all covered regions and an on-target rate of -70%. The LTC panel was able to detect the CFTR mutations present in all three families and was effective at differentiating affected children from unaffected children for CFTR F508del and R553X in both heterozygous and homozygous WT backgrounds at as low as a 4% fetal fraction.
[0201] Example 7Atty. Dkt. No.: N.059.W0.01
[0202] Objective: To report the first milestone readout of the EXPAND study, an ongoing prospective multicenter study evaluating a novel single gene non-invasive prenatal screening test (sgNIPT). The test uses the methods described herein to assess autosomal recessive conditions: cystic fibrosis (CFTR), spinal muscular atrophy (SMN1), alpha-thalassemia (HBA1 / HBA2), and beta-hemoglobinopathies (HBB) including sickle cell disease.
[0203] Study Design: Maternal blood was collected from pregnant carriers of a monogenic condition. Fetal / neonatal variant status was confirmed by clinical diagnostic genetic testing or postnatal buccal swab. Cell-free DNA from maternal plasma was analyzed using hybrid capture-based, next-generation short-read sequencing. Maternal genomic DNA was analyzed using long-range PCR to derive the maternal haplotype associated with the variant of interest using common single nucleotide polymorphisms. Relative variant dosage analysis, relative haplotype dosage analysis, and allelic bias correction algorithms assessed fetal risk. Single nucleotide variants, small indels, and copy number variants were evaluated.
[0204] Results: Test performance was evaluated in 90 pregnant carriers of 101 variants: 44 CFTR, 18 SMN1, 21 HBB, and 18 HBA1 / 2 variants. The median maternal age was 31 years (range 20-44). The most common race / ethnicities were White (31%), Black (28%), multiple ethnicities (11%), and Hispanic (7%). The sensitivity for affected fetuses was 91% (10 / 11) and the within-cohort PPV was 53%. The sensitivity included 5 / 5 homozygous fetuses correctly identified, including one fetus affected by homozygous non-F508del CFTR variants. The no-call rate was 5.9% (6 / 101).
[0205] Conclusions: This readout suggests the sgNIPT assay described herein can accurately assess fetal risk for recessive conditions from a maternal blood sample. The test successfully detected difficult-to-identify homozygous fetuses and variants more common in non- White ethnicities. While carrier screening of both parents remains the ACOG-recommended strategy, sgNIPT can be a useful tool when paternal testing is not feasible.
[0206] Example 8 - LTC CFTR 508del and R553X Analysis
[0207] The goal of this experiment is to detect changes in the allele fraction of CFTR mutations (F508del and R553X) produced by artificial mother-child DNA mixtures. ThisAtty. Dkt. No.: N.059.W0.01 includes measuring the coefficient of variation of allele fraction for each target locus as well as the change attributable to the fetal fraction.
[0208] Cell lines representing 3 independent Cystic Fibrosis (CF) families were used. These families represent F508del and R553X, which are a 3 bp deletion and a point mutation respectively, allowing for detection of multiple variant types. Artificial fetal fractions of 0%, 4%, 8%, and 12% were used. 4 LTC replicates were performed from each sample and fetal fraction for a total of 96 LTC reactions.
[0209] Mononucleosomal DNA derived from families of cell lines were used to create parent-child mixtures at different fetal fractions. Specifically, each mononucleosomal DNA were diluted to a concentration of 0.825 ng / uL. The following steps produced enough mixture for this experiment plus enough additional material to create 4 additional libraries. 600 uL 12% “parent / child” mixtures were prepared by adding 72 uL “child” mononucleosomal DNA to 528 uL “mother” mononucleosomal DNA. 270 uL 8% “parent / child” mixtures were prepared by adding 180 uL of 12% “parent / child” to 90 uL “mother” mononucleosomal DNA. 270 uL 4% “parent / child” mixtures were prepared by adding 90 uL of 12% “parent / child” to 180 uL “mother” mononucleosomal DNA.
[0210] These mixtures as well as healthy cfDNA controls were used as input to the standard LTC-compatible library preparation. Additionally, PDPs were prepared using the sgNIPT POC LTC panel described in Examples 5-6. LTC were performed using the POC panel PDPs. LTC libraries were quantified and pooled. These pools were then sequenced on a NovaSeq S2 flowcell and the data analyzed.
[0211] Data was analyzed to detect CFTR pathogenic variants F508Del and R553X. The LTC data were compared to relevant hybrid capture data from a parallel assay. For F508Del, LTC performed at least as well as hybrid capture in separating affected, unaffected and heterozygous children in a heterozygous parental background (Fig. 7) with clear separation beginning around 4% fetal fraction. It also provided good signal in a homozygous WT background (Fig. 8). Similar results were achieved for at R553X (Fig. 9, Fig. 10).Atty. Dkt. No.: N.059.W0.01
[0212] Tn summary, the LTC panel is effective at differentiating affected children from unaffected children for CFTR F508del and R553X in both heterozygous and homozygous WT backgrounds at as low as a 4% fetal fraction.
Claims
Atty. Dkt. No.: N.059.W0.01WHAT IS CLAIMED IS:
1. A method for preparing a non-naturally occurring composition, comprising: extracting cell-free DNA from a plasma fraction of a blood sample of a pregnant person, wherein the pregnant person is a heterozygous carrier of at least one pathogenic variant in at least one target gene associated with autosomal recessive disorders, wherein the extracted cell-free DNA comprises a mixture of maternal cell-free DNA and fetal cell-free DNA; performing targeted enrichment on the extracted cell-free DNA or DNA derived therefrom to enrich a plurality of target variant loci in a plurality of target genes and generating enriched DNA, wherein the target variant loci comprise the at least one pathogenic variant; performing high-throughput sequencing on the enriched DNA or DNA derived thereof and generating sequence reads, and determining fetal genotypes of one or more of the target genes from the sequence reads.
2. A method for preparing a non-naturally occurring composition, comprising: genotyping cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof to identify the pregnant person as a heterozygous carrier of at least one pathogenic variant in at least one target gene associated with autosomal recessive disorders; extracting cell-free DNA from a plasma fraction of a blood sample of the pregnant person, wherein the extracted cell-free DNA comprises a mixture of maternal cell-free DNA and fetal cell-free DNA; performing targeted enrichment on the extracted cell-free DNA or DNA derived therefrom to enrich a plurality of target variant loci in a plurality of target genes and generating enriched DNA, wherein the target variant loci comprise the at least one pathogenic variant; and performing high-throughput sequencing on the enriched DNA or DNA derived thereof and generating sequence reads, and determining fetal genotypes of one or more of the target genes from the sequence reads.
3. The method of 1 or 2, wherein the at least one pathogenic variant comprises a single nucleotide variant (SNV), an indel. a copy number variation (CNV), a gene fusion, a chromosomal arrangement, or a combination thereof.Atty. Dkt. No.: N.059.W0.
014. The method of any of claims 1 -3, wherein the target genes comprise one or more of CFTR, HBA1, HBA2, HBB, or SMN1.
5. The method of claim 4, wherein the autosomal recessive disorders comprise one or more of cystic fibrosis, alpha thalassemia, beta thalassemia, sickle cell disease, or spinal muscular atrophy.
6. The method of any of claims 1-5, wherein the target genes comprise one or more of MEFV, GBA, PAH, HEXA, ATP7B, DHCR7, IKBKAP, CPT2, ASPA, GALC, ACADM, GAA, PKHD1, or GALT.
7. The method of claim 6, wherein the autosomal recessive disorders comprise one or more of familial mediterranean fever, Gaucher disease, phenylalanine hydroxylase deficiency, hexosaminidase A deficiency (Tay-Sachs disease), Wilson disease, Smith-Lemil-Optiz syndrome, familial dysautonomia, carnitine palmitoyltransferase II deficiency, Canavan disease, Krabbe disease, medium chain acyl-CoA dehydrogenase deficiency, Pompe disease, autosomal recessive polycystic kidney disease, or galactosemia.
8. The method of any of claims 1-7, further comprises performing long-read sequencing on cellular DNA extracted from a blood sample of the pregnant person or a buffy-coat fraction thereof and generating phased haplotypes of the pregnant person at one or more of the target genes.
9. The method of any of claims 1-8, wherein the targeted enrichment further enriches a plurality of phasing SNP loci within and / or flanking one or more of the target genes, wherein the plurality of phasing SNP loci each has a minor allele frequency (MAF) of at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, or at least 10%.
10. The method of claim 9, wherein the plurality of phasing SNP loci are located up to 20,000 to 500,000 bases from at least one of the target variant loci.Atty. Dkt. No.: N.059.W0.0111 . The method of claim 9 or 10, wherein the plurality of phasing SNP loci comprise at least 100, at least 200, or at least 400 phasing SNP loci per target gene.
12. The method of any of claims 9-11, wherein the method comprises using (i) the phased haplotypes of the pregnant person at one or more of the target genes, and (ii) the sequence reads of the phasing SNP loci that are heterozygous in the pregnant person, to determine the most likely fetal genotype and / or the most likely haplotype the fetus inherited from the pregnant person.
13. The method of any of claims 9-12, wherein the method further comprises identifying deletion in one or more of the target genes using sequence reads at the plurality of phasing SNP loci.
14. The method of any of claims 1-13. wherein the method further comprises quantifying fetal fraction in the plasma fraction using the sequence reads.
15. The method of any of claims 1-14, wherein the method further comprises appending an adapter comprising a universal primer binding site to the extracted cell-free DNA or DNA derived therefrom and generating adapted DNA prior to the targeted enrichment.
16. The method of claim 15, wherein the adapter further comprises a molecular index / barcode sequence, and wherein sequence reads generated from the high-throughput sequencing are grouped together using the molecular index / barcode sequence.
17. The method of claim 15, wherein the adapter does not comprise a molecular index / barcode sequence, and wherein sequence reads generated from the high-throughput sequencing are grouped together using the fragment-end sequences of the extracted cell-free DNA or DNA derived therefrom.
18. The method of any of claims 15-17, wherein the method further comprises amplifying the adapted DNA using a primer that binds to the universal primer binding site and generating adapted- amplified DNA prior to the targeted enrichment.Atty. Dkt. No.: N.059.W0.0119. The method of claim 18, wherein the adapted-amplified DNA further comprises a sample index / barcode sequence and / or a sequencing primer binding site.
20. The method of any of claims 1-19, wherein the targeted enrichment comprises preforming targeted multiplex amplification or linked target capture to enrich the target loci.
21. The method of any of claims 1-19, wherein the targeted enrichment comprises preforming targeted probe capture to enrich the target loci.
22. The method of claim 21, wherein the targeted probe capture is performed using probes covering the entire exons of one or more of the target genes, and / or wherein the targeted probe capture is performed using probes covering the entire exons and introns of one or more of the target genes.
23. The method of claim 21, wherein the targeted probe capture is performed using probes covering the entire exons of CFTR.
24. The method of claim 21, wherein the targeted probe capture is performed using tiling probes providing at least 2X, at least 3X, or at least 4X tiled coverage of the entire exons of one or more of the target genes, and / or wherein the targeted probe capture is performed using tiling probes providing at least 2X, at least 3X, or at least 4X tiled coverage of the entire exons and introns of one or more of the target genes.
25. The method of claim 21, wherein the targeted probe capture is performed using tiling probes providing at least 4X tiled coverage of the entire exons of CFTR.
26. The method of any of claims 21-25, wherein the targeted probe capture is performed using hybrid capture probes of 80-200 nucleotides in length, 100-150 nucleotides in length, or 110-130 nucleotides in length.
27. The method of any of claims 1-26, wherein the high-throughput sequencing is performed with a median depth of read of at least 30,000, at least 40,000, at least 50,000, at least 60,000, at least 70,000, 30,000-150,000, or 40,000-100,000 per target variant locus.Atty. Dkt. No.: N.059.W0.0128. The method of any of claims 1 -27, wherein the method comprises determining fetal genotypes of one or more target variant loci at which the pregnant person is a heterozygous carrier of a pathogenic variant.
29. The method of any of claims 1-28, wherein the method comprises determining fetal genotypes of one or more target variant loci at which the pregnant person is a homozygous wildtype, wherein the pregnant person is a heterozygous carrier of a pathogenic variant at a different locus of the same target gene.
30. The method of any of claims 1-29, wherein the method does not comprise adding a known quantity of an internal control nucleic acid to the extracted cell-free DNA or DNA derived thereof prior to performing high-throughput sequencing.
Citation Information
Patent Citations
Linked duplex target capture
WO2017168332A1
Compositions, methods, and kits for isolating nucleic acids
WO2018156418A1
Methods for isolating nucleic acids with size selection
WO2019161244A1
Universal haplotype-based noninvasive prenatal testing for single gene diseases
CN109996894A
Compositions and methods for digital polymerase chain reaction
US20210301328A1