Non-invasive prenatal testing for autosomal recessive diseases

A non-invasive prenatal testing method using next-generation sequencing and in-silico size selection in maternal plasma enriches fetal DNA fraction, addressing the invasiveness of current tests and improving diagnostic accuracy for autosomal recessive diseases.

WO2025254838A1PCT designated stage Publication Date: 2025-12-11RGT UNIV OF CALIFORNIA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/US2025/030571
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-03
Filing Date
2025-05-22
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Current prenatal testing methods for autosomal recessive diseases like sickle cell disease and beta-thalassemia involve invasive procedures that pose risks to the fetus, necessitating a non-invasive alternative with high accuracy.

Method used

Next-generation sequencing is used to analyze maternal plasma for fetal DNA, employing probe capture and in-silico size selection to enrich fetal fraction, leveraging parental haplotypes for accurate fetal genotype prediction.

Benefits of technology

This method provides a non-invasive, high-accuracy prenatal test for autosomal recessive diseases, enhancing diagnostic precision and reducing fetal risk.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025030571_11122025_PF_FP_ABST
    Figure US2025030571_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Compositions, methods, kits, systems, and software are provided for non-invasive prenatal testing for autosomal recessive diseases. Next generation sequencing is used to sequence maternal and fetal DNA isolated from maternal plasma by probe capture. The fetal fraction of the sequencing reads for DNA isolated from maternal plasma is estimated by counting single nucleotide polymorphisms (SNPs) for which an allele is detected that is present in the paternal haplotype but absent in the maternal haplotype, based on the assumption that SNPs having a paternal allele belong to the fetal DNA. The fetal fraction is bioinformatically enriched by excluding sequencing reads over a specified length via in-silico size selection, which increases fetal genotype prediction accuracy. Parental haplotype information together with the read ratios observed at the linked SNPs is used to predict the fetal genotype at a site of a mutation linked to the autosomal recessive disease.
Need to check novelty before this filing date? Find Prior Art

Description

NON-INVASIVE PRENATAL TESTING FOR AUTOSOMAL RECESSIVE DISEASESCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims benefit of U.S. Provisional Patent Application No. 63 / 655,271 , filed June 3, 2024, which application is incorporated herein by reference in its entirety.BACKGROUND OF THE INVENTION

[0002] The beta-hemoglobinopathies, sickle cell disease (SCD) and beta-thalassemia ((3-Thal), are autosomal-recessive monogenic diseases that result from mutations in the beta-globin (HBB) gene and are the most common monogenic diseases in the world. SCD results from homozygosity for the sickle hemoglobin (HbS) variant, a Glu6Val missense mutation in the HBB gene (Piel et al., N Engl J Med (2017) 377(3): 305), while [3-Thal results from HBB mutations that reduce or prevent (3-globin synthesis (Frangoul et al., N Engl J Med (2021) 384(3): 252-260). While SCD results from a single variant, many single-nucleotide, poly-nucleotide, frameshift and deletion mutations can result in [3- Thal (Thein, Cold Spring Harb Perspect Med (2013) 3(5): a011700).

[0003] Prenatal testing for these autosomal recessive diseases using amniocentesis and chorionic villus sampling has been standard clinical practice for decades but confers some risk to the fetus (Kuliev et al., Prenat Diagn (1993) 13(3): 197-209; Report of National Institute of Child Health and Human Development Workshop on Chorionic Villus Sampling and Limb and Other Defects, October 20, 1992, Am J Obstet Gynecol, (1993) 169(1 ): 1 -6). A non-invasive test for these diseases would have significant clinical benefits (Rather et al., Heliyon (2023) 9(3): e13923).SUMMARY OF THE INVENTION

[0004] Compositions, methods, kits, systems, and software are provided for non-invasive prenatal testing for autosomal recessive diseases. Next generation sequencing (NGS) is used to sequence maternal and fetal DNA isolated from maternal plasma by probe capture. Maternal and paternal DNA is also sequenced to determine maternal and paternal haplotypes. The fetal fraction of the sequencing reads for DNA isolated from maternal plasma is estimated by counting single nucleotide polymorphisms (SNPs) for which an allele is detected that is present in the paternal haplotype but absent in the maternal haplotype, based on the assumption that SNPs having a paternal allele belong to the fetal DNA. To improve signal to noise, the fetal fraction is bioinformatically enriched by excluding sequencing reads over a specified length via in-silico size selection to separate the longer maternal sequencing reads from the shorter fetal sequencing reads, which increases fetal genotype prediction accuracy. The parental haplotype information together with the read ratios observed atthe linked SNPs is used to predict the fetal genotype at a site of a mutation linked to the autosomal recessive disease. In particular, the compositions, methods, kits, systems, and software can be used for prenatal testing for hemoglobinopathies such as sickle cell disease and thalassemia.

[0005] In one aspect, a method of non-invasive prenatal testing for diagnosing an autosomal recessive disease in a fetus is provided, the method comprising:(a) obtaining parental DNA from both parents of the fetus; (b) sequencing at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease in the parental DNA to determine paternal and maternal haplotypes; (c) obtaining a maternal blood plasma sample at a stage of prenatal development of the fetus, wherein the maternal blood plasma sample comprises a DNA mixture of maternal DNA and fetal DNA; (d) isolating the DNA mixture from the blood plasma sample using a capture probe that specifically binds to an allelic sequence in the gene linked to the autosomal recessive disease, wherein the capture probe captures the maternal DNA and the fetal DNA comprising the allelic sequence; (e) sequencing the DNA mixture after said isolating to generate a plurality of allelic sequence reads; (f) performing in silico size selection on the plurality of allelic sequence reads for the DNA mixture to increase fetal fraction of the allelic sequence reads; (g) calculating the fetal fraction after said in silico size selection, wherein the fetal fraction is calculated based on known genotypes at single nucleotide polymorphisms (SNPs) of the paternal and maternal haplotypes and observed ratios of the genotypes of the SNPs in the allelic sequence reads for the DNA mixture after said in silico size selection, wherein any SNP for which an allele is detected that is absent in the maternal haplotype and present in the paternal haplotype is assumed to belong to the fetal DNA; and (h) determining the fetal genotype at a site of a mutation linked to the autosomal recessive disease by comparing an observed ratio of genotypes for the allelic sequence reads of the DNA mixture after said in silico size selection to an expected ratio for each possible genotype of the fetal DNA based on the paternal and maternal haplotypes and the calculated fetal fraction, wherein the fetus is diagnosed as having the autosomal recessive disease if the fetal genotype is homozygous for the mutation linked to the autosomal recessive disease, wherein the fetus is diagnosed as being a carrier of the autosomal recessive disease if the fetal genotype is heterozygous for the mutation linked to the autosomal recessive disease, and wherein the fetus is diagnosed as not having the autosomal recessive disease if the fetal genotype does not have the mutation linked to the autosomal recessive disease.

[0006] In certain embodiments, the autosomal recessive disease is a monogenic autosomal recessive disease.

[0007] In certain embodiments, the autosomal recessive disease is a hemoglobinopathy such as, but not limited to sickle cell disease or thalassemia (e.g., alpha-thalassemia or beta thalassemia). Inan exemplary embodiment, in which the autosomal recessive disease is a beta-hemoglobinopathy, the gene linked to the autosomal recessive disease is a beta-globin gene (HBB) on chromosome 11 .

[0008] In certain embodiments, the allelic sequence comprises a mutation at a single-nucleotide polymorphism linked to the autosomal recessive disease.

[0009] In certain embodiments, the fetal fraction of the allelic sequence reads before performing in silico size selection is less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, or less than 3%.

[0010] In certain embodiments, the method further comprises amplifying the sequence of at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease to generate an amplicon, wherein the sequencing of step (b) comprises sequencing the amplicon to determine the paternal and maternal haplotypes. In some embodiments, the amplification of the sequence of at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease comprises performing polymerase chain reaction or isothermal nucleic acid amplification. In some embodiments, the amplicon comprises or consists of the sequence of the gene linked to the autosomal recessive disease. In some embodiments, the amplicon has a length in a range from 1 kilobase (kb) to 5 kb, 1 .5 kb to 4 kb, or 2 kb to 3 kb, including any length within these ranges such as 0.5 kb, 0.6 kb. 0.7 kb. 0.8 kb, 0.9 kb, 1 .0 kb, 1 .2 kb, 1 .4 kb, 1 .6 kb, 1 .8 kb, 2.0, 2.2 kb, 2.4 kb, 2.6 kb, 2.8 kb, 3.0 kb, 3.2 kb, 3.4 kb, 3.6 kb, 3.8 kb, 4.0 kb, 4.2 kb, 4.4 kb, 4.6 kb, 4.8 kb, or 5.0 kb. In some embodiments, the amplicon has a length of at least 1 kilobases, at least 2 kb, or at least 3 kb. In some embodiments, the amplicon has a length of 2.2 kb. In some embodiments, the sequencing of step (b) comprises performing nanopore sequencing to determine the paternal and maternal haplotypes.

[0011] In certain embodiments, the sequencing of step (e) comprises performing reversible- terminator sequencing by synthesis to generate the plurality of allelic sequence reads.

[0012] In certain embodiments, the maternal DNA and the fetal DNA from the maternal plasma is not sheared prior to performing the sequencing of step (e).

[0013] In certain embodiments, performing in silico size selection comprises excluding allelic sequence reads having a length greater than 155 bases, greater than 156 bases, greater than 157 bases, greater than 158 bases, greater than 159 bases, or greater than 160 bases.

[0014] In certain embodiments, the method further comprises amplifying the maternal DNA and the fetal DNA prior to performing the sequencing of step (e). In some embodiments, amplification of the maternal DNA and the fetal DNA comprises performing polymerase chain reaction (e.g., digital polymerase chain reaction or quantitative polymerase chain reaction), isothermal nucleic acid amplification, or clonal amplification.

[0015] In certain embodiments, the method further comprises adding adapters to the 5’ and 3’ ends of the maternal DNA and the fetal DNA prior to performing the sequencing of step (e). In some embodiments, the adapters are linear adapters, Y-adapters, stubby adapters, or hairpin adapters.

[0016] In certain embodiments, the method further comprises using a plurality of capture probes, wherein each probe specifically binds to a different allelic sequence linked to a different autosomal recessive disease to allow multiplexed prenatal testing of the fetus for different autosomal recessive diseases.

[0017] In certain embodiments, the method further comprises treating the fetus for the autosomal recessive disease if the fetus is diagnosed as having the autosomal recessive disease. In some embodiments, the treatment comprises transplanting stem cells into the fetus in utero or postnatally, or a combination thereof. For example, the stem cells may be used to tolerize the fetus for a postnatal transplant or comprise a gene-editing system to repair or activate the gene linked to the autosomal recessive disease. In some embodiments, the stem cells have a wild-type copy of the gene linked to the autosomal recessive disease. In some embodiments, the treatment comprises performing gene therapy on the fetus in utero or postnatally, or a combination thereof.

[0018] In another aspect, a computer implemented method for diagnosing an autosomal recessive disease in a fetus is provided, the computer performing steps comprising: (a) receiving maternal and paternal sequences of at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease; (b) genotyping a plurality of single nucleotide polymorphisms (SNPs) in the maternal and paternal sequences to determine paternal and maternal haplotypes; (c) receiving sequences of maternal blood plasma DNA, wherein the maternal blood plasma DNA comprises a DNA mixture of maternal DNA and fetal DNA, and wherein the sequences comprise a plurality of allelic sequence reads of an allelic sequence in a gene linked to the autosomal recessive disease; (d) performing in silica size selection on the plurality of allelic sequence reads for the DNA mixture to increase fetal fraction of the allelic sequence reads; (e) calculating the fetal fraction after said in silico size selection, wherein the fetal fraction is calculated based on known genotypes at single nucleotide polymorphisms (SNPs) of the paternal and maternal haplotypes and observed ratios of the genotypes of the SNPs in the allelic sequence reads for the DNA mixture after said in silico size selection, wherein any SNP for which an allele is detected that is absent in the maternal haplotype and present in the paternal haplotype is assumed to belong to the fetal DNA; and (f) determining the fetal genotype at a site of a mutation linked to the autosomal recessive disease by comparing an observed ratio of genotypes for the allelic sequence reads of the DNA mixture after said in silico size selection to an expected ratio for each possible genotype of the fetal DNA based on the paternal and maternal haplotypes and the calculated fetal fraction, wherein the fetus is diagnosed as having theautosomal recessive disease if the fetal genotype is homozygous for the mutation linked to the autosomal recessive disease, wherein the fetus is diagnosed as being a carrier of the autosomal recessive disease if the fetal genotype is heterozygous for the mutation linked to the autosomal recessive disease, and wherein the fetus is diagnosed as not having the autosomal recessive disease if the fetal genotype does not have the mutation linked to the autosomal recessive disease; and (g) displaying information regarding the fetal genotype.

[0019] In certain embodiments, the method further comprises storing the information regarding the fetal genotype in a database.

[0020] In certain embodiments, the method further comprises instructing a sequencer to sequence at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease in the maternal and paternal DNA.

[0021] In certain embodiments, the method further comprises instructing a sequencer to sequence the maternal DNA and fetal DNA from the maternal blood plasma.

[0022] In another aspect, a system is provided, the system comprising: (a) a storage component for storing data, wherein the storage component has instructions for diagnosing an autosomal recessive disease in a fetus stored therein; (b) a computer processor programmed to analyze maternal and paternal sequences of at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease and sequences of maternal DNA and fetal DNA from maternal blood plasma using one or more algorithms, wherein the computer processor is coupled to the storage component and configured to execute the instructions stored in the storage component in order to receive the inputted sequences, and analyze the sequences according to a computer implemented method described herein; and (c) a display component for displaying information regarding the fetal genotype.

[0023] In certain embodiments, the system further comprises a sequencer to sequence at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease in maternal and paternal DNA. In some embodiments, the sequencer can perform nanopore sequencing of said at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease to determine the paternal and maternal haplotypes.

[0024] In certain embodiments, the system further comprises a sequencer to sequence the maternal DNA and fetal DNA from the maternal blood plasma. In some embodiments, the sequencer can perform reversible-terminator sequencing by synthesis of the maternal DNA and fetal DNA to generate the plurality of allelic sequence reads.

[0025] In another aspect, a non-transitory computer-readable medium is provided comprising program instructions that, when executed by a processor in a computer, causes the processor toperform a computer implemented method, described herein, for diagnosing an autosomal recessive disease in a fetus.

[0026] In another aspect, a kit comprising the non-transitory computer-readable medium, described herein, and instructions for diagnosing an autosomal recessive disease in a fetus is provided.

[0027] In certain embodiments, the kit further comprises a capture probe that specifically binds to an allelic sequence in the gene linked to the autosomal recessive disease.

[0028] In certain embodiments, the kit further comprises primers for amplifying the maternal DNA and the fetal DNA from the maternal blood plasma and the parental DNA.

[0029] In certain embodiments, the kit further comprises adapters (e.g., for ligation to the 5’ and 3’ ends of the maternal and fetal DNA isolated from maternal plasma for amplification and / or sequencing).

[0030] In certain embodiments, the kit further comprises primers for sequencing the maternal DNA and the fetal DNA from the maternal blood plasma and the parental DNA.

[0031] In certain embodiments, the kit further comprises a plurality of capture probes, wherein each probe specifically binds to a different allelic sequence linked to a different autosomal recessive disease to allow multiplexed prenatal testing of the fetus for different autosomal recessive diseases.BRIEF DESCRIPTION OF THE DRAWINGS

[0032] FIG. 1. In-silico size selection cut-off values. The plasma library was sequenced on an Illumina MiSeq using a v2 2x 300 kit run for 175 cycles. Reads were aligned, PCR duplicates were removed and in silico size selection was applied as described in Methods.

[0033] FIG. 2. ISS and haplotype data increase confidence and accuracy of fetal genotype prediction. Inferred fetal genotypes for family IBT-42 based on plasma read ratios, ISS and HBB haplotypes, are shown in the upper panel. For each of the 12 SNPs (SNP) detected in maternal plasma, the post-deduplication read-depth (DeDup Coverage), hg19 reference sequence, hg19 chromosome 11 reference coordinate (Pos), maternal (Mother) and paternal (Father) genotype, plasma ratios (Plasma-1 ), inferred fetal genotype (InferredFetus -140bp ISS and Haplotype), and the fetal genotype confirmed via ON Min ION sequencing of the CVS (Confirmed Baby) are shown. In addition, the fetal fraction (FF) after deduplication (Dedup) and after ISS at 140bp (140bp ISS) is shown at the top for both plasma ratios. The shaded plasma ratios were used to make the prediction. The HBB disease-causing SNP (NC_000011 .9:g.5248155C>G) is highlighted. Phase between SNPs, as determined by ONT MinlON sequencing, is shown for SNPs 1 -4 and Mut in the red box, with variants on the same chromosome connected via light gray or dark gray lines. Phased SNP variants derived from ONT MinlON sequencing of the 2.2 KB HBB amplicon for both parents areshown in the lower panel. The fraction of pertinent reads is shown (Haplotype %) for each haplotype. Each SNP variant detected is shown in phase with its neighbors (Phased Variants) on the same line, with SNPs that were present among the fetal genotype reads presented in light gray or dark gray.DETAILED DESCRIPTION OF THE INVENTION

[0034] Compositions, methods, kits, systems, and software are provided for non-invasive prenatal testing for autosomal recessive diseases. Next generation sequencing is used to sequence maternal and fetal DNA isolated from maternal plasma by probe capture. The fetal fraction of the sequencing reads for DNA isolated from maternal plasma is estimated by counting single nucleotide polymorphisms (SNPs) for which an allele is detected that is present in the paternal haplotype but absent in the maternal haplotype, based on the assumption that SNPs having a paternal allele belong to the fetal DNA. The fetal fraction is bioinformatically enriched by excluding sequencing reads over a specified length via in-silico size selection, which increases fetal genotype prediction accuracy. Parental haplotype information together with the read ratios observed at the linked SNPs is used to predict the fetal genotype at a site of a mutation linked to the autosomal recessive disease.

[0035] Before the present compositions, methods, kits, systems, and software are described, it is to be understood that this invention is not limited to particular methods or compositions described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.

[0036] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range where either, neither or both limits are included in the smaller ranges is also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0037] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, some potential and preferred methods and materials are now described. All publications mentioned herein are incorporated herein by reference todisclose and describe the methods and / or materials in connection with which the publications are cited. It is understood that the present disclosure supersedes any disclosure of an incorporated publication to the extent there is a contradiction.

[0038] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible.

[0039] It must be noted that as used herein and in the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a primer" includes a plurality of such primers and reference to "the probe" includes reference to one or more probes and equivalents thereof, known to those skilled in the art, and so forth.

[0040] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed.Definitions

[0041] The term "about," particularly in reference to a given quantity, is meant to encompass deviations of plus or minus five percent.

[0042] The terms “individual", “subject”, and “patient”, are used interchangeably herein and refer to any mammalian subject, particularly humans. Mammalian subjects include human and non-human mammals such as non-human primates, including chimpanzees and other apes and monkey species; laboratory animals such as mice, rats, rabbits, hamsters, guinea pigs, and chinchillas; domestic animals such as dogs and cats; and farm animals such as sheep, goats, pigs, horses, and cows.

[0043] “Isolated” refers to an entity of interest that is in an environment different from that in which it may naturally occur. “Isolated” is meant to include entities that are within samples that are substantially enriched for the entity of interest and / or in which the entity of interest is partially or substantially purified.

[0044] "Substantially purified" generally refers to isolation of a component such as a substance (compound, nucleic acid, DNA) such that the substance comprises the majority percent of the samplein which it resides. Typically in a sample, a substantially purified component comprises at least 50%, preferably at least 80%-85%, more preferably at least 90-95% of the sample.

[0045] The terms "polymorphism," "polymorphic nucleotide," "polymorphic site" or "polymorphic nucleotide position" refer to a position in a nucleic acid that possesses the quality or character of occurring in several different forms. A nucleic acid polymorphism is characterized by two or more "alleles," or versions of the nucleic acid sequence. Typically, an allele of a polymorphism that is identical to a reference sequence is referred to as a "reference allele" and an allele of a polymorphism that is different from a reference sequence is referred to as an "alternate allele," or sometimes a "variant allele." As used herein, the term "major allele" refers to the more frequently occurring allele at a given polymorphic site, and "minor allele" refers to the less frequently occurring allele, as present in the general or study population.

[0046] The term "single nucleotide polymorphism" or "SNP" refers to a polymorphic site occupied by a single nucleotide, which is the site of variation between allelic sequences. The site is usually preceded by and followed by highly conserved sequences of the allele (e.g., sequences that vary in less than 1 / 100 or 1 / 1000 members of the populations). A single nucleotide polymorphism usually arises due to substitution of one nucleotide for another at the polymorphic site. Single nucleotide polymorphisms can also arise from a deletion of a nucleotide or an insertion of a nucleotide relative to a reference allele.

[0047] SNPs generally are described as having a minor allele frequency, which can vary between populations, but generally refers to the sequence variation (A,T,G, or C) that is less common than the major allele. The frequency can be obtained from dbSNP or other sources, or may be determined for a certain population using Hardy-Weinberg equilibrium (See for details see Eberle MA, Rieder MJ, Kruglyak L, Nickerson DA (2006) Allele Frequency Matching Between SNPs Reveals an Excess of Linkage Disequilibrium in Genic Regions of the Human Genome. PLoS Genet 2(9): e142. doi:10.1371 / journal.pgen.0020142; herein incorporated by reference).

[0048] The term "single nucleotide variation" or "SNV" refers to a DNA sequence variation, wherein a single nucleotide (adenine, thymine, cytosine, or guanine) in the genome sequence is altered.

[0049] “Providing an analysis” is used herein to refer to the delivery of an oral or written analysis (i.e., a document, a report, etc.). A written analysis can be a printed or electronic document. A suitable analysis (e.g., an oral or written report) provides any or all of the following information: identifying information of the mother and father of the fetus (name, age, etc.), a description of the maternal blood sample (e.g., date it was collected, stage of pregnancy of the mother, technique used to isolate DNA from the sample and prepare the DNA for sequencing), a description of the maternal and paternal genotypes / haplotypes, the technique used to sequence the maternal and paternal DNA(e.g., nanopore sequencing), the technique used to sequence the maternal blood plasma DNA, including the maternal DNA and fetal DNA (e.g., reversible-terminator sequencing by synthesis), the results of genotyping the fetus (e.g., whether the fetal genotype is homozygous or heterozygous for the mutation linked to the autosomal recessive disease or the fetal genotype lacks the mutation), an assessment as to whether the fetus is diagnosed as having the autosomal recessive disease (i.e., the fetal genotype is homozygous for the mutation linked to the autosomal recessive disease), the fetus is a carrier of the autosomal recessive disease (i.e., the fetal genotype is heterozygous for the mutation linked to the autosomal recessive disease), or the fetus is diagnosed as not having the autosomal recessive disease (i.e., the fetal genotype does not have the mutation linked to the autosomal recessive disease), a recommendation of whether or not to provide a treatment to the fetus for the autosomal recessive disease based on the diagnosis, etc. The report can be in any format including, but not limited to printed information on a suitable medium or substrate (e.g., paper); or electronic format. If in electronic format, the report can be in any computer readable medium, e.g., diskette, compact disk (CD), flash drive, and the like, on which the information has been recorded. In addition, the report may be present as a website address which may be used via the internet to access the information at a remote site.

[0050] The terms "treatment", "treating", "treat" and the like are used herein to generally refer to obtaining a desired pharmacologic and / or physiologic effect. The effect can be prophylactic in terms of completely or partially preventing a disease or symptom(s) thereof and / or may be therapeutic in terms of a partial or complete stabilization or cure for a disease and / or adverse effect attributable to the disease. The term “treatment" encompasses any treatment of a disease in a mammal, particularly a human, and includes: (a) preventing the disease and / or symptom(s) from occurring in a subject who may be predisposed to the disease or symptom but has not yet been diagnosed as having it; (b) inhibiting the disease and / or symptom(s), i.e., arresting their development; or (c) relieving the disease symptom(s), i.e., causing regression of the disease and / or symptom(s). Those in need of treatment include those already inflicted (e.g., those with an autosomal recessive disease) as well as those in which prevention is desired (e.g., those who have a genetic predisposition to developing an autosomal recessive disease).

[0051] A therapeutic treatment is one in which the subject is inflicted prior to administration and a prophylactic treatment is one in which the subject is not inflicted prior to administration. In some embodiments, the subject has an increased likelihood of becoming inflicted or is suspected of being inflicted prior to treatment. In some embodiments, the subject is suspected of having an increased likelihood of becoming inflicted.

[0052] The term “user” as used herein refers to a person that interacts with a device and / system disclosed herein for performing one or more steps of the presently disclosed methods. The user may be a health care practitioner, such as, a physician, a neonatologist, or a pediatrician, or a genomic specialist, a medical geneticist, or a genetic counselor.

[0053] The terms “polynucleotide,” “oligonucleotide,” “nucleic acid” and “nucleic acid molecule” are used herein to include a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, the term includes triple-, double- and single-stranded DNA, as well as triple-, double- and single-stranded RNA. It also includes modifications, such as by methylation and / or by capping, and unmodified forms of the polynucleotide. More particularly, the terms “polynucleotide,” “oligonucleotide,” “nucleic acid” and “nucleic acid molecule” include polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D-ribose), any other type of polynucleotide which is an N- or C- glycoside of a purine or pyrimidine base, and other polymers containing nonnucleotidic backbones, for example, peptide nucleic acids (PNAs), morpholino nucleic acids, locked nucleic acids (LNAs), glycol nucleic acids (GNAs), threose nucleic acids (TNAs) and hexitol nucleic acids (HNAs). and other synthetic sequence-specific nucleic acid polymers providing that the polymers contain nucleobases in a configuration which allows for base pairing and base stacking, such as is found in DNA and RNA. There is no intended distinction in length between the terms “polynucleotide,” “oligonucleotide,” “nucleic acid” and “nucleic acid molecule,” and these terms will be used interchangeably. Thus, these terms include, for example, 3'-deoxy-2',5'-DNA, oligodeoxyribonucleotide N3' P5' phosphoramidates, 2 -O-alkyl-substituted RNA, double- and singlestranded DNA, as well as double- and single-stranded RNA, DNA:RNA hybrids, and hybrids between PNAs and DNA or RNA, and also include known types of modifications, for example, labels which are known in the art, methylation, “caps,” substitution of one or more of the naturally occurring nucleotides with an analog, internucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), with negatively charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), and with positively charged linkages (e.g., aminoalklyphosphoramidates, aminoalkylphosphotriesters), those containing pendant moieties, such as, for example, proteins (including nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), those with intercalators (e.g., acridine, psoralen, etc.), those containing chelators (e.g., metals, radioactive metals, boron, oxidative metals, etc.), those containing alkylators, those with modified linkages (e.g., alpha anomeric nucleic acids, etc.), as well as unmodified forms of the polynucleotide or oligonucleotide.

[0054] The term "primer" or "oligonucleotide primer" as used herein, refers to an oligonucleotide that hybridizes to the template strand of a nucleic acid and initiates synthesis of a nucleic acid strand complementary to the template strand when placed under conditions in which synthesis of a primer extension product is induced, i.e., in the presence of nucleotides and a polymerization-inducing agent such as a DNA or RNA polymerase and at suitable temperature, pH, metal concentration, and salt concentration. The primer is preferably single-stranded for maximum efficiency in amplification, but may alternatively be double-stranded. If double-stranded, the primer can first be treated to separate its strands before being used to prepare extension products. This denaturation step is typically effected by heat, but may alternatively be carried out using alkali, followed by neutralization. Thus, a "primer" is complementary to a template, and complexes by hydrogen bonding or hybridization with the template to give a primer / template complex for initiation of synthesis by a polymerase, which is extended by the addition of covalently bonded bases linked at its 3' end complementary to the template in the process of DNA or RNA synthesis. Typically, nucleic acids are amplified using at least one set of oligonucleotide primers comprising at least one forward primer and at least one reverse primer capable of hybridizing to regions of a nucleic acid flanking the portion of the nucleic acid to be amplified.

[0055] The term “amplicon” refers to the amplified nucleic acid product of a PGR reaction or other nucleic acid amplification process (e.g., rolling circle amplification or isothermal amplification methods such as recombinase polymerase amplification (RPA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase-dependent amplification (HDA), and nicking enzyme amplification reaction (NEAR), and the like). Amplicons may comprise RNA or DNA depending on the technique used for amplification.

[0056] “Recombinant” as used herein to describe a nucleic acid molecule means a polynucleotide of genomic, cDNA, viral, semisynthetic, or synthetic origin which, by virtue of its origin or manipulation is not associated with all or a portion of the polynucleotide with which it is associated in nature. The term “recombinant” as used with respect to a protein or polypeptide means a polypeptide produced by expression of a recombinant polynucleotide. In general, the gene of interest is cloned and then expressed in transformed organisms, as described further below. The host organism expresses the foreign gene to produce the protein under expression conditions.

[0057] As used herein, a “solid support” refers to a solid surface such as a magnetic bead, latex bead, microtiter plate well, glass plate, nylon, agarose, acrylamide, and the like.

[0058] As used herein, the term “target nucleic acid region” or “target nucleic acid” denotes a nucleic acid molecule with a “target sequence” to be amplified. The target nucleic acid may be either singlestranded or double-stranded and may include other sequences besides the target sequence, whichmay not be amplified. The term “target sequence” refers to the particular nucleotide sequence of the target nucleic acid which is to be amplified. The target sequence may include a probe-hybridizing region contained within the target molecule with which a probe will form a stable hybrid under desired conditions. The “target sequence” may also include the complexing sequences to which the oligonucleotide primers complex and extended using the target sequence as a template. Where the target nucleic acid is originally single-stranded, the term “target sequence” also refers to the sequence complementary to the “target sequence” as present in the target nucleic acid. If the “target nucleic acid” is originally double-stranded, the term “target sequence” refers to both the plus (+) and minus (-) strands (or sense and anti-sense strands).

[0059] As used herein, the term "probe" refers to a polynucleotide that contains a nucleic acid sequence complementary to a nucleic acid sequence present in the target nucleic acid analyte (e.g., at SNP location). The polynucleotide regions of probes may be composed of DNA, and / or RNA, and / or synthetic nucleotide analogs. Probes may be labeled in order to detect the target sequence. Such a label may be present at the 5’ end, at the 3’ end, at both the 5’ and 3’ ends, and / or internally. The “probe” may contain at least one fluorescer and at least one quencher. Quenching of fluorophore fluorescence may be eliminated by exonuclease cleavage of the fluorophore from the oligonucleotide or by hybridization of the oligonucleotide probe to the nucleic acid target sequence. Additionally, the oligonucleotide probe will typically be derived from a sequence containing a selected SNV that lies between the sense and the antisense primers used in a nucleic acid amplification.

[0060] As used herein, the term “capture oligonucleotide” or “capture probe” refers to an oligonucleotide that contains a nucleic acid sequence complementary to a nucleic acid sequence present in a target nucleic acid analyte such that the capture probe can “capture” the target nucleic acid. One or more capture probes can be used in order to capture a single target analyte or multiple target analytes. The polynucleotide regions of a capture probe may be composed of DNA, and / or RNA, and / or synthetic nucleotide analogs. By “capture” is meant that the analyte can be separated from other components of the sample by virtue of the binding of the capture probe to the analyte. The capture probe may be associated with a solid support, either directly or indirectly, a solid surface such as a magnetic bead, latex bead, microtiter plate well, glass plate, nylon, agarose, acrylamide, and the like.

[0061] The terms “hybridize” and “hybridization” refer to the formation of complexes between nucleotide sequences which are sufficiently complementary to form complexes via Watson-Crick base pairing. Where a primer “hybridizes” with a target (template), such complexes (or hybrids) are sufficiently stable to serve the priming function required by, e.g., a DNA polymerase to initiate DNA synthesis.

[0062] It will be appreciated that the hybridizing sequences need not have perfect complementarity to provide stable hybrids. In many situations, stable hybrids will form where fewer than about 10% of the bases are mismatches, ignoring loops of four or more nucleotides. Accordingly, as used herein the term “complementary” refers to an oligonucleotide that forms a stable duplex with its “complement” under assay conditions, generally where there is about 90% or greater homology.

[0063] The term "sample" as used herein relates to a material or mixture of materials, typically, although not necessarily, in liquid form, containing one or more analytes of interest.

[0064] The term “assaying” is used herein to include the physical steps of manipulating a sample to generate data related to the sample. As will be readily understood by one of ordinary skill in the art, a sample must be “obtained” prior to assaying the sample. Thus, the term “assaying” implies that the sample has been obtained. The terms “obtained” or “obtaining” as used herein encompass the act of receiving an extracted or isolated sample. For example, a testing facility can “obtain” a sample in the mail (or via delivery, etc.) prior to assaying the sample. In some such cases, the sample was “extracted” or “isolated” from an individual by another party prior to mailing (i.e., delivery, transfer, etc.), and then “obtained” by the testing facility upon arrival of the sample. Thus, a testing facility can obtain the sample and then assay the sample, thereby producing data related to the sample.

[0065] The terms “obtained” or “obtaining” as used herein can also include the physical extraction or isolation of a sample from a subject. Accordingly, a sample can be isolated from a subject (and thus “obtained”) by the same person or same entity that subsequently assays the sample. When a sample is “extracted” or “isolated” from a first party or entity and then transferred (e.g., delivered, mailed, etc.) to a second party, the sample was “obtained” by the first party (and also “isolated” by the first party), and then subsequently “obtained” (but not “isolated”) by the second party. Accordingly, in some embodiments, the step of obtaining does not comprise the step of isolating a sample.

[0066] The “melting temperature” or “Tm” of double-stranded DNA is defined as the temperature at which half of the helical structure of DNA is lost due to heating or other dissociation of the hydrogen bonding between base pairs, for example, by acid or alkali treatment, or the like. The Tmof a DNA molecule depends on its length and on its base composition. DNA molecules rich in GC base pairs have a higher Tmthan those having an abundance of AT base pairs. Separated complementary strands of DNA spontaneously reassociate or anneal to form duplex DNA when the temperature is lowered below the Tm. The highest rate of nucleic acid hybridization occurs approximately 25 degrees C below the Tm. The Tmmay be estimated using the following relationship: Tm= 69.3 + 0.41 (GC)% (Marmur et al. (1962) J. Mol. Biol. 5:109-118).

[0067] The term “Y-adapter” refers to an adapter that has a Y-shaped structure with a singlestranded 5’ arm, a single stranded 3’ arm, and a double-stranded stem. The arms of the Y-adapter have noncomplementary sequences, and the stem is double-stranded with complementary strands of DNA that hybridize to each other. Y-adapters are attached to double-stranded polynucleotides by ligating the double-stranded stem of the Y-adapters to the 5’ and 3’ ends of the polynucleotides. The use of Y adapters allows the addition of different adapter sequences to the 5' and 3' ends of a DNA library simultaneously.

[0068] The terms “hairpin adapter”, “loop adapter”, and “hairpin loop adapter” refer to an adapter that comprises a hairpin loop structure. After ligation of a hairpin adapter to a polynucleotide, the hairpin loop can be cleaved to produce strands that have non-complementary sequences on the ends. In some cases, the loop of a hairpin adapter may contain a uracil residue, wherein the loop can be cleaved using uracil DNA glycosylase and endonuclease VIII.

[0069] “Homology” refers to the percent identity between two polynucleotide or two polypeptide moieties. Two nucleic acid, or two polypeptide sequences are “substantially homologous” to each other when the sequences exhibit at least about 50% sequence identity, preferably at least about 75% sequence identity, more preferably at least about 80%-85% sequence identity, more preferably at least about 90% sequence identity, and most preferably at least about 95%-98% sequence identity over a defined length of the molecules. As used herein, substantially homologous also refers to sequences showing complete identity to the specified sequence.

[0070] The terms "modification" or "alteration" as used herein in relation to a position or amino acid mean that the amino acid in the specific position has been modified compared to the amino acid of the wild-type protein.

[0071] A "substitution" means that an amino acid residue is replaced by another amino acid residue. For example, the term "substitution" refers to the replacement of an amino acid residue by another selected from the naturally-occurring standard 20 amino acid residues, rare naturally occurring amino acid residues (e.g. hydroxyproline, hydroxylysine, allohydroxylysine, 6-N-methylysine, N- ethylglycine, N-methylglycine, N-ethylasparagine, allo-isoleucine, N-methylisoleucine, N- methylvaline, pyroglutamine, aminobutyric acid, ornithine, norleucine, norvaline), and non-naturally occurring amino acid residue, often made synthetically, (e.g. cyclohexyl-alanine).

[0072] Amino acids may be represented by their one-letter or three-letters code according to the following nomenclature: A: alanine (Ala); C: cysteine (Cys); D: aspartic acid (Asp); E: glutamic acid (Glu); F: phenylalanine (Phe); G: glycine (Gly); H: histidine (His); I: isoleucine (lie); K: lysine (Lys); L: leucine (Leu); M: methionine (Met); N: asparagine (Asn); P: proline (Pro); Q: glutamine (Gin); R:arginine (Arg); S: serine (Ser); T: threonine (Thr); V: valine (Vai); W: tryptophan (Trp) and Y: tyrosine (Tyr).

[0073] A substitution can be a conservative or non-conservative substitution. Examples of conservative substitutions are within the groups of basic amino acids (arginine, lysine and histidine), acidic amino acids (glutamic acid and aspartic acid), polar amino acids (glutamine, asparagine and threonine), hydrophobic amino acids (methionine, leucine, isoleucine, cysteine and valine), aromatic amino acids (phenylalanine, tryptophan and tyrosine), and small amino acids (glycine, alanine and serine).

[0074] As used herein, the terms "sequence identify" or "identity" refer to the number (or fraction expressed as a percentage %) of matches (identical amino acid residues) between two polypeptide sequences. The sequence identity is determined by comparing the sequences when aligned so as to maximize overlap and identity while minimizing sequence gaps. In particular, sequence identity may be determined using any of a number of mathematical global or local alignment algorithms, depending on the length of the two sequences. Sequences of similar lengths are aligned using a global alignment algorithm (e.g. Needleman and Wunsch algorithm; Needleman and Wunsch, 1970) which aligns the sequences optimally over the entire length, while sequences of substantially different lengths are aligned using a local alignment algorithm (e.g. Smith and Waterman algorithm (Smith and Waterman, 1981 ) or Altschul algorithm (Altschul et al., 1997; Altschul et al., 2005)). Alignment for purposes of determining percent amino acid sequence identity can be achieved in various ways that are within the skill in the art, for instance, using publicly available computer software available on internet web sites such as blast.ncbi.nlm.nih.gov / or ebi.ac.uk / Tools / emboss / . Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithm needed to achieve maximal alignment over the full length of the sequences being compared. For purposes herein, % amino acid sequence identity values refer to values generated using the pair wise sequence alignment program EMBOSS Needle, that creates an optimal global alignment of two sequences using the Needleman-Wunsch algorithm, wherein all search parameters are i5 set to default values, i.e. Scoring matrix=BLOSUM62, Gap open=10, Gap extend=0.5, End gap penalty=false, End gap open=10 and End gap extend=0.5.

[0075] Herein, the terms "peptide", "polypeptide", "protein", "enzyme", refer to a chain of amino acids linked by peptide bonds, regardless of the number of amino acids forming said chain.

[0076] “Recombinant” as used herein to describe a nucleic acid molecule means a polynucleotide of genomic, cDNA, viral, semisynthetic, or synthetic origin which, by virtue of its origin or manipulation is not associated with all or a portion of the polynucleotide with which it is associated in nature. The term “recombinant” as used with respect to a protein or polypeptide means a polypeptide producedby expression of a recombinant polynucleotide. In general, the gene of interest is cloned and then expressed in transformed organisms, as described further below. The host organism expresses the foreign gene to produce the protein under expression conditions.

[0077] The terms "connected" or "coupled" are used in an operational sense and are not necessarily limited to a direct connection or coupling. For example, two devices or components may be coupled directly, or via one or more intermediary media or devices. As another example, devices may be coupled in such a way that information or data can be passed between them, while not sharing any physical connection with one another. In some cases, two devices or components may be connected by a wire or wirelessly to each other.

[0078] "Diagnosis" as used herein generally includes determination as to whether a subject is likely affected by a given disease, disorder or dysfunction.

[0079] "Prognosis" as used herein generally refers to a prediction of the probable course and outcome of a clinical condition or disease. A prognosis of a patient may be made, for example, based on genotyping to determine the presence of one or more alleles which are indicative of the risk of developing a disease, disorder or dysfunction. Determining the prognosis of a patient may further involve evaluating factors or symptoms of a disease that are indicative of a favorable or unfavorable course or outcome of the disease. It is understood that the term "prognosis" does not necessarily refer to the ability to predict the course or outcome of a condition with 100% accuracy. Instead, the skilled artisan will understand that the term "prognosis" refers to an increased probability that a certain course or outcome will occur; that is, that a course or outcome is more likely to occur in a patient exhibiting a given condition, when compared to those individuals not exhibiting the condition.Non-invasive Prenatal Testing for an Autosomal Recessive Disease

[0080] Reagents and methods are provided for non-invasive prenatal testing for an autosomal recessive disease in a fetus. The presence of fetal DNA in maternal blood plasma allows genotyping of the fetus by obtaining and analyzing a blood sample from the mother during pregnancy without harming the fetus. The maternal blood plasma sample, which contains a DNA mixture of maternal DNA and fetal DNA, is obtained at a stage of prenatal development of the fetus. In addition, DNA is obtained from both parents of the fetus and sequenced to determine paternal and maternal haplotypes.

[0081] A capture probe that specifically binds to an allelic sequence in a gene linked to the autosomal recessive disease is used to capture the maternal DNA and the fetal DNA in the maternal blood plasma, and the captured maternal and fetal DNA is sequenced to generate a plurality of allelic sequence reads. In silico size selection is performed on the plurality of allelic sequence reads for theDNA mixture to increase the fetal fraction of the allelic sequence reads from the DNA mixture. After in silico size selection, the fetal fraction is calculated based on the known genotypes at SNPs of the paternal and maternal haplotypes based on sequencing the parental DNA and the observed ratios of the genotypes of the SNPs in the allelic sequence reads for the DNA mixture from the maternal blood plasma. Any SNP for which an allele is detected that is absent in the maternal haplotype and present in the paternal haplotype is assumed to belong to the fetal DNA. The fetal genotype at a site of a mutation linked to the autosomal recessive disease is inferred by comparing observed ratios of genotypes for the allelic sequence reads of the DNA mixture after in silico size selection to expected ratios for possible genotypes of the fetal DNA based on the paternal and maternal haplotypes and the calculated fetal fraction.

[0082] The subject methods can be applied to prenatal testing for any autosomal recessive disease. Exemplary autosomal recessive diseases include, without limitation, hemoglobinopathies such as, but not limited to sickle cell disease (defective hemoglobin subunit beta gene (HBB) on chromosome 11 ) and thalassemia, including alpha-thalassemia (defective hemoglobin subunit alpha 1 (HBA1 ) gene and / or hemoglobin subunit alpha 2 (HBA2) gene on chromosome 16) and beta thalassemia (defective hemoglobin subunit beta gene (HBB) on chromosome 11 ), cystic fibrosis (defective cystic fibrosis transmembrane conductance regulator (CFTR) gene on chromosome 7), autosomal recessive polycystic kidney disease (defective PKHD1 ciliary IPT domain containing fibrocystin / polyductin gene (PKHD1 ) on chromosome 6), medium-chain acyl-CoA dehydrogenase deficiency (defective acyl-CoA dehydrogenase medium chain gene (ACADM) on chromosome 1 ), Gaucher disease (defective glucosylceramidase beta 1 gene (GBA) on chromosome 1 ), Tay-Sachs disease (defective hexosaminidase subunit alpha (HEXA) gene on chromosome 15), Niemann-Pick disease (defective sphingomyelin phosphodiesterase 1 gene (SMPD1 ) on chromosome 11 , defective NPC intracellular cholesterol transporter 1 gene (NPC1 ) on chromosome 18, or NPC intracellular cholesterol transporter 2 gene (NPC2) on chromosome 14), spinal muscular atrophy (defective survival of motor neuron 1 , telomeric gene (SMN1 ) on chromosome 5), familial hypercholesterolemia (defective low density lipoprotein receptor gene (LDLR) on chromosome 19 or defective apolipoprotein B gene (ApoB) on chromosome 2), Wolman disease (defective lipase A gene (LIPA also known as LAL) on chromosome 10), phenylketonuria (defective phenylalanine hydroxylase gene (PAH) on chromosome 12), and Roberts syndrome (defective establishment of sister chromatid cohesion N-acetyltransferase 2 gene (ESCO2) on chromosome 8).

[0083] Prenatal testing for an autosomal recessive disease is typically performed if the fetus is suspected of being at risk of developing an autosomal recessive disease because one or both parents have the autosomal recessive disease or are genetic carriers of the autosomal recessivedisease, or if there is a family history including relatives known to have the disease. The prenatal testing is performed to determine if the fetus has any mutated gene(s) that may lead to development of the autosomal recessive disease. In cases in which the autosomal recessive disease is a monogenic autosomal recessive disease, the fetal genotype must contain two copies of a defective gene having a mutation linked to the monogenic autosomal recessive disease for the fetus to be diagnosed with the autosomal recessive disease. An individual who carries a single copy of the defective gene is a “genetic carrier” of the autosomal recessive disease but is not considered to have the disease. If both parents are genetic carriers of the autosomal recessive disease, each parent will carry one copy of the mutated gene, and the fetus will have a 25% risk of having the autosomal recessive disease. If one parent has the autosomal recessive disease and carries two copies of the mutated gene, and the other parent is a genetic carrier of the autosomal recessive disease and carries one copy of the mutated gene, the risk of the fetus having the autosomal recessive disease is 50%. The genetics of autosomal recessive diseases involving more than one gene are somewhat more complicated, but the subject methods are also applicable to diagnosing those diseases.

[0084] A maternal blood or plasma sample can be obtained by conventional techniques. For example, blood samples can be obtained by venipuncture according to methods well known in the art. The blood sample may be obtained at any stage of pregnancy at which fetal DNA is present in the maternal blood, including the first trimester, second trimester, or third trimester of pregnancy. Fetal DNA may be detectable in maternal blood as early as 4 weeks, 5 weeks, 6 weeks, or 7 weeks of gestation. In some embodiments, the maternal blood sample is obtained at 8 weeks to 40 weeks of gestation, at 10 weeks to 15 weeks of gestation, at 10 weeks to 24 weeks of gestation, at 15 weeks to 38 weeks of gestation, or at 20 weeks to 30 weeks of gestation, including at any time within these ranges such as at 8 weeks, 9 weeks, 10 weeks, 11 weeks, 12 weeks, 13 weeks, 14 weeks, 15 weeks, 16 weeks, 17 weeks, 18 weeks, 19 weeks, 20 weeks, 21 weeks, 22 weeks, 23 weeks, 24 weeks, 25 weeks, 26 weeks, 27 weeks, 28 weeks, 29 weeks, 30 weeks, 31 weeks, 32 weeks, 33 weeks, 34 weeks, 35 weeks, 36 weeks, 37 weeks, 38 weeks, 39 weeks, or 40 weeks of gestation. Samples of parental DNA for SNP haplotyping may be collected before or after the maternal blood sample is obtained. Any type of biological sample containing DNA from the parents may be used for SNP haplotyping, but typically the sample will be a blood sample or buccal swab.

[0085] The strategy of inferring fetal genotypes by comparing the observed ratios of allelic sequence reads to the expected ratios for the possible genotypes is limited by the fetal fraction because the expected values of the read ratios is determined by the fetal fraction. At low fetal fractions (e.g., 10% or less), it becomes difficult to accurately determine the genotype of the fetus unless size selection is used to enrich for fetal sequence reads. In certain embodiments, the fetal fraction of the allelicsequence reads before performing in silica size selection is less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, or less than 3%.

[0086] With in silica size selection, the genotype of the fetus can be determined with greater accuracy. The mean fragment size of the DNA in maternal plasma is typically between 145 bp and 201 bp. The maternal DNA has a longer length (e.g., typically > 162 bp) than that of the fetal DNA (e.g., typically < 156 bp). The fraction of the allelic sequence reads that are fetal sequence reads (i.e., fetal fraction) can be enriched by in silica size selection by excluding longer allelic sequence reads belonging to the maternal DNA. In certain embodiments, fetal sequence reads are enriched by in silica size selection by excluding allelic sequence reads having a length greater than 155 bases, greater than 156 bases, greater than 157 bases, greater than 158 bases, greater than 159 bases, or greater than 160 bases. Increasing the fetal fraction by raising the length cutoff increases the accuracy of genotyping but also reduces the read count. Therefore, the length cutoff threshold may be optimized to increase accuracy of genotyping while maintaining adequate signal to noise.

[0087] To determine the parental haplotypes, a sequence of at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease is amplified to generate an amplicon. Any suitable nucleic acid amplification method may be used for this purpose, such as polymerase chain reaction or isothermal nucleic acid amplification, as described further below. In some embodiments, the amplicon comprises or consists of the sequence of the gene linked to the autosomal recessive disease. In some embodiments, the amplicon further comprises a region of the chromosome outside of the gene linked to the autosomal recessive disease that includes additional SNPs for haplotyping. In some embodiments, the amplicon has a length in a range from 1 kilobase (kb) to 5 kb, 1 .5 kb to 4 kb, or 2 kb to 3 kb, including any length within these ranges such as 0.5 kb, 0.6 kb. 0.7 kb. 0.8 kb, 0.9 kb, 1 .0 kb, 1 .2 kb, 1 .4 kb, 1 .6 kb, 1 .8 kb, 2.0, 2.2 kb, 2.4 kb, 2.6 kb, 2.8 kb, 3.0 kb, 3.2 kb, 3.4 kb, 3.6 kb, 3.8 kb, 4.0 kb, 4.2 kb, 4.4 kb, 4.6 kb, 4.8 kb, or 5.0 kb. In some embodiments, the amplicon has a length of at least 1 kb, at least 2 kb, or at least 3 kb, or less than 5 kb, less than 4 kb, less than 3 kb, or less than 2 kb. In some embodiments, the amplicon has a length of 2.2 kb. In some embodiments, nanopore sequencing is used to determine the paternal and maternal haplotypes.

[0088] The maternal DNA and the fetal DNA from the maternal blood plasma may also be amplified prior to sequencing. In some embodiments, amplification of the maternal DNA and the fetal DNA comprises performing polymerase chain reaction (e.g., digital polymerase chain reaction or quantitative polymerase chain reaction), isothermal nucleic acid amplification, or clonal amplification.

[0089] In some embodiments, the subject methods include providing an analysis of the results of the non-invasive prenatal testing for an autosomal recessive disease. A suitable analysis (e.g., anoral or written report) provides any or all of the following information: identifying information of the mother and father of the fetus (name, age, etc.), a description of the maternal blood sample (e.g., date it was collected, stage of pregnancy of the mother, technique used to isolate DNA from the sample and prepare the DNA for sequencing), a description of the maternal and paternal genotypes / haplotypes, the technique used to sequence the maternal and paternal DNA (e.g., nanopore sequencing), the technique used to sequence the maternal blood plasma DNA, including the maternal DNA and fetal DNA (e.g., reversible-terminator sequencing by synthesis), the results of genotyping the fetus (e.g., whether the fetal genotype is homozygous or heterozygous for the mutation linked to the autosomal recessive disease or the fetal genotype lacks the mutation), an assessment as to whether the fetus is diagnosed as having the autosomal recessive disease (i.e., the fetal genotype is homozygous for the mutation linked to the autosomal recessive disease), the fetus is a carrier of the autosomal recessive disease (i.e., the fetal genotype is heterozygous for the mutation linked to the autosomal recessive disease), or the fetus is diagnosed as not having the autosomal recessive disease (i.e., the fetal genotype does not have the mutation linked to the autosomal recessive disease), a recommendation of whether or not to provide a treatment to the fetus for the autosomal recessive disease based on the diagnosis, etc.

[0090] As described above, an analysis can be an oral or written report (e.g., written or electronic document). The analysis can be provided to the parents of the fetus, to a physician, to a testing facility, etc. The analysis can also be accessible as a website address via the internet. In some such cases, the analysis can be accessible by multiple different entities (e.g., the parents of the fetus, a physician, a testing facility, etc.).

[0091] In some embodiments, the fetus is treated for an autosomal recessive disease if the fetal genotype indicates that the fetus is at risk of developing the autosomal recessive disease. For example, stem cells may be administered to the fetus to tolerize the fetus for a postnatal transplant. The fetus may be treated with stem cells that comprise a gene-editing system to repair or activate the gene linked to the autosomal recessive disease or treated with stem cells having a wild-type copy of the gene linked to the autosomal recessive disease. Alternatively or additionally, gene therapy may be performed on the fetus in utero and / or postnatally to compensate for the defective gene.Capture Probes

[0092] One or more capture probes can be used to capture one or more target nucleic acid analytes such as fetal DNA and maternal DNA in maternal blood plasma or parental DNA for SNP haplotyping. Each capture probe comprises a nucleotide sequence that is complementary to a target allelicsequence such that the capture probe can specifically bind to and “capture” DNA comprising an allelic sequence. By “capture” is meant that the target analyte (e.g., fetal DNA, maternal DNA, or paternal DNA comprising an allelic sequence) can be separated from other components of the sample by virtue of the binding of the capture probe to the target analyte. The polynucleotide regions of a capture probe may be composed of DNA, and / or RNA, and / or synthetic nucleotide analogs.

[0093] One or more capture probes can be used in order to capture a single target analyte or multiple target analytes. In some embodiments, multiple capture probes are used having specificity for different allelic sequences of genes linked to different autosomal recessive diseases to allow multiplexed screening for different diseases. Any combination of autosomal recessive diseases may be tested for in parallel by using a panel of capture probes. In some embodiments, capture probes are used to simultaneously screen a fetus for two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more autosomal recessive diseases.

[0094] For example, sickle cell disease and beta thalassemia can be tested for in parallel with capture probes specific for their different allelic sequences in the HBB gene on chromosome 11. Cystic fibrosis can be added to the screening by including a capture probe specific for its allelic sequence in the CFTR gene on chromosome 7. Autosomal recessive polycystic kidney disease can be added to the screening by including a capture probe specific for its allelic sequence in the PKHD1 gene on chromosome 6. Medium-chain acyl-CoA dehydrogenase deficiency can be added to the screening by including a capture probe specific for its allelic sequence in the ACADM gene on chromosome 1. Gaucher disease can be added to the screening by including a capture probe specific for its allelic sequence in the GBA gene on chromosome 1. Tay-Sachs disease can be added to the screening by including a capture probe specific for its allelic sequence in the HEXA gene on chromosome 15. Niemann-Pick disease can be added to the screening by including a capture probe specific for its allelic sequence in the SMPD1 gene on chromosome 11 , NPC1 gene on chromosome 18, or NPC2 gene on chromosome 14. Spinal muscular atrophy can be added to the screening by including a capture probe specific for its allelic sequence in the SMN1 gene on chromosome 5. Familial hypercholesterolemia can be added to the screening by including a capture probe specific for its allelic sequence in the LDLR gene on chromosome 19 or ApoB gene on chromosome 2. Wolman disease can be added to the screening by including a capture probe specific for its allelic sequence in the LIPA gene on chromosome 10. Phenylketonuria can be added to the screening by including a capture probe specific for its allelic sequence in the PAH gene on chromosome 12. Roberts syndrome can be added to the screening by including a capture probe specific for its allelic sequence in the ESCO2 gene on chromosome 8.

[0095] A capture probe may be associated with a solid support, either directly or indirectly. In certain embodiments, the biological sample (e.g., whole blood or plasma) containing target DNA analytes is contacted with a solid support in association with capture probes. The capture probes, which may be used separately or in combination, may be associated with the solid support, for example, by covalent binding of the capture moiety to the solid support, by affinity association, hydrogen binding, or nonspecific association. The capture probe may be attached to the solid support in a variety of manners. For example, the oligonucleotide may be attached to the solid support by attachment of the 3' or 5' terminal nucleotide of the probe to the solid support. In some cases, the capture probe is attached to the solid support by a linker which serves to distance the probe from the solid support. The linker is usually at least 10-50 atoms in length, more preferably at least 15-30 atoms in length. The required length of the linker will depend on the particular solid support used. For example, a six- atom linker is generally sufficient when highly cross-linked polystyrene is used as the solid support.

[0096] A wide variety of linkers are known in the art which may be used to attach a capture probe to a solid support. The linker may be formed of any compound which does not significantly interfere with the hybridization of the target sequence to the probe attached to the solid support. The linker may comprise a homopolymeric oligonucleotide, which can be readily added to the linker by automated synthesis. The homopolymeric sequence can be either 5’ or 3' to the allelic target sequence. In some embodiments, the capture probes include a homopolymer chain, such as, for example poly A, poly T, poly G, poly C, poly II, poly dA, poly dT, poly dG, poly dC, or poly dll in order to facilitate attachment to a solid support. The homopolymer chain can be from about 10 to about 40 nucleotides in length, or preferably about 12 to about 25 nucleotides in length, or any integer within these ranges, such as for example, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, or 24 nucleotides. The homopolymer, if present, can be added to the 3' or 5' terminus of the capture probes by enzymatic or chemical methods. This addition can be made by stepwise addition of nucleotides or by ligation of a preformed homopolymer. Capture probes comprising such a homopolymer chain can be bound to a solid support comprising a complementary homopolymer. Alternatively, biotinylated capture probes can be bound to avidin- or streptavidin-coated beads. See, e.g., Chollet et al., Nucl. Acids Res. (1985) 13:1529-1541.

[0097] Alternatively, polymers such as functionalized polyethylene glycol can be used as the linker Such polymers do not significantly interfere with the hybridization of probes to the target sequence. Examples of linkages include polyethylene glycol, carbamate and amide linkages. The linkages between the solid support, the linker and the probe are preferably not cleaved during removal of base protecting groups under basic conditions at high temperature.

[0098] The solid support may take many forms including, for example, nitrocellulose reduced to particulate form and retrievable upon passing the sample medium containing the support through a sieve; nitrocellulose or the materials impregnated with magnetic particles or the like, allowing the nitrocellulose to migrate within the sample medium upon the application of a magnetic field; beads or particles which may be filtered or exhibit electromagnetic properties; and polystyrene beads which partition to the surface of an aqueous medium. Examples of types of solid supports for immobilization of the capture probe include controlled pore glass, glass plates, polystyrene, avidin-coated polystyrene beads, cellulose, nylon, acrylamide gel and activated dextran.

[0099] In some embodiments, the solid support comprises magnetic beads. The magnetic beads may contain primary amine functional groups, which facilitate covalent binding or association of the capture probes to the magnetic support particles. Alternatively, the magnetic beads may have immobilized thereon homopolymers, such as poly T or poly A sequences. The homopolymers on the solid support will generally be complementary to any homopolymer on the capture probe to allow attachment of the capture probe to the solid support by hybridization. The use of a solid support with magnetic beads allows for a one-pot method of isolation, amplification and detection as the solid support can be separated from the biological sample by magnetic means.

[0100] The magnetic beads or particles can be produced using standard techniques or obtained from commercial sources. In general, the particles or beads may comprise magnetic particles, though they can also include other magnetic metal or metal oxides, whether in impure, alloy, or composite form, as long as they have a reactive surface and exhibit an ability to react to a magnetic field. Other materials that may be used individually or in combination with iron include, but are not limited to, cobalt, nickel, and silicon.

[0101] The association of the capture probes with the solid support is initiated by contacting the solid support with a medium containing the capture probes. In an exemplary embodiment, magnetic beads containing poly dT groups are hybridized with capture probes that comprise poly dA contiguous with the capture sequence (i.e. , sequence substantially complementary to an allelic sequence of a target DNA analyte). The poly dA on the capture probe and the poly dT on the solid support hybridize thereby immobilizing or associating the capture probes with the solid support.

[0102] In certain embodiments, the capture probes are combined with a biological sample (e.g., blood or plasma) under conditions suitable for hybridization with target nucleic acids (e.g., maternal DNA, fetal DNA, or paternal DNA comprising an allelic sequence) prior to immobilization of the capture probes on a solid support. The capture probe-target nucleic acid complexes formed are then bound to the solid support. In other embodiments, a solid support with associated capture probes is brought into contact with a biological sample under hybridizing conditions. The immobilized captureprobes hybridize to the target nucleic acids present in the biological sample. Typically, hybridization of capture probes to the targets can be accomplished in approximately 15 minutes but may take as long as 3 to 48 hours.

[0103] The solid support is then separated from the biological sample, for example, by filtering, centrifugation, passing through a column, or by magnetic means. The solid support may be washed to remove unbound contaminants and transferred to a suitable container (e.g., a microtiter plate). As will be appreciated by one of skill in the art, the method of separation will depend on the type of solid support selected. Since the targets are hybridized to the capture probes immobilized on the solid support, the target strands are thereby separated from the impurities in the sample. In some cases, extraneous nucleic acids, proteins, carbohydrates, lipids, cellular debris, and other impurities may still be bound to the support, although at much lower concentrations than initially found in the biological sample. Those skilled in the art will recognize that some undesirable materials can be removed by washing the support with a washing medium. The separation of the solid support from the biological sample preferably removes at least about 70%, more preferably about 90% and, most preferably, at least about 95% or more of the non-target nucleic acids present in the sample.Adapters

[0104] Adapter oligonucleotides comprising known sequences can be added to the 5' and 3' ends of nucleic acids (e.g., fetal DNA and maternal DNA from maternal blood or parental DNA for SNP haplotyping) to facilitate amplification and / or sequencing. In some embodiments, adapters are added to the maternal DNA, fetal DNA, and / or paternal DNA before amplification or sequencing. Adapters can be designed with a primer binding site having a sequence suitable for hybridizing to primers for primer-dependent amplification and / or sequencing, Adapters may also include sites that allow nucleic acids to attach to a solid support. To facilitate multiplexing, the adapters can be barcoded.

[0105] Adapter oligonucleotides can be ligated to the ends of a nucleic acid using a ligase. Any suitable ligase may be used, including, without limitation, phage ligases such as T4 or T7 DNA ligase, archaeal ligases, or bacterial ligases. The adapters that are ligated to either end of a nucleic acid may be the same or different. Adapters may be attached to the ends of DNA by either blunt end ligation or sticky end ligation. Ligation of the ends of two DNA molecules involves formation of a phosphodiester bond between the 3'-hydroxyl group at the 3’ end of one DNA molecule with the 5'- phosphoryl group at the 5’ end of another DNA molecule. The ends of DNA molecules may be prepared for ligation by blunting of the DNA ends and phosphorylation of the 5’ end. Blunting involves removing a single-stranded overhang (e.g., which may have been created by a restriction enzyme) by adding nucleotides to the complementary strand using the overhang as a template forpolymerization, or removing the overhang using an exonuclease. DNA ends may be "blunted" to allow non-compatible ends to be joined by ligation. DNA polymerases, such as the Klenow fragment of DNA polymerase I and T4 DNA polymerase can be used to fill in nucleotides or chew back a 3’ overhang. A nuclease, such as Mung Bean Nuclease may be used for removal of a 5' overhang. In some cases, adapters are designed with a poly T overhang, which allows an adapter to be ligated to the 3'-ends of DNA having a poly A-overhang. For example, the ends of nucleic acids may be blunted and phosphorylated at the 5’ ends by treating the DNA with T4 polynucleotide kinase, T4 DNA polymerase, and the Klenow large fragment. A poly A tail can be added to the 3' ends of nucleic acids using either Taq polymerase or the Klenow large fragment.

[0106] Solid phase amplification of polynucleotides is typically performed by first ligating known adapter sequences to each end of a target polynucleotide. The double-stranded polynucleotide is then denatured to form a single-stranded template molecule that is immobilized on a solid support (e.g., the surface of a flow-cell for the Illumina platform, the membrane surface proximal to the nanopore for Oxford Nanopore Technologies nanopore sequencing platform, or beads for the Ion Torrent platform). The adapter sequence on the 3' end of the template is hybridized to an extension primer, and amplification is performed by extending the primer. In certain aspects, a sequencing platform adapter construct includes one or more nucleic acid domains selected from: a domain (e.g., a "capture site" or "capture sequence") that specifically binds to a surface- attached sequencing platform oligonucleotide (e.g., the P5 or P7 oligonucleotides attached to the surface of a flow cell in an Illumina® sequencing system); a sequencing primer binding domain (e.g., a domain to which the Read 1 or Read 2 primers of the Illumina® platform may bind); a barcode domain (e.g., a domain that uniquely identifies the sample source of the nucleic acid being sequenced to enable sample multiplexing by marking every molecule from a given sample with a specific barcode or "tag"); a barcode sequencing primer binding domain (a domain to which a primer used for sequencing a barcode binds); or any combination of such domains.

[0107] The adapters chosen for the preparation of a sequencing library should be compatible with the sequencing system to be used. Polynucleotides are incorporated into the sequencing library by ligation to sequencing adapters containing specific sequences designed to work with the sequencing platform (i.e., sequencing platform adapter domain). A sequencing platform adapter domain, when present in an adapter, may include one or more nucleic acid domains of any length and sequence suitable for the sequencing platform of interest. In some embodiments, the nucleic acid domains are from 4 to 200 nucleotides in length. For example, the nucleic acid domains may be from 4 to 100 nucleotides in length, such as from 6 to 75, from 8 to 50, or from 10 to 40 nucleotides in length. According to certain embodiments, the sequencing platform adapter construct includes a nucleicacid domain that is from 2 to 8 nucleotides in length, such as from 9 to 15, from 16 to 22, from 23 to 29, or from 30 to 36 nucleotides in length.

[0108] The nucleic acid domains may have a length and sequence that enables a polynucleotide employed by the sequencing platform of interest to specifically bind to the nucleic acid domain, e.g., for solid phase amplification and / or sequencing. The nucleotide sequences of nucleic acid domains useful for sequencing on a sequencing platform of interest may vary and / or change over time. Adapter sequences are typically provided by the manufacturer of the sequencing platform (e.g., in technical documents provided with the sequencing system and / or available on the manufacturer's website). Based on such information, the sequence of any sequencing platform adapter domains, amplification primers, and / or the like, may be designed to include all or a portion of one or more nucleic acid domains in a configuration that enables sequencing the nucleic acid insert (e.g., maternal or fetal DNA from maternal blood plasma or paternal DNA for SNP haplotyping) on the platform of interest.

[0109] Any suitable type of adapter may be ligated to the DNA to be sequenced (e.g., maternal or fetal DNA from maternal blood plasma or paternal DNA for SNP haplotyping), including, without limitation, linear adapters, Y adapters, stubby adapters, and hairpin adapters. In some embodiments, two different linear adapters are ligated to the 5’ and 3’ ends of every member of a library. A disadvantage of using linear double stranded adapters is that if two different adapters (referred to as A and B) are ligated to double-stranded DNA in a library, a mixture of ligation products is generated including about 50% incorrectly ligated products with the same adapter on both ends (A-insert-A, B- insert-B), and only 50% having correctly ligated adapters (e.g., A-insert-B or B-insert-A). The DNA inserts having A-A and B-B adapters are unable to be amplified by PCR.

[0110] Y-adapters have a Y-shaped structure with a single-stranded 5’ arm, a single stranded 3’ arm, and a double-stranded stem. The arms of the Y have noncomplementary sequences, and the stem is double-stranded with complementary strands of DNA that hybridize to each other. Y-adapters are attached to the DNA being sequenced by ligating the double-stranded stem of the Y-adapters to the 5’ and 3’ ends of the DNA. The use of Y adapters allows the addition of different adapter sequences to the 5' and 3' ends of a DNA library simultaneously. An advantage of using Y adapters rather than linear adapters is that the Y adapters provide an efficient method of attaching two different adapter sequences on the 5’ and 3’ ends of every DNA molecule in a library with a single adapter. After ligation of a Y adapter to both DNA ends, the library can be amplified using primers having sequences that are complementary to the sequences of the 5’ and 3’ arms of the Y-adapters. The use of Y adapters also allows both DNA strands to be sequenced.

[0111] In some embodiments, a hairpin adapter is used. Hairpin adapters generally comprise a double-stranded "stem" region and a single stranded "loop" region. In some embodiments, the hairpin adapter comprises one strand (i.e., one continuous strand) capable of adopting a hairpin structure, wherein the hairpin adapter comprises a self-complementary palindromic region that forms the stem and a non-complementary region that forms the loop of the hairpin adapter. In addition, hairpin adapters may comprise various components of adapters, including, without limitation, amplification priming sites, barcode sequences, and specific sequencing platform adapter domain sequences (e.g., P5 and P7 or A and P1 adapter sequences). Hairpin adapters may further comprise one or more cleavage sites capable of being cleaved under cleavage conditions. In some embodiments, a cleavage site is located in the loop region. Cleavage at a cleavage site generates two separate strands from the hairpin adapter. In some embodiments, cleavage at a cleavage site in the loop region generates a partially double stranded adapter with a double stranded stem and two unpaired strands forming a "Y" structure. Cleavage sites may comprise, for example, uracil and / or deoxyuridine bases, which may be cleaved, for example, using DNA glycosylases, endonucleases, RNAses, and the like and combinations thereof. In some embodiments, the loop of a hairpin adapter contains a uracil residue and can be cleaved using uracil DNA glycosylase and endonuclease VIII. An advantage of using hairpin adapters is that such adapters allow contiguous sequencing of both strands of a double-stranded DNA molecule by covalently attaching one strand to the other. Hairpin adapters also help to minimize the formation of adaptor-dimers during adaptor ligation.

[0112] Hairpin adapters are commercially available from New England Biolabs (Ipswich, MA) such as the NEBNext® adaptor, which has sequences that are compatible with Illumina sequencing platforms. The Oxford Nanopore Min ION sequencing platform uses a Y-adapter and a hairpin adapter. The Y-adapter is ligated to the one end of a double-stranded DNA molecule, which provides for attachment of the DNA molecule to the sequencing nanopore. The hairpin-like adapter is ligated to the other end of the double-stranded DNA molecule to allow sequencing of both strands of the DNA molecule in series. Exemplary Oxford Nanopore Y adapters are described by Karamitros et al. (2015) Nucleic Acids Res. 43(22): e152; herein incorporated by reference.Amplification of Nucleic Acids

[0113] Any primer-dependent amplification method known in the art may be used for amplification of nucleic acids (e.g., fetal DNA and maternal DNA in maternal blood or parental DNA for SNP haplotyping) including, without limitation, polymerase chain reaction (PCR), rolling circle amplification, and isothermal amplification methods such as recombinase polymerase amplification(RPA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase-dependent amplification (HDA), and nicking enzyme amplification reaction (NEAR), and the like.

[0114] In some instances, amplification may be performed by polymerase chain reaction (PCR). PCR primers should be of sufficient length to provide for hybridization to complementary template DNA under annealing conditions. The primers will generally be at least 6 bp in length, including but not limited to e.g., at least 10 bp in length, at least 15 bp in length, at least 16 bp in length, at least 17 bp in length, at least 18 bp in length, at least 19 bp in length, at least 20 bp in length, at least 21 bp in length, at least 22 bp in length, at least 23 bp in length, at least 24 bp in length, at least 25 bp in length, at least 26 bp in length, at least 27 bp in length, at least 28 bp in length, at least 29 bp in length, at least 30 bp in length, and may be as long as 60 bp in length or longer, where the length of the primers will generally range from 18 to 50 bp in length, including but not limited to, e.g., from about 20 to 35 bp in length. In some instances, the template DNA may be contacted with a single primer or a set of two primers (forward and reverse primers), depending on whether primer extension, linear or exponential amplification of the template DNA is desired. Methods of PCR that may be employed in the subject methods include but are not limited to those described in U.S. Pat. Nos.: 4,683,202; 4,683,195; 4,800,159; 4,965,188 and 5,512,462, the disclosures of which are herein incorporated by reference.

[0115] In addition to the above components, a PCR reaction mixture may include a polymerase and deoxyribonucleoside triphosphates (dNTPs). The desired polymerase activity may be provided by one or more distinct polymerase enzymes. In many embodiments, the reaction mixture includes at least a Family A polymerase, where representative Family A polymerases of interest include, but are not limited to: Thermus aquaticus polymerases, including the naturally occurring polymerase (Taq) and derivatives and homologues thereof, such as Klentaq (as described in Proc. Natl. Acad. Sci USA (1994) 91 :2216-2220, the disclosure of which is incorporated herein by reference in its entirety); Thermus thermophilics polymerases, including the naturally occurring polymerase (Tth) and derivatives and homologues thereof, and the like. In certain embodiments where the amplification reaction that is carried out is a high fidelity reaction, the rea ction mixture may further include a polymerase enzyme having 3'-5' exonuclease activity, e.g., as may be provided by a Family B polymerase, where Family B polymerases of interest include, but are not limited to: Thermococcus litoralis DNA polymerase (Vent) (e.g., as described in Perler et aL, Proc. Natl. Acad. Sci. USA (1992) 89:5577, the disclosure of which is incorporated herein by reference in its entirety); Pyrococcus species GB-D (Deep Vent); Pyrococcus furiosus DNA polymerase (Pfu) (e.g., as described in Lundberg et aL, Gene (1991 ) 108: 1-6, the disclosure of which is incorporated herein by referencein its entirety), Pyrococcus woesei (Pwo) and the like.. Generally, the reaction mixture will include four different types of dNTPs corresponding to the four naturally occurring bases are present, i.e., dATP, dTTP, dCTP and dGTP and in some instances, may include one or more modified nucleotide dNTPs.

[0116] Alternatively, a polymerase that preferentially uses dUTP rather than dTTP can be used to perform PCR. Such polymerases include archaeal family B DNA polymerases such as Nanoarchaeum equitans B DNA polymerase, which can utilize deaminated bases such as uracil and hypoxanthine and performs PCR with higher fidelity than Thermus aquaticus (Taq) DNA polymerase (e.g., as described in Choi et al. (2008) Appl. Environ. Microbiol. 74(21 ): 6563-6569; herein incorporated by reference). In addition, engineered polymerases such as Q5U Hot Start High- Fidelity DNA Polymerase from New England Biolabs (Ipswich, MA) and Phusion U DNA polymerase from Thermo Fisher Scientific (Waltham, MA), which contain a mutation in the nucleotide-binding pocket that enables these polymerases to amplify templates containing uracil and inosine bases, may be used to perform PCR with dUTP. The use of polymerases that utilize UTP is useful for preventing carryover contamination in different PCR runs. The uracil-containing amplicon products of such polymerases can be digested by a uracil-DNA glycosylase to remove residual products from previous PCR amplifications and suppress template contamination between runs.

[0117] In addition, one or more PCR additives or enhancing agents may be included to improve the yield of the amplification reaction, for example, by reducing secondary structure in a nucleic acid or mispriming events. Such additives or enhancing agents include, but are not limited to, dimethyl sulfoxide (DMSO), N,N,N-trimethylglycine (betaine), formamide, glycerol, nonionic detergents (e.g., Triton X-100, Tween 20, and Nonidet P-40 (NP-40)), 7-deaza-2'-deoxyguanosine, bovine serum albumin, T4 gene 32 protein, polyethylene glycol, 1 ,2-propanediol, and tetramethylammonium chloride.

[0118] A PCR reaction will generally be carried out by cycling the reaction mixture between appropriate temperatures for annealing, elongation / extension, and denaturation for specific times. Such temperature and times will vary and will depend on the particular components of the reaction including, e.g., the polymerase and the primers as well as the expected length of the resulting PCR product. In some instances, e.g., where nested or two-step PCR are employed the cycling-reaction may be carried out in stages, e.g., cycling according to a first stage having a particular cycling program or using particular temperature(s) and subsequently cycling according to a second stage having a particular cycling program or using particular temperature(s).

[0119] Multistep PCR processes may or may not include that addition of one or more reagents following the initiation of amplification. For example, in some instances, amplification may be initiatedby elongation with the use of a polymerase and, following an initial phase of the reaction, additional reagent(s) (e.g., one or more additional primers, additional enzymes, etc.) may be added to the reaction to facilitate a second phase of the reaction. In some instances, amplification may be initiated with a first primer or a first set of primers and, following an initial phase of the reaction, additional reagent(s) (e.g., one or more additional primers, additional enzymes, etc.) may be added to the reaction to facilitate a second phase of the reaction. In certain embodiments, the initial phase of amplification may be referred to as “preamplification”.

[0120] In particular, the subject methods are applicable to digital PCR techniques. For digital PCR, a sample containing nucleic acids is separated into a large number of partitions before performing PCR. Partitioning can be achieved in a variety of ways known in the art, for example, by use of micro well plates, capillaries, emulsions, arrays of miniaturized chambers or nucleic acid binding surfaces. Separation of the sample may involve distributing any suitable portion including up to the entire sample among the partitions. Each partition includes a fluid volume that is isolated from the fluid volumes of other partitions. The partitions may be isolated from one another by a fluid phase, such as a continuous phase of an emulsion, by a solid phase, such as at least one wall of a container, or a combination thereof. In certain embodiments, the partitions may comprise droplets disposed in a continuous phase, such that the droplets and the continuous phase collectively form an emulsion.

[0121] The partitions may be formed by any suitable procedure, in any suitable manner, and with any suitable properties. For example, the partitions may be formed with a fluid dispenser, such as a pipette, with a droplet generator, by agitation of the sample (e.g., shaking, stirring, sonication, etc.), and the like. Accordingly, the partitions may be formed serially, in parallel, or in batch. The partitions may have any suitable volume or volumes. The partitions may be of substantially uniform volume or may have different volumes. Exemplary partitions having substantially the same volume are monodisperse droplets. Exemplary volumes for the partitions include an average volume of less than about 100, 10 or 1 mL, less than about 100, 10, or 1 nL, or less than about 100, 10, or 1 pL, among others.

[0122] After separation of the sample, PCR is carried out in the partitions. The partitions, when formed, may be competent for performance of one or more reactions in the partitions. Alternatively, one or more reagents may be added to the partitions after they are formed to render them competent for reaction. The reagents may be added by any suitable mechanism, such as a fluid dispenser, fusion of droplets, or the like.

[0123] In some embodiments, nucleic acids are amplified by emulsion PCR to compartmentalize the amplification reactions of individual DNA molecules. An aqueous PCR mixture with forward and reverse primers is mixed with an oil to create the emulsion. Preferably, each droplet of water in theoil emulsion contains one bead and one molecule of template DNA (e.g., a single fetal DNA or maternal DNA molecule of the sequencing library), such that individual molecules are amplified in separate emulsion droplets. After amplification, the emulsion is broken, e.g., using isopropanol and detergent with vortexing. In some embodiments, the gene fragment library and the sequencing library are bound to magnetic beads or superparamagnetic beads prior to amplification, wherein amplification and breaking of the emulsion is followed by magnetic separation of the beads. For a description of emulsion PCR, see, e.g., Kanagal-Shamanna et al. (2016) Methods Mol Biol. 1392:33- 42, Zhu et al. (2012) Anal Bioanal Chem. 403(8):2127-43, Zhang et al. (2020) Lab Chip 20(13):2328- 2333, Siu et al. (2021) Taianta 221 :121593, Zheng et al. (2011 ) Nat. Protoc. 6(9):1367-1376, and Kojima et al. (2015) Methods Mol. Biol. 2015;1347:87-100; herein incorporated by reference.

[0124] After PCR amplification, nucleic acids can be quantified by counting the partitions that contain PCR amplicons. Partitioning of the sample allows quantification of the number of different molecules by assuming that the population of molecules follows a Poisson distribution. For a description of digital PCR methods, see, e.g., Hindson et al. (2011 ) Anal. Chem. 83(22):8604-8610; Pohl and Shih (2004) Expert Rev. Mol. Diagn. 4(1 ) :41 -47; Pekin et al. (201 1 ) Lab Chip 11 (13): 2156-2166; Pinheiro et al. (2012) Anal. Chem. 84 (2): 1003-1011 ; Day et al. (2013) Methods 59(1 ):101 -107; herein incorporated by reference in their entireties.

[0125] In some instances, amplification may be carried out under isothermal conditions, e.g., by means of isothermal amplification. Methods of isothermal amplification generally make use of enzymatic means of separating DNA strands to facilitate amplification at constant temperature, such as, e.g., strand-displacing polymerase or a helicase, thus negating the need for thermocycling to denature DNA. Any convenient and appropriate means of isothermal amplification may be employed in the subject methods including but are not limited to: recombinase polymerase amplification (RPA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicasedependent amplification (HDA), nicking enzyme amplification reaction (NEAR), and the like.

[0126] LAMP generally utilizes a plurality of primers, e.g., 4-6 primers, which may recognize a plurality of distinct regions, e.g., 6-8 distinct regions, of target DNA. Synthesis is generally initiated by a strand-displacing DNA polymerase with two of the primers forming loop structures to facilitate subsequent rounds of amplification. LAMP is rapid and sensitive. In addition, the magnesium pyrophosphate produced during the LAMP amplification reaction may, in some instances be visualized without the use of specialized equipment, e.g., by eye.

[0127] RPA combines isothermal recombinase-mediated primer targeting with strand-displacement DNA synthesis (Piepenburg et al. (2006) PLOS Biology. 4 (7): e204; herein incorporated by reference). The technique uses two primers together with a recombinase, a single-stranded DNA-binding protein, and a strand-displacing polymerase for amplification. Unlike PCR, heat is not required for melting of the DNA strands. Instead, a recombinase-primer complex is used for localized strand exchange to place oligonucleotide primers at homologous sequences of the DNA template. The single-stranded DNA-binding protein binds to the displaced template strand to prevent the primers from being ejected by branch migration. Dissociation of the recombinase leaves the 3'-end of the primer accessible to the strand displacing DNA polymerase (e.g., the large fragment of Bacillus subtilis Pol I), which catalyzes primer extension. Cyclic repetition of this process results in exponential amplification.

[0128] SDA generally involves the use of a strand-displacing DNA polymerase (e.g., Bst DNA polymerase, Large (Klenow) Fragment polymerase, Klenow Fragment (3 -5' exo-), and the like) to initiate at nicks created by a strand-limited restriction endonuclease or nicking enzyme at a site contained in a primer. In SDA, the nicking site is generally regenerated with each polymerase displacement step, resulting in exponential amplification.

[0129] HDA generally employs: a helicase which unwinds double-stranded DNA unwinding to separate strands; primers, e.g., two primers, that may anneal to the unwound DNA; and a stranddisplacing DNA polymerase for extension.

[0130] NEAR generally involves a strand-displacing DNA polymerase that initiates elongation at nicks, e.g., created by a nicking enzyme. NEAR is rapid and sensitive, quickly producing many short nucleic acids from a target sequence.

[0131] In some instances, entire amplification methods may be combined or aspects of various amplification methods may be recombined to generate a hybrid amplification method. For example, in some instances, aspects of PCR may be used, e.g., to generate the initial template or amplicon or first round or rounds of amplification, and an isothermal amplification method may be subsequently employed for further amplification. In some instances, an isothermal amplification method or aspects of an isothermal amplification method may be employed, followed by PCR for further amplification of the product of the isothermal amplification reaction. In some instances, a sample may be preamplified using a first method of amplification and may be further processed, including e.g., further amplified or analyzed, using a second method of amplification. As a non-limiting example, a sample may be preamplified by PCR and further analyzed by qPCR.

[0132] In some instances, the method further comprises monitoring the amplification of a target DNA molecule such as is performed in, e.g., real-time PCR, also referred to herein as quantitative PCR (qPCR).DNA Synthesis

[0133] Adapters, primers, and probes are readily synthesized by standard techniques, e.g., solid phase synthesis via phosphoramidite chemistry, as disclosed in U.S. Patent Nos. 4,458,066 and 4,415,732, incorporated herein by reference; Adams et al., J. Amer. Chem. Soc. (1983) 105:661 - 663, Froehler et al., Tetrahedron Lett. (1983) 24:3171 -3174, Beaucage et al., Tetrahedron (1992) 48:2223-2311 ; and Applied Biosystems User Bulletin No. 13 (1 April 1987). Other chemical synthesis methods include, for example, the phosphotriester method described by Narang et al., Meth. Enzymol. (1979) 68:90 and the phosphodiester method disclosed by Brown et al., Meth. Enzymol. (1979) 68:109. Poly(A) or poly(C), or other non-complementary nucleotide extensions may be incorporated into polynucleotides using these same methods. Hexaethylene oxide extensions may be coupled to the polynucleotides by methods known in the art. Cload et al., J. Am. Chem. Soc. (1991) 113:6324-6326; U.S. Patent No. 4,914,210 to Levenson et al.; Durand et al., Nucleic Acids Res. (1990) 18:6353-6359; and Horn et al., Tet. Lett. (1986) 27:4705-4708. Polynucleotides, adapters, and primers can be synthesized in the laboratory, for example, using an automatic synthesizer or by enzymatic DNA synthesis (EDS), as described further below.

[0134] In some embodiments, polynucleotides having only nucleotides that naturally occur in DNA such as adenine, thymine, guanine and cytosine are synthesized. In other embodiments, unnatural bases, nucleotide analogs, or unnatural base pairs (UBPs) are incorporated into polynucleotides during synthesis. Unnatural nucleobases may have different base pairing and base stacking properties than natural nucleobases. Artificial nucleic acids include peptide nucleic acids (PNAs), morpholino nucleic acids, locked nucleic acid (LNAs), glycol nucleic acids (GNAs), threose nucleic acids (TNAs) and hexitol nucleic acids (HNAs). Modifications may include, but are not limited to, N3- methylation of cytosine, 06-methylation of guanine and N-acetylation of guanine. Exemplary nucleotide analogs include, without limitation, dideoxynucleotides, deaza-nucleotides, aminoallyl nucleotides, thiol containing nucleotides, biotin-containing nucleotides, furan-modified bases, and fluorescent base analogues (e.g., 2-aminopurine, 1 ,3-diaza-2-oxophenothiazine, 3-M 1 , 6-MI, 6-MAP, and pyrrolo-dC). Oligonucleotide phosphorothioates (OPS), having an oxygen atom in the phosphate moiety replaced by a sulfur, may also be synthesized. Examples of UBPs include d5SICS and dNaM, which have hydrophobic nucleobases with two fused aromatic rings that form a (d5SICS-dNaM) complex or base pair in DNA. Various modified nucleotides, which can be used in enzymatic synthesis of polynucleotides by terminal deoxynucleotidyl transferase, or variants thereof, without the presence of a template, are described in U.S. Patent Nos. 11 ,059,849 and 5,763,594; herein incorporated by reference in their entireties. In some cases, modifications to nucleotides may block the polymerization of the nucleotide and / or allow the interaction of the nucleotide with anothermolecule such as a protein. In some cases, Polynucleotides may be chemically modified after DNA synthesis.Sequencing

[0135] Any high-throughput technique for sequencing can be used in the practice of the invention. DNA sequencing techniques include dideoxy sequencing reactions (Sanger method) using labeled terminators or primers and gel separation in slab or capillary, sequencing by synthesis using reversibly terminated labeled nucleotides, pyrosequencing, 454 sequencing, sequencing by synthesis using allele specific hybridization to a library of labeled clones followed by ligation, real time monitoring of the incorporation of labeled nucleotides during a polymerization step, polony sequencing, SOLID sequencing, and the like.

[0136] Certain high-throughput methods of sequencing comprise a step in which individual molecules are spatially isolated on a solid surface where they are sequenced in parallel. Such solid surfaces may include nonporous surfaces (such as in Solexa sequencing, e.g. Bentley et al, Nature, 456: 53-59 (2008) or Complete Genomics sequencing, e.g. Drmanac et al, Science, 327: 78-81 (2010)), arrays of wells, which may include bead- or particle-bound templates (such as with 454, e.g. Margulies et al, Nature, 437: 376-380 (2005) or Ion Torrent sequencing, e.g., U.S. patent publication 2010 / 0137143 or 2010 / 0304982), micromachined membranes (such as with SMRT sequencing, e.g. Eid et al, Science, 323: 133-138 (2009)), or bead arrays (as with SOLiD sequencing or polony sequencing, e.g. Kim et al, Science, 316: 1481 -1414 (2007)). Such methods may comprise amplifying the isolated molecules either before or after they are spatially isolated on a solid surface. Prior amplification may comprise emulsion-based amplification, such as emulsion PGR or rolling circle amplification.

[0137] Of particular interest is sequencing on the Illumina MiSeq, NextSeq, and HiSeq platforms, which use reversible-terminator sequencing by synthesis technology (see, e.g., Shen et al. (2012) BMC Bioinformatics 13:160; Junemann et al. (2013) Nat. Biotechnol. 31 (4):294-296; Glenn (2011 ) Mol. Ecol. Resour. 11 (5):759-769; Thud! et al. (2012) Brief Funct. Genomics 11 (1):3-11 ; herein incorporated by reference); the Oxford Nanopore Technologies Inc. MinlON, GridlON, and PromethlON nanopore sequencing platforms, which can be used to determine the sequences of DNA or RNA by monitoring changes in electrical current as nucleic acids are passed through a protein nanopore (see, e.g., Lu et al. (2016) Genomics Proteomics Bioinformatics 14(5):265-279, Petersen et al. (2019) J. Clin. Microbiol. 58(1 ):e01315-19, Kono et al. (2019) Dev Growth Differ. 61 (5):316-326, Deamer et al. (2016) Nat. Biotechnol. 34(5):518-24, Madoui et al. (2015) BMC Genomics 16:327, Szalay et al. (2015) Nat. Biotechnol 33, 1087-1091 ; herein incorporated byreference); the PacBIO Single Molecule, Real-Time (SMRT) sequencing platforms, including the Sequel, HiFi, and RS II sequencing platforms (see, e.g., Ardui et al. (2018) Nucleic Acids Res. 46(5):2159-2168, An et al. (2018) Genes (Basel) 9(1 ):43), Nakano et al. (2017) Hum Cell. 30(3):149- 161 ; herein incorporated by reference), the Omniome sequencing by binding (SBB®) short-read sequencing platform using high fidelity plasmonic nanohole arrays (see, e.g., Cetin et al. (2018) ACS Sens. 3(3):561 -568; herein incorporated by reference), the Gynapsys compact DNA sequencer, which uses metal oxide semiconductor (CMOS) sequencing chips for electronic data detection and sequencing by synthesis (SBS) chemistry, the Singular Genomics G4 benchtop sequencing platform, which uses SBS chemistry, and the Element Biosciences AVITI™ benchtop sequencer, which uses a modified form of SBS chemistry that reduces reagent usage.System and Computer Implemented Methods

[0138] The subject methods are implemented using a DNA sequencer. Any suitable DNA sequencer may be used to acquire sequences, including any commercially available machine. Exemplary DNA sequencers suitable for the practice of the subject methods include nanopore sequencing devices for long-read sequencing such as the MinlON portable nanopore sequencing device and the PromethlON nanopore sequencing device from Oxford Nanopore Technologies (Oxford, United Kingdom), which can be used for sequencing parental DNA to determine paternal and maternal haplotypes; and reversible-terminator sequencing by synthesis devices for short-read sequencing such as the MiSeq, NextSeq, and HiSeq platforms from Illumina, Inc. (San Diego, CA), which can be used for sequencing maternal DNA and fetal DNA from maternal blood plasma.

[0139] In certain embodiments, a computer-implemented method is used for instructing a DNA sequencer to sequence DNA, such as maternal DNA, paternal DNA, or fetal DNA. In certain embodiments, the computer-implemented method instructs a nanopore sequencer to sequence at least a portion of a maternal or paternal chromosome comprising a gene linked to the autosomal recessive disease, which the fetus is at risk of inheriting. In certain embodiments, the computer- implemented method instructs a reversible-terminator sequencing by synthesis sequencer to sequence the maternal DNA and fetal DNA from maternal blood plasma.

[0140] In some embodiments, a computer implemented method is used for diagnosing an autosomal recessive disease in a fetus based on non-invasive prenatal testing. A processor can be programmed to perform steps of a computer implemented method comprising: (a) receiving maternal and paternal sequences of at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease; (b) genotyping a plurality of single nucleotide polymorphisms (SNPs) in the maternal and paternal sequences to determine paternal and maternal haplotypes; (c) receivingsequences of maternal blood plasma DNA, wherein the maternal blood plasma DNA comprises a DNA mixture of maternal DNA and fetal DNA, and wherein the sequences comprise a plurality of allelic sequence reads of an allelic sequence in a gene linked to the autosomal recessive disease; (d) performing in silica size selection on the plurality of allelic sequence reads for the DNA mixture to increase fetal fraction of the allelic sequence reads; (e) calculating the fetal fraction after said in silica size selection, wherein the fetal fraction is calculated based on known genotypes at single nucleotide polymorphisms (SNPs) of the paternal and maternal haplotypes and observed ratios of the genotypes of the SNPs in the allelic sequence reads for the DNA mixture after said in silica size selection, wherein any SNP for which an allele is detected that is absent in the maternal haplotype and present in the paternal haplotype is assumed to belong to the fetal DNA; and (f) determining the fetal genotype at a site of a mutation linked to the autosomal recessive disease by comparing an observed ratio of genotypes for the allelic sequence reads of the DNA mixture after said in silico size selection to an expected ratio for each possible genotype of the fetal DNA based on the paternal and maternal haplotypes and the calculated fetal fraction, wherein the fetus is diagnosed as having the autosomal recessive disease if the fetal genotype is homozygous for the mutation linked to the autosomal recessive disease, wherein the fetus is diagnosed as being a carrier of the autosomal recessive disease if the fetal genotype is heterozygous for the mutation linked to the autosomal recessive disease, and wherein the fetus is diagnosed as not having the autosomal recessive disease if the fetal genotype does not have the mutation linked to the autosomal recessive disease; and (g) displaying information regarding the fetal genotype. In certain embodiments, the method further comprises storing the information regarding the fetal genotype in a database.

[0141] The methods can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, a data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine- readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or any combination thereof.

[0142] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markuplanguage document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0143] In a further aspect, the system for performing the computer implemented method, as described, may include a processor, a storage component (i.e., memory), a display component, and other components typically present in general purpose computers. In some embodiments, the processor is provided by a computer or handheld device (e.g., a cell phone or tablet). The storage component stores information accessible by the processor, including instructions that may be executed by the processor and data that may be retrieved, manipulated or stored by the processor.

[0144] The storage component includes instructions. For example, the storage component includes instructions for diagnosing an autosomal recessive disease in a fetus according to the methods described herein. The computer processor is coupled to the storage component and configured to execute the instructions stored in the storage component in order to receive maternal and paternal sequences of at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease and sequences of maternal DNA and fetal DNA from maternal blood plasma and analyze the sequences according to one or more algorithms, as described herein (see, e.g., Examples).

[0145] The processor and / or memory may be operably connected to a display device, for example, via a wired, such as a Universal Serial Bus (USB) connection, or wireless connection, such as a Bluetooth connection. Any convenient display device, such as a liquid crystal display (LCD), lightemitting diode (LED) display, plasma (PDP) display, quantum dot (QLED) display or cathode ray tube display device may be used. The display component displays information regarding the fetal genotype.

[0146] The storage component may be of any type capable of storing information accessible by the processor, such as a hard-drive, memory card, ROM, RAM, DVD, CD-ROM, USB Flash drive, write- capable, and read-only memories. The processor may be a general purpose processor, a graphics processor unit, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or moremicroprocessors in conjunction with a DSP core, or any other such configuration. Although described herein primarily with respect to digital technology, a processor can also include primarily analog components. A computing environment can include any type of computer system, including, but not limited to, a computer system based on a microprocessor, a graphics processor unit, a mainframe computer, a digital signal processor, a portable computing device, a personal organizer, a device controller, and a computational engine within an appliance, to name a few.

[0147] The steps of a method, process, or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module, engine, and associated databases can reside in memory resources such as in RAM memory, PRAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of non-transitory computer-readable storage medium, media, or physical computer storage known in the art. An exemplary storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In the alternative, the processor and the storage medium can reside as discrete components in a user terminal.

[0148] The instructions may be any set of instructions to be executed directly (such as machine code) or indirectly (such as scripts) by the processor. In that regard, the terms "instructions," "steps" and "programs" may be used interchangeably herein. The instructions may be stored in object code form for direct processing by the processor, or in any other computer language including scripts or collections of independent source code modules that are interpreted on demand or compiled in advance.

[0149] Data may be retrieved, stored or modified by the processor in accordance with the instructions. For instance, although the system is not limited by any particular data structure, the data may be stored in computer registers, in a relational database as a table having a plurality of different fields and records, XML documents, or flat files. The data may also be formatted in any computer-readable format such as, but not limited to, binary values, ASCII or Unicode. Moreover, the data may comprise any information sufficient to identify the relevant information, such as numbers, descriptive text, proprietary codes, pointers, references to data stored in other memories (including other network locations) or information which is used by a function to calculate the relevant data.

[0150] In certain embodiments, the processor and storage component may comprise multiple processors and storage components that may or may not be stored within the same physicalhousing. For example, some of the instructions and data may be stored on removable CD-ROM and others within a read-only computer chip. Some or all of the instructions and data may be stored in a location physically remote from, yet still accessible by, the processor. Similarly, the processor may comprise a collection of processors which may or may not operate in parallel.

[0151] In some embodiments, the method can be performed using a cloud computing system. In these embodiments, the maternal and paternal sequences of at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease and sequences of maternal DNA and fetal DNA from maternal blood plasma and the programming can be exported to a cloud computer, which runs the program, and returns an output to the user.Kits

[0152] Also provided are kits containing any of the reagents, devices, systems, or software described herein for use in non-invasive prenatal testing for diagnosing an autosomal recessive disease in a fetus. In some embodiments, the subject kits include reagents for practicing the subject methods such as a capture probe that specifically binds to an allelic sequence in the gene linked to the autosomal recessive disease, primers for amplifying the maternal DNA and the fetal DNA from the maternal blood plasma and the parental DNA, adapters (e.g., for ligation to the 5’ and 3’ ends of the maternal and fetal DNA isolated from maternal plasma and the parental DNA for haplotyping), primers for sequencing the maternal DNA and the fetal DNA from the maternal blood plasma and the parental DNA, and the like. In certain embodiments, the kit further comprises a plurality of capture probes, wherein each probe specifically binds to a different allelic sequence linked to a different autosomal recessive disease to allow multiplexed prenatal testing of the fetus for different autosomal recessive diseases.

[0153] Kits may further comprise one or more containers of the compositions described herein. Suitable containers for the compositions include, for example, bottles, vials, syringes, and test tubes. Containers can be formed from a variety of materials, including glass or plastic.

[0154] In certain embodiments, the kit comprises software for carrying out the computer implemented methods, described herein. In some embodiments, the kit comprises a non-transitory computer-readable medium comprising program instructions that, when executed by a processor in a computer, causes the processor to perform a computer implemented method described herein.

[0155] In certain embodiments, the kit comprises a system for non-invasive prenatal testing for diagnosing an autosomal recessive disease in a fetus wherein the system comprises: (a) a storage component for storing data, wherein the storage component has instructions for diagnosing an autosomal recessive disease in a fetus stored therein; (b) a computer processor programmed toanalyze maternal and paternal sequences of at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease and sequences of maternal DNA and fetal DNA from maternal blood plasma using one or more algorithms, wherein the computer processor is coupled to the storage component and configured to execute the instructions stored in the storage component in order to receive the inputted sequences, and analyze the sequences according to a computer implemented method described herein; and (c) a display component for displaying information regarding the fetal genotype. In certain embodiments, the system further comprises a sequencer to sequence at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease in maternal and paternal DNA. In some embodiments, the sequencer can perform nanopore sequencing of said at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease to determine the paternal and maternal haplotypes. In certain embodiments, the system further comprises a sequencer to sequence the maternal DNA and fetal DNA from the maternal blood plasma. In some embodiments, the sequencer can perform reversible-terminator sequencing by synthesis of the maternal DNA and fetal DNA to generate the plurality of allelic sequence reads.

[0156] In addition to the above components, the subject kits may further include (in certain embodiments) instructions for practicing the subject methods. For example, the kit may include instructions for performing non-invasive prenatal testing for diagnosing an autosomal recessive disease in a fetus using the methods described herein. These instructions may be present in the subject kits in a variety of forms, one or more of which may be present in the kit. One form in which these instructions may be present is as printed information on a suitable medium or substrate, e.g., a piece or pieces of paper on which the information is printed, in the packaging of the kit, in a package insert, and the like. Yet another form of these instructions is a computer readable medium, e.g., diskette, compact disk (CD), DVD, Blu-ray, flash drive, and the like, on which the information has been recorded. Yet another form of these instructions that may be present is a website address which may be used via the internet to access the information at a removed site.Examples of Non-Limiting Aspects of the Disclosure

[0157] Aspects, including embodiments, of the present subject matter described above may be beneficial alone or in combination, with one or more other aspects or embodiments. Without limiting the foregoing description, certain non-limiting aspects of the disclosure numbered 1-49 are provided below. As will be apparent to those of skill in the art upon reading this disclosure, each of the individually numbered aspects may be used or combined with any of the preceding or followingindividually numbered aspects. This is intended to provide support for all such combinations of aspects and is not limited to combinations of aspects explicitly provided below:1. A method of non-invasive prenatal testing for diagnosing an autosomal recessive disease in a fetus, the method comprising:(a) obtaining parental DNA from both parents of the fetus;(b) sequencing at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease in the parental DNA to determine paternal and maternal haplotypes;(c) obtaining a maternal blood plasma sample at a stage of prenatal development of the fetus, wherein the maternal blood plasma sample comprises a DNA mixture of maternal DNA and fetal DNA;(d) isolating the DNA mixture from the blood plasma sample using a capture probe that specifically binds to an allelic sequence in the gene linked to the autosomal recessive disease, wherein the capture probe captures the maternal DNA and the fetal DNA comprising the allelic sequence;(e) sequencing the DNA mixture after said isolating to generate a plurality of allelic sequence reads;(f) performing in silica size selection on the plurality of allelic sequence reads for the DNA mixture to increase fetal fraction of the allelic sequence reads;(g) calculating the fetal fraction after said in silico size selection, wherein the fetal fraction is calculated based on known genotypes at single nucleotide polymorphisms (SNPs) of the paternal and maternal haplotypes and observed ratios of the genotypes of the SNPs in the allelic sequence reads for the DNA mixture after said in silico size selection, wherein any SNP for which an allele is detected that is absent in the maternal haplotype and present in the paternal haplotype is assumed to belong to the fetal DNA; and(h) determining the fetal genotype at a site of a mutation linked to the autosomal recessive disease by comparing an observed ratio of genotypes for the allelic sequence reads of the DNA mixture after said in silico size selection to an expected ratio for each possible genotype of the fetal DNA based on the paternal and maternal haplotypes and the calculated fetal fraction, wherein the fetus is diagnosed as having the autosomal recessive disease if the fetal genotype is homozygous for the mutation linked to the autosomal recessive disease, wherein the fetus is diagnosed as being a carrier of the autosomal recessive disease if the fetal genotype is heterozygous for the mutation linked to the autosomal recessive disease, and wherein the fetus is diagnosed as not having theautosomal recessive disease if the fetal genotype does not have the mutation linked to the autosomal recessive disease.2. The method of aspect 1 , wherein the autosomal recessive disease is a monogenic autosomal recessive disease.3. The method of aspect 1 or 2, wherein the autosomal recessive disease is a hemoglobinopathy.4. The method of aspect 3, wherein the hemoglobinopathy is sickle cell disease or thalassemia.5. The method of aspect 4, wherein the thalassemia is alpha-thalassemia or beta thalassemia.6. The method of any one of aspects 3-5, wherein the chromosome comprising the gene linked to the autosomal recessive disease is chromosome 11 .7. The method of aspect 6, wherein the gene linked to the autosomal recessive disease is a giobin gene.8. The method of aspect 7, wherein the globin gene is an alpha-globin gene or a betaglobin gene.9. The method of any one of aspects 1 -8, wherein the allelic sequence comprises a mutation at a single-nucleotide polymorphism linked to the autosomal recessive disease.10. The method of any one of aspects 1 -9, wherein the fetal fraction of the allelic sequence reads before said performing in silico size selection is less than 10%.11. The method of aspect 10, wherein the fetal fraction of the allelic sequence reads before said performing in silico size selection is less than 5%.12. The method of any one of aspects 1 -1 1 , further comprising amplifying said at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease to generate an amplicon, wherein said sequencing of step (b) comprises sequencing the amplicon to determine the paternal and maternal haplotypes.13. The method of aspect 12, wherein said amplifying said at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease comprises performing polymerase chain reaction or isothermal nucleic acid amplification.14. The method of aspect 12 or 13, wherein the amplicon comprises or consists of the sequence of the gene linked to the autosomal recessive disease.15. The method of any one of aspects 12-14, wherein the amplicon has a length in a range from 1 kilobase to 5 kilobases.16. The method of any one of aspects 1 -15, wherein the sequencing of step (b) comprises performing nanopore sequencing to determine the paternal and maternal haplotypes.17. The method of any one of aspects 1 -16, wherein the sequencing of step (e) comprises performing reversible-terminator sequencing by synthesis to generate the plurality of allelic sequence reads.18. The method of any one of aspects 1 -17, wherein the maternal DNA and the fetal DNA from the maternal plasma is not sheared prior to performing the sequencing of step (e).19. The method of any one of aspects 1 -18, wherein said performing in silico size selection comprises excluding allelic sequence reads having a length greater than 155 bases, greater than 156 bases, greater than 157 bases, greater than 158 bases, greater than 159 bases, or greater than 160 bases.20. The method of any one of aspects 1 -19, further comprising amplifying the maternal DNA and the fetal DNA prior to performing the sequencing of step (e).21 . The method of aspect 20, wherein said amplifying the maternal DNA and the fetal DNA comprises performing polymerase chain reaction, isothermal nucleic acid amplification, or clonal amplification.22. The method of aspect 21 , wherein the polymerase chain reaction is digital polymerase chain reaction or quantitative polymerase chain reaction.23. The method of any one of aspects 1 -22, further comprising adding adapters to the 5’ and 3’ ends of the maternal DNA and the fetal DNA prior to performing the sequencing of step (e).24. The method of aspect 23, wherein the adapters are linear adapters, Y-adapters, stubby adapters, or hairpin adapters.25. The method of any one of aspects 1 -24, further comprising treating the fetus for the autosomal recessive disease if the fetus is diagnosed as having the autosomal recessive disease.26. The method of aspect 25, wherein said treating comprises transplanting stem cells into the fetus in utero or postnatally, or a combination thereof.27. The method of aspect 26, wherein the stem cells tolerize the fetus for a postnatal transplant.28. The method of aspect 26, wherein the stem cells comprise a gene-editing system to repair or activate the gene linked to the autosomal recessive disease.29. The method of aspect 26, wherein the stem cells have a wild-type copy of the gene linked to the autosomal recessive disease.30. The method of aspect 25, wherein said treating comprising performing gene therapy on the fetus in utero or postnatally, or a combination thereof.31 . The method of any one of aspects 1 -30, further comprising using a plurality of capture probes, wherein each probe specifically binds to a different allelic sequence linked to a differentautosomal recessive disease to allow multiplexed prenatal testing of the fetus for different autosomal recessive diseases.32. A computer implemented method for diagnosing an autosomal recessive disease in a fetus, the computer performing steps comprising:(a) receiving maternal and paternal sequences of at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease;(b) genotyping a plurality of single nucleotide polymorphisms (SNPs) in the maternal and paternal sequences to determine paternal and maternal haplotypes;(c) receiving sequences of maternal blood plasma DNA, wherein the maternal blood plasma DNA comprises a DNA mixture of maternal DNA and fetal DNA, and wherein the sequences comprise a plurality of allelic sequence reads of an allelic sequence in a gene linked to the autosomal recessive disease;(d) performing in silico size selection on the plurality of allelic sequence reads for the DNA mixture to increase fetal fraction of the allelic sequence reads;(e) calculating the fetal fraction after said in silico size selection, wherein the fetal fraction is calculated based on known genotypes at single nucleotide polymorphisms (SNPs) of the paternal and maternal haplotypes and observed ratios of the genotypes of the SNPs in the allelic sequence reads for the DNA mixture after said in silico size selection, wherein any SNP for which an allele is detected that is absent in the maternal haplotype and present in the paternal haplotype is assumed to belong to the fetal DNA; and(f) determining the fetal genotype at a site of a mutation linked to the autosomal recessive disease by comparing an observed ratio of genotypes for the allelic sequence reads of the DNA mixture after said in silico size selection to an expected ratio for each possible genotype of the fetal DNA based on the paternal and maternal haplotypes and the calculated fetal fraction, wherein the fetus is diagnosed as having the autosomal recessive disease if the fetal genotype is homozygous for the mutation linked to the autosomal recessive disease, wherein the fetus is diagnosed as being a carrier of the autosomal recessive disease if the fetal genotype is heterozygous for the mutation linked to the autosomal recessive disease, and wherein the fetus is diagnosed as not having the autosomal recessive disease if the fetal genotype does not have the mutation linked to the autosomal recessive disease; and(g) displaying information regarding the fetal genotype.33. The computer implemented method of aspect 32, further comprising storing the information regarding the fetal genotype in a database.34. The computer implemented method of aspect 32 or 33, further comprising instructing a sequencer to sequence at least a portion of the chromosome comprising a gene linked to the autosomal recessive disease in the maternal and paternal DNA.35. The computer implemented method of aspect 34, wherein the sequencer performs nanopore sequencing of said at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease to determine the paternal and maternal haplotypes.36. The computer implemented method of any one of aspects 32-35, further comprising instructing a sequencer to sequence the maternal DNA and fetal DNA from the maternal blood plasma.37. The computer implemented method of aspect 36, wherein the sequencer performs reversible-terminator sequencing by synthesis of the maternal DNA and fetal DNA to generate the plurality of allelic sequence reads.38. A system comprising:(a) a storage component for storing data, wherein the storage component has instructions for diagnosing an autosomal recessive disease in a fetus stored therein;(b) a computer processor programmed to analyze maternal and paternal sequences of at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease and sequences of maternal DNA and fetal DNA from maternal blood plasma using one or more algorithms, wherein the computer processor is coupled to the storage component and configured to execute the instructions stored in the storage component in order to receive the inputted sequences, and analyze the sequences according to the computer implemented method of any one of aspects 32-37; and(c) a display component for displaying information regarding the fetal genotype.39. The system of aspect 38, further comprising a sequencer to sequence at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease in the maternal and paternal DNA.40. The system of aspect 39, wherein the sequencer can perform nanopore sequencing of said at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease to determine the paternal and maternal haplotypes.41 . The system of any one of aspects 38-40, further comprising a sequencer to sequence the maternal DNA and fetal DNA from the maternal blood plasma.42. The system of aspect 41 , wherein the sequencer can perform reversible-terminator sequencing by synthesis of the maternal DNA and fetal DNA to generate the plurality of allelic sequence reads.43. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor in a computer, causes the processor to perform the computer implemented method of any one of aspects 32-37.44. A kit comprising the non-transitory computer-readable medium of aspect 43 and instructions for diagnosing an autosomal recessive disease in a fetus.45. The kit of aspect 44, further comprising a capture probe that specifically binds to an allelic sequence in the gene linked to the autosomal recessive disease.46. The kit of aspect 44 or 45, further comprising primers for amplifying the maternal DNA and the fetal DNA from the maternal blood plasma and the parental DNA.47. The kit of any one of aspects 44-46, further comprising adapters.48. The kit of any one of aspects 44-47, further comprising primers for sequencing the maternal DNA and the fetal DNA from the maternal blood plasma and the parental DNA.49. The kit of any one of aspects 44-48, further comprising a plurality of capture probes, wherein each probe specifically binds to a different allelic sequence linked to a different autosomal recessive disease to allow multiplexed prenatal testing of the fetus for different autosomal recessive diseases.

[0158] It will be apparent to one of ordinary skill in the art that various changes and modifications can be made without departing from the spirit or scope of the invention.EXPERIMENTAL

[0159] The following examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the present invention, and are not intended to limit the scope of what the inventors regard as their invention nor are they intended to represent that the experiments below are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g. amounts, temperature, etc.) but some experimental errors and deviations should be accounted for. Unless indicated otherwise, parts are parts by weight, molecular weight is average molecular weight, temperature is in degrees Centigrade, and pressure is at or near atmospheric.

[0160] All publications and patent applications cited in this specification are herein incorporated by reference as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference.

[0161] The present invention has been described in terms of particular embodiments found or proposed by the present inventor to comprise preferred modes for the practice of the invention. It will be appreciated by those of skill in the art that, in light of the present disclosure, numerous modifications and changes can be made in the particular embodiments exemplified without departing from the intended scope of the invention. All such modifications are intended to be included within the scope of the appended claims.Example 1Non-lnvasive Prenatal Testing of Beta- Hemoglobinopathies using Next Generation Sequencing, In- Silico Seguence Size Selection and Haplotyping HE2.28INTRODUCTION

[0162] Following the discovery of fetal DNA in the maternal plasma[8] and the implementation of Next Generation Sequencing (NGS) and digital PGR, the development of non-invasive prenatal testing (NIPT) based on the analysis of plasma libraries has become well-established for chromosomal aneuploidies [9, 10]. The clonal property of NGS allows the quantitative resolution of mixtures, such as the fetal and maternal DNA in plasma, by sequencing the components separately.Recently, we and others have demonstrated that NGS analysis of plasma libraries can also be used for the more challenging diagnosis of autosomal recessive diseases, such as the hemoglobinopathies[11 , 12, 13, 14], In this approach, the fetal genotype is predicted by counting Illumina MiSeq sequence reads from plasma libraries and comparing the observed ratios of diseasecausing mutation (Mut) alleles and reference (Ref) alleles to those expected for the three possible genotypes (Mut / Mut, Mut / Ref and Ref / Ref. The expected ratios for the three possible genotypes are based on estimates of the fetal fraction (FF), which is determined by counting allelic sequence reads present in the plasma and absent in the mother.

[0163] This strategy for fetal genotype prediction has two limitations. First, when the FF is low, the expected ratios for the three possible genotypes are not far apart and the statistical confidence in the predicted fetal genotype is therefore correspondingly lower. Second, the prediction is based on the observed sequence read ratios at only one or two positions, limiting the statistical confidence. To specifically address the limitations of NIPT for hemoglobinopathies using individual SNPs, we have developed a strategy that 1 ) bioinformatically increases the FF and 2) uses parental HBB haplotypes to incorporate the allelic read ratios observed at linked SNPs in our prediction at the mutation site(s).

[0164] Based on the observation that fetal DNA may be slightly shorter than maternal DNA in plasma

[0015] , we and others have shown that it is possible to bioinformatically increase the FF by excluding longer sequence reads, via in-silico size selection (ISS)[11 , 14], Here, we demonstrate how the length distributions of fetal reads are skewed to shorter lengths, relative to maternal sequence reads, and how the FF increases as a function of the excluded read length.

[0165] Finally, we have addressed the issue of identifying phase between SNPs and mutations by PCR amplifying a 2.2 kb region that includes the HBB gene in parents, sequencing this amplicon using Oxford Nanopore MinlON single-molecule sequencing technology, and determining the parental HBB haplotypes using Soft Genetics’ NextGENe LR software. Having established phase for the linked SNPs, we can then use the observed sequence read ratios at these sites to infer which parental haplotype was transmitted to the fetus and thus predict with greater confidence the fetal genotype at the mutation site.Methods

[0166] The probe design, specimen collection, and probe capture / NGS methods applied in this work were performed as previously described by Erlich et. al. (2022)

[0011] . DNA library preparation, probe capture enrichment and sequencing methodologies were modified as described below.Subjects

[0167] Maternal and paternal whole blood and chorionic villus samples (CVS) for families in which both parents were known to carry HBB mutations were collected at the Post-Graduate Institute of Medical Education and Research (PGIMER) in Chandigarh, India (114 families in total) and at the UCSF Benioff Children’s Hospital Oakland (UBCHO), in Oakland CA,US. In all families, both parents were known to be carriers of a HBB mutation. All specimens were collected with Institutional Review Board Approval at UCSF Benioff Children’s Hospital Oakland and the Post Graduate Institute of Medical and Educational Research.DNA Library Preparation

[0168] We previously

[0011] sheared DNA extracted from whole blood and plasma to 250 bp prior to Illumina MiSeq sequencing. For the work described here, extracted plasma DNA was not sheared prior to sequencing. Given the mean fragment size of plasma DNA, between 145 and 201 bp

[0016] , shearing to 250 bp was not deemed necessary.

[0169] As described below, DNA was extracted from whole blood, PCR amplified, and sequenced using Oxford Nanopore Technologies’ (ONT) (Oxford, UK) MinlON platform. For ONT sequencing, library preparation for single-individual sequencing runs was performed using ONT Ligation Sequencing kits (SQK-LSK109), and multiplex library preparation was performed using Ligation Sequencing gDNA Native Barcoding kits (SQK-NBD1 14.96) according to the manufacturer’s specifications.Probe Capture Enrichment

[0170] Our previously described Probe Capture protocol^ 1 , 17] was modified to use a custom Twist Bioscience (South San Francisco, CA, USA) probe panel (PanHeme panel TE-95715506). This panel covers 4kb of chromosome 1 1 , and additionally includes 436 genomic (non-HBB) SNPs. This protocol involved 1 1 cycles of PCR amplification following adapter ligation, and 15 cycles following probe capture.Genomic DNA PCR Amplification and Quantification

[0171] To sequence the entire HBB gene and flanking SNPs, we amplified a 2,252 nucleotide-long segment using primers A and D described by Chan et al.

[0018] . PCR amplification was performed using 25 pL AmpliTaq Gold™ 360 Master Mix (1X, Applied Biosystems™), 1 pL of primers (0.2uM for each primer), and 150 ng of genomic DNA in PCR-grade water in 50 p.L reaction volumes. After 10 minutes denaturation at 95°C, 33 PCR cycles [15 seconds at 95°C, 58.5°C for 30 seconds, and72°C for 2 minutes] were followed by 7 minutes of extension at 72°C. PCR products were quantified via PicoGreen, using the Quant-iT™ PicoGreen™ dsDNA Assay kits (Invitrogen, Waltham, MA, UA), and BioAnalyzer, using Agilent DNA 7500 kits (Agilent).DNA Sequencing

[0172] Plasma DNA was sequenced on the Illumina MiSeq platform as previously described

[0011] . Some plasma libraries were sequenced with the MiSeq v3. 2 x 600 cycle reagent kits. Some were sequenced with the MiSeq v2 300-cycle reagent kits, run for either 150 or 175 cycles, given the short DNA fragments present in the plasma.

[0173] The 2.2 kb PCR product containing the HBB gene was PCR amplified from DNA extracted from parental whole blood and CVS specimens, and sequenced using the ONT MinlON system. Singleplex sequencing was performed on a MinlON Mk1 b instrument using R.9.4.1 and R10.4.1 flowcells, with MinKNOW software (version 5.0.5) over 4 hours. Multiplex sequencing was performed using the MinlON R10.4.1 flowcells, with MinKNOW software (version 5.0.5) over 3.5 hours.Data Analysis and Informatics

[0174] MiSeq read processing, elimination of PCR duplicates

[0011] (deduplication) and ISS was performed using the size_selection_and_deduplication.py script, available at github.com / kaeaton / NIPT_lnformatics, which controls the Fastp

[0019] , BWA-MEM

[0020] , Genome Analysis Toolkit 4

[0021] (GATK), and Gencore

[0022] command line applications. Zipped, MiSeq- generated, fastq-formatted paired end reads are first processed by Fastp, which trims any remaining MiSeq adapter sequences, manages unique molecular identifier

[0023] (UM I) sequences, performs overlap analysis of paired-end reads, replaces mismatched, low-quality nucleotides, and performs ISS when desired, through application of the user-specified length limit’ [LL] parameter. The end product of the Fastp process is paired ‘RT and ‘R2’ fastq read files.

[0175] Alignment of the fastq-formatted reads to the GRCh37 / hg19 reference was performed using the Burrows-Wheeler Alignment BWA) tool, implemented as BI / 144 MEM, which returns a SAM file. GA TK con verts this into a BAM file, which is inspected by GATKs FixMatelnformation tool, ensuring position matches within paired reads. Pair reads are then sorted by genomic coordinates by the GATK SortSAMtoo\ before removal of PCR duplicates, based on the UMI tags.

[0176] Removal of PCR duplicates (deduplication )was performed by Gencore, which builds consensus sequences using the duplicate reads, providing the longest, most accurate consensus read per UMI. The default minimum number of reads (1) is used to build consensus sequences, to ensure that all probed regions are included.

[0177] The resulting plasma reads were aligned and analyzed using Soft Genetics’ (State College, PA, USA) NextGENe software

[0024] version 2.4.2.2 using GRCh37 / hg19 genomic coordinates, and previously described settings

[0011] . Aligned reads were post-processed for variant calling, identifying HBB SNPs.

[0178] SoftGenetics’ NextGENe LR software (version 1.0.4.3) was used to analyze MinlON singleplex and multiplex FASTQ data files, identifying HBB SNP haplotypes for maternal, paternal and CVS specimens. In particular, NextGENe LR can phase HBB SNPs with the common 619bp intron 2 - exon 3 deletion variant (GRCh37 / hg19 chr1 1 : 5246486-5247107)

[0025] . The average readdepth per ONT sequencing run was ~1 million reads, so that the read depth for multiplexed sequence runs was -100,000 reads per specimen.

[0179] Sequence reads were aligned to the GRCh37 / hg19 human genome reference sequence positions chr11 :5246155-5248406, with 75% homology. NextGene LR indel percentage and variant effective percentage thresholds were set at 30% and 10%, respectively. Given the error rate of MinlON sequencing, variant allele calls were made for minor allele sequence read frequencies >16%. As MinKNOW generates multiple -10MB FASTQ files for each subject / barcode, all FASTQ files for a given subject / barcode were loaded into NextGENe LR and analyzed as a “single sample”.Results

[0180] To increase the accuracy of predicting the fetal genotype from sequence read ratios from the plasma libraries, we have 1) bioinformatically increased the fetal fraction by in silico size selection of plasma reads and 2) determined the beta-globin haplotypes by Oxford Nanopore sequencing of parental blood samples.In-silico size selection

[0181] The length distribution of Illumina MiSeq sequence reads from fetal DNA was compared to maternal reads for all informative SNPs in the plasma library. Informative SNPs are those for which an allele detected in the plasma is absent in the mother, and therefore assumed to have been transmitted from the father to the fetus. For example, if the mother is A / A, and 5% of the plasma reads are T, we can compare the length of fetal T reads with the length of the predominantly maternal A reads (we assume that 5% of the plasma A reads are from the fetus). Since half of the fetal reads will be A, the proportion of maternal A reads is 100% minus (FF% / 2).

[0182] This comparison was initially performed for all informative SNPs in a plasma library sequenced with an Illumina MiSeq v3 600 cycle kits, as previously described

[0011] . FIG. 1 shows the analysis with a v2 2x 300 cycle kit run for 175 cycles. The FF is estimated by counting the obligatefetal reads in a plasma library, calculating the average fetal proportion over all informative SNPs, and then multiplying by two. To systematically monitor the increase in the FF achieved by bioinformatically excluding sequence reads above a defined length, we excluded reads sequentially starting from 167bp (FIG. 1 ) The FF rises, peaking at around 135 bp. The number of informative reads for calculating the FF is shown on the right y-axis. These are reads for SNPs where the mother is homozygous. From a statistical perspective, the increase in FF needs to be balanced against the reduction in read count and the resulting increase in noise. (The reads informative for the most challenging fetal genotype prediction are derived from the SNPs where the mother is heterozygous. Predicting the fetal genotype for SNPs where the mother is homozygous is easier.)

[0183] The goal of bioinformatically increasing the FF by ISS is to increase the accuracy and statistical confidence in the fetal genotype predictions based on the comparing the observed and expected sequence read ratios in the plasma library. Increasing the FF will change the expected read ratios and, with the exception of a heterozygous fetus, can alter the observed read ratios as well. FIG. 2 illustrates a case in which ISS altered the plasma read ratios, leading to more accurate fetal genotype predictions. As illustrated for seven families in Table 1 , our application of ISS to increase the FF can help correct fetal genotype predictions, based solely on mutation site read ratios.HBB Haplotype Analysis

[0184] The application of parental SNP haplotype information, in concert with the plasma sequence read ratios observed at SNPs linked to the mutation site(s), could increase the accuracy and statistical confidence in fetal genotype predictions. We adopted an HBB haplotype strategy using ONT MinlON long read sequencing, to sequence a 2.2 kb HBB region amplified from parental DNA. NextGENe LR (v1 .0.4.3) identifies the two parental haplotypes for each parent from the FASTQ, typically using ~1 M reads per singleplex subject and ~75K reads per multiplex subject. NextGENE LR can phase SNPs with the common 619 deletion, allowing the use of linked SNPs alleles to predict whether this deletion was transmitted to the fetus (Table 1).

[0185] FIG. 2 illustrates the combined use of ISS and HBB haplotype analysis to incorporate sequence read ratio data at linked SNPs into the mutation site prediction for family IBT-42.ISS increases the FF from 8.04 to 14.66. For SNP1 , the observed read ratio changes from 55A / 44G to 59A / 41 G, making the correct prediction of a A / A fetal genotype more likely. For SNP2, the ratio changes from 48A / 52C to 44 / A / 57C, making the correct prediction of a C / C genotype more likely. For SNP3, however, the ratios both before and after ISS are 47C / 53G, consistent with either a C / G heterozygote or a G / G homozygote.

[0186] The haplotype information is required to correctly predict the GIG fetal genotype. At the mutation site, the correct C / C is predicted by the 57C / 43G ratio alone but the ISS and haplotype analysis provide additional confidence in the prediction. For SNP4, the observed 47A / 53G is consistent with either an A / G or a GG fetal genotype, but the ISS (43A / 57G) and the haplotype analysis predict the correct G / G genotype. Table 1 illustrates how the ISS and haplotype analysis can help prediction of fetal genotype at the clinical mutation sites of seven additional families, including a family carrying the -619 deletion.Discussion

[0187] A robust NIPT for the hemoglobinopathies, as well as other autosomal recessive monogenic diseases, could eliminate the risk posed to a fetus by amniocentesis and CVS biopsies. The discovery of fetal DNA in the maternal plasma has enabled the prediction of the fetal genotype by analyzing maternal plasma by Next Generation Sequencing. The strategy of inferring fetal genotypes by comparing the observed ratios of allelic sequence reads to the expected ratios for the possible genotypes is limited by the FF. Since the expected values of the read ratios is determined by the FF, this strategy becomes difficult at low FFs. We and others have shown that bioinformatically excluding longer reads can enrich the FF (11 ). ISS, however, involves the trade-off between increased FF and a reduced read count, which increases statistical noise.HBB Haplotypes

[0188] The accuracy in fetal genotype prediction, based solely on observed sequence read ratios at mutation site(s), is limited by the available data. We and others have demonstrated the utility of considering read ratios at linked SNPs for fetal genotype prediction^ 1 , 14, 26]. In our previous work, we used the availability of either an informative sibling or short-range haplotypes determined from overlapping short MiSeq reads to aid in the prediction^ 1], However, these are not reliable or generalizable strategies to determine parental HBB haplotype ratios.

[0189] Here, we have used Oxford Nanopore Min ION sequencing of a 2.2 kb amplicon and Soft Genetics' NextGENe LR software to determine the HBB haplotypes of both parents. A recent NIPT paper

[0027] reported the Oxford Nanopore sequencing of a 10 kb and a 20 kb HBB amplicon. We have focused on a shorter amplicon to ensure robust PCR amplification and to minimize the potential for the formation of hybrid PCR sequences created by “template switching”, an artifact that is more likely with long amplicons

[0028] . These artifacts could make the determination of haplotypes more difficult. Longer amplicons do allow more SNPs to be phased, but we have shown that a 2.2 kb amplicon is sufficient for these purposes.

[0190] Our current protocol is to sequence the plasma libraries, prepared by probe capture, on the Illumina MiSeq while sequencing the 2.2 kb HBB fragment PGR amplified from the parental DNA on the Oxford Nanopore MinlON. Though both our MiSeq and MinlON library preparation protocols each take two days, sequencing via ONT takes 3.5 hours vs 1 .5 days via MiSeq. Given this, library preparation and sequencing can be staggered so that parental HBB haplotypes are available to facilitate interpretation of the plasma read data as soon as the latter are generated.Conclusions

[0191] Where NIPT for autosomal recessive diseases may once have seemed very challenging, contemporary sequencing and informatic methods have made it a realistic clinical goal. The use of in silica size selection of plasma reads increases the fetal fraction and increases the accuracy of fetal genotype prediction. Determining the parental beta-globin haplotypes further increases the predictive accuracy.TABLE 1. Ratios and predictions of clinically significant mutations across 25 familiesFor each family, more than one HBB mutation may be present. In some cases, maternal and paternal genomes displayed mutations at different positions.For each position, the Plasma Variant Ratio identifies percentage of reads with the specific variants. For example, the 51 C I 49T ratio for position 5248159 indicates that 51% of the reads have a C sequence at this position, while 49% have a T sequence.The Predicted Fetal Genotype is the result returned by our algorithm, based solely on the Variant Ratios. The Fetal Fraction identifies the percent of reads for that position that are derived from the fetus. a) Fetal genotypes at the mutation sites were predicted based on plasma ratio and fetal fraction (FF) values. Incorrect predictions are indicated in salmon-colored cells. b) In-silico size selection (ISS) was performed for these 3 libraries with size limits at 190bp and 160bp to enrich the FF.For PST 5 at position 5246969, the plasma ratio after 190bp ISS indicated that the G / T genotype was the most probable, resulting in a correct prediction.For IBT 43, 160bp ISS provided higher confidence to predict the G / T genotype at the paternal mutation site. However, the prediction at the maternal mutation site was still incorrect after ISS. Oxford Nanopore sequencing will be performed to determine the haplotypes.For IBT31 , 160bp ISS provided higher confidence to predict Ins / C genotype at the paternal mutation site. However, the prediction at the maternal mutation site changed from inconclusive to incorrect after ISS. Nanopore sequencing will be performed to determine the haplotypes. c) Haplotypes were determined using SoftGenetics’ NextGENeLR software after performing super high accuracy base calling via the Oxford Nanopore Technologies MinKNOW platform (FIG. 2). d) Short-range haplotypes were inferred based on the sibling’s genotype. These inferred haplotypes changed the predictions at both maternal and paternal mutation sites from inconclusive to correct.REFERENCES

[0192] 1. Piel, F.B., M.H. Steinberg, and D.C. Rees, Sickle Cell Disease. N Engl J Med, 2017.377(3): p. 305.

[0193] 2. Williams, T.N. and D. J. Weatherall, World distribution, population genetics, and health burden of the hemoglobinopathies. Cold Spring Harb Perspect Med, 2012. 2(9): p. a011692.

[0194] 3. Frangoul, H., et aL, CRISPR-Cas9 Gene Editing for Sickle Cell Disease and fi-Thalassemia. N Engl J Med, 2021 . 384(3): p. 252-260.

[0195] 4. Thein, S.L., The molecular basis of ^-thalassemia. Cold Spring Harb Perspect Med,2013. 3(5): p. a011700.

[0196] 5. Kuliev, A.M., et aL, Risk evaluation of CVS. Prenat Diagn, 1993. 13(3): p. 197-209.

[0197] 6. Report of National Institute of Child Health and Human Development Workshop onChorionic Villus Sampling and Limb and Other Defects, October 20, 1992. Am J Obstet Gynecol, 1993. 169(1 ): p. 1 -6.

[0198] 7. Rather, R.A. and S.C. Saha, Reappraisal of evolving methods in non-invasive prenatal screening: Discovery, biology and clinical utility. Heliyon, 2023. 9(3): p. e13923.

[0199] 8. Lo, Y.M., et aL, Presence of fetal DNA in maternal plasma and serum. Lancet, 1997.350(9076): p. 485-7.

[0200] 9. Mavrou, A., et al., Identification of nucleated red blood cells in maternal circulation: a second step in screening for fetal aneuploidies and pregnancy complications. Prenat Diagn, 2007. 27(2): p. 150-3.

[0201] 10. Lo, Y.M., et aL, Plasma placental RNA allelic ratio permits noninvasive prenatal chromosomal aneuploidy detection. Nat Med, 2007. 13(2): p. 218-23.

[0202] 11. Erlich, H.A., et aL, Noninvasive Prenatal Test for / 3-Thalassemia and Sickle CellDisease Using Probe Capture Enrichment and Next-Generation Sequencing of DNA in Maternal Plasma. J Appl Lab Med, 2022. 7(2): p. 515-531 .

[0203] 12. D'Aversa, E., et aL, Droplet Digital PCR for Non-lnvasive Prenatal Detection of FetalSingle-Gene Point Mutations in Maternal Plasma. Int J Mol Sci, 2022. 23(5).

[0204] 13. Xiong, L., et aL, Non-invasive prenatal testing for fetal inheritance of maternal (3- thalassaemia mutations using targeted sequencing and relative mutation dosage: a feasibility study. Bjog, 2018. 125(4): p. 461 -468.

[0205] 14. van Campen, J., et aL, A novel non-invasive prenatal sickle cell disease test for all at- risk pregnancies. Br J Haematol, 2020. 190(1 ): p. 119-124.

[0206] 15. Lo, Y.M., et aL, Maternal plasma DNA sequencing reveals the genome-wide genetic and mutational profile of the fetus. Sci Transl Med, 2010. 2(61 ): p. 61 ra91 .

[0207] 16. Chan, K.C., et aL, Size distributions of maternal and fetal DNA in maternal plasma.Clin Chem, 2004. 50(1 ): p. 88-92.

[0208] 17. Bose, N., et aL, Target capture enrichment of nuclear SNP markers for massively parallel sequencing of degraded and mixed samples. Forensic Sci Int Genet, 2018. 34: p. 186-196.

[0209] 18. Chan, O.T., et aL, Comprehensive and efficient HBB mutation analysis for detection of beta-hemoglobinopathies in a pan-ethnic population. Am J Clin Pathol, 2010. 133(5): p. 700-7.

[0210] 19. Chen, S., et aL, fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics,2018. 34(17): p. i884-i890.

[0211] 20. Li, H., Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv, 2013. 1303.3997v2: p. 1 -3.

[0212] 21 . van der Auwera, G. and B.D. O'Connor, Genomics in the Cloud: Using Docker, GATK, and WDL in Terra. 2020: O'Reilly Media, Incorporated.

[0213] 22. Chen, S., et al., Gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data. BMC Bioinformatics, 2019. 20(Suppl 23): p. 606.

[0214] 23. Kivioja, T., et al., Counting absolute numbers of molecules using unique molecular identifiers. Nat Methods, 2011 . 9(1 ): p. 72-4.

[0215] 24. Sikkema-Raddatz, B., et al., Targeted next-generation sequencing can replaceSanger sequencing in clinical diagnostics. Hum Mutat, 2013. 34(7): p. 1035-42.

[0216] 25. Dabbag h Bag heri, S., et al., Revisiting a Complex Rearrangement Involving a 619Base Pairs Deletion, 6 Nucleotide Insertion Followed by a A > G Substitution Causing Thalassemia. Indian J Hematol Blood Transfus, 2016. 32(4): p. 500-503.

[0217] 26. Wu, R., C.-X. Ma, and G. Casella, Joint linkage and linkage disequilibrium mapping of quantitative trait loci in natural populations. Genetics, 2002. 160: p. 779-792.

[0218] 27. Jiang, F., et al., Noninvasive prenatal testing for ^-thalassemia by targeted nanopore sequencing combined with relative haplotype dosage (RHDO): a feasibility study. Sci Rep, 2021. 11 (1 ): p. 5714.

[0219] 28. Turner, D.J., C. Tyler-Smith, and M.E. Hurles, Long-range, high-throughput haplotype determination via haplotype-fusion PCR and ligation haplotyping. Nucleic Acids Res, 2008. 36(13): p. e82.

Claims

What is claimed is:

1. A method of non-invasive prenatal testing for diagnosing an autosomal recessive disease in a fetus, the method comprising:(a) obtaining parental DNA from both parents of the fetus;(b) sequencing at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease in the parental DNA to determine paternal and maternal haplotypes;(c) obtaining a maternal blood plasma sample at a stage of prenatal development of the fetus, wherein the maternal blood plasma sample comprises a DNA mixture of maternal DNA and fetal DNA;(d) isolating the DNA mixture from the blood plasma sample using a capture probe that specifically binds to an allelic sequence in the gene linked to the autosomal recessive disease, wherein the capture probe captures the maternal DNA and the fetal DNA comprising the allelic sequence;(e) sequencing the DNA mixture after said isolating to generate a plurality of allelic sequence reads;(f) performing in silica size selection on the plurality of allelic sequence reads for the DNA mixture to increase fetal fraction of the allelic sequence reads;(g) calculating the fetal fraction after said in silico size selection, wherein the fetal fraction is calculated based on known genotypes at single nucleotide polymorphisms (SNPs) of the paternal and maternal haplotypes and observed ratios of the genotypes of the SNPs in the allelic sequence reads for the DNA mixture after said in silico size selection, wherein any SNP for which an allele is detected that is absent in the maternal haplotype and present in the paternal haplotype is assumed to belong to the fetal DNA; and(h) determining the fetal genotype at a site of a mutation linked to the autosomal recessive disease by comparing an observed ratio of genotypes for the allelic sequence reads of the DNA mixture after said in silico size selection to an expected ratio for each possible genotype of the fetal DNA based on the paternal and maternal haplotypes and the calculated fetal fraction, wherein the fetus is diagnosed as having the autosomal recessive disease if the fetal genotype is homozygous for the mutation linked to the autosomal recessive disease, wherein the fetus is diagnosed as being a carrier of the autosomal recessive disease if the fetal genotype is heterozygous for the mutation linked to the autosomal recessive disease, and wherein the fetus is diagnosed as not having the autosomal recessive disease if the fetal genotype does not have the mutation linked to the autosomal recessive disease.

2. The method of claim 1 , wherein the autosomal recessive disease is a monogenic autosomal recessive disease.

3. The method of claim 1 or 2, wherein the autosomal recessive disease is a hemoglobinopathy.

4. The method of claim 3, wherein the hemoglobinopathy is sickle cell disease or thalassemia.

5. The method of claim 4, wherein the thalassemia is alpha-thalassemia or beta thalassemia.

6. The method of any one of claims 3-5, wherein the chromosome comprising the gene linked to the autosomal recessive disease is chromosome 11 .

7. The method of claim 6, wherein the gene linked to the autosomal recessive disease is a globin gene.

8. The method of claim 7, wherein the globin gene is an alpha-globin gene or a betaglobin gene.

9. The method of any one of claims 1 -8, wherein the allelic sequence comprises a mutation at a single-nucleotide polymorphism linked to the autosomal recessive disease.

10. The method of any one of claims 1 -9, wherein the fetal fraction of the allelic sequence reads before said performing in silica size selection is less than 10%.11 . The method of claim 10, wherein the fetal fraction of the allelic sequence reads before said performing in silica size selection is less than 5%.

12. The method of any one of claims 1 -11 , further comprising amplifying said at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease togenerate an amplicon, wherein said sequencing of step (b) comprises sequencing the amplicon to determine the paternal and maternal haplotypes.

13. The method of claim 12, wherein said amplifying said at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease comprises performing polymerase chain reaction or isothermal nucleic acid amplification.

14. The method of claim 12 or 13, wherein the amplicon comprises or consists of the sequence of the gene linked to the autosomal recessive disease.

15. The method of any one of claims 12-14, wherein the amplicon has a length in a range from 1 kilobase to 5 kilobases.

16. The method of any one of claims 1 -15, wherein the sequencing of step (b) comprises performing nanopore sequencing to determine the paternal and maternal haplotypes.

17. The method of any one of claims 1 -16, wherein the sequencing of step (e) comprises performing reversible-terminator sequencing by synthesis to generate the plurality of allelic sequence reads.

18. The method of any one of claims 1 -17, wherein the maternal DNA and the fetal DNA from the maternal plasma is not sheared prior to performing the sequencing of step (e).

19. The method of any one of claims 1-18, wherein said performing in silica size selection comprises excluding allelic sequence reads having a length greater than 155 bases, greater than 156 bases, greater than 157 bases, greater than 158 bases, greater than 159 bases, or greater than 160 bases.

20. The method of any one of claims 1 -19, further comprising amplifying the maternal DNA and the fetal DNA prior to performing the sequencing of step (e).21 . The method of claim 20, wherein said amplifying the maternal DNA and the fetal DNA comprises performing polymerase chain reaction, isothermal nucleic acid amplification, or clonal amplification.

22. The method of claim 21 , wherein the polymerase chain reaction is digital polymerase chain reaction or quantitative polymerase chain reaction.

23. The method of any one of claims 1 -22, further comprising adding adapters to the 5’ and 3’ ends of the maternal DNA and the fetal DNA prior to performing the sequencing of step (e).

24. The method of claim 23, wherein the adapters are linear adapters, Y-adapters, stubby adapters, or hairpin adapters.

25. The method of any one of claims 1-24, further comprising treating the fetus for the autosomal recessive disease if the fetus is diagnosed as having the autosomal recessive disease.

26. The method of claim 25, wherein said treating comprises transplanting stem cells into the fetus in utero or postnatally, or a combination thereof.

27. The method of claim 26, wherein the stem cells tolerize the fetus for a postnatal transplant.

28. The method of claim 26, wherein the stem cells comprise a gene-editing system to repair or activate the gene linked to the autosomal recessive disease.

29. The method of claim 26, wherein the stem cells have a wild-type copy of the gene linked to the autosomal recessive disease.

30. The method of claim 25, wherein said treating comprising performing gene therapy on the fetus in utero or postnatally, or a combination thereof.31 . The method of any one of claims 1 -30, further comprising using a plurality of capture probes, wherein each probe specifically binds to a different allelic sequence linked to a different autosomal recessive disease to allow multiplexed prenatal testing of the fetus for different autosomal recessive diseases.

32. A computer implemented method for diagnosing an autosomal recessive disease in a fetus, the computer performing steps comprising:(a) receiving maternal and paternal sequences of at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease;(b) genotyping a plurality of single nucleotide polymorphisms (SNPs) in the maternal and paternal sequences to determine paternal and maternal haplotypes;(c) receiving sequences of maternal blood plasma DNA, wherein the maternal blood plasma DNA comprises a DNA mixture of maternal DNA and fetal DNA, and wherein the sequences comprise a plurality of allelic sequence reads of an allelic sequence in a gene linked to the autosomal recessive disease;(d) performing in silico size selection on the plurality of allelic sequence reads for the DNA mixture to increase fetal fraction of the allelic sequence reads;(e) calculating the fetal fraction after said in silico size selection, wherein the fetal fraction is calculated based on known genotypes at single nucleotide polymorphisms (SNPs) of the paternal and maternal haplotypes and observed ratios of the genotypes of the SNPs in the allelic sequence reads for the DNA mixture after said in silico size selection, wherein any SNP for which an allele is detected that is absent in the maternal haplotype and present in the paternal haplotype is assumed to belong to the fetal DNA; and(f) determining the fetal genotype at a site of a mutation linked to the autosomal recessive disease by comparing an observed ratio of genotypes for the allelic sequence reads of the DNA mixture after said in silico size selection to an expected ratio for each possible genotype of the fetal DNA based on the paternal and maternal haplotypes and the calculated fetal fraction, wherein the fetus is diagnosed as having the autosomal recessive disease if the fetal genotype is homozygous for the mutation linked to the autosomal recessive disease, wherein the fetus is diagnosed as being a carrier of the autosomal recessive disease if the fetal genotype is heterozygous for the mutation linked to the autosomal recessive disease, and wherein the fetus is diagnosed as not having the autosomal recessive disease if the fetal genotype does not have the mutation linked to the autosomal recessive disease; and(g) displaying information regarding the fetal genotype.

33. The computer implemented method of claim 32, further comprising storing the information regarding the fetal genotype in a database.

34. The computer implemented method of claim 32 or 33, further comprising instructing a sequencer to sequence at least a portion of the chromosome comprising a gene linked to the autosomal recessive disease in the maternal and paternal DNA.

35. The computer implemented method of claim 34, wherein the sequencer performs nanopore sequencing of said at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease to determine the paternal and maternal haplotypes.

36. The computer implemented method of any one of claims 32-35, further comprising instructing a sequencer to sequence the maternal DNA and fetal DNA from the maternal blood plasma.

37. The computer implemented method of claim 36, wherein the sequencer performs reversible-terminator sequencing by synthesis of the maternal DNA and fetal DNA to generate the plurality of allelic sequence reads.

38. A system comprising:(a) a storage component for storing data, wherein the storage component has instructions for diagnosing an autosomal recessive disease in a fetus stored therein;(b) a computer processor programmed to analyze maternal and paternal sequences of at least a portion of a chromosome comprising a gene linked to the autosomal recessive disease and sequences of maternal DNA and fetal DNA from maternal blood plasma using one or more algorithms, wherein the computer processor is coupled to the storage component and configured to execute the instructions stored in the storage component in order to receive the inputted sequences, and analyze the sequences according to the computer implemented method of any one of claims 32-37; and(c) a display component for displaying information regarding the fetal genotype.

39. The system of claim 38, further comprising a sequencer to sequence at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease in the maternal and paternal DNA.

40. The system of claim 39, wherein the sequencer can perform nanopore sequencing of said at least a portion of the chromosome comprising the gene linked to the autosomal recessive disease to determine the paternal and maternal haplotypes.41 . The system of any one of claims 38-40, further comprising a sequencer to sequence the maternal DNA and fetal DNA from the maternal blood plasma.

42. The system of claim 41 , wherein the sequencer can perform reversible-terminator sequencing by synthesis of the maternal DNA and fetal DNA to generate the plurality of allelic sequence reads.

43. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor in a computer, causes the processor to perform the computer implemented method of any one of claims 32-37.

44. A kit comprising the non-transitory computer-readable medium of claim 43 and instructions for diagnosing an autosomal recessive disease in a fetus.

45. The kit of claim 44, further comprising a capture probe that specifically binds to an allelic sequence in the gene linked to the autosomal recessive disease.

46. The kit of claim 44 or 45, further comprising primers for amplifying the maternal DNA and the fetal DNA from the maternal blood plasma and the parental DNA.

47. The kit of any one of claims 44-46, further comprising adapters.

48. The kit of any one of claims 44-47, further comprising primers for sequencing the maternal DNA and the fetal DNA from the maternal blood plasma and the parental DNA.

49. The kit of any one of claims 44-48, further comprising a plurality of capture probes, wherein each probe specifically binds to a different allelic sequence linked to a different autosomal recessive disease to allow multiplexed prenatal testing of the fetus for different autosomal recessive diseases.

Citation Information

Cited By

  • Primer group and kit for detecting various mutations of human embryo beta-thalassemia and application of primer group and kit

    CN122168751A