Methods for reliable non-invasive preimplantation genetic testing

A transposon-based DNA amplification and Bayesian linkage analysis method enhances non-invasive preimplantation genetic testing accuracy by addressing maternal contamination and allele dropout, achieving nearly 100% detection of genetic diseases.

JP2026502019APending Publication Date: 2026-01-21BEIJING CHANGPING LAB +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024520732
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-20
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Current non-invasive preimplantation genetic testing (niPGT) methods for genetic diseases suffer from low amplification success rates, high false-positive and false-negative results, and diagnostic errors due to maternal contamination and allele dropout, making them unsuitable for clinical applications.

Method used

A transposon-based DNA amplification method combined with a Bayesian linkage analysis is employed to amplify low-input DNA samples from culture medium or blastocoelic fluid, addressing maternal contamination and allele dropout, and iteratively calculating likelihood ratios for accurate disease-carrying chromosome determination.

Benefits of technology

The method achieves nearly 100% detection accuracy for genetic diseases, reducing misdiagnosis rates and improving the reliability of non-invasive preimplantation genetic testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502019000001_ABST
    Figure 2026502019000001_ABST
Patent Text Reader

Abstract

The present invention relates to one or more accurate methods for non-invasive preimplantation genetic testing. Specifically, the present invention relates to systems and methods for amplifying minute amounts of DNA in solution, and linkage analysis methods for analyzing biological samples of culture medium from in vitro cultured embryos to determine susceptibility to genetic diseases in the embryos.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to one or more accurate methods for non-invasive preimplantation genetic testing. Specifically, the present invention relates to systems and methods for amplifying minute amounts of DNA in solution, and linkage analysis methods for analyzing biological samples of culture medium from in vitro cultured embryos to determine susceptibility to genetic diseases in the embryos. [Background technology]

[0002] Preimplantation genetic testing (PGT), known in the art (Handyside et al., first children born in 1989 and 1990), has been used to identify single-gene disorders (Kerem et al., 1989; Handyside et al., 1992; Liu et al., 1994; Harton et al., 1996; Ao et al., 1998; Sermon et al., 1998; Xu et al., 1999; Hussey et al., 1999; Ray et al., 2000; De Rycke et al., 2001; Moutou et al., 2001; Verlinsky et al., 2001; Sermon et al., 2001; Girardet et al., 2003; Fiorentino et al., 2006; Kahraman et al., 2007). al., 2014), aneuploidy (Munne et al., 1993; Wilton et al., 2001; Wells et al., 2002; Treff et al., 2010; Gutierrez-Mateo et al., 2011; Yang et al., 2012; Scott et al., 2013; Forman et al., 2013; Wells et al. al., 2014; Rubio et al., 2017), structural variation (Conn et al., 1998; Scriven et al., 1998; Coonen et al., 2000; Munne et al., 2000; Escudero et al., 2003; Melotte et al., 2004; Le Caignec et al., 2006; Traversa et al. al.,2010;Fiorentino et al. al., 2010 ; Fiorentino et al., 2010 ; Rius et al., 2011 ; Fiorentino et al., 2011 ; Treff et al., 2016 ; Zhang et al., 2017 ; Tan et al., 2013 ; Chow et al., 2018 ), or are widely used in clinical in vitro fertilization (IVF) to avoid the selection of embryos with polygenic disorders ( Treff et al., 2019 ; Kumar et al., 2022 ).It has been reported that the use of one NGS-based test, chromosome copy number analysis and linkage analysis, and targeted haplotyping for mutations can occur simultaneously (Yan et al., 2015).Currently, PGT requires a trophectoderm (TE) biopsy, which is invasive and subject to sampling bias (Dokras et al., 1990; Handyside et al., 1990; Kokkali et al., 2005).

[0003] The use of cell-free DNA in blastocoelic fluid (BF) for amplification and genetic testing has been reported as a minimally invasive preimplantation genetic test (Palini 2013, Gianaroli et al., 2014; Tobler et al., 2015; Magli et al., 2016; Zhang et al., 2016; Shangguan et al., 2017; Capalbo et al., 2018; Tsuiko et al., 2018; Magli et al., 2018).

[0004] Completely non-invasive methods for preimplantation genetic testing have been described for aneuploidies and monogenic disorders, using cell-free DNA in the culture medium of spent embryos (Galluzzi et al., 2015; Wu et al., 2015; Xu et al., 2016; Shamonki et al., 2016; Feichtinger et al., 2017; Lane et al., 2017; Liu et al., 2017; Ho et al., 2018; Vera-Rodriguez et al., 2018; Capalbo et al., 2018; Fang et al., 2019; Huang et al., 2019; Rubio et al., 2019; Yeung et al., 2019; Jiao et al., 2019; Ou et al., 2022).

[0005] Minimally invasive and non-invasive preimplantation genetic testing are two major methods for avoiding biopsy damage in preimplantation genetic testing. The embryo culture method and spent culture medium (SCM) collection approach, as well as the single-cell whole genome amplification (scWGA) method used for small amounts of cell-free DNA, can directly affect the results. The success rate of amplification of samples from blastocoelic fluid and spent embryo culture medium varies greatly among different clinical facilities. When using the same amplification method, the success rate of amplification in BF is much lower than that in spent culture medium (SCM), as reported by Galluzzi et al. (2015) as 62.5% for SCM vs. 44.4% for BF, and by Capalbo et al. (2018) as 89.7% for SCM vs. 27.4% for BF. Although the amplification success rate for SCM was higher than that for BF, SCM results were affected by maternal contaminants and the addition of exogenous proteins to the culture medium for embryo growth. Maternal DNA fragments in maternal blood are longer than fetal DNA fragments (Chan et al., 2004), and fragment length bias has been shown to make fetal DNA amplification ineffective. Furthermore, ensuring amplification accuracy while maintaining yield for clinical testing applications is crucial. Currently, commercially available single-cell amplification kits are based on DOP-PCR, MDA, or MALBAC technologies, and these kits vary in amplification efficiency and fidelity (Huang et al., 2015). All of these methods use exponential amplification, resulting in high false negative and false positive results. In 2017, a new amplification method was invented, using linear amplification (Chen et al., 2017; U.S. Patent No. 10,894,980, referred to as "LIANTI"). The false positive rate for SNP detection was much lower than with other methods. However, these kits have not been improved for DNA fragment length characteristics and the high protein content of the culture medium, resulting in a large reaction system and low amplification efficiency. An efficient and accurate amplification method to obtain more embryo DNA information in SCM and / or BF systems is urgently needed. However, the yield of the original LIANTI method is too low and reproducible, making it unacceptable for clinical use due to the precious nature of embryos.

[0006] Regarding analysis, while there are numerous reports on non-invasive preimplantation genetic testing (niPGT-A) for aneuploidies (reviewed by Leaver et al., 2020), there are only a few reports on non-invasive preimplantation genetic testing (niPGT-M) for monogenic disorders (Galluzzi et al., 2015; Zhang et al., 2016; Shangguan et al., 2018; Wu et al., 2015; Liu et al., 2017; Capalbo et al., 2018; Ou et al., 2022). Genotype concordance in successfully amplified SCMs compared with trophectoderm biopsy results also varied widely, ranging from 21% to 88% (Capalbo et al., 2018; Liu et al., 2017; Ou et al., 2022). False-positive and false-negative rates, important metrics for evaluating the accuracy of test methods, were not mentioned in these articles. However, based on the genotype concordance rates in the reported articles, the false-negative and false-positive rates were very high, making noninvasive PGT-M clinically unfavorable (Capalbo et al., 2018; Cimadomo et al., 2020). The most important criterion for evaluating PGT-M is the misdiagnosis rate, not the genotype concordance rate. The misdiagnosis rate published by the ESHRE PGT Consortium was very low (<0.1%) (De Rycke et al., 2017). The risk of misdiagnosis with new test methods should be evaluated and compared with conventional PGT-M.

[0007] In addition to detecting disease-associated mutation sites, the analytical method used in the reported paper was linkage analysis. When the allele dropout (ADO) rate is less than 5%, the minimum number of fully informative markers is recommended to be at least two STRs or SNPs proximal and distal to the mutation site (ESHRE 2020). Ou et al. used four informative SNPs, while Liu et al. used 10 informative SNPs. Because the ADO rate for SCM or BF is much higher than 5%, simply using a small number of informative sites for linkage analysis can easily lead to misdiagnosis. The distance of the informative SNP markers to the gene is important for residual recombination risk.

[0008] Furthermore, maternal contamination further led to diagnostic errors. Widely used linkage analysis methods are based on the Lander-Green algorithm and its variants, which do not consider maternal contamination and require high-quality sequencing data (Lander et al. 1987; Kruglyak et al. 1996; Idury et al. 1997; Kruglyak et al. 1998; Abecasis et al. 2002). Direct application of these methods results in high misdiagnosis rates (Capalbo et al. 2018; Cimadomo et al. 2020). Few linkage analysis methods take maternal cell admixture or contamination into account exist (Fan et al. 2002, Nabieva et al. 2020). However, they still require much higher-quality sequencing data than those from culture media, and their primary application is based on maternal blood samples during pregnancy or cells from the placenta or amniotic fluid (Fan et al. 2002, Nabieva et al. 2020). These existing methods, when applied to sequencing data from culture media, result in numerous diagnostic errors. The high false-positive and false-negative rates of niPGT-M are primarily related to the low initial DNA content and the presence of maternal contaminants in the used culture media, as well as the high ADO rate, making all existing niPGT-M methods undesirable for clinical application. The key to applying niPGT-M to clinical IVF is ensuring that the accuracy of this technology reaches the level of current PGT-M. For niPGT-M, accuracy should exceed 97% by detecting the presence of alleles carrying disease-causing mutations in embryos. To the best of our knowledge, no linkage analysis method for non-invasive procedures exists. The most important solution is to develop a more accurate linkage analysis method for samples with high ADO rates in the presence of maternal contaminants. Summary of the Invention

[0009] The present invention provides a DNA amplification method for amplifying low input amounts of DNA in solution, and a new linkage analysis method based on Bayesian models for PGT-M and preimplantation genetic testing for polygenic disorders (PGT-P) disease carriage analysis in the presence of maternal contaminants and haplotype loss.

[0010] Embodiments of the present disclosure provide improved wet-lab methods for DNA lysis, whole genome amplification (WGA), and next-generation sequencing (NGS) on limited samples (e.g., clinically spent culture medium, blastocoelic fluid, single cells, or limited numbers of cells) to achieve high coverage, high accuracy, and high amplification success rates. The systems and methods described herein are aimed at efficiently amplifying DNA with different fragment distributions in complex systems with high protein content.

[0011] This method is modified from the LIANTI method described in the art to efficiently increase amplification for higher product yield. The LIANTI sequencing data demonstrated high coverage, given sequencing depth, high accuracy indicated by low false-positive mTEs, and low chimera rates. High accuracy during single-cell genome amplification is a key requirement for single-nucleotide variant (SNV) detection, and low chimera rates are a key requirement for structural variant (SV) detection, both of which are critical for many applications based on single-cell genomics.

[0012] For example, in some embodiments, the present disclosure provides a method for lysing a clinical sample prior to amplification and a method for uniformly and efficiently amplifying a whole genome sample using a transposon-based method, the method comprising: a) removing serum proteins from the clinical sample and varying the lysis temperature and time to expose as much DNA as possible. In some embodiments, the cell-free DNA sample is in spent culture medium. In some embodiments, the cell-free DNA sample is in blastocoelic fluid. Aspects of the present disclosure also include b) improving subsequent amplification techniques, including, but not limited to, adjusting protease and / or transposome concentrations, adding DNA primers for reverse transcription, increasing the number of amplification cycles in the second-strand DNA amplification step, etc. As a result, accurate amplification can be achieved, and yields can be increased to suit sequencing analysis. The amplified products can be used in library preparation methods for next-generation sequencing, targeted amplification of disease locus regions, or chip sequencing.

[0013] According to one aspect, the present disclosure provides a linkage analysis method that addresses the problems of large fraction loss of haplotype (LFLoH), low coverage, and high maternal cell contamination (MCC) rates, which can hinder the estimation of inherited parental chromosomes at disease-causing mutation sites. According to the present disclosure, the linkage analysis described herein increases the number of SNPs or markers compared to conventional methods to address the high ADO rate, and further provides a set of linkage analysis methods based on Bayesian analysis, including: (a) calculating the MCC rate and likelihood incorporating the haplotype state of each SNP at each single SNP; (b) recursively estimating the MCC rate and haplotype state of each SNP; (c) calculating the likelihood of observing SNP data within regions flanking the disease-causing mutation site; and (d) determining whether an embryo carries the disease-causing chromosome and providing a confidence level.

[0014] Embodiments of the present disclosure provide accurate linkage analysis of samples with low DNA input, such as spent culture medium and / or blastocoelic fluid. The present disclosure also provides an amplification method for amplifying low amounts of DNA in samples such as spent culture medium or single cells. To improve accuracy and reduce the risk of misdiagnosis, the present disclosure also provides a Bayesian model for supporting PGT-M and PGT-P disease-carrying analysis. The linkage analysis iteratively calculates the likelihood ratio of inherited disease-carrying chromosomes relative to disease-free chromosomes, the MCC rate of SNPs, and the incorporation of haplotype status, and outputs a confidence level for the disease-carrying analysis. The analysis workflow is shown in Figure 1. The method described herein minimizes the false-negative rate of the output disease-carrying analysis and determines embryos most likely to carry healthy chromosomes. The Bayesian model takes as input the sequencing data and parental haplotypes staged by the Merlin algorithm.

[0015] To improve data quality, SNPs with extremely low quality were excluded. Several parameters were estimated before calculating likelihood ratios: sequencing error rate, MCC rate, and haplotype state. The MCC rate was defined as the proportion of DNA fragments originating from maternal chromosomes. The haplotype state was defined as the true parental origin of the DNA at each SNP and was divided into "paternal chromosome only," "maternal chromosome only," or "parental chromosomes." For example, parental chromosome only suggested that the DNA detected at this SNP originated only from the father. The MCC rate and haplotype state were iteratively calibrated until convergence.

[0016] We then recursively calculated the likelihood of sequencing data within a specific physical distance to the disease-causing mutation site. To complete the recursion, the recombination probability between adjacent SNPs was determined, and the likelihood at a single SNP was calculated. The recombination probability was obtained from the long-running DECODE dataset, while the single-SNP likelihood was obtained from a binomial model whose parameters were determined by the estimated sequencing error rate, MCC rate, and haplotype status.

[0017] In an embodiment of the present disclosure, instead of fixing the number of SNPs considered as in previous studies, we add SNPs one by one from the disease-causing mutation site and calculate the log-likelihood ratio for each SNP subset to obtain a curve of the log-likelihood ratio versus the physical distance of the terminal SNPs. Typically, after adding enough SNPs, the curve begins near zero on the vertical axis and eventually plateaus. Based on the characteristics of the curve, we determine which chromosome is inherited and classify the confidence level of our disease-carrying analysis into four categories: high confidence, medium confidence, possible, and uncertain. This method allows accurate allele discrimination in the presence of maternal contaminants. This method enables niPGT-M to achieve 100% detection accuracy. [Brief explanation of the drawings]

[0018] [Figure 1] 1 illustrates the workflow of one embodiment of an exemplary method for accurate non-invasive PGT. [Figure 2] 1 shows the mechanism of an exemplary amplification method involving DNA primer addition. [Figure 3A] The results of a single variable titration of the amplification method are shown, including the protease working concentration. [Figure 3B] Results of a single variable titration of the amplification method are shown, including the transposome working concentration. [Figure 3C] The results of a single variable titration of the amplification method are shown, including primers for first strand cDNA synthesis in reverse transcription. [Figure 3D] Results of a single variable titration of the amplification method are shown, including the PCR cycle for the second strand DNA step. [Figure 4] Graph showing genome coverage of each sample from Case 1, including patient / proband gDNA, amplified discarded embryo DNA, and amplified culture medium DNA. [Figure 5] Graph showing genome coverage of each sample from Case 2, including patient / proband gDNA, amplified discarded embryo DNA, and amplified culture medium DNA. [Figure 6]Schematic of genetic linkage analysis for family Case 1. The chromosome containing the slash harbors the disease-causing mutation. [Figure 7] PRS scores of three discarded embryos and their corresponding spent culture media are shown. [Figure 8A] Copy number (CN) plots of spent culture media are shown. Figure 8A: Aneuploid niPGT-A profile. [Figure 8B] Copy number (CN) plots of spent culture media are shown. Figure 8B: Mosaic niPGT-A profile. [Figure 8C] Copy number (CN) plots of spent culture media are shown. Figure 8C: Euploid niPGT-A profile. DETAILED DESCRIPTION OF THE INVENTION

[0019] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. All patents, patent applications, and publications mentioned herein are incorporated by reference.

[0020] As used herein, the term "genetic disease" refers to a disorder caused by an abnormality in an individual's DNA or genes. These disorders can be inherited from parents or can arise from naturally occurring genetic mutations.

[0021] As used herein, the term "monogenic disease" refers to a genetic disorder caused by mutations in a single gene. These diseases are typically passed on in a predictable manner from one generation to another, following Mendelian patterns of inheritance.

[0022] As used herein, the term "polygenic disease" refers to an inherited disorder caused by the combined effects of mutations in multiple genes. These diseases are often influenced by both genetic and environmental factors and do not follow simple Mendelian inheritance patterns.

[0023] As used herein, the term "solution" refers to a liquid mixture in which minor components (e.g., DNA) are distributed within a major component (e.g., a solvent). In particular, a solution is a complex system with a high protein content, but the DNA in the solution may be at the minimum picogram level. For example, a solution may be spent culture medium or blastocoelic fluid.

[0024] As used herein, the term "disease-causing chromosome" refers to a chromosome or chromosomal region that contains one or more genes that are responsible for a particular genetic disease.

[0025] As used herein, the term "disease-causing mutation site" refers to a specific location within a gene or chromosome where a mutation occurs, resulting in the development of a genetic disease.

[0026] The term "transposon," as used herein, refers to a nucleic acid sequence in DNA that can change its location within a genome. Transposons contain transposase binding sites and RNA polymerase promoter sequences, as described in U.S. Patent No. 10,894,980, entitled "Methods of Amplifying Nucleic Acid Sequences Mediated by Transposase / Transposon DNA Complexes," which relates to linear amplification via a transposon insertion amplification process, which may be referred to as "LIANTI," and is incorporated herein in its entirety.

[0027] As used herein, the term "transposome" refers to a set of transposase bound to transposon DNA, which is a transposase / transposon DNA complex dimer.

[0028] As used herein, the term "first strand cDNA synthesis" refers to the formation of single-stranded DNA by reverse transcribing an RNA transcript of the original DNA fragment.

[0029] As used herein, the term "second strand DNA synthesis" refers to the formation of a complementary strand to the first strand cDNA after RNA removal by RNase digestion.

[0030] As used herein, the term "allelic form" refers to a particular variant of a gene at a particular location (locus) on a chromosome. Different alleles can have different effects on a trait or disease.

[0031] As used herein, the term "allelic depth" refers to the number of reads (sequences) obtained for a particular allele at a given position, which indicates the abundance of that allele in a sample.

[0032] As used herein, the term "single nucleotide polymorphism (SNP)" refers to a single nucleotide base pair variation in a DNA sequence that occurs between members of a species. SNPs can influence traits, susceptibility to disease, and other genetic characteristics.

[0033] As used herein, the term "SNP density" refers to the number of SNPs obtained in a particular region of the genome from experimental data divided by the number of all human common SNPs in the database with a minor allele frequency greater than 0.01.

[0034] As used herein, the term "haplotype" refers to a combination of alleles (genetic variants) at multiple loci on the same chromosome that are inherited together due to physical proximity.

[0035] As used herein, the term "haplotype status" refers to the absence of paternally or maternally inherited alleles in experimental data at a particular SNP position.

[0036] As used herein, the term "haplotype loss" refers to a situation in which a particular allele of a haplotype is lost in a given experimental sample.

[0037] As used herein, the term "mapping" refers to aligning sequence reads to a reference genome to determine their locations.

[0038] As used herein, the term "SNP calling" refers to the process of identifying SNPs from sequencing data.

[0039] As used herein, the term "phasing" refers to the process of distinguishing between paternal and maternal chromosomes within a sequencing sample from any individual or embryo. This distinction allows the determination of the specific alleles derived from each parent, which is important for accurately understanding the inheritance patterns of genetic variants and how alleles of a haplotype are inherited together.

[0040] As used herein, the term "prephasing" refers to phasing parental haplotypes before disease carriage analysis is performed on a biological sample.

[0041] As used herein, the term "recombination probability" refers to the probability that genetic recombination (exchange of genetic material) will occur between two particular loci on a chromosome during meiosis.

[0042] The term "recursive algorithm" refers to a method of solving a problem by repeatedly applying the same steps until convergence occurs.

[0043] The term "a," "an," or "the" is intended to mean "one or more" unless specifically indicated to the contrary.

[0044] Wet Lab In one aspect, the present invention refers to a method for amplifying DNA in a solution, wherein the DNA content in the solution may be a minimum of picogram amounts, the method comprising the steps of: 1) Sample lysis: adding lysis buffer to the solution, heating, then adding protease and incubating to completely expose DNA, followed by inactivation under high temperature to obtain lysate; 2) DNA amplification: a. adding a transposome containing a transposase and a transposon containing a transposase binding site and an RNA polymerase promoter sequence to the lysate and inserting it into the DNA fragment; b. Filling the 9-bp gap caused by the Tn5 transposon insertion and extending to both ends of each fragment; c. In vitro transcription using the original sequence as a template; d. Adding DNA primers complementary to each RNA strand of the Tn5 transposition junction sequence to the reverse transcription system to form first strand cDNA via reverse transcription; and e. Forming second strand DNA via cycling amplification.

[0045] The transposon-based amplification method is based on that described in US Pat. No. 10,894,980.

[0046] In one general embodiment, a transposition system having an RNA polymerase promoter sequence is combined with an RNA polymerase and a reverse transcriptase along with a DNA primer to form a first-strand cDNA through reverse transcription, the DNA primers being complementary to each RNA strand of the Tn5 transposition junction sequence. In one embodiment, a transposon is inserted into DNA within a sample. An RNA polymerase is used to generate an RNA amplicon, which is then reverse transcribed into DNA along with the DNA primer. A complement to the DNA is generated, forming a double-stranded genomic DNA sequence.

[0047] According to one embodiment, a method for amplifying double-stranded DNA is provided, which includes contacting the DNA with a transposase (e.g., TN5 transposase) bound to transposon DNA, where the transposon DNA contains a double-stranded transposase (Tnp) binding site and an RNA polymerase promoter sequence, to form a transposase / transposon DNA complex dimer called a transposome. The transposon DNA can be in the form of a single-stranded extension or a loop in which each end is linked to the corresponding strand of the double-stranded transposase binding site. The transposome binds to a DNA target location along the double-stranded DNA and cleaves the double-stranded DNA into multiple double-stranded fragments, each of which has a first complex bound to the top strand by a Tnp binding site and a second complex bound to the bottom strand by a Tnp binding site. A transposon binding site is bound to each 5' end of the double-stranded fragment. According to one embodiment, the transposase is removed from the complex. The double-stranded fragment is extended along the transposon DNA to generate a double-stranded extension product with a T7 promoter at each end. According to one embodiment, gaps that may result from the binding of the transposase binding site to the double-stranded genomic DNA fragment can be filled and extended. The double-stranded extension product is contacted with an RNA polymerase (such as T7 RNA polymerase) to generate RNA transcripts of the double-stranded extension product. According to one embodiment, multiple RNA transcripts of the double-stranded extension product are generated using the RNA polymerase. The RNA transcripts are reverse-transcribed into multiple corresponding single-stranded DNAs using a reverse transcriptase and DNA primers complementary to each RNA strand of the Tn5 transposition binding site sequence. Strands complementary to the single-stranded DNAs are generated, forming multiple double-stranded DNAs. The double-stranded DNA can then be amplified and sequenced, for example, using high-throughput sequencing methods known to those skilled in the art.

[0048] According to a particular embodiment, an exemplary transposon system is the Tn5 transposon system.Other useful transposon systems are known to those skilled in the art and include the Tn3 transposon system (see Maekawa, T., Yanagihara, K., and Ohtsubo, E. (1996), A cell-free system of Tn3 transposition and transposition immunity, Genes Cells 1, 1007-1016), the Tn7 transposon system (see Craig, N.L. (1991), Tn7: a target site-specific transposon, Mol. Microbiol. 5, 2569-2573), the Tn10 transposon system (see Chalmers, R., Sewitz, S., Lipkow, K., and Crellin, P. (2000), Complete nucleotide sequence of Tn10, J. Bacteriol. 182, 2970-2972), the Piggybac transposon system (see Li, X., Burnight, ER, Cooney, AL, Malani, N., Brady, T., Sander, JD, Staber, J., Wheelan, SJ, Joung, JK, McCray, PB, Jr., et al. (2013), PiggyBac transposase tools for genome engineering, Proc. Natl. Acad. Sci. USA 110, E2279-2287), the Sleeping beauty transposon system (see Ivics, Z., Hackett, PB, Plasterk, RH, and Izsvak, Z. (1997), Molecular reconstruction of Sleeping Beauty, a Tc1-like transposon from fish, and its transposition in human cells, Cell 91, 501-510), and the Tol2 transposon system (see Kawakami, K. (2007), Tol2: a versatile gene transfer vector in vertebrates, Genome Biol. 8 Suppl. 1, S7).

[0049] In certain embodiments, an exemplary RNA polymerase is T7 RNA polymerase. Other useful RNA polymerases are known to those skilled in the art and include T3 RNA polymerase (see Jorgensen, ED, Durbin, RK, Risman, SS and McAllister, WT (1991) Specific contacts between the bacteriophage T3, T7, and SP6 RNA polymerases and their promoters, J. Biol. Chem. 266, 645-651) and SP6 RNA polymerase (see Melton, DA, Krieg, PA, Rebagliati, MR, Maniatis, T., Zinn, K., and Green, MR (1984) Efficient in vitro synthesis of biologically active RNA and RNA hybridization probes from plasmids containing a bacteriophage SP6 promoter, Nucleic Acids Res. 12, 7035-7056).

[0050] In one embodiment, the transposon DNA is designed to contain a double-stranded 19-bp Tn5 transposase (Tnp) binding site at one end linked or connected to a strong T7 promoter sequence. Upon transposition, Tnp and the transposon DNA bind to each other and dimerize to form a transposome. The transposome then randomly captures or otherwise binds to DNA in the sample. Next, the transposases in the transposome cleave the DNA, with one transposase cleaving the top strand and one transposase cleaving the bottom strand, to create DNA fragments. Thus, the transposon DNA is randomly inserted into the DNA, leaving a 9-bp gap on either side of the transposition / insertion site. The resulting DNA fragment has the transposon DNA Tnp binding site attached to the 5' position of the top strand and the transposon DNA Tnp binding site attached to the 5' position of the bottom strand.

[0051] After transposition, gap extension is performed to fill the 9-bp gap and complement the T7 promoter sequence originally engineered into the transposon DNA. This results in active, double-stranded T7 promoter sequences attached to both ends of each DNA fragment. Next, T7 RNA polymerase is added, and the DNA fragments are linearly amplified using in vitro transcription, generating multiple RNAs containing the same sequence as the original double-stranded DNA template. Finally, the amplified RNA is converted back into a double-stranded DNA molecule by reverse transcription using reverse transcriptase and a DNA primer as described above to form DNA, followed by second-strand synthesis. The DNA fragments can then be further processed for standard library preparation, amplification, and sequencing.

[0052] Specific Tn5 transposition systems have been described and are available to those skilled in the art. See Goryshin, IY and W.S. Reznikoff, Tn5 in vitro transposition. The Journal of biological chemistry, 1998. 273(13): pp. 7367-74; Davies, D.R., et al., Three-dimensional structure of the Tn5 synaptic complex transposition intermediate. Science, 2000. 289(5476): pp. 77-85; Goryshin, IY, et al., Insertional transposon mutagenesis by electroporation of released Tn5 transposition complexes. Nature biotechnology, 2000. 18(1): pp. 97-100 and Steiniger-White, M., I. Rayment, and W.S. Reznikoff, Structure / function insights into Tn5 transposition. Current opinion in structural biology, 2004. 14(1): pp. 50-7. Kits that utilize the Tn5 transposition system for DNA library preparation and other uses are known, each of which is incorporated herein by reference in its entirety for all purposes.See Adey, A., et al., Rapid, low-input, low-bias construction of shotgun fragment libraries by high-density in vitro transposition. Genome biology, 2010. 11(12): p. R119; Marine, R., et al., Evaluation of a transposase protocol for rapid generation of shotgun high-throughput sequencing libraries from nanogram quantities of DNA. Applied and environmental microbiology, 2011. 77(22): p. 8071-9; Parkinson, N. J., et al., Preparation of high-quality next-generation sequencing libraries from picogram quantities of target DNA. Genome research, 2012. 22(1): p. 125-33; Adey, A. and J. Shendure, Ultra-low-input, tagmentation-based whole-genome bisulfite sequencing. Genome research, 2012. 22(6): p. 1139-43; Picelli, S., et al., Full-length RNA-seq from single cells using Smart-seq2. Nature protocols, 2014. 9(1): p. 171-81 and Buenrostro, J. D., et al., Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position. Nature methods, 2013.each of which is incorporated herein by reference in its entirety for all purposes. See also WO 98 / 10077, EP 2527438, and EP 2376517, each of which is incorporated herein by reference in its entirety. A commercially available transposition kit is available from Illumina, marketed under the name NEXTERA.

[0053] In vitro transcription (IVT) with T7 RNA polymerase is useful for generating RNA amplicons. Van Gelder,RN,et al.,Amplified RNA synthesized from limited quantities of heterogeneous cDNA.Proceedings of the National Academy of Sciences of the United States of America,1990.87(5):p.1663-7;Kawasaki,ES,Microarrays and the gene expression profile of a single cell.Annals of the New York Academy of Sciences,2004.1020:p.92-100;Livesey,FJ,Strategies for microarray analysis of limiting amounts of RNA.Briefings in functional genomics&proteomics,2003.2(1):p.31-6;Tang,F., K.Lao, and MASurani,Development and applications of single-cell transcriptome analysis.Nature methods,2011.8(4 Suppl):p.S6-11; Hashimshony, T.,et See, e.g., Shankaranarayanan, P., et al., Single-tube linear DNA amplification for genome-wide studies using a few thousand cells. Nature protocols, 2012.7(2):328-38, each of which is incorporated herein by reference in its entirety.According to the present disclosure, IVT advantageously provides linear amplification, with all copies generated from the original DNA template. The resulting RNA molecules are reverse transcribed into single-stranded DNA, followed by the formation of a complementary strand, resulting in double-stranded DNA, which is the amplicon linearly amplified from the original DNA template.

[0054] In one embodiment, the solution is culture medium or blastocoelic fluid. In particular, culture medium or blastocoelic fluid has a high protein content and a low DNA input. According to one aspect described below, the method is applied to culture medium by way of example only.

[0055] In one embodiment, the lysis buffer is 2x to 10x lysis buffer. The original volume of culture medium is much larger than the volume of a single cell. Amplification of the highest possible percentage of culture medium was achieved by using a very small amount of highly concentrated lysis buffer. This maximizes the percentage of culture medium DNA used as a template in this amplification system.

[0056] In one embodiment, the working concentration of the protease in the lysis mixture is greater than 0.6 mg / mL, preferably in the range of 0.8-2 mg / mL. For example, the working concentration of the protease is 0.6 mg / mL, 0.8 mg / mL, 1 mg / mL, 1.2 mg / mL, 1.4 mg / mL, 1.6 mg / mL, 1.8 mg / mL, 2 mg / mL, or more. The amount of protease is increased to remove excess exogenous serum proteins.

[0057] In one embodiment, the incubation is carried out at 45 to 60°C for 1 to 3 hours.

[0058] In one embodiment, inactivation is carried out at 80-90° C. Reducing the protease reaction time and increasing the protease inactivation temperature ensures that there are no excess protease residues in the subsequent reaction.

[0059] In one embodiment, the working concentration of the transposome is greater than 6 nM, preferably in the range of 10 to 40 nM, for example, the working concentration of the transposome is 6 nM, 6.25 nM, 8 nM, 10 nM, 12.5 nM, 14 nM, 16 nM, 18.75 nM, 20 nM, 22 nM, 24 nM, 26 nM, 28 nM, 30 nM, 32 nM, 34 nM, 36 nM, 37.5 nM, 40 nM, or more.

[0060] In one embodiment, the DNA primer has a length of 15 to 25 bp. Preferably, the 3' end of the DNA primer contains the sequence 5'-GACAG-3' or 5'-CTGTC-3'. For example, random primers 5'-NNNNNNNNNNNNNNGACAG-3' or 5'-NNNNNNNNNNNNNNCTGTC-3' can be used. More preferably, the DNA primer contains the sequence shown as 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO: 1) or 5'-TCTACACATATTCTCTGTC-3' (SEQ ID NO: 2), or has at least 70%, preferably 80%, preferably 90%, preferably 95%, and preferably 100% identity to SEQ ID NO: 1 or SEQ ID NO: 2. For example, primer 5'-AGATGTGTATAAGAG-3' can be used.

[0061] The distribution of DNA in the culture medium was measured, confirming the presence of both long and short fragments of cell-free DNA in the culture medium. Because all currently available kits tend to amplify long fragments, the transposon-based amplification method described herein, based on that described in U.S. Patent No. 10,894,980, as amended herein, is designed to amplify both long fragments (e.g., in some embodiments, greater than 1 kb) and short fragments (e.g., in some embodiments, 100 bp to 600 bp, or less than 1 kb). The transposon amplification method described herein does not use self-priming; for highest efficiency, it uses DNA primers added at appropriate concentrations during the reverse transcription step. RNA generated by in vitro transcription is characterized by the sequence of the complete transposon DNA sequence at the 3' end. The two primers designed for reverse transcription, designated 19_1 and 19_2, and their binding positions are shown in Figure 2.

[0062] In one embodiment, cycling amplification is performed for more than four cycles, e.g., 4 to 20 cycles, preferably 6 to 15 cycles. After the formation of the first strand of cDNA in reverse transcription, additional amplification cycles are used to amplify the second strand of DNA. For example, linear amplification, as described in U.S. Pat. No. 10,894,980 and C Chen, D Xing, L Tan, H Li, G Zhou, L Huang, and X X Xie, Science 2017;356(6334):189-194, is a highly accurate amplification, but the product yield of linear amplification is lower than that of exponential amplification. To facilitate clinical testing, the present disclosure provides a method for increasing the amplification yield. In one embodiment, this is achieved during the reaction, in which the second strand is formed by increasing the number of amplification cycles to obtain sufficient DNA for sequencing library preparation.

[0063] The methods disclosed herein achieve nearly 100% amplification success rate while maintaining amplification fidelity.

[0064] Dry Lab In another aspect, the present invention refers to a method for determining a susceptibility of an embryo to a genetic disease by analyzing a biological sample of culture medium of an in vitro cultured embryo, the method comprising: 1) amplifying DNA molecules from biological samples of culture medium of in vitro cultured embryos and / or from biological samples of discarded embryos within the family; 2) preparing a DNA library and sequencing DNA molecules from a biological sample to obtain multiple sequence reads; 3) Performing mapping and SNP calling, such as by a computer system, to generate and quality control SNP data; 4) Performing haplotype prephasing using a computer system or other method on blood samples from both parents or one parent with the disease; 5) Calculating, by a computer system or the like, the likelihood of four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of both parents, taking into account the maternal cell contamination (MCC) rate and haplotype status; 6) Estimating the MCC rate, such as by a computer system, and determining the haplotype status of each SNP to indicate whether DNA from only one parent is present in the culture medium; 7) calculating, such as by using a computer system, the likelihood of observing SNP data within the region adjacent to the disease-causing mutation site, and combining the likelihood and recombination probability at each single SNP; 8) To determine whether an embryo carries a disease-causing chromosome and provide a confidence level.

[0065] In some embodiments disclosed herein, steps 5 and 6 above provide specific technical improvements over some existing techniques that use SNP data to determine whether an embryo carries a disease-causing chromosome. For example, if a wet lab provides very high-quality biological samples that yield highly precise and accurate SNP data free of maternal contaminants from the sample, one or more of the calculations performed in steps 5 and 6 may only result in a slight improvement in accuracy. For example, methods for amplifying trace amounts of DNA in solution (described in step 1 above) are sometimes used to improve the quality of the resulting SNP data. Nevertheless, despite attempts to improve the quality of biological samples (e.g., by keeping them contaminant-free, amplifying them, etc.), this is not always possible to the desired degree or confidence level. Furthermore, high ADO rates and maternal contaminants are inherent properties of the DNA content in certain biological samples, such as culture medium from in vitro-cultured embryos, and are unavoidable even if wet labs can reach near-perfect performance. In contrast to prior art systems, the novel and non-obvious calculations performed in steps 5 and 6 improve the accuracy of analyzing biological samples of culture medium from in vitro-cultured embryos to more accurately determine genetic disease susceptibility in the embryos. Rather than relying solely on wet labs and their known techniques to prepare very high-quality biological samples for linkage analysis and / or non-invasive preimplantation genetic testing, the present disclosure includes one or more additional computer operations performed in steps 5 and / or 6 to overcome various shortcomings of the prior art. In other words, prior art systems lack linkage analysis methods for non-invasive procedures similar to those disclosed herein, e.g., the significantly more accurate linkage analysis methods for samples with high ADO rates in the presence of maternal contaminants that some embodiments disclosed herein enable. Steps 1 through 8 are further detailed as follows:

[0066] Step 1): Amplify DNA molecules from biological samples of culture medium from in vitro cultured embryos and / or from biological samples of discarded embryos within the family.

[0067] In one embodiment, the embryo has contaminants and / or haplotype losses that originate from maternal cells.

[0068] If neither the proband nor the grandparents are present, biological samples from discarded embryos within the family are used for inference after prephasing, as described in detail below.

[0069] Step 2): Prepare a DNA library and sequence DNA molecules from the biological sample to obtain multiple sequence reads.

[0070] Amplicons were sheared and library preparation was performed. DNA sequencing libraries were sequenced on an Illumina platform to generate raw sequencing data.

[0071] Step 3): Performing mapping and SNP calling to generate SNP data and controlling its quality by a computer system.

[0072] Raw end read pairs were trimmed from Illumina sequencing adapters, inline barcodes, junction and T7 promoter sequences, and low-quality ends from cutadapt. Clean reads were mapped to the human reference genome hg19 using BWA mem. Duplicates were annotated in each BAM file and sorted using MarkDuplicatesSpark in GATK 4.2.0. Base quality score recalibration (BQSR) was performed using BaseRecalibrator to generate a recalibration table, followed by ApplyBQSR, which adjusted the base quality accordingly. For variant calling, GATK HaplotypeCaller generated a genomic variant call format (GVCF) file and a VCF file for each sample. The GVCF was utilized to perform joint genotyping across all samples. SNPs were filtered using variant quality score recalibration (VQSR) using HapMap 3.3, Omni 2.5, 1000 Genome phase I, and dbSNP 151 as the SNP training set. A 99% sensitivity threshold was selected to accurately filter SNPs. Biallelic SNPs were retained for subsequent analysis. The VCF includes the chromosome number, physical location, identification number (ID), and allele type for each SNP in each blood sample, biological sample of in vitro cultured embryo culture medium, and / or biological sample of discarded embryos within the family, as well as sequencing quality, sequencing depth (DP), and allele depth (AD).

[0073] SNPs with two different alleles are considered, and all SNPs with a DP value of less than 5 in blood samples, or a sequencing quality value (QUAL) of less than 30, or a gene quality (GQ) value of less than 10 are excluded, as their quality is considered too low and the signal-to-noise ratio is too small.

[0074] After removing low-quality SNPs, the flanking region of the disease-causing mutation site in the biological sample is defined as 10 MB (1 million base pairs) upstream and downstream of the disease-causing mutation site. The sequencing error rate (SER) in the flanking region is estimated. SNPs with either a paternal or maternal genotype of 00 or a paternal or maternal genotype of 11 are selected. Ideally, only genotypes identical to those of the parents should be detected at such SNPs. Therefore, sequencing data with genotypes different from those of the parents must be erroneous reads. Therefore, the estimated sequencing error rate (SER) is defined as the ratio of the total number of error reads to the total number of reads.

[0075] The relative density of SNPs detected in the flanking regions of disease-causing mutation sites was estimated and defined as the ratio of the total number of human common SNPs detected in this region to the total number of human common SNPs in this region. Here, human common SNPs are defined as all SNPs with a minor allele frequency (MAF) of 0.01 or greater, and the minor allele frequency of a population of alleles refers to the frequency of the second most common allele in the population. Relative SNP density for each sample was calculated using the human common SNP data updated in 2018 from dbSNP [Sherry et al., 2001].

[0076] After the above steps are completed, culture medium samples with abnormal MCC rates (>0.65 or <-0.5), low SNP density (<0.001), or high sequencing error rates (>0.2) are considered to be of extremely low quality and are therefore excluded. Such culture medium samples are classified as "indeterminate" and the reason for the low data quality is output. All samples that pass quality control proceed to the next stage.

[0077] Step 4): Haplotype prephasing is performed on blood samples from both parents or one of the parents carrying the disease by a computer system.

[0078] To estimate parental haplotypes, genomic data from one or more of the following samples within the family are used: a blood sample from the proband, a biological sample from a discarded embryo, or a blood sample from at least one of the grandparents. If neither the proband nor the grandparents are present, discarded embryos are used for estimation. Because discarded embryos are amplified samples that may contain errors in the amplification process, to ensure accuracy of estimation, estimation is performed using data from at least two discarded embryos in parallel to correct for Mendelian errors.

[0079] Plink 1.9 with the option "--Mendel" is used to identify sites with Mendelian errors. BEAGLE 4.0 is then used to phase the genotype data, taking into account parentage and simultaneously excluding sites with Mendelian errors.

[0080] Step 5): Calculate the likelihood of four possible inheritance scenarios at each single SNP, indicating whether the inheritance strand is inherited from the paternal or maternal allele of both parents, taking into account the maternal cell contaminant (MCC) rate and haplotype status.

[0081] For single SNP likelihood, a binomial distribution was used to build the model (see the appendix section of Nabieva et al.

[2020] ), and the parameters of the binomial distribution used here are the sequencing error rate,

number

number

number

number

[0082] Here, ○ indicates a product by position. The paternal genotype at the nth SNP is expressed as fGT = fGT1 / fGT2, and the maternal genotype is expressed as mGT = mGT1 / mGT2, where fGT1, fGT2, mGT1, mGT2∈{0,1}.

number

number

number

number

[0083] In some embodiments, the same computer system can perform all of the steps of the method for determining susceptibility to a genetic disease in an embryo. In other embodiments, one or more separate computer processors within the computer system (e.g., computer subsystems) may perform one or more of the steps of the aforementioned method, without necessarily performing all of them. In other words, the computer system can include multiple subsystems, each with its own computer processor, thereby distributing the computational load of performing the multiple steps described in the aforementioned method.

[0084] Step 6): Estimate the MCC rate and determine the haplotype state of each SNP, which represents whether DNA from only one parent is present in the culture medium.

[0085] Step 6) comprises the following steps: Step (a): further comprising first estimating the MCC rate and haplotype status of each SNP in the region flanking the disease-causing mutation site.

[0086] The maternal cell contaminant concentration (MCC) rate is defined as the ratio of maternal DNA to the total amount of DNA in the embryo culture medium; theoretically, this value should be between 0 and 1. All SNPs containing paternal genotype 00 and maternal genotype 11, or containing paternal genotype 11 and maternal genotype 00 within adjacent regions, are selected. For these SNPs, the paternal or maternal origin of each read detected in the culture medium is theoretically determined, and the ratio of maternal reads to total reads is calculated and used as an initial estimate of the MCC rate. In contrast to prior art systems that omit linkage analysis for non-invasive procedures similar to those described herein, the present disclosure provides various computational operations on SNP data that enable more accurate linkage analysis, for example, for samples with high ADO rates in the presence of maternal contaminants, as described herein. In particular, in wet laboratories using non-invasive procedures, the missing DNA fragments from one parent are often longer and contain dozens of SNPs, which can sometimes cause data quality degradation, thereby adversely affecting and reducing sequencing capacity.The computer operations disclosed herein solve a real technical problem by enabling otherwise unusable biological samples to be useful for determining susceptibility to genetic diseases in embryos.

[0087] Disclosed herein is a method for preparing a biological sample of culture medium from an in vitro-cultured embryo, where the underlying biological sample is of low quality, particularly due to a high ADO rate in the presence of maternal contaminants, or otherwise unsequenceable. This method converts SNP data from an otherwise unsequenceable biological sample into a state that can be analyzed to determine the susceptibility of a genetic disease in an embryo. For the purpose of determining the susceptibility of a genetic disease in an embryo, a low-quality DNA sample results in data that is nearly unsequenceable. Furthermore, because embryonic DNA enters the culture medium as fragments, if certain longer DNA fragments do not enter the culture medium, sequencing becomes impossible (e.g., a phenomenon defined herein as Large Fragment Loss of Haplotypes (LFLoH)). The present disclosure describes a method for converting seemingly unsequenceable data into a state in which the resulting SNP data overcomes the drawbacks of LFLoH, making previously unusable biological samples readily usable. Because SNP data obtained from wet labs are based on biological samples of less than ideal quality due to the aforementioned conditions, some embodiments, particularly for the purpose of determining whether embryos awaiting implantation carry disease-carrying chromosomes, perform novel and nontrivial computational steps on the SNP data to convert seemingly unusable data into usable data. The innovative computational steps disclosed herein can avoid another round of sequencing in the wet lab, thus saving time, materials, labor, and other resources. As detailed in step 6(b) below, analysis of LFLoH provides insight into whether parental DNA fragments are detected at each SNP, so that the final haplotype status labels at all SNPs can be iteratively updated to become more accurate until convergence of MCC rates or other final estimation criteria is achieved.

[0088] It should be noted that the preparation methods described above are distinct from processing methods. The preparation methods recited herein do not simply identify a natural correlation between two naturally occurring characteristics of a biological sample. For example, simply using a natural correlation between elevated cfDNA levels in a collected bodily sample and organ transplant health status to identify potential organ rejection may not be sufficient to recite patent-eligible subject matter. In contrast, in some examples disclosed herein, the computational methods disclosed in steps 5 and / or 6 improve the quality of previously unusable biological samples, transforming them into a state that provides prospective parents with accurate, non-invasive verification of their embryo's susceptibility to genetic disease. By avoiding invasive procedures, such as inserting a large needle into a pregnant woman's amniotic sac, many deaths of pregnant women and fetuses can be avoided. The disclosed preparation methods provide a technical solution to a real-world problem that previously posed higher risks to fetuses, women, and their health.

[0089] Ideally, it would be advantageous to be able to detect DNA from both parents in embryo culture medium. However, in practice, it has been observed that a significant portion of the detectable genome in embryo culture medium originates from only one parent. This type of DNA loss differs from the traditional meaning of allelic dropout (ADO), which occurs due to amplification heterogeneity. It results in some short DNA fragments that are not successfully amplified in a particular amplification round and are therefore significantly less amplified than other DNA fragments that should be detected. It reflects the efficiency of the amplification method and typically requires short DNA fragments containing only about one or two SNPs. However, in culture medium samples obtained by non-invasive methods, the missing DNA fragments from one parent are often longer and contain dozens or more SNPs. This haplotype loss is not due to low amplification efficiency, but rather to poor data quality. Because embryonic DNA enters the culture medium as fragments, the absence of certain longer DNA fragments makes sequencing impossible. We define Large Fragment Loss of Haplotypes (LFLoH) as the phenomenon in which long, contiguous DNA fragments from one or both parents do not enter the culture medium and are therefore undetectable by sequencing. This disclosure provides a model for addressing LFLoH in a two-step process.

[0090] First, SNPs that can be determined directly from the AD value are labeled as containing either paternal DNA only (F), maternal DNA only (M), or both (FM), while others remain unlabeled.

[0091] In the second step, the SNPs upstream and downstream of these labeled SNPs are searched for, and then it is determined whether the surrounding SNPs can be labeled with the same label as the SNP labeled in the first step. Specifically, in the first step, T is assumed to be a predetermined integer threshold (the default value is 3).

[0092] For SNPs in the flanking regions of disease-causing mutation sites, those that satisfy one of the following conditions are labeled F: The paternal genotype is 00 and the maternal genotype is 11. AD0>T and AD1=0 The paternal genotype is 11 and the maternal genotype is 00. AD1>T and AD0=0 The paternal genotype is 01 and the maternal genotype is 00. AD0=0 and AD1>T The paternal genotype is 01 and the maternal genotype is 11. AD0>T and AD1=0

[0093] Some SNPs labeled F can be further labeled F1 or F2, indicating that the detected SNP originates from the father's paternally derived or maternal chromosome. Specifically, an SNP labeled F is labeled F1 if it meets one of the following conditions: The paternal genotype is 1|0 and the maternal genotype is 00. AD1>T and AD0=0 The paternal genotype is 0|1 and the maternal genotype is 11. AD0>T and AD1=0, In the formula, 0|1 and 1|0 represent the paternal genotype, and the numbers to the left and right of | represent the paternal-origin gene and maternal-origin gene of the father at this SNP, respectively.

[0094] Similarly, a SNP labeled F is labeled F2 if it meets one of the following two conditions: The paternal genotype is 0|1 and the maternal genotype is 00. AD1>T and AD0=0 The paternal genotype is 1|0 and the maternal genotype is 11. AD0>T and AD1=0, SNPs that satisfy one of the following conditions are labeled M: The paternal genotype is 00 and the maternal genotype is 11. AD1>T and AD0=0 The paternal genotype is 11 and the maternal genotype is 00. AD0>T and AD1=0

[0095] No SNP is labeled M1 or M2 due to the presence of MCC. SNPs that meet one of the following conditions are labeled FM: · The paternal genotype is 00 and the maternal genotype is 11. AD1>[T / 2] and AD0>[T / 2]; · The paternal genotype is 00 and the maternal genotype is 11. AD1>[T / 2] and AD0>[T / 2]; · The paternal genotype is 01 and the maternal genotype is 11. AD1>[T / 2] and AD0>[T / 2]; · The paternal genotype is 01 and the maternal genotype is 00. AD1>[T / 2] and AD0>[T / 2]; In the formula, [ ] represents the floor function.

[0096] After completing the initial labeling procedure, there are five possible labels: F, M, FM, F1, and F2. The ordinal numbers - N1 ≤ s1 ≤ s2 <... D Suppose there are several labeled SNPs with N2 ≤ . Then, the specific labeled SNPs s d For ordinal s d -1 and s d A search is performed among the SNPs with +1. d The th SNP is called the central SNP, and the s d -1st and sth d The area between the first and second points is called the search area. The search area is divided into two parts: d The nth SNP is divided into an upstream portion and a downstream portion of the nth SNP. In the downstream search, a label is assigned to each SNP in the downstream search region. The nth SNP is then assigned a label during the downstream search from the nearest labeled SNP upstream of it.

number

number

[0097] Only the search process in the downstream region is described as an example below: First,

number

number

number

[0098] where:

number

[0099] Finally, as a result of the search procedure, there will be a maximum of two labels for each SNP. For those SNPs with less than two labels, the label will be added as "unknown". Then, the two labels for each SNP are merged to complete the LFLoH labeling procedure. According to the principle of final labeling, only after sufficient confidence that only paternal or maternal DNA fragments are detected, the SNP will be labeled as F or M; otherwise, the SNP will be labeled as FM. Specifically, the final label for the nth SNP is

number

[0100] Step (b): Recursively update the MCC rate and haplotype state of each SNP. The initial calculation of the MCC rate in the previous section was based on SNPs with homozygous but distinct parental genotypes, and it was assumed that both paternal and maternal DNA fragments were detected at these SNPs. However, analysis of large fraction loss of haplotype (LFLoH) can provide new insight into whether parental DNA fragments are detected at each SNP, facilitating a re-estimation of the MCC rate. The new estimate of the MCC rate should change the likelihood and update the final labeling at all SNPs. Because the labeling at SNPs is closely intertwined with the MCC rate estimate and both cannot be estimated simultaneously, an iterative estimation method is employed, whereby each estimate is updated by the current estimate of the other, and this process is repeated until convergence is reached. In each stage, the MCC rate is first estimated using all SNPs labeled as FM, and then the LFLoH for each SNP is estimated using the new MCC rate estimate. This process is repeated until the difference between two adjacent MCC rate estimates is within a threshold (set at 0.1 or 0.05), and the estimate in the final iteration is used as the final estimate of the MCC rate, and the corresponding LFLoH state is taken as the final indicator of haplotype state. If the MCC rate estimates do not converge after 10 rounds of iteration, the average of the 10 MCC rate estimates is considered the final estimate, and this estimate is used to update the haplotype state.

[0101] At least one advantage of the recursive procedure described herein is that it reduces the calculation of likelihood function to the calculation of recombination probability and single SNP likelihood.The computational load of the processor in a computer system is improved due to the technological advances resulting from the recursive procedure.In addition to reducing the processor load, the recursive procedure can reduce memory usage because it integrates the equations to be calculated.The recursive procedure described herein is novel and non-obvious, not well understood, not conventional, and not conventional.

[0102] Step 7: Combine the likelihood at each single SNP with the recombination probability to calculate the likelihood of observing SNP data in the region adjacent to the mutation site that causes the disease.

[0103] The present disclosure provides a method for recursively calculating likelihoods. First, if the mutation site causing the disease is on the paternal chromosome, the final haplotype state is labeled as F or FM, and all SNPs with heterozygous paternal genotypes are selected for recursion. If the mutation site causing the disease is on the maternal chromosome, the final haplotype state is labeled as M or FM, and all SNPs with heterozygous maternal genotypes are selected. The mutation site causing the disease is still shown as the 0th SNP. Among all the retained ones, the SNPs downstream of it have ordinal numbers 1, 2,..., N2, while the SNPs upstream of it have ordinal numbers -1, -2,... -N1. For the i-th SNP,

Equation

[0104] AD i is represented as the allele depth (AD) at the i-th SNP. When i < j, AD i : j ={AD i ,AD i+1 ,...,AD i+j}. Given the genotype and haplotype label at the mutation site causing the disease, the upstream and downstream recombination and sequencing data are conditionally independent, and the upstream and downstream analyses are treated separately. Here, the recursive formulation for the downstream region will be further explained. Given the observation of all detected AD values and the inferred haplotype states for SNPs in the adjacent region, the probability that the paternal embryonic chromosome is inherited from the grandparents, and the probability that the maternal embryonic chromosome is inherited from the grandparents are calculated. Specifically, the goal is

Equation

[0105] The Bayes formula below

number

number

[0106] The next step is

number

number

[0107] The single SNP likelihood is

number

number

[0108] P(F n+1 =k,M n+1 =l|F n =i,M n =j) is the recombination probability.

[0109] The above recursive procedure reduces the calculation of the likelihood function to the calculation of recombination probability and single SNP likelihood. The estimated recombination probability between any two positions on the human genome is obtained from a high-precision database of recombination rates described in Kong et al.

[2002] . The single SNP likelihood was described in step (5).

[0110] Step 8): Determine whether the embryo carries the disease-causing chromosome and provide a confidence level.

[0111] After calculating the likelihood, the log-likelihood ratio (LLR) is calculated to determine which of the two paternal chromosomes is inherited by the embryo, and similarly for the maternal chromosome. If the disease-causing mutation site is on the paternal chromosome, then:

number

[0112] A positive LLR indicates that the data considered support the allele being inherited from the father of the affected parent; otherwise, the data support it being inherited from the mother of the affected parent.

[0113] However, it is recognized that the number of SNPs used in the likelihood calculation cannot be determined in advance: if a specific fixed number of SNPs are used in the calculation, or if all SNPs within a specific fixed physical distance to the site of the disease-causing mutation are used in the calculation, it cannot be determined with certainty whether the information implied by these SNPs is sufficient to support a conclusion, and to what extent the likelihood will change depending on the number of SNPs used.

[0114] Therefore, a log-likelihood ratio curve is plotted as a function of the number of SNPs used, and the properties of this curve are used to determine whether an embryo has a disease-causing mutation site. SNPs are added sequentially to the likelihood calculation, and for each N2 = 1, 2, ..., the likelihood is calculated based on which parent's chromosome the disease-causing mutation site is located on, and the results are

number

number

[0115] Finally, if the number of SNPs used is large enough to determine whether an embryo carries a disease-carrying allele, adding new SNPs will no longer significantly change the log-likelihood ratio. Therefore, when the vertical coordinate of the curve tends to stabilize, the sign of the stable value can be used as a criterion for determining whether an embryo carries a disease-carrying allele, and its absolute value serves as a measure of the confidence level. In practice, the steady-state log-likelihood ratio (SLLR) downstream of the disease-causing mutation site can be used as a measure of the confidence level.

number

[0116] Here, ε0 is a predetermined threshold. In our experiments, ε = 0.1. Here, N0 is the number of SNPs required to confirm stability, which is 30 by default. If there are not 30 consecutive SNPs with a log-likelihood range within 0.1, N0 is reduced to 10 to detect whether the curve is stable. If either the upstream or downstream region has fewer than 10 SNPs in the data, this region is ignored, and a confidence interval based on the opposite curve is determined. This occurs when the disease-causing mutation site is located at either end of a chromosome or near the spindle. Similarly, we calculate the upstream stationary log-likelihood ratio (SLLRup) and denote their arithmetic mean as the average stationary log-likelihood ratio (ASLLR).

[0117] According to the SLLR, ASLLR, and other characteristics of the log-likelihood curve, the confidence level of the decision is classified into four categories: highly confident, moderately confident, possible, and indeterminate. Specifically, if SLLRup·SLLRdown<0 for a sample, the sample is immediately classified as "indeterminate" because opposite conclusions are obtained from the upstream and downstream regions. For the remaining samples, the maximum differential log-likelihood ratio (MDLLR) is defined as follows:

number

[0118] Finally, the confidence level is

number

[0119] Following the above protocol, the genotype and which alleles are inherited from the parents at any single SNP can be determined. This information can then be used to perform disease prevalence analysis for monogenic diseases, and multiple identified SNPs in the embryo can be combined to determine susceptibility to polygenic diseases.

[0120] The present invention also refers to a method for analyzing chromosomal aneuploidy by means of a biological sample of culture medium of an in vitro cultured embryo for performing preimplantation genetic testing for aneuploidy.

[0121] In some examples, an optional additional step 9 may be performed after the determination of whether the embryo carries a disease-causing chromosome is made in step 8. In one example, a determination made by the computer system having a confidence level above a threshold amount that the embryo is likely to carry a disease-carrying chromosome can cause the computer system to output a warning and advise the prospective parents to discard the embryo. Conversely, if the embryo is determined to carry a disease-carrying chromosome but the determination is below a threshold confidence level, the computer system can generate a warning to appropriate parties. In one example, a specific treatment method (e.g., preserve the embryo using the same, similar, or other non-invasive or invasive methodology, or other options; preserve the embryo and transfer it; discard the embryo; retest the embryo) for a particular patient whose medical history is known to the medical facility / provider can be generated by the computer system for use by a medical professional to advise the prospective parents. In another example, one or more method steps described herein can be used to select a patient's embryo for a specific treatment or additional testing or preparation to achieve a particular result. For example, for some embryos that do not pass quality control, the computer system may not determine whether they carry a disease-carrying chromosome, in which case the embryo may be tested again using one or more non-invasive or invasive techniques, or may be discarded. Meanwhile, for embryos that pass quality control, the computer system may output a confidence level indicating whether they carry a disease-carrying chromosome. Furthermore, embryos with a high confidence level above a threshold level and diagnosed as not carrying a disease-carrying chromosome may be selected for implantation. Meanwhile, for embryos with a confidence level below the high confidence level, the computer system may generate a warning advising a medical professional of several options, including implantation with some assumed risks, additional testing of the embryo, and other options.

[0122] Computer Systems Any of the computer systems referred to herein may utilize any suitable number of subsystems. In some embodiments, a computer system may include a single computer device, where the subsystems may be components of the computer device. In other embodiments, a computer system may include multiple computer devices, each of which is a subsystem, having internal components.

[0123] A computer system may include, for example, multiple identical components or subsystems connected to each other by external or internal interfaces. In some embodiments, computer systems, subsystems, or devices may communicate over a network. In such cases, one computer may be considered a client and another computer a server, each of which may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0124] It should be understood that any of the embodiments of the present invention can be implemented in the form of control logic using hardware (e.g., application specific integrated circuits or field programmable gate arrays) and / or computer software with processors that are generally programmable in a modular or integrated manner. Based on the disclosure and teachings provided herein, those skilled in the art will know and understand other ways and / or methods to implement embodiments of the present invention using hardware and combinations of hardware and software.

[0125] Any of the software components or functions described in this application may be implemented as software code executed by a processor using any suitable computer language, such as Java, C++, or Perl, using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission, suitable media including random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive or floppy disk, or optical media such as a compact disk (CD) or DVD (digital versatile disk), flash memory, etc. The computer-readable medium may also be any combination of such storage or transmission devices.

[0126] Such a program may also be encoded and transmitted using a carrier signal adapted for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Thus, a computer-readable medium according to an embodiment of the present invention may be created using a data signal encoded with such a program. A computer-readable medium encoded with program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable medium may reside on or within a single computer program product (e.g., a hard drive, CD, or entire computer system), or may reside on or within different computer program products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results mentioned herein.

[0127] Any of the methods described herein can be performed in whole or in part on a computer system including one or more processors, which can be configured to perform the steps. Accordingly, embodiments can be directed to a computer system configured to perform the steps of any of the methods described herein, potentially with various components performing each step or group of steps. While presented as numbered steps, steps of the methods herein can be performed simultaneously or in a different order. Furthermore, some of these steps may be used with some of other steps from other methods. Also, all or some of the steps may be optional. Furthermore, any of the steps of any of the methods can be performed using modules, circuits, or other means for performing these steps. [Example]

[0128] The present invention will now be illustrated, but not limited, by reference to specific embodiments described in the following examples.

[0129] Example 1 Test subjects The first case was an autosomal dominant disorder. The wife had a family history of autosomal dominant disorders and suffered from hereditary osteogenesis imperfecta (OI), a condition characterized by easy fractures without obvious cause or minimal trauma. The couple already had an affected son who also had OI. The wife's genetic diagnosis showed a deletion c.299-14_c.302del in COL1A1 (collagen, type I, alpha 1), already known to cause the disease. The couple underwent PGT-M treatment. Eighteen metaphase II oocytes were retrieved and fertilized by intracytoplasmic sperm injection (ICSI). Twelve embryos developed to the blastocyst stage, and several trophectoderm cells were biopsied from each embryo for PGT-M.

[0130] The second case was an autosomal recessive disorder. Both parents had a recessive hearing loss disorder. Genetic testing for the wife revealed an insertion c.99_100insT in the USH2A gene. Genetic testing for the husband revealed a duplication c.991dupA and a mutation c.5699G>T in the USH2A gene. The couple already had a deaf daughter, whose genetic testing revealed an insertion c.99_100insT, a duplication c.991dupA, and a mutation c.5699G>T in the USH2A gene inherited from both parents. These mutations cause hearing loss. The couple underwent PGT-M treatment. Nine metaphase II oocytes were retrieved and fertilized by ICSI. Four embryos developed to the blastocyst stage, and several trophectoderm cells were biopsied from each embryo for PGT-M.

[0131] Blastocyst biopsy and PGT-M testing. Standard protocols were used for ICSI. Embryos were individually cultured in 15 μL microdroplets of equilibrated Vitrolife G series medium containing human serum albumin (HSA) (LifeGlobal) overlaid with mineral oil in a Miri incubator (ESCO) at 37°C in a dry atmosphere of 6-7% CO balanced with 5% O and N . Embryos were evaluated on day 5 or 6 for trophectoderm (TE) biopsy. A small number of TE cells from each hatched blastocyst were biopsied and sent to an external PGT laboratory for PGT-M testing.

[0132] Example 2 Sample collection Collection of culture medium samples and discarded embryos. Immediately after transferring the blastocysts to another dish for biopsy, 13 μL of spent culture medium was removed from each remaining spent medium droplet. Discarded embryos and their corresponding spent culture medium samples were also collected for analysis. Each discarded embryo was gently moved with a pipette tip to the edge of its microdroplet, then removed and transferred to an RNase-DNase-free PCR tube containing 3 μL of lysis buffer. Pipette tips were changed between each sample collection to avoid cross-contamination. A negative control was a medium droplet incubated and collected under the same conditions as those used for blastocyst culture, but without embryos.

[0133] All samples were frozen immediately after collection and stored at −80°C until analysis.

[0134] Example 3 Lysis Protocol Spent culture medium lysis. A 13 μL sample of frozen spent culture medium was thawed, gently mixed, and 10.8 μL was transferred to a PCR tube (Maximum Recovery, Axygen) containing 1.2 μL of 10× lysis buffer (1× lysis buffer: 60 mM Tris-Ac pH 8.3, 2 mM EDTA pH 8.0, 15 mM DTT, 0.5 μM carrier ssDNA (5′-TCAGGTTTTCCTGAA-3′, PAGE-purified Thermo Fisher Scientific oligo)) and spun down. This was heated to 75°C for 30 minutes and then incubated at 75°C for 30 minutes. 0.5 μL of 35 mg / mL QIAGEN protease (dissolved in water and stored at 4°C) was added, spun down, and incubated at 55°C for 2 hours, followed by 80°C for 30 minutes. The resulting lysate in the PCR tube can be immediately subjected to the modified linear transposon-based amplification described herein or stored in the freezer for later use.

[0135] Lysis of discarded embryos. Frozen discarded embryos were thawed and heated at 75°C for 30 minutes. Discarded embryo samples were then lysed by adding 0.5 μL of 5 mg / mL Qiagen protease and heating in 3 μL of lysis buffer (55°C for 2 hours and 80°C for 30 minutes).

[0136] Example 4 Whole genome amplification Preparation of transposomes. Transposon DNA (5' / Phos / CTGTCTCTTATACACATCTGAACAGAATTTAATACGACTCACTATAGGGAGATGTGTATAAGAGACAG-3', PAGE-purified Thermo Fisher Scientific oligo) was annealed into a 1.5 μM self-looping structure by slow cooling in annealing buffer (20 mM Tris-Ac pH 8.3, 50 mM NaCl, 2 mM EDTA pH 8.0). The 1.5 μM annealed transposon DNA was then mixed with an equal volume of approximately 1 μM Tn5 transposase (Lucigen, EZ-Tn5™ transposase) and incubated at room temperature for 30 minutes to dimerize into transposomes at a final concentration of approximately 0.25 μM. Transposomes can be stored at -20°C for extended periods, and approximately 0.5 µL of transposomes is required for each single-cell transposon-based amplification.

[0137] Whole genome amplification using a modified LIANTI (Linear Amplification by Transposon Insertion) method. Starting with culture media lysate, 20 μL of transposition mixture was assembled in a buffer containing 2.5 mM MgCl2 and 18.75 nM transposomes. The transposition reaction was carried out at 55°C for 12 min. Then, 0.84 μL of a mixture of EDTA and ssDNA was added to each tube and incubated at 68°C for 30 min. After removing the transposase reaction, 0.2 μL of Q5 HF DNA polymerase (New England Biolabs) was added and heated to 73°C for 45 s in the presence of 2 mM MgCl2 and 200 μM dNTPs for end-filling and extension of the fragments. 0.5 μL of 4 mg / mL Qiagen protease was added, and the PCR tube was heated to 50°C for 1 hour in the presence of 4 mM EDTA to inactivate the DNA polymerase, followed by protease heat inactivation at 77°C for 20 minutes in the presence of 450 mM NaCl in a total volume of approximately 25 μL. DNA fragments were amplified to RNA overnight at 37°C in 90 μL of T7 in vitro transcription reaction mixture, as described in U.S. Patent No. 10,894,980 and C Chen, D Xing, L Tan, H Li, G Zhou, L Huang, X S Xie Science 2017;356(6334):189-194.

[0138] The next day, 10 μL of 0.5 M EDTA was added to each tube, and the RNA was column-purified (Zymo Research). 18 μL of RNA was transferred to a PCR tube containing 3.1 μL of a mixture of 6.5 mM each dNTP, 9.7 μM carrier ssDNA (5'-TCAGGTTTTCCTGAA-3', PAGE-purified Thermo Fisher Scientific oligo), and 3.2 U / μL SUPERase In RNase inhibitor (Invitrogen). Denaturation incubation was performed at 70°C for 1 minute and 90°C for 15 seconds, followed by cooling on ice. The first run of reverse transcription was performed in a 30 μL SuperScript IV reverse transcription system containing 0.67 mM of each dNTP, 0.6 U / μL SUPERase In RNase inhibitor (Invitrogen), and 6 U / μL SuperScript IV Reserve Transcriptase (Invitrogen) in SuperScript IV buffer, using the primer 5'-AGATGTGTATAAGAGACAG-3' and PAGE-purified Thermo Fisher Scientific oligos. The incubation program was 25°C for 1 minute, 37°C for 1 minute, 42°C for 1 minute, 50°C for 1 minute, 55°C for 15 minutes, 60°C for 10 minutes, 65°C for 12 minutes, 70°C for 8 minutes, 75°C for 5 minutes, and 80°C for 10 minutes. RNA was then removed by incubation with 10 ng / μL affinity-purified RNase A (Invitrogen) and 0.08 U / μL RNase H (New England Biolabs) for 30 min at 37°C. Second-strand synthesis was performed with primer 5'-NNNNNNNNGGGAGATGTGTATAAGAGACAG-3' and PAGE-purified Thermo Fisher Scientific oligos in 100 μL of the Q5 DNA polymerase system (New England Biolabs, 1× Q5 reaction buffer, 1× Q5 High GC enhancer, 200 μM dNTPs, 0.5 μM primers, 0.02 U / μL Q5 DNA polymerase).Each tube was heated to 98°C for 30 seconds, followed by 10 cycles of 98°C for 10 seconds, 58°C for 30 seconds, 60°C for 30 seconds, 65°C for 30 seconds, 70°C for 2.5 minutes, and then 72°C for 6 minutes for chain extension. The resulting amplicons were column purified in 23 μL of elution buffer and stored at -20°C.

[0139] The DNA concentration of the amplified product was measured using a Qubit 2.0 fluorometer (ThermoFisher Scientific) equipped with a Qubit dsDNA HS Assay Kit (Life Technologies). In the first case, 17 spent culture media were successfully amplified. In the second case, 9 spent culture media were amplified using this method. Eight culture media were successfully amplified, and one failed. The amplification success rate for spent culture media was 96.15%.

[0140] DNA amplification of discarded embryos. DNA was amplified according to the LIANTI amplification method described in U.S. Patent No. 10,894,980 and C Chen, D Xing, L Tan, H Li, G Zhou, L Huang, and XS Xie. Science 2017;356(6334):189-194, except that DNA primer 19_1 5'-AGATGTGTATAAGAGACAG-3' and PAGE-purified Thermo Fisher Scientific oligos were added to the first-strand cDNA synthesis reaction. All four discarded embryos from the first case and five discarded embryos from the second case were successfully amplified.

[0141] Example 5 Comparison of variables in lysis and amplification methods As described in the LIANTI paper (Chen 2017) and a paper comparing single-cell whole genome amplification methods (Huang 2015), a diploid human cell line, BJ primary human foreskin fibroblasts (ATCC CRL-2522), was used. The BJ cell line served as a reference standard for comparing different lysis and amplification methods. To compare the performance of various amplification methods on spent culture medium, the BJ cell line was cultured in Eagle's Minimum Essential Medium (ATCC) supplemented with 10% fetal bovine serum, 100 IU / mL penicillin, and 100 μg / mL streptomycin in a 5% CO2 incubator at 37°C. After 48 hours, the culture medium from the BJ cell line was harvested. The BJ cells were then trypsinized, washed with 1x PBS, and resuspended in 1x PBS. A single BJ cell was then pipetted into a PCR tube containing 3 μL of lysis buffer.

[0142] Using the LIANTI procedure described in U.S. Patent No. 10,894,980 and C Chen, D Xing, L Tan, H Li, G Zhou, L Huang, XS Xie. Science 2017;356(6334):189-194, three single-cell and three culture medium samples were amplified using QIAGEN protease at 0.5 mg / ml (working concentration) and transposome at 6.25 nM (working concentration). DNA was not successfully amplified.

[0143] The culture medium DNA was then amplified using the amplification method described in Examples 3-4, except that one variable was adjusted each time. Specifically, the variables included the working concentration of protease, working concentration of transposome, PCR cycles, and primer sequence. Each variable was repeated three times.

[0144] The comparative results are shown in Table 1 and Figure 3. Using 0.5 mg / ml (working concentration) of QIAGEN protease or 6.25 nM (working concentration) of transposome in combination with DNA primer 19_1, significant amounts of DNA could be successfully amplified. Increasing the working concentration of QIAGEN protease or transposome resulted in a corresponding increase in the amount of amplified DNA.

[0145] Similar to the LIANTI procedure described in U.S. Patent No. 10,894,980 and C Chen, D Xing, L Tan, H Li, G Zhou, L Huang, and XS Xie. Science 2017;356(6334):189-194, DNA was not successfully amplified without DNA primers. However, adding primers with specific sequences, 19_1, 19_2, and 15_1, or random primer 3+14N, successfully amplified a significant amount of DNA. The amplification results with primer 25_1 were not as good as those with primers 19_1 or 19_2, likely due to increased mismatch between primer 25_1 and the RNA strand at the Tn5 transposition junction. Preferably, the primer length is 15-20 bp. Compared to 3+14N, adding random primer 5+14N resulted in poorer DNA amplification, demonstrating the importance of the 3'-terminal sequence of the DNA primer. [Table 1]

[0146] Example 6 Library preparation for sequencing Once the libraries were prepared, the amplicons were subjected to a conventional sonication process (Covaris S2) to the length required for the sequencing platform. The final insert size was 2 × 150 bp, suitable for paired-end Illumina sequencing. Library preparation was performed using the NEBNext Ultra II DNA Library Prep Kit for Illumina (New England Biolabs) according to the manufacturer's instructions, omitting the optional size selection step. For blood sample gDNA, all samples were subjected to a conventional sonication process (Covaris S2). Library preparation was performed using a PCR-free Library Prep Kit, NEBNext Ultra II DNA Library Prep Kit for Illumina (NEB#E7645S / L).

[0147] Illumina sequencing. Bulk samples, culture medium samples, and discarded embryo samples were sequenced using 2 × 150 bp paired-end sequencing on an Illumina NovaSeq 6000 System or an Illumina HiSeq 4000 platform. The sequencing lane was shared by 24 samples with NEBNext Indexes. Each gDNA sample was sequenced to generate approximately 90 Gb of raw data, and each culture medium sample or discarded embryo was sequenced to generate approximately 30 Gb of raw data.

[0148] Example 7 Mapping and SNP calling The genome coverage and allele dropout rate for each sample were calculated. The LIANTI process, described in U.S. Patent No. 10,894,980 and C Chen, D Xing, L Tan, H Li, G Zhou, L Huang, and X X Xie. Science 2017;356(6334):189-194, provides the highest genome coverage compared to other single-cell whole genome amplification methods (Chen et al., 2017). However, the genome coverage of DNA in spent culture medium showed very low genome coverage and high allele dropout rates. In the first case, as shown in Figure 4, the genome coverage of 17 spent culture medium samples ranged from 16.38% to 70.34%. In the second case, as shown in Figure 5, the genome coverage of nine successfully amplified spent culture medium samples ranged from 4.79% to 34.97%. Their ADO rates were much higher than 5%. Genome coverage is far from reaching standards that would allow linkage analysis with several loci (ESHRE 2020).

[0149] Example 8 Bayesian linkage and disease prevalence analysis The dry lab method described above was used for linkage analysis.

[0150] In the first case, heterozygous SNPs on chromosome 17 in the wife and homozygous SNPs in the husband were analyzed. Figure 6 shows the linkage analysis plots for the four SCMs in Case 1. Among these four SCMs, 1-SCM-02 had the largest and tightest number of SNP loci, resulting in a "high confidence" level of confidence in determining the absence of the maternal disease-causing allele in the embryo. Although large fragments of the haplotype were missing in 1-SCM-05 and 1-SCM-07, the detected SNPs were calculated using a Bayesian model, supporting the "high confidence" level of disease-carrying analysis in determining the absence of the disease-causing allele in the embryo. In contrast, 1-SCM-06 detected fewer SNP loci than the other three SCMs, resulting in a "medium confidence" level of confidence in determining the absence of the inherited disease-causing allele in the disease-carrying analysis.

[0151] In the second case, we sequenced the genomes of blood samples from the wife, husband, and affected daughter to perform linkage analysis for each embryo. All SNPs from the three individuals were called at a sequencing depth of >10x within 2 Mb of the USH2A gene. The heterozygous SNP on chromosome 1 in the wife and the homozygous SNP in the husband for the mutation site (c.99_100insT) carried by the wife were included in the analysis, as were the heterozygous SNP on chromosome 1 in the husband and the homozygous SNP in the wife for the mutation site (c.991dupA and c.5699G>T) carried by the husband.

[0152] The results for all spent culture media in cases 1 and 2 are summarized in Table 2. [Table 2]

[0153] In the first case, the culture medium from discarded embryo 1-SCM-11 was excluded due to the lack of a reliable standard for that sample, as the corresponding discarded embryo was not recovered. Of the remaining 16 culture media, two were excluded due to abnormal maternal cell contaminant (MCC) rates. Both MCC rates were outside the reference range before calibration (the rate for 1-SCM-12 was -0.53, and the rate for 1-SCM-15 was +0.663) and were therefore excluded from subsequent disease-prevalence analysis. The only culture medium considered "indeterminate" in Example 1 was 1-SCM-10, which is likely due to the divergent signatures of the steady log-likelihood ratios on either side of the disease-causing mutation site. The log-likelihood ratios converged around +10.1 in the upstream region and around -11.3 in the downstream region. Despite this, disease-prevalence analysis was performed on 13 samples, classifying eight as high-confidence, two as medium-confidence, and three as possible. These results were consistent with the biopsy findings.

[0154] In the second case, all 14 samples (7 culture media from both parental lines) were counted in the final result. 2-SCM-03 was rejected during the quality control stage due to its abnormal MCC rate (-0.501) after calibration, precluding further analysis of both disease-causing mutation sites. Additionally, 2-SCM-04 (maternal) and 2-SCM-08 (maternal) were labeled "indeterminate," indicating uncertain results regarding whether the maternally inherited gene was the disease-causing gene. Of the remaining 10 samples (6 paternal and 4 maternal), 3 were classified as "high confidence," 2 as "moderate confidence," and 5 as "possible." Again, these results were consistent with the biopsy results.

[0155] As an example, the culture medium 1-SCM-01 from the first case was analyzed. Eligible SNPs were selected within a 10-Mbp region upstream and downstream of the disease-causing mutation site, allowing for estimation of basic statistics such as the sequencing error rate and maternal cell contamination (MCC) rate. This involved removing all SNPs with a parental allele depth of less than 5 or a QUAL score of less than 30. All SNPs with three or more alleles were removed. In this region, 11,913 of the 74,935 SNPs had a genotype quality (GQ) score of 3 or less. All SNPs with GQ scores of 6 or less, which correspond to the lowest 20% of all SNP GP scores, were excluded. Among the remaining SNPs, 4,911 SNPs were identified for which both parents showed the same homozygous genotype (i.e., both were 00 or both were 11). Within this subset, 220 SNPs showed at least one false read. These 4911 SNPs generated a total of 45499 reads, including 629 erroneous reads, yielding an estimated sequencing error rate of 0.0138.

[0156] After reviewing all eligible SNPs, 484 SNPs were identified where the parents exhibited different but identical genotypes (i.e., either the paternal genotype was 00 and the maternal genotype was 11, or vice versa). The total number of reads at these SNPs was 4779, with 3097 coming from the mother and the remainder coming from the father. This data yielded an MCC rate of 0.296, calculated as (3097*2-4779) / 4779. After calibration, 332 of these 484 SNPs were classified as either "containing only paternal chromosomes" or "containing only maternal chromosomes." The remaining 152 SNPs, deemed "containing both parental chromosomes," were used to calculate the estimated MCC rate after calibration. These SNPs generated a total of 1761 reads, with 955 coming from the mother and the remainder coming from the father, yielding an estimated MCC rate of 0.0846. Using the calibrated MCC rate, log-likelihood values ​​were calculated relative to the disease-causing mutation site. The log-likelihood ratio for any segment of the SNP, whether upstream or downstream, was negative. From the 21st SNP upstream (48,003,174 base pairs), the log-likelihood ratio stabilized around -15.1. Meanwhile, in the downstream region, from the 7th SNP (48,321,621 base pairs), the log-likelihood ratio stabilized around -11.2. Based on this analysis, it was estimated that this sample did not inherit the disease-carrying gene from its mother, and this result was labeled "high confidence." Comparison of the SCM allele classification results with those of TE biopsies or discarded embryos yielded a 100% concordance rate. The accuracy of the method described herein is sufficiently high for PGT-M.

[0157] Example 9 Prevalence analysis for PGT-P To demonstrate the effectiveness of the method described herein in determining susceptibility to polygenic diseases, a polygenic risk score (PRS) for type 2 diabetes was calculated for the second case. The PRS is derived from the combination of multiple genetic variants, specifically single nucleotide polymorphisms (SNPs), directly associated with the disease. Of the 139 common SNPs associated with type 2 diabetes mentioned in Xue et al. (2018), 57 SNPs were identified as present in the family's parental samples. For each SNP, a Bayesian model was used separately to determine the genotype of each disease-associated SNP. Possible scores for an SNP are 0, 1, or 2, representing a homozygous unaffected allele, a heterozygous allele, and a homozygous affected allele, respectively. The PRS is defined as the weighted sum of these scores, with weights equal to the absolute value of the log odds ratio for the minor allele at each SNP. We examined three discarded embryos from the second case along with their corresponding culture media. The results, detailed in Table 3 and Figure 7, show alignment and correlation between the PRS scores of discarded embryos and the PRS scores derived from the culture medium. [Table 3]

[0158] Example 10 Noninvasive pre-implantation genetic testing for aneuploidy In addition to niPGT-M and niPGT-P, our sequencing data will be used for non-invasive preimplantation genetic testing for aneuploidy (niPGT-A). High-quality reads were aligned to the human reference genome hg19, followed by copy number variation (CNV) detection for each sample. Mapped reads were subjected to GC bias correction resulting from the GC content of DNA. The read counts for each bin of 1000 kb were defined as the copy number (CN). CNV analysis was performed at 1 Mb resolution using HMMcopy (v1.28.0) and DNAcopy (v1.58.0). Segmentation was performed by HMMcopy with an e-value of 0.9999 and DNAcopy with an alpha value of 0.0001. CNVs were identified by using results from both HMMcopy and DNA to identify aneuploid segments larger than 10 Mb, while retaining double-positive counts as true events. Cytoband files for these aneuploid segments were calculated using the cytoband file for hg19 from the University of California, Santa Cruz database (available at the world wide website hgdownload.cse.ucsc.edu / downloads.html).

[0159] In the first case, the CNV plots of spent culture medium were compared with those of TE biopsies. The CNV results of spent culture medium were abnormal if the corresponding TE biopsy results were abnormal. For example, in Figure 8(a), the niPGT-A results of 1-SCM-02 showed monosomy 10 aneuploidy, which was consistent with the corresponding TE results. In Figure 8(b), the niPGT-A results of 1-SCM-04 showed a chromosomal deletion of the q arm of chromosome 9, which was also consistent with the TE biopsy results. In Figure 8(c), the niPGT-A results of 1-SCM-14 were normal, and the corresponding TE biopsy results were also normal. These results demonstrate that the method described herein can analyze the copy number of spent culture medium.

[0160] The features disclosed in the foregoing description or in the following claims may be utilized to realize the invention in various of its forms, either in their specific form, or in connection with means for performing a disclosed function, or methods or processes for achieving a disclosed result, either separately or in any combination of such features, as appropriate.

[0161] The foregoing invention has been described in some detail by way of illustration and example, for purposes of clarity and understanding. It will be apparent to those skilled in the art that changes and modifications may be practiced within the scope of the appended claims. It is therefore to be understood that the foregoing description is intended to be illustrative, and not limiting. The scope of the invention should, therefore, be determined not with reference to the above description, but should instead be determined with reference to the following appended claims, along with the full scope of equivalents to which such claims are entitled.

[0162] The patents, published applications, and scientific literature referred to herein establish the knowledge of those skilled in the art and are hereby incorporated by reference in their entirety to the same extent as if each were specifically and individually indicated.

[0163] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the invention, however, other embodiments of the invention may be directed to specific embodiments of each individual aspect or particular combinations of these individual aspects.

[0164] The sequences used in this invention are summarized in Table 4. [Table 4]

[0165] Embodiment An embodiment of the present disclosure relates to a method for amplifying DNA in a solution, wherein the DNA content in the solution can be at minimum picogram levels, comprising the steps of: combining a lysis buffer with the solution, heating the solution, then combining a protease with the solution, incubating the solution, followed by inactivating the protease under elevated temperature to obtain a lysate; combining a transposome with the lysate, the transposome comprising a transposase and a transposon comprising a transposase binding site and an RNA polymerase promoter sequence, wherein the transposome binds to the DNA in the solution, the transposase cleaves the DNA in the solution to generate DNA fragments, and the transposon is integrated or inserted into the DNA fragments; filling 9 bp gaps resulting from transposon insertion and extending to both ends of each fragment; performing in vitro transcription using the original sequence as a template to form RNA; combining the RNA with DNA primers complementary to each RNA strand at the transposition junction sequence and a reverse transcription system to form first-strand cDNA via reverse transcription; and forming second-strand DNA via cycling amplification. In one embodiment, the solution is culture medium or blastocoelic fluid. In one embodiment, the DNA primer has a length of 15 to 20 bp. In one embodiment, the 3' end of the DNA primer comprises the sequence 5'-GACAG-3' or 5'-CTGTC-3', preferably the DNA primer comprises the sequence set forth as 5'-GATGTGTATAAGAGACAG-3' (SEQ ID NO: 1) or 5'-TCTACACATATTCTCTGTC-3' (SEQ ID NO: 2), or is at least 70%, preferably 80%, preferably 90%, preferably 95%, preferably 100% identical to SEQ ID NO: 1 or SEQ ID NO: 2. In one embodiment, the lysis buffer is a 2x to 10x lysis buffer. In one embodiment, the working concentration of the protease in the lysis mixture is greater than 0.6 mg / mL, preferably in the range of 0.8 to 2 mg / mL. In one embodiment, the incubation is performed at 45 to 60°C for 1 to 3 hours. According to one embodiment, the inactivation is carried out at 80-90° C. According to one embodiment, the working concentration of the transposome is higher than 6 nM, preferably in the range of 10-40 nM.In one embodiment, the cycling amplification is carried out for more than 4 cycles, for example, 4 to 20 cycles, preferably 6 to 15 cycles.

[0166] The present disclosure provides a method for amplifying DNA in a culture medium, the DNA content of which can be as low as picogram levels, comprising the steps of: combining a 2x to 10x lysis buffer with the culture medium, heating the culture medium, combining a protease with the culture medium to a working concentration of 0.8 to 2 mg / mL, and incubating the culture medium at 45 to 60°C for 1 to 3 hours, followed by inactivating the protease at 80 to 90°C to obtain a lysate; combining a transposome containing a transposase and a transposon comprising a transposase binding site and an RNA polymerase promoter sequence with the lysate to a working concentration of 10 to 40 nM, wherein the transposome binds to the DNA in the solution, the transposase cleaves the DNA in the solution to generate DNA fragments, and the transposon is integrated into or binds to the DNA fragments. or inserted; filling the 9 bp gap resulting from transposon insertion and extending to both ends of each fragment; performing in vitro transcription using the original sequence as a template to form RNA; combining the RNA with a DNA primer that pairs with the nucleic acid sequence conjugated to the transposome and a reverse transcription system to form first strand cDNA via reverse transcription, wherein the DNA primer comprises the sequence shown as 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO: 1) or 5'-TCTACACATATTCTCTGTC-3' (SEQ ID NO: 2), or is at least 70%, preferably 80%, preferably 90%, preferably 95%, preferably 100% identical to SEQ ID NO: 1 or SEQ ID NO: 2; and forming second strand DNA via 6 to 15 cycles of amplification.

[0167] The present disclosure provides a method for determining susceptibility to a genetic disease in an embryo by analyzing a biological sample of culture medium from an in vitro cultured embryo, the method comprising: 1) amplifying DNA molecules from the biological sample of culture medium from the in vitro cultured embryo and / or a biological sample from a family member's discarded embryo; 2) preparing a DNA library and sequencing the DNA molecules from the biological sample to obtain multiple sequence reads; 3) performing mapping and SNP calling by a computer system to generate SNP data and control its quality; and 4) performing haplotype prephasing by a computer system on blood samples from both parents or one parent carrying the disease. The method includes: 1) performing a maternal cell contamination (MCC) mapping; 2) calculating the likelihood of four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of each parent, taking into account the maternal cell contamination (MCC) rate and haplotype status; 3) estimating the MCC rate and determining the haplotype status of each SNP to indicate whether DNA from only one parent is present in the culture medium; 4) calculating the likelihood of observing SNP data in the region flanking the disease-causing mutation site and combining the likelihood at each single SNP with the recombination probability; and 5) determining whether the embryo carries the disease-causing chromosome and providing a confidence level. According to one embodiment, the embryo has maternal cell contamination and / or haplotype loss. According to one embodiment, in step 3), the mapping and SNP calling outputs the physical location, allele type, allele depth, and sequencing quality for each SNP in a biological sample of the culture medium of the in vitro-cultured embryo and / or a biological sample of a discarded embryo within the family. According to one embodiment, in step 3) quality control of the SNP data is performed by filtering out SNPs with low Genetic Quality (GQ) values ​​from biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos within the family. According to one embodiment, this comprises filtering out SNPs with the lowest 20% GQ values ​​from biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos within the family.According to one embodiment, in step 3), the sequencing error rate and relative density of SNPs in the region flanking the disease-causing mutation site are estimated, and culture medium samples with extremely low SNP density or high sequencing error rate are excluded. According to one embodiment, the excluded SNP density is less than 0.001, or the excluded sequencing error rate is greater than 0.2. According to one embodiment, in step 4), one or more of the following samples within the family are used: a blood sample from the proband, a biological sample from a discarded embryo, or a blood sample from at least one grandparent. According to one embodiment, in step 5), the estimated sequencing error rate is integrated when calculating the likelihood of each individual SNP. According to one embodiment, step 6) further comprises the following steps: a): initially estimating the MCC rate and haplotype state of each SNP in the region flanking the disease-causing mutation site; and b) recursively updating the MCC rate and haplotype state of each SNP. According to one embodiment, step b) comprises using a label search method. According to one embodiment, after step a), samples with abnormal MCC rates are excluded. According to one embodiment, the abnormal MCC rate is greater than 0.65 or less than -0.5. According to one embodiment, in step 7), a Bayesian model and a recursive algorithm are used. According to one embodiment, in step 8), log-likelihood curves are first plotted for both the upstream and downstream regions of the disease-causing mutation site. According to one embodiment, the genetic disease includes both monogenic and polygenic diseases. According to one embodiment, after step 8), multiple identified SNPs of the embryo are combined to determine its susceptibility to polygenic diseases.

[0168] The present disclosure provides a method for determining a susceptibility of an embryo to a genetic disease by analyzing a biological sample of culture medium of an in vitro cultured embryo, the method comprising: 1) amplifying DNA molecules from the biological sample of culture medium of an in vitro cultured embryo and / or amplifying DNA molecules from a biological sample of a discarded embryo in a family using the method of any one of claims 1 to 12; 2) preparing a DNA library and sequencing the DNA molecules from the biological sample to obtain multiple sequence reads; 3) performing mapping and SNP calling by a computer system to generate SNP data and control its quality; and 4) performing computer-generated SNP analysis on blood samples of both parents or one parent carrying the disease. The method includes: 1) performing haplotype prephasing using a data system; 2) calculating the likelihood of four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of each parent, taking into account the maternal cell contamination (MCC) rate and haplotype status; 3) estimating the MCC rate and determining the haplotype status of each SNP to indicate whether DNA from only one parent is present in the culture medium; 4) calculating the likelihood of observing SNP data in the region flanking the disease-causing mutation site and combining the likelihood at each single SNP with the recombination probability; and 5) determining whether the embryo possesses the disease-causing chromosome and providing a confidence level. According to one embodiment, the embryo has maternal cell contamination and / or haplotype loss. According to one embodiment, in step 3), the mapping and SNP calling outputs the physical location, allele type, allele depth, and sequencing quality for each SNP in a biological sample of the culture medium of the in vitro-cultured embryo and / or a biological sample of a discarded embryo within the family. According to one embodiment, in step 3) quality control of the SNP data is performed by filtering out SNPs with low Genetic Quality (GQ) values ​​from biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos within the family.According to one embodiment, the method includes filtering out SNPs with the lowest 20% GQ values ​​from biological samples of culture medium from in vitro-cultured embryos and / or biological samples from discarded embryos within the family. According to one embodiment, in step 3), the sequencing error rate and relative density of SNPs in the region adjacent to the disease-causing mutation site are estimated, and culture medium samples with extremely low SNP density or high sequencing error rate are excluded. According to one embodiment, the SNP density to be excluded is less than 0.001, or the sequencing error rate to be excluded is greater than 0.2. According to one embodiment, in step 4), one or more of the following samples within the family are used: a blood sample from the proband, a biological sample from a discarded embryo, or a blood sample from at least one grandparent. According to one embodiment, in step 5), the estimated sequencing error rates are integrated when calculating the likelihood of each individual SNP. According to one embodiment, step 6) further comprises the following steps: a) initially estimating the MCC rate and haplotype state of each SNP in the region adjacent to the disease-causing mutation site; and b) recursively updating the MCC rate and haplotype state of each SNP. According to one embodiment, step b) uses a label search method. According to one embodiment, after step a), samples with abnormal MCC rates are excluded. According to one embodiment, the abnormal MCC rate is greater than 0.65 or less than -0.5. According to one embodiment, step 7) uses a Bayesian model and a recursive algorithm. According to one embodiment, step 8) first plots log-likelihood curves for both the upstream and downstream regions of the disease-causing mutation site. According to one embodiment, the genetic disease includes both monogenic and polygenic diseases. According to one embodiment, step 8) comprises combining multiple identified SNPs in an embryo to determine its susceptibility to a polygenic disease.

[0169] The present disclosure provides a computer program comprising a plurality of instructions executable by a computer system, the computer program being adapted, when executed, to control the computer system to analyze a biological sample of culture medium of an in vitro cultured embryo to determine susceptibility to a genetic disease in the embryo, the computer program comprising: 1) amplifying DNA molecules from the biological sample of culture medium of the in vitro cultured embryo, and / or amplifying DNA molecules from a biological sample of a discarded embryo in a family; 2) preparing a DNA library and sequencing the DNA molecules from the biological sample to obtain a plurality of sequence reads; 3) performing mapping and SNP calling by the computer system to generate SNP data and control its quality; and 4) analyzing blood samples of both parents or one of the parents carrying the disease. and performing haplotype prephasing using a computer system; 5) calculating the likelihood of four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of each parent, taking into account the maternal cell contamination (MCC) rate and haplotype state; 6) estimating the MCC rate and determining the haplotype state of each SNP to indicate whether DNA from only one parent is present in the culture medium; 7) calculating the likelihood of observing SNP data within regions flanking the disease-causing mutation site and combining the likelihood at each single SNP with the recombination probability; and 8) determining whether the embryo carries the disease-causing chromosome and providing a confidence level. According to one embodiment, the embryo has maternal cell contamination and / or haplotype loss. According to one embodiment, in step 3), mapping and SNP calling outputs the physical location, allele type, allele depth and sequencing quality for each SNP in the biological sample of culture medium of in vitro cultured embryos and / or the biological sample of discarded embryos within the family. According to one embodiment, quality control of the SNP data in step 3) is performed by filtering out SNPs with low Genetic Quality (GQ) values ​​in the biological sample of culture medium of in vitro cultured embryos and / or the biological sample of discarded embryos within the family.According to one embodiment, the method includes filtering out SNPs with the lowest 20% GQ values ​​from biological samples of culture medium from in vitro-cultured embryos and / or biological samples from discarded embryos within the family. According to one embodiment, in step 3), the sequencing error rate and relative density of SNPs in the region adjacent to the disease-causing mutation site are estimated, and culture medium samples with extremely low SNP density or high sequencing error rate are excluded. According to one embodiment, the SNP density to be excluded is less than 0.001, or the sequencing error rate to be excluded is greater than 0.2. According to one embodiment, in step 4), one or more of the following samples within the family are used: a blood sample from the proband, a biological sample from a discarded embryo, or a blood sample from at least one grandparent. According to one embodiment, in step 5), the estimated sequencing error rates are integrated when calculating the likelihood of each individual SNP. According to one embodiment, step 6) further comprises the following steps: a) initially estimating the MCC rate and haplotype state of each SNP in the region adjacent to the disease-causing mutation site; and b) recursively updating the MCC rate and haplotype state of each SNP. According to one embodiment, step b) uses a label search method. According to one embodiment, after step a), samples with abnormal MCC rates are excluded. According to one embodiment, the abnormal MCC rate is greater than 0.65 or less than -0.5. According to one embodiment, in step 7), a Bayesian model and a recursive algorithm are used. According to one embodiment, in step 8), first, log-likelihood curves are plotted for both the upstream and downstream regions of the disease-causing mutation site. According to one embodiment, genetic diseases include both monogenic and polygenic diseases. According to one embodiment, after step 8), multiple identified SNPs of an embryo are combined to determine its susceptibility to a polygenic disease.

[0170] The present disclosure provides a computer system for determining a susceptibility of an embryo to a genetic disease by analyzing a biological sample of culture medium of an in vitro cultured embryo, the computer system comprising: 1) means for amplifying DNA molecules from the biological sample of culture medium of the in vitro cultured embryo and / or amplifying a biological sample of a discarded embryo within a family; 2) means for preparing a DNA library and sequencing the DNA molecules from the biological sample to obtain a plurality of sequence reads; 3) means for performing mapping and SNP calling by the computer system to generate SNP data and control the quality thereof; and 4) means for performing haplotyping by the computer system on blood samples of both parents or one parent carrying the disease. The method includes: 1) a means for performing genomic mapping; 2) a means for performing genomic mapping; 3) a means for calculating the likelihood of four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of each parent, taking into account the maternal cell contamination (MCC) rate and haplotype state; 4) a means for calculating the likelihood of four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of each parent, taking into account the maternal cell contamination (MCC) rate and haplotype state; 5) a means for calculating the likelihood of four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of each parent, taking into account the maternal cell contamination (MCC) rate and haplotype state; 6) a means for determining the haplotype state of each SNP to estimate the MCC rate and indicate whether DNA from only one parent is present in the culture medium; 7) a means for calculating the likelihood of observing SNP data in regions flanking the disease-causing mutation site and combining the likelihood at each single SNP with the recombination probability; and 8) a means for determining whether the embryo carries the disease-causing chromosome and providing a confidence level. According to one embodiment, the embryo has maternal cell contamination and / or haplotype loss. According to one embodiment, in step 3), the mapping and SNP calling outputs the physical location, allele type, allele depth, and sequencing quality for each SNP in a biological sample of the culture medium of the in vitro-cultured embryo and / or a biological sample of a discarded embryo within the family. According to one embodiment, in means 3) quality control of SNP data is performed by filtering out SNPs with low Genetic Quality (GQ) values ​​from biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos within the family.According to one embodiment, the method includes filtering out SNPs with the lowest 20% GQ values ​​from biological samples of culture medium from in vitro-cultured embryos and / or biological samples from discarded embryos within the family. According to one embodiment, in means 3), the sequencing error rate and relative density of SNPs in regions adjacent to the disease-causing mutation site are estimated, and culture medium samples with extremely low SNP density or high sequencing error rate are excluded. According to one embodiment, the SNP density to be excluded is less than 0.001, or the sequencing error rate to be excluded is greater than 0.2. According to one embodiment, in means 4), one or more of the following samples within the family are used: a blood sample from the proband, a biological sample from a discarded embryo, or a blood sample from at least one grandparent. According to one embodiment, in means 5), the estimated sequencing error rate is integrated when calculating the likelihood of each individual SNP. According to one embodiment, means 6) further comprises the following functions: a) initially estimating the MCC rate and haplotype state of each SNP in the region adjacent to the disease-causing mutation site; and b) recursively updating the MCC rate and haplotype state of each SNP. According to one embodiment, function b) uses a label search method. According to one embodiment, after function a), samples with abnormal MCC rates are excluded. According to one embodiment, the abnormal MCC rate is greater than 0.65 or less than -0.5. According to one embodiment, means 7) uses a Bayesian model and a recursive algorithm. According to one embodiment, means 8) first plots log-likelihood curves for both the upstream and downstream regions of the disease-causing mutation site. According to one embodiment, the genetic disease includes both monogenic and polygenic diseases. According to one embodiment, the method further comprises means for combining multiple identified SNPs in an embryo to determine its susceptibility to a polygenic disease.

[0171] The present disclosure provides a method for preparing a biological sample of culture medium from an in vitro cultured embryo, which would otherwise be unsequenceable, into a converted state that can be analyzed to determine susceptibility to a genetic disease in the embryo, the method comprising the steps of: 1) amplifying DNA molecules from the biological sample of culture medium from the in vitro cultured embryo and / or from a biological sample of a discarded embryo within a family; 2) preparing a DNA library and sequencing the DNA molecules from the biological sample to obtain multiple sequence reads; 3) performing mapping and SNP calling by a computer system to generate and quality control SNP data; 4) performing haplotype prephasing by a computer system on blood samples from both parents or one parent carrying the disease; and 5) determining whether the genetic strand is inherited from the paternal or maternal alleles of the parents. 5) calculating, by a computer system, the likelihood of four possible inheritance scenarios at each single SNP, representing whether or not DNA from only one parent is present in the culture medium, taking into account the maternal cell contaminant (MCC) rate and haplotype state; 6) transforming, by a computer system, the SNP data by estimating the MCC rate and determining the haplotype state of each SNP to indicate whether or not DNA from only one parent is present in the culture medium; 7) calculating, by a computer system, the likelihood of observing SNP data within regions flanking the disease-causing mutation site and combining the likelihood and recombination probability at each single SNP; and 8) determining, by a computer system, whether or not the embryo carries the disease-causing chromosome and providing a confidence level, that the biological sample has a high ADO rate in the presence of maternal contaminants.

[0172] References The following full cited references correspond to the respective abbreviated forms of the citations used above and are incorporated as if fully set forth herein: JPEG2026502019000040.jpg193161 JPEG2026502019000041.jpg224162 JPEG2026502019000042.jpg224162 JPEG2026502019000043.jpg224162 JPEG2026502019000044.jpg216161 JPEG2026502019000045.jpg216161 JPEG2026502019000046.jpg224162 JPEG2026502019000047.jpg216162 JPEG2026502019000048.jpg216161 JPEG2026502019000049.jpg208161 JPEG2026502019000050.jpg107161

Claims

1. 1. A method for amplifying DNA in a solution, wherein the DNA content in the solution can be at a minimum picogram level, comprising the steps of: combining a lysis buffer with the solution, heating the solution, then combining a protease with the solution, incubating the solution, followed by inactivating the protease under elevated temperature to obtain a lysate; a transposase, a transposon containing a transposase binding site and an RNA polymerase promoter sequence; combining a transposome comprising the transposon with a lysate, wherein the transposome binds to DNA in solution, the transposase cleaves the DNA in solution to generate DNA fragments, and the transposon integrates or inserts into the DNA fragments; filling the 9 bp gaps resulting from transposon insertions and extending to both ends of each fragment; performing in vitro transcription using the original sequence as a template to form RNA; combining the RNA with DNA primers complementary to each RNA strand at the translocation junction sequence and a reverse transcription system to form first strand cDNA via reverse transcription; and forming second strand DNA via cycling amplification;

2. 2. The method of claim 1, wherein the solution is culture medium or blastocoelic fluid.

3. The method of claim 1, wherein the DNA primer has a length of 15 bp to 20 bp.

4. The method of claim 2, wherein the 3' end of the DNA primer comprises the sequence 5'-GACAG-3 or 5'-CTGTC-3', preferably the DNA primer comprises the sequence shown as 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO: 1) or 5'-TCTACACATATTCTCTGTC-3' (SEQ ID NO: 2), or is at least 70%, preferably 80%, preferably 90%, preferably 95%, preferably 100% identical to SEQ ID NO: 1 or SEQ ID NO:

2.

5. 2. The method of claim 1, wherein the lysis buffer is a 2x to 10x lysis buffer.

6. 2. The method of claim 1, wherein the working concentration of the protease in the lysis mixture is greater than 0.6 mg / mL, preferably in the range of 0.8 to 2 mg / mL.

7. 2. The method of claim 1, wherein the incubation is carried out at 45 to 60°C for 1 to 3 hours.

8. 2. The method of claim 1, wherein the inactivation is carried out at 80 to 90°C.

9. The method of claim 1, wherein the working concentration of transposome is higher than 6 nM, preferably in the range of 10-40 nM.

10. The method of claim 1, wherein the cycling amplification is carried out for more than 4 cycles, for example 4 to 20 cycles, preferably 6 to 15 cycles.

11. 1. A method for amplifying DNA in a culture medium, wherein the DNA content in the culture medium can be at a minimum picogram level, comprising the steps of: combining 2x-10x lysis buffer with culture medium, heating the culture medium, combining protease with the culture medium to a working concentration of 0.8-2 mg / mL, incubating the culture medium at 45-60°C for 1-3 hours, followed by inactivating the protease at 80-90°C to obtain a lysate; a transposase, a transposon containing a transposase binding site and an RNA polymerase promoter sequence; combining transposomes comprising: filling the 9 bp gaps resulting from transposon insertions and extending to both ends of each fragment; performing in vitro transcription using the original sequence as a template to form RNA; combining the RNA with a DNA primer that pairs with the nucleic acid sequence conjugated to the transposome and a reverse transcription system to form first strand cDNA via reverse transcription, wherein the DNA primer comprises the sequence set forth as 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO:1) or 5'-TCTACACATATTCTCTGTC-3' (SEQ ID NO:2), or is at least 70%, preferably 80%, preferably 90%, preferably 95%, preferably 100% identical to SEQ ID NO:1 or SEQ ID NO:2; and c. Forming second strand DNA through 6 to 15 cycles of amplification.

12. 1. A method for determining a susceptibility of an embryo to a genetic disease by analyzing a biological sample of culture medium of an in vitro cultured embryo, the method comprising: 1) amplifying DNA molecules from biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos of the family; 2) preparing a DNA library and sequencing DNA molecules from the biological sample to obtain multiple sequence reads; 3) Performing mapping and SNP calling by a computer system to generate and quality control the SNP data; 4) performing haplotype prephasing on blood samples from both parents or one of the parents carrying the disease by a computer system; 5) Calculating the likelihood of the four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of both parents, taking into account the maternal cell contamination (MCC) rate and haplotype status; 6) Estimating the MCC rate and determining the haplotype status of each SNP to indicate whether DNA from only one parent is present in the culture medium; 7) Calculating the likelihood of observing SNP data within the region flanking the disease-causing mutation site and combining the likelihood and recombination probability at each single SNP; 8) To determine whether an embryo carries a disease-causing chromosome and provide a confidence level.

13. 13. The method of claim 12, wherein the embryo has maternal cell contamination and / or haplotype loss.

14. 13. The method of claim 12, wherein in step 3), the mapping and SNP calling outputs the physical location, allele type, allele depth and sequencing quality for each SNP in biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos within the family.

15. 13. The method of claim 12, wherein in step 3) quality control of the SNP data is performed by filtering out SNPs with low genetic quality (GQ) values ​​from biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos within the family.

16. 16. The method of claim 15, wherein SNPs with the lowest 20% GQ values ​​are filtered out from biological samples of culture medium from in vitro cultured embryos and / or biological samples from discarded embryos within the family.

17. 13. The method of claim 12, wherein in step 3) the sequencing error rate and relative density of SNPs in the regions flanking the disease-causing mutation site are estimated, and culture medium samples with extremely low SNP density or high sequencing error rate are excluded.

18. 18. The method of claim 17, wherein the excluded SNP density is less than 0.001 or the excluded sequencing error rate is greater than 0.

2.

19. The method of claim 12, wherein in step 4) one or more of the following samples within the family are used: a blood sample from the proband, a biological sample from a discarded embryo, or a blood sample from at least one grandparent.

20. 13. The method of claim 12, wherein in step 5) the estimated sequencing error rate is integrated when calculating the likelihood of each individual SNP.

21. Step 6) is the following step: a) first estimating the MCC rate and haplotype status of each SNP within the region flanking the disease-causing mutation site; and b) Recursively updating the MCC rate and haplotype state of each SNP 13. The method of claim 12, further comprising:

22. 22. The method of claim 21, wherein step b) uses a label-finding method.

23. 22. The method of claim 21, wherein after step a), samples with abnormal MCC rates are excluded.

24. 24. The method of claim 23, wherein the abnormal MCC rate is greater than 0.65 or less than −0.

5.

25. 13. The method of claim 12, wherein step 7) uses a Bayesian model and a recursive algorithm.

26. The method according to claim 12, wherein in step 8), log-likelihood curves are first plotted for both the upstream and downstream regions of the disease-causing mutation site.

27. 13. The method of claim 12, wherein the genetic disease includes both monogenic and polygenic diseases.

28. The method according to any one of claims 12 to 27, wherein after step 8), the identified SNPs of the embryo are combined to determine its susceptibility to a polygenic disease.

29. 1. A method for determining a susceptibility of an embryo to a genetic disease by analyzing a biological sample of culture medium of an in vitro cultured embryo, the method comprising: 1) Using the method according to any one of claims 1 to 12 to amplify DNA molecules from a biological sample of the culture medium of an in vitro cultured embryo and / or amplify DNA molecules from a biological sample of a discarded embryo within a family; 2) preparing a DNA library and sequencing DNA molecules from the biological sample to obtain multiple sequence reads; 3) Performing mapping and SNP calling by a computer system to generate and quality control the SNP data; 4) performing haplotype prephasing on blood samples from both parents or one of the parents carrying the disease by a computer system; 5) Calculating the likelihood of the four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of both parents, taking into account the maternal cell contamination (MCC) rate and haplotype status; 6) Estimating the MCC rate and determining the haplotype status of each SNP to indicate whether DNA from only one parent is present in the culture medium; 7) Calculating the likelihood of observing SNP data within the region flanking the disease-causing mutation site and combining the likelihood and recombination probability at each single SNP; 8) To determine whether an embryo carries a disease-causing chromosome and provide a confidence level.

30. 30. The method of claim 29, wherein the embryo has maternal cell contamination and / or haplotype loss.

31. 30. The method of claim 29, wherein in step 3), the mapping and SNP calling outputs the physical location, allele type, allele depth and sequencing quality for each SNP in biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos within the family.

32. 30. The method of claim 29, wherein in step 3) quality control of the SNP data is performed by filtering out SNPs with low genetic quality (GQ) values ​​from biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos within the family.

33. 33. The method of claim 32, wherein SNPs with the lowest 20% GQ values ​​are filtered out from biological samples of culture medium from in vitro cultured embryos and / or biological samples from discarded embryos within the family.

34. 30. The method of claim 29, wherein in step 3) the sequencing error rate and relative density of SNPs in regions adjacent to the disease-causing mutation site are estimated, and culture medium samples with extremely low SNP density or high sequencing error rate are excluded.

35. 35. The method of claim 34, wherein the excluded SNP density is less than 0.001 or the excluded sequencing error rate is greater than 0.

2.

36. 30. The method of claim 29, wherein step 4) uses one or more of the following samples within the family: a blood sample from the proband, a biological sample from a discarded embryo, or a blood sample from at least one grandparent.

37. 30. The method of claim 29, wherein in step 5), the estimated sequencing error rate is integrated when calculating the likelihood of each individual SNP.

38. Step 6) is the following step: a) first estimating the MCC rate and haplotype status of each SNP within the region flanking the disease-causing mutation site; and b) Recursively updating the MCC rate and haplotype state of each SNP 30. The method of claim 29, further comprising:

39. 39. The method of claim 38, wherein step b) uses a label-finding method.

40. 39. The method of claim 38, wherein after step a), samples with abnormal MCC rates are excluded.

41. 41. The method of claim 40, wherein the abnormal MCC rate is greater than 0.65 or less than −0.

5.

42. 30. The method of claim 29, wherein step 7) uses a Bayesian model and a recursive algorithm.

43. 30. The method of claim 29, wherein in step 8), log-likelihood curves are first plotted for both the upstream and downstream regions of the disease-causing mutation site.

44. 30. The method of claim 29, wherein the genetic disease includes both monogenic and polygenic diseases.

45. The method of any one of claims 29 to 44, wherein after step 8), the identified SNPs of the embryo are combined to determine its susceptibility to a polygenic disease.

46. 1. A computer program comprising a plurality of instructions executable by a computer system and adapted, when executed, to control the computer system to analyze a biological sample of culture medium of an in vitro cultured embryo to determine a susceptibility of the embryo to a genetic disease, the computer program comprising: 1) Using the method according to any one of claims 1 to 10 to amplify DNA molecules from a biological sample of the culture medium of an in vitro cultured embryo and / or amplify DNA molecules from a biological sample of a discarded embryo within a family; 2) preparing a DNA library and sequencing DNA molecules from the biological sample to obtain multiple sequence reads; 3) Performing mapping and SNP calling by a computer system to generate and quality control the SNP data; 4) performing haplotype prephasing on blood samples from both parents or one of the parents carrying the disease by a computer system; 5) Calculating the likelihood of the four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of both parents, taking into account the maternal cell contamination (MCC) rate and haplotype status; 6) Estimating the MCC rate and determining the haplotype status of each SNP to indicate whether DNA from only one parent is present in the culture medium; 7) Calculating the likelihood of observing SNP data within the region flanking the disease-causing mutation site and combining the likelihood and recombination probability at each single SNP; 8) To determine whether an embryo carries a disease-causing chromosome and provide a confidence level.

47. The program of claim 46, wherein the embryo has contaminants and / or haplotype losses originating from maternal cells.

48. 47. The program of claim 46, wherein in step 3), the mapping and SNP calling outputs the physical location, allele type, allele depth and sequencing quality for each SNP in a biological sample of culture medium of an in vitro cultured embryo and / or a biological sample of a discarded embryo within the family.

49. 47. The program of claim 46, wherein in step 3) quality control of SNP data is performed by filtering out SNPs with low genetic quality (GQ) values ​​from biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos within the family.

50. 50. The program of claim 49, wherein SNPs having the lowest 20% GQ values ​​are filtered out from biological samples of culture medium from in vitro cultured embryos and / or biological samples from discarded embryos within the family.

51. 47. The program of claim 46, wherein in step 3) the sequencing error rate and relative density of SNPs in regions adjacent to disease-causing mutation sites are estimated, and culture medium samples with extremely low SNP density or high sequencing error rate are excluded.

52. 52. The program of claim 51, wherein the excluded SNP density is less than 0.001 or the excluded sequencing error rate is greater than 0.

2.

53. 47. The program of claim 46, wherein step 4) uses one or more of the following samples within the family: a blood sample from the proband, a biological sample from a discarded embryo, or a blood sample from at least one grandparent.

54. 47. The program of claim 46, wherein in step 5), estimated sequencing error rates are integrated when calculating the likelihood of each individual SNP.

55. Step 6) is the following step: a) first estimating the MCC rate and haplotype status of each SNP within the region flanking the disease-causing mutation site; and b) Recursively updating the MCC rate and haplotype state of each SNP 47. The program of claim 46, further comprising:

56. 56. The program of claim 55, wherein step b) uses a sign-finding method.

57. 56. The program of claim 55, wherein after step a), samples with abnormal MCC rates are excluded.

58. 58. The program of claim 57, wherein the abnormal MCC rate is greater than 0.65 or less than −0.

5.

59. 47. The program of claim 46, wherein step 7) uses a Bayesian model and a recursive algorithm.

60. The program according to claim 46, wherein in step 8), log-likelihood curves are first plotted for both the upstream and downstream regions of the disease-causing mutation site.

61. 47. The program of claim 46, wherein the genetic disease includes both monogenic and polygenic diseases.

62. The program according to any one of claims 46 to 61, wherein after step 8), the identified SNPs of the embryo are combined to determine its susceptibility to a polygenic disease.

63. 1. A computer system for analyzing a biological sample of culture medium of an in vitro cultured embryo to determine a susceptibility of the embryo to a genetic disease, the computer system comprising: 1) means for amplifying DNA molecules from a biological sample of the culture medium of an in vitro cultured embryo and / or a biological sample of a discarded embryo within a family, using the method according to any one of claims 1 to 10; 2) means for preparing a DNA library and sequencing DNA molecules from a biological sample to obtain multiple sequence reads; 3) A means to perform mapping and SNP calling by a computer system to generate and quality control the SNP data; 4) A means for performing haplotype prephasing on blood samples from both parents or one parent carrying the disease by a computer system; 5) A means to calculate the likelihood of the four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of both parents, taking into account the maternal cell contamination (MCC) rate and haplotype status; 6) A means to determine the haplotype status of each SNP to estimate MCC rates and indicate whether DNA from only one parent is present in the culture medium; 7) A means to calculate the likelihood of observing SNP data within regions flanking the disease-causing mutation site and combine the likelihood and recombination probability at each single SNP; 8) A means of determining whether an embryo carries a disease-causing chromosome and providing a confidence level.

64. 64. The computer system of claim 63, wherein the embryo has contaminants and / or haplotype losses originating from maternal cells.

65. 64. The computer system of claim 63, wherein in means 3), the mapping and SNP calling outputs the physical location, allele type, allele depth and sequencing quality for each SNP in a biological sample of culture medium of an in vitro cultured embryo and / or a biological sample of a discarded embryo within the family.

66. 64. The computer system of claim 63, wherein in means 3) quality control of SNP data is performed by filtering out SNPs with low genetic quality (GQ) values ​​from biological samples of culture medium of in vitro cultured embryos and / or biological samples of discarded embryos within the family.

67. 64. The computer system of claim 63, wherein SNPs having the lowest 20% GQ values ​​are filtered out from biological samples of culture medium from in vitro cultured embryos and / or biological samples from discarded embryos within the family.

68. 64. The computer system of claim 63, wherein in step 3) the sequencing error rate and relative density of SNPs in regions adjacent to the disease-causing mutation site are estimated and culture medium samples with extremely low SNP density or high sequencing error rate are excluded.

69. 69. The computer system of claim 68, wherein the excluded SNP density is less than 0.001 or the excluded sequencing error rate is greater than 0.

2.

70. The computer system of claim 63, wherein in step 4), one or more of the following samples within the family are used: a blood sample from the proband, a biological sample from a discarded embryo, or a blood sample from at least one grandparent.

71. 64. The computer system of claim 63, wherein in means 5), estimated sequencing error rates are integrated when calculating the likelihood of each individual SNP.

72. Means 6) has the following functions: a) first estimating the MCC rate and haplotype status of each SNP within the region flanking the disease-causing mutation site; and b) Recursively updating the MCC rate and haplotype state of each SNP 64. The computer system of claim 63, further comprising:

73. 73. The computer system of claim 72, wherein function b) uses a landmark search method.

74. 73. The computer system of claim 72, wherein after function a), samples with abnormal MCC rates are excluded.

75. 75. The computer system of claim 74, wherein the abnormal MCC rate is greater than 0.65 or less than −0.

5.

76. 64. The computer system of claim 63, wherein means 7) uses a Bayesian model and a recursive algorithm.

77. The computer system according to claim 63, wherein in step 8), log-likelihood curves are first plotted for both the upstream and downstream regions of the disease-causing mutation site.

78. 64. The computer system of claim 63, wherein the genetic disorders include both monogenic and polygenic disorders.

79. 79. The computer system of any one of claims 63 to 78, further comprising means for combining a plurality of identified SNPs of an embryo to determine its susceptibility to a polygenic disease.

80. 1. A method for preparing a biological sample of culture medium from an otherwise unsequenceable in vitro cultured embryo into a transformed state that can be analyzed to determine susceptibility to a genetic disease in the embryo, comprising: 1) amplifying DNA molecules from biological samples of culture medium of in vitro cultured embryos and / or from biological samples of discarded embryos within the family; 2) preparing a DNA library and sequencing DNA molecules from the biological sample to obtain multiple sequence reads; 3) performing mapping and SNP calling by a computer system to generate and quality control the SNP data; 4) performing haplotype prephasing on blood samples from both parents or one parent carrying the disease by a computer system; 5) calculating the likelihood of four possible inheritance scenarios at each single SNP, representing whether the genetic strand is inherited from the paternal or maternal allele of both parents, by a computer system, taking into account the maternal cell contamination (MCC) rate and haplotype status; 6) A computer system transforms the SNP data by estimating MCC rates and determines the haplotype status of each SNP to indicate whether DNA from only one parent is present in the culture medium; 7) calculating the likelihood of observing SNP data in regions flanking the disease-causing mutation site using a computer system, and combining the likelihood and recombination probability at each single SNP; and 8) A computer system will determine whether an embryo carries a disease-causing chromosome and provide a confidence level. Including, A method wherein the biological sample has a high ADO rate in the presence of maternal contaminants.

81. The method of claim 80, further comprising the step of any one of claims 13 to 27.

Citation Information

Patent Citations

  • Nucleic acid sequence amplification method

    JP2018527947A