Methods for identifying genetic variants in embryos

The method for genetic screening of embryos by analyzing multiple analytes and using deep learning-based tools addresses the failure of current tests to detect disease-causing variants, enhancing pregnancy success by selecting healthy embryos for implantation.

JP2026506491APending Publication Date: 2026-02-25エンブリオーム インコーポレイテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025543172
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-26
Filing Date
2024-01-26
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Current prenatal tests, carrier screening, and preimplantation genetic testing fail to detect the majority of disease-causing genetic variants in embryos, leading to high miscarriage rates and genetic defects in offspring, as they focus on aneuploidies and known genetic disorders, missing potentially lethal mutations.

Method used

A method for genetic screening of embryos involves analyzing multiple sources of analytes from an embryo, comparing them to a reference genome, and using a variant caller to identify genetic variants, including SNVs, CNVs, SVs, and epigenetic markers, with deep learning-based tools to predict embryo health and rank embryos for implantation based on genetic risk.

Benefits of technology

Accurately identifies genetic variants in embryos, improving pregnancy success rates by selecting embryos with the highest likelihood of full-term pregnancy and lowest lifetime genetic risk, reducing the incidence of genetic defects and miscarriages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026506491000001_ABST
    Figure 2026506491000001_ABST
Patent Text Reader

Abstract

The present disclosure provides, in part, a method for identifying genetic variants in an embryo, the method including the steps of: (a) obtaining two or more sources of analytes from an embryo; (b) analyzing the two or more sources of analytes to obtain genetic information for each source; (c) comparing the genetic information for each source against one or more reference genomes using at least one variant caller, where the variant caller identifies variants between each source and the reference genome; and (d) combining the variants to identify differences present only in the sources, where the differences are genetic variants in the embryo.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 441,291, filed January 26, 2023, the entire contents of which are incorporated herein by reference.

[0002] FIELD OF THE INVENTION The present disclosure relates generally to the field of reproduction. More specifically, the present disclosure relates to methods for identifying genetic variants in embryos to assess disease risk. [Background technology]

[0003] Background of the Invention Currently available prenatal tests, carrier screening, and preimplantation genetic testing ("PGT") typically focus on detecting aneuploidies and known genetic disorders for which the parents are carriers, but fail to detect the majority of disease-causing genetic variants in individual embryos. Even if a family is able to rule out one or several known genetic risks, the fertilized egg still carries a potentially lethal genetic defect.

[0004] Natural conception presents many unique developmental challenges. Approximately half of all fertilized eggs carry potentially lethal genetic mutations that can affect pregnancy, childhood, or adulthood. As a result, approximately one in three pregnancies ends in miscarriage (Wilcox AJ, et al. N. Engl J Med, 1988; 319: 189-194), approximately one in 14 children is born with a genetic defect (Ceyhan-Birsoy O, et al. Am J. Human Genet, 2019; 104: 76-93), and approximately one in 20 adults harbors genetic disease variants that dramatically increase the risk of many cancers, sudden cardiac death, aneurysms, autoimmune disorders, and neurodegenerative diseases such as Huntington's disease (Abul-Husn NS, et al. Science, 2016; 354: 6319).

[0005] Rare, high-impact variants, not polygenic risk, guide clinical decisions. Polygenic risk is a good way to understand populations but is inadequate for understanding individuals or the developmental success of a single zygote. For example, for common diseases such as breast cancer, the lifetime risk is approximately 15%. Women in the top 1% of polygenic risk scores ("PRS") have a 31% lifetime risk of breast cancer, twice the lifetime risk of normal women. Meanwhile, carriers of rare BRCA1 / 2 mutations have an 85% lifetime risk of developing breast cancer (six times normal). For very rare diseases such as brain cancer, the lifetime risk is approximately 0.01%. Common variants pose a 1.2- to 3.5-fold lifetime risk (Melin et al., 2017, Nat. Genet. 49(5): 789-794), while rare variants confer a 5,000-fold lifetime risk (Bainbridge, et al., 2015, J. Natl. Cancer Inst. 107(1): 1-4). Most clinical tests, such as carrier screening and preimplantation genetic testing (including PGT-monogenic disorders (PGT-M) and PGT-A aneuploidies), use targeted approaches to balance sensitivity and specificity to avoid missing disease (i.e., avoiding "false negatives") while simultaneously avoiding the anxiety and discarding of potentially viable embryos due to "false positives." PGT-A accounts for approximately 50% of failed in vitro fertilization (IVF) cycles and pregnancy losses (Zhao C, et al. 2021, Gen in Med. 23:435-442). PGT next-generation sequencing (NGS) can diagnose whole chromosome aneuploidies (WCA) in 95% of embryos, but does not accurately detect events below the chromosome level (Cascante SD, et al. 2023, Fert. Ster. 120(6):1161-1169). Carrier testing reports 200–800 genes, while PGT-M may report only one or two. Prenatal testing, such as noninvasive prenatal testing (NIPT), focuses on the most likely genetic errors, which are easiest to accurately detect. Most NIPT assays report 5–13 conditions. However, it is important to note that the only action that can be taken based on prenatal testing is to continue or terminate the pregnancy. Preimplantation is the most actionable period for genetic findings. For example, BRCA2 mutation carriers have a 70% lifetime risk of developing breast cancer. Risk reduction depends on the life stage. If diagnosed during adulthood, risk can be reduced by increased monitoring or preventive surgery. Prenatal diagnosis is not currently available because breast cancer risk can only be reduced by terminating the pregnancy. In contrast, embryo screening does not pose a risk if parents choose to implant another embryo. Whole exome sequencing (WES) allows comprehensive screening of the protein-coding region of the embryo's genome (the exome), which represents only a small portion of the entire genome (approximately 2%) but contains many variants known to cause disease.A recent study applying exome sequencing showed that 60–75% of all sporadic cases tested could be explained by de novo mutations, i.e., changes not found in either parent (Acuna-Hidalgo R, et al. 2016, Genome Biology, 17:241). Generally, every embryo formed after fertilization carries approximately 100 de novo genetic changes. The older the parents, the more likely they are to accumulate changes in their sperm and eggs that can be passed on to the embryo. Interestingly, trio testing (parent-child trio) has been found to increase the success rate of genomic diagnosis for rare pediatric diseases by nearly fivefold (Wright CF, et al. 2023, N Engl. J. Med. 388:1559–1571). In the same study, approximately 76% of 3,599 diagnosed affected individuals who underwent trio testing were found to have pathogenic de novo variants (ibid.). Therefore, there is a need for accurate embryo screening methods that involve the analysis of more pregnancy success determinants prior to embryo transfer and implantation, resulting in an accurate assessment of genetic health and embryo outcome throughout development and beyond. [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] Wilcox AJ, et al. N. Engl J Med, 1988; 319: 189-194 [Non-patent document 2] Ceyhan-Birsoy O, et al. Am J. Human Genet, 2019; 104: 76-93 [Non-patent document 3] Abul-Husn NS, et al. Science, 2016; 354: 6319 [Non-patent document 4] Melin et al., 2017, Nat. Genet. 49(5): 789-794 [Non-patent document 5] Bainbridge, et al., 2015, J. Natl. Cancer Inst. 107(1): 1-4 [Non-patent document 6] Zhao C, et al. 2021, Gen in Med. 23:435-442 [Non-Patent Document 7] Cascante SD, et al. 2023, Fert. Ster. 120(6):1161-1169 [Non-patent document 8] Acuna-Hidalgo R, et al. 2016, Genome Biology, 17:241 [Non-Patent Document 9] Wright CF, et al. 2023, N Engl. J. Med. 388:1559-1571 Summary of the Invention [Means for solving the problem]

[0007] Summary of the Invention In part, the present disclosure provides methods for genetic screening of embryos, particularly prior to embryo transfer and implantation, in which an individual source of the analyte is analyzed and compared to a reference genome to identify genetic variants in the embryo.

[0008] Accordingly, one aspect of the present disclosure provides a method for identifying genetic variants in an embryo, the method comprising obtaining two or more sources of analytes from the embryo, analyzing the two or more sources of analytes to obtain genetic information for each source, comparing the genetic information for each source against one or more reference genomes using at least one variant caller that identifies variants between each source and the reference genome, and combining the variants to identify differences that exist only in the sources, thereby identifying genetic variants in the embryo.

[0009] In some embodiments, genetic variants are single nucleotide variants (SNVs), multinucleotide variants (MNVs), copy number variants (CNVs), structural variants (SVs), or changes in epigenetic markers.As a non-limiting example, epigenetic markers are DNA methylation markers that are used to identify imprinting defects.As a non-limiting example, genetic variants are identified as apparent de novo variants (DNVs).

[0010] In some embodiments, the two or more sources of analytes include cells biopsied from an embryo or blastocyst, cell culture medium, cells and / or genetic material isolated from the blastocoel, or combinations thereof.

[0011] In some embodiments, the biopsy is collected using laser pulse-assisted pipette delivery. In some embodiments, two or more cells are isolated from the biopsy and separated using physical or enzymatic methods. Non-limiting examples of physical methods include laser dissection or micropipette isolation. Non-limiting examples of enzymatic methods include extraction using digestive enzymes, including but not limited to dispase, collagenase, hyaluronidase, papain, DNase-I, Accutase, or trypsin.

[0012] In some embodiments, the genetic information obtained from each source may include nucleic acid sequences of a panel, exome, whole genome, or transcriptome, with or without mitochondrial sequences, with or without epigenetic markers. As a non-limiting example, the nucleic acid may include DNA, RNA, or both.

[0013] In some embodiments, nucleic acid sequences are obtained by isolating DNA from sources, generating libraries by whole genome amplification of DNA from each source, and sequencing the libraries to generate sequence data files for the sources.

[0014] In some embodiments, before carrying out complete sequencing of the library, embryos can be sequenced at lower coverage, and aneuploidy can be identified.The embryos can then be excluded from further sequencing.By way of non-limiting example, aneuploidy can be detected by unexpectedly high or low coverage in genomic regions.

[0015] In some embodiments, RNA is isolated and sequenced.In some embodiments, RNA sequence data is used to characterize transcriptome and determine gene expression level.In some embodiments, gene expression level is correlated with patient outcome.As a non-limiting example, patient outcome includes, but is not limited to, successful pregnancy.

[0016] In some embodiments, a panel of genetic information is established that correlates gene expression levels with patient outcome.

[0017] In some embodiments, RNA-seq data is used to identify aberrant splicing events, thereby identifying transcriptional dysregulation and one or more variants that affect one or more transcriptional pathways. By way of non-limiting example, aberrant splicing events can include cryptic splicing, aberrant exon usage, or upstream open reading frames.

[0018] In some embodiments, RNA-seq data is used to identify biased allelic expression, thereby identifying imprinting, copy number determination, and / or variants that affect expression.

[0019] In some embodiments, RNA-seq data is used in combination with DNA sequencing or alone to call genomic variants.

[0020] In some embodiments, the reference genome is a genomic sequence from a public or private repository, one or more parents, genetic relatives, and / or embryos originating from the same parents.

[0021] In some embodiments, genetic variants are further identified by comparing the genetic information of each source with the genetic information of one or more parents to estimate haplotype phasing, impute missing data, determine inheritance or lack thereof, and identify apparent de novo variants. In some embodiments, genetic variants are further annotated for phenotypic associations using databases of variants associated with diseases, phenotypes, disease risk or protection, and traits. In some embodiments, genetic variants are further annotated for their potential to cause adverse effects on gene function via in silico, in vivo, or in vitro methods.

[0022] In some embodiments, the variant caller is a deep learning-based variant caller. In some embodiments, the deep learning-based variant caller is constructed based on genetic information from one or more parents, genetic relatives, and / or embryos derived from the same parents, and is capable of identifying apparent de novo variants for the purpose of preimplantation genetic screening. In some embodiments, the genetic variants are present in genomes with known or suspected contributions to adult or childhood genetic diseases, or genomes that are not currently associated with a genetic disease but are under genetic constraints.

[0023] In some embodiments, the combining step of the methods described above and herein comprises identifying genetic variants present in at least two sources, or identifying genetic variants that can only be detected by combining genetic information from multiple sources. As a non-limiting example, genetic variants are identified in multiple sources, or are determined to be errors, artifacts, or mosaics.

[0024] Some aspects of the present disclosure provide a method for screening multiple embryos. Such a method includes performing genetic analysis on each embryo as described above and in the methods described herein, and then ranking the embryos based on the results of the genetic analysis. The embryo with the highest likelihood of full-term pregnancy and the lowest lifetime genetic risk is ranked highest for transfer and implantation. In some embodiments, the embryos are ranked based on the known harmfulness of the detected variants or the predicted harmfulness of the detected variants. Here, such predictions are made by various methods. As non-limiting examples, prediction methods may include in silico harmful prediction tools, allele evolutionary conservation, and / or allele frequencies in selected or unselected population databases.

[0025] In some embodiments, the effect of genetic variants on embryo rank is weighted based on input from the phenotype selected for investigation. As a non-limiting example, the input may be a customer's desire to avoid a particular pathological condition.

[0026] In some embodiments, the influence of genetic variants on embryo rank is weighted based on automated phenotypic overlap between disease-associated phenotypes and their associated genes and the phenotypes selected for investigation. Overlap can be computed using an ontology that can determine related terms as overlaps rather than exactly the same terms. As a non-limiting example, the phenotypes selected for investigation may be derived from the family history of one or both parents, where the parents may wish to avoid breast cancer risk, and therefore genes associated with risk of other cancers are also weighted more heavily.

[0027] In some embodiments, the impact of a genetic variant on embryo rank is weighted based on variant zygosity, the presence of other variants in the genome, the predicted deleteriousness of the variant, whether the variant is benign, unknown, likely pathogenic, or previously curated as pathogenic, and known disease severity, penetrance, and / or inheritance patterns associated with the gene.

[0028] In some embodiments, the effect of a genetic variant on the rank of an embryo is weighted based on the presence or absence of the variant in one or both parents and its zygosity in the embryo or parent.

[0029] In some embodiments, the effect of genetic variants on embryo rank is weighted based on the age of onset of known or suspected disease.

[0030] In some embodiments, the influence of gene variants on embryo rank is weighted based on previously curated gene or variant rankings.

[0031] In some embodiments, the effect of genetic variants on embryo rank is weighted based on treatment options for a given disease.

[0032] In some embodiments, the effect of genetic variants on embryo rank is weighted based on their effect on gene function or expression.

[0033] In some embodiments, the screening method further comprises generating a report regarding the ranking of the embryos, wherein the ranked embryos are qualitatively categorized as recommended for transfer, recommended not for transfer, or unknown for transfer.

[0034] In order to better understand the subject matter disclosed herein and to illustrate how it may be carried out in practice, embodiments will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which: [Brief explanation of the drawings]

[0035] [Figure 1] Figure 1 illustrates carrier screening for reproductive risk.

[0036] [Figure 2] Figure 2 is a flow chart outlining whole genome sequencing (WGS) and whole genome amplification (WGA) of cell lines, where analytes from individual cells and bulk cells are analyzed and compared.

[0037] [Figure 3] Figure 3 is a flow chart outlining a method for WGS and WGA in embryo samples and subsequent identification of genetic variants, where multiple sources of analytes are analyzed.

[0038] [Figure 4] FIG. 4 is a flowchart outlining a method for identifying genetic variants in embryos (using WGA and WGS) that takes into account parental genetic data.

[0039] [Figure 5-1] Figure 5 is a flowchart illustrating the classification of variants into broad categories, including known pathogenic variants, non-coding variants, coding variants, loss-of-function (LoF) variants, synonymous variants, and variants not worthy of further consideration (discard), as well as "panic" situations that require manual review. [Figure 5-2] Same as above.

[0040] [Figure 6-1] FIG. 6 is a flow chart illustrating classification of variants into broad categories, such as "No go," where the embryo will not be considered for implantation, for inherited coded LoF variants, as well as weighting of variants to aid in embryo ranking. [Figure 6-2] Same as above.

[0041] [Figure 7] FIG. 7 is a flowchart illustrating classification of variants into broad categories such as "No go," weighting, and manual review of known pathogenicity (e.g., based on the ClinVar database).

[0042] [Figure 8-1] Figure 8 is a flowchart illustrating variant classification (the "autosomal recessive process") for "pathogenic" (i.e., likely to cause disease, known to cause disease, or possible to cause disease) variants in genes associated with autosomal recessive diseases. Determining whether multiple variants are present in a gene and whether those variants are in cis or trans, or whether variants are present in a gene in a homozygous or hemizygous state, broadly categorizes embryos as "No go," weighted, or subjected to manual review. [Figure 8-2] Same as above.

[0043] [Figure 9] FIG. 9 is a flow chart illustrating the classification of de novo LoF code variants into the broad categories described above.

[0044] [Figure 10-1] FIG. 10A is a flow chart illustrating structural variant classification into "No go" for embryos or reclassifying as LoF in the appropriate gene for SV deletions.

[0045] FIG. 10B is a flow chart illustrating structural variant classification into "No go" for the embryo or reclassification as LoF in the appropriate gene for SV acquisition.

[0046] [Figure 10-2]FIG. 10C is a flow chart illustrating structural variant classification into "No go" for the embryo or reclassifying as LoF in the appropriate gene for the inversion.

[0047] FIG. 10D is a flow chart illustrating structural variant classification into "No go" for the embryo or reclassifying as LoF in the appropriate gene for the translocation.

[0048] [Figure 11-1] FIG. 11 is a flow chart illustrating repeat expansion categorization based on repeat expansions of known pathogenic and "premutation" length as well as repeat length in the parents, to broadly categorize embryos as described above. [Figure 11-2] Same as above.

[0049] [Figure 12-1] FIG. 12A is a flow chart illustrating classification of large regions of homozygous variants as treated as normal (recessive disease), treated as normal (deletion), and classified as manual review.

[0050] FIG. 12B is a flow chart illustrating the categorization of synonymous and intronic variants that are classified as treated as LoF.

[0051] [Figure 12-2] FIG. 12C is a flow chart illustrating categorization of large regions of uniparental inheritance into classification as No go or manual review.

[0052] FIG. 12D is a flow chart illustrating potential categorization of variants identified, for example, from the spinal muscular atrophy pipeline.

[0053] [Figure 13]Figure 13 is a flowchart illustrating an example of categorizing SNVs or small indel variants into various broad categories. The initial steps, DeepVariant, VCF Info, and Population Frequency, are detailed, leading in some cases to broad categories of discard or later extension.

[0054] [Figure 14-1] Figure 14A is a flowchart illustrating the reference sequence step of the analysis. Categorization can proceed through the following decision points: coding variant, intronic variant, intergenic variant, nonsense or LoF variant, missense variant, synonymous variant, splicing-affecting variant (X out of 4), splicing-affecting variant (clinical X out of 4), and evidence of pathogenicity (ClinVar), potentially leading to broad categories of panic, discard, or later expansion.

[0055] Figure 14B is a flow chart illustrating the phenotyping steps in the analysis, where categorization can proceed through decision points of meeting ACMG criteria, meeting ClinVar pathogenicity criteria, meeting OMNIM criteria, and meeting embryonic "Intolerome" criteria, potentially leading to broad categories of discard or later expansion.

[0056] Figure 14C is a flowchart illustrating the steps of trio genetic analysis, where categorization can proceed through the following decision points: De Novo dominant, inherited dominant, homozygous recessive, X-linked, inherited compound heterozygosity, and manual review capture net, potentially leading to the broad category of panic, or later expansion. Note that each of the De Novo dominant, inherited dominant, homozygous recessive, X-linked, and inherited compound heterozygosity decision points has a sub-flowchart detailed in Figures 16A-16E. [Figure 14-2] Same as above.

[0057] [Figure 15] FIG. 15 is a flow chart illustrating the reporting / occlusion subset step of the process.

[0058] [Figure 16-1] Figure 16A is a sub-flowchart for the De Novo Dominant decision point. The ksw analysis can proceed through the decision points of De Novo (not from parent), from parent, dominant mechanism, and recessive mechanism, potentially leading to the broad categories of PANIC or other pipeline references.

[0059] Figure 16B is a sub-flowchart illustrating the X-linked decision point, where analysis can proceed through the decision points of non-PAR region, proband male and hemizygous, and proband female and homozygous, potentially leading to the broad categories of manual review or other pipeline reference.

[0060] [Figure 16-2] Figure 16C is a sub-flowchart illustrating the inherited dominant decision point, where analysis can proceed through the decision points of De Novo (not from parent), parental, (P1)-hetero / (P2)-WT / (E)-hetero, rare (P1 / 2) and (E)-hetero, dominant mechanism, and recessive mechanism, potentially leading to the broad category of panic or other pipeline reference.

[0061] [Figure 16-3]Figure 16D is a sub-flowchart illustrating the homozygous recessive decision point, where analysis can proceed through the following decision points: De Novo (not from parent), parental, (P1)-hetero / (P2)-hetero / (E)-homo, (P1)-hetero / (P2)-WT / (E)-homo, rare (P1 / 2) and (E)-homo, dominant mechanisms, and recessive mechanisms, potentially leading to the broad categories of manual review, PANIC, or other pipeline references.

[0062] [Figure 16-4] Figure 16E is a subflowchart illustrating the inherited compound heterozygosity decision point, where analysis can proceed through the following decision points: genes with two or more SNVs, SNV phasing locations on separate chromosomes, SNV phasing unknown, LoF / LoF phased SNVs, LoF / other pathogenic phased SNVs, LoF / LoF SNVs, Lof / other pathogenic SNVs, potentially leading to the broad categories of discard, manual review, and panic.

[0063] [Figure 17A] FIG. 17A shows an example of aneuploidy calling for a set of individual cells derived from an aneuploid female embryo that showed a consistent loss of chromosome 19.

[0064] [Figure 17B] Figure 17B shows a different example of aneuploidy calling for a set of individual cells derived from a female embryo that was euploid, where some copy number variants were observed in individual cells but were not consistently represented throughout the embryo. DETAILED DESCRIPTION OF THE INVENTION

[0065] Detailed Description To facilitate understanding of the principles of the present disclosure, reference will now be made to embodiments and specific language will be used to describe the embodiments, but no limitation of the scope of the present disclosure is intended, and it will be understood that such changes and further modifications of the present disclosure exemplified herein would normally occur to one skilled in the art to which the present disclosure pertains.

[0066] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0067] As used herein, the terms "comprise," "comprising," "include," "including," "have," "having," and their cognates mean "including but not limited to." The term "consisting of" means "including and limited to." The term "consisting essentially of" means that a composition, method, or structure may include additional components, steps, and / or portions, but only if the additional components, steps, and / or portions do not materially alter the basic and novel characteristics of the claimed composition, method, or structure.

[0068] As used herein, the term "method" or "methods" refers to manners, means, techniques and procedures for accomplishing a given task, including, but not limited to, manners, means, techniques and procedures that are known or readily derivable from known manners, means, techniques and procedures to practitioners in the arts of chemistry, pharmacology, biology, biochemistry and medicine.

[0069] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the present invention. Thus, the description of a range should be considered to include all possible subranges specifically disclosed as well as individual numerical values ​​within that range. For example, the description of a range such as 1 to 6 should be considered to include specifically disclosed subranges, such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numbers within that range, such as 1, 2, 3, 4, 5, and 6. This is true regardless of the breadth of the range.

[0070] Whenever a numerical range is given herein, it is intended to include every recited number (fractional or integer) that falls within the range given. The phrases "ranging / ranges between" a first recited number and "ranges from" a first recited number and "to" a second recited number are used interchangeably herein and are intended to include the first recited number and the second recited number, and all fractional and integer numbers therebetween.

[0071] As used herein, the term "analyte" or "analytes" refers to biological macromolecules, e.g., proteins, DNA, and / or RNA, necessary for the function, growth, and development of cells and organisms.

[0072] As used herein, the phrase "a source of analyte" or "sources of analytes" refers to a source from which the genetic information of an embryo can be obtained, including, but not limited to, a single cell or group of cells biopsied from an embryo or blastocyst, cell culture medium, cells or genetic information from the blastocoel, or a combination thereof.

[0073] As used herein, the phrase "genetic information" refers to the nucleic acid sequences of a panel, exome, whole genome, or transcriptome, with or without mitochondrial sequences, and with or without epigenetic markers.

[0074] As used herein, the phrase "reference genome" refers to the whole genome sequence from a public or private repository, one or more parents, or genetically related individuals, e.g., biological siblings and / or embryos derived from the same parents, from which reference genetic information can be obtained.

[0075] As used herein, the term "whole genome amplification" or "WGA" refers to a technique, method, or kit used to generate large amounts of DNA from small amounts of starting material. Unlike traditional PCR, WGA aims to amplify the entire genome of an organism, rather than a specific region.

[0076] As used herein, the terms "whole genome sequencing," "whole-genome sequencing," and "WGS" are used interchangeably and refer to a laboratory process in which the entire or nearly entire DNA sequence of an organism's genome is determined at once. In this disclosure, whole genome sequencing is performed on each embryo and each parent to determine all 6 billion bases of each individual embryo's genome and parent's genome, respectively, where the genome may or may not include the mitochondrial genome. The result of such sequencing is the complete DNA sequence of an individual, including non-coding sequences and which may include epigenetic markers.

[0077] As used herein, the term "coverage" refers to the average number of times a particular sequence is explored by sequence reads. For example, "30xWGS" means that the entire genome is sequenced at an average overlap depth of 30 times.

[0078] As used herein, the term "coverage analysis" refers to assessing the genome for variability, including, but not limited to, whether there are regions with higher or lower coverage than expected. When expected coverage is based on total sequence coverage, known biases in coverage based on local genome composition or embryo sex can be empirically determined.

[0079] As used herein, the term "variant calling" refers to the process of identifying the difference between the sequencing reads of a target individual and the sequencing reads of a reference genome.Similarly, as used herein, the term "variant caller" refers to a tool for variant calling of sequencing data from one or more nucleic acid sequencing datasets (DNA, RNA, etc.).As a non-limiting example, the variant caller used herein may include DeepVariant, a new TensorFlow machine learning-based variant caller.

[0080] As used herein, the phrase "training DeepVariant" refers to feeding a training algorithm a dataset with expected results known as a "truth set," thereby penalizing or rewarding the machine learning algorithm based on performance. The algorithm is then used on a dataset with no known truth set. Similar performance is expected.

[0081] As used herein, the term "merged variant calling" refers to merging the results of multiple single-source variant calls from a single embryo or calls that look at all embryo data together with parental data simultaneously. The term "merged variant caller" refers to a tool or method for variant calling from multiple sources.

[0082] As used herein, the phrase "genetic variant" refers to differences in genetic information between sources of an analyte derived from an embryo, or between the embryo genome and a reference genome. Various types of genetic variants are described below.

[0083] As used herein, the term "single nucleotide variant," "single-nucleotide variant," or "SNV" refers to a variation in a single base that occurs when a single base (adenine, thymine, cytosine, or guanine) in a genomic sequence is changed. SNVs may be rare or common in one population and common or rare in a different population. Sometimes SNVs are known as single nucleotide polymorphisms (SNPs), but SNVs and SNPs are not interchangeable. To be considered an SNP, the variant must be present in at least 1% of a population.

[0084] As used herein, the term "indel" is an abbreviation for insertion / deletion, and refers to the simultaneous insertion and deletion of short lengths of DNA (usually less than 50 base pairs) into and from a genome relative to a reference genome. A specific example of an indel is the expansion or contraction of a repeat, whereby a repetitive element of DNA becomes longer or shorter.

[0085] As used herein, the term "copy number variant" or "CNV" refers to duplication or deletion that changes the number of copies of a specific DNA segment in the genome.Aneuploidy is one type of CNV, and typically consists of a very large segment of a chromosome or an entire chromosome.CNV contributes to 9% of the pathogenesis of hereditary retinal degeneration (Zampaglione et al., 2020; Genetics in Medicine, 22: 1079-1087).

[0086] As used herein, the term "structural variant" or "SV" refers to a rearrangement of a portion of the genome, which may be a deletion, duplication, insertion, inversion, translocation, or a combination thereof, and may or may not be copy-neutral. SVs have been implicated in several conditions, including polycystic kidney disease, cardiomyopathy, amyotrophic lateral sclerosis (ALS), and some cases of intellectual disability.

[0087] As used herein, the term "disease-causing genetic variant" refers to a genetic change in a DNA sequence that has been shown to cause disease or is associated with an increased risk of developing disease.

[0088] As used herein, the term "phenotype" refers to any measurable characteristic of an individual that is attributable at least in part to genetics. Phenotypes may include, but are not limited to, traits disease risk or protection.

[0089] As used herein, the terms "de novo variant," "apparent de novo variant," "DNV," or "apparent DNV" refer to a genetic variant in an embryo that is not clearly present in sequence data generated from the parents. In some cases, the variant is present in the parent's egg or sperm cells, but not in any of their other cells. In other cases, the variant may be present in a subset of the parent's cells, including the gonads, but not in the sequenced sample (mosaicism). In still other cases, the variant is present in the embryo after fertilization. As the embryo develops, some or all of the resulting cells in the developing embryo contain the variant. A de novo variant is one explanation for a genetic disorder that is present in an affected child but not in either parent.

[0090] As used herein, the term "parent" refers to the gamete donor (i.e., biological parent) who contributes to the genetic makeup of the embryo.

[0091] As used herein, the terms "patient outcome," "patient outcomes," and the like refer to the outcome of a pregnancy, the development of a disease, or the occurrence of other desirable or undesirable conditions in a parent, embryo, or child.

[0092] As used herein, terms such as "favorable outcome," "favorable outcomes," "successful pregnancy," and the like refer to the avoidance of certain negative or pathological conditions. For example, a favorable outcome may be carrying the fetus to term. In another non-limiting example, a favorable outcome may be that the embryo does not have a genetic predisposition to one or more pathological conditions. The one or more pathological conditions may vary depending on the desires of the client or biological parents.

[0093] As used herein, the terms "pathologic condition," "pathologic conditions," and the like refer to abnormal anatomical or physiological conditions to be avoided when screening embryos at the request of a client or biological parent.

[0094] In part, the disclosure provides a method for identifying genetic variants in an embryo, the method comprising obtaining two or more sources of analytes from the embryo, analyzing the two or more sources of analytes to obtain genetic information for each source, comparing the genetic information for each source against one or more reference genomes using at least one variant caller to identify variants between each source and the reference genome, and combining the variants to identify differences present only in the sources, thus identifying genetic variants in the embryo.

[0095] The genetic variants identified can be single nucleotide variants (SNVs), multinucleotide variants (MNVs), copy number variants (CNVs), structural variants (SVs), or alterations in epigenetic markers.

[0096] MNV includes, but is not limited to, insertion, deletion, combination of insertion and deletion, or repeat expansion involving multiple nucleotides.MNV calling can include, but is not limited to, identifying insertion, deletion, and simultaneous insertion-deletion (indel), as well as repeat expansion and repeat shortening (special cases of insertion and deletion, respectively).By way of non-limiting example, MNV calling can identify single-base deletion / insertion / indel or multiple-base deletion / insertion / indel.

[0097] SV includes, but is not limited to, inversion, translocation, repeat expansion and repeat shortening (as a special case of insertion or deletion).SV calling can include, but is not limited to, identifying genomic rearrangements, including rearrangements that cause copy number changes and copy-neutral changes.By way of non-limiting example, copy-neutral changes can include translocations and inversions.

[0098] The epigenetic marker can be a DNA methylation marker used to identify imprinting defects. Methylation markers can be used to identify imprinting defects, including but not limited to, biallelic imprinting and biased X-chromosome inactivation.

[0099] Parental data can be used to identify genetic mutations that exist in embryos but do not exist in parental WGS, i.e., de novo variants (DNVs).Parental data can also be used to determine inheritance, including haplotype phasing and isodisomy, and to determine whether inherited variants are cis or trans, i.e., whether one is maternal and the other paternal.A non-limiting example is determining whether one variant is paternal and the other variant is maternal, which may cause recessive diseases.

[0100] Parental data can also be used to impute missing data in the embryo. Sometimes, it is not possible to assay parts of the embryo's genome by amplification or sequencing, but from the inheritance of certain haplotypes, other variants can be inferred to be present in the embryo even if they are not directly observed.

[0101] In some embodiments, the methods for identifying genetic variants in embryos described above and disclosed herein can be used for preimplantation genetic testing ("PGT"), in which embryo genetic testing is performed, and for screening embryos before transfer and implantation, in which the genomes of all embryos and individual cells isolated from both parents are sequenced and analyzed. Generally, prospective parents and / or clients create a set of their embryos using IVF technology. The embryos are then biopsied using standard protocols. Before the embryos are transferred for pregnancy, all 3 billion base pairs (or 6 billion bases) of each embryo's genome are sequenced along with the genomes of both biological parents. For whole-genome sequencing of each embryo, individual cells are isolated from the embryo and sequenced. The sequencing data are then merged and analyzed to identify embryos with the lowest genetic risk for embryo transfer and implantation. PGT and the methods described above and disclosed herein detect genetic risks not only before birth but also after birth.

[0102] In some embodiments, variations of the methods or related techniques for identifying genetic variants in embryos described above and disclosed herein can be applied to free-floating or biopsied embryonic cells collected from the uterus or bloodstream of a pregnant woman. That is, in addition to PGT of embryos in assisted reproductive technologies (ART) pregnancies, similar methods or techniques can also be used with improved genetic screening for naturally occurring pregnancies.

[0103] In some embodiments, variations of the methods or related techniques for identifying genetic variants in embryos described above and disclosed herein can be applied to one or more biological parents. That is, in addition to PGT of embryos in assisted reproduction pregnancies, similar methods or techniques can also be used to inform natural conception planning with improved genetic screening.

[0104] To identify genetic variants in embryos, analytes can be obtained from cells biopsied from embryos or blastocysts, cell culture medium, cells and / or genetic material isolated from the blastocoel, or combinations thereof. Analytes include, but are not limited to, analytes from culture medium, blastocoel fluid, or parental samples (e.g., blood, buccal swabs, saliva, semen). The bulk cells are then separated / dissociated into single cells. For any of the methods described above and disclosed herein, 1 to 10 single cells can be isolated from each embryo for single-cell analysis. As a non-limiting example, 2, 3, 4, 5, 6, 7, 8, 9, or 10 single cells can be isolated from an embryo.

[0105] In some embodiments, the biopsy is taken using laser pulse-assisted pipette delivery. Embryo biopsies can be performed after 5-6 days of laboratory culture, when the embryo reaches the blastocyst or hatched blastocyst stage. In some embodiments, two or more cells are isolated from the biopsy and separated using physical or enzymatic methods. By way of non-limiting example, physical methods may include laser dissection or micropipette isolation. The cells may then be dissociated into individual cells. By way of non-limiting example, enzymatic methods may include extraction using digestive enzymes, including, but not limited to, dispase, collagenase, hyaluronidase, papain, DNase-I, Accutase, or trypsin.

[0106] Because the amount of DNA contained in single cells is too small to be put into a sequencer, it is necessary to carry out amplification (for example, WGA) before sequencing.Amplification can induce errors, such as failing to amplify some parts of genome, over-amplifying some parts of genome, and failing to amplify both alleles of some parts of genome.These technical artifacts can mimic true and harmful variants.However, by sequencing a large number of single cells, the above-mentioned and methods disclosed herein can infer true variants from artifacts by multiple independent observations of variants.

[0107] The genetic information obtained from each source may include nucleic acid sequences of panels, exomes, whole genomes, or transcriptomes, with or without mitochondrial sequences, with or without epigenetic markers. As a non-limiting example, nucleic acids may include DNA, RNA, or both.

[0108] In some embodiments, nucleic acid sequences are obtained by isolating DNA from sources, generating libraries by whole genome amplification of DNA from each source, and sequencing the libraries to generate sequence data files for the sources.

[0109] In some embodiments, before carrying out complete sequencing of the library, embryos can be sequenced at lower coverage, and aneuploidy can be identified.The embryos can then be excluded from further sequencing.By way of non-limiting example, aneuploidy can be detected by unexpectedly high or low coverage in genomic regions.

[0110] Whole-genome sequencing (WGS) of both the biological parents and each embryo provides the highest possible resolution and sensitivity for genetic analysis. In contrast, current offerings for carrier screening or prenatal testing analyze only about 1% of the genome, where clinical diagnosis is most certain, and analyze only one parent (carrier screening) or the fetus (prenatal testing) at a time, further limiting potential insights. Current offerings for IVF use limited-coverage specialized sequencing or SNP arrays, which generate 1000 times less data than whole-genome sequencing, primarily due to the difficulty of obtaining and interpreting more extensive genetic data, as well as efforts to contain costs.

[0111] In some embodiments, sequencing data, e.g., sequencing data from separate cells, is used to call variants, and the results are then combined for analysis. In some embodiments, sequencing data, e.g., sequencing data from separate cells, is first combined and then used to call variants.

[0112] In some embodiments, RNA is isolated and sequenced. In some embodiments, RNA sequence data is used to characterize the transcriptome and determine gene expression levels. In some embodiments, gene expression levels are correlated with patient outcomes. Non-limiting examples of patient outcomes include, but are not limited to, successful pregnancy. In this method, a panel of genetic information is established that shows the correlation between gene expression levels and patient outcomes.

[0113] RNA sequence data can be used to identify aberrant splicing events, dysregulation of transcriptional regulation, and one or more variants that affect one or more transcriptional pathways. By way of non-limiting example, aberrant splicing events can include cryptic splicing, aberrant exon usage, or upstream open reading frames.

[0114] RNA-seq data can be used to identify biased allelic expression, and thereby also identify imprinting, copy number determination, and / or variants that affect expression.

[0115] In some embodiments, RNA sequence data is used alone or in combination with DNA sequencing to call genome variants.In some embodiments, RNA is used to confirm aneuploidy call.In some embodiments, multiple data points from each cell (DNA, RNA) can be integrated, and / or multiple data points from multiple cells can be integrated.

[0116] To identify genetic variants in embryos, the genetic information of each source is compared with the genetic information of one or more reference genomes. In some embodiments, the reference genome is a genomic sequence from a public or private repository, one or more parents, genetic relatives, and / or an embryo derived from the same parent. When using a genomic sequence from one or more parents as a reference genome, genetic variants can be further identified by comparing the genetic information of each source with the genetic information of one or more parents to estimate haplotype phasing, impute missing data, determine inheritance or lack thereof, and identify apparent de novo variants. In some embodiments, genetic variants are further annotated for association with phenotypes by using databases of variants associated with diseases, phenotypes, disease risk or protection, and traits. In some embodiments, genetic variants are further annotated for their potential to cause adverse effects on gene function through in silico, in vivo, or in vitro methods.

[0117] Identified variants can be annotated based on their presence or absence in other datasets, including cohorts of healthy adults, children with early-onset diseases, and trio datasets. These datasets include private and public datasets. To date, academic research has deployed useful data in silos. Research into childhood and adult genetic diseases is conducted separately from research into the genetic causes of miscarriage and infertility. These existing datasets can be added to embryo research data and customer data to expand the reference sequence database.

[0118] In some embodiments, the variant caller may be a deep learning-based variant caller. In some embodiments, the deep learning-based variant caller is constructed based on the genetic information of one or more parents or genetic relatives and is capable of identifying apparent de novo variants for the purpose of preimplantation genetic screening. Genetic variants known or suspected to contribute to adult or childhood genetic diseases (i.e., disease-causing genetic variants) may be present in the genome. In other situations, genetic variants may be present in a genome that is not currently associated with a genetic disease but is under genetic constraints.

[0119] For the methods of identifying genetic variants in embryos described above and herein, the combining step can include identifying genetic variants present in at least two sources, or identifying genetic variants that can only be detected by combining genetic information from all sources. As a non-limiting example, genetic variants can be identified in multiple sources (not necessarily the majority of sources), or can be identified as mosaic.

[0120] Some aspects of the present disclosure provide a method for screening multiple embryos. Such a method includes carrying out genetic analysis on each embryo as described above and in the methods described herein, and then ranking the embryos based on the results of the genetic analysis. The embryo with the highest probability of full-term pregnancy and the lowest postnatal genetic risk is ranked highest for transfer and implantation. In some embodiments, the embryos are ranked based on the known harmfulness of the detected variants or the predicted harmfulness of the detected variants, where such prediction is made by various methods. As non-limiting examples, the prediction method may include in silico harmful prediction tools, allele evolutionary conservation, and / or allele frequency in selected or unselected population databases.

[0121] The effect of genetic variants on embryo rank can be weighted based on input from the phenotype selected for investigation. As a non-limiting example, the input can be a customer's desire to avoid a particular pathological condition.

[0122] Alternatively, the influence of genetic variants on embryo rank can be weighted based on automated phenotypic overlap between disease-associated phenotypes and their associated genes and the phenotypes selected for investigation. The overlap can be computed using an ontology that can determine related terms as overlaps rather than exactly the same terms. As a non-limiting example, the phenotypes selected for investigation may be derived from the family history of one or both parents, where the parents may wish to avoid breast cancer risk, and therefore genes associated with risk of other cancers are also weighted more heavily.

[0123] Additionally, the impact of a gene variant on embryo rank can be weighted based on the presence or absence of the variant in one or both parents and its zygosity in the embryo or parent, or the age of onset of a known or suspected disease, previously curated gene or variant rankings, treatment options for a given disease, or impact on gene function or expression.

[0124] The screening method described above and herein can further comprise the step of generating a report using the ranking of embryos.Instead of quantitative ranking, embryo risk can be classified into two or more categories, including: red light (high risk; embryo transfer is not recommended); yellow light (some identifiable risk; parents can weight differently based on their respective priorities; whether to transfer is up to the parents' decision); green light (no specific risk is identified; transfer and implantation are recommended); and a set of no data generated.That is, ranked embryos can be qualitatively categorized based on recommended transfer, not recommended transfer, transfer with some specific identifiable risk, or uncertainty about risk. [Example]

[0125] The following examples are offered by way of illustration and not by way of limitation. Example 1 Embryo screening

[0126] An early version of the pipeline was developed to identify embryos that may have excessively high genetic risk (labeled "No go") and to stratify the remaining embryos based on their assessment of residual genetic risk: embryos with a low probability of disease were labeled "Go."

[0127] Simulations were performed using random "matings" of 14 men and 14 women, with 20 "embryos" per mating. This resulted in 196 possible parental combinations with 20 embryos each, generating 3,920 simulated embryos. For a typical couple in their early to mid-30s, we determined that the probability of implanting an embryo that passes all disease screening tests in the first cycle is approximately 96-98% (Table 1). [Table 1]

[0128] This pipeline was tested to analyze a cohort of children with rare diseases. Only about 30% of children with rare, early-onset diseases receive a genetic diagnosis, but using the disclosed methodology, about 80% of these children could have been successfully screened because they had rare deleterious variants in genes that are restricted for deleterious variants.

[0129] In conclusion, an early version of this analysis pipeline identified a 2.5-fold increased childhood risk compared with current best clinical diagnosis, with a 96-98% probability of implanting at least one "healthy" embryo. Example 2 Carrier testing

[0130] To fully characterize parental genetic risk, we generated a first-stage pipeline. We screened 900 parent genomes for loss-of-function, known pathogenic, or highly deleterious mutations (top 0.1%) in all known recessive disease genes and known pathogenic mutations in autosomal dominant genes. These genomes were then compared with the Invitae carrier panel (the "Invitae Comprehensive Carrier Screen," which contains 287 genes). As shown in Figure 1, this pipeline identified approximately six times more genetic risk than a conventional comprehensive panel. Furthermore, this pipeline identified not only autosomal recessive genes (rare clinical variants and LoFs), but also rare variants in autosomal dominant genes from the ClinVar database, as well as variants that were rare in the population and had a high combined annotation-dependent depletion score (CADD), indicating a high probability of being deleterious autosomal recessive genes. Identified genetic risks include CRTAP (osteogenesis imperfecta type VII) and PIEZO2 (distal arthrogryposis with proprioceptive and tactile impairment). Example 3 cell line research

[0131] First, we characterized the whole genome sequencing ("WGS") and whole genome amplification ("WGA") of cell lines (Figure 2). These cell lines were purchased from Corriell and have publicly available sequence data from the Genome in a Bottle (GIAB) Consortium (NA24385, NA24632, and NA18278). The goal of these studies is to evaluate WGA kits and define merging strategies to be applied in embryo screening.

[0132] Materials: Two distinct groups of samples were used. The first group consisted of 10 individual cells from three GIAB samples sequenced with 15x WGS, and the second group consisted of three "bulk" samples of 5-6 cells from each GIAB. Variant calling had previously been performed on each GIAB sample, generating a ground truth dataset of genetic variants that could be used to compare with results from this WGS approach.

[0133] As a first step, assess the coverage variability of the three GIABs by comparing the average coverage of appropriately large regions (10-100 Kbp) with the expected (average) coverage to determine the number of regions with higher than expected coverage, if any. For each single-cell sample, obtain a subset of data (e.g., 10x, 4x, 5x, 1x, 0.1x) for preprocessing. Repeat the same WGS for the "bulk" sample.

[0134] Then, variants in each cell can be identified using any available variant caller (such as platypus, GATK's Haplotype caller, Freebayes, Octopus, Delly, etc.), either alone or in combination with other genetic data. In addition, variants can be identified using a custom-trained artificial intelligence variant caller optimized for detecting CNVs and SNVs from single-cell data. This caller is trained by providing a training algorithm with a dataset (ground truth dataset) with known expected results, thereby penalizing or rewarding the machine learning algorithm based on performance. This algorithm is then used on a dataset that does not have a known ground truth, but similar performance is expected.

[0135] SNV calling: For every dataset from the first group of samples (i.e., 3 x 10 individual cells), variants were independently called using platypus with downsampling of 15x and 10x (60 Variant Call Format (VCF)-SET A). For each cell line, 10 cells were merged to 4x, 8 cells to 5x, and 4 cells to 10x (repeat with different cells) and variants were called using platypus (12 VCF-SET B). The key is to sequence everything at a total of 40x to determine whether cell number is important. For every dataset from the second group of samples (3 "bulk" samples), variants were called using platypus (3 VCF-SET C).

[0136] Next, 5, 7, and 10 random VCFs from SET A for each WGA are merged by two techniques: majority rule (i.e., seen 3, 4, or 6 times, respectively) and "seen more than once" rule (i.e., any variant seen in at least two VCFs). Merging results in 18 VCFs (SET D).

[0137] Calls are then compared between 12 VCFs (SET B), 3 VCFs (SET C), and 18 VCFs (SET D) by constructing a confusion matrix for each sample, which records the true positive, true negative, false positive, and false negative rates between the sample and the known gold standard dataset.

[0138] Aneuploidy: Aneuploidy is determined by acquiring low-coverage whole-genome data (0.1×) and comparing the expected coverage across large regions (10–100 Mbp) with the measured coverage. Regions with higher than expected coverage likely represent copy gains, which may involve entire chromosomes or portions of chromosomes, while regions with reduced coverage may represent copy losses of all or portions of chromosomes. Expected coverage can theoretically be determined based on sequencing depth, GC content, and other local genomic factors, and may also be based on empirical data from sequencing euploid samples.

[0139] SV calling: For every dataset from the first group of samples (i.e., 3 x 10 individual cells), call variants independently at 15x, 10x, 5x, 1x, and 0.1x (special callers may be required for 1x and 0.1x) (150 VCF-SET E). For each cell line, merge 10 cells to 4x, 8 cells to 5x, and 4 cells to 10x (repeat with different cells) and call variants using delly (12 VCF-SET F).

[0140] Next, for each WGA, 5, 7, and 10 random VCFs from SET A are merged by two techniques: majority rule and "seen more than once" rule. The merging results in 9 VCFs (SET G).

[0141] Use SURVIVOR to determine the number of false deletions / gains called by each VCF in SET E and how frequently they are consistent (i.e., non-random dropouts or over-amplifications).

[0142] Once de novo calls are generated from isolated cells, the results help inform WGA and merging strategies. Additionally, data generated from cell line experiments are combined with data from GIAB to form a trio used to train DeepVariant. Example 4 Research on donated embryos

[0143] The methodology for cell lines described in Example 3 above is extended to embryos (Figure 3). First, each embryo is biopsied, and one or more single cells are isolated from each embryo using physical or enzymatic methods. The starting source material can be multiple sources of embryonic material, such as embryos, blastocysts, cell culture media, etc.

[0144] The WGA kit and merging strategy defined by the cell line experiments in Example 3 are applied to embryos. These methods include steps such as very low-pass WGS (0.1x), using this low-pass WGS to exclude some types of aneuploidy, and methods for analyzing single cells versus merged cells. The sequencing step can be DNA (genome), RNA (transcriptome), or methylation site identification. Each has its own library preparation method. The goal of each sequencing output is to identify variants associated with disease or disease risk / protection.

[0145] The data obtained from WGA and WGS is then analyzed to determine the match between individual cells and embryos to assess reproducibility and consistency. Results from each set of approximately six individual cells are compared to results from the whole embryo. Results from each cluster are then compared to each other. The same software pipeline was used, but with streamlined data management. Additionally, the previously trained DeepVariant algorithm is applied to the embryo data for evaluation. Once analyzed, embryos are classified and ranked based on likelihood of full-term pregnancy and postnatal genetic risk, and a report is generated with recommendations regarding which embryos should be selected for transfer and implantation. Classification and report generation involves manual and automated steps (Figures 5-12). For example, classification of some variants can be fully automated because the variant has previously been reported as deleterious by multiple sources. In other cases, a weighting can be applied to a variant because it is suspected to be deleterious by various means but has not previously been reported as pathogenic, and may be present in a gene but does not necessarily cause disease. In some situations, for example, when multiple de novo mutations in a single gene are associated with an autosomal recessive disease, manual review may be necessary. Example 5 Genetic analysis and embryo ranking

[0146] The results from cell lines and embryos are further expanded into a proposed end-to-end plan for a minimum viable product (MVP) (Figure 4). First, each embryo is biopsied, and one or more single cells are isolated from each embryo using physical or enzymatic methods. Next, the WGA kit and merging strategy defined by the cell line experiments in Example 3 and finalized during the embryo experiments provided in Example 4 are applied to this study. Standard WGS methods are used for parental WGS.

[0147] Develop a variant calling method for structural variant (SV) and single nucleotide variant (SNV) calling that utilizes parental genetic information. For example, SNV calling for germ cells includes haplotype phasing, missing data imputation, and de novo calling, all of which utilize parental genetic information. The current approach involves exploring modifications of DeepVariant to accommodate these steps, as described in Example 3.

[0148] Once analyzed, embryos are classified and ranked based on the likelihood of full-term pregnancy and postnatal genetic risk, and a report is generated with recommendations regarding which embryos should be selected for transfer and implantation. Classification and reporting involves manual and automated steps (Figures 5-12). Specifically, variants can be classified into broad categories, including known pathogenic variants, non-coding variants, coding variants, loss-of-function (LoF) variants, synonymous variants, and variants not worthy of further consideration (discard), as well as "panic" variants that require manual review (Figure 5). For SNVs or small insertions / deletions (indels) less than 50 bp, the workflow can include assessing the characteristics of the variants according to the following decision points: blacklisted variants; whitelisted variants (i.e., known pathogenicity, potentially requiring special consideration, and potentially resulting in panic), ClinVar variants; protein-coding and non-synonymous; and intronic and synonymous. As categorization progresses, ClinVar decision points may further include considering whether the variant is benign or likely benign (B or LB) and pathogenic or likely pathogenic (P or LP). For pathogenic P or LP, the ClinVar Path, discussed in more detail below with respect to Figure 7, is followed. Protein-coding and non-synonymous decision points may further include decision points that consider whether the variant has a start codon loss, a stop codon gain, and a splice + / -2, a deeper splice + / -10 > 0.5 splice AI loss, or a frameshift to arrive at whether the variant is LoF or non-LoF. If the variant is not protein-coding and non-synonymous, the decision process may proceed to a decision point that further considers whether the variant is intronic and synonymous. If the variant is intronic and synonymous, it is classified as such; if the variant is not intronic and synonymous, it is discarded.

[0149] Regarding heredity, LoF codes, SNV variants, broad categories such as "No go" where embryos are not considered for transfer, and the weighting of variants to assist in embryo ranking can be used (Figure 6). The workflow may include categorizing variants by the following decision points: minor allele frequency (MAF) < X%; variants corresponding to American College of Medical Genetics (ACMG) criteria (e.g., ACMG 73); variants described as Online Mendelian Inheritance in Man (OMIM) morbid; and high constraint (pli > XX). For the MAF < X% decision point, in this categorization, variants with a minor allele frequency less than a certain percentage (X%) are further evaluated, and one can panic or continue categorization. The ACMG decision point may further include examining LoF known mechanisms and then weighting in that case. The OMIM morbid decision point may further include examining LoF known mechanisms, X-linked, autosomal recessive, autosomal dominant, penetrance / severity / family history (FHx) of the variant. If the variant is X-linked, the categorization may further include examining X-linked dominant (XLD) and whether it is male (in which case it is No go). If it is XLD, the penetrance, severity, or FHx of the variant is examined. In the weighting at this stage, one can consider that a female with XLR inherited from the mother may have reached this variant or that XLD with variable penetrance. If the embryo is male, the workflow may further include examining whether the variant was inherited from the father, and if so, whether it is XLR inherited from the father or XLD that should always cause the disease but is the result of the parent being alive, and thus panic. Similarly, during the decision process, when reaching high constraint, there may be a possibility of LoF inherited from the parent rather than OMIM morbid, which is somewhat special.

[0150] For known pathogenic variants (e.g., based on the ClinVar database), broad categories such as "No go," "weighting," and "under-go Manual Review" can be used (Figure 7). Categorization can include assessing the characteristics of the variant with X-linked (XL), autosomal recessive (AR), and autosomal dominant (AD) call points. If the variant is XL, categorization can further include considering whether it is XLR and whether it is in females. If so, weight it; if not, go No go. If the variant is AD, categorization can further include considering variable penetrance or severity, family history (FHx or Fam Hx), and whether the variant is a de novo variant potentially reaching No go, weighting, or manual review. If the variant is AR, the workflow can proceed to the AR workflow, described in more detail below with respect to Figure 8.

[0151] Referring to Figure 8, for a variant to reach this workflow, it must have already been determined to be "pathogenic" in the AR gene.For these pathogenic variants in genes related to autosomal recessive (AR) disease, determine whether multiple variants exist in the gene, and whether these variants are cis or trans, or whether the variants exist in the gene in a homozygous or hemizygous state.Embryos can be broadly categorized as "No go" to be weighted or undergo manual review.More specifically, the decision point of whether a variant is homozygous or hemizygous can further include determining whether the variant is homozygous or hemizygous in both parents, and then considering whether it is trans (check whether the multiple variants in a gene all come from one parent ("cis") or from different parents ("trans").This assumes that almost all recessive genes are fully penetrant, so if the variant does not exist in surviving parents, the variant is likely non-pathogenic. If the variant exists in both parents, or if it is a variant on the X chromosome of a boy, the variant is No go, and is homozygous inheritance from both parents (consanguineous), or hemizygous in a boy.At the decision point of further considering whether the variant exists on the X chromosome in a boy, if the answer is negative, manual review is required.Such a variant indicates non-Mendelian inheritance, and may be overlapping with deletion, or may be missed in parents or allele dropout in children.The decision point of whether there is more than one variant in a gene can further include considering whether there are more than zero variants from the father, and if applicable, whether there are more than zero variants from the mother.If so, it is No go.This is the most common type of autosomal recessive.If there are not more than zero variants from the father, categorization can further include considering whether there are more than zero variants from the mother.If there are no more than 0 maternal variants, consider whether there are more than 0 variants of unknown inheritance. Getting here indicates there is one inherited variant and one variant of unknown inheritance (likely de novo). If the variant cannot be phased, the variant is likely a no go, but this is a rare event.

[0152] Figure 9 illustrates the classification of de novo LoF coding variants (SNVs) into the broad categories previously shown. More specifically, categorization can include the following decision points: ACMG; OMIM morbid; and highly constrained.

[0153] Referring now to Figures 10A-D, categorization may include the following call points for SV deletions: de novo and >XX MBp, and overlapping the coding region of a gene (Figure 10A). For SV gains in Figure 10B, categorization may include the call points of being at the beginning or end of the coding region of a gene and being triplosensitive. For inversion calling in Figure 10C, categorization may include the call points of being de novo and at the beginning or end of the coding region of a gene. Finally, for translocation calling in Figure 10D, categorization may include the call points of being de novo and at the beginning or end of the coding region of a gene.

[0154] Figure 11 illustrates repeat expansion categorization based on repeat expansions of known pathogenic and "premutation" length as well as repeat length in the parents to broadly categorize embryos as No go or for manual review.

[0155] Figures 12A-D illustrate the classification of large regions of homozygosity, uniparental isodisomy, synonymous and intronic variants, and variants identified from the spinal muscular atrophy pipeline. More specifically, in Figure 12A, large region variants of homozygosity can be classified as Treat as normal (recessive disease), Treat as normal (deletion), and Manual Review. In Figure 12B, synonymous and intronic variants can be classified as Treat as LoF. In Figure 12C, large regions of uniparental inheritance can be classified as No go or Manual Review. Finally, Figure 12D is a flowchart illustrating potential categorization of variants identified, for example, from the spinal muscular atrophy pipeline. Example 6 Genetic analysis and embryo ranking

[0156] Figures 13-14C illustrate categorization of SNVs or small indel variants using DeepVarian, VCF information, population frequency, reference sequence, phenotyping, and trio genetic steps to classify SNVs or small indel variants into various broad categories.

[0157] More specifically, referring to Figure 13, the first steps of DeepVariant, VCF information, and population frequency are detailed. This analysis can lead to broad categories of discard or later extension. Note that population frequency analysis can include querying the 1KG and Gnomad databases, for example.

[0158] Referring now to FIG. 14A, in the reference sequence step, categorization may proceed through the decision points of coding variant, intronic variant, intergenic variant, variant is nonsense or LoF, variant is missense, variant is synonymous, variant affecting splicing (X out of 4), variant affecting splicing (clinical X out of 4), and evidence of pathogenicity (ClinVar), potentially leading to broad categories of panic, discard, or later expansion.

[0159] Referring to Figure 14B, during the phenotyping step, categorization can proceed through decision points of Met ACMG Criteria, Met ClinVar Pathogenicity Criteria, Met OMNIM Criteria, and Met Embryo "Intolerome" Criteria Met, potentially leading to broad categories of Discard or later Expansion.

[0160] Referring to Figure 14C, in the trio inheritance step, categorization may proceed through the decision points of De Novo dominant, inherited dominant, homozygous recessive, X-linked, inherited compound heterozygosity, and manual review capture net, potentially leading to the broader category of panic or later extension. Note that De Novo dominant, X-linked, inherited dominant, homozygous recessive, and inherited compound heterozygosity each have sub-flowcharts detailed in Figures 16A-E.

[0161] 15, the categorizations of the various pipelines can be aggregated in the report / occlusion subset step of the process. In this example, a category may be arrived at from the SNV annotation pipeline and progress through the decision points of put on the occlusion list and put on the report list, leading to broader categories of discard, manual review, or later expansion before the report is finalized.

[0162] Referring to Figure 16A, in this sub-flowchart for the De Novo Dominant decision point, the analysis can proceed through the decision points of De Novo (not from parent), parent, dominant mechanism, and recessive mechanism, potentially leading to the broad category of panic or other pipeline reference.

[0163] Referring to Figure 16B, in this sub-flowchart for the X-linked decision point, the analysis can proceed through the decision points of non-PAR region, proband male and hemizygous, and proband female and homozygous, potentially leading to the broad categories of manual review or other pipeline reference.

[0164] Referring to Figure 16C, in this sub-flowchart for the genetic dominance decision point, the analysis can proceed through the decision points of De Novo (not from parents), parental, (P1)-hetero / (P2)-WT / (E)-hetero, rare (P1 / 2) and (E)-hetero, dominant mechanism, and recessive mechanism, potentially leading to the broad category of panic or other pipeline reference.

[0165] Referring to Figure 16D, in this sub-flowchart for the homozygous recessive decision point, the analysis can proceed through the decision points of De Novo (not from parents), Parental, (P1)-hetero / (P2)-hetero / (E)-homo, (P1)-hetero / (P2)-WT / (E)-homo, Rare (P1 / 2) and (E)-homo, Dominant mechanism, and Recessive mechanism, potentially leading to the broad categories of Manual Review, PANIC, or other pipeline reference.

[0166] Referring to Figure 16E, in this sub-flowchart illustrating inherited compound heterozygosity, the analysis can proceed through decision points for genes with two or more SNVs, phasing locations of SNVs on separate chromosomes, phasing unknown SNVs, SNVs phased to LoF / LoF, SNVs phased to LoF / other pathogenic, LoF / LoF SNVs, Lof / other pathogenic SNVs, potentially leading to the broad categories of discard, manual review, and panic. Example 7 Aneuploidy calling

[0167] Figures 17A and 17B show examples of aneuploidy calling for sets of individual cells from two different embryos. Analysis reveals both distinct and consistent signal differences between cells. All cells in each figure were from the same embryo. Cells labeled "B" were isolated from biopsies, cells labeled "M" were isolated from the inner cell mass, and samples labeled "R" were bulk samples of the remaining cells from each embryo (to demonstrate consistency). Copy number calls were performed using CNVkit software based on a custom reference set composed of single-cell sequences amplified using the same WGA method. Figure 17A shows a female embryo that was aneuploid, demonstrating consistent loss of chromosome 19. Figure 17B shows a female embryo that was euploid, showing several copy number variants in individual cells, but none that were consistently present throughout the embryo.

[0168] Discussion: The genetic analysis and embryo screening methods described above and disclosed herein involve using comprehensive genomic data and analysis, along with the most sensitive screening criteria, to identify embryos with the highest likelihood of full-term pregnancy and the lowest postnatal genetic risks across all 6 billion bases of the human genome. The disclosed analysis is more thorough because it considers constrained genes not currently considered disease-causing, small CNVs below the detection limit of most aneuploidy screens, structural variants, and intergenic and intronic mutations. Rather than pooling cells first, multiple single cells are individually sequenced and then combined, reducing areas of low or absent coverage. Furthermore, genetic information from the embryo is compared with genetic information from one or more parents and / or one or more genetic relatives. This allows for distinguishing between sequencing artifacts and actual mutations, as any variants in multiple cells are likely to be actual mutations.

[0169] Some key challenges associated with whole genome sequencing using single cells include genome-wide coverage, inter-reaction reproducibility, and accuracy measurement. These challenges are effectively resolved by the method of the present disclosure. First, by obtaining embryo data using multiple sources (cells and / or cell-free sources) and then merging the data, the method of the present disclosure can achieve better genome-wide coverage, thus improving accuracy through depth and comparison. In some embodiments, sequencing data, for example, sequencing data from separate cells, is used to call variants, and the results are then integrated for analysis. In some embodiments, sequencing data, for example, sequencing data from separate cells, is first integrated and then used to call variants. Second, the method of the present disclosure can offset the biases of each method by using multiple WGA methods. Third, the method of the present disclosure improves the accuracy of variant calling by identifying de novo variants using data from multiple embryo sources and parent / family genomic data.

[0170] Genetic analysis of multiple embryos during in vitro fertilization (IVF) requires a method for comparing the disease risk of each embryo to select the embryo for transfer and implantation. Each embryo has distinct, even if overlapping, genetic makeup and therefore a different disease risk. The disclosed method involves automated classification of variants and subsequent embryo ranking, which has not been previously performed or suggested. Preimplantation allows for easy addressing of all genetic variants, meaning parents can choose not to transfer an embryo due to genetic susceptibility to disease. This is particularly true for adult-onset diseases. For example, pathogenic BRCA2 mutations can lead to breast cancer within 30 to 50 years, but treatments exist and are likely to improve in the future. This risk may need to be assessed in the presence or absence of childhood disease. Therefore, genetic information from the entire genome must be taken into account. Automation allows for rapid and cost-effective embryo ranking. The ranking process and report generation summarizes the whole genome analysis and empowers parents to select embryos that meet their desired goals after being informed by the analysis results provided by the methods described above and disclosed herein.

[0171] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated herein by reference. Furthermore, citation or identification of any reference in this application should not be construed as an admission that such reference is available as prior art to the present invention. To the extent section headings are used, they should not be construed as necessarily limiting.

[0172] Although the present invention and its advantages have been described in detail, it should be understood that various changes, substitutions, and alterations can be made therein without departing from the spirit and scope of the invention as defined in the appended claims.

[0173] It will be readily apparent to those skilled in the art that the present invention is well adapted to carry out the objects and obtain the ends and advantages mentioned above, as well as the advantages inherent therein. The examples and methods described herein are presently representative of preferred embodiments and are exemplary and are not intended to limit the scope of the invention. Modifications and other uses of the examples and methods described herein will occur to those skilled in the art and are encompassed within the spirit of the invention as defined by the appended claims.

Claims

1. 1. A method for identifying genetic variants in an embryo, comprising: (a) obtaining two or more sources of analytes from said embryo; (b) analyzing the two or more sources of analyte to obtain genetic information for each source; (c) comparing the genetic information of each source against one or more reference genomes using at least one variant caller, wherein the variant caller identifies variants between each source and the reference genome; (d) combining the variants to identify differences present only in the source, wherein the differences are genetic variants in the embryo; A method comprising:

2. 10. The method of claim 1, wherein the genetic variant is a single nucleotide variant (SNV), a multinucleotide variant (MNV), a copy number variant (CNV), a structural variant (SV), or an alteration in an epigenetic marker.

3. The method of claim 2, wherein the epigenetic marker is a DNA methylation marker for identifying defects or transcriptional dysregulation.

4. 3. The method of claim 2, wherein the genetic variant is identified as an apparent de novo variant (DNV).

5. 10. The method of claim 1, wherein the two or more sources of analytes comprise cells biopsied from the embryo or blastocyst, cell culture medium, cells and / or genetic material isolated from the blastocoel, or a combination thereof.

6. 6. The method of claim 5, wherein the biopsy is taken using laser pulse assisted pipette delivery.

7. 7. The method of claim 6, further comprising isolating two or more cells from the biopsy and dissociating the two or more cells using physical or enzymatic methods.

8. 8. The method of claim 7, wherein the physical method comprises laser dissection or micropipette isolation.

9. 8. The method of claim 7, wherein the enzymatic method comprises dissociation using a digestive enzyme, wherein the enzyme is selected from the group consisting of dispase, collagenase, hyaluronidase, papain, DNase-I, accutase, and trypsin.

10. 10. The method of claim 1, wherein the genetic information comprises nucleic acid sequences of a panel, exome, whole genome, or transcriptome, with or without mitochondrial sequences, with or without epigenetic markers.

11. The method of claim 10 , wherein the nucleic acid comprises DNA, RNA, or both.

12. wherein the nucleic acid is DNA and the DNA sequence is synthesized by the steps of: (a) generating a library of DNA from each source; (b) performing low-pass whole genome sequencing of the DNA; (c) applying an aneuploidy filter to the sequencing data from (b) to filter out embryos with aneuploidy; and (d) performing additional sequencing of the DNA from embryos that are not excluded to generate a sequence data file for the embryos. The method of claim 11 obtained by

13. 13. The method of claim 12, wherein application of the aneuploidy filter results in the detection of chromosomal regions with unexpectedly increased or decreased coverage compared to expected coverage based on total sequence coverage, known biases in coverage based on local genomic composition, and the sex of the embryo.

14. 12. The method of claim 11, wherein RNA is isolated, amplified and sequenced.

15. 15. The method of claim 14, wherein RNA-seq data is used to characterize the transcriptome and determine gene expression levels.

16. 16. The method of claim 15, wherein gene expression levels correlate with patient outcome.

17. 17. The method of claim 16, wherein the patient outcome comprises a successful pregnancy.

18. 17. The method of claim 16, wherein genetic information and / or gene expression correlates with patient outcome.

19. 15. The method of claim 14, wherein RNA-seq data is used in combination with or alone with DNA sequencing to call genomic variants.

20. 15. The method of claim 14, wherein RNA-seq data is used to identify aberrant splicing events and dysregulation of transcriptional regulation and one or more variants affecting one or more transcriptional pathways.

21. 21. The method of claim 20, wherein the aberrant splicing event is cryptic splicing, aberrant exon usage, or an upstream open reading frame.

22. 15. The method of claim 14, wherein RNA-seq data is used to identify biased allelic expression, thereby identifying imprinting, copy number determination, and / or variants that affect expression.

23. 10. The method of claim 1, wherein the reference genome is a genomic sequence from a public or private repository, one or more parents, or one or more genetic relatives.

24. 24. The method of claim 23, wherein the genetic variants are further identified by comparing the genetic information of each source to genetic information of one or more parents to estimate haplotype phasing, impute missing data, determine inheritance or lack thereof, and identify apparent de novo variants.

25. 25. The method of claim 24, further comprising annotating the genetic variants for association with a phenotype by using a database of variants associated with diseases, phenotypes, disease risk or protection and traits.

26. 25. The method of claim 24, further comprising annotating the genetic variants for their potential to cause deleterious effects on gene function by using gene function or biological pathway databases or via in silico, in vivo, or in vitro methods.

27. 10. The method of claim 1, wherein the variant caller is a deep learning based variant caller.

28. 28. The method of claim 27, wherein the deep learning-based variant caller is built based on genetic information of one or more parents or genetic relatives and is capable of identifying apparent de novo variants for the purpose of preimplantation genetic screening.

29. 10. The method of claim 1, wherein the genetic variant is present in a genome with a known or suspected contribution to adult or childhood genetic disease, or a genome with no currently associated genetic disease but under genetic constraint.

30. 10. The method of claim 1, wherein the combining step comprises identifying genetic variants present in at least two sources or identifying genetic variants that can only be detected by combining genetic information from all sources.

31. 31. The method of claim 30, wherein the genetic variants are identified in multiple sources.

32. 31. The method of claim 30, wherein the genetic variant is identified as mosaic.

33. 1. A method of screening a plurality of embryos, comprising: (a) performing a genetic analysis on each embryo according to any of the preceding claims; (b) ranking the embryos based on the results of the genetic analysis, wherein embryos with the highest likelihood of full-term pregnancy and the lowest lifetime genetic risk are ranked highest for transfer and implantation; A method comprising:

34. 34. The method of claim 33, wherein embryos are ranked based on known deleterious or benign variants, or based on variants predicted to be deleterious or benign by various methods.

35. 35. The method of claim 34, comprising in silico deleterious prediction tools, allele evolutionary conservation, and / or allele frequencies in selected or unselected population databases.

36. 34. The method of claim 33, wherein the influence of genetic variants on embryo rank is weighted based on input from customer-based phenotyping.

37. 37. The method of claim 36, wherein the input is a desire to avoid a particular pathological condition.

38. 34. The method of claim 33, wherein the influence of genetic variants on embryo rank is weighted based on automated phenotypic overlap between disease-associated phenotypes and their associated genes and customer-provided phenotypes.

39. 39. The method of claim 38, wherein the phenotype is derived from the family history of one or both parents.

40. 39. The method of claim 38, wherein the overlap is computed using an ontology that allows for weighting associated phenotypes relative to embryo rank.

41. 34. The method of claim 33, wherein the effect of a genetic variant on embryo rank is weighted based on known disease severity, penetrance, and / or inheritance patterns associated with the gene.

42. 34. The method of claim 33, wherein the effect of a genetic variant on embryo rank is weighted based on the presence or absence of the variant in one or both parents and the zygosity in the parents.

43. 34. The method of claim 33, wherein the effect of genetic variants on embryonic order is weighted based on the age of onset of known or suspected disease.

44. 34. The method of claim 33, wherein the influence of genetic variants on embryo rank is weighted based on previously curated gene or variant rankings.

45. 34. The method of claim 33, wherein the effect of genetic variants on embryo rank is weighted based on treatment options for a given disease.

46. 34. The method of claim 33, further comprising generating a report regarding the plurality of embryos, wherein the ranked embryos are qualitatively categorized based on recommended for transfer, recommended not for transfer, or unknown for transfer.