Method for identifying genetic variant in an embryo
The method for identifying genetic variants in embryos through multiple analyte sources and advanced sequencing techniques addresses the limitations of current testing, enhancing IVF success by accurately screening for genetic risks and improving pregnancy outcomes.
Patent Information
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- EMBRYOME INC
- Filing Date
- 2026-05-20
- Publication Date
- 2026-07-17
AI Technical Summary
Current prenatal and preimplantation genetic testing methods fail to accurately detect rare and high-impact genetic variants in embryos, leading to potential miscarriages and birth defects, and lack comprehensive screening before embryo transfer.
A method for identifying genetic variants in embryos by obtaining multiple analyte sources, analyzing genetic information, comparing with reference genomes, and combining variants to detect differences, using techniques like whole-genome sequencing, RNA sequencing, and deep learning-based variant calling.
Accurately identifies genetic variants in embryos, enabling ranking of embryos for optimal pregnancy outcomes and reducing the risk of genetic diseases, thereby improving IVF success rates and reducing miscarriages.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202480021727.X (22) Application Date 2024.01.26 (30) Priority Data 63 / 441,291 2023.01.26 US (85) PCT International Application Entering National Phase Date 2025.09.24 (86) PCT International Application Application Data PCT / IB2024 / 050754 2024.01.26 (87) PCT International Application Publication Data WO2024 / 157222 EN 2024.08.02 (71) Applicant: Inbromy Company Address: Ontario, Canada (72) Inventors: E. Wenner, M. Bainbridge, J. Gorushau (74) Patent Agency: King & Wood Mallesons, Beijing 11256 Patent Attorney Chen Wenping (51) Int.Cl. G16B 20 / 20 (2019.01) C12Q 1 / 68 (2018.01) C12Q 1 / 6809 (2018.01) G16B 30 / 00 (2019.01) G16B 40 / 00 (2019.01) G16H 50 / 30 (2018.01) (54) Invention Title Method for Identifying Genetic Variations in Embryos (57) Abstract This disclosure provides in part a method for identifying genetic variations in embryos. The method includes: (a) obtaining two or more analyte sources from the embryo; (b) analyzing the two or more analyte sources to obtain genetic information from each source; (c) comparing the genetic information from each source with one or more reference genomes using at least one variant invocation procedure, wherein the variant invocation procedure identifies variants between each source and the reference genome; and (d) combining the variants to identify differences presented only by the source, wherein the differences are genetic variants in the embryo. Claims 3 pages, Description 18 pages, Drawings 28 pages, CN 121100382 A 2025.12.09 CN 1 21 10 03 82 A 1. A method for identifying genetic variants in an embryo, the method comprising: (a) obtaining two or more analyte sources from the embryo; (b) analyzing the two or more analyte sources to obtain genetic information of each source; (c) comparing the genetic information of each source with one or more reference genomes using at least one variant calling procedure, wherein the variant calling procedure identifies variants between each source and the reference genome; and (d) combining the variants to identify differences presented only by the sources, wherein the differences are genetic variants in the embryo.2. The method of claim 1, wherein the genetic variant is a single nucleotide variant (SNV), a multinucleotide variant (MNV), a copy number variant (CNV), a structural variant (SV), or an alteration of an epigenetic marker. 3. The method of claim 2, wherein the epigenetic marker is a DNA methylation marker used to identify defects or transcriptional disorders. 4. The method of claim 2, wherein the genetic variant is identified as a distinct de novo variant (DNV). 5. The method of claim 1, wherein the two or more analyte sources comprise cells from embryonic or blastocyst tissue biopsies, cells and / or genetic material isolated from cell culture media, blastocyst cavities, or combinations thereof. 6. The method of claim 5, wherein the tissue biopsy is acquired using laser pulse-assisted pipetting delivery. 7. The method of claim 6, further comprising isolating two or more cells from the tissue biopsy and dissociating the two or more cells using physical or enzymatic methods. 8. The method of claim 7, wherein the physical method comprises laser cutting or micropipette separation. 9. The method of claim 7, wherein the enzymatic method comprises dissociation using a digestive enzyme selected from: dispersase, collagenase, hyaluronidase, papain, DNase-I, Accutase, and trypsin. 10. The method of claim 1, wherein the genetic information comprises nucleic acid sequences of a subset, exome, whole genome, or transcriptome, having or not having mitochondrial sequences and having or not having epigenetic markers. 11. The method of claim 10, wherein the nucleic acid comprises DNA, RNA, or both. 12. The method of claim 11, wherein the nucleic acid is DNA, and wherein the DNA sequence is obtained by the following steps: (a) generating a DNA library from each source; (b) performing low-pass whole-genome sequencing on the DNA; (c) applying an aneuploidy filter to the sequencing data from (b) to eliminate embryos with aneuploidy; and (d) performing additional sequencing on the DNA from the un-eliminated embryos to generate a sequence data file of the embryos. 13. The method of claim 12, wherein the application of the aneuploidy filter results in the detection of chromosomal regions with unexpectedly increased or decreased coverage compared to expected coverage based on total sequence coverage, known coverage bias based on local genomic composition, and the sex of the embryo. 14. The method of claim 11, wherein RNA is isolated, amplified, and sequenced. 15. The method of claim 14, wherein RNA sequence data is used to characterize the transcriptome and determine gene expression levels. 16. The method of claim 15, wherein gene expression levels are correlated with patient outcomes.17. The method of claim 16, wherein the patient outcome includes a successful pregnancy. 18. The method of claim 16, wherein genetic information and / or gene expression are associated with the patient outcome. Claims 1 / 3 Page 2 CN 121100382 A 19. The method of claim 14, wherein RNA sequence data is used to invoke genomic variants in conjunction with DNA sequencing or alone. 20. The method of claim 14, wherein RNA sequence data is used to identify aberrant splicing events and transcriptional dysregulation, and one or more variants affecting one or more transcriptional pathways. 21. The method of claim 20, wherein the aberrant splicing event is cryptic splicing, aberrant exon use, or upstream open reading frames. 22. The method of claim 14, wherein RNA sequence data is used to identify biased allele expression, thereby identifying imprinting, copy number determination, and / or variants affecting expression. 23. The method of claim 1, wherein the reference genome is a genomic sequence from a public or private repository, one or more parents, or one or more genetic relatives. 24. The method of claim 23, wherein the genetic variant is further identified by comparing the genetic information from each source with the genetic information of one or more parents to infer haplotype phasing, imputation of deletion data, determination of inheritance or deletion thereof, and identification of significant de novo variants. 25. The method of claim 24, further comprising annotating the correlation between the genetic variant and the phenotype using a database of variants associated with disease, phenotype, disease risk or protection, and traits. 26. The method of claim 24, further comprising annotating the possibility that the genetic variant causes an adverse effect on gene function by using a database of gene function or biological pathways or by computer, in vivo, or in vitro methods. 27. The method of claim 1, wherein the variant calling procedure is a deep learning-based variant calling procedure. 28. The method of claim 27, wherein the deep learning-based variant calling procedure is constructed based on the genetic information of one or more parents or genetic relatives and is capable of identifying significant de novo variants for preimplantation genetic screening purposes. 29. The method of claim 1, wherein the genetic variant occurs in a genome known or suspected of contributing to genetic diseases in adults or children, or in a genome without currently associated genetic diseases but subject to genetic constraints. 30. The method of claim 1, wherein the combining step comprises identifying a genetic variant present in at least two sources, or identifying a genetic variant that can only be detected by combining genetic information from all sources. 31. The method of claim 30, wherein the genetic variant is identified in multiple sources.32. The method of claim 30, wherein the genetic variant is identified as a chimera. 33. A method for screening multiple embryos, the method comprising: (a) performing a genetic analysis on each embryo according to any of the preceding claims; and (b) ranking the embryos based on the results of the genetic analysis, wherein embryos having the best chance of full-term pregnancy and the lowest lifetime genetic risk are ranked highest in terms of transfer and implantation. 34. The method of claim 33, wherein the embryos are ranked based on known harmful or benign variants, or based on variants predicted to be harmful or benign according to various methods. 35. The method of claim 34, wherein the method comprises a computer-aided harmful prediction tool, allele evolutionary conservation, and / or allele frequencies in a selected or unselected population database. 36. The method of claim 33, wherein the influence of genetic variants on embryo ranking is derived from a weighted average based on client phenotypic input. 37. The method of claim 36, wherein the input is a desired avoidance of a specific pathological condition. Claims 2 / 3 Page 3 CN 121100382 A 38. The method of claim 33, wherein the effect of genetic variants on embryo sequencing is weighted based on automatic phenotypic overlap between disease-related phenotypes and their associated genes and phenotypes provided by the client. 39. The method of claim 38, wherein the phenotype originates from the family history of one or two parents. 40. The method of claim 38, wherein the overlap is calculated using ontology, which allows related phenotypes to influence embryo sequencing. 41. The method of claim 33, wherein the effect of genetic variants on embryo sequencing is weighted based on known disease severity, penetrance, and / or gene-related inheritance patterns. 42. The method of claim 33, wherein the effect of genetic variants on embryo sequencing is weighted based on the presence or absence of variants in one or two parents and conjugation in said parents. 43. The method of claim 33, wherein the effect of genetic variants on embryo sequencing is weighted based on known or suspected age of onset. 44. The method of claim 33, wherein the effect of genetic variants on embryo sequencing is weighted based on pre-assessed gene or variant sequencing. 45. The method of claim 33, wherein the effect of genetic variants on embryo sequencing is weighted based on treatment options for a given disease. 46. The method of claim 33, further comprising a report of the generation of said plurality of embryos, wherein the sequenced embryos are qualitatively classified based on recommendations for transplantation, non-transplantation, or uncertain transplantation. Claims 3 / 3 Page 4 CN 121100382 A Method for Identifying Genetic Variants in Embryos Cross-Reference to Related Applications
[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 441,291, filed January 26, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure generally relates to the field of reproduction. More specifically, this disclosure relates to methods for identifying genetic variants in embryos to assess disease risk. Background Art
[0003] Currently available prenatal testing, carrier screening, and preimplantation genetic testing (“PGT”) typically focus on detecting aneuploidy and known genetic conditions in parents who are carriers, but cannot detect most pathogenic genetic variants in a single embryo. Even if a family can rule out one or more known genetic risks, the fertilized egg may still carry a potentially fatal genetic defect.
[0004] Natural conception is fraught with inherent developmental challenges. Nearly half of fertilized eggs carry potentially fatal genetic mutations that may occur during pregnancy, childhood, or adulthood. Therefore, approximately one-third of pregnancies end in miscarriage (Wilcox AJ et al., N. Engl J Med, 1988; 319: 189-194), approximately one-fourteenth of children are born with genetic defects (Ceyhan-Birsoy O et al., Am J. Human Genet, 2019; 104: 76-93), and approximately one-twentieth of adults carry variants of inherited diseases that significantly increase the risk of many cancers, sudden cardiac death, aneurysms, autoimmune diseases, and neurodegenerative diseases such as Huntington's disease (Abul-Husn NS et al., Science, 2016; 354: 6319).
[0005] Rare and high-impact variants, rather than polygenic risks, drive clinical decisions. While this is a good way to understand the population, polygenic risks are a poor way to understand the success of development in an individual or a single fertilized egg. For example, for a common disease like breast cancer, there is a ~15% lifetime risk. Women in the top 1% of the polygenic risk score (“PRS”) ranking have a 31% lifetime risk of developing breast cancer, which is twice that of normal. However, rare BRCA1 / 2 mutation carriers have a lifetime risk of developing breast cancer as high as 85% (six times that of normal). For very rare diseases, such as brain cancer, there is about a 0.01% lifetime risk. Common variants result in a 1.2–3.5 times lifetime risk (Melin et al., 2017, Nat. Genet. 49(5): 789–794), while rare variants can result in a 5,000 times lifetime risk (Bainbridge et al., 2015, J. Natl. Cancer Inst. 107(1): 1–4).
[0006] Most clinical tests, such as carrier screening and preimplantation genetic testing (including PGT monogenic disease (PGT-M) andPGT-A aneuploidy uses targeted approaches to balance sensitivity and specificity to avoid missing the disease (i.e., avoiding "false negatives") while avoiding "false positive" panic or discarding potentially viable embryos. PGT-A accounts for approximately [percentage missing] of IVF cycle failures and miscarriages (Zhao C et al., 2021, Gen in Med. 23:435-442). Although next-generation sequencing (NGS) of PGT can diagnose whole-chromosome aneuploidy (WCA) in 95% of embryos, it cannot accurately detect subchromosome events (Cascante SD et al., 2023, Fert. Ster. 120(6):1161-1169). Carrier tests report 200-800 genes, while PGT-M may only report 1 or 2 genes. Prenatal testing, such as non-invasive prenatal testing ("NIPT"), focuses on the most definitive genetic errors that are most easily and accurately detected. Most NIPT assays report 5-13 cases. However, it is worth noting that, based on prenatal testing, the only action that can be taken is to maintain or terminate the pregnancy. Preimplantation genetic testing is the most actionable time; for example, BRCA2 mutation carriers have a 70% lifetime risk of developing breast cancer. Methods to reduce risk depend on the stage of life. Diagnosis can reduce risk through enhanced monitoring or prophylactic surgery. Prenatal diagnosis is not currently available because termination of pregnancy is the only potentially effective way to mitigate the risk of breast cancer. In contrast, embryo screening poses no risk if the parents choose to transfer different embryos. Whole-exome sequencing (WES) allows for comprehensive screening of the protein-coding regions (exomes) of the embryonic genome. Although the exome only comprises a small portion (~2%) of the entire genome, it contains many known pathogens. Recent studies using exome sequencing have shown that 60% to 75% of sporadic cases across all tests can be explained by de novo mutations, i.e., variations not found in either parent (Acuna-Hidalgo R et al., 2016, Genome Biology, 17:241). Typically, each embryo formed after fertilization contains approximately 100 de novo genetic variations. The older the parents, the greater the tendency for these variations to accumulate in sperm and eggs, and these variations can be passed on to ligands. Interestingly, triple testing (parent-offspring triple) has been found to increase the probability of successful genomic diagnosis of rare pediatric diseases by nearly five times (Wright CF et al., 2023, N Engl. J. Med. 388:1559-1571). The same study found that among 3599 individuals who underwent triple testing for diagnosis, approximately 76% had pathogenic de novo variants (i.e., the same parent-offspring triple variant).(Above). Therefore, there is a need for accurate embryo screening methods, including analyzing more pregnancy success determinants before embryo transfer and implantation, to accurately assess genetic health and embryonic outcomes through development and beyond. Summary of the Invention
[0007] This disclosure provides in part a method for genetic screening of embryos, particularly before embryo transfer and implantation, wherein a single source of an analyte is analyzed and compared with a reference genome to identify genetic variants in the embryo.
[0008] Accordingly, one aspect of this disclosure provides a method for identifying genetic variants in an embryo. The method involves obtaining two or more sources of an analyte from an embryo; analyzing the two or more sources of the analyte to obtain genetic information for each source; comparing the genetic information for each source with one or more reference genomes using at least one variant caller, wherein the variant caller identifies variants between each source and the reference genome; and combining the variants to identify differences presented only by the source, thereby identifying genetic variants in the embryo.
[0009] In some embodiments, the genetic variant is a single nucleotide variant (SNV), a multinucleotide variant (MNV), a copy number variant (CNV), a structural variant (SV), or an alteration of an epigenetic marker. As a non-limiting example, an epigenetic marker is a DNA methylation marker used to identify imprinting defects. As a non-limiting example, a genetic variant is identified as a distinct neovariant (DNV).
[0010] In some embodiments, the sources of two or more analytes include cells from embryonic or blastocyst tissue biopsies, cells and / or genetic material isolated from cell culture media, blastocyst cavities, or combinations thereof.
[0011] In some embodiments, the tissue biopsy is collected using laser pulse-assisted pipetting delivery. In some embodiments, two or more cells are isolated from the tissue biopsy and the two or more cells are dissociated using physical or enzymatic methods. As a non-limiting example, physical methods may include laser cutting or micropipette dissociation. As a non-limiting example, enzymatic methods may include dissociation using digestive enzymes, including but not limited to: dispersants, collagenases, hyaluronidases, papain, DNase-I, Accutase, or trypsin.
[0012] In some embodiments, the genetic information obtained from each source may include nucleic acid sequences of panels, exomes, whole genomes, or transcriptomes, with or without mitochondrial sequences and with or without epigenetic markers. As a non-limiting example, nucleic acids may comprise DNA, RNA, or both.
[0013] In some embodiments, DNA is isolated from the source, and the whole genome amplification of DNA from each source yields... (The rest of the text appears to be a reference to a specification, page 2 / 18, reference 6, CN 121100382 A, and is not directly related to the previous paragraph.)A live library is prepared and the library is sequenced to generate a sequence data file of the source, thereby obtaining a nucleic acid sequence.
[0014] In some embodiments, embryos may be sequenced to reduce coverage and aneuploidy may be identified before the library is fully sequenced. Subsequently, the embryos may be eliminated from further sequencing. As a non-limiting example, aneuploidy may be detected by unexpectedly high or low coverage in genomic regions.
[0015] In some embodiments, RNA is isolated and sequenced. In some embodiments, RNA sequence data is used to characterize the transcriptome and determine gene expression levels. In some embodiments, gene expression levels are associated with patient outcomes. As a non-limiting example, patient outcomes include, but are not limited to, successful pregnancy.
[0016] In some embodiments, a genetic information panel is established to associate gene expression levels with patient outcomes.
[0017] In some embodiments, RNA sequence data is used to identify aberrant splicing events, thereby identifying dysregulation of transcriptional regulation and one or more variants affecting one or more transcriptional pathways. As a non-limiting example, aberrant splicing events may include cryptic splicing, aberrant exon use, or upstream open reading frames.
[0018] In some embodiments, RNA sequence data is used to identify biased allele expression, thereby identifying imprinting, copy number determination, and / or variants affecting expression.
[0019] In some embodiments, RNA sequence data is used in conjunction with DNA sequencing or alone to invoke genomic variants.
[0020] In some embodiments, the reference genome is a genomic sequence from a public or private repository, one or more parents, genetic relatives, and / or embryos produced from the same parent.
[0021] In some embodiments, the genetic variants are further identified by comparing genetic information from each source with genetic information from one or more parents to infer haplotype phasing, imputation of deletion data, determination of inheritance or deletion, and identification of distinct new variants. In some embodiments, the correlation between genetic variants and phenotypes is further annotated using a database of variants associated with disease, phenotype, disease risk, or protection and traits. In some embodiments, the likelihood that genetic variants will cause detrimental effects on gene function is further annotated by computer simulation, in vivo, or in vitro methods.
[0022] In some embodiments, the variant invocation procedure is a deep learning-based variant invocation procedure. In some implementations, the deep learning-based variant calling procedure is constructed based on the genetic information of one or more parents, genetic relatives, and / or embryos produced from the same parent, and is capable of identifying significant neoplasms for preimplantation genetic screening purposes. In some implementations, the genetic variant occurs in a genome known or suspected of contributing to genetic diseases in adults or children, or in a genome without currently associated genetic diseases but subject to genetic constraints.
[0023] In some embodiments, the combined steps of the methods described above and herein include identifying genetic variants that appear in at least two sources, or identifying genetic variants that can only be detected by combining genetic information from multiple sources. As a non-limiting example, the genetic variant is identified in multiple sources, or is determined to be erroneous, artificially created, or mosaic.
[0024] Some aspects of this disclosure provide a method for screening multiple embryos. This method includes performing genetic analysis on each embryo according to the methods described above and herein; and then ranking the embryos based on the results of the genetic analysis. Embryos with the best chance of full-term pregnancy and the lowest lifetime genetic risk are ranked highest in terms of transfer and implantation. In some embodiments, embryos are ranked based on the known harmfulness of detected variants, or based on the predicted harmfulness of detected variants predicted by various methods. As a non-limiting example, the prediction method may include computer-aided harmfulness prediction tools, allele evolution conservation, and / or allele frequencies in a selected or unselected population database.
[0025] In some embodiments, the influence of genetic variants on embryo ranking is weighted based on phenotypic inputs selected for examination. As a non-limiting example, said input is a specific pathological condition that the client wishes to avoid.
[0026] In some embodiments, the effect of genetic variants on embryo sequencing is weighted based on disease-related phenotypes and the automatic phenotypic overlap between their genes and phenotypes used for examination, as specified in the specification 3 / 18 page 7 CN 121100382 A. Overlap can be calculated using an ontology that allows related rather than strictly identical terms to determine overlap. As a non-limiting example, the phenotypes selected for examination are derived from the family history of one or two parents, where the parents may wish to avoid the risk of breast cancer, and therefore genes associated with other cancer risks will also be weighted more heavily.
[0027] In some embodiments, the effect of genetic variants on embryo sequencing is weighted based on the conjugability of the variant, the presence of other variants in the genome, the predicted harmfulness of the variant, whether the variant has been prediagnosed as benign, unknown, potentially pathogenic, or pathogenic, and the known severity of disease, penetrance, and / or the genetic pattern associated with the gene.
[0028] In some embodiments, the effect of genetic variants on embryo sequencing is weighted based on the presence or absence of the variant in one or two parents and the conjugability in said parents.
[0029] In some embodiments, the effect of genetic variants on embryo sequencing is weighted based on known or suspected age of onset.
[0030] In some embodiments, the effect of genetic variants on embryo sequencing is weighted based on pre-curated gene or variant sequencing.
[0031] In some embodiments, the effect of genetic variants on embryo sequencing is based on treatment options for a given disease.Weighted.
[0032] In some embodiments, the effect of genetic variants on embryo sorting is weighted based on their effect on gene function or expression.
[0033] In some embodiments, the screening method further includes generating a report with embryo sorting, wherein the sorted embryos are qualitatively classified based on recommendations for transplantation, non-transplantation, or uncertain transplantation. Brief Description of the Drawings
[0034] To better understand the subject matter disclosed herein and to illustrate how it can be implemented in practice, embodiments are now described by way of non-limiting example only with reference to the accompanying drawings, in which: Figure 1 illustrates carrier screening vs. reproductive risk.
[0035] Figure 2 is a flowchart outlining whole-genome sequencing (WGS) and whole-genome amplification (WGA) of cell lines, in which analytes from single cells and analytes from a large number of cells are analyzed and compared.
[0036] Figure 3 is a flowchart outlining the methods of WGS and WGA and the subsequent identification of genetic variants in embryo samples, in which multiple analyte sources are analyzed.
[0037] Figure 4 is a flowchart outlining the methods for identifying genetic variants in embryos (using WGA and WGS) considering parental genetic data.
[0038] Figure 5 is a flowchart illustrating variant classification into broad categories, including known pathogenic variants, non-coding variants, coding variants, loss-of-function (LoF) variants, synonymous variants, and variants that will not require further consideration (Trash), as well as "Panic" cases requiring manual review.
[0039] Figure 6 is a flowchart illustrating variant classification into broad categories for genetic, coding, and LoF variants, such as "No go," where embryos will not be considered for transplantation, nor will variant weighting be considered to assist in embryo sorting.
[0040] Figure 7 is a flowchart illustrating variant classification into broad categories for known pathogens (e.g., from ClinVar data), such as "No go," weighting, and manual review.
[0041] Figure 8 is a flowchart illustrating variant classification for "pathogenic" (i.e., possibly, known, or very likely pathogenic) variants in genes associated with autosomal recessive diseases ("autosomal recessive processes"), as per the specification page 4 / 18 8 CN 121100382 A. Determining whether multiple variants exist in a gene, whether these variants are cis or trans, or whether variants exist in a gene in a homozygous or hemizygous state, to broadly classify embryos as "prohibited" for weighting or manual review.
[0042] Figure 9 is a flowchart illustrating the classification of variants encoding neonatal LoFs into the broad categories shown above.
[0043] Figure 10A is a flowchart illustrating the classification of structural variants as "prohibited" for embryos or their reclassification as LoFs in the appropriate gene for SV deletion.
[0044] Figure 10B is a flowchart illustrating the classification of structural variants as “prohibited” for embryos or reclassification as LoF in the appropriate gene obtained for SV.
[0045] Figure 10C is a flowchart illustrating the classification of structural variants as “prohibited” for embryos or reclassification as LoF in the appropriate gene for inversion.
[0046] Figure 10D is a flowchart illustrating the classification of structural variants as “prohibited” for embryos or reclassification as LoF in the appropriate gene for translocation.
[0047] Figure 11 is a flowchart illustrating the classification of repeat amplifications into the broad categories of embryos as shown above, based on known pathogenicity and the length of the “pre-mutation” repeat amplifications and the repeat length in the parent.
[0048] Figure 12A is a flowchart illustrating the classification of homozygous variants into broad regions for normal treatment (recessive disease), normal treatment (deletion), and manual review.
[0049] Figure 12B is a flowchart illustrating the classification of synonymous and intron variants into those considered as LoF.
[0050] Figure 12C is a flowchart illustrating the classification of broad regions of uniparental genetic characteristics into prohibited or manual review.
[0051] Figure 12D is a flowchart illustrating the potential classification of variants identified from, for example, the spinal muscular atrophy pipeline.
[0052] Figure 13 is a flowchart illustrating examples of classifying SNVs or small insertion / deletion variants into various broad categories. Initial steps involving DeepVariant, VCF Info, and population frequency are detailed, which in some cases can lead to junk or late-expanding broad categories.
[0053] Figure 14A is a flowchart illustrating the reference sequence steps of the analysis, where classification can be performed through decision points such as coding variant, intron variant, intergenic variant, variant is meaningless or LoF, variant is missense, variant is synonymous, variant splicing (X is 4), variant splicing (clinical X is 4), and pathogenicity evidence (ClinVar), which can potentially lead to junk, junk, or late-expanding broad categories.
[0054] Figure 14B is a flowchart illustrating the phenotypic classification steps of the analysis, where classification can be performed through decision points of ACMG criterion compliance, ClinVar pathogenicity criterion compliance, OMNIM criterion compliance, and embryonic "Intolerome" criterion compliance, which can potentially lead to large classes of waste or late expansion.
[0055] Figure 14C is a flowchart illustrating the triple genetic steps of the analysis, where classification can be performed through decision points of neonatal dominant, genetic dominant, homozygous recessive, X-linked, genetic complex heterozygote, and artificial review capture net, which can potentially lead to large classes of waste or late expansion. Note that the decision points of neonatal dominant, genetic dominant, homozygous recessive, X-linked, and genetic complex heterozygote all have sub-flowcharts described in detail in Figures 16A-16E.
[0056] Figure 15 is a flowchart illustrating the report / mask subset steps of this process.
[0057] Figure 16A is a sub-flowchart illustrating the neonatal dominant decision point, where analysis can be performed through decision points of neonatal (non-parental), parental, dominant mechanism, and recessive mechanism, which can potentially lead to a class of panic or other pipelines.
[0058] Figure 16B is a sub-flowchart illustrating the X-linked decision point, where analysis can be performed through decision points of non-PAR region, proband male and hemizygote, and proband female and homozygote, which can potentially lead to a class of artificial review or other pipelines. Specification 5 / 18 pages 9 CN 121100382 A
[0059] Figure 16C is a sub-flowchart illustrating the genetic dominant decision point, where analysis can be performed through decision points of neonatal (non-parental), parental, (P1)-Het / (P2)-WT / (E)-Het, rare (P1 / 2) and (E)-Het, dominant mechanism, and recessive mechanism, which can potentially lead to a class of panic or other pipelines.
[0060] Figure 16D is a sub-flowchart illustrating the decision point for recessive homozygotes, where analysis can be performed through decision points for newborn (non-parental), parental, (P1)-Het / (P2)-Het / (E)-Homo, (P1)-Het / (P2)-WT / (E)-Homo, rare (P1 / 2) and (E)-Homo, dominant and recessive mechanisms, which can potentially lead to large classes of censorship, panic, or other pipelines.
[0061] Figure 16E is a sub-flowchart illustrating the decision point for genetic complex heterozygotes, where analysis can be performed through decision points for genes with >= 2 SNVs, SNV phased on a separate chromosome, phased SNV unknown, phased SNV LoF / LoF, phased SNV LoF / other pathogenicity, SNV LoF / LoF, and SNV LoF / other pathogenicity, which can potentially lead to large classes of censorship, censorship, and panic.
[0062] Figure 17A shows an example of aneuploidy invoked from an individual cell set of an aneuploid female embryo, which shows a loss of homology of chromosome 19.
[0063] Figure 17B shows a different example of aneuploidy invoked from an individual cell set of an euploid female embryo, wherein some copy number variants are present in individual cells, but no copy number variants are present homologously throughout the embryo. Detailed Description
[0064] In order to facilitate an understanding of the principles of this disclosure, embodiments will now be referenced and described using specific language. However, it should be understood that this is not intended to limit the scope of the disclosure, and such changes and further modifications to the disclosure as shown herein are commonly thought of by those skilled in the art.
[0065] The terminology used herein is for describing particular embodiments only and is not intended to limit the invention.
[0066] As used herein, the terms “comprise,” “comprising,” “include,” “including,” “have,” “having,” and variations thereof mean “including, but not limited to.” The term “composed of” means “including and limited to.” The term “consistently composed of” means that a composition, method, or structure may include additional ingredients, steps, and / or portions, provided that such additional ingredients, steps, and / or portions do not materially alter the basic and novel features of the claimed composition, method, or structure.
[0067] As used herein, terms refer to the manner, means, techniques, and procedures for accomplishing a given task, including but not limited to manner, means, techniques, and procedures known to or readily developed from known methods, means, and procedures by those skilled in the art of chemistry, pharmacology, biology, biochemistry, and medicine.
[0068] In this application, various embodiments may be presented in the form of a scope. It should be understood that the description in the form of a scope is merely for convenience and brevity and should not be construed as a rigid limitation on the scope of the invention. Therefore, the description of a scope should be considered to have disclosed all possible sub-scopes and the various numerical values within that scope. For example, a description of a range such as 1 to 6 should be considered as having specifically disclosed ranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., and individual numerical values within that range, such as 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0069] Whenever a numerical range is indicated herein, it means to include any referenced number (fraction or integer) within the indicated range. The phrases “range between the first indicated value and the second indicated value” and “range from the first indicated value to the second indicated value” are used interchangeably herein and are intended to include the first and second indicated values and all fractions and integers therebetween. Specification 6 / 18 pages 10 CN 121100382 A
[0070] As used herein, the term “analyte” or “analytes” refers to biological macromolecules, such as proteins, DNA, and / or RNA, that are essential for the function, growth, and development of cells and organisms.
[0071] As used herein, the phrase “a source of analyte” or “sources of analytes” refers to a source from which genetic information of an embryo can be obtained. Sources include, but are not limited to, single cells or groups of cells from tissue biopsies of the embryo or blastocyst, cells or genetic information from cell culture media, the blastocoel, or combinations thereof.
[0072] As used herein, the phrase “genetic information” refers to the nucleic acid sequence of a subgroup, exome, whole genome, or transcriptome, which may or may not have a mitochondrial sequence and epigenetic markers.
[0073] As used herein, the phrase “reference genome” refers to the whole genome sequence from a public or private repository, one or more parents, or one or more genetically related individuals (such as biological siblings and / or embryos derived from the same parent), where reference genetic information is available.
[0074] As used herein, the terms “whole genome amplification” or “WGA” refer to a technique, method, or kit for producing large amounts of DNA from a small amount of initial raw material. Unlike conventional PCR, WGA aims to amplify the entire genome of an organism, rather than a specific region.
[0075] As used herein, the terms “whole genome sequencing,” “whole-genome sequencing,” and “WGS” are used interchangeably to refer to a laboratory procedure that determines all or almost all of an organism’s genomic DNA sequence at once. In this disclosure, whole-genome sequencing is performed on each embryo and each parent to determine the full 6 billion bases of the genome of each embryo and parent, which may or may not include the mitochondrial genome. The result of this sequencing is a complete DNA sequence of the individual, including non-coding sequences, and may include epigenetic markers.
[0076] As used herein, the term “coverage” refers to the average number of times a particular sequence is queried by a sequence read. For example, “30 x WGS” means that the entire genome will be sequenced with an average redundancy depth of 30 times.
[0077] As used herein, the term “coverage analysis” refers to assessing the variability of the genome, including but not limited to the presence of regions with coverage higher or lower than expected. Where the expected coverage is based on total sequence coverage, known biases in coverage based on local genomic composition or embryonic sex can be determined empirically.
[0078] As used herein, the term “variant call” refers to the process of identifying differences between sequencing reads of a target individual and sequencing reads of a reference genome. And as used herein, the term “variant call procedure” refers to a tool for calling variants of sequencing data from one or more nucleic acid sequencing datasets (DNA, RNA, etc.). As a non-limiting example, the variant invocation tool used herein may include DeepVariant, a novel variant invocation program based on TensorFlow machine learning.
[0079] As used herein, the phrase “training DeepVariant” refers to the dataset provided for training the algorithm, where the expected outcome is referred to as the “truth set”, so that the machine learning algorithm is penalized or rewarded based on performance. The algorithm is then used...For datasets without a known truth set, similar performance is expected.
[0080] As used herein, the term “merge variant call” refers to merging variant call results from multiple single sources from a single embryo, or viewing call results from all embryo data and parental data simultaneously. The term “merge variant call procedure” refers to a tool or method for calling variants from multiple sources.
[0081] As used herein, the phrase “genetic variant” refers to differences in genetic information between embryo analyte sources, or differences in genetic information between an embryo and a reference genome. Different types of genetic variants are described below.
[0082] As used herein, the terms “single nucleotide variant,” “single-nucleotide variant,” or “SNV” refer to a variant in a single nucleotide, which occurs when a single nucleotide (adenine, thymine, cytosine, or guanine) in the genome sequence is altered. SNVs may be rare or common in one population, but common or rare in different populations. SNVs are sometimes referred to as single nucleotide polymorphisms (SNPs), although SNVs and SNPs are not used interchangeably. To qualify as an SNP, the variant must be present in at least 1% of the population.
[0083] As used herein, the term “indel” is an abbreviation for insertion / deletion, which refers to the simultaneous insertion and deletion of a small segment of DNA (typically less than 50 base pairs) in the genome relative to a reference genome. A special case of insertion / deletion is duplication amplification or contraction, where the DNA repeating element becomes longer or shorter.
[0084] As used herein, the term “copy number variant” or “CNV” refers to a duplication or deletion that alters the copy number of a specific DNA segment within the genome. Aneuploidy is a type of CNV that typically consists of a very large chromosomal segment or an entire chromosome. CNVs contribute 9% of pathogenicity in hereditary retinal degeneration (Zampaglione et al., 2020; Genetics in Medicine, 22: 1079-1087).
[0085] As used herein, the term “structural variant” or “SV” refers to a re-sampling of a genomic portion, which may be a deletion, duplication, insertion, inversion, translocation, or a combination thereof, and may or may not be copy-neutral. SVs are associated with a variety of conditions, including polycystic kidney disease, cardiomyopathy, amyotrophic lateral sclerosis (ALS), and some cases of intellectual disability.
[0086] As used herein, the term “pathogenic genetic variant” refers to a genetic change in the DNA sequence that has been shown to cause disease or be associated with an increased risk of disease.
[0087] As used herein, the term “phenotype” refers to any measurable characteristic of an individual that is at least partially attributable to genetics. It may include, but is not limited to, disease risk or trait protection.
[0088] As used herein, the terms “de novo variant,” “obvious de novo variant,” “DNV,” or “obvious DNV” refer to a genetic variant in the embryo that is not obviously present in the sequence data derived from the parent. In some cases, the variant occurs in the egg or sperm cells of the parent, but not in any of its other cells. In other cases, the variant may be present in a subset of parental cells, including gonads, but not in the sequenced sample (chimeric). In still other cases, the variant occurs in the embryo after fertilization. As the embryo grows, some or all of the resulting cells in the growing embryo contain this variant. A de novo variant is an explanation of a genetic condition that affects the child but not both parents.
[0089] As used herein, the term “parent” refers to the gamete donor (i.e., the biological parent) that contributes to the genetic composition of the embryo.
[0090] As used herein, the terms “patient outcome”, “patient outcomes”, etc., refer to the outcome of pregnancy, the onset of disease, or the initiation of other desired or unwanted conditions in the parent, embryo, or child.
[0091] As used herein, the terms “favorable outcome”, “favorable outcomes”, “successful pregnancy”, etc., refer to the achievement of avoiding a specific negative or pathological condition. For example, a favorable outcome could be maintaining the fetus to full term. In other non-limiting instances, a favorable outcome could be that the fetus does not have a genetic susceptibility to one or more pathological conditions. One or more pathological conditions may vary according to the wishes of the client’s biological parents.
[0092] As used herein, the terms “a pathologic condition”, “pathologic conditions”, etc., refer to an abnormal anatomical or physiological condition that needs to be avoided when screening an embryo according to the wishes of the client or biological parents. Specification 8 / 18 pages 12 CN 121100382 A
[0093] This disclosure provides in part a method for identifying genetic variants in an embryo. The method involves obtaining two or more analyte sources from an embryo; analyzing the two or more analyte sources to obtain genetic information from each source; comparing the genetic information from each source with one or more reference genomes using at least one variant calling procedure, wherein the variant calling procedure identifies variants between each source and the reference genome; and combining the variants to...The identification is based solely on the differences presented by the source, wherein the differences are genetic variants in the embryo.
[0094] The genetic variant to be identified may be a single nucleotide variant (SNV), a multinucleotide variant (MNV), a copy number variant (CNV), a structural variant (SV), or an alteration of an epigenetic marker.
[0095] MNVs include, but are not limited to, insertions, deletions, combinations of insertions and deletions, or duplications involving multiple nucleotides. MNV calls may include, but are not limited to, identifying insertions, deletions, simultaneous insertion-deletion (indel), and duplications and contractions (specific examples of insertions and deletions, respectively). As a non-limiting example, single-base deletions / insertions / insertion-deletion or multi-base deletions / insertions / insertions can be identified by MNV calls.
[0096] SVNs include, but are not limited to, inversions, translocations, duplications and contractions (specific examples of insertions or deletions). SV calls may include, but are not limited to, identifying genomic rearrangements, including rearrangements that result in copy number changes and copy-neutral changes. As a non-limiting example, copy-neutral changes may include translocations and inversions.
[0097] Epigenetic markers can be DNA methylation markers used to identify imprinting defects. Methylation markers can be used to identify imprinting defects, including, but not limited to, when both alleles are imprinted and the X is inactivated by a distorted pattern.
[0098] Parental data can be used to identify gene mutations that occur in the embryo but not in the parental WGS, i.e., de novo variants (DNVs). Parental data can also be used to determine genetic traits, including haplotype phasing and allelic disomies, and to identify whether a genetic variant is cis or trans, i.e., one from the mother and one from the father. As a non-limiting example, determining whether one variant comes from the father and the other from the mother may result in a recessive disease.
[0099] Parental data can also be used to infer missing data in the embryo. Sometimes amplification or sequencing may fail to determine a portion of the embryonic genome, and one can infer the presence of other variants in the embryo from the inheritance of a particular haplotype, even if they are not directly observed.
[0100] In some embodiments, the methods for identifying genetic variants in embryos described above and disclosed herein can be used for preimplantation genetic testing (“PGT”), in which embryos are genetically tested, and embryos are screened prior to transfer and implantation, wherein the genomes of individual cells isolated from each embryo and both parents are sequenced and analyzed. Typically, the future parents and / or client will use IVF technology to create a set of their embryos. However, tissue biopsies of the embryos are performed according to standard protocols. Prior to embryo transfer for pregnancy, the entire 3 billion base pairs (or 6 billion bases) of the genome of each embryo are sequenced, as are the genomes of both biological parents. For whole-genome sequencing of each embryo, from the embryo…Single cells are isolated from the fetus and sequenced. The sequencing data are then merged and analyzed to identify the embryos with the lowest genetic risk for embryo transfer and implantation. The PGT and methods disclosed above and herein detect not only prenatal genetic risk but also postnatal genetic risk.
[0101] In some embodiments, changes to the methods or related techniques disclosed above and herein for identifying genetic variants in embryos can be applied to free-floating or tissue biopsy embryonic cells collected from the uterus or blood of a pregnant woman. That is, in addition to embryo PGT in assisted pregnancy, similar methods or techniques can be used for natural pregnancy through improved genetic screening.
[0102] In some embodiments, changes to the methods or related techniques disclosed above and herein for identifying genetic variants in embryos can be applied to one or more biological parents. That is, in addition to embryo PGT in assisted pregnancy, similar methods or techniques can be used to provide information for natural pregnancy plans through improved genetic screening. (Instruction manual 9 / 18 pages 13 CN 121100382 A)
[0103] To identify genetic variants in an embryo, analytes can be obtained from cells from an embryonic or blastocyst tissue biopsy, cells and / or genetic material isolated from a cell culture, the blastocyst cavity, or a combination thereof. This includes, but is not limited to, analytes from culture media, blastocyst cavity fluid, or parental samples (e.g., blood, oral swabs, saliva, semen), followed by separation / dissociation of a large number of cells into single cells. For any of the methods described above and disclosed herein, 1-10 single cells can be isolated from each embryo for single-cell analysis. As a non-limiting example, 2, 3, 4, 5, 6, 7, 8, 9, or 10 single cells can be isolated from an embryo.
[0104] In some embodiments, a laser pulse-assisted pipette delivery method is used to collect tissue biopsies. Embryonic tissue biopsies can be performed after 5-6 days of culture in a laboratory, at which point the embryo has reached the blastocyst or hatching blastocyst stage. In some embodiments, two or more cells are isolated from the tissue biopsy using physical or enzymatic methods. As a non-limiting example, physical methods may include laser cutting or micropipette separation. The cells can then be dissociated into single cells. As a non-limiting example, enzymatic methods may include extraction using digestive enzymes, including but not limited to dispersases, collagenases, hyaluronidases, papain, DNase-I, Accutase, or trypsin.
[0105] Single cells need to be amplified (e.g., WGA) before sequencing because their DNA load on the sequencer is too low. Amplification can lead to errors such as failure to amplify parts of the genome, over-amplification of genomic parts, and failure to amplify two alleles of a genomic part. These technical artifacts can mimic real and harmful variants.However, by sequencing multiple single cells, the methods described above and disclosed herein can infer the true variant from the output product through multiple independent observations of the variant.
[0106] The genetic information obtained from each source may include nucleic acid sequences of a group, exome, whole genome, or transcriptome, which may or may not have mitochondrial sequences and epigenetic markers. As a non-limiting example, the nucleic acid may comprise DNA, RNA, or both.
[0107] In some embodiments, nucleic acid sequences are obtained by isolating DNA from the source, generating a library from whole genome amplification of DNA from each source, and sequencing the library to generate a sequence data file of the source.
[0108] In some embodiments, embryos may be sequenced to reduce coverage and aneuploidy may be identified before the library is fully sequenced. Subsequently, the embryos may be eliminated from further sequencing. As a non-limiting example, aneuploidy may be detected by unexpectedly high or low coverage in genomic regions.
[0109] Whole-genome sequencing (WGS) of both biological parents and each embryo provides the highest possible resolution and sensitivity for genetic analysis. In contrast, current carrier screening or prenatal testing analyzes only about 1% of the genome for which a clinical diagnosis is most certain; analyzing one parent at a time (carrier screening) or only the fetus (prenatal testing) further limits the possible insights. Current IVF products use specialized sequencing or SNP arrays with limited coverage, which produce 1000 times less data than whole-genome sequencing, primarily due to the difficulty in obtaining and interpreting a wider range of genetic data, and due to efforts to control costs.
[0110] In some embodiments, for example, sequencing data from individual cells are used to call variants, and the results are then combined for analysis. In some embodiments, for example, sequencing data from individual cells are first combined and then used to call variants.
[0111] In some embodiments, RNA is isolated and sequenced. In some embodiments, RNA sequence data is used to characterize the transcriptome and determine gene expression levels. In some embodiments, gene expression levels are correlated with patient outcomes. As a non-limiting example, patient outcomes include, but are not limited to, successful pregnancy. In this approach, a set of genetic information associated with gene expression levels and patient outcomes is established.
[0112] RNA sequence data can be used to identify aberrant splicing events and to identify dysregulation of transcriptional regulation and one or more variants affecting one or more transcriptional pathways. As non-limiting examples, aberrant splicing events may include cryptic splicing, aberrant exon use, or upstream open reading frames.
[0113] RNA sequence data can also be used to identify biased allele expression, thereby identifying imprinting, copy number determination, and / or variants affecting expression.
[0114] In some embodiments, RNA sequence data is used to call genomic variants alone or in combination with DNA sequencing. In some embodiments, RNA is used to confirm aneuploidy calls. In some embodiments, multiple data points from each cell (DNA, RNA) may be combined, and / or data points from multiple cells may be combined.
[0115] To identify genetic variants in an embryo, genetic information from each source is compared with genetic information from one or more reference genomes. In some embodiments, the reference genome is a genomic sequence from a public or private repository, one or more parents, genetic relatives, and / or embryos produced from the same parent. When using genomic sequences from one or more parents as a reference genome, genetic variants can be further identified by comparing genetic information from each source with genetic information from one or more parents to infer haplotype phasing, imputation of missing data, determination of inheritance or deletion, and identification of distinct new variants. In some embodiments, the correlation between genetic variants and phenotypes is further annotated using a variant database associated with disease, phenotype, disease risk, or protection and traits. In some embodiments, the potential for genetic variants to have detrimental effects on gene function is further annotated using computer simulations, in vivo or in vitro methods.
[0116] Identified variants can be annotated based on their presence or absence in other datasets, including cohorts of healthy adults, children with early-onset diseases, and triple datasets. These datasets include both private and public datasets. Academic research to date has developed useful data in silos. Research on genetic diseases in children and adults is conducted separately from research on the genetic causes of miscarriage and infertility. These existing datasets can be added to embryo research data and client data to expand the reference sequence database.
[0117] In some embodiments, the variant calling procedure can be a deep learning-based variant calling procedure. In some embodiments, the deep learning-based variant calling procedure is built on genetic information from one or more parents or genetic relatives and is capable of identifying obvious new variants for preimplantation genetic screening purposes. There may be genetic variants in the genome that are known or suspected to contribute to genetic diseases in adults or children (i.e., pathogenic genetic variants). In other cases, there may be genetic variants in the genome that are not currently associated with any genetic disease but are genetically restricted.
[0118] For the methods described above and herein for identifying genetic variants in embryos, the combination step may include identifying genetic variants present in at least two sources, or identifying genetic variants that can only be detected by combining genetic information from all sources. As a non-limiting example, genetic variants may be identified in multiple sources (not necessarily most sources), or identified as mosaics.
[0119] Some aspects of this disclosure provide a method for screening multiple embryos. This method includes performing a genetic analysis on each embryo according to the methods described above and herein; and then ranking the embryos based on the results of the genetic analysis. Embryos with the best chance of full-term pregnancy and the lowest postpartum genetic risk are ranked highest in terms of transfer and implantation. In some embodiments, embryos are ranked based on the known harmfulness of detected variants or based on the predicted harmfulness of detected variants predicted by various methods. As a non-limiting example, the prediction method may include computer-aided harmfulness prediction tools, allele evolution conservation, and / or allele frequencies in a selected or unselected population database.
[0120] The influence of genetic variants on embryo ranking may be weighted based on phenotypic inputs selected for examination. As a non-limiting example, the input may be a client's desire to avoid a specific pathological condition.
[0121] Alternatively, the influence of genetic variants on embryo ranking may be weighted based on disease-related phenotypes and automatic phenotypic overlap between the selected relevant genes and phenotypes for examination. Overlap can be calculated using an ontology, as described on pages 11 / 18 of this specification, CN 121100382 A, which allows for the determination of overlap using related rather than strictly identical terms. As a non-limiting example, the phenotypes selected for examination are derived from the family history of one or two parents, where the parents may wish to avoid the risk of breast cancer, and therefore genes associated with other cancer risks will also be weighted more heavily.
[0122] Alternatively, the influence of genetic variants on embryo sequencing can be weighted based on the presence or absence of variants in one or two parents and their conjugation in the embryo or parent, known or suspected age of onset, predicted gene or variant sequencing, treatment options for a given disease, or their influence on gene function or expression.
[0123] The screening methods described above and herein may also include generating an embryo sequencing report. Embryo sorting can be divided into two or more categories, rather than quantitative sorting, including: red light (high risk; embryo transfer not recommended), yellow light (some identifiable risk; different parents may weight it differently based on their own preferences; the parents decide whether to transfer), green light (no specific risk identified; transfer and implantation recommended), and a set of no data generated. That is, sorted embryos can be qualitatively classified based on transfer, no transfer, transfer with some specific identifiable risk, or recommendation for uncertain risk.
[0124] The following examples are for illustrative purposes only and not for limitation.
[0125] Example 1: Embryo Screening An early version of the pipeline was developed to identify embryos that may have unwanted high genetic risk (marked "prohibited") and stratify the remaining embryos based on an assessment of the remaining genetic risk. Embryos with a lower probability of disease were marked "released".
[0126] Simulations were performed using random "crossovers" of 14 males and 14 females, with 20 "embryos" per crossover. That is, 196 parental combinations, each with 20 embryos, produced 3920 simulated embryos. It has been determined that a typical couple in their 30s has approximately 96-98% chance of embryos being transferred in their first cycle, passing all disease screenings (Table 1).
[0127] Table 1
[0128] This pipeline was tested to analyze a group of children with rare diseases. Although only ~30% of children with rare early-onset diseases receive genetic diagnosis, the method of this disclosure can successfully screen ~80% of children because of rare harmful variants in their genes, which are limited by harmful variants.
[0129] In summary, compared to the best current clinical diagnosis, the early form of the analytical pipeline identifies more than 2.5 times the risk of children and has a 96-98% probability of at least one "healthy" embryo being transferred.
[0130] Example 2: Carrier testing produces a first-pass pipeline to fully characterize parental genetic risk. A screening of 900 parental genomes was conducted to identify loss-of-function or known pathogenic or extremely harmful (top 0.1%) mutations in all known recessive disease genes, as well as known pathogenicity in autosomal dominant genes, and then compared to the Invitae carrier cohort (“Invitae Comprehensive Carrier Screening”, which includes 287 genes). As shown in Figure 1, the genetic risks identified by this pipeline were approximately 6-fold higher compared to the conventional comprehensive cohort. Furthermore, this pipeline identified not only autosomal recessive genes (rare, clinical variants, and LoF), but also rare variants of autosomal dominant genes from the Clin Var database, and autosomal recessive variants rare in populations with high combinatorial annotation-dependent exhaustion scores (CADD), indicating a high likelihood of being harmful. Identified genetic risks included CRTAP (osteogenesis imperfecta, type VII) and PIEZO2 (distal joint contractures, impaired proprioception and tactile sensation). Instructions for Use, Pages 12 / 18, 16 CN 121100382 A
[0131] Example 3: Cell Line Studies First, whole-genome sequencing (“WGS”) and whole-genome amplification (“WGA”) of the cell lines were characterized. (Figure 2). These cell lines were purchased from Corriell and had publicly available sequence data from the Genome in a Bottle (GIAB) Consortium (NA24385, NA24632, and NA18278). The aim of these studies was to evaluate the WGA kit and determine a pooling strategy for use in embryo screening.
[0132] Materials: Two different sets of samples were used. The first set consisted of ten (10) single cells from three GIAB samples, at 15xWGS sequencing, the second group being “bulk” samples of three (3) 5-6 cells from each GIAB. Variant calls had previously been performed on each GIAB sample to produce a truth set of genetic variants that could be used to compare the results of this WGS method.
[0133] As a first pass, the coverage variability of the three GIABs was assessed to identify regions with higher-than-expected coverage (if any), and to determine how many low-coverage regions were identified by comparing the average coverage of appropriate large regions (10-100 Kbp) with the expected (average) coverage. For each single-cell sample, a subset of data (e.g., 10x, 4x, 5x, 1x, 0.1x) was preprocessed. The same WGS was repeated for the “bulk” samples.
[0134] Next, variants could be identified individually or in combination with other genetic data in each cell using any available variant call program (platypus, haplotype call program for GATK, Freebayes, Octopus, Delly, etc.). Furthermore, variants can be identified using a custom-trained AI variant invoking program optimized for detecting CNVs and SNVs from single-cell data. This invoking program is trained on a dataset (the truth set) with known expected outcomes, thus penalizing or rewarding the machine learning algorithm based on performance. However, the algorithm is then used on datasets without known correct answers but with similar expected performance.
[0135] SNV Invoking: For each dataset in the first set of samples (i.e., 3 x 10 single cells), variants are independently invoked using platypus when sampling at 15x and 10x (60 variant invoking forms (VCFs) – set A). For each cell line, 4x from 10 cells, 5x from 8 cells, and 10x from 4 cells (repeated using different cells) are combined, along with variants invoked using platypus (12 VCFs – set B). The aim is to sort everything to a total of 40x to determine if cell number is important. For each dataset from the second set of samples (3 “large” samples), platypus call variants (3 VCFs – set C) are used.
[0136] Next, 5, 7, and 10 random VCFs from each WGA from set A are merged using two techniques: the majority rule (i.e., seen 3, 4, or 6 times respectively) and the “seen more than once” rule (i.e., any variant seen in at least 2 VCFs). The result of the merging will produce 18 VCFs (set D).
[0137] Then, calls are compared between the 12 VCFs (set B), 3 VCFs (set C), and 18 VCFs (set D) by constructing a confusion matrix for each sample. The confusion matrix records the true positives and true negatives between the sample and the known truth set.Sex, false positives, and false recessive rates.
[0138] Aneuploidy: Aneuploidy is determined by acquiring whole-genome data with low coverage (0.1X) and comparing the expected coverage of large regions (10-100 Mbp) with the measured coverage. Regions with higher-than-expected coverage may indicate copy increases, which may include the entire chromosome or a part of a chromosome, while regions with low coverage may indicate copy loss of all or part of the chromosome. Expected coverage can be determined theoretically based on sequencing depth, GC content, and other local genomic factors, or based on empirical data from sequencing euploidy samples.
[0139] SV calls: For each dataset from the first set of samples (i.e., 3 x 10 single cells), variants are called independently at 15x, 10x, 5x, 1x, and 0.1x (specific calling procedures may be required for 1x and 0.1x) (150 VCF-sets E). For each cell line, 4x from 10 cells, 5x from 8 cells, and 10x from 4 cells (using different cell repeats, page 13 / 18 of the specification, CN 121100382 A) were merged, along with a delly call variant (12 VCFs – set F).
[0140] Next, 5, 7, and 10 random VCFs from each WGA from set A were merged using two techniques: the majority rule and the “see more than once” rule. The result of the merges would produce 9 VCFs (set G).
[0141] SURVIVOR was used to determine how many non-true deletions / gains were called for each VCF in set E, and to examine the frequency of their consistency (i.e., non-random deletions or over-amplifications).
[0142] Once nascent calls were generated from the isolated cells, the results would help inform the WGA and merge strategy. Furthermore, data from the cell line experiments were combined with data from GIAB to prepare a triplet, which was used to train DeepVariant.
[0143] Example 4: Donated Embryo Research extends the cell line methodology from Example 3 above to embryos (Figure 3). First, a tissue biopsy is performed on each embryo, and one or more single cells are isolated from each embryo using physical or enzymatic methods. The initial raw materials can be embryonic materials from various sources, such as embryos, blastocysts, cell culture media, etc.
[0144] The WGA kit and pooling strategy defined in the cell line experiments of Example 3 are applied to the embryos. These methods include steps such as very low-pass WGS (0.1x), which is used to eliminate some types of aneuploidy, and methods for analyzing single cells vs. pooled cells. Sequencing steps can be sequencing DNA (genome), RNA (transcriptome), or identifying methylation sites. Each has its own library preparation method. The goal of each sequencing output is to identify variants associated with disease or disease risk / protection.
[0145] Data from WGA and WGS are then analyzed to determine the consistency of individual cells with embryos, thereby evaluating reproducibility and consistency. Results for approximately six individual cells from each cluster are compared with results for the entire embryo. Furthermore, results from each cluster are compared with each other. The same software pipeline is used, but with simplified data management. Additionally, previously trained DeepVariant is applied to the embryo data for evaluation. After analysis, embryos are classified and ranked based on the chance of full-term pregnancy and postnatal genetic risk, and reports are generated to recommend which embryo(s) to select for transfer and implantation. Classification and report generation involve both manual and automated steps (Figures 5-12). For example, the classification of some variants can be fully automated because the variant has previously been reported as a harmful variant from multiple sources; in other cases, variants can be weighted because they may appear harmful in various ways but have not previously been reported as pathogenic, and they may be present in genes that do not always cause disease. Some cases may require manual review, such as multiple new mutations in a single gene associated with autosomal recessive genetic diseases.
[0146] Example 5: Genetic analysis and embryo sequencing will further extend the results from cell lines and embryos into an end-to-end implementation plan of a minimum feasible protocol (MVP) (Figure 4). First, a tissue biopsy will be performed on each embryo, and one or more single cells will be isolated from each embryo using physical or enzymatic methods. Next, the WGA kit and pooling strategy defined in the cell line experiments of Example 3 and finalized in the donor embryo experiments of Example 4 will be applied in this study. Parental WGS will use standard WGS methods.
[0147] Variant calling methods for structural variants (SVs) and single nucleotide variants (SNVs) that utilize parental genetic information will be developed. For example, SNV calling for embryonic cells will include haplotype phasing, imputation of missing data, and neonatal calling—all of which utilize parental genetic information. The current approach involves exploring modifications to DeepVariant to accommodate these steps, as described in Example 3.
[0148] After analysis, embryos are classified and ranked based on the chance of full-term pregnancy and postpartum genetic risk, and a report is generated to recommend which embryo(s) to select for transfer and implantation. Classification and report generation involve both manual and automated steps (Figures 5-12). Specifically, variants can be categorized into broad classes, including known pathogenic variants, non-coding variants, coded variants (see specification page 14 / 18, CN 121100382 A), loss-of-function (LoF) variants, synonymous variants, and variants that will not require further consideration (junk) as well as “panic” cases that will require manual review (Figure 5). For SNVs or <50For small insertions / deletions (indels) of <bp, the workflow can include evaluating variant features through the following decision points: variants in the blacklist; variants in the whitelist; (i.e., known pathogenicity, possible special considerations, and potential to cause panic), ClinVar variants; variants in the protein-coding region and not synonymous variants; and intronic and synonymous variants; as a classification process, the ClinVar decision point can further include considering whether the variant is benign or likely benign (B or LB) and pathogenic or likely pathogenic (P or LP). If it is pathogenic P or LP, it follows the ClinVar pathway discussed in more detail below with reference to Figure 7. The decision point for variants in the protein-coding region and not synonymous variants can further include considering whether the variant has a start loss, stop gain, and splice + / −2, deep splice + / −10, splice AI loss >.5, or frameshift to determine whether the variant is LoF or non-LoF. If the variant is not in the protein-coding region and not a synonymous variant, then further considering whether the variant is an intronic and synonymous variant, the investigation can progress to the decision point. If the variant is an intronic and synonymous variant, it is classified as such, and if the variant is not an intronic and synonymous variant, it is junk.
[0149] For inherited, LoF-coding, SNV variants, a broad category such as "prohibited" can be used, where transplant embryos will not be considered, and the variants can be weighted to help rank the embryos (Figure 6). The workflow can include classifying the variants through the following decision points: minor allele frequency (MAF) < X%; variants according to the American College of Medical Genetics (ACMG) criteria (e.g., ACMG 73); variants of Online Mendelian Inheritance in Man (OMIM) disease types; and high constraint (pli > XX). For the MAF < X% decision point, the classification can further evaluate variants with a minor allele frequency below a certain percentage (X%) and panic or continue classification. The ACMG decision point can further include considering known LoF mechanisms, and if so, weighting is performed. The OMIM disease type decision point can further include considering known LoF mechanisms, X-linked, autosomal recessive, autosomal dominant, variant penetrance / severity / family history (FHx). If the variant is X-linked, the classification can further include considering X-linked dominant inheritance (XLD) and males (if so, prohibited). If it is XLD, consider variant penetrance, severity, or FHx. In the weighting at this stage, females who inherit XLR from their mothers may be considered to have caused this variant or XLD with variable penetrance. If the embryo is male, the workflow can further include considering whether the variant is inherited from the father, and if so, arriving here may be the result of XLR inherited from the father, or XLD that should always cause disease, but the parents survived, so it is classified as panic. Similarly, if the investigation leads to high constraint, thenIt is not an OMIM disease type and may have LoF inherited from the parent, which is quite unique.
[0150] For variants known to be pathogenic (e.g., from the ClinVar database), "ban," weighting, and manual review can be used (Figure 7). Classification may include assessing variant characteristics through decision points of X-linked (XL), autosomal recessive (AR), and autosomal dominant (AD). If the variant is XL, the classification may further include considering whether it is XLR and in females. If yes, it is weighted, and if not, it is banned. If the variant is AD, the classification may further include considering variable penetrance or severity, family history (FHx or Fam Hx), and whether the variant is a new variant to reach the possibility of banning, weighting, or manual review. If the variant is AR, the workflow may proceed to the AR workflow described in more detail in Figure 8 below.
[0151] Referring now to Figure 8, in order for a variant to reach this workflow, it has been determined that it is "pathogenic" in the AR gene. For these pathogenic variants in genes associated with autosomal recessive (AR) disorders, it can be determined whether multiple variants exist in a given gene, whether these variants are cis or trans, or whether the variants are present in the gene in a homozygous or hemizygous state. Embryos can be broadly classified as “prohibited” for weighting or manual review. More specifically, the decision point for whether a variant is homozygous or hemizygous can further include considering whether it is also homozygous or hemizygous in the parent, and then trans (checking whether multiple variants in the gene all come from one parent (“cis”), or from different parents (“trans”). This assumes that almost all recessive genes are fully penetrating, so if it is not a living parent, it is likely not pathogenic. If it is present in two parents, or if it is a variant on the X chromosome of a boy, it is prohibited, and it is homozygous inherited from two parents (blood relatives), or hemizygous in a boy. If the decision point further considers whether it is on the X chromosome in boys, and the answer is no, then manual review will be required. This type of variant shows non-Mendelian inheritance and may overlap with deletions; it may also be deleted in the parents or have a missing allele in the child. The decision point of whether there is >1 variant in the gene can further include considering whether the father has >0 variants, and if so, whether the mother has >0 variants. If so, it is forbidden. This is the most common type of autosomal recessive inheritance. If the father does not have >0 variants, the classification can further include considering whether the mother has >0 variants. If the mother does not have >0 variants, then consider whether there are >0 variants with unknown genetic characteristics. Reaching this point indicates there is 1.A genetic variant and one variant with unknown genetic characteristics (possibly newborn). If the variant cannot be phased, the variant is likely to be prohibited, but this would be a rare event.
[0152] Figure 9 shows the classification of newborn LoF-coded variants (SNVs) into the broad categories shown above. More specifically, the classification may include the following decision points: ACMG; OMIM disease type and high constraint.
[0153] Referring now to Figures 10A-D, the classification may include newborn and >XX MBp as well as decision points for overlapping coding regions of genes with SV deletions (Figure 10A). In SV acquisition in Figure 10B, the classification may include decision points for the start or end coding regions of genes and triploid-sensitive regions. For inversion calls in Figure 10C, the classification may include newborn and decision points for the start or end coding regions of genes. Finally, for translocation calls in Figure 10D, the classification may include newborn and decision points for the start or end coding regions of genes.
[0154] Figure 11 illustrates the classification of embryos into prohibited or manually reviewed categories based on known pathogenicity and the length of the "pre-mutation" repeat amplification, as well as the repeat length in the parent.
[0155] Figures 12A-D illustrate the classification of homozygous large regions, uniparental allelic disomies, synonymous and intronic variants, and variants identified from the spinal muscular atrophy pipeline. More specifically, in Figure 12A, large regions of homozygous variants can be classified as normal treatment (recessive disease), normal treatment (deletion), and manually reviewed. In Figure 12B, synonymous and intronic variants can be classified as LoF. In Figure 12C, large regions of uniparental genetic characteristics can be classified as prohibited or manually reviewed. Finally, Figure 12D is a flowchart illustrating the potential classification of variants identified, for example, from the spinal muscular atrophy pipeline.
[0156] Example 6: Genetic Analysis and Embryo Sequencing Figures 13-14C illustrate the classification of SNVs or small insertion / deletion variants into various major categories through DeepVarian, VCF information, population frequency, reference sequence, phenotypic classification, and triple inheritance steps.
[0157] More specifically, referring to Figure 13, the initial steps of DeepVarian, VCF information, and population frequency are described in detail. The analysis can potentially lead to trash or late-expanding major categories. It is worth noting that population frequency analysis may include, for example, queries to the 1KG and Gnomad databases.
[0158] Referring now to Figure 14A, in the reference sequence step, classification can be performed through decision points such as coding variant, intron variant, intergenic variant, variant is meaningless or LoF, variant is missense, variant is synonymous, variant splicing (X is 4), variant splicing (clinical X is 4), and pathogenicity evidence (ClinVar), which can potentially lead to trash, trash, or late-expanding major categories.
[0159] Referring to Figure 14B, in the phenotypic classification step, classification can be performed through decision points of ACMG criterion compliance, ClinVar pathogenicity criterion compliance, OMNIM criterion compliance, and embryonic "Intolerome" criterion compliance, which can potentially lead to large classes of waste or later expansion.
[0160] Referring to Figure 14C, in the triple inheritance step, classification can be performed through decision points of neonatal dominant, genetic dominant, homozygous recessive, X-linked, genetic complex heterozygote, and artificial review capture net, which can potentially lead to large classes of waste or later expansion. It is worth noting that neonatal dominant, X-linked, genetic dominant, homozygous recessive, and genetic complex heterozygote each have sub-flowcharts described in detail in Figures 16A-E.
[0161] Referring to Figure 15, in the report / mask subset step of the process, classifications of various pipelines may cluster. In this embodiment, the category arriving from the SNV annotation pipeline can be determined through decision points on the mask list and the report list, which can lead to a major category that is spam, manually reviewed, or expanded later before the report is finalized.
[0162] Referring to Figure 16A, in this sub-flowchart for the neonatal dominant decision point, analysis can be performed through decision points for neonatal (non-parental), from parental, dominant mechanism, and recessive mechanism, which can potentially lead to a major category that is spam or sees other pipelines.
[0163] Referring to Figure 16B, in this sub-flowchart for the X-linked decision point, analysis can be performed through decision points for non-PAR zone, proband male and hemizygote, and proband female and homozygote, which can potentially lead to a major category that is manually reviewed or sees other pipelines.
[0164] Referring to Figure 16C, in this sub-flowchart for the genetic dominant decision point, analysis can be performed through decision points for newborn (non-parental), parental, (P1)-Het / (P2)-WT / (E)-Het, rare (P1 / 2) and (E)-Het, dominant mechanism and recessive mechanism, which can potentially lead to a class of panic or seeing other pipelines.
[0165] Referring to Figure 16D, in this sub-flowchart for the homozygous recessive decision point, analysis can be performed through decision points for newborn (non-parental), parental, (P1)-Het / (P2)-Het / (E)-Homo, (P1)-Het / (P2)-WT / (E)-Homo, rare (P1 / 2) and (E)-Homo, dominant mechanism and recessive mechanism, which can potentially lead to a class of artificial censorship, panic or seeing other pipelines.
[0166] Referring to Figure 16E, in this sub-flowchart showing a genetic complex heterozygote, analysis can be performed using genes with >= 2 SNVs, phasing SNVs onto a separate chromosome, phasing of unknown SNVs, phasing of phased SNVs LoF / LoF, and phasing of SNVs.Decision points for LoF / Other Pathogenicity, SNV LoF / LoF, and SNV Lof / Other Pathogenicity are performed, which can potentially lead to large categories of junk, manual review, and panic.
[0167] Example 7: Aneuploidy Call Figures 17A and 17B show examples of aneuploidy calls on single cell sets from two different embryos. Analysis shows the difference between intercellular dissimilar signals and intercellular congruent signals. All cells in each figure are from the same embryo. Cells labeled “B” are isolated from tissue biopsies, cells labeled “M” are isolated from the inner cell mass, and samples labeled “R” are large samples of the remaining cells from each embryo (to show congruence). Copy number calls were performed using CNVkit software based on a custom reference set consisting of single-cell sequences amplified using the same WGA method. Figure 17A is an aneuploid female embryo showing loss of congruence of chromosome 19. Figure 17B shows an euploid female embryo in which some copy number variants are present in a single cell, but no variants are present consistently throughout the embryo.
[0168] Discussion: The genetic analysis and embryo screening methods described above and disclosed herein involve using comprehensive genomic data and analysis, along with the most sensitive screening criteria, to identify embryos with the best chance of term pregnancy and the lowest postnatal genetic risk among the 6 billion bases of the human genome. The analysis in this disclosure will be more thorough because it includes studying constraint genes not currently considered pathogenic factors, small CNVs below the detection limits of most aneuploidy screenings, structural variants, and intergenic and intronic mutations. Because multiple single cells are sequenced individually and then combined, rather than pooling cells first, there will be fewer areas with low or no coverage. Furthermore, the genetic information from the embryo will be compared with genetic information from one or more parents and / or one or more genetic relatives. This will allow for the differentiation of predicted artifacts from actual mutations, as any variant in multiple cells could be an actual mutation. Specification page 17 / 18 21 CN 121100382 A
[0169] Some key challenges associated with single-cell whole-genome sequencing include coverage of the entire genome, reproducibility between responses, and measurement of accuracy. The method of this disclosure effectively answers these challenges. First, by using multiple sources (cell and / or cell-free sources) to obtain embryonic data and then merging the data, the method of this disclosure can achieve better coverage of the entire genome, thereby improving accuracy through depth and comparison. In some embodiments, for example, sequencing data of individual cells are used to call variants, and then the results are combined for analysis. In some embodiments, for example, sequencing data of individual cells are first combined and then used to call variants. Second, by using a multiplex WGA method, this disclosure...The method disclosed herein compensates for biases in each method. Third, by using data from multiple embryo sources as well as parental / family genomic data to identify neonatal variants, the method of this disclosure improves the accuracy of variant recall.
[0170] Genetic analysis of multiple embryos during in vitro fertilization (IVF) requires a method to compare the disease risk of each embryo in order to select one or more embryos for transfer and implantation. Each embryo has different (if overlapping) genetics and therefore different disease risks. The method of this disclosure involves the automated classification of variants and subsequent embryo sequencing, which has never been done or suggested before. It is easy to act on all genetic variants before implantation, i.e., the parents may choose not to transfer an embryo due to genetic susceptibility to a disease. This is especially true for adult-onset diseases, for example, BRCA2 mutations may lead to breast cancer within 30 to 50 years, but there are still treatments and potentially improved treatments in the future. It may be necessary to evaluate this risk in the presence or absence of childhood disease. Therefore, genetic information from the entire genome must be considered. Automation enables rapid and cost-effective embryo sequencing. The sequencing process and report generation summarize the whole-genome analysis to enable parents to select embryos that meet their desired goals after receiving the analysis results provided by the methods described above and disclosed herein.
[0171] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference in their entirety to the same extent that each individual publication, patent or patent application is specifically and individually indicated to be incorporated herein by reference. Furthermore, any reference cited or identified in this application should not be construed as an admission that such reference is prior art to the invention. The use of section headings should not be construed as a necessary limitation.
[0172] Although the invention and its advantages have been described in detail, it should be understood that various changes, substitutions and alterations may be made herein without departing from the spirit and scope of the invention as defined in the appended claims.
[0173] It will be readily understood by those skilled in the art that the invention is well adapted to achieve the stated objects and obtain the stated objects and advantages, as well as the objects and advantages inherent therein. The embodiments and methods described herein represent preferred embodiments and are exemplary and are not intended to limit the scope of the invention. Variations and other uses therein will occur to those skilled in the art and are included within the spirit of the invention as defined by the scope of the claims. Instruction manual 18 / 18 page 22 CN 121100382 A Figure 1 Instruction manual drawing 1 / 28 page 23 CN 121100382 A Figure 2 Instruction manual drawing 2 / 28 page 24 CN 121100382 A Figure 3 Instruction manual drawingPage 3 / 28, 25 CN 121100382 A, Figure 4, Instruction Manual Appendix; Page 4 / 28, 26 CN 121100382 A, Figure 5, Instruction Manual Appendix; Page 5 / 28, 27 CN 121100382 A, Figure 5, Continued from above, Instruction Manual Appendix; Page 6 / 28, 28 CN 121100382 A, Figure 6, Instruction Manual Appendix; Page 7 / 28, 29 CN 121100382 A, Figure 6, Continued from above, Instruction Manual Appendix; Page 8 / 28, 30 CN 121100382 A, Figure 7, Instruction Manual Appendix; Page 9 / 28, 31 CN 121100382 A, Figure 8, Instruction Manual Appendix; Page 10 / 28, 32 CN 121100382 A, Figure 8, Continued from above, Instruction Manual Appendix; Page 11 / 28, 33 CN 121100382 A, Figure 9, Instruction Manual Appendix; Page 12 / 28, 34 CN 121100382 A Figure 10A Figure 10B Instruction Manual Drawings 13 / 28 Page 35 CN 121100382 A Figure 10C Figure 10D Instruction Manual Drawings 14 / 28 Page 36 CN 121100382 A Figure 11 Instruction Manual Drawings 15 / 28 Page 37 CN 121100382 A Figure 11 Continued from above Figure Instruction Manual Drawings 16 / 28 Page 38 CN 121100382 A Figure 12A Figure 12B Instruction Manual Drawings 17 / 28 Page 39 CN 121100382 A Figure 12C Figure 12D Instruction Manual Drawings 18 / 28 Page 40 CN 121100382 A Figure 13 Instruction Manual Drawings 19 / 28 Page 41 CN 121100382 A Instruction Manual Drawings 20 / 28 Page 42 CN 121100382 A Instruction manual figures 21 / 28, page 43, CN 121100382 A, Figure 15; Instruction manual figures 22 / 28, page 44, CN 121100382 A, Figure 16A, Figure 16B; Instruction manual figures 23 / 28, page 45, CN 121100382 A, Figure 16C; Instruction manual figures 24 / 28, page 46, CN 121100382 A, Figure 16D; Instruction manual figures 25 / 28, page 47, CN 121100382 A, Figure 16E; Instruction manual figures 26 / 28, page 48, CN 121100382 A, Figure 17A; Instruction manual figures 27 / 28, page 49, CN121100382 A Figure 17B Instruction Manual Appendix 28 / 28 Page 50 CN 121100382 A
Claims
1. A method of identifying a genetic variant in an embryo, the method comprising: (a) obtaining two or more sources of analyte from the embryo; (b) analyzing the two or more sources of analyte to obtain genetic information for each source; (c) comparing the genetic information for each source to one or more reference genomes using at least one variant caller, wherein the variant caller identifies variants between each source and the reference genome; and (d) combining the variants to identify a difference that is presented by only the source, wherein the difference is a genetic variant in the embryo.
2. The method of claim 1, wherein the genetic variant is a single nucleotide variant (SNV), a multi-nucleotide variant (MNV), a copy number variant (CNV), a structural variant (SV), or an alteration in an epigenetic marker.
3. The method of claim 2, wherein the epigenetic marker is a DNA methylation marker used to identify defects or transcriptional dysregulation.
4. The method of claim 2, wherein the genetic variant is identified as a de novo variant (DNV).
5. The method of claim 1, wherein the two or more sources of analyte comprise cells from an embryo or blastocyst tissue biopsy, cells and / or genetic material isolated from cell culture media, blastocyst cavity, or a combination thereof.
6. The method of claim 5, wherein the tissue biopsy is collected using laser pulse assisted pipette delivery.
7. The method of claim 6, further comprising isolating two or more cells from the tissue biopsy and dissociating the two or more cells using a physical or enzymatic method.
8. The method of claim 7, wherein the physical method comprises laser cutting or micropipette isolation.
9. The method of claim 7, wherein the enzymatic method comprises dissociation using a digestive enzyme, wherein the enzyme is selected from the group consisting of: dispase, collagenase, hyaluronidase, papain, DNase-I, Accutase, and trypsin.
10. The method of claim 1, wherein the genetic information comprises nucleic acid sequences of a panel, exome, whole genome, or transcriptome, with or without mitochondrial sequences and with or without epigenetic markers.
11. The method of claim 10, wherein the nucleic acid comprises DNA, RNA, or both.
12. The method of claim 11, wherein the nucleic acid is DNA, and wherein the DNA sequences are obtained by the steps of: (a) generating a DNA library from each source; (b) performing low-pass whole genome sequencing on the DNA; (c) applying an aneuploidy filter to the sequencing data from (b) to eliminate embryos with aneuploidy; and (d) performing additional sequencing on the DNA from the embryos that were not eliminated to generate sequence data files for the embryos.
13. The method of claim 12, wherein application of the aneuploidy filter results in detection of chromosomal regions with unexpected increases or decreases in coverage compared to expected coverage based on total sequence coverage, known coverage biases based on local genomic composition, and the sex of the embryo.
14. The method of claim 11, wherein RNA is isolated, amplified, and sequenced.
15. The method of claim 14, wherein RNA sequence data is used to characterize the transcriptome and determine gene expression levels.
16. The method of claim 15, wherein gene expression levels are correlated with patient outcomes.
17. The method of claim 16, wherein the patient outcomes include successful pregnancy.
18. The method of claim 16, wherein genetic information and / or gene expression is correlated with patient outcomes.
19. The method of claim 14, wherein RNA sequence data is used to call genomic variants in conjunction with DNA sequencing or alone.
20. The method of claim 14, wherein RNA sequence data is used to identify aberrant splicing events and transcriptional dysregulation, and one or more variants affecting one or more transcriptional pathways.
21. The method of claim 20, wherein the aberrant splicing events are cryptic splicing, aberrant exon usage, or upstream open reading frames.
22. The method of claim 14, wherein RNA sequence data is used to identify allele-biased expression, thereby identifying imprints, copy number determinations, and / or variants affecting expression.
23. The method of claim 1, wherein the reference genome is a genomic sequence from a public or private repository, one or more parents, or one or more genetic relatives.
24. The method of claim 23, wherein the genetic variants are further identified by comparing the genetic information from each source to genetic information of one or more parents to infer haplotype phasing, impute missing data, determine inheritance or lack thereof, and identify de novo variants.
25. The method of claim 24, further comprising annotating the genetic variants for correlation with phenotypes by using a database of variants associated with disease, phenotype, disease risk or protection, and traits.
26. The method of claim 24, further comprising annotating the genetic variants for likelihood of deleterious impact on gene function by using a database of gene function or biological pathways or by in silico, in vivo, or in vitro methods.
27. The method of claim 1, wherein the variant caller is a deep learning-based variant caller.
28. The method of claim 27, wherein the deep learning-based variant caller is constructed based on genetic information of one or more parents or genetic relatives and is capable of identifying de novo variants for the purposes of preimplantation genetic screening.
29. The method of claim 1, wherein the genetic variants occur in the genome known or suspected to contribute to an adult or childhood genetic disease, or occur in the genome without a currently relevant genetic disease but are constrained by inheritance.
30. The method of claim 1, wherein the step of combining comprises identifying genetic variants that occur in at least two sources, or identifying genetic variants that can only be detected by combining genetic information from all sources.
31. The method of claim 30, wherein the genetic variants are identified in multiple sources.
32. The method of claim 30, wherein the genetic variants are identified as chimeras.
33. A method of screening a plurality of embryos, the method comprising: (a) performing a genetic analysis on each embryo according to any of the preceding claims; and (b) ranking the embryos based on the results of the genetic analysis, wherein the embryo with the best chance of term pregnancy and lowest lifetime genetic risk is ranked highest for transfer and implantation.
34. The method of claim 33, wherein the embryos are ranked based on known deleterious or benign variants, or based on variants predicted to be deleterious or benign according to various methods.
35. The method of claim 34, wherein the method comprises computer deleterious prediction tools, allelic evolutionary conservation, and / or allelic frequency in selected or unselected population databases.
36. The method of claim 33, wherein the impact of genetic variants on embryo ranking is weighted from customer-based phenotypic input.
37. The method of claim 36, wherein the input is a desire to avoid a particular pathological condition.
38. The method of claim 33, wherein the impact of genetic variants on embryo ranking is weighted based on automatic phenotypic overlap between a disease-related phenotype and its associated genes and phenotypes provided by the customer.
39. The method of claim 38, wherein the phenotype stems from a family history of one or both parents.
40. The method of claim 38, wherein the overlap is computed using ontology, which allows related phenotypes to have an impact on embryo ranking.
41. The method of claim 33, wherein the impact of genetic variants on embryo ranking is weighted based on known disease severity, penetrance, and / or genetic mode of inheritance associated with the gene.
42. The method of claim 33, wherein the impact of genetic variants on embryo ranking is weighted based on presence or absence of the variant in one or both parents and zygosity in the parents.
43. The method of claim 33, wherein the impact of genetic variants on embryo ranking is weighted based on known or suspected age of onset.
44. The method of claim 33, wherein the impact of genetic variants on embryo ranking is weighted based on pre-ranked gene or variant ranking.
45. The method of claim 33, wherein the impact of genetic variants on embryo ranking is weighted based on treatment options for a given disease.
46. The method of claim 33, further comprising generating a report of the plurality of embryos, wherein the ranked embryos are qualitatively classified based on a recommendation to implant, not to implant, or to implant with uncertainty.