Method for identifying genetic variant in an embryo
Patent Information
- Application Number
- EP2024747033
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-26
- Filing Date
- 2024-01-26
- Publication Date
- 2025-12-03
AI Technical Summary
Current prenatal and preimplantation genetic testing methods are inadequate in detecting genetic variants in embryos, leading to high miscarriage rates and increased risks of genetic disorders, as they focus on aneuploidy and known genetic disorders, failing to identify potentially lethal de novo mutations and rare variants that can cause severe diseases.
A method involving the analysis of multiple sources of analytes from embryos, comparing genetic information against reference genomes using variant callers to identify genetic variants such as SNVs, CNVs, and structural variants, and combining data to detect de novo mutations, which can only be present in the embryo, thereby providing a comprehensive assessment of genetic health before transfer.
This approach enables accurate identification of genetic variants, allowing for the selection of embryos with the lowest genetic risk for transfer, potentially reducing miscarriage rates and the risk of inherited diseases, and providing actionable genetic information for parents.
Smart Images

Figure IB2024050754_02082024_PF_FP
Abstract
Description
Attorney Docket No.21-2014-WO METHOD FOR IDENTIFYING GENETIC VARIANT IN AN EMBRYO CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application Serial Number 63 / 441,291, filed January 26, 2023, the contents of which are hereby incorporated by reference in their entirety. FIELD OF THE INVENTION
[0002] The present disclosure generally relates to the field of reproduction. More specifically, the present disclosure relates to methods of identifying genetic variant(s) in embryos to assess disease risk. BACKGROUND OF THE INVENTION
[0003] Currently available prenatal testing, carrier screening and preimplantation genetic testing (“PGT”) typically focus on the detection of aneuploidy and known genetic disorders for which the parents are carriers, but are unable to detect the majority of disease-causing genetic variants in an individual embryo. Even when families are able to rule out one or a handful of known genetic risks, a fertilized egg still carries potentially lethal genetic defects.
[0004] Natural conception is fraught with inherent developmental challenges. Nearly half of fertilized eggs carry a potentially lethal genetic mutation that could strike during pregnancy, childhood, or in adulthood. As a result, approximately 1 in 3 pregnancies end in miscarriage (Wilcox AJ, et al. N. Engl J Med, 1988; 319: 189-194), approximately 1 in 14 children are born with a genetic defect (Ceyhan-Birsoy O, et al. Am J. Human Genet, 2019; 104: 76-93), and approximately 1 in 20 adults harbor genetic disease variants that can dramatically increase the risk of many cancers, sudden cardiac death, aneurysms, autoimmune disorders and neurodegenerative diseases like Huntington’s (Abul-Husn NS, et al. Science, 2016; 354: 6319).
[0005] Rare and high impact variants, not polygenic risk, drive clinical decision making. Although it is a good way to understand populations, polygenic risk is a poor way to understand individuals or the developmental success of a single fertilized egg. For instance, for a common disease like breast cancer, there is ~15% lifetime risk. Women in the top 1% of polygenic risk scores (“PRS”) have a 31% lifetime risk of breast cancer, which is 2 x normal. Whereas, rare BRCA1 / 2 mutation carriers have up to 85% lifetime risk of getting breast cancer (6 x normal). For a very rare disease, such as brain cancer, there is about 0.01% of lifetime risk. Common variants cause 1.2-3.5x lifetime risk (Melin et al., 2017, Nat. Genet.49(5): 789-794), whereas rare variants impart a 5,000x lifetime risk (Bainbridge, et al., 2015, J. Natl. Cancer Inst.107(1): 1-4).
[0006] Most clinical tests like carrier screening and preimplantation genetic testing ( including PGT-Monogenic disease (PGT-M) and PGT-A aneuploidy) use a targeted approach to balance 1 93176811Attorney Docket No.21-2014-WO sensitivity and specificity - to avoid missing disease (i.e., avoid “false negative”), while avoiding “false positive” scares or discarding potentially viable embryos. PGT-A accounts for ~50% of in vitro fertilization (“IVF”) cycle failure and pregnancy loss (Zhao C, et al. 2021, Gen in Med. 23:435-442). While PGT next-generation sequencing (NGS) can diagnose whole chromosome aneuploidy (WCA) in 95% of embryos, it does not accurately detect sub-chromosomal events (Cascante SD, et al.2023, Fert. Ster.120(6):1161-1169). Carrier tests report on 200 - 800 genes while PGT-M may report on only 1 or 2 genes. Prenatal tests like non-invasive prenatal testing (“NIPT”) focus on the most certain genetic errors that are easiest to accurately detect. Most NIPT assays report on 5-13 conditions. However, it is noted that the only action that can be taken based on prenatal tests is to keep or end a pregnancy. Preimplantation is when genetic findings are most actionable. For instance, BRCA2 mutation carriers have a 70% lifetime risk of developing breast cancer. The means for mitigating risk is dependent on life-stage. Diagnosis during adulthood would reduce risk by increased monitoring or performing prophylactic surgery. Diagnosis at prenatal stage is not offered at this time because ending the pregnancy would be the only possible mitigation breast cancer risk. In contrast, embryo screening would provide no risk if parents choose to transfer a different embryo. Whole Exome Sequencing (WES) allows for comprehensive screening of the protein-coding regions (exome) of an embryo’s genome. While the exome represents only a small fraction of the entire genome (~2%), it contains many known disease-causing variants. Recent studies that applied exome sequencing showed that of all sporadic cases tested, between 60 and 75% could be explained by de novo mutations, i.e., changes that are not found in either parent (Acuna-Hidalgo R, et al.2016, Genome Biology, 17:241). In general, every embryo that is formed after fertilization has about 100 de novo genetic changes. The older the parents, the higher the tendency to accumulate changes in their sperm and eggs that can be passed down to the embryo. Interestingly, it has been found that trio testing (a parent- offspring trio) increases the probability of successful genomic diagnosis of a rare pediatric disease by nearly five times (Wright CF, et al. 2023, N Engl. J. Med. 388:1559-1571). The same study found that out of 3599 afflicted individuals that were tested in trios who received a diagnosis, approximately 76% had a pathogenic de novo variant (Id.). As such, there is a need for accurate embryo screening methods that include analysis of more pregnancy success determinants prior to embryo transfer and implantation to provide accurate assessment of genetic health and embryo outcome through development and beyond. 2 93176811Attorney Docket No.21-2014-WO SUMMARY OF THE INVENTION
[0007] The present disclosure provides, in part, methods for genetic screening of embryos, particularly prior to embryo transfer and implantation, wherein individual sources of analytes are analyzed and compared against reference genome to identify genetic variant(s) in the embryos.
[0008] Accordingly, one aspect of the present disclosure provides a method of identifying a genetic variant in an embryo. This method involves obtaining two or more sources of analytes from the embryo; analyzing the two or more sources of analytes to obtain genetic information of each source; comparing the genetic information of each source against one or more reference genome(s) using at least one variant caller that identifies variant(s) between each source and the reference genome; and combining the variant(s) to identify a difference that is present only from the sources, thus identifying a genetic variant in the embryo.
[0009] In some embodiments, the genetic variant is a single nucleotide variant (SNV), a multi- nucleotide variant (MNV), a copy number variant (CNV), a structural variant (SV), or an alteration in an epigenetic marker. By way of non-limiting example, the epigenetic marker is a DNA methylation marker used for identifying imprinting defects. By way of non-limiting example, the genetic variant is identified as an apparent De novo variant (DNV).
[0010] In some embodiments, the two or more sources of analytes include cells biopsied from the embryo or blastocyst, cells and / or genetic material isolated from cell culture media, blastocoel, or a combination thereof.
[0011] In some embodiments, the biopsy is collected using laser pulse aided pipette delivery. In some embodiments, two or more cells are isolated from the biopsy and separated using a physical or enzymatic method. By way of non-limiting example, the physical method may include laser dissection or micropipette isolation. By way of non-limiting example, the enzymatic method may include extraction using a digestive enzyme, including, but not limited to, Dispase, Collagenase, Hyaluronidase, Papain, Dnase-I, Accutase, or Trypsin.
[0012] In some embodiments, the genetic information obtained from each source may include nucleic acid sequence of panel, exome, whole genome, or transcriptome with or without mitochondrial sequence and with or without epigenetic markers. By way of non-limiting example, the nucleic acid may include DNA, RNA or both.
[0013] In some embodiments, the nucleic acid sequence is obtained by isolating DNA from the source(s), generating libraries from whole genome amplification of DNA from each source and sequencing the libraries to generate a sequence data file of the source.
[0014] In some embodiments, prior to sequencing the libraries fully, embryos may be sequenced to lower coverage and aneuploidies may be identified. Embryos may subsequently be eliminated 3 93176811Attorney Docket No.21-2014-WO from further sequencing. By way of non-limiting example, aneuploidies may be detected by unexpectedly high or low coverage in a genomic region.
[0015] In some embodiments, RNA is isolated and sequenced. In some embodiments, RNA sequence data is used to characterize the transcriptome and determine gene expression levels. In some embodiments, gene expression levels are correlated to patient outcomes. By way of non- limiting example, the patient outcomes include, but are not limited to, successful pregnancy.
[0016] In some embodiments, a panel of genetic information correlating gene expression levels and patient outcomes is established.
[0017] In some embodiments, RNA sequence data is used to identify an aberrant splicing event thereby identifying dysregulation of transcriptional regulation and one or more variants that impact one or more transcriptional pathways. By way of non-limiting example, the aberrant splicing event may include cryptic splicing, aberrant exon usage, or upstream open reading frames.
[0018] In some embodiments, RNA sequence data is used to identify biased allele expression thereby identifying imprinting, copy number determination, and / or variants that affect expression.
[0019] In some embodiments, RNA sequence data is used to call genomic variants in combination with DNA sequencing or alone.
[0020] In some embodiments, the reference genome is a genome sequence from a public or private repository, one or more parents, a genetic relative and / or an embryo generated from the same parents.
[0021] In some embodiments, the genetic variant is further identified by comparing the genetic information of each source to genetic information of one or more parents to deduce haplotype phasing, impute missing data, determine inheritance or lack thereof, and identify apparent De novo variants. In some embodiments, the genetic variant is further annotated for its association with a phenotype by using a database of variants associated with diseases, phenotypes, disease risk or protection and traits. In some embodiments, the genetic variant is further annotated for its potential to cause deleterious effects on gene function through in silico, in vivo, or in vitro methods.
[0022] In some embodiments, the variant caller is a deep learning-based variant caller. In some embodiments, the deep learning-based variant caller is built based on genetic information of one or more parents, a genetic relative and / or an embryo generated from the same parents, and is capable of identifying an apparent De novo variant for the purpose of preimplantation genetic screening. In some embodiments, the genetic variant occurs in the genome with a known or suspected contribution to an adult or childhood genetic disease or in the genome with no current associated genetic disease but under genetic constraint.
[0023] In some embodiments, the combining step of the method described above and herein comprises identifying a genetic variant that occurs in at least two sources, or identifying a genetic 4 93176811Attorney Docket No.21-2014-WO variant that can only be detected by combining genetic information from multiple sources. By way of non-limiting example, the genetic variant is identified in multiple sources, or determined to be error, artifact or mosaicism.
[0024] Some aspects of the present disclosure provide a method of screening multiple embryos. Such method comprises conducting genetic analysis on each embryo according to the method described above and herein; and then ranking the embryos based on the results from the genetic analysis. The embryo with the best chance of a term pregnancy and the lowest lifetime genetic risk is ranked the highest for transfer and implantation. In some embodiments, embryos are ranked based on known deleteriousness of detected variants, or based on predicted deleteriousness of detected variants where such prediction is made by various methods. By way of non-limiting example, the predicting methods may include in silico deleteriousness prediction tools, allele evolutionary conservation, and / or allele frequency in a selected or unselected population database.
[0025] In some embodiments, a genetic variant’s impact on embryo rank is weighted based on input from a phenotype selected for examination. By way of non-limiting example, the input may be the desire of a customer to avoid a particular pathologic condition.
[0026] In some embodiments, a genetic variant’s impact on embryo rank is weighted based on automated phenotypic overlap between the phenotypes associated with a disease, and its associated gene and phenotypes selected for examination. Overlap may be computed using an ontology that allows related, as opposed to strictly identical terms, to determine overlap. By way of non-limiting example, the phenotypes selected for examination are derived from family history of one or both parents where parents may wish to avoid breast cancer risk, and thereby genes associated with risk of other cancers will also be weighted more heavily.
[0027] In some embodiments, a genetic variant’s impact on embryo rank is weighted based on variant zygosity, the presence of other variants in the genome, the predicted deleteriousness of a variant, whether the variant has been precurated as benign, unknown, likely pathogenic or pathogenic as well as known disease severity, penetrance, and / or inheritance pattern associated with a gene.
[0028] In some embodiments, a genetic variant’s impact on embryo rank is weighted based on presence or absence of the variant in one or both parents and its zygosity in the embryo or the parents.
[0029] In some embodiments, a genetic variant’s impact on embryo rank is weighted based on known or suspected age of onset of disease.
[0030] In some embodiments, a genetic variant’s impact on embryo rank is weighted based on pre-curated gene or variant rankings. 5 93176811Attorney Docket No.21-2014-WO
[0031] In some embodiments, a genetic variant’s impact on embryo rank is weighted based on treatment options for a given disease.
[0032] In some embodiments, a genetic variant’s impact on embryo rank is weighted based on its impact on gene function or expression.
[0033] In some embodiments, the screening method further comprises generating a report with the ranking of the embryos, wherein ranked embryos are qualitatively categorized based on recommendation to transfer, not to transfer, or unsure to transfer. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to better understand the subject matter that is disclosed herein and to exemplify how it may be carried out in practice, embodiments will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which:
[0035] FIG.1 illustrates carrier screening vs. reproductive risk.
[0036] FIG. 2 is a flow chart outlining whole genome sequencing (WGS) and whole genome amplification (WGA) of cell lines, wherein analytes from individual cells and analytes from bulk of cells are analyzed and compared.
[0037] FIG.3 is a flow chart outlining a method of WGS and WGA and subsequent identification of genetic variants in an embryo sample, wherein multiple sources of analytes are analyzed.
[0038] FIG. 4 is a flow chart outlining a method of identifying genetic variants in an embryo (using WGA and WGS) taking into consideration of parental genetic data.
[0039] FIG. 5 is a flow chart illustrating variant classification into broad categories, including known pathogenic variants, non-coding variants, coding variants, loss of function (LoF) variants, synonymous variants and variants that will not warrant further consideration (Trash) as well as “Panic” conditions where Manual review will be required.
[0040] FIG.6 is a flow chart illustrating variant classification for inherited, coding, LoF variants into broad categories such as “No go” wherein an embryo would not be considered for transfer as well as Weighting of variants to aid in ranking an embryo.
[0041] FIG.7 is a flow chart illustrating variant classification of known pathogenics (e.g., from ClinVar database) into broad categories such as “No go” Weighting and Manual review.
[0042] FIG. 8 is a flow chart illustrating variant classification for “pathogenic” (i.e., likely, known, or possibly disease-causing) variants in genes associated with autosomal recessive disease (“autosomal recessive process”). It is determined whether multiple variants exist in a gene and whether those variants are in cis or trans, or whether a variant exists as a homozygous or hemizygous state in a gene to broadly categorize the embryo as “No go”, to weight it, or under-go Manual review. 6 93176811Attorney Docket No.21-2014-WO
[0043] FIG. 9 is a flow chart illustrating variant classification of a De novo LoF coding variant into broad categories as previously indicated.
[0044] FIG. 10A is a flow chart illustrating structural variant classification into “No go” for embryo or to reclassify as LoF in the appropriate gene for SV Deletion.
[0045] FIG. 10B is a flow chart illustrating structural variant classification into “No go” for embryo or to reclassify as LoF in the appropriate gene for SV Gain.
[0046] FIG. 10C is a flow chart illustrating structural variant classification into “No go” for embryo or to reclassify as LoF in the appropriate gene for Inversion.
[0047] FIG. 10D is a flow chart illustrating structural variant classification into “No go” for embryo or to reclassify as LoF in the appropriate gene for Translocation.
[0048] FIG. 11 is a flow chart illustrating repeat expansion categorization based on known pathogenic and “pre-mutation” lengths of repeat expansions as well as the repeat length in the parents to broadly categorize an embryo as previously stated.
[0049] FIG. 12A is a flow chart illustrating classification of large regions of homozygosity variants to classify as Treat normally (recessive disease), Treat normally (deletion), and Manual review.
[0050] FIG.12B is a flow chart illustrating categorization of synonymous and intronic variants to classify as Treat as LoF.
[0051] FIG. 12C is a flow chart illustrating categorization of large regions of uniparental inheritance to classify as No go or Manual Review.
[0052] FIG.12D is a flow chart illustrating a potential categorization of variants identified from, for example, a spinal muscular atrophy pipeline.
[0053] FIG.13 is a flow chart illustrating an example of the categorization of SNV or small indel variants into various broader categories. The initial steps of DeepVariant, VCF Info, and Population Frequency are detailed, leading in some instances to the broad categories of Trash or Later Expansion.
[0054] FIG.14A is a flow chart illustrating the Reference Sequence step of the analysis where the categorization may flow through the decision points of Coding Variant, Intronic Variant, Intergenic Variant, Varian is a Nonsense or LoF, Variant is Missense, Variant is Synonymous, Variant Splicing Effect (X of 4), Variant Splicing Effect (Clinical X of 4), and Pathogenic Evidence (ClinVar), which may potentially lead to the broader categories of Panic, Trash, or Later Expansion.
[0055] FIG. 14B is a flow chart illustrating the step of Phenotype Classification in the analysis where the categorization may flow through the decision points of ACMG Criteria Met, ClinVar 7 93176811Attorney Docket No.21-2014-WO Pathogenicity Criteria Met, OMNIM Criteria Met, and Embryo “Intolerome” Criteria Met, which may potentially lead to the broader categories of Trash, or Later Expansion.
[0056] FIG. 14C is a flow chart illustrating the step of Trio Heredity in the analysis where the categorization may flow through the decision points of De Novo Dominant, Inherited Dominant, Homozygous Recessive, X-Linked, Inherited Compound Heterozygous, and Manual Review Catch Net, which may potentially lead to the broader categories of Panic or Later Expansion. Note that De Novo Dominant, Inherited Dominant, Homozygous Recessive, X-Linked, and Inherited Compound Heterozygous decision points each have a sub-flow chart detailed in FIGS.16A-16E.
[0057] FIG.15 is a flow chart illustrating the Reporting / Masking Subsets step of the process.
[0058] FIG.16A is a sub-flow chart for the De Novo Dominant decision point where the analysis may flow through the decision points of De Novo (Not Parental), From Parent, Dominant Mechanism and Recessive Mechanism, which may potentially lead to the broader categories of Panic or See Other Pipeline.
[0059] FIG. 16B is a sub-flow chart illustrating the X-linked decision point where the analysis may flow through the decision points of Non-PAR region, Proband Male and Hemizygous, and Proband Female and Homozygous, which may potentially lead to the broader categories of Manual review or See Other Pipeline.
[0060] FIG.16C is a sub-flow chart illustrating the Inherited Dominant decision point where the analysis may flow through the decision points of De Novo (Not Parental), From Parent, (P1)- Het / (P2)-WT / (E)-Het, Rare (P1 / 2) with (E)-Het, Dominant Mechanism, and Recessive Mechanism, which may potentially lead to the broader categories of Panic or See Other Pipeline.
[0061] FIG.16D is a sub-flow chart illustrating the Homozygous Recessive decision point where the analysis may flow through the decision points of De Novo (Not Parental), From Parent, (P1)- Het / (P2)-Het / (E)-Homo, (P1)-Het / (P2)-WT / (E)-Homo, Rare (P1 / 2 with (E)-Homo, Dominant Mechanism, and Recessive Mechanism, which may potentially lead to the broader categories of Manual Review, Panic, or See Other Pipeline.
[0062] FIG.16E is a sub-flow chart illustrating the Inherited Compound Heterozygous decision point where the analysis may flow through the decision points of Gene with >= 2 SNVs, Phasing Place SNVs on separate chromosome, Phasing of SNVs Unknown, Phased SNVs LoF / LoF, Phased SNVs LoF / Other Pathogenic, SNVs LoF / LoF, SNVs Lof / other pathogenic, which may potentially lead to the broader categories of Trash, Manual Review, and Panic.
[0063] FIG.17A shows an example of aneuploidy calling on sets of individual cells from a female embryo that was aneuploid, which showed a consistent loss of Chromosome 19. 8 93176811Attorney Docket No.21-2014-WO
[0064] FIG.17 B shows a different example of aneuploidy calling on sets of individual cells from a female embryo that was euploid, where there were some copy number variants in individual cells but none that appeared consistently throughout the embryo. DETAILED DESCRIPTION
[0065] For the purposes of promoting an understanding of the principles of the present disclosure, reference will now be made to embodiments and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the disclosure is thereby intended, such alteration and further modifications of the disclosure as illustrated herein, being contemplated as would normally occur to one skilled in the art to which the disclosure relates.
[0066] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0067] As used herein, the terms “comprise”, “comprising”, “include”, “including”, “have”, “having” and their conjugates mean “including but not limited to”. The term “consisting of” means “including and limited to”. The term “consisting essentially of” means that the composition, method or structure may include additional ingredients, steps and / or parts, but only if the additional ingredients, steps and / or parts do not materially alter the basic and novel characteristics of the claimed composition, method or structure.
[0068] As used herein, the term “method” or “methods” refers to manners, means, techniques and procedures for accomplishing a given task including, but not limited to, those manners, means, techniques and procedures either known to, or readily developed from known manners, means, techniques and procedures by practitioners of the chemical, pharmacological, biological, biochemical and medical arts.
[0069] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0070] Whenever a numerical range is indicated herein, it is meant to include any cited numeral (fractional or integral) within the indicated range. The phrases “ranging / ranges between” a first 9 93176811Attorney Docket No.21-2014-WO indicate number and a second indicate number and “ranging / ranges from” a first indicate number “to” a second indicate number are used herein interchangeably and are meant to include the first and second indicated numbers and all the fractional and integral numerals therebetween.
[0071] As used herein, the term “analyte” or “analytes” refers to biological macromolecule(s) such as protein(s), DNA, and / or RNA that are necessary for the function, growth, and development of cells and organisms.
[0072] As used herein, the phrase “a source of analyte” or “sources of analytes” refers to a source where genetic information of an embryo can be obtained. The sources include, but are not limited to, a single cell or group of cells biopsied from the embryo or blastocyst, cells or genetic information from cell culture media, blastocoel, or a combination thereof.
[0073] As used herein, the phrase “genetic information” refers to nucleic acid sequence of panel, exome, whole genome, or transcriptome with or without mitochondrial sequence and with or without epigenetic markers.
[0074] As used herein, the phrase “a reference genome” refers to a whole genome sequence from a public or private repository, one or more parents, or genetically related individual(s) such as a biological sibling and / or an embryo generated from the same parents, wherein reference genetic information can be obtained.
[0075] As used herein, the term “whole genome amplification” or “WGA” refers to a technique, a method, or a kit that is used to produce large quantities of DNA from a small amount of starting material. Unlike conventional PCR, WGA is aimed at amplifying the entire genome of an organism rather than a specific region.
[0076] As used herein, the terms “whole genome sequencing”, “whole-genome sequencing”, and “WGS” are interchangeable and refer to a laboratory process of determining the entirety, or nearly the entirety, of the DNA sequence of an organism’s genome at a single time. In the present disclosure, whole genome sequencing is performed on each embryo and each parent to determine all 6 billion bases of each individual embryo’s and parent’s genome which may or may not include the mitochondrial genome. The result of such sequencing is an individual’s complete DNA sequence, including non-coding sequence and may include epigenetic markers.
[0077] As used herein, the term “coverage” refers to the average number of times a particular sequence is interrogated by a sequence read. For example, “30 x WGS” means the entire genome will be sequenced an average redundant depth of 30 times.
[0078] As used herein, the term “coverage analysis” refers to assessment of genomes for variability, including, but not limited to, whether there are regions with higher or lower than expected coverage. Where expected coverage is based on total sequence coverage, known biases 10 93176811Attorney Docket No.21-2014-WO in coverage based on local genomic composition or gender of the embryo may be determined empirically.
[0079] As used herein, the term “variant calling” refers to a process of identifying differences between the sequencing reads of a target individual and of a reference genome. Also as used herein, the term “variant caller” refers to a tool for variant calling of sequencing data from one or more nucleic acid sequencing datasets (DNA, RNA, etc.). By way of non-limiting example, the variant caller used herein may include DeepVariant, which is a new TensorFlow machine learning- based variant caller.
[0080] As used herein, the phrase “train DeepVariant” refers to supplying the training algorithm with datasets where the expected outcome is known as a “truth set”, thus the machine-learning algorithm is penalized or rewarded based on performance. The algorithm is then used on datasets with no known truth set. Similar performance is expected.
[0081] As used herein, the term “merged variant calling” refers to merging the results of variant calls from multiple single sources from a single embryo or calls from looking at all the embryo data along with parental data simultaneously. The term “merged variant caller” refers to a tool or method for variant calling from multiple sources.
[0082] As used herein, the phrase “a genetic variant” refers to a difference in genetic information among sources of analytes from an embryo or a difference in genetic information between the embryo and reference genomes. Different types of genetic variants are described below.
[0083] As used herein, the term “single nucleotide variant”, “single-nucleotide variant” or “SNV” refers to a variation in a single nucleotide, which occurs when a single nucleotide (adenine, thymine, cytosine, or guanine) in the genome sequence is altered. A SNV may be rare or common in one population but common or rare in a different population. Sometimes SNVs are known as single nucleotide polymorphisms (SNPs), although SNV and SNPs are not interchangeable. To qualify as a SNP, the variant must be present in at least 1% of the population.
[0084] As used herein, the term “indel” is short for insertion / deletion, which refers to simultaneous insertion of a small length of DNA (usually less than 50 base pairs) into and deletion of a small length of DNA from the genome relative to a reference genome. A special case of an indel is a repeat expansion or contraction whereby a repetitive element of DNA becomes longer or shorter.
[0085] As used herein, the term “copy number variant” or “CNV” refers to a duplication or deletion that changes the number of copies of a particular DNA segment within the genome. Aneuploidies are a type of CNV typically consisting of a very large segment of a chromosome, or an entire chromosome. CNVs contribute 9% of pathogenicity in the inherited retinal degenerations (Zampaglione et al., 2020; Genetics in Medicine, 22: 1079-1087). 11 93176811Attorney Docket No.21-2014-WO
[0086] As used herein, the term “structural variant” or “SV” refers to rearrangement of part of the genome, which can be a deletion, duplication, insertion, inversion, translocation or a combination thereof and may or may not be copy neutral. SVs have been implicated in a number of conditions, including polycystic kidney disease, cardiomyopathies, amyotrophic lateral sclerosis (ALS) and some cases of intellectual disability.
[0087] As used herein, the term “disease-causing genetic variants” refers to a genetic change in the DNA sequence that have been shown to cause a disease or been associated with an increased risk of developing a disease.
[0088] As used herein, the term “phenotype” refers to any measurable characteristic of an individual due at least in part to genetics. It may include, but not be limited to, disease risk or protection of a trait.
[0089] As used herein, the term “De novo variant”, “apparent De novo variant”, “DNV”, or “apparent DNV” refers to a genetic variant in the embryo that is not apparently present in the sequence data generated from the parents. In some cases, the variant occurs in a parent’s egg or sperm cell but is not present in any of their other cells. In other cases, the variant may be present in a subset of the parental cells, including the gonads, but not in the sample that was sequenced (mosaicism). Still in other cases, the variant occurs in the embryo after fertilization. As the embryo grows, some or all of the resulting cells in the growing embryo contain the variant. De novo variants are one explanation for genetic disorders in an affected child but not in either parents.
[0090] As used herein, the term “parent” refers to the gamete donors that contribute to the genetic make-up of the embryo (i.e., biological parent).
[0091] As used herein, the terms “patient outcome”, “patient outcomes”, and the like refer to outcome of pregnancy, onset of disease, or onset of other desirable or undesirable condition in parent, embryo, or child.
[0092] As used herein, the terms “favorable outcome”, “favorable outcomes”, “a successful pregnancy”, and the like refer to achieving the avoidance of a particular negative or pathologic condition. For example, a favorable outcome may be carrying a fetus to term. In other non- limiting examples, a favorable outcome may be an embryo that does not have a genetic predisposition to one or more pathologic conditions. The one or more pathologic conditions may vary depending on the desires of the customer or biological parents.
[0093] As used herein, the terms “a pathologic condition”, “pathologic conditions”, and the like refer to abnormal anatomical or physiological condition(s) to be avoided when screening embryos depending on the desires of the customer or biological parents.
[0094] The present disclosure provides, in part, a method of identifying a genetic variant in an embryo. This method involves obtaining two or more sources of analytes from the embryo; 12 93176811Attorney Docket No.21-2014-WO analyzing the two or more sources of analytes to obtain and genetic information of each source; comparing the genetic information of each source against one or more reference genome using at least one variant caller that identifies variant(s) between each source and the reference genome; and combining the variant(s) to identify a difference that is present only from the sources, thus identifying a genetic variant in the embryo.
[0095] The genetic variant to be identified may be a single nucleotide variant (SNV), a multi- nucleotide variant (MNV), a copy number variant (CNV), a structural variant (SV), or an alteration in an epigenetic marker.
[0096] MNV includes, but is not limited to, insertions, deletions, insertions in combination with deletions, or repeat expansions wherein multiple nucleotides are involved. A MNV calling may include, but is not limited to, identification of insertions, deletions, and simultaneous insertion- deletions (indels), as well as repeat expansions and contractions (a special case of insertion and deletion, respectively). By way of non-limiting example, single base deletions / insertions / indels or multiple base deletions / insertions / indels may be identified through MNV calling.
[0097] SV includes, but is not limited to, inversions, translocations, repeat expansions and contractions (as a special case of insertion or deletion). An SV calling may include, but is not limited to, identification of rearrangements of the genome including rearrangements that cause copy number changes as well as copy-neutral changes. By way of non-limiting example, a copy- neutral change may include translocations and inversions.
[0098] An epigenetic marker may be a DNA methylation marker used for identifying imprinting defects. Methylation markers may be used to identify imprinting defects, including but not limited to, when both alleles are imprinted and skewed X-inactivation.
[0099] Parent data may be used to identify genetic mutations that occur in the embryo but not in parental WGS, i.e., De novo variants (DNV). Parental data may also be used for determining inheritance, which includes haplotype phasing and isodisomy and to identify whether inherited variants are in cis or trans, i.e., one from mother and one from father. By way of non-limiting example, determining whether one variant comes from the father and the other from the mother which can cause recessive disease.
[0100] Parental data may also be used to impute missing data in the embryo. Sometimes amplification or sequencing may fail to assay a portion of the embryo’s genome, one may infer from inheritance of particular haplotypes that other variants are present in the embryo even though they have not been directly observed.
[0101] In some embodiments, the method of identifying a genetic variant in an embryo disclosed above and herein can be used for preimplantation genetic testing (“PGT”), wherein genetic testing of embryos is performed, as well as for screening of embryos prior to transfer and implantation, 13 93176811Attorney Docket No.21-2014-WO wherein the genome of individual cells isolated from every embryo and both parents are sequenced and analyzed. In general, future parents and / or customers create a set of their embryos using IVF technology. The embryos are then biopsied per standard protocols. Before transferring an embryo for pregnancy, all 3 billion base pairs (or 6 billion bases) of each embryo’s genome are sequenced, along with both biological parents. For the whole genome sequencing of each embryo, individual cells are isolated from the embryo and sequenced. The sequencing data are subsequently merged and analyzed to identify the embryos with the lowest genetic risk for embryo transfer and implantation. The PGT and methods disclosed above and herein detect genetic risks, not only antenatal, but also postnatal.
[0102] In some embodiments, variations of the method of identifying a genetic variant in an embryo disclosed above and herein or variations of the associated technology could be applied to free floating or biopsied embryonic cells collected in utero or in the bloodstream of a pregnant woman. That is, in addition to PGT of embryos in assisted pregnancies, similar methods or technologies could also be used to natural pregnancies with improved genetic screening.
[0103] In some embodiments, variations of the method of identifying a genetic variant in an embryo disclosed above and herein or variations of the associated technology could be applied to one or more biological parents. That is, in addition to PGT of embryos in assisted pregnancies, similar methods or technologies could also be used to inform planning for natural pregnancies with improved genetic screening.
[0104] To identify a genetic variant in an embryo, analytes may be obtained from cells biopsied from the embryo or blastocyst, cells and / or genetic material isolated from cell culture media, blastocoel, or a combination thereof. This includes, but is not limited to, analytes from culture media, blastocoel fluid, or parental samples (e.g., blood, buccal swab, saliva, semen), and then separating / dissociating the bulk of cells into single cells. For any of the methods disclosed above and herein, 1-10 single cells may be isolated from each embryo for single cell analysis. By way of non-limiting examples, 2, 3, 4, 5, 6, 7, 8, 9, or 10 single cells may be isolated from the embryo.
[0105] In some embodiments, the biopsy is collected using laser pulse aided pipette delivery. Embryo biopsy may be performed after 5-6 days of culture in the laboratory, when the embryos have reached the blastocyst or hatching blastocyst stage. In some embodiments, two or more cells are isolated from the biopsy and separated using a physical or enzymatic method. By way of non- limiting example, the physical method may include laser dissection or micropipette isolation. Cells may then be dissociated into individual cells. By way of non-limiting example, the enzymatic method may include extraction using a digestive enzyme, including, but not limited to, Dispase, Collagenase, Hyaluronidase, Papain, DNase-I, Accutase, or Trypsin. 14 93176811Attorney Docket No.21-2014-WO
[0106] Single cells require amplification (e.g., WGA) prior to sequencing as they contain too little DNA load onto a sequencer. Amplification may induce errors such as a failure to amplify some portions of the genome, over amplification of portions of the genome, and failure to amplify both alleles of portions of the genome. These technical artifacts may mimic true and deleterious variants. However, by sequencing multiple single cells, the methods disclosed above and herein can infer true variants from artifacts by multiple independent observations of the variants.
[0107] The genetic information obtained from each source may include nucleic acid sequence of panel, exome, whole genome, or transcriptome with or without mitochondrial sequence and with or without epigenetic markers. By way of non-limiting example, the nucleic acid may include DNA, RNA or both.
[0108] In some embodiments, the nucleic acid sequence is obtained by isolating DNA from the source(s), generating libraries from whole genome amplification of DNA from each source and sequencing the libraries to generate a sequence data file of the source.
[0109] In some embodiments, prior to sequencing the libraries fully, embryos may be sequenced to lower coverage and aneuploidies may be identified. Embryos may subsequently be eliminated from further sequencing. By way of non-limiting example, aneuploidies may be detected by unexpectedly high or low coverage in a genomic region.
[0110] Whole genome sequencing (WGS) of both biological parents and each embryo provides the highest possible resolution and sensitivity for genetic analysis. In contrast, current offerings in Carrier Screening or Prenatal Testing only analyze about 1% of the genome where clinical diagnosis is most certain;and either analyze one parent at a time (carrier screening) or only the fetus (prenatal testing), which further limits possible insight. Current offerings in IVF use specialized sequencing with limited coverage or SNP arrays that generate 1000x less data than whole-genome sequencing, primarily due to the difficulty of obtaining and interpreting more extensive genetic data as well as due to efforts to constrain cost.
[0111] In some embodiments the sequencing data of, for example, separate cells is used for calling variants and the results are subsequently combined for analysis. In some embodiments, the sequencing data of, for example, separate cells is combined first and then used for calling variants.
[0112] In some embodiments, RNA is isolated and sequenced. In some embodiments, RNA sequence data is used to characterize the transcriptome and determine gene expression levels. In some embodiments, gene expression levels are correlated to patient outcomes. By way of non- limiting example, the patient outcomes include, but are not limited to, successful pregnancy. In this method, a panel of genetic information correlating gene expression levels and patient outcomes is established. 15 93176811Attorney Docket No.21-2014-WO
[0113] RNA sequence data may be used to identify an aberrant splicing event and identify dysregulation of transcriptional regulation and one or more variants that impact one or more transcriptional pathways. By way of non-limiting example, the aberrant splicing event may include cryptic splicing, aberrant exon usage, or upstream open reading frames.
[0114] RNA sequence data may also be used to identify biased allele expression thereby identifying imprinting, copy number determination, and / or variants that affect expression.
[0115] In some embodiments, RNA sequence data is used to call genomic variants alone or in combination with DNA sequencing. In some embodiments, RNA is used to confirm aneuploidy calls. In some embodiments, multiple data points from each cell (DNA, RNA) may be combined, and / or data points from multiple cells may be combined.
[0116] To identify a genetic variant in an embryo, the genetic information of each source is compared to genetic information of one or more reference genome(s). In some embodiments, the reference genome is a genome sequence from a public or private repository, one or more parents, a genetic relative and / or an embryo generated from the same parents. When using a genome sequence from one or more parents as a reference genome, the genetic variant may be further identified by comparing the genetic information of each source to genetic information of one or more parents to deduce haplotype phasing, impute missing data, determine inheritance or lack thereof, and identify apparent De novo variants. In some embodiments, the genetic variant is further annotated for its association with a phenotype by using a database of variants associated with diseases, phenotypes, disease risk or protection and traits. In some embodiments, the genetic variant is further annotated for its potential to cause deleterious effects on gene function through in silico, in vivo, or in vitro methods.
[0117] Identified variants may be annotated based on their presence or absence in other data sets including cohorts of healthy adults, children with early onset disease, and trio datasets. These datasets include private and public datasets. Academic studies to date have developed useful data in silos. Studies of childhood and adult genetic diseases are done separately from studies of the genetic causes of miscarriage and infertility. These existing datasets may be added to embryo research data and customer data to expand the reference sequence database.
[0118] In some embodiments, the variant caller may be a deep learning-based variant caller. In some embodiments, the deep learning-based variant caller is built based on genetic information of one or more parents or a genetic relative, and is capable of identifying an apparent De novo variant for purpose of preimplantation genetic screening. A genetic variant may occur in the genome with a known or suspected contribution to an adult or childhood genetic disease (i.e., disease-causing genetic variant). In other situations, a genetic variant may occur in the genome with no current associated genetic disease but under genetic constraint. 16 93176811Attorney Docket No.21-2014-WO
[0119] For the method of identifying a genetic variant in an embryo described above and herein, the combining step may include identifying a genetic variant that occurs in at least two sources, or identifying a genetic variant that can only be detected by combining genetic information from all sources. By way of non-limiting example, the genetic variant may be identified in multiple sources (not necessarily a majority of sources), or identified as mosaic.
[0120] Some aspects of the present disclosure provide a method of screening multiple embryos. Such method comprises conducting genetic analysis on each embryo according to the method described above and herein; and then ranking the embryos based on the results from the genetic analysis. The embryo with the best chance of a term pregnancy and the lowest postnatal genetic risk is ranked the highest for transfer and implantation. In some embodiments, embryos are ranked based on known deleteriousness of detected variants, or based on predicted deleteriousness of detected variants where such prediction is made by various methods. By way of non-limiting example, the predicting methods may include in silico deleteriousness prediction tools, allele evolutionary conservation, and / or allele frequency in a selected or unselected population database.
[0121] A genetic variant’s impact on embryo rank may be weighted based on input from a phenotype selected for examination. By way of non-limiting example, the input may be the desire of a customer to avoid a particular pathologic condition.
[0122] Alternatively, a genetic variant’s impact on embryo rank may be weighted based on automated phenotypic overlap between the phenotypes associated with a disease, and its associated gene and phenotypes selected for examination. Overlap may be computed using an ontology that allows related, as opposed to strictly identical terms, to determine overlap. By way of non-limiting example, the phenotypes selected for examination are derived from family history of one or both parents where parents may wish to avoid breast cancer risk, and thereby genes associated with risk of other cancers will also be weighted more heavily.
[0123] Still alternatively, a genetic variant’s impact on embryo rank may be weighted based on presence or absence of the variant in one or both parents and its zygosity in the embryo or the parents, known or suspected age of onset of disease, precurated gene or variant rankings, treatment options for a given disease, or its impact on gene function or expression.
[0124] The screening method described above and herein may further comprise generating a report with the ranking of the embryos. Instead of a quantitative ranking, embryo risk may be classified into two or more categories, including: red light (high risk; embryo transfer not recommended), yellow light (some identifiable risk; different parents may weight differently based on their own preferences; parents’ decision whether to transfer), green light (no specific risk identified; recommended for transfer and implantation), and a set for which no data was generated. That is, 17 93176811Attorney Docket No.21-2014-WO the ranked embryos may be qualitatively categorized based on recommendation to transfer, not to transfer, transfer with some specific ascertainable risk, or uncertainty about the risk.
[0125] The following examples are offered by way of illustration and not by way of limitation. Example 1: Embryo Screening
[0126] An early version of pipeline was developed to identify embryos that might possess unnecessarily high genetic risk (labeled “No go”) and to stratify remaining embryos based on assessment of the residual genetic risk. The embryos with low likelihood of disease were labeled “Go”.
[0127] Simulations were run using random “crossing” of 14 males and 14 females with 20 “embryos” per crossing. That is, 196 parent combinations with 20 embryos each generated 3920 simulated embryos. It was determined that a typical couple in their early-mid 30’s would have about 96-98% probability of having an embryo to transfer in their first cycle that passes all disease screenings (Table 1). Table 1
[0128] The pipeline was tested to analyze a cohort of children with rare diseases. Although only ~30% of children with rare, early onset disease would receive a genetic diagnosis, the methodology of the present disclosure could have successfully screened ~80% of the children because they had rare, deleterious variants in genes that were constrained for deleterious variants.
[0129] In conclusion, the early version of the analytic pipeline identified 2.5 x more childhood risk compared to the best current clinical diagnosis, and 96-98% probability of having at least one “healthy” embryo to transfer. Example 2: Carrier Test
[0130] The first pass pipeline was generated to fully characterize parental genetic risk. The genomes of 900 parents were screened for loss of function or known pathogenic, or extremely deleterious (top 0.1%) mutations in all known recessive disease genes, as well as known pathogenic in autosomal dominant genes and subsequently compared to Invitae’s carrier panel (“Invitae Comprehensive Carrier Screen” containing 287 genes). As shown in FIG.1, this pipeline identified about 6 x as much genetic risk as compared to a traditional comprehensive panel. In 18 93176811Attorney Docket No.21-2014-WO addition, this pipeline identified not only autosomal recessive gene (rare, clinical variant, and LoF), but also autosomal dominant gene rare variant from the Clin Var data base, and autosomal recessive gene variants rare in the population with a high combined annotation depended depletion score (CADD) that indicates a high probability of being deleterious. The identified genetic risks include CRTAP (Osteogenesis imperfecta, type VII) and PIEZO2 (Arthrogryposis, distal, with impaired proprioception and touch). Example 3: Cell Line Studies
[0131] Whole genome sequencing (“WGS”) and whole genome amplification (“WGA”) of cell lines are first characterized. (FIG. 2). These cell lines are purchased from Corriell and have publicly available sequence data from the Genome in a Bottle (GIAB) Consortium (NA24385, NA24632 and NA18278). The goal of these studies is to evaluate WGA kits and define merging strategy to be applied to embryo screenings.
[0132] Materials: Two different groups of samples are used. First group is ten (10) individual cells from three GIAB samples sequenced at 15x WGS, and second group is three (3) “bulk” samples of 5-6 cells from each GIAB. Variant calling has been previously performed on each GIAB sample, producing a truth set of genetic variants that can be used to compare the results from this WGS approach.
[0133] As a first pass, the three GIAB are assessed for coverage variability to determine regions with higher than expected coverage, if any, and how many low coverage regions by comparing the average coverage of a suitably large region (10-100 Kbp) to the expected (average) coverage. For each single cell sample, a subset of data (e.g., 10x, 4x, 5x, 1x, 0.1x) is taken for pre-processing. Same WGS is repeated for the “bulk” samples.
[0134] Next, variants may be identified in each cell alone or in combination with other genetic data using any available variant callers (platypus, GATK’s Haplotype caller, Freebayes, Octopus, Delly, etc.). Further, a custom trained artificial intelligent variant caller that was optimized for detection of CNVs and SNVs from single cell data may also be used to identify variants. This caller is trained by supplying the training algorithm with datasets where the expected outcome is known (a truth set), thus the machine-learning algorithm is penalized or rewarded based on performance. The algorithm is then used on datasets with no known right answer but with similar performance expected.
[0135] SNV Calling: For every data set from the first group of samples (i.e., 3 x 10 individual cells), variants are independently called with platypus at 15x and 10x down sampling (60 Variant Call Formats (VCFs) - SET A). For each cell-line merge 4x from 10 cells, 5x from 8 cells and 10x from 4 cells (repeat with different cells) and variants are called with platypus (12 VCFs - SET B). 19 93176811Attorney Docket No.21-2014-WO The point of this is to sequence everything to 40x total to determine if the number of cells matter. For every dataset from the second group of samples (3 “bulk” samples), variants are called with platypus (3 VCFs - SET C).
[0136] Next, 5, 7 and 10 random VCFs from SET A for each WGA are merged by two techniques: majority rule (i.e., seen 3, 4 or 6 times, respectively) and “seen more than once” rule (i.e., any variant seen in at least 2 VCFs). The results of the merging will generate 18 VCFs (SET D).
[0137] The calls are then compared among the 12 VCFs (SET B), 3 VCFs (SET C) and 18 VCFs (SET D) by building confusion matrix for each sample. A confusion matrix records true positive, true negative, false positive, and false negative rates between a sample and a known truth set.
[0138] Aneuploidy: aneuploidy is determined by taking low coverage whole genome data (0.1X) and comparing the expected coverage over a large region (10-100 Mbp) to the measured coverage. Regions with higher than expected coverage likely indicate a copy gain which may include the entire chromosome or a part of a chromosome, whereas regions with reduced coverage may indicate copy loss of all or part of a chromosome. Expected coverage may be determined theoretically, based on sequencing depth, GC content and other local genomic factors, or may be based on empirical data from sequencing euploid samples.
[0139] SV Calling: For every data set from the first group of samples (i.e., 3 x 10 individual cells), variants are independently called with 15x, 10x, 5x, lx, 0.1x (may need special caller for lx and 0.1x) (150 VCFs - SET E). For each cell-line merge 4x from 10 cells, 5x from 8 cells and 10x from 4 cells (repeat with different cells) and variants are called with delly (12 VCFs - SET F).
[0140] Next, 5, 7 and 10 random VCFs from SET A for each WGA are merged by two techniques: majority rule and “seen more than once” rule. The results of the merging will generate 9 VCFs (SET G).
[0141] SURVIVOR will be used to determine how many not-real deletions / gains are called by each VCF in SET E and see how often it is consistent (i.e., non-random drop-out or over- amplification).
[0142] Once De novo calls are generated from the isolated cells, the results will help inform the WGA & merging strategy. In addition, the data generated from the cell line experiments in combination with the data from GIAB to make trios will be used to train DeepVariant. Example 4: Donated Embryo Studies
[0143] The cell line methodologies from Example 3 above will be extended to embryos (FIG.3). First, each embryo undergoes biopsy and one or more single cells are isolated from each embryo using physical or enzymatic methods. Starting source material could be multiple sources of embryonic material such as embryo, blastocyst, cell culture media, etc. 20 93176811Attorney Docket No.21-2014-WO
[0144] The WGA kits and merging strategy defined from the cell line experiments in Example 3 will be applied to the embryos. These approaches include steps such as a very low pass WGS (0.1x), using this low pass WGS to eliminate some types of aneuploidies, and approaches to analyzing single vs. merged cells. The sequencing step could be sequencing DNA (genome), RNA (transcriptome), or identification of methylation sites. Each has its own library preparation method. The goal of each sequencing output is to identify variants that have been associated with disease or disease risk / protection.
[0145] Data obtained from WGA and WGS will then be analyzed to determine the concordance of individual cells with the embryo to evaluate reproducibility and consistency. Results from each set of ~6 individual cells will be compared to results from the whole embryo. In addition, results from each of the clusters will be compared to each other. The same software pipeline will be used but with streamlined data management. Furthermore, the previously trained DeepVariant will be applied to embryo data for evaluation. Upon analysis, the embryos will be classified and ranked based on the chance of a term pregnancy and postnatal genetic risk, a report will be generated with recommendations as to which embryo(s) to select for transfer and implantation. The classification and report generation involve manual and automatic steps (FIGS. 5-12). For instance, the classification of some variants may be entirely automatic, because the variant has been previously reported as deleterious from multiple sources, in other cases a variant may be weighted as it may appear deleterious through various means but has not been previously reported as pathogenic, and may occur in a gene that does not always cause disease. Some situations may require Manual review such as multiple De novo mutations in a single gene associated with autosomal recessive disease. Example 5: Genetic Analysis and Embryo Ranking
[0146] The results from the cell lines and embryos will be further expanded to a proposed end-to- end plan for minimum viable product (MVP) (FIG.4). First, each embryo undergoes biopsy and one or more single cells are isolated from each embryo using physical or enzymatic methods. Next, the WGA kits and merging strategy defined from the cell line experiments in Example 3 and finalized during the donated embryo experiments in Example 4 will be applied in this study. Parental WGS will use standard WGS methods.
[0147] Variant calling methods will be developed for structural variant (SV) and single nucleotide variant (SNV) calling that utilize parental genetic information. For instance, SNV calling for the embryo cells will include haplotype phasing, imputation for missing data, and De novo calling - all leveraging the parental genetic information. Current approach involves exploring modifications of DeepVariant to accommodate these steps, as described in Example 3. 21 93176811Attorney Docket No.21-2014-WO
[0148] Upon analysis, the embryos will be classified and ranked based on the chance of a term pregnancy and postnatal genetic risk, a report will be generated with recommendations as to which embryo(s) to select for transfer and implantation. The classification and report generation involve manual and automatic steps (FIGS. 5-12). In particular, variants may be classified into broad categories, including known pathogenic variants, non-coding variants, coding variants, loss of function (LoF) variants, synonymous variants and variants that will not warrant further consideration (Trash) as well as “Panic” conditions where manual review will be required (FIG. 5). For SNV or small insertion / deletion (indel) of <50 bp, the workflow may include the assessing variant characteristics through the decision points of: Variant in Blacklist; Variant in Whitelist; (i.e., well known pathogenic, possibly with special considerations, and which may lead to Panic), ClinVar Variant; In Protein Coding Region and not Synonymous; and Intronic and Synonymous; As the categorization progresses, the ClinVar decision point may further include considering if the variant is benign of likely benign (B or LB) and pathogenic or likely pathogenic (P or LP). If pathogenic P or LP, then it follows ClinVar Path discussed in more detail in reference to FIG.7 below. The In Protein Coding Region and not Synonymous decision point may further include the decision point of considering if the variant has Start Loss, Stop Gain, and Splice + / -2, Deeper Splice + / -10 SpliceAI loss >.5, or Frameshift to arrive at whether or not the variant is LoF or No- LoF. If the variant is not In Protein Coding Region and not Synonymous, then the inquiry may progress to the decision point further considering if the variant is Intronic and Synonymous. If the variant is intronic and synonymous, then it is classified as so, if the variant is not intronic and synonymous, then Trash.
[0149] For inherited, LoF coding, SNV variants, broad categories such as “No go” wherein an embryo would not be considered for transfer and weighting of variants to aid in ranking an embryo may be used (FIG.6). The workflow may include categorizing the variant through the decision points of: Minor Allele Frequency (MAF) < X%; Variant in American College of Medical Genetics (ACMG) criteria (e.g., ACMG 73); Variant in Online Mendelian Inheritance in Man (OMIM) morbid; and High Constraint (pli > XX). For the MAF < X% decision point, the categorization may further evaluate variants with a minor allele frequency less than a certain percentage (X%) and Panic or continue categorization. The ACMG decision point may further include considering LoF know mechanism, if so then Weighting. The OMIM morbid decision point may further include considering LoF known mechanism, X-linked, Autosomal Recessive, Autosomal Dominant, Variant Penetrance / Severity / Family History (FHx). If the variant is X- linked, then the categorization may further include considering X-linked dominant (XLD), and Male (if so No go). If XLD then considering Variant penetrance, severity, or FHx. In Weighting at this stage, it may be considered that a female with an XLR received from the mother may have 22 93176811Attorney Docket No.21-2014-WO led to this variant or an XLD that has variable penetrance. If the embryo is a male, then the workflow may include further considering if the variant was inherited from the father, and if this is the case, arriving here may be the result of an XLR that was inherited from the father, or an XLD that should always cause disease, but the parent survived, therefore Panic. Similarly, if the inquiry leads to High Constraint, then it is not an OMIM morbid and a LoF may be present that was inherited from a parent, which is rather peculiar.
[0150] For variants of known pathogenic (e.g., from ClinVar database), broad categories such as “No go” Weighting and under-go Manual Review may be used (FIG.7). The categorization may include assessing variant characteristics through the decision points of X-linked (XL), Autosomal Recessive (AR), and Autosomal Dominant (AD). If the variant is XL, then the categorization may further include considering if it is XLR and in female. If yes, then Weighting, if no, then it is No go. If the variant is AD, then the categorization may further include considering Variable penetrance or Severity, Family history (FHx or Fam Hx), and whether the variant is a De novo variant to potentially arrive at No go, Weighting, or Manual Review. If the variant is AR, then the workflow may proceed to the AR workflow described in more detail in FIG.8 below.
[0151] Referring now to FIG. 8, in order for a variant to arrive at this workflow, it has already been determined that it is “pathogenic” in an AR gene. For these pathogenic variants in genes associated with autosomal recessive (AR) disease, it is determined whether multiple variants exist in a gene and whether those variants are in cis or trans, or whether a variant exists as a homozygous or hemizygous state in a gene. The embryo may be broadly categorized as “No go”, to Weight it, or under-go Manual Review. More specifically, the decision point of Is Variant Homo or Hemizygous may further include considering if it is also Homo or Hemizygous in parent, then trans (checking whether multiple variants in a gene are all from one parent (“Cis”), or from different parents (“Trans”). This assumes that almost all recessive genes are fully penetrant, and therefore if it isn’t a parent that is alive, it is probably not pathogenic. If it is in both parents, or it is a variant on X chromosome of a boy, then it is a No go and it is homozygous inherited from both parents (consanguineous), or hemizygous in a boy. If in the decision point further considering if it is on chromosome X in a boy the answer is negative, then it will require Manual Review. Such a variant shows nonmendelian inheritance and could overlap a deletion, of maybe it was missed in the parent or allelic drop out in the child. The decision point of Are there >1 variant in the gene may further include considering if there is >0 variants from the father and if yes, then if there is >0 variant from the mother. If so, then it is No go. This is the most common type of autosomal recessive. If there are no >0 variant from the father, then the categorization may further include considering if there is >0 variant from the mother. If there are no >0 variants from the mother, then consider if there is >0 variants with unknown inheritance. Arrival here indicates there is 1 23 93176811Attorney Docket No.21-2014-WO inherited variant and 1 variant with unknown inheritance (likely De novo). If the variant cannot be phased, then the variant is probably an No go, but this would be a rare event.
[0152] FIG.9 illustrates variant classification of a De novo LoF coding variant (SNV) into broad categories as previously indicated. More specifically, the categorization may include the decision points of: ACMG; OMIM morbid; and High Constraint.
[0153] Referring now to FIGS.10A-D, the categorizations may include the decision points of De novo and >XX MBp, and Overlaps coding region of the gene for SV Deletions (FIG.10A). In the SV Gain in FIG.10B, the categorization may include the decision points of Starts or Ends coding region of a gene and Triplosensitive. For Inversion calling in FIG.10C, the categorization may include the decision points of De novo, and Starts or Ends coding region of a gene. Finally, for Translocation calling in FIG.10D, the categorization may include the decision points of De novo and Starts or Ends coding region of a gene.
[0154] FIG.11 illustrates repeat expansion categorization based on known pathogenic and “pre- mutation” lengths of repeat expansions as well as the repeat length in the parents to broadly categorize an embryo as No go or conduct Manual Review.
[0155] FIGS. 12A-D illustrate classification of large regions of homozygosity, uniparental isodisomy, synonymous and intronic variants, and variants identified from a spinal muscular atrophy pipeline. More specifically, in FIG. 12A large regions of homozygosity variants maby classified as Treat normally (recessive disease), Treat normally (deletion), and Manual review. In FIG. 12B, synonymous and intronic variants may be classified as Treat as LoF. In FIG. 12C, large regions of uniparental inheritance may be classified as No go or Manual Review. Finally, FIG. 12D is a flow chart illustrating a potential categorization of variants identified from, for example, a spinal muscular atrophy pipeline. Example 6: Genetic Analysis and Embryo Ranking
[0156] FIGS.13 to 14C illustrate categorization of SNV or small indel variants through the steps of DeepVarian, VCF Information, Population Frequency, Reference Sequence, Phenotype Classification, and Trio Heredity to classify them into various broader categories.
[0157] Referring more specifically to FIG. 13, the initial steps of DeepVariant, VCF Info, and Population Frequency are detailed. The analysis may potentially lead to the broad categories of Trash or Later Expansion. Note that the Population Frequency analysis may include, for example, 1KG and Gnomad database inquiries.
[0158] Referring now to FIG.14A, in the Reference Sequence step the categorization may flow through the decision points of Coding Variant, Intronic Variant, Intergenic Variant, Varian is a 24 93176811Attorney Docket No.21-2014-WO Nonsense or LoF, Variant is Missense, Variant is Synonymous, Variant Splicing Effect (X of 4), Variant Splicing Effect (Clinical X of 4), and Pathogenic Evidence (ClinVar), which may potentially lead to the broader categories of Panic, Trash, or Later Expansion.
[0159] Referring to FIG.14B, in the step of Phenotype Classification step, the categorization may flow through the decision points of ACMG Criteria Met, ClinVar Pathogenicity Criteria Met, OMNIM Criteria Met, and Embryo “Intolerome” Criteria Met, which may potentially lead to the broader categories of Trash, or Later Expansion.
[0160] Referring to FIG.14C, in the step of Trio Heredity the categorization may flow through the decision point of De Novo Dominant, Inherited Dominant, Homozygous Recessive, X-Linked, Inherited Compound Heterozygous, and Manual Review Catch Net, which may potentially lead to the broader categories of Panic or Later Expansion. Note that De Novo Dominant, X-Linked, Inherited Dominant, Homozygous Recessive, and Inherited Compound Heterozygous each have a sub-flow chart detailed in FIG.16A-E.
[0161] Referring to FIG.15, in the Reporting / Masking Subsets step of the process, categorizations of various pipelines may aggregate. In this example, the category from the SNV Annotation pipeline arrives and may flow through the decision points of On Mask List and On Report List, which may lead to the broader categories of Trash, Manual Review, of Later Expansion before the report is finalized.
[0162] Referring to FIG. 16A, in this sub-flow chart for the De Novo Dominant decision point, the analysis may flow through the decision points of De Novo (Not Parental), From Parent, Dominant Mechanism and Recessive Mechanism, which may potentially lead to the broader categories of Panic or See Other Pipeline.
[0163] Referring to FIG.16B, in this sub-flow chart for the X-linked decision point, the analysis may flow through the decision points of Non-PAR region, Proband Male and Hemizygous, and Proband Female and Homozygous, which may potentially lead to the broader categories of Manual review or See Other Pipeline.
[0164] Referring to FIG. 16C, in this sub-flow chart for the Inherited Dominant decision point, the analysis may flow through the decision points of De Novo (Not Parental), From Parent, (P1)- Het / (P2)-WT / (E)-Het, Rare (P1 / 2) with (E)-Het, Dominant Mechanism, and Recessive Mechanism, which may potentially lead to the broader categories of Panic or See Other Pipeline.
[0165] Referring FIG.16D, in this sub-flow chart for the Homozygous Recessive decision point, the analysis may flow through the decision points of De Novo (Not Parental), From Parent, (P1)- Het / (P2)-Het / (E)-Homo, (P1)-Het / (P2)-WT / (E)-Homo, Rare (P1 / 2 with (E)-Homo, Dominant Mechanism, and Recessive Mechanism, which may potentially lead to the broader categories of Manual Review, Panic, or See Other Pipeline. 25 93176811Attorney Docket No.21-2014-WO
[0166] Referring to FIG. 16E, in this sub-flow chart illustrating the Inherited Compound Heterozygous, the analysis may flow through the decision points of Gene with >= 2 SNVs, Phasing Place SNVs on separate chromosome, Phasing of SNVs Unknown, Phased SNVs LoF / LoF, Phased SNVs LoF / Other Pathogenic, SNVs LoF / LoF, SNVs Lof / other pathogenic, which may potentially lead to the broader categories of Trash, Manual Review, and Panic. Example 7: Aneuploidy Calling
[0167] FIGS.17A and 17B show examples of aneuploidy calling on sets of individual cells from 2 different embryos. The analyses show the differences between signals that are different among cells vs signals that are consistent across cells. All the cells in each figure were from the same embryo. Cells labeled "B" were isolated from biopsies, cells labeled "M" were isolated from the inner cell mass, and samples labeled "R" were bulk samples of the remaining cells from each embryo (to show concordance). The copy number calls were made using CNVkit software based on a custom reference set made up of single cell sequences amplified with the same WGA method. FIG. 17A is for a female embryo that was aneuploid, which showed a consistent loss of Chromosome 19. FIG. 17B was for a female embryo that was euploid, where there were some copy number variants in individual cells but none that appeared consistently throughout the embryo.
[0168] Discussion: the genetic analysis and embryo screening methods disclosed above and herein involves using comprehensive genomic data and analysis alongside the most sensitive screening criteria to identify the embryo(s) with the best chance of a term pregnancy and minimum postnatal genetic risk over all of the 6 billion bases in the human genome. The analytics of the present disclosure will be more thorough as it includes looking at constrained genes that are not currently considered disease causing, small CNVs below the detection limit of most aneuploidy screens, structural variants and intergenic and intronic mutations. Since multiple single cells are individually sequenced then combined rather than pooling cells first, there will be fewer regions of low or no coverage. In addition, genetic information from the embryo will be compared to genetic information from one or more parents and / or one or more genetic relatives. This will allow differentiation of sequencing artifacts from real mutations as any variant in multiple cells is likely real mutations.
[0169] Some key challenges associated with whole-genome sequencing with single cells include coverage across the genome, reproducibility between reactions, and measurement of accuracy. These challenges are effectively answered by the methods of the present disclosure. First, by using multiple sources (cells and / or cell-free sources) to obtain embryonic data and then merging the 26 93176811Attorney Docket No.21-2014-WO data, the methods of the present disclosure can achieve better coverage across the entire genome thus improve accuracy through depth and comparisons. In some embodiments the sequencing data of, for example, separate cells is used for calling variants and the results are subsequently combined for analysis. In some embodiments, the sequencing data of, for example, separate cells is combined first and then used for calling variants. Second, by using multiple WGA methods, the methods of the present disclosure can compensate for biases in each of the methods. Third, by using data from multiple embryonic sources as well as parents / familial genomic data to identify De novo variants, the methods of the present disclosure improves accuracy of calling variants.
[0170] The genetic analysis of multiple embryos during in vitro fertilization (IVF) will require a way to compare the disease risk of each embryo in order to select embryo(s) for transfer and implantation. Each embryo will have distinct, if overlapping, genetics and thus, different disease risks. The methods of the present disclosure involve automated classification of variants and subsequent ranking of embryos, which have not been done or suggested before. Prior to implantation, all genetic variants are easily actionable, that is, a parent may choose not to transfer an embryo because of a genetic susceptibility to disease. This is especially true for adult-onset diseases, e.g., a pathogenic BRCA2 mutation may lead to breast cancer in 30 to 50 years, but there are treatments and likely improved treatments in the future. This risk may need to be evaluated in the presence or absence of a childhood disease. Thus, genetic information from the entire genome must be taken into account. Automation allows rapid and cost-effective ranking of embryos. The ranking process and report generation summarizes the whole genome analysis to empower the parents to choose an embryo that meets their desired goal after being informed by the analytics provided by the methods disclosed above and herein.
[0171] All publications, patents and patent applications mentioned in this specification are herein incorporated in their entirety by reference into the specification, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated herein by reference. In addition, citation or identification of any reference in this application shall not be construed as an admission that such reference is available as prior art to the present invention. To the extent that section headings are used, they should not be construed as necessarily limiting.
[0172] Although the present invention and its advantages have been described in detail, it should be understood that various changes, substitutions and alterations can be made herein without departing from the spirit and scope of the invention as defined in the appended claims.
[0173] One skilled in the art will readily appreciate that the present invention is well adapted to carry out the objects and obtain the ends and advantages mentioned, as well as those inherent therein. The examples along with the methods described herein are presently representative of 27 93176811Attorney Docket No.21-2014-WO preferred embodiments, are exemplary, and are not intended as limitations on the scope of the invention. Changes therein and other uses will occur to those skilled in the art which are encompassed within the spirit of the invention as defined by the scope of the claims. 28 93176811
Claims
Attorney Docket No.21-2014-WO WHAT IS CLAIMED IS:
1. A method of identifying a genetic variant in an embryo, the method comprising: (a) obtaining two or more sources of analytes from the embryo; (b) analyzing the two or more sources of analytes to obtain genetic information of each source; (c) comparing the genetic information of each source against one or more reference genome using at least one variant caller, wherein the variant caller identifies variant(s) between each source and the reference genome; and (d) combining the variant(s) to identify a difference that is present only from the sources, wherein the difference is a genetic variant in the embryo. 2 The method of claim 1, wherein the genetic variant is a single nucleotide variant (SNV), a multi-nucleotide variant (MNV), a copy number variant (CNV), a structural variant (SV), or an alteration in an epigenetic marker. 3 The method of claim 2, wherein the epigenetic marker is a DNA methylation marker for identification of defects or transcriptional dysregulation. 4 The method of claim 2, wherein the genetic variant is identified as an apparent De novo variant (DNV). 5 The method of claim 1, wherein the two or more sources of analytes include cells biopsied from the embryo or blastocyst, cells and / or genetic material isolated from cell culture media, blastocoel, or a combination thereof. 6 The method of claim 5, wherein the biopsy is collected using laser pulse aided pipette delivery. 7 The method of claim 6, further comprising isolating two or more cells from the biopsy and dissociating the two or more cells using a physical or enzymatic method. 8 The method of claim 7, wherein the physical method includes laser dissection or micropipette isolation. 29 93176811Attorney Docket No.21-2014-WO 9. The method of claim 7, wherein the enzymatic method includes dissociation using a digestive enzyme, wherein the enzyme is elected from the group consisting of Dispase, Collagenase, Hyaluronidase, Papain, DNase-I, Accutase, and Trypsin.
10. The method of claim 1, wherein the genetic information includes nucleic acid sequence of panel, exome, whole genome, or transcriptome with or without mitochondrial sequence and with or without epigenetic markers.
11. The method of claim 10, wherein the nucleic acid includes DNA, RNA or both.
12. The method of claim 11, wherein the nucleic acid is DNA, and wherein the DNA sequence is obtained by the following steps: (a) generating libraries of DNA from each source; (b) conducting low-pass whole genome sequencing of the DNA; (c) applying an aneuploidy filter to the sequencing data from (b) to eliminate an embryo that has aneuploidies; and (d) conducting additional sequencing of the DNA from an embryo that is not eliminated to generate a sequence data file of the embryo.
13. The method of claim 12, wherein application of the aneuploidy filter leads to detection of chromosomal regions with unexpectedly increased or decreased coverage compared to the expected coverage based on total sequence coverage, known biases in coverage based on local genomic composition, and gender of the embryo.
14. The method of claim 11, wherein RNA is isolated, amplified and sequenced.
15. The method of claim 14, wherein RNA sequence data is used to characterize the transcriptome and determine gene expression levels.
16. The method of claim 15, wherein gene expression levels are correlated to patient outcomes.
17. The method of claim 16, wherein the patient outcomes include successful pregnancy. 30 93176811Attorney Docket No.21-2014-WO 18. The method of claim 16, wherein genetic information and / or gene expression is correlated to patient outcomes.
19. The method of claim 14, wherein RNA sequence data is used to call genomic variants in combination with DNA sequencing or alone.
20. The method of claim 14, wherein RNA sequence data is used to identify an aberrant splicing event and dysregulation of transcriptional regulation and one or more variants that impact one or more transcriptional pathways.
21. The method of claim 20, wherein the aberrant splicing event is cryptic splicing, aberrant exon usage, or upstream open reading frames.
22. The method of claim 14, wherein RNA sequence data is used to identify biased allele expression thereby identifying imprinting, copy number determination, and / or variants that affect expression.
23. The method of claim 1, wherein the reference genome is a genome sequence from a public or private repository, one or more parents, or one or more genetic relatives.
24. The method of claim 23, wherein the genetic variant is further identified by comparing the genetic information of each source to genetic information of one or more parents to deduce haplotype phasing, impute missing data, determine inheritance or lack thereof, and identify apparent De novo variants.
25. The method of claim 24, further comprising annotating the genetic variant for its association with a phenotype by using a database of variants associated with diseases, phenotypes, disease risk or protection and traits.
26. The method of claim 24, further comprising annotating the genetic variant for its potential to cause deleterious effects on gene function by using a database of gene functions or biological pathways or through in silico, in vivo, or in vitro methods.
27. The method of claim 1, wherein the variant caller is a deep learning-based variant caller. 31 93176811Attorney Docket No.21-2014-WO 28. The method of claim 27, wherein the deep learning-based variant caller is built based on genetic information of one or more parents or a genetic relative, and is capable of identifying an apparent De novo variant for purpose of preimplantation genetic screening.
29. The method of claim 1, wherein the genetic variant occurs in the genome with a known or suspected contribution to an adult or childhood genetic disease or in the genome with no current associated genetic disease but under genetic constraint.
30. The method of claim 1, wherein the combining step comprises identifying a genetic variant that occurs in at least two sources, or identifying a genetic variant that can only be detected by combining genetic information from all sources.
31. The method of claim 30, wherein the genetic variant is identified in multiple sources.
32. The method of claim 30, wherein the genetic variant is identified as mosaic.
33. A method of screening multiple embryos, the method comprising: (a) conducting genetic analysis on each embryo according to any preceding claim; and (b) ranking the embryos based on the results from the genetic analysis, wherein the embryo with the best chance of a term pregnancy and the lowest lifetime genetic risk is ranked the highest for transfer and implantation.
34. The method of claim 33, wherein embryos are ranked based on known deleterious or benign variants, or based on variants predicted to be deleterious or benign by various methods.
35. The method of claim 34, wherein the methods include in silico deleterious prediction tools, allele evolutionary conservation, and / or allele frequency in a selected or unselected population database.
36. The method of claim 33, wherein a genetic variant’s impact on embryo rank is weighted based on input from a customer-based phenotype. 32 93176811Attorney Docket No.21-2014-WO 37. The method of claim 36, wherein the input is the desire to avoid a particular pathologic condition.
38. The method of claim 33, wherein a genetic variant’s impact on embryo rank is weighted based on automated phenotypic overlap between the phenotypes associated with a disease, and its associated gene and phenotypes provided by a customer.
39. The method of claim 38, wherein the phenotypes are derived from family history of one or both parents.
40. The method of claim 38, wherein the overlap is computed using an ontology which allows related phenotypes to weigh on an embryo ranking.
41. The method of claim 33, wherein a genetic variant’s impact on embryo rank is weighted based on known disease severity, penetrance, and / or inheritance pattern associated with a gene.
42. The method of claim 33, wherein a genetic variant’s impact on embryo rank is weighted based on presence or absence of the variant in one or both parents and the zygosity in the parents.
43. The method of claim 33, wherein a genetic variant’s impact on embryo rank is weighted based on known or suspected age of onset of disease.
44. The method of claim 33, wherein a genetic variant’s impact on embryo rank is weighted based on pre-curated gene or variant rankings.
45. The method of claim 33, wherein a genetic variant’s impact on embryo rank is weighted based on treatment options for a given disease.
46. The method of claim 33, further comprising generating a report of the multiple embryos, wherein ranked embryos are qualitatively categorized based on recommendation to transfer, not to transfer, or unsure to transfer. 33 93176811