Multi-gene risk score for in vitro fertilization

The embryonic genome is constructed through whole-genome sequencing and phase separation technology, combined with the multigene risk scoring model, and the problem of inaccurate prediction of genetic disease risk in the existing technology is solved, and efficient risk assessment and selection of embryos and potential children is achieved.

CN120473129APending Publication Date: 2025-08-12マイオームインコーポレイテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510355629.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-08-06
Filing Date
2020-09-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the risk of genetic disease in individuals and potential children, especially in the context of couples having family history and sperm donors.

Method used

Through whole-genome sequencing and phase separation technology, the genome of the embryo is constructed, combined with multigene risk scoring models, the disease risk of the embryo and potential children is determined, and the genetic information of the paternal, maternal and ancestral families are used to evaluate the risk.

Benefits of technology

It improves the accuracy of predicting genetic disease risks, helps to select low-risk embryos and potential children, and reduces the probability of genetic disease transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120473129A_ABST
    Figure CN120473129A_ABST
Patent Text Reader

Abstract

Relates to multi-gene risk scores for in vitro fertilization. Provided is a method for determining the risk of a disease associated with an embryo, comprising constructing a genome of the embryo based on (i) one or more genetic variants in the embryo, (ii) a paternal haplotype, (iii) a maternal haplotype, (iv) a probability of transmission of the paternal haplotype, and (v) a probability of transmission of the maternal haplotype; assigning a polygene risk score for the embryo based on the constructed genome of the embryo; determining a disease risk associated with the embryo based on the polygene risk score; and determining delivery of the genetic variants causing the disease and / or haplotypes from the paternal and / or maternal genomes to the embryo. Also provided are methods of determining a range of disease risks for a mother and a potential child of a potential sperm donor. Also provided are methods of determining the risk of disease in an individual.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention application with the application date of September 30, 2020, Chinese application number 202080080085.2, and invention name “Polygenic risk score for in vitro fertilization”.

[0002] Cross-reference to related applications

[0003] This application claims the benefit of U.S. Provisional Application No. 62 / 908,374, filed September 30, 2019, and U.S. Provisional Application No. 63 / 062,044, filed August 6, 2020, each of which is incorporated herein by reference in its entirety. Technical Field

[0004] Described are methods for determining disease risk. Background Art

[0005] Currently, IVF clinics test for aneuploidies and single-gene disorders known to run in families. However, one in two couples has a family history of common diseases influenced by a combination of genetic, environmental, and lifestyle risk factors. Furthermore, sperm donor clinics currently test for a predisposition to a subset of diseases caused by single-gene disorders. There is a need in the art for improved prediction of inherited disease risk in individuals and in potential future children. Summary of the Invention

[0006] Provided is a method for determining disease risk associated with an embryo, the method comprising: performing whole genome sequencing on a biological sample obtained from a paternal subject to identify a genome associated with the paternal subject; performing whole genome sequencing on a biological sample obtained from a maternal subject to identify a genome associated with the maternal subject; phasing the genome associated with the paternal subject to identify a paternal haplotype; phasing the genome associated with the maternal subject to identify a maternal haplotype; performing rare genotyping on an embryo to identify one or more genetic variants in the embryo; and performing a genotyping test based on (i) one or more genetic variants in the embryo, (ii) the paternal haplotype, and (iii) the maternal haplotype. The invention also provides a method for constructing a genome of an embryo based on the transmission probability of the maternal haplotype, the maternal haplotype, the paternal haplotype, the maternal haplotype, the paternal haplotype, and the maternal haplotype; assigning a polygenic risk score to the embryo based on the constructed genome of the embryo; determining a disease risk associated with the embryo based on the polygenic risk score; determining the transmission of genetic variants and / or haplotypes from the paternal genome and / or the maternal genome that cause single-gene diseases to the embryo; and determining a combined disease risk associated with the embryo based on the polygenic disease risk and the transmission of genetic variants and / or haplotypes from the paternal genome and / or the maternal genome that cause single-gene diseases to the embryo.

[0007] Also provided is a method for outputting a disease risk score associated with an embryo, the method comprising: receiving a first data set comprising paternal genomic data and maternal genomic data; aligning sequence reads with a reference genome and determining genotypes across the genome using the paternal genomic data and the maternal genomic data; receiving a second data set comprising paternal and maternal sparse genomic data; phasing the paternal genomic data and the maternal genomic data to identify paternal haplotypes and maternal haplotypes; receiving a third data set comprising sparse genomic data of the embryo, paternal transmission probabilities, and maternal transmission probabilities; applying an embryo reconstruction algorithm to (i) the paternal haplotypes and maternal haplotypes, (ii) the sparse genomic data of the embryo and (iii) the transmission probabilities of each of the paternal haplotypes and maternal haplotypes to determine a constructed genome of the embryo; applying a polygenic model to the constructed genome of the embryo; outputting a disease risk associated with the embryo; determining the transmission of disease-causing genetic variants and / or haplotypes from the paternal genome and / or maternal genome to the embryo; and outputting the presence or absence of disease-causing variants and / or haplotypes in the embryo. Some methods further comprise outputting a combined disease risk associated with the embryo based on the transmission of polygenic disease risks and genetic variants causing single-gene diseases and / or haplotypes from the paternal genome and / or maternal genome to the embryo.

[0008] In some aspects, the method further includes using the paternal genomic data and / or the maternal genomic data to determine the paternal haplotype and / or the maternal haplotype. In some aspects, the method further includes using population genotype data and / or population allele frequencies to determine the disease risk of the embryo. In some aspects, the method further includes using family history of the disease and / or other risk factors to predict the disease risk.

[0009] In some aspects, whole-genome sequencing is performed using standard, PCR-free, ligated read (i.e., synthetic long read) or long-read protocols. In some aspects, rare genotyping is performed using microarray technology; next-generation sequencing of embryonic biopsies; or sequencing of cell culture media. In some aspects, phasing is performed using population-based and / or molecular-based methods (e.g., ligated reads). In some aspects, polygenic risk scores are determined by summing the effects across loci in a disease model.

[0010] In some aspects, the population genotype data comprises the allele frequencies and individual genotypes of at least about 300,000 unrelated individuals in the UK Biobank. In some aspects, the population phenotype data comprises both the self-report and clinical report (such as ICD-10 code) phenotypes of at least about 300,000 unrelated individuals in the UK Biobank. In some aspects, the population genotype data comprises population family history data, which comprises the self-report data of at least about 300,000 unrelated individuals in the UK Biobank and the information derived from the relatives of those individuals in the UK Biobank. In some aspects, disease risk is further determined by the score of the genetic information shared by the affected individual.

[0011] Also provided is a method for determining disease risk for one or more potential children, the method comprising: performing whole genome sequencing on (i) the intended mother and one or more potential sperm donors or (ii) the intended father and one or more potential egg donors; phasing the genomes of (i) the intended mother and one or more potential sperm donors or (ii) the intended father and one or more potential egg donors; simulating gametes based on recombination rate estimates; combining the simulated gametes to generate genomes of one or more potential children; assigning a polygenic risk score; and determining a distribution of disease probabilities based on the polygenic risk score.

[0012] Also provided is a method for outputting a probability distribution of disease risks for potential children, the method comprising: receiving a first dataset comprising genomic data of the intended mother; receiving one or more datasets comprising genomic data from one or more intended sperm donors; simulating gametes using estimated recombination rates (e.g., derived from the HapMap Consortium); generating genomes of one or more potential children using potential gamete combinations; estimating a polygenic risk score for the genome of each of the one or more potential children; and outputting a distribution of disease probabilities based on the polygenic risk scores.

[0013] Also provided are methods for determining the risk of a potential offspring of (i) an intended mother and a potential sperm donor or (ii) an intended father and a potential egg donor for a range of diseases, the methods comprising: (a) performing whole genome sequencing on the intended mother and one or more potential sperm donors to obtain a maternal genotype and one or more sperm donor genotypes or (ii) performing whole genome sequencing on the intended father and one or more potential egg donors to obtain a paternal genotype and one or more egg donor genotypes; (b) estimating a possible genotype for one or more potential offspring using (i) the maternal genotype and the potential sperm donor genotype or (ii) the intended father's genotype and the potential egg donor genotype; and (c) estimating a minimum possible polygenic risk score for the potential offspring using the possible genotypes of the potential offspring; and (d) estimating a maximum possible polygenic risk score for the potential offspring using the possible genotypes of the potential offspring.

[0014] Also provided is a method for outputting a series of disease risks for a potential child of (i) an intended mother and a potential sperm donor or (ii) an intended father and a potential egg donor, the method comprising: (a) receiving a first dataset comprising genomic data of the intended mother or genomic data of the intended father; (b) receiving one or more datasets comprising genomic data from one or more intended sperm donors or one or more intended egg donors; (c) deriving possible genotypes for the potential child using the genotypes of (i) the intended mother and the potential sperm donor or (ii) the intended father and the potential egg donor; (d) estimating a minimum polygenic risk score for the potential child by selecting the genotypes (of those derived in (c)) at each site in the model that minimize the score; (e) estimating a maximum polygenic risk score for the potential child by selecting the genotypes (of those derived in (c)) at each site in the model that maximize the score; and (f) outputting a series of disease risks using the minimum and maximum scores calculated in (d) and (e).

[0015] In some aspects, the method uses dense genotyping arrays on sperm donors, followed by genotype imputation for loci of interest that were not directly genotyped. In some aspects, the method uses family history of the disease and other relevant risk factors to determine disease risk.

[0016] In some aspects, whole genome sequencing is performed using standard, PCR-free, ligated reads (i.e., synthetic long reads), or long read protocols. In some aspects, phasing is performed using population-based and / or molecular-based methods (e.g., ligated reads). In some aspects, a polygenic risk score is determined by summing the effects across all loci in a disease model.

[0017] In some aspects, the population genotype data comprises allele frequencies and individual genotypes of at least about 300,000 unrelated individuals in the UK Biobank. In some aspects, the population phenotypic data comprises both self-reported and clinically reported (e.g., ICD-10 codes) phenotypes of at least about 300,000 unrelated individuals in the UK Biobank. In some aspects, the population family history comprises self-reported data of at least about 300,000 unrelated individuals in the UK Biobank and information derived from relatives of those individuals in the UK Biobank. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 An exemplary method for predicting and reducing disease risk is described.

[0019] Figure 2 A flowchart is depicted providing an exemplary method for determining a polygenic risk score.

[0020] Figure 3 An exemplary method for determining disease risk in a child is described.

[0021] Figure 4 Depicts illustrative inputs that may be used to determine disease probability.

[0022] Figure 5 Depicts a flow chart showing an exemplary method for selecting embryos based on disease likelihood.

[0023] Figure 6 Provides a graphical presentation of the risk reduction curves associated with specific diseases.

[0024] Figure 7 A flow chart is depicted providing an exemplary method for selecting a sperm donor.

[0025] Figure 8 Provides a graphical presentation of risk reduction curves generated for multiple donors for some autoimmune conditions.

[0026] Figure 9 An exemplary disease risk distribution associated with multiple sperm donors is provided.

[0027] Figure 10 A graphical representation of the ROC curve is provided, showing the improvement in predictive ability associated with determining the risk of prostate cancer.

[0028] Figure 11 An exemplary method for predicting disease risk associated with an embryo is illustrated.

[0029] Figure 12 An exemplary disease risk prediction graph associated with HLA typing for rheumatoid arthritis is shown.

[0030] Figure 13 An exemplary scaffold for identifying chromosome length phasing blocks to improve disease risk prediction is provided.

[0031] Figure 14 A graphical representation of the distribution of PRS (mean normalized to 0, standard deviation 1) for rheumatoid arthritis cases and controls is provided.

[0032] Figure 15 Shown are the ORs by deciles for rheumatoid arthritis.

[0033] Figures 16A-16C Shows lifetime risk of multiple conditions in several embryos, including Figure 16A The risk to the first embryo (called "embryo 2") is shown. Figure 16B The risk of the second embryo (called "embryo 3") is shown, and Figure 16CThe risk to the third embryo (called "embryo 4") is shown.

[0034] Figure 17A Shows lifetime risk and hazard ratio in several embryos compared with general population risk; Figure 17B The lifetime risk of the embryo is shown as a function of the polygenic risk score.

[0035] Figure 18 A diagram is provided of an exemplary parent support method for determining fetal disease risk.

[0036] Figure 19 Illustration of a potential workflow for genome-wide prediction of embryos.

[0037] Figure 20 Provides a diagram of how the whole chromosome phase of an individual can be obtained by performing whole genome sequencing of the individual, their partner, and two or more children and determining which loci are inherited by each child.

[0038] Figure 21 is a block diagram of an example computing device. DETAILED DESCRIPTION

[0039] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Unless otherwise indicated, the materials mentioned in the following description and examples were obtained from commercial sources.

[0040] As used herein, the singular forms "a", "an", and "the" refer to both the singular and the plural, unless expressly specified to refer to the singular only.

[0041] The term "about" means that the numerical value is understood not to be limited to the exact numerical value listed herein and is intended to refer to a numerical value substantially surrounding the numerical value without departing from the scope of the present invention. As used herein, "about" will be understood by those of ordinary skill in the art and will vary to some extent depending on the context in which it is used. If the use of the term is unclear to those of ordinary skill in the art based on the context in which it is used, "about" will mean up to plus or minus 10% of the particular term.

[0042] The term "gene" refers to a segment of DNA or RNA that encodes a polypeptide or performs a functional role in an organism. A gene can be a wild-type gene, or a variant or mutation of a wild-type gene. A "gene of interest" refers to a gene or a variant of a gene that may or may not be known to be associated with a particular phenotype or the risk of a particular phenotype.

[0043] "Expression" refers to the process of transcribing a polynucleotide from a DNA template (such as into mRNA or other RNA transcripts) and / or the subsequent translation of the transcribed mRNA into a peptide, polypeptide, or protein. Gene expression encompasses not only cellular gene expression, but also the transcription and translation of nucleic acids in cloning systems and in any other context. In cases where a nucleic acid sequence encodes a peptide, polypeptide, or protein, gene expression involves the production of nucleic acid (e.g., DNA or RNA, such as mRNA) and / or peptide, polypeptide, or protein. Thus, "expression level" can refer to the amount of nucleic acid (e.g., mRNA) or protein in a sample.

[0044] "Haplotype" refers to a group of genes or alleles that are inherited or expected to be inherited together from a single ancestor (such as a father, mother, grandfather, grandmother, etc.). The term "ancestor" refers to a person of whom a subject is a descendant, or, in the case of an embryo, a potential subject is a descendant. In preferred aspects, an ancestor refers to a mammalian subject, such as a human subject.

[0045] Diseases and methods

[0046] Provided are methods for identifying diseases that are entirely or partially genetically caused, or the risk of having or inheriting a disease. Genetic disorders can be caused by mutations in one gene (monogenic disorders), mutations in multiple genes (polygenic disorders), a combination of genetic mutations and environmental factors (multifactorial disorders), or chromosomal abnormalities (changes in the number or structure of entire chromosomes (the structures that carry genes)). In some aspects, the disease is a polygenic disorder, a multifactorial condition, or a rare monogenic disorder (e.g., one that has not been previously identified in the family).

[0047] Some aspects include determining whether the embryo is a carrier of a hereditary disorder. Some aspects include determining whether the embryo will develop into a subject having or likely having a hereditary disorder. Some aspects include determining whether the embryo will develop into a subject having or likely having one or more phenotypes relevant to a hereditary disorder.

[0048] Some aspects include selecting embryos based on their genetic makeup. For example, some aspects include selecting embryos with a low risk of carrying a genetic condition. Some aspects include selecting embryos that, if they develop into children or adults, will have a low risk of carrying a genetic condition. Some aspects include implanting the selected embryos into the uterus of a subject. Such methods are described in great detail, for example, in Balaban et al., “Laboratory Procedures for Human In Vitro Fertilization,” Semin. Reprod. Med., 32(4):272-82 (2014), which is incorporated herein by reference in its entirety.

[0049] Some aspects include assessing disease risks associated with embryos created using one or more sperm donors. Some aspects include selecting sperm donors based on disease risk. Some aspects include fertilizing eggs in vitro with the selected sperm.

[0050] Some aspects include determining an individual's health profile, such as based on the presence or absence of polygenic or rare single-gene variants. Some aspects include determining a distribution of disease probabilities, such as based on a polygenic risk score.

[0051] Screenable diseases are not limited. In some aspects, the disease is an autoimmune condition. In some aspects, the disease is associated with a specific HLA type. In some aspects, the disease is cancer. Exemplary conditions include coronary artery disease, atrial fibrillation, type 2 diabetes, breast cancer, age-related macular degeneration, psoriasis, colorectal cancer, deep vein thrombosis, Parkinson's disease, glaucoma, rheumatoid arthritis, celiac disease, vitiligo, ulcerative colitis, Crohn's disease, lupus, chronic lymphocytic leukemia, type 1 diabetes, schizophrenia, multiple sclerosis, familial hypercholesterolemia, hyperthyroidism, hypothyroidism, melanoma, cervical cancer, depression, and migraines. Some exemplary diseases include single gene disorders (e.g., sickle cell disease, cystic fibrosis), chromosome copy number disorders (e.g., Turner syndrome, Down syndrome), repeat expansion disorders (e.g., Fragile X syndrome), or more complex polygenic disorders (e.g., type 1 diabetes, schizophrenia, Parkinson's disease, etc.). Other exemplary diseases are described in PHYSICIANS' DESK REFERENCE (PRD Network 71st ed. 2016) and THE MERCK MANUAL OF DIAGNOSIS AND THERAPY (Merck 20th ed. 2018), each of which is incorporated herein by reference in its entirety. By definition, genetically complex diseases have multiple genetic loci that contribute to disease risk. In these cases, a polygenic risk score can be calculated and used to stratify embryos into high-risk and low-risk categories.

[0052] Embryonic genome construction

[0053] Provided is a novel and creative method for relating to embryonic genome construction. In some aspects, construction uses chromosome length parent haplotype and the rare genotyping (such as using SNP arrays or low coverage DNA sequencing) of parent and embryo to realize the full genome prediction in embryo. Such heterozygous approach can combine the genetic information from parent and other relatives (if available) (such as grandparents and siblings (i.e., brothers and sisters)) and use molecular methods (such as long fragment reading technology, 10X Chromium technology, Minion system) to directly obtain haplotype (such as dense haplotype block) from DNA. Chromosome length haplotype can be used for predicting the genome of the embryo in the setting of in vitro fertilization. The genomic sequence of such prediction can be used for predicting disease risk, both by directly measuring the transmission of the variant causing Mendelian disease, and by building polygenic risk score to predict disease risk.

[0054] In some aspects, the embryonic genome is constructed using the haplotype from two or more ancestors. In some aspects, the embryonic genome is constructed using paternal haplotype and maternal haplotype. In some aspects, the haplotype is the grandfather's haplotype. In some aspects, the haplotype is the grandmother's haplotype. In some aspects, the embryonic genome is constructed using paternal haplotype, maternal haplotype, and one or both of the grandfather's haplotype and the grandmother's haplotype. In some aspects, the rare embryonic genotype is obtained by sequencing the DNA obtained by the cell-free DNA in the embryonic culture medium, blastocyst fluid or the trophectoderm cell biopsy of the embryo.

[0055] Some aspects include determining one or more haplotypes for building embryonic genome.For example, such haplotype can be determined based on the genome sequence of ancestral subject.Some aspects include identifying the genome relevant to ancestral subject.Some aspects include implementing whole genome sequencing to identify the genome of ancestral subject to the biological sample obtained from ancestral subject.Some aspects include using one or more sibling embryos to determine haplotype.Such whole genome sequencing can be implemented using any one of multiple technologies, such as standard, without PCR, connection reading (for example, synthetic long read), or long read scheme. Exemplary sequencing techniques are disclosed, for example, in Huang et al., "Recent Advances in Experimental Whole Genome Haplotyping Methods," Int'l. J. Mol. Sci., 18 (1944): 1-15 (2017); Goodwin et al., "Coming of age: ten years of next-generation sequencingtechnologies," Nat. Rev. Genet., 17:333-351 (2016); Wang et al., "Efficient and unique cobarcoding of second-generation sequencing reads from long DNAmolecules enabling cost-effective and accurate sequencing, haplotyping, and denovo assembly," Genome Res., 29(5):798-808 (2019); and Chen et al. al., “Ultralow-inputsingle-tube linked-read library method enables short-read second-generationsequencing systems to routinely generate highly accurate and economical long-range sequencing information,” Genome Res., 30(6):898-909(2020), each of which is fully incorporated into this article by citation.

[0056] Genome phasing

[0057] Some aspects include phasing or imputing the ancestral genome to identify one or more haplotypes. For example, such phasing can be performed using population-based and / or molecular-based methods (such as linked read methods). Exemplary phasing techniques are disclosed in, for example, Choi et al., “Comparison of phasing strategies for whole human genomes,” PLoS Genetics, 14(4): e1007308 (2018); Wang et al., “Efficient and unique cobarcoding of second-generation sequencing reads from long DNA molecules enabling cost-effective and accurate sequencing, haplotyping, and denovo assembly,” Genome Res., 29(5): 798-808 (2019); and Chen et al., “Ultralow-input single-tube linked-read library method enables short-read second-generation sequencing systems to routinely generate highly accurate and economical long-range sequencing information,” Genome Res., 30(6): 898-909 (2020), each of which is incorporated herein by reference in its entirety.

[0058] In some aspects, phasing uses self-ligated read sequencing, long fragment reads, fosmid pool-based phasing, continuity-preserving transposon sequencing, whole genome sequencing, Hi-C methods, dilution-based sequencing, targeted sequencing (including HLA typing), or microarray-generated data.

[0059] Some aspects include using independently obtained rare phased genotypes to provide a scaffold to guide phasing. Computer software such as HapCUT, SHAPEIT, MaCH, BEAGLE or EAGLE can be used to phase the genotype of an ancestor. In some cases, computer programs use reference panels such as 1000Genomes or the Haplotype Reference Consortium to phase the genotype. In some cases, phasing accuracy can be improved by adding genotype data from relatives such as grandparents, siblings, or children.

[0060] Predicting embryonic genome sequence

[0061] Some aspects comprise using the rare genotyping combination of phased parental genome and embryo to predict the genome of embryo, which can allow determining the presence / absence of clinically relevant variants identified in parents and in embryo. This can be extended to include the risk / susceptibility alleles identified in parents and HLA types. In some aspects, rare genotyping is obtained using next generation sequencing. Rare genotyping is described in great detail in Kumar et al., “Whole genome prediction for preimplantation genetic diagnosis,” Genome Med., 7(1): Article 35, pages 1-8 (2015); Srebniak et al., “Genomic SNP array as a gold standard for prenatal diagnosis of foetal ultrasound abnormalities,” Molceular Cytogenet., 5: Article 14, pages 1-4 (2012); and Bejjani et al., “Clinical Utility of Contemporary Molecular Cytogenetics,” Annu. Rev. Genomics Hum. Genet., 9: 71-86 (2008), each of which is incorporated herein by reference in its entirety.

[0062] Rare genotyping can be implemented to the part of embryo extraction.Therefore, some aspects include extracting or obtaining one or more cells (such as via biopsy) from embryo.Some aspects include extracting or obtaining nucleic acid (such as DNA) from embryo or from one or more cells from embryo.Some aspects include extracting embryo material from embryo culture medium.

[0063] Some aspects use rare embryonic genotypes as scaffolds to phase the ancestral subject genome. Some aspects use information from one or more grandparent subjects (e.g., grandfather and / or grandmother subjects) to phase the parental genome. Some aspects use information from large reference panels (e.g., population-based data) to phase the parental genome.

[0064] In some aspects, the embryo is reconstructed using biological samples obtained from one or more ancestral subjects. Exemplary biological samples include one or more tissues selected from brain, heart, lung, kidney, liver, muscle, bone, stomach, intestine, esophagus, and skin tissue; and / or one or more biological fluids selected from urine, blood, plasma, serum, saliva, semen, sputum, cerebrospinal fluid, mucus, sweat, vitreous humor, and milk. Some aspects include obtaining a biological sample from a subject.

[0065] Some aspects include determining the transmission probability of one or more ancestral haplotypes. In some aspects, the transmission of variants from one or more maternal heterozygous loci can involve sequencing the maternal genome, sequencing or genotyping one or more biopsies from embryos, assembling or phasing the maternal DNA sample into haplotype blocks, utilizing information from multiple embryos (e.g., parental support techniques) to construct parental chromosome length haplotypes, and using statistical methods such as HMMs to predict the inheritance or transmission of these haplotype blocks. In some aspects, HMMs can also predict transitions between haplotype blocks or correct errors in maternal phasing.

[0066] Methods for predicting the transmission of variants from one or more paternally heterozygous loci can involve sequencing the paternal genome, sequencing or genotyping one or more biopsies from embryos, assembling or phasing the paternal DNA sample into haplotype blocks, using information from multiple embryos to improve the contiguity of haplotype blocks with chromosome length, and using statistical methods such as HMMs to predict the inheritance or transmission of these haplotype blocks. In some aspects, HMMs can also predict transitions between haplotype blocks or correct errors in maternal phasing.

[0067] The situation that both the mother and the father are heterozygous can be predicted in the above-mentioned manner.In the situation that both parents are homozygous for the same allele or different allele, the genotype of embryo is easily predicted.

[0068] In some aspects, the transmission probability is determined using the methods described in U.S. Application Serial Nos. 11 / 603,406; 12 / 076,348; or 13 / 110,685; or PCT Application Nos. PCT / US09 / 52730 or PCT / US10 / 050824, each of which is incorporated herein by reference in its entirety. In some aspects, regions with a transmission probability of 95% or greater are used to construct the embryonic genome.

[0069] In some respects, the embryonic genome is constructed using one or more genes or genetic variants in the embryo. In some respects, one or more genes or genetic variants are identified using rare genotyping of the embryo. In some respects, rare genotyping is implemented using microarray technology.

[0070] In some aspects, the embryonic genome is constructed using (i) one or more genetic variants in the embryo, (ii) one or more ancestral haplotypes (e.g., paternal haplotypes and maternal haplotypes), and (iii) the transmission probability of one or more haplotypes (e.g., paternal haplotypes and maternal haplotypes). In some aspects, rare genotyping is implemented using next generation sequencing.

[0071] Some aspects include embryonic genomic prediction using 1) whole-genome sequences from both grandparents on each side of the family, 2) phased whole-genome sequences from each parent, 3) rare genotypes of the parents measured by arrays, and 4) rare genotypes of the embryo. Without being bound by theory, it is believed that for a well-studied CEPH family, using this approach, 99.8% prediction accuracy can be achieved across 96.9% of the embryonic genome.

[0072] Some aspects include phasing the parental genome using 1) WGS of a single grandparent, 2) rare parental genotypes measured by array, and 3) a haplotype-resolved reference panel. Some aspects include phasing the parental genome using 1) rare parental genotypes measured by array, and 2) a haplotype-resolved reference panel (e.g., 1000 Genomes). Some aspects include phasing the parental genome using only a haplotype-resolved reference panel (e.g., 1000 Genomes).

[0073] Risk determination

[0074] Also provided is a method for determining the disease risk associated with the embryo (e.g., based on a genome of the structure of the embryo). Some aspects include determining whether the genetic variant causing the disease from the ancestral genome has been delivered to the embryo. Some aspects include determining whether a haplotype (e.g., related to the genetic variant causing the disease) has been delivered to the embryo. Some aspects include determining the presence or absence of a genetic variant causing the disease or improving disease susceptibility, including, but not limited to, single nucleotide variants (SNVs), small insertions / deletions, and copy number variants (CNVs). Some aspects include determining the presence or absence of a disease-related HLA type in the embryo.

[0075] In some aspects, one or more diseases (e.g., a panel of diseases) that can be ranked based on age of onset and disease severity can be used to determine phenotypic risk in an embryo. In some aspects, disease ranking can be combined with polygenic risk prediction to rank embryos according to potential disease risk.

[0076] Some aspects include determining that an embryo has a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, or more risk of a disease. Some aspects include determining that an embryo has a 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, 1%, or less risk of a disease. Some aspects include selecting embryos based on disease risk (e.g., selecting embryos with a relatively low risk of a disease) and / or based on the presence or absence of specific genetic variants (e.g., SNVs, haplotypes, insertions / deletions, and / or CNVs).

[0077] In some aspects, the disease risk associated with an embryo is determined using a polygenic risk score. In some aspects, a polygenic risk score (also known as a "PRS") is determined by summing the effects across loci in a disease model. In some aspects, a polygenic risk score is determined using population data. For example, population data can involve allele frequencies, individual genotypes, self-reported phenotypes, clinically reported phenotypes (e.g., ICD-10 codes), and / or family history information (e.g., derived from related individuals in one or more population databases). Such population data can be obtained from any of a number of databases, including the United Kingdom (UK) Biobank (which has information on approximately 300,000 unrelated individuals); various genotype-phenotype datasets as part of the Database of Genotypes and Phenotypes (dbGaP) maintained by the National Center for Biotechnology Information (NCBI); the European Genome-Phenome Archive; OMIM; GWASdb; PheGen1; the Genetic Association Database (GAD); and PhenomicDB.

[0078] In some aspects, disease risk is determined based on a polygenic risk score cutoff. For example, such a cutoff can include the top approximately 1% of a PRS distribution, the top approximately 2% of a PRS distribution, the top approximately 3% of a PRS distribution, the top approximately 4% of a PRS distribution, or the top 4% of a PRS distribution. Preferably, the cutoff is based on the top 3% of a PRS distribution. Polygenic risk score cutoffs can also be determined based on an absolute risk increase, such as approximately 5%, approximately 10%, or approximately 15%. Preferably, the polygenic risk score cutoff is determined based on an absolute risk increase of 10%.

[0079] Some aspects include using the predicted genome of the embryo to estimate phenotypic risk. In some aspects, the risk estimate uses 1) the predicted genome of the embryo, 2) the genotypes of the parents at loci of interest that were not predicted in the embryo (i.e., variants included in the polygenic risk score), and 3) allele frequencies in a reference cohort (e.g., UKBB) at loci of interest that were not predicted in the embryo (e.g., variants included in the polygenic risk score).

[0080] Some aspects include determining risk based on the transmission probability of one or more genetic variants (e.g., based on ancestral haplotypes). Some aspects include determining the combined risk associated with the embryo based on the transmission probability of a polygenic disease risk and one or more genetic variants (e.g., causing the transmission of a genetic variant of a monogenic disease and / or haplotype from paternal genome and / or maternal genome to the embryo).

[0081] A non-limiting exemplary system for predicting and reducing disease risk is Figure 1 A non-limiting exemplary polygenic risk score workflow is shown in Figure 2 Displayed in.

[0082] Donor selection

[0083] Also provided are methods for selecting sperm and / or egg donors. An estimated risk of a subject passing a disease to their offspring can be calculated by simulating the genomes of virtual children and calculating the disease risk for each child. Some aspects include determining the disease risk of the intended mother and one or more potential sperm donors. Some aspects include determining the disease risk of the intended father and one or more potential egg donors.

[0084] Some aspects include simulating gametes from potential mothers and fathers using phased parental genomes and simulated haplotype recombination sites (e.g., as determined using the HapMap database). Some aspects consider the respective recombination rates during meiosis that produced these gametes. In some aspects, these simulated gametes are combined with each other to generate numerous combinatorial possibilities to approximate the range of potential child genomes. Such child genome arrays can be converted into disease probability arrays to predict the distribution of disease risks across each child. See Figure 3 .

[0085] Risk estimates as described herein (e.g., in the Embryo Genome Construction section and / or the Examples section) can be used in the context of family planning for embryo selection during an IVF cycle and / or sperm donor selection. In some embodiments, potential parents receive a report containing individual risk estimates for multiple phenotypes across all available embryos or a series of risk values for each potential sperm donor. In some aspects, sperm donors are ranked based on disease risk for a condition or a group of conditions. In some aspects, donors are selected using a python script disclosed in U.S. Provisional Application No. 63 / 062,044, filed on August 6, 2020, or a modified version thereof.

[0086] Some aspects include selecting embryos based on a risk score. Some aspects include selecting an egg donor based on a risk score. Some aspects include selecting a sperm donor based on a risk score.

[0087] Implementation system

[0088] The methods described herein can be implemented on a variety of systems. For example, in some aspects, a system (e.g., for genomic embryo construction, donor selection, risk determination, and / or implementation of health reporting) includes one or more processors coupled to a memory. The methods can be implemented using code and data stored and executed on one or more electronic devices. Such electronic devices can store and communicate (internally and / or with other electronic devices via a network) code and data using computer-readable media, such as non-transitory computer-readable storage media (e.g., magnetic disks; optical disks; random access memory; read-only memory; flash memory devices; phase-change memory) and transient computer-readable transmission media (e.g., electrical, optical, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, or digital signals).

[0089] The memory can be loaded with computer instructions to train the model as needed (e.g., for identifying disease risk). In some aspects, the system is implemented on a computer, such as a personal computer, a laptop computer, a workstation, a computer terminal, a network computer, a supercomputer, a massively parallel computing platform, a television, a mainframe, a server farm, a widely distributed collection of loosely networked computers, or any other data processing system or user device.

[0090] The method may be implemented by processing logic comprising hardware (e.g., circuitry, dedicated logic, etc.), firmware, software (e.g., embodied on a non-transitory computer-readable medium), or a combination of the two. The operations described may be performed in any sequential order or in parallel.

[0091] Typically, a processor can receive instructions and data from read-only memory or random access memory, or both. A computer typically contains a processor capable of performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes or is operatively coupled to receive or transfer data from or to one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, optical disks, or solid-state drives) for storing data, or both. However, a computer need not have such devices. Furthermore, a computer may be embedded in other devices, such as smartphones, mobile audio or media players, game consoles, global positioning system (GPS) receivers, or portable storage devices (e.g., universal serial bus (USB) flash drives), to name a few. Suitable devices for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard drives or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.

[0092] A system of one or more computers can be configured to perform a specific operation or action by having software, firmware, hardware, or a combination thereof installed on the system that, when operated, causes the system to perform the action. One or more computer programs can be configured to perform a specific operation or action by including instructions that, when executed by a data processing device, cause the device to perform the action.

[0093] An exemplary implementation system is Figure 21 Such systems may be used to implement one or more of the operations described herein. The computing device may be connected to other computing devices in a LAN, an intranet, an extranet, and / or the Internet. The computing device may operate in the capacity of a server machine in a client-server network environment or in the capacity of a client machine in a peer-to-peer network environment.

[0094] The following examples are provided to illustrate the invention, but it should be understood that the invention is not limited to the specific conditions or details of these examples.

[0095] Example

[0096] Example 1: Parental Genomic Phasing for Parental Recurrence Risk Assessment and Disease Prediction in Embryos for Preimplantation Genetic Testing - Use of Predicted Embryonic Genomic Sequences in In Vitro Fertilization (IVF)

[0097] Embryo coverage and accuracy were calculated using three different approaches. According to the first approach, embryonic genome predictions were made using 1) whole genome sequences (WGS) from both grandparents on each side of the family, 2) phased WGS from each parent, 3) rare genotypes of the parents measured by array, and 4) rare genotypes of the embryo ( Figure 4 For one well-studied CEPH family, the protocol achieved 99.8% prediction accuracy across 96.9% of the embryonic genome. (Also covered is a protocol using 1) WGS of a single grandparent, 2) sparse parental genotypes measured by array, and 3) a haplotype-resolved reference panel.)

[0098] According to the second approach, embryo predictions use 1) sparse parental genotypes measured by array and 2) a haplotype-resolved reference panel (e.g., 1000 Genomes).

[0099] According to the third approach, embryo predictions are made using only a haplotype-resolved reference panel (e.g., 1000 Genomes).

[0100] The results from all three schemes are shown below in Table 1. The PRS showed results for approximately 1.4 million sites that were important in disease risk prediction.

[0101] Table 1: Embryo coverage and accuracy achieved with various phasing strategies

[0102]

[0103] Example 2: Using Predicted Embryonic Genomes to Estimate Phenotypic Risk

[0104] The probability of using a possible genotype (AA, AB, BB) given a parental genotype (M, D) at an unpredicted site in the embryo's genome (see Equation 1 below). In cases where parental genotypes are not available, the cohort-affected allele frequencies (AFs) are used. EA )(Equation 2).

[0105] Equation 1: β*P(AA|M,D)+β*P(AB|M,D)+β*P(BB|M,D)

[0106] Equation 2: 2*β*AF EA

[0107] Predicted embryos fell within the risk score percentile of the true score for 27 of the 30 models (90%).

[0108] Another approach involves using 1) the embryo's predicted genome and 2) allele frequencies in a reference cohort (e.g., UKBB) at loci of interest (i.e., variants included in the polygenic risk score) that were not predicted in the embryo. Allele frequencies are used as described in Equation 2 above. Using this approach, the predicted embryos fell within the risk score percentile for 23 of the 30 models (77%). When parental genotypes were included, all 30 predicted scores fell within 5% of the true score.

[0109] Example 3: Estimating and Improving Phenotypic Risk Estimates Using Polygenic Risk Models

[0110] Statistical framework

[0111] The robust model used for disease simulation and empirical analysis is the threshold liability model. Diseases are assumed to have a genetic component g~N(0,h 2 )(where h 2 is the narrow sense heritability) and the error component∈~N(0,1-h 2 ). The assumed responsibility l is given by:

[0112] l=g+∈~N(0,1)

[0113] is called potential liability, and it is assumed that the sample has risk on the potential liability scale. The threshold T is estimated based on the disease prevalence p, so that This is calculated based on the distribution of a standard normal random variable. Without being bound by theory, it is believed that everyone affected by the disease has l>T.

[0114] Modeling families involves modeling genetic liability as the sum of three components: two genetic components—the part measured by the PRS, the “unmeasured” part that is simply the residual genetic risk, and the irreducible non-genetic error. The potential genetic risk g from above can be decomposed into

[0115] g=g R +g U

[0116] Defined as

[0117] gU=g-gE

[0118] This last component is uncorrelated among family members. On the other hand, if the variance explained by the PRS on the responsibility scale is σ 2 , and g R,i and g R,j is the PRS component of the liability of the two first-degree relatives, then the covariance is given by:

[0119]

[0120] If g U,i and g U,j is the remaining unmeasured component of the liability of the two first-degree relatives, and h 2 is the heritability of the trait, then the covariance is given by:

[0121]

[0122] If g i are the children of g1 and g2, then

[0123]

[0124] For two first-degree relatives i and j with the following responsibilities,

[0125] l i =g R,i +g U,i +∈ i

[0126] l j =g R,j +g U,j +∈ i

[0127] We can see

[0128]

[0129] Because the error terms are uncorrelated.

[0130] IVF Embryo Selection Simulation

[0131] An IVF simulation is performed to answer the following question: Given a set of n embryos and a clinical phenotype of interest, how much less likely is the embryo with the minimum polygenic risk score to develop the disease during its lifetime than a randomly selected embryo? In other words, what is the relative risk reduction of the selection?

[0132] To answer this question, we used a two-step procedure to generate parameters for parents and subsequently their offspring. This procedure, or a variation of it, was used in simulations testing the effectiveness of donor selection and IVF embryo selection.

[0133] The following inputs are used in the embryo selection model: 2 , the variance explained by the polygenic risk score on the liability scale; h 2 , the additive heritability of traits on the liability scale; p, the lifetime prevalence of traits.

[0134] The output from this simulation is the risk reduction across different numbers of available embryos, which allows prospective couples undergoing IVF to find out which diseases can be meaningfully screened for.

[0135] Regulations

[0136] Step 1. For each parent, generate a distribution N(0, σ) if drawn from the general population or some other distribution (such as a variation of the mean or a truncated normal distribution) to represent elevated risk from family history. 2 ) of PRSg R . The distribution is N(0,h 2 -σ 2 ) or the remaining unmeasured genetic risk for the other conditions above g U .

[0137] Step 2. By calculating l1,…,l n To simulate n children:

[0138] Calculate the mean PRS of the parents:

[0139]

[0140] Calculate the mean residual genetic risk of the middle parent:

[0141]

[0142] For each child, the distribution is calculated as N(0,1-h 2 )’s independent error∈ i .

[0143] For each child, calculate the independent PRS recombination

[0144]

[0145] For each child, calculate the independent unmeasured / residual risk from the recombination

[0146]

[0147] Calculate the responsibility of child i by summing

[0148] l i =M R +M U +R P,i +R U,i +∈ i

[0149] Step 3. To determine the risk reduction, simulate millions of families with n = 3, 4, ..., 10. For each family, look at the responsibility l for the embryo with the smallest PRS. min Whether it exceeds the threshold t=Φ -1 (1-p), where Φ is the cumulative distribution function of the standard normal distribution.

[0150] Statistical Description

[0151] As a supplement, it can be shown that R P,i and R U,i To show that the covariances between siblings and between children and parents are correct, note that

[0152]

[0153] Because the last two terms are 0. The same calculation applies to the unmeasured genetic risk, that is

[0154]

[0155] So for g i =g R,i +g U,i ,

[0156]

[0157] A similar set of calculations shows that the parent-child covariance also satisfies the correct equation.

[0158] This procedure can be found in Figure 5 An example of a risk reduction curve with inputs is shown in Figure 6 The variance explained by the polygenic risk score is shown in Table 2 below, where “h2_lee” is the variance.

[0159] Table 2: Variance explained by polygenic risk scores for various conditions

[0160] Phenotype <![CDATA[h 2 _lee]]> Prevalence Disease type Heritability AMD 0.017064 0.0655 other 0.50 Breast cancer 0.026747 0.1240 cancer 0.31 Prostate cancer 0.051717 0.1160 cancer 0.58 CLL 0.045575 0.0057 cancer 0.60 psoriasis 0.079081 0.0400 Autoimmunity 0.75 Rheumatoid arthritis 0.017422 0.0140 Autoimmunity 0.60 Celiac disease 0.246643 0.0100 Autoimmunity 0.80 Crohn's disease 0.021475 0.0050 Autoimmunity 0.80 Type 1 diabetes 0.098359 0.0050 Autoimmunity 0.72 Type 2 diabetes 0.022617 0.2570 other 0.50 Atrial fibrillation 0.014569 0.2720 other 0.67 Bipolar disorder 0.030115 0.0250 Mental illness 0.55 Schizophrenia 0.035857 0.0050 Mental illness 0.80 Vitiligo 0.062567 0.0200 Autoimmunity 0.50 Inflammatory bowel disease 0.022788 0.0200 Autoimmunity 0.50

[0161] There are simulated donor families

[0162] To identify a lower-risk donor, the following is performed: (1) calculate the polygenic risk score for the intended mother, (2) calculate the polygenic risk score across N donors, and (3) select the donor with the lowest polygenic risk score. The procedure is essentially the same as above, with two changes: First, multiple donors (n = 10, 20, 30, ..., 100) are simulated, and the polygenic risk score is minimized over the polygenic risk score of the donors, rather than minimizing recombination. The flowchart of this method is in Figure 7 Displayed in.

[0163] Use the following input: 2 ,The variance explained by PRS on the responsibility scale;h 2 , the additive heritability of the trait on the liability scale; p, the lifetime prevalence of the trait. The output from this simulation is the risk reduction across the different numbers of available donors to be minimized, which allows clients using sperm or egg donors to find out which diseases can be meaningfully screened for. With the same example input as above, risk reduction curves are generated for different numbers of donors for some autoimmune conditions, which are Figure 8 Displayed in.

[0164] Additional embryo selection after donor selection

[0165] Another application of donor selection involves first selecting a donor and then selecting embryos with a lower disease risk. More specifically, a subject (e.g., a female subject) who is interested in using donor sperm for her children is provided with disease risk information. First, using her genetic test results and family history, multiple gametes are simulated and combined with the simulated sperm sample to obtain the risk of a known genetic cause of heart disease. This is her "personalized risk" of having a child with that condition and is a refinement of the "baseline risk." Second, using genetic information from each donor and information about which variants are phased with each other, a range of disease probabilities are calculated for gametes assuming they came from an individual donor. Finally, assuming that the donor was selected, multiple embryos (E1, E2, E3) fall within the disease risk distribution. See Figure 9 .

[0166] These methods can be used in the context of family planning during sperm donor selection. Potential parents can indicate which phenotypes they are particularly interested in, and risk scores for those phenotypes can be generated for each donor. These scores are used to predict disease risk in potential children of each sperm donor. Parents can be given a report containing these risk values, allowing them to select donors who will reduce their risk for the phenotypes of interest.

[0167] Family history

[0168] Family history can be incorporated to predict disease risk. In UK Biobank, there are several diseases for which parents and siblings self-report disease status: diabetes, heart disease, Alzheimer's disease, Parkinson's disease, breast cancer, and a few others. In addition, there are over 10,000 sibling pairs and a large number of half-sibling or other second-degree relative pairs. A model was constructed with a binary variable for family history, meaning: (i) for the set of diseases with a self-reported family history in UK Biobank, either a sibling or parent has the disease; or (ii) for any other disease, all samples have a first-degree relative in UK Biobank. Given this definition of the "family_history_presence" dummy, a logistic regression was run for each condition in the appropriate cohort using the following formula: log(P / (1-P)) = β_1*PRS + β_2*sex_male + β_3*family_history_presence.

[0169] In summary, the inputs include: data from a biobank containing self-reported family histories of disease, as well as pairs of first-degree relatives with medical records. The output includes: a model derived from logistic regression that incorporates PRS and family history to improve the accuracy of our predictions. The model is used to prioritize patients at higher risk of developing disease during their lifetime. An example output is listed in Table 3 below, which estimates β_1 (PRS), β_2 (sex dummy), and β_3 (family history dummy) for various conditions.

[0170] Table 3: Data from the logistic regression model incorporating the PRS

[0171]

[0172] When the family history dummy is added to the logistic regression, the improvement in prediction is quantified by the ROC curve for prostate cancer, as shown in Figure 10 Displayed in.

[0173] Increased model complexity

[0174] The model becomes more complex by including second and third degree relatives, more complex family trees, and / or related phenotypes. The above shows how to simulate first degree relatives. To allow for the inclusion of second degree family history, two additional family members can also be simulated for each parent. If P1 is a relative with R 1,iIf the parent is one, then we can generate the second-level family members by the following assumptions:

[0175]

[0176] where σ 2 Is PRS or unmeasured genetic risk g U The variance components of the potential responsibility scale.

[0177] Another layer of complexity can be added to the simulation: thresholds based on age and sex. If the prevalence of the disease varies depending on these variables, the thresholds can be adjusted to determine whether a sample from a family has the disease. For example, suppose that for type 2 diabetes, the prevalence is 20% in men over 80 years old, while the prevalence is 4% in women aged 55 years old. By substituting the empirical lifetime risk of the disease in the model above, the lifetime prevalence can be replaced by the lifetime risk. The thresholds for such samples would be 1-Φ(0.20) and 1-Φ(0.04), respectively, where Φ is the cumulative distribution function of the standard normal random variable. When one imposes a condition on a family pedigree, they are imposing a condition on a set of samples.

[0178] s i =g R,i +g U,i +∈ i >T i

[0179] Exceeding their age- and sex-specific thresholds T i .

[0180] Given a family tree Ped with information about the disease history, such as a father and grandfather with the disease, and three siblings without the disease, we can calculate

[0181] E(g U |ped)

[0182] The goal is to verify theoretical predictions about the quantity:

[0183] P(g R +g U +∈>T|g U =x)

[0184] This allows calculation of odds ratios.

[0185] HLA phenotype

[0186] Risk determination may involve phenotypes with a strong HLA component and associated HLA alleles that are not well-marked by SNVs. However, this approach can be applied to any condition with known disease associations with HLA alleles of significant effect sizes and involving other loci. Examples of complex phenotypes with HLA involvement include (but are not limited to) psoriasis, multiple sclerosis, type 1 diabetes, inflammatory bowel disease, Crohn's disease, ulcerative colitis, vitiligo, celiac disease, and systemic lupus erythematosus.

[0187] These methods can be applied in a variety of situations, including but not limited to individual disease risk prediction, risk reduction in both embryo selection and sperm donor selection scenarios, and prescribing guidelines for certain drugs where multiple genetic factors (including HLA type) influence the likelihood of response or adverse drug reaction.

[0188] HLA typing results are obtained from DNA-based methods, such as Sanger sequencing-based typing, or derived from whole-genome sequencing (WGS). First, a polygenic risk score is determined, for example, using the genome-wide association study (GWAS) effect size. One example is to sum the effect size and the effect allele dosage for all associated variants not in the MHC region. Second, associated HLA alleles are combined or pooled based on the HLA typing results (not the tag SNPs) using one of the following methods:

[0189] Combined PRS and HLA OR: Calculate the polygenic risk score for all individuals in the validation cohort to obtain metadata (e.g., mean, standard deviation, etc.). Obtain the odds ratio (OR) for HLA alleles with established associations with the phenotype of interest. Combine the OR derived from the individual's PRS compared to the validation cohort and HLA typing as follows:

[0190] OR=OR HLA *OR PRS *OR 入口统计学

[0191] The hazard ratio (RR) was calculated using the OR derived above and the prevalence of the disease in the validation cohort. This was then used to estimate the lifetime risk of the disease.

[0192] Incorporating HLA directly into the PRS: Incorporating HLA effect alleles directly into the polygenic risk score by adding the product of the effect size and the dose of each effect allele to the base PRS. This will be called the PRS HLA+ . PRS was calculated for all individuals in the validation cohort HLA+ , and obtain metadata (such as mean, standard deviation, etc.). HLA+ The model-derived ORs and the disease prevalence in the validation cohort were used to calculate the RR, which was then used to estimate the lifetime risk of disease.

[0193] Example 4: A method for ranking disease risk profiles for embryo and sperm donor selection

[0194] Provided are exemplary methods for ranking disease risk profiles, such as Figure 11 First, a weight w is calculated for each disease in a set of d diseases. d , that is, age of onset w a and disease severity s The sum of the weights of birth-onset diseases (e.g., celiac disease) a Greater than the w for diseases that typically don't appear until adulthood (like coronary artery disease) a Similarly, more serious diseases (like breast cancer) s Greater than w for diseases with milder phenotypes (like vitiligo) s .

[0195] Next, family history and polygenic risk scores are combined to generate a predicted risk for each condition of interest for each embryo.

[0196] Finally, the disease ranking and risk prediction are combined to generate a single score S for each embryo using the following equation: T , where RR is the relative risk derived from the combination of the polygenic risk score and family history for a given disease:

[0197]

[0198] Assume that the onset of w in adulthood, childhood, or at birth s are 0.5, 1, or 2, respectively. Similarly, assuming that w a The weights for a small set of conditions are given in Table 4 below, which are 0.5, 1, or 2, respectively, with the ability to select intermediate values for diseases with variable phenotypes:

[0199] Table 4: Weights of various conditions

[0200] disease Age of onset <![CDATA[w a ]]> Severity <![CDATA[w s ]]> <![CDATA[w d ]]> Breast cancer adult 0.5 Moderate-severe 1.5 2 Celiac disease born 2 Moderate 1 3 psoriasis childhood 1 Mild-moderate 0.75 1.75

[0201] Assuming three embryos have the following RRs for each of the above conditions, calculate the overall score for each embryo and rank them accordingly. For embryo 1, the score is calculated as follows:

[0202] S T =(2*2.4)+(3*1.4)+(1.75*2.7)=24.85

[0203] The disease risk for each of the three embryos is listed in Table 5 .

[0204] Table 5: Disease risk profile of three embryos

[0205] disease RR embryo 1 RR Embryo 2 RR embryo 3 Breast cancer 2.4 1.1 0.7 Celiac disease 1.4 1.6 1.4 psoriasis 2.7 7.3 2.7 <![CDATA[S T ]]> 13.7 19.8 10.3 Ranking 2 3 1

[0206] The same procedure is applied to sperm donor selection, where each donor receives a ranking across all diseases of interest. In both the embryo and donor selection contexts, scores are calculated for subsets of diseases (e.g., conditions for which the intended parents have a family history) or across all diseases for which the polygenic model is implemented.

[0207] Alternatively, this method can be used without summing all conditions of interest to prioritize the results for individual embryos / individuals. Each condition receives a score, and the condition with the highest score is prioritized. Using embryo 1 above as an example, the scores and rankings listed in Table 6 were generated.

[0208] Table 6: Embryo scores and rankings

[0209] disease RR embryo 1 <![CDATA[Disease score (RR*w d )]]> Disease ranking Breast cancer 2.4 4.8 1 Celiac disease 1.4 4.2 3 psoriasis 2.7 4.7 2

[0210] Example 5: Predicting the transmission of disease susceptibility variants to embryos

[0211] A copy of a colorectal cancer susceptibility variant (APC c.3920T>a) (and / or insertion, deletion, and / or copy number variant) was found in the father's WGS. This allele is not present in the mother. This variant is not directly measured in the rare genotyping of the embryo. The whole chromosome haplotype of the parents is obtained by any single or combined method described above. The reconstruction of the embryonic genome determines that the haplotype block containing the risk allele is passed from the father to one of the embryos. The risk allele is annotated as "present" in the embryo.

[0212] Example 6: Polygenic Risk of Common Diseases Using Embryo Prediction

[0213] Breast cancer has a common genetic component. A genetic risk score uses 69 variants to assess breast cancer risk. Of these variants, only 13% (9 / 69) were directly genotyped in the embryo. The percentile for the embryonic genetic risk score based on these variants was 84.6%. After embryo reconstruction, genotypes were imputed / inferred for 98.6% (68 / 69) of the embryos, resulting in a new embryonic genetic risk score percentile of 77.7%. After the embryos were born, the offspring's DNA was genotyped, resulting in a PRS percentile of 76.2%. This suggests that genetic risk scores derived from genome-wide embryo reconstruction have greater accuracy and less uncertainty due to the information from the additional variants.

[0214] Example 7: Predicting the transmission of disease-associated HLA types to embryos

[0215] The mother has rheumatoid arthritis (RA). HLA typing results (from WGS, PCR+Sanger sequencing, or any other appropriate method) reveal that she carries one copy of the HLA-DRB1*01:02 allele, which is associated with an increased risk for this condition. The father is homozygous for HLA-DRB1*04:02, an allele not known to be associated with an increased risk for RA. Based on complete phasing of chromosome 6 in each parent and reconstruction of the embryonic genome, it is determined that both the mother's haplotype 2 (HM2) and the father's haplotype 2 (HF2) are transmitted to the embryo. The RA risk allele is carried on the mother's haplotype 1 (HM1), and therefore the embryo is predicted not to carry the risk allele. See e.g. Figure 12 .

[0216] Example 8: Providing Families with a Profile of Disease Risk in Their Children

[0217] Two parents have expressed their interest in the risk of multiple genetic diseases in their future children. Using the method described above, the mean and recombinant values are calculated based on the genomes of the two parents to predict the disease risk range of their children, thereby providing guidance for future IVF treatment. Figure 9 .

[0218] Similarly, in the case of sperm donation, polygenic risk score distributions based on WGS of the mother and potential sperm donors can be simulated by recombination (see Figure 9 ).

[0219] Example 9: Incorporating Family History (FHx) to Improve Risk Estimates

[0220] Based on family history of the disease, the risk of developing psoriasis is estimated to be 10-30%. Using the polygenic model alone in embryos where one parent has psoriasis revealed only small differences in risk across embryos. Incorporating family history provided much better separation between embryo 1 and embryos 2 and 3, which clearly had other risk factors besides FHx, as shown in Table 7.

[0221] Table 7: Embryonic risk scores incorporating family history

[0222]

[0223] Similarly, family history can be incorporated to improve risk estimates for predicting the transmission of disease-associated HLA types.

[0224] Example 10: Incorporating HLA typing into psoriasis disease risk estimation

[0225] The presence or absence of two HLA types associated with the risk of developing psoriasis across embryos has a significant impact on overall disease risk. This example can be extended to the context of sperm donor selection or personal genome reporting, as shown in Table 8.

[0226] Table 8: Lifetime risk of psoriasis in multiple embryos

[0227] HLA-C*06:02 HLA-C*12:03 <![CDATA[OR prs ]]> RR Lifetime risk Embryo 1 Missing 1 copy 0.67 0.83 3.3% Embryo 2 1 copy 1 copy 0.75 2.91 11.6% Embryo 3 1 copy Missing 0.88 2.49 10.0%

[0228] Family history can be incorporated to further refine risk estimates in predicting the transmission of disease-associated HLA types. This technology can be extended to predict blood type from the embryonic genome, including the Rh status of the resulting fetus.

[0229] Example 11: Improving trait prediction accuracy

[0230] When the genotype of the variant in the polygenic model is unknown in the embryo, the parental genotype can be used to improve the trait prediction accuracy. The possible genotype probability given the parental genotype at the site is used, rather than the population allele frequency (AF) or the imputed genotype. Using the probabilities in Table 9 below, the dose of each possible genotype is added to the risk score. In practice, this improves the prediction accuracy measured by the predicted percentile of the polygenic risk, as shown in Table 10 below, which shows the improvement of the prediction of the polygenic model for Crohn's disease, in which 4 variants were not predicted in the embryo. The true polygenic risk score percentile ("true value") is determined using direct genotyping from WGS.

[0231] Table 9: Probability of embryo genotype based on parental genotype

[0232] Mother Father P(AA|M,D) P(AT|M,D) P(TT|M,D) AT TT 0 0.25 0.75

[0233] Table 10: Percentiles of polygenic risk scores

[0234] truth value Group AF dose 73.9% 62.5% 71.2%

[0235] Example 12: Haplotype Disease Risk

[0236] Some disease risks are based on phased haplotypes rather than individual variants. Embryo reconstruction generates phased haplotypes for more accurate prediction of trait risk. Table 11 below lists haplotypes in the gene APOE and their association with Alzheimer's disease risk (Corder et al., 1994).

[0237] Table 11: Haplotypes in APOE and associated risk of Alzheimer's disease

[0238] Haplotype rs429358 allele rs7412 allele Risk of Alzheimer's disease ε2 T T Protect ε3 T C neutral ε4 C C risk

[0239] The two variants are separated by 138 base pairs in the APOE gene. Neither rs429358 nor rs7412 were measured in the rare embryos. This precludes estimating Alzheimer's disease risk in the embryos. However, embryo reconstruction methods, using the parents' genotypes, predict a fully phased embryonic genome, which can be used to infer that the embryo is ε3 / ε3. This result was later validated by whole-genome sequencing of the offspring.

[0240] Table 12: Risk of Alzheimer's disease in reconstructed embryos

[0241] APOE haplotype Risk of Alzheimer's disease Mother ε3 / ε3 neutral Father ε3 / ε3 neutral Reconstructing embryos ε3 / ε3 neutral Embryos without reconstruction Unavailable Unavailable

[0242] Thus, embryo reconstruction enables APOE haplotype and Alzheimer's disease risk prediction, and disease status in general, based on haplotypes.

[0243] Example 13: Rare genotype scaffolds

[0244] Use of rare genotypes as scaffolds for whole genome phasing (see e.g. Figure 13 ) improves performance compared to the reference panel alone, as measured by the switch error rate (SER). Applying this technique to the well-studied sample NA12878, we saw an overall SER decrease from 0.6% when using the 1000 Genomes reference panel alone to 0.54% when using a set of approximately 140k high-confidence phased genotypes as scaffolds in combination with the reference panel. This difference is largely due to a reduction in long switch errors. For example, on chromosome 1, the raw number of long switch errors decreased by >60% (169 vs. 60). Overall, the combined approach (scaffold + reference panel) resulted in a reduction in the long switch error rate from 0.12% to 0.04%. This is important in embryo reconstruction, as long switch errors can lead to incorrectly predicted blocks.

[0245] Example 14: Polygenic Risk Score

[0246] Large-scale genome-wide association studies (GWAS) have identified genetic variants associated with a wide variety of diseases. These associations have paved the way for functional studies of disease biology, drug target discovery, and improved disease risk prediction. While individual common genetic variants may have little predictive value, combining these variants into genetic risk scores can explain a greater proportion of the genetic risk for disease. These multi-locus genetic risk scores, also known as polygenic risk scores (PRS), are most commonly calculated as a weighted sum of disease-associated genotypes.

[0247]

[0248] PRS indis the polygenic risk score for a given individual and disease with n associated variants, w i is the weight of the ith variant, usually taken from GWAS effect size, and G i is the individual's genotype for the risk allele of the i-th variant. PRSs have recently been investigated for their potential to predict the risk of a variety of diseases, including cardiovascular disease, breast cancer, and type 2 diabetes. These approaches have demonstrated the ability to stratify individuals according to their risk for these diseases.

[0249] Described is a method for validating and implementing polygenic models and visualizing risk estimates in consumer reports.

[0250] Choosing a polygenic risk model

[0251] Priority was given to previously published polygenic models for each condition of interest that had been tested on at least 1,000 individuals from a broad population. This excluded small studies with limited statistical power and studies tested on isolated populations that might not translate to other populations. Models using data from individuals in the UKBB study set were also excluded. Models that reported an area under the curve (AUC) greater than 0.65 and / or an odds ratio (OR) greater than 2 for individuals in the top quantile versus the bottom quantile were selected (see below for more information). A list of traits and published models and their evaluation statistics are shown in Table 13.

[0252] Table 13: Published disease models

[0253]

[0254]

[0255] When published models were not available, scores were constructed using SNPs from the GWAS catalog that met a genome-wide significance p-value threshold (p < 5e-8) as previously described (PMID: 30309464).

[0256] Defining each phenotype in UK Biobank

[0257] Each model was validated and standardized using data from the UK Biobank cohort. This resource includes genetic and disease information for 500,000 individuals. The following analysis used only unrelated individuals. As shown in Table 14, a combination of ICD-9 and ICD-10 codes, self-reported diseases, and procedure codes was used to define each phenotype of interest.

[0258] Table 14: UKBB phenotype definitions for each trait evaluated

[0259]

[0260]

[0261] A subset of diseases is shown in Table 15 below.

[0262] Table 15: Frequency of disease subsets in UK Biobank

[0263] disease frequency disease frequency Celiac disease 0.62% Atrial fibrillation 4.29% Coronary artery disease 6.64% Breast cancer 3.66%

[0264] Individuals were stratified according to their polygenic risk score (PGS) and the prevalence of disease in this group was investigated.

[0265] Evaluate the model using the UKBB dataset

[0266] A polygenic risk score was calculated as a weighted sum of disease-associated genotypes. The score was calculated for each individual in the UKBB, and the performance of the model was evaluated using multiple metrics.

[0267] PRS distribution across cases and controls

[0268] The dataset was decomposed into cases and controls for each trait, and distributions of scores were generated separately for cases and controls. Visual inspection of these distributions gave a rough idea of how well each model could distinguish cases from controls. For example, Figure 14 The distribution of PRS for rheumatoid arthritis cases and controls is shown (mean scaled to 0, standard deviation 1).

[0269] Receiver Operating Curve (ROC)

[0270] The ROC and area under the curve (AUC) were calculated by plotting the sensitivity and specificity of the model at different risk thresholds.

[0271] Stratified into PRS deciles

[0272] Individuals in the UK Biobank were stratified into groups with different risk profiles for disease. Individuals at highest risk (the top decile of PRS) were compared with those at median risk (those with a PRS in the middle 40-60 percentiles of the distribution). Disease prevalence was plotted across deciles for each disease, and the ratio of high risk to median risk was calculated across diseases. Figure 15 Shown are the ORs for rheumatoid arthritis by deciles.

[0273] Regression analysis including age and sex

[0274] Logistic regression was applied to each model after calculating the PRS across all unrelated individuals in the UK Biobank dataset. PGSis the regression coefficient of the PRS, corresponding to the odds ratio when the PRS is standardized to mean 0 and standard deviation 1. Age and sex were included where available and applicable.

[0275] LOR|GS=β0+β PRS PRS+β 年龄 Mean (age)

[0276] The odds ratios were then used to determine the threshold for high risk for the intermediate outcome for reporting purposes.

[0277] OR / SD by disease (centralized z-transformed means)

[0278] Based on the logistic model presented above, the PRS OR / SD was obtained by standardizing the PRS variables (to a mean of 0 and a standard deviation of 1), and then calculating the effect size. This process serves two purposes. First, it allows for direct comparison of the risk stratification capabilities of PRSs across diseases. PRSs for different diseases vary in the number of SNPs and their respective effect sizes, and therefore vary on very different scales. Without standardization, their corresponding effect sizes would not be directly comparable. By standardizing all PRSs, models can be directly ranked based on their OR / SD, resulting in a ranking that reflects their ability to stratify populations based on disease risk. Second, it allows for statistically accurate application of UKBB effect estimates to the US population. Effect sizes are estimated using the UKBB and then converted to odds ratios. When estimating relative risks from these odds ratios (see below), the US population disease prevalence is used to accurately capture the relative risk for individuals in the US with a given PRS. Standardization of the UKBB PRS (using the UKBB mean and standard deviation) allows for the use of PRSs from US individuals in the model (after adjusting for the US PRS mean and standard deviation). Due to random assortment in genetics, similar means and standard deviations of PRSs across the population can be expected, at least for individuals of European ancestry. The results from this analysis are shown in Table 16.

[0279] Table 16: Model validation statistics

[0280]

[0281]

[0282] PRS stratification of disease by age

[0283] After stratifying individuals into risk groups, we use UKBB data to estimate the percentage of individuals within these different groups who will be diagnosed with the disease. This information is visualized across the different strata, including high-risk (individuals in the top 5% of PRS) and average-risk (across the population). This is displayed as the predicted percentage of individuals diagnosed with the disease in a group with similar genetic risk to our individual of interest, assuming they have a PRS at the 75th percentile.

[0284] These figures help illustrate the utility of the PRS in stratifying individuals based on disease risk. Seeing a clear distinction in the proportion of groups diagnosed within different PRS strata confirms the model's ability to distinguish individuals based on their risk.

[0285] Calculate adjusted lifetime risk for an individual

[0286] We can start with the average lifetime risk for the sex of Americans. Next, we assess the risk markers in the genome and calculate a polygenic score based on the markers. This information is converted into an "odds ratio" using data from the UKBB described above. Finally, we use the formula that combines this odds ratio with the average lifetime risk to estimate the lifetime risk of an individual with this variation:

[0287]

[0288] Adjusted lifetime risk = c0*RR

[0289] Where p0 is the prevalence of the condition in the UKBB, c0 is the average lifetime risk of the condition in the United States, and OR is the odds ratio calculated above. The result is an estimate of an individual's lifetime risk compared to the population average. For some conditions, the average lifetime risk is not available. In these cases, indicate whether the genetics being analyzed indicate an increased risk.

[0290] Defining the threshold for “high risk”

[0291] In some cases, a threshold for high genetic risk is set based on known risk factors. For example, the relative risk of developing type 1 diabetes in an individual with an affected first-degree relative is 6.6. Therefore, the high-risk threshold for the PRS for type 1 diabetes is set to correspond to this relative risk. For phenotypes where this threshold is unavailable or the model cannot achieve it, we designate individuals at a 2-fold increased relative risk or a 10% increased absolute risk as high risk. The assessment measures for the subset of phenotypes where lifestyle or clinical factors inform the high-risk threshold are shown in Table 17.

[0292] Table 17: Evaluation of the model in a subset of unrelated UKBB individuals

[0293] disease Risk Factor (RR) PPV NPV % High Risk (%) Rheumatoid arthritis Smoking (1.9) 2.9% 98.9% 3.5% Coronary heart disease Family history (1.4) 9.8% 93.4% 3.7% Type 1 diabetes Family history (6.6) 1.9% 99.8% XX (4.9%)

[0294] Example 15: Multifactorial Condition (Polygenic Risk Score)

[0295] Genomic DNA obtained from the submitted samples was sequenced using Illumina or BGI technology. Reads were aligned to a reference sequence (hg19) and sequence variations were identified. For some genes, only specific variations were analyzed. Deletions and duplications were not examined unless otherwise noted above. In some cases, independent verification of HLA typing may have been performed by an external laboratory. Selected variants were annotated and interpreted according to ACMG (American College of Medical Genetics) guidelines. Only pathogenic or likely pathogenic variants were reported. Embryo and parental genotyping, followed by a "parent support" analysis, was performed. A genome reconstruction algorithm was used to reconstruct the embryonic genome using the embryonic genotypes and the parental whole-genome sequences. Only variants observed in the parental genomes that were predicted to affect the embryo were examined in the reconstructed embryonic genome. Polygenic risk scores were calculated for a subset of conditions. The model for each condition was evaluated on the UK Biobank cohort. Some polygenic risk scores can be refined using HLA typing. An individual's lifetime risk was calculated by adjusting their baseline risk (in a US cohort) based on their demographic information and polygenic risk score. Models that resulted in a 10% lifetime risk difference or a 1.9-fold increase in lifetime risk from the first to the last decile were included in the report. Some conditions (e.g., bipolar disorder) were retained in the experimental section at the investigators' discretion based on available evidence for model and genomic reconstruction performance. The lifetime risk of each condition for a given embryo was Figures 16A-16C Listed in.

[0296] Using psoriasis as a specific example, Figures 17A-17B Risk scores associated with susceptibility to psoriasis in three exemplary embryos are shown.

[0297] Example 16: Whole-genome prediction of embryos using haplotype-resolved genomic sequences

[0298] Haplotype-resolved genomic sequencing is combined with a panel of rare genotypes from single-cell or small-cell embryo biopsies to predict the embryo's full genome sequence. Specifically, stLFR technology is used for haplotype-resolved genomic sequencing of the father. Performance is evaluated at rare heterozygous positions (defined as allele frequencies of 1% or less). The inheritance of 230,117 loci in the embryo was predicted with 89.5% accuracy.

[0299] The material used in this study was retrospectively obtained from participants who had undergone a previous successful IVF cycle with preimplantation genetic diagnosis (PGD) (Table 16). Trophectoderm biopsies from a total of 10 embryos (day 5) were genotyped across a panel of 300,000 common SNPs using an accelerated 24-hour microarray protocol. In addition, each parent and all four grandparents were genotyped across the same panel.

[0300] Table 16: Tissue samples used for proof of concept

[0301]

[0302] Genomic DNA was extracted from whole blood or saliva samples. Neonatal and maternal DNA was processed using 30X WGS on the BGI platform. Paternal samples were processed using stLFR. DNA was extracted, amplified, and genotyped along with the parents and grandparents from a trophectoderm biopsy of one of ten day 5 embryos using a rapid microarray protocol across all samples using the Illumina CytoSNP-12 chip. Sibling embryo and parental SNP array measurements were combined using the "Parental Support" (PS) method ( Figure 18 , Figure 19 ), as detailed in Kumar et al., 2015. The embryo's whole genome sequence is predicted by combining the PS embryo genotype with the parental haplotype blocks (see Figure 18 ).

[0303] Example 17: Construction of whole chromosome haplotypes from haplotype blocks and parental information

[0304] To construct chromosome-length haplotypes in the IVF setting, haplotype-resolved genome sequencing of both parents is combined with information from sparse genotypes from sibling embryos. As part of the "Parent Support" (PS) approach, maximum likelihood estimates (MLEs) of heterozygous SNVs in each parent are created by combining recombination frequencies from the HapMap database with SNP array measurements from the parents and from the sibling embryos. This sparse, chromosome-length haplotype is insufficient to predict the embryo's genome but can be combined with dense molecularly derived haplotypes from parental samples (e.g., using long-read technology, 10x Genomics, CPT-seq, Pacific Biosciences, Hi-C) to predict the inherited genome sequence.

[0305] This information was obtained using several data streams. To generate dense haplotype blocks, shotgun sequencing was first performed at a median coverage of 34x for the mother and 30x for the father. Next, by sequencing a haploid subset of genomic DNA obtained by in vitro dilution pool amplification, 94.2% of the 1.94 million heterozygous SNVs in the mother and 92.4% of the 1.89 million heterozygous SNVs in the father were directly phased into long haplotype blocks. These molecularly derived "dense haplotype blocks" were combined with the sparse, but chromosome-length haplotypes to construct chromosome-length haplotype-resolved genomic sequences for the parents. This sequence information was subsequently used to predict the inherited genomic sequence of the embryo, but it can also be used to predict the potential offspring of both parents (for example, by simulating the potential eggs and sperm that would produce future children).

[0306] A potential workflow for genome-wide prediction of embryos is Figure 19 At the initial visit, the patient gives blood, which is used to generate a whole-genome sequence for each parent and predict possible conditions the couple is at risk for. After counseling, the parents undergo IVF, and the embryos are genotyped using conventional IVF PGD technology. This information is combined with the parents' whole-genome sequence information (haplotype-resolved) to predict the embryo's inherited genome and assess disease risk.

[0307] The chromosome length parental haplotypes were constructed using sibling embryo and parental genotypes. Parental phase was determined using statistical methods (e.g., maximum likelihood estimation) based on noise information obtained from each sibling embryo and a database of meiotic recombination frequencies.

[0308] Whole chromosome haplotype construction

[0309] Whole chromosome haplotypes are constructed by sequencing the genomes of an individual's relatives (including but not limited to parents, grandparents, or children). If an individual has two or more children from the same person, a whole chromosome haplotype for the individual can be obtained by performing whole genome sequencing on the individual, their partner, and two or more children and determining the loci inherited by each child. Figure 20 This provides whole-chromosome haplotype information without modifying the DNA sequencing process. This can be important, for example, if a couple already has two children and wishes to have another, and does not have any grandparent DNA samples.

[0310] Chromosomal haplotypes from individual sperm

[0311] The method of Example 17 was performed using whole chromosome haplotypes obtained by sequencing DNA obtained from individual sperm.

[0312] Example 18: Calculating Polygenic Risk Scores for Complex Genetic Diseases Using Embryonic Genomic Predictions

[0313] Genome-wide association studies can model polygenic risk scores for conditions such as type 1 diabetes, schizophrenia, Crohn's disease, celiac disease, and Alzheimer's disease. These approaches involve obtaining a list of genome-wide significant SNPs with observed odds ratios for disease-associated SNPs and calculating a "risk score" for each individual based on the set of SNPs observed in that individual. This approach was used to calculate polygenic risk scores for siblings to mimic the polygenic risk scores observed when comparing sibling embryos in IVF cycles. Genome sequences from a publicly available family with 12 siblings, two parents, and four grandparents were used. Each genomic variant file (VCF file) was converted into a PLINK file, and the plink –score command was used on the variant table to calculate a polygenic risk score for each individual in the family. Polygenic risk scores were calculated for each sibling and both parents. Polygenic risk scores were also calculated for each individual in the 1000Genomes cohort (approximately 2500 individuals) and for a subset of Caucasian individuals (approximately 200-300 individuals). The polygenic risk score of each family member was compared to that of a group of population-matched (European) individuals to determine whether the individual was at high or low risk.

[0314] A polygenic risk score for celiac disease has been developed in a Caucasian population that incorporates multiple SNPs (Abraham et al., 2014; PMC PMC3923679). This model has high sensitivity for celiac disease, allowing calculation of its negative predictive value at a specific PRS threshold. Assuming a family history of celiac disease, we estimated the negative predictive value to be 99.4% at a specific PRS (less than -1). After calculating the PRS for each individual, two individuals had PRSs below this threshold. In the IVF setting, we estimate that selecting these two embryos for implantation would reduce disease risk by approximately 10-fold.

[0315] A polygenic risk score for Alzheimer's disease has been previously established and found to be associated with earlier onset of Alzheimer's disease (Desikan et al., 2017; PMC5360219; Table 2). The parental PRS is shown as a dark blue dashed line. The PRS for each embryo is shown as a gray dashed line. After calculating the PRS for each individual, individuals with the lowest polygenic risk score were predicted to have a reduced risk of Alzheimer's disease (median age of onset 87 years compared to 80 years) compared to embryos with the highest polygenic risk score.

[0316] Table 17: Single nucleotide polymorphisms used to construct a polygenic risk score for Alzheimer's disease

[0317]

[0318]

[0319] Example 19: Correlation Calculation

[0320] Use the embryonic genotype to calculate an individual's correlation index with the adverse genetic trait. For example, consider a maternal grandparent with schizophrenia. Step 1: After inferring the embryonic genomes from Examples 1 and 2, calculate the correlation between each embryo and the genome of the affected individual. Step 2: Select the embryo with the lowest correlation with the affected individual.

[0321] Example 20: Predicting Disease Risk Using Computational Genetic Correlations via IBD (Identity by Descent)

[0322] An extension of Example 3 uses IBD instead of the genetic relatedness of the affected individual in disease prediction. Because different sibling embryos may have different IBDs than affected family members, this information can be used in addition to the PRS score to further refine the probability of the embryo's disease risk. The following example assumes that the risk of disease is evenly spread across the entire genome of the affected individual, so the risk is linear with the degree of IBD in the affected individual.

[0323] log(P / (1-P))=β_1*PRS+β_2*sex_male+β_3*family history_presence+β_4*IBD_affected individuals.

[0324] Example 21: Regions of shared genomic information

[0325] Identify regions of shared genetic information between two individuals and select embryos that do not contain homozygous regions that increase the likelihood of a Mendelian condition. In consanguineous couples or couples with a shared genetic background, offspring are likely to be homozygous for the disease-causing region. Because genes with known disease associations are heterogeneously distributed throughout the genome, disease can be minimized by avoiding homozygous regions within known disease-causing regions of the genome. Step 1: Identify regions of shared genetic information between the two parents. Step 2: Calculate the fraction of homozygous regions in each embryo. Step 3: Select embryos with the lowest homozygous regions overall or across known disease-causing regions.

Claims

1. A method for determining disease risk associated with an embryo, the method comprising: (a) performing whole genome sequencing on a biological sample obtained from the paternal subject to identify a genome related to the paternal subject; (b) performing whole genome sequencing on a biological sample obtained from the maternal subject to identify a genome associated with the maternal subject; (c) phasing the genome associated with the paternal subject to identify the paternal haplotype; (d) phasing the genome associated with the maternal subject to identify maternal haplotypes; (e) performing rare genotyping on the embryo to identify one or more genetic variants in the embryo; (f) constructing a genome of the embryo based on (i) one or more genetic variants in the embryo, (ii) the paternal haplotype, (iii) the maternal haplotype, (iv) the transmission probability of the paternal haplotype, and (v) the transmission probability of the maternal haplotype; (g) assigning a polygenic risk score to the embryo based on the constructed genome of the embryo; (h) determining embryo-associated disease risks based on polygenic risk scores; (i) determining the transmission of genetic variants causing single gene diseases and / or haplotypes from the paternal genome and / or maternal genome to the embryo; and (j) Determining the combined disease risk associated with the embryo based on polygenic disease risk and the transmission of genetic variants causing single-gene diseases and / or haplotypes from the paternal genome and / or maternal genome to the embryo.

2. A method for outputting a disease risk score associated with an embryo, the method comprising: (a) receiving a first dataset comprising paternal genomic data and maternal genomic data; (b) aligning sequence reads to a reference genome and determining genotypes across the genome using paternal genome data and maternal genome data; (c) receiving a second dataset comprising paternal and maternal sparse genomic data; (d) phasing the paternal genomic data and the maternal genomic data to identify the paternal haplotype and the maternal haplotype; (e) receiving a third dataset comprising sparse genomic data of the embryo, a paternal transmission probability, and a maternal transmission probability; (f) applying an embryo reconstruction algorithm to (i) the paternal haplotype and the maternal haplotype, (ii) the sparse genomic data of the embryo, and (iii) the transmission probability of each of the paternal haplotype and the maternal haplotype to determine a constructed genome of the embryo; (g) applying the polygenic model to the constructed genome of the embryo; (h) output of disease risks associated with the embryo; (i) determining the transmission of disease-causing genetic variants and / or haplotypes from the paternal genome and / or maternal genome to the embryo; and (j) The presence or absence of disease-causing variants and / or haplotypes in the output embryos.

3. The method of claim 2, further comprising outputting a combined disease risk associated with the embryo based on the transmission of polygenic disease risks and genetic variants causing single-gene diseases and / or haplotypes from the paternal genome and / or maternal genome to the embryo.

4. The method of any one of claims 1 to 3, wherein the method further comprises using the grandfather's genomic data and / or the grandmother's genomic data to determine the paternal haplotype and / or the maternal haplotype.

5. The method of any one of claims 1 to 4, wherein the method further uses population genotype data and / or population allele frequencies to determine the disease risk of the embryo.

6. The method of any one of claims 1 to 5, wherein the method further uses family history of the disease and / or other risk factors to predict disease risk.

7. The method of any one of claims 1 or 4-6, wherein whole genome sequencing is performed using a standard, PCR-free, ligated read (e.g., synthetic long read), or long read protocol.

8. The method of any one of claims 1 or 4-7, wherein rare genotyping is performed using microarray technology; next generation sequencing technology of embryo biopsy; or cell culture medium sequencing.

9. The method of any one of claims 1 to 8, wherein phasing is performed using a population-based and / or molecular-based method (e.g., ligation reading).

10. The method of any one of claims 1 to 9, wherein the polygenic risk score is determined by summing the effects across multiple loci in the disease model.