Polygenic risk scores for in vitro fertilization
By constructing an embryo genome using haplotypes and transmission probabilities, the method addresses the challenge of predicting genetic disease risk in IVF, enabling accurate embryo selection and donor choice to reduce inherited disease likelihood.
Patent Information
- Application Number
- JP2022519991
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-08-06
- Filing Date
- 2020-09-30
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2040-09-30
AI Technical Summary
Current IVF clinics lack the ability to accurately predict the risk of genetic diseases in individuals and their prospective children due to the complexity of genetic, environmental, and lifestyle risk factors, particularly for common diseases with a family history.
A method involving whole genome sequencing, haplotype phasing, sparse genotyping, and polygenic risk scoring is employed to determine disease risk in embryos by constructing an embryo genome using paternal and maternal haplotypes and transmission probabilities, and predicting the presence of disease-causing variants.
This approach enables precise prediction of genetic disease risk in embryos, allowing for the selection of embryos with lower disease risk and informed donor selection, thereby reducing the likelihood of inherited diseases in future children.
Smart Images

Figure 0007775188000052 
Figure 0007775188000053 
Figure 0007775188000054
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 62 / 908,374, filed September 30, 2019, and U.S. Provisional Application No. 63 / 062,044, filed August 6, 2020, each of which is incorporated by reference herein in its entirety. Technical Field
[0002] Methods for determining risk of disease are described. [Background technology]
[0003] Currently, IVF clinics test for aneuploidies and monogenic disorders that are known to run in families. However, one in two couples has a family history of common diseases that are affected by a combination of genetic, environmental, and lifestyle risk factors. In addition, sperm donor clinics currently test for a propensity to develop a subset of diseases caused by monogenic disorders. There is a need in the art to improve the ability to predict the risk of genetic diseases in individuals and their prospective children. Summary of the Invention
[0004] A method for determining disease risk associated with an embryo is provided, the method comprising: performing whole genome sequencing on a biological sample obtained from the paternal subject to identify a genome associated with the paternal subject; performing whole genome sequencing on a biological sample obtained from the maternal subject to identify a genome associated with the maternal subject; phasing the genome associated with the paternal subject to identify a paternal haplotype; phasing the genome associated with the maternal subject to identify a maternal haplotype; and performing sparse genotyping on the embryo to identify one or more genetic variants in the embryo. performing genotyping; constructing a genome for the embryo based on (i) one or more genetic variants in the embryo, (ii) a paternal haplotype, (iii) a maternal haplotype, (iv) a transmission probability of the paternal haplotype, and (v) a transmission probability of the maternal haplotype; assigning a polygenic risk score to the embryo based on the constructed genome of the embryo; determining a disease risk associated with the embryo based on the polygenic risk score; determining the transmission of monogenic disease-causing genetic variants and / or haplotypes from the paternal genome and / or maternal genome to the embryo; and determining a combined disease risk associated with the embryo based on the polygenic disease risk and the transmission of monogenic disease-causing genetic variants and / or haplotypes from the paternal genome and / or maternal genome to the embryo.
[0005] Also provided is a method for outputting a disease risk score associated with an embryo, the method comprising: receiving a first dataset including paternal genomic data and maternal genomic data; aligning sequence reads to a reference genome and genotyping the genome using the paternal genomic data and the maternal genomic data; receiving a second dataset including paternal sparse genomic data and maternal sparse genomic data; phasing the paternal genomic data and maternal genomic data to identify paternal and maternal haplotypes; and outputting a third dataset including the sparse genomic data of the embryo, paternal transmission probabilities, and maternal transmission probabilities. The method includes receiving a dataset; applying an embryo reconstruction algorithm to (i) paternal and maternal haplotypes, (ii) the embryo's sparse genome data, and (iii) the respective transmission probabilities of the paternal and maternal haplotypes to determine a constructed genome of the embryo; applying a polygenic model to the constructed genome of the embryo; outputting a disease risk associated with the embryo; determining the transmission of disease-causing genetic variants and / or haplotypes from the paternal and / or maternal genome to the embryo; and outputting the presence or absence of disease-causing variants and / or haplotypes in the embryo. Some methods further include outputting a composite disease risk associated with the embryo based on the polygenic disease risk and the transmission of monogenic disease-causing genetic variants and / or haplotypes from the paternal and / or maternal genome to the embryo.
[0006] In some embodiments, the method further comprises determining paternal and / or maternal haplotypes using paternal and / or grandmother genomic data. In some embodiments, the method further comprises determining embryonic disease risk using population genotype data and / or population allele frequencies. In some embodiments, the method further comprises predicting disease risk using family history of disease and / or other risk factors.
[0007] In some embodiments, whole genome sequencing is performed using standard, PCR-free, linked-read (i.e., synthetic long-read) or long-read protocols. In some embodiments, sparse genotyping is performed using microarray technology, next-generation sequencing technology of embryo biopsy, or sequencing of cell culture medium. In some embodiments, phasing is performed using population-based and / or molecular-based methods (e.g., linked-read). In some embodiments, polygenic risk scores are determined by summing the effects across sites in disease models.
[0008] In some embodiments, the population genotype data comprises the allele frequencies and individual genotypes of at least about 300,000 unrelated individuals in UK Biobank. In some embodiments, the population phenotype data comprises both self-reported and clinically reported (e.g., ICD-10 code) phenotypes for at least about 300,000 unrelated individuals in UK Biobank. In some embodiments, the population genotype data comprises the self-reported data of at least about 300,000 unrelated individuals in UK Biobank, and population family history data, including information obtained from the relatives of those individuals in UK Biobank. In some embodiments, disease risk is further determined by the proportion of genetic information shared by affected individuals.
[0009] Also provided is a method for determining disease risk for one or more future children, the method including: (i) performing whole genome sequencing on the expected mother and one or more future sperm donors, or (ii) the expected father and one or more future egg donors; phasing the genomes of (i) the expected mother and one or more future sperm donors, or (ii) the expected father and one or more future egg donors; simulating gametes based on an estimate of recombination rate; combining the simulated gametes to generate genomes for one or more future children; assigning a polygenic risk score; and determining a distribution of disease probabilities based on the polygenic risk score.
[0010] Also provided is a method for outputting a probability distribution of disease risks for future children, the method including: receiving a first dataset including genomic data from a prospective mother; receiving one or more datasets including genomic data from one or more prospective sperm donors; simulating gametes using estimated recombination rates (e.g., obtained from the HapMap Consortium); generating genomes for one or more future children using future combinations of gametes; estimating a polygenic risk score for the genome of each of the one or more future children; and outputting a distribution of disease probabilities based on the polygenic risk scores.
[0011] Also provided is a method for determining a range of disease risks for future children of (i) a prospective mother and future sperm donor, or (ii) a prospective father and future egg donor, the method comprising: (a) performing whole genome sequencing on the prospective mother and one or more prospective sperm donor(s) to obtain (i) a maternal genotype and one or more sperm donor(s) genotypes, or (ii) a paternal genotype and one or more egg donor(s) genotypes; (b) using (i) the maternal genotype and the future sperm donor(s), or (ii) the prospective paternal genotype and the future egg donor(s) genotype(s), to estimate the likely genotypes of one or more future children; (c) using the likely genotypes of the future children to estimate the lowest possible polygenic risk score of the future children; and (d) using the likely genotypes of the future children to estimate the highest possible polygenic risk score of the future children.
[0012] Also provided is a method for outputting a disease risk range for a future child of (i) a prospective mother and a prospective sperm donor, or (ii) a prospective father and a prospective egg donor, the method comprising: (a) receiving a first dataset including genomic data of the prospective mother or genomic data of the prospective father; (b) receiving one or more datasets including genomic data from one or more prospective sperm donors or one or more prospective egg donors; and (c) outputting a disease risk range for (i) the prospective mother and the prospective sperm donor(s), or (ii) the prospective father and the prospective egg donor(s). (d) deriving possible genotypes for future children using the genotypes of the genotypes (possibly several) of the genomic DNA; (d) estimating a minimum polygenic risk score for the future children by selecting a genotype (of those derived in (c)) at each site in a model that minimizes the score; (e) estimating a maximum polygenic risk score for the future children by selecting a genotype (of those derived in (c)) at each site in a model that maximizes the score; and (f) outputting a range of disease risk using the minimum and maximum scores calculated in (d) and (e).
[0013] In some embodiments, the method uses high-density genotyping arrays on the sperm donor(s), followed by genotyping imputation at sites of interest that have not been directly genotyped. In some embodiments, the method uses family history of disease and other relevant risk factors to determine disease risk.
[0014] In some embodiments, whole genome sequencing is performed using standard, PCR-free, linked-read (i.e., synthetic long-read) or long-read protocols. In some embodiments, phasing is performed using population-based and / or molecular-based methods (e.g., linked-read). In some embodiments, polygenic risk scores are determined by summing the effects across all sites in a disease model.
[0015] In some embodiments, the population genotype data includes allele frequencies and individual genotypes for at least about 300,000 unrelated individuals in the UK Biobank. In some embodiments, the population phenotype data includes both self-reported and clinically reported (e.g., ICD-10 codes) phenotypes for at least about 300,000 unrelated individuals in the UK Biobank. In some embodiments, the population family history includes self-reported data for at least about 300,000 unrelated individuals in the UK Biobank and information obtained from relatives of those individuals in the UK Biobank. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 illustrates an exemplary methodology for predicting and reducing the risk of disease. [Figure 2] FIG. 1 shows a flowchart providing an exemplary methodology for determining a polygenic risk score. [Figure 3] FIG. 1 illustrates an exemplary methodology for determining disease risk in children. [Figure 4] FIG. 1 illustrates exemplary inputs that can be used to determine the probability of disease. [Figure 5] FIG. 1 shows a flowchart illustrating an exemplary methodology for selecting embryos based on disease likelihood. [Figure 6] FIG. 1 is a graphical representation of risk reduction curves associated with specific diseases. [Figure 7] FIG. 1 shows a flowchart providing an exemplary methodology for selecting a sperm donor. [Figure 8] FIG. 1 is a graphical representation of risk reduction curves generated for multiple donors for several autoimmune disorders. [Figure 9] FIG. 1 illustrates an example of disease risk distribution associated with various sperm donors. [Figure 10] FIG. 1 is a graphical representation of an ROC curve showing the improvement in predictive ability associated with determining prostate cancer risk. [Figure 11]FIG. 1 shows an exemplary method for predicting embryo-associated disease risk. [Figure 12] FIG. 1 shows an exemplary disease risk transmission prediction chart associated with HLA typing for rheumatoid arthritis. [Figure 13] FIG. 1 provides an exemplary scaffold for identifying chromosome-length phased blocks to improve disease risk prediction capabilities. [Figure 14] Graphical representation of the distribution of PRS (scaled to a mean of 0 and a standard deviation of 1) for rheumatoid arthritis cases and controls. [Figure 15] FIG. 1 shows ORs per decile for rheumatoid arthritis. [Figure 16] 16A shows the lifetime risk of various conditions in several embryos. Figure 16A shows the risk for the first embryo (called "Embryo 2"), Figure 16B shows the risk for the second embryo (called "Embryo 3"), and Figure 16C shows the risk for the third embryo (called "Embryo 4"). [Figure 17A] FIG. 1 shows lifetime risks and risk ratios for several embryos compared with general population risk. [Figure 17B] FIG. 1 shows embryonic lifetime risk as a function of polygenic risk score. [Figure 18] FIG. 1 provides an illustration of an exemplary parental support method for determining embryonic disease risk. [Figure 19] FIG. 1 illustrates the future workflow for embryonic whole genome prediction. [Figure 20] FIG. 1 illustrates how an individual's entire chromosomal phase can be obtained by performing whole genome sequencing of the individual, their partner, and two or more children, and determining which loci each child has inherited. [Figure 21] FIG. 1 is a block diagram of an exemplary computing device. DETAILED DESCRIPTION OF THE INVENTION
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Materials referred to in the following description and examples are available from commercial sources unless otherwise noted.
[0018] As used herein, the singular forms "a," "an," and "the" refer to both the singular and the plural unless expressly stated to specify only the singular.
[0019] The term "about" means that the understood number is not limited to the exact number stated herein, but is intended to refer to a number substantially around the cited number without departing from the scope of the present invention. As used herein, "about" will be understood by those of ordinary skill in the art and will vary to some extent depending on the context in which it is used. If there are uses of the term that are not clear to those of ordinary skill in the art given the context in which it is used, "about" will mean up to ±10% of the particular term.
[0020] The term "gene" refers to a sequence of DNA or RNA that encodes a polypeptide or plays a functional role in an organism. A gene can be a wild-type gene, or a variant or mutation of a wild-type gene. A "gene of interest" refers to a gene or genetic variant that may or may not be known to be associated with a particular phenotype or risk of a particular phenotype.
[0021] "Expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. Expression of a gene encompasses not only cellular gene expression, but also the transcription and translation of a nucleic acid(s) in cloning systems and any other context. When a nucleic acid sequence encodes a peptide, polypeptide, or protein, gene expression relates to the production of the nucleic acid (e.g., DNA or RNA, such as mRNA) and / or peptide, polypeptide, or protein. Thus, "expression level" can refer to the amount of nucleic acid (e.g., mRNA) or protein in a sample.
[0022] A "haplotype" refers to a group of genes or alleles that are inherited or expected to be inherited together from a single ancestor (father, mother, grandfather, grandmother, etc.). The term "ancestor" refers to the person from whom a subject is descended, or, in the case of an embryo, the embryo from which a future subject is descended. In a preferred embodiment, an ancestor refers to a mammalian subject, such as a human subject.
[0023] Diseases and Methods Provided is a method for identifying diseases caused entirely or in part by genetics, or the risk of having or inheriting a disease.Heritage disorders can be caused by mutations in one gene (monogenic disorders), mutations in multiple genes (polygenic disorders), a combination of gene mutations and environmental factors (multifactorial disorders), or chromosomal abnormalities (changes in the overall number or structure of chromosomes, the structure that carries genes).In some embodiments, the disease is a polygenic disorder, a multifactorial condition, or a rare monogenic disorder (for example, a disorder that has not been previously identified in a family).
[0024] Some embodiments involve determining whether an embryo carries a genetic disorder. Some embodiments involve determining whether an embryo will develop into a subject that has or is likely to have a genetic disorder. Some embodiments involve determining whether an embryo will develop into a subject that has or is likely to have one or more phenotypes associated with a genetic disorder.
[0025] Some embodiments include selecting embryos based on their genetic makeup. For example, some embodiments include selecting embryos that have a low risk of carrying a genetic disorder. Some embodiments include selecting embryos that have a low risk of having a genetic disease if they develop into a child or adult. Some embodiments include implanting the selected embryos into the uterus of a subject. Such methods are described in more detail, for example, in Balaban et al., "Laboratory Procedures for Human In Vitro Fertilization," Semin. Reprod. Med., 32(4):272-82 (2014), which is incorporated herein by reference in its entirety.
[0026] Some embodiments involve assessing disease risk associated with embryos formed using one or more sperm donors. Some embodiments involve selecting sperm donors based on disease risk. Some embodiments involve fertilizing eggs in vitro with the selected sperm.
[0027] Some embodiments include determining an individual's health report based on, for example, the presence or absence of polygenic or rare single genetic variants. Some embodiments include determining a distribution of disease probabilities based on, for example, a polygenic risk score.
[0028] The diseases that can be screened for are not limited. In some embodiments, the disease is an autoimmune condition. In some embodiments, the disease is associated with a specific HLA type. In some embodiments, the disease is cancer. Exemplary conditions include coronary artery disease, atrial fibrillation, type II diabetes, breast cancer, age-related macular degeneration, psoriasis, colon cancer, deep vein thrombosis, Parkinson's disease, glaucoma, rheumatoid arthritis, celiac disease, vitiligo, ulcerative colitis, Crohn's disease, lupus, chronic lymphocytic leukemia, type I diabetes, schizophrenia, multiple sclerosis, familial hypercholesterolemia, hyperthyroidism, hypothyroidism, melanoma, cervical cancer, depression, and migraine. Some exemplary diseases include monogenic disorders (e.g., sickle cell disease, cystic fibrosis), chromosome copy number disorders (e.g., Turner syndrome, Down syndrome), repeat expansion disorders (e.g., Fragile X syndrome), or more complex polygenic disorders (e.g., Type I diabetes, schizophrenia, Parkinson's disease, etc.). Other exemplary diseases are described in the Physicians' Desk Reference (PRD Network 71st ed. 2016); and The Merck Manual of Diagnosis and Therapy (Merck 20th ed. 2018), each of which is incorporated herein by reference in its entirety. Diseases whose genetic traits are complex by definition have multiple genetic loci that contribute to disease risk. In these situations, a polygenic risk score can be calculated and used to stratify embryos into high-risk and low-risk categories.
[0029] Embryonic genome construction A novel and original method for constructing an embryo genome is provided. In some embodiments, the construction enables whole-genome prediction in embryos using chromosome-length parental haplotypes and sparse genotyping of parents and embryos (e.g., using SNP arrays or low-coverage DNA sequencing). Such hybrid approaches can combine genetic information from parents and possibly other relatives (e.g., grandparents and siblings) with haplotypes obtained directly from DNA (e.g., high-density haplotype blocks) using molecular methods (e.g., Long Fragment Read technology, 10X Chromium technology, Minion system). Chromosome-length haplotypes can be used to predict the genome of embryos in the context of in vitro fertilization. Such predicted genome sequences can be used to predict disease risk both by directly measuring the transmission of variants that cause Mendelian genetic diseases and by constructing polygenic risk scores to predict disease risk.
[0030] In some embodiments, the embryo genome is constructed using haplotypes from two or more ancestors. In some embodiments, the embryo genome is constructed using both paternal haplotypes and maternal haplotypes. In some embodiments, the haplotype is a paternal haplotype. In some embodiments, the haplotype is a grandmother haplotype. In some embodiments, the embryo genome is constructed using paternal haplotypes, maternal haplotypes, and one or both of grandfather and grandmother haplotypes. In some embodiments, the sparse embryo genotype is obtained by sequencing embryo culture medium, cell-free DNA in blastocoelic fluid, or DNA obtained from a trophectoderm cell biopsy of the embryo.
[0031] Some embodiments include determining one or more haplotypes used to construct an embryo genome. Such haplotypes can be determined, for example, based on the genome sequence of an ancestral subject. Some embodiments include identifying a genome associated with the ancestral subject. Some embodiments include performing whole genome sequencing on a biological sample obtained from the ancestral subject to identify the ancestral subject's genome. Some embodiments include using one or more sibling embryos to determine the haplotypes. Such whole genome sequencing can be performed using any of a variety of techniques, such as standard, PCR-free, linked-read (e.g., synthetic long-read), or long-read protocols.Exemplary sequencing techniques are described, for example, in Huang et al., "Recent Advances in Experimental Whole Genome Haplotyping Methods," Int'l. J. Mol. Sci., 18(1944):1-15(2017):1-15(2017); Goodwin et al., "Coming of age: ten years of next-generation sequencing technologies," Nat. Rev. Genet., 17:333-351 (2016); Wang et al., "Efficient and unique cobarcoding of second-generation sequencing reads from long DNA molecules enabling cost-effective and accurate sequencing, haplotyping, and de novo assembly," Genome Res., 29(5):798-808 (2019); and Chen et al., "Ultralow-input single-tube linked-read library method enables short-read second-generation sequencing systems to routinely generate highly accurate and economical long-range sequencing information," Genome Res. Res., 30(6):898-909 (2020), each of which is incorporated herein by reference in its entirety.
[0032] Genome phasing Some embodiments involve phasing or inferring ancestral genomes to identify one or more haplotypes. Such phasing can be performed, for example, using population-based and / or molecular-based methods (such as linked-read methods). Exemplary phasing techniques are disclosed, for example, in Choi et al., "Comparison of phasing strategies for whole human genomes," PLoS Genetics, 14(4):e1007308 (2018); Wang et al., "Efficient and unique cobarcoding of second-generation sequencing reads from long DNA molecules enabling cost-effective and accurate sequencing, haplotyping, and de novo assembly," Genome Res., 29(5):798-808 (2019); and Chen et al., "Ultralow-input single-tube linked-read library method enables short-read second-generation sequencing systems to routinely generate highly accurate and economical long-range sequencing information," Genome Res., 30(6):898-909 (2020), each of which is incorporated herein by reference in its entirety.
[0033] In some embodiments, phasing uses linked-read sequencing, long fragment reads, fosmid-pool-based phasing, contiguity-preserving transposon sequencing, whole-genome sequencing, Hi-C methodology, dilution-based sequencing, targeted sequencing (such as HLA typing), or data generated from microarrays.
[0034] Some embodiments include using independently obtained sparse phased genotypes to provide a base for phasing. Ancestor genotypes can be phased using computer software such as HapCUT, SHAPEIT, MaCH, BEAGLE, or EAGLE. In some cases, the computer program uses a reference panel such as the 1000 Genomes or Haplotype Reference Consortium to phase genotypes. In some cases, adding genotype data from relatives such as grandparents, siblings, or children can improve phasing accuracy.
[0035] Prediction of embryo genome sequences Some embodiments include using the phased parental genomes in combination with sparse phased genotyping of the embryo to predict the genome of the embryo, allowing for the determination of the presence or absence of clinically relevant variants identified in the parents and embryo. This can be expanded to include risk / susceptibility alleles identified in the parents and HLA types. In some embodiments, sparse genotyping is obtained using next-generation sequencing. Sparse genotyping is described in detail in Kumar et al., "Whole genome prediction for preimplantation genetic diagnosis," Genome Med., 7(1):Article 35, pp. 1-8 (2015); Srebniak et al., "Genomic SNP array as a gold standard for prenatal diagnosis of foetal ultrasound abnormalities," Molceular Cytogenet., 5:Article 14, pages 1-4 (2012); and Bejjani et al., "Clinical Utility of Contemporary Molecular Cytogenetics," Annu. Rev. Genomics Hum. Genet., 9:71-86 (2008), each of which is incorporated herein by reference in its entirety.
[0036] Sparse genotyping can be performed on extracted portions of an embryo. Thus, some embodiments involve extracting or obtaining one or more cells from an embryo (e.g., via biopsy). Some embodiments involve extracting or obtaining nucleic acid (e.g., DNA) from an embryo or from one or more cells of an embryo. Some embodiments involve extracting embryo material from embryo culture medium.
[0037] Some embodiments use sparse embryo genotypes as a scaffold for phasing ancestral subject genomes. Some embodiments use information from one or more grandparent subjects (e.g., grandparent and / or grandmother subjects) to phase parent genomes. Some embodiments use information from large reference panels (e.g., population-based data) to phase parent genomes.
[0038] In some embodiments, the embryo is reconstructed using biological sample(s) obtained from one or more ancestral subject(s). Exemplary biological samples include one or more tissues selected from brain, heart, lung, kidney, liver, muscle, bone, stomach, intestine, esophagus, and skin tissue, and / or one or more bodily fluids selected from urine, blood, plasma, serum, saliva, semen, sputum, cerebrospinal fluid, mucus, sweat, vitreous humor, and milk. Some embodiments include obtaining the biological sample from the subject.
[0039] Some embodiments include determining the probability of transmission of one or more ancestral haplotypes. In some embodiments, the transmission of variants from one or more maternal heterozygous sites can include sequencing the maternal genome, sequencing or genotyping one or more biopsies from embryos, assembling or phasing maternal DNA samples into haplotype blocks, utilizing information from multiple embryos (e.g., parental support techniques) to construct parental chromosome-length haplotypes, and predicting the inheritance or transmission of these haplotype blocks using statistical methods such as HMMs. In some embodiments, HMMs can also predict transitions between haplotype blocks or correct errors in maternal phasing.
[0040] Approaches to predict the transmission of variants from one or more paternal heterozygous sites can include sequencing the paternal genome; sequencing or genotyping one or more biopsies from embryos; assembling or phasing the paternal DNA sample into haplotype blocks; utilizing information from multiple embryos to improve the contiguity of the haplotype blocks to chromosome lengths; and predicting the inheritance or transmission of these haplotype blocks using statistical methods such as HMMs. In some embodiments, HMMs can also predict transitions between haplotype blocks or correct errors in maternal phasing.
[0041] Situations where both the mother and father are heterozygous can be predicted using the methods described above. The genotype of the embryo is easily predicted when both parents are homozygous for either the same allele or different alleles.
[0042] In some embodiments, transmission probability is determined using the methods described in U.S. Patent Application Nos. 11 / 603,406; 12 / 076,348; or 13 / 110,685; or PCT Application Nos. PCT / US09 / 52730 or PCT / US10 / 050824 (each of which is incorporated by reference in its entirety). In some embodiments, regions with a transmission probability of 95% or greater are used to construct the embryonic genome.
[0043] In some embodiments, embryo genome is constructed using one or more genes or genetic variants in embryo.In some embodiments, one or more genes or genetic variants are identified using sparse genotyping in embryo.In some embodiments, sparse genotyping is carried out using microarray technology.
[0044] In some embodiments, the embryo genome is constructed using (i) one or more genetic variants in the embryo, (ii) one or more ancestral haplotype(s) (e.g., paternal haplotypes and maternal haplotypes), and (iii) transmission probabilities of one or more haplotypes (e.g., paternal haplotypes and maternal haplotypes). In some embodiments, sparse genotyping is performed using next-generation sequencing.
[0045] Some embodiments include embryo genome prediction using 1) the whole genome sequences of both grandparents on each side of the family, 2) the phased whole genome sequences from each parent, 3) the sparse genotypes measured by the parental arrays, and 4) the sparse genotype of the embryo. Without being bound by theory, it is believed that a prediction accuracy of 99.8% for 96.9% of the embryo genome can be achieved using such methods for well-studied CEPH families.
[0046] Some embodiments include phasing the parent genomes using 1) WGS of one grandparent, 2) sparse parent genotypes measured by array, and 3) a haplotype-resolved reference panel. Some embodiments include phasing the parent genomes using 1) sparse parent genotypes measured by array, and 2) a haplotype-resolved reference panel (e.g., 1000 Genomes). Some embodiments include phasing the parent genomes using only a haplotype-resolved reference panel (such as 1000 Genomes).
[0047] Determining Risk Methods for determining disease risk associated with an embryo are also provided (e.g., based on a genome constructed for the embryo). Some embodiments involve determining whether a disease-causing genetic variant from an ancestral genome has been transmitted to the embryo. Some embodiments involve determining whether a haplotype (e.g., associated with a disease-causing genetic variant) has been transmitted to the embryo. Some embodiments involve determining the presence or absence of disease-causing or increased disease susceptibility genetic variants, including (but not limited to) single nucleotide polymorphisms (SNVs), small insertions / deletions, and copy number variations (CNVs). Some embodiments involve determining the presence or absence of disease-associated HLA types in the embryo.
[0048] In some embodiments, the phenotypic risk in the embryo can be determined using one or more diseases (e.g., a range of diseases) that can be ranked based on age of onset and disease severity. In some embodiments, disease ranking can be combined with polygenic risk prediction to rank the embryo by future disease risk.
[0049] Some embodiments include determining that the embryo has a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or higher disease risk. Some embodiments include determining that the embryo has a 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, 1% or lower disease risk. Some embodiments include selecting embryos based on disease risk (e.g., selecting embryos with a relatively low disease risk) and / or based on the presence or absence of particular genetic variants (e.g., SNVs, haplotypes, insertions / deletions, and / or CNVs).
[0050] In some embodiments, the disease risk associated with the embryo is determined using a polygenic risk score. In some embodiments, the polygenic risk score (also referred to as "PRS") is determined by summing the effects of all loci in a disease model. In some embodiments, the polygenic risk score is determined using population data. For example, the population data can include allele frequencies, individual genotypes, self-reported phenotypes, clinically reported phenotypes (e.g., ICD-10 codes), and / or family history information (e.g., obtained from relatives in one or more population databases). Such population data can be obtained from any of a variety of databases, such as the United Kingdom (UK) Biobank (which contains information on approximately 300,000 unrelated individuals), the National Center for Biotechnology Information (NCBI), the European Genome-phenome Archive; OMIM; GWASdb; PheGenl; Genetic Association Database (GAD); and various genotype-phenotype datasets that are part of the Genotype and Phenotype Database (dbGaP) maintained by PhenomicDB.
[0051] In some embodiments, disease risk is determined based on a cutoff value of the polygenic risk score.For example, such cutoffs can include a maximum of about 1% PRS distribution, a maximum of about 2% PRS distribution, a maximum of about 3% PRS distribution, a maximum of about 4% PRS distribution, or a maximum of 4% PRS distribution.Preferably, the cutoff is based on a maximum of 3% PRS distribution.The cutoff of the polygenic risk score can also be determined based on, for example, an absolute risk increase of about 5%, about 10%, or about 15%.Preferably, the cutoff of the polygenic risk score is determined based on an absolute risk increase of 10%.
[0052] Some embodiments involve using a predicted embryonic genome to estimate phenotypic risk. In some embodiments, risk estimation uses 1) the predicted genome of the embryo, 2) parental genotypes at sites of interest where prediction is not made in the embryo (i.e., variants included in the polygenic risk score), and 3) allele frequencies in a reference cohort (e.g., UKBB) at sites of interest where prediction is not made in the embryo (e.g., variants included in the polygenic risk score).
[0053] Some embodiments involve determining risk based on the probability of transmission of one or more genetic variants (e.g., based on ancestral haplotypes). Some embodiments involve determining a combined risk associated with an embryo based on the risk of a polygenic disease and the probability of transmission of one or more genetic variants (e.g., transmission of a monogenic disease-causing genetic variant(s) and / or haplotype from the paternal and / or maternal genome to the embryo).
[0054] A non-limiting exemplary system for predicting and reducing the risk of disease is shown in Figure 1. A non-limiting exemplary polygenic risk score workflow is shown in Figure 2.
[0055] Provider Selection Methods for selecting sperm and / or egg donors are also provided. An estimate of a subject's risk of passing a disease to their offspring can be computed by simulating the genomes of hypothetical children and calculating the disease risk for each child. Some embodiments include determining the disease risk of the expected mother and one or more future sperm donors. Some embodiments include determining the disease risk of the expected father and one or more future egg donors.
[0056] Some embodiments involve simulating gametes from future mothers and fathers using the phased parent genomes and simulated haplotype recombination sites, for example, as determined using the HapMap database. Some embodiments take into account the respective recombination rates during meiosis in the production of these gametes. In some embodiments, these simulated gametes are combined with each other to generate a large number of combination possibilities for estimating the range of future child genomes. This array of child genomes can be transferred to an array of disease probabilities to predict the distribution of disease risk for each child. See Figure 3.
[0057] The risk estimates described herein (e.g., in the Embryo Genome Construction section and / or the Examples section) can be used in the context of family planning in embryo selection and / or sperm donor selection during an IVF cycle. In some embodiments, prospective parents receive a report containing either individual risk estimates for multiple phenotypes in all available embryos or a range of risk values for each prospective sperm donor. In some aspects, sperm donors are ranked based on disease risk for a condition or set of conditions. In some aspects, donors are selected using the python script disclosed in U.S. Provisional Application No. 63 / 062,044, filed August 6, 2020, or a modification thereof.
[0058] Some embodiments include selecting an embryo based on the risk score, some embodiments include selecting an egg donor based on the risk score, some embodiments include selecting a sperm donor based on the risk score.
[0059] Mounting System The methods described herein can be implemented in a variety of systems. For example, in some embodiments, a system (e.g., for performing genomic embryo construction, donor selection, risk determination, and / or health reporting) comprises one or more processors coupled to memory. The methods can be implemented using code and data stored and embodied in one or more electronic devices. Such electronic devices can store and communicate (internally and / or with other electronic devices over a network) code and data using computer-readable media such as non-transitory computer-readable storage media (e.g., magnetic disks, optical disks, random access memory, read-only memory, flash memory devices, phase-change memory), and transient computer-readable transmission media (e.g., electrical, optical, acoustic, or other forms of propagated signals (carrier waves, infrared signals, digital signals, etc.)).
[0060] Computer instructions can be loaded into the memory as needed to train the model (e.g., to identify disease risk). In some embodiments, the system is implemented on a computer, such as a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a supercomputer, a massively parallel computing platform, a television, a mainframe, a server farm, or a widely distributed set of loosely networked computers, or any other data processing system or user device.
[0061] The method may be performed by processing logic comprising hardware (e.g., circuitry, dedicated logic, etc.), firmware, software (e.g., embodied on a non-transitory computer-readable medium), or a combination of both. The operations described may be performed in any order, or in parallel.
[0062] Generally, a processor can receive instructions and data from read-only memory or random-access memory, or both. A computer generally includes a processor capable of performing actions according to instructions and one or more memory devices for storing instructions and data. A computer also generally includes or is operably connected to one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, optical disks, or solid-state drives, for receiving or transferring data. However, a computer need not have such devices. Furthermore, a computer can be incorporated into another device, such as a smartphone, a mobile audio or media player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices, e.g., EPROMs, EEPROMs, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0063] One or more computer systems can be configured to perform particular operations or actions by installing software, firmware, hardware, or a combination thereof on the system and causing the system to perform the actions during operation. One or more computer programs can be configured to perform particular operations or actions by containing instructions that, when executed by a data processing device, cause the device to perform the actions.
[0064] An exemplary implementation system is shown in Figure 21. Such a system can be used to perform one or more of the operations described herein. The computing device may be connected to other computing devices in a LAN, an intranet, an extranet, and / or the Internet. The computing device may operate in the capacity of a server machine in a client-server network environment or in the capacity of a client in a peer-to-peer network environment.
[0065] The following examples are provided to illustrate the present invention, but it should be understood that the invention is not limited to the particular conditions or details of these examples.
[0066] Example Example 1: Phasing parental genomes for parental recurrence risk assessment and disease prediction in embryos for preimplantation genetic testing - Use in predicting embryo genomic sequence in in vitro fertilization (IVF).
[0067] Embryo coverage and accuracy were calculated using three different protocols. According to the first protocol, embryo genome prediction used 1) whole-genome sequences from both grandparents on each side of the family, 2) phased WGS from each parent, 3) sparse genotypes measured by arrays on the parents, and 4) sparse genotypes on the embryos (Figure 4). This protocol achieved a prediction accuracy of 99.8% in 96.9% of the embryo genomes of well-studied CEPH families. (Similar protocols using 1) WGS from one grandparent, 2) sparse genotypes measured by arrays on the parents, and 3) haplotype-resolved reference panels are also contemplated.)
[0068] According to the second protocol, embryonic predictions used 1) parental sparse genotypes measured by arrays, and 2) a haplotype-resolved reference panel (e.g., 1000 Human Genomes).
[0069] According to the third protocol, embryonic predictions used only haplotype-resolved reference panels (e.g., 1000 Human Genomes).
[0070] The results for all three protocols are shown below in Table 1. The PRS shows results for approximately 1.4 million sites that are significant in predicting disease risk. [Table 1]
[0071] Example 2: Using predicted embryonic genomes to estimate phenotypic risk The probability of a possible genotype (AA, AB, BB) given the parental genotypes (M, D) is used at the unpredicted site in the embryo's genome (see Equation 1 below). If parental genotypes are not available, the cohort affected allele frequency (AF) is used. EA ) (Equation 2)
number
number
[0072] Twenty-seven of the 30 models (90%) predicted a percentile risk score for embryos that fell within 3% of the true score.
[0073] An alternative process involves using 1) the embryo's predicted genome and 2) allele frequencies within a reference cohort (such as UKBB) at sites of interest for which no predictions are made in the embryo (i.e., variants included in the polygenic risk score). Allele frequencies were used as shown in Equation 2 above. This process was used to predict risk score percentiles for embryos that fell within 23 of 30 (77%) models. When parental genotypes were incorporated, all 30 predicted scores fell within 5% of the true score.
[0074] Example 3: Estimating and improving phenotypic risk estimates using polygenic risk models Statistical Framework The mainstay model for disease simulation and empirical analysis is the threshold liability model. A disease is characterized by a genetic component g~N(0,h 2 ), where h 2 is the narrow-sense heritability and error component ∈~N(0,1-h 2 ) The assumed liability l is given by:
number
[0075] The family simulation includes three components: two genetic components—the portion measured by the PRS, an "unmeasured" portion that is simply the residual genetic risk, and a simulation of genetic liability modeled as the sum of irreducible non-genetic error. The potential genetic risk, g, above is
number
number
[0076] This last component is not correlated among family members. On the other hand, the variance explained by the PRS in the liability scale is σ 2 and g R , i and g R , j If are the PRS components of liability for two first-degree relatives, the covariance is given by:
number
[0077] g U , i and g U , j is the remaining unmeasured component of the liability of the two first-degree relatives, and h 2 If is the heritability of the trait, the covariance is given by:
number
number
[0078] For two first-degree relatives i and j with liability,
number
number
[0079] IVF Embryo Selection Simulation An IVF simulation was performed to answer the following question: Given a set of n embryos and a desired clinical phenotype, how much less likely is the embryo with the lowest polygenic risk score to develop the disease over its lifetime than a randomly selected embryo? In other words, by how much is the relative risk of selection reduced?
[0080] To answer this question, a two-step procedure was used to generate parameters for parents and subsequent children. This procedure, or modifications of it, are used in simulations to test the effectiveness of donor selection and IVF embryo selection.
[0081] The following input values were used in the embryo selection model: σ 2 , variance explained by the polygenic risk score of the Liability Scale; h 2 , additive heritability of the trait on the liability scale; p, lifetime prevalence of the trait.
[0082] The output from this simulation is the risk reduction for different numbers of embryos available, which allows prospective IVF couples to be targeted for which diseases they can meaningfully screen for.
[0083] procedure Step 1. For each parent, to represent the elevated risk from family history, we use the distribution N(0,σ 2 ), or PRSg with several other distributions, such as mean-shifted or truncated normal R Generates the remaining unmeasured genetic risk g u is the distribution N(0,h 2 -σ 2 ) or any of the others listed above. Step 2. l1,…,l n Simulate n children by computing: Average midparent PRS from two parents:
number
number
number
number
number
[0084] Step 3. To determine the risk reduction, simulate millions of families in the range n = 3, 4, …, 10. For each family, determine the liability of the embryo with the smallest PRS. min However, the threshold t=Φ -1 Check whether it exceeds (1-p), where Φ is the cumulative distribution function of the standard normal distribution.
[0085] Statistical Notes As an addendum, R p,i and R U,i To show that the covariances between siblings and between children and parents are accurate, note that:
number
number
number
[0086] A similar series of calculations shows that the parent-child covariance also satisfies the correct equation.
[0087] This procedure can be seen schematically in Figure 5. An example of a risk reduction curve using the inputs is shown in Figure 6. The variance explained by the polygenic risk score is shown in Table 2 below, where "h2_lee" is the variance. [Table 2]
[0088] Donor family simulation To identify low-risk donors, we performed the following: (1) calculate the expected maternal polygenic risk score, (2) calculate the polygenic risk scores for N donors, and (3) select the donor with the lowest polygenic risk score. The procedure is essentially the same as above, except for two changes: first, we simulate the number of donors (n = 10, 20, 30, ..., 100) and minimize the polygenic risk score over the donor's polygenic risk score rather than minimizing recombination. A flowchart of this method is shown in Figure 7.
[0089] The following input values were used: σ 2 , variance explained by PRS in the liability scale; h 2 , the additive heritability of the trait on the liability scale; p, the lifetime prevalence of the trait. The output from this simulation is the risk reduction for various numbers of donors that are available to be minimized, allowing clients to target which diseases they can meaningfully screen for using sperm or egg donors. Using the same example inputs as above, we generated risk reduction curves for various numbers of donors for several autoimmune disorders, which are shown in Figure 8.
[0090] Additional embryo selection after donor selection An additional application of donor selection involves first selecting a donor and then selecting embryos with a low disease risk. More specifically, disease risk information is provided to a subject (e.g., a female subject) interested in using donor sperm for a child. First, using the woman's genetic test results and family history, multiple gametes are simulated and combined with the simulated sperm sample to obtain the risk of known genetic causes of heart disease. This is the woman's "personal risk" of having a child with this condition, a subdivision of the "baseline risk." Second, using genetic information from various donors and information on which variants phase with each other, a range of disease probability is calculated for gametes from each individual donor. Finally, assuming a donor is selected, multiple embryos (E1, E2, E3) fall within the disease risk distribution. See Figure 9.
[0091] This method can be used in the selection of sperm donors in the context of family planning. Prospective parents can indicate phenotypes of particular concern to them, and risk scores for those phenotypes can be generated for each donor. These scores are used to predict the risk of disease in each of the sperm donor's future children. Providing parents with a report containing these risk values can allow them the option of selecting a donor who reduces their risk of the phenotypes of concern.
[0092] Family history Family history can be incorporated into predicting disease risk. UK Biobank has several disease conditions self-reported by parents and siblings, including diabetes, heart disease, Alzheimer's disease, Parkinson's disease, breast cancer, and various others. In addition, there are over 10,000 sibling pairs and numerous half-sibling or other second-degree relative pairs. Models were constructed using a binary variable for family history, which means: (i) the set of diseases in UK Biobank with self-reported family history, siblings or parents with that disease; or (ii) for any other disease, all samples of first-degree relatives in UK Biobank. For each condition in the appropriate cohort, given this definition of the "has_family_history" dummy, a logistic regression was performed using the following equation: log(P / (1-P))=beta_1*PRS+beta_2*sex_male+beta_3*has_family_history
[0093] In summary, inputs include: self-reported family history of disease, and data from a biobank containing pairs of first-degree relatives with medical records. Outputs include: a logistic regression model incorporating PRS and family history to improve the accuracy of our predictions. The model was used to prioritize which patients were at higher risk of developing the disease in their lifetime. An example output is shown in Table 3 below, where beta_1 (PRS), beta_2 (gender dummy), and beta_3 (family history dummy) are estimated for several conditions. [Table 3]
[0094] As shown in Figure 10, the improvement in prediction when the has_family_history dummy was added to the logistic regression was quantified in the ROC curve for prostate cancer.
[0095] Increasing model complexity The model becomes more complex by incorporating second- and third-degree relatives, more complex pedigrees, and / or associated phenotypes. Methods for simulating close relatives are shown above. Two additional family members can also be simulated for each parent to allow for the incorporation of second-degree family history. P1 is the relative R 1,i If you are a parent with, you can generate second-degree family members by assuming:
number
[0096] We can also add an additional layer of complexity to the simulation, namely thresholds based on age and sex. If the incidence of the disease varies with these variables, we can adjust the threshold at which samples in families are judged to have the disease. As an example, consider type II diabetes, where the prevalence in men over 80 years old is 20%, while the prevalence in women aged 55 is 4%. We can replace lifetime prevalence with lifetime risk by substituting the empirical lifetime risk of the disease in the above model. The thresholds for such samples are 1-Φ(0.20) and 1-Φ(0.04), respectively, where Φ is the cumulative distribution function of a standard normal random variable. When conditioning on pedigrees, we condition on the sample set.
number
[0097] Given a family tree Ped with information about the medical history of a father with the disease and his paternal grandfather, three siblings without the disease, etc., the following can be calculated by a computer;
number
number
[0098] HLA phenotype Risk determination may include phenotypes with a strong HLA component where the relevant HLA alleles are not fully tagged by SNVs. However, this method can be applied to any condition where there is a known disease association with a significant effect size of an HLA allele and additional genetic loci are involved. Examples of complex phenotypes involving HLA include (but are not limited to) psoriasis, multiple sclerosis, type 1 diabetes, inflammatory bowel disease, Crohn's disease, ulcerative colitis, vitiligo, celiac disease, and systemic lupus erythematosus.
[0099] This method can be applied in multiple contexts, including, but not limited to, individual disease risk prediction, risk reduction in both embryo selection and sperm donor selection scenarios, and guidance on prescribing certain medications where multiple genetic factors, such as HLA type, influence the likelihood of response or drug side effects.
[0100] HLA typing results are obtained from DNA-based methods such as Sanger sequencing-based typing or derived from whole genome sequencing (WGS). First, a polygenic risk score is determined, for example, using effect sizes from genome-wide association studies (GWAS). One example is to sum the product of the effect sizes of all relevant variants not in the MHC region and the dosage of the effect alleles. Next, relevant HLA alleles are combined or integrated based on the HLA typing results (not tag SNPs) using one of the following methods:
[0101] Combining ORs with PRS and HLA: Polygenic risk scores are calculated for all individuals in the validation cohort, and metadata (e.g., mean, standard deviation, etc.) is obtained. Odds ratios (ORs) are obtained for HLA alleles for which an association with the phenotype of interest has been established. ORs derived from an individual's PRS compared to the validation cohort and HLA typing are combined as follows:
number
[0102] Incorporating HLA directly into the PRS: HLA effect alleles are incorporated directly into the polygenic risk score by adding the product of the effect size and the dosage of each effect allele to the base PRS. This allows for a PRS HLA+ It is called PRS. HLA+ The RR is calculated for all individuals in the validation cohort and metadata (e.g., mean, standard deviation, etc.) is obtained. The RR is calculated using the OR derived from the PRS HLA+ model and the prevalence of the disease in the validation cohort. This is used to estimate the lifetime risk of the disease.
[0103] Example 4: Methods for ranking disease risk profiles with application to embryo and sperm donor selection An exemplary method for ranking disease risk profiles is provided, as shown in Figure 11. First, the weights w d is calculated for each disease in the set of diseases d, which is expressed as the age at onset w a and disease severity w s is the sum of the weights of a is greater for diseases that are present at birth, such as celiac disease, than for diseases that do not typically appear until adulthood, such as coronary artery disease. s is greater for more severe diseases such as breast cancer than for diseases with milder phenotypes such as vitiligo.
[0104] Family history and polygenic risk scores are then combined to generate a predicted risk for each condition of interest for each embryo.
[0105] Finally, we combine the disease ranking and risk prediction to generate a single score S for each embryo using the formula: T where RR is the relative risk derived from the combination of family history of a particular disease and the polygenic risk score.
number
[0106] Given three embryos with the following RRs for each of the above states, a total score is calculated for each embryo and ranked accordingly. For embryo 1, the score is calculated as follows:
number
[0107] The disease risk for each of the three embryos is shown in Table 5. [Table 5]
[0108] The same procedure applies to sperm donor selection, with each donor being ranked for all diseases of interest. In both the embryo and donor selection contexts, scores are calculated for a subset of diseases (e.g., conditions for which the prospective parents have a family history) or for all diseases for which a polygenic model is implemented.
[0109] Alternatively, this method can be used to prioritize outcomes for a single embryo / individual without summing all states of interest. Each state will receive a score, and the state with the highest score(s) will be prioritized. Using Embryo 1 above as an example, we generated the scores and rankings shown in Table 6. [Table 6]
[0110] Example 5: Prediction of transmission of disease susceptibility variants to embryos. One copy of the CRC susceptibility variant (APC c.3920T>A) (and / or insertion, deletion, and / or copy number variant) is found in the father's WGS. The allele is not present in the mother. This variant is not directly measured by sparse genotyping of the embryo. Parental whole-chromosome haplotypes are obtained from any one or combination of the methods described above. Embryonic genome reconstruction determines that a haplotype block containing the risk allele is transmitted from the father to one of the embryos. The risk allele is noted as "present" in the embryo.
[0111] Example 6: Polygenic risk of common diseases using embryonic prediction. Breast cancer has a common genetic component. The genetic risk score assesses breast cancer risk using 69 variants. Of these variants, only 13% (9 / 69) have been directly genotyped in the embryo. The percentile of the embryo's genetic risk score based on these variants is 84.6%. After embryo reconstruction, 98.6% (68 / 69) of the embryo's genotypes were estimated / inferred, and the new percentile of the embryo's genetic risk score is 77.7%. After the embryo was born, the child's DNA was genotyped, and the PRS percentile was 76.2%. This indicates that the genetic risk score from whole-genome embryo reconstruction has greater precision and less uncertainty due to information on additional variants.
[0112] Example 7: Prediction of transmission of disease-associated HLA types to the embryo. The mother has rheumatoid arthritis (RA). HLA typing results (from WGS, PCR + Sanger sequencing, or any other suitable method) reveal that the mother carries one copy of the HLA-DRB1*01:02 allele, which is associated with an increased risk of this condition. The father is homozygous for HLA-DRB1*04:02, an allele not known to be associated with an increased risk of RA. Based on complete phasing of chromosome 6 of each parent and reconstruction of the embryo genome, it is determined that the maternal haplotype 2 (HM2) and the paternal haplotype 2 (HF2) will be transmitted to the embryo. Because the RA risk allele is carried on the maternal haplotype 1 (HM1), the embryo is predicted not to carry the risk allele. See, for example, Figure 12.
[0113] Example 8: Providing families with a range of disease risks in their children. Two parents present to a physician their concerns about the risk of various genetic diseases in their prospective child. Using the above method, the mean value and recombination of the midparents are specifically calculated to predict the range of disease risk for the child given the genomes of the two parents, and to guide the prospective IVF treatment. See Figure 9.
[0114] Similarly, in the case of sperm donation, the distribution of WGS-based polygenic risk scores for the mother and future sperm donor(s) can be simulated by recombination (see Figure 9).
[0115] Example 9: Incorporating family history (FHx) to improve risk estimation The risk of developing psoriasis is estimated to be 10-30% based on family history of the disease. In embryos where one parent has psoriasis, using polygenic models alone shows only minor differences in risk between embryos. As shown in Table 7, incorporating family history significantly improves the separation of Embryo 1 from Embryos 2 and 3, demonstrating that Embryos 2 and 3 have additional risk factors beyond FHx. [Table 7]
[0116] Similarly, family history can be incorporated to improve risk estimates in predicting transmission of disease-associated HLA types.
[0117] Example 10: Incorporation of HLA typing into psoriasis disease risk estimates The presence or absence of two HLA types associated with the risk of developing psoriasis clearly influences the overall disease risk to the embryo. This example can be extended to the context of sperm donor selection or personal genome reporting, as shown in Table 8. [Table 8]
[0118] Family history can be incorporated to further refine risk estimates in predicting the inheritance of disease-associated HLA types. This technique can be extended to predict blood type from the embryonic genome, including the Rh status of the resulting fetus.
[0119] Example 11: Improving trait prediction accuracy If the genotypes of variants in a polygenic model are unknown in the embryo, parental genotypes can be used to improve the accuracy of trait prediction. Instead of the population allele frequency (AF) or estimated genotype, the probability of the possible genotype is used, given the parental genotypes at that site(s). The dosage of each possible genotype is added to the risk score using the probabilities in Table 9 below. This improves the prediction accuracy, as measured by the predicted percentile of the polygenic risk, as shown in Table 10 below, which shows an improvement in prediction for a polygenic model of Crohn's disease where four variants were not predicted in the embryo. The true polygenic risk score percentile ("True") is determined using direct genotyping from WGS. [Table 9] [Table 10]
[0120] Example 12: Haplotype disease risk Some disease risks are based on phased haplotypes rather than individual variants. To more accurately predict trait risk, embryo reconstruction generates phased haplotypes. Table 11 below shows haplotypes of the APOE gene and the associated risk of Alzheimer's disease (Corder et al., 1994). [Table 11]
[0121] The two variants are 138 bp apart within the APOE gene. Neither rs429358 nor rs7412 were measured in the embryonic sparse assay. This does not include estimating Alzheimer's disease risk in the embryo. However, the embryo reconstruction method uses parental genotypes to predict a fully phased embryonic genome that can be used to infer whether the embryo is ε3 / ε3. This result is later verified by whole-genome sequencing of the child. [Table 12] Thus, embryonic reconstruction allows for risk prediction of APOE haplotypes and Alzheimer's disease, and generally, haplotype-based disease states.
[0122] Example 13: Sparse genotype scaffolding Using sparse genotyping as a scaffold for genome-wide phasing (see, for example, Figure 13) improves performance over reference panels alone, as measured by switch error rate (SER). Applying this approach to the well-studied sample NA12878, we found that the overall SER decreased from 0.6% using the 1000 Genomes reference panel alone to 0.54% using a set of approximately 140k high-confidence phased genotypes as a scaffold combined with the reference panel. This difference is primarily due to a reduction in long switch errors. For example, on chromosome 1, the raw count data for long switch errors decreased by more than 60% (169 vs. 60). Overall, the combined approach (scaffold + reference panel) reduced the long switch error rate from 0.12% to 0.04%. Long switch errors are important in embryo reconstruction because they result in erroneous blocks that are predicted to be propagated.
[0123] Example 14: Polygenic risk scores Large-scale genome-wide association studies (GWAS) have identified genetic variants associated with a wide variety of diseases. These associations have paved the way for functional studies of disease biology, discovery of drug targets, and improved disease risk prediction. While individual common genetic variants may have little predictive value, combining these variants into genetic risk scores can explain a larger proportion of the genetic risk of disease. These multilocus genetic risk scores, also known as polygenic risk scores (PRS), are most commonly computed as a weighted sum of disease-associated genotypes.
number
[0124] We describe how to validate and implement polygenic models and visualize risk estimates in consumer reports.
[0125] Selection of polygenic risk model Priority was given to previously published polygenic models for each condition of interest that had been tested in at least 1,000 individuals from a broad population. This excluded small studies with limited statistical power and studies that tested in isolated populations that could not be translated to other populations. Models that used data from individuals in the UKBB study setting were also excluded. Models were selected that reported an area under the curve (AUC) greater than 0.65 and / or an odds ratio (OR) greater than 2 for individuals in the upper and lower quantiles (see below for details). A list of published model characteristics and their evaluation statistics is provided in Table 13. [Table 13] TIFF0007775188000040.tif69162
[0126] When published models were unavailable, SNPs meeting a genome-wide significant p-value threshold (p<5e-8) from the GWAS catalog were used to construct scores as previously described (PMID: 30309464).
[0127] UK Biobank definitions of each phenotype Each model was validated and standardized using data from the UK Biobank cohort. This resource contains both genetic and disease information for 500,000 individuals. Only unrelated individuals were used for the following analyses. As shown in Table 14, we used a combination of ICD-9 and ICD-10 codes, as well as self-reported disease and procedure codes to define each phenotype of interest. [Table 14] TIFF0007775188000042.tif210162TIFF0007775188000043.tif211162TIFF0007775188000044.tif122162
[0128] A subset of diseases is shown in Table 15 below. [Table 15]
[0129] Individuals were stratified by polygenic risk score (PGS) to investigate the incidence of disease in this population.
[0130] Model evaluation using the UKBB dataset. Polygenic risk scores were calculated as the weighted sum of disease-associated genotypes. Each individual's score on UKBB was calculated, and various metrics were used to evaluate model performance.
[0131] Distribution of PRS across cases and controls: The dataset was split into cases and controls for each trait, and distributions of scores were generated separately for cases and controls. Visual inspection of these distributions gave a general idea of how well each model could distinguish between cases and controls. As an example, Figure 14 shows the distribution of PRSs (scaled to a mean of 0 and a standard deviation of 1) for rheumatoid arthritis cases and controls.
[0132] Receiver operating curve (ROC): ROC and area under the curve (AUC) were calculated by plotting the sensitivity and specificity of the model at various risk thresholds.
[0133] PRS stratification into deciles: Individuals from the UK Biobank were stratified into groups with different disease risk profiles. Individuals at highest risk (those in the top decile of PRS) were compared with those at median risk (those with PRS in the middle 40th-60th percentile of the distribution). Disease prevalence for each disease in deciles was plotted, and the ratio of high-risk to median risk was calculated across diseases. Figure 15 shows the OR per decile for rheumatoid arthritis.
[0134] Regression analysis incorporating age and gender: After calculating the PRS for all unrelated individuals in the UK biobank dataset, logistic regression was applied to each model. PGS is the regression coefficient of the PRS, corresponding to the odds ratio when the PRS is standardized to a mean of 0 and a standard deviation of 1. Age and sex were incorporated where available and applicable.
number
[0135] Odds ratios were then used to determine thresholds for high risk and intermediate outcome for reporting purposes.
[0136] OR / SD (mean-centered vs. z-transformed) by disease According to the logistic model described above, the OR / SD of the PRS was obtained by standardizing the PRS variables (mean 0, SD 1) before computing the effect sizes. This process is useful for achieving two goals. First, the risk stratification ability of PRSs can be directly compared across diseases. PRSs for various diseases are on widely different scales due to the different number of SNPs and their respective effect sizes. Their corresponding effect sizes cannot be directly compared if they are not standardized. Standardizing all PRSs allows models to be directly ranked based on OR / SD, which reflects their ability to separate populations based on disease risk. Second, it enables the statistically accurate application of UKBB effect estimates to the US population. UKBB was used to estimate effect sizes and then converted to odds ratios. When relative risks were estimated from these odds ratios (see below), the US population disease prevalence rates were used to accurately capture the relative risk of individuals with a particular PRS in the US. Standardizing the UKBB PRS (using the UKBB mean and SD) allows the PRS of US individuals (after adjusting for the US PRS mean and SD) to be used in the model. Due to the random assortment of genetics, one would expect similar means and SDs of PRS across populations, at least for individuals of European ancestry. The results of the analysis are shown in Table 16. [Table 16]
[0137] PRS stratification for disease vs. age: After stratifying individuals into different risk groups, UKBB data is used to estimate the proportion of the population diagnosed with disease in these various groups.This information is visually plotted in various strata, such as high-risk group (the top 5% of individuals according to PRS) and average-risk group (the entire population).Assuming that the individual of interest has PRS at the 75th percentile, the predicted percentage of individuals diagnosed with the group of individuals with similar genetic risk of our specific individual of interest is shown.
[0138] This plot is useful in illustrating the utility of the PRS in stratifying individuals based on disease risk: seeing a clear separation of the proportions of the population diagnosed within the different PRS strata confirms the model's ability to separate individuals based on risk.
[0139] Computer calculation of individual adjusted lifetime risk: We can start with the average lifetime risk for people of our gender in the US. We then assess the risk markers in the genome and calculate a polygenic score based on those markers. We convert this information into an "odds ratio" using the UKBB data above. Finally, we factor this odds ratio and the average lifetime risk using the formula to estimate the lifetime risk for an individual with this variation:
number
[0140] where P0 is the prevalence of the condition in the UKBB, C0 is the average lifetime risk of the condition in the US, and OR is the odds ratio calculated above. The result is an estimate of an individual's own lifetime risk compared to the population average. For some conditions, the average lifetime risk is not available. In these cases, it is shown whether the analyzed genetics indicate an increased risk.
[0141] Defining the "high risk" threshold In some cases, a high genetic risk threshold was established based on known risk factors. For example, the relative risk of developing type 1 diabetes for an individual with an affected first-degree relative is 6.6. Therefore, a high-risk threshold for the type 1 diabetes PRS was established that corresponded to that relative risk. For phenotypes where this was unavailable or the model failed to achieve a threshold, individuals with a 2-fold relative risk or a 10% increase in absolute risk were designated as high-risk. The assessment metrics for the subset of phenotypes for which lifestyle or clinical factors demonstrated a high-risk threshold are shown in Table 17. [Table 17]
[0142] Example 15: Multifactorial Conditions (Polygenic Risk Score) Genomic DNA obtained from submitted samples was sequenced using either Illumina or BGI technology. Reads were aligned to a reference sequence (hg19) to identify sequence variations. For some genes, only specific variations were analyzed. Deletions and duplications were not investigated unless otherwise noted above. In some scenarios, independent validation of HLA typing may be performed by an external laboratory. Selected variants were annotated and interpreted according to ACMG (American College of Medical Genetics) guidelines. Only pathogenic or likely pathogenic variants were reported. Genotyping of the embryos and parents was performed, followed by a "parent support" analysis. The embryo genome was reconstructed using the embryo genotype, and the whole-genome sequences of the parents were reconstructed using a genome reconstruction algorithm. Only variants observed in the parent genome that were predicted to affect the embryo were examined in the reconstructed embryo genome. Polygenic risk scores were calculated for a subset of conditions. Models for each condition were evaluated in the UK Biobank population. Some polygenic risk scores may be refined using HLA typing. Individual lifetime risks were calculated by adjusting baseline risk (US population) according to demographic information and polygenic risk scores. Models in which the upper and lower deciles resulted in a 10% lifetime risk difference or a 1.9-fold increase in lifetime risk were included in the report. Based on available evidence of model and genome reconstruction performance, certain conditions (e.g., bipolar disorder) were retained in the experimental section according to the investigator's discretion. The lifetime risks of various conditions for specific embryos are shown in Figure 16A-C.
[0143] Using psoriasis as a specific example, Figures 17A-B show risk scores associated with a predisposition to psoriasis in three exemplary embryos.
[0144] Example 16: Whole genome prediction of embryos using haplotype-resolved genome sequencing Haplotype-resolved genome sequencing was combined with a sparse set of genotypes from single- or few-cell embryo biopsies to predict the whole genome sequence of the embryo. Specifically, stLFR technology was used for paternal haplotype-resolved genome sequencing. Performance was evaluated at rare heterozygous positions (defined as allele frequencies below 1%). Inheritance of 230,117 sites was predicted in the embryo with 89.5% accuracy.
[0145] The material used in this study was obtained retrospectively from participants who had previously undergone successful rounds of IVF with preimplantation genetic diagnosis (Table 16). Trophectoderm biopsies from a total of 10 embryos (day 5) were each genotyped for a panel of 300,000 common SNPs using a rapid 24-hour microarray protocol. In addition, each parent and all four grandparents were genotyped with the same panel. [Table 18]
[0146] Genomic DNA was extracted from whole blood or saliva samples. Neonatal and maternal DNA was processed using 30XWGS on the BGI platform. Paternal samples were processed using stLFR. Trophectoderm biopsies from 10 day-5 embryos were subjected to DNA extraction, amplification, and genotyping with both parents and grandparents using a rapid microarray protocol using Illumina CytoSNP-12 chips for all samples. Sibling embryo and parental SNP array measurements were combined using the "parental support" (PS) method (Figures 18 and 19), as detailed in Kumar et al. (2015). The whole genome sequences of embryos were predicted by combining the genotypes of PS embryos with parental haplotype blocks (see Figure 18).
[0147] Example 17: Construction of whole chromosome haplotypes from haplotype blocks and parental information To construct chromosome-length haplotypes in an IVF setting, we combined haplotype-resolved genome sequencing of both parents with information from sparse genotypes derived from sibling embryos. As part of the "parental support" (PS) method, we generated maximum likelihood estimates (MLE) of each parent's heterozygous SNVs by combining recombination frequencies from the HapMap database with SNP array measurements from both parents and sibling embryos. While this sparse chromosome-length haplotype was not sufficient for predicting the embryo's genome, it can be combined with molecularly derived high-density haplotypes from parental samples (e.g., using long-fragment read technology, 10x Genomics, CPT-seq, Pacific Biosciences, Hi-C) to predict inherited genome sequence.
[0148] Several data streams were used to obtain this information. To generate high-density haplotype blocks, initial shotgun sequencing was performed at a median coverage of 34x and 30x for the mother and father, respectively. Next, by sequencing a haploid subset of genomic DNA obtained by in vitro dilution pool amplification, 94.2% of the 1.94 million heterozygous SNVs in the mother and 92.4% of the 1.89 million heterozygous SNVs in the father were directly phased into long haplotype blocks. These molecularly derived "high-density haplotype blocks" were combined with sparse, chromosome-wide haplotypes to construct chromosome-wide, haplotype-resolved genome sequences for the parents. This sequence information was then used to predict the inherited genome sequence of the embryo and could also be used to predict the future offspring of the two parents (e.g., by simulating future eggs and sperm that will produce future children).
[0149] The future workflow for embryo whole genome prediction is shown in Figure 19. During the first visit, the patient's blood will be drawn and used to generate whole genome sequences for each parent and predict the couple's possible risk for disorders. After counseling, the parents undergo IVF, and the embryos are genotyped using traditional IVF PGD techniques. This information is combined with the parent's whole genome sequence information (haplotype resolution) to predict the embryo's inherited genome and assess disease risk.
[0150] The genotypes of the sibling embryos and parents are used to construct parental haplotypes of chromosome lengths. Statistical approaches (such as maximum likelihood estimation) are used to determine parental phases from the noisy information obtained from each sibling embryo and a database of meiotic recombination frequencies.
[0151] Construction of whole-chromosome haplotypes Whole-chromosome haplotypes are constructed by sequencing the genomes of an individual's relatives, including, but not limited to, parents, grandparents, or children. For individuals with two or more children, the phase of an individual's entire chromosomes can be obtained by performing whole-genome sequencing of the individual, their partner, and two or more children, and determining the loci inherited by each child (Figure 20). This provides whole-chromosome-based haplotype information without modifying the DNA sequencing process. This would be appropriate, for example, if a couple already has two children, is seeking another child, and does not have DNA samples from any grandparents.
[0152] Chromosomal haplotypes from individual sperm The method of Example 17 is carried out using whole chromosome haplotypes obtained by sequencing DNA obtained from individual sperm.
[0153] Example 18: Using embryonic genomic predictions to calculate polygenic risk scores for genetically complex diseases. Genome-wide association studies have enabled the construction of polygenic risk score models for conditions such as type 1 diabetes, schizophrenia, Crohn's disease, celiac disease, and Alzheimer's disease. These approaches involve obtaining a genome-wide list of significant SNPs, including the observed odds ratios for SNPs associated with disease, and calculating a "risk score" for each individual depending on the configuration of SNPs found in that individual. Using this approach, we calculated polygenic risk scores for siblings and simulated the polygenic risk scores seen when comparing sibling embryos in IVF cycles. We used publicly available genome sequences from pedigrees of 12 siblings, two parents, and four grandparents. Each genome variant file (VCF file) was converted to a PLINK file, and the plink-score command was used with the variant table to calculate the polygenic risk score for each individual in the family. Polygenic risk scores were calculated for each sibling and for the two parents. Polygenic risk scores were calculated for each individual in the 1000 Genomes Cohort (approximately 2500 individuals) and for a subset of Caucasian individuals (approximately 200-300 individuals). The polygenic risk score for each family member was compared with the polygenic risk scores of a population-matched group of (European) individuals to determine whether the individual was at high or low risk.
[0154] A polygenic risk score for celiac disease has been developed within a Caucasian population incorporating multiple SNPs (Abraham et al., 2014; PMC PMC3923679). This model is highly sensitive to celiac disease, and the negative predictive value of this approach can be calculated at a specific PRS threshold. Given a family history of celiac disease, we estimate a negative predictive value of 99.4% at a specific PRS (<-1). After calculating the PRS for each individual, two individuals had PRSs below this threshold. In the context of IVF, we estimate that these two embryos could be selected for implantation, reducing the risk of disease by approximately 10-fold.
[0155] Polygenic risk scores for Alzheimer's disease have previously been developed and found to be associated with earlier onset of Alzheimer's disease (Desikan et al., 2017; PMC5360219; Table 2). The parental PRS is shown as a dark blue dashed line. Each of the embryonic PRSs is shown as a gray dashed line. After calculating the PRS for each individual, individuals with the lowest polygenic risk scores are predicted to have a reduced risk of Alzheimer's disease (median age of onset of 87 years instead of 80 years) compared to embryos with the highest polygenic risk scores. [Table 19]
[0156] Example 19: Relevance calculations The genotype of the embryo is used to calculate a relatedness index to individuals with undesirable genetic traits. For example, consider a maternal grandparent with schizophrenia. Step 1: After inferring the genome of the embryo from Examples 1 and 2, calculate the relatedness of each embryo to the genome of the affected individual. Step 2: Select the embryos with the lowest relatedness to the affected individual.
[0157] Example 20: Using genetic relatedness calculated via identity by descent to predict disease risk This is an extension of Example 3, where ancestry identity (IBD) is used in disease prediction instead of genetic relatedness to affected individuals. Because different sibling embryos have different IBD than affected family members, this information can be used in addition to PRS score to further increase the probability of embryo disease risk. The following example assumes that disease risk is evenly distributed throughout the affected individual's genome. Therefore, risk is proportional to the degree of IBD in affected individuals. log(P / (1-P))=beta_1*PRS+beta_2*sex_male+beta_3*has_family_history+beta_4*IBD_affected_individual.
[0158] Example 21: Regions of shared genomic information Identify regions of shared genetic information between two individuals and select embryos that do not contain regions of homozygosity that may increase the likelihood of Mendelian patterns. In consanguineous couples or couples with a shared genetic background, offspring may be homozygous for disease-causing regions. Because genes with known disease associations are unevenly spread throughout the genome, disease can be minimized by avoiding regions of homozygosity within genomic regions that cause known diseases. Step 1: Determine the regions of shared genetic information between the two parents. Step 2: Calculate the percentage of homozygous regions for each embryo. Step 3: Select embryos with the lowest regions of homozygosity across all or the total number of regions known to cause disease.
Claims
1. 1. A method for determining an embryo-associated complex disease risk, comprising: (a) performing whole genome sequencing on a biological sample obtained from the paternal subject to identify a genome associated with said paternal subject; (b) performing whole genome sequencing on a biological sample obtained from the maternal subject to identify a genome associated with the maternal subject; (c) phasing the genome associated with the paternal subject to identify a paternal haplotype; (d) phasing the genome associated with the maternal subject to identify maternal haplotypes; (e) performing sparse genotyping on the embryo to identify one or more genetic variants within the embryo; (f) constructing the genome of the embryo based on each combination of (i) the one or more genetic variants in the embryo, (ii) the paternal haplotype based on a transmission probability of the paternal haplotype, and (iii) the maternal haplotype based on a transmission probability of the maternal haplotype, wherein the one or more genetic variants are variants predicted to have an effect on disease risk of the embryo; and (g) assigning a polygenic risk score to the embryo based on the constructed genome of the embryo, wherein the polygenic risk score is based on a weighted combination of two or more disease-associated genotypes; (h) determining the transmission of monogenic disease-causing genetic variants and / or haplotypes from the paternal genome and / or the maternal genome to the embryo; (i) determining a composite disease risk associated with the embryo based on the polygenic risk score and the transmission of monogenic disease-causing genetic variants and / or haplotypes from the paternal and / or maternal genome to the embryo.
2. 1. A method for outputting a disease risk associated with an embryo, comprising: (a) receiving a first dataset including paternal genomic data and maternal genomic data; (b) aligning sequence reads to a reference genome and determining the genotype of the genome using the paternal genome data and the maternal genome data; (c) receiving a second dataset comprising paternal sparse genomic data and maternal sparse genomic data; (d) phasing the paternal genomic data and the maternal genomic data to identify paternal and maternal haplotypes; (e) receiving a third dataset comprising sparse genomic data of paternal transmission probability and maternal transmission probability for the embryo; (f) applying an embryo reconstruction algorithm to each combination of (i) the paternal haplotype based on the paternal transmission probability and the maternal haplotype based on the maternal transmission probability, and (ii) the sparse genomic data of the embryo to determine a constructed genome of the embryo; (g) applying a polygenic model to the constructed genome of the embryo; (h) outputting the disease risk associated with the embryo, wherein the disease risk is based on a weighted combination of two or more disease-associated genotypes; (i) determining the transmission of disease-causing genetic variants and / or haplotypes from the paternal genome and / or the maternal genome to the embryo; (j) outputting the presence or absence of disease-causing variants and / or haplotypes in said embryo.
3. 3. The method of claim 2, further comprising outputting a composite disease risk associated with the embryo based on the disease risk and the transmission of monogenic disease-causing genetic variants and / or haplotypes from the paternal and / or maternal genome to the embryo.
4. The method of any one of claims 1 to 3, further comprising determining paternal and / or maternal haplotypes using paternal and / or maternal genomic data.
5. 5. The method of any one of claims 1 to 4, further using population genotype data and / or population allele frequencies to determine said disease risk for said embryo.
6. The method of any one of claims 1 to 5, wherein a family history of the disease and / or other risk factors are further used to predict disease risk.
7. 2. The method of claim 1, wherein the whole genome sequencing is performed using a standard, PCR-free, linked-read (e.g., synthetic long-read), or long-read protocol.
8. 10. The method of claim 1, wherein the sparse genotyping is performed using microarray technology, next generation sequencing technology of embryo biopsies, or sequencing of cell culture media.
9. The method of any one of claims 1 to 8, wherein the phasing is performed using population-based and / or molecular-based methods (e.g. linked reads).
10. 10. The method of claim 1, wherein the polygenic risk score is determined by summing effects across sites in a disease model.
11. 6. The method of claim 5, wherein the population genotype data comprises allele frequencies and individual genotypes for at least about 300,000 unrelated individuals in the UK Biobank.
12. 6. The method of claim 5, wherein the population phenotype data comprises both self-reported and clinically reported (e.g., ICD-10 codes) phenotypes for at least about 300,000 unrelated individuals in the UK Biobank.
13. 6. The method of claim 5, wherein the population genotype data comprises self-reported data for at least about 300,000 unrelated individuals in the UK Biobank and population family history data comprising information obtained from relatives of those individuals in the UK Biobank.
14. 14. The method of claim 13, wherein the disease risk is further determined by the proportion of genetic information shared by affected individuals.
Citation Information
Patent Citations
Processing and analysis of complex nucleic acid sequence data
JP2017184742A
Genetic analysis
US20090299645A1
Methods for non-invasive prenatal ploidy calling
US20140154682A1