Polygenic risk score for in vitro fertilization

By employing whole-genome sequencing and polygenic risk scoring of paternal and maternal haplotypes, the method improves the prediction of genetic disease risk in embryos and future children, enabling the selection of low-risk embryos and reducing the occurrence of genetic disorders.

JP2026021564APending Publication Date: 2026-02-10マイオームインコーポレイテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025191899
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-08-06
Filing Date
2025-11-12
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Current IVF clinics struggle to accurately predict the genetic disease risk in embryos and future children due to limitations in genetic testing, particularly for complex diseases influenced by multiple genes and environmental factors, leading to a significant portion of couples being diagnosed with genetic or environmental disorders post-birth.

Method used

A method involving whole-genome sequencing and phasing of paternal and maternal haplotypes, sparse genotyping of embryos, and constructing polygenic risk scores to predict disease risk by analyzing genetic variants and haplotypes, combined with population data and family history to determine the genetic basis of single-gene and complex disorders.

Benefits of technology

Enhances the accuracy of predicting genetic disease risk in embryos and future children, allowing for the selection of embryos with lower disease risk, thereby reducing the incidence of genetic disorders in offspring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021564000001_ABST
    Figure 2026021564000001_ABST
Patent Text Reader

Abstract

To provide a polygenic risk score for in vitro fertilization.SOLUTION: A method for determining a disease risk associated with an embryo, the method comprising: (i) constructing a genome of the embryo based on one or more genetic variants in the embryo, (ii) a paternal haplotype, (iii) a maternal haplotype, (iv) a probability of transmission of the paternal haplotype, and (v) a probability of transmission of the maternal haplotype; assigning a polygenic risk score to the embryo based on the constructed genome of the embryo; Provided is a method comprising: determining a disease risk associated with an embryo; and determining transmission of a disease-causing genetic variant and / or haplotype from a paternal genome and / or a maternal genome to the embryo. Also provided is a method of determining the range of risk of a disease in a plurality of future offspring of a mother and a future sperm donor.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Provisional Application No. 62 / 908,374, filed September 30, 2019. and the benefit of U.S. Provisional Application No. 63 / 062,044, filed August 6, 2020. No. 6,239,999, each of which is incorporated herein by reference in its entirety. Technical Field

[0002] Methods for determining risk of disease are described. [Background technology]

[0003] Currently, IVF clinics treat aneuploidies and single genes known to occur in families. Testing for genetic disorders is underway, but one in two couples are diagnosed with a genetic or environmental disorder. Families with common diseases affected by a combination of cognitive and lifestyle risk factors In addition, sperm donor clinics currently offer treatment for conditions caused by monogenic disorders. In the art, individuals are tested for their propensity to develop a subset of known diseases. There is a need to improve our ability to predict the risk of genetic diseases in individuals and their expected future children. There is a need. Summary of the Invention

[0004] A method for determining disease risk associated with an embryo is provided, the method comprising determining a disease risk associated with a paternal subject. To identify the genome involved, whole-genome analyses were performed on biological samples obtained from the paternal subject. and performing sequencing; and from the maternal subject to identify the genome associated with the maternal subject. Whole genome sequencing was performed on the biological samples obtained; paternal haplotyping To identify the paternal haploid gene, we performed phasing of the genome associated with the paternal haploid gene; phasing the genome relative to the maternal subject to identify the phenotype; and Sparse genotyping of embryos to identify one or more genetic variants in (i) performing sparse genotyping of one or more embryos; The above genetic variants, (ii) paternal haplotype, (iii) maternal haplotype, (i based on (v) the probability of transmission of the paternal haplotype, and (v) the probability of transmission of the maternal haplotype. Constructing the genome of the embryo; and a polygenic risk score based on the constructed genome of the embryo. and assigning a disease risk associated with the embryo based on a polygenic risk score. and determining the genetic basis of single-gene disorders from the paternal and / or maternal genome to the embryo. Determining the transmission of causative genetic variants and / or haplotypes; Genetic disease risk and single inheritance from the paternal and / or maternal genome to the embryo Based on the transmission of genetic variants and / or haplotypes that cause childhood disorders and determining the combined disease risk associated with the embryo.

[0005] Also provided is a method for outputting a disease risk score associated with an embryo, the method comprising: receiving a first dataset including reference genomic data and maternal genomic data; Align sequence reads to the matched genome and compare the paternal and maternal genome data. Genotyping the genome using paternal sparse genome data and maternal receiving a second dataset including sparse genomic data; and paternal and maternal genomic data to identify paternal and maternal haplotypes and phasing of the sparse genomic data of embryos, paternal transmission probability and maternal transmission probability. receiving a third dataset including a probability; and configuring an embryo reconstruction algorithm to (i) identify the paternal embryo; (ii) sparse genomic data of embryos, and (i ii) Applying the paternal and maternal haplotypes to the respective transmission probabilities of the embryo Determining the assembled genome; and applying the polygenic model to the assembled genome of the embryo. and to generate disease risks associated with the embryo; and to generate genetic information for the paternal and / or maternal genomes. Transmission of disease-causing genetic variants and / or haplotypes from genome to embryo and determining the disease-causing variants and / or haplotypes in the embryo. and outputting the presence or absence of the polygenic disease risk. Genetic inheritance from the paternal and / or maternal genome to the embryo, causing a monogenic disorder Embryo-associated composite disease risk based on variant and / or haplotype transmission and outputting the

[0006] In some embodiments, the method further comprises the step of: The data can be used to further determine paternal and / or maternal haplotypes. In some embodiments, the method comprises using population genotype data and / or population pairwise data. and determining the disease risk of the embryo using the allele frequencies. In the method, a family history of disease and / or other risk factors are used to predict disease risk. The method further includes measuring the

[0007] In some embodiments, whole genome sequencing is performed using standard, PCR-free, linked performed using a long-read protocol (i.e., synthetic long reads) or In some aspects, sparse genotyping is performed using techniques such as microarray technology, embryo biopsy, and the like. This is performed using conventional sequencing techniques, or by sequencing the cell culture medium. In this paper, phasing is performed using population-based and / or molecular-based methods (e.g., linked In some embodiments, the polygenic risk score is performed using a disease model. The effect is determined by summing the effects across regions in the model.

[0008] In some embodiments, the population genotype data is obtained from at least one of the populations in UK Biobank. Each contains allele frequencies and individual genotypes for approximately 300,000 unrelated individuals. In some embodiments, the phenotypic data of the population is obtained from at least one of the populations in UK Biobank. Self-reported and clinically reported data from approximately 300,000 unrelated individuals (e.g., In some embodiments, the genotype data for a population includes both phenotypes (e.g., ICD-10 codes) and genotypes. The study relied on the autopsy data of at least approximately 300,000 unrelated individuals in the UK Biobank. Reported data and information obtained from relatives of those individuals in UK Biobank In some embodiments, the disease risk is assessed by a population of affected individuals. Therefore, it is further determined by the proportion of shared genetic information.

[0009] Also provided is a method for determining the disease risk of one or more future children, the method comprising: , (i) the prospective mother and one or more prospective sperm donors, or (ii) the prospective father. and performing whole genome sequencing on one or more prospective egg donors; and (i) predicting (ii) the expected mother and one or more prospective sperm donors, or (iii) the expected father and one or more prospective sperm donors. Genome phasing of prospective egg donors and mating based on recombination rate estimates Simulating offspring; combining simulated gametes to produce one or more offspring Generate the genome of the future child; assign a polygenic risk score; and and determining a distribution of disease probabilities based on the child risk scores.

[0010] A method for outputting a probability distribution of future child disease risk is also provided, the method comprising: receiving a first dataset including genomic data from a mother to be analyzed; receiving one or more datasets containing genomic data from prospective sperm donors; and; using estimated recombination rates (e.g., obtained from the HapMap Consortium) and simulating gametes; and using future combinations of gametes to develop one or more generating the genomes of the one or more future children; and Estimating genetic risk scores and estimating disease probabilities based on polygenic risk scores and outputting the fabric.

[0011] and (i) the prospective mother and future sperm donor, or (ii) the prospective father. Also provided is a method for determining the range of disease risk for future children of a prospective egg donor. The method includes: (a) (i) comparing the maternal genotype and the genotype of one or more sperm donors; To obtain the type, the prospective mother and one or more prospective sperm donors, or or (ii) to obtain the paternal genotype and the genotypes of one or more egg donors. , whole genome sequencing for the expected father and one or more prospective egg donors. (b) (i) conducting maternal genotype and prospective sperm donor genotype(s); ), or (ii) the expected paternal genotype and the genotype(s) of the future egg donor. (c) using the gene to predict the likely genotype of one or more future children; The potential genotypes of future children are used to determine the lowest possible polygenicity of future children. (d) estimating a child risk score using the likely genotypes of future children. and estimating the highest possible polygenic risk score for future children.

[0012] and (i) the prospective mother and future sperm donor, or (ii) the prospective father. Also provided is a method for outputting a range of disease risks for future children of a prospective egg donor. The method includes (a) obtaining predicted maternal genomic data or predicted paternal genomic data; (b) receiving a first dataset including one or more prospective sperm donors; or one or more datasets containing genomic data from one or more prospective egg donors. (c) (i) the prospective mother and prospective sperm donor(s), and (ii) using the genotypes of the expected father and future egg donor(s) to identify the future (d) deriving the child's likely genotypes; and (e) minimizing the score in the model. At each site, the genotype (derived in (c)) is selected to determine the potential of the future offspring. (e) estimating the child's minimum polygenic risk score; and (f) fitting the score-maximizing model. In this case, by selecting the genotype (derived in (c)) at each site, (f) Estimate the highest polygenic risk score of the child calculated in (d) and (e). The minimum and maximum scores generated are used to generate a range of risk for the disease. , including.

[0013] In some embodiments, the method comprises high density genotyping of sperm donor(s). Use arrays and then perform genotyping on sites of interest that have not been directly genotyped. In some embodiments, the methods use family history of the disease and other relevant risk factors. and determine disease risk.

[0014] In some embodiments, whole genome sequencing is performed using standard, PCR-free, linked performed using a long-read protocol (i.e., synthetic long reads) or In some embodiments, phasing is performed using population-based and / or molecular-based methods. In some embodiments, the method is performed using a polygenic list (e.g., linked reads). The score is calculated by summing the effects across all sites in the disease model. It is decided.

[0015] In some embodiments, the population genotype data is obtained from at least one of the populations in UK Biobank. Each contains allele frequencies and individual genotypes for approximately 300,000 unrelated individuals. In some embodiments, the phenotypic data of the population is obtained from at least one of the populations in UK Biobank. Self-reported and clinically reported data from approximately 300,000 unrelated individuals (e.g., In some embodiments, the family history of the population includes both UK and ICD-10 codes (phenotypes). self-reported data from at least approximately 300,000 unrelated individuals in Biobank and This includes information obtained from relatives of those individuals in the UK Biobank. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 illustrates an exemplary methodology for predicting and reducing the risk of disease. [Figure 2] FIG. 1 shows a flowchart providing an exemplary methodology for determining a polygenic risk score. [Figure 3] FIG. 1 illustrates an exemplary methodology for determining disease risk in children. [Figure 4] FIG. 1 illustrates exemplary inputs that can be used to determine the probability of disease. [Figure 5] FIG. 1 shows a flowchart illustrating an exemplary methodology for selecting embryos based on disease likelihood. [Figure 6] FIG. 1 is a graphical representation of risk reduction curves associated with specific diseases. [Figure 7] FIG. 1 shows a flowchart providing an exemplary methodology for selecting a sperm donor. [Figure 8] FIG. 1 is a graphical representation of risk reduction curves generated for multiple donors for several autoimmune disorders. [Figure 9] FIG. 1 illustrates an example of disease risk distribution associated with various sperm donors. [Figure 10] FIG. 1 is a graphical representation of an ROC curve showing the improvement in predictive ability associated with determining prostate cancer risk. [Figure 11] FIG. 1 shows an exemplary method for predicting embryo-associated disease risk. [Figure 12] FIG. 1 shows an exemplary disease risk transmission prediction chart associated with HLA typing for rheumatoid arthritis. [Figure 13] FIG. 1 provides an exemplary scaffold for identifying chromosome-length phased blocks to improve disease risk prediction capabilities. [Figure 14] Graphical representation of the distribution of PRS (scaled to a mean of 0 and a standard deviation of 1) for rheumatoid arthritis cases and controls. [Figure 15] FIG. 1 shows ORs per decile for rheumatoid arthritis. [Figure 16] 16A shows the lifetime risk of various conditions in several embryos. Figure 16A shows the risk for the first embryo (called "Embryo 2"), Figure 16B shows the risk for the second embryo (called "Embryo 3"), and Figure 16C shows the risk for the third embryo (called "Embryo 4"). [Figure 17A] FIG. 1 shows lifetime risks and risk ratios for several embryos compared with general population risk. [Figure 17B] FIG. 1 shows embryonic lifetime risk as a function of polygenic risk score. [Figure 18] FIG. 1 provides an illustration of an exemplary parental support method for determining embryonic disease risk. [Figure 19] FIG. 1 illustrates the future workflow for embryonic whole genome prediction. [Figure 20] FIG. 1 illustrates how an individual's entire chromosomal phase can be obtained by performing whole genome sequencing of the individual, their partner, and two or more children, and determining which loci each child has inherited. [Figure 21] FIG. 1 is a block diagram of an exemplary computing device. DETAILED DESCRIPTION OF THE INVENTION

[0017] Unless otherwise defined, all technical and scientific terms used herein are It has the same meaning as commonly understood by a person skilled in the art to which this invention pertains. Materials referenced in the description and examples below are obtained from commercial sources unless otherwise noted. Available.

[0018] As used herein, the singular forms "a," "an," and "the" refer to the singular Unless expressly stated to specify only the singular and the plural, both are indicated.

[0019] The term "about" is understood to limit a number to the exact number stated herein. It is understood that the present invention is not limited to the above-mentioned embodiments, and that the present invention is not limited to the above-mentioned embodiments. As used herein, "about" is intended to refer to a number that will be understood by a person of ordinary skill in the art. It will vary to some extent depending on the context in which it is used. Sometimes, when there are uses of a term that are not clear to persons of ordinary skill in the art, "about" may be used to refer to the minimum extent of a particular term. This means ±10% at most.

[0020] The term "gene" refers to a gene that encodes a polypeptide or has a functional role in an organism. A gene is a sequence of DNA or RNA that plays a role in a gene's function. A "gene of interest" may be a gene variant or mutation that induces a particular phenotype, or genes that may or may not be known to be associated with risk of a particular phenotype Refers to a genetic variant.

[0021] "Expression" refers to the process by which a polynucleotide is converted from a DNA template (e.g., into mRNA or other RNA). The process by which a molecule is transcribed (into a transcript), and / or the transcribed mRNA is subsequently converted into a peptide, The process by which a gene is translated into a polypeptide or protein. Gene expression is a process by which a gene is translated into a polypeptide or protein. not only gene expression in cloning systems but also nucleic acids (complexes) in cloning systems and in any other context. The term "transcription and translation" also encompasses transcription and translation of a nucleic acid sequence (or sequences thereof). If the gene encodes a protein, gene expression is initiated by a nucleic acid (e.g., DNA such as mRNA or RNA). NA) and / or peptide, polypeptide, or protein production. Thus, the "expression level" refers to the level of a nucleic acid (e.g., mRNA) or protein in a sample. It can refer to the quantity of quality.

[0022] A "haplotype" is a set of genes inherited together from a single ancestor (father, mother, grandfather, grandmother, etc.). refers to a group of genes or alleles that are expected to occur together or be inherited together The term "ancestor" refers to the person from whom the subject descends, or, in the case of an embryo, to whom the future subject In a preferred embodiment, an ancestor refers to a mammalian subject, such as a human subject. vinegar.

[0023] Diseases and Methods may have a disease or condition caused in whole or in part by genetics Genetic disorders are caused by a single gene. Mutations (monogenic disorders), mutations in multiple genes (polygenic disorders), and gene mutations A combination of natural mutations and environmental factors (multifactorial disorders), or chromosomal abnormalities (the total number of chromosomes) or structural changes, gene-carrying structures). In , diseases are classified as polygenic disorders, multifactorial conditions, or rare monogenic disorders (e.g. , a disorder not previously identified in the family).

[0024] Some embodiments include determining whether the embryo carries a genetic disorder. This aspect of the invention allows for determining whether an embryo will develop into a subject that has or may have a genetic disorder. Some embodiments include determining whether the embryo exhibits one or more phenotypes associated with a genetic disorder. This includes determining whether a subject will develop into a subject that has or may have the condition.

[0025] Some embodiments involve selecting embryos based on the genetic makeup of the embryo. Some embodiments involve selecting embryos that are at low risk of carrying a genetic disorder. Some embodiments may produce embryos that have a low risk of having a genetic disorder if they develop into a child or adult. Some embodiments include implanting the selected embryo into the uterus of the subject. Such methods are described, for example, in Balaban et al., "Laboratory Proc. edures for Human In Vitro Fertilization” , Semin.Reprod.Med.,32(4):272-82(2014) No. 6,299,499, which is incorporated herein by reference in its entirety.

[0026] Some embodiments involve the use of one or more sperm donors to identify disease risks associated with embryos formed using the donor. Some embodiments involve selecting sperm donors based on disease risk. Some embodiments involve using the selected sperm to produce eggs in vitro. This includes fertilizing the plant.

[0027] Some embodiments may involve, for example, polygenic or rare single genetic variants based on their presence or absence. Some embodiments include determining a health profile for an individual based on, for example, a polygenic risk. and determining a distribution of disease probabilities based on the scores.

[0028] The diseases that can be screened for are not limited. In some embodiments, the disease is an autoimmune disease. In some embodiments, the disease is associated with a particular HLA type. In some embodiments, the disease is cancer. Exemplary conditions include coronary artery disease, atrial fibrillation, I Type 1 diabetes, breast cancer, age-related macular degeneration, psoriasis, colon cancer, deep vein thrombosis, Parkinson's disease disease, glaucoma, rheumatoid arthritis, celiac disease, vitiligo, ulcerative colitis, Crohn's disease, lupus, chronic Lymphoblastic leukemia, type I diabetes, schizophrenia, multiple sclerosis, familial hypercholesterolemia diabetes, hyperthyroidism, hypothyroidism, melanoma, cervical cancer, depression, and migraines Some exemplary diseases include monogenic disorders (e.g., sickle cell disease, cystic fibrosis), chromosome copy number disorders (e.g., Turner syndrome, Down syndrome), Peat elongation disorders (e.g., fragile X syndrome), or more complex polygenic disorders (e.g., Other exemplary diseases include Physiological disorders such as type 1 diabetes, schizophrenia, and Parkinson's disease. sicians'Desk Reference(PRD Network 71st ed. 2016); and The Merck Manual of Diagnosis is and Therapy (Merck 20th edition, 2018), Each of these is incorporated herein by reference in its entirety. Complex diseases have multiple genetic loci that contribute to disease risk. In these situations, Polygenic risk scores are calculated and used to assign embryos to high- and low-risk categories. It can be layered into layers.

[0029] Embryonic genome construction Novel and inventive methods relating to the construction of embryonic genomes are provided. In some embodiments, the construction using chromosome-length parental haplotypes and sparse genotyping of parents and embryos (e.g., using SNP arrays or low-coverage DNA sequencing) Such hybrid approaches allow for genomic prediction. ong Fragment Read technology,10X Chromiu m technology, Minion system) to Genetic information from other relatives (e.g., grandparents and siblings), if available, and directly from DNA The haplotypes obtained (such as high-density haplotype blocks) can be combined. Using chromatomorph length haplotypes to predict embryonic genomes in in vitro fertilization situations These predicted genome sequences can be used to identify genes that cause Mendelian diseases. Direct measurement of variant transmission and polygenic recombination for predicting disease risk Both the risk of disease can be predicted by constructing a risk score.

[0030] In some embodiments, the embryonic genome is constructed using haplotypes from two or more ancestors. In some embodiments, the embryo genome is a mixture of paternal and maternal haplotypes. In some embodiments, the haplotype is constructed using both the paternal haplotype and the paternal haplotype. In some embodiments, the haplotype is a maternal haplotype. In an embodiment, the embryo genome comprises a paternal haplotype, a maternal haplotype, and a paternal haplotype. The haplotypes are constructed using one or both of the paternal and maternal haplotypes. In this study, sparse embryo genotypes were determined by analyzing the cell-free DNA in the embryo culture medium, the blastocoelic fluid, or the embryo's nutrition. This is obtained by sequencing DNA obtained from an ectodermal cell biopsy.

[0031] Some embodiments determine one or more haplotypes used to construct the embryo genome. Such haplotypes may be determined, for example, based on the genomic sequences of ancestral subjects. Some embodiments can be determined by identifying genomes associated with the ancestry of a subject. Some embodiments include a method for identifying the genome of an ancestral subject by using a gene encoding a gene obtained from the ancestral subject. In some embodiments, whole genome sequencing is performed on the biological sample. In this case, one or more sibling embryos may be used to determine the haplotype. Genome sequencing is standard, PCR-free, linked-read (e.g., synthetic long read) This can be done using any of a variety of techniques, such as a genomic DNA sequencing (GDNA) or long read protocol. Exemplary sequencing techniques are described, for example, in Huang et al., "Recent Advances in n Experimental Whole Genome Haplotyping Methods” Int'l. J. Mol. Sci., 18 (1944): 1-15 ( 2017):1-15(2017); Goodwin et al., "Coming of ag e:ten years of next-generation sequence g technologies”, Nat.Rev.Genet.,17:333-35 1 (2016);Wang et al., "Efficient and unique co barcoding of second-generation sequence g reads from long DNA molecules enabling cost-effective and accurate sequencing, haplotyping, and de novo assembly”, Genom e Res.,29(5):798-808(2019); and Chen et al., "Ul tralow-input single-tube linked-read lib rary method enables short-read second-ge neration sequencing systems to routine generate highly accurate and economical "long-range sequencing information", Geno Me Res., 30(6):898-909 (2020), and these each of which is incorporated herein by reference in its entirety.

[0032] Genome phasing Some embodiments involve phasing an ancestral genome to identify one or more haplotypes. Such fading may involve, for example, population-based and and / or can be performed using molecular-based methods (such as linked-read methods). Exemplary fading techniques are described, for example, in Choi et al., "Comparison of p hasing strategies for whole human genome s”, PLoS Genetics, 14(4):e1007308 (2018) Wa ng et al., "Efficient and unique cobarcoding of second-generation sequencing reads from long DNA molecules enabling cost-effecti ve and accurate sequencing,haplotyping,a Genome Res., 29(5):79 8-808(2019); and Chen et al. "Ultralow-input sin gle-tube linked-read library method enab les short-read second-generation sequences ing systems to routinely generate highly accurate and economical long-range sequence encing information”, Genome Res., 30(6):89. 8-909(2020), each of which is incorporated by reference in its entirety. incorporated herein.

[0033] In some embodiments, phasing is performed using linked-read sequencing. ad sequencing, long fragment reads t reads), fosmid pool-based phasing (fosmid-pool- based phasing), adjacent conserved transposon sequencing (contiguit y preserving transposon sequencing), whole genome Sequencing, Hi-C methodology, dilution-based sequencing sequencing), targeted sequencing (e.g., HLA typing) or microarray Use the data generated from

[0034] In some embodiments, the phasing inhibitors may be independently administered to provide a scaffold for inducing phasing. These include using the resulting sparse-phased genotypes. Computer software such as PEIT, MaCH, BEAGLE, or EAGLE can be used to phase the ancestral genotypes. The computer program is based on the 1000 Genomes or Haplotype Reference Consortium. Genotype phasing is performed using a reference panel such as a can be further refined by adding genotype data for relatives such as grandparents, siblings, or children. The accuracy of the sampling can be improved.

[0035] Prediction of embryo genome sequences Some embodiments involve phasing in combination with sparse-phased genotyping of embryos. and using the parent genomes determined to be the same as the genome of the embryo, thereby determining the genome of the parent and the embryo. This allows for the determination of the presence or absence of clinically relevant variants identified in Risk / susceptibility alleles identified in the parents and HLA types can be included. In some aspects, sparse genotyping is obtained using next generation sequencing. Persistence genotyping is based on the findings of Kumar et al., "Whole genome predictive on for preimplantation genetic diagnosis ”, Genome Med., 7(1):Article 35, pp. 1-8 (201 5); Srebniak et al., “Genomic SNP array as a go ld standard for prenatal diagnosis of fo etal ultrasound abnormalities”, Molceular Cytogenet.,5:Article 14,pages 1-4(2012 and Bejjani et al., "Clinical Utility of Conte mporary Molecular Cytogenetics”, Annu.Rev. This is described in detail in Genomics Hum. Genet., 9:71-86 (2008). No. 6,239,999, each of which is incorporated herein by reference in its entirety.

[0036] Sparse genotyping can be performed on extracted portions of the embryo. Thus, some embodiments , involves extracting or obtaining one or more cells from an embryo (e.g., via biopsy). Some embodiments involve extracting or purifying nucleic acid (e.g., DNA) from an embryo or from one or more cells of an embryo. Some embodiments include extracting the embryo material from the embryo culture medium.

[0037] Some embodiments utilize the phasing of sparse embryos as a scaffold for phasing of ancestral target genomes. Some embodiments use genotypes from one or more grandparent subjects (e.g., grandparents and Phase the parental genomes using information from the grandparental subjects. Some aspects rely on information from large reference panels (e.g., population-based data). ) to phase the parent genome.

[0038] In some embodiments, the embryo is a biological sample obtained from one or more ancestral subject(s). Exemplary biological samples include brain, heart, lung, and , one or more tissues selected from kidney, liver, muscle, bone, stomach, intestine, esophagus, and skin tissue. , and / or urine, blood, plasma, serum, saliva, semen, sputum, cerebrospinal fluid, mucus, sweat, vitreous In some embodiments, the biological fluids include one or more selected from the group consisting of: This involves obtaining a biological sample from a subject.

[0039] Some embodiments involve determining the transmission probability of one or more ancestral haplotypes. In some embodiments, transmission of variants from one or more maternal heterozygous sites is considered to be a maternal genotype. sequencing of the embryo, sequencing or genotyping of one or more biopsies from the embryo, maternal DNA samples Assembling or phasing the pairs into haplotype blocks, and phasing the parental chromosome length haplotypes. Use of information from multiple embryos to construct a population (e.g., parent-supported techniques), and use statistical methods such as HMMs to identify the inherited sequences of these haplotype blocks. In some embodiments, the HMM may include prediction of the degree of similarity between haplotype blocks. It is also possible to predict transitions or correct errors in maternal fading .

[0040] An approach to predicting transmission of variants from one or more paternal heterozygous sites is paternal Sequencing the genome; and sequencing or genotyping one or more biopsies from the embryo. and assembling or phasing paternal DNA samples into haplotype blocks. and multiple sizing to improve the contiguity of haplotype blocks to chromosome lengths. and using statistical methods such as HMMs to analyze these haptics. and predicting the inheritance or propagation of the prototype block. In this paper, HMMs are used to predict transitions between haplotype blocks, or maternal phases. It is also possible to correct errors in the coding.

[0041] Situations in which both the mother and father are heterozygous can be predicted using the methods described above. The genotype of the embryo is determined by whether both parents have the same allele or different alleles. This is easily predicted when the patient is homozygous for the gene.

[0042] In some embodiments, the propagation probability is calculated using the method described in U.S. Patent Application Nos. 11 / 603,406; 12 / 023,232; / 076,348; or PCT application PCT / U S09 / 52730 or PCT / US10 / 050824 (each of which is (which is incorporated herein by reference in its entirety) In some embodiments, regions with a 95% or greater probability of transmission are selected to construct the embryonic genome. Used for:

[0043] In some embodiments, the embryo genome comprises one or more genes or genetic variants in the embryo. In some embodiments, the gene is constructed using one or more genes or genetic variants. The genotypes are identified using sparse genotyping in embryos. Genotyping is performed using microarray technology.

[0044] In some embodiments, the embryo genome is determined by analyzing (i) one or more genetic variants in the embryo, ii) one or more ancestral haplotype(s) (e.g., paternal and maternal haplotypes) (iii) one or more haplotypes (e.g., paternal haplotype and In some embodiments, the superset is constructed using the transmission probabilities of the genetic haplotypes (parental and maternal haplotypes). Genotyping is performed using next generation sequencing.

[0045] In some embodiments, the genome sequences of 1) both grandparents on each side of the family are sequenced; 3) the phased whole genome sequence from the parental array; genotypes, and 4) embryo genome prediction using sparse embryo genotypes. Although this is not a reliable prediction, the 99.8% accuracy rate for 96.9% of the embryo genome is sufficient. It is believed that this could be achieved using such a method for the CEPH family studied previously. It is being done.

[0046] In some embodiments, the genomic DNA fragments are obtained by 1) WGS of one grandparent, 2) SEQ ID NO: 1, and 3) SEQ ID NO: 2. and 3) parental genotypes using a haplotype-resolved reference panel. Some embodiments include: 1) performing phasing of the signals measured by the array; 1) sparse parent genotypes, and 2) a haplotype-resolved reference panel (e.g., 10 This involves phasing the parent genome using the 2000 human genome. For example, using only a haplotype-resolved reference panel (e.g., 1000 Genomes) , which involves phasing the parent genome.

[0047] Determining Risk Methods for determining disease risk associated with an embryo are also provided (e.g., methods for determining disease risk associated with an embryo, e ... Genome-based. Some aspects involve the identification of disease-causing genetic variants from ancestral genomes. In some embodiments, the method further comprises determining whether the ant is transmitted to the embryo. whether a type (e.g., associated with a disease-causing genetic variant) has been transmitted to the embryo Some embodiments include, but are not limited to, determining whether a single nucleotide polymorphism is present. Disease-causing mutations include small non-uniform variants (SNVs), small insertions / deletions, and copy number variations (CNVs). This involves determining the presence or absence of genetic variants that cause an increased risk of or susceptibility to a disease. Some embodiments involve determining the presence or absence of disease-associated HLA types in the embryo.

[0048] In some embodiments, the phenotypic risk in the embryo is determined based on age at onset and severity of the disease. Determined using one or more diseases (e.g., a range of diseases) that can be ranked according to In some embodiments, disease ranking can be combined with polygenic risk prediction. Together, embryos can be ranked by future disease risk.

[0049] In some embodiments, the embryos are 10%, 20%, 30%, 40%, 50%, 60%, 70%, Determining that you have an 80%, 90%, 95%, 99% or greater risk of disease In some embodiments, the embryos are 90%, 80%, 70%, 60%, 50%, 40%, 30%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 1 Determining that a person has a 0%, 20%, 10%, 5%, 1% or less risk of disease Some embodiments include determining whether a patient is at higher risk for a disease (e.g., a relatively low risk of a disease). Selecting embryos with specific genetic variants (e.g., SNVs, Select embryos based on the presence or absence of haplotypes, insertions / deletions, and / or CNVs Includes:

[0050] In some embodiments, the disease risk associated with the embryo is determined using a polygenic risk score. In some embodiments, the polygenic risk score (also referred to as "PRS") is determined by: The effect is determined by summing the effects across sites in the disease model. In

[0003] , polygenic risk scores are determined using population data. For example, are allele frequencies, individual genotypes, self-reported phenotypes, and clinically reported phenotypes (e.g., IC D-10 code), and / or family history (e.g., parental history in one or more population databases) Such population data may include information from individuals in the same family. bank (which has information on approximately 300,000 unrelated individuals), Nationa l Center for Biotechnology Information(N CBI), The European Genome-phenome Archive ;OMIM;GWASdb;PheGenl;Genetic Association Database (GAD); and PhenomicDB Various genotype-phenotypes that are part of the Database of Genotypes and Phenotypes (dbGaP) The data can be obtained from any of a variety of databases, such as a dataset.

[0051] In some embodiments, the disease risk is determined based on a cutoff value of the polygenic risk score. For example, such cutoffs may include a maximum of approximately 1% in the PRS distribution, up to about 2% in PRS distribution, up to about 3% in PRS distribution, up to about 4% in PRS distribution, or The maximum 4% may be included. Preferably, the cutoff is based on the maximum 3% of the PRS distribution. The genetic risk score cutoff can be, for example, about 5%, about 10%, or about 15% absolute. It may also be determined based on increased risk. Preferably, the risk is determined based on the count of a polygenic risk score. The cut-off is based on a 10% absolute risk increase.

[0052] Some embodiments use predicted embryonic genomes to estimate phenotypic risk. In some embodiments, risk estimation is based on 1) the predicted genome of the embryo; is a site of interest where no prediction is made (i.e., variants included in the polygenic risk score). ) parental genotypes at the site of interest where no prediction is made in the embryo (e.g., multi-gene Reference cohorts (e.g., UK) in which variants are included in the genetic risk score BB) are used.

[0053] Some embodiments may be based on the probability of transmission of one or more genetic variants (e.g., ancestry). Some embodiments involve determining risk based on polygenic Risk of disease and probability of transmission of one or more genetic variants (e.g., paternal genome and and / or the maternal genome to the embryo, a genetic variant causing a monogenic disease ( Determine the combined risk associated with the embryo based on the genetic information (multiple genes) and / or haplotypes (transmission of the genetic information) This includes determining

[0054] A non-limiting exemplary system for predicting and reducing the risk of disease is shown in FIG. A non-limiting exemplary polygenic risk score workflow is shown in FIG.

[0055] Provider Selection Methods for selecting sperm and / or egg donors are also provided. The inherited risk estimates are derived by simulating the genomes of hypothetical children and comparing the disease risk of each child. The risk of developing a disease can be calculated by computing the risk of developing a disease. This involves determining the disease risk of the prospective mother and one or more future sperm donors. Some embodiments may include determining the disease risk of the prospective father and one or more future egg donors. This includes determining the

[0056] Some embodiments include, for example, determining the presence of a phenotype of a gene in a gene encoding a phenotype, as determined using the HapMap database. Using age-aged parental genomes and simulated haplotype recombination sites , including simulating gametes from a future mother and father. , taking into account the respective recombination rates during meiosis in the production of these gametes. In some embodiments, these simulated gametes are combined with one another to form This provides a large number of combinatorial possibilities for estimating the genome coverage of future offspring. The genome arrays of such children are transferred to disease probability arrays to estimate the disease risk for each child. The fabric can be predicted, see Figure 3.

[0057] The risk estimates described herein (e.g., sections and / or experiments on embryo genome construction) The Examples section provides information on family planning in embryo selection and / or sperm donor selection during IVF cycles. In some embodiments, prospective parents can use all available Individual risk estimates for multiple phenotypes in all embryos or the risk of each prospective sperm donor In some embodiments, the sperm donor receives a report containing one of a range of values ​​for the sperm count. The ranking is based on the disease risk of a condition or set of conditions. In this case, the provider is referring to U.S. Provisional Application No. 63 / 062,044, filed August 6, 2020. The selected python script is disclosed in, or a modification of, it.

[0058] Some embodiments include selecting embryos based on the risk score. Some embodiments involve selecting an egg donor based on a risk score. This includes selecting sperm donors based on their sperm scores.

[0059] Mounting System The methods described herein can be implemented in a variety of systems. For example, in some aspects systems (e.g., genomic embryo construction, donor selection, risk determination, and / or health reporting) The system (for executing the advertisement) comprises one or more processors coupled to a memory. The law is implemented using code and data stored and carried out on one or more electronic devices. Such electronic devices may store non-transitory computer-readable storage media (e.g., magnetic Disks, optical disks, random access memory, read-only memory, flash memory device, phase change memory), and a temporary computer-readable transmission medium (e.g., , optical, acoustic, or other form of propagated signal (carrier wave, infrared signal, digital signal, etc.) ) to store code and data, and and / or other electronic devices via a network).

[0060] To train models as needed (e.g., to identify disease risk) In some embodiments, the system may load computer instructions into the memory. Computers, e.g., personal computers, portable computers, workstations tion, computer terminals, network computers, supercomputers, large-scale Parallel computing platforms, televisions, mainframes, server farms, etc. Any computer, a loosely networked set of widely distributed computers, or any or other data processing system or user device.

[0061] This method can be implemented using hardware (e.g., circuits, dedicated logic, etc.), firmware, or software. software (e.g., embodied on a non-transitory computer-readable medium), or both The operations described may be implemented by processing logic, including combinations. They can be performed in any order or in parallel.

[0062] Generally, a processor may have read-only or random access memory, or It can receive instructions and data from both. a processor capable of performing operations and one or more memories for storing instructions and data Generally, a computer also receives data or transmits data to a computer. For transfer, e.g., magnetic disk, magneto-optical disk, optical disk, or solid One or more mass storage devices, such as a state drive, for storing data. The computer may be connected to or operably linked to such devices. Furthermore, the computer can be connected to another device, just a few Examples include smartphones, mobile audio or media players. , game console, Global Positioning System (GPS) receiver, or portable storage embedded in a storage device (for example, a Universal Serial Bus (USB) flash drive) Suitable for storing computer program instructions and data. Some devices include, for example, semiconductor memory devices, such as EPROMs and EEPROMs. , and flash memory devices, magnetic disks, e.g., internal hard disks or Removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks This includes all forms of non-volatile memory, media and memory devices, such as flash memory. The processor and memory may be supplemented by or incorporated into special purpose logic circuits. can be done.

[0063] A system of one or more computers, including software, firmware, hardware, , or a combination of these, installed on the system to prevent system actions during operation. can be configured to perform a specific operation or action by having the instructions that, when executed by a data processing device, cause the device to perform an action By including it, you can instruct one or more computers to perform a particular operation or action. You can configure a data program.

[0064] An exemplary implementation system is shown in Figure 21. Such a system may be used for the operations described herein. The computing device may be used to implement one or more of the following: Other computers within an intranet, extranet, and / or the Internet The computing device may be connected to a client server. within the capacity of the server machine in a server network environment, or in a peer-to-peer network It can operate within the capacity of the client of the environment.

[0065] The following examples are provided to illustrate the invention, but the invention is not limited to these implementations. It should be understood that no limitation is intended to the particular conditions or details of the examples.

[0066] Example Example 1: Parental recurrence risk assessment and disease prediction in embryos for preimplantation genetic testing Phasing of parental genomes for embryo genomics in in vitro fertilization (IVF) Use in predicting system sequences.

[0067] Embryo coverage and accuracy were calculated using three different protocols. According to the protocol, embryonic genome prediction involves 1) the complete genomes of both grandparents on each side of the family; 1) the phased WGS from each parent; 2) the phased WGS from each parent; and 3) the phased WGS measured by the parental arrays. 1) sparse genotype, and 2) sparse genotype of embryos were used (Figure 4). predicted accuracy of 99.9% in 96.9% of the well-studied CEPH family embryo genomes. achieved 9.8% (also measured by 1) WGS of one grandparent and 2) array sparse genotypes of the parents, and 3) a probabilistic model using a haplotype-resolved reference panel. Protocols are also being considered.

[0068] According to the second protocol, embryonic predictions are based on 1) parental sparseness measured by arrays; genotype, and 2) haplotype-resolved reference panels (e.g., 1000 Genomes). ) was used.

[0069] According to the third protocol, embryonic predictions are performed using a haplotype-resolved reference panel (e.g., For example, only the 1000 human genomes were used.

[0070] The results for all three protocols are shown in Table 1 below. The PRS is important for predicting disease risk. Results for approximately 1.4 million key sites are shown. [Table 1]

[0071] Example 2: Using predicted embryonic genomes to estimate phenotypic risk Possible genotypes (AA, AB, BB) given parental genotypes (M, D) ) probability is used at unpredicted sites in the embryo's genome (see Equation 1 below). If parental genotypes are not available, cohort affected allele frequencies (AF EA ) (Formula 2)

number

number

[0072] The risk of embryos falling within 3% of the true score was observed in 27 of the 30 models (90%). Score percentiles were predicted.

[0073] In a separate process, 1) the predicted genome of the embryo and 2) the part of interest for which no prediction is made in the embryo are identified. Reference cohort at each stage (i.e., variants included in the polygenic risk score) This involves using allele frequencies within a gene (such as UKBB). Allele frequencies are calculated using the above The embryos were used as shown in Equation 2. Using this process, 23 of the 30 models (77 % predicted risk score percentiles falling within the model. Parental genotypes were incorporated. In this case, all 30 predicted scores are within 5% of the true score.

[0074] Example 3: Estimating and improving phenotypic risk estimates using polygenic risk models Statistical Framework The workhorse model for disease simulation and empirical analysis is the threshold liability model. The disease is caused by a genetic factor g~N(0,h 2 ), where h 2 is a legacy in the narrow sense Transmission and error factors ∈~N(0,1-h 2 The assumed liability l is given below: Therefore, it is required,

number

[0075] The family simulations consisted of three components: two genetic components - measured by PRS; the "unmeasured" portion which is simply residual genetic risk, and the irreducible non-genetic This involves the simulation of genetic susceptibility, which is modeled as the sum of genetic errors. The potential genetic risk g above is

number

number

[0076] This last factor is not correlated among family members. The variance explained by RS is σ 2 and g R , i and g R , j But two first-degree relatives If the PRS components of the family's liability are:

number

[0077] g U , i and g U , j is the remaining unmeasured component of liability in two first-degree relatives. and h 2 If is the heritability of the trait, the covariance is given by:

number

number

[0078] For two first-degree relatives i and j with liability,

number

number

[0079] IVF Embryo Selection Simulation IVF simulations were performed to answer the following questions: What is the set of n embryos and If the desired clinical phenotype is obtained, the minimal polygenic risk is higher than in randomly selected embryos. How unlikely is it that an embryo with a CRSC will develop the disease in its lifetime? If so, how much does this reduce the relative risk of the choice?

[0080] To answer this question, we used a two-step procedure to measure parent and subsequent child parameters. This procedure, or modifications thereof, may improve the effectiveness of donor selection and IVF embryo selection. Used in simulations to check.

[0081] The following input values ​​were used in the embryo selection model: σ 2 , Polygenic Risk Liability Scale Variance explained by score; h 2 , additive heritability of traits on the liability scale; p, Lifetime prevalence of traits.

[0082] The output from this simulation is the risk reduction for different numbers of embryos available. This allows prospective IVF couples to be meaningfully screened for any disease. You can target what you can do.

[0083] procedure Step 1. For each parent, a sample of individuals from the general population is selected to represent elevated risk from family history. If you do, the distribution N(0,σ 2 ), or some other methods such as mean shift or truncated normal. PRSg with other distributions R Generates the remaining unmeasured genetic risk g u minutes Cloth N(0,h 2 -σ 2 ) or any of the others listed above. Step 2. l1,…,l n By computer calculation, n children are simulated. Support: Average midparent PRS from two parents:

number

number

number

number

number

[0084] Step 3. To determine risk reduction, millions of families are analyzed in the range n=3, 4, …, 10. For each family, the liability of the embryo with the smallest PRS is calculated. min However, the threshold t=Φ -1Check whether it exceeds (1-p), where Φ is the standard positive This is the cumulative distribution function of the normal distribution.

[0085] Statistical Notes As an addendum, R p,i and R U,i The form of fraternity can be justified. To show that the covariance between children and parents is accurate, note the following:

number

number

number

[0086] A similar series of calculations shows that the parent-child covariance also satisfies the correct equation.

[0087] This procedure can be seen schematically in Figure 5. An example of a risk reduction curve using the inputs The variance explained by the polygenic risk score is shown in Figure 6. The variance explained by the polygenic risk score is shown in Table 2 below, where , "h2_lee" is the variance. [Table 2]

[0088] Donor family simulation To identify low-risk donors, we performed the following: (1) expected maternal multiplication; (2) Calculate the polygenic risk score for N donors. and (3) select the donor with the lowest polygenic risk score. This is essentially the same as above, except two steps have been changed: first, provide We simulated the number of participants (n = 10, 20, 30, ..., 100) and tried to minimize recombination. Instead, the polygenic risk score is minimized relative to the donor's polygenic risk score. A flow chart of this method is shown in FIG.

[0089] The following input values ​​were used: σ 2 , variance explained by PRS in the liability scale; h 2 , the additive heritability of the trait on the liability scale; p, the lifetime prevalence of the trait. The output from the analysis is available to minimize risk in different numbers of providers. This allows clients to use sperm or egg donors to The goal is to meaningfully screen for disease. Using the same example input as above, Risk reduction curves were generated for various numbers of donors for several autoimmune disorders. Shown in Figure 8.

[0090] Additional embryo selection after donor selection An additional application of donor selection is to first select the donor and then select embryos with low disease risk. More specifically, disease risk information may be used to support the use of donor sperm for the purpose of raising a child. First, genetic testing of this woman is provided to a subject of particular interest (e.g., a female subject). Using test results and family history, multiple gametes are simulated and the simulated sperm are analyzed. This, combined with a child sample, provides risk for known genetic causes of heart disease. This is the "individual risk" of women having children with the condition, and is a subdivision of the "baseline risk." Second, genetic information from various donors and the ability to compare any variants are important. Using information on how gametes from individual donors are phased, A range of disease probabilities is calculated. Finally, assuming a donor has been selected, multiple embryos (E E1, E2, E3) fall within the disease risk distribution. See Figure 9.

[0091] This method can be used in the selection of sperm donors in the context of family planning. can indicate phenotypes of particular interest to them and provide risk scores for those phenotypes. A score can be generated for each donor. These scores are These risks are used to predict the risk of disease in a person's future children. By providing parents with a report containing the value, parents can reduce their risk of the phenotype of interest. This may provide an option to select the person.

[0092] Family history Family history can be incorporated into disease risk prediction. UK Biobank has Diabetes, heart disease, Alzheimer's, Parkinson's, breast cancer, and many other There are several self-reported disease conditions by parents and siblings, such as: There are more than 100 pairs of siblings and many half-siblings or other pairs of second-degree relatives. We constructed it using a binary variable for history, which means: (i) self-reported a range of diseases from UK Biobank with a family history of the disease, a sibling or parent with the disease; or (ii) in the case of any other disease, all first-degree relatives in UK Biobank For each state in the appropriate cohort, Given this definition of the dummies, a logistic regression was conducted using the following equation: log(P / (1-P))=beta_1*PRS+beta_2*sex_male+beta_3*has_family_history

[0093] In summary, inputs include: self-reported family history of disease, and medical Data from a biobank containing first-degree relative pairs with records. Output includes: We have incorporated PRS and family history to improve the accuracy of our predictions. The model is used to estimate the risk of any patient developing a disease during their lifetime. An example output is shown in Table 3 below. Here, beta_ 1 (PRS), beta_2 (gender dummy), and beta_3 (family history dummy) It has been estimated in several ways. [Table 3]

[0094] As shown in Figure 10, the has_family_history dummy is a logistic regression The improvement in prediction when added to the regression was quantified with a ROC curve for prostate cancer.

[0095] Increasing model complexity Incorporating second- and third-degree relatives, more complex pedigrees, and / or associated phenotypes We make the model more complex by including the following: Two additional family members for each parent to allow for the inclusion of second-degree family history. It is also possible to simulate the 1,iIf you are a parent with We can generate second-degree family members by assuming:

number

[0096] We added an additional layer of complexity to the simulation: age- and sex-based thresholds. If the incidence of this disease varies with these variables, The threshold at which samples in families are judged to be of a similar quality can be adjusted. If we assume urinary disease, the prevalence rate for men over 80 years old is 20%, while that for women aged 55 years old is 10%. The prevalence of is 4%. By substituting the empirical lifetime risk of the disease in the above model, , lifetime prevalence can be replaced by lifetime risk. Such a sample threshold is are 1-Φ(0.20) and 1-Φ(0.04), respectively, where Φ is the standard normal The cumulative distribution function of a random variable. When conditioning on pedigrees, the sample set It is a condition on

number

[0097] Information about the patient's medical history, including the patient's father and paternal grandfather, and three unaffected siblings Given a pedigree Ped with information: ;

number

number

[0098] HLA phenotype Risk determination has a strong HLA component, with relevant HLA alleles determined by SNVs. However, this method does not provide a significant effect size. Any condition with a known disease association with an HLA allele and in which additional loci are involved Examples of complex phenotypes involving HLA involvement include psoriasis, multiple sclerosis, Type 1 diabetes, inflammatory bowel disease, Crohn's disease, ulcerative colitis, vitiligo, celiac disease, and all These include (but are not limited to) systemic lupus erythematosus.

[0099] This method may be used in applications including, but not limited to, individual disease risk prediction, embryo selection and sperm donation. Reduced risk in both scenarios of donor selection, multiple genetic factors such as HLA type, and response guidance on prescribing specific medicines that affect the likelihood of side effects or side effects of the medicines; It can be applied in multiple situations.

[0100] HLA typing results can be compared with DNA-based typing, such as Sanger sequencing-based typing. These are obtained from various methods or derived from whole genome sequencing (WGS). Genetic risk scores are calculated using effect sizes from genome-wide association studies (GWAS), for example. One example is the effect size and efficacy of all relevant variants not in the MHC region. The goal is to sum the products of the doses of the resulting alleles. Based on HLA typing results (not tag SNPs) using one of the following methods: Combine or incorporate based on.

[0101] Combined ORs of PRS and HLA: Polygenic for all individuals in the validation cohort Calculate risk scores and obtain metadata (e.g., mean, standard deviation, etc.). Odds ratios ( ORs) are obtained for HLA alleles for which an association with the phenotype of interest has been established. ORs derived from the individual PRS compared with the validation cohort and HLA typing were: are combined as follows:

number

[0102] Directly incorporate HLA into PRS: HLA effect alleles are analyzed using the effect size and each effect allele. Incorporate directly into the polygenic risk score by adding the product of doses of This is a PRS HLA+ It is called PRS. HLA+ to all individuals in the validation cohort Calculate the PRS H Using ORs derived from the LA+ model and disease prevalence in the validation cohort This is used to estimate the lifetime risk of disease.

[0103] Example 4: Ranking disease risk profiles with application to embryo and sperm donor selection How to attach An exemplary method for ranking disease risk profiles is provided, as shown in FIG. First, the weights w d is calculated for each disease in the set of diseases d, which is the age at onset w a and disease severity w s is the sum of the weights of a Like coronary artery disease, more common in people with conditions that appear at birth, such as celiac disease, than in those that do not commonly manifest in the Similarly, w s is more likely to cause breast cancer than diseases with milder phenotypes such as vitiligo. The incidence of more severe diseases such as glaucoma is higher.

[0104] Family history and polygenic risk scores are then combined to identify each condition of interest for each embryo. Generate predicted risks.

[0105] Finally, the disease ranking and risk prediction were combined to calculate a single β for each embryo using the following formula: Score of S T where RR is a function of family history of a particular disease and polygenic risk score The relative risk is derived from the combination of

number

[0106] Consider three embryos with the following RRs for each of the above states: The scores are calculated and ranked accordingly. For embryo 1, the score is calculated as follows: To be:

number

[0107] The disease risk for each of the three embryos is shown in Table 5. [Table 5]

[0108] The same procedure is applied to sperm donor selection, with each donor being ranked for all diseases of interest. In the context of both embryo and donor selection, the score is based on a subset of diseases. (e.g., conditions with a family history in the expected parents) or where polygenic models are implemented Calculations are made for all diseases that are considered.

[0109] Alternatively, this method can be used without summing all states of interest to identify single embryos / Individual results can be prioritized. Each state receives a score, with the highest score Preference will be given to conditions with (or multiple) embryos. Using Embryo 1 above as an example, see Table 6. The scores and rankings shown were generated. [Table 6]

[0110] Example 5: Prediction of transmission of disease susceptibility variants to embryos. CRC susceptibility variant (APC c.3920T>A) (and / or insertion, deletion) One copy of the chromosome (a deletion and / or copy number variant) is found in the father's WGS. The allele is not present in the mother. This variant is not directly identified by sparse genotyping of the embryo. The parental whole chromosome haplotypes can be determined by any one of the above methods or by a combination of these. The embryo's genome is reconstructed to identify the haplotype containing the risk allele. It is determined that the type block is transmitted from the father to one of the embryos. The risk allele is It is stated as "present" in the embryo.

[0111] Example 6: Polygenic risk of common diseases using embryonic prediction. Breast cancer has a common genetic component. Genetic risk scores include 69 variants. used to assess breast cancer risk. Of these variants, 13% (9 / 69) Only 16 variants have been directly genotyped in the embryo. Genetic risk to the embryo based on these variants The percentile score is 84.6%. After embryo reconstruction, the embryo genotype is 98.6%. % (68 / 69) were estimated / inferred, and the new percentile for the embryo's genetic risk score was After the embryo is born, the child's DNA is genotyped and the PRS percentage is calculated. The mean variance was 76.2%, which is the mean variance for the genetic risk score from whole-genome embryo reconstruction. , with information on additional variants, it has higher precision and lower uncertainty. is doing.

[0112] Example 7: Prediction of transmission of disease-associated HLA types to the embryo. The mother has rheumatoid arthritis (RA). HLA typing results (WGS, P CR+ Sanger sequencing, or any other appropriate method), the mother has a high risk of this condition Carrying one copy of the HLA-DRB1*01:02 allele, which is associated with increased risk The father is homozygous for HLA-DRB1*04:02. This is an allele not known to be associated with an increased risk of RA. Based on the complete phasing of chromosome 6 of each parent and the reconstruction of the embryonic genome, Haplotype 2 (HM2) and paternal haplotype 2 (HF2) transmitted to the embryo The RA risk allele is carried on maternal haplotype 1 (HM1). Therefore, the embryo is predicted not to carry the risk allele. See, e.g., Figure 12. I want to be done that.

[0113] Example 8: Providing families with a range of disease risks in their children. Two parents tell their doctor that they are concerned about the risks of various genetic diseases in their expected child. Using the above method, the mean values ​​of the midparents and the recombination specifically calculates the risk of a child's disease when the genomes of both parents are taken into account. This will measure and guide the expected IVF treatment. See Figure 9.

[0114] Similarly, in the case of sperm donation, polygenicity based on WGS of the mother and future sperm donor(s) The distribution of offspring risk scores can be simulated by recombination (see Figure 9) .

[0115] Example 9: Incorporating family history (FHx) to improve risk estimation The risk of developing psoriasis is estimated to be 10-30% based on family history of the disease. In embryos with psoriasis parents, the polygenic model alone does not significantly alter the risk between embryos. As shown in Table 7, incorporating family history significantly improved the chances of embryo 1 vs. embryo 2 and embryo 3. The separation of embryos 2 and 3 was significantly improved, and embryos 2 and 3 were identified as having additional risk factors other than FHx. It is clear that [Table 7]

[0116] Similarly, family history can be incorporated to estimate risk in predicting transmission of disease-associated HLA types. The determination can be improved.

[0117] Example 10: Incorporation of HLA typing into psoriasis disease risk estimates The presence or absence of two HLA types associated with the risk of developing psoriasis may affect the overall risk of disease for the embryo. Examples of this include sperm donor selection or personal It can be extended to the context of genomic reports. [Table 8]

[0118] Incorporating family history to further improve risk estimation in predicting the inheritance of disease-associated HLA types This technique allows for the analysis of the embryonic genome, including the Rh status of the resulting fetus. This can be extended to predict blood type from

[0119] Example 11: Improving trait prediction accuracy If the genotype of a variant in a polygenic model is unknown in the embryo, the parental genotypes should be used. can be used to improve the accuracy of trait prediction. Instead of the estimated genotype, the parental genotypes at that site(s) are considered to determine the possible Use the probabilities in Table 9 below to determine the probability of each possible genotype. The dose is added to the risk score. In fact, four variants were not predicted in the embryo. This resulted in improved predictions for the polygenic model of Crohn's disease, as shown in Table 10 below. This improves the predictive accuracy as measured by the predicted percentile of polygenic risk. Polygenic risk score percentiles ("true") are based on direct genotyping from WGS. is determined using [Table 9] [Table 10]

[0120] Example 12: Haplotype disease risk Some disease risks are associated with phased haplotypes rather than individual variants. To better predict trait risk, embryo reconstruction is used to analyze phasing. Table 11 below shows the haplotypes and and associated risk of Alzheimer's disease (Corder et al., 1994). [Table 11]

[0121] The two variants are separated by 138 bp within the APOE gene. Neither rs429358 nor rs7412 were measured in the assay. However, embryo reconstruction methods do not involve estimating the risk of Zheimer's disease. can be used to infer that an embryo is ε3 / ε3 using This result can be later confirmed by whole genome sequencing of the child. It will be verified. [Table 12] Therefore, embryonic reconstruction can be used to identify APOE haplotypes and Alzheimer's disease, in general. , allowing for haplotype-based risk prediction of disease states.

[0122] Example 13: Sparse genotype scaffolding Using sparse genotyping as a scaffold for genome-wide phasing (See, e.g., FIG. 13), as measured by the switch error rate (SER), Improved performance over reference panels alone. By applying it to NA12878, the overall SER was The scaffold combined with the reference panel increased from 0.6% when using the scaffold alone. Using a set of approximately 140k high-confidence phased genotypes, the This difference is mainly due to the reduction in long switch errors. For example, on chromosome 1, the raw count data for long switch errors was reduced by more than 60%. Overall, the combined approach (scaffolding + reference paper) The error rate for long switches was reduced from 0.12% to 0.04%. A switch error results in an incorrect block that is predicted to be propagated, thus It is important in reconstruction.

[0123] Example 14: Polygenic risk scores Large-scale genome-wide association studies (GWAS) have identified genes associated with a wide range of diseases. These associations are useful for functional studies of disease biology, as well as for drug discovery. This paves the way for the discovery of common genetic barriers and improved disease risk prediction. Although these variants may have little predictive value, Combined into a score, they can explain a larger proportion of the genetic risk of disease. These multilocus genetic risk scores are also called polygenic risk scores (PRS). It is most commonly computed as a weighted sum of disease-associated genotypes.

number

[0124] How to validate and implement polygenic models and visualize risk estimates in consumer reports This article describes the following.

[0125] Selection of polygenic risk model At least 1000 individuals from a broad population are being tested for each condition of interest. We prioritized previously published polygenic models for the condition, which had limited statistical power. small studies that have been conducted in isolated populations that cannot be translated to other populations We excluded studies and models that used data from individuals in the UKBB study setting. Area under the curve (AUC) greater than 0.65 and / or upper and lower quantiles (details (see below) reported an odds ratio (OR) of >2 for individuals The characteristics of the published models and their evaluation statistics are listed in Table 13. [Table 13] TIFF2026021564000041.tif69162

[0126] If no published models are available, genome-wide significant p-value thresholds are selected from GWAS catalogs. SNPs meeting the p < 5e-8 value were used to construct a score as previously described (PMI D:30309464)

[0127] UK Biobank definitions of each phenotype Each model was validated and standardized using data from the UK Biobank cohort. This resource contains both genetic and disease information for 500,000 individuals. For the following analyses, only unrelated individuals were used. As shown in Table 14, It includes a combination of ICD-9 and ICD-10 codes, as well as self-reported disease We used disease and procedure codes to define each phenotype of interest. [Table 14] TIFF2026021564000043.tif210162 TIFF2026021564000044.tif211162 TIFF2026021564000045.tif122162

[0128] A subset of diseases is shown in Table 15 below. [Table 15]

[0129] Individuals were stratified by polygenic risk scores (PGS) to determine the risk of disease in this population. The incidence was investigated.

[0130] Model evaluation using the UKBB dataset. A polygenic risk score was calculated as the weighted sum of disease-associated genotypes. We calculate a score for each individual on the Ta.

[0131] Distribution of PRS across cases and controls: The dataset was divided into cases and controls for each trait, and the distribution of scores was calculated for cases and controls. A visual inspection of these distributions allowed us to determine the relative likelihood of each model. This gave a general idea of ​​how well the was able to distinguish between cases and controls. As an example, Figure 14 shows the distribution of PRS (mean = 0) for rheumatoid arthritis cases and controls. (The standard deviation is scaled to 1 in the figure.)

[0132] Receiver operating curve (ROC): ROC and area under the curve (AUC) are used to evaluate the sensitivity and specificity of the model at various risk thresholds. The degree was calculated by plotting the

[0133] PRS stratification into deciles: Stratifying individuals from UK Biobank into groups with different disease risk profiles The highest risk individuals (top decile of PRS) were compared with those with median risk. (individuals with PRS in the middle 40th to 60th percentile of the distribution). The disease prevalence for each disease was plotted, and the ratio of high risk to median risk was calculated across diseases. Figure 15 shows the OR per decile for rheumatoid arthritis.

[0134] Regression analysis incorporating age and gender: After calculating the PRS for all unrelated individuals in the UK biobank dataset, Stick regression was applied to each model. PGS is the regression coefficient of PRS, and PRS is Corresponds to odds ratios when standardized to a mean of 0 and a standard deviation of 1. Age and sex are Incorporated where available and applicable.

number

[0135] Odds ratios were then used to define high-risk and intermediate-outcome thresholds for reporting purposes. value was determined.

[0136] OR / SD by disease (mean-centered vs. z-transformed) According to the logistic model above, the OR / SD of the PRS is the effect size calculated by computer. This was obtained by standardizing the PRS variables (mean 0, SD 1) before calculating the The process is useful for achieving two goals: first, to improve the risk stratification ability of the PRS; PRSs for various diseases can be directly compared between patients. The effect sizes are different and therefore on very different scales. Their corresponding effect sizes are standardized. If the PRS is not standardized, it cannot be directly compared. This allows models to be directly ranked based on OR / SD, which allows for a better understanding of disease risk. Second, the UK population to the US population is ranked according to its ability to separate the populations. The aim is to enable the statistically accurate application of UKBB effect estimation. The relative risks were estimated from these odds ratios. (see below), using disease prevalence in the US population to identify the prevalence of specific PRS in the US The individual relative risk was accurately determined. Standardization of the UKBB PRS (UKBB mean and S D) to calculate the PRS of US individuals (after adjusting for the US PRS mean and SD). Random combinations of genetics will be available for use in models, at least in Europe. For individuals with ancestry, similar means and SDs of PRS can be expected in the population. The analytical results are shown in Table 16. [Table 16]

[0137] PRS stratification for disease vs. age: After stratifying individuals into different risk groups, UKBB data were used to assess the risk of these various groups. This information was used to estimate the proportion of the population diagnosed with the disease within the high-risk group (PRS). Visual plots are displayed across various strata, including the top 5% of people and the average risk group (the entire population). Assuming that the individual of interest has a PRS at the 75th percentile, the inventors The predicted patterns were determined for a group of individuals with similar genetic risk to those of the specific individuals of interest. -Percentages shown.

[0138] This plot illustrates the utility of the PRS in stratifying individuals based on disease risk. It is useful to ensure that the proportions of the population diagnosed within the different PRS strata are clearly separated. This confirms the model's ability to separate individuals based on risk.

[0139] Computer calculation of individual adjusted lifetime risk: We can start with the average lifetime risk for people of our gender in the United States. Then we can look at risk markers in the genome. This information is used to calculate a polygenic score based on the markers. Convert this to an "odds ratio" using the UKBB data. Finally, use the formula Factor the risk ratio and the average lifetime risk to estimate the individual's lifetime risk with this change:

number

[0140] where P0 is the prevalence of the condition in the UKBB and C0 is the average birth rate of the condition in the US. The lifetime risk, OR, is the odds ratio calculated above. The results are the individual risk compared to the population mean. It is an estimate of a person's own lifetime risk. For some conditions, average lifetime risk is not available. In these cases, the genetics analyzed will indicate whether or not there is an increased risk.

[0141] Defining the "high risk" threshold In some cases, a high genetic risk threshold was set based on known risk factors. For example, the relative risk of developing type 1 diabetes for an individual with an affected first-degree relative is 6.6. Therefore, the high-risk threshold for the PRS of type 1 diabetes corresponding to that relative risk is Representations that are not available or cannot achieve the threshold with this model For types, individuals with a 2-fold increase in relative risk or a 10% increase in absolute risk are considered high risk. A subset of phenotypes for which lifestyle or clinical factors indicated a high risk threshold was identified. The evaluation metrics for the project are shown in Table 17. [Table 17]

[0142] Example 15: Multifactorial Conditions (Polygenic Risk Score) Genomic DNA obtained from submitted samples will be analyzed by Illumina or BGI t Sequencing was performed using one of the following technologies. Reads were compared with the reference sequence. The sequence was aligned to the hg19 column and sequence variations were identified. Only alterations were analyzed. Deletions and duplications were not investigated unless otherwise stated above. In some scenarios, independent verification of HLA typing was performed by an external laboratory. The selected variant may be used in accordance with the ACMG (American College of Medical Annotated and interpreted according to the Guidelines for Medical Genetics Only pathogenic or likely pathogenic variants should be reported. The embryo genome was then subjected to typing and subsequent "parental support" analysis. The parent whole genome sequences are reconstructed using a genome reconstruction algorithm. Only variants observed in the parental genomes that are predicted to have an effect on the embryo were included. Polygenic risk scores were calculated for a subset of conditions in the reconstructed embryo genome. The models for each condition were evaluated in the UK Biobank population. Child risk scores may be refined using HLA typing. Individual lifetime risk is calculated based on population trends. Adjust baseline risk (US population) according to clinical information and polygenic risk score The upper and lower deciles were calculated by dividing the difference in lifetime risk by 10% or the lifetime risk by 10%. The model that resulted in a 1.9-fold increase was included in the report. Based on available evidence of efficacy, and at the discretion of the investigator, specific conditions (e.g., bilateral The lifetime risk of various conditions in specific embryos is shown in Figure 16A-C. show.

[0143] Using psoriasis as a specific example, Figures 17A-B show the progression of psoriasis in three exemplary embryos. Predictive risk scores are shown.

[0144] Example 16: Whole genome prediction of embryos using haplotype-resolved genome sequencing Haplotype-resolved genome sequencing was performed on a single embryo to predict the whole genome sequence of the embryo. Combined with a sparse set of genotypes from one or a few cell embryo biopsies. stLFR technology was used to sequence the father's haplotype-resolved genome. The heterozygous position (defined as an allele frequency of 1% or less) was evaluated. Inheritance of 17 sites was predicted in embryos with 89.5% accuracy.

[0145] The material used in this study was from women who had previously undergone a successful round of IVF with preimplantation genetic diagnosis. Trophectomees from a total of 10 embryos (day 5) were obtained retrospectively from participants (Table 16). Leaf biopsies were microarrayed using a rapid 24-hour microarray protocol to identify 300,000 common Each parent and four ancestors were genotyped for a panel of common SNPs. All parents were genotyped with the same panel. [Table 18]

[0146] Genomic DNA was extracted from whole blood or saliva samples. Neonatal and maternal DNA was , processed using 30XWGS on the BGI platform. Paternal samples were Trophectoderm biopsies from 10 day 5 embryos were processed using LFR. High-speed microarray analysis using the Illumina CytoSNP-12 chip in a pool DNA extraction, amplification, and genotyping with parents and grandparents using the RAY protocol Sibling embryos and parental SNP arrays were used as detailed in Kumar et al. (2015). The measurements were combined using the "parent support" (PS) method (Figures 18 and 19). Whole-genome sequencing of embryos allows combining the genotypes of PS embryos with parental haplotype blocks. This was predicted by (see Figure 18).

[0147] Example 17: Construction of whole chromosome haplotypes from haplotype blocks and parental information Haplotype resolution analysis of both parents was performed to construct chromosome-length haplotypes in an IVF setting. Genome sequencing was combined with information from sparse genotypes derived from sibling embryos. As part of the maximum support (PS) method, maximum likelihood estimates (Ma Maximum Likelihood Estimate (MLE) phase Recombination frequencies from the Map database were calculated using SNP array measurements from parents and sibling embryos. This sparse chromosome length haptic analysis is created by combining it with SNP array measurements. Although the lotype was not sufficient to predict the embryonic genome, it did predict the inherited genome sequence. To do this, molecularly derived high-density haplotypes (e.g., long-flag haplotypes) from parental samples are used. Segment Read Technology, 10x Genomics, CPT-seq, Pacific BioScience iences, Hi-C).

[0148] Information was obtained using several data streams. High-density haplotype blocks were generated. To achieve this, initial shotgun sequencing was performed on the maternal and paternal median 34x and 30x coverage. By sequencing a haploid subset of the resulting genomic DNA, 1.94 million copies of the mother's genomic DNA were identified. 94.2% of heterozygous SNVs in the mother and 92.4% of 1.89 million heterozygous SNVs in the father. These molecularly derived "high Combining "dense haplotype blocks" with sparse but chromosome-wide haplotypes The parental chromosome-length, haplotype-resolved genome sequences were then constructed. This sequence information was then used to was used to predict the inherited genome sequence of a person, but also to predict the future offspring of two parents. It could also be used to generate future eggs and sperm that will give rise to future children (e.g. by simulating

[0149] The future workflow for embryonic whole genome prediction is shown in Figure 19. At the first visit, blood samples are taken from the patient. This blood is used to generate the whole genome sequence of each parent, and the couple It is used to predict the disorders that may be at risk for. After counseling, parents undergo IVF and use conventional IVF PGD techniques to determine the genotype of the embryos and use this information Combined with the whole genome sequence information of the parents (haplotype resolution), the embryo's inherited genome predicts risk of disease and assesses disease risk.

[0150] The sibling embryo and parental genotypes were used to construct parental haplotypes for chromosome lengths. A statistical approach (such as maximum likelihood estimation) is used to eliminate the noise from each sibling embryo. The parent phases are determined from the database of meiotic recombination frequencies and information on the chromosomes.

[0151] Construction of whole-chromosome haplotypes Whole chromosome haplotypes are used to identify individuals, including but not limited to parents, grandparents, or children. It is constructed by sequencing the genomes of a person's relatives. For individuals with a genetic disorder, whole genome sequencing of the individual, their partner, and two or more children is performed. by administering the gene and determining the locus inherited by each child. This allows us to obtain the phase of all chromosomes of an individual (Figure 20). It provides haplotype information on a whole chromosome basis without altering the sequence determination process. This is the case, for example, if a couple already has two children and wants another, This would be appropriate in some cases where there are no DNA samples from any grandparents.

[0152] Chromosomal haplotypes from individual sperm The method of Example 17 involves sequencing DNA obtained from individual sperm. It is performed using whole chromosome haplotypes.

[0153] Example 18: Using embryonic genomic predictions to generate polygenic risk scores for genetically complex diseases Calculate. Genome-wide association studies have identified a number of conditions associated with type 1 diabetes, schizophrenia, Crohn's disease, celiac disease, It has become possible to build polygenic risk score models for conditions such as Alzheimer's disease. Their approach includes genome-wide analyses that include observed odds ratios for disease-associated SNPs. Obtain a list of significant SNPs for each individual and, depending on the location of the SNPs found in that individual, This approach involves calculating a "risk score" for each individual. Calculate the polygenic risk score of the younger brother and compare the polygenic risk observed when comparing sibling embryos in IVF cycles. Genetic risk scores were simulated for 12 siblings, two parents, and four grandparents. We used genome sequences from publicly available pedigrees. Each genome variant file (VCF) file) to a PLINK file and use the plink-score command as a variant. The polygenic risk scores for each individual in the family were calculated using the table. Scores were calculated for each sibling and for both parents. Polygenic risk scores were calculated as follows: Each individual in the 1000 Genomes Cohort (approximately 2500 individuals) and the subset of individuals who are Caucasian Polygenic risk scores for each family member were also calculated for a group of approximately 200-300 people. The data were compared with polygenic risk scores of a population-matched (European) group of individuals. Individuals were determined to be at high or low risk.

[0154] A polygenic risk score for celiac disease was developed within a Caucasian population incorporating multiple SNPs. (Abraham et al., 2014; PMC PMC3923679). The sensitivity for celiac disease was high, and the negative predictive value of this approach was high at certain PRS thresholds. Given a family history of celiac disease, we have calculated the specific PRS(-1 After calculating the PRS for each individual, the two individuals In the context of IVF, we consider these two Embryos can be selected for implantation, reducing the risk of disease by an estimated 10-fold.

[0155] Polygenic risk scores for Alzheimer's disease have been previously developed and has been found to be associated with early onset of disease ( Desikan et al., 2017 ; PM C5360219; Table 2). Parental PRSs are indicated by dark blue dashed lines. Each of the embryonic PRSs After calculating the PRS for each individual, the lowest polygenic risk Individuals with the highest polygenic risk score were more likely to develop Alzheimer's disease than those with the highest polygenic risk score. The risk of Heimer's disease is predicted to be reduced (median age of onset is 8 years instead of 80 years). 7 years old). [Table 19]

[0156] Example 19: Relevance calculations The genotype of the embryos is used to calculate an index of relatedness to individuals with undesirable genetic traits. For example, consider a maternal grandparent with schizophrenia. Step 1: Example 1 and Example After inferring the genome of the embryos from 2, the relatedness of each embryo to the genome of the affected individual is calculated. Step 2: Select the embryo with the lowest relatedness to the affected individual.

[0157] Example 20: Via Identity by Descent Using calculated genetic associations to predict disease risk An extension of Example 3, which substitutes genetic relatedness to affected individuals for disease prediction. To identify affected family members, identity by descent (IBD) is used. Different sibling embryos are not related to the affected family members. Because different IBDs exist, this information, in addition to the PRS score, can be used to assess the risk of disease in the embryo. In the following example, the risk of disease can be further increased by The assumption is that the disease is spread evenly throughout the population. Therefore, the risk is proportional to the number of affected individuals. It is proportional to the severity of a person's IBD. log(P / (1-P))=beta_1*PRS+beta_2*sex_male+beta_3*has_family_history+beta_4*IBD_aff ected_individual.

[0158] Example 21: Regions of shared genomic information Identifying regions of shared genetic information between two individuals, increasing the likelihood of Mendelian patterns Select embryos that do not contain homozygous regions. In couples with a genetic mutation, the offspring are likely to be homozygous for the disease-causing region. Genes with known disease associations are unevenly spread across the genome, By avoiding regions of homozygosity within known disease-causing genomic regions, disease can be minimized. Step 1: Determine the region of genetic information shared between the two parents. Step 2: Calculate the percentage of homozygous regions in each embryo. Step 3: Determine the genes that cause the disease. have the lowest region of homozygosity across the total or entire region known to be homozygous Select the embryo.

Claims

[Claim 1] The invention described in the specification.