Method for simulating expected embryo genotypes and estimating their risk of developing disease

JP2024536848A5Pending Publication Date: 2025-10-20マイオームインコーポレイテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024518661
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-27
Filing Date
2022-09-27
Publication Date
2025-10-20

AI Technical Summary

Technical Problem

Current IVF clinics and sperm donation centers cannot accurately predict the risk of complex polygenic diseases in embryos due to the inability to account for genetic, environmental, and lifestyle factors, leading to uninformed decision-making regarding IVF procedures.

Method used

A method using a concatenated approximation to simulate embryonic genotypes by phasing parental chromosomes and applying polygenic risk models to determine the probability of disease distribution in expected embryos, considering meiotic recombination sites and parental haplotypes.

Benefits of technology

This approach provides more accurate predictions of disease risk in embryos, allowing for informed decision-making in IVF procedures and reducing the financial and physical costs associated with unnecessary cycles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000035_0000
    Figure 00000035_0000
  • Figure 00000035_0001
    Figure 00000035_0001
  • Figure 00000035_0002
    Figure 00000035_0002
Patent Text Reader

Abstract

Disclosed herein is a method for determining the probability of disease distribution associated with an expected embryo by generating phased parent chromosomes, determining one or more meiotic recombination sites of interest, and generating one or more simulated embryo genotypes. A polygenic risk model can be applied to the genotype of each simulated embryo to generate a polygenic risk score and determine the probability of disease distribution of one or more diseases in the expected embryo.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 248,749, filed September 27, 2021, which is incorporated by reference herein in its entirety.

[0002] Technical Field The present disclosure relates generally to determining disease risk, and more specifically to methods for determining an expected embryo's risk of developing disease. [Background technology]

[0003] Currently, in vitro fertilization (IVF) clinics test for aneuploidies and single-gene disorders that are known to run in families. However, one in two couples have a family history of common diseases that are influenced by a combination of genetic, environmental, and lifestyle risk factors. Additionally, sperm donation clinics currently test for a propensity to develop a range of diseases caused by single-gene disorders but fail to consider the possibility that future embryos may develop complex polygenic diseases. Summary of the Invention

[0004] As described herein, exemplary embodiments determine a probability of a disease distribution associated with an expected embryo. A large number of simulated embryo genotypes are generated and subsequently used to generate a set of polygenic risk scores, which in turn allows for the determination of a disease distribution probability for one or more diseases for the expected embryo.

[0005] The embodiment examples described herein allow for the prediction of the occurrence or recurrence of disease associated with future expected embryos produced by IVF. Currently, certain methods can infer the genotype of existing embryos based on phased parent genomes and using microarray genotyping of existing embryos, but such methods do not provide information on the possibility of the occurrence or recurrence of certain diseases associated with future embryos. Considering the financial and physical costs associated with IVF, it may be advantageous to consider the possibility of disease(s) for expected embryos before undergoing IVF, so that all parties involved can make a more informed decision on the course of action to take. This may be particularly interesting for individuals who have a complex personal and / or familial history of diseases.

[0006] One possible approach to predicting disease risk is to simulate possible embryos using unlinked approximation methods, which allow each site in the embryo genotype to be treated independently from other sites. For example, to simulate the genotype of an embryo, the probability of each genotype from each parent can be determined and used to construct the genotype of the simulated embryo.

[0007] Unlinked approximation methods are relatively computationally simple and fast, and in some cases may produce satisfactory results. However, unlinked approximation methods have some drawbacks when linkage between sites is important in the embryo's genotype. In particular, unlinked approximation methods do not take into account that parental chromosomes are inherited in large segments, so such approximation methods underestimate the genetic variability between sibling embryos, which in turn leads to an underestimation of the variability in disease risk and the variability in genetic ancestry between sibling embryos. Because embryos inherit half of their DNA from one parent, their ancestry, quantified by the principal components, falls, on average, halfway between the ancestry of their parents. However, there is a large variation around this average, resulting in variation in genetic ancestry between siblings, which is missed by unlinked approximation methods.

[0008] In addition, non-additive effects between nearby gene sites (e.g., epistasis) or between haplotypes (e.g., dominance) play an important role in contributing to disease risk. One special case of this, for example in an oligogenic context, is compound heterozygosity, where two recessive alleles can either have no effect or be disease causative, depending on whether the two alleles are inherited from the same parent or from different parents.

[0009] The embodiments described herein advantageously allow for determining expected embryo disease risk using linkage approximation. In particular, parental chromosomes (e.g., paternal and maternal chromosomes) may be phased to obtain paternal and maternal genotypes. In some embodiments, genomic information from sibling embryos (e.g., previous IVF rounds) may also be determined. In linkage approximation, meiotic recombination sites of interest may be inferred based on parental chromosomes and genomic information from sibling embryos. Parental gametes may then be simulated based on the respective phased parental chromosomes and meiotic recombination sites of interest, which may then be used to generate the genotype of the simulated embryo. Thus, the linkage approximation method advantageously allows for the simulation of the genotype of an embryo inheriting chromosomal segments from each parent, and allows chromosome-length parental haplotypes to be determined across the genome for the simulated embryo. This allows for the preservation of parental ancestry, resulting in increased accuracy in the genetic variation of the simulated embryo. Linkage considerations may be particularly important when considering polygenic risk models that include highly effective linked single nucleotide polymorphisms (SNPs), such as autoimmune conditions. Therefore, subsequent polygenic risk scoring can be performed on the simulated embryos to generate more accurate probabilities of disease distribution for the expected embryos.

[0010] Thus, a method for determining the probability of disease distribution associated with an expected embryo is provided herein. The method includes generating a phased maternal chromosome set and a phased paternal chromosome set, and determining one or more meiotic recombination sites of interest. The method further includes generating one or more simulated embryo genotypes based on the phased maternal chromosome set, the phased paternal chromosome set, and the one or more meiotic recombination sites of interest. The method further includes applying a polygenic risk model to the one or more simulated embryo genotypes to generate a polygenic risk score set, where the polygenic risk score set includes a polygenic risk score for each simulated embryo genotype of the one or more simulated embryo genotypes, and determining the probability of disease distribution of one or more diseases for the expected embryo based on the polygenic risk score set.

[0011] In some embodiments, the method further comprises converting each polygenic risk score to a relative risk of the disease based on the polygenic risk score. In some embodiments, converting each polygenic risk score to a relative risk of the disease further comprises calculating an odds ratio for the polygenic risk score using an effect size model, and determining the relative risk of the disease based on the odds ratio and the prevalence of the disease associated with the particular disease.

[0012] In some embodiments, the method further comprises determining one or more risk thresholds for each disease. In some embodiments, the method further comprises determining a percentage of the probability of the disease distribution for the disease that meets the one or more risk thresholds corresponding to the disease.

[0013] In some embodiments, the method further comprises normalizing each polygenic risk score in the set of polygenic risk scores based on population data to generate a normalized set of polygenic risk scores, and determining the probability of the disease distribution is based on the normalized set of polygenic risk scores. In some embodiments, the population data comprises ancestry-specific population data.

[0014] In some embodiments, the method further comprises generating maternal gametes using a meiotic recombination model based on the phased maternal chromosome set and one or more meiotic recombination sites of interest. In some embodiments, the method further comprises generating paternal gametes using a meiotic recombination model based on the phased paternal chromosome set and one or more meiotic recombination sites of interest. In some embodiments, the method further comprises generating one or more simulated embryo genotypes based on the paternal gametes and the maternal gametes.

[0015] In some embodiments, the method further comprises obtaining a maternal genome from a maternal subject and a paternal genome from a paternal subject. In some embodiments, the method further comprises phasing the maternal genome to generate a phased set of maternal chromosomes. In some embodiments, the method further comprises phasing the paternal genome to generate a phased set of paternal chromosomes. In some embodiments, phasing the maternal genome or the paternal genome is performed using one or more of a population-based method or a molecular-based method.

[0016] In some embodiments, the method further comprises performing whole genome sequencing on a biological sample obtained from the maternal subject to determine the maternal genome. In some embodiments, the method further comprises performing whole genome sequencing on a biological sample obtained from the paternal subject to determine the paternal genome.

[0017] In some embodiments, the method further comprises determining sibling genomic information. In some embodiments, the method further comprises generating a phased set of maternal chromosomes based on the maternal genome and sibling genomic information. In some embodiments, the method further comprises generating a phased set of paternal chromosomes based on the paternal genome and sibling genomic information.

[0018] In some embodiments, chromosome-length parental haplotypes are obtained genome-wide for each simulated embryo.

[0019] In some embodiments, the method further comprises obtaining population genotype data comprising individual genotypes for a plurality of unrelated individuals. In some embodiments, the method further comprises generating a phased set of maternal chromosomes based on the maternal genome and the population genotype data. In some embodiments, the method further comprises generating a phased set of paternal chromosomes based on the paternal genome and the population genotype data.

[0020] In some embodiments, the method further comprises determining sibling genomic information. In some embodiments, the method further comprises determining one or more meiotic recombination sites of interest based on the sibling genomes, the maternal genome, and the paternal genome.

[0021] In some embodiments, the sibling genomic information is determined using at least one of array measurement, next generation sequencing, or whole genome sequencing, and the sibling genomic information is obtained from at least one of sibling embryos, full biological siblings, or half biological siblings.

[0022] In some embodiments, the method further comprises generating a recommendation for an additional in vitro fertilization (IVF) cycle based on the probability of disease distribution of one or more diseases of the expected embryo. In some embodiments, the method further comprises outputting the recommendation for the IVF cycle. In some embodiments, the recommendation for the additional IVF cycle indicates whether to perform an additional IVF round. In some embodiments, the method further comprises determining a disease occurrence risk based on the probability of the disease distribution, where the recommendation for the IVF cycle is based on the disease occurrence risk.

[0023] Similarly, a device for determining a probability of disease distribution associated with an expected embryo is disclosed herein. An exemplary device includes a processor and a memory, the memory storing software instructions that, when executed by the processor, cause the device to generate a phased maternal chromosome set and a phased paternal chromosome set and determine one or more meiotic recombination sites of interest. The processor and memory storing software instructions that, when executed by the processor, cause the device to further generate one or more simulated embryo genotypes based on the phased maternal chromosome set, the phased paternal chromosome set, and the one or more meiotic recombination sites of interest. The processor and memory storing software instructions that, when executed by the processor, cause the device to further apply a polygenic risk model to the one or more simulated embryo genotypes to generate a polygenic risk score set, where the polygenic risk score set includes a polygenic risk score for each simulated embryo genotype of the one or more simulated embryo genotypes, and cause the device to determine a probability of disease distribution of one or more diseases for the expected embryo based on the polygenic risk score set.

[0024] In some embodiments, the processor and memory, the memory storing software instructions that, when executed by the processor, cause the device to further convert each polygenic risk score to a relative risk of disease based on the polygenic risk score. In some embodiments, the processor and memory, the memory storing software instructions that, when executed by the processor, cause the device, when converting each polygenic risk score to a relative risk of disease, to further calculate odds ratios for the polygenic risk scores using an effect size model, and determine a relative risk of disease based on the odds ratios and the prevalence of disease associated with a particular disease.

[0025] In some embodiments, the processor and memory, the memory storing software instructions that, when executed by the processor, cause the device to further determine one or more risk thresholds for each disease. In some embodiments, the processor and memory, the memory storing software instructions that, when executed by the processor, cause the device to further determine a percentage of a probability of a disease distribution for a disease that meets one or more risk thresholds corresponding to the disease.

[0026] In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, further cause the apparatus to normalize each polygenic risk score in the polygenic risk score set based on population data to generate a normalized polygenic risk score set, where determining the probability of the disease distribution is based on the normalized polygenic risk score set. In some embodiments, the population data comprises ancestry-specific population data.

[0027] In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, further causes the device to generate maternal gametes based on the phased set of maternal chromosomes and the one or more meiotic recombination sites of interest using a meiotic recombination model. In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, further causes the device to generate paternal gametes based on the phased set of paternal chromosomes and the one or more meiotic recombination sites of interest using a meiotic recombination model. In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, further causes the device to generate one or more simulated embryo genotypes based on the paternal gametes and the maternal gametes.

[0028] In some embodiments, the processor and memory store software instructions that, when executed by the processor, cause the device to further obtain a maternal genome from the maternal subject and a paternal genome from the paternal subject. In some embodiments, the processor and memory store software instructions that, when executed by the processor, cause the device to further phase the maternal genome to generate a phased set of maternal chromosomes. In some embodiments, the processor and memory store software instructions that, when executed by the processor, cause the device to further phase the paternal genome to generate a phased set of paternal chromosomes. In some embodiments, phasing the maternal genome or the paternal genome is performed using one or more of a population-based method or a molecular-based method.

[0029] In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, further cause the device to perform whole genome sequencing of a biological sample obtained from the maternal subject to determine the maternal genome. In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, further cause the device to perform whole genome sequencing of a biological sample obtained from the paternal subject to determine the paternal genome.

[0030] In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, cause the device to further determine sibling genomic information. In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, cause the device to further generate a phased set of maternal chromosomes based on the maternal genome and sibling genome information. In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, cause the device to further generate a phased set of paternal chromosomes based on the paternal genome and sibling genome information.

[0031] In some embodiments, chromosome-length parental haplotypes are obtained across the genome for each simulated embryo.

[0032] In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, cause the device to further obtain population genotype data including individual genotypes for a plurality of unrelated individuals. In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, cause the device to further generate a phased set of maternal chromosomes based on the maternal genome and the population genotype data. In some embodiments, a processor and a memory, the memory storing software instructions that, when executed by the processor, cause the device to further generate a phased set of paternal chromosomes based on the paternal genome and the population genotype data.

[0033] In some embodiments, the processor and memory, the memory storing software instructions that, when executed by the processor, cause the device to further determine sibling genomic information. In some embodiments, the processor and memory, the memory storing software instructions that, when executed by the processor, cause the device to further determine one or more meiotic recombination sites of interest based on the sibling genomes, the maternal genome, and the paternal genome.

[0034] In some embodiments, the sibling genomic information is determined using at least one of array measurement, next generation sequencing, or whole genome sequencing, and the sibling genomic information is obtained from at least one of sibling embryos, full biological siblings, or half biological siblings.

[0035] In some embodiments, the processor and memory, the memory storing software instructions that, when executed by the processor, cause the device to further generate a recommendation for an additional in vitro fertilization (IVF) cycle based on a probability of disease distribution of one or more diseases for the expected embryo. In some embodiments, the processor and memory, the memory storing software instructions that, when executed by the processor, cause the device to further output an IVF cycle recommendation. In some embodiments, the recommendation for an additional IVF cycle indicates whether to perform an additional IVF round. In some embodiments, the processor and memory, the memory storing software instructions that, when executed by the processor, cause the device to further determine a risk of disease occurrence based on the probability of disease distribution, where the recommendation for an IVF cycle is based on the risk of disease occurrence.

[0036] Further disclosed herein is a computer program product for determining the probability of disease distribution associated with an expected embryo. The computer program product comprises at least one non-transitory computer-readable storage medium that stores software instructions that, when executed by the device, cause the device to generate a phased maternal chromosome set and a phased paternal chromosome set, and determine one or more meiotic recombination sites of interest. The computer program product further comprises at least one non-transitory computer-readable storage medium that stores software instructions that, when executed by the device, cause the device to generate one or more simulated embryo genotypes based on the phased maternal chromosome set, the phased paternal chromosome set, and one or more meiotic recombination sites of interest. The computer program product further comprises at least one non-transitory computer-readable storage medium storing software instructions that, when executed by the device, cause the device to apply the polygenic risk model to the one or more simulated embryo genotypes to generate a polygenic risk score set, where the polygenic risk score set includes a risk score for each polygene of the one or more simulated embryo genotypes, and determine a probability of a disease distribution for one or more diseases for the expected embryo based on the polygenic risk score set.

[0037] In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further convert each polygenic risk score into a relative risk of the disease based on the polygenic risk scores. In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device and when converting each polygenic risk score into a relative risk of the disease, further cause the device to calculate an odds ratio for the polygenic risk score using an effect size model, and determine a relative risk of the disease based on the odds ratio and a prevalence of the disease associated with a particular disease.

[0038] In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further determine one or more risk thresholds for each disease. In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further determine a percentage of a probability of a disease distribution for a disease that meets one or more risk thresholds corresponding to the disease.

[0039] In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further normalize each polygenic risk score in the set of polygenic risk scores based on population data to generate a normalized set of polygenic risk scores, where determining the probability of the disease distribution is based on the normalized polygenic risk score set. In some embodiments, the population data comprises ancestry-specific population data.

[0040] In some embodiments, the computer program product further comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to generate maternal gametes using a meiotic recombination model based on the phased set of maternal chromosomes and one or more meiotic recombination sites of interest. In some embodiments, the computer program product further comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to generate paternal gametes using a meiotic recombination model based on the phased set of paternal chromosomes and one or more meiotic recombination sites of interest. In some embodiments, the computer program product further comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to generate one or more simulated embryo genotypes based on the paternal gametes and the maternal gametes.

[0041] In some embodiments, the computer program product further comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to obtain a maternal genome from a maternal subject and a paternal genome from a paternal subject. In some embodiments, the computer program product further comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to phase the maternal genome to generate a phased set of maternal chromosomes. In some embodiments, the computer program product further comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to phase the paternal genome to generate a phased set of paternal chromosomes. In some embodiments, phasing the maternal genome or the paternal genome is performed using one or more of a population-based method or a molecular-based method.

[0042] In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to perform whole genome sequencing of the biological sample obtained from the maternal subject to determine the maternal genome. In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to perform whole genome sequencing of the biological sample obtained from the paternal subject to determine the paternal genome.

[0043] In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further determine sibling genomic information. In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further generate a phased set of maternal chromosomes based on the maternal genome and sibling genome information. In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further generate a phased set of paternal chromosomes based on the paternal genome and sibling genome information.

[0044] In some embodiments, chromosome-length parental haplotypes are obtained across the genome for each simulated embryo.

[0045] In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to obtain population genotype data including individual genotypes for a plurality of unrelated individuals. In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to generate a phased set of maternal chromosomes based on the maternal genome and the population genotype data. The computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to generate a phased set of paternal chromosomes based on the paternal genome and the population genotype data.

[0046] In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further determine sibling genomic information. In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further determine one or more meiotic recombination sites of interest based on the sibling genomes, the maternal genome, and the paternal genome.

[0047] In some embodiments, the sibling genomic information is determined using at least one of array measurements, next generation sequencing, or whole genome sequencing, and is obtained from at least one of sibling embryos, full biological siblings, or half biological siblings.

[0048] In some embodiments, the computer program product comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further generate a recommendation for an additional in vitro fertilization (IVF) cycle based on a probability of disease distribution of one or more diseases of the expected embryo. The computer program product further comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further output a recommendation for an IVF cycle. In some embodiments, the recommendation for an additional IVF cycle indicates whether to perform an additional IVF round. The computer program product further comprises at least one non-transitory computer readable storage medium storing software instructions that, when executed by the device, cause the device to further determine a risk of disease occurrence based on a probability of disease distribution, where the recommendation for an IVF cycle is based on the risk of disease occurrence.

[0049] The above brief summary is provided only for the purpose of summarizing some example embodiments described herein. The above-described embodiments are merely examples and should not be construed as narrowing the scope of the present invention in any way. It will be understood that the scope of the present disclosure encompasses many potential embodiments in addition to those summarized above, some of which are described in further detail below.

[0050] Having described in general terms above examples of specific embodiments, reference is now made to the accompanying drawings, which are not necessarily drawn to scale. Some embodiments may include fewer or more components than shown in the figures. [Brief description of the drawings]

[0051] [Figure 1] FIG. 1 shows an overview of an exemplary process for generating a polygenic risk score for a simulated embryo genotype that may be used in accordance with certain exemplary embodiments described herein. [Figure 2A-2B]FIG. 1 shows an exemplary process for phasing parent genomes using a parental support model that may be used in accordance with some embodiments described herein. [Diagram 3] 1 illustrates an example of a Hidden Markov Model configuration, according to some embodiments described herein. [Figure 4] 1 illustrates an example of a Hidden Markov Model computation in accordance with certain embodiments described herein. [Diagram 5] 1 illustrates an example of a framework for a parent support model, according to some embodiments described herein. [Figure 6] 1 illustrates an example of manipulating probability of disease distribution, according to some embodiments described herein. [Figure 7A-7L] 13 illustrates an example of the manipulation of disease distribution probabilities for various diseases determined using disjoint and connected approximations in accordance with some embodiments described herein. [Figure 8] 1 shows examples of polygenic risk score distributions for unlinked and linked approximations according to some embodiments described herein. [Figure 9] 13A-13C show examples of ancestry information contained in simulated embryo genotypes using unlinked and linked approximations according to some embodiments described herein. [Figure 10] 1 shows a schematic block diagram of an example device capable of performing various operations in accordance with some example embodiments described herein. [Figure 11] 1 shows an exemplary process for phasing parent genomes according to some embodiments described herein. [Figure 12] 1 illustrates an exemplary process for generating a simulated embryo genotype according to some embodiments described herein. [Figure 13A-13D] 13 shows examples of disease distribution probabilities for various diseases corresponding to Example 7. [Figure 14A-14B] 13 shows disease odds ratios by polygenic risk score decile corresponding to Example 6. [Figure 15]Correlation of embryo predictions with polygenic risk scores from born children is shown. [Figure 16] An example plot of transmitted haplotypes for sibling embryos is shown. [Figure 17] 13 illustrates an example flowchart for performing one or more actions based on an output using coupled approximations. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0052] Certain exemplary embodiments are described in more detail below with reference to the accompanying drawings, in which some, but not necessarily all, embodiments are shown. Because the invention described herein may be embodied in many different forms, the invention is not limited to only the embodiments described herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements.

[0053] Many modifications and other embodiments of the disclosures described herein will come to mind to one skilled in the art to which this disclosure pertains having the benefit of the teachings set forth in the foregoing description and the associated drawings. Accordingly, it is to be understood that the embodiments are not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Furthermore, while the foregoing description and the associated drawings describe exemplary embodiments in the context of certain exemplary combinations of elements and / or features, it is to be understood that different combinations of elements and / or features may be provided in alternative embodiments without departing from the scope of the appended claims. In this regard, for example, combinations of elements and / or features different from those expressly described herein are also contemplated as set forth in some of the appended claims. Although specific terms are used herein, they are used in a general and descriptive sense only and not for purposes of limitation.

[0054] Definitions of Certain Terms Unless otherwise defined, technical and scientific terms used herein have the same meaning as more commonly understood by one of ordinary skill in the art to which this invention belongs. Materials referred to in the following description and examples may be obtained from commercial sources unless otherwise noted.

[0055] The terms "computer-readable medium" and "memory" refer to non-transitory storage hardware, non-transitory storage devices, or non-transitory computer system memory that can store computer-executable instructions or software programs that can be accessed by a controller, a microcontroller, a computing system, or a module of a computing system. A non-transitory computer-readable medium can be accessed by a computing system or a module of a computing system to obtain and / or execute computer-executable instructions or software programs stored on the medium. Exemplary non-transitory computer-readable media include, but are not limited to, one or more types of hardware memory, non-transitory tangible media (e.g., one or more magnetic storage disks, one or more optical disks, one or more USB flash drives), computer system memory or random access memory (DRAM, SRAM, EDORAM, etc.), and the like.

[0056] The term "computing device" may refer to any computer implemented in hardware, software, firmware, and / or combinations thereof. Non-limiting examples of computing devices include personal computers, servers, laptops, mobile devices, smartphones, fixed terminals, personal digital assistants ("PDAs"), kiosks, custom hardware devices, wearable devices, smart home devices, IoT ("Internet-of-Things") enabled devices, and network-linked computing devices.

[0057] The term "about" means that the understood numerical value is not limited to the exact numerical value described herein, but refers to a numerical value that is substantially close to the described numerical value without departing from the scope of the present invention. As used herein, "about" is understood by those of ordinary skill in the art and will vary to some extent depending on the context in which it is used. If there are uses of the term that are not clear to those of ordinary skill in the art from the context in which the term is used, "about" will mean up to plus or minus 10% of the particular term.

[0058] The term "gene" refers to a stretch of DNA or RNA that encodes a polypeptide or plays a functional role in an organism. A gene can be a wild-type gene, or a variant or mutation of a wild-type gene. A "gene of interest" refers to a gene, or a variant of a gene, that is known or not known to be associated with a particular phenotype, or the risk of a particular phenotype.

[0059] The term "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (such as into an mRNA or other RNA transcript) and / or by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. Expression of a gene encompasses not only the expression of cellular genes, but also the transcription and translation of a nucleic acid(s) in cloning systems and other contexts. Where a nucleic acid sequence encodes a peptide, polypeptide, or protein, gene expression relates to the production of the nucleic acid (e.g., DNA or RNA, such as mRNA) and / or the peptide, polypeptide, or protein. Thus, "expression level" can refer to the amount of nucleic acid (e.g., mRNA) or protein in a sample.

[0060] The term "haplotype" refers to a group of genes or alleles that are inherited or expected to be inherited together from a single ancestor (father, mother, grandfather, grandmother, etc.). The term "ancestor" refers to a person from whom the subject is descended, or from whom, in the case of an embryo, the potential subject will be descended. In a preferred embodiment, ancestor refers to a mammalian subject, such as a human subject.

[0061] Data collection Genetic material for analysis by the methods described herein may be obtained from a variety of sources, including somatic cells (e.g., white blood cells, cells from tissue biopsies), germ cells (e.g., sperm, eggs, polar bodies). Genetic material may be collected from genetic relatives of the expected embryo (e.g., biological mother, biological father, biological siblings, sibling embryos, grandparents, etc.). In some embodiments, genomic DNA may be extracted from whole blood or saliva samples provided by paternal subjects, maternal subjects, sibling subjects (e.g., newborn children), grandparent subjects, etc.

[0062] Generation of simulated embryo genotypes - a concatenation approach As described above, it may be advantageous to use a linkage approach to generate simulated embryo genotypes, allowing parental haplotypes of chromosome lengths to be determined across the entire genome of the simulated embryo. Furthermore, using a linkage approach more accurately simulates the range of possible genotypes (and therefore PRS scores) between sibling embryos and preserves the composition of genome ancestry (which is lost when using unlinked genotypes), thereby allowing a local ancestry approach to be applied to risk scoring. In some embodiments, certain operations of the linkage approximation may be performed according to the method of "Whole-genome risk prediction of common diseases in human preimplantation embryos." (Nat Med 28, 513-516 (2022). Kumar et al., published March 21, 2022, and incorporated herein by reference in its entirety).

[0063] FIG. 1 shows an overview of various operations performed to generate simulated embryo genotypes and then predict the probability of disease distribution for the expected embryo. These operations are outlined in more detail below. Operations 102-106 may be performed to generate simulated embryo genotypes that represent possible genotypes for the expected embryo. Operation 108 may then be performed on the simulated embryo genotypes to determine a PRS score (e.g., disease risk) for the simulated embryo genotype. Operations 102-108 are repeated as many times as desired such that one or more simulated embryo genotypes may be generated for the expected embryo. In some embodiments, a threshold number of simulated embryo genotypes may be required. In some embodiments, at least 10 or more simulated embryo genotypes may be required. A PRS may then be generated for each simulated embryo genotype, and the PRS may be used to determine the probability of disease distribution for the expected embryo.

[0064] Sequencing A variety of molecular-based phasing methods are well known in the art and can be used to perform the methods described herein unless otherwise indicated by the context. Shotgun sequencing refers to a method of sequencing random DNA strands from genomes or large-scale genetic samples. DNA is randomly cut into many small segments and sequenced (e.g., using chain termination method) to obtain reads. Multiple overlapping reads for target DNA are obtained by performing several rounds of this fragmentation and sequencing. A computational algorithm then uses the overlapping ends of different reads to integrate the random segment reads into a continuous sequence. Shotgun sequencing can be used for whole genome sequencing. Any suitable form of sequencing, including those described herein, can be used to identify variants (e.g., SNPs) in a subject, which can then be used as a basis for measuring the genetic signal that indicates the ploidy state for the chromosome segment that contains the variant, as described elsewhere herein. According to certain aspects of the present invention, stratified sequencing can be used for whole genome sequencing. In some embodiments, phasing of parental genomic sequences may be performed according to the methods of International Publication No. WO 2021 / 067417 (Kumar et al.), published April 8, 2021, the contents of which are incorporated by reference herein in their entirety.

[0065] In some embodiments, DNA sequencing may include, for example, Sanger sequencing (chain termination sequencing). DNA sequencing may include the use of next-generation sequencing (NGS) or second-generation sequencing technology, which is typically characterized by being highly scalable and allows the entire genome to be sequenced at once. NGS technology usually allows multiple fragments to be sequenced at once, allowing "massively parallel" sequencing in an automated process. DNA sequencing may include third-generation sequencing technology (e.g., nanopore sequencing or SMRT sequencing), which generally allows longer reads to be obtained than those obtained by second-generation sequencing technology. Sequencing may include paired-end sequencing, where feasible, which sequences both ends of a DNA fragment, which may improve the ability to align reads to longer sequences. DNA sequencing may include sequencing by synthesis / ligation (e.g., ILLUMINA® sequencing), single molecule real-time (SMRT) sequencing (e.g., PACBIO® sequencing), nanopore sequencing (e.g., OXFORD NANOPORE® sequencing), ion semiconductor sequencing (Ion Torrent sequencing), combinatorial probe anchored synthesis sequencing, pyrosequencing, and the like.

[0066] In some embodiments, phasing uses linked-read sequencing, long fragment reads, fosmid pool-based phasing, contiguous conserved transposon sequencing, whole genome sequencing, Hi-C, dilution-based sequencing, targeted sequencing (such as HLA typing), or data generated from microarrays.

[0067] Some embodiments include the use of independently derived sparse phased genotypes to provide a scaffold to guide the phasing. Ancestral genotypes can be phased using computer software such as HapCUT, SHAPEIT, MaCH, BEAGLE, or EAGLE.

[0068] Population-based phasing may use a reference panel, such as 1000 Genomes or the Haplotype Reference Consortium, to phase genotypes. In some cases, the accuracy of phasing may be improved by the addition of genotype data for relatives, such as grandparents, siblings, or children.

[0069] Generating Phased Parent Genotypes To begin the process of generating simulated embryo genotypes, a phased maternal chromosome set and a phased paternal chromosome set may be generated for the maternal and paternal subjects, respectively. Each chromosome set may include one or more chromosomes corresponding to a homologous chromosome pair. The phased maternal chromosome set and the phased paternal chromosome set may be generated by phasing the genomes associated with the maternal and paternal subjects, respectively, using various methods, such as the population-based and / or molecular-based methods described above.

[0070] Both maternal and paternal genomes can be fully phased. Each parent genome can be phased using whole genome sequencing (WGS). In some embodiments, each parent genome is phased using a parent support model. The parent support model describes a method to combine SNP array measurements from one or more existing embryos and parents with recombination frequencies from a database (e.g., HapMap) to enable accurate prediction of chromosome copy number, insertions and deletions, embryo genotype, parent haplotype, and embryo parent haplotype origin hypothesis, using a method similar to that described in U.S. Pat. No. 8,515,679 (Rabinowitz et al.,) (incorporated herein by reference in its entirety). The parent support model can include one or more meiotic recombination models that simulate meiotic recombination sites during meiosis of each parent gamete.

[0071] For the reconstruction of the entire parent genome, two data sources are required. First, the whole genome sequence of the expected parent is required, as described below. Second, sibling genomic information is also required. Sibling genomic information can be obtained in various ways. In some embodiments, sibling genomic information can be obtained from sibling embryos by SNP microarray genotyping, next generation sequencing (NGS), etc. In some embodiments, sibling genomic information can be obtained from full biological siblings or half biological siblings by WGS, etc. Although sibling genomic information may be described herein as being determined for sibling embryos in some exemplary embodiments, it will be understood by those skilled in the art that alternative sources of sibling genomic information, such as full biological siblings and / or half biological siblings, can be used in addition to or instead of sibling embryos. When using SNP microarray genotyping to determine sibling genomic information, amplification is required because embryo biopsy only produces a limited amount of DNA.

[0072] The sibling genome data is shown in Figure 2A, where the allele measurements at each SNP are pattern coded based on the haplotype of parental origin in this example.

[0073] As shown in FIG. 2B, the parent support model may receive and process data sources (e.g., WGS from parents and SNP microarray genotyping (e.g., genomic information) from one or more sibling embryos) to generate one or more outputs. The one or more outputs may include phased parent genomes (e.g., both phased maternal genomes and phased paternal genomes), parental origin hypotheses, and genotypes of sibling embryos. The parent support model may be a hidden Markov model (HMM) that takes into account measurements of sibling genotypes and parental genotypes to improve accuracy across hundreds of thousands of positions. Table 1 outlines the inputs and outputs of the parent support model in further detail. [Table 1]

[0074] Figure 3 shows an example of the configuration of the parent support model, and Figure 4 shows an example of the output of the parent support model. In some embodiments, a complete implementation of the parent support model that supports meiotic crossover includes an HMM with a forward-backward (FBA) algorithm implemented.

[0075] HMM is a statistical Markov model, where the system being modeled is assumed to be a Markov process Xt with unobservable (i.e. hidden) states {x} over "time" t. The approach assumes that there is another process Yt with observable states {y}, whose behavior over time depends on X. The goal is to understand X by observing Y. In HMM, it can be assumed that for each time instance t, the conditional distribution of Yt depends only on Xt, via the probability P(y|x)=P(Yt=y|Xt=x). This probability is the output probability. The probability of an observable sequence Y=(Y1,...,Yn) can be written as P(Y)=ΣXP(Y|X)P(X) by Bayes' theorem.

[0076] Figure 4 further shows the posterior probability P(Xt=x|(y1,..,yn)), i.e. the probability of an unobservable state x at time t given the observed states (y1,..,yn). The forward algorithm computes the joint probability A(x,t)=P(Xt=x,y1,..,yt) of hidden states x and (y1,..,yt) as A(x,t)=P(Yt=yt|Xt=x)*ΣzP(x|z,t)*A(z,t-1), thus reducing the problem of order t to a problem of order t-1, as shown in Figure 3. P(x|z,t) is called the hidden state transition probability at time t. The posterior probability of any hidden state x at time n is therefore P(x|(y1,..,yn))~A(x,n).

[0077] Considering the above, Figure 5 shows the HHM framework for the parental support model. For the parental support model, we incorporate the fact that an embryo inherits an allele from the same parental homolog on consecutive SNPs unless meiotic recombination (probability estimated from databases such as HapMap) has occurred between the two SNPs. A joint distribution of genotype probabilities thus combines the array data, the genotypes of the individual embryos implied by the array data, and the parental haplotyping that can generate a distribution of genotypes among the various embryos. Consecutive SNPs represent "time" t. This approach is applied across each chromosome individually, at all sites on the array. The number of SNPs per chromosome ranges from about 4,300 (e.g., chromosome 21) to 23,700 (e.g., chromosome 2). This approach can be done across entire chromosomes instead of small regions of the genome. This allows for crossover within and between bins, and inference of problematic genomic sections.

[0078] Table 2 below further illustrates various parameters and outputs from the parent support model shown in FIG. [Table 2]

[0079] The transition probabilities shown in Figure 4 can be used to model meiotic recombination between consecutive SNPs. The transition probability from state z at SNP t-1 to state x at SNP t is modeled as follows:

number

[0080] In Equation 1, P(MG,t) and P(FG,t) are the prior distributions of parental haplotype populations at SNP t obtained from large-scale training datasets and public allele frequency databases. P(MH|MHz,t) and P(FH|FHz,t) are hypothetical transition probabilities, derived from a database (e.g., HapMap) that simulates the possibility of meiotic crossover between SNPs via the crossover probability between SNPs t-1 and t. Specifically, the transition probabilities can be expressed as P(H1|H1,t)=P(H2|H2,t)=1-ct (e.g., crossover does not occur) and P(H1|H2,t)=P(H2|H1,t)=ct (e.g., crossover has occurred), where ct is the crossover probability between SNPs t-1 and t.

[0081] The output probability, also shown in Figure 4, can be used to account for noise in microarray measurements in sequencing parent or sibling samples. Specifically, the output probability is the SNP-wise product of the data likelihood per channel for the true genotype G: P(data|genotype=G)=P(channel A data|G)*P(channel B data|G). Two different approaches can be used to model the likelihood of channel data. In the first approach, a simplified discrete output model is used.

[0082] For the discrete output model, the channel independent matrix product is obtained using Equation 2:

number

[0083] where din is the drop-in rate and dout is the drop-out rate. This product is based on the number of alleles A, B of the true genotype G and the measured genotype g, as shown in Table 3. The drop-in (din) and drop-out (dout) rate parameters are fitted separately using the microarray intensity data. In some embodiments, the drop-in rate of genomic data may be set to 0.1%, and the drop-out rate of genomic data may be set to 0.15%. [Table 3]

[0084] The second approach is a more complex continuous output model. For the continuous output model, a two-dimensional likelihood P(data|G)=P(channel A measurements|G)*P(channel B measurements|G) is used, where each channel likelihood is parameterized by a known continuous distribution of a particular genotype G. The distribution parameters are fitted to each couple using embryonic microarray measurements of the parental context that result in genotype G.

[0085] The output from the parental support model may be a phased set of maternal chromosomes and a phased set of paternal chromosomes.

[0086] Generation of simulated parent gametes Once the phased maternal chromosome set and the phased paternal chromosome set are generated, a meiotic recombination model can be used to generate maternal gametes based on the phased maternal chromosome set and paternal gametes based on the phased paternal chromosome set. Further, the meiotic recombination model can generate maternal gametes and paternal gametes based on one or more meiotic recombination sites of interest.

[0087] In some embodiments, maternal and paternal gametes may be simulated using a software-based approach, such as by using a parental support model as described above in Figures 2A-2B and / or by using one or more meiotic recombination models, which may be included in the parental support model. In some embodiments, meiotic recombination sites of interest (e.g., represented as breakpoints) may be derived using a software-based approach. Each phased parental chromosome set (e.g., maternal or paternal chromosome set) is then crossed at the meiotic recombination site of interest to generate the corresponding parental gamete (e.g., maternal or paternal gamete).

[0088] Generation of simulated embryo genotypes Once the maternal and paternal gametes are generated, these gametes may be combined to generate a simulated embryo genotype. As mentioned above, the above operations may be repeated as many times as desired, thereby generating one or more simulated embryo genotypes for the expected embryo. In some embodiments, a threshold number of simulated embryo genotypes may be required to increase confidence in determining downstream disease probabilities. For example, in some embodiments, at least 10 or more simulated embryo genotypes may be required. A PRS may then be generated for each simulated embryo genotype, and the PRS may be used to determine the probability of expected embryo disease distribution, as further described below.

[0089] Determination of polygenic risk score Polygenic risk scoring Once one or more simulated embryo genotypes are generated as described above, a polygenic risk model may be applied to each simulated embryo genotype to generate a polygenic risk score (PRS) (also referred to as a polygenic score (PGS) or genetic risk score (GRS)) for the corresponding simulated embryo genotype. One or more PRSs may be stored in a PRS set. The PRS may indicate the risk of a particular condition for an embryo having the genetic makeup of the simulated embryo genotype. The PRS determines whether a disease-causing variant is present in the simulated embryo genotype (inherited from a prior genome). The presence or absence of a particular disease-causing variant may increase disease susceptibility. Disease-causing variants include, for example, single nucleotide variants (SNVs), small DNA base insertions or deletions (indels), and / or copy number variants (CNVs).

[0090] In particular, the polygenic risk model may generate a polygenic risk score for the simulated embryo genotypes using Equation 3, described below.

number

[0091] In Equation 3, βi is the log odds ratio of the associated allele of SNPi, xi is the allele dosage of SNPi, and n is the total number of SNPs included in the polygenic risk model.

[0092] Table 4 shows examples of log-odds ratios associated with various disease-causing variants used to calculate the vitiligo PRS. [Table 4]

[0093] Normalization In some embodiments, each PRS may be normalized using one or more normalization methods. In some embodiments, each PRS is normalized based on population data. In some embodiments, the population data may be ancestry-specific population data. The ancestry-specific population data may be population data collected for a particular ancestor. In some embodiments, one or more haplotypes of the simulated embryo genotype may be evaluated to identify a corresponding ancestor for each haplotype. The ancestor with the largest portion (e.g., the largest proportion) may be selected for the simulated embryo genotype, and the ancestry-specific population data corresponding to that ancestor may be selected for the simulated embryo genotype. The genotype of each simulated embryo may then be normalized using the ancestry-based data.

[0094] One example of a method of normalization is standard score normalization, which can be expressed as Equation 4:

number

[0095] In Equation 4, z is the normalized PRS, x is the raw PRS (determined using Equation 1), μ is the mean for the matched population, and σ is the standard deviation for the matched population.

[0096] Additionally or alternatively, the PRS may be normalized by centering the PRS and dividing the centered PRS by the standard deviation, as shown in Equation 5 below:

number

[0097] In Equation 5, z is the normalized PRS, PRScentered is the centered PRS, and σ is the standard deviation of the population most closely related to the genotype of the simulated embryo, such as the population described in the 1000 Genomes Project. The centered PRS value may be determined by subtracting the PRS value predicted from a linear regression of the PRS on the first four principal component (PC) scores in control individuals (e.g., individuals without the phenotype of interest), as shown in Equations 6 and 7.

number

[0098] In Equation 7, βi is the log odds ratio for the associated allele of SNPi, and (PC)i is the corresponding principal component score determined using linear regression. In Equation 6, x is the PRS value, and xpred is the predicted PRS value.

[0099] Probability of determining disease distribution After determining one or more PRS (and in some embodiments, after normalization), each PRS for the simulated embryo genotypes may be used to determine a probability of disease distribution for the expected embryo. In some embodiments, a threshold number of simulated embryo genotypes may be required to determine an accurate probability of disease distribution. In some embodiments, at least 10 or more simulated embryo genotypes may be required.

[0100] Furthermore, one or more risk thresholds can be determined for each disease of interest. In some embodiments, the risk threshold can be a PRS value (or a relative risk value, discussed further below) that is associated with a higher than average disease risk. The risk threshold can be determined using clinical or other data.

[0101] Conversion to relative risk After determining one or more PRSs, and after normalization, each PRS can be converted into a relative risk (RR) of disease. The RR can be determined using an effect size model, which receives each PRS and determines the corresponding odds ratio for the PRS according to Equation 8.

number

[0102] In Equation 8, zscore is the normalized PRS as above, and B_PRS is the log odds ratio of the PRS. The effect size model can then determine the RR according to Equation 9.

number

[0103] In Equation 9, prev is the prevalence of the disease. Once the PRS is converted to RR, the probability of the disease distribution can be expressed using the RR instead of the PRS.

[0104] 7A-7K show additional examples of examples of disease distribution probabilities using RR for various diseases. Furthermore, in FIG. 7A-7K, both the unlinked approach and the linked approach are used to generate the disease distribution probabilities. Similarly, here, the arrows represent the predicted risk of each disease determined for the actual embryo. As shown in FIG. 7A-7K, in some examples, the unlinked approach brings the disease distribution probabilities closer to the linked approach, such as in FIG. 7A showing the disease distribution probabilities for Crohn's disease. However, in many other cases, the disease distribution probabilities determined by the unlinked approach diverge significantly from the disease distribution probabilities determined by the linked approach, such as in FIG. 7J showing the disease distribution probabilities for type 1 diabetes. As mentioned above, this divergence is due to the unlinked approach's failure to consider risk-contributing variants that are linked on the same haplotype, which are transmitted together and thus cooperatively increase the risk.

[0105] FIG. 8 shows an example of score distributions for unlinked and linked approximations. To better illustrate the impact of linked approximations on the determination of PRS, a simplified model with two sites may be considered. Each parent may be heterozygous at both sites (0 / 1). In the unlinked approximation, the probabilities that a child has genotypes 0 / 0, 0 / 1, and 1 / 1 are 0.25, 0.5, and 0.25, respectively. The weight of each risk allele may be assumed to be 0.5 to obtain the unlinked score distribution shown in FIG. 8. In the linked approach, it is assumed that these two sites are linked and can be reduced to a single site, where the weight of the risk allele is 1 to obtain the linked score distribution shown in FIG. 8. As shown in FIG. 8, the average PRS may be the same, but the distribution of PRS changes when linkage is taken into account.

[0106] In addition, Figure 9 further illustrates the impact of the unlinked and linked approaches on the genotypes of the simulated embryos in terms of the transmission of ancestry information in context. In the unlinked approach, the paternal contribution to the genotypes of the simulated embryos is obscured, which may produce artificial changes in the PRS predicted risk. Conversely, in the linked approach, local ancestry is maintained, thus allowing the PRS model to take the local ancestry approach into account when determining risk scoring.

[0107] Risk of disease development Once the probability of disease distribution is determined for the expected embryo, the development risk for one or more diseases can also be determined for the expected embryo. The development risk can be determined based on the probability of disease distribution and one or more thresholds. The one or more thresholds can be one or more PRS thresholds and / or RR thresholds that classify PRSs associated with high risk of disease. The proportion of simulated embryo genotypes that meet (e.g., exceed) a threshold can be used to determine the development risk for the expected embryo. The development risk can indicate the likelihood of a particular disease occurring in the expected embryo based on the simulated embryo genotypes determined using the linkage approximation.

[0108] For example, FIG. 6 shows an example of a disease distribution probability using RR for vitiligo. As shown in FIG. 6, a disease risk distribution (e.g., vitiligo) for an expected embryo determined using simulated embryo genotypes was generated and processed as described above. FIG. 6 further shows a triangle, which represents a parent RR calculated based on the provided and sequenced samples. Furthermore, an arrow shows a predicted risk of vitiligo for an actual embryo. The dotted line is a threshold used to distinguish PRS associated with high disease risk. The resulting disease distribution probability shown in FIG. 6 suggests that 93% of the embryos will have a RR below the RR threshold of 3, and 7% of the embryos will have a RR above the threshold. Thus, the development risk for the expected embryo may be 7%. Thus, families, medical providers, and other stakeholders may be informed that the risk of the expected embryo having a genotype associated with high risk vitiligo is relatively low.

[0109] Example implementation One example of the implementation of the linkage approximation is within a clinical setting, particularly one that performs preimplantation genetic testing (PGT-P) for polygenic disorders. Typically, women undergoing IVF often have more embryos available for transfer than are needed. This not only maximizes the chances of a successful pregnancy, but also provides an opportunity to minimize the chances of transmitting a disease affecting the mother or any of her family members to the child. Predicting disease risk in an embryo is possible for any disease that has a genetic component, including most common and rare diseases.

[0110] Preimplantation genetic testing is already routinely performed for aneuploidy screening (PGT-A), which involves obtaining an embryo biopsy. Embryonic cells collected in this process can then be genotyped by sequencing or microarray technology to gather base pair level information needed to predict general disease risk (PGT-P) for a particular embryo. Based on these predictions, IVF clinics can then select embryos for transfer that do not carry elevated disease risk.

[0111] However, in some cases, a particular round of IVF may produce embryos that are all determined to have an elevated risk of disease. As shown in the example flowchart of FIG. 17, a first IVF cycle (e.g., cycle 1) may be performed on the couple, and PGT-P may be used to infer the risk of disease for each embryo, as shown in operation 1702. In operation 1704, based on the results of the PGT-P, it may be determined whether all of the embryos are at elevated risk for one or more diseases. If one or more of the embryos are determined not to be at elevated risk for one or more diseases, those embryos may be selected for transfer and no additional IVF cycles may be required, whereby the process may proceed to operation 1712.

[0112] If all embryos are high risk, the process proceeds to operation 1706 where the expected embryo may be simulated using the concatenation approach described above. At operation 1708, it may be determined whether the developmental risk for the expected embryo meets one or more thresholds. For example, a threshold of 50% may be set such that a developmental risk having a value equal to or less than 50% meets the threshold. If a developmental risk of more than 50% is determined for the expected embryo, the threshold is not met.

[0113] If the developmental risk does not meet one or more thresholds, the process proceeds to operation 1712. At operation 1710, an additional IVF round (e.g., cycle 2) may not be recommended. This recommendation may occur if there is little chance of success for the expected embryos that do not have a high disease risk (e.g., as determined from PGT-P).

[0114] If the risk of occurrence meets one or more thresholds, the process proceeds to operation 1710. At operation 1710, it may be determined that a second cycle (eg, cycle 2) of IVF with PGT-P is recommended.

[0115] Regardless of the outcome, any recommendations for additional rounds of IVF may be output to clinical personnel (doctors, nurses, obstetricians, etc.), geneticists, patients, etc., so that the parties involved in determining the next course of action may be better informed about the risks and potential success rates of another round of IVF. The concatenated approach may be particularly beneficial when there is a large discrepancy in predicted risk between the genotypes of the simulated embryos.

[0116] Example of an implementation system The methods described herein may be implemented in a variety of systems. For example, in some embodiments, the systems may be used to generate phased parent chromosome sets, determine recombination sites of interest, generate one or more simulated embryo genotypes, apply polygenic risk models to one or more simulated embryo genotypes, determine disease distribution probabilities, etc.

[0117] The system may include one or more system devices, which may be realized by one or more computing devices or servers, as shown in FIG. 10 as device 1000. As shown in FIG. 10, device 1000 includes a processor 1002, memory 1004, and communication hardware 1006, each of which is described in detail below. Although the various components are shown in FIG. 10 as simply being connected to device 1000, it will be understood that device 1000 may further include a bus (not explicitly shown in FIG. 10) for passing information between any combination of the various components of device 1000. Device 1000 may be configured to perform the various operations described above.

[0118] The processor 1002 (and / or a co-processor or any other processor that assists or is otherwise associated with the processor) may communicate with the memory 1004 via a bus to pass information between components of the apparatus. The processor 1002 may be realized in a number of different ways, and may comprise one or more processing devices configured to perform independently. Additionally, the processor may comprise one or more processors configured in tandem via a bus to enable independent execution of software instructions, pipeline processing, and / or multi-threaded processing. Use of the term "processor" may be understood to include a single core processor, a multi-core processor, multiple processors in the apparatus 1000, a remote or "cloud" processor, or any combination thereof.

[0119] The processor 1002 may be configured to execute software instructions stored in the memory 1004 or otherwise accessible to the processor (e.g., software instructions stored on another storage device). In some cases, the processor may be configured to execute hard-coded functions. Thus, whether configured by hardware or software methods, or a combination of hardware and software, the processor 1002 represents an entity (e.g., physically embodied in circuitry) that can perform operations in accordance with various embodiments of the present invention, as configured accordingly. Alternatively, as another example, if the processor 1002 is embodied as an executor of software instructions, the software instructions may specifically configure the processor 1002 to perform the algorithms and / or operations described herein when the software instructions are executed.

[0120] The memory 1004 may be non-transitory and may comprise, for example, one or more volatile and / or non-volatile memories. In other words, for example, the memory 1004 may be an electronic storage device (e.g., a computer-readable storage medium). The memory 1004 may be configured to store information, data, content, applications, software instructions, and the like to enable the device to perform various functions in accordance with example embodiments contemplated herein.

[0121] The communications hardware 1006 may be any means, such as devices or circuits, implemented in either hardware or a combination of hardware and software, configured to receive and / or transmit data from / to a network and / or any other device, circuit, or module in communication with the apparatus 1000. In this regard, the communications hardware 1006 may include, for example, a network interface to enable communication with a wired or wireless communication network. For example, the communications hardware 1006 may comprise one or more network interface cards, antennas, buses, switches, routers, modems, and supporting hardware and / or software, or any other devices suitable for enabling communication over a network. Additionally, the communications hardware 1006 may comprise processing circuitry for causing such signals to be transmitted to the network or for processing receipt of signals received from the network.

[0122] The communications hardware 1006 may be configured to provide output to a user, and in some embodiments may be configured to receive indications of user input. The communications hardware 1006 may comprise a user interface, such as a display, and may further comprise components that control use of the user interface, such as a web browser, a mobile application, a dedicated user device, etc. In some embodiments, the communications hardware 1006 may comprise a keyboard, a mouse, a touch screen, a touch area, soft keys, a microphone, a speaker, and / or other input / output mechanisms. The communications hardware 1006 may use the processor 1002 to control the functionality of one or more of these user interface elements via software instructions (e.g., system software such as application software and / or firmware) stored in memory accessible to the processor 1002 (e.g., memory 1004).

[0123] Working Example Example 1 Preimplantation genetic testing (PGT) of in vitro fertilized embryos was performed to infer the inherited genomic sequence in 110 embryos across 10 couples and model susceptibility across 12 common conditions. The simulated expected embryos were then compared to the genomic sequence of the born child and also used to predict common disease risk by calculating polygenic risk scores and inferring inheritance of rare variants with large effects on disease risk.

[0124] Table 5 provides an overview of each couple (each assigned a respective case identifier). Performance for each case was determined by comparing the simulated embryo genotypes to the DNA genotypes of the born children. As shown in Table 5, the sites used for polygenic prediction yielded accuracies ranging from 99.0 to 99.4% for day 5 embryos and 97.2 to 99.1% for day 3 embryos. Case 1 included only day 3 embryos, and case 2 included both day 3 and day 5 embryos. All other cases included only day 5 embryos. Statistics are broken down by genotype (heterozygous or homozygous) in the born children. PGT from embryo biopsies was performed by a commercial laboratory (e.g., Natera, formerly Gene Security Network) on HumanCytoSNP-12 BeadChip arrays across 3 to 33 embryos. Coverage and accuracy were evaluated at genomic locations with high confidence genotype calls in parents and born children. [Table 5]

[0125] Additionally, Figure 15 shows the correlation between simulated embryo predictions and PRS from born children. The first graph in Figure 15 shows a close correlation between predicted raw PRS and measured (of born children) raw PRS, consistent with genotype concordance between predicted and measured polygenic risks.

[0126] The second graph in Figure 15 shows the correlation between predicted and measured Z scores derived from the raw PRS (r2 = 0.947). Families 5 and 9 were excluded from this analysis because the approach to mean-centered polygenic risk using population ancestry fails to account for admixture.

[0127] In the four cases where fresh blood samples were available, synthetic long-read sequencing was also performed on both maternal and paternal samples. Modifications to the above protocol included further high molecular weight DNA and library preparation using the TELL-Seq library using standard protocols, except for reduced transferase.

[0128] Example 2 For whole-genome sequencing of parental genotypes and born child genotypes, we targeted an average depth of 30x. The actual average coverage for all samples ranged from 29x to 111x. Table 6 shows the actual average depth used for each case in sequencing the corresponding mother, father, and child. Percentages greater than 20x (% ≥ 20x) indicate the percentage of genomic bases covered by at least 20 sequence reads. [Table 6]

[0129] WGS primary and secondary analyses were performed following the Broad Institute's best practices pipeline (GATK) implemented by Sentieon Software. The human reference genome sequence (GRCh37) was mapped using Burrow-Wheeler Aligner (bwa) version 0.7.17. Genotyping for each parent and actual child was then performed using two steps. First, joint variant calling for parents and born children was performed using Sentieon's GVCFtyper to capture sequences and filter them based on internal quality control thresholds. Joint variant calling allows all samples (e.g., maternal samples, paternal samples, born child samples) to be considered simultaneously to generate genotypes at many variant positions as opposed to the variant positions detected from a given sample. The internal quality control thresholds may include base quality control, median depth (DP), Fisher strand (FS), and quality score normalized by allele depth (QD). These internal quality control thresholds may be used to identify sequencing errors. Specifically, internal quality control thresholds were set as follows: BP ≥ 20, DP ≥ 8, FS < 30, and QD > 4. Genotypes were then called at sites specific to the polygenic model at a read depth of at least 8x.

[0130] Example 3 Embryo biopsies were genotyped by extracting and amplifying embryo DNA followed by genotyping using a rapid SNP microarray protocol (e.g., with Illumina's HumanCytoSNP-12 BeadChip). Sibling embryo and parental SNP microarray measurements were combined using a parental support model to determine maximum likelihood estimation (MLE) phases of heterozygous SNVs in each parent by combining recombination frequencies from the HapMap database with parent-derived SNP array measurements and sibling embryo-derived SNP array measurements. This combination can generate parental support haplotypes.

[0131] We then used a parental support model HMM to determine the parental haplotypes most likely to be transmitted to each embryo, given the SNP array measurements from the embryos and the MLE stage for each parent. The output of the HMM was used to inform meiotic recombination sites.

[0132] Example 4 To phase the WGS-derived variants in each parent, another simulation model was used to estimate haplotypes for the parents (e.g., using SHAPEIT4). Default parameters were used with additional database data, such as those available in the UK10K imputation cohort + 1000 genomes phase 3 (dataset EGAD00001000776), which served as a reference panel and parental support haplotype scaffold. This scaffold consists of approximately 200,000 phased variants and serves to anchor the phasing performed using the reference panel. Figure 11 shows the process of obtaining phased parental genotypes. Each chromosome is processed independently and in parallel, and all chromosomes are then combined. Multi-allelic sites were filtered out and discarded. To obtain additional performance for rare variants not represented by the reference panel, concatenated read sequences of high molecular weight DNA can be used.

[0133] In particular, concatenated read sequence data was generated for case IDs 5, 8, 9, and 10 using the TELL-Seq library preparation method. After read alignment and variant calling using the same process as above, with the addition of maintaining the molecular barcode information of each read, molecular phase was inferred using another model (e.g., HapCut2 model), which is a maximum likelihood-based tool for integrating haplotypes from DNA sequence reads. The positions of these haplotypes can be annotated by their global allele frequencies using the gnomad database.

[0134] Figure 16 shows a plot of transmitted haplotypes on chromosomes 3-8 for sibling embryos from family 5. The transmitted haplotypes are output from the parental support and form the basis of the PS embryo genotypes at the microarray sites. Green and red lines indicate parental haplotypes 1 and 2, respectively, for the maternal (MH) and paternal (FH) haplotypes in each embryo (regions of uncertainty are shown in yellow).

[0135] Example 5 To predict the whole genome sequence of each sibling embryo, the genotypes were combined with the parental genomes phased with the addition of haplotypes across chromosomes using the parental support model HMM. The parental haplotypes transmitted to the embryos were obtained by comparing the haplotypes with the genotypes of the sibling embryos. This was a process carried out across the maternal and paternal chromosomes, respectively. Figures 11 and 12 show this process in more detail.

[0136] Low quality sites in the parent and offspring genomes, as well as sites corresponding to multiallelic sites and Mendelian errors in each family's sequence data, can be filtered to form a set of "confident sites," which are used to assess coverage and accuracy. Predicted embryo genotype calls (obtained from the reconstruction) are compared to variants called by sequencing the offspring's DNA.

[0137] High-confidence sites were annotated by population allele frequency from the gnomADv2.1 dataset, which consists of approximately 15,000 whole genomes and 125,000 exomes from seven populations (African, Latino, Ashkenazi Jewish, East Asian, European, South Asian, and others). Variants with allele frequency <0.1% or not present in the gnomAD database were considered rare.

[0138] Table 7 shows the accuracy of the sites predicted by the reference panel and using concatenated read sequences. [Table 7]

[0139] Example 6 Polygenic risk scores and ancestry principal components were calculated for each simulated embryo genotype using a similar approach. In some cases, the prediction of the embryo genotype could not be determined, so the population allele frequencies were used to adjust the PRS score. The PRS score was centered and standardized as above and converted to odds ratios of disease taking into account the PRS. Specifically, Equation 3 was used, where β is the PRS effect size (i.e., log odds per standard deviation) obtained from UK Biobank, and PRS is the centered and standardized PRS. Figures 14A-14B show the disease odds ratios by polygenic risk score decile.

[0140] Example 7 Using a concatenation approach, simulated embryo genotypes were generated by starting with the phased genomes of both parents, adding recombinations between two maternal or two paternal chromosomes (to approximate meiotic recombination of gametes), and randomly combining these "virtual gametes". Haplotypes derived using the parental support model were combined with whole genome sequences to generate phased parent genomes. Sites of recombination were simulated using a meiotic recombination model (e.g., ped-sim, which includes a pedigree (two parents and one child) and a genetic map). Breakpoints (e.g., meiotic recombination sites) obtained from the meiotic recombination model were crossed with the phased parent genomes to generate "virtual gametes". Virtual gametes from the mother and father were then combined to generate simulated embryo genotypes. PRS was performed on these simulated embryo genotypes as described above. This process was repeated 500 times for each couple to generate a distribution of risk scores. In the unlinked approach, simulated embryo genotypes are generated by randomly selecting one allele from each parent and making no assumptions about whether adjacent variants are linked or not. Figures 13A-13D show the distribution of risk scores for various diseases using both the unlinked and linked approaches.

[0141] Furthermore, clinical PGT for aneuploidy (PGT-A) performed on the simulated embryo genotypes revealed that 69 of 110 embryos were euploid and 41 of 110 were aneuploid, as shown in Table 8. Whole genome reconstruction of the embryos was achieved by performing high coverage genome sequencing of both parents and by array measurements of sibling embryos as described above. [Table 8]

[0142] conclusion Many variations and other embodiments of the inventions described herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing description and the associated drawings. It is therefore to be understood that the invention is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, while the foregoing description and the associated drawings describe exemplary embodiments in the context of certain example combinations of elements and / or features, it is to be understood that different combinations of elements and / or features may be provided in alternative embodiments without departing from the scope of the appended claims. In this regard, for example, it is contemplated that combinations of elements and / or features other than those expressly described above may be set forth in some of the appended claims. Although specific terms are employed herein, they are used in a general and descriptive sense only and not for purposes of limitation.

[0143] All patents and publications mentioned in this specification are indicative of the level of those skilled in the art to which this invention pertains. All patents and publications in this specification are incorporated by reference to the same extent as if each individual publication was specifically and individually indicated to be incorporated by reference.

Claims

1. 1. A method for determining the probability of an expected embryo-associated disease distribution, comprising: generating a phased maternal chromosome set and a phased paternal chromosome set; determining one or more meiotic recombination sites of interest; and generating one or more simulated embryo genotypes based on the phased maternal chromosome set, the phased paternal chromosome set, and the one or more meiotic recombination sites of interest; applying at least one polygenic risk model to the one or more simulated embryo genotypes to generate a polygenic risk score set, the polygenic risk score set comprising a polygenic risk score for each simulated embryo genotype of the one or more simulated embryo genotypes; determining a probability of a disease distribution for one or more diseases for the expected embryo based on the polygenic risk score set; A method comprising:

2. 10. The method of claim 1, further comprising converting each polygenic risk score into a relative risk of disease based on the polygenic risk score.

3. converting each polygenic risk score into a relative risk of said disease; calculating an odds ratio for the polygenic risk score using an effect size model; determining the relative risk of a particular disease based on the odds ratio associated with the disease and the prevalence of the disease; The method of claim 2 further comprising:

4. 10. The method of claim 1, further comprising determining one or more risk thresholds for each disease.

5. 5. The method of claim 4, further comprising determining a percentage of a probability of a disease distribution for a disease that meets the one or more risk thresholds corresponding to the disease.

6. 2. The method of claim 1, further comprising normalizing each polygenic risk score in the polygenic risk score set based on population data to generate a normalized polygenic risk score set, wherein determining a probability of disease distribution is based on the normalized polygenic risk score set.

7. The method of claim 6 , wherein the population data comprises ancestry-specific population data.

8. generating maternal gametes based on the phased maternal chromosome set and the one or more meiotic recombination sites of interest using a meiotic recombination model; generating paternal gametes based on the phased paternal chromosome set and the one or more meiotic recombination sites of interest using the meiotic recombination model; generating one or more simulated embryo genotypes based on said paternal gametes and said maternal gametes; The method of claim 1 further comprising:

9. Obtaining the maternal genome from the maternal subject and the paternal genome from the paternal subject; phasing the maternal genome to generate a phased maternal chromosome set; phasing the paternal genome to generate a phased paternal chromosome set; further comprising: The method of claim 1.

10. 10. The method of claim 9, wherein the phasing of the maternal genome or the paternal genome is performed using one or more of a population-based method or a molecular-based method.

11. performing whole genome sequencing on a biological sample obtained from the maternal subject to determine the maternal genome; performing whole genome sequencing on a biological sample obtained from said paternal subject to determine the paternal genome; 10. The method of claim 9, further comprising:

12. Determining the genomic information of siblings; generating a phased maternal chromosome set based on the maternal genome and the sibling's genome information; generating a phased paternal chromosome set based on the paternal genome and the sibling's genome information; 10. The method of claim 9, further comprising:

13. obtaining population genotype data including individual genotypes for a plurality of unrelated individuals; generating the phased maternal chromosome set based on the maternal genome and the population genotype data; generating the phased paternal chromosome set based on the paternal genome and the population genotype data; 10. The method of claim 9, further comprising:

14. Determining the genomic information of siblings; determining one or more meiotic recombination sites of interest based on the sibling genomes, the maternal genome, and the paternal genome; 10. The method of claim 9, further comprising:

15. The genomic information of the siblings is determined using at least one of array measurement, next generation sequencing, or whole genome sequencing; The sibling genome is obtained from at least one of a sibling embryo, a full biological sibling, or a half biological sibling. The method of claim 12.

16. 2. The method of claim 1, wherein chromosome-length parental haplotypes are obtained across the genome for each simulated embryo.

17. generating a recommendation for an additional in vitro fertilization (IVF) cycle based on the disease distribution probability for one or more diseases for the expected embryo; outputting the IVF cycle recommendation; The method of claim 1 further comprising:

18. 18. The method of claim 17, further comprising determining a risk of disease development based on the probability of the disease distribution, and wherein the IVF cycle recommendation is based on the risk of disease development.

19. 19. The method of claim 18, wherein the additional IVF cycle recommendation indicates whether to perform an additional round of IVF.

20. 20. An apparatus for determining the probability of an expected embryo-associated disease distribution, the apparatus comprising: a processor; and a memory storing software instructions that, when executed by the processor, cause the apparatus to perform the steps of any one of claims 1 to 19.

21. 20. A computer program product for determining the probability of an expected embryo-associated disease distribution, comprising at least one non-transitory computer-readable storage medium storing software instructions that, when executed by said device, cause said device to perform the steps of any one of claims 1 to 19.