Widespread targets of non-mendelian epigenetic inheritance
A genome-scale analysis using long-read sequencing and a novel pipeline identifies non-Mendelian epigenetic inheritance patterns, addressing the lack of systematic study in mammalian systems and revealing patterns critical for disease susceptibility.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- JOHNS HOPKINS UNIVERSITY
- Filing Date
- 2025-11-20
- Publication Date
- 2026-05-28
Smart Images

Figure US2025056441_28052026_PF_FP_ABST
Abstract
Description
PATENT ATTORNEY DOCKET NO. JHU4760-1WO WIDESPREAD TARGETS OF NON-MENDELIAN EPIGENETIC INHERITANCECROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims benefit of priority under 35 U. S. C. § 119(e) to U. S. Provisional Application No. 63 / 723,094, filed November 20, 2024. The disclosure of the prior applications is considered part of and is herein incorporated by reference in the disclosure of this application in its entirety.STATEMENT REGARDING GOVERNMENT FUNDING
[0002] This invention was made with government support under grant DPI DK119129 awarded by the National Institute of Diabetes and Digestive and Kidney Diseases. The government has certain rights in the invention.BACKGROUND OF THE INVENTION FIELD OF THE INVENTION
[0003] The present invention relates generally to epigenetic inheritance patterns, and more specifically to the identification of non-Mendelian epigenetic inheritance patterns including emergent epigenetic patterns, and paramutation.BACKGROUND INFORMATION
[0004] It has been hypothesized that the inheritance of epigenetic traits, such as DNA methylation, may be more prevalent than is currently established. This phenomenon could help to explain poorly understood concepts in genetics such as (1) incomplete penetrance, (2) symptomatic heterozygotes for traits with recessive monogenic inheritance, and (3) complex phenoty pes which are influenced by environmental exposure and transmitted to offspring in a non-Mendelian manner. Notably, epigenetic traits can be inherited following several distinct patterns, some of which have been shown to violate Mendelian inheritance. These non-Mendelian patterns of epigenetic inheritance include: (1) parent-of-origin-specific DNA methylation patterns established by genomic imprinting; (2) polar overdominance, a combined genetic and parent-of-origin emergent effect which has been identified in rare instances such as the callipyge locus in sheep; and (3) transvection and paramutation, in which the methylation pattern of an allele is altered due to the presence of its homolog, previously observed in mammals only following the introduction of engineered mutations or transgenes. Furthermore, genomic loci have been identified which exhibit complex interactions between several distinct patterns of inheritance, such as the agouti viable yellow allele (Avy), regulated by an 11626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO intracistemal A particle (IAP), a subclass of endogenous retroviruses (ERV), for which epigenetic inheritance is influenced by parent-of-origin, strain effects, and environmental exposures. Despite the identification and characterization of isolated examples of these complex inheritance patterns, the role of non-Mendelian effects such as paramutation, transvection, and various forms of emergent epigenetic inheritance patterns have not been systematically studied in mammalian systems.
[0005] Here, a comprehensive genome-scale analysis of the inheritance of DNA methylation patterns in two tissues in relation to genotype, parent-of-origin, and sex was performed to lay a foundation necessary to ultimately understand the complex relationship between genetics, epigenetics, the environment, and disease. As there is some ambiguity in the literature regarding terminology used when discussing intergenerational epigenetic inhentance, a glossary of terms incorporating the definitions used is provide. In performing this analysis, DNA methylation levels at distinct loci throughout the genome is treated as traits to characterize their inheritance patterns across multiple generations in crosses of inbred mice. For this study, two strains from the Collaborative Cross (CC) mouse model were chosen, a panel of recombinant inbred strains derived from crosses of eight founder strains, to increase the level of inter-strain heterozygosity while minimizing hybrid dysgenesis and other genetic regulator}' incompatibilities which are frequently observed in inter-specific and subspecific crosses. Using long-read Oxford Nanopore Technologies (ONT) sequencing of these inbred CC strains as well as Fl and F2 crosses derived from them, direct measure both genetic sequence and DNA methylation together on the same sequenced strand of DNA were obtained, allowing for haplotype phasing to assign a parent-of-origin to each read, a tool to map methylation data from these highly divergent CC strains into a single pseudo-hybrid intermediate genome to enable the rigorous analysis of allele-specific DNA methylation without the bias introduced by the choice of a single reference genome was developed. To investigate the effect of these epigenetic inheritance patterns on gene expression, methylation analysis was coupled with measurements of allele-specific expression for the genes associated with these epigenetic inheritance patterns in which there exists a transcribed polymorphism between the two chosen CC strains. Finally, as DNA methylation is highly tissue-specific, and, as such, the intergenerational inheritance of DNA methylation patterns may also be tissuespecific, this inheritance genome-wide in two distinct tissues to gain insight into the tissuespecificity of these patterns was comprehensively analyzed.21626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0006] By designing the experiments in this manner, at least 522 autosomal examples of distinct forms of non-Mendelian inheritance, including the first paramutation event identified in a mammalian genome in the absence of transgenic manipulation, 54 emergent epigenetic inheritance patterns, at least 51 regions under the control of distal trans-acting methylation quantitative trait loci (meQTLs), and at least five new imprinted genes with one located on the X chromosome were identified. Two additional highly likely examples of paramutation over IAP elements were also identified. Furthermore, many examples of Mendelian inheritance of DNA methylation patterns, including 7,081 regions under the control of local cis-acting meQTLs. Notably, of the 6,604 and 2,059 autosomal regions which exhibit one of these intergenerational inheritance patterns in liver and muscle, respectively, at least 1,038 exhibit the same pattern in both tissues, representing 15.6% and 50.3% of the total patterns for each tissue were characterized and mapped (FIG. 8). This suggests the existence of both common and distinct mechanisms of regulation of intergenerational epigenetic inheritance patterns across tissues.SUMMARY OF THE INVENTION
[0007] The present invention is based on the seminal identification of non-Mendelian epigenetic inheritance patterns including emergent epigenetic patterns, and paramutation.
[0008] In one embodiment, the present invention provides a method for screening a subject for susceptibility to a disease or condition including: detecting a methylation status of a plurality of nucleic acid sequences in a sample from the subject, wherein detecting the methylation status comprises: (i) analyzing inheritance of epigenetic patterns in a nucleic acid sample from the subject, (ii) identifying the human subject as having an increased risk of developing the disease or condition.
[0009] In one aspect, the method further includes detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences. In another aspect, combining the genetic variations and the methylation status of the plurality of nucleic acid sequences includes using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genome-wide association studies (GWAS) and epigenome- wide association studies (EWAS). In one aspect, analyzing inheritance of epigenetic patterns includes identifying identifying Mendelian and non-Mendelian patterns of epigenetic inheritance. In some aspects, non-Mendelian patterns of epigenetic inheritance comprises non-dominant trans- acting meQTLs; imprinted genes; sex- 31626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO specific methylation patterns; emergent epigenetic patterns including over dominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation. In other aspects, analyzing includes using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome- wide in a nucleic acid sample from the subject. In some aspects, the long-read sequencing includes genetic polymorphisms and methylation status. In various aspects, the long-read sequencing includes long-read nanopore sequencing. In one aspect, wherein the subject is a mammal. In some aspects, the subject is human. In other aspects, the mammal is a cattle, a sheep, or a horse.
[0010] In another embodiment, the invention provides a method for generating a prognosis of a subject's susceptibility to a disease or condition including: (i) analyzing inheritance of epigenetic patterns, (ii) identifying the subject as having an increased risk of developing the disease or condition, wherein non-Mendelian epigenetic inheritance comprises emergent epigenetic patterns, and paramutation, (iii) storing Mendelian and non-Mendelian epigenetic inheritance data in a database that includes a set of information related to said subject; (iv) correlating the data with an association between the data and susceptibility to the disease or condition in the database; (v) generating a prognosis of the subject's susceptibility to the disease or condition; and (vi) communicating the prognosis of susceptibility to a medical practitioner.
[0011] In one aspect, the set of information related to said subject includes family medical history, diet, exercise and medical history of said subject. In another aspect, analyzing inheritance of epigenetic patterns includes identifying Mendelian and non-Mendelian patterns of epigenetic inheritance. In some aspects, non-Mendelian epigenetic inheritance includes nondominant trans- acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation. In other aspects, analyzing includes using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome- wide in a nucleic acid sample from the subject. In one aspect, the method further includes detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences. In one aspect, combining the genetic variations and the methylation status of the plurality of nucleic acid sequences includes using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenofype-associated genes missed by traditional genome-wide association studies (GW AS) and epigenome-wide association studies (EWAS). In some aspects, the long-read sequencing 41626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO includes genetic polymorphisms and methylation status. In various aspects, the long-read sequencing includes long-read nanopore sequencing. In one aspect, wherein the subject is a mammal. In some aspects, the subject is human. In other aspects, the mammal is a cattle, a sheep, or a horse.
[0012] In an additional embodiment, the invention provides a system for screening a subject for susceptibility to a disease or condition including: one or more processors; and a program having programming instructions stored thereon, which, when executed by the one or more processors, causes the system to perform operations including: detecting a methylation status of a plurality of nucleic acid sequences in a sample from the subject, wherein detecting the methylation status comprises: (i) analyzing inheritance of epigenetic patterns in a nucleic acid sample from the subject, and (ii) identifying the subject as having an increased risk of developing the disease or condition.
[0013] In one aspect, the operations further include: detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences. In another aspect, combining the genetic variations and the methylation status of the plurality of nucleic acid sequences includes using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genomewide association studies (GWAS) and epigenome-wide association studies (EWAS). In one aspect analyzing inheritance of epigenetic patterns includes: identifying Mendelian and non-Mendelian patterns of epigenetic inheritance. In some aspect, non-Mendelian epigenetic inheritance includes non-dominant / ram-acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation. In other aspects, analyzing the inheritance of epigenetic patterns in the nucleic acid sample from the subject includes: using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome- wide in a nucleic acid sample from the subject. In some aspects, the long-read sequencing includes genetic polymorphisms and methylation status. In various aspects, the long-read sequencing includes long-read nanopore sequencing. In one aspect, wherein the subject is a mammal. In some aspects, the subject is human. In other aspects, the mammal is a cattle, a sheep, or a horse.
[0014] In a further embodiment, the invention provides a system for generating a prognosis of a subject's susceptibility to a disease or condition including: one or more processors; and a 51626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO program having programming instructions stored thereon, which, when executed by the one or more processors, causes the system to perform operations including (i) analyzing inheritance of epigenetic patterns, (ii) identifying the subject as having an increased risk of developing the disease or condition, wherein non-Mendelian epigenetic inheritance includes non-dominant trans-acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation, (iii) storing Mendelian and non-Mendelian epigenetic inheritance data in a database that includes a set of information related to said subject, (iv) correlating the data with an association between the data and susceptibility to the disease or condition in the database, (v) generating a prognosis of the subject's susceptibility to the disease or condition, and (vi) causing display of an indication of the prognosis of susceptibility to a medical practitioner.
[0015] In one aspect, the set of information related to said subject includes family medical history, diet, exercise and medical history of said subject. In another aspect, analyzing the inheritance of epigenetic patterns includes identifying Mendelian and non-Mendelian patterns of epigenetic inheritance s. In one aspect, non-Mendelian epigenetic inheritance includes nondominant trans- acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation. In another aspect, analyzing the inheritance of epigenetic patterns includes: using long-read nanopore sequencing in conjunction with analysis of allele specific methylation patterns genome- wide in a nucleic acid sample from the subject. In one aspect, the operations further include: detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences. In one aspect, combining the genetic variations and the methylation status of the plurality of nucleic acid sequences includes using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genome-wide association studies (GWAS) and epigenome-wide association studies (EWAS). In some aspects, the long-read sequencing includes genetic polymorphisms and methylation status. In various aspects, the long-read sequencing includes long-read nanopore sequencing. In one aspect, wherein the subject is a mammal. In some aspects, the subject is human. In other aspects, the mammal is a cattle, a sheep, or a horse.61626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0016] In one embodiment, the invention provides a system for analyzing allele-specific DNA methylation patterns in in crosses of inter-strain and inter-specific crosses of organisms including one or more processors; and a memory having programming instructions stored thereon, which, when executed by the one or more processors, causes the system to perform operations including (a) generating a pairwise graphical genome between two reference genomes, (b) creating a pseudo-hybrid intermediate genome, wherein the pseudo-hybrid intermediate genome is a unified reference coordinate system including strain / allele-specific CpGs from both reference genomes; (c) aligning the pseudo-hybrid intermediate genome and the corresponding edges of the graphical genome for coordinate mapping; (d) identifying heterozygous genetic variants between the two provided reference genomes: (e) extracting consensus haplotyping decisions; and (f) mapping methylation coordinate from the two reference genomes into the ‘pseudo-hybrid intermediate genome, thereby analyzing allelespecific DNA methylation patterns.
[0017] In one aspect, generating the pairwise graphical genome includes dividing the two input reference genomes into smaller, corresponding pieces. In another aspect, the operations further include processing long-read sequencing data. In various aspect, the long-read sequencing data includes Oxford Nanopore Technologies (ONT) sequencing data. In one aspect, analyzing allele-specific DNA methylation patterns includes identifying intergenerational epigenetic inheritance patterns. In another aspect, the intergenerational epigenetic inheritance patterns include intergenerational paramutation. In some aspects, the inheritance patterns violate the laws of Mendelian inheritance. In one aspect, the organism is a mammal. In some aspects, the mammal is a human. In other aspects, the mammal is a cattle, a sheep, or a horse. In another aspect, the organism is a model organism. In one aspect, the references genomes are non-engineered mammalian genomes.
[0018] In a further embodiment, the invention provides a method for generating a prognosis of a subject's susceptibility to a disease or condition including (i) analyzing non-Mendelian patterns of epigenetic inheritance, (ii) identifying the human subject as having an increased risk of developing the disease or condition, (iii) storing the non-Mendelian epigenetic inheritance data in a database that includes a set of information related to said subject; (iv) correlating the data with an association between the data and susceptibility to the disease or condition in the database; (v) generating a prognosis of the subject's susceptibility to the disease or condition; and (vi) communicating the prognosis of susceptibility to a medical practitioner.71626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0019] In one aspect, the set of information related to said subject includes family medical history, diet, exercise and medical history of said subject. In another aspect, the non-Mendelian epigenetic inheritance includes non-dominant trans-acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation. In some aspects, analyzing comprises using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome-wide in a nucleic acid sample from the subject. In one aspect, the method further includes detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences. In some aspects, combining the genetic variations and the methylation status of the plurality of nucleic acid sequences comprises using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genomewide association studies (GWAS) and epigenome-wide association studies (EWAS). In another aspect, the long-read sequencing includes genetic polymorphisms and methylation status. In some aspects, the long-read sequencing includes long-read nanopore sequencing. In one aspect, the subject does not present the phenotype of a disease or disorder but is known for having inherited a mutation associated with a Mendelian disease or disorder from a family member affected with the Mendelian disease or disorder. In another aspect, the subject present the phenotype of a disease or disorder but is known for not having inherited a mutation associated with a Mendelian disease or disorder from a family member affected with the Mendelian disease or disorder. In some aspects, the subject is a mammal. In various aspects, the subject is human. In other aspects, the mammal is a cattle, a sheep, or a horse.BRIEF DESCRIPTION OF THE DRAWINGS
[0020] FIGs 1A-1C illustrate the identification and characterization of Mendelian and non-Mendelian epigenetic inheritance patterns. FIG. 1A shows that a long-read ONT sequencing was performed on DNA from inbred samples of two genetically divergent CC strains and their Fl crosses. Allele-specific methylation patterns were analyzed to identify and map both Mendelian and non-Mendelian patterns of epigenetic inheritance. Matched RNA was also sequenced and used to associate identified epigenetic inheritance patterns with gene expression. Additionally, F2s from the same CC strains ere sequenced to further investigate the genetic and non-genetic factors regulating the identified epigenetic inheritance patterns.81626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO FIG. IB is a schematic of a cis-acting meQTL in an F2 mouse, indicating that the local haplotype of the DMR will define the identity of the meQTL, as recombination is unlikely to occur between the SNP and the CpGs over which it regulates methylation when they are in close proximity (i.e. less than 50cM apart). A cis- acting regulatory' mechanism will establish an F2 population in which the local haplotype of the DMR is associated with the methylation level (e.g. CC019 alleles are unmethylated while CC037 alleles are methylated), as shown in (a). FIG. 1C is a schematic of a dominant rraws-acting meQTL in an F2 mouse, indicating that the local haplotype of the DMR will not necessarily define the identity of the meQTL, as there is a high likelihood that independent assortment of chromosomes or meiotic recombination will occur between the SNP and the CpGs over which it regulates methylation if they are greater than 50cM apart. A / ra -acting regulatory mechanism will establish an F2 population in which the local haplotype over the DMR is not associated with the methylation level (i.e. there is no clear methylation pattern associated with the CC019 or CC037 alleles), as shown for a dominant trans- acting meQTL in (a).
[0021] FIGs 2A-2E illustrate Czs-acting meQTLs. FIG. 2A is a generic representation of methylation in the context of a cis-acting meQTL in inbred CC019 and CC037 samples and their Fl crosses. FIG. 2B is a graph illustrating Cis-acting meQTL identified in both the liver (left) and muscle (right) over Loxhdl. FIG. 2C is a graph illustrating methylation pattern of the tissue-specific cis-acting meQTL identified over Tpcnl in the liver (left) and not in the muscle (right). FIG. 2D shows graphs illustrating Cis-acting meQTL identified in the liver over Vasp and Opa3, shown for inbred and Fl samples (left) and F2 samples (right). FIG. 2E is a graph illustrating allele-specific expression of Haao in the liver, exhibiting regulation by a cis-acting eQTL and overlapping a cis-acting meQTL in both the liver and muscle. For the methylation plots shown in FIGs. 2B, 2C and 2D, bold lines represent coverage-weighted mean methylation of the respective group and CpG sites included in the final analysis are denoted by tick marks on the x-axis.
[0022] FIGs 3A-3F illustrate non-dominant trans-acting meQTLs. FIG. 3A is a generic representation of methylation in the context of an incomplete dominant / ra -acting meQTL in inbred CC019 and CC037 samples and their Fl crosses.. FIG.3B is a generic representation of methylation in the context of a codominant trans-acting meQTL. Methylation patterns in FIGs 3A and 3B indicate that read-level methylation analyses are required to distinguish incomplete dominant from codominant trans-acting meQTLs, as their methylation patterns will appear identical at the sample-level.. FIG. 3C illustrates non-dominant trans- acting meQTL 91626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO identified in both the liver (left) and the muscle (right) over Fggy. FIG. 3D illustrates tissuespecific non-dominant trans-acting meQTL identified over Ctdspl in the liver (left) and not in the muscle (right). FIG.3E illustrates non-dominant trans-acting meQTL identified in the liver over Isoc2b, show n for inbred and Fl samples (left) and F2 samples (right). FIG.3F illustrates allele-specific expression of Socs5 in the liver, exhibiting regulation by a tissue-specific non-dominant trans-acting eQTL identified in the liver. For the methylation plots shown in (c), (d), and (e), bold lines represent coverage-weighted mean methylation of the respective group and CpG sites included in the final analysis are denoted by tick marks on the x-axis.
[0023] FIGs 4A-4J illustrate dominant tram-acting meQTLs, transvection, and paramutation. FIG. 4A is a generic representation of methylation in the context of a dominant trans-acting meQTL, transvection, or paramutation in inbred CC019 and CC037 samples and their Fl crosses. FIG. 4B illustrates dominant trans-acting meQTL / transvection / paramutation methylation pattern identified in both the liver (left) and muscle (right) over Vps37c. FIG.4C is a generic representation of methylation in the context of a dominant trans-acting meQTL in the inbred (left), Fl (middle), and F2 (right) generations. FIG. 4D is a generic representation of methylation in the context of transvection in the inbred (left), Fl (middle), and F2 (right) generations. FIG. 4E is a generic representation of methylation in the context of paramutation in the inbred (left), Fl (middle), and F2 (right) generations. Methylation patterns in (FIGs 4C-4E) indicate that sequencing of the F2 generation is required to distinguish dominant transacting meQTLs, transvection, and paramutation. FIG.4F illustrates paramutation methylation pattern identified in the liver over Capnll, shown for Fl samples (left) and F2 samples (right).FIG. 4G is a generic representation of methylation in the context of paramutation driven by the presence of a strain-specific IAP element in the inbred (left). Fl (middle), and F2 (right) generations. FIG. 4H illustrates, methylation over the CC037-specific Vps37c IAP shown for the inbred and Fl samples of the liver (left), inbred and Fl samples of the muscle (center), and F2 samples of the liver (left). For the liver F2s, only those samples / alleles with an average coverage of at least lx over the region shown are included. The IAP is highlighted in blue and the original DMR (identified without the inclusion of the IAP in the reference genome) is highlighted in red. As these plots no longer conform to the onginal pseudo-hybrid intermediate coordinate system due to the addition of the IAP, a single anchor coordinate indicating the start of the IAP is shown. FIG. 41 illustrates methylation over the CC037-specific Capnll IAP shown for the inbred and Fl samples of the liver (left), inbred and Fl samples of the muscle (center), and F2 samples of the liver (left). The IAP is highlighted in blue. FIG. 4J illustrates 101626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO total expression of Drapl in the liver, exhibiting a dominant trans -acting eQTL / transvection / paramutation expression pattern and proximal to a dominant trans- acting meQTL / transvection / paramutation methylation pattern in the liver. For the methylation plots show n in FIG. 4B, FIG. 4F, FIG. 4H, and FIG. 4I, bold lines represent coverage-weighted mean methylation of the respective groups and CpG sites included in the final analysis are denoted by tick marks on the x-axis.
[0024] FIGs 5A-5G illustrate emergent epigenetic inheritance patterns. FIG. 5A is a generic representation of methylation in the context of over- and underdominance in inbred CC019 and CC037 samples and their Fl crosses. FIG. 5B is a generic representation of methylation in the context of allele-specific over- and underdominance in inbred CC019 and CC037 samples and their F 1 crosses. FIG.5C is a generic representation of methylation in the context of biallelic dominance in inbred CC019 and CC037 samples and their Fl crosses. FIG.5D illustrates overdominant methylation pattern identified in the liver over Uspl8. FIG. 5E illustrates allele-specific underdominant methylation pattern identified in the liver over Jak3.FIG.5F illustrates biallelic dominant methylation pattern identified in the liver over Ccdc85c.FIG. 5G illustrates allele-specific expression of Ccdc85c in the liver, exhibiting a biallelic dominant expression pattern. The promoter of Ccdc85c overlaps the biallelic dominant methylation pattern shown in (f). For the methylation plots shown in FIG. 5D, FIG. 5E, and FIG. 5F, bold lines represent coverage-weighted mean methylation of the respective groups and CpG sites included in the final analysis are denoted by tick marks on the x-axis.
[0025] FIGs 6A-6F illustrate genomic imprinting. FIG. 6A is a generic representation of methylation in the context of genomic imprinting, i.e. parent-of-origin-specific methylation, in inbred CC019 and CC037 samples and their Fl crosses in both cross directions. FIG. 6B illustrates imprinted methylation pattern identified over the promoter of Slc38a4 in the liver. Inbred samples are shown on both plots and Fl samples have been split by cross direction, CC019xCC037 (top) and CC037xCC019 (bottom). Fig.6C illsutartes parent-of-origin-specific methylation pattern found exclusively within the liver identified over Zswim9, shown in both the liver (left) and muscle (right). Inbred samples are shown on both plots and Fl samples have been split by cross direction, CC019xCC037 (top) and CC037xCC019 (bottom). FIG. 6D illustrates parent-of-origin-specific methylation pattern found exclusively within the muscle identified over Asb4, show n in both the liver (left) and muscle (right). Inbred samples are shown on both plots and Fl samples have been split by cross direction, CC019xCC037 (top) and CC037xCC019 (bottom). FIG. 6E illustrates novel parent-of-origin-specific methylation 111626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO pattern identified over Scn8a in the liver. Inbred samples are shown on both plots and Fl samples have been split by cross direction. CC019xCC037 (top) and CC037xCC019 (bottom).FIG. 6F illustrates allele-specific expression of Sgce in the liver, exhibiting a parent-of-origin-specific expression pattern and overlapping a parent-of-origin-specific methylation pattern. For the methylation plots shown in FIG. 6B, FIG. 6C, FIG. 6D, and FIG. 6E, bold lines represent coverage-weighted mean methylation of the respective groups and CpG sites included in the final analysis are denoted by tick marks on the x-axis.
[0026] FIGs 7A-7H illustrates sex-specific DNA methylation and X chromosomal epigenetic inheritance patterns. FIG.7A is a generic representation of sex-specific methylation male and female inbred samples and Fl crosses. FIG. 7B illustrates sex-specific methylation pattern found exclusively within the liver identified over Aox3. shown in both the liver (left) and muscle (right). FIG. 7C illustrates sex-specific DMR identified in the liver upstream of Rpsl4, shown for inbred and Fl samples (left) and F2 samples (right). FIG.7D illustrates total expression of Aox3 in the liver, split by sex, exhibiting a sex-specific expression pattern and overlapping the sex-specific methylation patern shown in (b). FIG. 7E illustrates generic representation of methylation in the context of skewed X chromosome inactivation in inbred CC019 and CC037 samples and their Fl crosses. FIG. 7F shws skewed XCI methylation patern identified in the liver over the promoter of Efnbl. FIG. 7G illustrates allele-specific expression of Efnbl in the liver, exhibiting skewed XCI and overlapping the skewed XCI methylation pattern shown in FIG. 7B. FIG. 7H illustrates novel parent-of-origin-specific methylation patern identified in the liver over Zfp92. Inbred samples are shown on both plots and Fl samples have been split by cross direction, CC019xCC037 (left) and CC037xCC019 (right). For the methylation plots shown in FIGs 7B, 7C, 7F and 7H, bold lines represent coverage-weighted mean methylation of the respective group and CpG sites included in the final analysis are denoted by tick marks on the x-axis..
[0027] FIG. 8 is a table illustrating autosomal epigenetic inheritance paterns. Definitions, representative paterns, and violations to Mendelian inheritance of each autosomal epigenetic inheritance pattern identified, as well as the number of each patern identified in both the liver and muscle and the number common to both tissues.
[0028] FIG. 9 illustrates the computational pipeline for the analysis of allele-specifi methylation and expressionin crosses of inbred mice. ONT sequencing w as performed on liver and muscle DNA from the inbred and Fl generations as well as liver DNA from the F2 generation of two CC mouse strains, CC019 and CC037. MatchedRNA-seq was also performed 121626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO on liver RNA from the inbred and Fl generations. ONT sequencing data was aligned and phased twice, once using each of the CC reference genomes. Only consensus phasing decisions, inwhich the assignment of each read is consistent independent of the choice of reference genome, were utilized inorder to maximize the accuracy of phased methylation estimates and minimize the introduction of reference bias. Phased methylation data was then mapped into a pseudo-hybrid intermediate genome, comprised of all anchorsand each of the longer respective edges of CC019 or CC037 from the Collaborative Cross Graphical Genome. Analysis of methylation in the pseudo-hybrid intermediate genome enables the retention of strain-specific methylation information in order to maximize the amount of analyzable data without the introduction of additional reference bias. RNA-seq data was analyzed by aligning reads to a diploid transcriptome containingtranscript information for both CC019 and CC037. Allele-specifi methylation and expression data were then used to identify various forms of epigenetic inheritance patterns and subsequently analyze their effect on the expression of associated genes.
[0029] FIG. 10 is a schematic representation of the pipeline.
[0030] FIG. 11 is a schematic representation of the allele-specific epigenome-wide association study (AS-EAWS).DETAILED DESCRIPTION OF THE INVENTION
[0031] The present invention is based on the seminal identification of non-Mendelian epigenetic inheritance patterns including emergent epigenetic patterns, and paramutation
[0032] Before the present compositions and methods are described, it is to be understood that this invention is not limited to particular compositions, methods, and experimental conditions described, as such compositions, methods, and conditions may vary. It is also to be understood that the terminology used herein is for purposes of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only in the appended claims.
[0033] As used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, references to “the method” includes one or more methods, and / or steps of the type described herein which will become apparent to those persons skilled in the art upon reading this disclosure and so forth.131626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0034] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0035] As used herein, the term ‘'about” in association with a numerical value is meant to include any additional numerical value reasonably close to the numerical value indicated. For example, and based on the context, the value can vary up or down by 5-10%. For example, for a value of about 100. means 90 to 110 (or any value between 90 and 110).
[0036] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0037] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the invention, it will be understood that modifications and variations are encompassed within the spirit and scope of the instant disclosure. The preferred methods and materials are now described.
[0038] Mendel’s laws of inheritance sen e as the cornerstone of modem genetics, shaping much of the understanding of the inheritance of monogenic and complex phenotypes. However, epigenetic information has been shown to violate Mendelian inheritance through genomic imprinting, i.e. parent-of-origin-specific DNA methylation patterns. Despite its established potential to violate Mendel’s laws, comparatively little is known about the intergenerational inheritance of DNA methylation patterns. Here, a comprehensive approach was devised to investigate the inheritance of DNA methylation patterns in mice using long-read sequencing (such as long-read nanopore sequencing) in conjunction with a novel computational pipeline designed to analyze allele-specific methylation patterns genome- wide. Utilizing this approach within two distinct tissues, liver and muscle, abundant Mendelian patterns of epigenetic inheritance, primarily composed of cis-acting methylation quantitative trait loci (meQTLs) and comprising -93% of the identified autosomal epigenetic inheritance patterns were observed. Additionally, many novel examples of non-Mendelian forms of epigenetic inheritance were identified. These include 54 examples of emergent epigenetic inheritance patterns, in which at least one allele of the heterozygous offspring exhibits a methylation pattern not observed in either parent. The first instance of intergenerational paramutation identified in anon-transgenic mammalian genome was also found.141626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0039] Paramutation is a non-Mendelian epigenetic inheritance pattern in which the methylation pattern of a paramutable allele is altered due to the presence of its homologous paramutant allele in a manner which is subsequently heritable independent of the paramutant allele and other trans-acting genetic factors. Paramutation was confirmed in the F2 generation over the gene Capnl 1, a regulator of calcium-dependent signaling in spermatocytes during the late stages of meiosis, and is highly likely over the gene Vps37c. Intriguingly, both of these genes contain strain-specific transposable elements of the intracistemal A-type particle family, suggesting a potential mechanism for intergenerational paramutation. At least four new autosomal imprinted genes were also discovered, including Scn8a, Pcdhb4, and Fry in the liver and Socs5 in the muscle; a new imprinted gene on the X chromosome in the liver, Zfp92; and 305 instances of sex-specific methylation in the liver. Overall, at least 522 autosomal non-Mendelian epigenetic inheritance patterns, representing ~7% of the autosomal epigenetic inheritance patterns identified were identified. These results have profound implications for the understanding of monogenic disorders as well as complex traits analyzed by genome-wide association studies performed in the absence of concurrent epigenetic analyses.
[0040] Definitions
[0041] Allele-specific overdominance'. An emergent epigenetic inheritance pattern in which there is an increase in DNA methylation on one allele of the offspring relative to their parents.
[0042] Allele-specific underdominance'. An emergent epigenetic inheritance pattern in which there is a decrease in DNA methylation on one allele of the offspring relative to their parents.
[0043] Biallelic dominance'. An emergent epigenetic inheritance pattern in which there is an increase in DNA methylation on one allele of the offspring relative to their parents and a decrease in DNA methylation on the other allele.
[0044] Bipolar dominance'. An emergent epigenetic inheritance pattern in which there is an increase in DNA methylation in one cross direction of the offspring relative to their parents and a decrease in DNA methylation in the other cross direction. Bipolar dominance likely involves a combination of genomic imprinting and additional complex regulatory factors.
[0045] Cis-acting: Factors which can only exert their influence over sites located on the same allele and in close genomic proximity to them.
[0046] Cis-acting meQTL An meQTL which acts in cis.151626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0047] Codominant trans-acting meQTL: A non-dominant trans-acting meQTL in which the lack of a dominant variant leads to two distinct epi-alleles within each genetic allele, one matching the methylation pattern of each meQTL.
[0048] Dominant trans-acting meQTL: A trans-acting meQTL for which one of the two alleles dominates over the other, producing a methylation pattern on both alleles which matches that of the dominant meQTL. Methylation patterns in the F2 generation are determined by the zygosity of the sample over the trans-acting factor.
[0049] Emergent epigenetic inheritance pattern'. Epigenetic inheritance patterns in which at least one allele of the heterozygous offspring exhibits a novel DNA methylation pattern which is not found in either parental strain. Emergent patterns must involve a trans-acting factor, though they may also involve additional cis-acting factors.
[0050] Epi-allele: The presence of two (or more) distinct patterns of DNA methylation at a particular locus. In the literature, this often refers to epigenetic variation between individuals. Here, this term was primarily used to refer to epigenetic variation within individuals, particularly within the same genetic allele.
[0051] Epigenetics: All factors within a cell which are heritable across cell division, other than the DNA sequence itself. Epigenetics includes DNA methylation, histone post-translational modifications, chromatin accessibility7, and broader 3D chromatin structure.
[0052] Epigenetic inheritance pattern: The inheritance pattern of a particular epigenetic trait between parents and their offspring. Here, this term was used to refer specifically to the intergenerational inheritance of the epigenetic trait.
[0053] Fl generation: The generation made by crossing two inbred strains. These individuals are heterozygous throughout their entire genome, as meiotic recombination and independent assortment of chromosomes occur across two identical alleles during meiosis in the inbred generation.
[0054] F2 generation: The generation made by crossing two members of the Fl generation. These individuals will have regions of their genome which are homozy gous for each strain and regions which are heterozygous due to meiotic recombination and independent assortment of chromosomes which occurs across the two heterozygous alleles during meiosis in the Fl generation. Each individual within this generation will be genetically distinct from one another.
[0055] Genomic imprinting: Parent-of-origin-specific DNA methylation, i.e. distinct DNA methylation patterns present on the maternal and paternal alleles. Genomic imprinting can be further influenced by cis- or trans-acting factors.161626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0056] Genomic proximity. The linear distance (at the level of DNA sequence) between two genomic regions located on the same chromosome.
[0057] Genomic region'. A particular segment of DNA within the genome.
[0058] Identical by descent (IBD): Genomic regions in which there exists no genetic divergence between the founder strains being considered. Conventionally, these have no clear size definition. Here, they are defined as regions which cannot be phased using read lengths of 25kb.
[0059] Incomplete dominant trans-acting meQTL: A non-dominant trans-acting meQTL in which the lack of a dominant variant leads to an intermediate level of methylation across both alleles of heterozygous offspring.
[0060] Methylation quantitative trait locus (meQTL): A genetic variant which is capable of influencing DNA methylation over particular CpG sites in the genome.
[0061] Non-dominant trans-acting meQTL. A trans-acting meQTL for which neither of the two alleles fully dominates over the other, producing a methylation pattern on both alleles which is a mixture of the patterns associated with each meQTL in heterozygous offspring. This pattern includes both incomplete dominant and codominant trans-acting meQTLs.
[0062] Over dominance: An emergent epigenetic inheritance pattern in which there is an increase in DNA methylation on both alleles of the offspring relative to their parents.
[0063] Paramutation: The heritable alteration of DNA methylation of a paramutable allele due to the presence of its homologous paramutant allele. Paramutation can be caused by cis-or trans-acting factors (i.e. initiated by trans vecti on or a dominant trans-acting meQTL, respectively) but are ultimately heritable in subsequent generations independent of the dominant (paramutant) allele. Methylation patterns in the F2 generation match the altered methylation state.
[0064] Polar overdominance: An emergent epigenetic inheritance pattern in which there is an increase in DNA methylation in one cross direction of the offspring relative to their parents. Polar overdominance likely involves a combination of genomic imprinting and additional complex regulatory’ factors.
[0065] Polar under dominance: An emergent epigenetic inheritance pattern in which there is a decrease in DNA methylation in one cross direction of the offspring relative to their parents. Polar underdominance likely involves a combination of genomic imprinting and additional complex regulatory factors.171626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0066] Sex-specific DNA methylation'. DNA methylation on both alleles is driven by the sex of the individual.
[0067] Skewed X chromosome inactivation'. Preferential inactivation of one X chromosomal allele in heterozy gous offspring.
[0068] Trans-acting'. Factors which can exert their influence over sites located anywhere in the genome and on either allele.
[0069] Trans-acting meQTL'. An meQTL which acts in trans. As each of the two alleles’ meQTLs can influence methylation over the same CpG sites on either allele, this introduces competition between the genetic variants to establish their associated methylation pattern, leading to several distinct forms of trans-acting meQTLs.
[0070] Transvection: The alteration of DNA methylation of an allele by its homolog, i.e. one dominant parental allele alters the methylation state of the other allele to match its own. Transvection events exhibit characteristics of both cis- and trans-acting mechanisms insofar as they can influence DNA methylation over CpG sites located on the opposite allele, however they can only do this at the homologous site in the genome and, as such, are bound to influence CpGs in close genomic proximity to the genetic variant. Methylation patterns in the F2 generation are determined by the zy gosity of the sample over the local region containing the DMR.
[0071] Under dominance: An emergent epigenetic inheritance pattern in which there is a decrease in DNA methylation on both alleles of the offspring relative to their parents.
[0072] In one embodiment, the present invention provides a method for screening a subject for susceptibility to a disease or condition including: detecting a methylation status of a plurality of nucleic acid sequences in a sample from the subject, wherein detecting the methylation status comprises: (i) analyzing inheritance of epigenetic patterns in a nucleic acid sample from the subject, (ii) identifying the human subject as having an increased risk of developing the disease or condition.
[0073] In one aspect, the method further includes detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences. In another aspect, combining the genetic variations and the methylation status of the plurality of nucleic acid sequences includes using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genome-wide association studies (GW AS) and epigenome- wide association studies (EWAS). In one aspect,181626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO analyzing inheritance of epigenetic patterns includes identifying identifying Mendelian and non-Mendelian patterns of epigenetic inheritance. In some aspects, non-Mendelian patterns of epigenetic inheritance comprises non-dominant transacting meQTLs; imprinted genes; sexspecific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation. In other aspects, analyzing includes using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome-wide in a nucleic acid sample from the subject. In some aspects, the long-read sequencing includes genetic polymorphisms and methylation status. In various aspects, the long-read sequencing includes long-read nanopore sequencing. In one aspect, wherein the subject is a mammal. In some aspects, the subject is human. In other aspects, the mammal is a cattle, a sheep, or a horse.
[0074] In another embodiment, the invention provides a method for generating a prognosis of a subject's susceptibility to a disease or condition including: (i) analyzing inheritance of epigenetic patterns, (ii) identifying the human subject as having an increased risk of developing the disease or condition, wherein non-Mendelian epigenetic inheritance comprises emergent epigenetic patterns, and paramutation, (iii) storing Mendelian and non-Mendelian epigenetic inheritance data in a database that includes a set of information related to said subject; (iv) correlating the data with an association between the data and susceptibility to the disease or condition in the database; (v) generating a prognosis of the subject's susceptibility to the disease or condition; and (vi) communicating the prognosis of susceptibility to a medical practitioner.
[0075] In one aspect, the set of information related to said subject includes family medical history7, diet, exercise and medical history7of said subject. In another aspect, analyzing inheritance of epigenetic patterns includes identify ing Mendelian and non-Mendelian patterns of epigenetic inheritance. In some aspects, non-Mendelian epigenetic inheritance includes nondominant / ram- acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation. In other aspects, analyzing includes using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome- wide in a nucleic acid sample from the subject. In one aspect, the method further includes detecting genetic variations of a plurality7of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences. In one aspect, combining the genetic variations and the methylation status of the plurality of nucleic acid sequences includes using Allele-Specific 191626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genome-wide association studies (GWAS) and epigenome-wide association studies (EWAS). In some aspects, the long-read sequencing includes genetic polymorphisms and methylation status. In various aspects, the long-read sequencing includes long-read nanopore sequencing. In one aspect, wherein the subject is a mammal. In some aspects, the subject is human. In other aspects, the mammal is a cattle, a sheep, or a horse.
[0076] In an additional embodiment, the invention provides a system for screening a subject for susceptibility to a disease or condition including: one or more processors; and a program having programming instructions stored thereon, which, when executed by the one or more processors, causes the system to perform operations including: detecting a methylation status of a plurality of nucleic acid sequences in a sample from the subject, wherein detecting the methylation status comprises: (i) analyzing inheritance of epigenetic patterns in a nucleic acid sample from the subject, and (ii) identifying the subject as having an increased risk of developing the disease or condition.
[0077] In one aspect, the operations further include: detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences. In another aspect, combining the genetic variations and the methylation status of the plurality of nucleic acid sequences includes using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genomewide association studies (GWAS) and epigenome-wide association studies (EWAS). In one aspect analyzing inheritance of epigenetic patterns includes: identifying Mendelian and non-Mendelian patterns of epigenetic inheritance. In some aspect, non-Mendelian epigenetic inheritance includes non-dominant trans-acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation. In other aspects, analyzing the inheritance of epigenetic patterns in the nucleic acid sample from the subject includes: using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome-wide in a nucleic acid sample from the subject. In some aspects, the long-read sequencing includes genetic polymorphisms and methylation status. In various aspects, the long-read sequencing includes long-read nanopore sequencing.201626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO In one aspect, wherein the subject is a mammal. In some aspects, the subject is human. In other aspects, the mammal is a cattle, a sheep, or a horse.
[0078] In a further embodiment, the invention provides a system for generating a prognosis of a subject's susceptibility to a disease or condition including: one or more processors; and a program having programming instructions stored thereon, which, when executed by the one or more processors, causes the system to perform operations including (i) analyzing inheritance of epigenetic patterns, (ii) identifying the subject as having an increased risk of developing the disease or condition, wherein non-Mendelian epigenetic inheritance includes non-dominant / ra / ?.s -acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation, (lii) storing Mendelian and non-Mendelian epigenetic inheritance data in a database that includes a set of information related to said subject, (iv) correlating the data with an association between the data and susceptibility to the disease or condition in the database, (v) generating a prognosis of the subject's susceptibility to the disease or condition, and (vi) causing display of an indication of the prognosis of susceptibility to a medical practitioner.
[0079] In one aspect, the set of information related to said subject includes family medical history, diet, exercise and medical history7of said subject. In another aspect, analyzing the inheritance of epigenetic patterns includes identifying Mendelian and non-Mendelian patterns of epigenetic inheritance s. In one aspect, non-Mendelian epigenetic inheritance includes nondominant / / Y / -actmg meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation. In another aspect, analyzing the inheritance of epigenetic patterns includes: using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome-wide in a nucleic acid sample from the subj ect. In one aspect, the operations further include: detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences. In one aspect, combining the genetic variations and the methylation status of the plurality of nucleic acid sequences includes using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genome-wide association studies (GWAS) and epigenome-wide association studies (EWAS). In some aspects, the long-read sequencing includes genetic polymorphisms and methylation status. In various aspects,211626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO the long-read sequencing includes long-read nanopore sequencing. In one aspect, wherein the subject is a mammal. In some aspects, the subject is human. In other aspects, the mammal is a cattle, a sheep, or a horse.
[0080] In one embodiment, the invention provides a system for analyzing allele-specific DNA methylation patterns in in crosses of inter-strain and inter-specific crosses of organisms including one or more processors; and a memory having programming instructions stored thereon, which, when executed by the one or more processors, causes the system to perform operations including (a) generating a pairwise graphical genome between two reference genomes, (b) creating a pseudo-hybrid intermediate genome, wherein the pseudo-hybrid intermediate genome is a unified reference coordinate system including strain / allele-specific CpGs from both reference genomes; (c) aligning the pseudo-hybrid intermediate genome and the corresponding edges of the graphical genome for coordinate mapping; (d) identifying heterozygous genetic variants between the two provided reference genomes; (e) extracting consensus haplotyping decisions; and (f) mapping methylation coordinate from the two reference genomes into the ‘pseudo-hybrid intermediate genome, thereby analyzing allelespecific DNA methylation patterns.
[0081] In one aspect, generating the pairwise graphical genome includes dividing the two input reference genomes into smaller, corresponding pieces. In another aspect, the operations further include processing long-read sequencing data. In various aspect, the long-read sequencing data includes Oxford Nanopore Technologies (ONT) sequencing data. In one aspect, analyzing allele-specific DNA methylation patterns includes identifying intergenerational epigenetic inheritance patterns. In another aspect, the intergenerational epigenetic inheritance patterns include intergenerational paramutation. In some aspects, the inheritance patterns violate the laws of Mendelian inheritance. In one aspect, the organism is a mammal. In some aspects, the mammal is a human. In other aspects, the mammal is a cattle, a sheep, or a horse. In another aspect, the organism is a model organism. In one aspect, the references genomes are non-engineered mammalian genomes.
[0082] Presented below are examples discussing the analytical methods described herein contemplated for the discussed applications. The following examples are provided to further illustrate the embodiments of the present invention but are not intended to limit the scope of the invention. While they are ty pical of those that might be used, other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.221626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0083] In a further embodiment, the invention provides a method for generating a prognosis of a subject's susceptibility to a disease or condition including (i) analyzing non-Mendelian patterns of epigenetic inheritance, (ii) identifying the human subject as having an increased risk of developing the disease or condition, (iii) storing the non-Mendelian epigenetic inheritance data in a database that includes a set of information related to said subject; (iv) correlating the data with an association between the data and susceptibility to the disease or condition in the database; (v) generating a prognosis of the subject's susceptibility to the disease or condition; and (vi) communicating the prognosis of susceptibility to a medical practitioner.
[0084] In one aspect, the set of information related to said subject includes family medical history, diet, exercise and medical history of said subject. In another aspect, the non-Mendelian epigenetic inheritance includes non-dominant / ram-acting meQTLs; imprinted genes; sexspecific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation. In some aspects, analyzing comprises using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome- wide in a nucleic acid sample from the subject. In one aspect, the method further includes detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences. In some aspects, combining the genetic variations and the methylation status of the plurality’ of nucleic acid sequences comprises using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genomewide association studies (GWAS) and epigenome-wide association studies (EWAS). In another aspect, the long-read sequencing includes genetic polymorphisms and methylation status. In some aspects, the long-read sequencing includes long-read nanopore sequencing. In one aspect, the subject does not present the phenotype of a disease or disorder but is known for having inherited a mutation associated with a Mendelian disease or disorder from a family member affected with the Mendelian disease or disorder. In another aspect, the subject present the phenotype of a disease or disorder but is known for not having inherited a mutation associated with a Mendelian disease or disorder from a family member affected with the Mendelian disease or disorder. In some aspects, the subject is a mammal. In various aspects, the subject is human. In other aspects, the mammal is a cattle, a sheep, or a horse.231626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO EXAMPLES EXAMPLE 1EXPERIMENTAL PROCEDURES
[0085] Animal, housing, and genotyping.
[0086] Collaborative Cross lines CC019 / TauUnc and CC037 / TauUnc were sourced from the Systems Genetics Core Facility at the University’ of North Carolina at Chapel Hill and underwent breeding and maintenance at Texas A& M University. Fl mice were generated by crossing CC019 / TauUnc females with CC037 / TauUnc males (CC019 x CC037) and CC037 / TauUnc females with CC019 / TauUnc males (CC037 x CC019). The Fl mice were subsequently intercrossed to produce three distinct F2 populations: [(CC019 x CC037) x (CC019 x CC037)], [(CC019x CC037) x (CC037 x CC019)], and [(CC037 x CC019) x (CC037 x CC019)]. These Collaborative Cross strains do not carry the agouti viable yellow (A^) allele.
[0087] All mice, including parental lines, Fl progeny, and the F2 offspring, were housed at a temperature of 22 ± 2 °C under 12 h light / 12 h dark cycle with unrestricted access to food and water, and were euthanized at approximately 4 months (±1 month) of age. Liver biopsies were obtained from mice post-euthanasia and were used for genotyping the parental CC lines, Fl, and F2 populations with the MiniMUGA genotyping array (Neogen, MI). The remaining liver tissue was flash-frozen in liquid nitrogen prior to processing for ONT and RNA sequencing (RNA-seq). Right femoral muscle was also collected from mice post-euthanasia, rinsed with lx PBS, and flash-frozen in liquid nitrogen prior to processing for ONT sequencing. All experiments were performed in accordance and approval by Texas A& M University' Institution of Animal Care and Use Committee.
[0088] After genotyping, a minimal set of F2s was chosen for targeted ONT sequencing which satisfied the following two conditions: (1) for each of the candidate dominant transacting meQTL / transvection / paramutation regions to be targeted there are more than five heterozygous samples and three homozygous samples for each parental allele, CC019 and CC037; and (2) for each region in the genome there is at least one sample which is homozygous for each parental strain, i.e. there are no contiguous three SNPs in the genome which harbor at least one copy of the same genetic variants in all selected mice. These conditions ensure sufficient genetic diversity’ to understand the molecular origins of the various epigenetic inheritance patterns analyzed. Furthermore, a dominant / rans-acting meQTL needs only one copy of the dominant allele to establish a particular methylation pattern, and this variant can 241626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO exist anywhere in the genome. Thus, it is possible that, given a chosen set of F2s, the dominant / ra s -acting meQTL is present in all of the samples. This would create a scenario in which, despite being caused by a dominant / ram-acting meQTL, the dominant methylation pattern is present in all of the F2s, similar to what would be expected in the case of paramutation. As such, the second condition ensures that this scenario does not occur, allowing for the distinction between dominant / ra -acting meQTLs and paramutation events.
[0089] ONT sequencing was performed on liver DNA from 6 CC019 inbred samples (3 males and 3 females), 9 CC037 inbred samples (4 males and 5 females), 9 CC019xCC037 Fl crosses (4 males and 5 females), 13 CC037xCC019 Fl crosses (6 males and 7 females), and 19 F2 crosses (12 males and 7 females) as well as muscle DNA from 5 CC019 inbred samples (2 males and 3 females, 6 CC037 inbred samples (3 males and 3 females), 6 CC019xCC037 Fl crosses (3 males and 3 females), and 6 CC037xCC019 Fl crosses (3 males and 3 females). One female CC019xCC037 Fl liver sample was removed from subsequent methylation and expression analyses as phasing of the X chromosome revealed only one parental allele, indicating monosomy of the X chromosome (X0). Liver and muscle were chosen due to their relatively high degree of cellular homogeneity, with hepatocytes representing -60% of the cells and -80% of the total mass of the liver and type I and II myofibers accounting for -70% of the nuclei in the muscle.
[0090] Matched RNA-seq was performed on liver RNA from 6 CC019 inbred (3 males and 3 females), 9 CC037 inbred (4 males and 5 females), 9 CC019xCC037 Fl crosses (4 males and 5 females), and 9 CC037xCC019 Fl crosses (4 males and 5 females). One male CC019xCC037 was removed due to incorrect sequencing read depth (-360M compared to 50M targeted), one female CC019xCC037 Fl sample was removed from subsequent analysis as an outlier (total expression of the outlier was more than 3 standard deviations away from the mean in 6,927 / 12,894 expressed genes), and one male CC037xCC019 Fl sample was removed due to poor mapping to the diploid transcriptome (68.8% mapping rate).
[0091] High molecular weight DNA extraction.
[0092] High molecular weight (HMW) DNA was extracted from ~20mg of flash-frozen median liver lobe using the Nanobind Tissue Kit (PacBio SKU 102-302-100) and ~35mg of flash-frozen right femoral muscle using the Nanobind PanDNA Kit (PacBio SKU 103-260-000) following to the manufacturer's instructions. Homogenization was performed using the TissueRuptor II (Qiagen Cat. 9002755) at maximum speed for 10s.
[0093] DNA library preparation and ONT sequencing.251626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0094] Whole-genome ONT sequencing of the liver samples was performed in two batches. For batch 1 of the liver samples, 2µg of HMW DNA diluted to 70µL with nuclease-free water was sheared to ~10kb using the Megaruptor 2 (Diagenode Cat. B06010002) with an additional AMPure XP (Beckman Coulter Cat. A63881) cleanup. Samples were eluted into a final volume of 50µL of nuclease-free water. All remaining DNA after shearing, cleanup, and quantification was used as input for library preparation. Libraries were prepared using the SQK-LSK109 Ligation Sequencing Kit (ONT) following the manufacturer’s instructions. ONT sequencing was performed using a PromethION 24 and R9.4.1 flow cells (ONT Cat. FLO-PRO002). Sequencing was run for 72 hours. Batch 1 contains samples from all relevant groups except the 19x37 Fl cross direction.
[0095] For batch 2 of the liver samples, 3.25µg of HMW DNA at a concentration of 50 ng / µL was sheared using the Megaruptor 3 (Diagenode Cat. B06010003) at speed 29 resulting in N50s in the range of 20-30kb. All remaining DNA after shearing and quantification was used as input for library preparation. Libraries were prepared using the SQK-LSK110 Ligation Sequencing Kit (ONT) following the manufacturer’s instructions with the following changes which improve the recovery of long reads: SPRIselect beads (Beckman Coulter Cat. B23318) were used instead of AMPure XP beads; the duration of all incubation and elution steps performed at RT or 37°C were tripled; during DNA repair and end-prep, thermocycler incubation w as performed at 20°C for 10 minutes and then 65°C for 10 minutes; and LFB w as used during adapter ligation and clean-up. ONT sequencing was performed using a PromethION 24 and R9.4.1 flow cells. Sequencing was stopped manually at the end of the life of the flow cells, which were washed and reloaded once or twice depending on the amount of library prepared, flushing when -20% of total pores were remaining. To reach a coverage sufficient for a phased analysis of DNA methylation, ONT sequencing data for the same samples were pooled across different flow cells, where relevant. Samples within each sequencing run were randomized while ensuring an even distribution of each group. Additionally, as liver ONT sequencing was performed in two distinct batches with differing read lengths and ligation sequencing kit chemistries, a check for potential batch-specific DMRs was performed, as discussed in the DNA methylation analysis section of the Methods, below.
[0096] For the muscle samples, library preparation and sequencing were performed following the same protocol as batch 2 of the liver samples with the following changes: libraries were prepared using the SQK-LSK114 Ligation Sequencing Kit (ONT) following the261626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO manufacturer’s instructions with the same long read recover}’ adjustments as used in batch 2 of the liver samples, and ONT sequencing was performed using R10.4.1 flow cells.
[0097] For the targeted ONT sequencing of the F2 samples, two BED files were created containing the chromosome, start, and end coordinates of each of the selected DMRs to be targeted, one in the CC019 reference coordinate system and the other in the CC037 reference coordinate system. Included in the target panel were three dominant traws-acting meQTL / transvection / paramutation DMRs, 20 c / .s-acting meQTL DMRs, 5 non-dominant irans-acimg meQTL DMRs, 20 sex-specific DMRs, 5 imprinted DMRs, 10 skewed XCI DMRs, and 21 regions which were included based on a preliminary analysis but are no longer of interest and have thus been removed from subsequent analysis. An additional 500kb of padding was added to the start and end coordinates of each of the dominant trans-acXmg meQTL / transvection / paramutation regions to be targeted and lOOkb of padding was added to the start and end coordinates of the remaining target regions. Strain-specific BED files containing these coordinates were then used to subset the CC019 and CC037 reference genome FASTA files to just the sequence contained within the padded DMRs using the bedtools (v2.30.0) getfasta function. These subset FASTA files were then combined to make a single multi-FASTA with the sequence of each of the padded target regions from both the CC019 and CC037 reference genome in order to minimize sequencing bias which may have been introduced into adaptive sampling by the use of just one reference genomes.
[0098] Library preparation and sequencing of the F2 samples were performed as described in batch 2 of the liver sequencing above with the following exception: adaptive sampling was activated for sequencing and set to enrich for the multi-FASTA file containing the desired target regions.
[0099] RNA extraction.
[0100] RNA was extracted from ~10-20mg of flash-frozen median liver lobe using the RNeasy Mini kit (Qiagen Cat. 74104) following the manufacturer’s instructions with on-column DNase digestion. Disruption and homogenization were performed using the TissueRuptor II (Qiagen Cat. 9002755) at maximum speed for 10s.
[0101] RNA library preparation and sequencing.
[0102] Strand-specific mRNA libraries were generated using the NEBNext Ultra II Directional RNA Library Prep Kit for Illumina (New England BioLabs & E7760), and mRNA isolation was performed using the Poly(A) mRNA Magnetic Isolation Module (New England BioLabs #E7490). Library preparation followed the manufacturer’s protocol (Version 2.2,271626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO05 / 19). A 1 µg input was used, and samples were fragmented for 15 minutes to achieve an RNA insert size of approximately 200 bp. The subsequent PCR cycling conditions were as follows: 98°C for 30 seconds followed by 8 cycles of 98°C for 10 seconds and 65°C for 75 seconds, with a final step at 65°C for 5 minutes. The stranded mRNA libraries were sequenced on an Illumina NovaSeq 6000 using vl.5 chemistry with an S4 flow cell, generating 150 bp paired-end dual-indexed reads. The mRNA sequencing depth was 50 million reads per sample. RNA was sequenced in two batches, with all inbred samples included in one and all Fl samples included in the other. Six inbred samples were included in both batches to allow for an analysis of batch effects. Principal component analysis (PCA) plots of these six samples across suggested little-to-no effect of batch on expression (see expression analysis ). A single CC019xCC037 sample was sequenced to a read depth of -363.6 million reads (50M targeted) and was excluded from subsequent expression analysis.EXAMPLE 2COMPUTATIONAL PROCEDURES
[0103] Collaborative Cross Graphical Genomes.
[0104] Genome sequence and anchor information for CC019 and CC037 were extracted from the Collaborative Cross Graphical Genome (CCGG; v2.0) utilizing its API. The genome sequence was written into a separate FASTA file for each strain and the anchor information for each strain was written into a separate CSV with the following columns: source anchor, chromosome, outgoing edge, founder on that edge, source anchor coordinate - GRCm38, source anchor coordinate - GRCm39, and source anchor coordinate - respective strain (i.e. CC019 for the CC019 anchors or CC037 for the CC037 anchors).
[0105] Overall summary of methylation processing and analysis.
[0106] Briefly, ONT sequencing data for all samples was aligned and phased to both the CC019 and CC037 genomes and only reads with consensus phasing decisions, in which the allelic assignment was consistent independent of the reference genome, were utilized. Allelespecific methylation information was subsequently extracted from these reads and converted into a common coordinate system for analysis. DMR finding was then run using data from the inbred and Fl generations to identify regions of the genome in which DNA methylation exhibits at least one of the 12 intergenerational epigenetic inheritance patterns analyzed here. DMR finding was performed separately for liver and muscle, and separately within each tissue281626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO for the autosomes, X chromosomes of female samples, and maternal X chromosomal alleles for male samples. After identifying this set of genome- wide significant DMRs, the DMRs were subsequently filtered for quality control and categorized to best match them to an inheritance pattern utilizing methylation differences between the relevant groups. A subset of these DMRs were targeted for sequencing within the F2 generation to further validate / determine the effects of cis and tram -acting genetic variants on these methylation patterns.
[0107] ONT sequencing processing.
[0108] As ONT sequencing for the liver and muscle samples were performed using different versions of flow cells and library preparation kits, they are processed using different pipelines. Liver samples, sequenced using R9.4.1 flow cells and LSK109 / LSK110 Ligation Sequencing Kits, were basecalled using Guppy and methylation was called using Nanopolish. Canonical and modified basecalling for muscle samples, sequenced using R10.4.1 flow cells, were performed using Dorado and methylation information was subsequently extracted using modkit. Of note, Nanopolish accounts for supplementary alignments when performing methylation calling whereas modkit excludes these alignments from the methylation extraction. The use of multiple techniques for basecalling and methylation calling across these tissues allows for a limited degree of validation that the identified patterns do not arise due to technical artifacts introduced by these processes. Furthermore, methylation calling (liver) and extraction (muscle) is performed utilizing the reference genome of the strain / allele being considered.
[0109] Liver samples were processed using the following pipeline. Basecalling was performed using Guppy (v6.1.2) with the following parameters: — num callers 8 — gpu_runners_per_device 8 -chunks_per_runner 1024 -chunk_size 1000 -config dna r9.4.1 450bps sup_prom. cfg, implementing the default Guppy minimum q-score of 10 for super high accuracy basecalling models. The SeqKit (v2.2.0) split2 command with the parameter -p 10 was used to split the FASTQ files passing the minimum q-score cutoff into 10 parts for parallelization of subsequent alignment, phasing, and methylation calling. Alignment was performed using MiniMap2 (v2.24) with the following parameters: -a -x map-ont. Alignment of each FASTQ file was performed twice, once relative to the CC019 genome and once relative to the CC037 genome. Heterozygous genetic variants were identified using the nucmer command from MUMmer4 (v4.0.0rcl), converted to a delta file with only SNPs using the show-snps function with the following parameters: -Clr -I -T, and subsequently converted to a VCF file using the my-mummer-2-vcf.py script with the following parameters: -n —output- 291626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO header. Heterozygous variants were identified once with CC019 as the reference genome and once with CC037 as the reference genome. Phasing information was then added the VCF file to define all heterozygous variants originating from the same parental genome as in phase with one another in the heterozygous Fl genome. Haplotype phasing of aligned reads was performed using the WhatsHap (vl.6) haplotag function with the following parameter: — ignore-read-groups. Haplotype phasing was performed twice for each BAM file, once using each of the VCF files produced using MUMmer4. Consensus haplotype phasing information was extracted by identifying reads for which haplotype assignment was the same using both of the VCF files. Aligned BAM files were split into CC019-phased and CC037-phased files based on the consensus haplotype phasing information, with CC019-phased BAM file information originating from the alignment using CC019 as the reference and with the CC037-phased BAM file information originating from the alignment using CC037 as the reference. Prior to methylation calling, the FASTQ files were indexed using the Nanopolish (v0.14.0) index function using the raw- FAST5 files and the sequencing summary- file produced by Guppy basecalling. Methylation was called with the Nanopolish call-methylation function once for each set of phased reads using the indexed FASTQ files, phased BAM files, and strain-specific genomes as inputs. Methylation calls were converted to methylation frequency values for each CpG using the Nanopolish calculate_methylation_frequency.py script with the following parameters: -c 2.5 -s. As Nanopolish groups CpGs within lObp of each other together during methylation calling, this methylation conversion splits CpGs which are grouped together. Samtools (vl.7) was used for FASTQ and BAM indexing and sorting.
[0110] Muscle samples were processed as described in the above pipeline, with the following exceptions: (1) basecalling and methylation calling were performed using Dorado (v0.8.3) using the model dna_r10.4.1_e8.2_400bps_sup <7,v5.0.0 with the following parameters: -modified-bases 5mCG_5hmCG — min-qscore 10, (2) the resulting unaligned, modified BAM file was then converted to FASTQ format using the samtools (vl.19.2) fastq command with the following parameters to preserve the methylation calls: -T MM, ML, (3) alignment was performed using MiniMap2 (v2.28) with the following parameters: -a -y -x map-ont, (4) reads with a MAPQ quality- score of less than 20 w ere filtered out after alignment using the samtools view command with the following parameters: -bSq 20 (note that a MAPQ cutoff of 20 is also used by Nanopolish), (5) haploty pe phasing was performed using the WhatsHap (v2.3) haplotag function with the same VCF files produced above using MUMmer4 and the following parameters: — ignore-read-groups -skip-missing-contigs, and (6) methylation data 301626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WOwas extracted from the final BAM files using the modkit (vO.4.1) pileup command with the following parameters: --cpg --combine-strands -mod-thresholds m:0.8 -filter-threshold C:0.8 and given the appropriate reference genome. Samtools was also used for FASTQ and BAM indexing, sorting, and merging.
[0111] N50 calculations. N50s for each sample were calculated using the Proch-n50 (v1.4.2) tool.
[0112] Pairwise alignment of respective CCGG edges.
[0113] Using the CCGG FASTA and anchor information files, pairwise alignment was performed between all respective edges (those edges bounded by the same anchors) of CC019 and CC037 less than 3kb in length using the pairwise2 function of the Bio python library (v1.78) in order to create edge-specific CIGAR strings to be used in coordinate mapping. This pairwise alignment was also performed between the respective edges of the mmlO reference genome and both CC019 and CC037 to be used for mapping mmlO coordinates into the CC019 and CC037 coordinate systems and vice-versa.
[0114] Coordinate conversion.
[0115] The use of multiple strain-specific reference genomes necessitates the ability to map particular genome feature coordinates between the genomes. Furthermore, as most reference annotations are exclusively found in the standard mouse reference genome coordinates (e.g. mmlO), the ability to map coordinates from mmlO into strain-specific reference genomes is also crucial. Lastly, to further minimize the introduction of reference bias into the data, all methylation analyses were performed in a pseudo-hybrid intermediate genome created from the CC019 and CC037 reference genomes, rather than picking one of these two genomes in which to perform the analysis. This intermediate genome is created by constructing the maximal length sequences from CC019 and CC037 alignment between each pair of anchors. As such, this intermediate genome will have all CC019-specific and CC037-specific insertions represented and any coordinate which exists in either strain-specific genome will necessarily exist in the intermediate genome. As such, there is also a necessity to be able to convert all coordinates to and from this intermediate genome.
[0116] To do so, tools which are capable of mapping coordinates between the CCGG genomes and the intermediate genome were created, as well as between mmlO and the CCGG genomes. Briefly, these mapping tools utilize the edge-specific CIGAR strings created by the aforementioned pairwise alignment of respective edges of mmlO, CC019, and CC037. As the start and end coordinates of each anchor are known for mmlO. CC019, CC037, and the 311626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO intermediate genome, coordinate mapping between any of these coordinate systems can be performed by iterating through each of these CIGAR strings and adding to the anchor position of the respective coordinate system for every insertion and deletion present in the CIGAR string. By doing so, a mapping table can be created with each position in the starting genome and its corollary in the genome to which coordinates are being mapped.
[0117] When mapping reference information (e.g. gene, promoter, and enhancer coordinates, etc.) from mmlO into the intermediate coordinate system, the coordinates are first mapped into the CC019 and / or the CC037 coordinate system prior to being mapped into the intermediate coordinate system. When mapping between mmlO and CC019 / CC037, there will be coordinates which are present in one genome but not in the others. For this reference information, the nearest existing coordinate in the genome being mapped to is used as the corollary site. This method is not used for mapping of CpG coordinates prior to methylation analyses.
[0118] The mapping of coordinates between coordinate systems and all analyses in R are performed using coordinates which are one-based and closed.
[0119] Identification of identical by descent (IBD) regions.
[0120] IBD regions between CC019 and CC037 were defined as those which cannot be phased using 25kb reads. Briefly, the heterozygous SNPs between CC019 and CC037 from MUMmer4 were utilized and regions in which a 25kb read could not overlap two such heterozygous variants extracted. To do so, each SNP was iterated through in the MUMmer4 VCF files and its position analyzed as well as the positions of the three following SNPs relative to each other. For each set of four contiguous SNPs, if the distances betw een SNP 1 and SNP 3, SNP 2 and SNP 4, and SNP 2 and SNP 3 are all >= 25kb and the distance between SNP 1 and SNP 4 is >= 50kb, then the region between SNP 2 and SNP 3 is defined as IBD. For the first four SNPs of each chromosome, if the position of SNP 2 is >= 25kb aw ay from SNP 1 and the position of SNP 2 is <= 25kb away from SNP 3, the region between SNP 1 and SNP 2 is defined as IBD, whereas if the distances betw een SNP 1 and SNP 3, SNP 2 and SNP 4, and SNP 2 and SNP 3 are all >= 25kb and the distance between SNP 1 and SNP 4 is >= 50kb, then the region between SNP 1 and SNP 3 is defined as IBD. Finally, for the final 4 SNPs of each chromosome, if the position of SNP 3 is >= 25kb away from SNP 4 and the position of SNP 2 is <= 25kb aw ay from SNP 3, the region between SNP 3 and SNP 4 is defined as IBD, whereas if the distances between SNP 1 and SNP 3, SNP 2 and SNP 4, and SNP 2 and SNP 3 are all >=321626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 25kb and the distance between SNP 1 and SNP 4 is >= 50kb, then the region between SNP 2 and SNP 4 is defined as IBD. This analysis was performed excluding the Y chromosome.
[0121] These calculations were performed for both CC019 SNPs relative to the CC037 genome and CC037 SNPs relative to the CC019 genome. IBD coordinates were then converted from the CC reference genome coordinate systems into the intermediate coordinate system and merged to create a unified list of IBD regions using the GenomicRanges (v1.54.1) reduce function. These IBD coordinates were then overlapped with the mapped coordinates of common CpGs present in both the CC019 reference genome and the CC037 reference genome to estimate the percentage of the methylome which is IBD.
[0122] Within the CC019 and CC037 genomes, there are 19,258,931 and 19.417,691 CpGs, respectively. Of these, 17,911,438 and 17,969,433 can be mapped into the intermediate genome and 4,773,219 and 4,778,508 overlap IBD regions. Additionally, of the 16,966,618 common CpGs which can be mapped into the intermediate genome, 4,718,665 (27.8%) overlap IBD regions, leaving 11,873,848 autosomal and 374,105 X chromosomal common, non-IBD CpGs which can be mapped into the intermediate genome.
[0123] DNA methylation analysis.
[0124] For the inbred and Fl liver dataset, a single female CC019xCC037 Fl sample was removed from subsequent methylation and expression analyses as phased ONT reads of the X chromosome were assigned to only one allele and the coverage ratio of the X chromosome to the autosomes was roughly half that of other female samples, indicating that this sample was X0 and thus only contains 1 copy of the X chromosome, a phenomenon which has been reported in inbred mice.
[0125] Methylation data was processed in R (v4.3.1) using the bsseq package (v 1.38.0). For liver, coordinate-converted methylation frequency files for each sample were imported into R and converted into BS objects, specifying the chromosome, genomic position mapped into the pseudo-hybrid intermediate genome, Nanopolish called_sites_methylated, and Nanopolish called_sites as chr, pos, M, and C, respectively. For muscle, the read.modkit function from the bsseq package (v1.42.0) was utilized to import methylation data from the modkit output files into R (v4.4.0). The resulting BS objects with combined 5mC and 5hmC calls were utilized for subsequent analyses in R (v4.3.1) using the bsseq package (v1.38.0).
[0126] Methylation values were smoothed using the BSmooth function from bsseq. This function smooths the methylation data of each sample individually to decrease methylation variance and facilitate the identification of differentially methylated regions, which are more 331626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO commonly associated with functionally relevant methylation changes than single-CpG differences. Furthermore, methylation smoothing allows for a better tradeoff between coverage and sample size. BSmooth was run using all CpGs present in each sample, including strainspecific CpGs, with the default parameters except the following: maxGap = 100000.
[0127] Smoothed methylation data was subsequently subset into three groups: autosomal, X chromosome female samples, and X chromosome maternal allele of male samples, as proper analysis of the X chromosome necessitates a more in-depth consideration of sex imbalance due to X chromosome inactivation and because male samples only have maternal X chromosomes. Within each of these groups, the BS objects are subset to just those CpGs which are covered in at least 75% of the CC019 inbred samples and CC019 phased Fl alleles and 75% of the CC037 inbred samples and CC037 phased Fl alleles. This effectively removes strain-specific CpGs from the analysis, although the methylation values of strain-specific CpGs have been used to impute methylation values over strain-agnostic CpGs during smoothing. CpGs for which the smoothed methylation estimate for at least one sample was ‘NA’ were also filtered.
[0128] The following number of CpGs were analyzed for each of the conditions - liver autosomal: 11,501,000; muscle autosomal: 11,903,947; liver chrX female: 341,941; muscle chrX female: 352457; liver chrX male maternal allele: 324,311; muscle chrX male maternal allele: 346,487. Of these, the following number of CpGs overlapped non-IBD regions of the genome (1) liver autosomal: 10.747,060 (90.5%) of autosomal common non-IBD CpGs; (2) muscle autosomal: 10,784,692 (90.8%) of common non-IBD autosomal CpGs; (3) liver chrX female: 324,579 (86.8%) of common non-IBD chrX CpGs; (4) muscle chrX female: 329,943 (88.2%) of common non-IBD chrX CpGs; (5) liver chrX male maternal allele: 311,733 (83.3%) of common non-IBD chrX CpGs; and (6) muscle chrX male maternal allele: 325,500 (87%) of common non-IBD chrX CpGs.
[0129] DMR finding was performed utilizing an F-statistic test, which allows for a genomewide search for multiple distinct DNA methylation patterns using a single test. This identifies regions of the genome where one or more of these contrasts is different from zero while controlling the family -wise error rate. Following identification of DMRs, post-hoc categorization were done as described below. The following contrasts were considered:1- Psex = [inbred females + Fl females either allele] — [inbred males + Fl males either allele]a. This contrast is only included for the autosomal comparison.2. Poverunder_Dom = [inbreds] - [FIs either allele]341626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 3. ^imprinting = [CC019xCC037 Fl CC019 alleles + CC037xCC019 Fl CC037 alleles] - [CC019xCC037 Fl CC037 alleles + CC037xCC019 Fl CC019 alleles]a. This contrast is only included for the autosomal and X chromosome female sample comparisons.4- Pds = [inbred CC019s + Fl CC019 alleles] — [inbred CC037s + Fl CC037 alleles] 3- P trans=[inbred C CO 19s] — [inbred CC037s]6. Pbiaiieiic = [Fl CC019 alleles] - [Fl CC037 alleles]7- PAS emerge = [Fl CC019 alleles] — [inbreds + Fl CC037 alleles]8- PAS emergj.7 = [Fl CC037 alleles] — [inbreds + Fl CC019 alleles]9- Pparamut 19 = [Inbred CC019s] — [inbreds CC037s + FIs either allele]10. PParamut_37=[Inbred CC037s] — [inbreds CC019s + FIs either allele]11- Ppoiar_^x37 = [CC019xCC037 FIs] - [inbreds + CC037xCC019 FIs]a. This contrast is only included for the autosomal and X chromosome female sample comparisons.12. Ppoiar_37xw = [CC037xCC019 FIs] - [inbreds + CC019xCC037 FIs]a. This contrast is only included for the autosomal and X chromosome female sample comparisons.13. Pbipoiar = [CC019xCC037 FIs] - [CC037xCC019 FIs]a. This contrast is only included for the autosomal and X chromosome female sample comparisons.
[0130] A genome-wide search was performed using a single F-statistic on the autosomes, which simultaneously tests for these 13 contrasts. This research was supplemented with an additional search across the X chromosome for female samples (excluding contrast 1) and an additional search across the X chromosome for the maternal allele of male samples (excluding contrasts 1, 3, 11-13). The search was based on the fstat. pipeline function from bsseq, which runs the BSmooth.fstat, smoothSds, computeStat, and dmrFinder functions sequentially prior to running permutation testing using the permuteAll, getNullDistribution_BSmooth.fstat, and getFWER_fstat functions. Briefly, this pipeline identifies DMRs passing the F-statistic cutoff and then performs permutation testing to adjust for multiple testing and control the family -wise error rate (FWER). The fstat. pipeline function was modified to accept a k value as input to be used in the smoothSds and getNullDistribution_BSmooth.fstat functions. Furthermore, the pipeline has been updated to filter DMRs as well as null DMRs from permutation testing based on the following three criteria:1. Outlier samples are driven by low coverage• All hypermethylated samples with Cook's distance > 4 * mean(Cook's distances') or all hypomethylated samples with351626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO Cook’s distance > 4 * mean(Cook' s distances) have an average coverage over the DMR of < 3x.2. Groups with very low or very high coverage• Average coverages of any of the following four groups are < 3x or > 15 Ox:inbred CC019s. inbred CC037s, CC019 alleles of FIs, and CC037 alleles of FIs.3. Disproportionate coverages• \log2(Coverageinbred l9 / Coverageinbrgd 37)\ > 2• \log2(^CoverageF1 19 / CoverageF1 37)\ > 2• \log2(Coveragembred Vil Cover ageFA^)\ > 3 or\log2{Coverageinbred 37ICoverageFV -i7')\ > 3o Note this higher cutoff adjusts for the expected coverage differences between inbreds and phased FIs.
[0131] The pipeline was run using the following parameters: nperm = 1000, cutoff = 4.6², k = 21, maxGap.sd = 10⁸, maxGap.dmr = 1000. The cutoff of 4.6² is the recommended F-statistic bsseq cutoff based upon the default T-statistic cutoff of 4.6.Significant DMRs were defined as those with at least 3 CpGs and an FWER <= 0.1.
[0132] The three lists of output DMRs and smoothed BS objects were then used to calculate smoothed methylation differences between the groups defined in the contrast matrices used for testing as well as eight additional contrasts, all defined below:1. Ase= [inbred females + Fl females either allele] — [inbred males + Fl males either allele]a. This difference is only calculated for the autosomal comparison. 2. overunder _Dom = [inbreds] - [FIs either allele]3. ^imprinting = [CC019xCC037 Fl CC019 alleles + CC037xCC019 Fl CC037 alleles] - [CC019xCC037 Fl CC037 alleles + CC037xCC019 Fl CC019 alleles]a. This difference is only calculated for the autosomal and X chromosome female sample comparisons.4. Acis= [inbred CC019s + Fl CC019 alleles] — [inbred CC037s + Fl CC037 alleles] 5. trans = [inbred CC019s] — [inbred CC037s]6. ^biaiieiic = [Fl CC019 alleles] - [Fl CC037 alleles]7- AS emerge = [Fl CC019 alleles] — [inbreds + Fl CC037 alleles]8- ^AS emer _37=[Fl CC037 alleles] — [inbreds + Fl CC019 alleles]9- paramut 19 = [Inbred CC019s] — [inbreds CC037s + FIs either allele] 10. paramut _37 — [Inbred CC037s] — [inbreds CC019s + FIs either allele] 11. Δpolar_19x37= [CC019xCC037 FIs] - [inbreds + CC037xCC019 FIs]a. This contrast is only included for the autosomal and X chromosome female sample comparisons.12.polar 37zl9= [CC037xCC019 FIs] - [inbreds + CC019xCC037 FIs]361626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO a. This contrast is only included for the autosomal and X chromosome female sample comparisons.13. bipolar = [CC019xCC037 FIs] - [CC037xCC019 FIs]a. This contrast is only included for the autosomal and X chromosome female sample comparisons.14.Inbr_ I _i9 = [inbred CC019s] — [Fl CC019 alleles]15.Inbr_F1_37= [inbred CC037s] - [Fl CC037 alleles]16.lnbr_19_F1_37= [inbred CC019s] - [Fl CC037 alleles]17. A&r-37-F1-19= [inbred CC037s] - [Fl CC019 alleles]18. Abr-F1-19x37 19= [inbred CC019s] - [CC019x CC037 Fl CC019 alleles] 19.Inbr_F1-19x37_37= [inbred CC037s] - [CC019x CC037 Fl CC037 alleles] 20.Inbr_F1-37xl9_19= [inbred CC019s] - [CC037x CC019 Fl CC019 alleles] 21.Inbr_F1-37xl9_37= [inbred CC037s] - [CC037x CC019 Fl CC037 alleles]
[0133] These groupwise methylation differences were then used to assign each statistically significant DMR to its proper category. The defined categories are mutually exclusive and. as such, each DMR is assigned to only one. However, as these methylation patterns do not necessarily function independently of one another, additional categories have been defined which exhibit characteristics of multiple inheritance patterns, as well as a one for those DMRs which cannot be clearly categorized. DMRs assigned to these additional categories are not considered in subsequent analyses. For the liver data, 358 autosomal DMRs were in categories 14-19 and 8 DMRs in category 20. For the muscle data, 20 autosomal DMRs were in categories 14-19 and 0 DMRs in category 20. These numbers are after application of the smoothing filter described below. The maximum absolute categorical difference is defined as Amax.1. Sex-specific DMR:• Ifmax= |Asez| and |Asex| > 10%o This condition is only checked for the autosomal comparison.2. Imprinted DMR:* If A77lax\^imprinting\ and \ imrpinting \ — 10 / Oo This condition is only checked for the autosomal and X chromosome female sample comparisons.3. Overdominant and underdominant emergent:* If ^max \^OverUnder_Dom\ ^nd \ overUnder_Dom\ — 10%• OR \^OverUnder_Dom\ 10% & I Inbr_Fl_191 — 10% & | lnbr_Fl_371 — 10% & |A / nb,-_37-Fi_i 91 > 10% & | A[nbr_19_Fl_371 — 10% & sign(A / n&r-F1 19) = sign(A / nbr F1 37) & sign(A / nhr19_F1_37 ) sign(A / „br-37-Fl-19) 4. Polar emergent - CC019xCC037:371626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO* Apolar 19x37 — 10 / 6 & ^polar_37xl9^ < 10 / 6 & Ap(p0 ar| > 10 / 6 & l^transl < 10% & sign(A / nbr F1-19x37 19) = sign(A / nbr F1-19z37 37)o This condition is only checked for the autosomal and X chromosome female sample comparisons.5. Polar emergent - CC037xCC019:* \^polar _37xl9\ — 10 / & |^po(ar_19x3?l < 10 / 6 & bipolar — 10 / 6 & |Ajrcms| < 10% & sign(A / nbrF1-37X19.19 ) = sign(A;nbrFl-37xl9_37)o This condition is only checked for the autosomal and X chromosome female sample comparisons.6. Bipolar emergent:* l^&ipotarl — 10 / 6 & l^bipofarl — \^polar_19x37\ & ^bipolar^ — ^polar_37xl9^ ■ l^trcmsl < 10% & sign(A;nbr-F1-19x37 19) = sign(AZ>rFl-19x37_37) & sign(A / ndrFl- 37xl9_19)—sign(A / )lZ„._F1-37xl9-37)o This condition is only checked for the autosomal and X chromosome female sample comparisons.7. Biallelic emergent:* ^biallelic^ — 10 / 6 & |At).a?ls| < 0.5 * & l^ / n&r_Fl_19l 5 / 6 & l^ / )ib7-_Fl_37l — 5% & |A / nbr-19 F1 37| > 5% & |A / nZ,r-37-F1-19| > 5% & sign(A / nbr-F1 19)!= sign(A / ni)r-F1-37) 8. Allele-specific emergent - CC019:* lAis, emerg_19 | > 10% & I &biallelic\ > 10% & | A / nbr F1 19| > 10% & \^Inbr_37_Fl_191 — 10% & |Atrans| < 10% & | &inbr_Fl_371 < 10% &|A / nbr_19_Fl_37l < 10%9. Allele-specific emergent - CC037:* IAAS emerg_37l — 10% & | A / ,iaiiei ic| > 10% & |A / ni,r F1 37| > 10% & I^Znbr_19_Fl_37l — 10% & I A trans | < 10% & |A / ni)r F1 19| < 10% &\^Inbr_37_Fl_191 < 10%10. Dominant / ram-acting meQTL | transvection | paramutation - CC019:* \^p ar amut _19 | > 10% & |Atrans| > 10% & |AMbr F1 19| > 10% & A / )ibr_19_F1_37 10% & \&biallelic\ < 10% & \& Inbr_Fl_371 < 10% &A / n£>r_37_F1_19 < 10% 11. Dominant / ra -acting meQTL | transvection | paramutation - CC037:* \^paramut_37\ — 10 / 6 & |Ajrans| > 10 / 6 & | &inbr_Fl_371 — 10 / 6 & A / ;i£>r_37_Fl_19 — 10 / 6 & \^biallelic\ < 10 / 6 & |A;nj,r-F1.191<10% &|A / nbr_19_Fl_37l < 10% 12. Non-dominant irony-acting meQTL:* l^transl — 10 / 6 & | biallelic\ < 0.5 * |Atrcms| & | inbr_Fl_l 91 5 / 6 & | / nbr_Fl_37l 5% & | ^inbr_19_F1 371 5% & |A / nZ,r 37-F1 19| > 5% & sign(A / ni„._F1-19)!= sign(A / nbr-F1-37)13. cA-acting meQTL:381626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO*l^cisl — 10 / 6 & l^trcms | > 10% & > 10% & |AZnbr 19 F1 37| > 10% & lA / nbj- gy pi igl > 10% & |A / nbr-F1 19| < 10% & |A / 7li„._F1-37| < 10% 14. c / .s-acting meQTL / biallelic emergent:* ^biallelic^ l^ti-ajisl & | / nfcr_Fl_19l < 10% & I A Inbr_Fl_37 | < 10% & sign(A / nb7-_F1 19)!= sign(A / ni,r-F1 37)15. cw-acting meQTL / allele-specific emergent - CC019* l^transl |A / nbr_Fl_37l < 10 / 6 & sign(A7)lj9 / ._Fi'| g) — sign(A / ni),._37-F1-19) & sign(A / nbr-37-F1 19)!= sign(A;,li)r19. Fi.37) 16. c / .s-acting meQTL / allele-specific emergent - CC037* I ^biallelic\ > |Afrcmsl & |A / 77b7-_Fi_i | < 10 / 6 & sign(A / 71j7r-F^_37) — sign(A / 7jbr-19-F1 37) & sign(A / 7lbr-l 9- 1 37)!= sign(A / 7i;7r-37-F1-l9) 17. czs-acting meQTL / non-dominant trans- acting meQTL* l^trcmsl l^t>ia!(e(icl & l^ / nZ>r_Fl_19l < 10 / 6 & \^inbr_Ft_Tl I < 10 & sign(A / n&r F1 19)!= sign(A / riZ)r-F1-37)18. c / .s-acting meQTL / dominant trans-acting meQTL | transvection | paramutation - CC019* lAfransI > ^biallelic^ IAznbr. Fl.37l < 10 / 6 & sign(A7lj77._F^_9) — sign(A / nZ,,._19-F1-37) & sign(Abr 37 F1_19)!= sign(A / n / )r19. Fi.37) 19. c / .s-acting meQTL / dominant frans -acting meQTL | transvection | paramutation - CC037* lAfrcmsl > I AfriaUeZicI & |Aznbr _F1_19| < 10 / 6 & sign(A / „,7-F1 37) — sign(A / nb,._37-F1 19) & sign(Abr l9 F1 37)!= sign(A / ni)r37. Fi.19) 20. Uncategorized* All remaining DMRs
[0134] After categorization, a final filter was implemented to ensure that sufficient raw methylation differences are present over identified DMRs. To do so, raw methylation differences between the relevant groups for each category, w ith strain-specific CpGs included, were calculated over each DMR as well as Ikb and 3kb windows up and downstream of the DMRs, as smoothing takes into account regional methylation effects. DMRs in which at least one of these windows exhibits a raw methylation difference of at least 10% or half of the smoothed methylation difference and in the same direction as the smoothed methylation difference are retained. As nearby ch-acting meQTLs will create a methylation difference between the relevant groups of DMRs categorized as allele-specific emergent, dominant transacting meQTL / transvection / paramutation, non-dominant / ra / z.s-acting meQTL, or biallelic emergent, the relevant methylation difference for these DMRs was also checked to be greater than AC7Swithin the window.391626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0135] Whole-genome ONT sequencing of the inbred and Fl liver samples was completed in two batches with differing library preparation kits (batch 1: LSK109. batch 2: LSK110) and target read lengths (batch 1: lOkb, batch 2: 20-30kb), factors which may influence alignment, phasing, and base / methylation calling. Although each batch includes inbred samples from each strain as well as heterozygous Fl crosses, ensuring that none of the autosomal methylation comparisons are fully confounded by batch, autosomal DMRs were searched to identify any which may still be influenced by a batch effect. To do so, smoothed methylation differences between the relevant groups for each category were calculated over each DMR both with and without the samples from batch 1 included. Those DMRs for which the relevant methylation difference without the inclusion of the batch 1 samples is less than half of the relevant methylation difference with all samples included are noted as potentially influenced by sequencing batch. Of the 6,970 autosomal DMRs identified in the liver, only 29 (0.4%) fall into this category'. This additional check was not performed for the chrX analyses as all male inbred CC019 samples originated from batch 1. As such, male maternal allele comparisons relative to the inbred CC019 samples cannot be tested excluding the batch 1 samples.
[0136] For the liver F2 dataset, coordinate-converted methylation frequency files for each sample were imported into R and converted into BS objects. Methylation data was split by autosomal and X chromosomal CpGs. The local haploty pe of each sample over each of the 84 targeted DMRs was then assigned by comparing the phased coverage of each allele over the adaptive sampling target regions. Samples / regions for which the coverage of the phased CC019 allele is > 75% of the total coverage for the sample over the region are defined as homozygous CC019. Samples / regions for which the coverage of the phased CC037 allele is > 75% of the total coverage for the sample over the region are defined as homozygous CC037. Samples / regions for which the coverage of the phased CC019 and CC037 alleles are both <= 75% of the total coverage for the sample over the region are defined as heterozygous. Methylation data was then subset to those CpGs within 500kb of the candidate transvection / dominant / ram-acting meQTL / paramutation regions, or lOOkb of the remaining targeted regions. Smoothing was performed separately on each of these regions with the same parameters used for the inbred and Fl dataset. After smoothing, only the proper alleles were extracted based on the assigned local haploty pe: CC019 allele only for homozy gous CC019 samples, CC037 allele only for homozy gous CC037 samples, and both alleles for heterozy gous samples. These samples / alleles were then plotted over the autosomal DMRs, and just the female samples / alleles were plotted over the X chromosomal DMRs. Notably, only a single 401626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO female F2 sample had heterozygosity over any of the X chromosomal DMRs. As heterozygosity is necessary to identify skewed XCI, these DMRs were removed from subsequent analysis.
[0137] Analysis of methylation over the Vps37c and Capnll IAPS.
[0138] Analysis of methylation over each of these CC037-specific IAPs requires additional re-processing, as the lps37c IAP is not present in the CC037 reference genome and coordinates within the Capnll IAP, which is present within the CC037 reference genome, cannot be mapped into the intermediate genome as they reside within an edge which is longer than 3kb. Furthermore, IGV investigation of reads overlapping the Capnll IAP reveals additional supplementary alignments which are highly likely to belong to other IAP elements in the CC037 genome that are not present within the reference sequence. As such, due to the high degree of sequence similarity between these IAPs, MiniMap2 assigns the sections of these reads which directly overlap the IAPs missing from the reference genome as supplementary alignments and incorrectly places them over an IAP which is present within the reference genome, in this case the Capnll IAP. These supplementary alignments do not extend past the region containing the IAP.
[0139] Within the Integrative Genomics Viewer (IGV; v2.19.1), the sequence of the ~5kb insertion present in the CC037-phased reads of the dominant / ra -acting meQTL / transvection / paramutation region overlapping the gene Vps37c was copied from read 80ce626d-cfl3-4e01-bdb9-a78ce5876b04 of the CC037 inbred muscle sample 37-104 to be used as a representative example. The sequence of this insertion was run through the Dfam (v3.9) sequence search tool using Mus musculus as the model organism and the Dfam curated threshold cutoff. This query returned six matches, all of which fall into the IAP category. The longest two matches with the lowest E-values (both 0) are of the family lAPEz-int (overlapping 4,299bp of the 5,054bp insertion).
[0140] The CC037 reference genome sequence between the two flanking CpGs which can be mapped into the intermediate genome around the Capnll IAP region was also run through Dfam using the parameters defined above. This query returned 12 matches of which 6 fall into the IAP category, overlapping 5,274bp of the 5,841bp sequence. The longest two matches with the lowest E-values (both 0) are of the family lAPEz-int (overlapping 4,302bp of the 5,841bp sequence).
[0141] To investigate the methylation pattern over the Vps37c IAP, which is not present in the original CC037 reference genome and, as such, did not have methylation called by 411626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO Nanopolish in the liver samples, anew version of the CC037 reference genome which includes the IAP was created by incorporating the entire 5.054bp sequence at the site of the insertion.
[0142] Then, methylation over both the Vps37c and Capnl 1 TAPs was locally re-processed for all samples for which at least one CC037 allele is present. Genotyping data was utilized to select those F2s which were heterozygous or homozygous for the CC037 allele in this region based upon matching of the nearest upstream and downstream called variants for which CC019 and CC037 differ to the known genotype of CC019 and CC037.
[0143] For the inbred, Fl, and F2 liver samples, the Vps37c IAP re-processing was performed as follows: (1) the samtools view function was used to extract the IDs for all CC037-phased reads which aligned to the site of the original DMR overlapping Vps37c with an additional lOOkb added up- and downstream, (2) the seqtk (vl.5-rl33) subseq function was used to extract these reads from the basecalled FASTQ files, (3) minimap2 was used to align these reads to the new CC037 reference genome including the new IAP sequence, (4) the nanopolish index and call-methylation functions were used to call the methylation of these newly aligned reads using the new CC037 reference genome with the IAP included, and (5) the Nanopolish provided calculate_methylation_frequency.py script was used to convert the methylation calls into methylation frequencies for each CpG site. For the inbred and Fl muscle samples, the Vps37c IAP re-processing was performed as follows: (1) the samtools view¬ function was used to extract the IDs for all CC037-phased reads which aligned to the site of the original DMR overlapping Vps37c with an additional lOOkb added up- and downstream, (2) the seqtk subseq function was used to extract these reads from the basecalled and methylation called FASTQ fdes, (3) minimap2 was used to align these reads to the new CC037 reference genome including the new IAP sequence, (4) the modkit pileup command was used to extract methylation frequencies for each CpG site. Note that the CC019 samples / alleles were not re-processed as this IAP is not present in the CC019 genome.
[0144] With the incorporation of the Vps37c IAP into the CC037 reference genome, all coordinates after the start of the IAP can no longer be directly mapped into the same pseudohybrid intermediate coordinate system which was for the analysis. As such, a new, local pseudo-hybrid intermediate coordinate system was designed around this IAP in which to analyze its methylation concordantly with CC019 samples / alleles. To do so, the coordinate at which the IAP was inserted into the CC037 reference genome (i.e. its start coordinate) was mapped into the original pseudo-hybrid intermediate coordinate system in order to determine where in the new intermediate coordinate system the IAP begins (this is the first coordinate at 421626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO which the two intermediate coordinate systems diverge). CpG coordinates from CC019 samples / alleles which were previously mapped into the original intermediate coordinate system were then mapped into the new intermediate coordinate system as follows: (1) those coordinates which are upstream of the coordinate in which the IAP was inserted (i.e. their linear coordinate in the original intermediate coordinate system is less than the IAP insertion point coordinate) are kept the same, and (2) those coordinates which are downstream of the coordinate in which the IAP was inserted (i.e. their linear coordinate in the original intermediate coordinate system is greater than the IAP insertion point coordinate) are mapped to the sum of their coordinate in the original intermediate coordinate system and the length of the IAP insertion (5,054bp). Note that this ensures that no coordinates of CC019 samples / alleles are located over the IAP itself as this IAP is not present in the CC019 reference genome. The re-processed CC037 samples / alleles were then mapped into the new intermediate coordinate system utilizing the previous mapping of each respective sample into the original intermediate coordinate system as follows: (1) those coordinates which are upstream of the coordinate in which the IAP was inserted (i.e. their linear coordinate in the original intermediate coordinate system is less than the IAP insertion point coordinate) are kept the same, (2) those coordinates which are downstream of the end coordinate of the IAP in the new intermediate coordinate system (i.e. the sum of their linear coordinate in the original intermediate coordinate system and the length of the IAP is greater than the end coordinate of the IAP in the new intermediate coordinate system) are mapped to the sum of their coordinate in the original intermediate coordinate system and the length of the IAP insertion (5,054bp), and (3) those coordinates which are within the IAP (i.e. all remaining coordinates) are linearly mapped between the final coordinate from upstream of the IAP and the first coordinate from downstream of the IAP by adding the linear distance between each respective coordinate and its preceding coordinate in the new CC037 reference genome with the IAP included to the preceding coordinate in the new intermediate coordinate system. After mapping the coordinates into the new intermediate coordinate system, methylation was smoothed following the same procedure as previously described and plotted separately for the liver inbred and Fl samples, the liver F2 samples, and the muscle inbred and Fl samples. Note that for the F2 samples, a minimum coverage cutoff of lx was implemented to remove very low-covered samples as this region was not included in the original capture target regions.
[0145] Capnll IAP re-processing was performed as described above with the following exceptions: (1) CC037-phased reads which align to only the IAP region were excluded from 431626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO analysis to avoid analysis of incorrectly mapped supplementary alignments, taking only reads which overlap the lOOkb regions beginning Ikb up- and down-stream of the IAP, (2) a new genome was not created as the Capnl 1 IAP is already present within the CC037 reference genome, (3) an IAP length of 5,841bp was used, and (4) mapped coordinates from the original processing pipeline w ere used for all CpGs which are downstream of the end coordinate of the IAP. As the Capnl 1 IAP region was included in the original capture target regions, a minimum coverage cutoff was not implemented.
[0146] Note that for each of the re-processing and re-analysis steps above, the same versions of each tool and the same parameters used in the original processing were maintained. Additionally, the IAP coordinates used in the re-processing pipelines are estimates based upon the available sequencing and reference genome data, with the Cps37c IAP being considered as the full length of the 5,054bp insertion and the Capnl 1 IAP being considered as the full 5,841bp region between the flanking CpGs which can be mapped into the intermediate genome. Lastly, for each of these analyses, methylation over these estimated IAP regions is only considered for the allele which contains the IAP.
[0147] Cross-tissue DMR overlapping.
[0148] All cross-tissue DMR overlaps reported here show the difference between the total number of DMRs identified in at least one tissue (i.e. the sum of the number of liver and muscle DMRs) and the total number of DMRs remaining after reducing this superset of DMRs using the GenomicRanges reduce function. However, all DMR overlapping percentages (shown within-tissue) represent the percentage of DMRs within that tissue which overlap DMRs assigned the same category7within the other tissue. This is done to account for small changes in the DMR coordinates between tissues.
[0149] DMR gene tracks'.
[0150] Gene tracks for the DMR plots were made using the dmrseq (vl.22.1) dmrPlotAnnotations function with the mmlO annotation from annotatr (vl.28.0).
[0151] DMR-gene associations.
[0152] To identify genes which may be regulated by these DMRs. three levels of DMR-gene associations were defined: (1) promoter overlap, (2) enhancer overlap, and (3) genomic proximity (linear distance betw een the corresponding gene and DMR). Promoters w ere defined as the regions 3kb up- and down-stream of transcriptional start sites (TSSs), downloaded from version 102 of the mmusculus_gene_ensembl mart via BiomaRt (v2.58.0). Enhancer-gene interaction annotations were downloaded from EnhancerAtlas 2.0. Mus Musculus Liver 441626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO annotations were used for DMR-gene associations with liver DMRs and Mus Musculus Limb E14.5 annotations were used for the muscle DMRs. Enhancers were mapped from mm9 to mml 0 using the UCSC LiftOver tool. All TSSs and enhancers were then mapped from mml 0 into the intermediate genome coordinate system.
[0153] Using these mapped coordinates, DMR-gene associations w ere defined via the direct overlap of a DMR with defined (1) promoters or (2) enhancers, or (3) the location of a gene within lOOkb of a DMR. The level of association is noted for each gene considered.
[0154] RNA-seq pre-processing.
[0155] RNA sequencing reads were processed using a modified version of the SEESAW pipeline. Briefly, two strain-specific transcriptomes for CC019 and CC037 were constructed using the gffread (vO.12.7) function, strain-specific genomes, and strain-specific genome annotation files. The two strain-specific transcriptomes were then concatenated into a single diploid transcriptome, designating the origin strain for each gene in the FASTA definition lines. The diploid transcriptome was indexed using Salmon (vl.9.0) with the following parameter: --keep Duplicates. Sequenced reads were mapped to the diploid transcriptome and allelic expression levels were quantified using the Salmon quant function with the following parameters: -1 A, -validateMappings, and -numBootstraps 30. One male CC037xCC019 Fl sample exhibited poor mapping to the diploid transcriptome (68.8% mapping rate) and, as such, was removed from subsequent expression analysis.
[0156] Expression analysis.
[0157] Allelic expression values were then calculated in R using the fishpond package (v2.8.0). Briefly, a transcript-to-gene mapping file was constructed using the BiomaRt package (v2.58.0) and the nimusculus gene ensembl mart, extracting the following information for each transcript: ensembl transcript id. external gene name, ensembl gene id, chromosome_name, start_position, end_position. This mapping file was then converted to a GRanges object using the GenomicRanges package (vl.54.1). Salmon quantification output files were then imported using the importAllelicCounts function with the mapping file and the following parameter: format = "wide".
[0158] For each sample, allelic counts were converted to counts per million (CPM) by dividing by the total sum of the allelic counts for that sample, combined across all genes in the transcript-to-gene mapping file and both alleles, and multiplying by le6. Genes were then filtered such that those w ith a CPM less than 0.5 in at least half of the samples w ere removed from subsequent analyses. For subsequent allelic analyses, an additional level of filtering was 451626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO applied in which the allelic counts of inbred parental strains were used to identify genes for which allelic assignments were accurate, using a cutoff of 90% of reads being assigned to the proper allele matching the inbred strain in all inbred samples. For the analysis of XCI-associated genes, the above fdtering steps were performed using only the female samples.
[0159] To gauge the effect of batch on the RNA-seq data, PCA plots showing the allelespecific expression of autosomal genes with correct allelic assignment and expression above the aforementioned cutoff for the six inbred samples which were included in both RNA-seq batches. Clear clustering of these samples by sample and allele, rather than batch, suggest that the batch effect is much smaller than the biological difference between samples of different strains.
[0160] The allelic counts of the filtered genes, using only inbred samples from a single batch, were then analyzed with the swish function from the fishpond package using one of two methods: (1) the default global allelic imbalance analysis or (2) a differential allelic imbalance analysis. Additional analyses were performed on these filtered genes using total expression data, obtained by summing the allelic counts from both alleles and then converting to CPM. Total expression levels were analyzed using the DESeq2 package (vl.42.0). One female CC019xCC037 Fl sample had atotal expression CPM of more than 3 standard deviations away from the mean in 6,927 / 12,894 expressed genes and, as such, was removed from subsequent expression analyses as an outlier.
[0161] To compare methylation and expression, all genes associated with the identified DMRs via direct promoter / enhancer overlap or by genomic proximity were analyzed using a set of tests and conditions specific to each epigenetic inheritance pattern from which the DMR originates, outlined below. Tests for allele-specific overdominance and allele-specific underdominance were not designed as these necessitate a direct comparison in expression between a single phased allele of the FIs and the total expression of the inbreds, which cannot be performed.
[0162] Tests:1- Tstrain_inbred =DEseq2([total CC019] - [total CC037])2. 7’CCOI9_FI =DEseq2([total CC019] - [total Fl])3. TCCQ37 P1=DEseq2([total CC037] - [total Fl])4- TSex=DEseq2([ total female either allele] — [total male either allele])5. TAllele F1=Swishgiobai([Fl CC019 allele] - [Fl CC037 allele])461626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO a. Paired comparison6- Tlmprinting =Swishdifferentiai(([Fl CC019 allele] — [Fl CC037 allele]))a. Paired differential comparison based on cross direction
[0163] For the specified DMR categories listed below, the corresponding tests were used to identify the differentially expressed genes (DEGs), with TRUE being defined as a Benj amini -Hochberg (BH) adjusted p-value < 0.05. BH corrections were perfomied independently for each test within each associated DMR category:1. CA-acting meQTLs• Tstrain inbred= TRUE & TAtlele F1= TRUE &[sign(log2FC(Tst7.cl7n-777i)reri) sign(log2FC(fy[{e ejTi)]2. Non-dominant rran.y-acting meQTLs• Tstrainjnbred ~ TRUE & TAueie Fi— FALSE 3. Dominant / ram-acting meQTLs|transvection|paramutation - 19 • Tstrain_inbred=TRUE & TCC019_F1= TRUE & TCC037_F1= FALSE & [sign(iog2FC(Tstrain inbred) = sign(log2FC(TCCOi9_Fi)] 4. Dominant trara-acting meQTLs|transvection|paramutation - 37 • Tstrain_inbred= TRUE & TCC019 P1= FALSE & TCC037 P1= TRUE &[sign(log2FC(7’strain i7lbred)!= sign(log2FC(Tcc037_Fi)] 5. Genomic imprinting„ „r■, Avg CPM of CC037 allele of 19x37 + 0.0 l,x, • T Ilm„pruinttiinnga= TRUE & [ Lsig sn( \Zoa2(v-Avg CPM of CCOI 9 allele of 19x37 0.01 )7)! =., Avg CPM of CC037 allele of 37x19 + 0.0k signtvZo q2(v- ))] Avg CPM of CC019 allele of 37x19 + 0.0kJo A pseudocount of 0.01 is added to ensure that the denominator is not 06. Sex-specific DMRs• TSex= TRUE7. Overdominance or underdominance• ^CCO37_ I=TRUE & TCC019 F1= TRUE & [sign(log2FC(Tcc037 F1) = sign(log2F C(Tcco 19Fl)]8. Biallelic dominance• Tstrain Jnbred=FALSE & TAueie F1= TRUE9. Skewed XCI• Tstrain inbred— FALSE & TAueie F— TRUE• This comparison is only performed on female samples471626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0164] As the overdominance / underdominance and dominant / ra -acling meQTLs / transvection / paramutation comparisons require direct comparisons of expression across batches, the associated DEGs may in part be driven by batch effects.
[0165] Transcriptome variant analysis.
[0166] Heterozygous genetic variants were identified using the nucmer command from MUMmer4 run between the CC019 and CC037 strain-specific transcriptomes, converted to a delta file with SNPs and insertion-deletions (indels) using the show-snps function with the following parameters: -CTlr, and subsequently converted to a VCF file using the my-mummer-2-vcf.py script with the following parameter: -n. This process was run once with CC019 as the reference genome and once with CC037 as the reference genome. Genes having at least one transcript with one or more heterozygous transcribed variant were then identified using the transcript-to-gene mapping file and subsequently overlapped with the list of previously defined expressed genes.
[0167] MGI gene-disease associations.
[0168] Gene-disease associations were extracted from the following Mouse Genome Informatics (MGI) databases: Associations of Mouse Genes with DO Diseases and Mouse Models of Human Disease by Human Gene.
[0169] Read-level methylation plots.
[0170] Read-level methylation plots were generated for non-dominant tra s -acting meQTLs and skewed XCI DMRs using methylartist (vl.2.10). db-nanopolish was run on the Nanopolish methylation call output file for each sample using the following parameter: -t 2.5. DMRs were mapped from the intermediate coordinate system into the CC019 and CC037 coordinate systems and subsequently used as the input coordinates (with a 5kb extension) for the methylartist locus function. The CC019 coordinates were used for plotting CC019 inbred and CC019 phased alleles and the CC037 coordinates were used for plotting CC037 inbred and CC037 phased alleles.
[0171] Data Availability
[0172] All ONT and RNA-seq FASTQ files have been deposited to the NIH Sequence Read Archive (BioProject: PRJNA1107693, reviewer link: dataview.ncbi.nlm.nih.gov / object / PRJNAl 107693?reviewer=oadleaicl7kfqum3h07imh46gs)
[0173] Methylation frequency files, Guppy sequencing summary files, Salmon quantification files, and allele-specific gene counts have been deposited to the NCBI Gene 481626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO Expression Omnibus (accession code: GSE266668, reviewer token: spuvykoctlkjnsb). F2 genotyping data has been deposited to Zenodo and is available using the following link: zenodo.org / records / 17148323?token=eyJhbGciOiJIUzUxMiJ9.eyJpZCI6ImYyMGFlYzc4L WI4YzUtNDhjYy04YjJkLTIyZGYwY2UxOGQwZCIsImRhdGEiOnt9LCJyYW5kb20iOiIO YWJhMWRiZDc4MzQ2NWNiOWMyZWM3YzJiMGU5ZDJlYyJ9.slzgk2PKFXzKjxD7Yp jSt7djZUreDSo_RdnwlFLqDAfk_Y6_lakpLlatla7NXQEsRJOGiDfsV7u-aaT13nZIHA.
[0174] Code Availability
[0175] The processing and analysis code used in this manuscript is available through Zenodo using the following link: zenodo.org / records / 17148003?token=eyJhbGciOiJIUzUxMiJ9.eyJpZCI6ImZhODHZmYxL TM5OTltNDFjYyliODBjLTclYTlmMTVhYzA0MiIsImRhdGEiOnt9LCJyYW5kb20iOiI4 MTg5Mzk4NGNmNmM2N2JhYWVlNGQ4NmQwMzkyNzc4ZSJ9. ISBPDeQ4sSQuUOFiA KQX_jdNcKo9hNU_5qY-JN-WTqgtLPxy5FKvr3rNhNvxZl-jPlAdZvj3KHU6BahAtg8Wzw.
[0176] All the data accessible on the websites listed above, at paragraphs
[0172] ,
[0173] , and
[0175] are part of the present disclosure and are hereby incorporated in this disclosure in their entirety.EXAMPLE 3INTEGRATED GENETIC AND EPIGENETIC ANALYSIS OF THE INTERGENERATIONAL INHERITANCE OF DNA METHYLATION PATTERNS
[0177] A combined genetic and epigenetic analysis to identify DNA methylation patterns related to genotype, sex, intergenerational inheritance, tissue, and parent-of-origin effects was designed (FIG. 1A). This was achieved using genome-wide ONT long-read sequencing of liver and muscle DNA from 26 (15 liver, 11 muscle) male and female mice of two genetically distinct inbred CC strains, CC019 / TauUnc and CC037 / TauUnc, as well as 34 (22 liver, 12 muscle) of their Fl hybrids in both cross directions. Muscle and liver w ere profded in distinct mice. Targeted ONT sequencing of liver DNA from 19 F2 crosses was subsequently performed, generated by crossing members of the F 1 generation, to further analyze candidate genomic regions exhibiting several of the epigenetic inheritance patterns identified using the inbred and Fl datasets. Independent assortment of chromosomes and meiotic recombination in the Fl generation produce the unique genomes of the F2 generation in which the / ra - acting factors influencing the methylome are inherited independent of the local haplotype (i.e.491626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO homozygous CC019 or CC037, or heterozygous CC019 / CC037 alleles) of each region over which they influence methylation. As such, analysis of the F2 generation is essential to distinguish epigenetic inheritance patterns mediated by cis- and / ra / i.s- acting regulatory factors, as described in FIG. IB and 1C. The median coverage across both alleles for the wholegenome ONT sequencing of the inbred and Fl generations was 24x (ranging from 8x to 41x), while the median coverage across both alleles over the target regions in the F2 generation was 68x (ranging from 56x to 94x). Liver and muscle were chosen for analysis in part because these are relatively homogenous tissues, which reduces the possibility of cell type confounding.
[0178] To process and analyze the sequencing data, a computational pipeline optimized for the phasing and characterization of allele-specific methylation in genetically divergent mouse strains was developed (FIG. 8) Inter-strain analyses of sequencing data can be severely hampered by strain and reference biases. As such, computational pipeline was designed to minimize these biases while simultaneously maximizing the amount of data available for analysis. To do so, the pipeline performs alignment and phasing relative to strain-specific genomes and extracts only consensus phasing decisions which are consistent irrespective of the chosen reference genome. Methylation data is then mapped into a pseudo-hybrid intermediate genome, which incorporates genetic variation from both the CC019 and CC037 genomes, for subsequent analysis. The use of the intermediate genome avoids the introduction of bias from the choice of a single reference genome while maintaining strain-specific methylation information, which has been shown to increase statistical power in inter-strain comparisons of DNA methylation from isogenic mouse models.
[0179] Within the CC019 and CC037 reference genomes, there are ~19M CpGs of which -4.8M overlap regions of the genome which are identical by descent (IBD), i.e. lacking genetic variation between the two strains. Utilizing the experimental and computational approach, ~12M autosomal CpGs and ~350k X chromosomal CpGs were analyzed in both tissues, accounting for -90% and -85% of the CpGs which are present in both CC strains, can be mapped into the intermediate genome, and do not lie within IBD regions.
[0180] To explore the functional implications of the identified methylation patterns and their impact on gene expression, liver RNA from the inbred and Fl samples wree sequenced on which ONT sequencing w as performed. Outputs w ere evaluated for both allele-specific and total expression in the context of each of the epigenetic inheritance patterns identified, as described further in the Methods section. Only 9.5% (4,092 / 42,880 of all genetic elements considered. 3.956 / 19.655 of all protein coding genes) of genes contain transcribed 501626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO polymorphisms which differ between CC019 and CC037 with expression levels exceeding a minimum threshold. Consequently, allele-specific expression comparisons are feasible for only a subset of the genes over which allele-specific methylation analyses can be performed.
[0181] 12 distinct patterns of epigenetic inheritance related to genotype, sex, intergenerational inheritance, and parent-of-origin effects were defined and searched for. Whereas most of these patterns can be distinguished from each other using data from the inbred and Fl generations, three patterns are indistinguishable in these generations. These three patterns therefore require additional data from the F2 generation to be differentiated from one another (FIG. 8 and discussion below). Representations of each of the epigenetic inheritance patterns identified on the autosomes as well as the number of regions exhibiting each pattern are shown in FIG.8. Each of the patterns are discussed in detail below. It is important to note that the overall difference in the number of observed patterns between the liver and muscle is likely, in part, due to the lower sample size of muscle included in this analysis.EXAMPLE 4PATTERN 1; CIS-ACTING MEQTL
[0182] Among the observed epigenetic inheritance patterns, the most prevalent within both tissues were regions of the genome in which methylation stratifies by genotype and allele (FIG.2A). These methylation patterns are established by cis-acting meQTLs, genetic variants capable of regulating DNA methylation levels of proximal CpGs located on the same allele. In genomic regions under the control of these cis-acting meQTLs, Fl crosses between inbred strains will inherit one allele from each strain, and the CpGs on each allele will exhibit a methylation pattern matching the strain from which the allele was inherited. As such, the cis -acting meQTL and its respective methylation pattern are transmitted together from the inbred parental strains to each respective allele of the FIs, creating the observed genotype-specific methylation pattern. 7,081 genomic regions were identified which follow this pattern in at least one tissue, accounting for -93% of the autosomal intergenerational epigenetic inheritance patterns (FIG. 2B and 2C. FIG. 8). it was found that 16% and 50% of the cis-acting meQTLs identified in liver and muscle, respectively, are present in both tissues. Examples of tissueindependent and tissue-specific cis-acting meQTLs are shown in FIG. 2B and 2C.
[0183] To confirm that these differentially methylated regions (DMRs) are under the control of cis-acting regulatory elements, targeted ONT sequencing was performed on a set of 20 candidate regions in F2 mice. As expected, all 20 regions exhibit methylation levels that 511626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO segregate with the local haplotype over the target region (FIG. 2D), validating that the methylation patterns in these regions are established by cis-acting meQTLs.
[0184] To determine the effect of these cis-acting meQTLs on gene expression, the focus was on genes whose promoters / enhancers they overlap as well as genes located within 100kb of the corresponding DMR as potential targets of transcriptional regulation. As described above, only allele-specific expression in a small fraction of genes (2,610 / 13,013 genes associated with cis-acting meQTLs) was assessed. As such, many of the genes whose expression levels are associated with these DMRs will not be identifiable. Specifically, genes with an allelic imbalance in the FIs which mirrors the expression difference observed between the two parental strains (FIG. 2E) with no limitation on which allele is more highly expressed were searched for. 700 genes that exhibit cis-acting regulation of both expression and DNA methylation were found, suggesting that many of these cis-acting meQTLs may also function as expression quantitative trait loci (eQTLs).
[0185] In addition, the potential relationship between these cis-acting meQTLs and inherited disease was investigated by intersecting the set of genes associated with these DMRs via direct promoter / enhancer overlap or by genomic proximity with curated databases of mouse and human disease-gene associations from Mouse Genome Informatics (MGI). This analysis identified 2,201 such overlaps, suggesting that many of the genes associated with these cis-acting meQTLs may be relevant to human disease.EXAMPLE 5PATTERN 2: NON-DOMINANT TRANS ACTING meQTL
[0186] Regions in which methylation stratifies by genotype only in the inbred samples were also identified, while both alleles of the FIs exhibit intermediate levels of methylation (Fig. Sabi. These patterns are likely caused by / ra s-acting meQTLs, genetic variants which influence methylation levels of distal CpGs and are capable of regulating both alleles. 7 / vzm-acting meQTLs can manifest as dominant, in which the presence of a single copy of the dominant allele establishes the dominant methylation pattern on both alleles (FIG. 1C): incomplete dominant, in which individuals heterozygous for the / ram-acting meQTL exhibit an intermediate methylation level on both alleles (FIG. 3A); or codominant, in which heterozygous individuals exhibit both of the methylation patterns obser ed in the homozy gous individuals on both alleles (FIG. 3B). Within a population of inbred and Fl individuals, incomplete dominant and codominant / / v / m-acting meQTLs will each establish intermediate 521626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO levels of methylation on both alleles of the FIs which fall between the levels observed in the two parental strains, matching pattern 2. Conversely, dominant rram-acting meQTLs give rise to pattern 3, discussed further in the next section. Due to the binary nature of DNA methylation at a single CpG in a single cell, intermediate levels of methylation can arise either from variability' of methylation at the cellular level or at the CpG level within each cell. Using the long-read data, the variability of methylation at both the cellular and CpG levels can further be assessed and allowing distinguishing between incomplete dominance and codominance. Nondominant (incomplete dominant or codominant) / ram-acting meQTLs violate Mendelian inheritance, as neither allele is completely dominant nor recessive in its regulation of methylation.
[0187] 51 genomic regions which exhibit this methylation pattern in at least one tissue were identified (FIG. 3C, FIG. 8). While there were fewer non-dominant / ra - acting meQTLs identified in the muscle than in the liver, only 22% (2 / 9) of the muscle non-dominant transacting meQTLs were also observed in the liver (FIG. 3C and 3D), suggesting that these non-dominant / ram-acting meQTLs may be more tissue-specific than their czs-acting counterparts.
[0188] While incomplete dominant and codominant / ram-acting meQTLs establish identical methylation patterns at the sample level, they can be distinguished from one another at the read-level. In regions under the regulation of an incomplete dominant / ram- cting meQTL, each sequenced strand of DNA from both parental alleles of the Fl crosses will exhibit a level of methylation between that of the two inbred parental strains, indicating variability of methylation at the CpG-level within each cell (FIG. 3A). Conversely, in regions regulated by codominant / ram-acting meQTLs, each of the phased Fl alleles will exhibit two clear epialleles, which were defined here as groups of reads which are either fully methylated or fully unmethylated across the region, indicating variability of methylation at the cellular-level (FIG.3B). As such, the difference in the methylation patterns of the two inbred strains and the Fl crosses in these regions will be driven by differences in the distributions of the two epi-alleles, rather than a difference in the methylation level of each strand of DNA, as is the case with the incomplete dominant tram-acting meQTLs.
[0189] Five candidate regions under the control of non-dominant tram-acting meQTLs were chosen for subsequent targeted analysis in the F2s. Three of these five regions exhibit no clear relationship between the local haplotype of the F2s and the observed methylation states (FIG. 3E), further supporting that these regions are under the control of non-dominant transacting genetic variants. The remaining two DMRs exhibit a genotype- and allele-specific 531626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO methylation pattern in the F2 generation which shows a smaller difference than that observed in the inbred generation and roughly the same as that observed in the F 1 generation. This pattern suggests that these regions are likely under the regulation of both cis- and non-dominant / raw.s-acting meQTLs, and that, as would be expected, the methylation differences associated with the homozygous non-dominant trans-acting meQTLs in the inbred generation are removed in both the Fl and F2 generations, while the effects of the czs-acting meQTLs are maintained.
[0190] Among the genes associated with these DMRs via direct promoter / enhancer overlap or by genomic proximity7, 47 have been implicated in the etiology of diseases in the MGI databases. In the case of non-dominant / ra s-acting meQTLs which contribute to disease phenotypes, incomplete penetrance and variable expressivity may occur as heterozygous individuals with intermediate levels of methylation may not always reach molecular thresholds for disease manifestation or may exhibit more mild symptoms than homozy gous individuals or even other heterozy gous individuals. As such, these non-dominant / ra - acting meQTLs may be missed by purely genetic analyses attempting to identify genetic variants contributing to disease. Indeed, several of the associated diseases exhibit these complex forms of inheritance. Among these are trichodentoosseous syndrome, associated with Dlx3, which exhibits clinical variability7in the absence of genetic heterogeneity for which the authors hypothesize potential epigenetic or environmental contributions to its etiology; Currarino syndrome, associated with Mnxl, which exhibits both variable expressivity as well as reduced penetrance; and nail-patella syndrome, associated with Lmxlb, for which the effect of the driving mutation may be influenced by variants on the non-mutant allele acting in trans. For each of these three genes, a DMR identified in the liver exhibiting regulation by a non-dominant fra s -acting meQTL overlaps the promoter region.
[0191] Subsequently allele-specific and total expression levels of genes associated with these DMRs were analyzed via direct promoter / enhancer overlap or by genomic proximity to identify instances in which expression levels are consistent with regulation by non-dominant Zraws-acting meQTLs. These genes would exhibit expression levels which stratify by genotype in the inbred samples but lack significant allelic imbalance in the F 1 s. Nine such instances were identified(FIG. 3F), suggesting that these non-dominant / ram-acting meQTLs may also have the potential to function as eQTLs, thereby providing an additional link between these regulatory mechanisms.541626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WOEXAMPLE 6PATTERNS 3-5: DOMINANT TRANS ACTING MEQTL, TRANSVECTION, AND PARAMUTATION
[0192] The analysis further revealed epigenetic inheritance patterns in which methylation stratifies by genotype in the inbred samples while methylation on both alleles of the F Is match just one of the parental strains (FIG.4A-4B). 23 such regions were identified present in at least one tissue (FIG.8). Four of these 23 regions were identified in both tissues, representing 17% and 67% of these regions identified in the liver and muscle, respectively. Included among these four tissue-independent DMRs is a region overlapping the gene lps37c (FIG.4B), involved in the vesicular sorting of endocytic cargo and subject to further investigation below.
[0193] This pattern can conceivably arise as a result of three distinct underlying mechanisms: (a) dominant / ra -acting meQTLs, (b) transvection, or (c) paramutation, and additional data from the F2 generation is required to distinguish these potential mechanisms. The three distinct mechanisms are defined as follows:(a) Dominant / ra / rs-acting meQTLs require only a single copy of the dominant genetic variant to establish a particular methylation pattern on both alleles. Consequently, in regions under the control of this regulatory mechanism, inbred strains will exhibit distinct methylation patterns, while both alleles of the heterozygous Fl crosses will inherit the dominant methylation pattern (FIG.4C). In the F2 generation, the methylation over the region will be determined by the zygosity of the individual over the / ra -acting factor. Dominant / ra -acting meQTLs do not violate Mendelian inheritance.(b) In the case of transvection, one allele in a heterozygous individual acquires the methylation state of its dominant homolog. Thus, though the parental strains exhibit distinct methylation patterns, both alleles of the FIs will exhibit a methylation pattern matching that of parental strain from which the dominant homolog originates (FIG. 4D). In the F2 generation, the methylation over the region will be determined by the local zygosity of the individual. Transvection events do not violate Mendelian inheritance. Transvection naturally occurs in plants, Drosophila, and fungi, and has been observed within transgenic mammalian genomes under certain conditions. However, it has not previously been documented in a non-transgenic mammalian genome, nor has it been mapped genome-wide.551626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO (c) Paramutation involves the alteration of DNA methylation on a paramutable allele due to the presence of its homologous paramutant allele. However, unlike dominant transacting meQTLs and transvection, this altered methylation state is subsequently heritable independent of the dominant paramutant allele or other / ra / ?.s -acting factors (FIG. 4E). In the F2 generation, methylation over the region will match the altered methylation state in all individuals, regardless of their local or distal genotype. As such, the initiation of paramutation is a hybrid effect, and the heritable methylation pattern can be established by either a transvection event or a dominant / ram-acting factor. Paramutation violates Mendelian inheritance, as the dominant methylation pattern and associated phenotypes can be inherited even in the absence of the dominant allele. Similar to transvection, paramutation has been previously identified in plants. Drosophila, and transgenic mice, though it has not been previously observed in non-transgenic mammals nor mapped genome-wide.
[0194] To differentiate genomic regions regulated by paramutation, transvection, and dominant trans-acting meQTLs, a set of three candidate regions in the F2 generation was profiled. To do so, 19 F2 samples sequenced based on their genotyping were selected in order to ensure that, at each region in the genome, there is at least one sample which is homozygous for each of the two parental alleles - CC019 and CC037. This ensures that, if the methylation of a candidate region is driven by a dominant fra s-acting meQTL, at least one sample will exhibit the non-dominant methylation pattern, allowing for its distinction from paramutation.
[0195] it was observed that a candidate region overlapping the gene Capnll exhibits only one methylation pattern in the F2s, matching that of both F 1 alleles (FIG. 4F). This indicates that this region exhibits intergenerational paramutation, the first instance of this form of non-Mendelian intergenerational epigenetic inheritance to be identified in non-transgenic mice. Capnll is a calcium-dependent protease which has been shown to be expressed primarily in the testis during later stages of meiosis in both mice and humans, and decreased expression in humans has been shown to be associated with infertility and azoospermia.
[0196] Among the remaining two candidate regions profiled in the F2s which were not classified as paramutation, neither exhibited a methylation pattern which indicates clear regulation by a dominant trans-acting meQTL or a transvection event. Rather, these regions exhibit a methylation pattern which cannot be definitively classified into any of these three regulator}7mechanisms, as the presence of a single methylation pattern across all members the F2 generation (suggesting paramutation), nor an identical pattern in the F2 generation as is observed in the inbred and Fl generations (suggesting transvection), nor a clear lack of 561626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO association between the local genetics and the methylation over the region was not observed (suggesting a dominant trans- acting meQTL). As such, it is likely that these regions are under the control of additional complex regulatory mechanisms.
[0197] Within the aforementioned DMR overlapping Vps37c (FIG. 4B), it was noticed the presence of a~5kb insertion inside the DMR relative to CC037 reference genome. Investigation of this insertion using the Dfam database reveals that it is a transposable element (TE) belonging to the IAP family of ERVs. IAPS, as well as many other forms of TEs, are ty pically repressed via epigenetic mechanisms including DNA methylation, suggesting that the presence of the identified methylation pattern may serve to repress the IAP in the CC037 alleles of the heterozygous Fl generation in both liver and muscle, and that this IAP may not be repressed in the inbred CC037 samples for which methylation is significantly lower. Notably, IAPs have been shown to be capable of maintaining their methylation state through the waves of DNA demethylation which occur during gametogenesis and embry ogenesis, allowing for this repressive methylation pattern to be maintained in subsequent generations. As such, it was suspected that this DMR may be an additional example of intergenerational paramutation, in which the highly methylated pattern observed in the CC037 alleles of the Fl generation in both liver and muscle would be subsequently inherited in all CC037 alleles of the F2 generation. While this region was not included among the target regions for the sequencing of the F2 generation, there exists another IAP ~13kb downstream of the Capnll paramutation region which was included as a target region. The Capnll IAP has sufficiently high sequence homology to the Vps37c IAP such that the CC037 allele was captured by adaptive sampling with sufficient frequency for regional methylation analysis (average coverage of the CC037 allele with the IAP: ~8x, average coverage of the CC019 allele without the IAP: ~0.5x). As such, the methylation of CC037 samples and alleles were re-analyzed over this region utilizing anew CC037 reference genome with the IAP sequence included. Methylation over the Vps37c IAP is much lower for the inbred CC037 samples than the CC037 alleles of the heterozygous FIs in both the muscle and liver (FIG. 4G, 4H). Furthermore, a methylation pattern was observed in all of the CC037 alleles of the F2 generation which more closely matches that of the CC037 alleles of the Fl generation than the inbred CC037 samples, independent of the zygosity of the F2 samples in this region. While this result strongly suggests that this is an additional example of intergenerational paramutation, there remains the possibility that the observed methylation pattern over the Vps37c IAP is driven by a dominant / ra - cting meQTL. This is because methylation over the IAP cannot be analyzed within the five F2571626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO samples which are homozygous for the CC019 allele over this region, and the possibility that each of the 14 remaining F2s carry at least one copy of a dominant / ram-acting meQTL originating from the CC019 strain which regulates methylation over this TAP cannot be excluded.
[0198] Given this observation, it was sought to further analyze DNA methylation over the IAP ~13kb downstream of the Capnll DMR, also present only in the CC037 genome and included within the F2 capture target region with an average coverage of ~5 lx per sample / allele due to its proximity to the Capnll DMR. Methylation over this IAP shows a pattern similar to that of the Vps37c IAP, with inbred CC037 samples exhibiting a much lower methylation level than the CC037 alleles of the Fl generation. Furthermore, all CC037 alleles within the F2 generation exhibit a methylation pattern nearly identical to that of the Fl CC037 alleles, strongly suggesting that this is an additional example of intergenerational paramutation. However, as with the Vps37c IAP, there remains the possibility that methylation of this IAP is driven by a dominant / ra -acting meQTL for which none of the F2 samples which contain the IAP have the non-dominant allele.
[0199] 25 genes associated with these DMRs via direct promoter / enhancer overlap or by genomic proximity- have been implicated in the etiology of diseases in the MGI databases, several of which exhibit atypical patterns of inheritance. Among these are the variable phenotypes observed in the absence of genetic heterogeneity in Timothy syndrome and common variable immunodeficiency 8, associated with Cacnalc (for which a DMR identified in the liver overlaps the promoter) and Lrba, respectively. This suggests unexplained genetic or environmental factors associated with these diseases that may be mediated through the epigenome or via other forms of transcriptional regulation. In the investigation of the complex inheritance patterns of these disorders, paramutation, transvection, and dominant trans-acting meQTLs were likely not considered, nor was the epigenome sequenced and evaluated as a potential avenue of intergenerational inheritance of the disease phenotype. Given the results reported here, these additional modes of inheritance should be considered in family studies.
[0200] To investigate the link between these methylation patterns and gene expression, genes associated with these DMRs were analyzed via direct promoter / enhancer overlap or by genomic proximity to identify those for which the total Fl expression levels match that of only one of the parents. This analysis revealed 14 such genes (FIG. 4J), suggesting a potential role for these methylation patterns in the regulation of gene expression.581626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WOEXAMPLE 7PATTERNS 6-10: Emergent epigenetic inheritance paterns
[0201] Three distinct sets of emergent epigenetic inheritance patterns were were also identified, in which at least one allele of the F 1 generation exhibits a novel methylation pattern not observed in either of the parental strains. These patterns can be further categorized into (1) overdominance and underdominance, in which both alleles of heterozygous offspring exhibit higher or lower methylation levels than the inbred parental strains, respectively (FIG. 5A); (2) allele-specific overdominance and allele-specific underdominance, in which just one allele of heterozygous offspring exhibit higher or lower methylation levels than inbred parental strains as well as the other allele of the cross, respectively (FIG.5B); (3) bi all el ic dominance, in which one allele of the heterozygous offspring exhibit higher methylation than the inbred parental strains while the other allele exhibits lower methylation (FIG. 5C); (4) polar overdominance and polar underdominance, in which both alleles of one cross direction of the heterozygous offspring exh i bi t higher or lower methylation levels than inbred parental strains as well as both alleles of the other cross direction; and (5) bipolar dominance, in which both alleles of one cross direction of the heterozygous offspring exhibit higher methylation than the inbred parental strains while both alleles of the other cross direction exhibit lower methylation. Each of these patterns violate Mendelian inheritance, as the offspring inherit methylation states which do not match either parental allele. These patterns can be caused by many distinct regulatory7mechanisms and the methylation state in subsequent generations depends on the exact mechanism driving the pattern. For this reason, these patterns were not selected as targets in the F2 generation.
[0202] The analysis revealed 54 instances of these emergent epigenetic inheritance patterns present in at least one tissue (FIG. 8), of which one was overdominant (FIG. 5D), 22 were allele-specific overdominant or allele-specific underdominant (FIG.5E), and 31 were biallelic dominant (FIG. 5F). This represents a substantial increase over what has been previously reported in mammals in the literature - seven examples identified in mice in addition to the callipyge locus in sheep. Notably, only a single emergent epigenetic inheritance pattern is found in both liver and muscle, suggesting that these patterns are primarily regulated in a tissuespecific manner.591626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0203] Emergent epigenetic inheritance patterns which contribute to disease phenotypes are likely to violate typical Mendelian inheritance, as the disease will be associated with the emergent pattern in addition to or in instead of the genetic variant(s) being considered. As such, a higher or lower proportion of individuals may be affected than can be explained purely by Mendelian genetics. Furthermore, as the emergent epigenetic inheritance pattern may be susceptible to additional factors such as environmental exposures, the associated disorders may exhibit a high degree of phenotypic variability. 54 genes associated with these emergent epigenetic inheritance patterns via direct promoter / enhancer overlap or by genomic proximity have been implicated in the etiology of diseases in the MGI databases and several of these disorders demonstrate these complex patterns of inheritance. Among these are combined pituitary hormone deficiency-5, associated with Hesxl, for which both humans and mice exhibit incomplete penetrance of the heterozy gous phenotype; as well as developmental and epileptic encephalopathy 94, associated with Chd2 for which an allele-specific overdominance DMR identified in the liver overlaps the promoter, which exhibits clinical heterogeneity within and between families with the same mutations.
[0204] To further explore the functional role of these emergent epigenetic inheritance patterns, the expression of genes associated with these DMRs via was analyzed promoter / enhancer overlap or genomic proximity and 19 which exhibit expression patterns consistent with biallelic dominance were identified (FIG.5G).EXAMPLE 8PATTERN 11; GENOMIC IMPRINTING
[0205] Parent-of-origin-specific methylation, or genomic imprinting (FIG.6A), is currently the most thoroughly studied example of non-Mendelian inheritance of DNA methylation in mammals. Genomic imprinting violates Mendelian inheritance, as the inheritance of methylation and associated phenotypes is influenced by the parent from which the allele originates, rather than being determined solely by the identity of the allele. 111 autosomal regions which exhibit parent-of-origin-specific methylation patterns in at least one tissue were identified (FIG. 6B, FIG. 8). A single region exhibiting parent-of-origin-specific methylation on the X chromosome was also identified. This region is discussed further alongside the other X chromosomal epigenetic inheritance patterns. This list was compared to a previously reported curated list of DNA methylation-regulated imprinting control regions (ICRs) to 601626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO evaluate the analysis approach on known regions. Of the 13 known ICRs which are analyzable within the two selected strains, imprinted methylation patterns was identified in at least one tissue within Ikb of all 13.
[0206] While -54% of the autosomal imprinted regions identified in each tissue directly overlap one another (54.5% and 53.9% for liver and muscle, respectively), 59 (76.6%) of the imprinted regions identified in liver are within lOOkb of an imprinted region identified in the muscle. This indicates that, while these DMRs may not he over the exact same CpGs in different tissues, they tend to exist in the same regions / domains. However, several exceptions were observed in the form of tissue-specific imprinted genes, including Zswim9 and Asb4, which exhibit parent-of-origin-specific methylation only in the liver and muscle, respectively (FIG. 6C-D).
[0207] In addition, parent-of-origin-specific DMRs overlapping four autosomal genes were identified which have not previously been identified as imprinted: Scn8a (FIG. 6E), a sodium voltage-gated channel subunit involved in membrane depolarization during the formation of action potentials in neurons; Pcdhb4, a protocadherin involved in neuronal cell-cell connections; Fry, a microtubule binding protein involved in maintaining the integrity of mitotic centrosomes; and Socs5, involved in the suppression of cytokine signaling. Furthermore, parent-of-origin-specific methylation nearby four autosomal genes were identified which are variably reported as imprinted in the literature: Zswim9. Nav2, Cascl, and Cntnapl which has been predicted as a potential imprinted gene based on a DNA sequence analysis, though this was not validated in subsequent studies performed in the brain. Additionally, although not statistically significant, a parent-of-origin-specific methylation pattern was observed in the muscle over the novel imprinted gene Scn8a, identified in the liver. A parent-of-origin-specific methylation pattern was also observed in the liver over the novel imprinted gene Socs5, identified in the muscle.
[0208] The analysis also revealed parent-of-origin-specific expression in two genes associated with these parent-of-origin-specific DMRs via promoter / enhancer overlap or genomic proximity, including the known imprinted gene Sgce (FIG. 6F). Unexpectedly, another well-documented imprinted gene, Slc38ct4, did not exhibit parent-of-origin-specific expression despite exhibiting clear parent-of-origin-specific methylation over its promoter (FIG. 6B). This finding indicates that methylation over an ICR, while necessary' to establish imprinting of certain genes, may be insufficient to cause parent-of-origin-specific expression in all cases, signifying the involvement of additional regulatory mechanisms.611626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WOEXAMPLE 9PATTERN 12: SEX-SPECIFIC DNA METHYLATION
[0209] 305 autosomal regions which exhibit sex-specific methylation patterns were also identified (FIG. 7A-B, FIG.8). Sex-specific regulator}’ effects violate Mendelian inheritance, as the inheritance of the methylation pattern and associated phenotypes occurs in a manner dependent only on the sex of the offspring and independent of the identity of the allele. Notably, methylation differences in all but one of these regions (304; 99.7%) exhibit female hypermethylation relative to males. Furthermore, all 305 of these sex-specific methylation patterns were identified in the liver, whereas none were observed in the muscle. Sex-specific DNA methylation and gene expression patterns in the liver have been previously reported both in humans and mice and, as such, were expected to be far less abundant in the muscle.
[0210] To validate the sex-specific nature of these methylation patterns. 20 candidate regions in the F2 generation were subsequently examined. The observed sex-specific methylation is established in a manner independent of the identity of local and distal autosomal genetic variants, although an X-dosage or Y chromosome effect cannot be ruled out. As such, these regions are expected to retain their sex-specific methylation patterns in the F2s. As anticipated, all 20 targeted regions exhibit methylation patterns which segregate by sex in the F2 generation (FIG. 7C), further validating that methylation in these regions is regulation by a sex-dependent regulatory mechanism.
[0211] Furthermore, the analysis identified 233 genes associated with these sex-specific DMRs via promoter / enhancer overlap or genomic proximity which exhibit concordant sex-biased gene expression (FIG. 7D) Notably, 182 (78.1%) of these genes are more highly expressed in male samples than female samples, consistent with the observed methylation patterns. This finding highlights the extensive role of sex-specific methylation in the regulation of gene expression throughout the genome.
[0212] 198 genes associated with these DMRs via direct promoter / enhancer overlap or by genomic proximity have been implicated in the etiology of diseases from the MGI databases, many of which exhibit sex disparities in their clinical features. For example, skewed sex ratios are observed in type II diabetes mellitus and nephroblastoma, associated with Pik3c2g and Pou6I2, respectively. Furthermore, distinct clinical symptoms are observed between men and women with 621626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO colorectal cancer as well as inflammatory bowel diseases, associated with Rael (for which a sexspecific DMR identified in the liver overlaps an enhancer) and Ano6, respectively. Sex-specific regulation of DNA methylation and gene expression have also been implicated in the etiology of several disorders such as non-alcoholic fatty’ liver disease (NAFLD) and hepatocellular carcinoma (HCC) as w ell as critical biological processes such as drug metabolism. Given the abundance of these sex-specific methylation patterns, their effect on disease phenotypes and presentation should be considered when examining disorders which exhibit sexual dimorphism.EXAMPLE 10X CHROMOSOMAL EPIGENETIC INHERITANCE PATTERNS
[0213] DMR finding was performed separately for the autosomes and the X chromosome. Due to expected differences betw een the male and female X chromosomal methylation patterns resulting from X chromosome inactivation (XCI), two separate analyses were performed for the X chromosome - one using only female samples and one using only the maternal alleles of male samples. For both liver and muscle, few DMRs on the maternal alleles of the male samples were identified, whereas many DMRs were found in the female samples, the majority of which were categorized as skewed XCI.
[0214] XCI is a dosage compensation mechanism in which females inactivate one copy of the X chromosome in each cell. In mice, XCI has been shown to have a genetic contribution which has been mapped to a region (QTL) on the X chromosome over which local genetics contribute to the determination of which parental copy’ of the X chromosome will be inactivated. This can lead to skewed XCI, in which one X chromosomal allele is preferentially inactivated and, thus, more highly methylated than the other allele (FIG. 7E). Skewed XCI has been previously reported among several of the CC founder strains. Over the aforementioned QTL, called Ace. CC019 originates from WSB / EiJ, which carries the A e6allele, while CC037 originates from NOD / ShiLtJ, which carries the Xce8allele. The Xce8allele has been shown to be preferentially inactivated when combined with Xceballele.
[0215] Of the 218 and 226 regions which indicate skewed XCI in the liver and muscle, respectively, 209 (95.9% and 92.5% for liver and muscle, respectively) exhibit hypermethylation of the CC037 allele relative to the CC019 allele (FIG. 7F). This result indicates a skewing of XCI towards the CC037 allele within crosses of these two strains, consistent with the expected preferential inactivation of the Xce8allele. This pattern of skewed XCI can be clearly observed from the read-level methylation data over these DMRs. in which 631626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO the CC037 allele of female Fl samples typically exhibits a much higher proportion of the inactive methylation pattern (highly methylated) than the CC019 allele. Additionally, hypermethylation of the active CC019 allele was observed over Xist and its antisense RNA Tsix, two of the primary regulators of XCI which are known to be expressed exclusively from the inactive X chromosomal allele to mediate its inactivation. Furthermore, expression analysis of genes associated with these DMRs via promoter / enhancer overlap or genomic proximity reveals that all 40 (100%) of the genes showing expression patterns consistent with skewed XCI exhibit higher expression from the CC019 allele compared to the CC037 allele, in agreement with the preferential inactivation of the CC037 X chromosomal allele (FIG. 7G).
[0216] The analysis also identified a previously unreported region with parent-of-origin-specific methylation on the X chromosome in the liver of the female samples. This region overlaps a CpG island within the gene Zfp92 (FIG. 7H), a clinically -relevant gene involved in the suppression of TEs and implicated in X-linked intellectual disability and non-obstructive azoospermia. Notably, analyses of disorders including Turner's syndrome and autism spectrum disorder have suggested the presence of an imprinted locus on the X chromosome which influences cognitive function, though such a locus had not previously been identified. Furthermore, although not statistically significant, parent-of-origin-specific methylation over this region in the muscle was also observed.EXAMPLE 11DISCUSSION
[0217] It was found that the intergenerational inhentance of DNA methylation patterns shows extensive complexity beyond what is currently recognized in mammals. In addition to the identification of abundant c / s-acting meQTLs, driving -93% of autosomal epigenetic inheritance, and the validation of known imprinted genes, many novel non-Mendelian epigenetic inheritance patterns were also found in both the liver and muscle. These include the first instance of intergenerational paramutation identified in a naturally occurring mammalian genome as well as two additional highly likely examples of paramutation identified over IAP elements, 54 emergent epigenetic inheritance patterns, widespread sex-specific methylation in the liver, and at least five new imprinted genes with one located on the X chromosome. Remarkably, at least 522 autosomal regions, or -7% of the autosomal DMRs identified, exhibited non-Mendelian patterns of epigenetic inheritance. Furthermore, 1,038 regions in which the intergenerational inheritance of DNA methylation patterns is regulated in the same 641626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO manner in both liver and muscle were identified (15.6% and 50.4% of identified regions, respectively), indicating that this regulation can be both tissue-specific as well as tissueindependent.
[0218] While the analysis identified these DNA methylation patterns using a comparison of two mouse strains, they have been identified throughout the genome of two CC strains which are mosaics of eight distinct founder strains. As such, these patterns are not limited to the alleles of a specific pair of strains. It is likely that there are many additional loci in the mouse genome which exhibit these complex epigenetic inheritance patterns but could not be discovered using only two CC strains. Despite having a smaller fraction of IBD genome between pairs of CC strains than classical laboratory’ strains, methylation could not be phased in almost one-third of the genome. Moreover, the present study has analyzed these epigenetic inheritance patterns in the context of two tissues and in a stable environment. Considerations of additional tissues and environments are likely to increase the abundance of these complex inheritance patterns.
[0219] Additionally, in large part due to a lack of transcribed polymorphisms between the two strains analyzed, only the allele-specific expression of 9.5% of genes in the genome could be analyzed. The use of ONT long-read RNA sequencing could improve the ability to quantify allele-specific expression. Between the two selected CC strains, there are 8,024 expressed genes with at least one transcribed polymorphism, whereas 4,092 (51.0%)could successfully be analyzed.
[0220] While a few of the non-Mendelian inheritance patterns identified, such as genomic imprinting and sex-specific methylation, have been observed in mice and other species, the scale and diversify’ of methylation patterns reported here is unprecedented. Furthermore, many of these non-Mendelian patterns, such as paramutation, have not been previously observed in mammals in the absence of transgenic manipulation. A significant consequence of these non-Mendelian inheritance patterns is that phenotypes associated with the methylation pattern in these regions will not necessarily track with the genotype. As described above, human phenofypes associated with many of the genes associated with these non-Mendelian epigenetic inheritance patterns exhibit unconventional patterns of intergenerational inheritance.
[0221] Similar studies performed using allele-specific epigenomic datasets from human families and populations will indicate the degree to which these patterns are conserved across species and will also indicate the extent to which these patterns influence the intergenerational inheritance of phenotypes and disorders. To date, a comprehensive analysis of epigenetic inheritance in humans has yet to be performed. This is in large part due to difficulties associated 651626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO with controlling external factors which are known to influence the epigenome, such as environmental exposures and age. in human populations when compared to model organisms. Furthermore, the increased levels of genetic variation present in human populations, an issue which does not affect model organisms for which genetics and breeding can be tightly controlled, adds significant additional variability' and complexity to the analysis. However, an analysis of methylation in human trios identified more than 6.5 million CpGs which exhibit inheritance patterns that are incompatible with Mendel’s laws. As such, studies of pedigrees with complex inheritance that do not fit traditional Mendelian patterns should also include epigenetic analyses.
[0222] Numerous mechanisms have been identified which regulate genomic imprinting, including insulators, non-coding RNAs, histone modifications. DNA methylation, and chromatin dynamics. In contrast, little is known regarding the mechanisms driving the other non-Mendelian inheritance patterns discussed here. Notably, each of the regions which were identified as exhibiting intergenerational paramutation have associations with TEs of the IAP family, with two of the three regions directly overlapping IAP elements. This suggests that the methylation-based repression of these IAPS may be mechanistically linked with the observed paramutation events. IAPs, the subclass of ERV associated with the agouti locus, are subject to complex epigenetic regulation. DNA methylation of IAP elements can be influenced not only by genetic effects but also parent-of-origin and environmental exposures. Previous work has even identified an IAP under the regulation of a polymorphic cluster of dominant transacting KRAB zinc finger proteins. Furthermore, IAPs can be protected from the genome-wide demethylation events which occur during gametogenesis and embry ogenesis. Here, presented are the first examples of IAPs associated with paramutation in naturally occurring mammalian genomes. Additionally, while IAPs are more commonly found in mice than in humans, they have been identified to be clinically relevant to human health, although methylation patterns over these human IAPs have not been previously considered. Other classes of ERVs present in humans have been shown to exhibit methylation-associated repression, which may be regulated in a similar manner leading to complex inheritance patterns such as paramutation. Identifying which mechanisms establish the other epigenetic inheritance patterns will require further investigation.
[0223] Additionally, the large number of non-Mendelian epigenetic inheritance patterns identified here have potentially important implications in evolution. In particular, emergent epigenetic inheritance patterns generate diversify in the absence of mutation and may 661626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO contribute to phenomena such as hybrid vigor and hybrid dysgenesis, as well as the emergence of other novel traits within hybrid strains. Furthermore, these emergent epigenetic inheritance patterns may provide a mechanism to support genetic assimilation, a key hypothesis of Waddington which has not yet been explained, in which the environment introduces an epigenetic perturbation that eventually becomes fixed genetically, providing a survival advantage. Genetic assimilation need not invoke Lamarckian (non-Darwinian) selection if an epigenetic change can be ‘'assimilated” by the genome, generating alleles with the abi lity to transmit the epigenetic information. Notably, 42 of the 54 (78%) autosomal emergent epigenetic inheritance patterns identified exhibited de novo methylation in at least one of the alleles of the Fl crosses. Methylated cytosines are mutagenic to C to T transitions, with mutation rates up to 3-fold higher than non-methylated bases. Thus, the "genetic fixation” aspect of Waddington’s hypothesis may be the eventual acquisition of mutated bases at the alleles which acquire this de novo methylation, preventing reversion to the original phenoty pe. This model is particularly intriguing in the context of IAPS associated with some of the non-Mendelian epigenetic inheritance patterns. IAPs are a known source of new mutations in the mouse genome and their transposition is modulated by their epigenetic status.
[0224] Finally, it would be interesting to combine the allele-specific analysis with phenoty pic measurements, as this could improve the understanding of the link between genetics, epigenetics and phenotype. The use of allele-specific expression and allele-specific DNA methylation has long been felt to hold promise by us and others, but it has not been clear how to perform this at a genome scale. As it has now been shown that epigenetic variants can be comprehensively phased across the genome, this work opens the door to Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), in which phased allele-specific DNA methylation patterns are associated with a phenotype or disorder of interest. This method could identify novel phenotype-associated genes which would be missed by traditional genome- and epigenome-wide association studies (GW AS and EWAS), techniques which independently consider only genetic or epigenetic trait associations, respectively..EXAMPLE 12PIPELINE DESCRIPTION
[0225] Brief overview of the pipeline - purpose: the initial goal in the design of this pipeline was to enable the analysis of allele-specific DNA methylation patterns in crosses of inter-strain and inter-specific crosses of model organisms. The critical challenges associated with this form 671626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO of analysis are the consideration of strain / allele-specific CpGs and the potential for the introduction of reference biases. At a high level, the aim is to create a pipeline which would process long-read Oxford Nanopore Technologies (ONT) sequencing data in a manner which is not substantially biased by the choice of a single reference genome while still allowing for the analysis of these strain / allele-specific CpGs. To do so, a novel approach which utilizes a pairwise graphical genome instead of a single reference genome was implemented. The graphical genome effectively divides the two input reference genomes into smaller, corresponding pieces which can be utilized to simplify this analysis and to enable fast and accurate coordinate mapping functionalities. Furthermore, the pairwise graphical genome was used to create the ‘pseudo-hybrid intermediate genome’, a unified reference coordinate system which includes nearly (-99.9%) all of the strain / allele-specific CpGs from both reference genomes in a manner suitable for allele-specific analysis. With this processed allele-specific methylation data, then a variety of intergenerational epigenetic inheritance patterns were identified, some of which (such as intergenerational paramutation) have not previously been identified in non-engineered mammalian genomes. Surprisingly, -10% of these inheritance patterns violate the laws of Mendelian inheritance. FIGs 9 and 10 illustrate the pipeline described herein.
[0226] Brief overview' of the pipeline - steps: PLASMA runs in two main steps. The first step needs to be run only once per set of strains / species being analyzed. In this step, heterozygous genetic variants between the two provided reference genomes are identified (once with each genome treated as the reference), a graphical genome is created between the reference genomes, and the corresponding edges of the graphical genome are aligned for coordinate mapping. In the second step, which needs to be run once for each sample, base and methylation called ONT sequenced reads are aligned to both reference genomes and phased using the heterozy gous genetic variants identified in the first step. From here, consensus haplotyping decisions (reads assigned to the same allele independent of the choice of reference genome) are extracted, enabling the subsequent extraction of allele-specific methylation data in the proper reference genome coordinate system. Finally, methylation coordinates are mapped from the two reference genomes into the 'pseudo-hybrid intermediate genome’ for analysis.
[0227] When designing the pipeline, the graphical genome already existed (the Collaborative Cross Graphical Genome), in the improved version of it, the step of creating the graphical genome has been removed. Additionally, it was designed for older sequencing 681626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO chemistries and used some older tools (e.g. Nanopolish was used for methylation calling, now Remora is the standard).
[0228] Steps involving existing tools: The following steps of the pipeline utilize existing tools: basecalling and methylation calling (tools that can be used include Nanopolish and ONT’s Guppy; or ONT’s Dorado / Remora / modkit), alignment (uses MiniMap2), phasing (uses WhatsHap), variant calling (uses MUMmer4), k-mer counting (uses Jellyfish), and analysis (uses bsseq).
[0229] Challenge: Creation of pairwise graphical genome
[0230] Solution: Graph genomes are designed to efficiently represent multiple genomes within a single data structure. Typically, graph genomes are made up of nodes, which represent common sequence between all of the genomes represented in the graph, and edges, which harbor all of the genetic diversity between these genomes. Each of the genomes which comprise the graph genome can be effectively and efficiently reconstructed by tracing that particular genome’s pre-defined path through the nodes and edges.
[0231] Within the first version of the pipeline, the data processing was built upon the use of the Collaborative Cross Graphical Genome (CCGG), a graph genome representation of the 96 Collaborative Cross (CC) strains (of which two were used). However, this limits the utilization of this pipeline to hybrid crosses of CC strains. In order to generalize the pipeline to enable the analysis of additional strains and species, a pre-processing step which creates a pairwise graph genome representation of the two reference genomes provided into the pipeline was implemented. This process involves taking the output from the k-mer counting tool Jellyfish (note that this needs to be run several times and in particular ways in order to work and, critically, to work efficiently with reasonable RAM usage). The output k-mers are further processed to identify those which are unique within both of the provided reference genomes, present on the same chromosome across the two reference genomes, and monotonically increasing in both reference genomes (meaning that they appear in the same order in both). These k-mers serve as candidate nodes. Tools which collapse and remove overlapping nodes in order to maximize the efficiency of the graph genome representation have also been implemented. This graph genome creations has now been shown to effectively work between inter-strain and inter-specific crosses of varying levels of genetic diversity (successfully tested up to 1 SNP every7~40bp on average).
[0232] Challenge: Creation of a unified coordinate system with strain-specific CpGs of both original reference genomes691626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO
[0233] Solution: The strain-specific CpGs that was anticipated to be included in the final reference genome will either be caused by reference CpGs being mutated away / in (e.g. CpG -> TpG) or by deletion of at least one of the CpG bases in one of the genomes. Mutated CpGs have coordinates which exist in both reference genomes, allowing for their analysis without the creation of a novel reference coordinate system. However, CpGs which are deleted in at least one genome will have coordinates which exist in only the non-deleted coordinate system. As such, the CpGs from both strains which have a coordinate in only one genome cannot both be analyzed simultaneously using a single existing reference genome, leading to the design of the ‘pseudo-hybrid intermediate genome’ to solve this issue.
[0234] The pairwise graph genome created by the pipeline (or the CCGG) provides a view of the genome which is divided into small pieces, typically in the hundreds to thousands of base pairs, which are much more suitable for consideration and analysis than the genome-scale. Each node in the graph genome has a know n position in both genomes, and the order of the nodes are monotonically increasing in both genomes as well. As such, edges in different genomes which are bound between the same pair of nodes are corresponding edges. As all genetic diversity of the two genomes is represented on the edges of the graph genome, all strainspecific insertions / deletions (note that these are the same, depending on the genome being considered as ‘reference’) will also be represented on these edges. Critically, the strain containing the specific insertion (rather than the deletion) will have the longer edge between these two nodes. With this knowledge, it can be iterated through the nodes of the graph genome and continuously select the longer edge, building what was deemed the ‘pseudo-hybrid intermediate genome’. This genome will have all strain-specific insertions and none of the strain-specific deletions. As such, nearly all CpG coordinates from both genomes (-99.9%) will have corresponding coordinates in this hybrid genome.
[0235] Challenge: Coordinate mapping between the reference genomes and between the reference and pseudo-hybrid intermediate genome
[0236] Solution: pairwise alignment of the corresponding edges between the two genomes was performed to create edge-specific CIGAR strings for each genome. This is made possible by the small size of the edges defined in the graph genome. Furthermore, this highly granular view of the genome provides very high-fidelity base-level mapping accuracy. The graph genome is specifically created with the end goal of mapping optimization, as the edges of the graph genome are capped in length (where possible) at the maximum size allowing for pairwise alignment of the corresponding edges (3kb). This enables consistent mapping of -99.9% of 701626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO coordinates into the intermediate reference coordinate system (mapping efficiency between the reference genome is dependent on genetic variation between the two genomes as well as genome quality). Utilizing the node positions, node sizes, and these CIGAR strings, mapping tables between the coordinate systems of both reference genomes as well as that of the intermediate reference genome can be created. These mapping tables can be utilized to efficiently map millions of coordinates between the reference and intermediate genomes.
[0237] Challenge: How to perform phasing and integrate phasing decisions from across both reference genomes
[0238] Solution: As this is the first phasing pipeline which accounts for both parental reference genomes, the phasing had to be optimized. Initially, phasing was performed using all genetic variants called from MUMmer4, however it was found that, although a slightly larger portion of the genome was phasing, there was much more variability in phasing decisions depending on the choice of reference genome (likely due to lower accuracy of ONT basecalls over indels and homopolymers). As such, now phasing is only utilizing single nucleotide polymorphisms (SNPs).
[0239] Furthermore, the pipeline integrates the phasing decisions across the two reference genomes. In order to maximize the accuracy of the pipeline, only reads which are phased to the same genome independent of the choice of reference genome wree utilized. While this again causes a slight decrease in the amount of analyzable data, it serves to minimize the introduction of reference bias and enable the required inter-allele comparisons.
[0240] Challenge: Methylation analysis with the inclusion of both shared and allele-specific CpGs.
[0241] Solution: As allele-specific methylation analyses performed using only a single reference genome do not need to consider CpGs specific to both alleles, the analysis of methylation data which includes all of these allele-specific CpGs required optimization. Methylation levels of allele-specific CpGs cannot be compared directly (because one allele cannot have methylation at the missing CpG site). However, by including these CpGs in methylation smoothing (a technique which iteratively averages the methylation values in a sliding window applied across the genome) and then removing them prior to methylation testing, their methylation values have an influence on the methylation values of nearby shared CpGs without their direct inclusion in the testing. It was previously shown that this form of methylation analysis improves statistical power for comparisons between inbred strains of mice. This is the first time this form of analysis was applied to allele-specific methylation data 711626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO and its utility in identifying not only allele-specific methylation patterns but also a wide variety of more complex epigenetic inheritance patterns was demonstrated.
[0242] Identification of non-Mendelian patterns of epigenetic inheritance: As epigenetic inheritance patterns not previously been analyzed genome-wide in a haplotype-resolved manner, the tests required to identity' each of these patterns from the phased methylation data output by the PLASMA pipeline needed to be designed. The patterns as well as the tests used to identify each pattern are described in detail in the Methods section of the manuscript. While these patterns have been previously identified, the specific testing procedure (utilizing the bsseq F-stat test w ith a complex 10-contrast design with post-testing assignment to a particular inheritance pattern utilizing 14 conditions based upon methylation differences between the various contrasts) is entirety novel.EXAMPLE 13ALLELE-SPECIFIC EPIGENOME-WIDE ASSOCIATION STUDIES (AS-EWAS)
[0243] Allele-Specific Epigenome-Wide Association Studies (AS-EWAS) in which phased allele-specific DNA methylation patterns are associated with a phenotype or disorder of interest was investigated. This method could identify novel phenotype-associated genes which would be missed by traditional genome- and epigenome-wide association studies (GWAS and EWAS), techniques which independently consider only genetic or epigenetic trait associations, respectively.
[0244] At its core, AS-EWAS is about performing a haplotype- and DNA-methylation specific analysis (see FIG. 11).
[0245] As illustrated in FIG. 11, for each type of analysis there are the same four possible combined genetic and epigenetic alleles. The disease-risk or phenotype genetic allele is in light grey, and the non-risk is in dark grey. In this case, the risk methylation state is shown as a dark lollipop and the non-risk state as a white lollipop. In GWAS, the association is with the risk allele (both light grey alleles show the ailing mouse). In the EWAS, the association is with the risk methylation state (both methylated alleles showing the ailing mouse). In AS-EWAS, the association is with the risk alleles (methylation of the light grey allele in the ailing mouse). Note that in the case of a meQTL the risk allele will virtually always be associated with the risk methylation states (second combined genetic-epigenetic variant under AS-EWAS), and the fourth combined variant in AS-EWAS above will not appear. However, if a known or unknown environmental exposure is also necessary for the phenotype or disease-associated state the 721626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO fourth variant will appear but the second variant will be specific, or highly enriched, in those individuals carrying the phenotype or disease.
[0246] For a moment, let us assume that methylation is a binary variable (methylated or unmethylated). Then, two alleles can be expanded into four states which is all combinations between methylation and DNA allele.
[0247] An analysis of the relationship between these 4 states is performed, with the addition of covari ates of interest, with the phenotype of interest. It can be done by fitting an interaction between allele and methylation and specific interest is on the loci where the interaction (and not just the two main effects) are present. This is equivalent to the following thinking: there are 4 states (two alleles by two methylation states), and each state has a certain association with the phenotype of interest. There was then an interest in loci where one (or more) of the four states have a different association from the rest. This is the same as when a GWAS is performed except ther are 4 states instead of 2 alleles. In practice, it is likely to be too restrictive to model DNA methylation as a binary variable in which case DNA methylation using a restricted cubic splice, a flexible family of functions which is easy to use will be modeled. This results in a model like:Phenotype ~ Allele * spline(Methylation) + Other Covariates.
[0248] When the interaction is present, it can take many forms. For example, there could be an effect of methylation on either allele, as long as the effect is different between alleles. Another possibility is that methylation only has an effect on one allele. Other covariates include the aforementioned cell type proportions, but also population stratification (usually handled by PCs of the genoty pe matrix), age and technical covariates. The actual model used will depend on the type of variable encoding the phenotype information, which can either be binary or continuous or possibly ordinal. For binary phenotypes logistic regression will be used, for continuous phenotypes linear regression will be used (possibly after a transformation), and for ordinary variables a proportional odds model will be used. These are all standard statistical models which can be fitted using a variety of software.
[0249] References1. Jablonka, E. & Raz, G. Transgenerational epigenetic inheritance: prevalence, mechanisms, and implications for the study of heredity and evolution. The Quarterly Review of Biology 84, 131-176 (2009).2. Rakyan, V. K. & Beck, S. Epigenetic variation and inheritance in mammals. Current Opinion in Genetics & Development 16, 573-577 (2006).731626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 3. Cooper, D. N., Krawczak, M., Polychronakos, C., Tyler-Smith, C. & Kehrer-Sawatzki, H. Where genotype is not predictive of phenotype: towards an understanding of the molecular basis of reduced penetrance inhuman inherited disease. Human Genetics 132, 1077-1130 (2013).4. Wright, C. F. Incomplete penetrance and variable expressivity: from clinical studies to population cohorts. Frontiers in Genetics 13, 920390 (2022).5. Kalyta, K. et al. The Spectrum of the Heterozy gous Effect in Biallelic Mendelian Diseases — The Symptomatic Heterozygote Issue. Genes 14, 1562 (2023).6. Anway, M. D., Cupp, A. S., Uzumcu, M. & Skinner, M. K. Epigenetic transgenerational actions of endocrine disruptors and male fertility. Science 308, 1466-1469 (2005).7. Skinner, M. K. Endocrine disruptor induction of epigenetic transgenerational inheritance of disease. Molecular and Cellular Endocrinology 398, 4-12 (2014).8. Manikkam, M., Guerrero-Bosagna, C., Tracey, R., Haque, M. M. & Skinner, M. K. Transgenerational actions of environmental compounds on reproductive disease and identification of epigenetic biomarkers of ancestral exposures. PloS One 7, e31901 (2012).9. Yeshurun, S. & Hannan, A. J. Transgenerational epigenetic influences of paternal environmental exposures on brain function and predisposition to psychiatric disorders.Molecular Psychiatry 24, 536-548 (2019).10. Bohacek. J. & Mansuy, I. M. Epigenetic inheritance of disease and disease risk.Neuropsychopharmacology 38, 220-236 (2013).11. Reik, W. & Walter, J. Genomic imprinting: parental influence on the genome. Nature Reviews Genetics 2, 21-32 (2001).12. Ferguson-Smith, A. C. Genomic imprinting: the emergence of an epigenetic paradigm. Nature Reviews Genetics 12, 565-575 (2011).13. Cockett, N. E. et al. Polar overdominance at the ovine callipyge locus. Science 273, 236-238 (1996).14. Chandler, V. L. Paramutation: from maize to mice. Cell 128, 641-645 (2007).15. Hollick, J. B. Paramutation and related phenomena in diverse species. Nature Reviews Genetics 18, 5-23 (2017).16. Rassoulzadegan, M., Magliano, M. & Cuzin, F. Transvection effects involving DNA methylation during meiosis in the mouse. The EMBO Journal 21, 440-450 (2002).17. Rodriguez, J. D. et al. A model for epigenetic inhibition via transvection in the mouse. Genetics 207, 129-138 (2017).18. Wolff, G. L., Kodell, R. L., Moore, S. R. & Cooney, C. A. Maternal epigenetics and methyl supplements affect agouti gene expression in Avy / a mice. The FASEB Journal 12, 949-957 (1998).19. Morgan, H. D., Sutherland, H. G., Martin, D. I. & Whitelaw, E. Epigenetic inheritance at the agouti locus in the mouse. Nature Genetics 23, 314-318 (1999).20. Muller, H. Recessive genes causing interspecific sterility and other disharmonies between Drosophila melanogaster and simulans. Genetics 27, 157 (1942).21. Vrana, P. B. et al. Genetic and epigenetic incompatibilities underlie hybrid dysgenesis in Peromyscus. Nature Genetics 25, 120-124 (2000).22. Wulfridge, P., Langmead, B., Feinberg, A. P. & Hansen, K. D. Analyzing whole genome bisulfite sequencing data from highly divergent genotypes. Nucleic Acids Research 47, ell7-ell7 (2019).23. Tycko, B. Allele-specific DNA methylation: beyond imprinting. Human Molecular Genetics 19, R210-R220 (2010).741626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 24. Price. J., Wright, J., Kula, K., Bowden, D. & Hart. T. A common DLX3 gene mutation is responsible for tricho-dento-osseous syndrome in Virginia and North Carolina families. Journal of Medical Genetics 35, 825-828 (1998).25. Ashcraft, K. W. & Holder, T. M. Hereditary presacral teratoma. Journal of Pediatric Surgery 9, 691-697 (1974).26. Kim, I.-S. et al. Clinical and genetic analysis of HLXB9 gene in Korean patients with Currarino syndrome. Journal of Human Genetics 52, 698-701 (2007).27. Renwick, J. Nail-patella syndrome: Evidence for modification by alleles at the main locus. Annals of Human Genetics 21, 159-169 (1956).28. Stelzer, G. et al. The GeneCards suite: from gene data mining to disease genome sequence analyses. Current Protocols in Bioinformatics 54, 1.30. 1-1.30. 33 (2016).29. Fukaya, T. & Levine, M. Transvection. Current Biolog) 27, R1047-R1049 (2017).30. Matzke, M., Matzke, A. J. & Scheid, O. M. Inactivation of repeated genes — DNA-DNA interaction? Homologous Recombination and Gene Silencing in Plants, 271-307 (1994).31. Hopmann, R., Duncan, D. & Duncan, I. Transvection in the iab-5, 6, 7 region of the bithorax complex of Drosophila: homology independent interactions in trans. Genetics 139, 815-833 (1995).32. Duncan, I. W. Transvection effects in Drosophila. Annual Review of Genetics 36, 521-556 (2002).33. Mellert, D. J. & Truman, J. W. Transvection is common throughout the Drosophila genome. Genetics 191. 1129-1141 (2012).34. Chen, J.-L. et al. Enhancer action in trans is permitted throughout the Drosophila genome. Proceedings of the National Academy of Sciences 99, 3723-3728 (2002).35. Aramayo, R. & Metzenberg, R. L. Meiotic transvection in fungi. Cell 86, 103-113 (1996).36. Chandler, V. L. & Stam, M. Chromatin conversations: mechanisms and implications of paramutation. Nature Reviews Genetics 5, 532-544 (2004).37. De Vanssay, A. et al. Paramutation in Drosophila linked to emergence of a piRNA-producing locus. Nature 490, 112-115 (2012).38. Ben- Aharon, I., Brown, P R., Shalgi, R. & Eddy, E. M. Calpain 11 is unique to mouse spermatogenic cells. Molecular Reproduction and Development: Incorporating Gamete Research 73, 767-773 (2006).39. Dear, T. & Boehm, T. Diverse mRNA expression patterns of the mouse calpain genes Capn5, Capn6 and Capnl 1 during development. Mechanisms of Development 89, 201-209 (1999).40. Dear, T. N., Moller, A. & Boehm, T. CAPN11: A calpain with high mRNA levels in testis and located on chromosome 6. Genomics 59, 243-247 (1999).41. Malcher, A. et al. Potential biomarkers of nonobstructive azoospermia identified in microarray gene expression analysis. Fertility and Sterility 100, 1686-1694. e7 (2013). 42. Wheeler, T. J. et al. Dfam: a database of repetitive DNA based on profile hidden Markov models. Nucleic Acids Research 41, D70-D82 (2012).43. Deniz, 0., Frost, J. M. & Branco, M. R. Regulation of transposable elements by DNA modifications. Nature Reviews Genetics 20, 417-431 (2019).44. Lane, N. et al. Resistance of IAPS to methylation reprogramming may provide a mechanism for epigenetic inheritance in the mouse. Genesis 35, 88-93 (2003).45. Osipovich, A. B. et al. ZFP92, a KRAB domain zinc finger protein enriched in pancreatic islets, binds to Bl / Alu SINE transposable elements and regulates retroelements and genes. PLoS Genetics 19, el010729 (2023).751626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 46. Bauer, R. Update on the Molecular Genetics of Timothy Syndrome. Frontiers in Pediatrics 9(2021).47. Alangari, A. et al. LPS-responsive beige-like anchor (LRBA) gene mutation in a family with inflammatory bowel disease and combined immunodeficiency. Journal of Allergy and Clinical Immunology 130, 481-488. e2 (2012).48. Wolf, J. B., Cheverud, J. M., Roseman, C. & Hager, R. Genome-wide analysis reveals a complex pattern of genomic imprinting in mice. PLoS Genetics 4, el000091 (2008).49. Lawson, H. A., Cheverud, J. M. & Wolf, J. B. Genomic imprinting and parent-of-origin effects on complex traits. Nature Reviews Genetics 14, 609-617 (2013).50. Wolf, J. B., Hager, R. & Cheverud, J. M. Genomic imprinting effects on complex traits. Epigenetics 3, 295-299 (2008).51. Thomas, P. Q. et al. Heterozygous HESX1 mutations associated with isolated congenital pituitary' hypoplasia and septo-optic dysplasia. Human Molecular Genetics 10, 39-45 (2001).52. Petersen, A. K., Streff, H., Tokita, M. & Bostwick, B. L. The first reported case of an inherited pathogenic CHD2 variant in a clinically affected mother and daughter. American Journal of Medical Genetics Part A 176, 1667-1669 (2018).53. Hong, K., Bjerregaard. P., Gussak. I. & Brugada, R. Short QT syndrome and atrial fibrillation caused by mutation in KCNH2. Journal of Cardiovascular Electrophysiology 16, 394-396 (2005).54. Gigante, S. et al. Using long-read sequencing to detect imprinted DNA methylation. Nucleic Acids Research 47, e46-e46 (2019).55. Xie, W. et al. Base-resolution analyses of sequence and parent-of-origin dependent DNA methylation in the mouse genome. Cell 148, 816-831 (2012).56. Hanna, C. W. et al. Pervasive polymorphic imprinted methylation in the human placenta. Genome Research 26, 756-767 (2016).57. Luedi, P. P., Hartemink, A. J. & Jirtle, R. L. Genome-wide prediction of imprinted murine genes. Genome Research 15. 875-884 (2005).58. Tuskan, R. G. et al. Real-time PCR analysis of candidate imprinted genes on mouse chromosome 11 shows balanced expression from the maternal and paternal chromosomes and strain-specific variation in expression levels. Epigenetics 3, 43-50 (2008).59. Garcia-Calzon, S., Perfilyev, A., de Mello, V. D., Pihlajamaki. J. & Ling, C. Sex differences in the methylome and transcriptome of the human liver and circulating HDL-cholesterol levels. The Journal of Clinical Endocrinology & Metabolism 103, 4395-4408 (2018).60. Oliva, M. et al. The impact of sex on gene expression across human tissues. Science 369, eaba3066 (2020).61. Zhuang, Q. K.-W. et al. Sex chromosomes and sex phenoty pe contribute to biased DNA methylation in mouse liver. Cells 9, 1436 (2020).62. Kautzky-Willer, A., Harreiter, J. & Pacini, G. Sex and gender differences in risk, pathophysiology and complications of type 2 diabetes mellitus. Endocrine Reviews 37, 278-316 (2016).63. Okbah, A. A. & Al-Shamahy, H. A. Nephroblastoma (Wilms' Tumor): Sex and Age Distribution and Correlation Rate with Ages, Sex, And Kidney Side in Sana’a City, Yemen. J. Clinical Oncology Case Reports 1(2022).64. Kim, S.-E. et al. Sex-and gender-specific disparities in colorectal cancer risk. World Journal of Gastroenterology: WJG 21, 5167 (2015).761626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 65. Goodman, W. A., Erkkila, I. P. & Pizarro, T. T. Sex matters: impact on pathogenesis, presentation and treatment of inflammatory bowel disease. Nature Reviews Gastroenterology & Hepatology 17, 740-754 (2020).66. Vachher, M., Bansal, S., Kumar, B., Yadav, S. & Burman, A. Deciphering the role of aberrant DNA methylation in NAFLD and NASH. Heliyon (2022).67. Tryndyak, V. P. et al. Non-alcoholic fatty liver disease-associated DNA methylation and gene expression alterations in the livers of Collaborative Cross mice fed an obesogenic high-fat and high-sucrose diet. Epigenetics 17, 1462-1476 (2022).68. Ye, W., Siwko, S. & Tsai, R. Y. Sex and race-related DNA methylation changes in hepatocellular carcinoma. International Journal of Molecular Sciences 22, 3820 (2021). 69. Yang, L. et al. Sex differences in the expression of drug-metabolizing and transporter genes in human liver. Journal of Drug Metabolism & Toxicology 3(2012).70. Penaloza, C. G. et al. Sex-dependent regulation of cytochrome P450 family members Cyplal, Cyp2el, and Cyp7bl by methylation of DNA. The FASEB Journal 28, 966 (2014).71. Chadwick, L. H., Pertz, L. M., Broman, K. W., Bartolomei, M. S. & Willard, H. F. Genetic control of X chromosome inactivation in mice: definition of the Xce candidate interval. Genetics 173, 2103-2110 (2006).72. Calaway. J. D. et al. Genetic architecture of skewed X inactivation in the laboratory mouse. PLoS Genetics 9, el003853 (2013).73. Sun, K. Y. et al. Bayesian modeling of skewed X inactivation in genetically diverse mice identifies a novel Xce allele associated with copy number changes. Genetics 218, iyab034 (2021).74. Augui, S., Nora, E. P. & Heard, E. Regulation of X-chromosome inactivation by the X-inactivation centre. Nature Reviews Genetics 12, 429-442 (2011).75. Schwartz, C. E. et al. X-Linked intellectual disability update 2022. American Journal of Medical Genetics Part A 191, 144-159 (2023).76. Schwartz, C., Norris, J., Harr, M., Orrico, A. & Zackai, E. Mutations in ZFP92, a novel KRAB Zine-finger protein, results in an X-linked intellectual disability and mitochondrial dysfunction disorder, in American Society of Human Genetics Annual Meeting Vol. 68 (2018).77. Zhuang, X. & Liu, P. Mutation in ZFP92 gene is associated with NOA. Fertility and Sterility 108, el 33 (2017).78. Skuse, D. H. et al. Evidence from Turner's syndrome of an imprinted X-linked locus affecting cognitive function. Nature 387, 705-708 (1997).79. Skuse, D. H. Imprinting, the X-chromosome, and the male brain: explaining sex differences in the liability to autism. Pediatric Research 47, 9-9 (2000).80. Yang, H. et al. Subspecific origin and haplotype diversity in the laboratory mouse. Nature Genetics 43, 648-655 (2011).81. Roberts, A., Pardo-Manuel de Villena, F., Wang, W., McMillan, L. & Threadgill, D. W. The polymorphism architecture of mouse genetic resources elucidated using genomewide resequencing data: implications for QTL discovery and systems genetics. Mammalian Genome 18, 473-481 (2007).82. Diez-Villanueva, A. et al. Identification of intergenerational epigenetic inheritance by whole genome DNA methylation analysis in trios. Scientific Reports 13, 21266 (2023). 83. Barlow, D. P. & Bartolomei, M. S. Genomic imprinting in mammals. Cold Spring Harbor Perspectives in Biology 6, a018382 (2014).84. Bertozzi, T. M., Elmer, J. L., Macfarlan, T. S. & Ferguson-Smith, A. C. KRAB zinc finger protein diversification drives mammalian interindividual methylation variability. Proceedings of the National Academy of Sciences 117, 31290-31300 (2020).771626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 85. Garry, R. F. et al. Detection of a human intracistemal A-type retroviral particle antigenically related to HIV. Science 250, 1127-1129 (1990).86. Sander, D. M. et al. Involvement of human intracistemal A-type retroviral particles in autoimmunity. Microscopy Research and Technique 68, 222-234 (2005).87. Turelli, P. et al. Interplay of TRIM28 and DNA methylation in controlling human endogenous retroelements. Genome Research 24, 1260-1270 (2014).88. Reiss, D., Zhang, Y. & Mager, D. L. Widely variable endogenous retroviral methylation levels in human placenta. Nucleic Acids Research 35, 4743-4754 (2007).89. Xia, J., Han, L. & Zhao, Z. Investigating the relationship of DNA methylation with mutation rate and allele frequency in the human genome. BMC genomics 13, 1-9 (2012). 90. Maksakova, I. A. et al. Retroviral elements and their hosts: insertional mutagenesis in the mouse germ line. PLoS genetics 2, e2 (2006).91. Zhou, W., Liang, G., Molloy, P. L. & Jones, P. A. DNA methylation enables transposable element-driven genome expansion. Proceedings of the National Academy of Sciences 117, 19359-19366 (2020).92. Bell, C. G. & Beck, S. Advances in the identification and analysis of allele-specific expression. Genome medicine 1, 56 (2009).93. Feinberg. A. P. Hie key role of epigenetics in human disease prevention and mitigation. New England Journal of Medicine 378, 1323-1334 (2018).94. Gimelbrant, A., Hutchinson, J. N., Thompson, B. R. & Chess, A. Widespread monoallelic expression on human autosomes. Science 318, 1136-1140 (2007).95. The Collaborative Cross, a community resource for the genetic analysis of complex traits. Nature Genetics 36, 1133-1137 (2004).96. Sigmon, J. S. et al. Content and performance of the MiniMUGA genotyping array: a new tool to improve rigor and reproducibility in mouse research. Genetics 216, 905-930 (2020).97. Ben-Moshe, S. & Itzkovitz, S. Spatial heterogeneity in the mammalian liver. Nature Reviews Gastroenterology & Hepatology 16, 395-410 (2019).98. Lai, Y. et al. Multimodal cell atlas of the ageing human skeletal muscle. Nature 629, 154-164 (2024).99. Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841-842 (2010).100. Su, H. et al. The Collaborative Cross Graphical Genome. BioRxiv. 858142 (2019). 101. Shen, W., Sipos, B. & Zhao, L. SeqKit2: A Swiss army knife for sequence and alignment processing. Imeta 3, el 91 (2024).102. Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094-3100 (2018).103. Marcais. G. et al. MUMmer4: A fast and versatile genome alignment system. PLoS Computational Biology 14, el005944 (2018).104. Martin, M. et al. WhatsHap: fast and accurate read-based phasing. BioRxiv, 085050 (2016).105. Simpson, J. T. et al. Detecting DNA cytosine methylation using nanopore sequencing. Nature Methods 14, 407-410 (2017).106. Li, H. et al. The sequence alignment / map format and SAMtools. Bioinformatics 25, 2078-2079 (2009).107. Telatin, A., Fariselli, P. & Birolo, G. SeqFu: a suite of utilities for the robust and reproducible manipulation of sequence files. Bioengineering 8, 59 (2021).108. Cock, P. J. et al. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics 25, 1422 (2009).781626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 109. Lawrence, M. el al. Software for computing and annotating genomic ranges. PLoS Computational Biology 9, el003118 (2013).110. Halter, K, Chen, J., Priklopil, T., Monfort, A. & Wutz, A. Cdk8 and Hira mutations trigger X chromosome elimination in naive female hybrid mouse embryonic stem cells. Chromosome Research 32, 12 (2024).111. Hansen, K. D., Langmead, B. & Irizarry, R. A. BSmooth: from whole genome bisulfite sequencing reads to differentially methylated regions. Genome Biology 13, 1-10 (2012). 112. Ziller, M. J., Hansen, K. D., Meissner, A. & Aryee, M. J. Coverage recommendations for methylation analysis by whole-genome bisulfite sequencing. Nature methods 12, 230-232 (2015).113. Hansen, K. D. et al. Increased methylation variation in epigenetic domains across cancer types. Nature Genetics 43, 768-775 (2011).114. Robinson, J. T. et al. Integrative genomics viewer. Nature Biotechnology 29, 24-26 (2011).115. seqtk, Toolkit for processing sequences in FASTA / Q formats.116. Korthauer, K., Chakraborty. S., Benjamini, Y. & Irizarry, R. A. Detection and accurate false discovery' rate control of differentially methylated regions from whole genome bisulfite sequencing. Biostatistics 20. 367-383 (2019).117. Cavalcante, R. G. & Sartor, M. A. Annotatr: genomic regions in context.Bioinformatics 33, 2381-2383 (2017).118. Durinck, S. et al. BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis. Bioinformatics 21, 3439-3440 (2005).119. Gao, T. & Qian, J. EnhancerAtlas 2.0: an updated resource with enhancer annotation in 586 tissue / cell types across nine species. Nucleic Acids Research 48. D58-D64 (2020). 120. Wu, E. Y. et al. SEESAW: detecting isoform-level allelic imbalance accounting for inferential uncertainty. Genome Biology 24, 165 (2023).121. Pertea, G. & Pertea, M. GFF utilities: GffRead and GffCompare. FlOOOResearch 9(2020).122. Patro, R., Duggal, G., Love, M. I., Irizarry, R. A. & Kingsford, C. Salmon provides fast and bias-aware quantification of transcript expression. Nature methods 14, 417-419 (2017).123. Zhu, A., Srivastava, A., Ibrahim, J. G., Patro, R. & Love, M. I. Nonparametric expression analysis using inferential replicate counts. Nucleic Acids Research 47, e!05-e!05 (2019).124. Love, M. I., Huber, W. & Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biology 15, 1-21 (2014).125. Blake, J. A. et al. Mouse Genome Database (MGD): knowledgebase for mousehuman comparative biology. Nucleic Acids Research 49, D981-D987 (2021).126. Cheetham, S. W., Kindlova, M. & Ewing, A. D. Methylartist: tools for visualizing modified bases from nanopore sequence data. Bioinformatics 38, 3109-3112 (2022).
[0250] Although the invention has been described with reference to the above examples, it will be understood that modifications and variations are encompassed within the spirit and scope of the invention. Accordingly, the invention is limited only by the following claims.791626378753 1
Claims
PATENT ATTORNEY DOCKET NO. JHU4760-1WO What is claimed is:
1. A method for screening a subject for susceptibility to a disease or condition comprising:detecting a methylation status of a plurality of nucleic acid sequences in a sample from the subject, wherein detecting the methylation status comprises:(i) analyzing inheritance of epigenetic patterns in a nucleic acid sample from the subject,(ii) identifying the subject as having an increased risk of developing the disease or condition.
2. The method of claim 1, further comprising detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences.
3. The method of claim 2, wherein combining the genetic variations and the methylation status of the plurality of nucleic acid sequences comprises using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genome-wide association studies (GW AS) and epigenome-wide association studies (EWAS).
4. The method of claim 1, wherein analyzing inheritance of epigenetic patterns comprises identifying Mendelian and non-Mendelian patterns of epigenetic inheritance.
5. The method of claim 4, wherein non-Mendelian patterns of epigenetic inheritance comprises non-dominant trans-acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation.
6. The method of claim 1, wherein analyzing comprises using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome-wide in a nucleic acid sample from the subject.
7. The method of claim 6, wherein the long-read sequencing comprises genetic polymorphisms and methylation status.801626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 8. The method of claim 6. wherein the long-read sequencing comprises long-read nanopore sequencing.
9. The method of claim 1, wherein the subject is a mammal.
10. The method of claim 9, wherein the subject is human.
11. The method of claim 9, wherein the mammal is a cattle, a sheep, or a horse.
12. A method for generating a prognosis of a subject's susceptibility to a disease or condition comprising:(i) analyzing inheritance of epigenetic patterns,(ii) identifying the subject as having an increased risk of developing the disease or condition, wherein non-Mendelian epigenetic inheritance comprises emergent epigenetic patterns, and paramutation,(iii) storing Mendelian and non-Mendelian epigenetic inheritance data in a database that includes a set of information related to said subject;(iv) correlating the data with an association between the data and susceptibility to the disease or condition in the database;(v) generating a prognosis of the subject's susceptibility to the disease or condition;and(vi) communicating the prognosis of susceptibility to a medical practitioner.
13. The method of claim 12, wherein the set of information related to said subject comprises family medical history, diet, exercise and medical history of said subject.
14. The method of claim 12, wherein analyzing inheritance of epigenetic patterns comprises identifying Mendelian and non-Mendelian patterns of epigenetic inheritance.
15. The method of claim 14, wherein non-Mendelian epigenetic inheritance comprises non-dominant trans- acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation.
16. The method of claim 12, wherein analyzing comprises using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome-wide in a nucleic acid sample from the subject.811626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 17. The method of claim 16, further comprising detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences.
18. The method of claim 17, wherein combining the genetic variations and the methylation status of the plurality of nucleic acid sequences comprises using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS). thereby identifying novel phenotype-associated genes missed by traditional genome-wide association studies (GWAS) and epigenome-wide association studies (EWAS).
19. The method of claim 1, wherein the long-read sequencing comprises genetic polymorphisms and methylation status.
20. The method of claim 16, wherein the long-read sequencing comprises long-read nanopore sequencing.
21. The method of claim 12, wherein the subject is a mammal.
22. The method of claim 21, wherein the subject is human.
23. The method of claim 21, wherein the mammal is a cattle, a sheep, or a horse.
24. A system for screening a subject for susceptibility to a disease or condition comprising:one or more processors; anda program having programming instructions stored thereon, which, when executed by the one or more processors, causes the system to perform operations comprising:detecting a methylation status of a plurality of nucleic acid sequences in a sample from the subject, wherein detecting the methylation status comprises:(i) analyzing inheritance of epigenetic patterns in a nucleic acid sample from the subject, and(ii) identifying the subject as having an increased risk of developing the disease or condition.
25. The system of claim 24, wherein the operations further comprise:detecting genetic variations of a plurality of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences.821626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 26. The system of claim 24, wherein combining the genetic variations and the methylation status of the plurality of nucleic acid sequences comprises using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genome-wide association studies (GWAS) and epigenome-wide association studies (EWAS).
27. The system of claim 24, wherein analyzing inheritance of epigenetic patterns comprises:identifying Mendelian and non-Mendelian patterns of epigenetic inheritance.
28. The system of claim 27, wherein non-Mendelian epigenetic inheritance comprises non-dominant trans- acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation.
29. The system of claim 24, wherein analyzing the inheritance of epigenetic patterns in the nucleic acid sample from the subject comprises:using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome- wide in a nucleic acid sample from the subject.
30. The system of claim 29, wherein the long-read sequencing includes genetic polymorphisms and methylation status.
31. The system of claim 29, wherein the long-read sequencing comprises long-read nanopore sequencing.
32. The system of claim 24, wherein the subject is a mammal.
33. The system of claim 32, wherein the subject is human.
34. The system of claim 32, wherein the mammal is a cattle, a sheep, or a horse.
35. A system for generating a prognosis of a subject's susceptibility to a disease or condition comprising:one or more processors; anda program having programming instructions stored thereon, which, when executed by the one or more processors, causes the system to perform operations comprising(i) analyzing inheritance of epigenetic patterns,831626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO (ii) identifying the subject as having an increased risk of developing the disease or condition, wherein non-Mendelian epigenetic inheritance comprises non-dominant trans-acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation, (iii) storing Mendelian and non-Mendelian epigenetic inheritance data in a database that includes a set of information related to said subject,(iv) correlating the data with an association between the data and susceptibility to the disease or condition in the database,(v) generating a prognosis of the subject's susceptibility to the disease or condition, and(vi) causing display of an indication of the prognosis of susceptibility to a medical practitioner.
36. The system of claim 35, wherein the set of information related to said subject comprises family medical history, diet, exercise and medical history’ of said subject.37 The system of claim 35, wherein analyzing the inheritance of epigenetic patterns comprises:identifying Mendelian and non-Mendelian patterns of epigenetic inheritance.
38. The system of claim 37, wherein non-Mendelian epigenetic inheritance comprises non-dominant trans-acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation.
39. The system of claim 35, wherein analyzing the inheritance of epigenetic patterns comprises:using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome- wide in a nucleic acid sample from the subject.
40. The system of claim 39, wherein the operations further comprise:detecting genetic variations of a plurality’ of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality’ of nucleic acid sequences.841626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 41. The system of claim 40, wherein combining the genetic variations and the methylation status of the plurality of nucleic acid sequences comprises using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genome-wide association studies (GWAS) and epigenome-wide association studies (EWAS).
42. The system of claim 39, wherein the long-read sequencing includes genetic polymorphisms and methylation status.
43. The system of claim 42, wherein the long-read sequencing comprises long-read nanopore sequencing.
44. The system of claim 35, wherein the subject is a mammal.
45. The system of claim 44, wherein the subject is human.
46. The system of claim 44, wherein the mammal is a cattle, a sheep, or a horse.
47. A system for analyzing allele-specific DNA methylation patterns in in crosses of inter-strain and inter-specific crosses of organisms comprising:one or more processors; anda memory having programming instructions stored thereon, which, when executed by the one or more processors, causes the system to perform operations comprising:(a) generating a pairwise graphical genome betw een two reference genomes;(b) creating a pseudo-hybrid intermediate genome, wherein the pseudo-hybrid intermediate genome is a unified reference coordinate system including strain / allele-specific CpGs from both reference genomes;(c) aligning the pseudo-hybrid intermediate genome and the corresponding edges of the graphical genome for coordinate mapping;(d) identifying heterozygous genetic variants between the two provided reference genomes;(e) extracting consensus haplotyping decisions; and(1) mapping methylation coordinate from the two reference genomes into the ‘pseudohybrid intermediate genome,thereby analyzing allele-specific DNA methylation patterns.851626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 48. The system of claim 47, wherein generating the pairwise graphical genome comprises dividing the two input reference genomes into smaller, corresponding pieces.
49. The system of claim 47, wherein the operations further comprise processing long-read sequencing data.
50. The system of claim 49, wherein the long-read sequencing data comprises Oxford Nanopore Technologies (ONT) sequencing data.
51. The system of claim 47, wherein analyzing allele-specific DNA methylation patterns comprises identifying intergenerational epigenetic inheritance patterns.
52. The system of claim 51, wherein the intergenerational epigenetic inheritance patterns include intergenerational paramutation.
53. The system of claim 51, wherein the inheritance patterns violate the laws of Mendelian inheritance.
54. The system of claim 47, wherein the organism is a mammal.
55. The system of claim 54, wherein the mammal is a human.
56. The system of claim 54, wherein the mammal is a cattle, a sheep, or a horse.
57. The system of claim 47, wherein the organism is a model organism.
58. The system of claim 47, wherein the references genomes are non-engineered mammalian genomes.
59. A method for generating a prognosis of a subject's susceptibility to a disease or condition comprising:(i) analyzing non-Mendelian patterns of epigenetic inheritance,(ii) identifying the human subject as having an increased risk of developing the disease or condition,(iii) storing the non-Mendelian epigenetic inheritance data in a database that includes a set of information related to said subject;(iv) correlating the data with an association between the data and susceptibility to the disease or condition in the database;(v) generating a prognosis of the subject's susceptibility to the disease or condition;and861626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO (vi) communicating the prognosis of susceptibility to a medical practitioner.
60. The method of claim 59, wherein the set of information related to said subject comprises family medical history, diet, exercise and medical history of said subject.
61. The method of claim 59, wherein the non-Mendelian epigenetic inheritance comprises non-dominant trans- acting meQTLs; imprinted genes; sex-specific methylation patterns; emergent epigenetic patterns including overdominance, underdominance, polar overdominance, polar underdominance, and bipolar dominance; and paramutation.
62. The method of claim 59, wherein analyzing comprises using long-read sequencing in conjunction with analysis of allele specific methylation patterns genome-wide in a nucleic acid sample from the subject.
63. The method of claim 62, further comprising detecting genetic variations of a plurality' of nucleic acid sequences in the sample from the subject and combining the genetic variations and the methylation status of the plurality of nucleic acid sequences.
64. The method of claim 63, wherein combining the genetic variations and the methylation status of the plurality of nucleic acid sequences comprises using Allele-Specific Epigenome-Wide Association Studies (AS-EWAS), thereby identifying novel phenotype-associated genes missed by traditional genome-wide association studies (GWAS) and epigenome-wide association studies (EWAS).
65. The method of claim 62, wherein the long-read sequencing includes genetic polymorphisms and methylation status.
66. The method of claim 65, wherein the long-read sequencing comprises long-read nanopore sequencing.
67. The method of claim 59, wherein the subject does not present the phenotype of a disease or disorder but is known for having inherited a mutation associated with a Mendelian disease or disorder from a family member affected with the Mendelian disease or disorder.
68. The method of claim 59, wherein the subject present the phenotype of a disease or disorder but is know n for not having inherited a mutation associated with a Mendelian disease or disorder from a family member affected with the Mendelian disease or disorder.
69. The method of claim 59, wherein the subject is a mammal.871626378753 1PATENT ATTORNEY DOCKET NO. JHU4760-1WO 70. The method of claim 69, wherein the subject is human.
71. The method of claim 69, wherein the mammal is a cattle, a sheep, or a horse.881626378753 1
Citation Information
Patent Citations
DNA methylation analysis to identify cell type
US12110559B2
Compositions and methods for detecting predisposition to cardiovascular disease
US20240360513A1
Methods to identify structural variations that cause diseases and the regions to repair with gene editing
WO2019226951A1
Fragmentation for measuring methylation and disease
WO2023147783A1
Methods for identifying risk of autism
WO2023164243A1