Methods for identifying epigenetic modifications in sperm, diagnosing infertility, and identifying treatment strategies
By using panel probe sets and targeted enrichment solid supports to assess H3K4me3 enrichment in sperm, the method addresses the inadequacies of current sperm health assessment and infertility diagnosis, providing a more accurate and effective approach to improving fertility outcomes.
Patent Information
- Application Number
- PCT/CA2024/051563
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
Current methods for assessing sperm health and diagnosing male infertility are outdated and ineffective, failing to accurately predict the clinical outcome of assisted reproductive therapies and identify appropriate treatment strategies.
The development of panel probe sets and targeted enrichment solid supports that assess sperm health by identifying enrichment differences in histone H3 lysine 4 trimethylation (H3K4me3) at specific genomic regions, allowing for the discrimination between fertile and infertile men and guiding treatment or lifestyle changes.
This approach enables more accurate diagnosis of infertility, predicts clinical outcomes of assisted reproductive therapies, and identifies effective treatment strategies, thereby improving fertility outcomes for couples seeking treatment.
Smart Images

Figure CA2024051563_30052025_PF_FP_ABST
Abstract
Description
METHODS FOR IDENTIFYING EPIGENETIC MODIFICATIONS IN SPERM, DIAGNOSING INFERTILITY, AND IDENTIFYING TREATMENT STRATEGIESRELATED FAMILY DETAIL
[0001] This application claims the benefit of priority of U.S. Provisional Patent Application 63 / 602,077 filed November 22, 2023, which is incorporated herein in its entirety.FIELD
[0002] The present disclosure relates to methods for assessing sperm health and fertility in a subject. Also disclosed herein are panel probe sets and targeted enrichment solid supports comprising the panel probe sets for assessing sperm health, screening for male infertility, assessing environmental exposures and predicting the clinical outcome of assisted reproductive therapy, and / or identifying treatment strategies that are likely to improve clinical outcomes.INTRODUCTION
[0003] Experts in male fertility have raised the alarm that we are facing a potential crisis. Sperm counts have reportedly declined by 50% in the last 40 years with low levels causing infertility in 1 / 20 young men. This decline is due to factors such as sedentary lifestyle, poor nutrient intake, being overweight, exposure to chemicals, and alcohol / tobacco / cannabis use. Genetic factors are implicated in ~15% of infertile men, with half of infertile men having normal semen analysis and are therefore categorized as idiopathic. Infertility affects 17% of couples seeking to have children in Canada (Statistics Canada) and the USA. Assisted reproduction technology (ART) such as in vitro fertilization (IVF) costs about CAD$13K per cycle with an average of 4 cycles needed to achieve a pregnancy (Canadian Fertility and Andrology Society). The number of treatment cycles has increased 10% from 2004-2013, with about 50-75%% of cycles not resulting in pregnancy. The emotional and financial toll of using ART is profound with couples spending as much as 30% of their annual income on ART. Male factor infertility has been the primary reason couples sought infertility treatment in Canada (CARTR Annual Report 2019). Yet despite the infertility industry being worth $28 billion and predicted to rise 8-10% a year, the essential role of men has been forgotten. This is evidenced by the lack of progress in methods of semen analysis and diagnosing male infertility which have not advanced in >50 yrs (World Health Organization, standard clinical analysis) and are based on sperm count, motility and morphology. This emphasizes the urgent need to develop new technology for the identification of fertility in men.
[0004] Sperm is a highly unique cell as it is the only cell that leaves the body, enters the female reproductive tract to penetrate the oocyte and deliver its epigenetic and genetic material. To acquire these unique functions, it undergoes a tightly controlled transcriptional program that facilitates the complex cell differentiation process. The packaging of the genetic and epigenetic material forthe safe transport to the oocyte results in a chromatin structure that is unlike any other cell. During mammalian spermatogenesis, a dramatic chromatin remodeling occurs whereby testis-specific histone variants are incorporated throughout meiosis and the vast majority of nucleosomes are evicted and replaced by protamines in the spermatid (Kimmins & Sassone- Corsi, 2005). Interestingly, sperm retain about 1 % and 15% of nucleosomes in mice and men respectively(Tanphaichitr et al., 1978; Balhorn et al., 1977; Jung et al., 2017). The nucleosomes that remain contain histones that can be modified on their n-terminus by acetylation and methylation, among other modifications. Depending on the modification it can act as either gene activating or silencing. The inventors have focused on histone H3K4me3 (histone3 lysine 4 tri-methylation) in sperm, as they and others have shown that it is responsive to the environment, including diet, toxicants and obesity. It also associated with reproductive outcomes and male fertility (Lismer and Kimmins 2023). Recent studies in mice have assessed histone H3 lysine 4 methylation (H3K4me) and early gene expression in the developing embryo (Siklenka et al., 2015; Liu et al., 2016; Dahl et al., 2016; Zhang et al., 2016). Liu et al, 2016 demonstrated that H3K4me3 is associated with zygotic genome activation and is found at transcription start sites (TSS) in early embryos, and Siklenka et al., 2015 and Lismer et al., 2020 showed that changes in the amount of methylation on H3K4me2 / 3 in sperm can alter gene expression in the embryo and result in abnormal offspring development. Lismer et al., 2021 demonstrated in a mouse model that sperm H3K4me3 is transmitted to the embryo to alter gene expression and embryo development.
[0005] Diet and physical activity are key modifiable lifestyle factors that affect weight and male fertility(Maleki & Tartibian, 2017a; Maleki & Tartibian, 2018; Maleki & Tartibian, 2017b; Kimmins et al. 2023). Men make a million sperm with every heartbeat and it takes 3 months to progress from a stem cell to a mature sperm, providing a 3-month window to optimize and improve fertility. Emerging evidence indicates that paternal diets high in fat or low in folate or protein, can alter the phenotype in offspring, including metabolism (Carone et al. 2010; Lambrot et al. 2013; Ng et al. 2010; Pepin et al., 2022) through alterations of the sperm epigenome. Lismer et al. (2021) show that H3K4me3 is altered at developmental genes and putative enhancers in the sperm of mice exposed to a folate-deficient diet, and that a subset of H3K4me3 alterations are retained in the preimplantation embryo and are associated with deregulated embryonic gene expression.
[0006] Thus there is a need for methods for assessing sperm health and infertility in men, predicting the clinical outcome of assistive reproductive therapies, and the identification of treatment strategies that are likely to improve clinical outcomes (Kimmins et al. 2023). Environmental exposures (Diet, BMI and toxicants) can all alter the sperm, H3K4me3 indicating it is an environmental sensor.SUMMARY
[0007] The inventors have identified enrichment of H3K4me3 at specific regions of the genome in human sperm. In semen samples, these regions are shown herein to be associated with infertility in humans. For example, it was found, as demonstrated herein, that purified sperm from overweight or obese infertile men show differential H3K4me3 enrichment at a plurality of genomic regions. These regions comprise a molecular signature in sperm that the inventors identified as being able to for example discriminate in overweight men, fertile men from men with idiopathic infertility (men with normal semen analysis according to WHO criteria with unexplained infertility). These regions were used to identify fertility profiles . Of these, a subset were selected as targets for a panel probe set. As further demonstrated herein, enrichment differences in H3K4me3 were identified at 8979 regions in sperm from men that were either fertile sperm donors or infertile (defined as havingone abnormal parameter identified by standard semen analysis, failed IVF, and / or with poor embryo development, and no female factors identified). Of these, based on enrichment differences, additional regions were selected as targets for the panel probe set (total targets related to infertility on the panel =482). As described herein, a total of 18 regions were selected as internal controls. These regions (or targets) are listed in Table 1 and Table 2, respectively.
[0008] In some embodiments, the probes are labelled.
[0009] The probes of the panel probe sets can be attached to a solid support, for example as a targeted enrichment solid support. The panel probe sets with or without attachment to a solid support can for example be used to profile semen samples (e.g. profile sperm chromatin) or to diagnose infertility and guide treatment or lifestyle changes in men.
[0010] Also described herein are sperm customized methods for ChlP-seq and a bioinformatic pipeline that utilizes for example a reference population data set. The Examples section shown herein demonstrate the feasibility of using the sperm epigenome to improve diagnosis and treatment guidance for couples undergoing ART (assisted reproduction technology). As shown herein, a plurality of regions were identified as being differentially enriched for H3K4me3 (deH3K4me3) between overweight or obese fertile and infertile men, and 8979 regions were identified as being differentially enriched between fertile and infertile men. A subset of these regions can be used, for example those that show the greatest differences in enrichment between fertile and infertile sperm, and / or localize to genes or functional genomic regions that are associated with fertility, developmental processes or disease. A number of genomic regions highly enriched or lowly enriched for H3K4me3 in sperm that are not altered in any of the groups studied in association with fertility status (e.g. that do not change with fertility) can be used to serve as internal quality controls.
[0011] An aspect includes a method for profiling sperm or determining sperm histone 3 lysine 4 trimethyl (H3K4me3) nucleosome methylation levels in a subject, the method comprising: a) obtaining a first plurality of sequence reads from a semen sample or a sperm sample obtained from the subject, wherein the sequence reads are from H3K4me3 nucleosome fragments; and b) determining a H3K4me3 level at each of a plurality of genomic regions using the first plurality of sequence reads, the plurality of genomic regions comprising at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 482, or any number in between and including 5 and 482 genomic regions identified in Table 1 , and optionally at least one genomic region selected from Table 2.
[0012] In various embodiments of the aspects described herein, the plurality of genomic regions comprises: a) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 153, or any number in between and including 5 and 153 genomic regions that are categorized in table 1 as Category 1 ; b) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 59, or any number in between and including 5 and 59 genomic regions that are categorized in table 1 as Category 2; c) at least 5, 10, 15, 20, 25, 30, 35, 40, or any number in between and including 5 and 40 genomic regions that are categorized in table 1 as Category 3; d) at least 5, 10, 13, or anynumber in between and including 5 and 13 genomic regions that are categorized in table 1 as Category 4; e) at least 5, 10, 15, 20, 25, 30, 35, 40, 41 , or any number in between and including 5 and 41 genomic regions that are categorized in table 1 as Category 5; f) at least 5 or 6 genomic regions that are categorized in table 1 as Category 6; g) at least 5, 9, or any number in between and including 5 and 9 genomic regions that are categorized in table 1 as Category 7; h) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 51 , or any number in between and including 5 and 51 genomic regions that are categorized in table 1 as Category 8; i) at least 5, 10, 15, 20, 25, 29, or any number in between and including 5 and 29 genomic regions that are categorized in table 1 as Category 9; j) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 68, or any number in between and including 5 and 68 genomic regions that are categorized in table 1 as Category 10; k) at least 5, 10, 13, or any number in between and including 5 and 13 genomic regions that are categorized in table 1 as Category 11 ; optionally I) at least 1 , at least 2, at least 3, at least 4, at least 5, or 6 genomic regions that are categorized in table 2 as Category A; m) at least 1 , at least 2, at least 3, at least 4, at least 5, or 6 genomic regions that are categorized in table 2 as Category B; and / or n) at least 1 , at least 2, at least 3, at least 4, at least 5, or 6 genomic regions that are categorized in table 2 as Category C. Category A in some embodiments are excluded.
[0013] In an embodiment, the plurality of genomic regions comprises: at least 5 genomic regions that are categorized in table 1 as Category 1 ; at least 5 genomic regions that are categorized in table 1 as Category 2; at least 5 genomic regions that are categorized in table 1 as Category 3; at least 5 genomic regions that are categorized in table 1 as Category 4; at least 5 genomic regions that are categorized in table 1 as Category 5; at least 5 genomic regions that are categorized in table 1 as Category 6; at least 5 genomic regions that are categorized in table 1 as Category 7; at least 5 genomic regions that are categorized in table 1 as Category 8; at least 5 genomic regions that are categorized in table 1 as Category 9; at least 5 genomic regions that are categorized in table 1 as Category 10; at least 5 genomic regions that are categorized in table 1 as Category 11 ; and optionally at least 5 genomic regions that are categorized in table 2 as Category C, and at least 2 genomic regions that are categorized in table 2 as Category B.
[0014] In an embodiment, the sperm sample is at least 99% purified sperm. In an embodiment the sperm sample that is at least 99% purified sperm is obtained by a method comprising the steps of a) obtaining a semen sample comprising sperm from the subject, b) pelleting the sperm, c) resuspending the sperm in buffer, d) flash-freezing the resuspended sperm, e) thawing the sperm, f) determining the purity of the sperm, and if necessary g) repeating steps b) to f) until the sample is at least 99% sperm. In an embodiment, the semen sample is a neat semen sample. In an embodiment, the semen sample or sperm sample comprises or is mixed with a cryoprotectant. In an embodiment, the semen sample or sperm sample is a fresh sample (e.g. unfrozen). In an embodiment, the semen sample or sperm is a frozen sample, optionally with or without cryoprotectant. In an embodiment, the histone 3 lysine 4 trimethylated (H3K4me3) nucleosome fragments are mono-nucleosomal fragments. In some embodiments, the mono-nuclesomal fragments are isolated by a method comprising one or more of the following steps: obtaining a semen sample comprising sperm from the subject or that has been obtained from the subject; purifying the semen sample to obtain a sperm sample, treating the sperm samplewith MNase to obtain substantially mono-nucleosome fragments (e.g. wherein at least 90% or at least 95% are mononucleosome fragments); and precipitating H3K4me3 nucleosome fragments using an H3K4me3-specific antibody optionally in combination with a magnetic antibody capture agent. In an embodiment, the H3K4me3- specific antibody is C42D8, available from Cell Signaling Technology (Danvers, MA, USA). In an embodiment, the magnetic antibody capture agent is a magnetic bead. In an embodiment, the method further comprises purifying DNA from the H3K4me3 nucleosome fragments and subjecting the purified DNA to a size-selection step.
[0015] In an embodiment, obtaining the sequence reads (e.g. the read counts) comprises assaying by next-generation sequencing (NGS) and optionally the method further comprises generating a sequencing library and polymerase chain reaction (PCR)-amplifying the library prior to sequencing.
[0016] In an embodiment, the method further comprises targeted enrichment of H3K4me3 nucleosome fragments from the plurality of genomic regions.
[0017] In an embodiment, the targeted enrichment comprises hybridization capture-based enrichment, for example using a panel probe set or targeted enrichment solid support described herein.
[0018] In an embodiment, the subject is a sperm donor. The methods can be used to select a sperm donor, for example for intrauterine insemination (IUI) or in vitro fertilization (IVF).
[0019] An aspect includes a method for predicting if a subject is infertile, the method comprising: a) determining the H3K4me3 nucleosome methylation levels in the subject according to an aspect described herein; and b) using the determined H3K4me3 nucleosome methylation levels to predict if the subject is infertile. In an embodiment, step b) comprises i) comparing the H3K4me3 level at each genomic region of the plurality of genomic regions to a control H3K4me3 level or reference value for each respective genomic region in the plurality of genomic regions, wherein the H3K4me3 level at one or more regions relative to control H3K4me3 level or reference value is predictive of infertility; and / or ii) calculating the risk of infertility of the subject based on statistical modeling; and optionally classifying or identifying the subject as likely fertile or infertile. The method can for example comprise assigning a fertility score.
[0020] In various embodiments of the aspects described herein, predicting if a subject is infertile may further comprise semen analysis techniques which measure semen volume, pH, sperm count, sperm motility, and / or sperm morphology.
[0021] A further aspect includes a method for assessing sperm epigenomic health of a subject, wherein the sperm epigenome is assessed at the level of histone H3K4me3, the method comprising: a) determining the H3K4me3 nucleosome methylation levels in the subject according to an aspect described herein; and b) using the determined H3K4me3 nucleosome methylation levels to assess the sperm epigenomic health of the subject. In an embodiment, step b) comprises i) comparing the H3K4me3 level at each genomic region of the plurality of genomic regions to a control H3K4me3 level or reference value for each respective genomic region in the plurality of genomic regions, wherein the H3K4me3 level at one or more regions relativeto control H3K4me3 level or reference value is indicative of sperm epigenomic health; and / or ii) determining the sperm epigenomic health of the subject based on statistical modeling.
[0022] In various embodiments of the aspects described herein, the control H3K4me3 level or reference value for each respective genomic region in the plurality of genomic regions is obtained from a reference population, optionally a fertile reference population and / or infertile reference population.
[0023] In various embodiments of the aspects described herein, assessing sperm epigenomic health of the subject comprises one or more of: identifying or classifying the subject as fertile or infertile; predicting the clinical outcome of assisted reproductive therapy (ART); predicting the clinical outcome of embryo development; and / or predicting the clinical outcome of lifestyle intervention on fertility, ART outcome, and / or embryo development, optionally lifestyle intervention comprises cessation of cannabis use, dietary changes and / or weight loss.
[0024] In various embodiments of the aspects described herein, the H3K4me3 nucleosome methylation levels and / or the sperm epigenomic health of the subject is compared to H3K4me3 nucleosome methylation levels and / or sperm epigenomic health determined previously, optionally 3-6 months prior.
[0025] In various embodiments of the aspects described herein, assessing sperm epigenomic health of the subject may be carried out in conjunction with semen analysis techniques which measure semen volume, pH, sperm count, sperm motility, and / or sperm morphology.
[0026] Another aspect includes a panel probe set which can for example be used in a method described herein. In an embodiment, the probe set comprises at least one or at least two probes for each of a plurality of genomic regions, the plurality of genomic regions comprising at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 482, or any number in be-tween and including 5 and 482 of the genomic regions identified in Table 1 and optionally at least two probes for each of at least one of the genomic regions identified in Table 2.
[0027] In an embodiment, the panel probe set comprises at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11 , at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 probes for each of the plurality of genomic regions. The probes of the panel probe set can for example hybridize to discrete segments (e.g. be non-overlapping) of the genomic region. The probes can hybridize in a tiling fashion or interspaced within the genomic region or a portion thereof. In some embodiments, the probes for a genomic region hybridize within a portion of the genomic region (e.g. a capture area) of about 400 bp, 500 bp or 600 bp (e.g. about 400 bp, 500 bp or 600 bp of a genomic region listed in Table 1 and / or 2). The capture area can also be larger, for example 1 kb, 2 kb, or 3 kb or the entire genomic region as listed in Table 1 and / or 2. The capture area can be proximal to or overlapping of a transcription start site (TSS), for example it can be 400 bp, 500 bp or 600 bp upstream or downstream of a TSS or a portion of the capture area may be upstream and a portion may be downstream.
[0028] In an embodiment, each or a plurality of the probes (e.g. the hybridization portion of each thereof) is 100-150 nucleotides, 120-150 nucleotides, 130-150 nucleotides, 140-150 nucleotides, or any number in between and including 100-150 nucleotides, optionally 140 nucleotides in length or about 140 nucleotides or about 145 nucleotides in length.
[0029] In an embodiment, each or a plurality of the probes of comprises one or more labels. A label can for example be additional nucleotides or a non-nucleotide moiety, such as radiolabel or a small molecule. The additional nucleotides and may for example be 10-20 nucleotides. They may be added at one or both ends of a probe.
[0030] A label can for example be biotin, for example one or more biotin molecules, which is / are introduced for example using a biotinylated analog of a nucleotide (e.g. biotin dUTP label) and allows purification using beads such as Streptavidin Magnetic Dynabeads (Thermofisher).
[0031] An additional aspect includes a targeted capture solid support comprising a solid support and the panel probe set. The panel probe set can be immobilized on the solid support. The immobilization can be non- covalent as in the case of biotin and streptavidin or similar strong non-covalent biological interactions, for example with a dissociation constant in the region of 1013M , 1014M or 1015M.
[0032] In an embodiment, the solid support comprises a plurality of beads, optionally magnetic beads, wherein each bead comprises a distinct probe or group of probes of the probe set.
[0033] In various embodiments of the aspects described herein, the H3K4me3 nucleosome methylation level is determined using a panel probe set or targeted capture solid support described herein.
[0034] The preceding section is provided by way of example only and is not intended to be limiting on the scope of the present disclosure and appended claims. Additional objects and advantages associated with the compositions and methods of the present disclosure will be appreciated by one of ordinary skill in the art in light of the instant claims, description, and examples. For example, the various aspects and embodiments of the disclosure may be utilized in numerous combinations, all of which are expressly contemplated by the present description. These additional advantages objects and embodiments are expressly included within the scope of the present disclosure. The publications and other materials used herein to illuminate the background of the disclosure, and in particular cases, to provide additional details respecting the practice, are incorporated by reference, and for convenience are listed in the appended reference section.BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Further objects, features and advantages of the disclosure will become apparent from the following detailed description taken in conjunction with the accompanying figures showing illustrative embodiments of the disclosure, in which:
[0036] Fig. 1A and 1 B. Identification of regions bearing differential H3K4me3 signatures between fertile and infertile men stratified by BMI.
[0037] Fig. 2A and B. Reproducibility of ChlP-seq targeting H3K4me3 on human sperm. A. Pearson correlation heatmap of average RPKM-normalized scores (computed per 25 bases windows covering the genome) between sperm samples obtained from fertile men. Samples from two different sequencing experiments are indicated (exp2 and 3). For exp2, ChlPs were done on a different day for each sample. For exp3, ChlPs were done on 2 samples at a time. This heatmap shows a high correlation for all samples indicating the remarkable degree of reproducibility of ChlP-seq on human sperm. This is confirmed by a genome browser screenshot (B) of a large portion of chromosome 9 (63MBases) that shows for 5 human sperm samples highly identical RPKM normalized tracks.
[0038] Fig. 3A Reproducibility of ChlP-seq experiments in Study 2. Spearman correlation heatmap between H3K4me3 counts in the 50,117 reference peaks. The heatmap shows a high degree of reproducibility of ChlP-seq for H3K4me3 in this data with the exception of a few samples.
[0039] Fig. 3B. H3K4me3 median count distribution under reference peaks in sperm samples profiled in the study (n =55).
[0040] Figure 4. Panel testing: Sperm samples (n=44) from sperm donors and clinical samples from men with various categories of infertility were subjected to chromatin immunoprecipitation (ChIP) followed by capture of the target regions then sequenced. Contiguous regions (overlapping regions) were removed, and duplicates were marked. Bioinformatic analysis identified 468 regions that differed in enrichment between samples. The samples clustered into 4 categories that were termed fertile 1 , fertile 2, infertile 1 , and infertile 2.DESCRIPTION OF VARIOUS EMBODIMENTS
[0041] H3K4me3 in sperm can be altered by diet, and these alterations are transmitted to the embryo and associated with abnormal embryo gene expression and development. For example, H3K4me3 in embryos on paternal chromatin highly resembles that of sperm and that altered H3K4me3 in sperm via for example a folate deficient paternal diet, results in a pattern of alterations that persist on embryonic chromatin (see Lismer et al. 2021).
[0042] Infertility refers to an inability to conceive after 6-12 months (depending on age of patients), or more of regular unprotected sexual intercourse. In particular, infertility as used herein refers to male infertility in which the subject experiences infertility without any obvious cause (i.e. subject has normal standard semen analysis). Fertility refers to an ability to achieve pregnancy through either regular unprotected sexual intercourse or through assisted reproductive therapies (ART) including intrauterine insemination (IUI) or in vitro fertilization (IVF).
[0043] It is demonstrated herein that overweight or obese infertile men show differential H3K4me3 enrichment at a plurality of regions. It is also demonstrated herein that infertile men show differential enrichment at 8979 (FDR<0.05) regions compared to fertile men, regardless of BMI. The regions include genes implicatedin sperm formation and / or function, embryo development, neurodevelopment, epigenetics, and regions sensitive to paternal age, environmental exposures, and lifestyle. The inventors have identified herein H3K4me3 molecular signatures that are able to discriminate for example fertile from idiopathic infertile sperm (men with normal standard semen analysis with unexplained infertility) and / or men with one or more abnormal parameters for example using the WHO semen analysis (WHO laboratory manual for the examination and processing of human semen; Sixth edition, 27, July 2021).
[0044] The following is a detailed description provided to aid those skilled in the art in practicing the present disclosure. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the disclosure. All publications, patent applications, patents, figures and other references mentioned herein are expressly incorporated by reference in their entirety.L Definitions
[0045] As used herein, the following terms may have meanings ascribed to them below, unless specified otherwise. However, it should be understood that other meanings that are known or understood by those having ordinary skill in the art are also possible, and within the scope of the present disclosure. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In the case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.
[0046] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the description. Ranges from any lower limit to any upper limit are contemplated. The upper and lower limits of these smaller ranges which may independently be included in the smaller ranges is also encompassed within the description, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either both of those included limits are also included in the description.
[0047] It must be noted that as used herein and in the appended claims, the singular forms "a", "an", and "the" include plural references unless the context clearly dictates otherwise.
[0048] All numerical values within the detailed description and the claims herein are modified by “about” or “approximately” the indicated value, and take into account experimental error and variations that would be expected by a person having ordinary skill in the art.
[0049] The terms "about", “substantially” and “approximately” as used herein mean a reasonable amount of deviation of the modified term such that the end result is not significantly changed. These terms of degree should be construed as including a deviation of at least ±5% of the modified term if this deviation wouldnot negate the meaning of the word it modifies or unless the context suggests otherwise to a person skilled in the art.
[0050] The phrase "and / or," as used herein in the specification and in the claims, should be understood to mean "either or both" of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with "and / or" should be construed in the same fashion, i.e., "one or more" of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the "and / or" clause, whether related or unrelated to those elements specifically identified.
[0051] As used herein in the specification and in the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as "only one of or "exactly one of or, when used in the claims, "consisting of will refer to the inclusion of exactly one element of a number or list of elements. In general, the term "or" as used herein shall only be interpreted as indicating exclusive alternatives (i.e., "one or the other but not both") when preceded by terms of exclusivity, such as "either," "one of," "only one of," or "exactly one of."
[0052] In the claims, as well as in the specification above, all transitional phrases such as "comprising,""including," "carrying," "having," "containing," "involving," "holding," "composed of," and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases "consisting of and "consisting essentially of shall be closed or semi-closed transitional phrases, respectively
[0053] As used herein, the phrase "at least one," in reference to a list of one or more elements, should be understood to mean at least one element selected from anyone or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase "at least one" refers, whether related or unrelated to those elements specifically identified. The phrase “at least” preceding a series of numbers modifies each of the series of numbers by “at least”. For example, the phrase “at least 1 , 2, 3, 4, 5 ...” means “at least 1 , at least 2, at least 3, at least 4, at least 5 ...”.
[0054] The phrase “categorized” in the context of being categorized in a table as a specified category, means the genomic regions are listed in the table as being in the specified category.
[0055] The phrase “H3K4me3 nucleosome fragment”, refers to a nucleosome containing histone H3 tri-methylated on lysine 4 (H3K4me3) and / or the DNA that is associated with that histone modification. It includes and / or is used to mean a mononucleosomal fragment or enriched for mononucleosomal fragments (e.g. greater than 90%).
[0056] The phrases “level of H3K4me3 methylation”, “H3K4me3 methylation level”, or “H3K4me3 level” are intended to refer to the relative enrichment of nucleosomes bearing H3K4me3 that are associated with a genomic region (e.g. normalized using Reads Per Kilobase per Million mapped reads (RPKM)). The number of sequencing reads or read counts at a given region obtained by H3K4me3 ChlP-seq is indicative of the level of H3K4me3 present at that region in the sample. A region is considered to be enriched for H3K4me3 methylation if the level of H3K4me3 methylation is above an assigned read threshold, for example above the average level of H3K4me3 methylation across the genome, above a certain RPKM, above a certain percentage of normalized counts, or above a specified mean normalized count. For example, a cutoff threshold for low reads of Iog2 (median raw counts) >4.5 (as in Example 1) or a Iog2 (median(raw counts)+8) > 4.5 (as in Example 3) may be applied. Alternatively or additionally, the top 75% of normalized count for the significant regions, or a mean normalized count of >4.38 (as in Example 3) may be used as a cutoff for high / medium significant enrichment. Other cutoffs may be used, depending, for example, on the quality of the sequencing data.
[0057] The phrase “differential enrichment” is used to refer to a significant difference in the indicated feature detected in a sample compared to a reference or another sample. For example, differential enrichment of H3K4me3 at a specific region indicates that the level of H3K4me3 methylation detected in a sample is significantly different from a reference level. The differential enrichment may be positive or negative, i.e. the level of H3K4me3 methylation detected in a sample may be significantly higher or significantly lower than that of the reference.
[0058] The phrases “reference population” or “reference population data set” is used to refer to a population of patients (or data derived therefrom) which serve as a control against which to compare a test population (e.g. infertile men) or a subject (e.g. a man suspected of being infertile). The reference population can be a random population (e.g. fertility status unknown), a population known to be healthy (e.g. fertile sperm donors), or a population known to be infertile, optionally for known or suspected reasons (e.g. known toxicant exposure). In preferred embodiments, the reference population consists of healthy sperm donors (and data e.g. H3K4me3 methylation levels, derived therefrom).
[0059] The term “genomic region” or “region” is used to refer to a discrete region of the genome. A genomic region may be indicated for example by identifying approximate chromosomal positions that define the boundaries of the genomic region (e.g. start and stop), or may be indicated as being within a specified distance of a chromosomal position such as a transcription start site. It should be understood that a “target region” or “target”, as used herein, refers to a specific genomic region of interest, for example one or more the regions listed in Table 1 or Table 2, as well as smaller portions, or sub-regions, thereof, for example a portion of a genomic region selected for targeted enrichment. Target region and genomic region when referring to a region of interest, can be used interchangeably. The genomic region can also include sequence immediately upstream or downstream of an identified target region. For example, the genomic regions identified in Table 1 and Table 2 may include a further 1 ,000 bp upstream and / or downstream of the identified start and stop locations.
[0060] The term “sperm epigenomic health” is a measure (e.g. score or probability) of sperm vigor based on sperm epigenomics, and includes indicators of sperm reproductive health, and can be used for example to assess likely clinical outcomes (e.g. paternal- related causes of miscarriage or embryo development) and lifestyle- and environmental exposure-related epigenomic changes, in addition to being a predictor of fertility rate or infertility. Markers of sperm epigenomic health can for example indicate whether overall health (e.g. BMI, toxicant exposure) is contributing to infertility, and may suggest whether one or more lifestyle changes could improve fertility.
[0061] The term “overweight” as used herein refers to an individual having a body mass index (BMI) of for example of > 24.9, optionally less than 30, such as a BMI of 25-29. The term “obese” as used herein refers to an individual having a BMI of 30 or greater than 30. Unless the context dictates otherwise, the term “overweight” encompasses both overweight and obese individuals.
[0062] The term “obtaining” used herein in the context of sequence reads includes measuring or determining using an assay, for example in the context of obtaining a first plurality of sequence reads of histone 3 lysine 4 trimethylated (H3K4me3) nucleosome fragments, the term obtaining includes for example performing next generation sequencing to determine the sequence and / or number of reads of the histone 3 lysine 4 trimethylated (H3K4me3) nucleosome fragments. It can also include receiving digital data (e.g. corresponding to sequence reads) of sperm histone 3 lysine 4 trimethyl (H3K4me3) nucleosome fragments.
[0063] It should also be understood that, in certain methods described herein that include more than one step or act, the order of the steps or acts of the method is not necessarily limited to the order in which the steps or acts of the method are recited unless the context indicates otherwise.
[0064] Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure, the preferred methods and materials are now described.IL Methods, Reagents, and Kits
[0065] Described herein are methods for determining sperm H3K4me3 nucleosome methylation levels of specific genomic regions in a subject, assessing the sperm epigenomic health of a subject, and / or predicting fertility or infertility of a subject. Also described herein are methods for predicting the clinical outcome of assisted reproductive therapy, and methods for identifying treatment strategies that are likely to improve clinical outcomes in a subject. The methods described herein can be used, for example, to establish robust preconception health advising tools for men to improve fertility rates and / or embryo development. Also described herein are methods for preparing a sperm sample, and methods for preparing H3K4me3 mononucleosome fragments from a semen sample or sperm sample. Also described herein are reagents and kits which can optionally be used in the methods described herein.
[0066] As shown herein, H3K4me3 is localized throughout the sperm genome and is differentially enriched 8979 genomic regions in infertile men compared to fertile sperm donors. Accordingly, H3K4me3 isdemonstrated herein to be a suitable candidate for identifying differences in the sperm epigenome, based on enrichment differences associated with fertility status (fertile versus infertile, BMI and response to the environment (e.g. toxicant exposure or diet)) and transmission to the embryo.
[0067] The methods herein are useful for determining the sperm H3K4me3 nucleosome methylation levels in a subject, and using the H3K4me3 level at each of a plurality of genomic regions to assess the sperm epigenomic health and / or fertility status of the subject.
[0068] Accordingly, one aspect of the disclosure is a method for determining sperm H3K4me3 nucleosome methylation levels in a subject (e.g. profiling sperm chromatin), the method comprising: obtaining a plurality of sequence reads from a semen sample or a sperm sample obtained from the subject, wherein the sequence reads are from H3K4me3 nucleosome fragments; and determining a H3K4me3 level using the first plurality of sequence reads at each of a plurality of genomic regions, the plurality of genomic regions comprising at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 482, or any number in between and including 5 and 482 genomic regions identified in Table 1 , and optionally at least one genomic region selected from the genomic regions identified in Table 2.
[0069] In some embodiments, the subject is a sperm donor. The method can be performed for example for selecting a sperm donor, for example for intrauterine insemination (IUI), in vitro fertilization (IVF), or for further analysis. In some embodiments, the method further comprises selecting a sperm donor. In yet other embodiments, the method further comprises performing IUI or in vitro fertilization. In other embodiments, the method further comprises further sperm testing.
[0070] Another aspect of the disclosure is a method for assessing sperm epigenomic health of a subject, the method comprising: determining the sperm H3K4me3 nucleosome methylation levels in a subject, wherein the sperm H3K4me3 nucleosome methylation levels are determined at a plurality of genomic regions comprising at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 482, or any number in between and including 5 and 482 genomic regions identified in Table 1 , and optionally at least one genomic region selected from the genomic regions identified in Table 2; and using the sperm H3K4me3 nucleosome methylation levels to assess the sperm epigenomic health of the subject.
[0071] Another aspect of the disclosure is a method for determining if a subject is infertile, the method comprising: determining the sperm H3K4me3 nucleosome methylation levels in a subject, wherein the sperm H3K4me3 nucleosome methylation levels are determined at a plurality of genomic regions comprising at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 482, or any number in between and including 5 and 482 genomic regions identified in Table 1 , and optionally at least one genomic region selected from the genomic regions identified in Table 2; and using the sperm H3K4me3 nucleosome methylation levels to determine if the subject is infertile.
[0072] Another aspect of the disclosure is a method for treating reduced fertility in a subject in need thereof, wherein the subject has been determined as having reduced sperm epigenomic health and / or fertilitybased on sperm H3K4me3 nucleosome methylation levels determined using a method described herein, and the method comprises administering a treatment for improving fertility based on the determined sperm H3K4me3 nucleosome methylation levels. The treatment plan can be considered a personalized treatment plan as it is based on the subject’s determined H3K4me3 nucleosome methylation levels.
[0073] Another aspect of the disclosure is a method of providing a fertility treatment plan for a subject in need thereof, wherein the subject has been determined as having reduced sperm epigenomic health and / or fertility based on sperm H3K4me3 nucleosome methylation levels determined using a method described herein, and the method comprises providing a fertility treatment plan based on the determined sperm H3K4me3 nucleosome methylation levels.
[0074] In some embodiments, the method further comprises determining a score.
[0075] In various embodiments of the aspects described herein, the step of using the determinedH3K4me3 nucleosome methylation levels (e.g. for profiling sperm chromatin, for assessing sperm epigenomic health, or determining fertility status) comprises i) comparing the H3K4me3 level at each genomic region of the plurality of genomic regions to a control H3K4me3 level or reference value for each respective genomic region in the plurality of genomic regions, wherein the H3K4me3 level at one or more regions relative to control H3K4me3 level or reference value is indicative of sperm epigenomic health and / or fertility status; and / or ii) determining the sperm epigenomic health and / or fertility status of the subject based on statistical modeling. In various embodiments, the subject is assigned a score (e.g. sperm epigenomic health or fertility status score) based on the determined H3K4me3 nucleosome methylation levels.
[0076] In an embodiment, the determined H3K4me3 nucleosome methylation levels are used for assessing BMI-associated infertility or sperm epigenomic health. In an embodiment, the genomic regions are used for assessing cannabis consumption-associated infertility or epigenomic health. In an embodiment, the genomic regions are used for assessing toxicant exposure-associated infertility or epigenomic health. In an embodiment, the genomic regions are used for assessing paternal epigenomic age.
[0077] In an embodiment the sperm epigenomic health and / or fertility score is determined after one or more lifestyle modifications, and optionally compared to previous score. In an embodiment, the one or more lifestyle modifications comprises weight-loss and / or diet modification, nicotine / cannabis cessation, and / or limiting or reducing toxicant exposure. In an embodiment, the sperm epigenomic health and / or fertility score is determined at least three months or at least 6 months after the one or more lifestyle modifications. In an embodiment, the previous score is determined at least three months or at least six months prior.
[0078] In various embodiments of the aspects described herein, the treatment for improving fertility comprises one or more of intrauterine insemination (IUI), in vitro fertilization (IVF), a dietary plan, a weight loss plan, an exercise plan, nicotine / cannabis cessation, and / or limiting or reducing exposure to one or more specific toxicants.
[0079] In various embodiments of the aspects described herein, the treatment plan comprises one or more of intrauterine insemination (IUI), in vitro fertilization (IVF), a dietary plan, a weight loss plan, an exercise plan, nicotine / cannabis cessation, and / or limiting or reducing exposure to one or more specific toxicants.
[0080] In various embodiments of the aspects described herein, dietary modifications or plans may include, for example and without limitation, increased intake of specific nutrients (e.g. folate) and decreased caloric intake. In various embodiments of the aspects described herein, weight-loss plans may include, for example and without limitation, interventions such as a weight loss drug (e.g. Bupropion-naltrexone (Contrave™), Liraglutide (Saxenda™), Orlistat (Xenical™, Alli™), Phentermine-topiramate (Qsymia™), Semaglutide (Wegovy™, Ozempic™), or Setmelanotide (Imcivree™)), dietary interventions (e.g. decreased caloric intake), and / or surgery.
[0081] In various embodiments of the aspects described herein, predicting if a subject is infertile or profiling sperm chromatin, or assessing sperm epigenomic health of the subject may further comprise or be carried out in conjunction with semen analysis techniques which measure semen volume, pH, sperm count, sperm motility, and / or sperm morphology.
[0082] The H3K4me3 nucleosome fragments may be obtained from a sperm sample (e.g. purified from a semen sample) using a suitable procedure for example the procedure detailed in Lambrot et al. 2021 , or Example 1. Typically, the nucleosome fragments are substantially mononucleosomal fragments. In embodiments, the nucleosome fragments isolated are at least 90% or at least 95% mononucleosomal fragments. If less than 90%, the isolation can be repeated.
[0083] In some embodiments, the procedure involves obtaining a semen sample comprising sperm from the subject, purifying the sperm to at least 95%, 96%, 97%, 98% or 99% purity (e.g. a sperm sample), preparing mononucleosomes from the sperm sample, isolating H3K4me3 mononucleosomes using a suitable procedure such as ChIP, and purifying H3K4me3 mononucleosome DNA fragments. Accordingly, in an embodiment, the obtaining histone 3 lysine 4 trimethylated (H3K4me3) nucleosome fragments from a sperm sample or a semen sample from the subject comprises: obtaining a semen sample comprising sperm from the subject that has been obtained from the subject, purifying the sperm to at least 95%, 96%, 97%, 98% or 99% purity (e.g. a sperm sample), preparing mononucleosomes from the sperm sample, isolating H3K4me3 mononucleosomes using a procedure such as ChIP, and purifying H3K4me3 mononucleosome DNA fragments.
[0084] In an embodiment, the purifying is at least 99% sperm e.g. less than 1% somatic cells.
[0085] As described herein, a reduction of background and improvement in specificity and detection of differential enrichment is achieved when the semen sample is purified to obtain a sperm sample e.g. purified sperm. As known in the art, a semen sample comprises seminal fluid and sperm. As used herein, the term “sperm sample” means a purified sample comprising sperm, that is purified from a semen sample (e.g. a sample purified from semen, for example to at least 95%, 96%, or 97%, 98%, or 99% sperm cells), unless the contextclearly dictates otherwise. The semen sample from the subject may be fresh or frozen, and may be neat semen (e.g. semen that has not previously been washed or otherwise processed prior to purification). For example, the semen sample can be neat, fresh or frozen, with cryoprotectant, washed fresh or frozen, with or with or without abstinence. The semen sample or the sperm sample can be mixed with or without cryoprotectant. For example, the sample can be neat frozen, with or without cryoprotectant. The cryoprotectant can be a freezing medium. Examples of suitable cryoprotectants include Sperm Freezing medium (Origio), Arctic Sperm Cryopreservation medium, CryoSperm™, (Origio), or SpermFreeze Solution TM, (Vitrolife).
[0086] Prior to isolating the H3K4me3 nucleosome fragments, the sperm sample can be purified to for example at least 95%, 96%, 97%, 98% or 99% purity. As used herein, sperm that is at least 99% pure means that at least 99% of the cells observed in a haemocytometer are spermatozoa, and less than 1 % of observed cells are somatic cells. Any suitable method may be used to purify sperm. For example, sperm may be purified using a density gradient method as described in Lambrot et al., or by cell sorting approaches based on parameters such as size, or by sperm-cell specific marker selection approaches, or by filtration. Alternately, as described herein in Example 1 , a sperm sample that is at least 99% sperm can be obtained by a method comprising the steps of a) obtaining a semen sample comprising sperm from the subject, b) pelleting the sperm, c) resuspending the sperm in buffer, d) flash-freezing the resuspended sperm, e) thawing the sperm, f) determining the purity of the sperm, and g) repeating steps b) to f) until the sample is at least 99% sperm.
[0087] The H3K4me3 nucleosome fragments may be isolated using any suitable technique, for example using native chromatin immunoprecipitation (ChIP), or targeted cleavage-based techniques such as TAM-ChIP™, CUT&RUN or CUT&Tag. Any suitable H3K4me3 antibody may be used, including polyclonals and monoclonals. Suitable antibodies include, for example, C42D8, available from Cell Signaling Technologies (cat#: 9751), and 9HCLC, RM340, and MABI 0304, available from Thermo Fisher Scientific.
[0088] For ChlP-based protocols, any suitable antibody capture agent may be used for isolation of antibody bound fragments. As shown herein, magnetic beads such as Dynabeads, available from Invitrogen, are suitable for this purpose. Other suitable magnetic or otherwise coated beads for purifying antibody linked cells or molecules are commercially available. Exemplary ChIP methods are provided in Example 1 .
[0089] As described herein, a reduction of background and improvement in specificity and detection of differential enrichment is achieved when mono-nucleosomal fragments are used in the methods and with the reagents described herein. Mono-nucleosome fragments can be obtained by treating the sperm sample with a suitable nuclease such as micrococcal nuclease (MNase) or by using a commercially available nucleosome preparation kit or by using nucleoplasmin. Accordingly, in an embodiment, the sperm sample is treated with Mnase to obtain mono-nucleosome fragments priorto nucleosome isolation. Incomplete digestion of chromatin by Mnase results in the presence of higher order nucleosomes such as dinucleosome and trinucleosomes. H3K4me3 mono-nucleosome fragments can also be obtained by targeted cleavage-based genomic fragmentation for example using Mnase (for example CUT&RUN protocols) or transposase-based cleavageand tagging (for example TAM-ChIP™ or CUT&Tag protocols). Accordingly, in an embodiment, H3K4me3 mono-nucleosome fragments are prepared using TAM-ChIP™, CUT&RUN or CUT&Tag.
[0090] The DNA fragments corresponding to incompletely digested chromatin can be removed using a size selection step following DNA purification. Accordingly, in an embodiment, DNA purified from nucleosomes is subjected to a size-selection step, for example to remove DNA greater than 500 bp, 400 bps, 300 bps, 250 bps, 210bps or 200 bps. Suitable size-selection methods and kits are readily available, including for example Agencourt AMPure XP beads. Other similar available beads include Mag-Bind or other next generation size selection beads.
[0091] The quantity and fragment size of purified and size-selected DNA can be monitored using any suitable technique. For example, a bioanalyzer instrument may be used in combination with a commercially available kit.
[0092] In some embodiments, the method further comprises generating a sequencing library from the nucleosomal DNA. Sequencing libraries can be generated from the purified mononucleosomal DNA using any suitable method and may depend on the sequencing platform being used. Suitable sequencing library preparation kits are commercially available, for example the QIAseq™ Ultralow Input Library Kit or the Kapa BioSystems HTP Library Preparation kit. Other library preparation methods and kits are available. These include Collibri kits from Invitrogen and kits for use with illumina sequencers such as TruSeq Chip library preparation kit. Sequencing libraries may also be prepared by a service that prepares sequencing libraries.
[0093] Obtaining the sequence reads is performed for example by sequencing the nucleosomal DNA.Sequencing may be performed on any suitable platform, for example Illumina HiSeq 2500, the NovaSEQ 5000, HiSeq3000, HiSeq4000, MiSeq, NextSeq, MiniSeq, or iSeq. Sequencing may be single-end or paired-end, and may be any suitable length, for example, 100bp reads.
[0094] The ChlP-seq methods shown herein have been used to generate a high depth sequencing data set of which was used to identify H3K4me3 peaks in a reference population, for example as shown in Example 1 .
[0095] Sequencing data may be pre-processed using any suitable methods. Short sequencing reads may be trimmed or not using suitable software tools such as eg Trimgalore or Trimmomatic. Trimmed reads may be aligned against the reference genome using any suitable software eg such as short-read DNA aligner bowtie2 or bwa using any suitable settings. For example, for paired-end reads aligned using BWA, the settings may be MEM -MP -t 10 -v 2 -c 100, or for single end reads aligned using Bowtie the settings may be -t -v 3 -m 100. Unaligned reads may be filtered, sorted, and / or converted to BAM files using any suitable software such as e.g. Samtools. Aligned reads may be filtered according to mismatch rate and / or alignment quality score. Exemplary pre-processing methods are described in the Examples.
[0096] Sequencing data may be normalized and analyzed using suitable methods. Aligned reads may be used to obtain counts at defined genomic regions (e.g. TSS + / - 1 kB, or specific regions of interest such asthose identified herein) or to de novo identify genomic intervals with significant binding enrichment. De novo identification of these regions interest may be performed using peak-based methods such as MACS2 or window-based methods using a sliding window approach whereby the number of reads inside each window is counted. Alternatively, H3K4me3 peaks identified in sperm in a reference population (see e.g. Lambrot et al. 2021) can be used to identify peaks in H3K4me3 data obtained from sperm for example from individual men. Various parameters may be changed to optimize e.g. peak number or peak width. Genomic intervals (identified de novo or not) with low counts e.g. below a certain threshold, may be filtered out as they may be background regions of non-specific binding. Once counts under genomic intervals of interest are obtained, differentially enriched regions between group of interest may be identified using generalized linear models or derivative approaches which will assign a statistics indicating confidence that they are differentially bound. The statistical approach may include the modeling, normalization and variance stabilization of count data. The specific methods used will depend on the distribution of the sample data and what provides the best fit. For example, normalization may be used to adjust for different library sizes and sequencing depths. Regions enriched for H3K4me3 can be identified using any suitable software such as for example R / Bioconductor package csaw (v1.18.0). Reads with a mapping quality score above a certain threshold, for example 20, may be counted in sliding windows of fer example 150 bp. Windows may then be filtered and the remaining contiguous windows merged. Windows with a Iog2 fold change over a specified threshold (e.g. 4, 5, 6, 7, 8) may be retained. Filtered f windows may be normalized for example by library size using counts per million (CPM) normalization, composition bias using trimmed means of M-values (TMM) normalization, and / or batch effects using the sva ComBat function (version 3.28.0). H3K4me3 positive windows can be compared across samples using for example EdgeR. Multiple testing adjustment can be achieved for example using Benjamini-Hochberg. Cut-offs may be used to identify regions associated with e.g. fertility status or BMI or other environmental exposure such as toxicant exposures. Regions may be identified as differentially enriched for H3K4me3 using any suitable cutoff such as for example using an FDR < 0.2. Similar comparisons can be achieved for example using dseq2 / diffbind. Enrichments may be identified for example using z-scores determined for example by the Bioconductor package regioneR. A threshold cut off may be used to distinguish between enrichment and background (anything below the cut off is not enriched and considered background). CpG density can be used for example as an internal control as H3K4me3 is most enriched here to have high coverage of TSS. Correlation of counts under the identified peaks may be computed pairwise for all samples, and the median correlation may be used to select one or more sets of samples for differential binding analysis. An empirical threshold which maximizes true positives and minimize false positives may be set to determine significance. If desired, Bigwig coverage tracks may be generated using for example DeepTools2, and visualized for example in IGV. Genomewide coverage may be calculated and normalized using Reads Per Kilobase per Million mapped reads (RPKM). Exemplary analysis methods are described in Lismer et al. 2020, Lismer et al. 2021 and the Examples.
[0097] In various embodiments of the aspects described herein, the subset of genomic regions comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 482, or any number in between and including 5 and 482 genomic regions that are identified in Table 1. In anembodiment, the subset of genomic regions comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 153, or any number in between and including 5 and 153 genomic regions that are categorized in table 1 as Category 1. In an embodiment, the subset of genomic regions comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 59, or any number in between and including 5 and 59 genomic regions that are categorized in table 1 as Category 2. In an embodiment, the subset of genomic regions comprises at least 5, 10, 15, 20, 25, 30, 35, 40, or any number in between and including 5 and 40 genomic regions that are categorized in table 1 as Category 3. In an embodiment, the subset of genomic regions comprises at least 5, 10, 13, or any number in between and including 5 and 13 genomic regions that are categorized in table 1 as Category 4. In an embodiment, the subset of genomic regions comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 41 , or any number in between and including 5 and 41 genomic regions that are categorized in table 1 as Category 5. In an embodiment, the subset of genomic regions comprises at least 5 or 6 genomic regions that are categorized in table 1 as Category 6. In an embodiment, the subset of genomic regions comprises at least 5, 9, or any number in between and including 5 and 9 genomic regions that are categorized in table 1 as Category 7. In an embodiment, the subset of genomic regions comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 51 , or any number in between and including 5 and 51 genomic regions that are categorized in table 1 as Category 8. In an embodiment, the subset of genomic regions comprises at least 5, 10, 15, 20, 25, 29, or any number in between and including 5 and 29 genomic regions that are categorized in table 1 as Category 9. In an embodiment, the subset of genomic regions comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 68, or any number in between and including 5 and 68 genomic regions that are categorized in table 1 as Category 10. In an embodiment, the subset of genomic regions comprises at least 5, 10, 13, or any number in between and including 5 and 13 genomic regions that are categorized in table 1 as Category 11. In an embodiment, the subset of genomic regions comprises any combination of the foregoing Categories, e.g., any number of genomic regions in between and including 5 and the total number in Category 1 and any number of genomic regions in between and including 5 and the total number in Category 3 or any number of genomic regions in between and including 5 and the total number in Category 2, any number of genomic regions in between and including 5 and the total number in Category 5, and any number of genomic regions in between and including 5 and the total number in Category 10.
[0098] In an embodiment, the subset of genomic regions comprises: at least 5 genomic regions that are categorized in table 1 as Category 1 ; at least 5 genomic regions that are categorized in table 1 as Category 2; at least 5 genomic regions that are categorized in table 1 as Category 3; at least 5 genomic regions that are categorized in table 1 as Category 4; at least 5 genomic regions that are categorized in table 1 as Category 5; at least 5 genomic regions that are categorized in table 1 as Category 6; at least 5 genomic regions that are categorized in table 1 as Category 7; at least 5 genomic regions that are categorized in table 1 as Category 8; at least 5 genomic regions that are categorized in table 1 as Category 9; at least 5 genomic regions that are categorized in table 1 as Category 10; and / or at least 5 genomic regions that are categorized in table 1 as Category 11 .
[0099] The H3K4me3 methylation levels of one or more control genomic regions may also be determined for normalization and / or comparison to a control H3K4me3 level or reference value. In an embodiment, the set of control genomic regions comprises at least 1 , 2, 3, 4, 5, 6, 10, 15, 18, or any number in between and including 1 and 18 of the genomic regions identified in Table 2. In an embodiment, the set of control genomic regions comprises 1 , 2, 3, 4, 5, or 6 genomic regions that are categorized in table 2 as Category A. In an embodiment, the set of control genomic regions comprises 1 , 2, 3, 4, 5, or 6 genomic regions that are categorized in table 2 as Category B. In an embodiment, the set of control genomic regions comprises 1 , 2, 3, 4, 5, or 6 genomic regions that are categorized in table 2 as Category C. In an embodiment, the set of control genomic regions comprises 3, 4, 5, or 6 genomic regions that are categorized in table 2 as Category C, and optionally at least 1 or at least 2 genomic regions categorized in table 2 as Category B. In an embodiment, the subset of genomic regions comprises any combination of any of the foregoing Categories, e.g., any number of genomic regions in Category A and any number of genomic regions in Category C or any number of genomic regions in Category A, any number of genomic regions Category B, and any number of genomic regions in Category C.
[0100] The H3K4me3 levels at the selected targets will have a statistical normal based on a reference population, and any deviation from the norm either up or down in enrichment will be associated with a profile of for example either fertile or infertile. For example, for the regions in Table 1 , deviation from the statistical normal of a healthy reference population can be associated with infertility. Similarly, deviation from the statistical normal of an infertile reference population, can be associated with fertility.
[0101] A naive Bayes’ classifier may be trained to predict fertility status using internal leave-one-out cross-validation. The distribution of accuracy significances that can be obtained from 100,000 predictors built using a varied number of regions sampled may be determined. The best set of regions and / or probes may be selected for further validation in an independent cohort. Combining coefficients with the relative binding levels of the selected region, a risk score for each individual will be calculated: Risk score = ^=1regioni x coefficient^ Individuals may be classified into low- and high-risk groups based on for example the optimal risk score cutoff value, which represents the point at which the Youden index (sensitivity + specificity - 1) reached a maximum value. Both the continuous and high / low risk score may be evaluated with relevant clinical variables using a multivariate logistic regression model. The decision on which predictors to include in the model may be based on clinical and statistical reasoning to address challenges such as confounders, interactions and multicollinearity. Other factors may also be considered. Forward-backward variable selection may be performed based on Bayesian Information Criterion (BIC). Based on the results of the of the multivariate analysis, a nomogram may be formulated to predict high to low risk of infertility. The predictive accuracy and discrimination ability of the nomogram may be determined using a calibration plot and C-index, respectively.
[0102] Alternatively or additionally, the methods described herein comprise comparing the H3K4me3 levels at one or more genomic regions to one or more controls or reference values. Controls may include, for example and without limitation, H3K4me3 levels in healthy patients (or a reference population H3K4me3 levelsobtained or determined from samples obtained from of a group of healthy patients) which can be used to create "control values". Control values may be obtained or derived from a pool of healthy patients (e.g. a negative control value) or from a pool of patients with known cause(s) of infertility and / or clinical outcome (e.g. a positive control value). Controls may also include an internal control, for example the H3K4me3 levels at a target region relative to the total number of reads, or the H3K4me3 levels relative to the H3K4me3 levels from a second region from the patient sample (e.g. the H3K4me3 level relative to the H3K4me3 level of a second region, or a relative to the H3K4me3 level of a control region). Similarly, a reference value may include for example a threshold value, above or below which (depending on the specific region and reference value) may indicate the patient has an increased or decreased risk of infertility (or other clinical outcome), or may indicate the patient does not have an increased or decreased risk thereof. In an embodiment, the reference population is a fertile reference population consisting of sperm donors and the H3K4me3 levels at one or more regions that do not differ in enrichment between fertile and infertile samples are used as internal controls for normalization.
[0103] The expressions “targeted enrichment” and “target enrichment” are used herein to refer to methods which increase the abundance of specific (target) nucleic acid sequences in a population of nucleic acids, for example prior to detection and / or sequencing. Sequencing libraries can be enriched for DNA fragments corresponding to the specific genomic regions of interest by targeted enrichment of selected genomic regions (or fragments thereof) using any suitable next generation sequencing (NGS) target enrichment approach, including but not limited to, hybridization capture-based enrichment, PCR-based enrichment, primer extension-based enrichment, amplicon sequencing, and targeted nanopore sequencing with Cas9 guided adapter ligation. Enriching for DNA fragments corresponding to the specific regions of interest may reduce the number of sequencing reads needed, for example, for detecting differential enrichment of H3K4me3 at target regions in a patient sample, predicting if a subject is infertile, orassessing sperm epigenomic health of a subject. For example the methods described herein may comprise obtaining fewer than or about 15 million sequence reads, fewer than or about 10 million sequence reads, fewer than or about 7.5 million sequence reads, fewer than or about 5 million sequence reads, fewer than or about 4 million sequence reads, fewer than or about 3 million sequence reads, fewerthan or about 2 million sequence reads, orfewerthan or about 1 million sequence reads, depending on the specific regions and / or number of regions being enriched.
[0104] Accordingly, in some embodiments, the selected genomic regions (or fragments thereof) may be enriched using one or more probes (or primers), typically at least two probes, complementary to all or a portion of the genomic region of interest prior to sequencing. The term “probe” as used herein encompasses capture probes for hybridization capture-based enrichment and primers for PCR- and primer extension-based enrichment, and minimally comprises a nucleic acid molecule. Capture probes (e.g. probes of the probe sets) for hybridization capture-based enrichment can comprise a nucleic acid of any suitable length, for example 50- 150 nucleotides, 100-150 nucleotides, 120-150 nucleotides, 130-150 nucleotides, 140-150 nucleotides, or any number in between and including 50 and 150 nucleotides in length. Optionally, the probes are 140 nucleotides in length, optionally 147 nucleotides in length. They may be labelled biotin. Primers for PCR-based and / orprimer extension-based enrichment can comprise a nucleic acid of any suitable length for example 10-50 nucleotides, 15-40 nucleotides, 20-30 nucleotides, or any number in between and including 10 and 50 nucleotides in length. Each probe of the probe sets described herein is complementary to a sequence and hybridizes as described herein to a genomic region in Table 1 or Table 2.
[0105] The term “nucleic acid” as used herein refers to a molecule or sequence of nucleoside or nucleotide monomers consisting of naturally occurring bases, sugars and intersugar (backbone) linkages. The term also includes modified or substituted sequences comprising non-naturally occurring monomers or portions thereof. The nucleic acid sequences of the present application may be deoxyribonucleic acid sequences (DNA) or ribonucleic acid sequences (RNA) and may include naturally occurring bases including adenine, guanine, cytosine, thymidine and uracil. The sequences may also contain modified bases. Examples of such modified bases include aza and deaza adenine, guanine, cytosine, thymidine and uracil; and xan-thine and hypoxanthine. The nucleic acid can be either double stranded or single stranded, and represents the sense or antisense strand. Further, the term "nucleic acid" includes the complementary nucleic acid sequences as well as codon optimized or synonymous codon equivalents. The term "isolated nucleic acid sequences" as used herein refers to a nucleic acid substantially free of cellular material or culture medium when produced by recombinant DNA techniques, or chemical precursors, or other chemicals when chemically synthesized. An isolated nucleic acid is also substantially free of sequences which naturally flank the nucleic acid (i.e. sequences located at the 5' and 3' ends of the nucleic acid) from which the nucleic acid is derived.
[0106] The probes may be labeled for example to aid in the capture of the nucleosomal fragments (hybridized capture probe, primer, PCR products, primer extension products, etc.) of interest. The label can be any suitable label, for example biotin or any another binding approach used to capture labelled probes, primers, PCR products, primer extension products, etc. Probes can be synthesized using any suitable technique and / or may be available from a commercial vendor, for example Integrated DNA technologies, NGS discovery pools, or Twist Bioscience custom panels. The probes can be the panel probe sets described herein.
[0107] In embodiments using labelled probes (e.g. a panel probe set where at least a subset of the probes are labelled), the methods can further comprise use of capture beads such as magnetic beads comprising a moiety that binds the label. For example, where the label comprises biotin (e.g. the probe is biotinylated), the moiety can be streptavidin or avidin. This can be used to enrich the target probe regions that are hybridized to the biotinylated probes. Also described herein are panel probe sets for targeted enrichment of one or more genomic regions of interest, for example for use in a method described herein. A panel probe set can comprise any number of probes for each region, for example at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, or more probes for each region of interest. Probe sets may be provided in any suitable format, for example as a composition comprising a plurality of probes.
[0108] The panel probe set can comprise probes for each of 5 or more genomic regions of interest. For example the probe set may comprise at least or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100, 200, 300, 400, 500, 750, 1000, 2000, 3000, 4000, 5000, 10,000 or more probes targeting at least 5, 10, 15, 20, 25, 30,35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 482, or any number in between and including 5 and 482 genomic regions that are identified in Table 1. Optionally, the probe set further comprises probes for at least 1 , 2, 3, 4, 5, 6, 10, 15, 18, or any number in between and including 1 and 18 of the genomic regions identified in Table 2.
[0109] In an embodiment, the panel probe set comprises probes for at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 153, or any number in between and including 5 and 153 genomic regions that are categorized in table 1 as Category 1 . In an embodiment, the probe set comprises probes for at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 59, or any number in between and including 5 and 59 genomic regions that are categorized in table 1 as Category 2. In an embodiment, the probe set comprises probes for at least 5, 10, 15, 20, 25, 30, 35, 40, or any number in between and including 5 and 40 genomic regions that are categorized in table 1 as Category 3. In an embodiment, the probe set comprises probes for at least 5, 10, 13, or any number in between and including 5 and 13 genomic regions that are categorized in table 1 as Category 4. In an embodiment, the probe set comprises probes for at least 5, 10, 15, 20, 25, 30, 35, 40, 41 , or any number in between and including 5 and 41 genomic regions that are categorized in table 1 as Category 5. In an embodiment, the probe set comprises probes for at least 5 or 6 genomic regions that are categorized in table 1 as Category 6. In an embodiment, the probe set comprises probes for at least 5, 9, or any number in between and including 5 and 9 genomic regions that are categorized in table 1 as Category 7. In an embodiment, the probe set comprises probes for at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 51 , or any number in between and including 5 and 51 genomic regions that are categorized in table 1 as Category 8. In an embodiment, the probe set comprises probes for at least 5, 10, 15, 20, 25, 29, or any number in between and including 5 and 29 genomic regions that are categorized in table 1 as Category 9. In an embodiment, the probe set comprises probes for at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 68, or any number in between and including 5 and 68 genomic regions that are categorized in table 1 as Category 10. In an embodiment, the probe set comprises probes for at least 5, 10, 13, or any number in between and including 5 and 13 genomic regions that are categorized in table 1 as Category 11 .
[0110] In an embodiment, the panel probe set comprises probes for: at least 5 genomic regions that are categorized in table 1 as Category 1 ; at least 5 genomic regions that are categorized in table 1 as Category 2; at least 5 genomic regions that are categorized in table 1 as Category 3; at least 5 genomic regions that are categorized in table 1 as Category 4; at least 5 genomic regions that are categorized in table 1 as Category 5; at least 5 genomic regions that are categorized in table 1 as Category 6; at least 5 genomic regions that are categorized in table 1 as Category 7; at least 5 genomic regions that are categorized in table 1 as Category 8; at least 5 genomic regions that are categorized in table 1 as Category 9; at least 5 genomic regions that are categorized in table 1 as Category 10; and / or at least 5 genomic regions that are categorized in table 1 as Category 11 .
[0111] In an embodiment, each probe (e.g. the hybridization portion thereof) is 10-150 nucleotides in length, or any number in between and including 10-150 nucleotides in length. Where the probe is a primer, itmay for example be 50 nucleotides or less, for example 25 nucleotides or more or less. Where the probe is not a primer, it may be 50 nucleotides or more, preferably 100 or more such as 140.
[0112] In some embodiments, each probe is 50-100 nucleotides, 100-150 nucleotides, 120-150 nucleotides, 130-150 nucleotides, 140-150 nucleotides, or any number in between and including 50-150 nucleotides in length. In an embodiment, each probe is 140 nucleotides in length.
[0113] In other embodiments, the probe is 10-50 nucleotides, 15-40 nucleotides, 20-30 nucleotides, or any number in between and including 10 and 50 nucleotides in length. When referring to nucleic acid molecule length, bases, base pairs and nucleotides can be used interchangeably, e.g., 140 bases, 140 bp or 140 nucleotides.
[0114] The panel probe sets as described can be used for targeted enrichment of genomic regions of interest. The panel probe sets are in some embodiments attached to a solid support Accordingly, an aspect of the disclosure includes a target enrichment solid support comprising a solid support and a panel probe set e.g. a plurality of nucleic acid probes complementary to all or a portion of a plurality of genomic regions attached to the solid support. In an embodiment, the solid support comprises a plurality of beads, optionally wherein each bead comprises a distinct nucleic acid probe of the plurality of nucleic acid probes. In an embodiment, the beads are magnetic beads.
[0115] Also described herein are kits optionally for targeted enrichment of one or more target regions comprising one or more probes, panel probe sets, and / or compositions described herein, and optionally one or more reagents for targeted enrichment such as capture beads e.g. streptavidin beads, and / or one or more buffers. The capture beads can for example be magnetic beads comprising a moiety that binds a label on the one or more or all of the probes of the panel probe set.HL Examples
[0116] Example 1 : Methods For The Detection Of Genomic Regions Associated With Fertility And Clinical Outcomes Based On Differential Enrichment For Histone H3 Lysine 4 Tri -Methylation (H3K4me3) In Human Sperm
[0117] Provided herein are the details of a sperm-specific ChlP-seq (chromatin immunoprecipitation followed by high depth sequencing), and bioinformatic analysis that were used to identify regions bearing altered H3K4me3 levels in sperm from infertile men in comparison to fertile men. These methods use data generated from the reference study population of Lambrot et al. 2021 . As shown herein for fertile and idiopathic infertile men (normal semen analysis, no female factor infertility), clinical data including body mass index (BMI) and clinical outcomes was used to identify regions of differential enrichment for H3K4me3 in sperm from men that were fertile versus infertile (Fig. 1).
[0118] Part A: H3K4me3 ChlP-sequencing on human sperm: Approaches detailed in the protocol of Hisano et al., were improved and modified to enhance the sensitivity of the ChIP technique for profilingH3K4me3 in sperm. This was followed by high depth sequencing that allows for 1) the discovery of genomic regions bearing H3K4me3 in sperm; and 2) for detecting differences in H3K4me3 enrichment between fertile and infertile men. These customized approaches produce high quality reproducible data with minimal variation between experiments (Fig. 2 and 3). The mean reads per sample are 33.4M (+ / - 7.9M) with 80-85 aligned.
[0119] I - Preparation of magnetic beads
[0120] Magnetic beads with protein A are used, for example Dynabeads Protein A, Invitrogen, cat#: 10002D.1) Prepare a 0.5% BSA (Bovine Serum Albumin) solution in 1X PBS (Phosphate Buffered Saline = Buffer-BSA). Maintain solution at 4°C throughout the experiment.2) Aliquot 50pl magnetic beads per sperm sample in a tube and 50p I in another tube to preclear the sample of any unspecific binding.3) Wash the beads with Buffer-BSA 3 times: i) Vortex the beads gently to resuspend in Buffer-BSA. ii) Pellet the beads using a magnetic rack and remove the supernatant iii) After the 3rdwash, resuspend the beads to be used for ChIP in Buffer-BSA and the second aliquot of beads for preclearing in Buffer-BSA.4) Add 4pg of H3K4me3 antibody to the tube containing the beads dedicated for the ChIP. H3K4me3 antibody: Tri methyl histone H3K4 (C42D8) Rabbit mAb, cat#: 9751 , Cell Signaling Technology (Danvers, MA, USA).5) Do not add anything to the preclearing tube.6) Place both tubes on a rotator for 6-8h at 4°C.
[0121] II - Cleaning and preparation of sperm7) Sperm can be used that is either fresh or stored at -80 as either:• Neat semen with or without cryoprotection.8) Thaw the sample on ice.9) T ransfer the sperm to a 1 ,5ml tube.10) Add cold PBS.11) Spin the tube in a microcentrifuge at 4000xg for 10 min at 4°C. A pellet containing spermatozoa will form.12) Remove the supernatant.13) Re-suspend gently in cold PBS.14) Spin at 4000xg for 10 min at 4°C.15) Discard the supernatant.16) Flash-freeze the tube containing the washed sperm in liquid nitrogen.17) Place the tube on ice and resuspend the pellet gently in 1 ml cold PBS.18) Repeat steps 14-17.19) Prepare a 1 / 10 dilution of the sperm solution in H2O for counting using an improved Neubauer haematocytometer (e.g. 4pL sperm solution in 36pl H2O).20) Add 10pl diluted sperm to each counting chamber (duplicate counting).21) Incubate for 1 min at RT to allow the sperm to settle.22) Count the 25 central squares of each chamber of the haematocytometer.23) If <99% of the cells are sperm, repeat steps 14-17 as many times as needed to obtain a preparation >99% spermatozoa. Contaminant somatic cells are mostly leukocytes and epithelial cells.24) 12 x106spermatozoa / sample are used for each ChIP experiment.25) Spin at 4000xg for 10 min at 4°C and discard supernatant. Leave the tube on ice. The rest of the sperm not used for the experiment can be frozen in cryoprotectant after centrifugation.
[0122] III - Chromatin preparation26) Prepare lysis Buffer-1.1 solution using Buffer 1 stock.• Buffer-1 stock contains: 1.5ml 1 M Tris-HCI (pH 7.5) (final cone. 15mM Tris-HCI), 6ml 1 M KCI (final cone. 60mM KCI), 0.5ml 1 M MgCI2(final cone. 5mM MgCI2), 20pl 0.5M EGTA (0.1 mM EGTA), 87ml MilliQ. Filter through 0.22pm filter and store at room temperature (RT) for several months.• Add sucrose to the Buffer-1 stock to a 0.3 M final concentration and DTT (dithiothreitol) to a 10 mM final concentration to make Buffer-1 .1 .27) Resuspend the sperm pellet in 300pL Buffer-1 .1 . Pipet vigorously up and down at least 20 times. If the sperm preparation has less than 12x106cells, add 50pl Buffer-1 .1 per 2x106sperm cells (e.g. 150pl Buffer-1.1 for 6x106cells).28) Aliquot the sperm suspension in 6 tubes of 50pl (or less tubes if the ChIP is done on less than 12x106spermatozoa). Leave the tubes on ice.29) Prepare Buffer-1 .2.Use Buffer-1.1 leftover.• Add NP-40 to final concentration of 0.5% (vol / vol) and sodium deoxycholate to a final concentration of 1 % (wt / vol).30) After adding the detergents, vortex Buffer 1 .2 for at least 1 min.31) Add 50pl Buffer-1 .2 to each tube containing 50pl of the sperm in Buffer 1.1.32) Mix by pipetting up and down, gently vortexing and flicking the tubes.33) Quickly spin down using a table centrifuge and incubate for 30 min on ice.
[0123] IV - Mnase digestion:34) Prepare complete Mnase-buffer-2.0 by adding:• sucrose (0.3M) to Mnase buffer stock• 30 Units of Mnase (Roche, cat#: 10107921001) for every 2x106sperm cells (30U / 100pL of cleaned and counted sperm)35) Digestion of the chromatin:Add 10OpI Mnase-buffer-2.0 to each of the 6 tubes simultaneously using a multichannel pipette and mix by pipetting up and down. Incubate at 37°C in a heat block for exactly 5 min.36) Stop reaction by adding 2pl of 0.5M EDTA to each tube simultaneously using a multichannel pipette.37) Vortex and place on ice for 10-20 mins or more.38) Combine tubes into one 1 ,5ml tube.39) Spin the tube at 17000xg for 10 min at RT.40) Transfer supernatant (1.2 mL mono-nucleosomal chromatin) to a clean 1.5ml.41) Add 48pl of 25X PIC (protease inhibitors cocktail).
[0124] V - Preclearinq Step42) Wash the magnetic beads dedicated for preclearing (no antibody) with Buffer-BSA 1 time.43) Resuspend in 50 pl of Combined Buffer (for 1 mL = 475p I Buffer 1 Stock, 475 p I Mnase Buffer Stock, 0.103g Sucrose).44) Add the bead slurry to the tube containing the mono-nucleosomal sperm chromatin (from step 41- 42).45) Rotate at 4°C for 30 mins
[0125] VI - Immunoprecipitation46) Place tube on magnetic stand to pellet preclearing beads.47) Transfer the supernatant (pre-cleared mono-nucleosomal sperm chromatin) to a 1.5ml LoBind Eppendorf tube.48) Leave this tube on ice.49) Stop the rotation of the tube containing the bead-antibody complex.50) Wash in 100pl Buffer-BSA 1 time.51) Wash in 10OpI Combined buffer 1 time.52) Resuspend bead-antibody complex in 50 pl Combined buffer.53) Add bead slurry to precleared mono-nucleosomal sperm chromatin.54) Incubate overnight at 4 °C using rotator.
[0126] VII - Washes55) Put the tube containing the beads-antibody / chromatin on a magnetic rack.56) Discard the supernatant.57) Resuspend beads in 1 ml wash buffer A and incubate at 4°C for 5 min on a rotator.• Wash buffer A contains 50mM Tris-HcL (pH 7.5), 10mM EDTA and 75mM NaCI.58) Centrifuge at 1000xg at 4°C for 1 min to bring beads to bottom of tube.59) Pellet beads on magnetic rack and discard supernatant.60) Add 1 ml wash buffer B and incubate at 4°C for 5 min on a rotator• Wash buffer B contains 50mM Tris-HCI (pH 7.5), 10mM EDTA and 125mM NaCI.61) Centrifuge at 1000xg at 4°C for 1 min. Pellet beads on magnetic rack.62) Discard supernatant and add 1 ml Wash buffer B. Transfer to a new LoBind tube.63) Incubate at 4°C for 5 min on a rotator64) Centrifuge at 1000xg at 4°C for 1 min. Pellet beads on magnetic rack.65) Discard supernatant.
[0127] VIII - Elution of chromatin from the beads66) Add 125pl elution buffer to washed bead-antibody-chromatin complex.• Elution buffer is prepared fresh and contains 0.1 M NaHCCh, 0.2% SDS and 5mM DTT.67) Put tube in a heat block at 65°C for 10 minutes while shaking (4000rpm). Vortex briefly every 3-4 min.68) Transfer tube to a magnetic stand to pellet beads.69) T ransfer supernatant to a new LoBind tube.70) Add another 125pl elution buffer to pelleted beads.71) Repeat step 68-70.72) Final volume is 250pl.
[0128] IX - DNA purification73) Add 5 pl Rnase A (1 Opg / pl) for 30 minutes at 37°C.74) Add 5 pl proteinase K (1 Opg / pl) and incubate overnight at 55°C.75) A kit is used to clean the DNA (ChIP DNA clean & concentrator, cat#: D5205 Zymo Research).76) Add 5X volume ChIP DNA binding bufferto the ChIP sample (=1250 pl Binding Buffer).77) Wash only 1X with 200 pl wash buffer.78) Elute with 15pl EB buffer (10 mM Tris-CI, pH 8.5, Qiagen).79) Repeat elution in same tube for final volume of 30pl.
[0129] X - Quality check and quantitation80) Using an Agilent 2100 Bioanalyzer Instrument with a High Sensitivity DNA Analysis kit (cat#: 5067- 4626).81) Follow manufacturer’s protocol and load 1 pl of ChIP DNA on the bioanalyzer chip.82) A peak of DNA should be observed at 147 base pairs (bp), which is the size of a nucleosome and indicates that the Mnase digestion was effective.83) Additional peaks representing DNA not entirely digested by the Mnase such as dinucleosomes, and trinucleosomes can be observed and this is a size selection step prior to preparing the sequencing libraries is necessary.
[0130] XI - Size selection: This step reduces potential background, improves specificity and detection of enrichment differences for the identification of individual profile differences in H3K4me3 between samples.84) A size selection is performed using size selection beads such as Agencourt AMPure XP beads (Beckman Coulter Inc., cat#: A63880).85) Adjust the volume of the ChIP DNA solution to 50pl with EB buffer.86) Add 1X (50 pl) resuspended Agencourt AMPure XP beads to each sample and mix well by pipetting. This allows for removal of DNA > 200bp with only a minimal loss of the DNA present at 147bp.87) Incubate the mixture at 5 min at room temperature.88) Pellet the beads on a magnetic stand for 2 min, then carefully transfer the supernatant to a clean tube. Discard the beads that are bound to the DNA>200bp.89) Add 2X (compared to original volume) resuspended Agencourt AMPure XP beads to each sample and mix well by pipetting. This allows for the binding of all the DNA that was present in the supernatant of step 89.90) Pellet the beads on the magnetic stand and carefully discard the supernatant.91) Wash the beads by adding 1 ml fresh 80% ethanol to each pellet (do not remove tube from magnetic stand). Carefully discard the supernatant.92) Repeat step 92 for a total of 2 ethanol washes.93) Incubate on the magnetic stand for 10 min or until the beads are dry.94) Elute by resuspending in elution buffer.95) Pellet beads on the magnetic stand.96) Carefully transfer 30p I of supernatant to a tube.97) Do a quality check and quantitation. Repeat the size selection protocol until the 147bp peak represents the desired percentage of the DNA in the sample, for example greater than 90%
[0131] XII - Library preparation
[0132] The QIAseq™ Ultralow Input Library Kit library preparation kit has been validated for this step for use with human sperm ChIP samples.98) the QIAseq™ Ultralow Input Library Kit from Qiagen (Cat#: 180495) is used, following manufacturer’s protocol.99) Start with for example 6ng of cleaned ChIP DNA and do for example 9 cycles of PCR amplification.
[0133] XIII- Deep sequencing
[0134] Libraries were sequenced on Illumina HiSeq machines with a target of >30 million reads per library.Part B: Bioinformatic Analysis for the detection of enrichment differences in sperm H3K4me3 associated with fertility and clinical outcomes
[0135] This analytical approach following the customized ChlP-Seq protocols (detailed above) was used for the detection of enrichment differences in H3K4me3 between sperm samples. By this developed method for analysis, the H3K4me3 profiles that are associated with either fertility and good clinical outcomes (pregnancy) or infertility and poor clinical outcomes (no pregnancy) were identified. The H3K4me3 reference population data set developed in Lambrot et al. 2021 was used as the reference profile data set for H3K4me3 in sperm that was compared to H3K4me3 sperm profiles from either fertile or infertile men (Fig. 1).
[0136] Study Participants: These studies were approved by the McGill Research Ethics Board and informed consent was obtained from all participants. Semen samples were collected by masturbation after at least 3 days of abstinence. Study 1 , participants representative of a reference population in Canada were recruited in three Canadian cities: Toronto, Montreal, and Ottawa (n=33). Information pertaining to the participants age, fertility status, semen analysis etc. is provided in tables S1-S3 below. Study 2, participants (n=132) were recruited from the CreATe Fertility Center (Toronto, Canada) and all had normal semen analysis based on WHO standards and were of either normal BMI (BMI <24.9 kg / m2), or overweight or obese (BMI >24.9 kg / m2). To reduce variability in the reference epigenome, participants were excluded if the sperm DNA Fragmentation Index (DFI) was greater than 30, were more than 50 yrs of age, were a smoker, or had the TT genotype for the C677T SNP of the MTHFR enzyme (a known confounder for the epigenome).Table S1. List of 30 men included in the reference population and their health and sperm parameters.MTHFR: Methylenetetrahydrofolate Reductase; MTHFR C677T SNP (rs1801133), M / mL: million spermatozoa per mL semen; na: not applicable.Table S2. List of 7 fertile men and their health and sperm parameters.M / mL: million spermatozoa per mL semen.Table S3. Sequencing and alignment statistics of the human sperm H3K4me3 ChlP-seq. All samples were mapped to genome assembly hg19.Map Efic =Mapping efficiency Dups = duplicates
[0137] This H3K4me3 differential analysis identified signature differences between fertile and infertile men and revealed that fertility status was associated with BMI. Study 2 focussed on a subset of participants (n=55) with clinical outcome data who met the inclusion criteria. Men were classified as infertile with poor clinical outcomes if there was no identifiable female factor, they had two normal semen analyses, failed IVF, or IUI, there was no pregnancy, or was pregnancy loss. Men were classified as fertile with good clinical outcomes ifpregnancy occurred. Men were stratified by body mass index (BMI) status (normal: BMI <24.9 kg / m2and overweight or obese: BMI > 24.9 kg / m2).Data preprocessingRaw reads were trimmed from the 3’ end to have a phred score of at least 30. Illumina sequencing adapters were removed from the reads. Trimming and clipping were performed using Trimmomatic version 0.36 (Bolger et al. 2014). Each read was mapped against the Feb. 2009 assembly of the human genome (hg19, GRCh37 Genome Reference Consortium Human Reference 37 (GCA_000001405.1) downloaded from UCSC genome browser using Bowtie 2 (Langmead et al. 2012) version 2.3.4.1 . Reads that exhibited more than 3 mismatches were excluded. SAMtools (Li et al. 2009) version 1 .9 was then used to sort and convert SAM files. Sample and H3K4me3 peaks selection based on the reference population data set
[0138] In the reference population data set 50,117 H3K4me3 peaks were identified in sperm (Lambrot et al. 2021). Correlation of counts under the 50K peaks are computed pairwise for all samples. Fig. 3A shows that there is a high degree of correlation in H3K4me3 counts within these reference peaks across samples profiled in this study. Based on the median correlation of counts underthe reference peaks, two sets of samples were selected for the differential binding analyses. The first set includes samples with a median correlation in H3K4me3 counts across samples > 0.66 (n = 50). The second set includes samples with a median correlation in H3K4me3 counts across samples > 0.68 (n = 46).
[0139] In order to ensure a robust detection of H3K4me3 enrichment, reference peaks with lower counts across the 55 samples profiled in this cohort were excluded. Overall, 40,773 peaks with Iog2 (median raw counts) >=4.5 were selected for differential enrichment analyses (Fig. 3B; n = 40,773 peaks).Differential enrichment analysis to identify markers of infertility
[0140] As BMI is an important confounder in the onset of men’s infertility, analyses were stratified on BMI status. Within each stratum (i. normal and ii. Overweight), the differences of H3K4me3 enrichment were investigated, adjusted for experimental batches using DESeq2 as implemented in the DiffBind R package. Differences in H3K4me3 enrichment were identified between fertile and infertile men in both analyses (samples : median corr > 0.66 or 0.68). In the first analysis including 27 samples from overweight men (median corr > 0.66 ; see above (27 samples of the 50 were from overweight men)), 3,634 peaks were identified with differential enrichment according to their fertility status (FDR < 0.2 ; Fig. 1A). In the second analysis including 24 samples from overweight men (median corr > 0.68; see above (24 of the 46 samples were from men that were overweight) 6,106 peaks were identified with differential enrichment according to their fertility status (Fig. 1 B). The majority of the significant peaks in this second list of markers (Fig. 1 B) exhibit a decrease in H3K4me3 binding in infertile men compared to fertile men.Example 2: Infertility and toxicant exposure
[0141] As described in Lismer et al. 2022, ChlP-seq was performed on sperm samples from 48 men with normal semen analysis with either low or high exposure to toxicants (DDT) to identify regions of differential enrichment.
[0142] Briefly, differential binding analysis was conducted to identify regions exhibiting different levels of H3K4me3 binding in sperm of men with either low or high toxicant exposures (lowest fertile of measured DDT levels vs highest fertile of measured DDT levels). To do so the Bioconductor / R package edgeR (version 3.26.0) (Robinson et al., 2010) was employed, using the negative binomial distribution and shrinkage estimates of the dispersions to model read counts. Regions with FDR (Benjamini and Hochberg, 1995) below 0.2 were defined as significantly differentially enriched regions (deH3K4me3) between the two exposure groups (low or high). 1 ,865 peaks with differentially enriched H3K4me3 (deH3K4me3; FDR < 0.2) were identified in men from the low vs high exposure groups.
[0143] Example 3: Markers of Infertility
[0144] Using the methods described in example 1 , ChlP-seq was performed on sperm samples from 25 healthy proven fertile sperm donors, and 41 men being treated for infertility. Sequencing experiments were combined and yielded at least 25 million reads per sample, with >98% mapped reads. This identified 50,117 H3K4me3 peaks. A cutoff of Iog2(median(raw counts)+8) >=4.5 for low reads was applied when all samples were included.
[0145] Differential binding analysis was conducted to identify regions exhibiting different levels of H3K4me3 binding in sperm of fertile or infertile men. To do so the Bioconductor / R package edgeR (version 3.26.0) (Robinson et al., 2010) was employed, using the negative binomial distribution and shrinkage estimates of the dispersions to model read counts. Between fertile and infertile men 8979 regions (FDR < 0.05) were identified as being differentially enriched for H3K4me3 (deH3K4me3). Of these, 4130 regions were at a promoter (+ / - 3kb proximal to the TSS) and categorized as being of medium or high enrichment. High / Medium significant regions were defined based on enrichment levels, with the cut-off being the top 75% of normalized count for these significant regions (mean normalized count greater than 4.38).Example 4: Methods for assessing fertility status
[0146] Based on the regions identified in Examples 1-3, a subset of 482 differentially enriched sperm H3K4me3 regions (Table 1) and 18 control regions (Table 2) were chosen to be targeted by a panel of 140 bp probes (~1 ,988 probes total). For inclusion on the panel, preference was given to differentially enriched regions containing a promoter (+ / - 3kb proximal to the TSS). Further preference was given to regions that were categorized as having medium or high enrichment in the fertile population, and / or regions showing very strong statistical difference between or fertile and infertile or BMI or toxicant exposure. A subset of regions was selected that are in promoters for genes implicated in idiopathic infertility (overlapped with deH3K4me3 genesidentified in example 1), and male infertility genes that were identified as being involved in environmental exposures, lifestyle, embryo development, epigenetics, sensitive to age and neurodevelopment. The regions selected for validation are listed in Table 1 . Probes are designed to capture genomic target regions of interest. Probes for the target regions in Table 1 cover a ~400 base pair region at the promoter of the selected genes. Control target regions were included for normalization and are listed in Table 2.
[0147] 96 sperm samples were subjected to ChlP-seq targeting H3K4me3 followed by hybridization to probes for the target regions, elution and then sequencing. Samples were used from fertile sperm donors and men that were clinically classified as idiopathic infertile, subfertile including mild athenozoosperia, infertile men, idiopathic infertile men with a range in BMI (normal or >24.9), men under 30 and over 40, men of unknown fertility status and fertile men. The bioinformatician was blind to sample status.Table 1 : selected genomic regions showing differential enrichment in infertile men compared to fertile men.Chromosomal locations are with reference to human genome assembly hg19. H3K4me3 methylation may be determined (and / or probes may be designed to enrich for H3K4me3 nucleosome fragments) anywhere within the identified region and an additional + / - 1 ,000bp of the identified start and end (e.g. an additional 1 ,000bp before the identified start location, and / or an additional 1 ,000bp after the identified end location).Table 2: Example control regions. Chromosomal locations are with reference to human genome assembly hg19. H3K4me3 methylation may be determined (and / or probes may be designed to enrich for H3K4me3 nucleosome fragments) anywhere within the identified region and an additional + / - 1 ,000bp of the identified start and end (e.g. an additional 1 ,000bp before the identified start location, and / or an additional 1 ,000bp after the identified end location).Validation and Performance Testing Methods:
[0148] Sample preparation, sequencing, and data processing: Sperm sample processing, ChIP, and library preparation were carried out as described in steps I - XI of Example 1 , Part A, or as described in Example 6, below.
[0149] Following standard protocols, libraries were prepared for hybridization with the amount of each indexed library corresponding to the concentration of each library pool. The capture probes were then mixed with the library primers and hybridized overnight (16 hrs) in a thermal cycler at 85°C. The hybridized targets were then bound to streptavidin beads and incubated at 68°C for 5 min, spun, placed on a magnetic stand for 1 min, washed several times, and the beads were prepared for post-PCR amplification. Each enriched library was validated and quantified using an Agilent Bioanlyzer or similar instrument followed by sequencing on an illumina platform.
[0150] Short sequencing reads were trimmed using Trimmomatic. Reads (optionally trimmed) and aligned against the reference genome using bowtie2 or bwa using any suitable settings. Aligned reads were filtered according to mismatch rate and / or alignment quality score.
[0151] Aligned reads were used to obtain counts under defined the target enriched genomic locations (e.g. TSS + / - 3kB of regions indicated in Table 1). Genomic intervals with low counts e.g. below a selected threshold such as 5 or 10 reads, were removed.
[0152] Building panel and assessing performance: Once counts under genomic intervals of interest were obtained, each probe’s performance to discriminate between a mean fertile and infertile enrichment level was assessed. Each target region was assessed using Mann Whitney tests to identify probes that do not discriminate between fertile and infertile samples. Directionality was also determined as the enrichment levels at the selected targets have a normal based on the reference population, and any deviation from the norm either up or down in enrichment was associated with a profile for example of either fertile or infertile. The statistical approach included the modeling, normalization and / or variance stabilization of countdata. An empirical threshold which maximizes true positives and minimize false positives was set to determine significance.
[0153] A naive Bayes’ classifier was trained to predict fertility status using internal leave-one-out cross-validation. The best set of regions were selected for further validation in an independent cohort.
[0154] Building a predictor and assigning a score to individual patients: Combining coefficients with the relative binding levels of the selected region, a risk score for each individual will be calculated: Risk score = ^=1regioni x coefficient^ Individuals will be classified into low- and high-risk groups based on the optimal risk score cutoff value, which represents the point at which the Youden index (sensitivity + specificity - 1) reached a maximum value. Both the continuous and high / low risk score will be evaluated with relevant clinical variables using a multivariate logistic regression model. The decision on which predictors to include in the model will be based on clinical and statistical reasoning to address challenges such as confounders, interactions and multicollinearity. Forward-backward variable selection will be performed based on Bayesian Information Criterion (BIC). Based on the results of the of the multivariate analysis, a nomogram formulated to predict high to low risk of infertility will be built. The predictive accuracy and discrimination ability of the nomogram will be determined using a calibration plot and C-index, respectively.
[0155] Other scoring metrics for indicating the risk of infertility may also be developed, for example taking into account the number of reads from the one or more targets from the patient sample (optionally normalized to one or more internal controls), the level of enrichment of the one or more targets relative to a control or one or more other targets, and / or whether the enrichment of one or more targets is above or below a threshold value specific for each of the one or more targets.
[0156] Nomograms or other scoring metrics will be used to predict the risk of infertility (for example on a scale of high to low) for each individual patient, and / or to classify individuals into low- and high-risk groups.
[0157] Results: The panel probe set targets were tested using samples from sperm donors and from men being treated for infertility. Sperm samples (n=96) from sperm donors and clinical samples from men with various categories of infertility were subjected to chromatin immunoprecipitation (ChIP) followed by capture of the target regions using a total of 1988 probes targeting each of the 482 regions of Table 1 and each of the 18 regions identified in Table 2, then sequenced. Contiguous regions (overlapping regions) were removed, and duplicates were marked. Bioinformatic analysis identified regions that differed in enrichment between samples. As shown in Figure 4, a subset of n=44 of the samples clustered into 4 categories that were termed fertile 1 (e.g. most fertile), fertile 2 (e.g. subgroup), infertile 1 (e.g. subgroup) and infertile 2 (e.g most infertile). This demonstrates the ability of the panel to discriminate differences in fertility status. Upon further testing, target enrichment levels will be assigned a fertility score that is associated with being fertile, or infertile and predicted clinical treatment outcomes (pregnant vs no pregnancy) that will be used to personalize fertility treatment.
[0158] Example 5: Targeted Cleavage-based H3K4me3 Nucleosome and Library Preparation
[0159] Sperm samples can be processed to purity and lysed as described in Example 1. Libraries comprising H3K4me3-associated genomic fragments can be prepared using a variety of techniques prior to enriching for target regions.I. Library Preparation using native ChIP
[0160] Chromatin cleavage and library preparation can be carried out according to the protocols described in Lambrot et al. 2021 , or Example 1 , Part A, or modified versions thereof, and using H3K4me3 antibody (e.g. Cell Signaling, catalogue #9751). Briefly, purified sperm are lysed and chromatin is digested with suitable nuclease e.g. Mnase to obtain mononucleosomes. Alternatively, monoucleosomes can be prepared from chromatin by sonication. Lysate with mononucleosomes is optionally precleared with unlabeled beads, and then incubated with an anti-H3K4me3 antibody and an antibody capture agent (e.g. protein A beads). The anti-H3K4me3 antibody may be preincubated with the antibody capture agent, or the antibody capture agent may be added subsequently. H3K4me3 mononucleosomes are precipitated, washed, and eluted. Mononucleosome DNA is optionally purified by treatment with an Rnase (e.g. Rnase A) and / or a proteinase (e.g. proteinase K) and purified using standard protocols. DNA fragments are optionally size-selected until the 147 bp peak represents the desired percentage of DNA in the sample (e.g. greater than 90%). Sequencing libraries are then prepared using standard library preparation techniques e.g. QIAseq™ Ultralow Input Library Kit. Sequencing libraries can optionally be enriched for specific targets as described in Example 3. Sequencing and data analysis can be carried out as described above.II. Library Preparation using Cleavage Under Targets and Tagmentation (CUT&Tag)
[0161] Cleavage Under Target and Tagmentation (CUT & TAG) is an immunotethering approach to profile chromatin. Chromatin cleavage and library preparation will be carried out according to the protocols described in Kaya-Okur et al., CUT&Tag for efficient epigenomic profiling of small samples and single cells. Nat Commun. 2019 Apr 29;10(1):1930. Doi: 10.1038 / s41467-019-09982-5, and using H3K4me3 antibody (e.g. Cell Signaling, catalogue #9751). The steps of this protocol include immobilizing cells, adding histone H3K4me3 antibody (e.g. Cell Signaling, catalogue #9751), and secondary antibody, followed by adding activated pAG- T n5 transposase to cleave target sites and add adaptors. This is followed by target enrichment for all or a subset of specific targets identified in Table 1 and optionally Table 2, then the tagmented DNA is amplified by PCR and sequenced. Sequencing and data analysis will be carried out as described above.III. Library preparation using CUT&RUN
[0162] Cleavage Under Targets and Release Using Nuclease (CUT & RUN) is an immunotethering approach to profile chromatin described in Skene & Henikoff. It is an efficient targeted nuclease strategy for high-resolution mapping of DNA binding sites. Elife. 2017 Jan 16;6:e21856. Doi: 10.7554 / eLife.21856) The steps of this protocol include immobilizing cells, adding histone H3K4me3 antibody (e.g. Cell Signaling, catalogue #9751), adding activated pAG-Mnase to cleave the target DNA complex, extract DNA, prepare libraryand sequence. Sequencing libraries will optionally be enriched for all or a subset of specific targets identified in Table 1 and optionally Table 2. Sequencing and data analysis will be carried out as described above.IV. Library preparation using TAM-ChIP™
[0163] TAM-ChIP includes chromatin immunoprecipitation (ChIP) followed by a secondary antibody- directed protein targeting of the primary antibody target followed by library preparation. The TAM secondary antibody is conjugated to TsTn5 transposase and an oligonucleotide (US Pat. Nos. 9,938,524 & 10,689,643; EP Pat. Nos. 2783001 & 2999784). ChIP is performed using histone H3K4me3 (e.g. Cell Signaling, catalogue #9751), then, a species-specific TAM-ChIP antibody conjugate containing Illumina-compatible sequencing adapters is added to bind the ChIP antibody. Activation of the TsTn5 transposase cuts the nearby DNA surrounding the genomic region of interest and pastes the antibody-associated adapters into the DNA sequence. Following washing, elution, Proteinase K treatment, reversal of cross-links and DNA purification, the ChIP library is ready for amplification. Following immunoprecipitation, cross-links are reversed, the proteins are removed by Proteinase K, and the DNA is recovered and purified. Sequencing libraries will optionally be enriched for all or a subset of specific targets as identified in Table 1 and optionally Table 2. Sequencing and data analysis will be carried out as described above.Example 6: Panels for predicting fertility status in individual patients
[0164] A sperm-customized targeted epigenome panel to sensitively detect differences in H3K4me3 enrichment associated with fertility status is designed. For example, based on the results from Example 4, probes targeting specified genomic regions (Table 1 and / or 2), or a subset thereof, optionally in combination with additional probes targeting other regions, will be used in a fertility panel for semen analysis.
[0165] The panel may be for 1) a more accurate fertility diagnosis in particular for men who would be identified as idiopathic infertile under current clinical approaches; 2) prediction of clinical outcomes; and / or 3) identification of modifiable infertility with adaptation of a healthy lifestyle including cessation of nicotine / cannabis use, weight loss and a healthy diet. Men will be assigned a fertility score across the categories of target probes. Based on this score they will be provided with a personalized treatment plan that could involve intervention to improve their fertility score and the results will allow them to make an informed decision regarding treatment options based on the predicted success of intrauterine insemination (IUI) vs in vitro fertilization (IVF) derived from that fertility score. Men may then be counseled to follow customized pre-conception plans, such as a dietary plan, a weight loss plan, an exercise plan, nicotine / cannabis cessation, and / or limiting or reducing exposure to one or more specific toxicants, optionally with repeat testing after 3-6 months which is equivalent to one or two spermatogenic cycles. Dietary plans may include for example increased intake of specific nutrients (e.g. folate) or decreased caloric intake. Weight loss plans may include for example interventions such as a weight loss drug (e.g. Bupropion-naltrexone (Contrave™), Liraglutide (Saxenda™), Orlistat (Xenical™, Alli™), Phentermine-topiramate (Qsymia™), Semaglutide (Wegovy™, Ozempic™), or Setmelanotide (Imcivree™)), dietary interventions (e.g. decreased caloric intake), and / or surgery.
[0166] Sample preparation, sequencing, and data processing for each individual patient will be carried out for example as described in the Examples above. The sequencing library will be enriched for target regions of interest prior to sequencing, for example as described in Example 4. Target regions of interest are enriched for example by being captured by hybridization to biotinylated probes and then isolated by magnetic pulldown. The number of probes used to capture a target region may depend in part on the size of the target region. In some cases as few as one or two probes could be used to enrich for each target region. In most cases 4 probes or fewer are used to enrich for the target region.
[0167] The nomogram or other scoring metrics developed in Example 4 will be applied to the sequencing data to assess the sperm epigenomic health of the patient, predict the risk of infertility for the patient or to classify the individual as a low- or high-risk group.
[0168] The methods described herein may also be used in conjunction with existing semen analysis techniques which measure for example semen volume, pH, sperm count, sperm motility, and / or sperm morphology.Example 7
[0169] The methods described herein can be performed using a two stage pipeline. In the first stage data is preprocessed and in the second stage data is analysed.For example, sequence reads can be raw reads obtained by paired-end sequencing and provided as a FASTQ format. Optionally a quality control check can be performed, for example by FASTQC. The raw reads can be filtered and trimmed (e.g. by Trimgalore). The sequences reads can be mapped for example using BowTie2, contiguous reads removed and a mismatch threshold set ex. >3. Duplicates can be marked for removal or retained and the data downsampled or not. The data is then normalized using the control probe set followed by visualization using for example heat maps, MA plots and PCA plots. Machine learning is then used (example Gradient boosting) to classify the sample as fertile or infertile.Example 8
[0170] Other genome reads other than GRCh37 can be used. For example, the genomic regions can readily be mapped to GRCh38. The methods have been repeated using GRCh38 with the same results.
[0171] While the present application has been described with reference to examples, it is to be understood that the scope of the claims should not be limited by the embodiments set forth in the examples, but should be given the broadest interpretation consistent with the description as a whole.
[0172] All publications, patents and patent applications are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety. Where a term in the present application is found to be defined differently in a document incorporated herein by reference, the definition provided herein is to serve as the definition for the term.References:Bolger, A.M., Lohse, M. and Usadel, B. (2014) Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 30, 2114-2120.Brykczynska, U., Hisano, M., Erkek, S., Ramos, L., Oakeley, E. J., Roloff, T. C., Beisel, C., Schubeler, D., Stadler, M. B. & Peters, A. H. Repressive and active histone methylation mark distinct promoters in human and mouse spermatozoa. Nat Struct Mol Biol 17, 679-687, doi:10.1038 / nsmb.1821 (2010).Carone, B. R., Fauquier, L., Habib, N., Shea, J. M., Hart, C. E., Li, R., Bock, C., Li, C., Gu, H., Zamore, P. D., Meissner, A., Weng, Z., Hofmann, H. A., Friedman, N. & Rando, O. J. Paternally induced transgenerational environmental reprogramming of metabolic gene expression in mammals. Cell 143, 1084-1096, doi:10.1016 / j.cell.2010.12.008 (2010).Hajizadeh Maleki, B. & Tartibian, B. Moderate aerobic exercise training for improving reproductive function in infertile patients: A randomized controlled trial. Cytokine 92, 55-67, doi:10.1016 / j.cyto.2017.01.007 (2017). (2107b)Hajizadeh Maleki, B. & Tartibian, B. Resistance exercise modulates male factor infertility through antiinflammatory and antioxidative mechanisms in infertile men: A RCT. Life Sciences 203, 150-160, doi: 10.1016 / j.lfs.2O18.04.039 (2018).Hammoud, S. S., Nix, D. A., Zhang, H., Purwar, J., Carrell, D. T. & Cairns, B. R. Distinctive chromatin in human sperm packages genes for embryo development. Nature 460, 473-478, doi:10.1038 / nature08162 (2009).Hisano, M., Erkek, S., Dessus-Babus, S. et al. Genome-wide chromatin analysis in mature mouse and human spermatozoa. Nat Protoc 8, 2449-2470 (2013).Jung, Y. H., Sauria, M. E. G., Lyu, X., Cheema, M. S., Ausio, J., Taylor, J. & Corces, V. G. Chromatin States in Mouse Sperm Correlate with Embryonic and Adult Regulatory Landscapes. Cell Rep 18, 1366-1382, doi:10.1016 / j.celrep.2017.01 .034 (2017).Kimmins, S., Anderson, R.A., Barratt, C.L.R. et al. Frequency, morbidity and equity — the case for increased research on male fertility. Nat Rev Urol (2023).Lambrot, R., Xu, C., Saint-Phar, S., Chountalos, G., Cohen, T., Paquet, M., Suderman, M., Hallett, M. & Kimmins, S. Low paternal dietary folate alters the mouse sperm epigenome and is associated with negative pregnancy outcomes. Nat Commun 4, 2889, doi:10.1038 / ncomms3889 (2013).Lambrot, R., Siklenka, K., Lafleur, C., and Kimmins, S. (2019). The Genomic Distribution of Histone H3K4me2 in Spermatogonia is Highly Conserved in Sperm. Biol Reprod. Jun 1 ;100(6):1661 -1672.Lambrot R, Chan D, Shao X, Aarabi M, Kwan T, Bourque G, Moskovtsev S, Librach C, Trasler J, Dumeaux V, Kimmins S. Whole-genome sequencing of H3K4me3 and DNA methylation in human sperm reveals regions of overlap linked to fertility and development. Cell Rep. 2021 Jul 20;36(3):109418. doi: 10.1016 / j.celrep.2021.109418. PMID: 34289352.Langmead, B., and Salzberg, S.L. (2012). Fast gapped-read alignment with Bowtie 2. Nat Methods 9, 357-359.Li, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N., Marth, G., Abecasis, G., Durbin, R., and Genome Project Data Processing, S. (2009). The Sequence Alignment / Map format and SAMtools. Bioinformatics 25, 2078-2079.Lismer A, Siklenka K, Lafleur C, Dumeaux V, Kimmins S. Sperm histone H3 lysine 4 trimethylation is altered in a genetic mouse model of transgenerational epigenetic inheritance. Nucleic Acids Res. 2020 Nov 18;48(20):1 1380-11393. doi: 10.1093 / nar / gkaa712. PMID: 33068438; PMCID: PMC7672453.Lismer A, Dumeaux V, Lafleur C, Lambrot R, Brind'Amour J, Lorincz MC, Kimmins S. Histone H3 lysine 4 trimethylation in sperm is transmitted to the embryo and associated with diet-induced phenotypes in the offspring. Dev Cell. 2021 Mar 8;56(5):671-686.e6. doi: 10.1016 / j.devcel.2021 .01 .014. Epub 2021 Feb 16. PMID: 33596408.Lismer, A., Shao, X., Dumargne, M.C. et al. Exposure of Greenlandic Inuit and South African VhaVenda men to the persistent DDT metabolite is associated with an altered sperm epigenome at regions implicated in paternal epigenetic transmission and developmental disease - a cross-sectional study. bioRxiv 2022.08.15.504029; doi: https: / / doi.org / 10.1101 / 2022.08.15.504029.Lismer A, Kimmins S. Emerging evidence that the mammalian sperm epigenome serves as a template for embryo development. Nat Commun. 2023 Apr 14; 14(1 ):2142. doi: 10.1038 / s41467-023-37820-2. PMID: 37059740.Maleki, B. H. & Tartibian, B. High-Intensity Exercise Training for Improving Reproductive Function in Infertile Patients: A Randomized Controlled Trial. J Obstet Gynaecol Can 39, 545-558, doi:10.1016 / j.jogc.2017.03.097 (2017). (2017a)Ng, S. F., Lin, R. C., Laybutt, D. R., Barres, R., Owens, J. A. & Morris, M. J. Chronic high-fat diet in fathers programs beta-cell dysfunction in female rat offspring. Nature 467, 963-966, doi:10.1038 / nature09491 (2010).Pepin AS, Lafleur C, Lambrot R, Dumeaux V, Kimmins S. Sperm histone H3 lysine 4 tri-methylation serves as a metabolic sensor of paternal obesity and is associated with the inheritance of metabolic dysfunction. Mol Metab. 2022 May;59:101463. doi: 10.1016 / j.molmet.2022.101463. Epub 2022 Feb 17. PMID: 33795.Siklenka, K. et al. Disruption of histone methylation in developing sperm impairs offspring health transgenerationally. Science. 350, aab2006 (2015).Zhang, Y., Liu, T., Meyer, C.A. et al. Model-based Analysis of ChlP-Seq (MACS). Genome Biol 9, R137 (2008). https: / / doi.Org / 10.1186 / gb-2008-9-9-r137.Zhang, B., Zheng, H., Huang, B. et al. Allelic reprogramming of the histone modification H3K4me3 in early mammalian development. Nature 537, 553-557 (2016).
Claims
CLAIMS:1 . A method for determining sperm histone 3 lysine 4 trimethyl (H3K4me3) nucleosome methylation levels in a subject, the method comprising: a) obtaining a first plurality of sequence reads from a semen sample or a sperm sample obtained from the subject, wherein the sequence reads are from H3K4me3 nucleosome fragments; and b) determining a H3K4me3 level at each of a plurality of genomic regions, the plurality of genomic regions comprising at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 482, or any number in between and including 5 and 482 genomic regions identified in Table 1 , and optionally at least one genomic region selected from Table 2, using the first plurality of sequence reads.
2. The method of claim 1 , wherein the plurality of genomic regions comprises: a) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 153, or any number in between and including 5 and 153 genomic regions that are categorized in table 1 as Category 1 ; b) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 59, or any number in between and including 5 and 59 genomic regions that are categorized in table 1 as Category 2; c) at least 5, 10, 15, 20, 25, 30, 35, 40, or any number in between and including 5 and 40 genomic regions that are categorized in table 1 as Category 3; d) at least 5, 10, 13, or any number in between and including 5 and 13 genomic regions that are categorized in table 1 as Category 4; e) at least 5, 10, 15, 20, 25, 30, 35, 40, 41 , or any number in between and including 5 and 41 genomic regions that are categorized in table 1 as Category 5; f) at least 5 or 6 genomic regions that are categorized in table 1 as Category 6; g) at least 5, 9, or any number in between and including 5 and 9 genomic regions that are categorized in table 1 as Category 7; h) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 51 , or any number in between and including 5 and 51 genomic regions that are categorized in table 1 as Category 8; i) at least 5, 10, 15, 20, 25, 29, or any number in between and including 5 and 29 genomic regions that are categorized in table 1 as Category 9; j) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 68, or any number in between and including 5 and 68 genomic regions that are categorized in table 1 as Category 10; k) at least 5, 10, 13, or any number in between and including 5 and 13 genomic regions that are categorized in table 1 as Category 11 ; l) optionally at least 1 , at least 2, at least 3, at least 4, at least 5, or 6 genomic regions that are categorized in table 2 as Category A;m) at least 1 , at least 2, at least 3, at least 4, at least 5, or 6 genomic regions that are categorized in table 2 as Category B; and / or n) at least 1 , at least 2, at least 3, at least 4, at least 5, or 6 genomic regions that are categorized in table 2 as Category C.
3. The method of claim 1 or claim 2, wherein the plurality of genomic regions comprises: at least 5 genomic regions that are categorized in table 1 as Category 1 ; at least 5 genomic regions that are categorized in table 1 as Category 2; at least 5 genomic regions that are categorized in table 1 as Category 3; at least 5 genomic regions that are categorized in table 1 as Category 4; at least 5 genomic regions that are categorized in table 1 as Category 5; at least 5 genomic regions that are categorized in table 1 as Category 6; at least 5 genomic regions that are categorized in table 1 as Category 7; at least 5 genomic regions that are categorized in table 1 as Category 8; at least 5 genomic regions that are categorized in table 1 as Category 9; at least 5 genomic regions that are categorized in table 1 as Category 10; at least 5 genomic regions that are categorized in table 1 as Category 1 1 ; and optionally at least 5 genomic regions that are categorized in table 2 as Category C, and at least 2 genomic regions that are categorised in table 2 as Category B.
4. The method of any one of claims 1 to 3, wherein the sperm sample is at least 99% purified sperm or the semen sample is neat semen, optionally fresh or frozen and / or wherein the subject is a sperm donor.
5. The method of claim 4, wherein the sperm sample that is at least 97%, 98% or 99% purified sperm is obtained by a method comprising the steps of a) obtaining the neat semen sample comprising sperm, optionally thawing if frozen, b) pelleting the sperm, c) resuspending the sperm in buffer, d) optionally flash-freezing the resuspended sperm, e) if frozen, thawing the sperm, f) determining the purity of the sperm, and g) repeating steps b) to f) until the sample is at least 97%, 98% or 99% sperm.
6. The method of any one of claims 1 to 5, wherein the histone 3 lysine 4 trimethylated (H3K4me3) nucleosome fragments are mono-nucleosomal fragments, and optionally are isolated by a method comprising one or more of the following steps: obtaining a semen sample comprising sperm from the subject or that has been obtained from the subject; purifying the semen sample to obtain a sperm sample, treating the sperm sample with MNase to obtain substantially mono-nucleosome fragments; and precipitating H3K4me3 nucleosome fragments using an H3K4me3-specific antibody optionally in combination with an antibody capture agent, optionally a magnetic antibody capture agent.
7. The method of claim 6, wherein the H3K4me3-specific antibody is C42D8.
8. The method of claim 6 or 7, wherein the magnetic antibody capture agent is a magnetic bead.
9. The method of any one of claims 1 to 8, wherein the method further comprises purifying DNA from the H3K4me3 nucleosome fragments and subjecting the purified DNA to a size-selection step.
10. The method of any one of claims 1 to 9, wherein obtaining the sequence reads comprises assaying by next-generation sequencing (NGS) and optionally the method further comprises generating a sequencing library and polymerase chain reaction (PCR)-amplifying the library prior to sequencing.11 . The method of any one of claims 1 to 10, wherein the method further comprises targeted enrichment of H3K4me3 nucleosome fragments from the plurality of genomic regions.
12. The method of claim 11 , wherein the targeted enrichment comprises hybridization capture-based enrichment.
13. The method according to any one of claims 1 to 12 further comprising a) after determining the H3K4me3 nucleosome methylation levels in the subject according to the methods of any one of claims 1 to 12; and b) selecting a sperm sample if the subject has a fertile profile and / or predicting if the subject is infertile using the determined H3K4me3 nucleosome methylation levels.
14. The method of claim 13, wherein step b) comprises i) comparing the H3K4me3 level at each genomic region of the plurality of genomic regions to a control H3K4me3 level or reference value for each respective genomic region in the plurality of genomic regions, wherein the H3K4me3 level at one or more regions relative to control H3K4me3 level or reference value is predictive of infertility; and / or ii) calculating the risk of infertility of the subject based on statistical modeling; and optionally classifying the subject as fertile or infertile.
15. The method according to any one of claims 1 to 12, further comprising a) after determining the H3K4me3 nucleosome methylation levels in the subject according to the methods of any one of claims 1 to 12; and b) assessing the sperm epigenomic health of the subject using the determined H3K4me3 nucleosome methylation levels.
16. The method of claim 15, wherein step b) comprises i) comparing the H3K4me3 level at each genomic region of the plurality of genomic regions to a control H3K4me3 level or reference value for each respective genomic region in the plurality of genomic regions, wherein the H3K4me3 level at one or more regions relative to control H3K4me3 level or reference value is indicative of sperm epigenomic health; and / or ii) determining the sperm epigenomic health of the subject based on statistical modeling.
17. The method of claim 14 or claim 16, wherein the control H3K4me3 level or reference value for each respective genomic region in the plurality of genomic regions is obtained from a reference population, optionally a fertile reference population and / or infertile reference population.
18. The method of any one of claims 15 to 17, wherein assessing sperm epigenomic health of the subject comprises one or more of: classifying the subject as fertile or infertile; predicting the clinical outcome of assisted reproductive therapy (ART); predicting the clinical outcome of embryo development; and / or predicting the clinical outcome of lifestyle intervention on fertility, ART outcome, and / or embryo development, optionally lifestyle intervention comprises cessation of cannabis use, dietary changes and / or weight loss.
19. The method of any one of claims 15 to 18, wherein the H3K4me3 nucleosome methylation levels and / or the sperm epigenomic health of the subject is compared to H3K4me3 nucleosome methylation levels and / or sperm epigenomic health determined previously, optionally 3-6 months prior.
20. A panel probe set for assessing sperm epigenomic health of a subject, the probe set comprising at least two probes for each of a plurality of genomic regions, the plurality of genomic regions comprising at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 482, or any number in between and including 5 and 482 of the genomic regions identified in Table 1 and optionally at least two probes for each of at least one of the genomic regions identified in Table 2.21 . The panel probe set of claim 20, wherein the plurality of genomic regions comprises: a) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 153, or any number in between and including 5 and 153 genomic regions that are categorized in table 1 as Category 1 ; b) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 59, or any number in between and including 5 and 59 genomic regions that are categorized in table 1 as Category 2; c) at least 5, 10, 15, 20, 25, 30, 35, 40, or any number in between and including 5 and 40 genomic regions that are categorized in table 1 as Category 3; d) at least 5, 10, 13, or any number in between and including 5 and 13 genomic regions that are categorized in table 1 as Category 4; e) at least 5, 10, 15, 20, 25, 30, 35, 40, 41 , or any number in between and including 5 and 41 genomic regions that are categorized in table 1 as Category 5; f) at least 5 or 6 genomic regions that are categorized in table 1 as Category 6; g) at least 5, 9, or any number in between and including 5 and 9 genomic regions that are categorized in table 1 as Category 7; h) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 51 , or any number in between and including 5 and 51 genomic regions that are categorized in table 1 as Category 8; i) at least 5, 10, 15, 20, 25, 29, or any number in between and including 5 and 29 genomic regions that are categorized in table 1 as Category 9; j) at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 68, or any number in between and including 5 and 68 genomic regions that are categorized in table 1 as Category 10; k) at least 5, 10, 13, or any number in between and including 5 and 13 genomic regions that are categorized in table 1 as Category 11 ;l) optionally at least 1 , at least 2, at least 3, at least 4, at least 5, or 6 genomic regions that are categorized in table 2 as Category A; m) at least 1 , at least 2, at least 3, at least 4, at least 5, or 6 genomic regions that are categorized in table 2 as Category B; and / or n) at least 1 , at least 2, at least 3, at least 4, at least 5, or 6 genomic regions that are categorized in table 2 as Category C.
22. The panel probe set of claim 20 or claim 21 , wherein the plurality of genomic regions comprises: at least 5 genomic regions that are categorized in table 1 as Category 1 ; at least 5 genomic regions that are categorized in table 1 as Category 2; at least 5 genomic regions that are categorized in table 1 as Category 3; at least 5 genomic regions that are categorized in table 1 as Category 4; at least 5 genomic regions that are categorized in table 1 as Category 5; at least 5 genomic regions that are categorized in table 1 as Category 6; at least 5 genomic regions that are categorized in table 1 as Category 7; at least 5 genomic regions that are categorized in table 1 as Category 8; at least 5 genomic regions that are categorized in table 1 as Category 9; at least 5 genomic regions that are categorized in table 1 as Category 10; at least 5 genomic regions that are categorized in table 1 as Category 11 ; and optionally at least 5 genomic regions that are categorized in table 2 as Category C, and at least 2 genomic regions that are categorised in table 2 as Category B.
23. The panel probe set of any one of claims 20 to 22, wherein the probe set comprises at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11 , at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 probes for each of the plurality of genomic regions.
24. The panel probe set of any one of claims 20 to 23, wherein each probe is 100-150 nucleotides, 120- 150 nucleotides, 130-150 nucleotides, 140-150 nucleotides, or any number in between and including 100-150 nucleotides, optionally about or 140 nucleotides in length, each optionally labelled.
25. A targeted capture solid support comprising a solid support and the panel probe set of any one of claims 20 to 24 or a kit comprising the panel probe set of any one of claims 20 to 24 or the targeted capture solid support.
26. The targeted capture solid support or kit of claim 25, wherein the solid support comprises a plurality of beads, optionally magnetic beads, wherein each bead comprises a distinct probe of the probe set.
27. The method of any one of claims 1 to 19, wherein the H3K4me3 nucleosome methylation level is determined using the panel probe set of any one of claims 20 to 24 or the targeted capture panel or kit of claim 25 or 26.
Citation Information
Patent Citations
Method for assessing fertility based on male and female genetic and phenotypic data
US20170351806A1
Diagnostic use of cell free DNA chromatin immunoprecipitation
US20210024994A1