Minor histocompatibility antigen markers associated with graft versus host disease and uses thereof

An analytic framework using SNP analysis and kits to predict GvHD risk by identifying specific mHAgs in donors and recipients, addressing the inconsistency in linking mHAg repertoires to clinical outcomes and enabling targeted donor selection to reduce GvHD.

WO2025174666A1PCT designated stage Publication Date: 2025-08-21DANA FARBER CANCER INSTITUTE INC +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/014995
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-06
Filing Date
2025-02-07
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing technologies fail to consistently link patient-specific minor histocompatibility antigen (mHAg) repertoires to clinical outcomes in graft-versus-host disease (GvHD) following allogeneic hematopoietic stem cell transplantation, particularly in predicting acute and chronic GvHD in specific organs like the liver and lungs.

Method used

An analytic framework utilizing whole exome sequencing and SNP analysis to identify specific SNPs associated with mHAgs, allowing for the prediction of GvHD risk by comparing donor and recipient polymorphisms, and employing kits with probes or primers to detect these SNPs for donor and recipient samples.

Benefits of technology

Enables the identification of transplant recipients at increased or decreased risk of developing organ-specific GvHD, facilitating donor selection to minimize GvHD through targeted immunosuppressive therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000055_0000
    Figure 00000055_0000
  • Figure 00000056_0000
    Figure 00000056_0000
  • Figure 00000057_0000
    Figure 00000057_0000
Patent Text Reader

Abstract

Minor histocompatibility antigens (mHAgs) associated with graft versus host disease (GvHD) clinical outcomes identified by single nucleotide polymorphisms (SNPS) and uses thereof are described. Further, mHAgs cross-reactive against gut-tropic viral epitopes associated with acute GvHD of the gut and uses thereof are described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] MINOR HISTOCOMPATIBILITY ANTIGEN MARKERS ASSOCIATED

[0002] WITH GRAFT VERSUS HOST DISEASE AND USES THEREOF

[0003] CONTINUING APPLICATION DATA

[0004] This application claims the benefit of U.S. Provisional Application Serial Nos. 63 / 554,484, filed February 16, 2024, and 63 / 728,785, filed December 6, 2024, each of which is incorporated by reference herein.

[0005] GOVERNMENT FUNDING

[0006] This invention was made with government support under Grant Nos. 1R01HL157174 and P01 CA229092, awarded by the National Institutes of Health. The Government has certain rights in the invention.

[0007] BACKGROUND

[0008] T cell alloreactivity against minor histocompatibility antigens (mHAgs), polymorphic peptides resulting from donor-recipient (D-R) disparity at sites of genetic polymorphisms (for example, SNPs, indels, and frameshifts), is at the core of the therapeutic effect of allogeneic hematopoietic stem cell transplantation (allo-HSCT). Despite the crucial role of mHAgs in graft- versus-leukemia (GvL) and graft-versus-host (GvHD) reactions, there is a need for more information to consistently link patient-specific mHAg repertoires to clinical outcomes.

[0009] SUMMARY

[0010] There are provided methods of identifying a transplant recipient at increased risk of developing acute liver graft versus host disease (GvHD) after receipt of donor tissue from a donor, the method including: determining for the donor the presence or absence of one or more single nucleotide polymorphisms (SNPs) from a panel of SNPs; determining for the recipient the presence or absence of one or more SNPs from the panel of SNPs; wherein the panel of SNPs includes the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NCOA7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q68OR SNP, the UBE2O p.G1207S SNP, and / or the VSTM4 p.F68S SNP; wherein the donor and recipient are HLA-matched; and wherein the presence in the recipient and the absence in the donor of one or more of the SNPs from the panel of SNPs is indicative of an increased risk of the recipient developing acute liver GvHD after receipt of donor tissue from the donor. In some aspects, the methods further include administering an immunosuppressive agent to the transplant recipient.

[0011] Also provided are methods for selecting a transplant donor, the method including: obtaining single nucleotide polymorphism (SNP) information for a potential donor for one or more SNPs from a panel of SNPs; wherein the panel of SNPs includes the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP; determining for the potential donor the presence or absence of one or more SNPs from the panel of SNPs; obtaining SNP information for a transplant recipient for one or more SNPs from the panel of SNPs; and determining for the transplant recipient the presence or absence of one or more SNPs from the panel of SNPs; wherein the potential donor and transplant recipient are HLA-matched; and selecting as the transplant donor a potential donor wherein one or more SNPs from the panel of SNPs is present in both the potential donor and the transplant recipient or one or more SNPs from the panel of SNPs is absent in both the potential donor and the transplant recipient.

[0012] In some aspects of the methods provided herein, the one or more SNPs from the panel of SNPs includes the ATXN2 p.L107V SNP and the HLA-matched donor and recipient includes HLA B0801, HLA B0702, HLA B55O1, HLA C0602, HLA B2703, and / or HLA A0301; the one or more SNPs from the panel of SNPs includes the C4orf48 p,17L SNP and the HLA-matched donor and recipient includes HLA A2402, HLA C0303, HLA A0201, HLA C0602, HLA C0701, HLA C0401, HLA Cl 203, and / or HLA B0801; the one or more SNPs from the panel of SNPs includes the NCOA7 p.S284A SNP and the HLA-matched donor and recipient includes HLA A0301, HLA C0303, HLA C0304, HLA B4402, and / or HLA C0802; the one or more SNPs from the panel of SNPs includes the TGFB1 p.PlOL SNP and the HLA-matched donor and recipient includes HLA A0201, HLA B0702, and / or HLA C0701; the one or more SNPs from the panel of SNPs includes the TNC p.Q680R SNP and the HLA-matched donor and recipient includes HLA A0101, HLA Bl 801, HLA Bl 501, HLA C0602, HLA A6801, HLA A0201, and / or HLA B4402; the one or more SNPs from the panel of SNPs includes the UBE2O p.G1207S SNP and the HLA-matched donor and recipient includes HLA C0501 , HLA B4101, HLA B4001, and / or HLA B3701; and / or the one or more SNPs from the panel of SNPs includes the VSTM4 p.F68S SNP and the HLA-matched donor and recipient includes HLA C0501, HLA C0701, HLA C0303, HLA C0304, HLA C0501, HLA C0401, HLA C0702, and / or HLA C0802.

[0013] Also provided herein are kits including a plurality of probes capable of binding to and / or identifying one or more of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP in a sample from the transplant donor and / or the transplant recipient. In some aspects, the kit includes a plurality of probes capable of binding to and / or identifying each of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and the VSTM4 p.F68S SNP in a sample from the transplant donor and / or the transplant recipient.

[0014] Also provided herein are kits including a plurality of primers for selectively amplifying one or more of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP in a sample from the transplant donor and / or the transplant recipient. In some aspects, the kit includes a plurality of primers for selectively amplifying each of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and the VSTM4 p.F68S SNP in a sample from the transplant donor and / or the transplant recipient.

[0015] Also provided herein are methods of determining a transplant recipient’s risk of developing chronic lung graft versus host disease (GvHD) after receipt of donor tissue from a donor, the method including: obtaining single nucleotide polymorphism (SNP) information for the donor for a panel of lung-specific minor histocompatibility antigen (mHAg) SNPs; wherein the panel of lung-specific mHAg SNPs includes one or more of the autosomal lung-specific mHAg SNPs (listed in FIG. 10); determining for the donor the presence or absence of each of the SNPs of the panel of lung-specific mHAg SNPs; obtaining SNP information for the recipient for the panel of lung-specific mHAg SNPs; and determining for the recipient the presence or absence of each of the SNPs of the panel of lung-specific mHAg SNPs; wherein the donor and recipient are HLA-matched; wherein the presence in the recipient and the absence in the donor of more than about 424 of the SNPs of the panel of lung-specific mHAg SNPs is indicative of an increased risk of the recipient developing chronic lung GvHD after receipt of donor tissue from the donor; and wherein the presence in the recipient and the absence in the donor of less than about 424 of the SNPs of the panel of lung-specific mHAg SNPs is indicative of a decreased risk of the recipient developing chronic lung GvHD after receipt of donor tissue from the donor. In some aspects, the method further includes administering an immunosuppressive agent to the transplant recipient.

[0016] Also provided herein are methods for selecting a transplant donor so that the risk of a recipient developing chronic lung graft versus host disease (GvHD) is reduced, the method including: obtaining single nucleotide polymorphism (SNP) information for the donor for a panel of lung-specific minor histocompatibility antigen (mHAg) SNPs; wherein the panel of lungspecific mHAg SNPs includes one or more of the autosomal lung-specific SNPs (listed in FIG. 10); determining for the donor the presence or absence of each of the SNPs of the panel of lungspecific mHAg SNPs; obtaining SNP information for the recipient for the panel of lung-specific mHAg SNPs; and determining for the recipient the presence or absence of each of the SNPs of the panel of lung-specific mHAg SNPs; wherein the donor and recipient are HLA-matched; wherein the presence in the recipient and the absence in the donor of more than about 424 of the SNPs of the panel of lung-specific mHAg SNPs is indicative of an increased risk of the recipient developing chronic lung GvHD if the donor is selected as the transplant donor; and wherein the presence in the recipient and the absence in the donor of less than about 424 of the SNPs of the panel of lung-specific mHAg SNPs is indicative of a decreased risk of the recipient developing chronic lung GvHD if the donor is selected as the transplant donor.

[0017] Also provided herein are methods for selecting a transplant donor, the method including: determining for the potential donor the presence or absence of one or more of the single nucleotide polymorphisms (SNPs) of a panel of lung-specific minor histocompatibility antigen (mHAg) SNPs; and determining for the transplant recipient the presence or absence of one or more of the SNPs of the panel of lung-specific mHAg SNPs; wherein the panel of lung-specific mHAg SNPs includes one or more of the autosomal lung-specific mHAg SNPs (listed in FIG. 10); and wherein the donor and recipient are HLA-matched; and selecting as the transplant donor the potential donor wherein less than about 424 of the SNPs of the panel of lung-specific mHAg SNPs are present in the recipient and absent in the potential donor. In some aspects, for the methods provided herein, an increased risk of chronic lung GvHD includes a 5-year cumulative incidence of chronic lung GvHD of about 12% or greater.

[0018] In some aspects, for the methods provided herein, a decreased risk of chronic lung GvHD includes about a 5-year cumulative incidence of chronic lung GvHD of about 0.93% or less.

[0019] Also provided herein are kits for identifying a transplant recipient at increased risk of developing chronic lung graft versus host disease (GvHD) after receipt of donor tissue from a donor and / or for selecting a transplant donor so that the risk of the recipient developing chronic lung GvHD is reduced, the kit including a plurality of probes capable of binding to and / or identifying one or more of the lung-specific mHAg SNPs (listed in FIG. 10) in a sample from the transplant donor and / or the transplant recipient. In some aspects, the kit includes a plurality of probes capable of binding to and / or identifying each of the lung-specific mHAg SNPs (listed in FIG. 10) in a sample from the transplant donor and / or the transplant recipient.

[0020] Also provided herein are kits for identifying a transplant recipient at increased risk of developing chronic lung graft versus host disease (GvHD) after receipt of donor tissue from a donor and / or for selecting a transplant donor so that the risk of the recipient developing chronic lung GvHD is reduced, the kit including a plurality of primers for selectively amplifying one or more of the lung-specific mHAg SNPs (listed in FIG. 10) in a sample from the transplant donor and / or the transplant recipient. In some aspects, the kit includes a plurality of primers for selectively amplifying each of the lung-specific mHAg SNPs (listed in FIG. 10) in a sample from the transplant donor and / or the transplant recipient.

[0021] In some aspects, for a method as provided herein, obtaining the SNP information for the donor and / or the recipient includes targeted next generation sequencing (NGS), whole exome sequencing (WES), whole genomic sequencing (WGS), and / or SNP array-based genotyping.

[0022] Also provided herein are methods of identifying a transplant recipient at increased risk of developing acute graft versus host disease (GvHD) of the gastrointestinal (GI) system after receipt of donor tissue from a donor, the method including: identifying peptide polymorphisms of minor histocompatibility antigen (mHAgs) between the donor and the recipient, wherein the mHAgs are Gl-specific mHAgs (expressed on Gl-resident non-hematopoietic cells), comparing the amino acid sequences of the Gl-specific mHAgs with the amino acid sequences of antigenic epitopes of one or more viruses relevant for allo-HCT recipients, identifying and quantifying the number of amino acid sequences of the Gl-specific mHAgs that are homologous (cross-reactive) to an amino acid sequence of the antigenic epitopes of the one or more viruses relevant for allo- HCT recipients; wherein a number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients greater than the median is indicative of an increased risk of the transplant recipient developing acute graft versus host disease (GvHD) of the gastrointestinal (GI) system after receipt of donor tissue. In some aspects, the median number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients is about 37 (range of about 13 to about 87).

[0023] Also provided herein are methods for selecting a transplant donor so that the risk of a recipient developing acute graft versus host disease (GvHD) of the gastrointestinal (GI) system is reduced, the method including: identifying peptide polymorphisms of minor histocompatibility antigen (mHAgs) between the donor and the recipient, wherein the mHAgs are Gl-specific mHAgs (expressed on Gl-resident non-hematopoietic cells), comparing the amino acid sequences of the Gl-specific mHAgs with the amino acid sequences of antigenic epitopes of one or more viruses relevant for allo-HCT recipients, identifying and quantifying the number of amino acid sequences of the Gl-specific mHAgs that are homologous (cross-reactive) to an amino acid sequence of the antigenic epitopes of the one or more viruses relevant for allo-HCT recipients; wherein a number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients greater than the median is indicative of a decreased risk of the transplant recipient developing acute graft versus host disease (GvHD) of the gastrointestinal (GI) system after receipt of donor tissue. In some aspects, the median number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients comprises about 37 (range of about 13 to about 87).

[0024] Also provided herein are methods of identifying a transplant recipient suitable for posttransplant administration of immunosuppressive agent to minimize acute graft versus host disease (GvHD) of the gastrointestinal (GI) system along with receipt of donor tissue from a donor, the method including: identifying peptide polymorphisms of minor histocompatibility antigen (mHAgs) between the donor and the recipient, wherein the mHAgs are Gl-specific mHAgs (expressed on Gl-resident non-hematopoietic cells), comparing the amino acid sequences of the Gl-specific mHAgs with the amino acid sequences of antigenic epitopes of one or more viruses relevant for allo-HCT recipients, identifying and quantifying the number of amino acid sequences of the Gl-specific mHAgs that are homologous (cross-reactive) to an amino acid sequence of the antigenic epitopes of the one or more viruses relevant for allo-HCT recipients; wherein a number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients greater than the median is indicative of a transplant recipient suitable for post-transplant administration of a prophylactic immunosuppressive agent to minimize acute graft versus host disease (GvHD) of the gastrointestinal (GI) system along with receipt of donor tissue from a donor. In some aspects, the median number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients is about 37 (range of about 13 to about 87). In some aspects, the prophylactic immunosuppressive agent targets GvHD of the gut. In some aspects, the prophylactic immunosuppressive agent comprises vedolizumab.

[0025] In some aspects, for the methods provided herein the viral antigen of a virus relevant for allo-HCT recipients is a gut-tropic virus.

[0026] In some aspects, for the methods provided herein the viral antigen of a virus relevant for allo-HCT recipients includes an antigen of an adenovirus (ADV), a cytomegalovirus (CMV) and / or an Epstein-Barr virus (EBV).

[0027] In some aspects, for the methods provided herein the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients comprises about 8 to about 11 amino acids. In some aspects, the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients comprises about 8, about 9, about 10, or about 11 amino acids. In some aspects, the amino acid sequence of the antigenic epitope of the virus relevant for allo- HCT recipients consists of 8 to 11 amino acids. In some aspects, 7, the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients consists of a peptide of 8, 9, 10, or 11 amino acids. In some aspects, the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients comprises aa HLA class I epitope.

[0028] In some aspects, the method further includes administering a prophylactic immunosuppressive agent to the transplant recipient. In some aspects, the prophylactic immunosuppressive agent targets GvHD of the gut. In some aspects, the prophylactic immunosuppressive agent is vedolizumab.

[0029] In some aspects of a method as provided herein, the donor and the recipient are human.

[0030] In some aspects of a method as provided herein, the transplantation of donor tissue is for the treatment of cancer, sickle cell disease, an inherited anemia, or bone marrow failure. In some aspects, the transplantation of donor tissue is for the treatment of a malignant or a nonmalignant hematologic disease with indication to allogeneic hematopoietic cell transplantation. In some aspects, the cancer includes a hematologic cancer. In some aspects, the hematologic cancer includes a leukemia. In some aspects, the cancer includes acute myeloid leukemia (AML), myelodysplastic syndrome (MDS), acute lymphoblastic leukemia (ALL), T-cell acute lymphoblastic leukemia (T-ALL), peripheral T cell lymphoma (PTCL), a myeloproliferative neoplasm (MPN), and / or non-Hodgkin's lymphoma (NHL).

[0031] In some aspects of a method as provided herein, the donor tissue includes hematopoietic cells. In some aspects, the donor tissue includes HLA-matched allogeneic hematopoietic stem cells.

[0032] Also provided herein are computer programs including computer program code configured to cause one or more physical computing devices to perform a method of analysis of sample results for a transplant donor and a transplant recipient when the code is run on the one or more physical computing devices, the method of analysis as provided herein.

[0033] As used herein, “isolated” refers to material removed from its original environment (e.g., the natural environment if it is naturally occurring), and thus is altered “by the hand of man” from its natural state.

[0034] The term “and / or” means one or all of the listed elements or a combination of any two or more of the listed elements.

[0035] The words “preferred” and “preferably” refer to embodiments that may afford certain benefits, under certain circumstances. However, other embodiments may also be preferred, under the same or other circumstances. Furthermore, the recitation of one or more preferred embodiments does not imply that other embodiments are not useful and is not intended to exclude other embodiments.

[0036] The terms “comprises” and variations thereof do not have a limiting meaning where these terms appear in the description and claims.

[0037] Unless otherwise specified, “a,” “an,” “the,” and “at least one” are used interchangeably and mean one or more than one.

[0038] Also herein, the recitations of numerical ranges by endpoints include all numbers subsumed within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, 5, etc.). For any method disclosed herein that includes discrete steps, the steps may be conducted in any feasible order. And, as appropriate, any combination of two or more steps may be conducted simultaneously.

[0039] Unless otherwise indicated, all numbers expressing quantities of components, molecular weights, and so forth used in the specification and claims are to be understood as being modified in all instances by the term "about." Accordingly, unless otherwise indicated to the contrary, the numerical parameters set forth in the specification and claims are approximations that may vary depending upon the desired properties. At the very least, and not as an attempt to limit the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.

[0040] Notwithstanding that the numerical ranges and parameters are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. All numerical values, however, inherently contain a range necessarily resulting from the standard deviation found in their respective testing measurements.

[0041] In several places throughout the application, guidance is provided through lists of examples, which examples can be used in various combinations. In each instance, the recited list serves only as a representative group and should not be interpreted as an exclusive list. It is to be understood that the particular examples, materials, amounts, and procedures are to be interpreted broadly.

[0042] All headings throughout are for the convenience of the reader and should not be used to limit the meaning of the text that follows the heading, unless so specified.

[0043] It is to be understood that the illustrative examples, materials, amounts, and procedures described are to be interpreted broadly.

[0044] BRIEF DESCRIPTION OF THE FIGURES

[0045] FIGS. 1A and IB. Building a pipeline for systematic mHAg discovery. FIG. lA is an overview of the analysis workflow for prediction of minor histocompatibility antigens (mHAgs). Starting from whole exome sequencing (WES) obtained from germline DNA of donor and recipient pairs, single-nucleotide polymorphisms (SNPs) altering the protein sequence and present only in the patient are identified. Through the application of tissue-specificity filters (see FIG. IB), the polymorphisms present in genes expressed in the tissues of interest are selected and all possible k-mers encompassing the SNPs are computed and then subjected to HLA class I binding prediction. Resulting candidate mHAgs can be used for downstream applications. FIG. IB presents the tissue-specificity filters of the pipeline. Top left, GvHD filter to comprehensively catalog the genes expressed in GvHD target organs. Single-cell RNA-Seq datasets were used to define an expression atlas of tissue-resident cell types from skin, liver, colon (GI), lung, oral mucosa and lacrimal gland (in addition to hematopoietic cell lineages); colored vertical bars on the upset plot - enumeration of genes with exclusive organ-specific expression; bottom left, GvL filter to identify genes with preferential expression in acute myeloid leukemia (AML): a single-cell based molecular classifier (van Galen et al., 2019, Cell,' 176:1265- 1281 el224) was applied to the Beat AML dataset (Tyner et al., 2018, Nature, 562:526-531) to define distinct AML expression clusters (ECs). AML genes with evidence of expression at the RNA or protein level in the GTEx repository of human adult healthy tissues were excluded, to define a list of 259 candidate genes with preferential expression in AML (heatmap). Right, Y mHAg filter, designed to identify mHAgs arising from genes in the male-specific region of the Y chromosome (MSY) which serve as additional targets of allorecognition in female-to-male transplants. The flowchart summarizes the principal steps of Y-chromosome pipeline.

[0046] FIGS. 2A-2E. mHAgs shape GvHD outcomes. FIG. 2A presents cumulative incidence (CI) of grade ILIV acute GvHD stratifying patients based on the overall autosomal mHAg load below (light red) or above (dark red) the median value (median autosomal mHAg load = 632). 1-year CI of acute GvHD was 6.3% (95% confidence interval: 2.8% - 12%) and 16% (95% confidence interval: 9.5% - 23%), respectively; p = 0.029 (Gray’s test). FIG. 2B shows Hazard Ratios (HR) from multivariate Cox proportional hazards regression modeling of variables influencing post-transplant outcomes. Statistically significant p values are in bold. FIG. 2C is a distribution of 14 patients experiencing lung chronic GvHD (NTH grade 2-3) across deciles of lung mHAgs (i.e., load of predicted mHAgs in genes expressed in the lung). FIG. 2D is the cumulative incidence (CI) of lung chronic GvHD stratifying patients based on the lung mHAg load below (grey) or above (light blue) the median value (median lung-specific mHAg load: 424). 5-year CI of lung chronic GvHD are 0.93% (95% confidence interval: 0.08% - 4.7%) versus 12% (95% confidence interval: 6.7% - 20%), p < 0.001 (Gray’s test). FIG. 2E shows odds ratios (OR) from logistic regression modeling of variables influencing post-transplant outcomes. For each variable, the reference category is indicated as ‘ref’ Statistically significant p values are in bold. AML: acute myeloid leukemia; MDS: myelodysplastic syndrome; CR: complete remission; MAC: myeloablative conditioning; RIC: reduced-intensity conditioning; PBSC: peripheral blood stem cell; BM: bone marrow; F to M: female-to-male transplant.

[0047] FIGS. 3A-3C. mHAgs in acute liver GvHD. FIG. 3A is a distribution of 13 patients experiencing liver acute GvHD (grade II-IV) across deciles of liver mHAgs (i.e., load of predicted mHAgs in genes expressed in the liver). FIG. 3B is a heatmap showing the expression profile of the 7 genes with recurrent SNPs in patients experiencing acute liver GvHD in singlecell RNA-Seq clusters from liver, colon (GI) and skin. FIG. 3C shows the number of cooccurring driver liver mHAgs in patients with acute GvHD with (green, n = 13) or without (grey, n = 21) liver involvement. Whiskers indicate min and max, with all individual values shown; p = 0.002 with Mann Whitney test.

[0048] FIG. 4. Characteristics of the DFCI-MRD Cohort (n = 20).

[0049] FIG. 5. A compilation of all HLAs found as binders of the peptides arising from the seven SNPs associated with acute liver GvHD.

[0050] FIG. 6. A schematic of viral cross-reactivity.

[0051] FIG. 7. A schematic for systematic evaluation of cross-reactivity.

[0052] FIGS. 8A and 8B. The load of cross-reactive mHAgs correlates with the incidence of severe GI acute GvHD. FIG. 8A shows ADV / CMV / EBV crossreactive mHAgs versus GI aGvHD grade III-IV. FIG. 8B shows grade III-IV GI aGvHD CI (%) versus time after transplant.

[0053] FIG. 9. Multivariable analysis confirms cross-reactive mHAg load as an independent risk factor for GI acute GvHD.

[0054] FIGS. 10-1 to 10-71. A compilation of SNPs associated with lung GvHD. For the genes with expression in the lung that in the DFCI-MRD cohort (SNPs discordant between donor and recipient (donor negative, recipient positive)) able to bind one of the HLAs expressed by the donor and the recipient, 6212 genes are informative (column A), with multiple SNPs (19,202 total) (column B). DETAILED DESCRIPTION

[0055] Minor histocompatibility antigens (mHAgs) composed of immunogenic peptides presented by HLA molecules can cause immune responses involved in graft-versus-host disease (GvHD) and graft-versus-leukemia (GvL) effects after allogeneic hematopoietic cell transplantation (alloHCT). The pathogenesis of GvHD, the most detrimental immune-related complication after allogeneic hematopoietic cell transplantation, is attributed to a donor-derived immune response directed against mHAgs either broadly expressed across tissues or expressed specifically in GvHD-affected tissues. With allogeneic hematopoietic cell transplantation between HLA-matched donors-recipient pairs, mHAgs present in recipient tissues are sensed as foreign by donor T cells and are expected to be highly immunogenic due to the lack of central tolerance against them. The role of minor histocompatibility antigens in mediating GvHD and GvL following allogeneic hematopoietic cell transplantation is recognized but not well- characterized. Despite the crucial role of mHAgs in GvL and GvHD reactions, it has not been possible thus far to consistently link patient-specific mHAg repertoires to clinical outcomes.

[0056] Provided herein is an analytic framework to systematically identify autosomal and Y- encoded mHAgs associated with GvHD clinical outcomes. This analytic framework, shown schematically in FIGS. 1A and IB and described in more detail in the examples section included herewith, is based on the integration of polymorphism detection by whole exome sequencing of germline DNA from donor-recipient (D-R) pairs together with organ-specific transcriptional- and proteome-level expression. Application of this pipeline to a cohort of 220 HLA-matched allo- HSCT D-R pairs uncovered novel associations with GvHD outcomes, including, but not limited to acute liver GvHD and chronic lung GvHD. Specifically, the analytic framework described herein identifies single nucleotide polymorphisms (SNPs) (both autosomal and Y chromosome- encoded) that produce amino acid coding differences between recipients and donors and are associated with GvHD outcomes.

[0057] A single-nucleotide polymorphism (SNP) is a germline substitution of a single nucleotide at a specific position in the genome, with a frequency of more than 1% in the population. For example, a G nucleotide present at a specific location in a reference genome may be replaced by an A in a minority of individuals. The two possible nucleotide variations of this SNP G or A are called alleles. Around 90% of genome variations are limited to SNPs. SNPs have proven to be of great value for medical diagnostics. Several SNP databases are maintained, including, but not limited to, dbSNP a SNP database maintained by the National Center for Biotechnology Information (NCBI) (available on the worldwide web at ncbi.nlm.nih.gov / snp / ). As of June 8, 2015, dbSNP listed 149,735,377 SNPs for the human genome.

[0058] Methods are provided utilizing information on SNP differences between allo-HCT recipients and donors associated to identify transplant recipients at an increased risk or a decreased risk of developing GvHD after receipt of donor tissue from a donor. GVHD is a potentially serious complication of allogeneic stem cell transplantation, occurring when the donor stem cells (“the graft”) attack healthy cells in the patient (“the host”).

[0059] GvHD can be acute or chronic, with acute GvHD (aGvHD) occurring early in the posttransplant period (within about 100 days after the transplant) and chronic GvHD (cGvHD) occurring after the 100-day mark post-transplant. The appearance of moderate to severe cases of aGvHD and / or cGVHD adversely influences long-term survival. The methods can utilize information on SNP differences between recipients and donors associated to identify transplant recipients at an increased risk or a decreased risk of developing acute GvHD after receipt of donor tissue from a donor. The methods can also utilize information on SNP differences between recipients and donors associated to identify transplant recipients at an increased risk or a decreased risk of developing chronic GvHD after receipt of donor tissue from a donor.

[0060] GVHD can affect many parts of the body, including the eyes, skin, and mouth, as well as the gastrointestinal (GI) tract, liver, and lungs. The methods described herein, utilizing information on SNP differences associated with GvHD outcomes, may be used to identify transplant recipients at an increased risk or a decreased risk of developing an organ specific GvHD after receipt of donor tissue from a donor. Such an organ specific GvHD includes, but is not limited to, acute liver GvHD and chronic lung GvHD.

[0061] Acute liver GvHD

[0062] Methods are provided for identifying a transplant recipient at increased risk of developing acute liver GvHD after receipt of donor tissue from a donor and / or for selecting a transplant donor so that the risk of the recipient developing acute liver GvHD is reduced. Such methods include determining for the donor the presence or absence of one or more single nucleotide polymorphisms (SNPs) from a panel of SNPs and determining for the recipient the presence or absence of one or more SNPs from the panel of SNPs, wherein the panel of SNPs includes the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE2O p.G1207S SNP, and / or the VSTM4 p.F68S SNP, and wherein the presence in the recipient and the absence in the donor of one or more of the SNPs from the panel of SNPs is indicative of an increased risk of the recipient developing acute liver GvHD after receipt of donor tissue from the donor.

[0063] The determination of the presence or absence of one or more single nucleotide polymorphisms (SNPs) from the panel of SNPs may include a determination of the presence or absence of any one, any two, any three, any four, any five, any six, or all seven of the ATXN2 P.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP.

[0064] In some aspects of the methods, the donor and the recipient may be HLA-matched, wherein HLA stands for human leukocyte antigens. In some aspects of the methods, when the one or more SNPs from the panel of SNPs includes the ATXN2 p.L107V SNP, the HLA type of HLA-matched donor and recipient includes the HLA type HLA B0801, HLA B0702, HLA B5501, HLA C0602, HLA B2703, and / or HLA A0301. In some aspects of the methods, when the one or more SNPs from the panel of SNPs includes the C4orf48 p,17L SNP, the HLA type of the HLA-matched donor and recipient includes HLA A2402, HLA C0303, HLA A0201, HLA C0602, HLA C0701, HLA C0401, HLA Cl 203, and / or HLA B0801. In some aspects of the methods, when the one or more SNPs from the panel of SNPs includes the NCOA7 p.S284A SNP, the HLA type of the HLA-matched donor and recipient includes HLA A0301, HLA C0303, HLA C0304, HLA B4402, and / or HLA C0802. In some aspects of the methods, when the one or more SNPs from the panel of SNPs comprises the TGFB1 p.PlOL SNP, the HLA type of the HLA-matched donor and recipient includes HLA A0201, HLA B0702, and / or HLA C0701. In some aspects of the methods, when the one or more SNPs from the panel of SNPs comprises the TNC p.Q680R SNP, the HLA type of the HLA-matched donor and recipient includes HLA A0101, HLA Bl 801, HLA Bl 501, HLA C0602, HLA A6801, HLA A0201, and / or HLA B4402. In some aspects of the methods, when the one or more SNPs from the panel of SNPs comprises the UBE2O p.G1207S SNP, the HLA type of the HLA-matched donor and recipient includes HLA C0501, HLA B4101, HLA B4001, and / or HLA B3701. In some aspects of the methods, when the one or more SNPs from the panel of SNPs comprises the VSTM4 p.F68S SNP, the HLA type of the HLA-matched donor and recipient includes HLA C0501, HLA C0701, HLA C0303, HLA C0304, HLA C0501, HLA C0401, HLA C0702, and / or HLA C0802.

[0065] Kits for use in the identification of risk for developing GvHD and / or for use in donor selection are also provided. Kits are also provided for use in identifying a transplant recipient at increased risk of developing acute liver GvHD after receipt of donor tissue from a donor and / or for selecting a transplant donor so that the risk of the recipient developing acute liver GvHD is reduced.

[0066] A kit is any manufacture (for example, a package or container) including at least one reagent for specifically detecting one or more of the SNPs described herein. The kit may be promoted, distributed, or sold as a unit for performing the methods of the present disclosure.

[0067] The kit may include a plurality of probes capable of binding to and / or identifying one or more of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP in a sample from the transplant donor and / or the transplant recipient. In some aspects, the kit may include probes capable of binding to and / or identifying any one, any two, any three, any four, any five, any six, or all seven of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP.

[0068] The kit may include a plurality of primers for selectively amplifying one or more of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP in a sample from the transplant donor and / or the transplant recipient. In some aspects, the kit may include a plurality of primers for selectively amplifying any one, any two, any three, any four, any five, any six, or all seven of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP.

[0069] Chronic lung GvHD

[0070] As described in the examples included herewith the analytic framework to systematically identify mHAgs associated with GvHD clinical outcomes identified novel associations with chronic lung GvHD. A compilation of SNPs associated with chronic lung GvHD for the 9260 genes with expression in the lung that in the DFCI-MRD cohort (SNPs discordant between donor and recipient (donor negative, recipient positive)) able to bind one of the HLAs expressed by the donor and the recipient is provided in FIG. 10. Only 6,212 genes were informative, but many had multiple SNPs, for a total of 19,202 SNPs.

[0071] Methods are provided for determining a transplant recipient’s risk of developing chronic lung GvHD after receipt of donor tissue from a donor. The method includes: determining for the donor the presence or absence of each of the SNPs of a panel of lung-specific mHAg SNPs; and determining for the recipient the presence or absence of each of the SNPs of the panel of lungspecific mHAg SNPs; wherein the presence in the recipient and the absence in the donor of more than about a median number of the SNPs of the panel of lung-specific mHAg SNPs is indicative of an increased risk of the recipient developing chronic lung GvHD after receipt of donor tissue from the donor; and wherein the presence in the recipient and the absence in the donor of less than about a median number of the SNPs of the panel of lung-specific mHAg SNPs is indicative of a decreased risk of the recipient developing chronic lung GvHD after receipt of donor tissue from the donor.

[0072] Methods are provided for selecting a transplant donor so that the risk of a recipient developing chronic lung GvHD is reduced. The method includes: determining for the donor the presence or absence of each of a SNPs of the panel of lung-specific mHAg SNPs; and determining for the recipient the presence or absence of each of the SNPs of the panel of lungspecific mHAg SNPs; wherein the presence in the recipient and the absence in the donor of more than about the median number of the SNPs of the panel of lung-specific mHAg SNPs is indicative of an increased risk of the recipient developing chronic lung GvHD after receipt of donor tissue from the donor; and wherein the presence in the recipient and the absence in the donor of less than about the median number of the SNPs of the panel of lung-specific mHAg SNPs is indicative of a decreased risk of the recipient developing chronic lung GvHD after receipt of donor tissue from the donor.

[0073] Methods for selecting a transplant donor are also provided. The method includes: determining for the donor the presence or absence of each of the SNPs of a panel of lung-specific mHAg SNPs; and determining for the recipient the presence or absence of each of the SNPs of the panel of lung-specific mHAg SNPs; wherein the presence in the recipient and the absence in the donor of more than about the median number of the SNPs of the panel of lung-specific mHAg SNPs is indicative of an increased risk of the recipient developing chronic lung GvHD after receipt of donor tissue from the donor; and wherein the presence in the recipient and the absence in the donor of less than about the median number of the SNPs of the panel of lungspecific mHAg SNPs is indicative of a decreased risk of the recipient developing chronic lung GvHD after receipt of donor tissue from the donor.

[0074] With the methods for a determination of clinical outcomes for chronic lung GvHD described herein, the panel of lung-specific mHAg SNPs includes one or more of the autosomal lung-specific mHAg SNPs (listed in FIG. 10). In some aspects, the panel of lung-specific mHAg SNPs may include about any 10, about any 50, about any 100, about any 500, about any 1,000, about any 2,000, about any 3,000, about any 4,000, about any 5,000, about any 6,000, about any 7,000, about any 8,000, about any 9,000, about any 10,000, about any 11,000, about any 12,000, about any 13,000, about any 14,000, about any 15,000, about any 16,000, about any 17,000, about any 18,000, or about any 19,000 of the lung-specific mHAg SNPs (listed in FIG. 10), or any range thereof. In some aspects, the panel of lung-specific mHAg SNPs may include all of the 19,202 SNPs.

[0075] With the methods for a determination of clinical outcomes for chronic lung GvHD described herein, a median number of the SNPs of the panel of lung-specific mHAg SNPs may statistically determined. In some aspects, a median number of the SNPs of the panel of lungspecific mHAg SNPs may be about 424 SNPs from the panel of 19,202 lung-specific mHAg SNPs listed in FIG. 10.

[0076] The methods provided herein may be used in association with transplantation of donor tissue for the treatment of a malignant or a nonmalignant hematologic disease with indication to allogeneic hematopoietic cell transplantation. Such disorders may include, but are not limited to, cancer, sickle cell disease, an inherited anemia, or bone marrow failure. In some aspects, the cancer may be a hematologic cancer. In some aspects, the hematologic cancer may be a leukemia. In some aspects, the cancer may be acute myeloid leukemia (AML), myelodysplastic syndrome (MDS), acute lymphoblastic leukemia (ALL), T-cell acute lymphoblastic leukemia (T- ALL), peripheral T cell lymphoma (PTCL), a myeloproliferative neoplasm (MPN), and / or nonHodgkin's lymphoma (NHL).

[0077] For the methods provided herein, any of a variety of donor tissues may be transplanted, including but not limited to, allogeneic hematopoietic stem cells. Such stem cells may be obtained from the bone marrow or the bloodstream. In some aspects, the donor tissue includes HLA-matched allogeneic hematopoietic stem cells.

[0078] For the methods provided herein, SNP information for a panel of SNPs may be obtained for a donor and / or a recipient by any of a variety of methods. As SNPs are allelic germline variations, SNPs information for a donor and / or a recipient may be determined from a genomic DNA sample. A panel of SNPs includes, but is not limited to, a panel of SNPs including one or more of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP; or a panel of lung-specific mHAg SNPs including autosomal lung specific mHAg SNPs listed in FIG. 10.

[0079] SNP information may be obtained, for example, by utilizing any of a variety of DNA sequencing approaches, including for example, targeted next generation sequencing (NGS), whole exome sequencing (WES), and whole genomic sequencing (WGS). The term “Next Generation Sequencing (NGS)” herein refers to sequencing methods that allow for massively parallel sequencing of clonally amplified molecules and of single nucleic acid molecules. Nonlimiting examples of NGS include sequencing-by-synthesis using reversible dye terminators, and sequencing-by-ligation. In some applications, SNP information may be obtained utilizing microarray-based SNP genotyping assays. In some applications, SNP information may be obtained by capillary electrophoresis, mass spectrometry, single-strand conformation polymorphism (SSCP), single base extension, electrochemical analysis, denaturating HPLC and gel electrophoresis, restriction fragment length polymorphism (RFLP), chip detection, various novel PCR or qPCR based technologies, such as amplification refractory mutation system PCR (ARMS-PCR), and kompetitive allele-specific PCR (KASP), or hybridization analysis.

[0080] SNP information for a panel of SNPs for the donor and / or SNP information for a panel of SNPs for the recipient may have been previously determined and may be provided as a computer-readable medium having stored thereon computer-readable SNP information for the panel of SNPs.

[0081] A donor may be a human donor as well as a non-human mammalian subject. A recipient may be a human recipient as well as a non-human mammalian subject. Although the examples herein concern humans and the language is primarily directed to human concerns, the concept of this disclosure is applicable to any mammal, and is useful in the fields of veterinary medicine, animal sciences, research laboratories and such.

[0082] Following a determination of an increased risk of a recipient developing GVHD, including, but not limited to acute liver GvHD and / or chronic lung GvHD, after receipt of donor tissue from the donor, the receipt may be treatment with one or more immunosuppressive drugs to suppress donor T-cell function. Such a drug may be administered before and / or after transplantation of the donor tissue. Such drugs include, but are not limited to a topical steroid ointment (for GvHD of the skin), steroid eye drops (for GvHD of the eye), a corticosteroid such as prednisone or methylprednisolone, cyclophosphamide, methotrexate, cyclosporine, tacrolimus, nycophenolate mofetil, sirolimus, and / or a biologies, such as for example, Abatacept (ORENCIA®), antithymocyte globulin (ATG), Alemtuzumab (CAMPATH®), Tocilizumab (ACTEMRA®), Ruxolitinib (JAKAFI®), Ibrutinib (IMBRUVICA®), and Belumosudil (REZUROCK®).

[0083] Kits for also provided use in identifying a transplant recipient at increased risk of developing chronic lung GvHD after receipt of donor tissue from a donor and / or for selecting a transplant donor so that the risk of the recipient developing chronic lung GvHD is reduced. A kit is any manufacture (for example, a package or container) including at least one reagent for specifically detecting one or more of the SNPs described herein. The kit may be promoted, distributed, or sold as a unit for performing the methods of the present disclosure.

[0084] The kit may include a plurality of probes capable of binding to and / or identifying one or more of the lung-specific mHAg SNPs listed in FIG. 10 in a sample from the transplant donor and / or the transplant recipient. In some aspects, the kit may include probes capable of binding to and / or identifying about any 10, about any 50, about any 100, about any 500, about any 1,000, about any 2,000, about any 3,000, about any 4,000, about any 5,000, about any 6,000, about any 7,000, about any 8,000, about any 9,000, about any 10,000, about any 11,000, about any 12,000, about any 13,000, about any 14,000, about any 15,000, about any 16,000, about any 17,000, about any 18,000, about any 19,000, any range thereof, or all of the lung-specific mHAg SNPs listed in FIG. 10.

[0085] The kit may include a plurality of primers for selectively amplifying one or more of the lung-specific mHAg SNPs listed in Fig. 10 in a sample from the transplant donor and / or the transplant recipient. In some aspects, the kit may include a plurality of primers for selectively amplifying about any 10, about any 50, about any 100, about any 500, about any 1,000, about any 2,000, about any 3,000, about any 4,000, about any 5,000, about any 6,000, about any 7,000, about any 8,000, about any 9,000, about any 10,000, about any 11,000, about any 12,000, about any 13,000, about any 14,000, about any 15,000, about any 16,000, about any 17,000, about any 18,000, about any 19,000, any range thereof, or all of the lung-specific mHAg SNPs listed in FIG. 10.

[0086] Acute GvHD of the Gut

[0087] The gastrointestinal system (GI) is often affected by acute Graft-versus-Host-Disease (aGvHD) and severe manifestations of aGvHD of the gut portend a poor prognosis after allogeneic hematopoietic cell transplantation (allo-HCT). Methods are provided to predict and mitigate such GI aGvHD risk. Building on the clinical observation that viral reactivations often coincide with the development of GvHD, it was determined that sequence homology of alloreactive peptides with gut-tropic viral epitopes contribute to the pathophysiology of GI aGvHD. A shown in Example 2, cross-reactivity against gut-tropic viruses is a key contributor to GI aGvHD pathophysiology and molecular mimicry between viral epitopes and mHAgs can be systematically assessed and used in GvHD risk prognostication.

[0088] Methods are provided of identifying a transplant recipient at increased risk of developing acute GvHD of the GI system after receipt of donor tissue from a donor. The method includes identifying peptide polymorphisms of minor histocompatibility antigen (mHAgs) between the donor and the recipient, wherein the mHAgs are Gl-specific mHAgs (expressed on Gl-resident non-hematopoietic cells), comparing the amino acid sequences of the Gl-specific mHAgs with the amino acid sequences of antigenic epitopes of one or more viruses relevant for allo-HCT recipients, identifying and quantifying the number of amino acid sequences of the Gl-specific mHAgs that are homologous (cross-reactive) to an amino acid sequence of the antigenic epitopes of the one or more viruses relevant for allo-HCT recipients; wherein a number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients greater than the median is indicative of an increased risk of the transplant recipient developing acute graft versus host disease (GvHD) of the GI system after receipt of donor tissue

[0089] Provided also are methods for selecting a transplant donor so that the risk of the recipient developing acute graft versus host disease of the GI system is reduced. The method includes identifying peptide polymorphisms of minor histocompatibility antigen (mHAgs) between the donor and the recipient, wherein the mHAgs are Gl-specific mHAgs (expressed on Gl-resident non-hematopoietic cells), comparing the amino acid sequences of the Gl-specific mHAgs with the amino acid sequences of antigenic epitopes of one or more viruses relevant for allo-HCT recipients, identifying and quantifying the number of amino acid sequences of the Gl-specific mHAgs that are homologous (cross-reactive) to an amino acid sequence of the antigenic epitopes of the one or more viruses relevant for allo-HCT recipients; wherein a number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients greater than the median is indicative of a decreased risk of the transplant recipient developing aGvHD of the GI system after receipt of donor tissue.

[0090] Also provided are methods of identifying a transplant recipient suitable for posttransplant administration of immunosuppressive agent to minimize acute GvHD of the GI system along with receipt of donor tissue from a donor, the method including: identifying peptide polymorphisms of minor histocompatibility antigen (mHAgs) between the donor and the recipient, wherein the mHAgs are Gl-specific mHAgs (expressed on Gl-resident non- hematopoietic cells), comparing the amino acid sequences of the Gl-specific mHAgs with the amino acid sequences of antigenic epitopes of one or more viruses relevant for allo-HCT recipients, identifying and quantifying the number of amino acid sequences of the Gl-specific mHAgs that are homologous (cross-reactive) to an amino acid sequence of the antigenic epitopes of the one or more viruses relevant for allo-HCT recipients; wherein a number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients greater than the median is indicative of a transplant recipient suitable for posttransplant administration of a prophylactic immunosuppressive agent to minimize acute GvHD of the GI system along with receipt of donor tissue from a donor. In some aspects, the prophylactic immunosuppressive agent to be delivered targets GvHD of the gut. In some aspects, the prophylactic immunosuppressive agent to be delivered comprises vedolizumab.

[0091] With the methods of identifying a transplant recipient at increased risk of developing acute GvHD of the GI system, of selecting a transplant donor so that the risk of a recipient developing acute graft versus host disease of the GI system is reduced, and of identifying a transplant recipient suitable for post-transplant administration of immunosuppressive agent to minimize acute GvHD of the GI system provided herein, the virus relevant for allo-HCT recipients may be a gut-tropic virus. In some aspects, the virus relevant for allo-HCT recipients may be, but is not limited to, adenovirus (ADV), a cytomegalovirus (CMV) and / or an Epstein- Barr virus (EBV).

[0092] With the methods of identifying a transplant recipient at increased risk of developing acute GvHD of the GI system, of selecting a transplant donor so that the risk of a recipient developing acute graft versus host disease of the GI system is reduced, and of identifying a transplant recipient suitable for post-transplant administration of immunosuppressive agent to minimize acute GvHD of the GI system provided herein, the median number of Gl-specific mHAgs homologous to an antigenic epitope of a virus relevant for allo-HCT recipients may range from about 5 to about 500, from about 5 to about 250, from about 10 to about 100, or from about 13 to about 87. In some aspects, the median number of a Gl-specific mHAg homologous to an antigenic epitope of a virus relevant for allo-HCT recipients may be about 5, about 10, about 15, about 20, about 25, about 30, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90 about 95, about 100, or any range thereof. In some aspects, the median number of a Gl-specific mHAg homologous to an antigenic epitope of a virus relevant for allo-HCT recipients may be about 37.

[0093] With the methods of identifying a transplant recipient at increased risk of developing acute GvHD of the GI system, of selecting a transplant donor so that the risk of a recipient developing acute graft versus host disease of the GI system is reduced, and of identifying a transplant recipient suitable for post-transplant administration of immunosuppressive agent to minimize acute GvHD of the GI system provided herein, the amino acid sequence of an antigenic epitope of a virus relevant for allo-HCT recipients comprises about 8 to about 11 amino acids. In some aspects, the amino acid sequence of an antigenic epitope of a virus relevant for allo-HCT recipients comprises about 8, about 9, about 10, or about 11 amino acids. In some aspects, the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients consists of 8 to 11 amino acids. In some aspects, 7, the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients consists of a peptide of 8, 9, 10, or 11 amino acids. In some aspects, the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients comprises aa HLA class I epitope.

[0094] With the methods of identifying a transplant recipient at increased risk of developing acute GvHD of the GI system, of selecting a transplant donor so that the risk of a recipient developing acute graft versus host disease of the GI system is reduced, and of identifying a transplant recipient suitable for post-transplant administration of immunosuppressive agent to minimize acute GvHD of the GI system provided herein, the method may further include administering a prophylactic immunosuppressive agent to the transplant recipient. In some aspects, the prophylactic immunosuppressive agent targets GvHD of the gut. In some aspects, the prophylactic immunosuppressive agent is vedolizumab (Chen et al., Nat Med 2024; 30(8): 2277-2287, “Vedolizumab for the prevention of intestinal acute GVHD after allogeneic hematopoietic stem cell transplantation: a randomized phase 3 trial”).

[0095] EXAMPLES

[0096] Example 1

[0097] Systematic identification of minor histocompatibility antigens informs outcomes after allogeneic stem cell transplantation

[0098] This example describes an analytic framework to systematically identify autosomal and Y-encoded mHAgs, including their detection on HLA class I ligandomes and functionally verifying their immunogenicity, based on the integration of polymorphism detection by whole exome sequencing of germline DNA from D-R pairs together with organ-specific transcriptional- and proteome-level expression. Application of this pipeline to a cohort of 220 HLA-matched allo-HSCT D-R pairs uncovered novel associations with GvHD outcomes for the prevention or treatment of post-transplant disease recurrence.

[0099] Genomic analyses to quantify neoantigens arising from somatic tumor mutations have been tremendously impactful towards advancing cancer immunology and immunotherapy and implementing personalized immune-based treatments in clinical practice. A widely used individualized form of immunotherapy that is potentially curative for many blood disorders is allogeneic hematopoietic stem cell transplantation (allo-HSCT), wherein a suitable stem cell donor is selected for each patient based on HLA matching. In this scenario, the driving principle underlying response is alloreactivity, primarily originating from immune responses against minor histocompatibility antigens (mHAgs), HLA-binding peptides derived from polymorphic protein sequences differing between donor and recipient (D-R) pairs. Like tumor neoantigens, mHAgs are sensed as foreign by donor T cells and are expected to be highly immunogenic due to the lack of central tolerance against them. However, mHAgs are inherited as germline traits encoded by polymorphic genes rather than presenting as somatic events, hence they are not tumor-specific antigens per se. The pathogenesis of graft-versus-host disease (GvHD), the most detrimental immune-related complication after allo-HSCT, can be thus attributed to a donor-derived immune response directed against mHAgs either broadly expressed across tissues or expressed specifically in GvHD-affected tissues. Conversely, the curative graft-versus-leukemia (GvL) effect can be conceptualized as the result of productive donor immune responses against mHAgs expressed on hematopoietic cells, including, but not limited to epitopes with hematopoietic tissue restriction. While more than 50,000 allogeneic transplants are performed annually - with numbers still rising - the beneficial effect of allo-HSCT is too often hampered by disease relapse or complicated by GvHD, which together account for >50% of post-transplant mortality. Hence, identification of molecular determinants to aid in predicting transplant outcomes is urgently needed.

[0100] Currently, only D-R HLA matching and the activity of GvHD prophylaxis strategies are available to help clinicians in this challenge. Post-transplant outcomes are likely impacted by genome-wide mHAg load, delineated in an organ- and malignancy-specific fashion. While >100 individual mHAgs have been identified thus far worldwide, only few have been linked to increased risk of GvHD, and such associations have been only inconsistently validated. More recently, the use of high-throughput sequencing technologies to catalogue the mHAg repertoire has been reported, but only one study linked mHAg burden to 1-year GvHD mortality in a large yet non-contemporary transplant cohort. Certainly, large-scale sequencing capabilities, coupled with the availability of robust HLA class I epitope prediction models, now offer an unprecedented opportunity to delineate the mHAg repertoire of allo-transplanted patients in a personalized fashion which can then be linked to clinical outcomes.

[0101] Results

[0102] Building a pipeline for systematic mHAg discovery

[0103] To identify single nucleotide germline variants (both autosomal and Y chromosome- encoded) present exclusively in the recipient and resulting in nonsynonymous alterations in protein-coding regions, whole-exome sequencing (WES) data from paired D-R DNA was analyzed. An analytic pipeline was devised that further incorporated: (i) fdtering for variants in genes of interest, either expressed in GvHD-targeted tissues (skin, liver, GI, lung, oral mucosa and lacrimal gland); or in GvL genes, i.e. genes preferentially expressed in malignant (and non- malignant) hematopoietic cells, but not in non-hematopoietic tissues, and (ii) predicting variantcontaining peptide 8-1 Imers (k-mers) binding to patient-specific HLA class I alleles using the tool HLAthena (FIG. 1A).

[0104] A critical component of the pipeline was building a reliable expression atlas for acute or chronic GvHD-targeted tissues (‘GvHD filter’). To capture less frequent albeit biologically relevant cell types consistently underestimated in bulk RNA expression profiles, multiple external single-cell datasets of healthy human skin (Reynolds et al., 2021, Science,' 371 :6527), liver (Aizarani et al., 2019, Nature, 572,199-204; and MacParland et al., 2018, Nat Commun; 9:4383), lung (Vieira Braga et al., 2019, Nat Med, 25: 1153-1163; and Laughney et al., 2020, Nat Med, 26.259-269), GI (Parikh et al., 2019, Nature,' 567:49-55; and Wang et al., 2020, J Exp Med, 217(2):e20191130), oral mucosa (Williams et al., 2021, Cell, 184:4090-4104.e4015), and lacrimal gland (Bannier-Helaouet et al., 2021, Cell Stem Cell, 28: 1221-1232. el227) was evaluated (FIG. IB). For each tissue, the corresponding single-cell data sets were merged and then clustered and annotated tissue-resident cell types, for which specific expressed genes were then identified. Using a cut-off of 5 counts per million for positive expression (based on the expression levels of lineage-defining markers, such as MLANA, SFTB and ALB), 13,512 genes expressed in >1 GvHD target tissue were identified.

[0105] Conversely, to predict candidate mHAgs with an acceptable safety profile if targeted therapeutically, we defined a set of genes with exclusive or near-exclusive expression within the hematopoietic compartment (‘GvL filter’). The focus was on acute myeloid leukemia [AML] and myelodysplastic syndromes [MDS], as they represent the most common indications for allo- HSCT in adults. To capture the transcriptional heterogeneity of malignant myeloid cells, a single-cell based classifier was applied to define genes expressed by leukemic cells in the Beat AML cohort (Tyner et al., 2018, Nature, 562:526-531), and then removed those expressed in non-hemopoietic tissues, per the GTEx database (both RNA (Consortium, The Genotype-Tissue Expression (GTEx) project, 2013, Nat Genet,' 45:580-585) and / or protein (Jiang et al., 2020, Cell; 183:269-283 e219)), applying a gender-specific tolerance for reproductive organs based on the inputted patient gender. Similarly, a broader ‘Hematopoietic filter’ was built by incorporating gene expression profiles from 18 mature hematopoietic lineages and hematopoietic stem and precursor cells.

[0106] Through this stringent process, 259 genes with preferential expression in AML (and 615 broadly expressed in the hematopoietic compartment) were identified. Notably, these genes were distributed: (i) across all chromosomes, thereby providing targetable options even in the presence of chromosomal aberrations, that frequently occur in myeloid malignancies; and (ii) across distinct biological pathways and cellular compartments. The pipeline was further tuned to incorporate RNA-Seq data from leukemic blasts, if available, to confirm patient-specific expression.

[0107] In the -25% of allo-HSCT cases consisting of male recipients paired with female donors (F— >M), chromosome Y-encoded genes represent an additional source of potentially immunogenic epitopes, since the female immune system lacks central tolerance to these. To systematically evaluate Y-encoded mHAgs, genes from the 78 harbored in the male-specific region (MSY) of the Y chromosome with expression >1 TPM in >1 adult GvHD target tissue per GTEx were selected, generated corresponding in silico proteomes, and filtered out all homologous peptides (100% BLAST identity) arising from other chromosomal locations (primarily gene paralogues). HLA class I binding prediction of the remaining set of unique k- mers resulted in a median of 62 (range: 24 - 107) predicted epitopes per allele. The number of predicted epitopes per MSY gene varied greatly across the different HLA alleles, when grouped based on their peptide binding motifs.

[0108] The ability of the pipeline to predict a set of known mHAgs (12 autosomal (Griffioen, et al., 2016, Front Immunol, 7: 1002) and 9 Y-encoded (Feng et al., 2008, Trends Immunol, 29:624-632)) was confirmed in a training dataset of 19 D-R pairs with available WES data (Bachireddy et al., 2021, Cell Rep; 37: 109992; and Bachireddy et al., 2020, Sci Transl Med 12: 12(561):eabb7661). The pipeline verified that: (i) genes harboring the causative SNPs were included in the GvHD and Y filters, respectively; (ii) the SNPs were detected in various combinations among the D-R pairs in accordance with their reported allelic frequencies (Sherry et al., 2001, Nucleic Acids Res; 29:308-311); (iii) in D-R pairs with correct SNP configuration (i.e. Dneg, Rpos), peptides corresponding to the mHAg epitopes were part of in silico generated -mers; and iv) mHAg-corresponding peptides were predicted as HLA binders. Overall, our pipeline predicted 18 of 21 known mHAgs. Detection failures were due to the epitope originating from an alternative open reading frame (not captured by our pipeline, which focuses on mHAgs deriving from SNPs, indels and frameshifts) or to HLA prediction rank above the threshold of positivity.

[0109] Antigenicity and immunogenicity of predicted mHAgs

[0110] To ascertain whether predictions were supported by evidence of HLA presentation, a previously generated HLA class I ligandome dataset of 60 single-HLA class I expressing B721.221 cell lines that encompassed all peptide-binding motifs identified by Sarkizova et al. (Sarkizova et al., 2020, Nat Biotechnol,' 38: 199-209) was evaluated. From whole-genome sequencing of parental B721.221 cells, exonic non-synonymous SNPs present in the B721.221 cells (surrogate ‘HSCT recipient’) that were absent in the reference genome (surrogate ‘donor’) were identified. Epitopes predicted from these sites were searched against the immunopeptidomes from these monoallelic cell lines, and 517 were confirmed to be presented across all 60 HLA alleles (median of 8 [range: 1-20] MS-supported peptides per allele). To evaluate Y mHAg antigenicity, male immunopeptidomes from 11 male cell lines were generated and interrogated and from the IEDB database for epitopes predicted from the 9 MSY genes (since B721.221 cells are of female origin. Indeed, 81 predicted Y mHAgs were presented across 37 HLA alleles.

[0111] To estimate the fraction of predicted binders that were immunogenic, Y-encoded mHAgs in the setting of F— >M transplants were analyzed, since SNP genotyping was not required for functional validation, and this facilitated ready and unbiased testing of all predicted epitopes for any given HLA allele. Peripheral blood mononuclear cells (PBMCs) were isolated from 3 healthy female donors carrying 6 common class I HLA alleles (HLA-A0201, -A0101, -B0702, - Bl 801, -C0501, and -C0702), and challenged purified T cells with pools of synthetic peptides encompassing all 410 predicted binders (HLAthena rank <0.5, n=53 for HLA-A0201; n=97 for HLA-A0101; n=47 for HLA-B0702; n=68 for HLA-B 1801; n=94 for HLA-C0501; and n=51 for HLA-C0702). Peptides were pulsed on CD3-depleted PBMCs, and cocultured with autologous female-derived T cells. Cells were similarly restimulated after 7 days; on day 14, they were screened for antigen specificity by dextramer staining.

[0112] Of 410 peptides tested, 110 (14, 21 and 75 HLA- A, -B and -C restricted, respectively) elicited antigen-specific T cell responses, with a preponderance of HLA-C-restricted epitopes. For 9 such Y mHAgs, antigen-specific recognition by CD137 upregulation and IFNy production in response to the cognate epitope presented on autologous immortalized B cell lines was confirmed. To investigate the unexpectedly high number of immunogenic HLA-C-restricted epitopes, epitope hydrophobicity, a property associated with immunogenicity, was evaluated). Observed inter-allele differences were mainly driven by the richness of hydrophobic residues in the allele-specific binding motifs, with HLA-A0201 having the highest frequency of hydrophobic predicted binders (Kyte-Doolittle hydrophobicity score. Comparison between peptides with or without experimental evidence of immunogenicity across HLA alleles revealed immunogenicity to correlate with higher hydrophobicity scores only for HLA-A0101 and HLA- C0501. Thus, hydrophobicity potentially contributed but was not the sole driver of HLA-C epitope immunogenicity. Additional contributing factors were considered, including: (i) the in vitro stimulation conditions potentially favoring HLA-C restricted responses, as antigen- presenting cells express higher cell surface levels of HLA-C than other cell types; and (ii) a less stringent central tolerance for HLA-C epitopes, since cortical and medullary thymic epithelial cells tend to express less HLA-C than HLA-A and -B. Nevertheless, by longitudinally tracking Y-encoded mHAg-specific T cells ex vivo in a patient undergoing a F^M transplant and experiencing severe chronic GvHD, a sizeable population (accounting for up to 3% of the circulating CD8+ cell pool) of T cells specific for HLA-C0501 -restricted epitopes with in vitro evidence of immunogenicity was detected, thereby supporting the in vivo impact of HLA-C- restricted mHAgs. mHAgs shape GvHD outcomes

[0113] To determine how mHAg load relates to GvHD risk, The mHAg pipeline was used to evaluate paired whole-exomes (>85% target base coverage at >20x depth) generated from 220 D- R pairs treated with a matched-related donor (MRD) allo-HSCT for AML / MDS (FIG. 4). The impact of diagnosis, prognostic risk score, status at transplant, conditioning regimen, and graft source against median mHAg load on risk for acute and chronic GvHD were considered. Across these patients, only total autosomal mHAg load > median was associated with increased risk of developing grade II-IV acute GvHD (HR = 2.54, 95% confidence interval: 1.03 - 6.26, p = 0.043. While the total autosomal mHAg load was not associated with development of chronic GvHD, subgroup analysis of 55 F^M transplants revealed that a combined score of autosomal and Y- encoded mHAgs > median was associated with development of NIH moderate / severe chronic GvHD (p = 0.042).

[0114] Whether organ-specific mHAg load could predict the risk for organ-restricted acute or chronic GvHD was analyzed. By evaluating organ-specific GvHD occurrence across deciles of mHAg load, two scenarios were identified with distinct patterns of association. In the case of chronic lung GvHD, all 14 positive cases occurred in patients with a lung-specific autosomal mHAg load > median (FIGS. 2C and 2D). By logistic regression modelling with other key clinical variables, lung mHAg load > median was the sole factor associated with increased risk of lung chronic GvHD (OR = 17.22, 95% confidence interval: 3.18 - 321.63, p = 0.008, FIG. 2E).

[0115] In contrast, for acute GvHD of the liver, an opposite pattern was observed, with 12 of 13 affected patients having a liver mHAg load < median (FIG. 3A). A limited set of immunodominant liver mHAgs likely drives the clinical manifestations in these patients. By analyzing the mHAg landscape of patients affected by liver acute GvHD, an enriched representation of 7 polymorphisms in genes with preferential expression in liver was observed (FIG. 3B). Notably, these SNPs were less frequent in patients developing acute GvHD without liver involvement (p = 0.002) and co-occurred at a lower frequency in patients experiencing chronic liver GvHD (p = 0.0002). This last observation could be partly explained by the presence of interferon-responsive elements in the promoter region of 4 of 7 of these genes, potentially leading to their increased expression (and presentation) in the inflammatory milieu characterizing the early peri-transplant period.

[0116] Extended characterization of circulating T cells from MRD044 (liver acute GvHD) was undertaken by full flow cytometry panel testing the presence of T cell specific for all the three predicted driver liver mHAgs for MRD044. The specificities tested were: ATXN2 p.L107P (HLA-B0702-restricted, TGFB1 p.PlOL (HLA-B0702-restricted) and VSTM4 p.F68S (HLA- C0701 -restricted). T cell reactivity against CMV pp65 was analyzed in parallel as control. Specifically, the TPRVTGGGAM HLA-B0702-restricted epitope was used (as it shared the same HLA restriction of 2 out of 3 liver mHAgs), and when additional sample was available (i.e., from day +236, +430 and +469 post-transplant) also the NLVPMVATV HLA-A0201 -restricted epitope was analyzed.

[0117] For one patient experiencing liver acute GvHD, ex vivo T cells specific for 1 (of the 3 predicted) driver liver mHAgs (i.e., an HLA-B0702-restricted ATXN2 p.L107P epitope could be traced. Overall, the data show the potential of mHAg burden for refining the risk prediction of GvHD.

[0118] FIG. 5 is a compilation of all HLAs found as binders of the peptides arising from the seven SNPs associated with acute liver GvHD.

[0119] Discussion

[0120] The analytic pipeline described herein overcomes the limitations of previous efforts to delineate the patient-specific mHAg repertoires due to the incorporation of multiple novel features. First, the pipeline provides genome-wide assessment of the entire mHAg landscape, as it uses WES as data input for mHAg prediction (rather than SNP genotyping, utilized in most prior studies). Second, mHAgs are more stringently filtered to be incorporated within the GvHD or GvL set than any prior study on the basis of expression on GvHD targeted tissues or on malignant myeloid cells. This was achieved through: (i) intensive incorporation of information from high resolution single-cell RNA-seq data (analyzing 354,606 individual cells across 9 datasets covering 6 GvHD-targeted organs) to curate an expression atlas of all resident cell types present in organs frequently targeted by acute and chronic GvHD; and (ii) usage of RNA and protein GTEx data from >15,000 samples across 49 non-hematopoietic organs to selected genes with preferential expression in malignant and non-malignant hematopoietic cells and limited - if not absent - representation in healthy tissues. Third, the pipeline was designed to predict not only autosomal but also Y-encoded mHAgs. Fourth, it substantially reduces false positive predictions by eliminating all possible redundant k-mers. These are peptide sequences that, in addition to arising from the discordant SNP, are also generated from other sites in donor or patient exomes. This was accomplished through extensive pruning via sequential BLAST searches against the patient- and donor-specific proteomes reconstructed in silica from WES; indeed, it was found that -40% of SNP-encompassing k-mers were redundant. As a result of this pruning step, the unbiased screening of Y-encoded mHAgs predicted a higher percentage of epitopes with detectable immunogenicity than similar screenings performed for cancer neoantigens.

[0121] The robust pipeline described in this example provided numerous novel insights with several clear translational implications through its application to a cohort of 220 patients transplanted from HLA-matched related donors. For one, if allowed for the identification of genomic attributes that could aid in the prognostication of GvHD risk. It was found that: (i) autosomal mHAg burden was predictive of grade II-IV acute GvHD incidence, in univariate and multivariable analysis; (ii) the presence of both autosomal and Y-encoded predicted mHAg load > median was associated with increased risk of chronic GvHD in F— >M transplants, with the caveat of the limited number of sex-mismatched patients available for this subgroup analysis; (iii) two mechanisms of organ-specific mHAg expression were associated with risk of distinct types of GvHD, namely predicted organ-specific mHAg burden (for chronic lung GvHD) and presence of immunodominant organ-specific mHAgs (for acute liver GvHD). Thes findings thus indicate that molecular characterization of D-R pairs to define GvHD risk could facilitate the design of personalized post-HCT treatments to minimize this highly morbid condition, including incorporation of additional immunosuppressive therapies, introduced early for high-risk patients, or reduced-intensity approaches in low-risk settings. This is particularly valuable for lung GvHD, which is associated with the highest morbidity and mortality.

[0122] A corollary to such molecular-based risk prognostication is the notion that quantification and qualitative characterization of mHAgs per D-R pair could inform donor selection, when more than one candidate is available.

[0123] The robustness and modular design of the pipeline facilitates the high throughput analysis of additional large-scale patient datasets, from which one can anticipate gaining greater sensitivity to further refine the algorithm for prognostication and prediction of response to HCT in future studies. Such datasets could include the analysis of ethnically diverse allo-HCT study populations to expand the list of actionable GvL mHAgs to improve population coverage - especially for individuals of Asian or African ancestry. Further studies could also evaluate other transplant modalities, such as different donor types (i.e., unrelated donors, versus our current study of related donors) and GvHD prophylaxis regimens, such as post-transplant cyclophosphamide, further expected to substantially impact the mHAg-specific T cell repertoire by in vivo purging of alloreactive T cells. Additional studies on larger cohorts would furthermore aid in providing increased power to potentially pinpoint other organ-specific driver mHAgs in settings beyond lung and liver GvHD.

[0124] In summary, this example demonstrates the potential applications of personalized mHAg prediction in allo-HCT. HCT represents a model setting of precision medicine: one individual donor is selected for one individual recipient based on genetic findings, and optimal donor matching to modulate the GvHD / GvL effects has been an inherent point of debate since its inception. Immunogenetic models, such as the one proposed herein, will have important implications for precision immuno-oncology to improve patient outcomes.

[0125] Methods

[0126] Patient Samples

[0127] Peripheral blood mononuclear cells (PBMCs) were collected from allo-HCT patients and donors following written informed consent through a sample collection protocol (DFCI / HCC 01- 206) approved by the Dana-Farber Institutional Review Board in accordance with the principles of the Declaration of Helsinki. All samples were processed by Ficoll-Paque PLUS (Fisher Scientific) density gradient centrifugation and then cryopreserved with FBS / 10% DMSO and stored in liquid nitrogen until time of analysis.

[0128] For the DFCI-AML-MRD cohort, matched donor and recipient DNA was analyzed, as well as PBMCs in some cases, collected from all adult patients who underwent first T-replete allo-HSCT from a matched related donor (MRD) between January 1st, 2013 and December 31st, 2020. Patients were considered in complete remission if disease activity could not be documented by BM evaluation. All other patients not falling within this definition were categorized as having active disease. Minimal residual disease was not considered for the definition of disease status at transplant, as the information was not available for all the study patients. Additionally, patients were stratified by disease-specific prognostic risk scores: ELN 2017 (Dohner et al., 2017, Blood, 129:424-447) and R-IPSS (Greenberg et al., 2012, Blood, 120:2454-2465) for AML and MDS, respectively. For multivariate analysis, the ELN and R- IPSS scores were consolidated in a single variable, termed ‘overall risk score’ and comprising 3 categories: i) favorable, including ELN favorable and R-IPSS very low / low; ii) intermediate, encompassing ELN intermediate and R-IPSS intermediate; and iii) adverse, including ELN adverse and R-IPSS high / very high. Clinical diagnosis and grading of acute GvHD were annotated according to consensus criteria (Przepiorka et al., 1995, Bone Marrow Transplant, 15:825-828; and Glucksberg et al., 1974, Transplantation,' 18:295-304). Chronic GvHD diagnosis and grading were based on the National Institutes of Health consensus criteria (Pavletic et al., 2010, Biol Blood Marrow Transplant, 16:871-890). Exome Sequencing, Processing, and Analysis

[0129] Library preparation and sequencing. A total of 440 DNA samples originating from 220 D-R pairs were processed and sequenced by whole-exome sequencing (>85% target base coverage at >20x depth; Genomics Platform, Broad Institute). For all D-R pairs, genomic DNA (250 ng) was provided by the HLA typing Lab at the Brigham and Women Hospital (Boston, MA). Library construction from double-stranded DNA was performed using the KAPA Library Prep kit (KAPA Hyper Prep with Library Amplification Primer Mix, product KK8504), with palindromic forked adapters from Integrated DNA Technologies. Libraries were then amplified by 10 cycles of PCRs, and enzymatic clean-ups were performed using AMPure XP beads (Beckman Coulter). Following the library construction, library quantification was performed using a standardized PicoGreen dsDNA Quantitation Reagent (Invitrogen) assay. All library construction, hybridization and capture steps were automated on the Agilent Bravo liquid handling system. Hybridization and capture were then performed using the XGen hybridization and wash kit (IDT) following the manufacturer’s recommendations on the Hamilton Starlet. After post-capture enrichment, library pools were quantified using a qPCR kit from KAPA Biosystem (automated assay on the Agilent Bravo). Based on qPCR quantification, pools were normalized using a Hamilton Starlet and sequenced on NovaSeq S4 platform using the NovaSeq 6000 Xp workflow with a paired-end reads of 2xl51bp. Quality control identification check was performed using fingerprint genotyping of 95 common SNPs by Fluidigm Genotyping.

[0130] Alignment and quality control. All DNA sequence data were processed through Broad Institute pipelines. Outputs from Illumina software were processed by the Picard data- processing pipeline to yield BAM / CRAM files. Raw sequence data were aligned to the human genome hgl9 genome assembly (v.b37, using BWA-MEM [v.0.7.15-rl 140]) provided by the Picard and Genome Analysis Toolkit (GATK) developed at the Broad Institute (Andreatta et al., 2019, Proteomics,' 19:el 800357), a process that involves marking duplicate reads, recalibrating base qualities, and realigning around indels.

[0131] Single-cell Analysis

[0132] For all publicly available datasets of skin (Reynolds et al., 2021, Science,' 371 :6527), liver (Aizarani et al., 2019, Nature; 572,199-204; and MacParland et al., 2018, Nat Commun 9:4383), GI (Parikh et al., 2019, Nature; 567:49-55; and Wang et al., 2020, J Exp Med; 217(2):e20191 130), lung (Vieira Braga et al., 2019, Nat Med, 25: 1153-1163; and Laughney et al., 2020, Nat Med, 26:259-269), lacrimal gland (Bannier-Helaouet et al., 2021, Cell Stem Cell; 28: 1221-1232.el227) and oral mucosa (Williams et al., 2021, Cell ; 184: 4090-4104. e4015), fdes were downloaded from the appropriate repository (GEO or EGAS). For each dataset, only data generated from healthy subjects were extracted and imported into Seurat-compatible objects. All quality control, normalization, and downstream analyses were performed using the R package Seurat (Hao et al., 2021, Cell; 184:3573-3587 e352 ;ver. 4.3.0, available on the world wide web at github.com / satijalab / seurat). Low quality cells were excluded from downstream analyses based on percentage of mitochondrial reads (<20), features per cell (>200 and <4,000), and number of reads per cell (<20,000). For GI, liver, and lung, 2 independent datasets were merged, and data integration and batch correction were performed using Harmony. For all datasets, Louvain clustering was then performed on all cells with the ‘FindClusters’ function using the first 50 PCs and a resolution of 1. Through manual annotation, clusters of hematopoietic origin were identified using standard lineage markers (PTPRC, CD3E, MS4A1, CD79B, KLRB1, KLRG1, LYZ, CD68, CD14, KIT) and excluded from further analysis (with the only exceptions of Langerhans cells in skin and Kupffer cells in liver). Clustering was repeated after the removal of immune cells, and non-immune cell types were manually annotated using the same set of lineage-defining genes used in the original publications. For each cluster, gene expression profiles were compiled, upon exclusion of non-coding genes as well as HLA genes. Genes with a sum count >5 counts per million (CPM) were retained to create the final list of genes to be used for the GvHD filter. For the thymic single-cell dataset from Park et al. (Park et al., 2020, Science; 367(6480):eaay3224), files were downloaded from the worldwide web at developmentcellatlas. ncl. ac.uk and loaded into a Seurat object as described above. EPCAM+ thymic epithelial cells were clustered using a resolution of 0.03 to define 3 macro-clusters (corresponding to cortical, medullary, and myo / neuroendocrine thymic epithelial cells). Pseudobulk expression of AIRE, HLA-A, HLA-B and HLA-C within the 3 clusters was calculated using the ‘ Aggregate Expression' function in Seurat and the resulting expression levels were normalized on the number of cells present in each cluster. Minor Histocompatibility Antigen Pipeline

[0133] Input files for the pipeline were donor and recipient exomes in the form of BAM / CRAM files aligned to the hg!9 reference genome, that were processed using Deep Variant (Poplin et al., 2018, Nat Biotechnol', 36:983-987) version 1.1.0 and Funcotator (part of the GATK package, v4.2.6.1) to define germline variants (including SNPs, indels and frameshifts) and proceed to their annotation, respectively. By comparing resulting VCF files from each D-R pair, only germline variants present exclusively in the recipient and producing nonsynonymous alteration in protein-coding regions were retained. Subsequent steps included filtering for variants present in genes of interest through the use of ‘GvHD,’ ‘GvL,’ and / or ‘Y mHAg’ filters. The ‘GvHD filter’ has been already described in the previous section. The GvL filter features 2 main components: the ‘AML filter’ (to define genes expressed by malignant myeloid cells) and the ‘Hematopoietic filter’ (to define genes expressed across the different hematopoietic lineages). With respect to the ‘AML filter’ design, to address the challenge posed by AML heterogeneity, a molecular classifier based on AML single-cell expression profiles (van Galen et al., 2019, Cell,' 176: 1265-1281 el224) is used to define 7 expression clusters in the Beat AML cohort (Tyner et al., 2018, Nature,' 562:526-531). Genes expressed with TPM >2 in each cluster are retained. Next, to select genes with preferential expression in AML, all genes with median expression >5 TPM and / or max expression >8 TPM in any of the normal tissues present in the Genotype-Tissue Expression (GTEx) Project RNA repository were filtered out. More specifically, based on the patient gender (a required input information when running the pipeline), the list of GTEx tissues varies to include either female reproductive organs (for male patients) or male reproductive organs (for female patients); for both female and male patients, whole blood, spleen, and lung are excluded given the inherent large proportion of hematopoietic cells present. The resulting list of genes is then subjected to a second filtering round using the GTEx proteomics dataset (Jiang et al., 2020, Cell,' 183:269-283 e219), with all genes with a tissue specificity score (TS) >2 in any GTEx tissue being removed; in instances when protein data was missing, we applied a second RNA-Seq filtering step, using a cutoff of log2(TPM)>5 in any GTEx tissue, a threshold shown to reliably correspond to protein expression (Jiang et al., 2020, Cell,' 183:269-283 e219). The same filtering steps are applied for the generation of the ‘Hematopoietic filter,’ using as starting point publicly available RNA-Seq expression profiles of 18 purified mature hemopoietic cell types (Uhlen et al., 2019, Science, 366:6472) and of hematopoietic stem and precursor cells (Drissen et al., 2019, Sci Immunol,' 4(35):eaau7148; and Cesana et al., 2018, Cell Stem Cell,' 22:575-588 e577). Overall, the resulting gene sets are composed by 259 genes for the ‘AML filter’, and 615 for the ‘Hematopoietic filter’ (with 224 overlapping genes). The list of variant-encompassing k- mers are then blasted against custom in silico proteomes inferred from the exomes from both donor and recipient, sequentially. All k-mers found elsewhere in the proteomes (100% homology) are discarded, and remaining unique k-mers are subjected to binding prediction to the patient-specific class I HLA alleles through HLA thenci (Sarkizova et al., 2020, Nat Biotechnol,' 38: 199-209) using a threshold of 0.5% prediction rank to define binders. Whenever available, pre-transplant AML RNA-Seq expression can be provided as an optional input file, in order to filter prediction results based on actual expression in the specific patient analyzed. The pipeline is entirely docked on Terra, one of the NCI Cloud Pilots systems, to ensure fully reproducible analyses.

[0134] Immunopeptidome Analysis

[0135] B721.221 monoallelic cell line immunopeptidome analysis. WGS from parental B721.221 was processed using BWA-MEM [v.0.7.15-rl 140], subsetted to coding regions, and annotated with DeepVariant. A modified version of the mHAg pipeline was used in order to predict all k-mers deriving from non-synonymous polymorphisms present in the B721.221 (serving as surrogate ‘recipient’) and not in the reference hgl9 genome (serving as surrogate ‘donor’). Genes were filtered based on RNA-Seq expression in B721.221 (TPM>10). Raw mass spectra from 60 HLA-monoallelic B721.221 cell lines were interpreted with the Spectrum Mill (SM) software package, version 8.01 (Broad Institute; proteomics.broadinstitute.org). MS / MS spectra were excluded from searching if they did not have a precursor sequence MH+ in the range of 700-2000 and a precursor charge of + 1 to +3. Similar spectra with the same precursor m / z acquired in the same chromatographic peak were merged. MS / MS spectra with a sequence tag length >1 (i.e. minimum of three masses separated by the in-chain masses of two amino acids) were searched with no-enzyme specificity. MS / MS spectra were searched against the B721.221 -specific SNPs appended to a base reference proteome composed by 98,298 entries, including all University of California Santa Cruz Genome Browser genes with hgl9 annotation of the genome and its protein-coding transcripts (63,691 entries), 602 common laboratory contaminants, 2043 curated smORFs (IncRNA and uORFs), 237,427 nuORF DB vl .037, and the JPT iRT peptides (JPT Peptide Technologies, Berlin, Germany, RTK-1 -10 pmol) for a total of 303,803 entries (Ouspenskaia et al., 2022, Nat Biotechnok, 40:209-217). Target-decoy FDR estimation was enabled by Spectrum Mill with on-the-fly generation of decoy sequences during searches. For each candidate sequence passing the precursor mass tolerance filter, the internal sequence was reversed, while holding fixed the second position and the peptide C terminus, to maintain not only equal size target and decoy search spaces, but also comparable HLA class I binding motifs among the sequence candidate population. Peptide spectrum matches (PSMs) within <1% false discovery rate (% FDR) were confidently assigned for individual spectra via the target decoy estimation of the SM Autovalidation module. PSMs were filtered for precursor charges of + 1 to +5, sequence lengths ranging between 9 to 40 amino acids, and a minimum backbone cleavage score of 5. PSMs were consolidated to the peptide level to generate lists of confidently observed peptides for each allele using the Spectrum Mill protein / peptide summary module’s peptide-distinct mode with filtering distinct peptides set to case sensitive.

[0136] IEDB search for Y-encoded mHAgs. A curated set of previously identified HLA class I ligands was downloaded from IEDB on the worldwide web at iedb.org / downloader.php?file_name=doc / mhc_ligand_full.zip (accessed on September 19, 2022). Records were filtered to Organism = homo sapiens, Epitope Object Type = Linear peptide, Parent Protein and Antigen Name = MSY genes of interest and Allele Name consistent with human HLA class I nomenclature.

[0137] AML cell lines. The following AML cell lines were purchased from DSMZ: Mono- MAC6, MUTZ-3, OCI-AML3 and SET2; MOLM-13 were kindly gifted by the Genovese Lab (DFCI). All cell lines are part of the LL-100 panel (Quentmeier et al., 2019, Sci Rep,' 9:8218), i.e., profiled at the genomic and transcriptomic level with publicly available results. For each cell line, paired WES and RNA-Seq data were downloaded from ENA (accession numbers: PRJEB30297 and PRJEB30312, respectively). WES was run through a modified version of the mHAg pipeline, using hgl9 as the surrogate ‘donor’ to define the GvL mHAg landscape of each AML cell line. RNA-Seq data was used to evaluate the expression levels of the genes harboring the GvL SNPs and to infer HLA class I typing with OptiType (Szolek et al., 2014, Bioinformatics, 30:3310-3316). Up to 50 million or 0.2 g of each AML cell line were immunoprecipitated, as previously described (Sarkizova et al., 2020, Nat Biotechnol, 38: 199- 209). Peptides of three immunoprecipitations were combined, acid eluted, and analyzed using LC / MS-MS on Orbitrap Exploris 480 equipped with a FAIMSpro interface (Thermo Fisher Scientific). MS spectra were interpreted with Spectrum Mill as detailed in the ‘B721.221 monoallelic cell line immunopeptidome analysis’ section, with the only difference being the proteome customization with the GvL SNP entries.

[0138] Immunogenicity Assays

[0139] Peptides. All synthetic 8-1 Imers peptides used throughout the study were purchased as microscale libraries with purity >70% (median purity: 93.8%, range: 70.1% - 99.8%) from GenScript, and dissolved in ultrapure DMSO (Sigma Aldrich) to a stock concentration of 10 mM.

[0140] Antigen-specific stimulation. PBMCs for immunogenicity testing were isolated by Ficoll- Paque PLUS (Fisher Scientific) density gradient centrifugation from peripheral blood samples of healthy donors. DNA was extracted with the DNA mini kit (Qiagen) following manufacturer’s instructions and used for HLA typing (through the Brigham and Hospital HLA typing lab), and for SRY PCR to select those from female donors with appropriate HLA restrictions. SRY PCR primers were used, following the protocol described by Cui et al. (Cui et al., 1994, Lancet, 343: 79-82). T cells were enriched from PBMCs using the PanT cell selection kit (Miltenyi Biotech) and then stimulated with autologous CD3-depleted PBMCs pulsed with pools of 5 pM synthetic Y-encoded mHAg peptides at 1 : 1 ratio, in RPMI-1640 supplemented with 5% AB-positive heat- inactivated human serum (Gemini Bioproduct), in presence of 30 ng / ml of IL-21 (Peprotech) till day 3, then replaced by 5 ng / ml of IL-7 and IL-15 from day 4 on (both from Peprotech). On day 7 cultures were restimulated with CD3-depleted PBMCs pulsed with 2 pM peptide pools, in the presence of 5 ng / ml of IL-7 and IL-15. Half-medium change and supplementation of cytokines were performed every 3 days.

[0141] Quantification of mHAg-specific T cells. The presence of antigen-specific T cells was determined on day 14 post stimulation, by staining with HLA-A0201, -A0101, -B0702, -B1801, -C0501 and -C0702 easYmers (Immunaware) first loaded with the relevant peptides, and then docked on U-Load dextramers (Immudex), following manufacturer’s instructions. mHAg specificity was assessed by analyzing 3 specificities at a time using triplets of FITC-, PE- and APC-conjugated U-Loads. Cells were then stained with anti-CD3 antibody (conjugated with BV510, clone UCHT1, Biolegend), CD8 (BV785, clone RPA-T8, Biolegend), CD4 (PerCP- Cy5.5, clone RPA-T4, Biolegend), and Zombie Violet (vitality dye, Biolegend). Samples were acquired on a high throughput sampler (HTS)-equipped Fortessa cytometer (BD Biosciences) and analyzed using Flowjo vl0.8 software (BD Biosciences).

[0142] Generation of EBV-immortalized B cell lines (EBV-LCLs). CD 19+ cells were isolated from PBMCs of the same female healthy subjects used for the immunogenicity studies through magnetic selection (Miltenyi Biotech). 0.2 - 0.5 x 106purified B cells were then incubated with a 1 : 1 mix of RPM1-1640 supplemented with 20% FBS and 1% penicillin / streptomycin, and EBV supernatant (ATCC) in a total volume of 200 uL. Every 3-4 days thereafter, cultures were examined for signs of transformation and fresh medium was added. EBV-immortalized cell lines were established in 3-4 weeks and were then maintained in culture with RPM1-1640 supplemented with 10% FBS at a maximum density of 1-2 x 106cells / mL.

[0143] CD137 and IFNy catch assays. For 9 mHAg specificities selected based on the expansion level of mHAg-specific T cells, functional validation of antigen-specificity was performed used 2 different assays: CD137 upregulation and IFNy production upon coculture with autologous EBV-LCLs pulsed with the relevant (or control) peptides. Peptide pulsing of target cells was performed by incubating EBV-LCLs in serum-free medium at a density of 2 * 106cells per ml for 2 h in the presence of 5 iiM peptides. Ovalbumin (OVA) peptide was used as control. For CD 137 assay, upon overnight co-incubation of effector and target cells, mHAg specificity was assessed by flow cytometric detection of CD137 upregulation on CD8+ T cells, using the following antibodies: anti-human CD8a (BV785, clone RPA-T8, Biolegend), CD4 (PerCP - Cy5.5, clone RPA-T4, Biolegend), Zombie Violet (vitality dye, Biolegend), CD19 (APC- Fire750, clone HIB19, Biolegend) and CD14 (APC-Fire750, clone M5E2, Biolegend). Data were acquired on a high throughput sampler (HTS)-equipped Fortessa cytometer (BD Biosciences) and analyzed using Flowjo vl0.8 software (BD Biosciences). To detect the antigen-specific production of IFNy, T cells were co-incubated with autologous EBV-LCL pulsed with the appropriate peptides for 6 hours, and then labeled with the IFNy secretion Assay detection kit in PE (Miltenyi Biotech) following manufacturer’s instructions. Cells were then counterstained with the following antibodies: anti-human CD8a (BV785, clone RPA-T8, Biolegend), CD4 (PerCP-Cy5.5, clone RPA-T4, Biolegend), Zombie Violet (vitality dye, Biolegend), CD19 (APC-Fire750, clone HIB19, Biolegend) and CD14 (APC-Fire750, clone M5E2, Biolegend). Data were acquired on a high throughput sampler (HTS)-equipped Fortessa cytometer (BD Biosciences) and analyzed using Flowjo vl0.8 software (BD Biosciences).

[0144] Peptide hydrophobicity calculation. The GRAVY hydrophobicity index with the Kyte- Doolittle scale was calculated for the 410 tested Y mHAgs using the ‘Peptides’ R package (v 2.4.4).

[0145] 1000 Genomes in silico allo-HCT Simulation

[0146] The WGS bam fdes for the 1000 Genomes Project (1000G), mapped to the GRCH37 reference genome, were obtained from the Google Brain Genomics repository. Suitable D-R pairs were selected based on the HLA typing information available for each individual in the 1000G dataset. The criterion for selection was that >5 alleles in the HLA class I genes (HLA- A, HLA-B, HLA-C) had to be matched between donor and recipient. A total of 844 individuals combined into 2270 unique D-R pairs satisfying this criterion were identified. The study cohort comprised 239 individuals from EUR, 216 from EAS, 160 from SAS, 140 from AFR and 89 from AMR ancestry. WGS bam files from the selected individuals were reduced to include only the coding region of the genes included in the AML and Hematopoietic filter. Variant calling was performed on the reduced bam files using DeepVariant (version 1.1.0) (Poplin et al., 2018, Nat Biotechnol,' 36:983-987), generating germline variant call files. Allelic frequency (AF) of each GvL SNP was calculated for the 1000G cohort and compared to the gnomAD allele frequency which was incorporated into the Funcotator task. Population coverage analysis was performed as previously described by Bui et al. (Bui et al., 2006, BMC Bioinformatics, 7: 153).

[0147] Statistical Analyses

[0148] All statistical analyses were conducted using Prism v.9.5 (GraphPad Software) or R v.4.2.2 (https: / / www.r-project.org / ). The following statistical tests were used in this study, unless otherwise indicated: paired t test, unpaired t test or Mann Whitney test, based on prior assessment of data normality (using Shapiro-Wilk test or Kolmogorov-Smirnov test, depending on sample size). The minimum threshold for significance was defined as p < 0.05, and all statistical tests were two-sided.

[0149] Clinical outcomes are reported as of June 2022 (lock date: 25 June 2022). Overall survival (OS) was defined as the time from stem cell infusion to death from any cause. Patients who were alive were censored at the time last seen alive. GvHD-free / rel apse-free survival (GRFS) was defined as time to first occurrence of grade III-IV acute GvHD, chronic GvHD requiring systemic treatment, relapse, or death, whichever occurred first. Probabilities of OS and GRFS were estimated with the Kaplan-Meier method using the ‘survival’ (version 3.4-0) R package, while cumulative incidences were estimated using the ‘tidycmprsk’ (version 0.2.0) R package. Cumulative incidence of non-relapse mortality (NRM), relapse, acute GvHD, and chronic GvHD were computed to take into account the presence of competing risks. Specifically, in calculating the cumulative incidence rates of NRM, the competing risk was relapse; when calculating relapse, the competing risk was NRM. For acute and chronic GvHD, the competing risks were relapse and death from other causes. All patients were considered evaluable for acute GvHD analysis, and those who had a documented engraftment and a followup >100 days were evaluated for chronic GvHD (n = 210). To link mHAg load with GvHD outcomes (i.e., overall acute and chronic, as well as organ-specific GvHD), initial evaluation of deciles of mHAg load was used to guide subsequent analyses and define stratification criteria (mHAg load median vs. analysis of individual SNPs). The log-rank test and the Gray test were used for group comparison of OS and cumulative incidence of relapse and GvHD, respectively. The risk factor analysis on acute GvHD-specific hazard was firstly assessed by the Cox proportional-hazards model, and hazard ratio (HR) with the associate 95% confidence interval were calculated for each variable. Forest plots were performed by the forest model function in the ‘forestmodel’ (version 0.6.2) R package. Analysis of promoter region for interferonresponsive elements in the genes associated with liver acute GvHD was performed using the open access interferome database (worldwide web at interferome.org) (Samarajiwa et al, 2009, Nucleic Acids Res; 37:D852-857), using the following search conditions: all interferon types, homo sapiens species (all systems and sample types).

[0150] This example has now published as Cieri et al., “Systematic identification of minor histocompatibility antigens predicts outcomes of allogeneic hematopoietic cell transplantation,” Nat Biotechnol. 2024 Aug 21. doi: 10.1038 / s41587-024-02348-3, which is herein incorporated by reference in its entirety. Example 2

[0151] Acute GvHD of the Gut Is Associated with Minor Histocompatibility Antigens Cross-Reactive Against Gut-Tropic Viral Epitopes

[0152] Severe manifestations of aGvHD of the gut portend a poor prognosis after allogeneic hematopoietic cell transplantation (allo-HCT) and new tools to predict and possibly mitigate GI aGvHD risk are urgently needed. In the context of HLA-matched allo-HCT, T cell alloreactivity is directed against minor histocompatibility antigens (mHAgs), polymorphic peptides resulting from donor-recipient (D-R) disparity at sites of genetic polymorphisms. The previous example describes an analytic framework (FIGS. 1A and IB) to systematically identify mHAgs based on integration of polymorphism detection by whole-exome sequencing (WES) of germline DNA from D-R pairs together with organ-specific expression and HLA class I epitope prediction.

[0153] Building on the clinical observation that viral reactivations often coincide with the development of GvHD, this example explored whether sequence homology of alloreactive peptides with gut-tropic viral epitopes (i.e., from adenovirus (ADV), cytomegalovirus (CMV), and Epstein Barr virus (EBV) may contribute to the pathophysiology of GI aGvHD.

[0154] Methods

[0155] A computational algorithm was devised to quantify the extent of cross-reactivity across the patient-specific mHAg landscape, inferred from analysis of WES of allo-HCT D-R pairs. Complete viral proteomes for ADV, CMV and EBV were retrieved from UniProt (available strains / virus: 51, 14 and 6, respectively) to create in silico all possible 8-1 Imers (n= 1,108,265) that were then subjected to HLA class I epitope prediction, resulting in a catalog of 184,889 viral epitopes across 105 HLA class I alleles. Predicted mHAgs from WES analysis of 220 HLA- matched D-R pairs were fdtered for expression in GLresident non-hematopoietic cells, as identified by single cell RNA-Seq datasets of GI tissue from healthy and post-transplant individuals. Homology with predicted viral epitopes was defined by either: (i) sliding window, wherein the mHAg and viral antigen were required to share >6 consecutive amino acids; or (ii) skipping window, wherein the Levenshtein distance between the mHAg and viral antigen was <2 amino acids, irrespective of their position within the epitope sequence. Predicted mHAgs were considered cross-reactive only if displaying sequence homology with viral epitopes presented on the same HLA restriction.

[0156] Results mHAg load is not sufficient to predict GI aGvHD. This example systematically evaluated the load of cross-reactive alloantigens (FIGS. 6 and 7) (Macdonald et al., Immunity 2009; Borbulevych et al., Immunity 2009 Amir et al., Blood 2010; Wang et al., PNAS 2017; Hamza et al., Blood 2024; Dolton et al., JCI 2024).

[0157] The DFCI-AML-MRD cohort is detailed in FIG. 4. Across the 220 HLA-matched allo- HCT D-R pairs analyzed (for recipients with AML / MDS), the median number of cross-reactive Gl-specific mHAgs was 37 (range: 13 - 87).

[0158] As shown in FIGS 8A and 8B, the load of cross-reactive mHAgs correlates with the incidence of severe GI aGvHD. Patients who developed severe GI aGvHD (grade III-IV, n=l 1) had a higher load of cross-reactive Gl-specific mHAgs than patients who did not (P = 0.038). The impact of diagnosis, prognostic risk score, status at transplant, conditioning regimen, graft source, donor age, and CMV status against median cross-reactive mHAg load on risk for GI aGvHD was considered. Across these patients, only cross-reactive mHAg load > median was associated with increased risk of developing severe GI acute GvHD (HR = 4.96, 95% CI: 1.13 - 19.44, P = 0.03).

[0159] As shown in FIG. 9, multivariable analysis confirms cross-reactive mHAg load as an independent risk factor for GI. To functionally verify cross-reactivity, attention was focused on the Gl-specific mHAg repertoire of 2 patients in the cohort who experienced severe GI aGvHD and shared the common HLA-A*03:01, B*07:02 and C*07:02 haplotype. T cells from healthy donors were first challenged with the same haplotype with synthetic peptides corresponding to predicted viral epitopes. After 2 rounds of stimulation, T cells were screened for antigen specificity by dextramer staining. Of the 21 peptides tested, 4 (1 HLA-B*07:02, 3 HLA- C*07:02; 2 ADV- and 2 EBV-derived epitopes) elicited viral-specific responses. T cells were then co-cultured with autologous antigen-presenting cells pulsed with the viral epitopes or cross- reactive mHAgs, and T cell activation was assessed by CD 137 upregulation. For all 4 viral epitopes, viral-specific T cells upregulated CD137 also in response to the cross-reactive GI- specific mHAgs, thereby indicating their potential for cross-reactivity. This example demonstrates that cross-reactivity against gut-tropic viruses is a key contributor to GI aGvHD pathophysiology. Molecular characterization of D-R pairs to quantify cross-reactivity provides a tool for pre-transplant prognostication and may facilitate the design of personalized post-HCT treatments to minimize this highly morbid condition.

[0160] This example also demonstrates that molecular mimicry between viral epitopes and mHAgs can be systematically assessed and used in GvHD risk prognostication. Evaluation of the cross-reactivity with microbiome-derived epitopes is warranted. Further, this example identifies potential biomarker to guide the use of prophylactic strategies such as vedolizumab (Chen et al., Nat Med 2024; 30(8): 2277-2287, “Vedolizumab for the prevention of intestinal acute GVHD after allogeneic hematopoietic stem cell transplantation: a randomized phase 3 trial”).

[0161] The complete disclosure of all patents, patent applications, and publications, and electronically available material (including, for instance, nucleotide sequence submissions in, e.g., GenBank and RefSeq, and amino acid sequence submissions in, e.g., SwissProt, PIR, PRF, PDB, and translations from annotated coding regions in GenBank and RefSeq) cited herein are incorporated by reference. In the event that any inconsistency exists between the disclosure of the present application and the disclosure(s) of any document incorporated herein by reference, the disclosure of the present application shall govern. The foregoing detailed description and examples have been given for clarity of understanding only. No unnecessary limitations are to be understood therefrom. Variations of the exact details shown and described, obvious to one skilled in the art, are included within the scope of the claims.

Claims

What is claimed is:

1. A method of identifying a transplant recipient at increased risk of developing acute liver graft versus host disease (GvHD) after receipt of donor tissue from a donor, the method comprising: determining for the donor the presence or absence of one or more single nucleotide polymorphisms (SNPs) from a panel of SNPs; determining for the recipient the presence or absence of one or more SNPs from the panel of SNPs; wherein the panel of SNPs comprises the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP; wherein the donor and recipient are HLA-matched; and wherein the presence in the recipient and the absence in the donor of one or more of the SNPs from the panel of SNPs is indicative of an increased risk of the recipient developing acute liver GvHD after receipt of donor tissue from the donor.

2. The method of claim 1, further comprising administering an immunosuppressive agent to the transplant recipient.

3. A method for selecting a transplant donor, the method comprising: obtaining single nucleotide polymorphism (SNP) information for a potential donor for one or more SNPs from a panel of SNPs; wherein the panel of SNPs comprises the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP; determining for the potential donor the presence or absence of one or more SNPs from the panel of SNPs; obtaining SNP information for a transplant recipient for one or more SNPs from the panel of SNPs; and determining for the transplant recipient the presence or absence of one or more SNPs from the panel of SNPs;wherein the potential donor and transplant recipient are HLA-matched; and selecting as the transplant donor a potential donor wherein one or more SNPs from the panel of SNPs is present in both the potential donor and the transplant recipient or one or more SNPs from the panel of SNPs is absent in both the potential donor and the transplant recipient.

4. The method of any one of claims 1 to 3, wherein: the one or more SNPs from the panel of SNPs comprises the ATXN2 p.L107V SNP and the HLA-matched donor and recipient comprise HLA B0801, HLA B0702, HLA B5501, HLA C0602, HLA B2703, and / or HLA A0301; the one or more SNPs from the panel of SNPs comprises the C4orf48 p.17L SNP and the HLA-matched donor and recipient comprise HLA A2402, HLA C0303, HLA A0201, HLA C0602, HLA C0701, HLA C0401, HLA Cl 203, and / or HLA B0801; the one or more SNPs from the panel of SNPs comprises the NCOA7 p.S284A SNP and the HLA-matched donor and recipient comprise HLA A0301, HLA C0303, HLA C0304, HLA B4402, and / or HLA C0802; the one or more SNPs from the panel of SNPs comprises the TGFB1 p.PlOL SNP and the HLA-matched donor and recipient comprise HLA A0201, HLA B0702, and / or HLA C0701; the one or more SNPs from the panel of SNPs comprises the TNC p.Q680R SNP and the HLA-matched donor and recipient comprise HLA A0101, HLA B1801, HLA B1501, HLA C0602, HLA A6801, HLA A0201, and / or HLA B4402; the one or more SNPs from the panel of SNPs comprises the UBE2O p.G1207S SNP and the HLA-matched donor and recipient comprise HLA C0501, HLA B4101, HLA B4001, and / or HLA B3701; and / or the one or more SNPs from the panel of SNPs comprises the VSTM4 p.F68S SNP and the HLA-matched donor and recipient comprise HLA C0501, HLA C0701, HLA C0303, HLA C0304, HLA C0501, HLA C0401, HLA C0702, and / or HLA C0802.

5. A kit for comprising a plurality of probes capable of binding to and / or identifying one or more of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NCOA7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE2O p.G1207S SNP, and / or the VSTM4 p.F68S SNP in a sample from the transplant donor and / or the transplant recipient.

6. The kit of claim 5, the kit comprising a plurality of probes capable of binding to and / or identifying each of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and the VSTM4 p.F68S SNP in a sample from the transplant donor and / or the transplant recipient.

7. A kit comprising a plurality of primers for selectively amplifying one or more of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and / or the VSTM4 p.F68S SNP in a sample from the transplant donor and / or the transplant recipient.

8. The kit of claim 7, the kit comprising a plurality of primers for selectively amplifying each of the ATXN2 p.L107V SNP, the C4orf48 p.P17L SNP, the NC0A7 p.S284A SNP, the TGFB1 p.PlOL SNP, the TNC p.Q680R SNP, the UBE20 p.G1207S SNP, and the VSTM4 p.F68S SNP in a sample from the transplant donor and / or the transplant recipient.

9. A method of determining a transplant recipient’s risk of developing chronic lung graft versus host disease (GvHD) after receipt of donor tissue from a donor, the method comprising: obtaining single nucleotide polymorphism (SNP) information for the donor for a panel of lung-specific minor histocompatibility antigen (mHAg) SNPs; wherein the panel of lung-specific mHAg SNPs comprises one or more of the autosomal lung-specific mHAg SNPs listed in FIG. 10; determining for the donor the presence or absence of each of the SNPs of the panel of lung-specific mHAg SNPs; obtaining SNP information for the recipient for the panel of lung-specific mHAg SNPs; and determining for the recipient the presence or absence of each of the SNPs of the panel of lung-specific mHAg SNPs; wherein the donor and recipient are HLA-matched; wherein the presence in the recipient and the absence in the donor of more than about 424 of the SNPs of the panel of lung-specific mHAg SNPs is indicative of an increased risk of the recipient developing chronic lung GvHD after receipt of donor tissue from the donor; andwherein the presence in the recipient and the absence in the donor of less than about 424 of the SNPs of the panel of lung-specific mHAg SNPs is indicative of a decreased risk of the recipient developing chronic lung GvHD after receipt of donor tissue from the donor.

10. The method of claim 9, further comprising administering an immunosuppressive agent to the transplant recipient.

11. A method for selecting a transplant donor so that the risk of a recipient developing chronic lung graft versus host disease (GvHD) is reduced, the method comprising: obtaining single nucleotide polymorphism (SNP) information for the donor for a panel of lung-specific minor histocompatibility antigen (mHAg) SNPs; wherein the panel of lung-specific mHAg SNPs comprises one or more of the autosomal lung-specific SNPs listed in FIG. 10; determining for the donor the presence or absence of each of the SNPs of the panel of lung-specific mHAg SNPs; obtaining SNP information for the recipient for the panel of lung-specific mHAg SNPs; and determining for the recipient the presence or absence of each of the SNPs of the panel of lung-specific mHAg SNPs; wherein the donor and recipient are HLA-matched; wherein the presence in the recipient and the absence in the donor of more than about 424 of the SNPs of the panel of lung-specific mHAg SNPs is indicative of an increased risk of the recipient developing chronic lung GvHD if the donor is selected as the transplant donor; and wherein the presence in the recipient and the absence in the donor of less than about 424 of the SNPs of the panel of lung-specific mHAg SNPs is indicative of a decreased risk of the recipient developing chronic lung GvHD if the donor is selected as the transplant donor.

12. A method for selecting a transplant donor, the method comprising: determining for the potential donor the presence or absence of one or more of the single nucleotide polymorphisms (SNPs) of a panel of lung-specific minor histocompatibility antigen (mHAg) SNPs; anddetermining for the transplant recipient the presence or absence of one or more of the SNPs of the panel of lung-specific mHAg SNPs; wherein the panel of lung-specific mHAg SNPs comprises one or more of the autosomal lung-specific mHAg SNPs listed in FIG. 10; and wherein the donor and recipient are HLA-matched; and selecting as the transplant donor the potential donor wherein less than about 424 of the SNPs of the panel of lung-specific mHAg SNPs are present in the recipient and absent in the potential donor.

13. The method of any one of claims 9 to 12, wherein an increased risk of chronic lung GvHD comprises a 5-year cumulative incidence of chronic lung GvHD of about 12% or greater.

14. The method of any one of claims 9 to 12, wherein a decreased risk of chronic lung GvHD comprises about a 5-year cumulative incidence of chronic lung GvHD of about 0.93% or less.

15. A kit for identifying a transplant recipient at increased risk of developing chronic lung graft versus host disease (GvHD) after receipt of donor tissue from a donor and / or for selecting a transplant donor so that the risk of the recipient developing chronic lung GvHD is reduced, the kit comprising a plurality of probes capable of binding to and / or identifying one or more of the lung-specific mHAg SNPs listed in FIG. 10 in a sample from the transplant donor and / or the transplant recipient.

16. The kit of claim 15, the kit comprising a plurality of probes capable of binding to and / or identifying each of the lung-specific mHAg SNPs listed in FIG. 10 in a sample from the transplant donor and / or the transplant recipient.

17. A kit for identifying a transplant recipient at increased risk of developing chronic lung graft versus host disease (GvHD) after receipt of donor tissue from a donor and / or for selecting a transplant donor so that the risk of the recipient developing chronic lung GvHD is reduced, the kit comprising a plurality of primers for selectively amplifying one or more of the lung-specificmHAg SNPs listed in FIG. 10 in a sample from the transplant donor and / or the transplant recipient.

18. The kit of claim 17, the kit comprising a plurality of primers for selectively amplifying each of the lung-specific mHAg SNPs listed in FIG. 10 in a sample from the transplant donor and / or the transplant recipient.

19. The method of any one of claims 1 to 4 or 9 to 14, wherein obtaining the SNP information for the donor and / or the recipient comprises targeted next generation sequencing (NGS), whole exome sequencing (WES), whole genomic sequencing (WGS), and / or SNP arraybased genotyping.

20. A method of identifying a transplant recipient at increased risk of developing acute graft versus host disease (GvHD) of the gastrointestinal (GI) system after receipt of donor tissue from a donor, the method comprising: identifying peptide polymorphisms of minor histocompatibility antigen (mHAgs) between the donor and the recipient, wherein the mHAgs are Gl-specific mHAgs (expressed on Gl-resident non-hematopoietic cells), comparing the amino acid sequences of the Gl-specific mHAgs with the amino acid sequences of antigenic epitopes of one or more viruses relevant for allo-HCT recipients, identifying and quantifying the number of amino acid sequences of the Gl-specific mHAgs that are homologous (cross-reactive) to an amino acid sequence of the antigenic epitopes of the one or more viruses relevant for allo-HCT recipients; wherein a number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients of greater than the median is indicative of an increased risk of the transplant recipient developing acute graft versus host disease (GvHD) of the gastrointestinal (GI) system after receipt of donor tissue.21 . The method of claim 20, wherein the median number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients comprises about 37 (range of about 13 to about 87).

22. A method for selecting a transplant donor so that the risk of a recipient developing acute graft versus host disease (GvHD) of the gastrointestinal (GI) system is reduced, the method comprising: identifying peptide polymorphisms of minor histocompatibility antigen (mHAgs) between the donor and the recipient, wherein the mHAgs are Gl-specific mHAgs (expressed on Gl-resident non-hematopoietic cells), comparing the amino acid sequences of the Gl-specific mHAgs with the amino acid sequences of antigenic epitopes of one or more viruses relevant for allo-HCT recipients, identifying and quantifying the number of amino acid sequences of the Gl-specific mHAgs that are homologous (cross-reactive) to an amino acid sequence of the antigenic epitopes of the one or more viruses relevant for allo-HCT recipients; wherein a number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients greater than the median is indicative of a decreased risk of the transplant recipient developing acute graft versus host disease (GvHD) of the gastrointestinal (GI) system after receipt of donor tissue.

23. The method of claim 22, wherein the median number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients comprises about 37 (range of about 13 to about 87).

24. A method of identifying a transplant recipient suitable for post-transplant administration of immunosuppressive agent to minimize acute graft versus host disease (GvHD) of the gastrointestinal (GI) system along with receipt of donor tissue from a donor, the method comprising: identifying peptide polymorphisms of minor histocompatibility antigen (mHAgs) between the donor and the recipient,wherein the mHAgs are Gl-specific mHAgs (expressed on Gl-resident non-hematopoietic cells), comparing the amino acid sequences of the Gl-specific mHAgs with the amino acid sequences of antigenic epitopes of one or more viruses relevant for allo-HCT recipients, identifying and quantifying the number of amino acid sequences of the Gl-specific mHAgs that are homologous (cross-reactive) to an amino acid sequence of the antigenic epitopes of the one or more viruses relevant for allo-HCT recipients; wherein a number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients greater than the median is indicative of a transplant recipient suitable for post-transplant administration of a prophylactic immunosuppressive agent to minimize acute graft versus host disease (GvHD) of the gastrointestinal (GI) system along with receipt of donor tissue from a donor.

25. The method of claim 24, wherein the median number of Gl-specific mHAgs homologous to an antigenic epitope of one or more viruses relevant for allo-HCT recipients comprises about 37 (range of about 13 to about 87).

26. The method of any one of claims 24 to 25, wherein the prophylactic immunosuppressive agent targets GvHD of the gut.

27. The method of any one of claims 24 to 25, wherein the prophylactic immunosuppressive agent comprises vedolizumab.

28. The method of any one of claims 20 to 27, wherein the virus relevant for allo-HCT recipients comprises a gut-tropic virus.

29. The method of any one of claims 20 to 28, wherein the virus relevant for allo-HCT recipients comprises an adenovirus (ADV), a cytomegalovirus (CMV) and / or an Epstein-Barr virus (EBV).

30. The method of any one of claims 20 to 29, wherein the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients comprises about 8 to about 11 amino acids.

31. The method of any one of claims 20 to 30, wherein the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients comprises about 8, about 9, about 10, or about 11 amino acids.

32. The method of any one of claims 20 to 31 , wherein the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients consists of 8 to 11 amino acids.

33. The method of any one of claims 20 to 32, wherein the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients consists of a peptide of 8, 9, 10, or 11 amino acids.

34. The method of any one of claims 20 to 33, wherein the amino acid sequence of the antigenic epitope of the virus relevant for allo-HCT recipients comprises a HLA class I epitope.

35. The method of any one of claims 20 to 34, further comprising administering a prophylactic immunosuppressive agent to the transplant recipient.

36. The method of claim 35, wherein the prophylactic immunosuppressive agent targets GvHD of the gut.

37. The method of claim 35 or 36, wherein the prophylactic immunosuppressive agent comprises vedolizumab.

38. The method of any one of claims 1 to 4, 9 to 14, 19, or 20 to 37, wherein the donor and the recipient are human.

39. The method of any one of claims 1 to 4, 9 to 14, 19, or 20 to 38, wherein the donor and recipient are HLA-matched.

40. The method of any one of claims 1 to 4, 9 to 14, 19, or 20 to 39, wherein the transplantation of donor tissue is for the treatment of a malignant or a nonmalignant hematologic disease with indication to allogeneic hematopoietic cell transplantation.

41. The method of any one of claims 1 to 4, 9 to 14, 19, or 20 to 39, wherein the transplantation of donor tissue is for the treatment of cancer, sickle cell disease, an inherited anemia, or bone marrow failure.

42. The method of claim 41, wherein the cancer comprises a hematologic cancer.

43. The method of claim 42, wherein the hematologic cancer comprises a leukemia.

44. The method of claim 41, wherein the cancer comprises acute myeloid leukemia (AML), myelodysplastic syndrome (MDS), acute lymphoblastic leukemia (ALL), T-cell acute lymphoblastic leukemia (T-ALL), peripheral T cell lymphoma (PTCL), a myeloproliferative neoplasm (MPN), and / or non-Hodgkin's lymphoma (NHL).

45. The method of any one of claims 1 to 4, 9 to 14, 19, or 20 to 44, wherein the donor tissue comprises hematopoietic cells.

46. The method of any one of claims 1 to 4, 9 to 14, 19, or 20 to 45, wherein the donor tissue comprises HLA-matched allogeneic hematopoietic stem cells.