Cell-free RNA liquid biopsy to monitor hematopoietic stem cell transplantation

cfRNA profiling in plasma or serum samples addresses the invasive nature of current HSCT monitoring by predicting GVHD and infections through biomarker analysis, facilitating early detection and treatment.

WO2025184473A1PCT designated stage Publication Date: 2025-09-04CORNELL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017796
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2025-02-28
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Current methods for monitoring hematopoietic stem cell transplantation (HSCT) complications such as graft-versus-host disease (GVHD) and infections are invasive and lack noninvasive molecular tests for early prediction and diagnosis.

Method used

A method involving the analysis of cell-free RNA (cfRNA) profiles from plasma or serum samples to determine biomarker panels, cell type of origin, and cell turnover, enabling monitoring of HSCT recipients for complications by nucleotide sequencing and RNA profiling.

Benefits of technology

Provides noninvasive monitoring of HSCT recipients for GVHD and infections, predicting complications before they occur, and guiding treatment decisions based on cfRNA dynamics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000048_0001
    Figure IMGF000048_0001
  • Figure IMGF000035_0001
    Figure IMGF000035_0001
  • Figure 00000062_0000
    Figure 00000062_0000
Patent Text Reader

Abstract

The current disclosure is directed to methods of monitoring a subject undergoing hematopoietic stem cell transplantation (HSCT) and methods for monitoring a microbiome of an HSCT recipient subject for infection.
Need to check novelty before this filing date? Find Prior Art

Description

CELL-FREE RNA LIQUID BIOPSY TO MONITOR HEMATOPOIETIC STEM CELL TRANSPLANTATIONCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application Serial No. 63 / 560,363, filed March 1, 2024, the entire contents of which are incorporated herein by reference.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] This disclosure was made with government support under a research project supported by National Institutes of Health (NIH) grant R01 Al 146165. The government has certain rights in this invention.BACKGROUND

[0003] Hematopoietic stem cell transplantation (HSCT) is one of the first successful immunotherapies. In allogeneic HSCT, the recipient’s hematopoietic and immune systems are rapidly replaced by donor cells. For patients with leukemia, the engraftment of donor immune cells contributes to a graft-versus-leukemia effect. However, the success of HSCT is limited by the frequent occurrence of complications, including graft-versus-host disease (GVHD), relapse, and infection. GVHD is one of the most common and severe complications of allogeneic HSCT. GVHD occurs when engrafted T cells attack the host’s organs, which can lead to tissue damage and in severe cases, death. GVHD is further classified into acute GVHD (aGVHD), which typically occurs within the initial 100 days post-transplantation, and chronic GVHD (cGVHD), which develops later. Most HSCT recipients experience at least one complication and complications can often co-occur, such as GVHD and infection. Complications can manifest with non-specific symptoms, such as rash or elevated liver enzymes, and progress rapidly, making early prediction and diagnosis critical. The current standard of care involves a battery of diagnostic tests including invasive tissue biopsies. There is a clear need for noninvasive molecular tests to monitor HSCT recipients and predict complications before they occur.

[0004] Cell-free DNA (cfDNA) and cell-free RNA (cfRNA) are released in plasma as a byproduct of active excretion and cell death, and therefore reflect cell turnover, immune activity,and solid tissue damage. Recent studies have shown that methylation profiling of cfDNA can be used to identify complications of allogeneic HSCT.

[0005] Several recent studies highlight applications of RNA liquid biopsy in inflammatory and infectious disease, including pediatric inflammatory diseases, tuberculosis, COVID-19, sepsis, and non-alcoholic fatty liver disease. There has been only one study on the use of cfRNA in HSCT, and it focused on the immediate post-transfusion period, showing that cfRNA reflects bone marrow reconstitution following allogeneic and autologous HSCT (A. Ibarra, et al., Nat. Commun. 11, 400 (2020)). However, that study did not investigate cfRNA dynamics in the context of HSCT complications over a long-time course, as provided herein.SUMMARY

[0006] The current disclosure is directed to methods of monitoring a subject undergoing hematopoietic stem cell transplantation (HSCT) and methods for monitoring a microbiome of an HSCT recipient subject for infection.

[0007] In one aspect, the current disclosure is directed to a method of monitoring a subject undergoing HSCT, the method comprising: a) obtaining cell-free RNA from at least one biological sample of the subject, wherein the biological sample is a plasma sample or serum sample; b) determining: i) a profile of a panel of biomarkers in the cell-free RNA; ii) cfRNA abundances; iii) distance to health; iv) cell type of origin of the cfRNA; and / or v) cell turnover; and c) determining the subject’s health status based on at least one of the modalities listed in b).

[0008] In some embodiments, the at least one biological sample of the subject is obtained on the day of infusion. In some embodiments, the at least one biological sample of the subject is obtained after the day of infusion. In some embodiments, the at least one biological sample of the subject is obtained less than 36 months after the day of infusion. In some embodiments, theat least one biological sample of the subject is obtained less than 2 months after the day of infusion, less than 3 months after the day of infusion, less than 6 months after the day of infusion, or less than 12 months after the day of infusion.

[0009] In some embodiments, the at least one biological sample of the subject comprises a set of longitudinally collected biological samples of the subject, wherein the samples are obtained at preconditioning, day of infusion, engraftment, and / or one or more post engraftment timepoints. In some embodiments, the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint being at less than 2 months after the day of infusion, less than 3 months after the day of infusion, less than 6 months after the day of infusion, less than 12 months after the day of infusion, less than 24 months after the day of infusion, and / or less than 36 months after the day of infusion.

[0010] In some embodiments, the subject is receiving immunosuppression treatment and will be tapered off the immunosuppression treatment at an initial tapering date, the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint occurs prior to the initial tapering date, less than three months after the initial tapering date, and / or less than six months after the initial tapering date.

[0011] In some embodiments, the step of determining a profde of a panel of biomarkers comprises performing nucleotide sequencing of the cell-free RNA.

[0012] In some embodiments, the step of determining a profde of a panel of biomarkers comprises measuring the levels of RNAs corresponding to the biomarkers in the panel.

[0013] In some embodiments, the determining the subject’s health status based on the profde of the panel of biomarkers comprises comparing the profde of the panel of biomarkers in the cell free RNA from the biological samples to a respective reference profde of the panel of biomarkers at the respective time points. In some embodiments, the subject’s health status comprises treatment response, treatment toxicity, likelihood of developing graft versus host disease (GVHD), and / or likelihood of developing complications due to the HSCT. In some embodiments, the complications due to the HSCT comprise microbial / bacterial infection and / or viral infection. In some embodiments, determining the subject’s health status treatment response, further comprises calculating average correlation of cfRNA abundances for each sample with respect to reference profde of the panel of biomarkers. In some embodiments, thesubject’s recovery health status further comprises the cell turnover values for neutrophils, leukocytes, and lymphocytes using the cfRNA cell type of origin fraction, total cfRNA concentration, and blood cell counts. In some embodiments, the total cfRNA concentration is calculated by normalizing the sample’s cDNA library concentration by the volume of input plasma and number of PCR cycles performed during library preparation to get the estimated total cfRNA concentration; multiplying the estimated total cfRNA concentration by the cfRNA cell type of origin fraction for a cell type of interest, resulting in the concentration of cfRNA derived from that cell type; and / or dividing the concentration of cfRNA derived from the cell type of interest by the blood cell counts for that respective cell type, providing an estimate of cell turnover in arbitrary units (a.u.), wherein the estimated cell turnover is comparable between cfRNA samples.

[0014] In some embodiments, the method further comprises quantifying distance to health, profiling cell types of origin, measuring immune cell turnover, and performing RNA biomarker analysis. In some embodiments, the method further comprises characterizing GVHD organ involvement, wherein the characterizing GVHD organ involvement comprises comparing the cell type of origin fraction of relevant cell types for each organ across the reference profile; and correlating organ specific damage in subjects with GVHD. In some embodiments, measuring immune cell turnover comprises comparing the cfRNA profile of leukocyte, lymphocyte, and neutrophil turnover in the subject sample to the reference profile and correlating an increase in leukocyte, lymphocyte, and neutrophil turnover to a prediction of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject. In some embodiments, performing RNA biomarker analysis comprises comparing the cfRNA profile of the subject sample to the reference profile at the pre-determined timepoint and correlating an increase in interferon alpha / beta and gamma signaling as an early indication of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject. In some embodiments, quantifying distance to health comprises comparing the cfRNA profile in the subject sample to the reference profile using either a correlation in RNA abundances or a mahalanobis distance of the cell type of origin and correlating an increase in distance to a prediction of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in thesubject. In some embodiments, the profiling cell types of origin comprises comparing the cfRNA profile of leukocyte, lymphocyte, neutrophil, dendritic cell, and / or endothelial cell in the subject sample to the reference profile and correlating an increase in cfRNA abundance of leukocyte, lymphocyte, neutrophil, dendritic cell, and / or endothelial cell cfRNA to a prediction of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject.

[0015] In some embodiments, the panel of biomarkers comprises one, more (including, e.g., at least 30, at least 50, at least 75, or at least 100), or all of genes selected from the group consisting of: NIPAL3, MPO, CD9, IDS, TMEM159, XYLT2, OSBPL5, APBA2, EDC4, SNX29, VPS13D, HIPK2, IDI1, PDK3, IFI35, INPP5A, GNB5, DGCR2, NDST1, ATP6AP1, FBXW11, IRAG1, ATP2A3, SCARF1, DGKD, OSTM1, CDC14B, SSH1, TXLNA, KIF3C, EPS15, B4GALT1, OAS1, CDIP1, SLC9A1, STRN4, PITPNM2, CRAT, ACOT7, KCNK6, PSMD8, MARCHF2, POLR2E, SNUB, MYH9, MLC1, CTSG, NIN, PYGB, PSMA7, CDS2, CST3, UBL4A, SLC25A15, USP10, USP31, H0MER2, BPNT2, PPP6R1, PDCD5, PTPRS, DDX49, CLIP2, TRIM14, DNM1, VSIR, ARHGAP21, FBXL20, FTSJ3, ALOX12, ABCC3, GRPEL1, ZBTB16, PANXI, VWF, MLEC, OAS3, OAS2, EIF2B1, CCND3, FBRSL1, DBN1, CISH, ABHD14B, KANSL3, LOXL3, ODC1, WIPF1, PLEK, NIDI, GBP1, Clorfl98, IFIT3, IFIT2, ACAT2, SNX19, PLXDC2, TNFSF10, ZMIZ2, DDX39A, NMI, MRS2, TMX4, DSTN, MMP24OS, SIN3B, VIL1, GNAZ, TPST2, TST, KRI1, SLC44A2, CDKN2D, BCL2L2, CLIP1, SLC35D2, PRR7, SLC6A6, DIAPH1, FMO5, PRKAB2, RAFI, DLG4, CTNNBL1, BTBD2, PDZD2, CTIF, IRAK2, SORT1, UBAC2, OASL, CD36, CCT7, KIAA0513, APPL2, TTYH3, TLN1, TPMT, HADH, PCSK6, KIFC3, EIF4A3, CERS2, PYCR2, VGLL4, CTDSPL, TCTA, PLAC8, TIFA, FBXL17, SNX12, DOCK5, NFIB, ENDOD1, FERMT3, MTMR12, NBAS, GBP5, PSD3, BRAE, SKI, GPAT4, FUR, FBXW5, PTGIR, KALRN, SLC37A1, GAB3, PCSK7, ORAI2, MGAT4B, EMC10, SELENON, H3-3A, COL6A3, AIMP1, ANKRD33B, CITED2, GALNT10, CRYL1, TTC7B, CMTM5, WBP1L, G0LM2, RBPMS2, SAMD14, FN3K, TPM4, GATAD2A, ANKRD11, TRAPPC9, SPINT2, ECU, MAP4K2, MLKL, CD2BP2, GP9, REPS2, CHD3, CRTAP, RGS19, FIBP, DCAKD, ABLIM3, GPR137, ATP2A2, RAB1B, PACS1, PPP2R2D, MCTP1, TOM1L2, RUFY1, SAMD9L, NR2C2, CD151, NDUFAF3, GP5, PACS2, TMEM64, COA4, UBA7, M0B2, ADI1, RBM10, EWSR1,PTTG1IP, CLDN5, SNN, BTBD6, SLC24A3, ZBTB37, METTL7A, IRF7, BRCC3, IFIT1 , RASA3, CD300E, PDE2A, TPCN1, ISG15, PEAR1, PLSCR1, PARVB, HACD4, TMEM63A, VK0RC1L1, MAML3, PDLIM7, FLNA, SRC, AD ARBI, HTT, MAP3K5, MY05A, PSAP, MYO1C, PAPSS2, UNC13B, CDC42BPB, PLXNB3, GRK5, AGP ATI, LY6G6F, HLA-C, HLA-E, PRR13, DOK6, HLA-H, HLA-A, RNF208, FGD5-AS1, FAM30A, GPX1, HLA-B, LINC00674, PEDS1, EIF6, CCDC71L, MTRNR2L8, ITGB3, RP11-4OE2, IKBKG, RN7SL1, and RP11-147L13.1 E In some embodiments, the panel of biomarkers comprises one, more, or all of RAFI, RFC2, ERCC6L2, IDI1, FIBP, EDC4, FBXL20, KANSL1, and POP4. In some embodiments, the panel of biomarkers comprises at least one of RAFI, IDI1, EDC4, KANSL1, RFC2, POP4, and ERCC6L2.

[0016] In some embodiments, the panel of biomarkers comprises one, more (including, e.g., at least 30, at least 50, at least 75, or at least 100), or all of genes selected from the group consisting of: RHD, CCL21, TMEM176B, CTD-2643I7.4, RP11-329A14.1, ESPN, CLEC10A, MTCO3P12, CD8B, HTATSF1P2, KCNQ1OT1, Y RNA, MSC, ANKRD36BP2, EVPL, TNNT1, TSPAN2, MYH11, ATOH8, TEAD4, LINS1, BRSK1, Y RNA, ZBED6, TMEM205, RNF182, MTRNR2L1, PVALB, RP11-678G14.3, ALB, RNA5SP145, MY0M2, RPL9P9, MYCL, RNA5SP149, CDC42EP1, ABO, SPOCD1, GALNT3, DDX11L10, PCDH17, RP11- 452G18.2, DNAJB5, SMIM24, MS4A1, IGHA1, MTND4P12, DDX11L2, HLA-J, Y RNA, Y RNA, C19orf33, MT-RNR2, USP54, GALT, LINC00504, ARL6IP6, MTRNR2L8, RN7SL767P, PDZK1IP1, KRT1, COL4A1, PROK2, SLC15A2, H2BC8, Y RNA, SBK1, DEPPI, BAHCC1, HBG2, CXCL9, TOX2, RALY-AS1, SNAI2, ALPK3, STAP1, ERAP2, PLA2G7, GSTM5, CD79A, SPRY4, IGHD, MYO10, AC114271.2, PARVA, CYP2E1, CCR2, STK33, IGHV1-3, XXbac-BPG299F 13.17, ST20-AS1, Y RNA, CFD, RP11-712B9.2, LILRA4, MEIS3P1, SRGAP1, HID1, and REEPL In some embodiments, the panel of biomarkers comprises one, more (including, e.g., at least 30, at least 50, at least 75, or at least 100), or all of genes selected from the group consisting of: RHD, CCL21, TMEM176B, CTD-2643I7.4, RP11- 329A14.1, ESPN, CLEC10A, MTCO3P12, CD8B, HTATSF1P2, KCNQ1OT1, Y RNA, MSC, ANKRD36BP2, EVPL, MYH11, ATOH8, TEAD4, Y RNA, ZBED6, RNF182, MTRNR2L1, RNA5SP145, RPL9P9, MYCL, RNA5SP149, CDC42EP1, SPOCD1, GALNT3, DDX11L10, PCDH17, DNAJB5, MTND4P12, C19orf33, MT-RNR2, PDZK1IP1, KRT1, SLC15A2, DEPPI,BAHCC1, ERAP2, IGHD, MYO 10, CFD, RP1 1-712B9.2, SRGAP1, RNF152, RN7SL736P, RN7SL665P, and TRIP13. In some embodiments, the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: RHD, CCL21, TMEM176B, CTD- 264317.4, RP11-329A14.1, ESPN, CLEC10A, MTCO3P12, HTATSF1P2, KCNQ1OT1, ANKRD36BP2, RNF182, and RNA5SP149. In some embodiments, the panel of biomarkers comprises RHD and CCL21.

[0017] In some embodiments, the subject is receiving immunosuppression treatment and will be tapered off the immunosuppression treatment, and the panel is predictive of immunosuppression treatment tapering failure.

[0018] In some embodiments, the panel of biomarkers comprises one, more, or all of Major Histocompatibility Complex, Class I - E, C, A, and B (HLA-E, HLA-C, HLA-A, and HLA-B, respectively) genes.

[0019] In some embodiments, the panel of biomarkers comprises at least one histone gene. In some embodiments, the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: Hl-1, Hl-2, Hl-3, Hl-4, Hl-5, Hl-6, H2AC1, H2AC4, H2AC6, H2AC7, H2AC8, H2AC11, H2AC12, H2AC13, H2AC14, H2AC15, H2AC16, H2AC17, H2AC18, H2AC19, H2AC20, H2AC21, H2AC25, H2BC1, H2BC3, H2BC4, H2BC5, H2BC6, H2BC7, H2BC8, H2BC9, H2BC10, H2BC11, H2BC12, H2BC13, H2BC14, H2BC15, H2BC17, H2BC18, H2BC21, H2BC26, H2BC12L, H3C1, H3C2, H3C3, H3C4, H3C6, H3C7, H3C8, H3C10, H3C11, H3C12, H3C13, H3C14, H3C15, H3-4, H4C1, H4C2, H4C3, H4C4, H4C5, H4C6, H4C7, H4C8, H4C9, H4C11, H4C12, H4C13, H4C14, H4C1 5H1-0, Hl-7, Hl-8, Hl-10, H2AZ1, H2AZ2, MACROH2A1, MACROH2A2, H2AX, H2AJ, H2AB1, H2AB2, H2AB3, H2AP, H2AL1Q, H2AL3, H2BK1, H2BW1, H2BW2, H2BW3P, H2BN1, H3-3A, H3-3B, H3-5, H3-7, H3Y1, H3Y2, CENPA, and H4C16. In some embodiments, the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: Hl-1, Hl-2, Hl- 3, Hl-4, Hl-5, Hl-6, H2AC1, H2AC4, H2AC6, H2AC7, H2AC8, H2AC11, H2AC12, H2AC13, H2AC14, H2AC15, H2AC16, H2AC17, H2AC18, H2AC19, H2AC20, H2AC21, H2AC25, H2BC1, H2BC3, H2BC4, H2BC5, H2BC6, H2BC7, H2BC8, H2BC9, H2BC10, H2BC11, H2BC12, H2BC13, H2BC14, H2BC15, H2BC17, H2BC18, H2BC21, H2BC26, H2BC12L, H3C1, H3C2, H3C3, H3C4, H3C6, H3C7, H3C8, H3C10, H3C11, H3C12, H3C13, H3C14,H3C15, H3-4, H4C1, H4C2, H4C3, H4C4, H4C5, H4C6, H4C7, H4C8, H4C9, H4C1 1, H4C12, H4C13, H4C14, and H4C15. In some embodiments, the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of H2BC21, Hl-2, H2BC4, H2AC6, Hl-4, H2BC12, H3C10, and Hl-5.

[0020] In some embodiments, the panel of biomarkers comprises one, more (including, e.g., at least 30, at least 50, at least 75, or at least 90), or all of genes selected from the group consisting of: RPA2, CLSPN, CDCA8, CDC20, KIF2C, NASP, USP1, PSRC1, ANP32E, CKS1B, NUF2, NEK2, DTL, CENPF, LBR, EXO1, RRM2, CENPA, MSH2, BUB1, CKAP2L, MCM6, CDCA7, HJURP, MCM2, SMC4, ECT2, SLBP, TACC3, CENPE, HMGB2, CENPU, CDC25C, HMMR, GMNN, TTK, CASP8AP2, ANLN, RFC2, CDCA2, MCM4, CCNE2, DSCC1, ATAD2, CKS2, TUBB4B, CDK1, KIF20B, KIF11, HELLS, MKI67, RRM1, E2F8, CKAP5, FEN1, POLD3, RAD51AP1, NCAPD2, CDCA3, CBX5, PRIM1, TMPC), GAS2L3, UNG, CKAP2, G2E3, DLGAP5, UBR7, RAD51, NUSAP1, WDR76, CCNB2, TIPIN, KIF23, BLM, CTCF, GINS2, PIMREG, AURKB, CDC6, TOP2A, BRIP1, JPT1, BIRC5, TYMS, NDC80, UHRF1, PCNA, TPX2, UBE2C, AURKA, CHAF1B, CDC45, MCM5, RANGAP1, GTSE1, and POLA1. In some embodiments, the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: NASP, USP1, ANP32E, CENPF, RRM2, MCM6, SMC4, SLBP, TACC3, CENPE, HMGB2, ATAD2, TUBB4B, KIF20B, MKI67, CKAP5, NCAPD2, CBX5, TMPC, CKAP2, CTCF, TOP2A, PCNA, TPX2, and RANGAPL In some embodiments, the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: ANP32E, CENPF, HMGB2, TUBB4B, and MKI67.

[0021] An aspect of the current disclosure is directed to a method for monitoring a microbiome of an HSCT recipient subject for infection, the method comprising: a) obtaining cell-free RNA from at least one biological sample of the subject, wherein the biological sample is obtained at the day of infusion or less than 36 months, or 24 months, or 12 months, or less than 6 months, or less than 3 months, or less than 2 months after the day of infusion, wherein the biological samples are plasma samples or serum samples; b) determining a profile of a panel of metagenomic biomarkers in the cell-free RNA; c) determining an inflammatory disease status in the subject based on the profile of the panel of metagenomic biomarkers; andd) removing background total viral counts of all non-mammalian infecting viral families from each sample.

[0022] In some embodiments, the at least one biological sample of the subject comprises a set of longitudinally collected biological samples of the subject at preconditioning, day of infusion, engraftment, and / or one or more post infusion timepoints. In some embodiments, the step of determining a profile of a panel of biomarkers comprises performing nucleotide sequencing of the cell-free RNA. In some embodiments, the step of determining a profile of a panel of biomarkers comprises measuring the levels of RNAs corresponding to the biomarkers in the panel. In some embodiments, monitoring the microbiome of the subject based on the profile of the panel of biomarkers comprises comparing the profile of the panel of biomarkers in the cell free RNA from the biological sample to a reference profile of the panel of biomarkers. In some embodiments, the reference profile of the panel of biomarkers is from a healthy donor or an HSCT subject without complications. In some embodiments, the panel of biomarkers comprises those for at least one of the Flaviviridae, Anelloviridae, Retroviridae, and Herpesviridae viral families. In some embodiments, the abundance of prokaryotic reads are quantified for the cfRNA from one or a combination of longitudinal time points and compared to the reference abundance of prokaryotic reads.

[0023] In some embodiments, a correlation is made when the subject sample comprises elevated levels of biomarkers when compared to the reference panel that the subject may be experiencing complications.

[0024] In some embodiments, at least two cfRNA measurements are used in determining the subject’s recovery trajectory and wherein the at least two cfRNA measurements used consist of the parameters included in the group consisting of microbial kingdom counts, genes counts, viral family counts, distance to health, and cell turnover. In some embodiments, the cfRNA measurements used consist of microbial kingdom counts, genes counts, viral family counts, distance to health, and cell turnover.

[0025] In some embodiments, the set of longitudinally collected biological samples includes more than one predetermined timepoints after infusion. In some embodiments, the set of longitudinally collected biological samples comprises a biological sample collected on the day of infusion. In some embodiments, the more than one predetermined timepoint comprises monthlyintervals after the day of infusion, i.e. at 1 Month, IM; at 2 Month, 2M; at 3 Month, 3M; and / or at 6 Month, 6M). In some embodiments, the profile of biomarkers at an early time point (e.g., 1, 2, or 3 months) post infusion is indicative of the likelihood of GVHD.

[0026] In some embodiments, the biological sample is a plasma sample. In some embodiments, the biological sample is a serum sample.

[0027] In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.

[0028] In some embodiments, the profile of the panel of biomarkers comprises subject-specific gene expression of the biomarkers in the panel. In some embodiments, the cell-free RNA is substantially free of contaminant RNA.

[0029] In some embodiments, the subject’s recovery trajectory is used in determining treatment for the subject. In some embodiments, test performance of the trained classification model is measured through one or more of accuracy, sensitivity, specificity, and / or area under the receiver operating characteristic curve (ROC-AUC). In some embodiments, the inflammatory syndrome status is determined at a greater than 80%, 85%, 90%, 91%, 92%, 93%, 94%, or 95% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 80%, 85%, 90%, 91%, 92%, 93%, 94%, or 95% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, or 95% specificity. In some embodiments, the trained classification model comprises one or more models selected from the group consisting of generalized linear models with Ridge and LASSO feature selection (GLMNETRIDGE and GLMNETLASSO), support vector machines with linear and radial basis function kernel (SVMLin and SVMRAD), random forest (RF), random forest ExtraTrees (EXTRATREES), neural networks (NNET), linear discriminant analysis (LDA), nearest shrunken centroids (PAM), C5.0 (C5), k-nearest neighbors (KNN), naive bayes (NB), CART (RPART), generalized linear model (GLM), and greedy forward search algorithm (GFS).

[0030] In some embodiments, the method further comprises treating the subject based on the recovery trajectory. In some embodiments, the treatment for aGVHD comprises administering an immunosuppressive to the subject. In some embodiments, the treatment for complications due to HSCT comprises administering treatment regimens known to combat the complication.BRIEF DESCRIPTION OF DRAWINGS

[0031] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the office upon request and payment of the necessary fee.

[0032] FIG. 1A-D. Shows an overview of the sample cohort. (A) Sample collection overview. (B) Day relative to transplantation samples were collected for each timepoint. (C) Timing of acute GVHD (aGVHD) and chronic GVHD (cGVHD) relative to transplantation. (D) Correlation of cfRNA abundance to healthy stem cell donors of samples from noCOMP and COMP patients.

[0033] FIG. 2A-D. (A) Upset plot of complications. (B) Timing of complications relative to transplantation. (C) Distance of cfRNA cell type of origin profiles to healthy stem cell donors of samples from noCOMP and COMP patients (Methods). (D) Within group average correlations of noCOMP and COMP patients across the sampling time course (Pearson correlation). Selfcomparisons were removed prior to average calculation.

[0034] FIG. 3A-D. cfRNA origins during and after HSCT. (A) Stacked area plot of mean fraction of top ten most abundant cell types across the time course. (B) Complete blood cell counts and cfRNA cell type of origin fraction for paired samples across the treatment time course. Samples are colored by grouping: those that did not experience a complication in the first 9 months (noCOMP) and those that did (COMP). (C) Scatterplots of cfRNA CTO estimates and complete blood cell counts with accompanying histograms for each axis for neutrophils, monocytes, and lymphocytes. (D) Counts per million reads assigned to the Anelloviridae and Pegivirus in noCOMP, COMP, and healthy patients.

[0035] FIG. 4. Assigned read counts of viral genera Betacoronavirus (coronaviruses), Betapolyomavirus (BKV), Cytomegalovirus (CMV), Lymphocryptovirus (EBV), and Mastadenovirus (adenoviruses).

[0036] FIG. 5A-G. cfRNA signature prior to aGVHD diagnosis. (A) Correlation of cfRNA abundance to healthy stem cell donors, (B) leukocyte fraction, and (C) cell type of origin diversity of samples from patients who develop aGVHD and those who do not. Samples from prior to aGVHD diagnosis. (D) Volcano plot of the results from the differential abundance analysis comparing engraftment time point samples from patients who develop aGVHD andthose who do not (DESeq2, BH-adjusted p-values). (E) Variance stabilization transformation values of significantly differentially abundant genes at the engraftment time point of samples from patients that develop aGVHD and those that do not develop (DESeq2, BH adjusted p-value < 0.1, AUC-ROC > 0.7). Samples and genes are clustered based on correlation. (F) Enriched pathways of significantly differentially abundant RNAs at the engraftment time point of samples from patients that develop aGVHD versus those that do not develop aGVHD. (G) Major Histocompatibility Complex, Class I - E, C, A, and B (HLA-E, HLA-C, HLA-A, and HLA-B transcript counts per million relative to days to aGVHD diagnosis in samples at the engraftment, one-month, and two-month time points.

[0037] FIG. 6. Platelet fraction of samples from patients who develop aGVHD and those who do not. Samples from prior to aGVHD diagnosis.

[0038] FIG. 7A-C. cfRNA signature prior to cGVHD / L-aGVHD diagnosis. (A) Volcano plot of the results from the differential abundance analysis comparing engraftment time point samples from patients who develop aGVHD and those who do not (DESeq2, BH-adjusted p-values). (B) ROC-AUC curves of the training and test set of the median performing model. (C) The top 15 most selected transcripts of the 101 repeated stratified train-test split modeling.

[0039] FIG. 8. cfRNA reflects GVHD organ involvement. Maximum cell type of origin measurement for Intrahepatic cholangiocyte, hepatocyte, and salivary gland cells for samples split by GVHD status and organ involvement.

[0040] FIG. 9A-H. Immunosuppression Taper. (A) Clinical cohort overview. (B) Days of taper start relative to HSCT and failure relative to taper start. (C) cfRNA CTO measurements of mature conventional dendritic cell and (D) endothelial cell at the start, end, and failure of taper, along with temporal dynamics in patients with more than one sample. (E) Volcano plot of the results from the differential abundance analysis comparing engraftment time point samples from patients who failed immunosuppression taper and those that did not (DESeq2, BH-adjusted p- values). (F) Longitudinal CPM values of TCF7 and MYOF in patients who fail immunosuppression taper and those that do not. (G) ROC-AUC curves of the training and test set of the median performing model. (H) The top 15 most selected transcripts of the 101 repeated stratified train-test split modeling.

[0041] FTG. 10A -E. Shared signatures between cohorts. (A) The number of models that selected RHD and CCL21 from FIG. 7C and FIG. 9H. (B) Transcript CPM of RHD and CCL21 in DFCI and MCC cohorts. (C) Values of CCL2HRHD CPM loglO scaled in DFCI and MCC cohorts. (D) ROC curves of the CCL2HRHD CPM values from MCC and DFCI. (E) KM survivorship curves of CCL21IRHD CPM values split into high and low groups based on medians separately calculated for each cohort.

[0042] FIG. 11A-B. Measuring immune cell turnover. (A) Paired cell counts, cfRNA cell type of origin fractions, and estimated turnover in arbitrary units for paired samples. (B) Median immune cell turnover measurements from samples of patients with cGVHD, aGVHD, or non- GVHD where all samples were obtained prior to GVHD diagnosis. Dot size and color represent median cell turnover in arbitrary units (a.u.).

[0043] FIG. 12A-E. Metagenomic dynamics of cfRNA in HSCT (A) Family count distribution of all viral reads assigned to healthy stem cell donors and schematic of background removal. (B) Counts per million background removed viral counts in noCOMP, COMP, and healthy patients.(C) Counts per million reads assigned to the Flaviviridae, Anelloviridae , Herpesviridae, and Retroviridae viral families in noCOMP, COMP, and healthy patients. (D) Counts per million reads assigned prokaryotic in origin. (E) Fraction of randomly sampled bacterial reads from each sample that aligned to a prokaryotic rRNA operon reference. Asterisks indicate statistical significance by Mann-Whitney U test unless otherwise noted using Benjamini -Hochberg adjusted p-values as follows: ns, non-significant; *, p < 0.05; **, p < 0.01; ***, p < 0.001; ****, p < 0.001.

[0044] FIG. 13A-D. Cell cycle transcript dynamics. (A) Sum of cell cycle transcript counts per million z-scores in samples from patients with and without complications over time. (B) Sum of cell cycle transcript counts per million z-scores in samples collected at Engraftment from patients who go on to develop aGVHD, patients who do not aGVHD, and healthy donors (HD). (C) Sum of histone transcript counts per million values in samples from patients with and without complications over time. (D) Sum of histone transcript counts per million values in samples collected at Engraftment from patients who go on to develop aGVHD, patients who do not aGVHD, and healthy donors.DETAILED DESCRIPTION

[0045] The current disclosure describes plasma cfRNA profiling performed on longitudinal samples from allogeneic HSCT recipients and samples from healthy stem cell donors using RNA-sequencing. cfRNA profiles from patients which are complication-free for 9 months after transplantation are similar, whereas cfRNA profiles of patients with recoveries complicated by frequent post allogeneic HSCT complications are all unique and very different from those of complication-free patients. Using multiple analysis techniques, cfRNA is shown to reflect immune reconstitution, viral reactivation, immune activation and solid tissue damage related to aGVHD and cGVHD, and dynamics of immunosuppression taper. This disclosure introduces new modalities to analyze cfRNA and demonstrates cfRNA as a highly versatile and informative analyte to monitor the clinical course of allogeneic HSCT recipients.

[0046] Although claimed subject matter will be described in terms of certain examples, other examples, including examples that do not provide all the benefits and features set forth herein, are also within the scope of this disclosure. Various structural, logical, and process step changes may be made without departing from the scope of the disclosure.

[0047] Ranges of values are disclosed herein. The ranges set out a lower limit value and an upper limit value. Unless otherwise stated, the ranges include the lower limit value, the upper limit value, and all values between the lower limit value and the upper limit value, including, but not limited to, all values to the magnitude of the smallest value (either the lower limit value or the upper limit value).

[0048] In the description that follows, certain conventions will be followed as regards to the usage of terminology. Generally, terms used herein are intended to be interpreted consistently with the meaning of those terms as they are known to those of skill in the art. In practicing the present disclosure, many conventional techniques in molecular biology, microbiology, cell biology, biochemistry, and immunology are used, which are within the skill of the art. These techniques are described in greater detail in, for example, Molecular Cloning: a Laboratory Manual 4th edition, J.F. Sambrook and D.W. Russell, ed. Cold Spring Harbor Laboratory Press 2012; Recombinant Antibodies for Immunotherapy, Melvyn Little, ed. Cambridge University Press 2009; “Oligonucleotide Synthesis” (M. J. Gait, ed., 1984); “Animal Cell Culture” (R. I. Freshney, ed., 1987); “Methods in Enzymology” (Academic Press, Inc.); “Current Protocols inMolecular Biology” (F. M. Ausubel et al., eds., 1987, and periodic updates); “PCR: The Polymerase Chain Reaction”, (Mullis et al., ed., 1994); “A Practical Guide to Molecular Cloning” (Perbal Bernard V., 1988); “Phage Display: A Laboratory Manual” (Barbas et al., 2001). The contents of these references and other references containing standard protocols, widely known to and relied upon by those of skill in the art, including manufacturers’ instructions are hereby incorporated by reference as part of the disclosure.

[0049] The term “biological sample” includes body samples from an animal, including biological fluids such as serum, plasma, vitreous fluid, lymph fluid, synovial fluid, follicular fluid, seminal fluid, amniotic fluid, milk, whole blood, urine, cerebro-spinal fluid, saliva, sputum, tears, perspiration, mucus, and tissue culture medium, as well as tissue extracts such as homogenized tissue, and cellular extracts. In some embodiments, the biological sample is a serum, plasma or urine sample. In some embodiments, the biological sample is a plasma sample. In some embodiments, the biological sample is a serum sample.

[0050] The term “subject” refers to mammals. Non-limiting examples of mammals include but are not limited to human, horse, camel, dog, cat, pig, cow, goat, and sheep. In some embodiments, the mammal is human.

[0051] The terms “cell-free RNA” or “cfRNA” refer to cfRNA released by cells in both the blood compartment and from vascularized solid tissues. cfRNA contains information about systemic immune dynamics and immune-tissue interactions. cfRNA is primarily released by dying cells; therefore, cfRNA may provide insights into pathways of cell death and mechanisms of cellular injury. cfRNA is also released into the blood by way of active secretion, cells release cfRNA to the subject’s bodily fluid and thus may increase the quantity of the specific cfRNA in the subject’s biological sample as compared to a healthy individual. cfRNA may include any types of RNA that are circulating in the bodily fluid of a person without being enclosed in a cell body or a nucleus.

[0052] A “gene panel” or “panel” refers to a defined selection of genes that are enriched or sequenced. A panel will include relevant pathogen associated genes and likely variants of selected genes. In some embodiments, the panel comprises a collection of signature genes. As used herein, the term “signature gene” refers to a gene whose expression is correlated, either positively or negatively, with disease extent or outcome or with another predictor of diseaseextent or outcome. In some embodiment, the gene panels are used in sequencing assays. Sequencing based assays are known in the art. Some non-limiting examples of sequencing based assays are next generation sequencing (NGS), Sanger sequencing, oxidative bisulfite sequencing, direct RNA sequencing, and others. In a next generation sequencing (NGS) panel test, only clinically important genes are examined to obtain genomic data in a timely and cost-effective manner. In some embodiments, the gene panels are used in NGS. In some embodiments, the gene panels are used in non-sequencing based assays. Non-sequencing based assays are known in the art. Non-limiting examples of non-sequencing based assays include microarrays, polymerase chain reaction (PCR), including real-time PCR, reverse transcription PCR (RT- PCR), quantitative reverse transcription PCR (RT-qPCR), quantitative PCR (qPCR), digital droplet PCR (ddPCR), and others (as described in Shemer, R. et al., Current Protocols in Molecular Biology, 127.1 (2019): e90; and Zemmour, Hai, et al., Nature Communications, 9.1 (2018): 1-9, both incorporated herein in their entirety)..

[0053] As used herein, the term “profile” generally refers to gene expression profile. A gene expression profile can be understood to mean a pattern of abundance of expression of genes. In some embodiments, a gene expression profile is a measurement of the expression of multiple genes at once. In some embodiments, gene expression is measured by cfRNA in a subject, e.g., cfRNA in a serum or plasma sample, in which case, the profile is that of a subject’s cfRNA. Gene expression profiles can be characteristic or unique to a status of a subject, e.g., a healthy status or a disease status (such as cancer or infectious disease) and can therefore be used to distinguish between different statuses. In some embodiments, the profile is a profile of genes identified in this disclosure to be associated with infectious or inflammatory syndromes, which are also referred to herein as signature genes. It is to be understood that the profile of signature genes in this disclosure have not all been associated with the infectious or inflammatory syndromes previously. Signature genes of this disclosure were identified using the methods developed and disclosed herein. Signature genes exhibit differential expression in subjects having the infectious or inflammatory syndromes relative to subjects without infectious or inflammatory syndromes, for example. The differential expression refers to the difference in abundance of a signature gene in subjects having infectious or inflammatory syndromes relative to subjects without infectious or inflammatory syndromes. In some embodiments, the expressionlevels of signature genes may be used to predict progression of infectious or inflammatory syndromes. A “signature nucleic acid” is a nucleic acid comprising or corresponding to, in case of cDNA, the complete or partial sequence of a RNA transcript encoded by a signature gene, or the complement of such complete or partial sequence. A signature protein is encoded by or corresponding to a signature gene of the disclosure.

[0054] As used herein, the term “reference profile” refers to the profile of genes (e.g., the same genes identified in this disclosure to be associated with an infectious or inflammatory disease state) in a subject not inflicted with an infectious or inflammatory disease state.

[0055] In some embodiments, machine learning and model training is performed using R (v4.1.3) with the DESeq2 (vl.34.0), Caret (v6.0.90), and pROC (vl.18.0) packages. In some embodiments, sample metadata and count matrices split 70 / 30 into a training set and a test set, controlling for disease status, HIV status, and cohort to minimize differences in the training and test datasets. In some embodiments, sample metadata and count matrices are split into a training set, validation set, and test set, controlling for disease status, HIV status, and cohort to minimize differences in the datasets.

[0056] In some embodiments, features for model training are selected by filtering and differential abundance analysis. First, differential abundance analysis was performed on the raw training counts using DESeq2. Then feature selection was conducted using the output from DESeq2 on retained genes. In some embodiments, genes are excluded that have a base-mean of less than 100 and a Benjamini-Hochberg adjusted p-value greater than 0.05. In some embodiments, genes are excluded that have a base-mean of less than 50 and a Benjamini- Hochberg adjusted p-value greater than 0.05.

[0057] In some embodiments, machine learning algorithms are trained using 5-fold cross validation and grid search hyperparameter tuning. In some embodiments, accuracy, sensitivity, specificity, and / or area under the receiver operating characteristic curve (ROC-AUC) are used to measure test performance. In some embodiments, the classification models used are generalized linear models with Ridge and LASSO feature selection (GLMNETRIDGE and GLMNETLASSO), support vector machines with linear and radial basis function kernel (SVMLin and SVMRAD), random forest (RF), random forest ExtraTrees (EXTRATREES), neural networks (NNET), linear discriminant analysis (LDA), nearest shrunken centroids (PAM),C5.0 (C5), k-nearest neighbors (KNN), naive bayes (NB), CART (RPART), and generalized linear models (GLM).

[0058] In one aspect, the current disclosure is directed to a method of monitoring a subject undergoing HSCT, the method comprising: a) obtaining cell-free RNA from at least one biological sample of the subject, wherein the biological sample is a plasma sample or serum sample; b) determining: i) a profde of a panel of biomarkers in the cell-free RNA; ii) cfRNA abundances; iii) distance to health; iv) cell type of origin of the cfRNA; and / or v) cell turnover; and c) determining the subject’s health status based on at least one of the modalities listed in b.

[0059] In some embodiments, the at least one biological sample of the subject is obtained at the day of infusion. In some embodiments, the at least one biological sample of the subject is obtained after the day of infusion. In some embodiments, the at least one biological sample of the subject is obtained less than 36 months after the day of infusion. In some embodiments, the at least one biological sample of the subject is obtained less than 2 months after the day of infusion. In some embodiments, the at least one biological sample of the subject is obtained less than 3 months after the day of infusion. In some embodiments, the at least one biological sample of the subject is obtained less than 6 months after the day of infusion. In some embodiments, the at least one biological sample of the subject is obtained less than 12 months after the day of infusion.

[0060] In some embodiments, the at least one biological sample of the subject comprises a set of longitudinally collected biological samples of the subject, wherein the samples are obtained at preconditioning, day of infusion, engraftment, and / or one or more post engraftment timepoints. In some embodiments, the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint being at less than 2 months after the day of infusion.. In some embodiments, the one or more post engraftment timepoints comprise multiple timepoints,and at least one timepoint being at less than 3 months after the day of infusion. In some embodiments, the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint being at less than 6 months after the day of infusion. In some embodiments, the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint being at less than 12 months after the day of infusion. In some embodiments, the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint being at less than 24 months after the day of infusion. In some embodiments, the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint being at less than 36 months after the day of infusion.

[0061] In the current disclosure, cfRNA dynamics were measured throughout the allogeneic HSCT treatment course. This approach introduced multiple novel challenges, as HSCT recipients can experience diverse complications that often co-occur at any point during the treatment course. Unlike previous cfRNA research, which primarily focused on diagnosing a specific condition, the present study demanded a more comprehensive and longitudinal approach to capture the intricate interplay of immune reconstitution, solid organ damage, and microbial dynamics. Quantification of “distance-to-normal” was used to deal with challenges of highly heterogeneous sample groups. As used herein, “distance to normal” refers to the Average Pearson correlation between the subject donor that has undergone HSCT and healthy stem cell donor sample. Distance-to-normal is calculated through two metrics: 1) the average correlation of transcript counts and 2) distance of cell type of origin profiles of each plasma cfRNA profile to those of the healthy stem cell donors. This measurement revealed that patients free of complications follow highly similar trajectories approaching normal, while patients who did experience complications follow unique, individualized trajectories. These observations follow the Anna Karenina principle in statistics and provide a theoretical framework for addressing similarly heterogeneous patient groups (W. B. Goh, L. Wong, Trends Biotechnol. 36, 488-498 (2018).

[0062] As used herein, “health status” refers to a status of a subject, e.g., a healthy status or a disease status (such as GVHD or infectious disease) and can therefore be used to distinguish between different statuses. Gene expression profiles can be characteristic or unique to the status of a subject, e.g., a healthy status or a disease status (such as cancer or infectious disease) andcan therefore be used to distinguish between different statuses. In some embodiments, a health status comprises healthy or disease free status, a latent disease status, an active disease status, or a progressing disease status. In some embodiments, the determining the subject’s health status based on the profile of the panel of biomarkers comprises comparing the profile of the panel of biomarkers in the cell free RNA from the biological samples to a respective reference profile of the panel of biomarkers at the respective time points. In some embodiments, the subject’s health status comprises treatment response, treatment toxicity, likelihood of developing graft versus host disease (GVHD), and / or likelihood of developing complications due to the HSCT.

[0063] In some embodiments, the subject is receiving immunosuppression treatment and will be tapered off the immunosuppression treatment at an initial tapering date, the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint occurs prior to the initial tapering date, less than three months after the initial tapering date, and / or less than six months after the initial tapering date.

[0064] “Immunosuppressive treatment” or “immunosuppression therapy” refers to a drug regimen given to subjects to lower their bodies’ immune response, so the immune system does not damage transplanted organs and tissues. Immunosuppression treatment is known in the art. Tapering immunosuppressive therapy refers to gradually reducing the dose of an immunosuppressive therapy over time until the subject is no longer receiving immunosuppressive therapy, rather than stopping the therapy suddenly. Tapering can occur over weeks, months, or years. Some subjects are never completely taken off immunosuppressants after a transplant; instead, immunosuppressants are slowly reduced over time.Immunosuppressant taper schedules are determined by the attending physician.Immunosuppression tapering is decided on a case-by-case basis by the attending physician. As used herein, “immunosuppression taper failure” refers to a situation where a subject that has undergone a transplant or autoimmune condition starts to gradually reduce their immunosuppressive medication dosage (tapering), but their body begins to reject the transplanted organ or experiences a flare-up of their autoimmune disease, indicating that the immune system is becoming too active due to the reduced medication level; essentially, the taper was not successful in maintaining adequate immune suppression.

[0065] In some embodiments, the samples are obtained from the subject at a time relative to the start of immunosuppression therapy tapering, i.e. the first day that the therapy is reduced or initial tapering date. In some embodiments, the samples are obtained from the subject on the last day of taper, or the day when the immunosuppressant regimen is completely removed, i.e. the day after the last dose of immunosuppressant is taken by the subject. In some embodiments, the subject is receiving immunosuppression treatment and will be tapered off the immunosuppression treatment at the initial tapering date, the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint occurs prior to the initial tapering date. In some embodiments, the at least one timepoint occurs less than three months after the initial tapering date. In some embodiments, the at least one timepoint occurs less than six months after the initial tapering date. In some embodiments, the at least one timepoint occurs on the first day of immunosuppression taper. In some embodiments, the at least one timepoint occurs on the last day of taper. In some embodiments, the at least one timepoint occurs on the day of taper failure. In some embodiments, the at least one timepoint occurs at two weeks after the end of taper. In some embodiments, the at least one timepoint occurs one month after the end of taper. In some embodiments, the at least one timepoint occurs three months after the end of taper. In some embodiments, the at least one timepoint occurs six months after the end of taper. In some embodiments, the multipole timepoints occur prior to the initial tapering date and less than three months after the initial tapering date. In some embodiments, the multipole timepoints occur prior to the initial tapering date and less than six months after the initial tapering date. In some embodiments, the multipole timepoints occur less than three months after the initial tapering date and between 3 and 6 months after the initial tapering date. In some embodiments, the multiple timepoints occur in some combination of on the first day of immunosuppression taper, the last day of taper, on the day of taper failure, and at two weeks, one month, three months, and six months after the end of taper. For example, in some embodiments, the multiple timepoints occur on the first day of immunosuppression taper and the last day of taper. In some embodiments, the multiple timepoints occur on the first day of immunosuppression taper, the last day of taper, and on the day of taper failure. In some embodiments, the multiple timepoints occur on the first day of immunosuppression taper, the last day of taper, and at two weeks after the end of taper. In some embodiments, the multiple timepoints occur on the first day ofimmunosuppression taper, the last day of taper, and at one month after the end of taper. Tn some embodiments, the multiple timepoints occur on the first day of immunosuppression taper, the last day of taper, and three months after the end of taper. In some embodiments, the multiple timepoints occur on the first day of immunosuppression taper, the last day of taper, and at six months after the end of taper. In some embodiments, the multiple timepoints occur on the last day of taper and on the day of taper failure. In some embodiments, the multiple timepoints occur on the last day of taper and at two weeks after the end of taper. In some embodiments, the multiple timepoints occur on the last day of taper and at one month after the end of taper. In some embodiments, the multiple timepoints occur on the last day of taper and three months after the end of taper. In some embodiments, the multiple timepoints occur on the last day of taper and at six months after the end of taper. In some embodiments, the multiple timepoints occur on the first day of immunosuppression taper and at two weeks after the end of taper. In some embodiments, the multiple timepoints occur on the first day of immunosuppression taper and at one month after the end of taper. In some embodiments, the multiple timepoints occur on the first day of immunosuppression taper and three months after the end of taper. In some embodiments, the multiple timepoints occur on the first day of immunosuppression taper and at six months after the end of taper.

[0066] In some embodiments, the step of determining a profile of a panel of biomarkers comprises performing nucleotide sequencing of the cell-free RNA. Methods of sequencing are described in Nucleic acid biomarkers of immune response and cell and tissue damage in children with COVID-19 and MIS-C. Cell Rep. Med. 4, 101034 (2023), which is incorporated herein by reference in its entirety.

[0067] In some embodiments, the step of determining a profile of a panel of biomarkers comprises measuring the levels of RNAs corresponding to the biomarkers in the panel.

[0068] In some embodiments, the complications due to the HSCT comprise microbial / bacterial infection and / or viral infection. In some embodiments, determining the subject’s health status treatment response, further comprises calculating average correlation of cfRNA abundances for each sample with respect to reference profile of the panel of biomarkers. In some embodiments, the subject’s recovery health status further comprises the cell turnover values for neutrophils, leukocytes, and lymphocytes using the cfRNA cell type of origin fraction, total cfRNAconcentration, and blood cell counts. In some embodiments, the total cfRNA concentration is calculated by normalizing the sample’s cDNA library concentration by the volume of input plasma and number of PCR cycles performed during library preparation to get the estimated total cfRNA concentration; multiplying the estimated total cfRNA concentration by the cfRNA cell type of origin fraction for a cell type of interest, resulting in the concentration of cfRNA derived from that cell type; and / or dividing the concentration of cfRNA derived from the cell type of interest by the blood cell counts for that respective cell type, providing an estimate of cell turnover in arbitrary units (a.u.), wherein the estimated cell turnover is comparable between cfRNA samples.

[0069] In some embodiments, the method further comprises quantifying distance to health, profding cell types of origin, measuring immune cell turnover, and performing RNA biomarker analysis.

[0070] In some embodiments, the step of determining or profiling the cell type of origin of cfRNA comprises performing nucleotide sequencing of the cell-free RNA. In some embodiments, the step of estimating the cfRNA cell types comprises using a reference RNA-seq data set with a deconvolution algorithm. Deconvolution algorithms are known in the art, for example, the BayesPrism and the Tabula Sapiens human cell atlas as a reference. In some embodiments, tissue damage is measured by comparing the cell type of origin measurements to a reference group. The CTO estimates are compared with known biomarkers and other indicators of specific organ injury or dysfunction. In some embodiments, the clinical decision support tool comprises the predication of inflammatory condition and measurements of tissue damage.

[0071] As used herein, “measuring immune cell turnover” is directed to the cell turnover rate for neutrophils, monocytes, leukocytes, and lymphocytes. The measurement is performed by dividing the product of the total cfRNA concentration and the cfRNA fraction derived from each cell type by the cell counts for that cell type. Such a calculation provides an estimate of cell turnover that is comparable between samples and can be trended over time.

[0072] In some embodiments, the method further comprises characterizing GVHD organ involvement, wherein the characterizing GVHD organ involvement comprises comparing the cell type of origin fraction of relevant cell types for each organ across the reference profde; and correlating organ specific damage in subjects with GVHD. In some embodiments, measuringimmune cell turnover comprises comparing the cfRNA profile of leukocyte, lymphocyte, and neutrophil turnover in the subject sample to the reference profile and correlating an increase in leukocyte, lymphocyte, and neutrophil turnover to a prediction of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject. In some embodiments, performing RNA biomarker analysis comprises comparing the cfRNA profile of the subject sample to the reference profile at the pre-determined timepoint and correlating an increase in interferon alpha / beta and gamma signaling as an early indication of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject. In some embodiments, quantifying distance to health comprises comparing the cfRNA profile in the subject sample to the reference profile using either a correlation in RNA abundances or a mahalanobis distance of the cell type of origin and correlating an increase in distance to a prediction of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject. In some embodiments, the profiling cell types of origin comprises comparing the cfRNA profile of leukocyte, lymphocyte, neutrophil, dendritic cell, and / or endothelial cell in the subject sample to the reference profile and correlating an increase in cfRNA abundance of leukocyte, lymphocyte, neutrophil, dendritic cell, and / or endothelial cell cfRNA to a prediction of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject.

[0073] In some embodiments, the step of determining a profile of a panel of biomarkers comprises performing nucleotide sequencing of the cell-free RNA. In some embodiments, the step of determining a profile of a panel of biomarkers comprises measuring the levels of RNAs corresponding to the biomarkers in the panel. In some embodiments, the step of determining a profile of a panel of biomarkers comprises measuring the levels of one, more, or all RNAs corresponding to the biomarkers included in the panel.

[0074] As used herein, “one, more, or all” refers to the panel of biomarkers comprising one of the listed genes, more than one of the listed genes, or all of the listed genes, respectively. In some embodiments, more than one includes at least 2 of the listed genes. In some embodiments, more than one includes at least 3 of the listed genes. In some embodiments, more than one includes at least 4 of the listed genes. In some embodiments, more than one includes at least 5 ofthe listed genes. In some embodiments, more than one includes at least 6 of the listed genes. In some embodiments, more than one includes at least 7 of the listed genes. In some embodiments, more than one includes at least 8 of the listed genes. In some embodiments, more than one includes at least 9 of the listed genes. In some embodiments, more than one includes at least 10 of the listed genes. In some embodiments, more than one includes at least 15 of the listed genes. In some embodiments, more than one includes at least 20 of the listed genes. In some embodiments, more than one includes at least 25 of the listed genes. In some embodiments, more than one includes at least 30 of the listed genes. In some embodiments, more than one includes at least 40 of the listed genes. In some embodiments, more than one includes at least 50 of the listed genes. In some embodiments, more than one includes at least 60 of the listed genes. In some embodiments, more than one includes at least 70 of the listed genes. In some embodiments, more than one includes at least 80 of the listed genes. In some embodiments, more than one includes at least 90 of the listed genes. In some embodiments, more than one includes at least 100 of the listed genes.

[0075] In some embodiments, the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: NIPAL3, MPO, CD9, IDS, TMEM159, XYLT2, OSBPL5, APBA2, EDC4, SNX29, VPS13D, HIPK2, IDI1, PDK3, IFI35, INPP5A, GNB5, DGCR2, NDST1, ATP6AP1, FBXW11, IRAG1, ATP2A3, SCARF1, DGKD, OSTM1, CDC14B, SSH1, TXLNA, KIF3C, EPS15, B4GALT1, OAS1, CDIP1, SLC9A1, STRN4, PITPNM2, CRAT, ACOT7, KCNK6, PSMD8, MARCHF2, POLR2E, SNUB, MYH9, MLC1, CTSG, NIN, PYGB, PSMA7, CDS2, CST3, UBL4A, SLC25A15, USP10, USP31, HOMER2, BPNT2, PPP6R1, PDCD5, PTPRS, DDX49, CLIP2, TRIM14, DNM1, VSIR, ARHGAP21, FBXL20, FTSJ3, ALOX12, ABCC3, GRPEL1, ZBTB16, PANXI, VWF, MLEC, OAS3, OAS2, EIF2B1, CCND3, FBRSL1, DBN1, CISH, ABHD14B, KANSL3, LOXL3, ODC1, WIPF1, PLEK, NIDI, GBP1, Clorfl98, IFIT3, IFIT2, ACAT2, SNX19, PLXDC2, TNFSF10, ZMIZ2, DDX39A, NMI, MRS2, TMX4, DSTN, MMP24OS, SIN3B, VIL1, GNAZ, TPST2, TST, KRI1, SLC44A2, CDKN2D, BCL2L2, CLIP1, SLC35D2, PRR7, SLC6A6, DIAPH1, FMO5, PRKAB2, RAFI, DLG4, CTNNBL1, BTBD2, PDZD2, CTIF, IRAK2, SORT1, UBAC2, OASL, CD36, CCT7, KIAA0513, APPL2, TTYH3, TLN1, TPMT, HADH, PCSK6, KIFC3, EIF4A3, CERS2, PYCR2, VGLL4, CTDSPL, TCTA, PLAC8, TIFA, FBXL17, SNX12, DOCK5, NFIB,END0D1, FERMT3, MTMR12, NBAS, GBP5, PSD3, BRAF, SKI, GPAT4, Fl 1R, FBXW5, PTGIR, KALRN, SLC37A1, GAB3, PCSK7, ORAI2, MGAT4B, EMC 10, SELENON, H3-3A, COL6A3, AIMP1, ANKRD33B, CITED2, GALNT10, CRYL1, TTC7B, CMTM5, WBP1L, GOLM2, RBPMS2, SAMD14, FN3K, TPM4, GATAD2A, ANKRD11, TRAPPC9, SPINT2, ECU, MAP4K2, MLKL, CD2BP2, GP9, REPS2, CHD3, CRTAP, RGS19, FIBP, DCAKD, ABLIM3, GPR137, ATP2A2, RAB1B, PACS1, PPP2R2D, MCTP1, TOM1L2, RUFY1, SAMD9L, NR2C2, CD151, NDUFAF3, GP5, PACS2, TMEM64, COA4, UBA7, MOB2, ADI1, RBM10, EWSR1, PTTG1IP, CLDN5, SNN, BTBD6, SLC24A3, ZBTB37, METTL7A, IRF7, BRCC3, IFIT1, RASA3, CD300E, PDE2A, TPCN1, ISG15, PEAR1, PLSCR1, PARVB, HACD4, TMEM63A, VKORC1L1, MAML3, PDLIM7, FLNA, SRC, AD ARBI, HTT, MAP3K5, MYO5A, PSAP, MYO1C, PAPSS2, UNC13B, CDC42BPB, PLXNB3, GRK5, AGP ATI, LY6G6F, HLA-C, HLA-E, PRR13, DOK6, HLA-H, HLA-A, RNF208, FGD5-AS1, FAM30A, GPX1, HLA-B, LINC00674, PEDS1, EIF6, CCDC71L, MTRNR2L8, ITGB3, RP11- 401.2, IKBKG, RN7SL1, and RP11-147L13.11.

[0076] In some embodiments, the panel of biomarkers comprises one, more, or all of the RAFI, RFC2, ERCC6L2, IDI1, FIBP, EDC4, FBXL20, KANSL1, and P0P4. In some embodiments, the panel of biomarkers comprises at least one of RAFI, IDI1, EDC4, KANSL1, RFC2, P0P4, and ERCC6L2.

[0077] In some embodiments, the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: RHD, CCL21, TMEM176B, CTD-2643I7.4, RP11- 329A14.1, ESPN, CLEC10A, MTCO3P12, CD8B, HTATSF1P2, KCNQ1OT1, Y RNA, MSC, ANKRD36BP2, EVPL, TNNT1, TSPAN2, MYH11, ATOH8, TEAD4, LINS1, BRSK1, Y RNA, ZBED6, TMEM205, RNF182, MTRNR2L1, PVALB, RP11-678G14.3, ALB, RNA5SP145, MY0M2, RPL9P9, MYCL, RNA5SP149, CDC42EP1, ABO, SPOCD1, GALNT3, DDXI ILIO, PCDH17, RP11-452G18.2, DNAJB5, SMIM24, MS4A1, IGHA1, MTND4P12, DDX11L2, HLA-J, Y RNA, Y RNA, C19orfi3, MT-RNR2, USP54, GALT, LINC00504, ARL6IP6, MTRNR2L8, RN7SL767P, PDZK1IP1, KRT1, COL4A1, PROK2, SLC15A2, H2BC8, Y RNA, SBK1, DEPPI, BAHCC1, HBG2, CXCL9, TOX2, RALY-AS1, SNAI2, ALPK3, STAP1, ERAP2, PLA2G7, GSTM5, CD79A, SPRY4, IGHD, MYO 10, AC114271.2, PARVA, CYP2E1, CCR2, STK33, IGHV1-3, XXbac-BPG299F13.17, ST20-AS1,Y RNA, CFD, RP11-712B9.2, LILRA4, MEIS3P1, SRGAP1, HID1, and REEP1. In some embodiments, the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: RHD, CCL21, TMEM176B, CTD-2643I7.4, RP11-329A14.1, ESPN, CLEC10A, MTCO3P12, CD8B, HTATSF1P2, KCNQ1OT1, Y RNA, MSC, ANKRD36BP2, EVPL, MYH11, AT0H8, TEAD4, Y RNA, ZBED6, RNF182, MTRNR2L1, RNA5SP145, RPL9P9, MYCL, RNA5SP149, CDC42EP1, SPOCD1, GALNT3, DDX11L10, PCDH17, DNAJB5, MTND4P12, C19orf33, MT-RNR2, PDZK1IP1, KRT1, SLC15A2, DEPPI, BAHCC1, ERAP2, IGHD, MYOIO, CFD, RP11-712B9.2, SRGAP1, RNF152, RN7SL736P, RN7SL665P, and TRIP13. In some embodiments, the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: RHD, CCL21, TMEM176B, CTD- 264317.4, RP11-329AI4.1, ESPN, CLEC10A, MTCO3P12, HTATSF1P2, KCNQ1OT1, ANKRD36BP2, RNF182, and RNA5SP149. In some embodiments, the panel of biomarkers comprises RHD and CCL21.

[0078] In some embodiments, the analysis of the biological sample obtained from the subject is predictive of immunosuppression taper failure. In some embodiments, analyzing the cell type of origin data predicts immunosuppression taper failure when elevated levels of mature conventional dendritic cell - and endothelial cell-derived cfRNA are observed in the sample as compared to a reference sample of a subject that successfully completed immunosuppression taper. In some embodiments, longitudinal analysis of subjects with at least one post-taper sample will be predictive of taper failure when the longitudinal cfRNA counts demonstrate a sustained elevation of these cfRNA.

[0079] In some embodiments, the panel of biomarkers comprises TCF7. In some embodiments, the panel of biomarkers comprises MYOF. In some embodiments, the panel of biomarkers comprises TCF7 and MYOF.

[0080] In some embodiments, the panel of biomarkers comprises one, more, or all of the genes MY0M2, CXCL9, and PVALB. In some embodiments, the panel of biomarkers comprises MY0M2. In some embodiments, the panel of biomarkers comprises CXCL9. In some embodiments, the panel of biomarkers comprises PVALB.

[0081] In some embodiments, the panel of biomarkers comprises one, more, or all of Major Histocompatibility Complex, Class I - E, C, A, and B (HLA-E, HLA-C, HLA-A, and HLA-B, respectively) genes.

[0082] In some embodiments, the panel of biomarkers comprises at least one histone gene. In some embodiments, the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: Hl-1, Hl-2, Hl-3, Hl-4, Hl-5, Hl-6, H2AC1, H2AC4, H2AC6, H2AC7, H2AC8, H2AC11, H2AC12, H2AC13, H2AC14, H2AC15, H2AC16, H2AC17, H2AC18, H2AC19, H2AC20, H2AC21, H2AC25, H2BC1, H2BC3, H2BC4, H2BC5, H2BC6, H2BC7, H2BC8, H2BC9, H2BC10, H2BC11, H2BC12, H2BC13, H2BC14, H2BC15, H2BC17, H2BC18, H2BC21, H2BC26, H2BC12L, H3C1, H3C2, H3C3, H3C4, H3C6, H3C7, H3C8, H3C10, H3C11, H3C12, H3C13, H3C14, H3C15, H3-4, H4C1, H4C2, H4C3, H4C4, H4C5, H4C6, H4C7, H4C8, H4C9, H4C11, H4C12, H4C13, H4C14, H4C1 5H1-0, Hl-7, Hl-8, Hl-10, H2AZ1, H2AZ2, MACROH2A1, MACROH2A2, H2AX, H2AJ, H2AB1, H2AB2, H2AB3, H2AP, H2AL1Q, H2AL3, H2BK1, H2BW1, H2BW2, H2BW3P, H2BN1, H3-3A, H3-3B, H3-5, H3-7, H3Y1, H3Y2, CENPA, and H4C16. In some embodiments, the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: Hl-1, Hl-2, Hl- 3, Hl-4, Hl-5, Hl-6, H2AC1, H2AC4, H2AC6, H2AC7, H2AC8, H2AC11, H2AC12, H2AC13, H2AC14, H2AC15, H2AC16, H2AC17, H2AC18, H2AC19, H2AC20, H2AC21, H2AC25, H2BC1, H2BC3, H2BC4, H2BC5, H2BC6, H2BC7, H2BC8, H2BC9, H2BC10, H2BC11, H2BC12, H2BC13, H2BC14, H2BC15, H2BC17, H2BC18, H2BC21, H2BC26, H2BC12L, H3C1, H3C2, H3C3, H3C4, H3C6, H3C7, H3C8, H3C10, H3C11, H3C12, H3C13, H3C14, H3C15, H3-4, H4C1, H4C2, H4C3, H4C4, H4C5, H4C6, H4C7, H4C8, H4C9, H4C11, H4C12, H4C13, H4C14, and H4C15. In some embodiments, the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: H2BC21, Hl-2, H2BC4, H2AC6, Hl-4, H2BC12, H3C10, and Hl-5.

[0083] In some embodiments, the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: RPA2, CLSPN, CDCA8, CDC20, KIF2C, NASP, USP1, PSRC1, ANP32E, CKS1B, NUF2, NEK2, DTL, CENPF, LBR, EXO1, RRM2, CENPA, MSH2, BUB1, CKAP2L, MCM6, CDCA7, HJURP, MCM2, SMC4, ECT2, SLBP, TACC3, CENPE, HMGB2, CENPU, CDC25C, HMMR, GMNN, TTK, CASP8AP2, ANLN, RFC2, CDCA2,MCM4, CCNE2, DSCC1, ATAD2, CKS2, TUBB4B, CDK1, KIF20B, KIF1 1, HELLS, MKI67, RRM1, E2F8, CKAP5, FEN1, P0LD3, RAD51AP1, NCAPD2, CDCA3, CBX5, PRIM1, TMPO, GAS2L3, UNG, CKAP2, G2E3, DLGAP5, UBR7, RAD51, NUSAP1, WDR76, CCNB2, TIPIN, KIF23, BLM, CTCF, GINS2, PIMREG, AURKB, CDC6, TOP2A, BRIP1, JPT1, BIRC5, TYMS, NDC80, UHRF1, PCNA, TPX2, UBE2C, AURKA, CHAF1B, CDC45, MCM5, RANGAP1, GTSE1, and POLA1. In some embodiments, the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: NASP, USP1, ANP32E, CENPF, RRM2, MCM6, SMC4, SLBP, TACC3, CENPE, HMGB2, ATAD2, TUBB4B, KIF20B, MKI67, CKAP5, NCAPD2, CBX5, TMPO, CKAP2, CTCF, TOP2A, PCNA, TPX2, and RANGAP1. In some embodiments, the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: ANP32E, CENPF, HMGB2, TUBB4B, and MKI67.

[0084] An aspect of the current disclosure is directed to a method for monitoring a microbiome of an HSCT recipient subject for infection, the method comprising: a) obtaining cell-free RNA from at least one biological sample of the subject, wherein the biological sample is obtained at the day of infusion or less than 36 months, or 24 months, or 12 months, or less than 6 months, or less than 3 months, or less than 2 months after the day of infusion, wherein the biological samples are plasma samples or serum samples; b) determining a profile of a panel of metagenomic biomarkers in the cell-free RNA; c) determining an inflammatory disease status in the subject based on the profile of the panel of metagenomic biomarkers; and d) removing background total viral counts of all non-mammalian infecting viral families from each sample.

[0085] In some embodiments, the at least one biological sample of the subject comprises a set of longitudinally collected biological samples of the subject at preconditioning, day of infusion, engraftment, and / or one or more post infusion timepoints. In some embodiments, the step of determining a profile of a panel of biomarkers comprises performing nucleotide sequencing of the cell-free RNA. In some embodiments, the step of determining a profile of a panel of biomarkers comprises measuring the levels of RNAs corresponding to the biomarkers in the panel. In some embodiments, monitoring the microbiome of the subject based on the profile ofthe panel of biomarkers comprises comparing the profile of the panel of biomarkers in the cell free RNA from the biological sample to a reference profile of the panel of biomarkers. In some embodiments, the reference profile of the panel of biomarkers is from a healthy donor or an HSCT subject without complications. In some embodiments, the panel of biomarkers comprises those for at least one of the Flaviviridae, Anelloviridae , Retroviridae, Pegivirus, and Herpesviridae viral families. In some embodiments, the panel of biomarkers comprises biomarkers for at least one of the viral genera: Betacoronavirus (coronaviruses), Betapolyomavirus (BKV), Cytomegalovirus (CMV), Lymphocryptovirus (EBV), and Mastadenovirus (adenoviruses). In some embodiments, the abundance of prokaryotic reads are quantified for the cfRNA from one or a combination of longitudinal time points and compared to the reference abundance of prokaryotic reads.

[0086] In some embodiments, a correlation is made that the subject may be experiencing complications when the subject sample comprises elevated levels of biomarkers when compared to the reference panel. As used herein, “elevated levels of biomarkers” refers to the count of biomarker of the subject sample being significantly higher statistically as compared to the reference panel of biomarker.

[0087] In some embodiments, at least two cfRNA measurements are used in determining the subject’s recovery trajectory and wherein the at least two cfRNA measurements used consist of the parameters included in the group consisting of microbial kingdom counts, gene counts, viral family counts, distance to health, and cell turnover. In some embodiments, the cfRNA measurements used consist of microbial kingdom counts, genes counts, viral family counts, distance to health, and cell turnover.

[0088] As used herein, “recovery trajectory” refers to the track or path of how a patient's health improves over time after a transplant. Recovery trajectory can be predicted based on methods presented herein.

[0089] In some embodiments, the set of longitudinally collected biological samples includes more than one predetermined timepoints after infusion. In some embodiments, the set of longitudinally collected biological samples comprises a biological sample collected on the day of infusion. In some embodiments, the more than one predetermined timepoint comprises monthly intervals after the day of infusion, i.e. at 1 Month, IM; at 2 Month, 2M; at 3 Month, 3M; and / orat 6 Month, 6M). In some embodiments, the profile of biomarkers at an early time point post infusion is indicative of the likelihood of GVHD. In some embodiments, an early time point post infusion refers to a time point less than 1 month post infusion. In some embodiments, an early time point post infusion refers to a time point less than 2 months post infusion. In some embodiments, an early time point post infusion refers to a time point less than 3 months post infusion.

[0090] In some embodiments, the biological sample is a plasma sample. In some embodiments, the biological sample is a serum sample.

[0091] In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.

[0092] In some embodiments, the profile of the panel of biomarkers comprises subject-specific gene expression of the biomarkers in the panel. In some embodiments, the cell-free RNA is substantially free of contaminant RNA.

[0093] In some embodiments, the subject’s recovery trajectory is used in determining treatment for the subject.

[0094] In some embodiments, test performance of the trained classification model is measured through one or more of accuracy, sensitivity, specificity, and / or area under the receiver operating characteristic curve (ROC-AUC).

[0095] As used herein, “test accuracy” or “accuracy” provides the final generalization power. In some embodiments, the inflammatory syndrome status is determined at a greater than 80% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 85% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 86% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 87% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 88% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 89% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 90% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 91% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 92% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greaterthan 93% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 94% accuracy. In some embodiments, the inflammatory syndrome status is determined at a greater than 95% accuracy.

[0096] As used herein, “sensitivity” describes the probability of a positive test result, conditioned on the individual truly being positive. Sensitivity is determined by the number of true positives divided by the sum of true positives and false negatives. As such, a test which reliably detects the presence of a condition, resulting in a high number of true positives and low number of false negatives, will have a high sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 80% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 85% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 86% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 87% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 88% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 89% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 90% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 91% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 92% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 93% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 94% sensitivity. In some embodiments, the inflammatory syndrome status is determined at a greater than 95% sensitivity.

[0097] As used herein, “specificity” refers to the probability of a negative test result, conditioned on the individual truly being negative. Specificity is determined by the number of true negatives divided by the sum of true negatives and false positives. As such, a test which reliably excludes individuals who do not have the condition, resulting in a high number of true negatives and low number of false positives, will have a high specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 70% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 75% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 80% specificity.In some embodiments, the inflammatory syndrome status is determined at a greater than 81% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 82% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 83% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 84% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 85% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 86% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 87% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 88% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 89% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 90% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 91% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 92% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 93% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 94% specificity. In some embodiments, the inflammatory syndrome status is determined at a greater than 95% specificity.

[0098] In some embodiments, the trained classification model comprises one or more models selected from the group consisting of generalized linear models with Ridge and LASSO feature selection (GLMNETRIDGE and GLMNETLASSO), support vector machines with linear and radial basis function kernel (SVMLin and SVMRAD), random forest (RF), random forest ExtraTrees (EXTRATREES), neural networks (NNET), linear discriminant analysis (LDA), nearest shrunken centroids (PAM), C5.0 (C5), k-nearest neighbors (KNN), naive bayes (NB), CART (RPART), generalized linear model (GLM), and greedy forward search algorithm (GFS).EXAMPLESExample 1. Clinical cohort 1 - HSCT dynamics

[0099] To test the utility of cfRNA to monitor HSCT, 449 samples were collected and analyzed from 96 subjects who underwent allogeneic HSCT at the Dana-Farber Cancer Institute for a variety of conditions, including leukemia, lymphoma, and aplastic anemia (Table SI). Sampleswere collected longitudinally at predetermined time points before and after HSCT (prior to conditioning therapy, PR; at the day of infusion prior to the transplant, DO; at hematopoietic engraftment, E; and at 1 Month, IM; 2 Month, 2M; 3 Month, 3M; and 6 Month, 6M after transplantation, FIG. 1A and B). In addition, 18 samples were collected from healthy stem cell donors to use as a reference. Complications were common within the first 9 months of HSCT in this cohort of recipients, including acute and chronic GVHD (aGVHD 53%, cGVHD 32%), BK virus reactivation (BKVR, 53%), cancer relapse (35%), and graft failure (5%). Recipients often had more than one complication, and the timing of complications differed between individuals and within complication categories (FIG. 1C, FIG. 2A-B).Table 1Abbreviations: M: male; F: female; AA: aplastic anemia; AML: acute myeloid leukemia; ALL: acute lymphoid leukemia; CLL: chronic lymphocytic leukemia; CMML: chronic myelomonocytic leukemia; MDS: myelodysplastic syndrome; MM: multiple myeloma; MPS: myelodysplastic syndrome; PBSC:peripheral blood stem cell; HLA: human leukocyte antigen; Tac: tacrolimus; Sir: sirolimus / rapamycin; MTX: methotrexate; MMF: mycophenolate mofetilExample 2. cfRNA reflects recovery and return to health.

[0100] cfRNA was used to measure a “distance to normal.” To this end, two metrics were calculated: the average correlation of transcript counts and distance of cell type of origin profiles of each plasma cfRNA profile to those of the healthy stem cell donors (FIG. ID, FIG. 2C Methods). Patients were then grouped into two broad categories: those who experienced at least one major complication within 9 months post-transplant (COMP, n=82) and those who did not (noCOMP, n=10) (Methods). cfRNA profiles at the engraftment time point were found to be the least similar to normal profiles, followed by a time-dependent increase in correlation to the normal profiles. A similar trend was observed in the distance of cell type of origin profiles. Within-group correlations were also measured at each timepoint, and showed that the cfRNA profiles of patients with complications have a lower within-group correlation than complication- free patients and the healthy stem cell donors, indicating that cfRNA profiles from patients with a complication -free course follow a similar trajectory while each patient with a complicated recovery has a unique trajectory, which can be thought of as an analogy to the Anna Karenina principle in statistics (75) (FIG. 2D).Example 3. cfRNA origins are temporally dynamic and distinct from whole blood cell counts.

[0101] To gain a deeper understanding of the dynamics of plasma cfRNA profiles during HSCT, the cell types that contribute cfRNA to the mixture in plasma were quantified. To this end, BayesPrism was used to computationally deconvolve cfRNA using the Tabula Sapiens singlecell atlas as a reference. This analysis revealed rich dynamic changes in the cell-types-of-origin (CTO) of cfRNA in response to conditioning and stem cell infusion (FIG. 3A). For example, the proportion of hematopoietic-lineage-cell derived cfRNA was much lower at the day of infusion compared to the engraftment time point, likely reflecting the effects of pre-transplant conditioning and subsequent recovery of donor hematopoiesis.

[0102] Next, the cfRNA CTO fractions were compared with whole blood cell counts obtained via automated image-based cell counting from paired samples. Similar trends were observed between the cell count and CTO for lymphocytes, but not for neutrophils or monocytes (FIG.IB). For example, highly elevated levels of neutrophil and monocyte derived cfRNA were observed, but not cell counts, in COMP patients. It was found that while log-normalized cell count data follows a unimodal distribution, the CTO measurements of neutrophils and monocytes were bimodal, with patients either having high or low CTO levels (FIG. 1C). Interestingly, lymphocyte CTO measurements were unimodal, which may indicate that cfRNA from lymphocytes is more dependent on the abundance of living cells, which in turn may be explained by the lower turnover rates for these cells compared to neutrophils and monocytes.

[0103] Allogeneic HSCT recipients are highly immunocompromised and consequently susceptible to infection. To test whether cfRNA from viruses can be identified in plasma, all sequences that did not align to the human genome to a microbial reference database were mapped. Reads were removed which align to viruses that are not known to infect humans and viruses that are known contaminants (methods). From the remaining data, high levels of cfRNA from Anellovirus and Pegivirus (FIG. 3D) were observed. Reactivation of Anelloviridae was observed in response to the immunosuppressive therapy, for both patients with and without complications. Similar blooms of Anelloviridae have previously been reported for solid-organ transplant patients through the lens of cfDNA. In addition, the burden of Pegivirus cfRNA was strikingly elevated throughout the post -transplant course for patients with complications compared to patients without complications. We further analyzed the burden of cfRNA from five viral genera known to affect HSCT recipients: Betacoronavirus (coronaviruses), Betapolyomavirus (BKV), Cytomegalovirus (CMV), Lymphocryptovirus (EBV), and Mastadenovirus (adenoviruses). Overall low viral burden was observed, but higher detection rates in COMP patient samples (FIG. 4).Example 4. Patients who develop acute GVHD have elevated cfRNA immune cell signatures at engraftment.

[0104] Acute GVHD (aGVHD) is a common complication of HSCT in which engrafted T cells attack the patient’s organs, typically within the initial 100 days post-transplantation. In view of the frequent occurrence of aGVHD in the initial period following transplantation, we examined the engraftment time point samples for signatures that may predict the onset of aGVHD. Only one patient was diagnosed with aGVHD prior to engraftment, and that patient's engraftment sample was removed from subsequent analyses. Results showed that patients who developaGVHD were significantly less similar to the healthy donor samples, as measured by the correlation with the plasma cfRNA profiles of healthy donors (FIG. 5A; Mann-Whitney U-test, p-value = 0.045). In addition, patients who develop aGVHD were found to have higher levels of cfRNA derived from leukocytes and lower levels of cfRNA derived from platelets than those who did not develop aGVHD (FIG. 5B, FIG. 6). Furthermore, patients who develop aGVHD were found to have significantly higher cell type of origin diversity (FIG. 5C; Simpson alpha diversity test, p-value = 0.016).

[0105] Next, transcript-level differences were analyzed between patients that develop aGVHD and those that do not at the engraftment time point. 843 differentially abundant transcripts were found (FIG. 5D; DESeq2, Benjamini-Hochberg (BH) adjusted p-value < 0.1). Using the most differentially abundant transcripts (BH adjusted p-value < 0.1, area under the receiver operating characteristic curve (AUC-ROC) > 0.7), samples clustered well based on future aGVHD status (FIG. 5E). Pathway analysis (Qiagen IP A) was performed and found pathways enriched in patients who develop aGVHD included “RHOGDI Signaling”, “Hedgehog ‘off state”, “Metabolism of polyamines”, “S Phase”, and “G1 to G2 / S transition” while pathways enriched in patients who do not develop aGVHD were related to platelet activity, integrin signaling, and neutrophil degranulation (FIG. 5F).

[0106] Finally, transcript abundance was analyzed in relation to the time to diagnosis of aGVHD using patient samples collected at the engraftment, one-month, and two-month time points before diagnosis of aGVHD. 132 transcripts were significantly associated with time to diagnosis (methods; DESeq2, BH adjusted p-value < 0.1). Notably, Major Histocompatibility Complex, Class I - E, C, A, and B (HLA-E, HLA-C, HLA-A, and HLA-B) were all negatively associated with time to aGVHD diagnosis - the closer a sample was taken to aGVHD diagnosis, the higher the abundance of HLA-E, HLA-C, HLA-A, and HLA-B (FIG. 5G).Example 5. Prognostic cfRNA signatures of cGVHD and L-aGVHD

[0107] Chronic GVHD (cGVHD) and late-acute GVHD (L-aGVHD) are later-onset complications of HSCT, and to identify cfRNA signatures predictive of its development, samples collected at the three- and six-month time points were analyzed. The analysis was limited to samples collected before any cGVHD / L-aGVHD diagnosis. At both the three- and six-month post transplantation time points, patients who developed cGVHD / L-aGVHD exhibitedsignificantly higher levels of dendritic cell-derived cfRNA compared to those who did not develop cGVHD / L-aGVHD. Next, to develop potential diagnostic signatures, differential abundance analysis (DESeq2) was performed. 106 significantly differentially abundant transcripts were identified (FIG. 7A, Benj ami ni -Hochberg adjusted p-value < 0.1).

[0108] Given these promising differences and the large sample size available (n=147), we analyzed the potential to train a model to predict cGVHD / L-aGVHD. 101 repeated stratified train-test splits were performed without replacement to train and test a LASSO regression model (Methods). The model with the median test performance achieved promising results (AUC- ROC train = 1.00, test = 0.75; FIG. 7B). Finally, leveraging the inherent feature selection of the LASSO algorithm, we explored which transcripts were most consistently chosen. A small subset of transcripts was repeatedly selected across iterations, with six transcripts (CTD-264317.4, DDX11L2, SMIM24, ALB, LINS1, TSPAN1) represented in all 101 models (FIG. 7C). Interestingly, ALB was found to be elevated in patients who do not develop cGVHD / L-aGVHD, despite those patients not having higher protein ALB or ALT levels (data not shown). Also, multiple genes used for blood typing, RHD, HLA-J, and ABO, were selected by many models.Example 6. cfRNA reflects organ specific GVHD damage.

[0109] Next, cfRNA profding was tested to pinpoint organ-specific damage due to acute and chronic GVHD, focusing on skin, GI, and liver involvement. Patients were stratified into four groups for each organ: GVHD+ with organ involvement, GVHD+ without organ involvement, GVHD-, and healthy stem cell donors. The maximum fraction of relevant cell types for each organ was then compared across the sampling time course, allowing control for time of diagnosis (FIG. 7). A significantly higher burden of intrahepatic cholangiocyte and hepatocyte-derived cfRNA was observed for patients with liver GVHD, and salivary gland-derived cfRNA for patients with GI GVHD. Notably, tissue-specific cfRNA was not observed for patients with skin GVHD. These results indicate that cfRNA reflects organ-specific damage in patients with acute and chronic GVHD.Example 7. Clinical cohort 2 - Immunosuppression Taper

[0110] Unlike solid organ transplantation, most allogeneic HCT recipients eventually discontinue immunosuppression to achieve immune tolerance. Predicting which patients can successfully taper IS without developing cGVHD or L-aGVHD is a major clinical challenge,with little published data on successful discontinuation rates and no established predictive biomarkers. To assess the utility of cfRNA signatures to predict immunosuppression taper outcomes following HSCT, 100 samples were collected and analyzed from 53 subjects who underwent allogeneic HSCT at the Moffitt Cancer Center for a range of conditions. Samples were obtained at following key timepoints: the first day of immunosuppression taper, the last day of taper, on the day of taper failure, and at two weeks, one month, three months, and six months after the end of taper (FIG. 9A). The initiation of taper varied among patients according to individual clinical progression, and the onset of taper failure occurred over a broad range of days following the start of taper (FIG. 9B). Based on these outcomes, patients were classified as either having failed (n = 30) or successfully completed (n = 23) taper.Example 8. cfRNA signature for risk of immunosuppression taper failure[OHl] Next, cfRNA signatures predictive of taper failure prior to immunosuppression tapering were identified. Analysis of the cell type of origin data revealed that patients who eventually failed taper exhibited elevated levels of mature conventional dendritic cell- and endothelial cell- derived cfRNA compared to those who succeeded (FIG. 9C-D). Notably, these elevated cell contributions were evident at the start of taper, at its conclusion, and during the event-driven collection at the time of failure. Furthermore, longitudinal analysis of patients with at least one post-taper sample demonstrated a sustained elevation of these cfRNA signals in individuals who ultimately experienced taper failure.

[0112] Next, transcript-level differences were analyzed between patients that succeed taper and those that do not use samples collected at the end of taper or during a follow up time point. 573 differentially abundant transcripts were found (FIG. 9E; DESeq2, BH adjusted p-value < 0. 1). Of these, the two most significantly differentially abundant were TCF7 and MYO F, which maintain differences across the longitudinal sampling time course (FIG. 9F).

[0113] Given these promising differences and the large sample size available (n=88), the potential to train a model to predict immunosuppression failure was analyzed. 101 repeated stratified train-test splits were performed without replacement to train and test a LASSO regression model. Given the longitudinal design of the study, stratified group splitting was employed for the train / test partitioning. This approach ensured that all samples from an individual patient were assigned exclusively to either the training or test set, thereby preventingdata leakage. The model with the median test performance achieved promising results (AUC- ROC train = 1.00, test = 0.77; FIG. 9G). Finally, leveraging the inherent feature selection of the LASSO algorithm, we explored which transcripts were most consistently chosen. Notably, a small subset of transcripts was repeatedly selected across iterations, with three transcriptsMY0M2, CXCL9, PVALB) represented in over 96 models (FIG. 9H).Example 9. Shared signatures across cohorts

[0114] Despite our cohorts being distinct, we sought to find signatures of cGVHD or immunosuppression failure that were shared across cohorts. For this, we analyzed overlap in the transcripts chosen by the repeated stratified machine learning modeling done for cGVHD in the DFCI cohort and for immunosuppression failure in the MCC cohort. We summed together the number of models that chose each gene in the DFCI and MCC modeling and found that the two genes with the highest combined number of models were RHD and CCL21 (FIG. 10A). When analyzing the counts of these markers at the DFCI three- and six-month time points, along with the MCC taper start and taper end timepoints we see conserved patterns - RHD is elevated in patients that do not develop the target complication while CCL21 is increased in patients that do develop the target complication (FIG. 10B). Given these promising findings, we created a new cfRNA score - a log 10 normalized score of CCL21 over RHD transcript counts per million - where we saw conserved patterns across sites (FIG. 10C). When we separate samples by site and calculate AUC we find moderate performance (AUC ROC MCC=0.71, DFCI=0.68), which is promising given the lack of tuning and that this only uses two transcripts (FIG. 10D). Finally, a survivorship analysis revealed that patients with high CCL2HRHD scores develop cGVHD or fail immunosuppression taper sooner than those with a low score (high score > median; FIG.10E)Example 10. Measuring cell turnover

[0115] Given that blood cell counts measure live cells and cfRNA is a surrogate for dead cells, we tested if this observed discrepancy can be explained by changes in cellular turnover: while the number of live cells may remain relatively unchanged, the lifespan and cell replacement rate may vary greatly for specific cell types. The cell turnover rate was calculated for neutrophils, monocytes, leukocytes, and lymphocytes by dividing the product of the total cfRNA concentration and the cfRNA fraction derived from each cell type by the cell counts for that celltype. This provided an estimate of cell turnover that is comparable between samples and can be trended over time. Using this measure, a striking increase in the turnover rate of neutrophils and monocytes at the engraftment time point was observed. This increase in turnover is expected given that neutrophils and monocytes are among the first cells to engraft. The turnover rate of neutrophils followed a bimodal distribution at the subsequent timepoints (FIG. 11A), which separates patients with an inflammatory immune response and those without. Similarly, the monocyte turnover rate displayed a bimodal distribution after the one-month timepoint. Based on these observations, cell turnover rates measured by cfRNA may provide insights into immune dynamics beyond what can be learned from blood cell counts alone.

[0116] We next determined whether cfRNA also captures the immune dynamics associated with GVHD prior to diagnosis. The turnover rate was quantified for lymphocytes, monocytes, neutrophils, and total leukocytes using the approach described above (FIG. 11B). To account for the impact of GVHD treatment on immune dynamics, only samples collected before diagnosis were used. Elevated immune cell turnover was found to be associated with both aGVHD and cGVHD, in particular: elevated leukocyte turnover at Engraftment in aGVHD patients, elevated lymphocyte turnover at the 3M timepoint in patients with late onset aGVHD, elevated monocyte turnover at the 3M timepoint for patients with late onset aGVHD and cGVHD, and elevated leukocyte and neutrophil turnover at the 6M time point for cGVHD patients (FIG. 11C). These findings highlight the immune activity preceding diagnosis and manifestation of GVHD symptoms.Example 11. cfRNA Metagenomics

[0117] Allogeneic HSCT recipients are highly immunocompromised and consequently susceptible to infection. To test whether cfRNA from viruses and bacteria can be identified in plasma, all sequences that did not align to the human genome were mapped to a microbial reference database (Methods). The plasma cfRNA virome of the healthy stem cell donors was first characterized to establish a healthy baseline. Four non-mammalian viral families accounted for 98% of all viral reads in these samples (Siphoviridae^ Tospoviridae Baculoviridae, and Alphafexiviridae, FIG. 12A). These viruses are not known to infect humans; therefore, these reads were suspected to be due to contamination or incorrect classification of human reads. To test this, the abundance of those 4 viruses were correlated to the total RNA concentration andhuman alignment rates within samples (Methods). Strong negative correlation was observed between 3 of the 4 viruses and total RNA concentration (Tospoviridae , Baculoviridae , and Alphafexiviridae). Such negative correlation is indicative of contamination: samples with lower RNA concentration are proportionally more susceptible to environmental contamination such as from reagents. The signal for Siphoviridae was strongly and positively correlated to human alignment rate, indicating misclassification of human reads. Therefore Siphoviridae, Tospoviridae, Baculoviridae , Alphafexiviridae, and all other non-mammalian viral families were excluded from our analyses. After removal of this background, remaining viral cfRNA was from I' laviviridae, Retroviridae , Annelloviridae, and other mammalian viruses that are known to be present in healthy populations, giving us confidence that these reads are bona fide. Next the virome for COMP and noCOMP patient samples was compared across the sampling time course and observed significantly higher viral counts at the preconditioning and day zero timepoints for all patient samples regardless of complication status (FIG. 12B). However, for timepoints after stem cell infusion, we observed significantly higher viral cfRNA in plasma of COMP patients relative to noCOMP patients and the healthy stem cell donors.

[0118] The longitudinal dynamics of the most abundant viral families in samples from patients with complications was analyzed (Flaviviridae , Anelloviridae, Herpesviridae, and Retroviridae, FIG. 12C) The burden of Flaviviridae cfRNA was strikingly elevated throughout the posttransplant course for patients with complications compared to patients without complications (FIG. 12C). Additionally, reactivation of Anelloviridae was observed in response to the immunosuppressive therapy, for both patients with and without complications (FIG. 12C). Similar blooms of Anelloviridae vQ previously been reported for solid-organ transplant patients through the lens of cfDNA. Last, significantly higher levels of Herpesviridae and Retroviridae cfRNA were observed at preconditioning and day of infusion relative to stem cell donors. For patients without complications, the burden of cfRNA from these viruses decreased posttransplantation, whereas for patients with complications, the burden of cfRNA from these viruses remained elevated (FIG. 12C). We further analyzed the burden of cfRNA from five viral genera known to affect HSCT recipients: Betacoronavirus (coronaviruses), Betapolyomavirus (BKV), Cytomegalovirus (CMV), Lymphocryptovirus (EBV), and Mastadenovirus (adenoviruses).Overall low viral burden was observed, but higher detection rates in COMP patient samples (FIG. 4)

[0119] The bacterial component of the cfRNA microbiome was quantified. Elevated burden of bacterial cfRNA was observed at preconditioning and the day of infusion relative to stem cell donors (FIG. 12D). After infusion we observed significantly higher bacterial cfRNA in COMP patient samples compared to noCOMP patient samples and healthy stem cell donors, while there was no difference in bacterial cfRNA for noCOMP patient samples and healthy stem cell donors. The majority (60%) of the reads matched to a prokaryotic rRNA operon reference (FIG. 12E). These results show that cfRNA may be a relevant analyte for measuring metagenomic dynamics and that the burden of viral and bacterial cfRNA may be used as an indicator of complications and immunosuppression in HSCT.Example 12. Cell cycle and histone transcript dynamics

[0120] Next, the abundance of cell cycle transcripts of genes previously annotated as markers of S or G2 / M phase was explored (Tirosh et al., 2016). The abundance of each transcript was z- score normalized, then summed together all z-scores of each sample. Analyzing trends over time, we see that the sum of z-scores is higher at engraftment relative to pre-infusion timepoints and decreases temporally from one to six months (FIG. 13A). Interestingly, some samples from patients with complications presented elevated sum of z-scores at multiple timepoints. Given the striking signature at Engraftment, we next explored the cell cycle marker trends relative to aGVHD. Samples at Engraftment showed that patients who go on to develop aGVHD have higher sum Z-scores of cell cycle markers (p-value = 0.07, Mann Whitney U-test, FIG. 13B). Interestingly, both groups had higher levels at engraftment than the healthy donors.

[0121] The abundance of histone transcripts was then examined. Similar to the cell cycle transcript, elevated sum of histone transcript counts were observed at engraftment relative to preinfusion timepoints and a temporal decrease in abundance from one to three months, with some samples from patients with complications being outliers much higher than others (FIG. 13C). Also, elevated levels of histone transcripts were found in samples at Engraftment from patients who go on to develop aGVHD compared to those that do not (p-value = 0.06, Mann Whitney U- test, FIG. 13D)Example 13. Materials and methods

[0122] Ethics Statement

[0123] The study was approved by the Dana-Farber / Harvard Cancer Center’s Office of Human Research Studies and Moffitt Cancer Center IRB. All patients provided written informed consent.

[0124] Clinical Cohort 1

[0125] Adult patients undergoing allogeneic HCT at Dana-Farber Cancer Institute were prospectively enrolled on a rolling basis from August 2018 to August 2019. Samples were collected longitudinally at predetermined time points: prior to conditioning therapy (Pre- Conditioning, PR), at the day of infusion prior to stem cell infusion (Day of Infusion, DO), at the time of myeloid engraftment (Engraftment, E) and at monthly intervals after the day of infusion (1 Month, IM; 2 Month, 2M; 3 Month, 3M; 6 Month, 6M). This cohort was initially established to monitor BKV after transplant and additional sampling was done at the onset of BK-related symptoms, disease, or reactivation. In the case of two timepoints overlapping, the sample was preferentially labeled as engraftment or month 1 / 2 / 3 / 6 (in that order). Samples from healthy stem cell transplant donors were used as a healthy reference control.

[0126] Clinical Cohort 2

[0127] Adult patients that had received HCT at the Moffitt Cancer Centre were enrolled on a rolling basis from October 2010 to February 2018. Samples were collected at the following timepoints: start of immunosuppression taper (Taper Start), end of immunosuppression taper (Taper End), day of taper failure (Failure) and at regular intervals after the end of taper (2 weeks, wk2; 1 month, mol; 3 months, mo3; 6 months, mo6).

[0128] Engraftment

[0129] Neutrophil engraftment was considered when blood samples contained an absolute neutrophil count greater than or equal to 500 cell per microliter blood on two separate measurements. Median day of engraftment was 19 days after transplant.

[0130] Patient Categorization

[0131] In cohort 1, patients were categorized into three groups: healthy stem cell donors, early complication-free (noCOMP), and early complication-present (COMP). Healthy stem cell donors were individuals who provided stem cell products for patients in this study. Complication-present patients were HSCT recipients that were diagnosed with any of the following complications within 9 months of infusion: relapse, graft failure, acute GVHD (aGVHD), chronic GVHD (cGVHD), and BK viral relapse (BKVR). Complication-free patients were HCT recipients that were not diagnosed with a complication within 9 months of infusion. Medical records of each patient in the complication-free cohort were reviewed to confirm that none of these patients experienced CMV or EBV reactivation, veno occlusive disease (VOD), hemolytic uremic syndrome (HUS) or serious infections after engraftment. If one of those complications were experienced, that patient was categorized as COMP.

[0132] In cohort 2, patients were grouped into those that succeeded taper (Success) and those that failed taper (Failure). Failure was characterized as chronic GVHD of late acute GVHD occurrence either during taper or after ending taper.

[0133] Complication Diagnosis Criteria

[0134] BK Viral Reactivation was identified using a commercial real-time PCR assay on urine and blood samples (Eurofins Viracor, LLC, Lenexa, KS). Relapse, graft failure, and GVHD: Medical records, including medical oncology notes, clinical laboratory tests and pathology reports were retrospectively reviewed. Data was captured in a clinical management system. Patients were monitored for relapse, graft failure, GVHD, infections and other complications during and after hematopoietic stem cell transplantation. Monitoring involved clinical histories, physical examinations, and regular laboratory assessments to track patients' progress and identify any potential complications. Supplementary biopsies were conducted to confirm diagnoses of GVHD or relapse. Data collection was conducted retrospectively.

[0135] Sample Collection

[0136] Cohort 1 (DFCI) - Blood samples were collected through standard venipuncture in ethylenediaminetetraacetic acid tubes (Becton Dickinson, reference No. 366643). Plasma was extracted through blood centrifugation (2,000 rpm for 10 min using a Beckman Coulter Allegra 6R centrifuge) and stored in 0.5- to 2-mL aliquots at -80 °C. Plasma samples were shipped from the Dana-Farber Cancer Institute to Cornell University on dry ice.

[0137] Cohort 2 (MCC) - Blood samples were collected through standard venipuncture in CPT- Na citrate tubes (Becton Dickinson, reference No. 362753). Plasma was extracted through double blood centrifugation (first: 1,500 g for 20 min at room temperature, second: 609 g for 5min at room temperature) and stored in ImL aliquots at -80 °C. Plasma samples were shipped from the Moffitt Cancer Center to Cornell University on dry ice.

[0138] Sample processing and sequencing.

[0139] Samples were processed using methods described previously(23), with the following modifications: sequencing was done on an Illumina NextSeq or NovaSeq and reads were trimmed to 61 base pairs.

[0140] Sample quality filtering.

[0141] Quality filtering was done as previously described(23), with the following modifications: Samples were removed from analysis if either the intron to exon ratio was greater than 3, if a sample had less than 10,000,000 reads, if the alignment rate was less than 10%, if the duplication rate was greater than 75%, or if the 5’ -3’ read alignment ratio bias was greater than 3 z-scores from the mean.

[0142] Correlation to healthy stem cell donors

[0143] Average Pearson correlation was calculated between each sample and each healthy stem cell donor sample using the top 500 most variable genes after variance stabilizing transformation in the Engraftment and 1, 2, 3, and 6 month follow up time point. For healthy stem cell donor samples, correlation to self was excluded from the average calculation. Variance stabilizing transformation was performed using the DESeq2 NarianceStabilizingTransformation function.

[0144] Distance to healthy stem cell donors

[0145] We introduce a method to measure distance of cell type of origin profiles in cfRNA that accurately reflects the unique characteristics of cell type profiles by considering two critical sources of variation: the inherent variability among different cell types and the correlation patterns between them. This approach ensures that the metric: 1) down-weights the difference for cell types with higher variability and 2) avoids overestimating differences caused by cell types that co-vary within the control population. The process involves several steps: 1) Computing the mean and covariance matrix for cell type fractions from healthy stem cell donors. Given the typical scenario where the number of cell types exceeds the sample size of controls, we apply the nearPD function from the R Matrix package to project the covariance matrix to the nearest positive-definite matrix. 2) Computing the square root of the Mahalanobis distance between thecell type fraction of the query sample and the mean of the healthy stem cell donors under the inverse of the projected covariance matrix:where x E Rkdenotes the cell type fraction of the query sample across k cell types, ylt...,ynE Rkdenote the cell type fraction of control samples, p, and are the mean and projected covariance matrix of samples y. The implementation is based on the Mahalanobis function provided by the R matrixcalc package.

[0146] Cell counts and cell type of origin deconvolution.

[0147] Cell counts were obtained from blood samples collected at the same time as the plasma for cfRNA analysis. Blood samples were collected in EDTA containing tubes and subsequently analyzed at the clinical laboratories of Brigham and Women’s Hospital or Dana-Farber Cancer Institute using automated image-based cell counters. Automated cell counters provided WBC differential counts as well as complete blood cell counts.

[0148] Deconvolution of cfRNA to obtain cell type fractions was done using BayesPrism(v3.0)(REF) using Tabula Sapiens as a single cell reference.

[0149] For lymphocytes, the sum of deconvolution results for B cells, T cells, plasma cells, NK cells, and innate lymphoid cells was used. For leukocytes, the sum of deconvolution results for all lymphocytes, along with neutrophils, basophils, mast cells, macrophages, mature conventional dendritic cells, and monocytes was used.

[0150] Metagenomic analysis

[0151] Reads unaligned to the human genome were taxonomically assigned using Kraken2 and the standard database with default parameters(dS). Species abundances were calculated using Braken and formatted using a custom Python script(59). Counts per million (CPM) reads were calculated from Braken outputs by dividing the new estimated counts by the number of sequenced reads and multiplying the value by one million. Log 10 scaled CPM values were correlated to human alignment rate and cfRNA quantification estimate. The following nonmammalian viral families were removed from analysis: Siphoviridae , Baculoviridae, Tospoviridae, Alphafl exiviridae Virgaviridae , Bromoviridae , Myoviridae, Sol moviridae. Genomoviridae , Endornaviridae, Tombusviridae.

[0152] Differential cell type of origin and transcript abundance analysis

[0153] Differentially abundant RNAs were calculated using DESeq2(36), genes with a baseMean < 5 were removed, and p-values were adjusted using the Benjamini -Hochberg method. For cGVHD, sampling time point was included as a covariate. Adjusted p-values and log2 fold changes were averaged and used as input for Qiagen Ingenuity Pathway Analysis.

[0154] Differential cell type of origin analysis was performed using unpaired Wilcoxon tests and p-values were adjusted using the Benjamini -Hochberg method.

[0155] Longitudinal differential abundance analysis was performed cursing DESeq2(36), and p- values were adjusted using the Benjamini -Hochberg methods. Sampling timepoint and day relative to aGVHD diagnosis were used as covariates.

[0156] GVHD diagnosis and organ involvement

[0157] Acute and chronic GVHD were diagnosed and graded using NIH consensus criteria. We also categorized patients with acute and chronic GVHD based on involvement of skin, liver, and GI tract. Patients were classified into three groups for each organ: GVHD+ with organ involvement, GVHD+ without organ involvement, GVHD-, and compared with healthy stem cell donors. GVHD+ patients had either acute and / or chronic GVHD. Cell type of origin results for relevant cell types were compared using unpaired Wilcoxon tests. For liver involvement, hepatocyte and intrahepatic cholangiocyte-derived cfRNA were compared. For skin involvement, melanocyte and keratinocyte-derived cfRNA were compared. For GI involvement, salivary gland, intestinal stem cell, acinary cell of salivary gland, intestinal epithelial cell, intestinal tuft cell, goblet cell, and duodenum glandular cell derived cfRNA were compared.

[0158] Machine learning modeling

[0159] For our machine learning classification models, we first normalized the counts using TMM CPM. We then conducted 101 repeated stratified train-test splits without replacement to train and evaluate a LASSO regression model, assigning 80% of the samples to training and 20% to testing in each split. For the immunosuppression taper analysis, we used stratified group splitting to ensure that all samples from the same patient were placed in either the training or testing set, thereby preventing data leakage.

[0160] Quantification and statistical analyses

[0161] The programming language R (v4.1.0) was utilized for all statistical analyses. Statistical significance was assessed through two-sided Wilcoxon signed-rank tests and Mann-Whitney Utests, unless specified otherwise. Machine learning algorithms were trained using the Caret R package and pipelines were run using the Snakemake workflow management system. In boxplots, boxes denote the 25th and 75th percentiles, the band within the box signifies the median, and whiskers extend to 1.5 times the interquartile range of the hinge. The alignment of all sequencing data was performed against the GRCh38 Gencode v38 Primary Assembly, with feature counting conducted using the GRCh38 Gencode v38 Primary Assembly Annotation.

[0162] Data availability

[0163] De-identified RNA-seq count matrices have been uploaded to the NCBI (National Center for Biotechnology Information) GEO (Gene Expression Omnibus) database and will be publicly available upon publication (GSE268164). All code has been deposited on GitHub (https: / / github.com / conorloy / cfmaHSCT) and will be available upon publication.

Claims

WHAT IS CLAIMED IS:

1. A method of monitoring a subject undergoing hematopoietic stem cell transplantation (HSCT), the method comprising: a) obtaining cell-free RNA from at least one biological sample of the subject, wherein the biological sample is a plasma sample or serum sample; b) determining: i) a profile of a panel of biomarkers in the cell-free RNA; ii) cfRNA abundances; iii) distance to health; iv) cell type of origin of the cfRNA; and / or v) cell turnover; and c) determining the subject’s health status based on at least one of the modalities listed in step b).

2. The method of claim 1, wherein the at least one biological sample of the subject is obtained at the day of infusion.

3. The method of claim 1, wherein the at least one biological sample of the subject is obtained after the day of infusion.

4. The method of claim 3, wherein the at least one biological sample of the subject is obtained less than 36 months after the day of infusion.

5. The method of claim 4, wherein the at least one biological sample of the subject is obtained less than less than 2 months, less than 3 months, less than 6 months, or less than 12 months after the day of infusion.

6. The method of claim 1, wherein the at least one biological sample of the subject comprises a set of longitudinally collected biological samples of the subject, wherein the samples are obtained at preconditioning, day of infusion, engraftment, and / or one or more post engraftment timepoints.

7. The method of claim 6, wherein the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint being at less than 2 months after the day of infusion, less than 3 months after the day of infusion, less than 6 months after the day of infusion, less than 12 months after the day of infusion, less than 24 months after the day of infusion, and / or less than 36 months after the day of infusion.

8. The method of claim 6, wherein the subject is receiving immunosuppression treatment and will be tapered off the immunosuppression treatment at an initial tapering date, the one or more post engraftment timepoints comprise multiple timepoints, and at least one timepoint occurs prior to the initial tapering date, less than three months after the initial tapering date, and / or less than six months after the initial tapering date.

9. The method of any one of the previous claims, wherein the step of determining a profde of a panel of biomarkers comprises performing nucleotide sequencing of the cell-free RNA.

10. The method of any one of the previous claims, wherein the step of determining a profde of a panel of biomarkers comprises measuring the levels of RNAs corresponding to the biomarkers in the panel.

11. The method of any one of the previous claims, wherein the determining the subject’s health status based on the profde of the panel of biomarkers comprises comparing the profde of the panel of biomarkers in the cell free RNA from the biological samples to a respective reference profde of the panel of biomarkers at the respective time points.

12. The method of any one of the previous claims, wherein the subject’s health status comprises treatment response, treatment toxicity, likelihood of developing graft versus host disease (GVHD), and / or likelihood of developing complications due to the HSCT.

13. The method of claim 12, wherein the complications due to the HSCT comprise microbial / bacterial infection and / or viral infection.

14. The method of claim 12 or 13, wherein determining the subject’s health status treatment response, further comprises calculating average correlation of cfRNA abundances for each sample with respect to reference profile of the panel of biomarkers.

15. The method of claim 12 or 13, wherein the subject’s recovery health status further comprises the cell turnover values for neutrophils, leukocytes, and lymphocytes using the cfRNA cell type of origin fraction, total cfRNA concentration, and blood cell counts.

16. The method of claim 15, wherein the total cfRNA concentration is calculated by: normalizing the sample’s cDNA library concentration by the volume of input plasma and number of PCR cycles performed during library preparation to get the estimated total cfRNA concentration; multiplying the estimated total cfRNA concentration by the cfRNA cell type of origin fraction for a cell type of interest, resulting in the concentration of cfRNA derived from that cell type; and dividing the concentration of cfRNA derived from the cell type of interest by the blood cell counts for that respective cell type, providing an estimate of cell turnover in arbitrary units (a.u.), wherein the estimated cell turnover is comparable between cfRNA samples.

17. The method of claim 12 or 13, further comprising quantifying distance to health, profiling cell types of origin, measuring immune cell turnover, and performing RNA biomarker analysis.

18. The method of claim 17, further comprising characterizing GVHD organ involvement, wherein the characterizing GVHD organ involvement comprises: comparing the cell type of origin fraction of relevant cell types for each organ across the reference profile; and correlating organ specific damage in subjects with GVHD.

19. The method of claim 17, wherein measuring immune cell turnover comprises comparing the cfRNA profile of leukocyte, lymphocyte, and neutrophil turnover in the subject sample to the reference profile and correlating an increase in leukocyte, lymphocyte, and neutrophil turnover toa prediction of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject.

20. The method of claim 17, wherein performing RNA biomarker analysis comprises comparing the cfRNA profde of the subject sample to the reference profile at the pre-determined timepoint and correlating an increase in interferon alpha / beta and gamma signaling as an early indication of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject.

21. The method of claim 17, wherein quantifying distance to health comprises comparing the cfRNA profile in the subject sample to the reference profile using either a correlation in RNA abundances or a mahalanobis distance of the cell type of origin and correlating an increase in distance to a prediction of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject.

22. The method of claim 17, wherein the profiling cell types of origin comprises comparing the cfRNA profile of leukocyte, lymphocyte, neutrophil, dendritic cell, and / or endothelial cell in the subject sample to the reference profile and correlating an increase in cfRNA abundance of leukocyte, lymphocyte, neutrophil, dendritic cell, and / or endothelial cell cfRNA to a prediction of treatment toxicity, likelihood of developing GVHD, and / or likelihood of developing complications due to the HSCT in the subject.23.. The method of any one of the previous claims, wherein the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: NIPAL3, MPO, CD9, IDS, TMEM159, XYLT2, OSBPL5, APBA2, EDC4, SNX29, VPS13D, HIPK2, IDI1, PDK3, IFI35, INPP5A, GNB5, DGCR2, NDST1, ATP6AP1, FBXW11, IRAG1, ATP2A3, SCARF1, DGKD, 0STM1, CDC14B, SSH1, TXLNA, KIF3C, EPS15, B4GALT1, OAS1, CDIP1, SLC9A1, STRN4, PITPNM2, CRAT, ACOT7, KCNK6, PSMD8, MARCHF2, POLR2E, SNUB, MYH9, MLC1, CTSG, NIN, PYGB, PSMA7, CDS2, CST3, UBL4A, SLC25A15, USP10, USP31, H0MER2, BPNT2, PPP6R1, PDCD5, PTPRS, DDX49, CLIP2, TRIM14, DNM1, VSIR, ARHGAP21, FBXL20, FTSJ3, AL0X12, ABCC3, GRPEL1, ZBTB16, PANXI, VWF, MLEC, OAS3, OAS2, EIF2B1, CCND3, FBRSL1, DBN1, CISH, ABHD14B,KANSU, L0XL3, 0DC1, WIPF1, PLEK, NIDI, GBP1, Clorfl98, IFIT3, IFIT2, ACAT2, SNX19, PLXDC2, TNFSF10, ZMIZ2, DDX39A, NMI, MRS2, TMX4, DSTN, MMP24OS, SIN3B, VIL1, GNAZ, TPST2, TST, KRI1, SLC44A2, CDKN2D, BCL2L2, CLIP1, SLC35D2, PRR7, SLC6A6, DIAPH1, FMO5, PRKAB2, RAFI, DLG4, CTNNBL1, BTBD2, PDZD2, CTIF, IRAK2, SORT1, UBAC2, OASL, CD36, CCT7, KIAA0513, APPL2, TTYH3, TLN1, TPMT, HADH, PCSK6, KIFC3, EIF4A3, CERS2, PYCR2, VGLL4, CTDSPL, TCTA, PLAC8, TIFA, FBXL17, SNX12, DOCK5, NFIB, ENDOD1, FERMT3, MTMR12, NBAS, GBP5, PSD3, BRAF, SKI, GPAT4, FUR, FBXW5, PTGIR, KALRN, SLC37A1, GAB3, PCSK7, ORAI2, MGAT4B, EMC 10, SELENON, H3-3A, COL6A3, AIMP1, ANKRD33B, CITED2, GALNT10, CRYL1, TTC7B, CMTM5, WBP1L, GOLM2, RBPMS2, SAMD14, FN3K, TPM4, GATAD2A, ANKRD11, TRAPPC9, SPINT2, ECU, MAP4K2, MLKL, CD2BP2, GP9, REPS2, CHD3, CRTAP, RGS19, FIBP, DCAKD, ABLIM3, GPR137, ATP2A2, RAB1B, PACS1, PPP2R2D, MCTP1, TOM1L2, RUFY1, SAMD9L, NR2C2, CD151, NDUFAF3, GP5, PACS2, TMEM64, COA4, UBA7, MOB2, ADI1, RBM10, EWSR1, PTTG1IP, CLDN5, SNN, BTBD6, SLC24A3, ZBTB37, METTL7A, IRF7, BRCC3, IFIT1, RASA3, CD300E, PDE2A, TPCN1, ISG15, PEAR1, PLSCR1, PARVB, HACD4, TMEM63A, VKORC1L1, MAML3, PDLIM7, FLNA, SRC, AD ARBI, HTT, MAP3K5, MYO5A, PSAP, MYO1C, PAPSS2, UNC13B, CDC42BPB, PLXNB3, GRK5, AGP ATI, LY6G6F, HLA-C, HLA-E, PRR13, DOK6, HLA-H, HLA-A, RNF208, FGD5-AS1, FAM30A, GPX1, HLA-B, LINC00674, PEDS1, EIF6, CCDC71L, MTRNR2L8, ITGB3, RP 11-401.2, IKBKG, RN7SL1, and RPl l-147L13.il.

24. The method of any one of the previous claims, wherein the panel of biomarkers comprises one, more, or all of RAFI, RFC2, ERCC6L2, IDI1, FIBP, EDC4, FBXL20, KANSL1, and P0P4.

25. The method of claim 24, wherein the panel of biomarkers comprises at least one of RAFI, IDI1, EDC4, KANSL1, RFC2, P0P4, and ERCC6L2.

26. The method of any one of the previous claims, wherein the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: RHD, CCL21, TMEM176B, CTD-2643I7.4, RP11-329A14.1, ESPN, CLEC10A, MTCO3P12, CD8B, HTATSF1P2, KCNQ1OT1, Y RNA, MSC, ANKRD36BP2, EVPL, TNNT1, TSPAN2,MYH1 1, AT0H8, TEAD4, LINS1, BRSK1, Y RNA, ZBED6, TMEM205, RNF I 82, MTRNR2L1, PVALB, RP11-678G14.3, ALB, RNA5SP145, MY0M2, RPL9P9, MYCL, RNA5SP149, CDC42EP1, ABO, SPOCD1, GALNT3, DDX11L10, PCDH17, RP11-452G18.2, DNAJB5, SMIM24, MS4A1, IGHA1, MTND4P12, DDX11L2, HLA-J, Y RNA, Y RNA, C19orf33, MT-RNR2, USP54, GALT, LINC00504, ARL6IP6, MTRNR2L8, RN7SL767P, PDZKHP1, KRT1, COL4A1, PROK2, SLC15A2, H2BC8, Y RNA, SBK1, DEPPI, BAHCC1, HBG2, CXCL9, TOX2, RALY-AS1, SNAI2, ALPK3, STAP1, ERAP2, PLA2G7, GSTM5, CD79A, SPRY4, IGHD, MYOIO, AC114271.2, PARVA, CYP2E1, CCR2, STK33, IGHV1-3, XXbac-BPG299F13.17, ST20-AS1, Y RNA, CFD, RP11-712B9.2, LILRA4, MEIS3P1, SRGAP1, HID1, and REEPL27. The method of any one of the previous claims, wherein the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: RHD, CCL21, TMEM176B, CTD-2643I7.4, RP11-329A14.1, ESPN, CLEC10A, MTCO3P12, CD8B, HTATSF1P2, KCNQ1OT1, Y RNA, MSC, ANKRD36BP2, EVPL, MYH11, AT0H8, TEAD4, Y RNA, ZBED6, RNF182, MTRNR2L1, RNA5SP145, RPL9P9, MYCL, RNA5SP149, CDC42EP1, SPOCD1, GALNT3, DDX11L10, PCDH17, DNAJB5, MTND4P12, C19orf33, MT-RNR2, PDZKHP1, KRT1, SLC15A2, DEPPI, BAHCC1, ERAP2, IGHD, MYOIO, CFD, RP11-712B9.2, SRGAP1, RNF152, RN7SL736P, RN7SL665P, and TRIP13.

28. The method of any one of the previous claims, wherein the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: RHD, CCL21, TMEM176B, CTD-2643I7.4, RP11-329A14.1, ESPN, CLEC10A, MTCO3P12, HTATSF1P2, KCNQ1OT1, ANKRD36BP2, RNF182, and RNA5SP149.

29. The method of any one of the previous claims, wherein the panel of biomarkers comprises RHD and CCL21.

30. The method of any one of claims 26-29, wherein the subject is receiving immunosuppression treatment and will be tapered off the immunosuppression treatment, and the panel is predictive of immunosuppression treatment tapering failure.31 . The method of any one of the previous claims, wherein the panel of biomarkers comprises one, more, or all of Major Histocompatibility Complex, Class I - E, C, A, and B (HLA-E, HLA-C, HLA-A, and HLA-B, respectively) genes.

32. The method of any one of the previous claims, wherein the panel of biomarkers comprises at least one histone gene.

33. The method of claim 32, wherein the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: Hl-1, Hl-2, Hl-3, Hl-4, Hl-5, Hl-6, H2AC1, H2AC4, H2AC6, H2AC7, H2AC8, H2AC11, H2AC12, H2AC13, H2AC14, H2AC15, H2AC16, H2AC17, H2AC18, H2AC19, H2AC20, H2AC21, H2AC25, H2BC1, H2BC3, H2BC4, H2BC5, H2BC6, H2BC7, H2BC8, H2BC9, H2BC10, H2BC11, H2BC12, H2BC13, H2BC14, H2BC15, H2BC17, H2BC18, H2BC21, H2BC26, H2BC12L, H3C1, H3C2, H3C3, H3C4, H3C6, H3C7, H3C8, H3C10, H3C11, H3C12, H3C13, H3C14, H3C15, H3-4, H4C1, H4C2, H4C3, H4C4, H4C5, H4C6, H4C7, H4C8, H4C9, H4C11, H4C12, H4C13, H4C14, H4C1 5H1-0, Hl -7, Hl -8, Hl -10, H2AZ1, H2AZ2, MACROH2A1, MACROH2A2, H2AX, H2AJ, H2AB1, H2AB2, H2AB3, H2AP, H2AL1Q, H2AL3, H2BK1, H2BW1, H2BW2, H2BW3P, H2BN1, H3-3A, H3-3B, H3-5, H3-7, H3Y1, H3Y2, CENPA, and H4C16.

34. The method of claim 32 or 33, wherein the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: Hl-1, Hl-2, Hl-3, Hl-4, Hl-5, Hl-6, H2AC1, H2AC4, H2AC6, H2AC7, H2AC8, H2AC11, H2AC12, H2AC13, H2AC14, H2AC15, H2AC16, H2AC17, H2AC18, H2AC19, H2AC20, H2AC21, H2AC25, H2BC1, H2BC3, H2BC4, H2BC5, H2BC6, H2BC7, H2BC8, H2BC9, H2BC10, H2BC11, H2BC12, H2BC13, H2BC14, H2BC15, H2BC17, H2BC18, H2BC21, H2BC26, H2BC12L, H3C1, H3C2, H3C3, H3C4, H3C6, H3C7, H3C8, H3C10, H3C11, H3C12, H3C13, H3C14, H3C15, H3-4, H4C1, H4C2, H4C3, H4C4, H4C5, H4C6, H4C7, H4C8, H4C9, H4C11, H4C12, H4C13, H4C14, and H4C15.

35. The method of any one of claims 32-34, wherein the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: H2BC21, Hl-2, H2BC4, H2AC6, Hl-4, H2BC12, H3C10, and Hl-5.

36. The method of any one of the previous claims, wherein the panel of biomarkers comprises one, more, or all of the genes selected from the group consisting of: RPA2, CLSPN, CDCA8, CDC20, KIF2C, NASP, USP1, PSRC1, ANP32E, CKS1B, NUF2, NEK2, DTL, CENPF, LBR, EX01, RRM2, CENPA, MSH2, BUB1, CKAP2L, MCM6, CDCA7, HJURP, MCM2, SMC4, ECT2, SLBP, TACC3, CENPE, HMGB2, CENPU, CDC25C, HMMR, GMNN, TTK, CASP8AP2, ANLN, RFC2, CDCA2, MCM4, CCNE2, DSCC1, ATAD2, CKS2, TUBB4B, CDK1, KIF20B, KIF11, HELLS, MKI67, RRM1, E2F8, CKAP5, FEN1, P0LD3, RAD51AP1, NCAPD2, CDCA3, CBX5, PRIM1, TMPO, GAS2L3, UNG, CKAP2, G2E3, DLGAP5, UBR7, RAD51, NUSAP1, WDR76, CCNB2, TIPIN, KIF23, BLM, CTCF, GINS2, PIMREG, AURKB, CDC6, TOP2A, BRIP1, JPT1, BIRC5, TYMS, NDC80, UHRF1, PCNA, TPX2, UBE2C, AURKA, CHAF1B, CDC45, MCM5, RANGAP1, GTSE1, POLA1.

37. The method of claim 36, wherein the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: NASP, USP1, ANP32E, CENPF, RRM2, MCM6, SMC4, SLBP, TACC3, CENPE, HMGB2, ATAD2, TUBB4B, KIF20B, MKI67, CKAP5, NCAPD2, CBX5, TMPO, CKAP2, CTCF, TOP2A, PCNA, TPX2, RAN GAPE38. The method of either claim 36 or 27, wherein the panel of biomarkers comprises one, more, or all histone genes selected from the group consisting of: ANP32E, CENPF, HMGB2, TUBB4B, MKI67.

39. A method for monitoring a microbiome of an HSCT recipient subject for infection, the method comprising: a) obtaining cell-free RNA from at least one biological sample of the subject, wherein the biological sample is obtained at the day of infusion or less than 36 months, or 24 months, or 12 months, or less than 6 months, or less than 3 months, or less than 2 months after the day of infusion, wherein the biological samples are plasma samples or serum samples; b) determining a profile of a panel of metagen omic biomarkers in the cell -free RNA; c) determining an inflammatory disease status in the subject based on the profile of the panel of metagenomic biomarkers; andd) removing background total viral counts of all non-mammalian infecting viral families from each sample.

40. The method of claim 39, wherein the at least one biological sample of the subject comprises a set of longitudinally collected biological samples of the subject at preconditioning, day of infusion, engraftment, and / or one or more post infusion timepoints.

41. The method of claim 39 or 40, wherein the step of determining a profile of a panel of biomarkers comprises performing nucleotide sequencing of the cell-free RNA.

42. The method of claim 39, 40, or 41, wherein the step of determining a profile of a panel of biomarkers comprises measuring the levels of RNAs corresponding to the biomarkers in the panel.

43. The method of any one of claims 39-42, wherein monitoring the microbiome of the subject based on the profile of the panel of biomarkers comprises comparing the profile of the panel of biomarkers in the cell free RNA from the biological sample to a reference profile of the panel of biomarkers.

44. The method of claim 43, wherein the reference profile of the panel of biomarkers is from a healthy donor or an HSCT subject without complications.

45. The method of any one of claims 39-44, wherein the panel of biomarkers comprises those for at least one of the Flaviviridae, Anelloviridae, Retroviridae, and Herpesviridae viral families.

46. The method of any one of claims 39-45, wherein the abundance of prokaryotic reads are quantified for the cfRNA from one or a combination of longitudinal time points and compared to the reference abundance of prokaryotic reads.

47. The method of any one of claims 39-46, wherein a correlation is made when the subject sample comprises elevated levels of biomarkers when compared to the reference panel that the subject may be experiencing complications.

48. The method of any one of the previous claims, wherein at least two cfRNA measurements are used in determining the subject’s recovery trajectory and wherein the at least two cfRNA measurements used consist of the parameters included in the group consisting of microbial kingdom counts, genes counts, viral family counts, distance to health, and cell turnover.

49. The method of 48, wherein the cfRNA measurements used consist of microbial kingdom counts, genes counts, viral family counts, distance to health, and cell turnover.

50. The method of claim 49, wherein the set of longitudinally collected biological samples includes more than one predetermined timepoints after infusion.

51. The method of claim 49, wherein the set of longitudinally collected biological samples comprises a biological sample collected on the day of infusion.

52. The method of either claim 50 or 51, wherein the more than one predetermined timepoint comprises monthly intervals after the day of infusion (1 Month, IM; 2 Month, 2M; 3 Month, 3M; 6 Month, 6M).

53. The method of claim 52, wherein the profde of biomarkers at an early time point (e.g., 1, 2 or 3 months) post infusion is indicative of the likelihood of GVHD.

54. The method of any one of the previous claims, wherein the biological sample is a plasma sample.

55. The method of any one of the previous claims, wherein the biological sample is a serum sample.

56. The method of any one of the previous claims, wherein the subject is a mammal.

57. The method of any one of the previous claims, wherein the subject is a human.

58. The method of any one of the previous claims, wherein the profile of the panel of biomarkers comprises subject-specific gene expression of the biomarkers in the panel.

59. The method of any one of the previous claims, wherein the cell-free RNA is substantially free of contaminant RNA.

60. The method of any one of the previous claims, wherein the subject’s recovery trajectory is used in determining treatment for the subject.

61. The method of any one of the previous claims, wherein test performance of the trained classification model is measured through one or more of accuracy, sensitivity, specificity, and / or area under the receiver operating characteristic curve (ROC-AUC).

62. The method of any one of the previous claims, wherein the inflammatory syndrome status is determined at a greater than 80%, 85%, 90%, 91%, 92%, 93%, 94%, or 95% accuracy.

63. The method of any one of the previous claims, wherein the inflammatory syndrome status is determined at a greater than 80%, 85%, 90%, 91%, 92%, 93%, 94%, or 95% sensitivity.

64. The method of any one of the previous claims, wherein the inflammatory syndrome status is determined at a greater than 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, or 95% specificity.

65. The method of any one of the previous claims, wherein the trained classification model comprises one or more models selected from the group consisting of generalized linear models with Ridge and LASSO feature selection (GLMNETRIDGE and GLMNETLASSO), support vector machines with linear and radial basis function kernel (SVMLin and SVMRAD), random forest (RF), random forest ExtraTrees (EXTRATREES), neural networks (NNET), linear discriminant analysis (LDA), nearest shrunken centroids (PAM), C5.0 (C5), k-nearest neighbors (KNN), naive bayes (NB), CART (RPART), generalized linear model (GLM), and greedy forward search algorithm (GFS).

66. The method of any one of the previous claims, wherein the method further comprises treating the subject based on the recovery trajectory.

67. The method of claim 66, wherein the treatment for aGVHD comprises administering an immunosuppressive to the subject.

68. The method of claim 66, wherein the treatment for complications due to HSCT comprises administering treatment regimens known to combat the complication.

Citation Information

Patent Citations

  • Methods for detecting tissue damage, graft versus host disease, and infections using cell-free DNA profiling

    US20230257822A1