Composite scoring methods
The integration of clinical symptoms and serum biomarkers in the MAGIC composite score addresses the limitations of current GVHD risk stratification systems by accurately predicting treatment response and non-relapse mortality, enabling personalized treatment strategies.
Patent Information
- Application Number
- PCT/US2025/026225
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-26
- Filing Date
- 2025-04-24
- Publication Date
- 2025-10-30
AI Technical Summary
Current GVHD risk stratification systems, such as the Minnesota system, lack a low-risk stratum necessary for treatment de-escalation and do not integrate both clinical and laboratory parameters, leading to suboptimal treatment strategies for steroid-refractory GVHD.
A computer-implemented method using a Mount Sinai Acute GVHD International Consortium (MAGIC) composite score that integrates clinical symptoms and serum biomarkers, employing a trained machine learning model and CART algorithm to assign a composite score, categorizing patients into three risk strata: low, intermediate, and high risk, improving prognostic accuracy.
The MAGIC composite score accurately predicts treatment response and non-relapse mortality, enabling personalized treatment strategies and minimizing steroid exposure by identifying patients with very low or high risk, thus improving clinical outcomes.
Smart Images

Figure IMGF000017_0001 
Figure IMGF000029_0001 
Figure IMGF000030_0001
Abstract
Description
COMPOSITE SCORING METHODSRELATED APPLICATIONS
[0001] This application claims the priority benefit of U.S. Provisional Patent Application No. 63 / 639,177, filed on April 26, 2024, said application is incorporated herein by reference in its entirety.RESEARCH STATEMENT
[0002] This invention was made with government support under grant number P01CA039542 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND
[0003] Recent advances in graft-versus-host disease (GVHD) prophylaxis shift towards less severe phenotype of GVHD but a significant minority of few patients develop steroid- refractory GVHD that leads to unfavorable outcomes. A risk stratification system that includes a low risk stratum is thus highly desirable. The inventors have now integrated these serum biomarkers with a new Manhattan risk model that uses clinical symptoms to establish Mount Sinai Acute graft-vs-host disease (GVHD) International Consortium (MAGIC) composite scores. These composite scores more accurately predict treatment response, and the risk of nonrelapse mortality (NRM) compared to other models and may better identify patients who could benefit from personalized primary treatment strategies.SUMMARY
[0004] One aspect of the disclosure is a computer-implemented method to assign a Mount Sinai Acute graft-vs-host disease (GVHD) International Consortium (MAGIC) composite score for a subject in need thereof, comprising: i) receiving, by a computing system, a first data set from the subject comprising clinical symptoms; ii) receiving, by a computing system, a second data set from the subject comprising serum biomarker data; iii) processing, by the computing system, the first data set using a trained machine learning model configured to generate a Manhattan risk model classification; iv) processing, by the computing system, the second data set to generate a serum biomarker score; and iv) assigning, by the computing system, a MAGIC composite score for the subject based on the Manhattan risk model classification and the serum biomarker score.
[0005] In an aspect, the trained machine learning model is trained on a training data set comprising clinical symptoms and non-relapse mortality data, wherein the training data set is evaluated using a classification and regression tree (CART) algorithm to generate a Manhattan risk model classification.
[0006] In an aspect, the Manhattan risk model classification is low risk, intermediate risk, or high risk.
[0007] In an aspect, the serum biomarker score is generated by applying a ST2 and REG3a algorithm to the second data set to generate a MAGIC algorithm probability (MAP) value.
[0008] In an aspect, if the MAP value is less than about 0.141, the biomarker score is 1. In an aspect, if the MAP is between about 0.141 and about 0.290, the biomarker score is 2. In an aspect, if the MAP is greater than about 0.290, the biomarker score is 3.
[0009] In an aspect, if the subject receives a Manhattan low risk or intermediate risk classification and a biomarker score of 1, the subject is assigned a MAGIC composite score of 1.
[0010] In an aspect, if the subject receives a Manhattan low risk classification and a biomarker score of 2, the subject is assigned a MAGIC composite score of 1. In an aspect, if the subject receives a Manhattan low risk classification and a biomarker score of 3, the subject is assigned a MAGIC composite score of 2. In an aspect, if the subject receives a Manhattan intermediate risk and a biomarker score of 2 or 3, the subject is assigned a MAGIC composite score of 2. In an aspect, if the subject receives a Manhattan high risk classification and a biomarker score of 1 or 2, the subject is assigned a MAGIC composite score of 2. In an aspect, if the subject receives a Manhattan high risk classification and a biomarker score of 3, the subject is assigned a MAGIC composite score of 3.
[0011] In an aspect, the method further comprises determining a prognosis of non-relapse mortality for the subject.
[0012] In an aspect, a MAGIC composite score of 1 correlates to a prognosis of low risk of non-relapse mortality. In an aspect, a MAGIC composite score of 2 correlates to a prognosis of intermediate risk of non-relapse mortality. In an aspect, a MAGIC composite score of 3 correlates to a prognosis of high risk of non-relapse mortality.
[0013] In an aspect, the method further comprises quantitatively assessing severity of GVHD in the subject.
[0014] In an aspect, the method further comprises monitoring treatment response by recalculating the MAGIC composite score after administration of a therapeutic intervention.
[0015] In an aspect, the therapeutic intervention comprises steroids, extracorporeal photopheresis (ECP), or biological therapies.
[0016] In an aspect, the subject is classified as a responder or a non-responder.
[0017] In an aspect, the MAGIC composite score is used to automatically adjust a treatment regime.
[0018] In an aspect, the adjustment is done via integration with an electronic health record system.
[0019] One aspect of the disclosure includes a computing system for assigning a MAGIC composite score for a subject in need thereof comprising: one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the computing system to: i) receive a first data set from the subject comprising clinical symptoms and nonrelapse mortality data; ii) receive a second data set from the subject comprising serum biomarker data; iii) apply a trained machine learning model to the first data set, wherein the model is trained on clinical symptoms and non-relapse mortality data to generate a Manhattan risk model classification for the subject; iv) generate a serum biomarker score for the second data set; and v) assign a MAGIC composite score for the subject based on the Manhattan risk model classification and the serum biomarker score.BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Aspects of the present disclosure will now be described, by way of example only, with reference to the attached Figures, wherein:
[0021] FIGs. 1A - 1C depict the association of AA scores on six-month NRM by Manhattan risk strata in the training cohort. FIG. 1A is a graph showing six-month cumulative incidence of NRM by Manhattan low risk. AA1 : 3.1% (95% CI, 1.5% to 5.5%), AA2:12.1% (95% CI, 6.6% to 19.4%), AA3: 27.8% (95% CI, 14.3% to 43.0%). FIG. IB is a graph showing six- month cumulative incidence of NRM by Manhattan intermediate risk. AA1 : 8.7% (95% CI, 5.5% to 12.6%), AA2: 18.4% (95% CI, 12.2% to 25.7%), AA3: 29.7% (95% CI, 19.7% to 40.4%). FIG. 1C is a graph showing six-month cumulative incidence of NRM by Manhattan high risk. AA1 : 20.4% (95% CI, 10.4% to 12.6%), AA2: 18.4% (95% CI, 12.2% to 25.7%), AA3: 29.7% (95% CI, 19.7% to 40.4%). Pie charts depict the percentage of each AA score. P values for pairwise comparisons were adjusted using the Bonferroni method.
[0022] FIGs. 2A - 2D depict MAGIC composite scores in the validation cohort. FIG. 2A is a graph showing six-month cumulative incidence of nonrelapse mortality (NRM). MAGICcomposite score 1 : 5.7% (95% CI, 3.3% to 8.9%), composite score 2: 28.8% (95% CI, 21.2% to 36.8%), composite score 3: 51.5% (95% CI, 33.1% to 67.2%). FIG. 2B is a graph showing six-month cumulative incidence of relapse. MAGIC composite score 1 : 8.3% (95% CI, 5.4% to 12.0%), composite score 2: 10.8% (95% CI, 6.2% to 16.9%), composite score 3: 6.7% (95% CI, 1.1% to 19.7%). FIG. 2C is a graph showing probability of overall survival (OS) of the six month; MAGIC composite score 1 : 90.6% (95% CI, 86.4% to 93.5%), composite score 2: 64.3% (95% CI, 55.3% to 71.9%), composite score 3: 42.4% (95% CI, 25.6% to 58.3%). Pie charts depict the percentage of each composite score. *P values for pairwise comparisons were adjusted using the
[0023] Bonferroni method. FIG. 2D is a time dependent area under curve (AUC) by receiver operating characteristic for NRM from the time of systemic treatment.
[0024] FIGs. 3A -3C depict long term outcomes of the MAGIC composite scores in the training cohort. FIG. 3A is a graph showing six-month cumulative incidence of nonrelapse mortality (NRM); MAGIC composite score 1 : 6.6% (95% CI, 4.9% to 8.7%), composite score 2: 23.8% (95% CI, 19.4% to 28.5%), composite score 3: 56.3% (95% CI, 43.6% to 67.1%). FIG. 3B is a graph showing six-month cumulative incidence of relapse; MAGIC composite score 1 : 11.5% (95% CI, 9.1% to 14.1%), composite score 2: 12.2% (95% CI, 9.0% to 16.0%), composite score 3: 10.0% (95% CI, 4.4% to 18.5%). FIG. 3C is a graph showing six-month overall survival (OS)
[0025] probability; MAGIC composite score 1 : 89.4% (95% CI, 86.8% to 91.6%), composite score 2: 70.7% (95% CI, 65.5% to 75.3%), composite score 3: 35.1% (95% CI, 24.1% to 46.3%). Pie charts depict the percentage of each composite score. *P values for pairwise comparisons were adjusted using the Bonferroni method.
[0026] FIGs. 4A and 4B are graphs showing AUC for 6-month NRM in each model. The comparisons of an area under curve (AUC) by receiver operating characteristic (ROC) for 6 month NRM among three models only in patients with serum samples in the training (FIG. 4 A) and validation (FIG. 4B) cohorts.
[0027] FIGs. 5A and 5B are graphs showing risk stratification of NRM using the Manhattan risk and MAGIC composite scores in patients with PTCy as GVHD prophylaxis. FIG. 5A is a graph showing six-month cumulative incidence of NRM by Manhattan risk. Low risk: 11.4% (95% CI, 6.7% to 17.5%), Intermediate risk: 10.6% (95% CI, 6.1% to 16.6%), High risk: 19.4% (95% CI, 8.4% to 33.9%). FIG. 5B is a graph showing six-month cumulative incidence of NRM by MAGIC composite scores. Score 1 : 7.7% (95% CI, 4.3% to 12.3%), score 2: 16.3%(95% CI, 8.6% to 26.0%), and score 3: 27.3% (95% CI, 5.8% to 55.2%). Pie charts depict the percentage of each risk. P values for pairwise comparisons were adjusted using the Bonferroni method.
[0028] FIGs. 6A-6C are graphs depicting day 28 ORR of each model in the validation cohort. Day 28 overall response rate (ORR) by the Minnesota risk (FIG. 6A), Manhattan risk (FIG. 6B), and MAGIC composite scores (FIG. 6C); Minnesota standard risk: 73.3%, Minnesota high risk: 49.5%, Manhattan low risk: 77.0%, Manhattan intermediate risk: 69.7%, Manhattan high risk: 48.5%, MAGIC composite score 1 : 79.8%, MAGIC composite score 2: 62.9%, MAGIC composite score 3: 30.3%. *P values for pairwise comparisons were adjusted using the Bonferroni method.
[0029] FIGs. 7A-7F are graphs comparing the 6-month NRM of responders and nonresponders in the validation set.
[0030] FIGs. 8A-8F are graphs comparing the 6-month NRM of responders and nonresponders in the training set.
[0031] FIGs. 9A and 9B are graphs depicting the evaluation of the ability of MAGIC Composite Response (MCR) to reclassify both clinical responders and clinical nonresponders.
[0032] FIG. 10 is a CONSORT diagram.
[0033] FIGs. 11A and 11B are graphs depicting areas under the curve (AUC) of the Minnesota risk. Glucksberg groups (Grades I / II vs. III / IV), and principal component-derived grading (4 grades) in the training cohort (n = 1306). (FIG. 11A) Minnesota risk vs. Glucksberg Grades (FIG. 11B) Minnesota risk vs. PCI 4 grades.
[0034] FIGs. 12A and 12B are graphs depicting AUC of the Minnesota risk and Manhattan risk system. Predicting 6-month NRM (FIG. 12A) in the training (n = 1306) and (FIG. 12B) validation (n = 557) cohorts. Patients without serum samples at treatment were included.
[0035] FIGs. 13A and 13B are graphs depicting the Manhattan risk model on relapse and OS in the validation cohort. FIG. 13A is a graph depicting six-month cumulative incidence of relapse; Manhattan low risk: 13.1% (95% CI, 9.1% to 17.8%), intermediate risk: 9.0% (95% CI, 5.7% to 13.2%), high risk: 11.2% (95% CI, 5.9% to 18.4%). FIG. 13B is a graph depicting six-month overall survival; Manhattan low risk: 87.3% (95% CI, 82.3% to 91.0%), intermediate risk: 80.2% (95% CI, 74.3% to 84.8%), high risk: 57.0% (95% CI, 46.5% to66.1%). Pie charts depict the percentage of each Manhattan risk. *P values for pairwise comparisons were adjusted using the Bonferroni method.
[0036] FIGs. 14A and 14B are graphs depicting NRM in the clinical risk models. Six-month cumulative incidence of NRM by Minnesota and Manhattan risk strata. FIG. 14A is a graph depicting the Training cohort. Minnesota standard risk: 10.2% (95% CI, 8.5% to 12.2%), Minnesota high risk: 36.8% (95% CI, 30.5% to 43.0%); Manhattan low risk: 7.1% (95% CI, 5.1% to 9.5%), Manhattan intermediate risk: 13.9% (95% CI, 11.1% to 16.9%), Manhattan high risk: 37.8% (95% CI, 31.2% to 44.4%). FIG. 14B is a graph depicting the Validation cohort. Minnesota standard risk: 11.0% (95% CI, 8.3% to 14. 1%), Minnesota high risk: 34.4% (95% CI, 25.3% to 43.6%); Manhattan low risk: 7.0% (95% CI, 4.2% to 10.8%), Manhattan intermediate risk: 14.9% (95% CI, 10.6% to 19.9%), Manhattan high risk: 35.8% (95% CI, 26.4% to 45.4%). Pie charts depict the percentage of each clinical risk. *P values for pairwise comparisons were adjusted using the Bonferroni method.DETAILED DESCRIPTIONI. Introduction
[0037] Acute graft-versus-host disease (GVHD) remains a substantial cause of morbidity and non-relapse mortality (NRM) and a major obstacle to successful outcomes after allogeneic hematopoietic cell transplantation (HCT) despite advances in prophylaxis. High doses of systemic steroids are used as first-line treatment for acute GVHD, but approximately 30% of patients develop steroid-refractory GVHD and experience poor outcomes. The long-term outcomes are variable in the patients who initially respond to steroid therapy and thus treatment courses tend to be long, intensive, and complicated by GVHD flares and steroid-related side effects. Treatment for GVHD thus may lead to both under-treatment of some patients and overtreatment of others.
[0038] The maximum severity of acute GVHD correlates with survival outcomes but can only be determined retrospectively and therefore cannot be used to guide treatment in real time. The Minnesota risk system, the only validated risk stratification based on GVHD symptoms at the initiation of treatment, possesses two strata, standard and high, but lacks the low-risk stratum necessary for treatment minimization. While several groups have reported GVHD biomarkers to predict outcomes, no model has yet integrated both clinical and laboratory parameters at treatment onset. The inventors hypothesized that the combination of clinical and biomarker values could create three separate acute GVHD grades with distinct prognoses. The inventorsused the Mount Sinai Acute International GVHD Consortium (MAGIC) database and biorepository to develop and validate a grading system with three strata solely based on clinical symptoms and then developed new MAGIC composite scores that integrate both clinical and biomarker parameters with improved prognostic accuracy.II. Definitions
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the methods described herein belong. Any reference to standard methods refers to the most recent available version of the method at the time of filing of this disclosure unless otherwise indicated.
[0040] For any method disclosed herein that includes discrete steps, the steps may be conducted in any feasible order. And, as appropriate, any combination of two or more steps may be conducted simultaneously.
[0041] All headings are for the convenience of the reader and should not be used to limit the meaning of the text that follows the heading, unless so specified.
[0042] The words "preferred" and "preferably" refer to embodiments of the invention that may afford certain benefits, under certain circumstances. However, other embodiments may also be preferred, under the same or other circumstances. Furthermore, the recitation of one or more preferred embodiments does not imply that other embodiments are not useful and is not intended to exclude other embodiments from the scope of the invention.
[0043] The term "comprises”, and variations thereof do not have a limiting meaning where these terms appear in the description and claims. Such terms will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements.
[0044] By "consisting of' is meant including, and limited to, whatever follows the phrase "consisting of." Thus, the phrase "consisting of' indicates that the listed elements are required or mandatory, and that no other elements may be present. By "consisting essentially of' is meant including any elements listed after the phrase, and limited to other elements that do not interfere with or contribute to the activity or action specified in the disclosure for the listed elements. Thus, the phrase "consisting essentially of' indicates that the listed elements are required or mandatory, but that other elements are optional and may or may not be present depending upon whether or not they materially affect the activity or action of the listed elements.
[0045] The singular form "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. These articles refer to one or to more than one (i.e., to at least one). As used herein, the term "or" is generally employed in its usual sense including "and / or" unless the content clearly dictates otherwise. The term "and / or" means any one or more of the items in the list joined by "and / or". As an example, "x and / or y" means any element of the three-element set {(x), (y), (x, y)}. In other words, "x and / or y" means "one or both of x and y". As another example, "x, y, and / or z" means any element of the seven-element set {(x), (y), (z), (x, y), (x, z), (y, z), (x, y, z)}. In other words, "x, y and / or z" means "one or more of x, y and z".
[0046] Where ranges are given, endpoints include all numbers subsumed within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, 5, etc.). Furthermore, unless otherwise indicated or otherwise evident from the context and understanding of one of ordinary skill in the art, values that are expressed as ranges can assume any specific value or subrange within the stated ranges in different embodiments of the disclosure, to the tenth of the unit of the lower limit of the range, unless the context clearly dictates otherwise. Herein, "up to" a number (for example, up to 50) includes the number (for example, 50). The term "in the range" or "within a range" (and similar statements) includes the endpoints of the stated range.
[0047] Reference throughout this specification to "one aspect,” “one embodiment,” "an aspect,” “an embodiment,” "certain aspects," “certain embodiments,” "some aspects," or “some embodiments,” etc., means that a particular feature, configuration, composition, or characteristic described in connection with the aspect is included in at least one aspect of the disclosure. Thus, the appearances of such phrases in various places throughout this specification are not necessarily referring to the same embodiment of the disclosure. Furthermore, the particular features, configurations, compositions, or characteristics may be combined in any suitable manner in one or more aspects.
[0048] Unless otherwise indicated, all numbers expressing quantities of components, molecular weights, and so forth used in the specification and claims are to be understood as being modified in all instances by the term "about." As used herein in connection with a measured quantity, the term "about" refers to that variation in the measured quantity as would be expected by the skilled artisan making the measurement and exercising a level of care commensurate with the objective of the measurement and the precision of the measuring equipment used. The term "about" as used in connection with a numerical value throughout the specification and the claims denotes an interval of accuracy, familiar and acceptable to a person skilled in the art. In general, such interval of accuracy is + / - 10%. Accordingly, unless otherwiseindicated to the contrary, the numerical parameters set forth in the specification and claims are approximations that may vary depending upon the desired properties sought to be obtained by the present invention. At the very least, and not as an attempt to limit the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.
[0049] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the invention are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. All numerical values, however, inherently contain a range necessarily resulting from the standard deviation found in their respective testing measurements.
[0050] The term "exemplary" means serving as a non-limiting example, instance, or illustration. As utilized herein, the terms "e.g.," and "for example" set off lists of one or more non-limiting aspects, examples, instances, or illustrations.
[0051] As used herein, the term "substantially" refers to the qualitative condition of exhibiting total or near-total extent or degree of a characteristic or property of interest. Biological and chemical phenomena rarely, if ever, go to completion and / or proceed to completeness or achieve or avoid an absolute result. The term "substantially" is therefore used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena. For example, "substantially" may refer to being within at least about 20%, alternatively at least about 10%, alternatively at least about 5% of a characteristic or property of interest.
[0052] The term “subject,” as used herein, refers to any animal. In some instances, the subject is a mammal. In some instances, the term “subject,” as used herein, refers to a human (e.g., a man, a woman, or a child).III. Composite Scoring Methods
[0053] Disclosed herein are computer-implemented methods to assign a Mount Sinai Acute graft-vs-host disease (GVHD) International Consortium (MAGIC) composite score for a subject in need thereof. Also disclosed here in are methods for quantitatively assessing severity of GVHD and monitoring treatment response in a subject in need thereof.
[0054] Current acute graft-vs-host disease (GVHD) risk stratification systems use either biomarkers (e.g., Ann Arbor scores (AA)) or clinical symptoms (e.g., Minnesota risk), but not both. The Minnesota system (standard or high risk) lacks a low risk stratum needed for treatment de-escalation. The disclosed MAGIC study had two goals: to determine whether arisk model based on clinical symptoms alone could define three separate groups and to determine whether inclusion of biomarker risks could further improve the prediction of risk. The inventors analyzed 1863 patients who underwent a first HCT for hematological disorders between 2014 - 2021 who received systemic treatment for acute GVHD. The inventors randomly divided them into training (n=1306) and validation (n=557) cohorts so that there were sufficient numbers of patients with less common clinical presentations (e.g., isolated stage 2 or stage 3 GI GVHD) in the training cohort. The inventors divided patients in the training cohort into groups based on clinical similarity and non-relapse mortality (NRM) and then used a classification and regression tree (CART) model to create three Manhattan risk groups that were independently confirmed by unsupervised K-means clustering. The inventors again applied a CART analysis to the training cohort to several combinations of Manhattan risk and AA scores to create a new MAGIC risk model of three strata that was also confirmed by K- means clustering.
[0055] The symptom-based Manhattan model, a clinical tool used to assess the risk of GVHD after allogeneic hematopoietic cell transplantation (HCT) contains three risk groups that accurately predicts treatment-response and NRM. The Manhattan model identifies three risk groups: low, standard, and high.
[0056] The MAGIC model integrates both clinical and biomarker parameters and is superior to the Manhattan model in predicting overall response rates to treatment at day 28 (Day 28 ORR) and NRM. These new systems offer more accurate risk assessments of patients with acute GVHD that may help to guide therapy.
[0057] High initial doses of corticosteroids and gradual tapers lasting for months have been the recommended treatment for GVHD for decades. Recent advances in GVHD prophylaxis, however, have reduced the overall incidence of severe GVHD and mild-moderate symptoms are now the dominant clinical phenotype. In this study the observed NRM for Minnesota standard risk patients (-11%) was half that of previous publications, reflecting a trend to less NRM from GVHD in these patients. The inventors first validated a Manhattan risk system using only clinical organ severity that identified significant numbers of patients with mild GVHD in a low risk stratum encompassing about 40% of patients. The size of this low risk stratum significantly increased to more than 60% of patients with the incorporation of biomarker values. The incidence of 6-month NRM for these patients (~6%) is almost half that of Minnesota standard risk (-11%) and represents a clinically important difference in outcomes. Patients with a MAGIC composite score of 1 thus possess a very low risk of NRM,which may serve to guide individual treatment strategies that minimize steroid exposure (NCT05090384).
[0058] Physicians often consider both clinical symptoms and laboratory findings in determining the treatment of individual patients. The incorporation of biomarkers and clinical symptoms in the MAGIC composite scores by creating a third risk group leverages the prognostic accuracy of the MAGIC serum biomarkers and resolves the dilemma clinicians face when the severity of clinical and laboratory parameters do not align. The risk of NRM changed for two notable subsets of patients when AA scores were integrated with Manhattan risk groups. First, nearly one-quarter of all patients who were classified as Manhattan intermediate risk but who had the lowest biomarker risk (AA1) experienced very low 6-month NRM of 8% and were therefore classified as the MAGIC composite score 1. Second, a small group (<5% of all patients) with Manhattan low risk and the highest biomarker risk (AA3) experienced 6- month NRM of 28% were therefore classified as the MAGIC composite score 2. The high risk of NRM for this small group is important to consider in treatment decisions.
[0059] Increasing numbers of patients are currently receiving PTCy-based GVHD prophylaxis in HLA matched donor HCT as well as HLA mismatched and haploidentical HCT. The Manhattan risk model using clinical symptoms alone did not distinguish between low and intermediate risk in such patients, but MAGIC composite scores did successfully stratify these patients into three groups for risk of NRM, however, further demonstrating the additive value of biomarker scores to clinical phenotypes.
[0060] When biomarker values are not readily available, the Manhattan risk model offers advantages compared to the Minnesota risk model. First, given the superior survival and response rate of patients with Manhattan low risk GVHD, they may be considered for clinical trials designed to minimize immunosuppressive treatment. Second, it may be desirable to exclude patients with Manhattan low risk who have excellent treatment responses to standard treatment along with low NRM from trials investigational treatments intended to improve response rates. Third, the inclusion of all patients with liver GVHD in the high-risk group regardless of other organ involvement resolves an anomaly of the Minnesota risk system that categorized some patients with both skin and liver GVHD as standard risk instead of high risk and these analyses were consistent, homogeneous, treatments were employed.
[0061] A new Manhattan risk system based on clinical symptoms alone at the initiation of systemic treatment, and new MAGIC composite scores that include biomarkers, are moreaccurate than current risk classification systems. The verification by a second statistical approach of both models lends confidence to the accuracy of their categorizations. These improved models offer the potential to better identify patients who might derive benefit from personalized primary treatment strategies for patients with both low risk and high-risk disease.
[0062] In some embodiments, a computing system is used to implements aspects of the present disclosure. In some aspects, the computing system includes at least one processing device, such as a central processing unit (CPU). The computing system may also include a system memory, and a system bus that couples various system components including the system memory to the processing device. The system bus is one of any number of types of bus structures including, but not limited to, a memory bus, or memory controller; a peripheral bus; and a local bus using any of a variety of bus architectures.
[0063] In some aspects, the system memory includes read only memory and random access memory. A basic input / output system containing the basic routines that act to transfer information within computing system, such as during start up, is typically stored in the read only memory.
[0064] In some aspects, the computing system also includes a secondary storage device for storing digital data. The secondary storage device is connected to the system bus by a secondary storage interface. The secondary storage devices and their associated computer readable media provide nonvolatile storage of computer readable instructions (including application programs and program modules), data structures, and other data for the computing system.
[0065] A number of program modules can be stored in secondary storage device or memory, including an operating system, one or more application programs, other program modules, and program data. In some embodiments, computing system includes input devices to enable a user to provide inputs to the computing system. The input devices are often connected to the processing device through an input / output interface that is coupled to the system bus. These input devices can be connected by any number of input / output interfaces, such as a parallel port, serial port, game port, or a universal serial bus. Wireless communication between input devices and interface is possible in some possible embodiments.
[0066] The computing system typically includes at least some form of computer- readable media. Computer readable media includes any available media that can be accessed by the computing system. By way of example, computer-readable media include computer readable storage media and computer readable communication media.
[0067] Computer readable storage media includes volatile and nonvolatile, removable and non-removable media implemented in any device configured to store information such as computer readable instructions, data structures, program modules or other data. Computer readable storage media includes, but is not limited to, random access memory, read only memory, electrically erasable programmable read only memory, flash memory or other memory technology, compact disc read only memory, digital versatile disks or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing system.
[0068] Computer readable communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term “modulated data signal” refers to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, computer readable communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency, infrared, and other wireless media. Combinations of any of the above are also included within the scope of computer readable media.
[0069] One embodiment of the disclosure includes a computer-implemented method to assign a Mount Sinai Acute graft-vs-host disease (GVHD) International Consortium (MAGIC) composite score for a subject in need thereof, comprising receiving, by a computing system, a first data set from the subject comprising clinical symptoms; receiving, by a computing system, a second data set from the subject comprising serum biomarker data; processing, by the computing system, the first data set using a trained machine learning model configured to generate a Manhattan risk model classification; processing, by the computing system, the second data set to generate a serum biomarker score; and assigning, by the computing system, a MAGIC composite score for the subject based on the Manhattan risk model classification and the serum biomarker score.
[0070] One embodiment of the disclosure includes a computing system for assigning a MAGIC composite score for a subject in need thereof comprising: one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the computing system to: i) receive a first data set from the subject comprising clinical symptoms and nonrelapse mortality data; ii) receive a second data set from the subject comprising serumbiomarker data; iii) apply a trained machine learning model to the first data set, wherein the model is trained on clinical symptoms and non-relapse mortality data to generate a Manhattan risk model classification for the subject; iv) generate a serum biomarker score for the second data set; and v) assign a MAGIC composite score for the subject based on the Manhattan risk model classification and the serum biomarker score.
[0071] In an aspect, the clinical symptoms may include acute GVHD symptoms or chronic GVHD symptoms. Acute GVHD symptoms include, but are not limited to, abdominal pain or cramps, nausea, vomiting, diarrheajaundice or other liver problems, skin rash, itching, redness on areas of the skin, and increased risk for infections. Chronic GVHD symptoms include, but are not limited to, dry eyes, burning sensation, or vision changes, dry mouth, white patches inside the mouth, and sensitivity to spicy foods, fatigue, muscle weakness, and chronic pain, joint pain or stiffness, skin rash with raised, discolored areas, as well as skin tightening or thickening, shortness of breath due to lung damage, vaginal dryness, weight loss, reduced bile flow from the liver, brittle hair and premature graying, damage to sweat glands, cytopenia, and pericarditis.
[0072] In an aspect the subject may have Grade I acute GVHD. This is typically characterized by maculopapular rash covering less than about 25% of body surface area (BSA), 2-3 mg / dL bilirubin levels, a stool output per day of 500-999 mL (Adult) or 10-19.9 mL / kg (child), and persistent nausea, vomiting, or anorexia, with a positive upper GI biopsy.
[0073] In an aspect the subject may have Grade II acute GVHD. This is typically characterized by maculopapular rash covering between about 25% to about 50% BSA, 3.1-6 mg / dL bilirubin levels, and a stool output per day of 1000-1500 mL (Adult) or 20-30 mL / kg (child).
[0074] In an aspect the subject may have Grade III acute GVHD. This is typically characterized by maculopapular rash covering greater than about 50% BSA, 6.1-15 mg / dL bilirubin levels, and a stool output per day of greater than about 1500 mL (Adult) or greater than about 30 mL / kg (child).
[0075] In an aspect the subject may have Grade IV acute GVHD. This is typically characterized by generalized erythroderma plus bullous formation and desquamation covering greater than about 5% BSA, greater than 15 mg / dL bilirubin levels, and severe abdominal pain with or without ileus, or grossly bloody stool (regardless of stool volume).
[0076] In an aspect, the serum biomarker data includes suppression of tumorigenicity 2 (ST2) biomarker data. In an aspect, the serum biomarker data includes regenerating islet-derivedprotein 3-a (REG3a) biomarker data. In an aspect, the serum biomarker data includes both ST2 and REG3a biomarker data. Biomarker data includes, but is not limited to, serum concentration levels.
[0077] In an aspect, the wherein the serum biomarker score is generated by applying a ST2 and REG3a algorithm to the second data set to generate a MAGIC algorithm probability (MAP) value. MAP was calculated as a single value between 0.001 and 0.999 according to the formula: log[— log(l - MAP)] = -11.263 + 1.844(logioST2) + 0.577(logioREG3a). If the MAP value is less than about 0.141, the biomarker score is 1 (or AA1). If the MAP is between about 0.141 and about 0.290, the biomarker score is 2 (or AA2). If the MAP is greater than about 0.290, the biomarker score is 3 (or AA3).
[0078] In an aspect, the first data set is processed using a trained machine learning model configured to generate a Manhattan risk model classification. In some embodiments, trained machine learning model includes a neural network or ensemble classifier. As disclosed in the examples, trained machine learning model is trained on a training data set comprising clinical symptoms and non-relapse mortality data, wherein the training data set is evaluated using a classification and regression tree (CART) algorithm to generate a Manhattan risk model classification.
[0079] The inventors hypothesized that the inclusion of serum biomarkers at the time of treatment onset would further improve the performance of the Manhattan clinical risk model. Serum samples at treatment onset were available in 80% (1050 / 1306) of the training and 78% (432 / 557) of the validation cohorts (Table 1).
[0080] Table 1. Patient characteristicsTraining Validation n=1306 n=557 P valuesMedian age at HCT, year [range] 56 [0, 79] 54 [0, 79] 0.206Recipient Age, category <18 146 (11.2) 83 (14.9) 0.07718-54 474 (36.3) 199 (35.7)>55 686 (52.5) 275 (49.4)Sex mismatch Female to Male 219 (16.8) 103 (18.6) 0.349Other 1087 (83.2) 452 (81.4)Primary disease Acute leukemia 677 (51.8) 299 (53.7) 0.349MDS / MPN 341 (26.1) 147 (26.4)Malignant lymphoma 117 (9.0) 36 (6.5)Other 171 (13.1) 75 (13.5)Disease risk Standard 1059 (81.1) 430 (77.2) 0.058High 247 (18.9) 127 (22.8)Donor type HLA matched related 267 (20.4) 111 (19.9) 0.058HLA matched unrelated 714 (54.7) 277 (49.7)HLA mismatched related 8 (0.6) 4 (0.7)HLA mismatched122 (9.3) 65 (11.7) unrelatedHaploidentical 148 (11.3) 65 (11.7)Umbilical cord blood 47 (3.6) 35 (6.3)GVHD prophylaxis CNI and MTX based 691 (52.9) 292 (52.4) 0.744CNI and MMF based 305 (23.4) 143 (25.7)PTCy 219 (16.8) 82 (14.7)Ex- vivo T-cell depletion 38 (2.9)Other 53 (4.1) 23 (4.1)HCT-CI 0-2 884 (67.7) 372 (66.8) 0.706>3 422 (32.3) 185 (33.2)In-vivo T-cell depletion No 809 (61.9) 341 (61.2) 0.795Yes 497 (38.1) 216 (38.8)Donor source Bone marrow 252 (19.3) 117 (21.0) 0.20Peripheral blood 1007 (77.1) 405 (72.7)Umbilical cord blood 47 (3.6) 35 (6.3)Conditioning MAC (TBI <8Gy) 532 (40.7) 226 (40.6) 0.923MAC (TBI >=8Gy) 204 (15.6) 91 (16.3)RIC 570 (43.6) 240 (43.1)Sample available at Tx No 256 (19.6) 125 (22.4) 0.168Yes 1050 (80.4) 432 (77.6)Median year of HCT 2018 [2014, 2017 [2014,0.117 [range] 2021] 2021]MDS / MPN, myelodysplastic syndromes / myeloproliferative neoplasms; CNI, calcineurin inhibitor; MTX, methotrexate; MMF, mycophenolate mofetil; PTCy, post-transplant cyclophosphamide; HCT-CI, hematopoietic cell transplantation- specific comorbidity index; MAC, myeloablative conditioning; TBI, total body irradiation; RIC, reduced intensity conditioning.
[0081] The 6-month NRM did not differ between patients with and without samples either cohort (16% vs. 13%, P = 0.296 and 16% vs. 12%, P = 0.329, respectively). As expected, the inventors found that AA scores independently stratified the risk of NRM in each risk group of the Manhattan risk model. The risk of NRM for each AA score increased with escalating Manhattan risk, further demonstrating improved prediction of outcome by combining clinical and biomarker assessments (FIGs. 1A - 1C). The inventors again applied a CART analysis to the nine combinations of the Manhattan risk and AA scores in the training cohort that created a new composite scoring system of three strata, which the inventors called the MAGIC composite scores (Table 2). The inventors confirmed the accuracy performance of the new model using an unsupervised K-means clustering algorithm.Table 2. Nine categories determined by the Manhattan risk and AA scores in the training cohortCART K-meansManhattan risk AA scores n (%) 6m NRM (%) MAGIC composite MAGIC compositeLow AA1 296(28.2) 3.1 1 1Low AA2 99(9.4) 12.1 1 1Low AA3 36(3.4) 27.8 2 2Intermediate AA1 247(23.5) 8.7 1 1Intermediate AA2 125(11.9) 18.4 2 2Intermediate AA3 74(7.0) 29.7 2 2High AA1 50(4.8) 20.4 2 2High AA2 52(5.0) 29.1 2 2High AA3 71(6.8) 56.3 3 3
[0082] In an embodiment, if the subject receives a Manhattan low risk or intermediate risk classification and a biomarker score of 1, the subject is assigned a MAGIC composite score of 1.
[0083] In an embodiment, if the subject receives a Manhattan low risk classification and a biomarker score of 2, the subject is assigned a MAGIC composite score of 1.
[0084] In an embodiment, if the subject receives a Manhattan low risk classification and a biomarker score of 3, the subject is assigned a MAGIC composite score of 2.
[0085] In an embodiment, if the subject receives a Manhattan intermediate risk and a biomarker score of 2 or 3, the subject is assigned a MAGIC composite score of 2.
[0086] In an embodiment, if the subject receives a Manhattan high risk classification and a biomarker score of 1 or 2, the subject is assigned a MAGIC composite score of 2.
[0087] In an embodiment, if the subject receives a Manhattan high risk classification and a biomarker score of 3, the subject is assigned a MAGIC composite score of 3.
[0088] The incidence of NRM within 6 months increased with each increase in MAGIC composite score but the incidence of relapse did not change, resulting in large differences in OS between each group in both the training and validation cohorts (FIGs. 2A-2C, FIGs. 3A- 3C). In the total population 356 / 1482 (24%) intermediate risk patients in the Manhattan modelbecame MAGIC composite score 1, with a 6-month NRM rate of only 8%. Furthermore, 46 / 1482 (3%) low Manhattan risk patients increased one risk stratum to MAGIC composite score 2, with a 6-month NRM of 28%, and 147 / 1428 (12%) of Manhattan high risk patients decreased one risk stratum to MAGIC composite score 2 with a 6-month NRM of 26%.
[0089] Using 6-month NRM as the outcome, the AUC of the MAGIC composite score model was significantly higher than that of Manhattan model in both the training (0.73 vs. 0.69, P = 0.019) and in the validation cohorts (0.76 vs. 0.70, P = 0.010) (FIGs. 4A and 4B). The inventors next assessed the prognostic efficacy of each model at several time points during the first year from GVHD treatment. The MAGIC composite scores were consistently superior to both Manhattan and Minnesota risk models (FIG. 2B). The Akaike information criterion (AIC) for predicting 6-month NRM based on MAGIC composite scores was also substantially lower (763.9) than the Manhattan (1006.9) or Minnesota risk model (1046.7).
[0090] The inventors evaluated the robustness of the MAGIC composite score model in two key subsets: patients with Glucksberg Grades II to IV acute GVHD and patients treated with >0.5 mg methylprednisolone per kg using the whole cohort. The MAGIC composite score model produced significantly higher AUCs in both subsets compared to the Manhattan risk model (0.72 vs. 0.66, P < 0.001, 0.74 vs. 0.68, P = 0.009).
[0091] The inventors also evaluated the model in a third key subset of patients, those developing acute GVHD after receiving prophylaxis that contained post-transplantation cyclophosphamide (PTCy) (n = 301). In the Manhattan risk model, there was no significant difference between groups for 6-month NRM in these patients (11% vs. 11% vs. 19%, P = 0.336) but incorporation of biomarkers into the MAGIC composite score model effectively stratified the risk of NRM in these patients (8% vs. 16% vs. 27%, P = 0.026, FIGs. 5A and 5B).
[0092] Finally, the inventors assessed these models for prediction of day 28 ORR, the standard endpoint for treatment response in clinical trials. The inventors did not observe significant differences in ORR between the low and intermediate Manhattan risk groups in the validation cohort (77% vs. 70%, P = 0. 182), but the inventors did observe significant differences between each MAGIC composite score (80% vs. 63% vs. 30%, FIGs. 6A-6C).
[0093] The validation cohort showed that MAGIC composite score model significantly improved prediction of NRM compared to Manhattan risk model as measured by the area under the receiver operating characteristic curve (AUC) (0.76 vs. 0.70, =0.010), which in turn wasbetter than Minnesota risk (0.69 vs. 0.64, =0.009). Each increase in MAGIC composite score correlated with an increase in 6 month NRM (6% vs. 29% vs. 52%, P < 0.001), and with a significant decrease in day 28 treatment response (80% vs. 63% vs. 30%, P<0.00l ).
[0094] The inventors concluded that the new MAGIC composite scores more accurately predict response to therapy and long term outcomes than systems based on clinical symptoms alone and may help guide clinical decisions and trial design.
[0095] In some embodiments, the methods may be used to identify subjects in need for medical care and to guide precision medical care for an individual, including but not limited to diagnostic assessment and therapy in asymptomatic, otherwise healthy-appearing persons with little or no apparent symptoms or evaluation of symptomatic patients with pre-existing mild, moderate or advanced symptoms.
[0096] In some embodiments, the methods may be used to quantitatively assess the severity of GVHD in a subject. In some embodiments, the methods may be used to determine a subject’s prognosis of non-relapse mortality. In some embodiments, a MAGIC composite score of 1 correlates to a prognosis of low risk of non-relapse mortality. In some embodiments, a MAGIC composite score of 2 correlates to a prognosis of intermediate risk of non-relapse mortality. In some embodiments, a MAGIC composite score of 3 correlates to a prognosis of high risk of non-relapse mortality.
[0097] In some aspects, a subject’s MAGIC composite score may be monitored to assess the effectiveness of a therapeutic intervention. For example, one embodiment includes monitoring treatment response by recalculating the MAGIC composite score after administration of a therapeutic intervention. In an embodiment, therapeutic intervention comprises steroids, extracorporeal photopheresis (ECP), or biological therapies. In an embodiment, the calculated MAGIC composite score is used to automatically adjust a treatment regime. In certain aspects, the adjustment may be done through integration with an electronic health record system.
[0098] The inventors identified 1135 patients in the MAGIC database who received systemic therapy for GVHD, who had serum samples available at the onset of therapy and on Day 28, the current gold standard for evaluating response to treatment. The inventors then divided patients into a training cohort (n=826) that underwent transplant from 2014-2020, and a validation cohort (n= 309) that underwent transplant from 2021-2023. The inventors chose a later validation set in order to reflect recent changes in practices of GVHD prophylaxis, particularly the increased use of post transplant cyclophosphamide. The median initial dose ofprednisolone was approximately 1.2 mg / kg / day, and approximately 20% of patients received second-line treatment before D28 of treatment.
[0099] The inventors categorized patients by both clinical severity and biomarker scores according to previous MCS criteria and then categorized by the same criteria at day 28 of treatment. Patients that had a complete resolution of symptoms were categorized as grade 0 (n=525, 63.6%) and divided into subgroups by biomarker Ann Arbor (AA) scores. Patients who were MCS1 at onset and were grade 0 AA1 or grade 0 AA2 on day 28 exhibited similar, low 6-month NRM (2.3% and 14% respectively) and were therefore grouped together. Likewise, patients who were MCS 2 / 3 at onset exhibited similar low 6-month NRM (7.4% and 14%) and were grouped together. Patients classified as MCS2 or MCS3 both at onset or at day 28 were combined in a single group because the number of MCS3 patients was small at both time points.
[0100] The inventors then applied the classification and regression tree algorithm to these eight groups to generate the MAGIC Composite Response (MCR) and categorized them as responders or nonresponders according to 6-month NRM. This process resulted in two groups with either low NRM (5-10%) or high NRM (23-50%) (Table 3).Table 3. Development of MAGIC Composite Response (MCR) in the training set (n=826)At treat .ment . A .t. Dm2o8 6-m . onth m NRM MCR* n ( zn% / \)MCS1 GradeO + AA1 / 2 4.6% (2.7-7.3) Response 324 (39.2)MCS1 GradeO + AA3 23.1% (5.1-48.6) Non-response 13 (1.6)MCS2 / 3 GradeO + AA1 / 2 9.4% (5.4-14.8) Response 149 (18.0)MCS2 / 3 GradeO + AA3 30.8% (17.0-45.6) Non-response 39 (4.7)MCS1 MCS1 3.0% (0.8-7.9) Response 99 (12.0)MCS2 / 3 MCS1 9.7% (2.4-23.2) Response 31 (3.8)MCS1 MCS2 / 3 27.5% (16.0-40.2) Non-response 51 (6.2)MCS2 / 3 MCS2 / 3 49.2% (39.9-57.8) Non-response 120 (14.5)MCS, MAGIC Composite Score; AA, Ann Arbor score; MCR, MAGIC Composite Response; MAP, MAGIC algorithm probability.*Categorized by Classification and regression tree (CART) analyses based on 6-month NRM.
[0101] The algorithm identified two subgroups whose response differed from that defined by clinical response only (CRO). First, patients with complete clinical resolution of symptoms (grade 0) and high biomarkers (AA3) at day 28 of treatment experienced a 6-month NRM of 23.1% and were classified as non responders despite their complete resolution of symptoms.Second, patients who were classified as MCS1 at both onset and at day 28 of treatment experienced a 6-month NRM of only 3.0% and were therefore classified as responders despite the persistence of mild symptoms that were primarily skin rashes (Table 3).
[0102] The inventors then compared the 6-month NRM of responders and nonresponders according to these two response criteria systems. As expected, both the MCR and CRO separated responders and nonresponders in both the training and validation cohorts (FIGs. 7A- 7F, 8A-8F). MCR produced a greater difference in 6 month mortality compared to CRO in the training cohort (33.7% vs 26.7%) and the validation cohort (38.9% vs 27.4%) (FIGs. 7A-7F, 8A-8F). This improved separation resulted in significantly improved area under the receiver operating curves (AUROC) in the training cohort (0.76 vs 0.67, <0.00l ) and the validation cohort (0.77 vs 0.69 =0.014) (FIGs. 7C, 8C). The proportion of patients categorized as responders was similar in both systems (68.9% v 69.3%) (FIGs. 7A-7F). There was no difference in relapse between response groups according to either system. As a result, the MCR produced a greater difference in overall survival for responders than CRO (42.2% vs 33.7%) (FIGs. 7D, 7E).
[0103] To better understand why MCR criteria were superior to CRO criteria in predicting long term mortality and survival, the inventors evaluated the ability of MCR to reclassify both clinical responders and clinical nonresponders. In the validation set, the application of MCR criteria to the CRO responder population identified a small group with a five fold greater incidence ofNRM (34.3% vs 6.8%, P<0.001 ) (FIG. 9A). Over 90% of these patients exhibited high biomarker scores at day 28 of treatment, and the majority (53%) died from uncontrolled acute GVHD. In CRO nonresponders, the MCR identified nearly one third of patients with a seven fold reduction in NRM (7.6% vs 50.7%, <0.001 ) (FIG. 9B) and only one patient died of uncontrolled acute GVHD. The MCR system is therefore both more sensitive and specific than the CRO in predicting NRM, highlighting the utility of biomarkers to improve prediction of long-term GVHD control and NRM.
[0104] The inventors next conducted analyses of key patient subsets using the entire cohort. The cohort included low-risk populations, such as patients with grade I acute GVHD for whom systemic treatment is often not used. The inventors therefore assessed the MCR criteria in patients with grade II-IV acute GVHD, which clearly stratified patients into two groups with significantly different 6-month NRM (7.3% vs 43.8%, P <0.001) and was again superior to CRO criteria with a greater AUROC (0.76 vs 0.69, <0.001 ).
[0105] The presently described technology and its advantages will be better understood by reference to the following examples. These examples are provided to describe specific implementations of the present technology. By providing these specific examples, it is not intended limit the scope and spirit of the present technology. It will be understood by those skilled in the art that the full scope of the presently described technology encompasses the subject matter defined by the claims appending this specification, and any alterations, modifications, or equivalents of those claims.
[0106] EXAMPLES
[0107] Patient selection
[0108] The inventors obtained clinical data and serum samples from the MAGIC database and biorepository that encompasses 23 HCT centers in North America, Europe, and Asia. Participating centers collected clinical information that focused on acute GVHD using a prospective-specimen-collection, retrospective -blinded-evaluation (PRoBE) study design and provided longitudinal serum samples. Patients were prospectively monitored weekly for acute GVHD symptoms according to institutional frequency. Informed consent from an institutional review board-approved protocol was obtained from all participants in accordance with the Declaration of Helsinki.
[0109] The inventors included both pediatric and adult patients who received a first HCT between 2014 and 2021 and who received systemic treatment for acute GVHD of at least 0.1 mg methylprednisolone per kilogram (kg) or equivalent steroid dose. The inventors excluded patients who developed primary relapse of malignancy or who received donor lymphocyte infusion or second HCT before systemic GVHD treatment. Acute GVHD was diagnosed and staged according to the published criteria. Minnesota risk, HCT-specific comorbidity index (HCT-CI) scores, and intensity of conditioning regimens were classified as previously reported. A complete response (CR) was defined as complete resolution of acute GVHD manifestations without secondary treatment, and a partial response (PR) was defined as improvement of less than CR but with a decrease in at least one organ stage without worsening of other organs and without secondary treatment. Overall response rate (ORR) was defined by CR or PR at day 28 after systemic treatment.
[0110] Serum samples
[0111] Serial serum samples were collected prospectively, shipped to a central laboratory, and cryopreserved. Serum concentrations of suppressor of tumorigeni city-2 (ST2) and regeneratingislet-derived protein 3-a (REG3a) (Zhao et al., 2018) were analyzed by enzyme-linked immunosorbent assays, as previously reported. The MAGIC algorithm probability (MAP) was calculated as a single value between 0.001 and 0.999 according to the formula: log[-log(l - MAP)] = -11.263 + 1.844(logioST2) + 0.577(logioREG3a). The inventors calculated Ann Arbor (AA) scores using previously validated thresholds (AA1 < 0.141; 0.141 <AA2 < 0.291; AA3 > 0.291).
[0112] Statistical analysis
[0113] The beginning of systemic treatment served as the starting point in all analyses. The primary endpoint was 6-month NRM and outcomes were censored at 6 months. The inventors estimated and plotted he cumulative incidence of NRM according to Gray’s method, and the inventors considered relapse and second allogeneic HCT as competing risks. The inventors used the Kaplan-Meier method and the log-rank test to estimate and compare overall survival (OS) probabilities. The inventors compared categorical variables using the Fisher's exact test, and continuous variables using the Mann- Whitney U test. The inventors used the area under curve (AUC) of receiver operating characteristic (ROC) analysis to compare the prognostic of the different models.
[0114] The inventors developed algorithms to predict 6-month NRM as follows: first, the inventors randomly divided patients into training and validation cohorts in a 7:3 ratio to provide sufficient numbers of patients with uncommon clinical presentations in the training cohort. Second, the inventors created groups with a minimum of 20 patients per group according to clinical similarities at the time of treatment and in 6-month NRM. Third, the inventors used a classification and regression tree (CART) algorithm to create three groups according to the risk of 6-month NRM after treatment onset. The criteria to separate groups included a maximum depth of 2 levels with a complexity parameter of 0.2 and at least 30 observations in each terminal node. The inventors also applied a K-means approach with Lloyd's algorithm as a sensitivity analysis for the accuracy of aggregation. The performance of each model was then evaluated in the validation cohort.
[0115] All statistical tests were 2-sided and a P- value < 0.05 was considered statistically significant. Statistical analyses were performed with R (The R Foundation for Statistical Computing, version 4.2.2, Vienna, Austria) or EZR version 1.61 (Jichi Medical University Saitama Medical Center, Saitama, Japan).
[0116] Patient characteristics
[0117] The inventors randomly divided 1863 patients who fulfilled all the inclusion criteria into a training (n=1306) and validation cohort (n=557) (FIG. 10) There were no significant differences in baseline characteristics between cohorts except for donor source (Table 1).
[0118] Severity of GVHD and organ involvement at the time of treatment were also similar between the training and validation cohorts (Table 4). The median follow-up of survivors after treatment initiation was 22 months (range, 1 to 58) and 23 months (range, 1 to 37) in the training and validation cohorts, respectively.
[0119] Table 4. GVHD characteristics at treatment
[0120]
[0121] Manhattan risk system
[0122] The inventors first categorized 76 combinations of all GVHD phenotypes that possessed at least one case in the training cohort (Table 5) into 24 groups based on individual organ severity at the time of treatment (Table 6).
[0123] Table 5. GVHD phenotype in the training cohortAll other combinations had zero patients.
[0124] Table 6. GVHD organ involvement categories in the training cohort
[0125] The inventors then combined groups with similar clinical characteristics and 6-month NRM to create 14 categories with at least 20 patients in each category (Table 6). Using the CART algorithm, the inventors further reduced the number of categories to three (low, intermediate, and high risk) which the inventors termed the Manhattan risk model. A sensitivity analysis using an unsupervised K-means clustering algorithm confirmed the accuracy of aggregation (Table 6).
[0126] Manhattan risk differed from Minnesota risk in two important subsets. First, approximately half of Minnesota standard risk patients became low risk: clinical symptoms included isolated stage 1 or 2 skin, isolated upper gastrointestinal (UGI), and stage 1 skin plusUGI. Second, patients with any liver involvement became high risk whereas by Minnesota criteria liver GVHD with stage 1 to 3 skin is considered standard risk.
[0127] The Glucksberg classification (Grades I / II vs. III / IV), and a recently proposed principal component-derived grading system possess similar AUCs to Minnesota risk for the prediction of 6-month NRM (FIGs. 11A and 11B). In the validation cohort the AUC of Manhattan model for 6-month NRM was significantly higher than that of Minnesota model (0.69 vs. 0.64, P = 0.009, FIGs. 12A and 12B). The Manhattan risk model did not predict relapse and thus differences in OS between groups were determined by differences in NRM (FIGs. 13A and 13B). The Manhattan model defined 40% of patients as low risk and the three Manhattan strata possessed distinctly different 6-month NRM in both the training and the validation cohorts FIGs. 14A and 14B). Comparison of risk categories by organ involvement are summarized for the two models in Table 6.
[0128] Table 6. Differences between Minnesota and Manhattan risk models in the whole cohort.
[0129] To evaluate the robustness of the Manhattan model, the inventors evaluated subsets limited to Glucksberg Grades II to IV acute GVHD or to treatment with >0.5 mg methylprednisolone per kg in the whole cohort. The AUCs of the Manhattan risk model remained superior to the Minnesota model for both groups (0.69 vs. 0.65, P = 0.028; 0.68 vs. 0.65, P = 0.005, respectively).
[0130] References1. Martin PJ: How I treat steroid-refractory acute graft-versus-host disease. Blood 135:1630-1638, 20202. Akahoshi Y, Spyrou N, Hogan WJ, et al: Incidence, clinical presentation, risk factors, outcomes, and biomarkers in de novo late acute GVHD. Blood Adv 7:4479-4491, 20233. Greinix HT, Eikema DJ, Koster L, et al: Improved outcome of patients with graft-versus-host disease after allogeneic hematopoietic cell transplantation for hematologic malignancies over time: an EBMT mega-file study. Haematologica 107: 1054-1063, 20224. Khoury HJ, Wang T, Hemmer MT, et al: Improved survival after acute graft-versus- host disease diagnosis in the modem era. Haematologica 102:958-966, 20175. Bolanos-Meade J, Hamadani M, Wu J, et al: Post-Transplantation Cyclophosphamide- Based Graft-versus-Host Disease Prophylaxis. N Engl J Med 388:2338-2348, 20236. Watkins B, Qayed M, McCracken C, et al: Phase II Trial of Costimulation Blockade With Abatacept for Prevention of Acute GVHD. J Clin Oncol 39: 1865-1877, 20217. Akahoshi Y, Igarashi A, Fukuda T, et al: Impact of graft-versus-host disease and graft- versus-leukemia effect based on minimal residual disease in Philadelphia chromosome-positive acute lymphoblastic leukemia. Br J Haematol 190:84-92, 20208. Martin PJ, Rizzo JD, Wingard JR, et al: First- and second-line systemic treatment of acute graft-versus-host disease: recommendations of the American Society of Blood and Marrow Transplantation. Biol Blood Marrow Transplant 18: 1150-63, 20129. Malard F, Holler E, Sandmaier BM, et al: Acute graft-versus-host disease. Nat Rev Dis Primers 9:27, 202310. Penack O, Marchetti M, Aljurf M, et al: Prophylaxis and management of graft-versus- host disease after stem-cell transplantation for haematological malignancies: updated consensus recommendations of the European Society for Blood and Marrow Transplantation. Lancet Haematol 11 :el47-el59, 202411. MacMillan ML, DeFor TE, Weisdorf DJ: The best endpoint for acute GVHD treatment trials. Blood 115:5412-7, 201012. Levine JE, Logan B, Wu J, et al: Graft-versus-host disease treatment: predictors of survival. Biol Blood Marrow Transplant 16: 1693-9, 201013. Saliba RM, Couriel DR, Giralt S, et al: Prognostic value of response after upfront therapy for acute GVHD. Bone Marrow Transplant 47: 125-31, 201214. Inamoto Y, Martin PJ, Storer BE, et al: Response endpoints and failure-free survival after initial treatment for acute graft-versus-host disease. Haematologica 99:385-91, 201415. Biavasco F, Ihorst G, Wasch R, et al: Therapy response of glucocorticoid-refractory acute GVHD of the lower intestinal tract. Bone Marrow Transplant 57:1500-1506, 202216. El Jurdi N, Rayes A, MacMillan ML, et al: Steroid-dependent acute GVHD after allogeneic hematopoietic cell transplantation: risk factors and clinical outcomes. Blood Adv 5: 1352-1359, 202117. Etra A, Capellini A, Alousi A, et al: Effective treatment of low-risk acute GVHD with itacitinib monotherapy. Blood 141 :481-489, 202318. Akahoshi Y, Kimura SI, Tada Y, et al: Cytomegalovirus gastroenteritis in patients with acute graft-versus-host disease. Blood Adv 6:574-584, 202219. Akahoshi Y, Spyrou N, Hoepting MDm, et al: Flares of Acute Graft- Versus-Host Disease (GVHD): A Mount Sinai Acute GVHD International Consortium (MAGIC) Analysis. Blood Adv, 202420. Weisdorf DJ, Hurd D, Carter S, et al: Prospective grading of graft-versus-host disease after unrelated donor marrow transplantation: a grading algorithm versus blinded expert panel review. Biol Blood Marrow Transplant 9:512-8, 200321. Przepiorka D, Weisdorf D, Martin P, et al: 1994 Consensus Conference on Acute GVHD Grading. Bone Marrow Transplant 15:825-8, 199522. Cahn JY, Klein JP, Lee SJ, et al: Prospective evaluation of 2 acute graft-versus-host (GVHD) grading systems: a joint Societe Francaise de Greffe de Moelle et Therapie Cellulaire (SFGM-TC), Dana Farber Cancer Institute (DFCI), and International Bone Marrow Transplant Registry (IBMTR) prospective study. Blood 106: 1495-500, 200523. Bayraktar E, Graf T, Ayuk FA, et al: Data-driven grading of acute graft-versus-host disease. Nat Commun 14:7799, 202324. MacMillan ML, Robin M, Harris AC, et al: A refined risk score for acute graft-versus- host disease that predicts response to initial therapy, survival, and transplant-related mortality. Biol Blood Marrow Transplant 21 :761-7, 201525. MacMillan ML, DeFor TE, Holtan SG, et al: Validation of Minnesota acute graft- versus-host disease Risk Score. Haematologica 105:519-524, 202026. Levine JE, Braun TM, Harris AC, et al: A prognostic score for acute graft-versus-host disease based on biomarkers: a multicentre study. Lancet Haematol 2:e21-9, 201527. Hartwell MJ, Ozbek U, Holler E, et al: An early-biomarker algorithm predicts lethal graft-versus-host disease and survival. JCI Insight 3, 201828. Holtan SG, DeFor TE, Panoskaltsis-Mortari A, et al: Amphiregulin modifies the Minnesota Acute Graft-versus-Host Disease Risk Score: results from BMT CTN 0302 / 0802. Blood Adv 2: 1882-1888, 201829. Etra A, Gergoudis S, Morales G, et al: Assessment of systemic and gastrointestinal tissue damage biomarkers for GVHD risk stratification. Blood Adv 6:3707-3715, 202230. Spyrou N, Akahoshi Y, Ayuk F, et al: The utility of biomarkers in acute GVHD prognostication. BloodAdv 7:5152-5155, 202331. Robin M, Porcher R, Michonneau D, et al: Prospective external validation of biomarkers to predict acute graft-versus-host disease severity. BloodAdv 6:4763-4772, 202232. McCurdy SR, Radojcic V, Tsai HL, et al: Signatures of GVHD and relapse after posttransplant cyclophosphamide revealed by immune profiling and machine learning. Blood 139:608-623, 202233. Luft T, Benner A, Jodele S, et al: EASIX in patients with acute graft-versus-host disease: a retrospective cohort analysis. Lancet Haematol 4:e414-e423, 201734. Socie G, Niederwieser D, von Bubnoff N, et al: Prognostic value of blood biomarkers in steroid-refractory or steroid-dependent acute graft-versus-host disease: a REACH2 analysis. Blood 141:2771-2779, 202335. Pepe MS, Feng Z, Janes H, et al: Pivotal evaluation of the accuracy of a biomarker used for classification or prediction: standards for study design. J Natl Cancer Inst 100: 1432-8, 200836. Harris AC, Young R, Devine S, et al: International, Multicenter Standardization of Acute Graft-versus-Host Disease Clinical Data Collection: A Report from the Mount Sinai Acute GVHD International Consortium. Biol Blood Marrow Transplant 22:4-10, 201637. Bacigalupo A, Ballen K, Rizzo D, et al: Defining the intensity of conditioning regimens: working definitions. Biol Blood Marrow Transplant 15: 1628-33, 200938. Sorror ML, Storer B, Storb RF: Validation of the hematopoietic cell transplantationspecific comorbidity index (HCT-CI) in single and multiple institutions: limitations and inferences. Biol Blood Marrow Transplant 15:757-8, 200939. Zhang J, Ramadan AM, Griesenauer B, et al: ST2 blockade reduces sST2-producing T cells while maintaining protective mST2-expressing T cells during graft-versus-host disease. Sci Transl Med 7:308ral60, 201540. Zhao D, Kim YH, Jeong S, et al: Survival signal REG3alpha prevents crypt apoptosis to control acute gastrointestinal graft-versus-host disease. J Clin Invest 128:4970-4979, 201841. Major-Monfried H, Renteria AS, Pawarode A, et al: MAGIC biomarkers predict longterm outcomes for steroid-resistant acute GVHD. Blood 131 :2846-2855, 201842. Aziz MD, Shah J, Kapoor U, et al: Disease risk and GVHD biomarkers can stratify patients for risk of relapse and nonrelapse mortality post hematopoietic cell transplant. Leukemia 34:1898-1906, 202043. Al Malki MM, London K, Baez J, et al: Phase 2 study of natalizumab plus standard corticosteroid treatment for high-risk acute graft-versus-host disease. Blood Adv 7:5189-5198, 202344. BzdokD, Krzywinski M, Altman N: Points of Significance: Machine learning: aprimer. Nat Methods 14: 1119-1120, 201745. Breiman L: Classification and regression trees. [Abingdon], [Routledge] [Abingdon], 201746. Hartigan JA, Wong MA: A K-Means Clustering Algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics) 28:100-108, 197947. Kanda Y: Investigation of the freely available easy-to-use software 'EZR' for medical statistics. Bone Marrow Transplant 48:452-8, 201348. Luznik L, O'Donnell PV, Symons HJ, et al: HLA-haploidentical bone marrow transplantation for hematologic malignancies using nonmyeloablative conditioning and high- dose, posttransplantation cyclophosphamide. Biol Blood Marrow Transplant 14:641-50, 200849. Meybodi MA, Cao W, Luznik L, et al: HLA-haploidentical vs matched-sibling hematopoietic cell transplantation: a systematic review and meta-analysis. Blood Adv 3:2581- 2585, 201950. Bolanos-Meade J, Reshef R, Fraser R, et al: Three prophylaxis regimens (tacrolimus, mycophenolate mofetil, and cyclophosphamide; tacrolimus, methotrexate, and bortezomib; or tacrolimus, methotrexate, and maraviroc) versus tacrolimus and methotrexate for prevention of graft-versus-host disease with haemopoietic cell transplantation with reduced-intensity conditioning: a randomised phase 2 trial with a non-randomised contemporaneous control group (BMT CTN 1203). Lancet Haematol 6:el32-el43, 201951. D'Souza A, Fretham C, Lee SJ, et al: Current Use of and Trends in Hematopoietic Cell Transplantation in the United States. Biol Blood Marrow Transplant 26:el77-el82, 202052. Rimando J, McCurdy SR, Luznik L: How I prevent GVHD in high-risk patients: posttransplant cyclophosphamide and beyond. Blood 141 :49-59, 202353. Mielcarek M, Furlong T, Storer BE, et al: Effectiveness and safety of lower dose prednisone for initial treatment of acute graft-versus-host disease: a randomized controlled trial. Haematologica 100:842-8, 2015
[0131] All features disclosed in the specification, including the claims, abstracts, and drawings, and all the steps in any method or process disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. Each feature disclosed in the specification, including the claims, abstract, and drawings, can be replaced by alternative features serving the same, equivalent, or similar purpose, unless expressly stated otherwise. Thus, unless expressly stated otherwise, each feature disclosed is one example only of a generic series of equivalent or similar features.
[0132] It will be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
Claims
CLAIMS1. A computer-implemented method to assign a Mount Sinai Acute graft-vs-host disease (GVHD) International Consortium (MAGIC) composite score for a subject in need thereof, comprising: i) receiving, by a computing system, a first data set from the subject comprising clinical symptoms; ii) receiving, by a computing system, a second data set from the subject comprising serum biomarker data; iii) processing, by the computing system, the first data set using a trained machine learning model configured to generate a Manhattan risk model classification; iv) processing, by the computing system, the second data set to generate a serum biomarker score; and iv) assigning, by the computing system, a MAGIC composite score for the subject based on the Manhattan risk model classification and the serum biomarker score.
2. The method of claim 1, wherein the trained machine learning model is trained on a training data set comprising clinical symptoms and non-relapse mortality data, wherein the training data set is evaluated using a classification and regression tree (CART) algorithm to generate a Manhattan risk model classification.
3. The method of claim 1 or claim 2, wherein the Manhattan risk model classification is low risk, intermediate risk, or high risk.
4. The method of any one of claims 1 to 3, wherein the serum biomarker score is generated by applying a ST2 and REG3a algorithm to the second data set to generate a MAGIC algorithm probability (MAP) value.
5. The method of claim 4, wherein if the MAP value is less than about 0.141, the biomarker score is 1.
6. The method of claim 4, wherein if the MAP is between about 0.141 and about 0.290, the biomarker score is 2.
7. The method of claim 4, wherein if the MAP is greater than about 0.290, the biomarker score is 3.
8. The method of any one of claim 1 to 7, wherein if the subject receives a Manhattan low risk or intermediate risk classification and a biomarker score of 1, the subject is assigned a MAGIC composite score of 1.
9. The method of any one of claim 1 to 7, wherein if the subject receives a Manhattan low risk classification and a biomarker score of 2, the subject is assigned a MAGIC composite score of 1.
10. The method of any one of claim 1 to 7, wherein if the subject receives a Manhattan low risk classification and a biomarker score of 3, the subject is assigned a MAGIC composite score of 2.
11. The method of any one of claim 1 to 7, wherein if the subject receives a Manhattan intermediate risk and a biomarker score of 2 or 3, the subject is assigned a MAGIC composite score of 2.
12. The method of any one of claim 1 to 7, wherein if the subject receives a Manhattan high risk classification and a biomarker score of 1 or 2, the subject is assigned a MAGIC composite score of 2.
13. The method of any one of claim 1 to 7, wherein if the subject receives a Manhattan high risk classification and a biomarker score of 3, the subject is assigned a MAGIC composite score of 3.
14. The method of any one of claims 1 to 13, wherein the method further comprises determining a prognosis of non-relapse mortality for the subject.
15. The method of claim 14, wherein a MAGIC composite score of 1 correlates to a prognosis of low risk of non-relapse mortality.
16. The method of claim 14, wherein a MAGIC composite score of 2 correlates to a prognosis of intermediate risk of non-relapse mortality.
17. The method of claim 14, wherein a MAGIC composite score of 3 correlates to a prognosis of high risk of non-relapse mortality.
18. The method of any one of claims 1 to 17, wherein the method further comprises quantitatively assessing severity of GVHD in the subject.
19. The method of any one of claims 1 to 18, wherein the method further comprises monitoring treatment response by recalculating the MAGIC composite score after administration of a therapeutic intervention.
20. The method of claim 19, wherein the therapeutic intervention comprises steroids, extracorporeal photopheresis (ECP), or biological therapies.
21. The method of claim 19 or claim 20, wherein the subject is classified as a responder or a non-responder.
22. The method of any one of claims 1 to 18, wherein the MAGIC composite score is used to automatically adjust a treatment regime.
23. The method of claim 22, wherein the adjustment is done via integration with an electronic health record system.
24. A computing system for assigning a MAGIC composite score for a subject in need thereof comprising: one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the computing system to: i) receive a first data set from the subject comprising clinical symptoms and nonrelapse mortality data; ii) receive a second data set from the subject comprising serum biomarker data; iii) apply a trained machine learning model to the first data set, wherein the model is trained on clinical symptoms and non-relapse mortality data to generate a Manhattan risk model classification for the subject; iv) generate a serum biomarker score for the second data set; and v) assign a MAGIC composite score for the subject based on the Manhattan risk model classification and the serum biomarker score.