Predicting Tolerability in Aggressive Non-Hodgkin's Lymphoma

The TRAIL model predicts DLBCL patients' tolerability to R-CHOP using clinical variables, enhancing treatment tolerance prediction and reducing adverse events, especially in elderly patients with comorbidities.

JP7823020B2Active Publication Date: 2026-03-03GENENTECH INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-26
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing models fail to predict whether a subject with diffuse large B-cell lymphoma (DLBCL) will tolerate treatments like R-CHOP, leading to adverse events such as death or hospitalization, especially in elderly patients with comorbidities, thus affecting treatment completion and clinical outcomes.

Method used

A machine learning model, TRAIL, uses a set of easily obtainable clinical variables to predict a subject's tolerability to treatments like R-CHOP, identifying those at high risk for intolerance, allowing for alternative therapies.

Benefits of technology

The model improves clinical outcomes by accurately predicting treatment tolerability, enabling personalized treatment plans and reducing adverse events in patients with DLBCL.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007823020000010
    Figure 0007823020000010
  • Figure 0007823020000011
    Figure 0007823020000011
  • Figure 0007823020000012
    Figure 0007823020000012
Patent Text Reader

Abstract

The systems and methods described herein improve outcomes for subjects with lymphoma. Subjects with lymphoma may be able to avoid traditional treatments that have a high probability of causing adverse events in the subject. The systems and methods allow for more accurate prediction of subjects who will not tolerate a particular treatment. The method may include accessing an input dataset including a plurality of input data values ​​associated with a particular subject with lymphoma. The method may further include inputting the input dataset into a machine learning model to generate a score corresponding to the degree to which the particular subject will tolerate the particular treatment. The method may include using the generated score to output a prediction of the particular subject's tolerance of the particular treatment.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 060,371, entitled "PREDICTING TOLERABILITY OF R-CHOP IN AGGRESSIVE NON-HODGKIN LYMPHOMA," filed August 3, 2020, and U.S. Provisional Patent Application No. 63 / 111,777, entitled "PREDICTING TOLERABILITY OF R-CHOP IN AGGRESSIVE NON-HODGKIN LYMPHOMA," filed November 10, 2020, the entire contents of which are incorporated herein by reference for all purposes.

[0002] The methods and systems disclosed herein generally relate to determining the likelihood that a particular subject will tolerate a particular treatment for diffuse large B-cell lymphoma, instead of determining the likelihood that the treatment will be effective.For example, machine learning models can be used to calculate outcome scores and predict the clinical outcome of a particular subject's treatment with R-CHOP (rituximab, cyclophosphamide, doxorubicin hydrochloride, vincristine, prednisolone). [Background technology]

[0003] R-CHOP (rituximab, cyclophosphamide, doxorubicin hydrochloride, vincristine, and prednisolone) remains the standard of care for initial (1L) treatment of diffuse large B-cell lymphoma (DLBCL), with long-term event-free survival rates of 60–70% for patients who complete the recommended 6–8 cycle course (Coiffier et al., Blood 2010). However, most patients are unable to tolerate R-CHOP due to comorbidities or vulnerabilities that confer susceptibility to treatment-related toxicities. These patients are less likely to achieve complete remission and are at higher risk of relapse due to treatment discontinuation, interruption, or reduction, resulting in reduced chemotherapy dose intensity (Dlugosz-Danecka et al., Cancer Med 2019). Proactive identification of these patients could facilitate treatment and thereby avoid the unfavorable risk / benefit profile of R-CHOP treatment. Identifying patients' tolerability of treatments other than R-CHOP is also desirable. The present disclosure addresses these and other needs. Summary of the Invention

[0004] In some embodiments, systems and methods are provided that relate to the construction and use of machine learning models that accurately predict whether a subject will tolerate a particular lymphoma treatment. Tolerating a particular treatment can mean not progressing the lymphoma and not suffering adverse events, including death and hospitalization, as a result of the particular treatment. Specifically, a model may be constructed and used to predict whether a subject diagnosed with diffuse large B-cell lymphoma (DLBCL) will tolerate anti-CD20 treatment using CHOP (cyclophosphamide, hydroxydaunorubicin [doxorubicin], Oncovin [vincristine], prednisone) and / or R-CHOP (rituximab, cyclophosphamide, doxorubicin hydrochloride, vincristine, prednisone) as therapeutic agents. The machine learning model, TRAIL (Tolerability of R-CHOP In Aggressive Non-Hodgkin Lymphoma), transforms the values ​​of a set of clinical variables to predict the tolerability of immunochemotherapy in previously untreated subjects with DLBCL. The results generated from the machine learning model can be output to help physicians identify subjects at high risk for intolerance to R-CHOP-based therapy. If the model predicts that a subject will not tolerate R-CHOP therapy, physicians can choose to consider, recommend, or offer alternative therapies. The model may be particularly important for elderly subjects with and without comorbidities. Elderly subjects are less likely to tolerate R-CHOP; however, identifying elderly subjects who can tolerate R-CHOP treatment can improve clinical outcomes across a range of ages. Instead of simply predicting the efficacy of a treatment, predicting the tolerability of a treatment for DLBCL provides improved clinical outcomes for subjects who would otherwise suffer from adverse events as a result of the treatment. A method for predicting the tolerability of a particular treatment, a system for implementing the method, and a computer program product for storing the method are described.

[0005] In one aspect, a method includes accessing an input dataset including a plurality of input data values ​​associated with a particular subject having lymphoma. Each input data value corresponds to a variable of a set of variables. The method includes inputting the input dataset into a machine learning model to generate a score corresponding to the degree to which the particular subject will tolerate a particular treatment. Tolerating a particular treatment may include not ending up with or reducing the particular treatment below a threshold amount within a certain period of time after initiating the particular treatment. The machine learning model includes a set of parameters determined using a plurality of training data elements. Each of the plurality of training data elements corresponds to a training subject. Each of the plurality of training data elements includes a training input dataset and a label. The label indicates the training subject's tolerance to the particular treatment. A function associates the received input dataset and parameters with a score. The method includes outputting a prediction of the particular subject's tolerance to the particular treatment using the generated score.

[0006] In some embodiments, the set of variables in the method may include results from a blood panel. The set of variables may include a characterization of a medical history of the particular subject. The method may include sending a request to a computer system that stores the medical history of the particular subject. The method may include receiving the medical history of the particular subject. The method may further include determining an index value that characterizes a comorbidity of the particular subject, the index value being an input data value of the plurality of data values.

[0007] In some embodiments, the set of variables may include results from an invasive diagnosis. In some embodiments, the set of variables may include albumin concentration, creatinine clearance, a comorbidity index, or the presence of a history of cardiovascular or diabetic conditions. The set of variables may further include bone marrow lymphocyte levels.

[0008] In some embodiments, the set of variables may include no more than 5, no more than 10, or no more than 15 variables. The set of variables may include hemoglobin level, red blood cell count, hematocrit level, number of concomitant medications, chloride level, total levels of CD3 and CD4 protein complexes and T-cell coreceptor, lymphocyte levels, or levels of CD3 protein complexes and T-cell coreceptor.

[0009] In some embodiments, the lymphoma may be diffuse large B-cell lymphoma. Specific treatment may include administration of rituximab, cyclophosphamide, doxorubicin hydrochloride, vincristine, and prednisolone. Specific treatment may include cyclophosphamide or doxorubicin hydrochloride.

[0010] In some embodiments, the method may further include processing a blood sample from a particular subject using a blood panel to determine one or more input data values ​​of the plurality of input data values.

[0011] In some embodiments, the period during which a particular treatment does not end or is not reduced below the threshold amount is 18 weeks or less.

[0012] In some embodiments, the prediction in a particular subject is that they are unlikely to tolerate a particular treatment.

[0013] In some embodiments, the method may further include comparing the score to a cutoff value, where the cutoff value is determined from a plurality of reference subjects, each of which may have a respective predicted score. The cutoff value may be determined to obtain a predetermined level of sensitivity or a predetermined level of specificity for predicting tolerability of a particular treatment in the plurality of reference subjects.

[0014] In some embodiments, the method may further include identifying a subset of the plurality of reference subjects. Each score of each reference subject of the subset may be lower than each score of each reference subject of the plurality of reference subjects other than the subset. The size of the subset may correspond to a predetermined percentage of the size of the plurality of reference subjects. The method may further include defining a cutoff value as a given value that separates each score of the reference subjects of the subset from the scores of reference subjects of the plurality of reference subjects other than the subset. Comparing the score to the cutoff value may include determining that the score exceeds the cutoff value. The prediction may be that a particular subject is unlikely to tolerate a particular treatment.

[0015] In some embodiments, the method may further include, in response to returning a prediction that the particular subject is unlikely to tolerate the particular treatment, outputting a recommendation to enroll the particular subject in a clinical study including a treatment different from the particular treatment.

[0016] In some embodiments, the method may include returning a prediction that a particular subject is unlikely to tolerate a particular treatment, and outputting a recommendation that the particular treatment not be administered to the particular subject.

[0017] In some embodiments, the score may be the probability that a particular subject will tolerate a particular treatment.

[0018] In one aspect, a computer-implemented method includes training a machine learning model to predict whether a particular subject with lymphoma will tolerate a particular treatment. The method includes accessing a training dataset including a plurality of training data elements. Each of the plurality of training data elements corresponds to a training subject with lymphoma. Each of the plurality of training data elements includes a training input dataset and a label. The label indicates the training subject's tolerance to the particular treatment. The training subject is said to have tolerated the particular treatment if, within a certain period of time after initiating the particular treatment, the particular treatment does not end up below a threshold amount or is not reduced below a threshold amount. The method includes training the machine learning model using the training dataset to generate a score corresponding to the degree to which the particular subject will tolerate the particular treatment. Training the machine learning model includes learning a set of parameters and determining a function that associates the set of parameters with a score.

[0019] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein.

[0020] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0021] Some embodiments of the present disclosure include a system comprising one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium containing instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0022] The terms and expressions which have been employed are used as terms of description and not of limitation, and the use of such terms and expressions is not intended to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, while the invention as claimed has been specifically disclosed by embodiments and optional features, it is to be understood that modifications and variations of the concepts disclosed herein may be left to those skilled in the art, and that such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.

[0023] The present disclosure is described in conjunction with the accompanying drawings, in which: [Brief explanation of the drawings]

[0024] [Figure 1] 1 is a process of using a machine learning model to predict the tolerability of a particular treatment in a particular subject according to some embodiments of the present disclosure. [Figure 2] 1 is a process for training a machine learning model according to some embodiments of the present disclosure. [Figure 3]1 is a computer network for using a machine learning model to predict the tolerability of a particular treatment in a particular subject according to some embodiments of the present disclosure. [Figure 4] 1 is a computer system according to some embodiments of the present disclosure. [Figure 5] FIG. 1 is a Venn diagram of hybrid tolerability endpoints according to some embodiments of the present disclosure. [Figure 6] 1 is a graph of overall survival versus time in months with and without an event according to some embodiments of the present invention. [Figure 7] 1 shows a table of odds ratios for baseline characteristics associated with hybrid outcomes according to some embodiments of the present disclosure, with most components listed having two categories of values ​​(e.g., for albumin,<LLNおよび> The variables are divided into four categories (LLN), but the values ​​of the variables are on a continuous scale, not a binary scale. [Figure 8] 1 shows a receiver operating characteristic (ROC) curve for the tolerability of R-CHOP in an aggressive non-Hodgkin's lymphoma (TRAIL) algorithm according to some embodiments of the present disclosure. The highlighted point represents one possible set point for the TRAIL model, trading sensitivity for specificity. [Figure 9] Models for predicting tolerability according to some embodiments of the present disclosure are listed. [Figure 10] 1 lists the top variables in a model with five variables according to some embodiments of the present disclosure. [Figure 11] Lists the top variables in a model with 15 variables according to some embodiments of the present disclosure. [Figure 12] Lists the top variables in a model with 10 variables according to some embodiments of the present disclosure. [Figure 13] Lists the top variables in a model with 15 variables according to some embodiments of the present disclosure. [Figure 14] Lists the top variables in a six-variable model according to some embodiments of the present disclosure. [Figure 15] 1 illustrates prediction of adverse events using a model with four variables on the GOYA holdout dataset according to some embodiments of the present disclosure. [Figure 16] 1 shows prediction of adverse events using a model with four variables on the MAIN R-CHOP-21 validation dataset according to some embodiments of the present disclosure. [Figure 17] 1 shows ROC curves for an example TRAIL model using GOYA training, GOYA holdout, and MAIN external validation data according to some embodiments of the present disclosure. [Figure 18A] 1 shows ROC curves for an example TRAIL model compared to defined prognostic factors using GOYA holdout data and MAIN external validation data according to some embodiments of the present disclosure. [Figure 18B] 1 shows ROC curves for an example TRAIL model compared to defined prognostic factors using GOYA holdout data and MAIN external validation data according to some embodiments of the present disclosure. [Figure 19A] 1 shows the results of a model combined with various IPI levels according to some embodiments of the present disclosure. [Figure 19B] 1 shows the results of a model combined with various IPI levels according to some embodiments of the present disclosure. [Figure 20] 1 shows the area under the curve using cutoff values ​​according to some embodiments of the present disclosure. [Figure 21] 1 shows the percentage of subjects experiencing intolerance events by risk category, according to some embodiments of the present disclosure. [Figure 22A] 1 shows Kaplan-Meier overall survival graphs stratified by risk category for the GOYA holdout dataset and the MAIN external validation data, according to some embodiments of the present disclosure. [Figure 22B] 1 shows Kaplan-Meier overall survival graphs stratified by risk category for the GOYA holdout dataset and the MAIN external validation data, according to some embodiments of the present disclosure. [Figure 23]1 shows Kaplan-Meier overall survival graphs stratified by risk category in the Flatiron RWD (real-world data) dataset, according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0025] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dashed line and a second label that distinguishes the similar components. When only a first reference label is used in this specification, the description is applicable to any one of the similar components having the same first reference label, regardless of the second reference label.

[0026] I. Overview Previously, models have been developed to predict the effectiveness of specific treatments for diseases such as lymphoma. These models can help determine whether a specific treatment will effectively treat a subject's lymphoma. However, in some instances, a specific treatment may actually be harmful to some subjects. For example, even if a treatment effectively treats a subject's lymphoma, it may induce adverse events, resulting in harm to the subject (e.g., net harm). In some subjects, a specific treatment may result in death, hospitalization, or serious illness. Avoiding such negative outcomes is desirable. No model has been developed to date that determines whether a specific subject can tolerate a specific treatment for lymphoma, instead of the effectiveness of the treatment.

[0027] A system and method for predicting whether a subject will tolerate a particular treatment has been developed. A machine learning model using an existing dataset is trained to predict whether a subject will tolerate a particular treatment for lymphoma. The machine learning model can include a relatively small number of easily obtainable variables and can accurately predict whether a subject will tolerate or not tolerate a particular treatment.

[0028] Accurately predicting the tolerability of a particular treatment in a subject can improve outcomes for subjects with lymphoma. Subjects with lymphoma who have a high rate of treatment-related adverse effects may be able to avoid conventional treatments and unfavorable clinical outcomes. The developed machine learning model provides a more accurate method for predicting adverse effects in subjects than conventional prognostic indicators such as the International Prognostic Index (IPI) or Eastern Cooperative Oncology Group Performance Status (ECOG PS).

[0029] For patients with newly diagnosed diffuse large B-cell lymphoma (DLBCL), the standard initial treatment is rituximab in combination with cyclophosphamide, doxorubicin, vincristine, and prednisolone (R-CHOP). Cyclic treatment with R-CHOP has proven effective in most patients.

[0030] Although R-CHOP has been the mainstay of DLBCL treatment for over 20 years, this regimen is associated with significant toxicity. Many DLBCL patients, particularly elderly patients, exhibit reduced physiologic reserve and often comorbidities, which can prevent adequate administration of R-CHOP. Among patients aged 66 years or older with newly diagnosed non-Hodgkin lymphoma in the Surveillance, Epidemiological, and End Results (SEER)-Medicare database, 52% had one or more comorbidities, and 26% had a Charlson comorbidity index greater than 2. The most common comorbidities were diabetes (25%), chronic obstructive pulmonary disease (16%), and congestive heart failure (12%). A retrospective cohort study of approximately 18,000 subjects aged 66 years or older diagnosed with DLBCL between 2001 and 2013 in the SEER-Medicare database reported that subjects aged >80 years were less likely to receive R-CHOP as initial treatment compared with subjects aged 66–80 years (46.5% of subjects >80 years vs. 71% of subjects aged 80 years or younger).

[0031] Although elderly subjects with DLBCL may be more susceptible to the toxicities of R-CHOP, not treating them with R-CHOP can also lead to poor outcomes. Outcomes have been found to be inferior for subjects who do not complete R-CHOP treatment compared with those who complete treatment. Subjects who are unable to complete R-CHOP treatment due to intolerance to R-CHOP face worse DLBCL-related outcomes than subjects who are able to complete R-CHOP treatment. Early identification of subjects unlikely to tolerate R-CHOP can improve the chances of avoiding R-CHOP-related adverse events and also allow these subjects the opportunity to consider clinical trials or other treatments. At the same time, subjects predicted to tolerate R-CHOP may have improved clinical outcomes if they receive complete R-CHOP treatment compared with those who receive partial R-CHOP treatment.

[0032] The model described herein provides a combination of superior prognostic performance, clinical usefulness, and simplicity compared to conventional methods.This particular model includes four simple variables, all of which can be easily obtained in routine clinical practice.Examples of variables include Charlson comorbidity index, the presence of cardiovascular disease or diabetes, low serum albumin, and low creatinine clearance, which are frequently found in subjects with DLBCL.

[0033] Predicting the outcome of subjects during induction treatment is a difficult task. It is clear that not all collected variables contain information useful for assessing induction success.

[0034] Models developed according to the methods disclosed herein can be used as clinical decision support to guide physician decisions. Physicians can use the model results, combined with clinical trial or other data, to determine a subject's course of treatment. Model results can also enable subjects to enroll in clinical trials for novel therapies. These novel therapies may differ from and / or not include all or some components of anti-CD20 or R-CHOP. Some embodiments can eliminate or reduce anti-CD20 or R-CHOP-associated toxicity for subjects with DLBCL who are predicted to be at risk of suffering from adverse effects or who are predicted to be unlikely to benefit from treatment. In some embodiments, models can be used for subjects undergoing a particular treatment to predict a subject's tolerability to the treatment and / or to customize monitoring procedures (e.g., customize tests performed and / or test frequency to improve the subject's comfort).

[0035] II. Definition CHOP refers to the combination of cyclophosphamide, hydroxydaunorubicin (doxorubicin), Oncovin (vincristine), and prednisone.

[0036] R-CHOP refers to CHOP and rituximab.

[0037] DLBCL refers to diffuse large B-cell lymphoma.

[0038] A "biological sample" refers to any sample obtained from a subject (e.g., a human subject, such as one with lymphoma). A biological sample may be a bodily fluid, such as blood, plasma, serum, urine, vaginal fluid, fluid from a testicular cyst (e.g., of a testicle), vaginal flushing fluid, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirates from various parts of the body (e.g., thyroid, breast), etc. Fecal samples can also be used.

[0039] "Adverse Event" or AE refers to an event including death, illness requiring hospitalization, a life-threatening event, or persistent disability or disability.

[0040] "Tolerating" in terms of a subject tolerating a treatment refers to not prematurely dropping below a threshold dose (e.g., 80%) or reducing the specific treatment below a threshold dose within a certain period of time after initiation of treatment as a result of death or other adverse events. The adverse events may be due to or be caused by some side effect of the specific treatment rather than the progression of the lymphoma. The period may be 1-2, 2-4, 4-8, 8-12, 12-16, 16-18, 18-24, 24-30, 30-40, 40-50, or 50-52 weeks. The threshold dose may be 50%-60%, 60%-70%, or 70%-80% of the planned dose. The planned dose may be determined for the subject at the time of diagnosis or prior to a subsequent treatment cycle. A subject "tolerating" a specific treatment often refers to a subject known to complete a specific treatment above a threshold dose within a certain period of time.

[0041] "Tolerance" refers to the degree to which a subject responds well to a particular treatment. Tolerance may be a binary measure indicating whether the treatment was tolerated or not. In some embodiments, tolerance may be a real or integer number along a scale. For example, a subject who dies from a particular treatment may be considered to have lower tolerance for the particular treatment than a subject whose dose of the treatment is reduced by 70% of the planned dose. Tolerance may be a numerical scale (e.g., 0 to 1) or a discrete classification (e.g., low, medium, high).

[0042] "Tolerance" refers to the likelihood or probability that a subject will tolerate a particular treatment. "Tolerance" often refers to the likelihood or probability for a subject before or during treatment before an adverse event occurs. For example, low "tolerance" may correspond to a high likelihood of unplanned hospitalization or emergency room visits.

[0043] "Disease progression" or PD refers to a worsening of the disease characterized by an increase in disease burden or new lymphoma lesions.

[0044] III. EXEMPLARY METHODS AND SYSTEMS FOR PREDICTING TOlerability A method for predicting the tolerability of a particular treatment in a particular subject is described. Figure 1 shows an exemplary process 100 for using a computational model to predict tolerability.

[0045] A subject who cannot tolerate a particular treatment will, within a certain period after the start of treatment, reduce the specific treatment below a threshold amount (for example, 80%) as a result of adverse events, or end the specific treatment as a result of such hospitalization or death.For example, a subject who cannot tolerate a treatment may have at least one of the following events during the first few treatment cycles: (1) adverse events (AE) that are not the result of disease progression and AE that lead to withdrawal from treatment before completion of treatment; (2) death (PD) that is not the result of disease progression; or (3) a reduction in dose intensity that results in a dose intensity below a certain threshold (this reduction is not related to disease progression).Chemotherapy is often administered at regular intervals called "cycles".Typical intervals include every 14, 21, or 28 days.

[0046] A particular treatment may include multiple treatment cycles, each of which may be 14 or 21 days long. The number of cycles may be 2, 3, 4, 5, 6, 7, 8, 9, or 10. For example, in the case of tolerability to CHOP-type treatment, intolerance to treatment may include an adverse event (AE) that leads to discontinuation of cyclophosphamide and / or doxorubicin before completion of a specific number of cycles below the default number. For example, intolerance to treatment may include failure to continue treatment after completing less than 90%, 80%, 70%, 60%, 50%, or 40% of the default number of treatment cycles. As another example, intolerance to treatment may include failure to continue treatment when the planned number of cycles of a particular treatment is greater than the number of completed cycles, or when the number of completed cycles is equal to 8, 6, 4, 2, or 1. Additionally, intolerance to treatment can include death not as a result of disease progression (PD) or a mean dose intensity of cyclophosphamide and / or doxorubicin reduced by less than 80%.

[0047] At block 110, an input dataset is accessed, the input dataset including a plurality of input data values ​​associated with a particular subject having lymphoma. The lymphoma may be diffuse large B-cell lymphoma (DLBCL). The particular subject may be human. The particular subject may range in age, including 60, 65, 70, 75, 80, or 85 years of age. The lymphoma may include (for example) non-Hodgkin's lymphoma, Hodgkin's lymphoma, chronic lymphocytic leukemia, cutaneous B-lipolymphoma, cutaneous T-cell lymphoma, or Waldenstrom's macroglobulinemia.

[0048] Each input data value corresponds to a variable in the set of variables. The set of variables may include results obtained from a blood panel, including a traditional blood panel, or a complete blood count. Blood panel variables may include levels of calcium, carbon dioxide, chloride, creatinine, glucose, potassium, sodium, and / or urea nitrogen. A complete blood count includes red blood cell, white blood cell, hemoglobin, hematocrit, and platelet counts. In some embodiments, the set of variables may include a characterization of a subject's medical history. For example, a variable may indicate the presence and type of any co-morbidities, such as cardiovascular disease or diabetes. A variable may represent results obtained from an invasive diagnostic, which may include a bone marrow sample or a biopsy. A model controller system, which may include a computer system that stores the machine learning model, has access to the input dataset.

[0049] The set of variables may include albumin concentration (ALBUMIN), creatinine clearance (CRCL), a comorbidity index (e.g., Charlson Comorbidity Index [CHARLSON]), the presence of a history of cardiovascular disease or diabetes (e.g., Heart, Vascular, and Diabetic Complications [HVD]), or a combination thereof. In some embodiments, the set of variables may include a combination of all four of these variables. In some embodiments, the set of variables may include a level of bone marrow lymphocytes. The level of bone marrow lymphocytes may be an absolute or normalized count (e.g., concentration, percentage), mass, or volume. The plurality of input data values ​​may include 1 to 5, 5 to 10, 10 to 15, or more than 15 values. The plurality of input data values ​​may include any value described herein. A value may be a count, a concentration, or an indication of the presence or absence of a component. Exemplary variables include hemoglobin (HGB) level, red blood cell count (RBC), hematocrit (HCRI), number of concomitant medications (CMCNT), chloride level (CHLOR), total levels of CD3 and CD4 protein complexes and T-cell coreceptor (F13005), total levels of CD3 and CD4 protein complexes and T-cell coreceptor and lymphocytes (F13005LY), age (AGE), and / or levels of CD3 protein complexes and T-cell coreceptor (CD3). In some embodiments, variables may include current quality of life health survey responses (e.g., QLQC30 obtained at eortc.org / app / uploads / sites / 2 / 2018 / 08 / Specimen-QLQ-C30-English.pdf (accessed July 12, 2021)).

[0050] Process 100 may include obtaining a plurality of input data values ​​for a particular subject. In some examples, one or more of the plurality of input data values ​​corresponds to one or more laboratory variables. Thus, obtaining the plurality of input data values ​​may include performing assays (e.g., a blood panel) on a biological sample obtained from the particular subject or receiving data from a computing system associated with a laboratory that performed assays on the biological sample.

[0051] Process 100 may include sending a request to a computing system (e.g., a care provider system associated with a care provider or a laboratory system associated with a testing device) to access an input dataset. Some or all of the input dataset may be stored on the laboratory system. The laboratory system may include a computer system that performs or is used to perform assays on a particular subject. The laboratory system may store data from assays performed by a technician or physician. In some embodiments, the care provider system may store some or all of the input dataset. For example, the care provider system may store a medical history of a particular subject. The medical history of a particular subject may be entered into the care provider system immediately before (e.g., within an hour or within a day) accessing the input dataset. In some embodiments, the medical history of a particular subject has been entered into the care provider system over weeks or years due to the particular subject's visits to a care provider. The medical history of a particular subject may be received by a model controller system. The received data may optionally be preprocessed. For example, the identification of one or more comorbidities may be detected in the medical record data, and an index value may be generated based on the detection. Index values ​​may include the Charlson Comorbidity Index or the Cardiac, Vascular, and Diabetic Complications Index.

[0052] Data from a laboratory system may be preprocessed by the laboratory system before being included in the input dataset. Data values, such as concentrations or counts, may be normalized by a reference value (e.g., to account for different sampling techniques) or multiplied by a calibration factor (e.g., to account for different biological sample testing equipment). Data values ​​may also be classified into different categories. Classifications may have binary categories (e.g., present or absent, normal or abnormal) or may have three or more categories (e.g., very low, low, normal, high, very high). Categories may be represented numerically (e.g., present may equal 1, and absent may equal 0). As an example, age may be divided into categories of 0-17, 18-64, and 65 or greater, with these categories represented by the numerical values ​​"1," "2," and "3." In some embodiments, data values ​​may be processed to generate index values ​​that characterize the data values. For example, medical history data may include dates, durations, and the severity of various comorbidities. Index values ​​(eg, the Charlson Comorbidity Index) can be calculated based on medical history data to characterize the risk of various comorbidities.

[0053] At block 120, the input dataset is input into the machine learning model to generate a score corresponding to the degree to which a particular subject is predicted to tolerate a particular treatment. The particular treatment may include administration of rituximab, cyclophosphamide, doxorubicin hydrochloride, vincristine, and prednisolone (R-CHOP). In other embodiments, the particular treatment may include administration of cyclophosphamide, doxorubicin hydrochloride, or a combination thereof. The treatment may include M-CHOP (mosuntuzumab and CHOP) treatment. The machine learning model may be stored in the model controller system.

[0054] The score generated by the machine learning model may be the predicted probability of a particular subject to tolerate a particular treatment. The probability may be weighted according to the severity of the adverse event. For example, the probability of death may be increased by a factor that takes into account the higher impact of death compared to other adverse events or by reducing the dose intensity.

[0055] The machine learning model may include a set of parameters trained using multiple training data elements. Each of the multiple training data elements corresponds to a training subject. Each of the multiple training data elements includes a training input dataset (e.g., including normalized data values, comorbidity indices, and data categories) and a label. The label indicates the training subject's tolerance to a particular treatment. Tolerance may be a binary classification corresponding to "tolerable" and "intolerant." In some embodiments, the label may indicate the severity of an adverse event suffered by the training subject. For example, a label of 1 may indicate death, a label of 0.5 may indicate hospitalization, and a label of 0 may indicate no adverse event. In some examples, the scale may be reversed, with labels indicating tolerance levels. For example, a label of 1 may indicate completion of treatment, and a label of 0 may indicate death.

[0056] The machine learning model may include a function that associates the received input dataset and parameters with a score. The function may be created by training the machine learning model to predict a tolerability label for a training subject with a training input dataset. The score generated for a particular subject may be the particular subject's predicted tolerability to a particular treatment.

[0057] In some examples, the training input dataset includes values ​​for each of a set of variables, and the machine learning model is configured to receive values ​​of only a subset of the set as inputs for conversion into a score. The subset can be identified by determining the degree to which each variable in the set predicts tolerance. The determination can include performing a univariate analysis (e.g., for each variable in the set individually) or a multivariate analysis (e.g., using some or all of the variables in the set). For example, a Pearson product-moment correlation can be evaluated between the variables and tolerance. Variables with a higher correlation can be identified as informative variables. Additionally or alternatively, these informative variables can be analyzed to evaluate the value of each variable in the predicted tolerance. The analysis may use variable transformation or ranking algorithms (e.g., GameRank (Huang TK, Lin CJ, Weng RC, Ranking individuals by group comparisons. Proceedings of the 23rd international conference on Machine learning. Pittsburgh, Pennsylvania, USA: Association for Computing Machinery, 2006: pp. 425-432), the entire contents of which are incorporated herein by reference for all purposes).

[0058] In some embodiments, the training input dataset may include training subjects with a lower median age and / or fewer median comorbidities than subjects in the dataset used for validation or actual use. The median age of the training input dataset may be 1-5 years, 5-10 years, 10-15 years, 15-20 years, or more than 20 years lower than the median age of the validation dataset or specific subjects. The median age of the training input dataset may be 30-40 years, 40-50 years, 50-60 years, or 60-70 years. Thus, generated scores are trained and developed in younger populations, population-matched, and then validated or practiced in older populations. Younger populations with fewer comorbidities can enable a deeper understanding of whether subjects will tolerate treatment by eliminating or reducing other known causes of hospitalization and death.

[0059] The machine learning model can use the informative variables to determine different functions (for predicting tolerability). Trial-and-error modeling methods can be used to improve the model's performance or overall accuracy. For example, the machine learning model can dynamically adjust parameters and / or configurations based on cross-validated area under the curve (AUC) or bootstrap optimism metrics obtained from previous iterations. The importance of variables in these functions can be evaluated using a ranking algorithm. The performance of functions in accurately predicting tolerability in the validation dataset can be compared with each other. The function selected for the machine learning model can be the model that most accurately predicts tolerability or that balances overall accuracy and computational efficiency. The function can be limited to a certain number of variables (e.g., 1-5, 5-10, 10-15) that may be the most informative, most correlated, or top-ranked variables.

[0060] Machine learning models may include, for example, regression models (e.g., linear and / or logistic regression models), artificial neural networks, decision trees, random forest models, support vector machines (SVMs), naive Bayes classification, clustering algorithms, principal component analysis (PCA), singular value decomposition (SVD), t-distributed stochastic neighbor embedding (tSNE), and ensemble models that classify new data points by building a set of classifiers and then weighted voting their predictions.

[0061] In block 130, the score is used to output a prediction of the particular subject's tolerance to the particular treatment. The prediction may be a non-numeric outcome. The prediction may be that the subject is unlikely to tolerate the treatment. For example, the prediction may include that the subject is likely to suffer an adverse event, such as death, illness requiring hospitalization, a life-threatening event other than lymphoma, or permanent disability or incapacity. The prediction may also be that the subject is likely to tolerate the treatment. In some embodiments, the prediction may be a numerical outcome. The prediction may be the likelihood or probability of the particular subject to tolerate the particular treatment. For example, if the score is a probability of tolerance for the particular subject, the prediction may be equal to the score. In some embodiments, the score may be a binary classification of whether the particular subject is likely to tolerate the particular treatment, and the output of the prediction corresponds to the binary classification. The prediction is not a prediction of the effectiveness of the treatment for lymphoma. The prediction may be output and transmitted to a care provider system.

[0062] Determining a prediction may require a machine learning model or post-processing technique that compares the score to a cutoff value. The cutoff value can be determined using the scores determined by the machine learning model for multiple reference subjects. For example, a subset of the multiple reference subjects can be identified, where the score of each reference subject in the subset is lower than the score of each reference subject not in the subset. The size of the subset can correspond to a predetermined percentage of the size of the multiple reference subjects. The cutoff value can be defined as a given value that separates the scores of reference subjects in the subset from the scores of reference subjects not in the subset. Thus, the cutoff value can correspond to the percentile of the reference subject with the lowest score. For example, the cutoff value can correspond to at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 90%, or 95% of the scores from the reference subjects. If the particular subject is determined to be in the subset of reference subjects with the lowest scores, the prediction may be that the particular subject will tend not to tolerate the particular treatment. If the particular subject is not in the subset of reference subjects with the lowest scores, the prediction may be that the particular subject will tend to tolerate the particular treatment.

[0063] The cutoff value can be determined by a classifier using a loss function that prioritizes identifying subjects who tolerate (or do not tolerate) treatment. For example, the cutoff value can be determined to avoid false negatives or false positives. The machine learning model can determine interface detachment clusters using a clustering model (e.g., PCA or independent component analysis [ICA]). The clusters can represent different degrees of tolerability (e.g., death, hospitalization, dose intensity reduction). The cutoff value can represent a linear, planar, or hyperplane between the clusters. Comparing the score to the cutoff value can include calculating a tolerability metric from the score and other components identified by the clustering model, and then comparing the tolerability metric to the cutoff value.

[0064] Comparing the score to the cutoff value may include determining that the score exceeds the cutoff value. Based on the comparison of the score above the cutoff value, the machine learning model (or post-processing analysis by the care provider system) can determine that the particular subject is included in or tends to be included in a particular subset of subjects with similar tolerance (e.g., a subset having subjects who tolerate the treatment, a subset having subjects who do not tolerate the treatment but do not die from the treatment). The machine learning model can then output a prediction of the particular subject's tolerance to a particular treatment by using the characteristics of the subset. For example, if the particular subject is included in a subset having subjects who do not die from the particular treatment, the output prediction may be that the particular subject is likely not to die from the particular treatment. In some embodiments, the score for the particular subject may be within the cutoff value, indicating or suggesting that the particular subject belongs to a particular subset.

[0065] The cutoff value can be selected to achieve a predetermined overall accuracy rate. For example, the cutoff value can be determined to achieve a predetermined level of sensitivity or specificity for predicting whether a subject will tolerate a particular treatment in a plurality of reference subjects. The predetermined level of sensitivity or specificity can be, independently, 60% to 70%, 70% to 80%, 80% to 90%, or 90% to 95%. The cutoff value can be selected to achieve a predetermined area under the curve (AUC) in a receiver operating characteristic (ROC) curve. The predetermined AUC can be 0.6 to 0.7, 0.7 to 0.8, 0.8 to 0.9, or 0.9 to 0.95.

[0066] The prediction of tolerability can be determined using only the score. In some examples, the prediction may use various prognostic indicators other than the score. For example, the prediction may use the average of various prognostic indicators, including the International Prognostic Index (IPI), ECOG PS (Eastern Cooperative Oncology Group Performance Status), age, or geriatric assessment.

[0067] In response to a prediction that a particular subject will not tolerate or is unlikely to tolerate a particular treatment, a recommendation to enroll the particular subject in a clinical trial may be output. The enrollment of the particular subject may be by a physician or by a treatment device, which may be any of the treatment devices described herein. A recommendation that the physician not administer the particular treatment to the particular subject may be output. The recommendation may include administering the particular treatment at a lower dose. The physician may be recommended to administer an alternative treatment, such as a different combination drug product, radiation therapy, bone marrow transplant, or palliative care. The recommendation may be transmitted to a care provider system, which may then output the recommendation. Process 100 may include administering the alternative treatment by a treatment device or a medical professional.

[0068] On the other hand, if the prediction indicates that the particular subject will tolerate or be prone to tolerate the particular treatment, a recommendation may be output to administer the particular treatment to the particular subject.

[0069] Process 100 may be repeated one or more times after a particular subject begins a particular treatment. The generated score may be monitored over time. If the score indicates a change from a previous score, the particular treatment for the particular subject may be interrupted or terminated. The change may be a change in score that results in a different output prediction. In some embodiments, the change may be a statistically significant change (e.g., more than 1, 2, or 3 standard deviations based on the score of a reference subject).

[0070] In an alternative embodiment, the input dataset may include raw data from an assay, and the model controller system may preprocess (e.g., normalize or classify) the raw data. Additionally, scores generated by the machine learning model may be compared to cutoff values ​​during post-processing steps performed by the care provider system. Furthermore, outputs and recommendations by the model controller system 320 may be displayed or generated by the care provider system.

[0071] III.A. Training method Some embodiments may include a computer-implemented method for training a machine learning model. Figure 2 shows a process 200 for training a model (e.g., the machine learning model used in block 120 of process 100).

[0072] At block 210, a training dataset is accessed. The training dataset includes a plurality of training data elements. Each of the plurality of training data elements corresponds to a training subject with lymphoma. Each of the plurality of training data elements includes a training input dataset and a label. The label indicates the training subject's tolerance to a particular treatment. The label may be any label described herein. For example, the label may be a binary indication of whether the subject tolerates a particular treatment (e.g., "yes" or "no"). In other examples, the label may reflect various severities of instances in which the subject did not tolerate a particular treatment. The training subjects may include subjects who tolerated a particular treatment and subjects who did not tolerate a particular treatment (e.g., suffered from an adverse event). In some embodiments, the training subjects may be at least 60, 65, 70, 75, 80, or 85 years old. The training dataset may have values ​​described for the input dataset in process 100 of FIG. 1.

[0073] In block 220, a machine learning model is trained to generate a score corresponding to the degree to which a particular subject with lymphoma will tolerate a particular treatment. The training uses a training dataset. Training the machine learning model includes learning a set of parameters. The set of parameters may include any of the parameters described herein. Training the machine learning model also includes determining a function relating the set of parameters to predicted tolerability.

[0074] Machine learning or deep learning algorithms can be used for training, including but not limited to support vector machines (SVMs), decision trees, naive Bayes classification, logistic regression, clustering algorithms, principal component analysis (PCA), singular value decomposition (SVD), t-distributed stochastic neighbor embedding (tSNE), artificial neural networks, and ensemble methods that build a set of classifiers and then weighted vote their predictions to classify new data points.

[0075] The function relating the set of parameters to predicted tolerability may include any of the formulas described herein. The parameters may include input data values ​​and coefficients and / or operators (e.g., log, square root, power). The function may be a linear or non-linear combination of the parameters.

[0076] An example of using the GOYA dataset to train a machine learning model is provided below.

[0077] III.B. System Example FIG. 3 illustrates a system 300 that may be used to perform the method or steps of process 100. System 300 includes a care provider system 310. Care provider system 310 may include a computer system and may be operated by a healthcare provider, including a clinic, hospital, or physician. The healthcare provider operating care provider system 310 may be responsible for treating a particular subject for lymphoma in process 100 of FIG. 1. The healthcare provider may identify a particular subject as eligible for a particular treatment. The healthcare provider may enter information about the particular subject and the particular treatment into care provider system 310. For example, the name of a particular subject and the type of treatment (e.g., R-CHOP or any treatment described herein) may be entered into care provider system 310.

[0078] The care provider system 310 can send a request to the model controller system 320 to provide a prediction of a particular subject's tolerance to a particular treatment. The model controller system 320 may include a computer system. The model controller system 320 may include a non-transitory computer-readable storage medium that stores the machine learning model and includes instructions for executing the steps of process 100. The model controller system 320 can access an input dataset that includes a plurality of input data values, as in block 110, where each input data value corresponds to a variable of the set of variables. The input dataset can be stored in the model controller system 320.

[0079] The model controller system 320 can send requests to the laboratory system 330 to access input data values. The laboratory system 330 may include multiple laboratory instruments. These laboratory instruments can perform assays (e.g., blood panels) on biological samples from particular subjects. Data from these assays can be stored in a data storage device in the laboratory system 330. The laboratory system 330 can send input data values ​​to the model controller system 320.

[0080] Similar to block 120 of process 100, an input data set can be input to a machine learning model in model controller system 320. Model controller system 320 can look up values ​​for particular variables in the input data set and then place the values ​​in appropriate locations in the machine learning model. The machine learning model can be the machine learning model described in process 100 of FIG. 1 and trained in process 200 of FIG. 2. Model controller system 320 can then run the machine learning model to generate a score corresponding to the degree to which a particular subject will tolerate a particular treatment. The score can be generated similar to block 120 of process 100.

[0081] The generated score may be used to determine a prediction of a particular subject's tolerance to a particular treatment, as in block 130 of process 100. The prediction may be transmitted from model controller system 320 to care provider system 310. Care provider system 310 may display the prediction or communicate the prediction to a physician. In some embodiments, care provider system 310 may use the generated score to determine a prediction of tolerance.

[0082] The care provider system 310 can determine a recommendation based on the tolerability prediction and output the recommendation. The recommendation may be the recommendation described in process 100. The recommendation may be displayed or communicated to a physician who can adopt the recommendation. In some embodiments, the recommendation may be determined by the model controller system 320 and sent to the care provider system 310.

[0083] It will be appreciated that in alternative embodiments, the care provider system 310 and the model controller system 320 can share components. For example, the instructions associated with the care provider system 310 may occupy one portion of a data storage device, while the instructions associated with the model controller system 320 may occupy another portion of the same data storage device.

[0084] III.C. Example Computer System Any of the computer systems referred to herein, including those for care provider system 310, model controller system 320, and laboratory system 330, can utilize any suitable number of subsystems. An example of such a subsystem is shown in computer system 10 of FIG. 4. In some embodiments, a computer system includes a single computer device, where the subsystems can be components of the computer device. In other embodiments, a computer system can include multiple computer devices, each with its own internal components, that are subsystems. Computer systems can include desktop and laptop computers, tablets, mobile phones, and other portable devices.

[0085] The subsystems shown in FIG. 4 are interconnected via a system bus 75. Additional subsystems are shown, such as a printer 74, a keyboard 78, a storage device 79, and a monitor 76 (e.g., a display screen, such as an LED) coupled to a display adapter 82. External and input / output (I / O) devices coupled to an I / O control device 71 may be connected to the computer system by any number of means known in the art, such as input / output (I / O) ports 77 (e.g., USB, Lightning). For example, the I / O ports 77 or external interface 81 (e.g., Ethernet, Wi-Fi, etc.) may be used to connect the computer system 1500 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus 75 not only enables information exchange between the subsystems, but also allows the central processor 73 to communicate with each subsystem and control instruction execution from the system memory 72 or storage device 79 (e.g., a fixed disk such as a hard drive or optical disk). The system memory 72 and / or storage device 79 may embody computer-readable media. The system memory 72 and / or storage device 79 can store input data sets, machine learning models, sets of parameters, functions, and results generated from the models. Another subsystem is a data collection device 85, such as a camera, microphone, accelerometer, etc. Any of the data mentioned herein can be output from one component to another, from one computer system to another, and to a user. Data can be output to a user through a monitor 76.

[0086] A computer system may include several identical components or subsystems connected together, for example, by an external interface 81, by an internal interface, or through a removable storage device that can be connected and disconnected from one component to another. In some embodiments, computer systems, subsystems, or devices may communicate over a network. In such an example, one computer may be considered a client and another computer may be considered a server, each of which may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0087] Aspects of some embodiments may be implemented in the form of control logic using hardware circuitry (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or using computer software, typically involving programmable processors in a modular or integrated fashion. As used herein, a processor may include a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board, or networked, as well as dedicated hardware. Based on the disclosure and teachings provided in this disclosure, those skilled in the art will know and understand other ways and / or methods of implementing some embodiments of the present disclosure using hardware and combinations of hardware and software.

[0088] Any software components or functions described in this application may be implemented as software code executed by a processor using any suitable computer language, such as, for example, Java, C, C++, C#, Objective-C, Swift, etc., or scripting languages, such as, for example, Perl or Python, using conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive, or optical media such as a compact disc (CD) or DVD (Digital Versatile Disc) or Blu-ray disc, flash memory, etc. The computer-readable medium may be any combination of such storage or transmission devices.

[0089] Such programs may also be encoded and transmitted using carrier signals adapted for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. In this manner, computer-readable media may be created using data signals encoded with such programs. Computer-readable media encoded with program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable medium may be provided on or within a single computer product (e.g., a hard drive, CD, or complete computer system) or may reside on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for outputting any of the results described herein to a user (e.g., a physician).

[0090] Any of the methods described herein can be implemented, in whole or in part, using a computer system including one or more processors that can be configured to perform the steps. Accordingly, some embodiments can be directed to a computer system configured to perform any of the steps of the methods described herein, potentially using different components that perform each step or group of steps. While the steps of the methods herein are presented as numbered steps, they may be performed simultaneously or at different times, or in different orders, where logically possible. Additionally, some of these steps may be used with some of the other steps from other methods. Also, all or some steps may be optional. Additionally, any of the steps of any of the methods may be performed using a system module, unit, circuit, or other means for performing these steps.

[0091] IV. Examples Subjects from the GOYA (Vitolo U, Trneny M, Belada D, et al. Obinutuzumab or rituximab plus cyclophosphamide, doxorubicin, vincristine, and prednisone in previously untreated diffuse large B-cell lymphoma. J Clin Oncol. 2017;35(31):3529-3537) and MAIN (Seymour JF, Pfreundschuh M, Trneny M, et al. R-CHOP with or without bevacizumab in patients with previously untreated diffuse large B-cell lymphoma: final MAIN study outcomes. Haematologica. 2014;99(8):1343-1349) trials were studied to illustrate factors associated with poor R-CHOP tolerability in subjects with DLBCL. The objective was to develop a model (TRAIL or R-CHOP Tolerance in Aggressive Lymphoma) to predict the risk of intolerance to R-CHOP induction therapy using variables captured at baseline and to validate the resulting predictions using data from GOYA and data from subjects receiving 21-day cycles of R-CHOP in the external independent MAIN trial.

[0092] For the purposes of this analysis, the composite tolerability endpoint (intolerance to CHOP) was defined as the occurrence of at least one of the following events within the first six cycles: discontinuation of cyclophosphamide and / or doxorubicin before completing six cycles, death without disease progression (PD), or a mean dose intensity of cyclophosphamide and doxorubicin <80% not associated with PD. The rationale for selecting an endpoint based on these components is to identify subjects who do not receive optimal CHOP chemotherapy dose intensity due to poor tolerability.

[0093] IV.A. Model Development with the GOYA Dataset Data for model development were obtained from the multicenter, open-label, randomized, phase 3 GOYA trial (NCT01287741). Subjects with DLBCL were recruited from 207 centers in 29 countries between July 2011 and June 2014. Eligible subjects were aged 18 years or older, previously untreated, and diagnosed with CD20-DLBCL-positive disease. Other inclusion criteria included: Eastern Cooperative Oncology Group Performance Status (ECOG-PS) of 0-2; adequate hematological, liver, and renal function; a left ventricular ejection fraction of 50% or greater; and no significant, uncontrolled comorbidities. Subjects in GOYA were treated with either six or eight 21-day cycles of R-CHOP or obinutuzumab plus CHOP (G-CHOP). The median follow-up period at the time of final analysis (January 2018) was 48 months. Because there were no statistical differences in outcomes detected between the obinutuzumab and rituximab treatment groups in GOYA, data from both treatment groups were pooled for this analysis.

[0094] Of the 1,407 evaluable subjects who initiated anti-CD20 + CHOP, the hybrid binary tolerability endpoint occurred in 188 (13.4%) subjects (Figure 5). 104 subjects (7.4%) had adverse events (AEs) leading to discontinuation of cyclophosphamide and / or doxorubicin, 162 subjects (11.5%) had an intentional reduction in the mean relative dose intensity of cyclophosphamide and doxorubicin below 80%, and 19 subjects (1.4%) died within the first six cycles unrelated to progressive disease (PD). Figure 6 shows a graph of overall survival versus time in months with and without events. The x-axis represents time in months. The y-axis represents overall survival. The top line on the graph (line 610) represents subjects without adverse events. The bottom line on the graph (line 620) represents subjects with adverse events. Figure 6 shows that subjects experiencing adverse events have poor overall survival.

[0095] Retrospective statistical modeling was performed using data from the GOYA safety-evaluable population (n = 1407). The GOYA data were partitioned into a 75% training set (n = 1051) used for model training and cross-validation, and a 25% holdout set (n = 356) blinded to the analysis team and serving as a test set for validation. The training set was divided into four quadrants for cross-validation, each containing 260–265 subjects. These quadrants and holdout sets were created in a manner that achieved balanced distributions of the following variables: immunotherapy arm, International Prognostic Index (IPI) category, PFS events, and two-way product sum of lymphoma disease. Further details on the generation of the balanced four-fold cross-validation are available in supplementary methods. Model performance was assessed using additional randomized 10-fold cross-validation and bootstrapping using the GOYA training data.

[0096] Model development included the following tasks: assessing correlations between individual predictors and outcomes, examining predictive distributions, creating transformed predictors if necessary, and finally a series of model evaluation steps. Pre-model validation steps were performed only on the GOYA training data and not on the holdout data.

[0097] III.A.1. Predictor Screening Associations between candidate covariates and the composite endpoint were first examined to identify advanced prognostic indicators and construct informative variables. Pearson product-moment correlations (95%, CI, and p-value) were assessed between each predictor and the composite endpoint. Correlations between predictors and individual components of the composite endpoint were used for predictor selection, and correlations between each pair of individual predictors were also assessed.

[0098] III.A.2. Constructing predictors As a result of the correlation analysis, univariate predictor distributions were further analyzed to construct informative variables. The distributions of candidate predictors were examined, and various transformations were considered (e.g., categorization of continuous variables, evaluation of nonlinear transformations (square root, logarithm, Box-Cox)). To evaluate the predictive value of each predictor, we implemented a ranking algorithm, GameRank (Huang TK, Lin CJ, Weng RC. Ranking individuals by group comparisons. Proceedings of the 23rd international conference on Machine learning. Pittsburgh, Pennsylvania, USA: Association for Computing Machinery; 2006: 425-432).

[0099] III.A.3. Determine the TRAIL model Determining the final model for prediction was an iterative process. Starting with a set of initial models based on highly correlated predictors, trial-and-error modeling methods were subsequently used to improve cross-validated area under the curve (AUC) and bootstrap optimism. Bootstrap estimates of model generalization error and optimism were obtained using the 632 method. The AUCs of the partitions were combined as a weighted average to obtain pooled estimates. GameRank was used to determine variable importance, and the included predictors were compared based on the AUC of the receiver operating characteristic (ROC). After model selection, the contribution of each predictor was assessed, and predictor availability in the clinical context was also considered for predictor selection.

[0100] The model was specified to receive as inputs: albumin level, bone marrow lymphocyte percentage, Charlson Comorbidity Index (CCI), creatinine clearance, and history of cardiac, vascular, or diabetic complications. "Event" refers to adverse events, not tolerating R-CHOP treatment. "All events" refers to all events, including situations that tolerate R-CHOP treatment. "Odds ratio" refers to the increased probability of having an adverse event as a result of a particular classification of a variable. For example, with a history of cardiovascular / diabetic conditions, the probability of having an adverse event versus no adverse event in the presence of that history is: JPEG0007823020000001.jpg13170. In the absence of such a medical history, the probability of an adverse event versus no adverse event is: JPEG0007823020000002.jpg13170. The odds ratio is 0.252 / 0.125=2.02, meaning that having a history of cardiovascular / diabetes increases the likelihood of adverse events by 2.02 times. The 95% confidence limits are shown in parentheses in the columns. LLN in Figure 7 represents the lower limit of normal. The Tolerability of R-CHOP in Aggressive Non-Hodgkin's Lymphoma (TRAIL) model was trained on 742 subjects with full data. Figure 7 shows higher albumin, creatinine clearance, CCI, and more total events with no history of cardiovascular / diabetes.

[0101] A receiver operating characteristic area under the curve (AUC) of 75.5% was achieved (sensitivity = 75%, specificity = 69.6%), as was a cross-validation (CV) AUC of 74.9% (Figure 8). The CIs in Figure 8 represent confidence intervals. To create a simpler model using information available in a routine clinical setting, bone marrow lymphocyte levels were removed from the input variable set, and the model was trained on 1,022 subjects with full data. A CV AUC of 70.1% (sensitivity = 68.1%, specificity = 67.2%) was achieved. Other top-ranked models using full data for all subjects achieved AUCs > 70% and included age, hemoglobin, T cell count, TH cell percentage, and lactate dehydrogenase (LDH). Implementation of CCI alone performed poorly, with a CV AUC of 62.9%.

[0102] An example of a resulting TRAIL model is of the form of formula (1): HYBRID.OUTCOME~I(NOIMP.CRCL^-1)+log(NOIMP.ALBUM)+CHARLSON+JUN20.HVD+sqrt(NOIMP.BMLYMP)+log(NOIMP.CRCLBSA) (1)

[0103] Equation (1) does not list coefficients before each term. Coefficients can be determined through logistic regression or other appropriate algorithms. In other words, terms such as CHARLSON are augmented by the coefficients in the model. The operator I is the inverse. NOIMP refers to no imputation (i.e., estimation of missing values). JUN20 refers to the date the data was acquired. The model need not be limited to only one date for acquiring data.

[0104] The output of the model (e.g., HYBRID.OUTCOME) may be a score. The score may be compared to a threshold. The threshold may be determined from one or more reference subjects known to tolerate or not tolerate treatment (e.g., R-CHOP) for DLBCL or other lymphoma. A score above the threshold may indicate that the subject will not tolerate the treatment or is predicted to experience an adverse event as a result of the treatment with greater than a certain probability (e.g., 30%, 40%, 50%, 60%, 70%, 80%, or 90%).

[0105] The model shown in Equation (1) and represented by the table in FIG. 1 is not the only model that can be used to determine susceptibility to treatment for DLBCL or other lymphomas. FIG. 9 lists other models of equations that do not have coefficients that can determine susceptibility to treatment and the AUC of each model. The terms used in FIG. 9 are listed in Table 5. The label auc.full_data refers to the AUC calculated and fitted to the full training data. Any of these equations in FIG. 9 may be functions used with the machine learning model of block 120. The model is not limited to the four-variable model example in Equation (1).

[0106] Missing values ​​for certain features in a sample can be imputed rather than submitted to model training. As a result, the sample size N can be larger than shown in Figure 9. For example, imputation was performed on the data, increasing N to 1051.

[0107] A model may use 1 to 5, 6 to 10, 10 to 20, 20 to 50, 50 to 100, 100 to 200, or more than 200 variables. The variables (i.e., features) may be any variables described herein, including any figures. The variables determined by training a model may be from a set of variables including 6 to 10, 10 to 20, 20 to 50, 50 to 100, 100 to 200, 200 to 500, 500 to 1,000, or more than 1,000 variables. The variables in the set of variables may be any variables described herein.

[0108] The top variables for various team sizes (i.e., a selection of variables for size m) are listed in Figures 10-14. The model can use any of these variables in any combination (or other combinations of variables). Figures 11 and 13 show models with the same team size and 50th rating ("r"), but with different variables as a result of the randomized nature of the ranking algorithm.

[0109] The four-variable TRAIL example model was evaluated on the GOYA holdout dataset, which was associated with subjects not in the training dataset.

[0110] IV.B. Clinical Validation Using the MAIN Dataset The four-variable TRAIL model example was also evaluated on the MAIN dataset. Data for external validation were obtained from the multicenter, randomized, double-blind, placebo-controlled, phase 3 MAIN study (NCT00486759). In total, 186 centers in 30 countries participated in MAIN, recruiting previously untreated DLBCL subjects aged 18 years or older between July 2007 and May 2010. The trial was designed to compare R-CHOP with or without the addition of bevacizumab. Investigators participating in MAIN at the study site level pre-determined whether a given subject would proceed with 14- or 21-day cycles of CHOP chemotherapy. Treatment with R-CHOP plus additional bevacizumab was blinded by placebo infusion. Further details of the study treatments were previously published in Seymour et al., "R-CHOP with or without bevacizumab in patients with previously untreated diffuse large B-cell lymphoma: final MAIN study outcomes." Haematologica, 2014;99(8):1343-1349, the contents of which are incorporated herein by reference for all purposes. The study was terminated early at the sponsor's discretion due to increased cardiac toxicity and prolonged progression-free survival (PFS) in the bevacizumab arm. The median follow-up periods for the R-CHOP and R-CHOP plus bevacizumab arms were 23.7 and 23.6 months, respectively. Due to safety signals observed in the bevacizumab arm and the known increased toxicity associated with 14-day cycles of CHOP (CHOP-14), only data from subjects treated with 21-day cycles of R-CHOP in MAIN (R-CHOP-21) were used for model validation.

[0111] IV.C. Various model validation results A set of models using specific variables was validated in two independent clinical trials. A composite binary endpoint was selected for analyzing the clinical trials. The endpoint was configured to qualify any of the following three events: (1) an adverse event (AE) leading to discontinuation of cyclophosphamide and / or doxorubicin; (2) a mean relative dose intensity of cyclophosphamide and doxorubicin less than 80% across six cycles, unrelated to progressive disease (PD); or (3) death within the first six cycles, unrelated to progressive disease (PD). The two independent clinical data sets used to validate the models are labeled GOYA holdout (or similar) and MAIN R-CHOP-21 (or similar). The "holdout" in GOYA holdout refers to data not used to train the model.

[0112] Table 1 shows the results of various models on the GOYA training and GOYA holdout datasets. The models are listed in the first column. The number of subjects, n, is in the second column. The cross-validation area under the curve (CV AUC) is shown in the third column. The training data AUC is in the fourth column. The holdout data AUC is in the fifth column. The training AUC after fitting to the full data is in the sixth column.

[0113] The rows indicate the three models used. The first model included the terms in Equation (1): creatinine clearance (CRCL); albumin (ALBUM); Charlson Comorbidity Index (CHARLSON); cardiac, vascular, and diabetic complications (HVD); bone marrow lymphocytes (BMLYMP); and BSA-corrected creatinine clearance (CRCLBSA). The second row is a model that excluded bone marrow lymphocytes. The third model excluded bone marrow lymphocytes and BSA-corrected creatinine clearance. The third model had the largest AUC for the GOYA holdout data. JPEG0007823020000003.jpg94170

[0114] The MAIN R-CHOP-21 data came from a study of bevacizumab (Avastin) in combination with rituximab (MabThera) and CHOP (cyclophosphamide, hydroxydaunorubicin [doxorubicin], Oncovin [vincristine], prednisone) chemotherapy in subjects with diffuse large B-cell lymphoma. This was a two-arm study designed to compare the efficacy and safety of bevacizumab (Avastin) in combination with rituximab (MabThera) and CHOP (cyclophosphamide, hydroxydaunorubicin [doxorubicin], Oncovin [vincristine], prednisone) chemotherapy (R-CHOP) with rituximab and CHOP chemotherapy (R-CHOP) in previously untreated subjects with CD20-positive diffuse large B-cell lymphoma (DLBCL). Subjects were randomized to receive eight cycles of R-CHOP plus bevacizumab or R-CHOP plus placebo. Treatment with bevacizumab / placebo and R-CHOP was given on either a 2-weekly or 3-weekly schedule, with bevacizumab given weekly at an average dose of 5 mg / kg (10 mg / kg for 2-week cycles and 15 mg / kg for 3-week cycles).

[0115] Table 2 shows the results of various models on the MAIN R-CHOP-21 dataset. The first column lists the model. The second column lists the AUC from the model fitted to the GOYA training data. The third column lists the AUC when the model was fitted to the full GOYA data (training and holdout). The rows indicate the three models used. The models included the same terms as those listed in Table 1. The MAIN R-CHOP-21 dataset did not have measurements of bone marrow lymphocytes, so the first model does not have an AUC. The model that excluded bone marrow lymphocytes and BSA-corrected creatinine clearance had the highest AUC on either training dataset. JPEG0007823020000004.jpg95170

[0116] IV.D. Results from the four-variable TRAIL model Figure 15 shows the prediction of a composite endpoint using a model with four variables (creatinine clearance (CRCL); albumin (ALBUM); Charlson Comorbidity Index (CHARLSON); and cardiac, vascular, and diabetic complications (HVD)). Figure 15 uses the GOYA holdout dataset. Overall survival is on the y-axis, and time in months is on the x-axis. The upper line (line 1510) indicates predicted event-free, and the lower line (line 1520) indicates predicted event-presence. The number of crossovers between predicted event-presence and predicted event-absence are listed in tabular form below the graph. Figure 15 also shows the hazard ratio (HR) indicating the predicted tolerable event presence group versus the predicted tolerable event-free group. Hazard ratios adjusted for baseline International Prognostic Index (IPI) > 2 are presented. The cutoff for determining the prediction of tolerable events was determined by finding the value that maximizes the Youden Index using the GOYA training data. FIG. 15 shows that the subjects with the predicted event had a lower overall survival than subjects without the predicted event.

[0117] Figure 16 shows the prediction of a composite endpoint using an example TRAIL model with four variables. Figure 16 presents the same type of results as Figure 15, but for the MAIN R-CHOP-21 dataset. Line 1610 indicates no predicted event. Line 1620 indicates predicted event. Figure 16 shows that subjects with a predicted event had lower overall survival than subjects without a predicted event.

[0118] Figure 17 shows the ROC curve for an example TRAIL model with four variables using the GOYA training, GOYA holdout, and MAIN external validation data. The model achieved an AUC (95% CI) of 0.700 (0.652 to 0.749) on the GOYA training data, 0.722 (0.647 to 0.797) on the GOYA holdout data, and 0.691 (0.604 to 0.778) on the MAIN external validation data.

[0119] Figure 18A shows the ROC curve for an example TRAIL model with four variables compared to defined prognostic factors, age, IPI, and Eastern Cooperative Oncology Group performance status (ECOG PS) using the GOYA holdout data. Figure 18B shows the ROC curve for an example TRAIL model compared to defined prognostic factors, age, IPI, and ECOG PS using the MAIN external validation. Lines 1810 and 1820 represent the example TRAIL model. Lines 1820 and 1860 represent age. Lines 1830 and 1870 represent IPI. Lines 1830 and 1870 represent ECOG PS. The AUC was consistently higher for the TRAIL model compared to clinical risk factors predicted to be associated with tolerability in both GOYA and MAIN.

[0120] In the GOYA holdout dataset, 74.1% of subjects reported treatment-emergent grade 3-5 AEs in the high-risk category, compared with 70.1% in the intermediate-risk and 65.6% in the low-risk categories. Corresponding rates in MAIN were 74.0%, 57.5%, and 48.4%, respectively. Higher proportions of subjects in the intermediate- and high-risk TRAIL categories experienced grade 4 and 5 AEs compared with the low-risk group.

[0121] IV.E. Sensitivity or Specificity-Based Cutoffs Figures 19A and 19B show the results of an example TRAIL model with four variables combined with various IPI levels. Figure 19A shows predicted survival for the GOYA holdout data, and Figure 19B shows predicted survival for the MAIN R-CHOP-21 data. The graphs present similar information as Figures 15 and 16. However, instead of two lines, four lines are plotted. The red lines (lines 1910 and 1950) indicate low IPI (IPI<2) without predicted tolerability events. The green lines (lines 1920 and 1960) indicate high IPI (IPI>2) without predicted tolerability events. The blue lines (lines 1930 and 1970) indicate low IPI with predicted tolerability events. The purple lines (lines 1940 and 1980) indicate high IPI with predicted tolerability events.

[0122] The cutoff can be used to improve prediction of whether a particular subject will tolerate a particular treatment. A value above (e.g., above or below) the cutoff can indicate that the subject is predicted to tolerate a particular treatment.

[0123] Table 3 shows the performance of using cutoffs in a model with four variables (creatinine clearance (CRCL); albumin (ALBUM); Charlson Comorbidity Index (CHARLSON); and cardiac, vascular, and diabetic complications (HVD)). The GOYA dataset was used. The first column lists the cutoff criteria. The second and third columns are training scores. The second column lists the cutoffs used. The third column lists the performance, expressed in terms of sensitivity and specificity. The fourth and fifth columns are CV scores. The fourth column lists the cutoffs used. The fifth column lists the performance, expressed in terms of sensitivity and specificity.

[0124] The data in the first row in Table 3 is for the criterion that maximizes the sum of sensitivity and specificity. The data in the second row is for the criterion that sets sensitivity to 0.80. The data in the third row is for the criterion that sets specificity to 0.80. The results show that the cutoff can be adjusted based on different criteria. Figure 20 shows the area under the curve graph using the cutoffs in Table 3. The two lines represent the training score (line 2010) and the CV score (line 2020). The labeled points are the sensitivity and specificity shown in Table 3. Table 4 shows the performance of using the cutoffs in a model using four variables: creatinine clearance (CRCL); albumin (ALBUM); Charlson Comorbidity Index (CHARLSON); and cardiac, vascular, and diabetic complications (HVD)). The MAIN R-CHOP dataset was used. The first column lists the cutoff criteria. The second and third columns are the CV scores (same as Table 3). The fourth and fifth columns are for the MAIN R-CHOP-21 dataset. The same criteria as in Table 4 are used. JPEG0007823020000006.jpg134170

[0125] IV.F. Risk Percentile-Based Cutoffs Figure 21 shows the percentage of subjects experiencing intolerable events by risk category. Cutoffs for low, medium, and high risk categories were defined based on predicted probability quartiles (low risk [0, 0.07], medium risk [0.07, 0.16], and high risk [0.16, 1]). In the GOYA holdout dataset, the proportions of subjects experiencing intolerable events were 3.3% in the low-risk group, 12.4% in the medium-risk group, and 32.9% in the high-risk group. The corresponding proportions in MAIN were 9.7%, 9.7%, and 34.2%, respectively. Figure 21 shows that a significant proportion of subjects with higher predicted probabilities suffered from adverse events at a significantly higher rate than subjects with lower predicted probabilities. The cutoffs in Figure 21 can be used for the cutoff values ​​described in process 100.

[0126] Figure 22A shows a Kaplan-Meier overall survival graph stratified by risk category in the GOYA holdout dataset. Figure 22B shows a Kaplan-Meier overall survival graph stratified by risk category in the MAIN external validation dataset. Lines 2210 and 2240 indicate low risk. Lines 2220 and 2250 indicate intermediate risk. Lines 2230 and 2260 indicate high risk. Similar to Figure 21, Figures 22A and 22B show that subjects classified as high risk by predicted probability tend to have lower overall survival. The steep decline in overall survival over time for intermediate risk (e.g., 72 months in Figure 22A and 48 months in Figure 22B) is an artifact resulting from the small sample size.

[0127] Four-year overall survival in the GOYA holdout dataset was 90.7% in the low-risk group, 80.2% in the intermediate-risk group, and 69.2% in the high-risk group.

[0128] In the GOYA holdout dataset, 74.1% of subjects reported treatment-emergent grade 3-5 AEs in the high-risk category, compared with 70.1% in the intermediate-risk and 65.6% in the low-risk categories. In the MAIN external validation dataset, the corresponding rates in MAIN were 74.0%, 57.5%, and 48.4%, respectively. A higher proportion of subjects in the intermediate- and high-risk TRAIL categories experienced grade 4 and 5 AEs compared with the low-risk group.

[0129] IV.G. Real-world data Figure 23 shows a Kaplan-Meier overall survival graph for the real-world dataset, Flatiron RWD. The Flatiron RWD dataset was not a clinical trial. Similar to Figures 21, 22A, and 22B, these lines are divided by scores from the example TRAIL model with four variables. Line 2310 indicates subjects with scores corresponding to the lowest probability of adverse events (highest predicted tolerability). Line 2320 indicates subjects with scores corresponding to a moderate probability of adverse events (moderate predicted tolerability). Line 2330 indicates subjects with scores corresponding to the highest probability of adverse events (lowest predicted tolerability). Line 2330 indicates subjects with lower predicted tolerability have the lowest overall survival. The steep decline in overall survival over time (e.g., 96 months) is an artifact of the small sample size. Figure 23 shows that the example TRAIL model with four variables is applicable to real-world subjects outside of clinical studies. It is expected that other models developed using the TRAIL algorithm will also be applicable to real-world subjects.

[0130] VI. Terminology Table 5 contains terms used in the figures and elsewhere in this specification, any of which may be variables in the machine learning models described in Figures 1 and 2. JPEG0007823020000007.jpg255170JPEG0007823020000008.jpg255170JPEG0007823020000009.jpg12170

[0131] VI. Additional Considerations Some embodiments of the present disclosure include a system comprising one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium containing instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0132] The terms and expressions which have been employed are used as terms of description and not of limitation, and the use of such terms and expressions is not intended to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0133] The description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the description of preferred exemplary embodiments provides those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.

[0134] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

[0135] All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes. None is admitted to be prior art.

Claims

1. accessing an input dataset comprising a plurality of input data values ​​associated with a particular subject having lymphoma, each input data value corresponding to one variable of a set of variables, the plurality of input data values ​​comprising index values ​​characterizing a comorbidity of the particular subject; inputting the input dataset into a machine learning model to generate a score corresponding to the degree to which a particular subject will tolerate a particular treatment, where tolerating a particular treatment comprises not ending up below a threshold dose of the particular treatment or not being reduced below a threshold dose within a period of time after initiating the particular treatment, and wherein the machine learning model: a set of parameters determined using a plurality of training data elements, each of the plurality of training data elements corresponding to a training subject, each of the plurality of training data elements including a training input data set and a label, the label indicating the training subject's tolerance to a particular treatment; and A function that associates the received input dataset and parameters with a score and and using the generated score to generate a prediction of the particular subject's tolerance to a particular treatment.

2. 10. The method of claim 1, wherein the set of variables comprises results from a blood panel.

3. The method of claim 1 or 2, wherein the set of variables comprises a characterization of a particular subject's medical history.

4. sending a request to a computer system that stores the medical history of a particular subject; receiving a medical history of a particular subject; determining an index value characterizing a comorbidity of a particular subject, the index value being an input data value of the plurality of data values; The method of claim 3 further comprising:

5. The method of claim 1 , wherein the set of variables comprises an outcome of an invasive diagnosis.

6. 6. The method of any one of claims 1 to 5, wherein the set of variables comprises albumin concentration, creatinine clearance, a comorbidity index, or the presence of a history of cardiovascular or diabetic disease.

7. 7. The method of claim 6, wherein the set of variables further comprises bone marrow lymphocyte levels.

8. 8. The method of claim 1, wherein the set of variables includes no more than 5, no more than 10, or no more than 15 variables.

9. 9. The method of any one of claims 1 to 8, wherein the set of variables comprises hemoglobin level, red blood cell count, hematocrit level, number of concomitant medications, chloride level, total levels of CD3 and CD4 protein complexes and T-cell coreceptor, lymphocyte levels, or levels of CD3 protein complex and T-cell coreceptor.

10. 10. The method of any one of claims 1 to 9, wherein the lymphoma is diffuse large B-cell lymphoma and the specific treatment comprises the administration of rituximab, cyclophosphamide, doxorubicin hydrochloride, vincristine, and prednisolone.

11. 11. The method of any one of claims 1 to 10, wherein the lymphoma is diffuse large B-cell lymphoma and the specific treatment comprises cyclophosphamide or doxorubicin hydrochloride.

12. 12. The method of any one of claims 1 to 11, further comprising processing a blood sample from a particular subject using a blood panel to determine one or more input data values ​​of the plurality of input data values.

13. 13. The method of any one of claims 1 to 12, wherein the duration is 18 weeks or less.

14. 14. The method of any one of claims 1 to 13, wherein the prediction is that a particular subject is unlikely to tolerate a particular treatment.

15. 14. The method of any one of claims 1 to 13, further comprising comparing the score to a cutoff value, wherein the cutoff value is determined from a plurality of reference subjects, each reference subject of the plurality of reference subjects having a respective predicted score.

16. 16. The method of claim 15, wherein the cutoff value is determined to obtain a predetermined level of sensitivity or a predetermined level of specificity for predicting the tolerability of a particular treatment in a plurality of reference subjects.

17. identifying a subset of a plurality of reference objects, wherein each score of each reference object in the subset is lower than each score of each reference object of the plurality of reference objects not in the subset, and the size of the subset corresponds to a predetermined percentage of the size of the plurality of reference objects; defining a cutoff value as a given value that separates each score of a reference subject in the subset from scores of reference subjects of the plurality of reference subjects that are not in the subset; 16. The method of claim 15, wherein comparing the score to a cutoff value comprises determining that the score exceeds the cutoff value, and the prediction is that the particular subject is unlikely to tolerate a particular treatment.

18. 18. The method of claim 14 or 17, further comprising, in response to returning a prediction that the particular subject is unlikely to tolerate the particular treatment, outputting a recommendation to enroll the particular subject in a clinical study including a treatment different from the particular treatment.

19. 18. The method of claim 14 or 17, further comprising, in response to returning a prediction that the particular subject is unlikely to tolerate the particular treatment, outputting a recommendation to treat the particular subject's lymphoma with another treatment different from the particular treatment.

20. 18. The method of claim 14 or 17, further comprising returning a prediction that the particular subject is unlikely to tolerate the particular treatment, and outputting a recommendation that the particular treatment not be administered to the particular subject.

21. 21. The method of any one of claims 1 to 20, wherein the score is the probability that a particular subject will tolerate a particular treatment.

22. accessing a training dataset including a plurality of training data elements, each of the plurality of training data elements corresponding to a training subject having lymphoma, each of the plurality of training data elements including a training input dataset and a label, the label indicating the training subject's tolerance to a particular treatment, wherein the training subject tolerates the particular treatment if the particular treatment does not end up below a threshold dose or is not reduced below a threshold dose within a certain period of time after initiating the particular treatment; training a machine learning model using the training dataset to generate a score corresponding to the degree to which a particular subject will tolerate a particular treatment, wherein training the machine learning model comprises: Learning a set of parameters; and determining a function that associates said set of parameters with a score; and Including, the score corresponds to the degree to which the particular subject tolerates the particular treatment, each input data value corresponds to one variable of a set of variables, and a plurality of input data values ​​include index values ​​characterizing comorbidities of the particular subject; Computer-implemented methods.

23. one or more data processors; a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform the method of any one of claims 1 to 22.

24. 23. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform the method of any one of claims 1 to 22.

Citation Information

Patent Citations

  • Method and apparatus for determining patterning process parameters

    JP2019508745A

  • Methods for reducing side effects of anti-CD30 antibody-drug conjugate therapy

    JP2020536916A

  • Use of Torque Teno Virus (TTV) as a Marker to Measure the Proliferative Potential of T Lymphocytes

    JP2023547961A