Machine learning prediction of treatment response
A machine learning model using biological signatures predicts therapy response, addressing therapy resistance by personalizing treatment strategies and improving outcomes in diseases like cancer.
Patent Information
- Application Number
- JP2022547780
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-08
- Filing Date
- 2021-02-07
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2041-02-07
AI Technical Summary
Existing therapies for diseases such as oncology face challenges due to therapy resistance, with the tumor microenvironment contributing to pro-tumorigenic and pro-metastatic processes that interfere with therapeutic effects, necessitating the identification of biomarkers to predict response to therapy.
A machine learning model trained on biological signatures from multiple time points to predict patient response to therapy, utilizing profiles like DNA, RNA, protein, and metabolomic data to identify key factors for improving treatment outcomes.
The model effectively predicts patient response to therapy, enabling personalized treatment decisions and identifying suitable alternative therapies, thereby enhancing treatment efficacy.
Smart Images

Figure 0007757291000003 
Figure 0007757291000004 
Figure 0007757291000005
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority under 35 U.S.C. §119(e) to U.S. Provisional Application Nos. 62 / 971,065, filed February 6, 2020, 63 / 022,736, filed May 11, 2020, and 63 / 089,304, filed October 8, 2020. The contents of all of the above applications are incorporated by reference as if fully set forth in their entireties herein.
[0002] The present invention relates to the field of machine learning. [Background technology]
[0003] One of the major complications in various diseases, including but not limited to oncology, is resistance to therapy. Many studies have focused on the involvement of tumor cell mutations and epigenetic changes in conferring drug resistance. However, in recent years, research has shown that the tumor microenvironment contributes to therapy resistance, and that in response to almost all types of anti-cancer therapy, the patient (i.e., the host) can generate pro-tumorigenic and pro-metastatic processes that can interfere with the therapeutic effect.
[0004] The host response to cancer therapy is a relatively newly described phenomenon that has resulted in a paradigm shift in understanding cancer progression and resistance to therapy, and the present invention suggests its use for early identification of non-responding patients and as a discovery tool for targets for medical intervention (e.g., selective inhibitors of key factors that can be used in combination with standard therapy to improve treatment outcomes in non-responding patients).
[0005] Therefore, there is a significant need to identify biomarkers that can predict response to therapy.
[0006] The foregoing examples of the related art and limitations associated therewith are intended to be illustrative and not exhaustive. Other limitations of the related art will become apparent to those skilled in the art upon reading this specification and studying the drawings. Summary of the Invention
[0007] The following embodiments and aspects thereof are described and illustrated in conjunction with systems, tools, and methods that are meant to be exemplary and illustrative, not limiting in scope.
[0008] In an embodiment, a system is provided comprising at least one hardware processor and a non-transitory computer-readable storage medium having program instructions stored therein, the program instructions being executable by the at least one hardware processor to receive, for each of a plurality of subjects having a specified type of disease and receiving a specified therapy to treat the disease, (a) a first biological signature associated with a biological sample collected at a first time point related to the specified therapy, and (b) a second biological signature associated with a biological sample collected at a second time point related to the specified therapy; calculate, for each of the plurality of subjects, a set of values representing a relationship between the first biological signature and the second biological signature associated with each subject; and, in a training phase, train a machine learning model with a training set including (i) the set of calculated values and (ii) a label associated with an outcome of the specified therapy in each of the subjects to generate a classifier suitable for predicting a response in a target patient to the specified therapy.
[0009] In an embodiment, there is also provided a method for predicting a response in a target patient to a specified therapy, the method including: receiving, for each of a plurality of subjects having a specified type of disease and receiving a specified therapy to treat the disease, (a) a first biological signature associated with a biological sample collected at a first time point related to the specified therapy, and (b) a second biological signature associated with a biological sample collected at a second time point related to the specified therapy; calculating, for each of the plurality of subjects, a set of values representing a relationship between the first biological signature and the second biological signature associated with the respective subject; and, in a training phase, training a machine learning model with a training set including (i) the set of calculated values and (ii) a label associated with an outcome of the specified therapy in each of the subjects, thereby generating a classifier suitable for predicting a response in the target patient to the specified therapy.
[0010] In an embodiment, there is further provided a computer program product including a non-transitory computer-readable storage medium having program instructions embodied therein, the program instructions being executable by at least one hardware processor to: receive, for each of a plurality of subjects having a specified type of disease and receiving a specified therapy to treat the disease, (a) a first biological signature associated with a biological sample collected at a first time point associated with the specified therapy, and (b) a second biological signature associated with a biological sample collected at a second time point associated with the specified therapy; calculate, for each of the plurality of subjects, a set of values representing a relationship between the first biological signature and the second biological signature associated with each subject; and, during a training phase, train a machine learning model with a training set including (i) the set of calculated values and (ii) a label associated with an outcome of the specified therapy in each of the subjects to generate a classifier suitable for predicting a response in the target patient to the specified therapy.
[0011] In some embodiments, the first biological signature and the second biological signature are each one of a DNA profile, an RNA profile, a protein profile, a metabolomic profile, a microbiome profile, a transcriptomic profile, a genomics profile, an epigenomics profile, a cellular profile, a post-translational modification-based profile, a single-cell-based analysis, and a regulatory RNA profile.
[0012] In some embodiments, the first biological signature and the second biological signature are each protein expression profiles, and the sets of values each include, for each protein in the protein expression profiles, a relationship between the expression levels of proteins in the first biological signature and the second biological signature.
[0013] In some embodiments, the protein expression profile comprises expression values of at least two proteins.
[0014] In some embodiments, the method further includes executing, and the program instructions are further executable to perform a dimensionality reduction step on the set of values to reduce the number of variables in at least one of the sets of values.
[0015] In some embodiments, the dimensionality reduction step identifies a subset of proteins that are dominant in each set of values, while in other embodiments, the dimensionality reduction generates new features that can predict response.
[0016] In some embodiments, dimensionality reduction involves considering all or some of the feature values as vector components and calculating their norms.
[0017] In some embodiments, the training set includes only a subset of the major proteins in each of the sets of values.
[0018] In some embodiments, the set of values is labeled with a label.
[0019] In some embodiments, each of the biological samples is one of plasma, whole blood, serum, cerebrospinal fluid (CSF), and peripheral blood mononuclear cells (PBMCs).
[0020] In some embodiments, the specified type of disease is a specified type of cancer, hi some embodiments, the cancer is selected from melanoma, non-small cell lung cancer (NSCLC), small cell lung cancer (SCLC), head and neck cancer, and genitourinary cancer.
[0021] In some embodiments, the training set further includes labels associated with the clinical data for at least some of the subjects.
[0022] In some embodiments, the prediction is represented as one of a set of binary values, continuous values, and discrete values.
[0023] In some embodiments, the prediction includes an indication of a secondary effect in the target subject.
[0024] In some embodiments, the method further comprises, in an inference stage, applying the classifier to the target set of values associated with a target subject, thereby predicting a response in the target subject to the specified treatment.
[0025] In some embodiments, the method further includes determining, and the program instructions are further executable to determine, based at least in part on the prediction, at least one of: continuing the specified therapy in the target subject; adjusting the specified therapy in the target subject; discontinuing the specified therapy in the target subject; and administering a different therapy to the target subject.
[0026] In some embodiments, the designated therapy is immunotherapy. In some embodiments, the designated therapy is a combination of immunotherapy and chemotherapy. In some embodiments, the designated therapy is a combination of immunotherapy and targeted therapy. In some embodiments, the designated therapy is a combination of two or more types of immunotherapy. In some embodiments, the immunotherapy is selected from anti-PD-1 / PD-L1 therapy, anti-CTLA-4 therapy, and both.
[0027] In some embodiments of the systems, computer program products, and methods provided herein, adjusting a designated therapy or administering a different therapy to the target subject is determined by a method including: (i) determining differentially expressed proteins (DEPs) between responders and non-responders; (ii) determining one or more resistance-associated proteins (RAPs) selected from the determined DEPs in a sample obtained from the subject; and (iii) selecting a therapy suitable for balancing the level of the one or more RAPs in the subject.
[0028] In some embodiments, determining one or more RAPs is by providing a probabilistic measure of the distance of DEP expression levels from a defined group of samples.
[0029] In some embodiments, determining one or more RAPs in a subject is by determining the expression distribution of each DEP in each of a responder group and a non-responder group, fitting a probability density function for each group, and calculating for each subject the probability that the DEP is associated with one of the response groups based on the DEP expression of the subject.In certain embodiments, determining one or more RAPs in a subject is by determining the probability that each DEP is associated with the responder distribution.In other embodiments, determining one or more RAPs in a subject is by determining the probability that each DEP is associated with the non-responder distribution.
[0030] In some embodiments, the therapy for balancing the level of one or more RAPs in the subject is selected from a list of approved or investigational drugs.
[0031] In addition to the exemplary aspects and embodiments described above, further aspects and embodiments will become apparent by reference to the figures and by study of the following detailed descriptions. [Brief explanation of the drawings]
[0032] [Figure 1] 1 is a flowchart of functional steps in a method for training a machine learning model to predict a patient's response to a therapy, according to some embodiments of the present disclosure. [Figure 2] 2 is a schematic diagram of the process steps of FIG. 1 according to some embodiments of the present disclosure. [Figure 3] FIG. 1 is a non-limiting schematic diagram of a quality control process based on a limit of detection (LOD) threshold, according to some embodiments of the present disclosure. [Figure 4] Illustrates experimental results according to several embodiments of the present disclosure (TP - true positive; FN - false negative; TN - true negative; FP - false positive, PPV - positive predictive value, NPV - negative predictive value). [Figures 5A-5D] Illustrates experimental results according to several embodiments of the present disclosure (TP - true positive; FN - false negative; TN - true negative; FP - false positive, PPV - positive predictive value, NPV - negative predictive value). [Figures 6A-6C] Illustrates experimental results according to several embodiments of the present disclosure (TP - true positive; FN - false negative; TN - true negative; FP - false positive, PPV - positive predictive value, NPV - negative predictive value). [Figure 7A-7C] Illustrates experimental results according to several embodiments of the present disclosure (TP - true positive; FN - false negative; TN - true negative; FP - false positive, PPV - positive predictive value, NPV - negative predictive value). [Figure 8] 1 is a flowchart of three filters for personalized analysis of potential targets for intervention, according to some embodiments of the present disclosure. Solid and dashed lines indicate positive and negative responses to survey questions, respectively. On the left, the analysis / data processing steps are shown, followed by the applied filters. The clinical filter appears three times in the flowchart. F1 specifies a cohort-based statistical filter, F2 specifies a personalized filter, and F3 specifies a clinical filter. [Figure 9]1 is a non-limiting example of RAP score calculation. The example in this figure shows the protein distribution of an exemplary protein ("Protein A") with R (light blue) and NR (orange) in the entire cohort (n=52). A patient with a Protein A expression level of 0.3, marked with a dashed line, has a P(NR) / P(R) ratio of 8, as calculated based on the area above 0.3 in the NR and R distributions (areas marked with solid colors). Following log2 transformation, the RAP score for this protein in this particular patient is 3. [Figures 10A-10B] Non-limiting examples of RAP score directionality are shown below. The choice of area for RAP score calculation depends on the relative positions of the R and NR distributions. A. If the median of the NR distribution for a given differentially expressed protein (DEP) is higher than the median of the R distribution, the area for Equation 1 is calculated based on the right tail. B. If the median of the R distribution for a given DEP is higher than the median of the NR distribution, the area for Equation 1 is calculated based on the left tail. [Figures 11A-11C] This shows that the distribution of RAP scores can depend on the difference between R and NR distributions. RAP scores are shown above the plots. A. R and NR distribution of Protein A expression levels or T1 / T0. B. R and NR distribution of Protein B expression levels or T1 / T0. C. Distribution of Protein B RAP scores among NR patients. [Figure 11D] We show that the RAP score can further be used to identify groups of non-responders who share similar RAP profiles. [Figures 12A-12B] 12A-12B show simulations of RAP perturbations. (FIG. 12A) A predictive signature based on the list of RAPs for the entire cohort is generated. (FIG. 12B) For a given patient, a specific RAP (or multiple RAPs) is perturbed. The baseline response probability is then compared to the perturbed response probability. [Figures 13A-13B]FIG. 13A shows the training (FIG. 13A) and validation (FIG. 13B) of a classifier based on the presented invention for predicting response to treatment in psoriasis patients. FIG. 13A: (Left) SVM yielded an AUC of 0.77. (Right) Accuracy = 0.7286, Sensitivity = 0.75, Specificity = 0.6818, PPV = 0.8372, NPV = 0.5556. FIG. 13B: (Left) SVM yielded an AUC of 0.751. (Right) Accuracy = 0.6714, Sensitivity = 0.6458, Specificity = 0.7273, PPV = 0.8378, NPV = 0.4848. TP - true positive; FN - false negative; TN - true negative; FP - false positive; PPV - positive predictive value; NPV - negative predictive value. [Figure 14] Figure 1 shows differential network analysis. Networks of correlation data were constructed separately for each group. From these group-specific maps, differential maps can be generated to identify differentially correlated proteins in each group. [Figure 15] 1 shows an example of a differential network between responders and non-responders based on an NSCLC dataset. [Figures 16A-16B] Predictions based on protein covariation are shown. As a non-limiting example, two proteins showing differential correlation fold-change values between responders and non-responders were examined. (16A) The correlation between responders and non-responders was positive (R = 0.37). The dashed line indicates a linear fit to the responder values. (16B) The residuals of the two protein pairs (i.e., the distance of each point from the linear fit in A) were calculated and used as input for an SVM classifier. The resulting predictor achieved an ROC AUC of 0.77. [Figure 17]Preliminary results of response prediction using naive predictors in n=67 NSCLC patients. Response quality is quantified by the area under the receiver operating curve (ROC) curve (AUC) for a training set consisting of n=37 responders and an independent validation set including n=15 responders and n=15 nonresponders. AUC was calculated for 1000 different sets of training and validation sets, where the n=37 responder training set was randomly sampled from a total of n=52 responders in the dataset. The resulting 1000 AUC values are shown in a histogram, with the median indicated by a solid gray vertical line. The mean value of the random classifier AUC=½ is indicated by a dashed vertical line. For comparison, the AUC distribution of the random classifier for n=15 responders and n=15 nonresponders is shown in white shading. DETAILED DESCRIPTION OF THE INVENTION
[0033] Disclosed are systems, methods, and computer program products that provide machine learning models configured to predict a patient's response to a treatment. Also disclosed are systems, methods, and computer program products that suggest suitable alternative or concomitant therapies for improving treatment outcomes in a patient.
[0034] In some embodiments, the present disclosure provides for training a machine learning model using a training dataset that includes biological profiles (e.g., protein expression profiles) of biological samples or biological signatures obtained from a plurality of subjects, e.g., a cohort or predetermined population, having a specified type of disease and receiving a specified type of treatment (e.g., a treatment associated with the specified type of disease).
[0035] In certain embodiments, the cohort or predetermined population of subjects is based on or determined according to any one of disease type, disease stage, disease treatment, treatment history, clinical profile, and any combination thereof.
[0036] In some embodiments, the trained machine learning model of the present disclosure may provide a prediction of response of target patients diagnosed with a specified disease to an associated specified treatment or therapy. In some embodiments, the machine learning model of the present disclosure may be trained on data from a cohort of subjects with a specified disease or type of disease, or a predetermined population, where biological samples are obtained from cohort participants at at least one time point relative to the treatment, e.g., T0 (e.g., before treatment) or T1 (e.g., during, throughout, or after treatment).
[0037] In some embodiments, the present disclosure further provides a process for identifying and characterizing a host response to a specified therapy. In some embodiments, the present disclosure is based, at least in part, on identifying one or more biological signatures that differ between two time points for a specified therapy to predict the efficacy and outcome of the therapy.
[0038] In some embodiments, the machine learning models of the present disclosure may be trained on data from a cohort of subjects with a specified disease or disease type, or a predetermined population, in which at least two biological samples are obtained from each cohort participant at two time points relative to treatment, e.g., T0 (e.g., before treatment) and T1 (e.g., during, throughout, or after treatment). In some embodiments, the biological samples are profiled to extract biological signatures, e.g., protein expression profiles.
[0039] Thus, in some embodiments, the present disclosure provides (i) computational approaches for training machine learning models to predict response in patients, and (ii) methods for selecting key proteins whose targeting may improve efficacy and / or response to a therapy.
[0040] This disclosure discusses aspects of the present invention related to predicting responses, e.g., host response, in cancer patients. As used herein, the term "host response" refers to a set of patient-driven factors that may limit or hinder the effectiveness of one or more cancer treatments or therapeutic modalities administered to a patient. However, the present method may be equally effective in predicting treatment and / or therapeutic response in the context of other diseases or disorders. Furthermore, the present method may be effective for enriching patient populations, such as for use in clinical trials. Furthermore, the present method may be effective in identifying novel combinations of therapeutics suitable for treating subjects.
[0041] In some embodiments, biological samples may be obtained from each subject, or at least a portion of the subjects, in a cohort of patients at identified times before, during, and / or after completion of a course of therapy. In some embodiments, biological samples may be obtained from each subject, or at least a portion of the subjects, at one or more specified stages and / or time points and / or steps before, during, and / or after completion of a course of therapy, e.g., before, during, and / or after treatment.
[0042] In some embodiments, a biological signature (e.g., protein expression profile) may be obtained from each of the biological samples. In some embodiments, the set of biological signatures may include statistically tested biological signatures obtained at multiple times (e.g., T0 and T1) from a cohort of subjects receiving a specified treatment. In some embodiments, a preprocessing stage may be performed to preprocess the biological signature data. In some embodiments, the preprocessing stage may include at least one of data cleaning and normalization, feature selection, feature extraction, dimensionality reduction, and / or any other suitable preprocessing method or technique.
[0043] In some embodiments, the paired biological signatures associated with each subject may be analyzed to determine differential expression within each pair, e.g., values associated with the differentially expressed factors (e.g., proteins) in the paired biological signatures. In some embodiments, this analysis provides a difference in the relationship between at least some of the proteins in each signature. In some embodiments, this analysis provides a set of values representing the difference in expression of at least some factors (e.g., proteins) in each or at least some paired biological signatures of the subjects. In some embodiments, the set of values representing the relationship between at least some factors (e.g., proteins) in the paired biological signatures may be based on one or more mathematical equations, such as multiplication of expression values or difference in the relationship between expression values. In some embodiments, the ratio is between the biological signature at T0 and the biological signature at T1. In some embodiments, the ratio is between the biological signature at T1 and the biological signature at T0. As used herein, the terms "paired biological signatures," "paired biological signatures," and variations thereof, refer to biological signatures obtained from multiple (i.e., two or more) biological samples received at multiple time points for a specified treatment. Thus, analysis may compare multiple biological signatures and provide a pattern of signatures over time. In some embodiments, monitoring the progression of a patient's disease state may require multiple sampling of biological signatures from the patient.
[0044] Thus, in some embodiments, a training dataset for a machine learning model of the present disclosure may include multiple sets of values associated with the differences and / or ratios in expression of at least some proteins in pairs of related biological signatures for each, or at least some, of cohorts of subjects each having a specified type of disease and each receiving a specified type of treatment and / or therapy associated with the specified type of disease.
[0045] In some embodiments, paired biological signatures may be correlated using the same factor (e.g., the same protein). In some embodiments, paired biological signatures may be correlated under multiple factors (e.g., various proteins) that define a network of factors (e.g., a protein network). As shown below (Figures 14-16), differential protein correlations can also provide a tool for feature engineering useful for predicting a subject's response to a specified treatment. As a non-limiting example, a protein network can be defined for each biological signature, and calculations are performed to define the overall behavior of each cohort (e.g., calculating the distance from the correlation trend line, as shown in Figure 16A).
[0046] In some embodiments, a training dataset for a machine learning model of the present disclosure may include multiple sets of values associated with differences (e.g., ratios) in expression of at least some factors (e.g., proteins) in an associated biological signature for each, or at least some, of a cohort of subjects having a specified type of disease and receiving a specified type of treatment and / or therapy associated with the specified type of disease, and at least some of the sets of values may be annotated with categorical labels indicative of the response and / or outcome of the treatment in each subject.
[0047] In some embodiments, a training dataset for a machine learning model of the present disclosure includes multiple sets of values associated with differences (e.g., ratios) in the expression of at least some factors (e.g., proteins) in the associated biological signature for each, or at least some, of a cohort of subjects having a specified type of disease and, e.g., receiving a specified type of treatment and / or therapy associated with the specified type of disease, where at least some of the sets of values may be annotated with categorical labels indicating the response and / or outcome of the treatment in each subject, and the annotations may be expressed as binary, e.g., positive / negative, responding / non-responding, continuous, and / or any numerical scale, e.g., 1-5 or complete response, partial response, overall response, duration of response, progression-free survival, adverse events, stable disease, or progressive disease. In some embodiments, additional and / or other annotation schemes may be employed and used for the training dataset. In some embodiments, the training dataset may be annotated with categorical labels indicating, for example, patient demographics and / or clinical data.
[0048] In some embodiments, the trained machine learning models of the present disclosure may provide predictions of response of patients diagnosed with a specified disease to an associated specified treatment or therapy.
[0049] In some embodiments, the trained machine learning model of the present disclosure provides a prediction of a patient's response to a specified treatment or therapy, e.g., as a binary value such as "yes / no," "response / non-response," or "favorable / unfavorable response." In some embodiments, the prediction may be expressed by a value indicating the probability of response (e.g., on a scale of 1 to 100%). In some embodiments, the prediction may be expressed on a scale and / or associated with a confidence parameter. Thus, in some embodiments, the machine learning model of the present disclosure may provide a prediction of the response rate and / or success rate of a specified treatment in a patient, e.g., the likelihood of a patient's favorable response to a specified treatment or therapy. For example, in some embodiments, the prediction may be expressed in discrete categories and / or on a scale including, e.g., "complete response," "partial response," "stable disease," "progressive disease," "pseudoprogression," and "hyperprogressive disease." In some embodiments, the prediction may indicate adverse or any other secondary effect, e.g., a side effect based on the host response. In some embodiments, the prediction may indicate whether the patient's response is associated with adverse or other secondary effects. In some embodiments, the prediction may indicate the patient's overall response to the specified treatment or therapy. In some embodiments, the prediction may indicate progression-free survival following treatment of the patient with a specified therapy or regimen. In some embodiments, the prediction may indicate duration of response rate for the patient. In some embodiments, additional and / or other scales and / or thresholds and / or response criteria may be used, for example, a graded scale from 1 (non-response) to 5 (response).
[0050] In some embodiments, the present disclosure may also provide predictions of adverse events associated with a specified treatment or therapy in a target patient. In some embodiments, the present disclosure may also provide predictions of metastases, metastasis location, and / or tumor burden in a target patient.
[0051] In some embodiments, the present disclosure may provide predictions of overall response, duration of response, and progression-free survival for target patients treated with a specified treatment or therapy.
[0052] In the context of cancer, the term "therapy" refers to any method of treating a specified disease in a subject. In the context of cancer, as used herein, the terms "therapy," "anti-cancer therapy," "cancer therapy modality," "therapeutic modality," "cancer treatment," or "anti-cancer treatment" refer to any method of treating cancer in a cancer patient, including radiation therapy; chemotherapy; targeted therapy, immunotherapy (immune checkpoint inhibition, immune checkpoint modulators, adoptive cell transfer therapy, oncolytic virotherapy, therapeutic vaccines, immune system modulation, and monoclonal antibodies), hormonal therapy, antiangiogenic therapy, and photodynamic therapy; hyperthermia and surgery, or a combination thereof. In some embodiments, the cancer therapy is immunotherapy. In some embodiments, the immunotherapy comprises immune checkpoint modulation. In some embodiments, the immunotherapy comprises immune checkpoint inhibition. In some embodiments, the inhibition comprises administering an immune checkpoint inhibitor. In some embodiments, the inhibitor is a blocking antibody. In some embodiments, the immunotherapy comprises immune checkpoint blockade. Immune checkpoint proteins are well known in the art and include, but are not limited to, PD-1, PD-L1, PD-L2, CTLA-4 (cytotoxic T lymphocyte-associated protein 4); A2AR (adenosine A2A receptor), also known as ADORA2A; B7-H3, also known as CD276; B7-H4, also known as VTCN1; B7-H5; LAG-3 (lymphocyte activation gene-3); BTLA (B and T lymphocyte attenuator), also known as C272; TIM-3 (T cell immunoglobulin and mucin domains); 3); IDO (indoleamine 2,3-dioxygenase); TDO (tryptophan 2,3-dioxygenase); KIR (killer cell immunoglobulin-like receptor); NOX2 (nicotinamide adenine dinucleotide phosphate NADPH oxidase isoform 2); SIGLEC7 (sialic acid-binding immunoglobulin-type lectin 7), also known as CD328; SIGLEC9 (sialic acid-binding immunoglobulin-type lectin 9), also known as CD329, TIGIT, and VISTA (V-domain Ig suppressor of T-cell activation).In some embodiments, the immunotherapy is an anti-PD-1 therapy. In some embodiments, the immunotherapy is an anti-PD-L1 therapy. In some embodiments, the immunotherapy is an anti-PD-L1 / PD-L2 therapy. In some embodiments, the immunotherapy is combined with another immunotherapy. In some embodiments, the immunotherapy is an anti-PD-1 and / or anti-PD-L1 therapy. In some embodiments, the immunotherapy is an anti-CTLA-4 therapy. In some embodiments, the immunotherapy is an anti-PD-1 and anti-CTLA-4 therapy. In some embodiments, the immunotherapy is an anti-PD-L1 and anti-CTLA-4 therapy. In some embodiments, the immunotherapy is combined with another therapeutic modality. In some embodiments, the therapeutic modality is another anti-cancer therapy. Examples of other anti-cancer therapies include, but are not limited to, chemotherapy, radiation, surgery, and targeted therapy. Any other anti-cancer therapy may be used in combination. In some embodiments, the immunotherapy is combined with chemotherapy. In some embodiments, the immunotherapy is combined with targeted therapy. In some embodiments, the immunotherapy is combined with two or more types of additional immunotherapy, hi some embodiments, the immunotherapy is selected from an anti-PD-1 / PD-L1 therapy, an anti-CTLA-4 therapy, and both.
[0053] In some embodiments, the additional treatment modality is a treatment for the side effects of immunotherapy. In general, the side effects of anti-cancer therapeutic agents and immunotherapy are well known. Any such anti-side effect treatment can be employed, including, but not limited to, steroids, folic acid, etc.
[0054] In certain embodiments, the term "treatment" or "therapy" refers to one or more sessions of a patient's treatment. In certain embodiments, the term "pre-treatment" refers to a time point before a specified treatment session, and the term "during treatment" refers to a time point after a treatment session and before the next treatment session. In alternative certain embodiments, the term "during treatment" refers to a time point between the second and third treatment sessions, between the third and fourth treatment sessions, between the fourth and fifth treatment sessions, etc. In some embodiments, the term "post-treatment" refers to a time point after completion of treatment. In certain embodiments, the term "post-treatment" refers to a time point after progression has been identified.
[0055] In certain embodiments, the term "pre-treatment" refers to a time point before the first session of a specified treatment, and the term "during treatment" refers to a time point after the first session of treatment and before the second session of treatment.
[0056] In some embodiments, aspects of the present invention further provide for monitoring a patient's responsiveness to treatment over time. In such embodiments, the analysis can provide a difference in the relationship between at least some of the proteins in each signature between two or more time points (e.g., T2 and T3) following treatment. The difference in the relationship between T1 and T0 is presented in this application for illustrative purposes only. Other differences or ratios are also applicable, such as T2 / T1, T3 / T2, T4 / T3, Tn+1 / Tn, Tn+x / Tn, and the like, as well as T2 / T0, T3 / T0, T4 / T0, and Tn / T0.
[0057] In some embodiments, the paired T0 and T1 expression profiles may correspond to before and after a designated one of the sessions of treatment, which may be the first, second, third, and / or another session of treatment. In such cases, the first expression profile refers to data obtained from a biological sample collected from the subject before receiving the designated one of the sessions of treatment, and the second expression profile refers to data obtained from a biological sample collected from the subject after receiving the designated one of the sessions of treatment.
[0058] Figure 1 is a flowchart of functional steps in a method for training a machine learning model to predict a patient's response to a therapy according to some embodiments of the present disclosure. Figure 2 is a schematic diagram of the process steps of Figure 1.
[0059] In some embodiments, in step 100, multiple biological samples may be received from a cohort of subjects, e.g., a predetermined population of patients with a specified type of disease. In some embodiments, the cohort assembled for purposes of the present disclosure may include multiple patients with the same and / or similar and / or related disease and / or syndrome and / or condition category, and / or related diseases, syndromes and / or conditions. In some embodiments, for at least some of the patients in the cohort, the specified disease and / or condition may be at different stages and / or may be combined with comorbidities and / or diseases. In some embodiments, the specified disease of the present disclosure may be expressed in terms of broad category (e.g., "cancer"), subtype (e.g., melanoma), and / or subcategory (e.g., specified type of melanoma).
[0060] In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a disease characterized by increased proliferation, decreased apoptosis, or both. In some embodiments, the disease is cancer. In some embodiments, the cancer is a solid cancer. In some embodiments, the cancer is a hematopoietic cancer. Types of cancer are well known in the art, and examples of cancer classes include, but are not limited to, sarcoma, melanoma, blastoma, carcinoma, leukemia, and lymphoma. Cancer types can also be classified by tissue / cell type of origin, including, for example, brain cancer, blood cancer, bone cancer, fat cancer, retinoblastoma, head and neck cancer, tongue cancer, nasopharyngeal cancer, pharyngeal cancer, throat cancer, esophageal cancer, stomach cancer, digestive cancer, intestinal cancer, lung cancer, colon cancer, colorectal cancer, liver cancer, pancreatic cancer, gallbladder cancer, penile cancer, thymus cancer, thyroid cancer, genitourinary cancer, prostate cancer, kidney cancer, ovarian cancer, cervical cancer, testicular cancer, skin cancer, glioblastoma multiforme (GBM), and uterine cancer. In some embodiments, the cancer is skin cancer. In some embodiments, the cancer is lung cancer. In some embodiments, the cancer is melanoma. In some embodiments, the cancer is small cell lung cancer (SCLC) or non-small cell lung cancer (NSCLC). In some embodiments, the cancer is genitourinary cancer. In some embodiments, the cancer is head and neck cancer. In some embodiments, the cancer is a solid tumor. In some embodiments, the cancer is a cancer treatable by immunotherapy.
[0061] In some embodiments, the disease is an autoimmune disease. In some embodiments, the autoimmune disease is psoriasis. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is an infectious disease. In some embodiments, the disease is a bacterial, viral, or fungal infection. In some embodiments, the disease is an inflammatory disease. In some embodiments, the disease is a respiratory disease. In some embodiments, the disease is a degenerative disease. In some embodiments, the disease is a neurodegenerative disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a cardiovascular disease. In some embodiments, the disease is a skeletal disease.
[0062] In some embodiments, the biological sample may include any type of biological sample obtained from an individual, including body tissue, body fluid, body excreta, exhaled breath, or other source. In some embodiments, the biological sample is a tumor. In some embodiments, the biological sample is a non-tumorigenic sample. The body fluid may be whole blood, plasma, serum, peripheral blood mononuclear cells (PBMCs), lymph, urine, saliva, semen, synovial fluid, and spinal fluid, and may be fresh or frozen. In certain embodiments of the methods according to the invention, the biological sample is plasma, whole blood, serum, cerebrospinal fluid (CSF), or PBMCs. In certain embodiments, the biological sample is plasma. In alternative certain embodiments, the biological sample is CSF. In some embodiments, the biological sample is a PBMC sample. In some embodiments, the biological sample is a blood sample.
[0063] In some embodiments, the cohort of the present disclosure comprises a group of subjects with similar phenotypes and undergo similar treatment.However, the definition of the cohort may vary depending on the classification of each cohort and the biological commonality of the participating subjects.In some embodiments, the cohort of the present disclosure may comprise patients with, for example, different demographics (e.g., gender, age, ethnicity), clinical measurements, disease stage, medical history, disease treatment history, general medical history (e.g., background diseases, including smoking history and drinking habits), genetic information, physical parameters, etc.
[0064] In some embodiments, patients in a cohort are and / or may be receiving different types of treatment, e.g., monotherapy, combination therapy, multi-stage or multi-session therapy, and / or multi-modality therapy.
[0065] In some embodiments, biological samples may be obtained from each subject in a cohort, or from at least a portion of the subjects, at identified times before, during, and / or after completion of a course of therapy. In some embodiments, biological samples may be obtained from each subject, or from at least a portion of the subjects, at one or more specified stages and / or time points and / or steps before, during, and / or after completion of a course of therapy, e.g., before, during, and / or after treatment.
[0066] In some embodiments, for at least a portion of the subjects, at least one pair of corresponding T0 and T1 biological samples are obtained at two or more different time points during the course of treatment, e.g., (i) pre-treatment, i.e., before the start of the course of treatment, and (ii) post-treatment, i.e., after the end of the entire course of treatment. (i) before treatment, i.e., before the start of the course of treatment, and (ii) during treatment, i.e., at specified time points during the course of treatment; In the case of multi-stage or multi-session treatment, (i) before treatment, i.e., before the start of the course of therapy, and (ii) after completion of a specified stage and / or session of the multi-stage or multi-session treatment; and / or In the case of multi-modality therapy, (i) pre-treatment, i.e., before the start of the course of therapy, and (ii) at designated time points and / or stages associated with one of the multiple treatment modalities. In the case of multi-modality therapy, data may be collected (i) pre-treatment, i.e., before the start of the course of therapy for each treatment modality, and (ii) at designated time points and / or stages associated with one of the treatment modalities.
[0067] In some embodiments, at step 102, each, or at least some, of the biological samples may be analyzed to identify multiple biomarkers and / or extract a biological signature. In some embodiments, the analysis may, for example, obtain a proteomic profile including protein expression for each of the samples. In some embodiments, the protein expression so obtained may identify proteins in each analyzed biological sample. In some embodiments, additional and / or other analyses may be performed on the biological samples to obtain one or more profiles selected from, for example, a DNA profile, an RNA profile, a circulating DNA profile, single-cell RNA sequencing, metabolomics, microbiome, transcriptome, genomics, epigenomics, cellular profiling, single-cell-based analysis, and microRNA. In some embodiments, the circulating DNA profile is a circulating tumor DNA profile. In some embodiments, the circulating DNA profile is a methylated circulating DNA profile.
[0068] In certain embodiments, the biological signature is selected from a DNA profile, an RNA profile, a metabolomic profile (e.g., glycomics, lipidomics), a microbial profile, a genomic profile, an epigenomic profile, a cellular profile, a post-translational modification-based profile, a single-cell-based analysis, and a regulatory RNA profile. In some embodiments, the expression is protein expression. In some embodiments, the expression is RNA expression. In some embodiments, the RNA is mRNA. In some embodiments, the RNA is a regulatory RNA. In some embodiments, the regulatory RNA is a microRNA. In some embodiments, the regulatory RNA is a long non-coding RNA. In some embodiments, the metabolomic profile is a lipid profile. In some embodiments, the metabolomic profile is a nucleic acid profile. In some embodiments, the metabolomic profile is a glycan profile. In some embodiments, the metabolomic profile is a vitamin profile. In some embodiments, the metabolomic profile is a fatty acid profile. In some embodiments, the metabolomic profile is an amino acid profile. In some embodiments, the metabolomic profile is a phenolic compound profile. In some embodiments, the metabolomic profile is an alkaloid profile. In some embodiments, the protein expression is metabolic protein expression. In some embodiments, the protein expression is membrane protein expression. In some embodiments, the protein expression is secreted protein expression. In some embodiments, the protein expression is cellular protein expression. In some embodiments, the biological signature is a genomic profile. In some embodiments, the biological signature is a genomic mutation profile. In some embodiments, the biological signature is a genomic epigenetic profile. In some embodiments, the biological signature is a methylome profile. In some embodiments, the epigenetic profile is a post-translational modification (PTM) profile.In some embodiments, the biological signature is a protein PTM profile. PTMs are well known in the art and include, but are not limited to, methylation, acetylation, phosphorylation, glycosylation, sumoylation, and ubiquitination. In some embodiments, the biological signature is a circulating DNA profile. In some embodiments, the biological signature is a circulating tumor DNA profile. In some embodiments, the biological signature is a methylated circulating tumor DNA profile. In some embodiments, the biological signature is the amount of a circulating tumor DNA profile. In some embodiments, the biological signature is the genotyping of mutations in a circulating tumor DNA profile. In some embodiments, the biological signature is an organism profile. In some embodiments, the biological signature is a microbiome profile. In some embodiments, the biological signature is an extracellular vesicle profile (either number or content). In some embodiments, the biological signature is a microparticle profile (either number or content). In some embodiments, the biological signature is an exosome profile (either number or content). In some embodiments, the biological signature is a circulating cell profile. In some embodiments, the biological signature is a circulating tumor cell profile. In some embodiments, the biological signature is a circulating immune cell profile. As used herein, the term "profile" is intended to encompass any variation in the determined entity, including presence or absence, as well as differences in type (e.g., genotype), amount, percentage, or expression, so long as it is suitable for predicting response to treatment.
[0069] Methods for performing expression profiling are well known in the art. RNA expression can be assayed by any known method, including polymerase chain reaction (PCR), real-time PCR, quantitative PCR, digital PCR, microarray, Northern blotting, and sequencing. In some embodiments, expression profiling involves PCR. In some embodiments, expression profiling involves hybridization to a microarray. In some embodiments, expression profiling involves sequencing. In some embodiments, the sequencing is next-generation sequencing. In some embodiments, the sequencing is deep sequencing. In some embodiments, the sequencing is massively parallel sequencing. Sequencing methods are well known in the art, and devices for sequencing are commercially available. Any known sequencing method can be used in accordance with the methods of the present invention.
[0070] Protein expression can be assayed by any known method, including immunoassays, immunoblotting, immunohistochemistry, FACS, ELISA, Western blotting, proteomic arrays, proteome sequencing, proximity extension assay (PEA)-based assays, aptamer-based assays, multiplex assays, and mass spectrometry. In some embodiments, expression profiling involves hybridizing to a proteomic array. In some embodiments, the proteomic array is an antibody array. In some embodiments, expression profiling involves whole proteome sequencing. In some embodiments, expression profiling involves targeted mass spectrometry. In some embodiments, expression profiling involves untargeted mass spectrometry. In some embodiments, expression profiling involves shotgun proteomics using mass spectrometry. In some embodiments, expression profiling involves top-down mass spectrometry. In some embodiments, expression profiling involves bottom-up mass spectrometry. In some embodiments, expression profiling involves data-independent acquisition (DIA) mass spectrometry. In some embodiments, expression profiling comprises data-dependent acquisition (DDA) mass spectrometry. Proteome / proteomics arrays are well known in the art and are commercially available. Examples of proteomics arrays include, but are not limited to, R&D Systems' Proteome Profiler Array, Creative Proteomics' CP Human Proteome Array, RPPA (reverse phase protein array), RayBiotech's Human Kiloplex Quantitative Proteomics Array, Olink Target 96, Olink Explore 96, and Integral Molecular's Membrane Proteome Array.
[0071] In some embodiments, step 104 may involve a pre-processing stage including at least one of data cleaning and normalization, data quality control, and / or any other suitable pre-processing method or technique.
[0072] Biological data derived from clinical samples may be subject to variations that may occur due to different sample collection or sample preparation procedures, quantification inaccuracies, batch effects, and / or any other technical biases that may lead to analytical errors. Thus, in some embodiments, preprocessing may include a quality control step in which at least some biological signatures may be removed based at least in part on measurable parameters of proteins expressed in the biological signatures.
[0073] In some embodiments, quality control and / or data cleaning and / or data normalization may include any one or more of the following: Data transformation: e.g., log2 transformation, Z-score transformation, median subtraction. · Statistical tests: Key statistical scales such as median, mean, first quartile of the data set (Q1), third quartile of the data set (Q3), variance, standard deviation, coefficient of variation (cv) etc. are calculated to assess data quality. Data visualization: A better understanding of the data becomes possible, such as whether the data is normally distributed or whether there are any technical biases, batch effects, or any outliers that behave significantly differently from the rest of the sample. · Data quality assessment: This involves defining which data should be included / removed / normalized in the analysis, thereby generating a new output containing only the required normalized results. Handling of quality control data: In certain cases, samples that differ significantly are considered for exclusion, mainly due to technical bias. In case of batch effects due to technical reasons, batch effect removal algorithms and / or data normalization can be applied. Batch effect removal: can be done in a variety of ways. Non-limiting examples are using a batch effect removal algorithm (e.g., limma), subtracting components with principal component analysis (PCA), median subtraction, Z-scoring, running the same reference sample in a different batch (a "bridge sample") and correcting based on those values. Handling data below the limit of detection (LOD): Approaches for handling values below the LOD level can be performed by data imputation. As a non-limiting example, T0 or T1 values below the LOD can be assigned to the LOD level of the investigated protein. If both time points are imputed, the T1 / T0 ratio, which is equal to 1, becomes equal to 0 after log2 transformation and, in some data analyses, can instead be assigned as a "not a number" (NaN) value. Other approaches for data imputation can also be used. Figure 3 is a schematic diagram of a quality control process that can be used to evaluate measurements below the limit of detection (LOD) threshold and / or above the maximum threshold. Handling Missing or Zero Values: Proteins with missing values (NaN), below the LOD value, or zero values in less than 0-100% of samples are filtered. Alternatively, missing (NaN), below the LOD value, or zero values can be imputed by other imputation methods. Following data imputation, several QC steps may be repeated.
[0074] Data normalization: If necessary, data is normalized prior to bioinformatics analysis. Data normalization can be performed at any level, for example, at the protein level, batch level, etc.
[0075] In some embodiments, at step 106, a differential expression value may be calculated for each pair of biological signatures. In some embodiments, for biological signatures that are protein expression profiles, the present disclosure provides for calculating the level of difference in expression values between each biological signature in a pair of signatures associated with a subject, e.g., the difference in expression values between biological signatures at at least two time points relative to a treatment, and / or a ratio, e.g., the T1 / T0 ratio. In some embodiments, this analysis does not take into account the biological function of the proteins and / or known interactions between proteins. In some embodiments, the T1 / T0 ratio is a numerical value determined by calculating the ratio of the baseline value (pre-treatment) and the duration of treatment. The T1 / T0 ratio may be used to predict a patient's responsiveness or non-responsiveness to a cancer treatment.
[0076] In some embodiments, additional, other, and / or alternative sets of values may be calculated, e.g., associated with biological processes, clinical data, and / or protein interaction-driven analyses between T1 / T0 signatures. 。
[0077] In some embodiments, in step 108, one or more feature selection, feature extraction, ensemble process, and / or dimensionality reduction steps may be performed on the value set.
[0078] In some embodiments, a feature selection and / or dimensionality reduction step may be performed to reduce the number of variables in each sample pair and / or to obtain a set of key variables, e.g., those variables that may have significant predictive power, such as protein expression levels. Thus, in some embodiments, the feature selection and / or dimensionality reduction step may result in a reduction in the number of proteins in each biological signature and / or value set. In some embodiments, dimensionality reduction selects key variables, e.g., proteins, based on the level of response predictive power they produce for the desired prediction. In some embodiments, dimensionality reduction generates a new feature or features that can predict response. In certain embodiments, dimensionality reduction involves treating all or some feature values as vector components and calculating their norm.
[0079] In some embodiments, any suitable feature selection and / or dimensionality reduction method or technique may be employed, including but not limited to: ANOVA with S0 parameter: Analysis of variance with an additional parameter (S0) that controls the relative importance of features based on the resulting test p-value and the difference between group means (see, for example, Tusher, Tibshirani and Chu, PNAS 98, pp. 5116-21, 2001). Scalable EMpirical Bayes Model Selection (SEMMS): An empirical Bayes feature selection method that applies parsimonious mixture models to identify significant predictors (see, e.g., Bar, Booth, and Wells, A scalable empirical Bayes approach to variable selection in generalized linear models, 2019). L2N: Differential expression analysis using a three-component mixture model. This model consists of two log-normal components (L2) for differentially expressed features, one component for under-expressed features and another component for over-expressed features, and a single normal component (N) for non-differentially expressed features (see, e.g., Bar and Schifano, Differential variation and expression analysis. Stat 8, e237, doi:10.1002 / sta4.237, 2019). Genetic Algorithms: A family of heuristic optimization algorithms that employ organic evolutionary techniques such as random mutation, recombination, and natural selection as a way to achieve an optimal configuration (see, for example, Popovic, Sifrim, Pavlopoulos, Moreau, and Bart De Moor, A Simple Genetic Algorithm for Biomarker Mining. 2012). Naive Classifier: A naive classifier evaluates response scores by reducing dimensionality to a single score. This is done by considering all features (e.g., a specific profile, such as protein expression levels) as components of a vector and calculating its norm. Dimensionality reduction reduces the potential risk of overfitting. In some embodiments, vector components are normalized according to typical component values among patients belonging to the same response group (e.g., responders), so that the normalized norm quantifies the amount of deviation from the typical respective class value. In additional embodiments, a naive classifier allows training using data from subjects belonging to only a portion of the response group.
[0080] In some embodiments, at step 110, a training dataset for training a machine learning model of the present disclosure may be constructed including a set of values representing the relationship (e.g., ratio or difference in expression values) of biological signatures at multiple time points with respect to treatment for at least a portion of the subjects in the cohort.
[0081] In some embodiments, the training dataset of the present disclosure may include additional information for training the machine learning model, such as clinical, demographic, and / or physical information about at least a portion of the subjects in the cohort. For example, in some embodiments, such data may include features obtained from the diseased tissue itself (e.g., from tumors in cancer patients). In some embodiments, such data may include demographic information (e.g., age, ethnicity, etc.); performance status; hematological and chemistry measurements; cancer history, e.g., cancer diagnosis date, primary cancer type and stage, disease biomarkers (e.g., PD-L1), disease treatment history, histology, TNM stage, assessment of measurable lesions, time to tumor progression, site of recurrence, proposed treatment; general medical history, including smoking history and drinking habits, background diseases, including hypertension, diabetes, ischemic heart disease, renal failure, chronic obstructive pulmonary disease, asthma, liver failure, inflammatory bowel disease, autoimmune diseases, endocrine diseases, etc.; family medical history; genetic information, such as genetic information, mutations, gene amplifications (e.g., EGF, These parameters may include, but are not limited to, R, BRAF, HER2, KRAS, MAP2K1, MET, NRAS, NTRK1, PIK3CA, RET, ROS1, TP53, ALK, MYC, NOTCH, PTEN, RB1, CDKN2A, KIT, and NF1; physical parameters such as body temperature, pulse, height, weight, BMI, blood pressure, complete blood count (including all investigated parameters), liver function, renal function, and electrolytes; medications (prescribed and nonprescribed); relative lymphocyte count; neutrophil-to-lymphocyte ratio; baseline protein levels in plasma (e.g., LDH); and / or marker staining (e.g., PD-L1 in tumors or on circulating tumor cells). In some embodiments, changes in one or more of the above information in response to a specified therapy may be analyzed and provided for training a machine learning model.
[0082] In some embodiments, one or more annotation schemes may be employed with respect to the training dataset. Thus, in some embodiments, a training dataset for a machine learning model of the present disclosure may include multiple sets of T1 / T0 ratios or differences in expression values for at least some of the subjects in a cohort, where at least some of these sets of values may be annotated with categorical labels indicating the response and / or outcome of treatment in each subject. In some embodiments, such annotations may be binary, e.g., positive / negative, and / or expressed in discrete categories, e.g., on a scale of 1 to 5. In some embodiments, binary-valued categorical labels may be expressed, e.g., as "yes / no," "response / non-response," or "favorable / unfavorable response." In some embodiments, discrete categorical labels and / or annotations may be expressed on a scale of, e.g., "complete response," "partial response," "stable disease," "progressive disease," "pseudoprogression," and "hyperprogressive disease." In some embodiments, additional and / or other scales and / or thresholds and / or response criteria may be used, e.g., a graded scale from 1 (non-response) to 5 (response). In some embodiments, the category label may be associated with adverse or any other secondary effect or response by the patient, for example, a side effect of the treatment.
[0083] In some embodiments, additional and / or other annotation schemes may be employed. In some embodiments, the training dataset may be annotated with patient demographic and / or clinical data, for example, as detailed above. In some embodiments, the training dataset may be annotated with overall response rate. In some embodiments, the training dataset may be annotated with progression-free survival rate. In some embodiments, the training dataset may be annotated with duration of response rate.
[0084] In some embodiments, in step 112, a machine learning model may be trained on the training dataset constructed in step 110. In some embodiments, any suitable machine learning algorithm or combination of methods may be employed, including but not limited to: Support Vector Machine (SVM): a non-parametric model that finds the optimal separating hyperplane that distinguishes between different classes. It can perform linear or non-linear classification. Penalized Logistic Regression (PLR) - a logistic model of regression that imposes penalties to reduce the influence of certain features. Generalized Linear Models (GLMs): A generalization of linear regression that integrates statistical models such as linear regression, logistic regression, and Poisson regression. GLMs extend linear regression by (1) supporting response variables with error distributions other than normal, and (2) by nonlinear relationships between predictors and the response variable. Random Forest (RF): involves generating multiple decision trees consisting of sequences of decision rules for protein expression values. These trees can be pruned to avoid overfitting. Each tree is constructed by randomly selecting a different sample. eXtreme Gradient Boosting (XGB): A gradient-boosted decision tree-based classification and regression algorithm. Decision trees are built one at a time, and each new tree corrects the errors of the previously trained decision tree.
[0085] In other embodiments, the machine learning model may be trained based on statistical measures, i.e., variance, median, mean, typical value, etc.
[0086] In some embodiments, in step 114, the machine learning model uses a classifier to predict response in target patients and generate a target set of T1 / T0 relationships (e.g., ratios or differences in expression values) that are suitable for receiving similar treatment as the patient cohort.
[0087] In some embodiments, at inference step 114, the trained machine learning model of the present disclosure may be applied to target data, e.g., a target set of T1 / T0 relationships (e.g., ratios or differences in expression values) for target patients with similar phenotypes and receiving similar treatments as the patient cohort. In some embodiments, inference of the trained machine learning model on the target data generates a treatment response prediction or response probability.
[0088] In some embodiments, the prediction is for a side effect or adverse event. In some embodiments, the prediction is for overall survival. In some embodiments, the prediction is for progression-free survival. In some embodiments, the prediction is for duration of response rate. In some embodiments, the prediction is for pseudoprogression. In some embodiments, the prediction is for hyperprogression. In some embodiments, the prediction is for disease progression. In some embodiments, predictions according to the present disclosure may be further supported and / or supplemented by identification of differentially expressed proteins and / or differentiation analysis of biological processes, as further detailed herein below.
[0089] In some embodiments, at step 116, a course of therapy for the target patient may be administered, adjusted, and / or modified based at least in part on inferring step 114. In some embodiments, such therapy adjustment may include prescribing follow-on and / or supplemental therapy for the target patient.
[0090] Identification of differentially expressed proteins Enrichment analyses, including enrichment analyses when focusing on a subset of features, or some network-based analyses when providing personalized potential targets for therapeutic intervention, as defined below, require first identifying DEPs between the investigated groups (e.g., responders vs. non-responders).
[0091] The term "differentially expressed protein" (DEP) refers to proteins whose distribution (expression level or change between two time points, e.g., TO and T1) differs between responders and non-responders (and possibly other groups, e.g., stable disease patients), including distribution differences detectable by numerical measurements (e.g., t-test, ANOVA, Kolmogorov-Smirnov test). In some cases, DEPs are defined as proteins whose median, mean, variance, or other statistical measure differs between responders and non-responders. In some cases, DEPs are defined as proteins whose distribution differs between responders and non-responders, where the mean or median does not change; however, the difference can be assessed by statistical tests. In some cases, DEPs are defined as proteins that do not follow the respective protein distribution between specific subgroups (i.e., responders or non-responders) in at least one patient.
[0092] Determining DEPs is an optional step, as some tools do not require a list of DEPs but rely on a selected metric calculated for all proteins in the proteomic profile, e.g., the fold change between the two investigated groups (i.e., responders and non-responders) or the p-value of a t-test.
[0093] The terms "protein profile," "protein expression profile," and "proteomic profile," used interchangeably herein, refer to the expression levels of proteins, e.g., cytokines, growth factors, and other proteins, or lists of proteins, expressed in plasma, CSF, or other body fluids, or tissues, at a particular time point. The number of proteins measured can vary between 1 and 20,000. Protein profiles can be used to diagnose diseases, conditions, or syndromes and determine the likelihood of therapeutic response. In some embodiments, the profile includes at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 500, 1000, 2000, 5000, 7000, or at least 20,000 proteins. Each possibility represents a separate embodiment of the present invention. In some embodiments, the protein profile is absolute protein expression. In other embodiments, the protein expression profile is a normalized or relative expression profile.
[0094] Differentiation of biological processes The differentiation of biological processes can be deciphered by using DEP as input. Another way is by performing per sample pathway enrichment and then aggregating the enrichment results.
[0095] The term "differentiation of biological processes" (DBP) refers to biological processes that occur in either responding or non-responding patients.
[0096] Differentiation of biological processes can be provided based on various databases, such as KEGG pathway analysis (https: / / www.genome.jp / kegg / ) and gene ontology (GO) analysis (geneontology.org).
[0097] Pathway enrichment analysis Proteome analysis at the pathway level can suggest biological processes affected by treatment. The purpose of this step is to translate host response-related changes at the protein expression level into changes at the biological process level. Therefore, the identified DEPs can be used as input for pathway enrichment analysis.
[0098] Pathway enrichment analysis translates proteomic changes into a list of putative biological processes that are down- or up-regulated in the body in response to cancer therapy treatment. The analysis can identify DBPs, or biological processes enriched in each group. The output of enrichment analysis is a list of DBPs and the proteins involved in each of them.
[0099] Network analysis An additional level of analysis based on proteomic data involves studying the co-changes in the expression levels of various proteins together, where proteins that exhibit correlations (either negative or positive correlations) between their proteomic profiles may indicate potentially interesting biological relationships. Correlations between proteins can arise for a variety of reasons, such as a common regulator (e.g., transcription factor, phosphatase, etc.) that affects both proteins, or one of the proteins being a regulator of the other. Protein networks can be constructed based on the correlations between the expression levels of different proteins. Because protein networks can differ between two different conditions (e.g., responders and non-responders), studying these differential networks (Figure 14) can decipher the biological mechanisms underlying different phenotypes (e.g., resistance to a therapy). An example of such a differential network is depicted in Figure 15.
[0100] Examination of differential networks between responders and non-responders may reveal new mechanisms associated with resistance to therapy and can be used to train classifiers. Network analysis can aid in feature engineering. For example, such analysis can help identify co-varying features, where engineering features capturing this relationship can be predicted by using the mathematical relationship between two or more proteins. Figure 16 shows a non-limiting example of a predictor based on correlation between protein pairs. Two proteins showing differential correlation fold-change values between responders and non-responders were examined. Figure 16A—The correlation between responders (triangles) and non-responders (circles) was positive (R = 0.37). The dashed line indicates a linear fit to the responder value. Figure 16B—The residuals of the two protein pairs (i.e., the distance of each point from the linear fit in Figure 16A) were calculated and used as input for an SVM-based classifier. The resulting predictor achieved a ROCAUC of 0.765. In addition, network analysis can potentially identify new approaches for intervention that would be highly valuable to pharmaceutical companies.
[0101] Potential personalized targets for intervention In another aspect, the present invention provides a personalized analysis of potential targets for intervention using a three-filter approach. The three-filter approach utilizes the strength of the cohort and can be visualized as a decision tree, such as the flowchart in Figure 8, which begins with a first filter that retains only DEPs. Proteins that are not DEPs do not continue to the next filter and are considered inactive proteins. The DEPs used for this analysis can be similar to or different from the DEPs provided for training the machine learning model. The second filter focuses on the patient-specific proteomic profile of the DEPs and retains only the patient's resistance-associated proteins (RAPs) based on their RAP score (described in more detail below). DEPs with a RAP score below a threshold (e.g., a RAP score of 1) are deemed not relevant to the particular patient and are excluded. Next, a third filter is applied to associate the RAPs with drugs or investigational drugs. The third filter can include two options: if the RAP has an approved drug (e.g., specific to the indication of interest or another indication), the RAP is considered an active protein. Alternatively, if there are clinical trials linking RAP with drug candidates (eg, for an indication of interest), RAP may also be considered an actionable protein.
[0102] The analysis provided herein may further include a simulation step in which a patient's RAP (or multiple RAP) expression values are modified toward an equilibrium value that may follow therapeutic intervention. Another predictive analysis is then performed to assess the impact of changes in protein expression values. This may help a physician determine which RAP to select for a patient (multiple RAPs that are potential targets for intervention may be received for a selected patient). Figure 12 shows a non-limiting simulation of a single RAP perturbation (although simulations of several RAP perturbations may also be performed). A predictive signature based on a list of RAPs for the entire cohort is first generated (Figure 12A). For a given patient, a particular RAP, or combination of RAPs, is balanced, and the baseline response probability is compared to the perturbation response probability (Figure 12B).
[0103] Cohort-based statistical filters First, it is essential to identify DEPs. As described herein, DEPs are proteins whose levels vary between responders and non-responders, and may also vary between T0 and T1 in responders and non-responders. In some cases, DEPs are proteins whose median values differ between responders and non-responders. In some cases, DEPs are proteins whose variances differ between responders and non-responders. In other cases, DEPs are proteins whose representations differ between responders and non-responders. This analysis narrows down the number of proteins to obtain a list of proteins that represent differences between two or more classes per cohort, and thus potentially play a role in resistance or response to treatment. DEPs can be the same as or independent of the DEPs identified for the machine learning aspect of the present invention.
[0104] Personalized Filters The list of proteins from the first filter (cohort-based statistical filter) is then interrogated in a patient-specific manner using the personalized filter, and the resistance-associated protein (RAP) score.
[0105] A RAP is defined for each patient. In some embodiments, a patient's RAP is a protein whose level or fold change deviates from the respective protein distribution in one of the response groups. Deviation can be quantified by numerical means, either by using the levels of multiple response groups (e.g., both responders and non-responders) or by using the distribution among specific response groups (e.g., responders or non-responders, or patients with stable disease). As a non-limiting example, the expression distribution of each DEP in the entire cohort can be examined for each response class (i.e., distributions for responders and non-responders are generated separately). In this case, a probability density distribution can be extracted for each group distribution with a total area under the curve of 1. For each patient, the upper or lower tail area (or other measure, such as the height in each of the response group distributions) of the DEP expression level (e.g., the choice of upper or lower may depend on the order of the medians of non-responders and responders, as described in detail below) is calculated for each distribution (Figure). A RAP score can then be calculated for each DEP for each patient, such as based on non-limiting Equation 1:
number
[0106] where P(NR) is the probability that a protein can be assigned to the non-responder distribution, and P(R) is the probability that a protein can be assigned to the responder distribution. In this example, RAP is a probabilistic measure of the distance of DEP expression levels from the responder distribution. A high RAP score indicates a protein that is more likely to be associated with the non-responder distribution rather than the responder distribution.
[0107] The RAP score is
number
[0108] For each patient, the DEPs following the first filter are examined, and a RAP score is calculated for each DEP. Optionally, a threshold for passing a personalized filter can be defined (e.g., a RAP score greater than 1.0). The list of RAPs for this patient continues with the next filter.
[0109] Calculating directionality A RAP score above a defined threshold may indicate that the protein can be assigned to the non-responder distribution (note that a high RAP score can assign the RAP to the responder distribution, respectively). When there is a difference in median between responders and non-responders, the selection of the tail direction of the area under the curve can be based on the order of the distribution medians. If the median of the non-responder group is higher than that of the responder group, the area selected for calculation is the right tail of the patient's DEP expression level (Figure 10A). If the median of the non-responder group is lower than that of the responder group, the area under the left tail of the distribution is used for the patient's DEP expression level (Figure 10B). Thus, proteins likely to be assigned to the non-responder distribution have a value greater than 1.0 (log2 scale), regardless of the DEP direction in the cohort (higher in responders or higher in non-responders). DEPs that are likely to be attributed to the responder group have a RAP score between 0 and 1 (log2 scale). If the median is not the differential but some other aspect of distribution variation, such as variance or typicality, the relative probability in a given range (not necessarily the tail) can also be used.
[0110] The impact of DEP statistics on the distribution of RAP scores Because the RAP score is based on the distribution of responders and non-responders across the cohort, differences in the distribution between responders and non-responders affect the RAP score distribution (Figures 11A-11D). DEPs with a larger difference between responders and non-responders indicate that patients are more likely to have a high RAP score compared to DEPs with a smaller gap between RAP scores and a smaller difference between responders and non-responders (Figures 11A and 11B). This indicates that patients are less likely to have a high RAP score due to a larger gap between RAP scores (Figure 11C).
[0111] RAP-based cluster This shows that the RAP score can be further used to identify groups of non-responders who share similar RAP profiles. For this analysis, various clustering algorithms can be used. A non-limiting example is the use of a consensus clustering algorithm to find the most robust cluster of samples following multiple iterations of clustering. An example of this is shown in Figure 11D. Various clusters can then be further characterized to explore the enrichment of various clinical parameters. For example, cluster #5 was enriched in patients who had recently quit smoking (Fisher's exact test p-value = 0.027).
[0112] Clinical Filters Following the first two filters described above, a list of personalized RAPs is generated for the patient. In the current filter (FIGS. 12A-B), potential drugs or investigational new drugs (INDs) targeting the patient's RAPs are identified. Clinical filtering may include the following steps: First, drugs / INDs targeting each RAP are searched for in a suitable database and associated with the RAP. Next, various clinical filters may be applied following data collection and analysis, which may include a research layer based on biological reasoning, a layer related to mechanism of action (direct / indirect, specific / non-specific, directional agreement, etc.), and a layer related to the drug's clinical relevance (drug development status / clinical relevance, etc.). RAPs associated with drugs / INDs based on the applied filters can be considered potential intervention targets for the patient. Alternatively, they can be considered the basis for potential collaboration with pharmaceutical companies.
[0113] Experimental results Example 1 - Melanoma Cohort The inventors conducted experiments to test the predictive ability of the machine learning model of the present disclosure.
[0114] The experimental training dataset consisted of biological samples from melanoma patients. Response to treatment for each patient was determined based on either Response Evaluation Criteria in Solid Tumors (RECIST) estimates or clinical benefit assessment. Patients with progressive disease (PD) were classified as non-responders (NR). Patients who demonstrated a partial response (PR) or complete response (CR) were classified as responders (R). Patients with stable disease (SD) were classified as SD patients.
[0115] In the bioinformatics analysis, SD samples were excluded to focus on the two more extreme and distinct groups: responders and non-responders. The exclusion of the SD group is often performed by other research groups because it behaves differently from the R and NR groups. Finally, the dataset included 33 samples for analysis. The clinical parameters of the training dataset are presented in Figure 4.
[0116] For each patient, pretreatment (T0) and early treatment (T1) plasma samples were collected. Using antibody array technology from RayBiotech (see Wilson, JJ et al., Antibody arrays in biomarker discovery. Adv Clin Chem 69, 255-324, doi:10.1016 / bs.acc.2015.01.002 (2015)), proteomic changes during anti-PD1 or anti-CTLA4 combined with anti-PD1 therapy were profiled. A total of 400 proteins were profiled per sample. Predictive biological signatures for response to treatment were extracted based on the log2 (log fold change) of the T1 / T0 ratio for each protein.
[0117] The data were processed in the following steps. First, T0 or T1 values below the limit of detection (LOD) were assigned an LOD value. Following log2 fold-change transformation, proteins with both T0 and T1 LOD values had a log2 fold-change value of 0. The data were filtered to retain proteins with a value of 0 in less than 50% of the samples (proteins for which both T0 and T1 values were below the limit of detection). Overall, this filtering step resulted in 330 valid proteins for downstream analysis. Additionally, in the QC analysis, large variations between samples were observed, while not all were centered around the 0 value. Therefore, the data were normalized by subtracting the overall median value from each sample.
[0118] To identify proteomic signatures that allow prediction of response, patients with relatively high or low overall variability were excluded from the training set. Thus, the first step consisted of differential expression analysis (based on log2 fold change values) between the group of responders (n = 5) and non-responders (n = 8).
[0119] To identify response-predicting proteins, the L2N method was applied to identify differentially expressed proteins (DEPs) between responders and nonresponders. One advantage of using this differential expression approach to identify predictive proteins is that it relies on a normalized model for continuous data, which is more powerful than binary classification methods and therefore requires a smaller sample size to obtain a set of predictors. A second advantage of using this approach is that it reduces the chance of overfitting because it provides separation between two steps: the first step is the process of fitting a model to find differentially expressed proteins, and the second step involves fitting a logistic model of the patient's response status using only the differentially expressed proteins as predictors. Using this approach, 10 differentially expressed proteins were identified. Response to treatment was determined based on clinical benefit (i.e., responder or nonresponder).
[0120] A logistic regression model (GLM) was generated using the 10 selected proteins as predictors and the true response status as the dependent variable. A good prediction was obtained with an area under the curve (AUC) of 0.84 in the receiver operating characteristic (ROC) plot for the entire dataset of 33 patients (Figure 5A), with a total of six misclassifications. Note that three of the misclassified patients were included in the 13 samples used in the first step, suggesting a low likelihood of overfitting. Of the 20 patients left behind, the predicted probability of responding to treatment for 17 patients (85%) was consistent with their true status. Points in the ROC plot were selected for sensitivity and specificity of at least 90% (Figure 5B), resulting in a sensitivity and specificity of 0.93 and 0.79, respectively.
[0121] Another approach to discovering predictive signatures for response to therapy based on host response using log fold change values was performed using a linear support vector machine (SVM) algorithm. In this approach, response to therapy was determined based on Response Evaluation Criteria in Solid Tumors (RECIST) estimates. Using this approach, various single-protein predictors were identified, and the top 25 protein predictors showed various AUCs ranging from 0.7 to 0.822. To maximize the predictive ROC AUC with a minimum number of proteins, a multi-protein model was also generated. The best prediction was obtained with a model consisting of three proteins, which yielded an ROC AUC of 0.88 (Figure 5C-5D).
[0122] A different cohort of patients was used to validate the results of the single-protein and multi-protein classifiers. This validation cohort dataset consisted of biological samples from 14 patients, with pre-treatment (T0) and early-treatment (T1) plasma samples collected. Proteomic changes during ICI treatment were profiled as previously described. Validation of previously obtained single-protein classifiers in the validation cohort demonstrated that not all single-protein-based models generalized well in the validation cohort dataset. Further investigation of the multi-protein classifier validation demonstrated that a three-protein model showed good predictive results with an AUC of 0.85.
[0123] Enrichment analysis Previous studies have shown that combination therapy improves immunotherapy response rates. In this cohort of melanoma patients, differences between anti-PD1 monotherapy and anti-PD1 and anti-CTLA4 combination therapy were examined by analyzing the key biological pathways and proteins underlying these treatment modalities and identifying differential biological processes. In contrast to classifier analyses, which do not consider the biological function of the proteins that are part of the classifier, this analysis aims to characterize the host response in the context of differential biological processes (DBPs) that change with treatment and differ between responders and non-responders. This exploration can be used as the basis for identifying driver proteins that could be potential targets for intervention as part of combination therapy with immunotherapy. To this end, proteomic data were analyzed using the MetaCore tool from Clarivate Analytics. A major advantage of using this tool is the highly curated database on which the protein maps are based.
[0124] Before performing enrichment analyses that could identify DBPs, statistical tests were applied to select differentially expressed proteins (DEPs; proteins whose levels change between responders and non-responders and / or between T0 and T1 in responders and non-responders) that could serve as input for the enrichment analysis.
[0125] To obtain the strongest proteins with the highest potential to capture biological differences between the two groups and the host response effect of therapy, the focus was on proteins that passed both a one-sample t-test (to identify proteins that change between T0 and T1) and a two-sample t-test (to identify proteins that differ between responders and nonresponders). Using this approach, DEPs were identified in the entire cohort and in two subsets of patients (two different treatment modalities: anti-PD1 monotherapy or anti-PD1 + anti-CTLA4 combination therapy). The DEPs for the entire cohort reveal distinct functional groups, particularly differences in MAPK signaling pathways or metabolism-related proteins. Figure 6A illustrates the differentially expressed proteins (DEPs) identified in plasma. Proteins are displayed using Proteomaps in a Voronoi plot, where each polygon specifies a DEP, and size correlates with change from T0 to T1. DEPs are grouped into KEGG functional groups.
[0126] Next, enrichment analysis was performed using MetaCore with DEP as input, following multiple comparison correction using an FDR-adjusted p-value <0.05. Non-responder enrichment analysis of the entire cohort revealed multiple pathways related to immunosuppression, such as regulatory T cell engagement, as well as pathways associated with cancer progression or skin sensitization (Figure 6B). The latter may be a potential side effect of immunotherapy in patients.
[0127] Additional enrichment analysis of non-responder DEPs across treatment modalities (monotherapy vs. combination therapy) revealed processes unique to each modality. Some of the significantly enriched pathways are involved in immunosuppression and may be part of the host response to immunotherapy that dampens response to treatment (Figure 6C).
[0128] Example 2 - Lung Cancer Cohort A cohort of 33 stage IV NSCLC patients treated with anti-PD1 therapy (nivolumab or pembrolizumab) was recruited. Response to treatment was determined using Response Evaluation Criteria in Solid Tumors (RECIST) 1.1 or estimated based on clinical assessment. Of the 33 samples, 15 patients were defined as responders (including complete responders, CR, partial responders, PR, stable disease, and SD), and 18 were defined as non-responders (including progressive disease and PD). For each patient, pretreatment (T0) and early treatment (T1) plasma samples were collected, and proteomic changes following anti-PD1 therapy were assessed. In total, 760 proteins were assessed per sample. Following data normalization and quality control, checks were performed to identify technical bias and technical outliers (outliers were not removed in this analysis). With the aim of extracting predictive signatures for response to treatment based on host response-related changes, the T1 / T0 ratio (fold change) for each protein following log2 transformation was examined. Proteins with values below the limit of detection (LOD) were then excluded, leaving a total of 418 proteins for analysis.
[0129] Bioinformatics analysis was performed using a multi-layer approach. Following data quality checks and normalization, the analysis continued in two parallel tracks. One track was directed toward classification, aiming to generate a classifier that would allow prediction of response to treatment based on host response data, as reflected by measuring changes between T0 and T1. The second track was aimed at identifying driver proteins. This involved investigating proteins in terms of their functional groups by applying advanced pathway enrichment tools. Further analysis with causal inference allowed the identification of driver proteins that could initiate the enriched processes identified in the analysis.
[0130] Next, a support vector machine (SVM) algorithm was used to discover potential predictive signatures for response based on host response. Overall, a three-protein signature was identified for NSCLC indications receiving anti-PD1 therapy. The three-protein signature had high predictive power, as indicated by the area under the curve (AUC) of the receiver operating characteristic (ROC) plot of 0.89 (Figure 7A). The confusion matrix results were displayed in a Sankey plot (Figure 7B). The cutoff point of the prediction probability scale was set to identify responders with 93% sensitivity, resulting in a specificity of 61%. The sensitivity threshold was set above 90% to avoid classifying responders as nonresponders.
[0131] To validate the results in an independent cohort, a validation cohort was assembled consisting of 54 samples from patients with stage IIIB-IV NSCLC receiving anti-PD1 therapy, including 15 responders and 39 non-responders. The three-protein signature was examined in a blinded manner, i.e., without response annotation for any of the samples. The AUC of the ROC curve for the validation set was 0.72, with a significant p-value of 0.013 (Figure 7C). Similar to the performance analysis of the training set, a high sensitivity (≥90%) cutoff was set to identify responders with a sensitivity of 93%, resulting in a specificity of 26%.
[0132] Enrichment analysis Next, we aimed to characterize the host response in the context of the differentiation of biological processes (DBPs) that change upon treatment and differ between responders and non-responders. This analysis can then be used as the basis for identifying proteins, referred to herein as "driver proteins," that may be potential targets for intervention, such as part of a combination therapy with a specified treatment. To this end, proteomic data can be analyzed using proteomic tools, including, but not limited to, the Key Pathway Advisor (KPA) commercial tool from Clarivate Analytics.
[0133] Before performing enrichment analysis that could identify DBPs, a statistical test was applied to select DEPs whose levels changed between T0 and T1 in responders and non-responders. DEPs served as the power for the enrichment analysis. To achieve this goal, separate one-sample Student's t-tests were performed for each group. Overall, 42 and 40 DEPs were identified in the non-responder and responder groups, respectively (p-value < 0.05).
[0134] Using the selected DEPS as input, enrichment analysis was performed using KPA under default settings of p-value thresholds of 0.05 and 0.01 for enrichment analysis and causal inference, respectively. Overall, there were 112 significantly enriched pathways in non-responders and one enriched pathway in responders. Of the 112 non-responder-enriched pathways, 21 pathways had a causal prediction for the direction of change (up- or down-regulated at T1 compared to T0). Among the differentiating biological pathways, multiple pathways related to immune response can be found, including either immune cell differentiation or interleukin-related signaling. In addition, there are multiple processes associated with cell adhesion and extracellular matrix (ECM) regulation. On the other hand, only a single pathway was enriched in the responder group.
[0135] Using KPA causal inference, it is possible to identify driver proteins potentially involved in host response-related DBPs. In non-responders, 979 driver proteins were identified, while in responder groups, 5 driver proteins were identified.
[0136] Example 3 - Predicted response in patients suffering from psoriasis A classifier based on the algorithm presented herein was trained to predict treatment response in psoriasis patients. The data used for this analysis were taken from Lewis E. Tomalin et al., "Early Quantification of Systemic Inflammatory Proteins Predicts Long-Term Treatment Response to Tofacitinib and Etanercept," Journal of Investigative Dermatology (2020) 140, 1026-1034. Blood samples were collected from 140 patients (96 responders and 44 non-responders) with moderate to severe chronic plaque psoriasis treated with the Janus kinase inhibitor compound tofacitinib (Xeljanz, 10 mg twice daily). A total of 92 inflammation-related (IFN) proteins and 65 proteins associated with cardiovascular disease (CVD) were assessed in blood samples collected immediately before treatment (W0, baseline) and 4 weeks after treatment (W4). Response to treatment is based on PASI75 (Psoriasis Area Severity Index [PASI] 75 at Week 12), a classic efficacy endpoint for psoriasis, where a patient is considered a responder if PASI is reduced by more than 75% after 12 weeks of treatment, otherwise a non-responder. A classifier for predicting whether a given patient will be a PASI-75 responder after 12 weeks of treatment was created using the methods described herein.
[0137] To this end, data from tofacitinib treatment of psoriasis were randomly divided into two subsets of equal size. The first was used to train a machine learning algorithm, and the second was used to validate the algorithm's results. A predictive signature was identified (Figure 13B, AUC ROC = 0.751) that passed validation (Figure 13A, AUC ROC = 0.772). This predictive signature was limited to three protein features to minimize the possibility of overfitting. This signature combined with week 4 (T1) and fold-change data demonstrates that the system and method of the present invention can be applied to assess the response of psoriasis (i.e., conditions other than cancer).
[0138] The present invention may be a system, a method, and / or a computer program product, which may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0139] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction-execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded device having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transitory signal, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or an electrical signal transmitted over a wire. Rather, the computer-readable storage medium is a non-transitory (ie, non-volatile) medium.
[0140] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.
[0141] The computer-readable program instructions for carrying out the operations of the present invention may be either source code or object code written in any combination of assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or one or more programming languages, such as object-oriented programming languages such as Java, Smalltalk, C++, and traditional procedural programming languages such as the "C" programming language, R, Python, etc. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a standalone software package, partly on the user's computer and partly on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the present invention.
[0142] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0143] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, cause the machine to generate means for implementing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored within a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture containing instructions that implement an aspect of the function / acts specified in a block or blocks of the flowcharts and / or block diagrams.
[0144] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device and cause the computer, other programmable apparatus, or other device to perform a series of operational steps to generate a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the function / act specified in a block or blocks of the flowcharts and / or block diagrams.
[0145] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing a specified logical function. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of dedicated hardware and computer instructions.
[0146] The description of a range of numerical values should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range of 1 to 6 should be considered to have specifically disclosed subranges of 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0147] The description of various embodiments of the present invention has been presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, practical applications to commercially available technologies, or technical improvements, or to enable persons other than those skilled in the art to understand the embodiments disclosed herein.
Claims
1. 1. A system comprising: at least one hardware processor; and a non-transitory computer-readable storage medium having program instructions stored therein, the program instructions comprising: receiving, for each of a plurality of subjects having a specified type of disease and receiving a specified therapy for treating the disease, (a) a first biological signature associated with a biological sample collected at a first time point associated with the specified therapy, and (b) a second biological signature associated with a biological sample collected at a second time point associated with the specified therapy; calculating, for each of the plurality of subjects, a set of values representing a relationship between the first biological signature and the second biological signature associated with the respective subject; and During the training stage, (i) a set of calculated values; and (ii) a label associated with the outcome of the specified treatment in each of the subjects; and A system executable by the at least one hardware processor to generate a classifier suitable for predicting a response in a target patient to the specified therapy.
2. 2. The system of claim 1, wherein the first biological signature and the second biological signature are each one of a DNA profile, an RNA profile, a protein profile, a metabolomic profile, a microbiome profile, a transcriptomic profile, a genomics profile, an epigenomics profile, a cellular profile, a post-translational modification-based profile, a single-cell-based analysis, and a regulatory RNA profile.
3. 3. The system of claim 1, wherein the first biological signature and the second biological signature are each protein expression profiles, and the set of values each comprises, for each protein in the protein expression profile, a ratio or difference in expression levels of the protein in the first biological signature and the second biological signature.
4. The system of claim 3 , wherein the protein expression profile comprises expression values of at least two proteins.
5. 5. The system of claim 1, wherein the program instructions are further executable to perform a dimensionality reduction step on the sets of values to reduce the number of variables in each of the sets of values.
6. The system of claim 5 , wherein the dimensionality reduction step identifies a subset of major proteins in each of the sets of values.
7. The system of claim 6 , wherein the training set includes only the subset of major proteins in each of the sets of values.
8. The system of any one of claims 1 to 7, wherein the set of values is labeled with the label.
9. The system of any one of claims 1 to 8, wherein each of the biological samples is one of plasma, whole blood, serum, cerebrospinal fluid (CSF), and peripheral blood mononuclear cells (PBMCs).
10. The system according to any one of claims 1 to 9, wherein the specified type of disease is a proliferative disease.
11. The system of claim 10 , wherein the proliferative disease is cancer.
12. The system of any one of claims 1 to 11, wherein the training set further comprises labels associated with clinical data for at least some of the subjects.
13. The system of any one of claims 1 to 12, wherein the prediction is represented as one of a set of binary values, continuous values, and discrete values.
14. The system of any one of claims 1 to 13, wherein the prediction comprises an indicator of secondary efficacy in the target patient.
15. 15. The system of any one of claims 1-14, wherein the program instructions are further executable to determine at least one of continuing the specified therapy in the target patient, adjusting the specified therapy in the target patient, discontinuing the specified therapy in the target patient, and administering a different therapy to the target patient based at least in part on the prediction.
16. The system according to any one of claims 1 to 15, wherein the specified treatment is immunotherapy.
17. The method of claim 16, wherein the immunotherapy is selected from an anti-PD-1 / PD-L1 therapy, an anti-CTLA-4 therapy, and both.
18. 1. A method for predicting a response in a target patient to a specified therapy, said method comprising, by at least one hardware processor, the following steps: receiving, for each of a plurality of subjects having a specified type of disease and receiving a specified therapy for treating the disease, (a) a first biological signature associated with a biological sample collected at a first time point associated with the specified therapy, and (b) a second biological signature associated with a biological sample collected at a second time point associated with the specified therapy; calculating, for each of the plurality of subjects, a set of values representing a relationship between the first biological signature and the second biological signature associated with the respective subject; During the training stage, (i) a set of calculated values; and (ii) a label associated with the outcome of the specified treatment in each of the subjects; and Train a machine learning model on a training set including thereby generating a classifier suitable for predicting a response in said target patient to said specified treatment; A method comprising:
19. 20. The method of claim 18, wherein the first biological signature and the second biological signature are each one of a DNA profile, an RNA profile, a protein profile, a metabolomic profile, a microbiome profile, a genomic profile, a transcriptomic profile, a cellular profile, an epigenomic profile, a post-translational modification-based profile, a single-cell-based analysis, and a regulatory RNA profile.
20. 20. The method of claim 18 or 19, wherein the first biological signature and the second biological signature are each protein expression profiles, and the sets of values each comprise, for each protein in the protein expression profiles, a ratio of the expression levels of the protein in the first biological signature and the second biological signature.
21. 21. The method of claim 20, wherein the protein expression profile comprises expression values of at least two proteins.
22. A method described in any one of claims 18 to 21, further comprising performing, by at least one hardware processor, a dimensionality reduction step on the set of values to reduce the number of variables in each of the set of values.
23. 23. The method of claim 22, wherein the dimensionality reduction step identifies a subset of major proteins in each of the sets of values.
24. 24. The method of claim 23, wherein the training set includes only the subset of major proteins in each of the sets of values.
25. The method of any one of claims 18 to 24, wherein the set of values is labelled with the label.
26. 26. The method of any one of claims 18 to 25, wherein each of the biological samples is one of plasma, whole blood, serum, cerebrospinal fluid (CSF), and peripheral blood mononuclear cells (PBMCs).
27. The method of any one of claims 18 to 26, wherein the specified type of disease is a proliferative disease.
28. 28. The method of claim 27, wherein the proliferative disease is cancer.
29. The method of any one of claims 18 to 28, wherein the training set further comprises labels associated with clinical data for at least some of the subjects.
30. The method of any one of claims 18 to 29, wherein the prediction is represented as one of a set of binary values, continuous values, and discrete values.
31. 31. The method of any one of claims 18-30, further comprising, during an inference stage, applying, by at least one hardware processor, the classifier to the target set of values associated with a target subject, thereby predicting a response in the target subject to the specified treatment.
32. 32. The method of claim 31 , further comprising determining, by at least one hardware processor, based at least in part on the prediction, at least one of: continuing the specified therapy in the target subject; adjusting the specified therapy in the target subject; discontinuing the specified therapy in the target subject; and administering a different therapy to the target subject.
33. The method of any one of claims 31 to 32, wherein the designated therapy is immunotherapy.
34. Adjusting the designated therapy or administering a different therapy to the target subject is performed by at least one hardware processor by the following steps: Determining differentially expressed proteins (DEPs) between responders and non-responders; determining one or more resistance-associated proteins (RAPs) selected from the DEPs in a sample obtained from the subject; identifying a treatment suitable for balancing the level of one or more RAPs in said subject; 33. The method of claim 32, wherein the determination is made by at least one hardware processor by executing:
35. The method of claim 32, wherein determining the one or more RAPs is performed by providing, by at least one hardware processor, a probabilistic measure of the distance of the expression level of the DEP from a defined group of samples selected from a responder group or a non-responder group.
36. 33. The method of claim 32, wherein determining the one or more RAPs is performed by at least one hardware processor by determining the expression distribution of each DEP in each of (i) a responder group and (ii) a non-responder group, fitting a probability density function for each group, and calculating, for each subject, the probability that the DEP is associated with one of the response groups based on the DEP expression of the subject.
37. 1. A computer program product including a non-transitory computer-readable storage medium having program instructions embodied therein, the program instructions comprising: receiving, for each of a plurality of subjects having a specified type of disease and receiving a specified therapy for treating the disease, (a) a first biological signature associated with a biological sample collected at a first time point associated with the specified therapy, and (b) a second biological signature associated with a biological sample collected at a second time point associated with the specified therapy; and calculating, for each of the plurality of subjects, a set of values representing a relationship between the first biological signature and the second biological signature associated with each of the subjects; During the training stage, (i) a set of calculated values; and (ii) a label associated with the outcome of the specified treatment in each of the subjects; and A computer program product executable by at least one hardware processor to generate a classifier suitable for predicting a response in a target patient to said specified treatment.
Citation Information
Patent Citations
Systems and methods for patient-specific prediction of drug responses from cell line genomics
JP2019016361A
Mutation-based disease diagnosis and tracking
JP2019509018A
Dasatinib response prediction models and methods therefor
US20180039732A1
Diagnosis tailoring of health and disease
US20190000349A1
Long non-coding RNA gene expression signatures in disease monitoring and treatment
US20190119730A1