Predicting patient response
Plasma proteomics and machine learning are used to calculate an overall resistance score for predicting patient response to ICI therapy, addressing the limitations of current biomarkers by integrating PD-L1 levels and plasma proteomic data for personalized treatment decisions.
Patent Information
- Application Number
- JP2025507666
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-09
- Filing Date
- 2023-08-10
- Publication Date
- 2025-09-09
AI Technical Summary
Current biomarkers for predicting patient response to immune checkpoint inhibitor (ICI) therapy in cancer, such as PD-L1 expression and tumor mutation burden (TMB), are not sufficiently accurate, failing to account for the complexity of tumor-immune system interactions and are limited by the need for tumor tissue samples.
A method using plasma proteomics and machine learning to calculate an overall resistance score from factor expression levels, integrating PD-L1 levels and plasma proteomic data to predict response to monotherapy or combination therapy in cancer patients.
Improves the prediction of patient response to ICI therapy by providing a comprehensive assessment of tumor microenvironment and immune cell dynamics, enabling personalized treatment decisions based on a single assay from circulating blood samples.
Smart Images

Figure 2025529765000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims the benefit of priority to International Patent Application No. PCT / IL2022 / 050881, filed August 11, 2022, U.S. Provisional Patent Application No. 63 / 423,551, filed November 8, 2022, U.S. Provisional Patent Application No. 63 / 442,174, filed January 31, 2023, and U.S. Provisional Patent Application No. 63 / 465,026, filed May 9, 2023, the entire contents of which are all incorporated herein by reference. [Background technology]
[0002] The present invention relates to the field of patient-specific diagnostics.
[0003] Background technology Immunotherapy based on immune checkpoint inhibitors (ICIs) represents a significant breakthrough in clinical oncology. ICIs enhance antitumor immune responses by targeting checkpoint proteins, such as PD-1, PD-L1, and CTLA-4, expressed on tumor and immune cells. While ICI therapy can achieve unprecedented long-term disease control across multiple tumor types, efficacy varies widely among patients, with the majority exhibiting primary or subsequent acquired resistance to treatment. In metastatic non-small cell lung cancer (NSCLC), for which ICI regimens are the standard of care, response rates range from 10% to 50%, depending on tumor PD-L1 expression and the type and choice of treatment. Identifying patients who may benefit from ICI therapy remains a major clinical challenge because available predictive biomarkers are not sufficiently accurate.
[0004] To date, tumor PD-L1 expression and tumor mutation burden (TMB) are the best biomarkers for predicting ICI response. Although immunohistochemistry to assess PD-L1 expression in tumor tissue is used as a companion diagnostic to inform treatment decisions for NSCLC, TMB is not yet routinely used. According to current guidelines for NSCLC patients lacking oncogenic driver mutations, patients with high tumor PD-L1 expression (defined as PD-L1 expression on at least 50% of tumor cells) are eligible for first-line ICI monotherapy, while ICI in combination with chemotherapy is the preferred option for patients with PD-L1 expression below 50%. However, clinical evidence demonstrates the limitations of PD-L1 biomarkers in predicting ICI response. For example, in the KEYNOTE-024 trial, approximately half of the patient cohort with high PD-L1 expression did not respond to pembrolizumab monotherapy. Furthermore, several clinical trials have reported clinical benefit from ICI therapy in some patients whose tumors have low PD-L1 expression. Notably, although both PD-L1 and TMB are involved in the mechanism of action of ICIs, these biomarkers do not account for the complexity of tumor-immune system interactions and the heterogeneous mechanisms underlying response and resistance to ICI therapy. Furthermore, PD-L1 testing requires tumor tissue, which may not be available.
[0005] To address this, a more comprehensive characterization of tumors, the tumor microenvironment (TME), peripheral immune cells, and other host factors is needed. Indeed, a growing number of predictive biomarkers for ICI outcomes are based on tumor genomic features and expression patterns, the abundance and phenotype of tumor-infiltrating lymphocytes in the TME, the dynamics of peripheral T cells, and other immune cell characteristics. Importantly, integrative models combining multiple biomarkers show promise for improving predictive capabilities, likely by better capturing the multifaceted nature of treatment efficacy. For example, combining PD-L1 and TMB biomarkers improves the prediction of response to ICI therapy in lung cancer patients. Furthermore, several studies have demonstrated improved prediction of ICI outcomes using integrated genomic, transcriptomic, and immune repertoire data. Although such models are promising, they are limited in that they are based on multiple assays and usually require tumor tissue specimens.
[0006] Plasma proteomics is a promising strategy for predictive biomarker discovery. Circulating blood contains thousands of proteins derived from developing tumors, the TME, peripheral immune cells, and other host cells. Therefore, the plasma proteome reflects tumor-intrinsic characteristics, immune cell dynamics, angiogenesis, extracellular matrix remodeling, and metabolic changes, making it a rich source of promising biomarkers that can be sampled minimally invasively and measured by a single assay. There is a great need for methods to determine patient-specific responses to immunotherapy that integrate PD-L1 levels and plasma proteomic data. Summary of the Invention
[0007] The present invention provides a method for predicting the response of a subject with a PD-L1 high, low, or negative cancer to a monotherapy or combination therapy, the method comprising calculating a resistance score for factors expressed by the subject and combining the resistance scores to generate an overall resistance score, wherein an overall resistance score above a predetermined threshold indicates that the subject is predicted to be resistant to the monotherapy or combination therapy.
[0008] In a first aspect, there is provided a method of predicting the response of a subject suffering from a PD-L1 high cancer to a monotherapy comprising anti-PD-1 / PD-L1 immunotherapy, the method comprising: a. i. In a population of subjects with cancer known to respond to the monotherapy (responders); ii. In a population of subjects known to be suffering from cancer and who do not respond to the above monotherapy (non-responders); and iii. In the above subject matter; receiving factor expression levels for a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, the resistance score comprising applying a machine learning algorithm trained on a training set including expression levels of the received factors in the responders and non-responders and the gender of each of the responders and non-responders to the expression levels of each received factor from the subject and the gender of the subject, the machine learning algorithm outputting the resistance score; c. combining the calculated resistance scores to generate an overall resistance score; Including, subjects with a composite resistance score above a predetermined threshold are predicted to not respond to the monotherapy, and subjects with a composite resistance score within the predetermined threshold are predicted to respond to the monotherapy; This predicts the subject's response to monotherapy.
[0009] In some embodiments, the overall resistance score is converted to an overall response score by the formula (1-overall resistance score), where an overall response score above a predetermined threshold indicates that the subject will respond to the monotherapy, and an overall response score below the predetermined threshold indicates that the subject will not respond to the monotherapy.
[0010] In some embodiments, the overall resistance score is converted to an overall response score by the formula (10-overall resistance score), where an overall response score above a predetermined threshold indicates that the subject will respond to the monotherapy, and an overall response score below the predetermined threshold indicates that the subject will not respond to the monotherapy.
[0011] In some embodiments, the training set further includes expression levels of the factors received in subjects with cancer who are known to respond to combination therapy comprising anti-PD-1 / PD-L1 immunotherapy and chemotherapy (combo-responders), and expression levels of the factors received in subjects with cancer who are known to not respond to combination therapy (combo-non-responders), and the gender of each of the combo-responders and combo-non-responders.
[0012] In another aspect, a method is provided for predicting the response of a subject suffering from a PD-L1 low or negative cancer to a combination therapy comprising anti-PD-1 / PD-L1 immunotherapy and chemotherapy, the method comprising: a. i. In a population of subjects (responders) who are suffering from cancer and known to respond to the combination therapy; ii. In a population of subjects known to be suffering from cancer and who do not respond to the combination therapy (non-responders); and iii. In the above subject matter; receiving factor expression levels for a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, the resistance score comprising applying a machine learning algorithm trained on a training set including expression levels of the received factors in the responders and non-responders and the gender of each of the responders and non-responders to the expression levels of each received factor from the subject and the gender of the subject, the machine learning algorithm outputting the resistance score; c. combining the calculated resistance scores to generate an overall resistance score; Including, subjects with a composite resistance score above a predetermined threshold are predicted to not respond to the combination therapy, and subjects with a composite resistance score within the predetermined threshold are predicted to respond to the combination therapy; This predicts the subject's response to the combination therapy.
[0013] In some embodiments, the overall resistance score is converted to an overall response score by the formula (1-overall resistance score), where an overall response score above a predetermined threshold indicates that the subject will respond to the combination therapy, and an overall response score below the predetermined threshold indicates that the subject will not respond to the combination therapy.
[0014] In some embodiments, the overall resistance score is converted to an overall response score by the formula (10-overall resistance score), where an overall response score above a predetermined threshold indicates that the subject will respond to the combination therapy, and an overall response score below the predetermined threshold indicates that the subject will not respond to the combination therapy.
[0015] In another aspect, there is provided a method of predicting a response of a subject afflicted with cancer to anti-PD-1 / PD-L1 immunotherapy, the method comprising: a. i. In a population of subjects (responders) who are suffering from cancer and known to respond to said immunotherapy; ii. In a population of subjects known to be suffering from cancer and not responding to the immunotherapy (non-responders); and iii. In the above subject receiving factor expression levels for a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, the resistance score comprising applying a machine learning algorithm trained on a training set including expression levels of the received factors in the responders and non-responders and the gender of each of the responders and non-responders to the expression levels of each received factor from the subject and the gender of the subject, the machine learning algorithm outputting the resistance score; c. Combining the calculated resistance scores to generate an overall resistance score. Including, Subjects with a composite resistance score above a predetermined threshold are predicted to not respond to this anti-PD-1 / PD-L1 immunotherapy; This predicts the subject's response to anti-PD-1 / PD-L1 immunotherapy.
[0016] In some embodiments, the training set further comprises expression levels of the factor received in subjects with cancer known to respond to monotherapy comprising anti-PD-1 / PD-L1 immunotherapy (mono-responders) and in subjects with cancer known to not respond to monotherapy (mono-non-responders), and the gender of each of the mono- and mono-non-responders.
[0017] In some embodiments, the overall resistance score is converted to an overall response score by the formula (1-overall resistance score), where an overall response score above a predetermined threshold indicates that the subject will respond to the immunotherapy and an overall response score below the predetermined threshold indicates that the subject will not respond to the immunotherapy.
[0018] In some embodiments, the overall resistance score is converted to an overall response score by the formula (10-overall resistance score), where an overall response score above a predetermined threshold indicates that the subject will respond to the immunotherapy and an overall response score below the predetermined threshold indicates that the subject will not respond to the immunotherapy.
[0019] In some embodiments, the plurality of factors comprises at least two factors selected from the factors provided in Table 4.
[0020] In some embodiments, the plurality of factors consists of factors selected from Table 4.
[0021] In some embodiments, responders and non-responders are determined based on progression-free survival (PFS) at 1 year after initiation of monotherapy or combination therapy.
[0022] In some embodiments, the method includes, prior to (b), selecting a subset of the plurality of factors, the subset including factors that best distinguish between responders and non-responders, and the calculating is for each factor of the subset.
[0023] In some embodiments, the selecting comprises applying a statistical test to the expression levels of the received factors, optionally the statistical test is a Kolmogorov-Smirnov test.
[0024] In some embodiments, the subset consists of at least 50 factors.
[0025] In some embodiments, the expression level of the factor is from a time point prior to administration of anti-PD-1 / PD-L1 immunotherapy to the subject.
[0026] In some embodiments, the combining is averaging
[0027] In some embodiments, combining includes determining the total number of factors having a resistance score above a predetermined threshold and generating an overall resistance score proportional to the total number.
[0028] In some embodiments, the method further comprises performing a dimensionality reduction step on the plurality of factors to reduce the number of factors in the plurality of factors.
[0029] In some embodiments, the cancer is selected from hepatobiliary cancer, cervical cancer, genitourinary cancer, anogenital cancer, testicular cancer, prostate cancer, thyroid cancer, ovarian cancer, nervous system cancer, eye cancer, lung cancer, soft tissue cancer, bone cancer, pancreatic cancer, bladder cancer, skin cancer, intestinal cancer, liver cancer, rectal cancer, colorectal cancer, esophageal cancer, gastric cancer, gastroesophageal cancer, breast cancer, renal cancer, skin cancer, head and neck cancer, and leukemia and lymphoma.
[0030] In some embodiments, the cancer is selected from lung cancer, skin cancer, anogenital cancer, cervical cancer, and head and neck cancer.
[0031] In some embodiments, the cancer is non-small cell lung cancer (NSCLC).
[0032] In some embodiments, the cancer is a tyrosine kinase inhibitor-resistant cancer.
[0033] In some embodiments, the predetermined threshold is determined by cross-validation within a training set, or is the median score of the training set.
[0034] In some embodiments, the plurality of factors is at least 200 factors.
[0035] In some embodiments, the expression level of the factor is the expression level of the factor in a biological sample provided by the subject.
[0036] In some embodiments, the biological sample is selected from plasma, whole blood, serum, or peripheral blood mononuclear cells.
[0037] In some embodiments, the biological sample is plasma or serum.
[0038] In some embodiments, the method further comprises administering a monotherapy to a subject predicted to respond to the monotherapy, or administering a combination therapy comprising anti-PD-1 / PD-L1 immunotherapy and chemotherapy to a subject predicted not to respond to the monotherapy.
[0039] In some embodiments, the method further comprises administering the combination therapy to a subject predicted to respond to the combination therapy, or administering an alternative therapy to a subject predicted not to respond to the combination therapy.
[0040] In some embodiments, the method further comprises administering the anti-PD-1 / PD-L1 immunotherapy to the subject predicted to respond to the anti-PD-1 / PD-L1 immunotherapy, or administering an alternative therapy to the subject predicted not to respond to the anti-PD-1 / PD-L1 immunotherapy.
[0041] In some embodiments, the anti-PD-1 / PD-L1 immunotherapy is selected from pembrolizumab, nivolumab, durvalumab, and atezolizumab.
[0042] In some embodiments, the chemotherapy is selected from carboplatin, paclitaxel, nab-paclitaxel, pemetrexed, vinorelbine, and cisplatin.
[0043] In some embodiments, the combination therapy comprises: Carboplatin, durvalumab, and paclitaxel; b. Atezolizumab, bevacizumab, carboplatin, and paclitaxel; c. Carboplatin, nab-paclitaxel, and pembrolizumab; d. Carboplatin, nivolumab, and paclitaxel; e. Carboplatin, nivolumab, and pemetrexed; f. Carboplatin, paclitaxel, pembrolizumab; g. Carboplatin, paclitaxel, pembrolizumab, and radiation; h. Carboplatin, and pembrolizumab; i. Carboplatin, pembrolizumab, and pemetrexed; j. Carboplatin, pembrolizumab, and vinorelbine; and k. Cisplatin, pembrolizumab, and pemetrexed is selected from.
[0044] In some embodiments, predicting response comprises predicting overall survival.
[0045] In some embodiments, predicting response comprises predicting progression-free survival.
[0046] In some embodiments, progression-free survival is one year after initiation of monotherapy or combination therapy.
[0047] In some embodiments, progression-free survival is one year after initiation of immunotherapy.
[0048] In some embodiments, the subject has a PD-L1 negative cancer.
[0049] In some embodiments, a PD-L1 high cancer comprises at least 50% of cancer cells that are positive for surface expression of PD-L1, and a PD-L1 low or negative cancer comprises less than 50% of cancer cells that are positive for surface expression of PD-L1.
[0050] In some embodiments, a PD-L1 low or negative cancer is a PD-L1 negative cancer that contains less than 1% of cells that are positive for surface expression of PD-L1.
[0051] In some embodiments, the trained machine learning algorithm comprises: During the training stage, (i) factor expression levels of resistance-associated factors in samples from subjects with cancer known to be responsive to anti-PD-1 / PD-L1 immunotherapy and factor expression levels of resistance-associated factors in samples from subjects with said cancer known to be unresponsive to said anti-PD-1 / PD-L1 immunotherapy; (ii) at least one clinical parameter of the known responsive subject and the known non-responsive subject; and (iii) a marker associated with the responsiveness of a subject suffering from the cancer; to generate a trained machine learning algorithm, wherein the trained machine learning algorithm is trained to output a resistance score or a response score.
[0052] In some embodiments, the expression level of the resistance-associated factor and the at least one clinical parameter are labeled with a label.
[0053] In some embodiments, the predetermined threshold for the overall resistance score is 5, and a resistance score greater than 5 indicates that the subject is resistant to the treatment, or the overall resistance score is converted to a response score by the formula (10-overall resistance score), and a response score greater than the predetermined threshold indicates that the subject is responsive to the treatment, optionally wherein the predetermined threshold for the response score is 5.
[0054] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description set forth below. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description. [Brief explanation of the drawings]
[0055] [Figure 1A-C] Illustrative examples of protein expression distribution in responder and non-responder populations at the single protein level. (1A-1C) Computer-generated examples of protein expression distribution for responder and non-responder populations. Shown are examples of protein expression levels that can be considered RAPs (light gray dashed line) or not (dark gray dashed line) based on the expression distribution data of the populations. [Figure 2] Illustration of the RAP score of Equation 2 implemented in Algorithm 1. The RAP score was calculated using synthetic data, where responder and non-responder populations were generated by sampling from a normal distribution. Expression levels of the synthetic population are shown in a histogram, with the responder population in dark gray and the non-responder population in light gray. Taking these distributions into account, the RAP score was calculated for each expression level. The resulting RAP scores are plotted as a blue curve, and the values are shown on the secondary Y-axis on the right. [Figure 3A-C] RAP score threshold determination based on AUC as a function of RAP score. The AUC at each RAP score was calculated, and the peak of the resulting curve was determined as the threshold (dotted line) for determining a protein as RAP or not RAP. (3A-3B). (3A) Graph showing determination of RAP score threshold using mathematical methods for protein measurements at T1 and (3B) T0. (3C) Graph showing determination of RAP score threshold using machine learning methods. [Figure 4A-B] (4A) Bar graph showing the number of RAPs for each patient in the study cohort (n=30). Responders - light grey, non-responders - dark grey. (4B) ROC curve for RAP analysis. [Figure 5] Heatmap of cancer features significantly enriched in six patients. Enrichment analysis was based on Fisher's exact test (FDR<0.05). Next to each patient's identifier, the number of RAPs for that patient is shown in parentheses. [Figure 6]Protein-protein network of significant RAPs in the current cohort. The network is based on the STRING database. Each node (protein) is color-coded based on the associated cancer feature. The black circular box indicates a targetable RAP. The size of each node correlates with the number of patients with the RAP studied. The compartment of each node is shown in the center (based on the Human Protein Atlas). I, intracellular. M, membrane. S, soluble. A protein may have multiple compartments. [Figure 7] Graph showing significantly enriched cancer features among 19 RAPs. Analysis was performed using Fisher's exact test. An enrichment factor greater than 1 indicates enrichment. [Figure 8] Heatmap of protein expression levels of 19 RAPs in healthy tissues. Expression data are based on the Human Protein Atlas (HPA) database. [Figure 9] Heatmap of the percentage of medium-high staining in patients with different cancer types, including NSCLC. Expression data is based on the Human Protein Atlas (HPA) database. [Figure 10A-F]Clinical description of the 184 patients included in the analysis. (10A) Heatmap depicting patient clinical characteristics: response to treatment (ORR1, ORR2, 1-year DCB); percent of cells expressing PD-L1 in biopsy immunohistochemistry, a prognostic marker of response to treatment; treatment type: ICI alone or ICI plus chemotherapy; treatment choice: first-line indicates ICI treatment given as initial systemic treatment for NSCLC, while advanced-line indicates previous non-ICI treatment was administered before the current ICI treatment. Gender indicates patient gender at birth. Histology indicates lung cancer histology (ADC - adenocarcinoma, SCC - squamous cell carcinoma). (10B-10C) Violin plots of the correlation between patient age and response at each time point: (10B) ORR1 and (10C) ORR2. ORR1 and ORR2 are overall responses determined 3 and 6 months after treatment initiation, respectively. (10D-10E) Graphical representation of response groups in (10D)ORR1 and (10E)ORR2. NR = non-responder. R = responder (partial or complete responder). SD = stable disease (these are included in the responder group in this model). (10F) Graphical representation of the population separated into development and validation sets. [Figure 11A-B] Performance of the classification model. ROC AUC was calculated using the overall tolerance score along with the actual overall response assessment at 3-month ORR, 6-month ORR, and 1-year duration of clinical response (DCB) for both T0 and T1. Results at T0 are shown for (11A, top panel) the development set and (11A, bottom panel; 11B, top panel) the validation set. (11B, bottom panel). A similar classification model was created based on T1. [Figure 12A-B](12A) Patients sorted by response probability score (calculated as 1 - resistance score) based on protein levels at TO. The actual observed response, ORR at 3 months, is indicated by color for each patient. (12B) Dot plot of the agreement between the predicted response probability based on protein expression at TO and the observed response probability at either 3 months, 6 months, or 1 year. Each point on the graph represents a specific patient, and different time points are indicated by different hues and marker types. The black diagonal line represents the straight line y = x, and the red diagonal line represents the regression line fitted to all points; the goodness of fit of the regression (R2) is shown. The horizontal line represents the average observed response probability for the three time points (color-coded) across the validation set. [Figure 13A-B] Survival analysis based on predicted results of ORR at 3 months based on T0 protein measurements for (13A) PFS and (13B) OS. [Figures 14A-14B] (14A) Functional network of all potential RAPs from this analysis. Each node represents a RAP, and edges between nodes indicate functional relationships. Nodes with large size and provided protein names represent investigational new drugs (INDs) in combination with immunotherapy. Nodes are colored based on protein function. (14B) Functional network for two patients: a predicted non-responder (top) and a predicted responder (bottom). RAPs detected for each patient are outlined in black. The non-responder patient had 44 RAPs detected, and the responder had 10 RAPs detected. [Figure 15] Functional differences between the top RAPs in each response group. Each polygon in the Voronoi plot represents a RAP, and the size correlates with the difference between responders and non-responders. RAPs in non-responders are involved in splicing, signaling, and cytoskeleton-related processes, whereas RAPs in responders are primarily involved in proteolysis and cell adhesion. Each color indicates a different overall function. [Figure 16] Table listing clinical parameters of the 339 patients included in the analysis. [Figure 17]Line graphs of the number of patients at each time point are shown by response group (NR for non-responders, R for responders) and overall. The patient cohort was divided into a development set and a validation set. [Figure 18] Association of CB with clinical parameters at 3, 6, and 12 months. Clinical parameters examined were age, sex, histology, treatment type, PD-L1 status, and ECOG performance status. NSCC, non-squamous cell carcinoma; SCC, squamous cell carcinoma; CB, clinical benefit; NCB, no clinical benefit; ICI, immune checkpoint inhibitor. [Figure 19A-B] Performance of predictive models based on clinical parameters. (19A) Receiver operating characteristic (ROC) plot of the PD-L1-based predictive model. (19B) ROC plot of the predictive clinical model based on PD-L1, gender, ECOG, and treatment selection. Area under the curve (AUC) values for each time point are shown. CI, confidence interval. [Figure 20A-B](20A) Development of a RAP Prediction Model. A cohort of patients with advanced NSCLC receiving ICI-based therapy was recruited. Pretreatment blood samples were obtained, and plasma proteomes were profiled. Clinical benefit (CB) was assessed 3, 6, and 12 months after the start of treatment, and patients were followed for 2 years. Prediction models for CB at each time point were developed as follows: Proteins exhibiting differential plasma levels in the CB and NCB patient populations were selected for model training using statistical tests. These proteins were collectively referred to as resistance-associated proteins (RAPs). Prediction models for CB were developed for each RAP using a machine learning algorithm. The CB prediction values estimated from each RAP were summed to calculate a RAP score for each patient. The RAP score (total number of active RAPs) was linearly scaled from 0 to 1, allowing for the conversion of a given patient's RAP score to a CB probability. (20B) Development and Validation of the RAP Model. The cohort was divided into a development set and a validation set (75% and 25%, respectively). The development set was randomly divided into a training set and a test set (75% and 25%, respectively). The training set was used to select a RAP, followed by model training, resulting in a predictive model for each RAP. A clinical benefit (CB) prediction was then made for each RAP for each patient in the test set. The CB predictions from all selected RAPs were summed to obtain a RAP score for each patient in the test set. This process was repeated 80 times, each time randomly dividing the patients in the development set into the training and test sets. The RAP scores were averaged for each patient in the development set and linearly scaled. The model output is a CB probability (value between 0 and 1). The model was then locked and tested on an independent validation set. [Figure 21] Effect of RAP number on model performance per time point. Different RAP numbers (ranging from 1 to 400) were selected. For each number, the model was run 10 times. Model performance was assessed by ROC analysis. AUC is shown. Based on this analysis, 50 was set as the cutoff for the number of RAPs selected. [Figures 22A-G]Identification of RAPs during model development. (22A) Histogram showing the number of identified RAPs grouped according to the number of times they were selected across 80 iterations. The top, middle, and bottom histograms are for the 3-, 6-, and 12-month time points, respectively. (22B) Total number of RAPs identified per time point. Some proteins measured by the SomaScan assay are redundant due to different aptamers binding to the same protein. The number of total and non-redundant RAPs is indicated by the open and dotted bars, respectively. The number of non-redundant RAPs identified at least 40 times across a total of 80 iterations is indicated by the solid bars. (22C) Venn diagram showing the number of RAPs identified per time point. (22D) Hierarchical clustering-based heat map showing the number of iterations in which the provided proteins were classified as RAPs. (22E) Cellular localization and potential cellular origin of RAPs. Data were obtained from the Human Protein Atlas. The same protein may be assigned to multiple cellular locations. (22F) Voronoi plot displaying the main biological functions of RAPs per time point. Each polygon represents a RAP, and the size correlates with the number of times the protein has been selected as a RAP. Proteins from the same KEGG biological process are grouped together (using the default settings of the Proteomaps tool). (22G) Enrichment analysis of RAPs per time point. Enrichment analysis was performed on RAPs selected with at least 10 replicates. Fisher's exact test (FDR<0.1) was used. [Figures 23A-E]Performance of the RAP prediction model. (23A) Bar graph showing predicted clinical benefit (CB) probabilities sorted from minimum to maximum. Patients with observed CB and no CB (NCB) are shown as light blue and dark blue bars, respectively. (23B) Overall survival analysis of patients stratified into high and low CB probability groups. The median CB probability per time point was used as the stratification threshold. HR, hazard ratio. CI, confidence interval. (23C) Progression-free survival analysis of patients stratified into high and low CB probability groups. The median CB probability per time point was used as the stratification threshold. HR, hazard ratio. CI, confidence interval. (23D) Predicted CB probability as a function of observed CB rate. Each point represents a patient. The observed CB rate for each predicted CB probability data point represents the proportion of observed CB patients within the patient group assigned a CB probability ±0.15. X=Y is shown as a black line. Goodness of fit is indicated. (23E) Receiver operating characteristic (ROC) plot for the RAP model by time point. Area under the curve (AUC) is shown. Dashed line indicates AUC=0.5. CI, confidence interval. [Figure 24] Enrichment analysis for CB probability and observed CB rate at each time point. Enrichment analysis was performed using the 2D-enrichment test. The X-axis shows the enrichment score for predicted CB probability. The Y-axis shows the enrichment score for observed rate (defined by the proportion of observed CB patients within the patient group assigned a CB probability ±0.15). Enrichment scores range from 1 to -1. Positive and negative enrichment scores indicate enrichment for high and low CB probability, and high and low observed CB rate, respectively. The solid line indicates the X=Y line. [Figure 25A-D](25A-25C) Comparison between CB probabilities at consecutive time points. Each point represents a patient in the cohort. CB probability at one time point is plotted against CB probability at subsequent time points. Color indicates the patient's CB signature at each time point and whether the signature of clinical benefit changed between time points. (25A) Comparison between 3 and 6 months. (25B) Comparison between 3 and 12 months. (25C) Comparison between 6 and 12 months. (25D) Sankey plot showing the progression of CB signature over time. CB, clinical benefit; NCB, no clinical benefit; NA, not available. [Figure 26A-B] The RAP model outperforms the PD-L1 model and the clinical parameter-based model. Predictive performance was compared among five models: the RAP model; the PD-L1-based model (PD-L1); the clinical model (CM); the combined RAP+PD-L1; and the combined RAP+CM. (26A) Receiver operating characteristic (ROC) plots of the five models at each time point. The area under the curve (AUC) is shown. The dashed line indicates AUC=0.5. CI, confidence interval. (26B) Forest plots comparing the five models. Top, Cox regression analysis based on overall survival (OS) data. Bottom, Cox regression analysis based on progression-free survival (PFS) data. [Figure 27] Performance of the RAP model in different patient subsets: NSCC, non-squamous cell carcinoma, SCC, and squamous cell carcinoma. [Figure 28] Kaplan-Meier plots for patients with high, low, and negative PD-L1 expression in the entire cohort. Left: overall survival (OS). Right: progression-free survival (PFS). Dashed lines indicate median survival. [Figure 29] The RAP model predicts differential survival outcomes in patients with PD-L1 expression ≥ 50%. PD-L1-high patients were stratified into high (left) and low (right) CB probability groups using the median CB probability of the cohort as the stratification threshold. Overall survival (OS; lower panel) and progression-free survival (PFS; upper panel) were assessed in patients receiving ICI-chemotherapy combination therapy versus ICI monotherapy. The dashed line indicates median survival. [Figure 30]The RAP model predicts differential survival outcomes in patients with PD-L1 <50%. PD-L1-low and PD-L1-negative patients (PD-L1-low-negative) were stratified into high (left) and low (right) CB probability groups using the median CB probability of the cohort as the stratification threshold. Overall survival (OS; lower panel) and progression-free survival (PFS; upper panel) were assessed in patients receiving ICI-chemotherapy combination therapy versus ICI monotherapy. The dashed line indicates median survival. [Figure 31A-C](31A) Patient Clinical Data. (31B) Development of the PROphet Prediction Model. A cohort of patients with advanced NSCLC receiving ICI-based therapy was recruited. Pretreatment blood samples were collected, and plasma proteomes were profiled using SomaScan technology. Clinical benefit (CB) was assessed 12 months after treatment initiation, and patients were followed for 2 years. A predictive model for CB was developed as follows: Proteins showing differential plasma concentrations in the CB and NCB patient populations were selected for model training using statistical tests. These proteins are collectively referred to as resistance-associated proteins (RAPs). A predictive model for CB was created for each RAP using a machine learning algorithm. The CB prediction values estimated from each RAP were summed to obtain a RAP score for each patient. The RAP score was linearly scaled from 0 to 1, allowing for the conversion of a given patient's RAP score into a CB probability, which determines the PROphet outcome as negative or positive on a scale of 0 and 10. (31C) Development and Validation of the RAP Model. The cohort was divided into a development set and a validation set. The development set was randomly divided into a training set and a test set (75% and 25%, respectively). The training set was used to select a RAP, followed by model training, resulting in a predictive model for each RAP. A clinical benefit (CB) prediction was then generated for each RAP for each patient in the test set. The CB predictions from all selected RAPs were summed to obtain a RAP score for each patient in the test set. This process was repeated 80 times, each time randomly dividing the patients in the development set into the training and test sets. The RAP scores were averaged for each patient in the development set and linearly scaled. The model output was a CB probability (value between 0 and 1), which was translated into a ROphet score. The model was then locked and tested on an independent validation set. [Fig. 32A-D]PROphet predicts overall survival for patients receiving ICI-based therapy and outperforms PD-L1-based predictions. (32A) Kaplan-Meier plot of PD-L1 ≥ 50% vs. PD-L1 < 50%. (32B) Predicted CB probability based on PD-L1 prediction as a function of observed CB rate. Each point represents a patient. The observed CB rate for each data point in predicted CB probability represents the observed proportion of CB patients within the patient group assigned a CB probability of ± 0.04. The goodness of fit (R2) is shown. (32C) Kaplan-Meier plot of patient stratification based on the PROphet model. (32D) Predicted CB probability based on the PROphet model as a function of observed CB rate. The observed CB rate for each data point in predicted CB probability represents the observed proportion of CB patients within the patient group assigned a CB probability of ± 0.05. [Figure 33A-B] PROphet is not predictive for chemotherapy patients. (33A) Kaplan-Meier plot for chemotherapy patients classified as PROphet positive or negative, with no significant difference between these two subgroups. (33B) Predicted CB probability based on the PROphet model as a function of observed CB rate. Each point represents a patient. The observed CB rate for each data point in predicted CB probability represents the proportion of observed CB patients within the patient group assigned a CB probability ±0.05. The goodness of fit (R2) is shown. [Figure 34] Flowchart of patients included in the PROphet+PD-L1 analysis. [Fig. 35A-H]The PROphet model, when combined with PD-L1 expression levels, predicts differential overall survival outcomes among different subgroups. (35A-35C) Kaplan-Meier plots for PROphet positivity prediction in patients with PD-L1 ≥ 50% (35A), PD-L1 1-49% (35B), and PD-L1 < 1% (35C). In 35A, patients with PD-L1 ≥ 50% received ICI-chemotherapy combination therapy or ICI monotherapy. In 35B and 35C, patients with PD-L1 1-49% and PD-L1 < 1% who received ICI-chemotherapy combination therapy were compared with patients who received chemotherapy alone. (35D-35F) Kaplan-Meier plots for PROphet®-negative prediction in patients with PD-L1 ≥ 50% (35D), PD-L1 1-49% (35E), and PD-L1 < 1% (35F). In 35D, patients with PD-L1 ≥ 50% received ICI-chemotherapy combination therapy or ICI monotherapy. In 35E and 35F, patients with PD-L1 1-49% and PD-L1 < 1% who received ICI-chemotherapy combination therapy were compared with patients who received chemotherapy alone. HR, hazard ratio. CI, confidence interval. (35G-35H) Kaplan-Meier plots for PROphet-positive patients (35G) or PROphet-negative patients (35H) with PD-L1 expression levels of 1-49%. ICI-chemotherapy combination therapy was compared with ICI monotherapy or chemotherapy monotherapy. HR, hazard ratio. CI, confidence interval. [Figure 36A-F]The PROphet model, when combined with PD-L1 expression levels, predicts differential progression-free survival between different subgroups. (36A-36C) Kaplan-Meier plots for PROphet positivity prediction in patients with PD-L1 ≥ 50% (36A), PD-L1 1-49% (36B), and PD-L1 < 1% (36C). In 36A, patients with PD-L1 ≥ 50% received ICI-chemotherapy combination therapy or ICI monotherapy. In 36B and 36C, patients with PD-L1 1-49% and PD-L1 < 1% who received ICI-chemotherapy combination therapy were compared with patients who received chemotherapy alone. (36D-36F) Kaplan-Meier plots for PROphet®-negative prediction in patients with PD-L1 ≥ 50% (36D), PD-L1 1-49% (36E), and PD-L1 < 1% (36F). In 36D, patients with PD-L1 ≥ 50% received ICI-chemotherapy combination therapy or ICI monotherapy. In 36E and 36F, patients with PD-L1 1-49% and PD-L1 < 1% who received ICI-chemotherapy combination therapy were compared with patients who received chemotherapy monotherapy. HR, hazard ratio. CI, confidence interval. [Figure 37A-C] Forest plot of multivariate analysis of the PROphet model. (37A) Patients with PD-L1 ≥ 50%. (37B) Patients with PD-L1 1-49%. (37C) Patients with PD-L1 < 1%. [Fig. 38A-D] Comparison between PROphet positive and negative results. (38A-38B) Comparison for patients with PD-L1≧50%. (38A) Patients receiving ICI monotherapy. (38B) Patients receiving ICI-chemotherapy combination therapy. (38C) Comparison for patients with PD-L1 1-49% receiving ICI-chemotherapy combination therapy. (38D) Comparison for patients with PD-L1<1% receiving ICI-chemotherapy combination therapy. [Figure 39A-H]Applicability of response prediction using the PROphet model in patients with melanoma, SCLC, and HPV-related malignancies. (39A-C) PROphet model predictions in the melanoma cohort. (39A) Model ROC AUC for 1-year sustainable clinical benefit. (39B) Predicted vs. observed response probability based on the PROphet model. Each point represents a specific patient. (39C) Kaplan-Meier plot for PROphet-positive and -negative patients. (39D) Kaplan-Meier curves showing survival for PROphet-positive and PROphet-negative patients with SCLC. (39E-G) Kaplan-Meier curves showing survival for PROphet-positive and PROphet-negative patients with HPV-related malignancies, including (39E) anogenital SCC, (39F) cervical cancer, and (39G) head and neck cancer. (39H) Kaplan-Meier curves for all HPV-related malignancies. Hazard ratios and p-values with 95% confidence intervals are shown. [Figure 40A-B] Applicability of response prediction using the PROphet model for NSCLC patients with targetable mutations. (40A) Correlation between PROphet score and overall survival for NSCLC patients with targetable mutations treated with PD-1 inhibitors. R2=0.41, p=0.0073. (40B) Kaplan-Meier curves showing survival for PROphet-positive and PROphet-negative patients, HR=0.36, p=0.07. DETAILED DESCRIPTION OF THE INVENTION
[0056] In some embodiments, the present invention provides methods for predicting the response of subjects with tumors containing high, low, or negative levels of PD-L1 to immunotherapy. Here, the inventors developed a novel and inherently robust machine learning (ML)-based model that analyzes pretreatment plasma proteomic profiles to predict benefit from ICI therapy in cancer patients. By integrating predictions from multiple proteomic biomarkers, the model accurately predicts clinical benefit at three time points along the treatment course, stratifies patients by survival outcome or PFS, and outperforms PD-L1-based predictions. Furthermore, the model shows the potential to further optimize treatment selection when used in conjunction with PD-L1 classification. Overall, the model provides clinically valuable information to aid in cancer treatment decisions.
[0057] The present invention is based, at least in part, on the discovery of a novel tool to support treatment decisions for cancer patients receiving ICI-based therapy. The RAP (PROphet) model provides two major clinical benefits. First, it successfully predicts therapeutic benefit at 12 months, demonstrating superior predictive ability to PD-L1-based models. Second, when used in combination with PD-L1 testing, the model helps determine whether patients should receive ICI therapy alone or an ICI-chemotherapy combination. Specifically, subjects with high PD-L1 levels and a high overall response score are predicted to respond to ICI therapy as monotherapy and need not be exposed to the adverse side effects of chemotherapy. Patients with high PD-L1 but a low overall response score are advised to proceed with ICI and chemotherapy combination therapy. Patients with low PD-L1 but a high overall response score are predicted to benefit from treatment with an ICI-chemotherapy combination, while patients with low PD-L1 and a low overall response score are advised to consider alternative therapies.
[0058] According to a first aspect, there is provided a method of predicting the response of a subject with a PD-L1 high cancer to a monotherapy, including immunotherapy, said method comprising: a. i. In a population of subjects known to respond to said treatment (responders); ii. In a population of subjects known to not respond to the above treatment (non-responders); and iii. In the above subject receiving expression levels for a plurality of factors; b. calculating a resilience score for a factor of the plurality of factors, the resilience score comprising applying a machine learning algorithm, the machine learning algorithm outputting the resilience score; c. combining the calculated resistance scores to generate an overall resistance score; Subjects with a composite resistance score above a predetermined threshold are predicted to not respond to said monotherapy, and subjects with a composite resistance score within a predetermined threshold are predicted to respond to said monotherapy; This predicts the subject's response to monotherapy.
[0059] According to another aspect, there is provided a method of predicting the response of a subject suffering from a PD-L1 low or negative cancer to a combination therapy comprising immunotherapy and chemotherapy, said method comprising: a. i. In a population of subjects known to respond to said treatment (responders); ii. In a population of subjects known to not respond to the above treatment (non-responders); and iii. In the above subject receiving expression levels for a plurality of factors; b. calculating a resilience score for a factor of the plurality of factors, the resilience score comprising applying a machine learning algorithm, the machine learning algorithm outputting the resilience score; c. combining the calculated resistance scores to generate an overall resistance score; Including, Subjects with a composite resistance score above a predetermined threshold are predicted to not respond to the combination therapy, and subjects with a composite resistance score within a predetermined threshold are predicted to respond to the combination therapy; This will predict the subject's response to the combination therapy.
[0060] According to another aspect, there is provided a method of predicting a subject's response to a treatment, said method comprising: a. i. In a population of subjects known to respond to said treatment (responders); ii. In a population of subjects known to not respond to the above treatment (non-responders); and iii. In the above subject receiving expression levels for a plurality of factors; b. calculating a tolerance score for at least one factor of said plurality of factors; c. Factors with a resistance score above the threshold were identified as being associated with resistance. Classifying as a factor Including, Subjects with a number of resistance-associated factors above a predetermined number are predicted to be resistant to the treatment, thereby predicting the subject's response to the treatment.
[0061] According to another aspect, there is provided a method of predicting a subject's response to a treatment, said method comprising: a. i. In a population of subjects known to respond to said treatment (responders); ii. In a population of subjects known to not respond to the above treatment (non-responders); and iii. In the above subject receiving expression levels for a plurality of factors; b. calculating a tolerance score for at least one factor of said plurality of factors; c. classifying factors with a resistance score above a threshold as factors associated with resistance; d. Summarizing the number of factors associated with resistance; and e. applying a trained machine learning algorithm to the number of factors and at least one clinical parameter associated with resistance, wherein the trained machine learning algorithm outputs an overall resistance score, wherein an overall resistance score above a predetermined threshold indicates that the subject is resistant to the treatment. thereby predicting the subject's response to treatment.
[0062] According to another aspect, there is provided a method of predicting a subject's response to a treatment, said method comprising: a. i. In a population of subjects known to respond to said treatment (responders); ii. In a population of subjects known to not respond to the above treatment (non-responders); and iii. In the above subject receiving expression levels for a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, wherein the resistance score is based on the similarity of the expression level of the factor in the subject to the expression level in the responder, and the similarity of the expression level of the factor in the subject to the expression level in the non-responder, and includes applying a trained machine learning algorithm that outputs a resistance score; c. Summing the calculated resistance scores to generate an overall resistance score Including, Subjects with a total resistance score above a predetermined threshold are predicted to be resistant to the treatment; This predicts the subject's response to treatment.
[0063] According to another aspect, in the training phase, (i) a number of resistance-associated factors expressed in samples from subjects known to have the disease and to respond to treatment, and from subjects known to have the disease and to not respond to treatment; (ii) at least one clinical parameter of the subject; and (iii) a label associated with the subject's response;
[0013] A method is provided that includes training a machine learning algorithm with a training set that includes:
[0064] According to another aspect, in the training phase, (i) factor expression levels of resistance-associated factors in samples from subjects known to have the disease and to respond to treatment, and from subjects known to have the disease and to not respond to treatment; (ii) at least one clinical parameter of the subject; and (iii) a label associated with the subject's response;
[0013] A method is provided that includes training a machine learning algorithm with a training set that includes:
[0065] In some embodiments, the method is a diagnostic method. In some embodiments, the method is an in vitro method. In some embodiments, the method is an ex vivo method. In some embodiments, the method is a computer-implemented method. In some embodiments, the method is a statistical method. In some embodiments, the method is a method that cannot be performed by the human mind. In some embodiments, the method is a computerized method. In some embodiments, the processor is a computer processor. In some embodiments, the processor is a computer.
[0066] In some embodiments, the method is for predicting a response to a treatment. In some embodiments, the method is for determining a response to a treatment. In some embodiments, the method is for determining a response score. In some embodiments, the method is for determining a response probability. In some embodiments, the response probability is a response score. In some embodiments, the method is for determining a probability of clinical benefit. In some embodiments, the method is for determining overall survival. In some embodiments, the method is for determining progression-free survival (PFS). In some embodiments, the method is for determining overall survival (OS). In some embodiments, the method is for determining a survival probability. In some embodiments, determining is predicting. In some embodiments, a resistance score is determined. In other embodiments, a prediction of resistance probability is determined. In some other embodiments, a resistance probability of less than 20% indicates that the subject will respond to the treatment. In some other embodiments, a resistance probability of less than 50% indicates that the subject will respond to the treatment. In some embodiments, a response score is determined. In other embodiments, a prediction of response probability is determined. In some other embodiments, a response probability of greater than 80% indicates that the subject will respond to the treatment. In some other embodiments, a probability of response of greater than 50% indicates that the subject will respond to treatment. In some embodiments, beyond is above. In some embodiments, beyond is below. Those skilled in the art will understand that above / below depends on the construction of the scale, as the scale can be designed to measure in either direction.
[0067] In some embodiments, the method is for determining whether a subject is a responder to a treatment. In some embodiments, the method is for determining whether a subject is a non-responder to a treatment. In some embodiments, the method is for predicting a subject's response to a treatment. In some embodiments, the method is for monitoring a response to a treatment. In some embodiments, the method is for determining whether a treatment should be continued, adjusted (e.g., by further treating the subject with an additional therapy including, but not limited to, an agent determined by the RAP analysis provided herein below), or changed. In some embodiments, the method is for determining whether a subject is a responder or non-responder to a treatment. In some embodiments, the method is for determining whether a subject is a responder to a treatment, a non-responder to a treatment, or has stable disease. In some embodiments, the method is for predicting whether a subject will respond or not to a treatment. In some embodiments, the responder is a responder to a monotherapy (monoresponder). In some embodiments, the responder is a responder to a combination therapy (combo-responder). In some embodiments, the non-responder is a non-responder to a monotherapy (mono-non-responder). In some embodiments, the non-responder is a non-responder to a combination therapy (combo-non-responder). In some embodiments, the method is for determining whether the subject will benefit from the treatment.
[0068] In some embodiments, non-response comprises progressive disease. In some embodiments, non-response comprises cancer progression. In some embodiments, non-response comprises stable disease. In some embodiments, non-response comprises worsening of disease symptoms. In some embodiments, non-response is not the onset of side effects. In some embodiments, non-response comprises cancer growth, metastasis, and / or continued growth. In some embodiments, non-response comprises lack of clinical benefit (NCB). In some embodiments, non-response is non-survival. In some embodiments, non-response is non-survival and / or cancer progression. In some embodiments, response is stable disease. In some embodiments, response comprises remission. In some embodiments, remission is minimal remission. In some embodiments, remission is partial remission. In some embodiments, remission is complete remission. In some embodiments, response is survival. In some embodiments, response is progression-free survival. In some embodiments, response is long-term progression-free survival. In some embodiments, response is measured using overall response rate (ORR). A trained physician will be familiar with methods for determining response, particularly ORR. In some embodiments, response is measured using RECIST (Response Evaluation Criteria In Solid Tumors). In some embodiments, response includes survival. In some embodiments, survival is overall survival. In some embodiments, survival is progression-free survival. In some embodiments, survival is overall survival. In some embodiments, response includes clinical benefit (CB). In some embodiments, response includes durable clinical benefit (DCB). In some embodiments, CB is DCB. In some embodiments, CB is PFS. In some embodiments, CB is PFS at 12 months after initiation of treatment. In some embodiments, CB is PFS at 7 months after initiation of treatment. In some embodiments, populations of subjects known to respond and known not to respond are determined based on PFS, and predicted response includes OS. In some embodiments, PFS is PFS at 12 months.In some embodiments, PFS is 7-month PFS. In some embodiments, PFS is 6-month PFS. In some embodiments, PFS is 3-month PFS. In some embodiments, OS is 12-month OS. In some embodiments, OS is 7-month OS. In some embodiments, OS is 6-month OS. In some embodiments, OS is 3-month OS. In some embodiments, no clinical benefit or non-clinical benefit is the absence of clinical benefit as described herein.
[0069] In some embodiments, the subject is a mammal. In some embodiments, the subject is human. In some embodiments, the subject has a disease. In some embodiments, the disease is treatable by therapy. In some embodiments, the disease is cancer. In some embodiments, the disease is treatable by an immune checkpoint inhibitor (ICI). In some embodiments, the cancer is a PD-L1 positive cancer. In some embodiments, the cancer is a PD-L1 high cancer. In some embodiments, the cancer is a PD-L1 low cancer. In some embodiments, the cancer is a PD-L1 negative cancer. In some embodiments, the cancer is a PD-L1 low or negative cancer. In some embodiments, the cancer is a solid cancer. In some embodiments, the cancer is a tumor. In some embodiments, the cancer is selected from hepatobiliary cancer, cervical cancer, genitourinary cancer (e.g., urothelial cancer), anogenital cancer, testicular cancer, prostate cancer, thyroid cancer, ovarian cancer, nervous system cancer, ocular cancer, lung cancer, soft tissue cancer, bone cancer, pancreatic cancer, bladder cancer, skin cancer, intestinal cancer, liver cancer, rectal cancer, colorectal cancer, esophageal cancer, gastric cancer, gastroesophageal junction cancer, breast cancer (e.g., triple-negative breast cancer), renal cancer (e.g., renal cell carcinoma), skin cancer, head and neck cancer, leukemia, and lymphoma. In some embodiments, the cancer is selected from skin cancer and lung cancer. In some embodiments, the cancer is skin cancer. In some embodiments, the cancer is lung cancer. In some embodiments, the skin cancer is melanoma. In some embodiments, the lung cancer is small cell lung cancer. In some embodiments, the lung cancer is non-small cell lung cancer. In some embodiments, the melanoma is unresectable melanoma. In some embodiments, the melanoma is metastatic melanoma. In some embodiments, the cancer is an HPV (human papillomavirus) positive cancer. In some embodiments, the cancer is an HPV-associated cancer. In some embodiments, the cancer is an anogenital cancer. In some embodiments, the anogenital cancer is an anogenital squamous cell carcinoma (SCC). In some embodiments, the anogenital cancer includes anal cancer, cervical cancer, penile cancer, vaginal cancer, and vulvar cancer.In some embodiments, the cancer is cervical cancer. In some embodiments, the cervical cancer is small cell cervical cancer. In some embodiments, the cancer is head and neck cancer. In some embodiments, the head and neck cancer is head and neck SCC (HNSCC). In some embodiments, the cancer is selected from lung cancer, skin cancer, anogenital cancer, cervical cancer, and head and neck cancer.
[0070] In some embodiments, the cancer is resistant to treatment. In some embodiments, the treatment is a non-immunotherapy. In some embodiments, the treatment is another treatment. In some embodiments, the treatment is a targeted therapy. In some embodiments, the treatment is not anti-PD-1 / L1 immunotherapy. In some embodiments, the cancer is resistant to targeted therapy. In some embodiments, the targeted therapy is a tyrosine kinase inhibitor (TKI). In some embodiments, the subject has previously been treated with a TKI. In some embodiments, the subject has been treated with a TKI and is found to be resistant to the TKI. In some embodiments, the method is a method of determining whether a subject resistant to targeted therapy will respond to PD-1 / L1 immunotherapy. In some embodiments, the subject has a TKI-resistant cancer. In some embodiments, the cancer is TKI-resistant NSCLC. In some embodiments, the cancer comprises a mutation in a tyrosine kinase receptor gene. In some embodiments, the tyrosine kinase receptor gene is selected from epidermal growth factor receptor (EGFR), anaplastic lymphoma kinase (ALK), and proto-oncogene tyrosine-protein kinase ROS (ROS1).
[0071] In some embodiments, the subject is naive to the treatment before the initial determination. In some embodiments, the subject has not received treatment before the initial determination. In some embodiments, the subject has previously received treatment. In some embodiments, the subject has previously been treated with a treatment other than the present treatment. In some embodiments, the subject is concurrently treated with a treatment other than the present treatment. In some embodiments, the other treatment is a TGFB-Trap fusion protein. In some embodiments, the other treatment is a tyrosine kinase inhibitor. In some embodiments, the subject is naive to any treatment. In some embodiments, the subject is naive to immunotherapy. In some embodiments, the treatment is a first-line treatment. In some embodiments, the treatment is an advanced-line treatment.
[0072] In some embodiments, the therapy is an anti-cancer therapy. In some embodiments, the anti-cancer therapy is radiation. In some embodiments, the anti-cancer therapy is chemotherapy. In some embodiments, the therapy is immunotherapy. In some embodiments, the anti-cancer therapy is immunotherapy. In some embodiments, the anti-cancer therapy is targeted therapy. In some embodiments, the anti-cancer therapy is selected from radiation, chemotherapy, immunotherapy, targeted therapy, hormone therapy, anti-angiogenic therapy and photodynamic therapy, hyperthermia, surgery, and combinations thereof. In some embodiments, the immunotherapy is selected from immune checkpoint inhibition, immune checkpoint modulation, immune checkpoint blockade, adoptive cell transfer therapy, oncolytic virotherapy, vaccine therapy, immune system modulation, and monoclonal antibody therapy. In some embodiments, the immunotherapy is selected from immune checkpoint inhibitors, immune checkpoint modulators, immune checkpoint blockers, adoptive cell transfer therapy, oncolytic virotherapy, therapeutic vaccines, immune system modulators, and monoclonal antibodies. In some embodiments, the immunotherapy is an immune checkpoint inhibitor. In some embodiments, the immunotherapy is an immune checkpoint blockade. In some embodiments, the targeted therapy is a tyrosine kinase inhibitor. In some embodiments, the targeted therapy is a TGFB-Trap fusion protein.
[0073] In some embodiments, immunotherapy is administered in combination with one or more conventional cancer therapies, including chemotherapy, targeted therapy, steroids, and radiation therapy. The combination of ICIs and chemotherapy / radiotherapy / targeted therapy is being studied in multiple clinical trials. Those skilled in the art will appreciate that the predictive proteins disclosed herein are predictive for immunotherapy as a monotherapy and as part of a combination therapy. In some embodiments, the treatment is a monotherapy. In some embodiments, the monotherapy includes immunotherapy. In some embodiments, the monotherapy consists of immunotherapy. In some embodiments, the monotherapy does not include chemotherapy. In some embodiments, the monotherapy is anti-PD-1 / PD-L1 immunotherapy. In some embodiments, the treatment is a combination therapy. In some embodiments, the combination therapy includes immunotherapy and another treatment. In some embodiments, the combination therapy includes immunotherapy and chemotherapy. In some embodiments, the combination therapy includes immunotherapy and targeted therapy. In some embodiments, the targeted therapy is a tyrosine kinase inhibitor. In some embodiments, the targeted therapy is an anti-transforming growth factor beta (TGFB) agent. In some embodiments, the TGFB agent is a TGFB trap fusion protein.TGFβ trap fusion proteins are well known in the art and are disclosed, for example, in Knudson et al., "M7824, a novel bifunctional anti-PD-L1 / TGFβ Trap fusion protein, promotes anti-tumor efficacy as monotherapy and in combination with a vaccine," Oncoimmunolog. 2018 Feb 14;7(5):e1426519 and Morris et al., "Bintrafusp alfa, an anti-PD-L1:TGF-β trap fusion protein, in patients with ctDNA-positive, liver-limited metastatic colorectal cancer," Cancer Res Commun. 2022 Sep;2(9):979-986, the contents of which are incorporated herein by reference in their entireties. In some embodiments, the combination therapy further comprises radiation. In some embodiments, the combination therapy further comprises a non-anti-PD-1 / PD-L1 immunotherapy. In some embodiments, the anti-PD-1 / PD-L1 immunotherapy is selected from pembrolizumab, nivolumab, durvalumab, and atezolizumab. In some embodiments, the anti-PD-1 / PD-L1 immunotherapy is selected from pembrolizumab, nivolumab, durvalumab, atezolizumab, and cemiplimab. In some embodiments, the immunotherapy comprises pembrolizumab. In some embodiments, the immunotherapy comprises nivolumab. In some embodiments, the immunotherapy comprises: In some embodiments, the immunotherapy comprises durvalumab. In some embodiments, the immunotherapy comprises atezolizumab. In some embodiments, the chemotherapy is selected from carboplatin, paclitaxel, nab-paclitaxel, pemetrexed, vinorelbine, and cisplatin. In some embodiments, the chemotherapy is selected from carboplatin, paclitaxel, nab-paclitaxel, pemetrexed, vinorelbine, cisplatin, dacarbazine, temozolomide, albumin-bound paclitaxel, and vinblastine. In some embodiments, the chemotherapy is carboplatin. In some embodiments, the chemotherapy is paclitaxel. In some embodiments, the chemotherapy is nab-paclitaxel. In some embodiments, the chemotherapy is pemetrexed. In some embodiments, the chemotherapy is vinorelbine. In some embodiments, the chemotherapy is cisplatin. In some embodiments, the combination therapy comprises carboplatin, durvalumab, and paclitaxel. In some embodiments, the combination therapy includes atezolizumab, bevacizumab, carboplatin, and paclitaxel. In some embodiments, the combination therapy includes carboplatin, nab-paclitaxel, and pembrolizumab. In some embodiments, the combination therapy includes carboplatin, nivolumab, and paclitaxel. In some embodiments, the combination therapy includes carboplatin, paclitaxel, and pembrolizumab. In some embodiments, the combination therapy includes carboplatin, nivolumab, and pemetrexed. In some embodiments, the combination therapy includes carboplatin, paclitaxel, pembrolizumab, and radiation. In some embodiments, the combination therapy includes carboplatin and pembrolizumab. In some embodiments, the combination therapy includes carboplatin, pembrolizumab, and pemetrexed. In some embodiments, the combination therapy comprises carboplatin, pembrolizumab, and vinorelbine. In some embodiments, the combination therapy comprises cisplatin, pembrolizumab, and pemetrexed. In some embodiments, the combination therapy comprises an anti-CTLA-4 antibody. In some embodiments, the CTLA-4 antibody is ipilimumab.In some embodiments, the CTLA-4 antibody is tremelimumab. In some embodiments, the combination therapy includes an anti-LAG3 antibody. In some embodiments, the LAG3 antibody is leratolimab. In some embodiments, the TKI is osimertinib, erlotinib, afatinib, gefitinib, dacomitinib, dacomitinib, amivantamab-vmjw, mobocertinib, sotrasib, adagrasib, alectinib, brigatinib, lorlatinib, ceritinib, crizotinib, entrectinib, dabrafenib, ceritinib, trametinib, vemurafenib, tepotinib, capmatinib, ceritinib, cemetinib, vemurafenib, tepotinib, capmatinib, cemet ... Selected from percatinib, pralsetinib, fam-trastuzumab, deruxtecan-nxki, ado-trastuzumab, emtansine, cabozantinib, ado-trastuzumab, emtansine, larotrectinib, alectinib, cetuximab, cobimetinib, encorafenib, binimetinib, lenvatinib, imatinib, dasatinib, nilotinib, and ripretinib.
[0074] The 2023 NCCN Guidelines provide the following list of treatments that may be used alone or in combination to treat NSCLC, melanoma, or SCLC: NSCLC-ICI: atezolizumab, pembrolizumab, durvalumab, nivolumab, ipilimumab, cemiplimab, cemiplimab-rwlc, and tremelimumab. TKIs: osimertinib, erlotinib, afatinib, gefitinib, dacomitinib, amivantamab-vmjw, mobocertinib, sotorasib, adagrasib, alectinib, brigatinib, lorlatinib, ceritinib, crizotinib, entrectinib, dabrafenib, ceritinib, trametinib, vemurafenib, tepotinib, capmatinib, selpercatinib, pralsetinib, fam-trastuzumab, deruxtecan-nxki, ado-trastuzumab emtansine, cabozantinib, ado-trastuzumab emtansine, larotrectinib, alectinib, and cetuximab. Anti-VEGF agents: ramucirumab and bevacizumab. Chemotherapy: carboplatin, paclitaxel, pemetrexed, gemcitabine, cisplatin, docetaxel, vinorelbine, etoposide, and albumin-bound paclitaxel. Melanoma-ICIs: nivolumab, pembrolizumab, ipilimumab, and leratlimab. Targeted therapy: dabrafenib, trametinib, vemurafenib, cobimetinib, encorafenib, binimetinib, and lenvatinib. KIT inhibitors: imatinib, dasatinib, nilotinib, and ripretinib. ROS1 fusion agents: crizotinib and entrectinib. NTRK fusion agents: larotrectinib and entrectinib. NRAS agents: binimetinib. Chemotherapy: dacarbazine, temozolomide, albumin-bound paclitaxel, carboplatin, paclitaxel, cisplatin, vinblastine, and dacarbazine. SCLC-Chemotherapy: cisplatin, etoposide, carboplatin, irinotecan, topotecan, lurbinectin, cyclophosphamide, doxorubicin, vincristine, docetaxel, gemcitabine, temozolomide, vinorelbine, bendamustine, platinum agents, and paclitaxel. ICI: atezolizumab, durvalumab, nivolumab, pembrolizumab, and ipilimumab.
[0075] In some embodiments, the immunotherapy is multiple immunotherapy. In some embodiments, the immunotherapy is immune checkpoint blockade. In some embodiments, the immunotherapy is immune checkpoint protein inhibition. In some embodiments, the immunotherapy is immune checkpoint protein modulation. In some embodiments, the immunotherapy comprises immune checkpoint inhibition. In some embodiments, the immunotherapy comprises immune checkpoint modulation. In some embodiments, the immune checkpoint blockade and / or immune checkpoint inhibition comprises administering an immune checkpoint inhibitor to the subject. In some embodiments, the inhibition comprises administering an immune checkpoint inhibitor. In some embodiments, the inhibitor is a blocking antibody. In some embodiments, the immunotherapy comprises immune checkpoint blockade. In some embodiments, the modulation comprises administering an immune checkpoint modulating agent. In some embodiments, the modulation of an immune checkpoint comprises administering an immune checkpoint modulating agent to the subject.
[0076] As used herein, the term "immune checkpoint inhibitor (ICI)" refers to a single ICI, a combination of ICIs, and a combination of an ICI with another cancer therapeutic. The ICI can be a monoclonal antibody, a bispecific antibody, a humanized antibody, a fully human antibody, a fusion protein, or a combination thereof, that is intended to block, inhibit, or modulate immune checkpoint proteins. In some embodiments, the immune checkpoint inhibitor is an immune checkpoint modulator. In some embodiments, the immune checkpoint inhibitor is an immune checkpoint blocker. In some embodiments, the immune checkpoint proteins are PD-1 (Programmed Death-1); PD-L1; PD-L2; CTLA-4 (Cytotoxic T Lymphocyte-associated protein 4); A2AR (Adenosine A2A Receptor), also known as ADORA2A; B7-H3 (also known as CD276); B7-H4 (also known as VTCN1); B7-H5; BTLA (B and T Lymphocyte Attenuator) (also known as CD272); IDO (Indoleamine 2,3-dioxygenase); KIR (Killer Cell Immunoglobulin-like Receptor); LAG-3 (Lymphocyte Activation Gene-3; TDO (tryptophan 2,3-dioxygenase); TIM-3 (T-cell immunoglobulin and mucin domain 3); VISTA (V-domain Ig suppressor of T-cell activation); NOX2 (nicotinamide adenine dinucleotide phosphate NADPH oxidase isoform 2); SIGLEC7 (sialic acid-binding immunoglobulin-type lectin 7), also known as CD328; SIGLEC9 (sialic acid-binding immunoglobulin-type lectin 9), also known as CD329; OX40 (tumor necrosis factor receptor superfamily, member 4), also known as CD134; and TIGIT. In some embodiments, the immune checkpoint protein is selected from PD-1, PD-L1, and PD-L2. In some embodiments, the immune checkpoint protein is selected from PD-1 and PD-L1. In some embodiments, the immune checkpoint protein is CTLA-4.In some embodiments, the immune checkpoint protein is PD-1. In some embodiments, the immune checkpoint blockade comprises anti-PD-1 / PD-L1 / PD-L2 immunotherapy. In some embodiments, the immune checkpoint blockade comprises anti-PD-1 immunotherapy. In some embodiments, the immune checkpoint blockade comprises anti-PD-1 immunotherapy and / or anti-PD-L1 immunotherapy. In some embodiments, the immune checkpoint blockade comprises anti-CTLA-4 immunotherapy. In some embodiments, the immune checkpoint blockade comprises anti-PD-1 immunotherapy and / or anti-PD-L1 immunotherapy and anti-CTLA-4 immunotherapy. In some embodiments, the immunotherapy is anti-PD-1 / PD-L1 immunotherapy. In some embodiments, the immunotherapy is anti-PD-1 / PD-L1 axis immunotherapy. In some embodiments, the immune checkpoint blockade comprises anti-LAG-3. In some embodiments, the immune checkpoint blockade comprises anti-PD-1 immunotherapy and / or anti-PD-L1 immunotherapy and anti-LAG-3 immunotherapy.
[0077] In some embodiments, the resistance-associated factor is a. i. In a population of subjects known to respond to said treatment (responders); ii. In a population of subjects known to not respond to the above treatment (non-responders); and iii. In the above subject receiving expression levels for a plurality of factors; b. calculating a tolerance score for at least one factor of said plurality of factors; c. Classifying factors with resistance scores above a threshold as factors associated with resistance. The method is determined by a method including:
[0078] In some embodiments, the resistance-associated factor is present in each subject. In some embodiments, the resistance-associated factor is present in responders. In some embodiments, the resistance-associated factor is present in non-responders. In some embodiments, the resistance-associated factor is labeled with a label. In some embodiments, the expression level of the resistance-associated factor is labeled with a label. In some embodiments, the resistance-associated factor is a resistance-associated protein.
[0079] In some embodiments, the immunotherapy is a blocking antibody. In some embodiments, the immunotherapy is the administration of a blocking antibody to a subject.
[0080] In some embodiments, the ICI is a monoclonal antibody (mAb) against PD-1 or PD-L1. In some embodiments, the ICI is a mAb that neutralizes / blocks / inhibits / modulates the PD-1 pathway. In some embodiments, the ICI is a mAb against PD-1. In some embodiments, the anti-PD-1 mAb is pembrolizumab (Keytruda; formerly known as lambrolizumab). In some embodiments, the anti-PD-1 mAb is nivolumab (Opdivo). In some embodiments, the anti-PD-1 mAb is pidilizumab (CT0011). In some embodiments, the anti-PD-1 mAb is cemiplimab (Libtayo, REGN2810). In some embodiments, the anti-PD-1 mAb is any one of AMP-224, MEDI0680, or PDR001. In some embodiments, the ICI is a mAb against PD-L1. In some embodiments, the anti-PD-L1 mAb is selected from atezolizumab (Tecentriq), avelumab (Bavencio), and durvalumab (Imfinzi). In some embodiments, the anti-PD-L1 mAb is atezolizumab. In some embodiments, the anti-PD-L1 mAb is durvalumab. In some embodiments, the ICI is a mAb against CTLA-4. In some embodiments, the anti-CTLA-4 mAb is ipilimumab. In some embodiments, the ICI is a mAb against LAG-3. In some embodiments, the anti-LAG-3 mAb is relatlimab.
[0081] As used herein, the term "factor" refers to a measurable biomolecule produced by a subject. In some embodiments, the factor is a protein. In some embodiments, the factor is RNA. In some embodiments, the factor is a gene. In some embodiments, the factor is a secreted factor. In some embodiments, the secreted factor is selected from a cytokine, a chemokine, a growth factor, a soluble receptor, and an enzyme. In some embodiments, the factor is a soluble factor. In some embodiments, the factor is a cellular factor. In some embodiments, the factor is a membrane factor. In some embodiments, the factor is a cell adhesion molecule. In some embodiments, the factor is a factor found in blood. In some embodiments, the factor is a factor produced by the host. In some embodiments, the factor is a resistance factor.
[0082] In some embodiments, expression is protein expression. In some embodiments, expression is secreted protein expression. In some embodiments, protein expression is soluble protein expression. In some embodiments, expression is cellular protein expression. In some embodiments, expression is membrane protein expression. In some embodiments, expression is mRNA expression. In some embodiments, expression is protein expression or mRNA expression. In some embodiments, the expression level is concentration. In some embodiments, the concentration is a concentration level. Those skilled in the art will understand that when the presence of a factor is measured in a liquid sample, expression can be provided as a concentration, such as mg / ml, or in any unit according to the method for determining factor expression. The any unit can be selected from relative fluorescence units (RFU) and NPX (Normalized Protein expression), or any other unit used as a measure of expression. The terms "expression" and "expression level" are used interchangeably herein and refer to the amount of a gene product present in a sample. In some embodiments, the gene product comprises a polynucleotide, e.g., tumor DNA, circulating tumor DNA, or circulating DNA. In some embodiments, the DNA is cell-free DNA. In some embodiments, determining comprises quantification of expression levels. In some embodiments, determining comprises normalization of expression levels. Determining the expression level of a factor can be performed by any method known in the art. Methods for determining protein expression include, for example, antibody arrays, immunoblotting, immunohistochemistry, flow cytometry (FACS), ELISA, proximity extension assay (PEA), aptamer-based assays, proteomics arrays, proteome sequencing, flow cytometry (CyTOF), multiplex assays, mass spectrometry, and chromatography. In some embodiments, determining the expression level of a protein comprises ELISA. In some embodiments, determining the expression level of a protein comprises protein array hybridization.In some embodiments, determining the expression level of a protein comprises mass spectrometry quantification. In some embodiments, determining the expression level of a protein comprises PEA. In some embodiments, determining the expression level of a protein comprises an aptamer. Methods for determining mRNA expression include, for example, RT-PCR, quantitative PCR, real-time PCR, microarray, Northern blotting, in situ hybridization, next-generation sequencing, and massively parallel sequencing.
[0083] In some embodiments, receiving the expression level of the factor is providing the expression level of the factor. In some embodiments, receiving the expression level of the factor is determining the expression level of the factor. In some embodiments, determining is measuring. In some embodiments, measuring is in a sample. In some embodiments, the expression level is detected in a sample. In some embodiments, the sample is a biological sample. In some embodiments, the sample is provided by the subject. In some embodiments, the sample is provided by the subject. In some embodiments, the sample is provided by a responder. In some embodiments, the sample is provided by a non-responder. In some embodiments, each subject in a population of responders provided a sample. In some embodiments, each subject in a population of non-responders provided a sample. In some embodiments, the sample is provided by the subject prior to receiving the treatment. In some embodiments, the expression level of the factor is from a time point prior to administering the treatment. In some embodiments, the treatment is monotherapy. In some embodiments, the treatment is anti-PD-1 / PD-L1 immunotherapy. In some embodiments, the treatment is combination therapy. In some embodiments, the treatment is anti-PD-1 / PD-L1 immunotherapy and chemotherapy. In some embodiments, the sample is provided by the subject after receiving treatment. In some embodiments, the determining is directly in the sample. In some embodiments, the determining is in an unprocessed sample. In some embodiments, the determining is in a processed sample. In some embodiments, the method further comprises processing the sample. In some embodiments, the processing comprises isolating protein from the sample. In some embodiments, the processing comprises isolating nucleic acid from the sample. In some embodiments, the nucleic acid is RNA. In some embodiments, the RNA is mRNA. In some embodiments, the processing comprises lysing cells in the sample. In some embodiments, the nucleic acid is cell-free DNA.In some embodiments, the nucleic acid is tumor cell DNA.
[0084] As used herein, the terms "peptide," "polypeptide," and "protein" are used interchangeably to refer to a polymer of amino acid residues. In another embodiment, as used herein, the terms "peptide," "polypeptide," and "protein" encompass native peptides, peptidomimetics (typically containing non-peptide bonds or other synthetic modifications), and peptide analogs peptoids and semipeptoids, or any combination thereof. In another embodiment, the described peptide polypeptides and proteins have modifications that make them more stable while in the body or more permeable to cells. In one embodiment, the terms "peptide," "polypeptide," and "protein" refer to naturally occurring amino acid polymers. In another embodiment, the terms "peptide," "polypeptide," and "protein" refer to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids.
[0085] In some embodiments, the sample is a biological sample. In some embodiments, the sample is tissue. In some embodiments, the tissue sample is a tumor sample. In some embodiments, the sample is a fluid. In some embodiments, the fluid is a biological fluid. In some embodiments, the sample is from a subject. In some embodiments, the sample is not a tumor sample. In some embodiments, the sample is a tumor sample. In some embodiments, the sample is not a hematopoietic cancer and the sample is a blood sample. In some embodiments, the sample does not contain cancer cells. In some embodiments, the blood sample includes a peripheral blood sample, a serum sample, and a plasma sample. In some embodiments, the sample is a plasma sample. In some embodiments, the sample is a serum sample. In some embodiments, processing comprises isolating plasma. In some embodiments, processing comprises isolating serum. In some embodiments, the biological fluid is selected from blood, plasma, serum, lymph, cerebrospinal fluid, urine, feces, semen, tumor fluid, and gastric fluid. In some embodiments, the samples obtained from the subject and the responder are the same type of sample. In some embodiments, the samples obtained from the subject and the responder are different types of samples. In some embodiments, the samples obtained from the subject and the non-responder are the same type of sample. In some embodiments, the samples obtained from the subject and the non-responder are different types of samples. In some embodiments, the samples obtained from the non-responder and the responder are the same type of sample. In some embodiments, the samples obtained from the non-responder and the responder are different types of samples. In some embodiments, the samples obtained from the subject, the non-responder, and the responder are the same type of sample. In some embodiments, the samples obtained from the subject, the non-responder, and the responder are blood samples. In some embodiments, the samples obtained from the subject, the non-responder, and the responder are plasma samples.In some embodiments, the samples obtained from the subjects, non-responders, and responders are serum samples. In some embodiments, the samples obtained from the subjects, non-responders, and responders are different types of samples.
[0086] In some embodiments, the factor is a factor of a plurality of factors. In some embodiments, expression levels of a plurality of factors are received. In some embodiments, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 41000, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6100, 6200, 6300, 6400, 100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 12000, 15000, 20000, 25000, 30000, 35000, or 40000 factors are received. Each possibility represents a separate embodiment of the present invention. In some embodiments, expression levels of at least 50 factors are received. In some embodiments, expression levels of at least 100 factors are received. In some embodiments, expression levels of at least 200 factors are received. In some embodiments, expression levels of at least 300 factors are received. In some embodiments, expression levels of at least 350 factors are received. In some embodiments, expression levels of at least 375 factors are received. In some embodiments, expression levels of at least 380 factors are received. In some embodiments, expression levels of at least 385 factors are received. In some embodiments, expression levels of at least 388 factors are received.In some embodiments, the plurality is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, The plurality may be 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 12000, 15000, 20000, 25000, 30000, 35000, or 40000. Each possibility represents a separate embodiment of the present invention. In some embodiments, the plurality is at least a factor of 50. In some embodiments, the plurality is at least a factor of 100. In some embodiments, the plurality is at least a factor of 200. In some embodiments, the plurality is at least a factor of 300. In some embodiments, the plurality is at least a factor of 350. In some embodiments, the plurality is at least a factor of 375. In some embodiments, the plurality is at least 380 factors. In some embodiments, the plurality is at least 385 factors. In some embodiments, the plurality is at least 388 factors. In some embodiments, expression levels of at least 50 factors are received. In some embodiments, expression levels of at least 100 factors are received. In some embodiments, expression levels of at least 200 factors are received. In some embodiments, expression levels of at least 300 factors are received. In some embodiments, expression levels of at least 350 factors are received. In some embodiments, expression levels of at least 375 factors are received. In some embodiments, expression levels of at least 380 factors are received. In some embodiments, expression levels of at least 385 factors are received. In some embodiments, expression levels of at least 388 factors are received. In some embodiments, expression levels of at least 400 factors are received.In some embodiments, expression levels of at least 1000 factors are received. In some embodiments, expression levels of at least 5000 factors are received. In some embodiments, expression levels of at least 6000 factors are received. In some embodiments, expression levels of at least 7000 factors are received. In some embodiments, expression levels of at least 8000 factors are received.
[0087] In some embodiments, the factor is selected from the factors provided in Table 4. In some embodiments, a plurality of factors is selected from the factors provided in Table 4. In some embodiments, the plurality of factors comprises at least two factors selected from the factors provided in Table 4. In some embodiments, the plurality of factors consists of factors selected from Table 4. In some embodiments, the factors provided in Table 4 are as follows: KCNAB2, IL12B, IL23A, MCL1, KIR2DS2, AGA, RPN1, LAT, MFAP2, PUF60, MPZ, ACE, RNF122, TXNDC5, CDH15, FGFBP3, COL11A2, INPP5E, ADH7, MVK, RNF146, SOCS3, RBFOX2, ARFGAP1, SRSF6, RBM23, DDR1, APOF, TRA2B, MCTS1, TBCA, RGS7, PTPN9, CSNK1G2, ILF3, TPPP2, ARHGEF2, SRSF7, EWSR1, FSTL1, SPP1, FLRT2, FLRT3, VTN, ATP1B1, WFIKKN2, NRAC, PKD 2, HSPA9, EMC4, ASAP2, NAP1L2, HTR7, DCUN1D3, RBL2, MAD1L1, GRB14, RBBP5, NAB2, CSF1R, CCN4, GPD1, KLK3, CXCL13, GZMA, C9, IL 12B, RAP1GAP, IGFBP1, DHX58, COPS2, IL1RAP, CCL25, HPX, ADM, CD93, ISG15, MYL6B, HSPA1A, MBD1, TRAPPC3, AKT2, CRLF1, FTL, R BBP4, BMPER, SERPINB5, PMP2, OTC, OTOR, AOC1, FGFBP1, ATRN, NAGLU, SAA1, SAA4, CLSTN1, GSS, DLD, EPHB4, PRSS27, MUC16, CFHR2 , HTRA1, KRT19, RBP4, SMOC2, BTD, TXLNA, MZB1, FADD, GSN, CDH17, LECT2, ADAMTSL1, RNASET2, SEMA4A, DDOST, BDH2, SNRPB2, GOL M1, RAB3A, CD46, SEPTIN6, WWOX, WDR5, HPCAL1, ALDH5A1, VAT1, SARS1, AFM, CDA, ITLN1, LRIG1, GREM1, PTGR2, UBE2L6, CLTA, GSR,<h2 style=";text-align:left;direction:ltr">PDCD6、SNCG、CRH、RGS21、UBE2R2、BASP1、GBP5、LMNB2、POP7、RAET1L、SEMA5 B、CNTN3、UBL3、MMACHC、GTF2B、GCHFR、LRATD2、SGK1、TSEN15、SAR1B、CDK5R AP3、HAUS1、NKIRAS1、PHOSPHO2、PCDH17、TRIM5、ALDH7A1、TXNL4A、CEP20、P DE1B、ITGA4、ITGB1、LRFN3、ADGRB1、SGSH、MGAT5、B3GAT1、MGAT5、FBLN7、APB B1IP、PON2、PPP2R5D、RBFOX1、TIMP1、GEMIN7、CSNK1A1L、PHF11、BTN2A2、SK P2、SPATA46、LIN7A、BORCS5、ARRDC5、PCYT1A、PHYH、ANKRD63、VCX、NTAN1、S TARD7、APOL2、FLT4、RCSD1、INIP、VMAC、XPNPEP3、IFNE、NELFA、KDM8、NCBP1 、USF2、LRRC75A、APCS、PLCD1、ESPN、RFX5、RPS6KB2、NOMO2、TCEAL2、CES3、DY RK1A, CYP2C19, CFI, IGFBP3, IL6, LEP, CRTC3, VEGFA, IL1RAP, HGF, PLA2G2A, CCL25, SERPINA7, POR, CCN3, HPX, IGFBP1, MMP3, FGA, FGB, FGG, BCAM, SPIN T1, HAT1, GHR, CFP, CNTN1, SERPINF2, IL19, MB, C9, IGHM, LBP, NAAA, HPLN1, IDS, NID1, ACAN, TGFBI, DLL4, FCGR3B, ACY1, IBSP, SERPINA4, POSTN, SELE, B 2M, HAMP, SERPINA1, AHSG, CKB, CKM, PROC, PROC, ANGPTL4, MBD4, PSMD7, IGHE, CXCL10, KLKB1, CFH, PFDN5, RBM39, DCTPP1, PRSS22, KYNU, IL6, AFM, SERPI NA6、ITIH4、SFN、CCL7、LYZ、MMP13、STC1、CAPG、PI3、GPC5、HRG、SCGB2A1、SI RT2、TNFAIP6、CD300C、GPNMB、KRT18、TNFSF14、LEPR、PRKCG、FGL1、PGLYRP2、NPFF, MFAP4, TMX3, PRKCSH, DEFB112, SEMA4D, ACP6, AFP, NGF, FTH1, FTL, DMKN, EPA10, CHRDL2, TP53, AOC1, IFNA8, CSH1, CSH2, TNC, PLTP, CCN1, CLSTN3, OIT3 , GGT2, FMOD, C5orf38, VWA1, INHBC, ADGRF5, C1QL2, PCYOX1, AOC2, CFHR4, LRRC15, POSTN, UBE2J1, GFRAL, IGF2, LILRB5, LILRA6, APOA2, VWA2, DEPP1, C1QTNF3 , SERPINA9, CFHR5, DLG3, GLTPD2, HBQ1, ENTPD1, AGGF1, NRG2, SPON2, FAM241B, JAML, BCHE, GPNMB, APOD, DLL1, PEAR1, RSPO4, LEP, ARL8B, PCDH10, MFAP3L, CD14, COL15A1, PCDH10, HAVCR1, ARHGEF10, MAN1A2, CRYZL1, TFPI2, PLXDC1, ACP2, BTD, MFAP2, ITIH2, EFCAB14, PLA1A, GZMK, YBX1, IDO1, NQO1, SPOCK3, and NXT1. The amino acid sequences of these factors can be found, for example, in the UNIPROT database, and the Uniprot accession number for each factor is provided in Table 4. Furthermore, methods, reagents, and assays for measuring the expression levels of these factors are well known in the art and commercially available.
[0088] In some embodiments, the population of responders has a disease. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is cancer. In some embodiments, the responders all have the same disease. In some embodiments, the population of non-responders has a disease. In some embodiments, the non-responders all have the same disease. In some embodiments, the population of responders and the population of non-responders all have the same disease. In some embodiments, the population of responders and the subject have the same disease. In some embodiments, the population of non-responders and the subject have the same disease. In some embodiments, the population of non-responders, the population of responders, and the subject have the same disease.
[0089] In some embodiments, the expression level is from a subject before receiving a treatment. In some embodiments, the expression level is determined for a subject before receiving a treatment. In some embodiments, the expression level is from time T0. In some embodiments, the expression level is a baseline expression level. In some embodiments, the sample is provided by a subject before receiving a treatment. In some embodiments, the expression level is from a subject before receiving a first treatment of a treatment. In some embodiments, the expression level is from a subject before receiving a first cycle of a treatment. In some embodiments, the treatment is a dose. In some embodiments, the treatment is a regimen. In some embodiments, the treatment is a combination of a dose and a regimen.
[0090] In some embodiments, "before" is at least 1 hour, 2 hours, 3 hours, 6 hours, 8 hours, 12 hours, 1 day, 2 days, 3 days, 5 days, 1 week, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, or 6 months before the treatment or administration of the treatment. Each possibility represents a separate embodiment of the present invention. In some embodiments, "before" is at least 1 hour. In some embodiments, "before" is immediately before the treatment or administration of the treatment. In some embodiments, "before" is at most 1 hour, 2 hours, 3 hours, 4 hours, 6 hours, 9 hours, 12 hours, 18 hours, 24 hours, 2 days, 3 days, 5 days, 1 week, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, or 6 months before the treatment or administration of the treatment. Each possibility represents a separate embodiment of the present invention. In some embodiments, before is up to 24 hours before the treatment or administration of the treatment. In some embodiments, the administration of the treatment is the first administration of the treatment. In some embodiments, the administration of the treatment is any administration of the treatment.
[0091] In some embodiments, the expression level is from a subject after receiving a treatment. In some embodiments, the expression level is from time T1. In some embodiments, the sample is provided by a subject after receiving a treatment. In some embodiments, the expression level is from a subject after receiving a first treatment with a treatment. In some embodiments, the expression level is from a subject after receiving any treatment with a treatment.
[0092] In some embodiments, later is a time point after initiation of therapy or administration of therapy sufficient to alter expression of at least one factor. In some embodiments, later is a time point after initiation of therapy or administration of the first treatment. In some embodiments, later is at least 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 4 weeks, 6 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, or 1 year. Each possibility represents a separate embodiment of the present invention. In some embodiments, later is at least 24 hours. In some embodiments, later is at least 2 weeks. In some embodiments, later is at least 3 weeks. In some embodiments, later is at least 6 weeks. In some embodiments, later is at most 1 week, 2 weeks, 3 weeks, 4 weeks, 6 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, or 1 year after initiation of therapy or administration of therapy. Each possibility represents a separate embodiment of the present invention.
[0093] In some embodiments, receiving the expression levels includes receiving expression levels of factors for a group of factors larger than a plurality of factors. In some embodiments, the expression levels received for the larger group are received for responders and non-responders. In some embodiments, a subgroup of proteins is selected from the group. In some embodiments, the subgroup is a subset. In some embodiments, the subgroup is designated for a plurality of factors. In some embodiments, the method includes designating. In some embodiments, receiving further includes applying a machine learning algorithm for each factor of the group. In some embodiments, the algorithm classifies the factors as responder-derived and non-responder-derived. In some embodiments, the algorithm outputs whether the subject who provided the sample having the measured expression level of the factor is a responder or a non-responder. In some embodiments, receiving further includes selecting a subgroup of factors for which the algorithm most evenly divides the subjects into responders and non-responders. In some embodiments, the subjects are all subjects in the population of responders and non-responders. In some embodiments, the factor processed by the algorithm that most evenly divides all responder and non-responder subjects into responder and non-responder groups (even if imprecisely designated) is selected as the subgroup. In some embodiments, the algorithm is trained on the expression levels of factors received in responders and non-responders. In some embodiments, the algorithm is trained on a training set. In some embodiments, training is on expression levels and tags indicating whether the expression levels were from responders or non-responders. In some embodiments, training is on expression levels, clinical information, and tags indicating whether the expression levels were from responders or non-responders. In some embodiments, training is on the number of factors associated with resistance.In some embodiments, training is on the number of factors associated with resistance and a tag indicating whether the number of factors associated with resistance was from a responder or non-responder.
[0094] In some embodiments, receiving further comprises determining a mean difference between responders and non-responders for each factor of the group. In some embodiments, receiving further comprises determining statistical significance between the levels of responders and non-responders for each factor of the group. In some embodiments, the statistical significance is between the means. In some embodiments, the statistical significance is a p-value. In some embodiments, receiving further comprises selecting a subgroup of factors with the greatest statistical significance. In some embodiments, a statistical test is applied to determine significance. In some embodiments, the test is a Kolmogorov-Smirnov test. In some embodiments, the subgroup includes a predetermined number of factors with the greatest significance. In some embodiments, the predetermined number is about 50 factors. In some embodiments, the predetermined number is at least 50 factors.
[0095] In some embodiments, the subgroups comprise factors that the algorithm most evenly divides the subjects into. In some embodiments, the even division is into responders and non-responders. In some embodiments, the subgroups are the top 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 750, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, or 5000. Each possibility represents a separate embodiment of the present invention. In some embodiments, the subgroups are the top 50. In some embodiments, the subgroups are the top 100. In some embodiments, the subgroups are the top 200. In some embodiments, the subgroups are the top 500.
[0096] In some embodiments, the method further comprises performing a dimensionality reduction step. In some embodiments, the reduction is with respect to a plurality of factors. In some embodiments, the reduction is to reduce the number of a plurality of factors. In some embodiments, the dimensionality reduction step identifies a subgroup or subset of factors. In some embodiments, the factors are main factors. In some embodiments, the training set includes expression levels of only a subset / subgroup of factors. In some embodiments, the subgroup or subset of factors is the factor that most evenly balances the number of predicted responders and non-responders. In some embodiments, the prediction is predicted by a machine learning algorithm. In some embodiments, the machine learning algorithm is a trained machine learning algorithm. In some embodiments, the machine learning algorithm is a machine learning algorithm in training.
[0097] In some embodiments, a preprocessing step may be performed to preprocess the received expression levels. In some embodiments, the preprocessing step may include at least one of data cleaning and normalization, feature selection, feature extraction, dimensionality reduction, and / or any other suitable preprocessing method or technique. Feature selection may be performed by a statistical test, such as the Kolmogorov-Smirnov (KS) test, or any other test known in the art.
[0098] In some embodiments, a factor selection and / or dimension reduction step may be performed to reduce the number of factors in each sample and / or to obtain a set of main factors, e.g., factors that may have significant predictive power. In some embodiments, the factor selection is RAP selection. Thus, in some embodiments, the factor selection and / or dimension reduction step may result in a reduction in the number of factors in each sample and / or set of values. In some embodiments, dimension reduction selects main factors, e.g., proteins, based on the level of response predictive power that the factors produce for the desired prediction. In specific embodiments, dimension reduction involves considering all or some factors as vector components and calculating their norms.
[0099] In some embodiments, any suitable factor selection and / or dimensionality reduction method or technique may be employed, for example, but not limited to, the following: ANOVA with S0 parameter: Analysis of variance with an additional parameter (S0) that controls for the relative importance of features based on the resulting test p-value and differences between group means (see, e.g., Tusher, Tibshirani and Chu, PNAS 98, pp5116-21, 2001). Scalable Empirical Bayesian Model Selection (SEMMS): An empirical Bayesian feature selection method that applies parsimonious mixture models to identify significant predictors (see, e.g., Bar, Booth, and Wells. A scalable empirical Bayesian approach to variable selection in generalized linear models, 2019). L2N: A method for differential expression analysis using a three-component mixture model consisting of two log-normal components (L2) for differentially expressed features (one for under-expressed features and one for over-expressed features) and a single normal component (N) for non-differentially expressed features (see, e.g., Bar and Schifano. Differential variation and expression analysis. Stat 8, e237, doi:10.1002 / sta4.237, 2019 Bar and Schifano. Differential variation and expression analysis. Stat 8, e237, doi:10.1002 / sta4.237, 2019). Genetic algorithm: a group of heuristic optimization algorithms that employ organic evolutionary techniques such as random mutation, recombination, and natural selection as a way to achieve an optimal configuration (see, for example, Popovic, Sifrim, Pavlopoulos, Moreau, and Bart De Moor. A Simple Genetic Algorithm for Biomarker Mining. 2012). Naive Classifier: A naive classifier evaluates response scores by reducing the dimensionality to a single score. This is done by considering all features (e.g., a specific profile, such as protein expression levels) as components of a vector and calculating its norm. Dimensionality reduction reduces the potential risk of overfitting. In some embodiments, the vector components are normalized according to typical component values among patients belonging to the same response group (e.g., responders), so that the normalized norm quantifies the amount of deviation from the typical respective class value. In further embodiments, a naive classifier allows training using data from subjects who belong to only a portion of the response group.
[0100] As used herein, the terms "responder" or "known to be responsive" are used interchangeably and refer to a subject who, upon administration of a treatment, shows an improvement in at least one criterion of the disease being treated with the treatment, or does not show an increase in the severity of the disease. In some embodiments, a responder is a subject who, upon administration of a treatment, shows an improvement in the disease being treated with the treatment. In some embodiments, a responder is a subject who shows a clinical benefit upon administration of the treatment. In some embodiments, a responder is a subject who does not show an increase in the severity of the disease upon administration of the treatment. In some embodiments, the increase in severity is over time. In some embodiments, the lack of an increase in severity is a stable state. In some embodiments, a responder is a subject who shows a mixed response upon administration of the treatment. In some embodiments, a responder is a subject who shows a mixed response upon administration of the treatment, where the mixed response is an improvement in at least one criterion of the disease but no improvement in other criteria of the disease. In some embodiments, the mixed response is a shrinkage of some lesions combined with the growth of new or existing lesions. In some embodiments, a responder is a subject in which the treatment results in an anti-disease response. In some embodiments, with respect to a subject with cancer, a responder is a subject in which the treatment results in an anti-cancer response. In some embodiments, the response is not a reduction in side effects. In some embodiments, the response is a reduction in side effects. In some embodiments, the response is a response to the disease itself. In some embodiments, the anti-cancer response is an anti-tumor response. In some embodiments, the anti-tumor response comprises tumor regression. In some embodiments, the anti-tumor response comprises tumor shrinkage. In some embodiments, the anti-tumor response comprises a lack of tumor growth. In some embodiments, the anti-tumor response comprises a lack of tumor metastasis. In some embodiments, the anti-tumor response comprises a lack of tumor hyperproliferation. In some embodiments, the improvement is in at least one symptom of the disease. In some embodiments, the response is a complete response. In some embodiments, the response is a minimal response.In some embodiments, the response is a partial response. In some embodiments, the response includes a stable state. In some embodiments, a responder is a subject who has a favorable response to treatment. In some embodiments, a non-responder is a subject who has an unfavorable response to treatment. In some embodiments, the unfavorable response is an increase in tumor burden. An increase in tumor burden can include an increase in either tumor size or total number of cancer cells, such as an increase in tumor size, an increase in tumor spread, an increase in metastasis, an increase in tumor cell proliferation, or other increase. In some embodiments, the response is a response to a monotherapy. In some embodiments, the response is a response to a combination therapy.
[0101] As used herein, a "favorable response" of a cancer patient refers to the "responsiveness" of the cancer patient to treatment with a therapy, i.e., treatment of a responsive cancer patient with a therapy results in a desirable clinical outcome, such as tumor regression, tumor reduction, or tumor necrosis; reduction in tumor burden; anti-tumor response by the immune system; or prevention or delay of tumor recurrence, tumor growth, or tumor metastasis. In some embodiments, the subject is a complete responder, or treatment with a cancer therapy results in a stable state. In some embodiments, a complete responder is a subject in whom there is no detectable cancer after treatment with a therapy. In this case, treatment of the responsive cancer patient with the therapy can be continued, or treatment can be discontinued if the patient is cancer-free, and is so advised. In some embodiments, the method further includes continuing treatment for a subject who is not a non-responder. In some embodiments, the subject is a non-responder, a minimal responder, a partial responder, or has stable disease, and the method further comprises continuing to administer the treatment to the subject and treating the subject with additional therapy to increase responsiveness (e.g., as determined using a resistance-associated protein (RAP) assay provided herein). In some embodiments, a subject that is not a non-responder is a responder.
[0102] As used herein, the terms "non-responder" and "known non-responsive" subject are used interchangeably and refer to a subject who does not show improvement or stabilization of their disease upon treatment. In some embodiments, a non-responder shows worsening of their disease upon treatment. In some embodiments, a non-responder is a subject who does not show clinical benefit upon treatment. In some embodiments, a non-responder is not a subject who experiences side effects of treatment. In some embodiments, a non-responder is a subject whose disease progresses. In some embodiments, a non-responder is a subject whose disease does not stabilize after treatment. In some embodiments, a non-responder is a subject whose disease does not improve after treatment. In some embodiments, a non-responder is a subject who is not a responder as defined herein above. In some embodiments, a non-responder is a subject who has an unfavorable response to treatment. In some embodiments, a non-responder is a subject who is resistant to treatment. In some embodiments, a non-responder is a subject who is refractory to treatment. In some embodiments, non-responder is a non-responsive to monotherapy. In some embodiments, non-responder is a non-responsive to combination therapy.
[0103] As used herein, an "unfavorable response" of a cancer patient refers to the "non-responsiveness" of the cancer patient to treatment with a therapy, such that treatment of a non-responsive cancer patient with a therapy does not result in a desirable clinical outcome, and likely results in undesirable outcomes such as tumor expansion, recurrence, or metastasis. In some embodiments, the method further comprises discontinuing administration of the therapy to the subject who is a non-responder. In some embodiments, the method further comprises continuing administration of the therapy to the subject in combination with an additional therapy. In some embodiments, the additional therapy increases the responsiveness of the non-responsive patient.
[0104] In some embodiments, the method is for determining whether the response is considered a durable response (e.g., progression-free survival of greater than 6 months). In some embodiments, the response is at least a 3-month response. In some embodiments, the response is a response at some point since treatment. In some embodiments, since treatment is from the start of treatment. In some embodiments, the response is a 3-month response. In some embodiments, the response is at least a 6-month response. In some embodiments, the response is a 6-month response. In some embodiments, the response is at least a 7-month response. In some embodiments, the response is a 7-month response. In some embodiments, the response is at least a 1-year response. In some embodiments, the response is at least a 2-year response. In some embodiments, the response is a 2-year response. In some embodiments, the response is at least a 3-year response. In some embodiments, the response is at least a 4-year response. In some embodiments, the response is a 4-year response. In some embodiments, the response is at least a 5-year response. In some embodiments, the response is a 5-year response. It will be understood by those skilled in the art that a response over at least a given period of time includes monitoring the response at least at that time, and possibly monitoring the response up to that time.
[0105] In some embodiments, the method further comprises administering the treatment to a subject predicted to respond to the treatment. In some embodiments, the method further comprises continuing to administer the treatment to a subject predicted to respond to the treatment. In some embodiments, the method further comprises not administering the treatment to a subject predicted not to respond to the treatment. In some embodiments, the method further comprises not administering the treatment to a subject predicted not to respond to the treatment. In some embodiments, the method further comprises discontinuing treatment to a subject predicted not to respond to the treatment. In some embodiments, the method further comprises administering an alternative therapy to a subject predicted to be a non-responder. In some embodiments, the alternative therapy is an additional therapy. In some embodiments, the additional therapy is chemotherapy. In some embodiments, the method further comprises administering or continuing to administer the treatment in combination with an agent or treatment that blocks or inhibits at least one factor associated with resistance in a subject predicted to be resistant to the treatment. In some embodiments, the agent or treatment that blocks or inhibits at least one factor associated with resistance is an additional therapy. In some embodiments, the agent or treatment that blocks or inhibits at least one signaling pathway of a factor associated with resistance is an additional therapy. In some embodiments, the combination therapy is administered to a subject predicted to be a non-responder.
[0106] In some embodiments, the method further comprises administering a monotherapy to a subject predicted to respond to the monotherapy. In some embodiments, the method further comprises administering a monotherapy to a subject with a PD-L1 high cancer predicted to respond to the monotherapy. In some embodiments, the method further comprises administering a combination therapy to a subject predicted not to respond to the monotherapy. In some embodiments, the method further comprises administering a combination therapy to a subject with a PD-L1 high cancer predicted not to respond to the monotherapy.
[0107] In some embodiments, the method further comprises administering the combination therapy to a subject predicted to respond to the combination therapy. In some embodiments, the method further comprises administering the combination therapy to a subject with a PD-L1 low or negative cancer predicted to respond to the combination therapy. In some embodiments, the method further comprises administering an alternative therapy to a subject predicted not to respond to the combination therapy. In some embodiments, the method further comprises administering an alternative therapy to a subject with a PD-L1 low or negative cancer predicted not to respond to the combination therapy. Examples of alternative therapies include, but are not limited to, other ICI combination therapies (e.g., with anti-CTLA-4) and non-chemotherapeutic therapies.
[0108] In some embodiments, the method further includes administering to the subject (e.g., the non-responder) an agent that modulates at least one factor. In some embodiments, modulating includes inhibiting, blocking, and modulating. In some embodiments, modulating is inhibiting. In some embodiments, the method further includes administering to the subject (e.g., the non-responder) an agent that modulates a pathway comprising at least one factor. In some embodiments, modulating at least one factor is modulating a pathway comprising at least one factor. In some embodiments, modulating a pathway includes modulating a driver protein / gene that controls the at least one factor. In some embodiments, modulating a pathway includes modulating a driver protein / gene that controls the pathway. In some embodiments, modulating a pathway comprising at least one factor is modulating a receptor for the factor (e.g., using a receptor agonist or antagonist), a ligand for the factor, a paralog of the factor, or a combination thereof. In some embodiments, modulating is modulating multiple factors. In some embodiments, modulating is modulating multiple factors in a signature. In some embodiments, modulating is modulating each factor in a signature. In some embodiments, the modulation achieves a better response to treatment. In some embodiments, the factor is a factor associated with resistance.
[0109] In some embodiments, the resistance score is a RAP score. In some embodiments, the resistance score is a response score. In some embodiments, the resistance score is a 1-response score. In some embodiments, the resistance score is a 10-response score. In some embodiments, the response score is a 1-resistance score. In some embodiments, the response score is a 10-resistance score. Those skilled in the art will understand that response scores and resistance scores are reciprocals. Thus, if the score scale is 0-1, converting one score to the other is a 1-score. Conversely, if the score scale is 0-10, converting one score to the other is a 10-score. The same applies to any scale used for the two scores. In some embodiments, the resistance score is an overall resistance score. In some embodiments, the response score is an overall response score. In some embodiments, the RAP score is an overall RAP score. In some embodiments, the resistance score is based on the similarity of the expression level of the factor in the subject to the expression level of the factor in the non-responder. In some embodiments, the resistance score is based on the similarity of the expression level of the factor in the subject to the expression level of the factor in the responder. In some embodiments, based on is calculated based on. In some embodiments, similarity is a lack of similarity. In some embodiments, similarity to a responder is a lack of similarity to a non-responder. In some embodiments, similarity to a non-responder is a lack of similarity to a responder. In some embodiments, similarity is measured on a scale.
[0110] In some embodiments, the scale is between 0 and 1, with 1 being completely similar to a non-responder and 0 being completely similar to a responder. In some embodiments, the resistance score is between 0 and 1, with 1 being completely similar to a non-responder and 0 being completely similar to a responder. In some embodiments, the resistance score is based on the similarity of the expression level of a factor in a subject to the expression level of the factor in a non-responder and the expression level of the factor in a responder. In some embodiments, the response score is between 0 and 1, with 1 being completely similar to a responder and 0 being completely similar to a non-responder. In some embodiments, the response score is a PROphet score. In some embodiments, a prophet-positive subject is a subject with a response score above a predetermined threshold. In some embodiments, a prophet-negative subject is a subject with a response score below a predetermined threshold. In some embodiments, the response score is based on the similarity of the expression level of a factor in a subject to the expression level of the factor in a non-responder and the expression level of the factor in a responder. In some embodiments, a response score of 0.5 to 1 indicates that the subject is a responder. In some embodiments, a response score greater than 0.5 indicates that the subject is a responder. In some embodiments, a response score between 0.5 and 0 indicates that the subject is a non-responder. In some embodiments, a response score less than 0.5 indicates that the subject is a non-responder.
[0111] In some embodiments, the scale is 0 to 10, with 10 being completely similar to the responder and 0 being completely similar to the non-responder. In some embodiments, the resistance score is 0 to 10, with 10 being completely similar to the non-responder and 0 being completely similar to the responder. In some embodiments, the resistance score is based on the similarity of the expression level of the factor in the subject to the expression level of the factor in the non-responder and the expression level of the factor in the responder. In some embodiments, the response score is 0 to 10, with 10 being completely similar to the responder and 0 being completely similar to the non-responder. In some embodiments, the response score is a PROphet score. In some embodiments, the response score is an overall response score. In some embodiments, a PROphet-positive subject is a subject with a response score above a predetermined threshold. In some embodiments, a PROphet-negative subject is a subject with a response score below a predetermined threshold. In some embodiments, the response score is based on the similarity of the expression level of the factor in the subject to the expression level of the factor in the non-responder and the expression level of the factor in the responder. In some embodiments, a response score of 5 to 10 indicates that the subject is a responder. In some embodiments, a response score greater than 5 indicates that the subject is a responder. In some embodiments, a response score of 5 to 0 indicates that the subject is a non-responder. In some embodiments, a response score less than 5 indicates that the subject is a non-responder.
[0112] In some embodiments, the method includes selecting a subset of factors prior to step (b). In some embodiments, the subset is a subset of a plurality of factors. In some embodiments, prior to step (b) is prior to calculating. In some embodiments, the subset is a subset of a plurality of factors. In some embodiments, the subset includes factors that best distinguish between responders and non-responders. In some embodiments, the factors that best distinguish are a top percentage. In some embodiments, the top percentage is the top 1, 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50% of the factors. Each possibility represents a separate embodiment of the present invention. In some embodiments, the top percentage is the top 20%. In some embodiments, the top factors are the top 10, 20, 25, 30, 40, 50, 60, 70, 75, 80, 90, or 100 factors. Each possibility represents a separate embodiment of the present invention. In some embodiments, the top factors are the top 50 factors. In some embodiments, the selecting comprises applying a Kolmogorov-Smirnov test. In some embodiments, a Kolmogorov-Smirnov test is applied to the expression levels of the received factors. In some embodiments, the Kolmogorov-Smirnov test determines how well the factors distinguish between responders and non-responders. In some embodiments, the Kolmogorov-Smirnov test outputs a measure of how well the factors distinguish, and the best factors are the factors with the highest scores. In some embodiments, the selecting comprises applying an XGBoost algorithm. In some embodiments, the calculating is for a subset. In some embodiments, the calculating is for each factor in the subset.
[0113] In some embodiments, the calculating comprises applying a machine learning algorithm. In some embodiments, the calculating comprises applying a machine learning model. In some embodiments, the machine learning model is a machine learning algorithm. In some embodiments, the machine learning model implements a machine learning algorithm. In some embodiments, the algorithm is a classifier. In some embodiments, the algorithm is a regression model. In some embodiments, the algorithm is supervised. In some embodiments, the algorithm is unsupervised. In some embodiments, the machine learning algorithm is trained on expression levels in responders. In some embodiments, the machine learning algorithm is trained on expression levels in non-responders. In some embodiments, the machine learning algorithm is trained on expression levels in responders and non-responders. In some embodiments, the machine learning algorithm is trained on a training set. In some embodiments, the machine learning algorithm is trained by a method of the present invention. In some embodiments, the machine learning algorithm is applied to a factor of the plurality of factors. In some embodiments, the machine learning algorithm is applied to each factor of the plurality of factors. In some embodiments, the machine learning algorithm is applied to a subset. In some embodiments, the machine learning algorithm is applied to a subset of factors. In some embodiments, the machine learning algorithm is applied to each factor of the subset of factors. In some embodiments, each factor is analyzed and calculated separately, and the machine learning algorithm does not use the expression levels of multiple factors as a training set. In some embodiments, the trained machine learning algorithm is applied to individual protein expression levels from the subject. In some embodiments, the machine learning algorithm trained for the expression levels of a specific factor in responders and non-responders is applied to the expression levels of that specific factor in the subject. Those skilled in the art will understand that for each factor of the multiple factors, a different algorithm is trained and then applied to each expression level of the subject.Thus, if three algorithms are trained separately on the expression levels of factors A, B, and C in responders and non-responders, the algorithm trained on the expression levels of factor A is applied to the expression levels of factor A in the subject, the algorithm trained on the expression levels of factor B is applied to the expression levels of factor B in the subject, and the algorithm trained on the expression levels of factor C is applied to the expression levels of factor C in the subject. In some embodiments, during the training phase, a machine learning model is trained on a training set containing expression data for a single factor from responders and non-responders, using the corresponding annotation of "responder" or "non-responder" to predict or classify the expression data of the factor into a "responder" class and a "non-responder" class. In some embodiments, during the inference phase, a machine learning model is applied to the expression data of a single factor from the subject to predict the classification of the factor as similar to a responder or non-responder. In some embodiments, the classification is a resistance score. In some embodiments, the classification is a response score. In some embodiments, the classification is a measure of how similar a factor is to a non-responder or how dissimilar it is to a responder.
[0114] In some embodiments, the trained machine learning algorithm is trained to predict the responsiveness of a subject suffering from a disease to a treatment. In some embodiments, the trained machine learning algorithm is trained to output a resistance score. In some embodiments, the trained machine learning algorithm is trained to output a resistance probability. In some embodiments, the trained machine learning algorithm is trained to output a probability of clinical benefit. In some embodiments, the trained machine learning algorithm is trained to output an activity score. In some embodiments, the trained machine learning algorithm is trained to predict the activity of a factor associated with resistance in a subject. In some embodiments, the trained machine learning algorithm is trained to predict whether a factor is associated with resistance in a subject. In some embodiments, the trained machine learning algorithm is trained to predict whether a factor in a subject is associated with resistance in a subject.
[0115] In some embodiments, the trained machine learning algorithm is trained to predict the responsiveness of a subject suffering from a disease to a treatment. In some embodiments, the trained machine learning algorithm is trained to output a response score. In some embodiments, the trained machine learning algorithm is trained to output a response probability. In some embodiments, the trained machine learning algorithm is trained to output a probability of clinical benefit. In some embodiments, the trained machine learning algorithm is trained to output an activity score. In some embodiments, the trained machine learning algorithm is trained to predict the activity of a factor associated with response in a subject. In some embodiments, the trained machine learning algorithm is trained to predict whether a factor is associated with response in a subject. In some embodiments, the trained machine learning algorithm is trained to predict whether a factor in a subject is associated with response in a subject.
[0116] In some embodiments, the training set includes expression levels of the received factor. In some embodiments, the training set includes expression levels of the received factor in both responders and non-responders. In some embodiments, the training set includes expression levels of the received factor in both mono-responders and mono-non-responders. In some embodiments, the training set includes expression levels of the received factor in both combo-responders and combo-non-responders. In some embodiments, the training set includes expression levels of the received factor in mono-responders, mono-non-responders, combo-responders, and combo-non-responders. In some embodiments, the training set includes expression levels of the received factor for only one factor. In some embodiments, the training set includes a number of resistance-related or response-related factors expressed in the samples. In some embodiments, the samples are from subjects suffering from the disease. In some embodiments, the samples are from responders. In some embodiments, the samples are from non-responders. In some embodiments, the training set includes at least one clinical parameter. In some embodiments, the clinical parameter is from a subject. In some embodiments, the subjects are responders and non-responders. In some embodiments, the training set comprises a label. In some embodiments, the label is related to the responsiveness of the subject. In some embodiments, the label is a responder or a non-responder. In some embodiments, a factor associated with resistance is labeled with a label. In some embodiments, the expression level of a factor associated with resistance is labeled with a label. In some embodiments, at least one clinical parameter is labeled with a label.
[0117] In some embodiments, the training set further includes at least one clinical parameter for each responder and non-responder, and the machine learning algorithm is applied to the expression levels of each received factor from the subject and at least one clinical parameter of the subject. In some embodiments, the at least one clinical parameter is the sex of the subject. In some embodiments, the training set further includes the sex of the subject. In some embodiments, the subject is each subject. In some embodiments, the sex is gender. In some embodiments, the at least one clinical parameter is sex. In some embodiments, the sex is the sex of the subject. In some embodiments, the sex is male or female. In some embodiments, the sex is sex at birth. In some embodiments, the training set includes the sex of each responder. In some embodiments, the training set includes the sex of each non-responder. In some embodiments, the training set includes the sex of each mono-responder. In some embodiments, the training set includes the sex of each mono-non-responder. In some embodiments, the training set includes the sex of each combo-responder. In some embodiments, the training set includes the gender of each combo-non-responder. In some embodiments, the clinical parameter is age. In some embodiments, the age is the age of the subject. In some embodiments, the clinical parameter is treatment selection. In some embodiments, the treatment selection parameter is whether the treatment was a first line treatment or an advanced treatment. In some embodiments, the treatment selection is a first line treatment. In some embodiments, the treatment selection is a secondary treatment. In some embodiments, the secondary treatment is an advanced treatment. Those skilled in the art will understand that an advanced treatment can be any treatment selection after a first line treatment, for example, a second line treatment, a third line treatment, a fourth line treatment, a fifth line treatment, etc. In some embodiments, the clinical parameter is whether the treatment is a first line treatment or an advanced treatment. In some embodiments, the clinical parameter is PD-L1 status.In some embodiments, the PD-L1 status is the PD-L1 status of the cancer. Methods for measuring PD-L1 levels in cancer cells (e.g., tumors) are well known in the art, and any such method may be used. In some embodiments, the PD-L1 status includes high PD-L1 or low PD-L1. In some embodiments, the PD-L1 status includes high PD-L1, low PD-L1, or no PD-L1. In some embodiments, the PD-L1 status includes high PD-L1, moderate PD-L1, or low PD-L1. In some embodiments, the PD-L1 level is a number between 0 and 100. In some embodiments, the PD-L1 level is a percentage between 0 and 100. In some embodiments, the PD-L1 status includes PD-L1 expression in less than 1% of cancer cells, in 1-49% of cancer cells, or in more than 50% of cancer cells. In some embodiments, PD-L1 expression in less than 1% of cancer cells represents no PD-L1 expression. In some embodiments, a PD-L1 low or negative cancer comprises fewer than 50% of cancer cells positive for PD-L1 expression. In some embodiments, the expression is surface expression. In some embodiments, a PD-L1 negative cancer comprises fewer than 1% of cancer cells positive for PD-L1 expression. In some embodiments, PD-L1 expression in less than 1% of cancer cells is low PD-L1 expression. In some embodiments, PD-L1 expression in 1-49% of cancer cells is low PD-L1 expression. In some embodiments, PD-L1 expression in 1-49% of cancer cells is moderate PD-L1 expression. In some embodiments, PD-L1 expression in 50% or more cancer cells is high PD-L1 expression. In some embodiments, a PD-L1 high cancer comprises expression in at least 50% of cancer cells. In some embodiments, a PD-L1 high cancer comprises at least 50% of cancer cells positive for PD-L1 expression. In some embodiments, a PD-L1 low cancer comprises expression in 1-49% of cells. In some embodiments, a PD-L1 non-cancer comprises expression in 0% of cells.In some embodiments, the non-PD-L1 cancer comprises expression in less than 1% of cells. In some embodiments, the PD-L1 low or negative cancer is a PD-L1 low cancer. In some embodiments, the PD-L1 low or negative cancer is a PD-L1 negative cancer. In some embodiments, the non-PD-L1 is a PD-L1 negative cancer.
[0118] In some embodiments, the clinical parameter is a known biomarker for the disease or a mutation in a known biomarker for the disease. In some embodiments, the biomarker is selected from MYC, NOTCH, EGFR, HER2, BRAF, KRAS, MAP2K1, MET, NRAS, NTRK1, NTRK2, NTRK3, PIK3CA, RET, ROS1, TP53, ALK, CDKN2A, KIT, NF1, BFAST, FGFR, LDH, PTEN, RB1, PD-L1, MSI (Microsatelite Instability), TMB (Tumor Mutational Burden), or a combination thereof. In some embodiments, the clinical parameter is biomarker expression. In some embodiments, expression is percent expression. In some embodiments, expression is mutation status.
[0119] In some embodiments, the training set further comprises the gender, age, and PD-L1 status of each responder and non-responder. In some embodiments, the training set further comprises the gender of each responder and non-responder. In some embodiments, the training set further comprises the age and PD-L1 status of each responder and non-responder. In some embodiments, the machine learning algorithm is applied to the expression levels of each received factor from the subject and the gender of the subject. In some embodiments, the machine learning algorithm is applied to the expression levels of each received factor from the subject, and the gender, age, and PD-L1 status of the subject. In some embodiments, the calculating comprises applying a machine learning algorithm trained with a training set comprising the expression levels of the received factors and at least one clinical parameter in responders and non-responders to the expression levels from the subject and at least one clinical parameter of the subject, whereby the machine learning algorithm outputs a resistance score. In some embodiments, the training comprises the expression levels of the received factors in responders and non-responders and the clinical parameters of each responder and non-responder, whereby the machine learning algorithm is applied to the expression levels of each received factor from the subject and the clinical parameter of the subject, whereby the machine learning algorithm outputs a response score. In some embodiments, the training includes expression levels of the received factors in the responders and non-responders and clinical parameters selected from gender, age, and PD-L1 expression of each responder and non-responder, or any combination thereof, and a machine learning algorithm is applied to the expression levels of each received factor from the subject and the subject's clinical parameters, and the machine learning algorithm outputs a response prediction. In some embodiments, the training set includes the number of resistance-associated factors and at least one clinical parameter in each responder and non-responder, and a machine learning algorithm is applied to the number of resistance-associated factors from the subject and at least one clinical parameter of the subject, and the machine learning algorithm outputs a response prediction.In some embodiments, the training set includes the number of resistance-associated factors in each responder and non-responder and the gender of each responder and non-responder, and a machine learning algorithm is applied to the subject-derived numbers of resistance-associated factors and the gender of the subject, and the machine learning algorithm outputs a response prediction. In some embodiments, the training set includes the number of resistance-associated factors in each responder and non-responder, the age and PD-L1 status of each responder and non-responder, and a machine learning algorithm is applied to the subject-derived numbers of resistance-associated factors and the age and PD-L1 status of the subject, and the machine learning algorithm outputs a response prediction.
[0120] In some embodiments, the training set comprises expression levels of the received factors in responders and non-responders. In some embodiments, the training set comprises expression levels of the received factors in responders and non-responders and clinical parameters. In some embodiments, the training set comprises expression levels of the received factors in responders and non-responders and the gender of each of the responders and non-responders. In some embodiments, the trained machine learning algorithm is applied to the expression levels of each received factor from the subject. In some embodiments, the trained machine learning algorithm is applied to the expression levels of each received factor from the subject. In some embodiments, the trained machine learning algorithm is applied to the expression levels of each received factor from the subject and clinical parameters from the subject. In some embodiments, the trained machine learning algorithm is applied to the expression levels of each received factor from the subject and the gender of the subject.
[0121] In some embodiments, the clinical parameter is the type of treatment. In some embodiments, the clinical parameter is expression of a therapeutic target. In some embodiments, the clinical parameter is expression of a protein in a process that is a therapeutic target. In some embodiments, the process is a process that includes the therapeutic target. In some embodiments, the expression is expression in a subject. In some embodiments, the expression is expression in diseased tissue. In some embodiments, the expression is expression in a diseased tissue sample. In some embodiments, the expression is expression in a tumor. In some embodiments, the expression is expression in a tumor sample. In some embodiments, the tumor sample is a biopsy. In some embodiments, the expression is expression that is not in a tumor. In some embodiments, the expression is expression that is not in a tumor sample. In some embodiments, the expression is expression in a liquid biopsy. In some embodiments, the expression is percent expression. In some embodiments, the percent is percent of cells. In some embodiments, the therapy is anti-PD-1 therapy and the protein in the process is PD-L1. In some embodiments, the therapy is anti-PD-L1 therapy and the target protein is PD-L1. In some embodiments, the clinical parameter is expression of PD-L1. In some embodiments, the training set includes at least one clinical parameter selected from treatment selection, PD-L1 expression, gender, and age. In some embodiments, the training set includes protein expression level and gender. In some embodiments, the training set includes number of RAPs, age, and PD-L1 status.
[0122] In addition, clinical parameters can also be included.Those skilled in the art will be able to select relevant clinical parameters for inclusion in the training set.Examples of additional clinical parameters include, but are not limited to, the histological type of sample (e.g., adenocarcinoma, squamous cell carcinoma, etc.), the location of metastasis, the location of tumor, cancer stage classification (e.g., tumor, lymph node and metastasis, TNM, stage classification, etc.), performance status (e.g., ECOG performance status, etc.), gene mutation, epigenetic status, general medical history, vital signs, blood measurements, renal and liver function, weight, height, pulse, blood pressure, and smoking history.
[0123] In some embodiments, in the inference stage, a trained machine learning algorithm is applied. In some embodiments, the trained machine learning algorithm is applied to the expression levels of each received factor. In some embodiments, the trained machine learning algorithm is applied to the expression levels of each received factor and at least one clinical parameter. In some embodiments, the trained machine learning algorithm is applied to the expression levels of each received factor from a subject and the gender of the subject. In some embodiments, the trained machine learning algorithm is applied to the number of proteins associated with resistance. In some embodiments, the trained machine learning algorithm is applied to the number of factors associated with resistance. In some embodiments, the trained machine learning algorithm is applied to the number of factors associated with resistance and at least one clinical parameter.
[0124] In some embodiments, in the inference step, an input is received. In some embodiments, the input includes a number of resistance-associated factors expressed in the sample. In some embodiments, the sample is from a subject. In some embodiments, the input includes at least one clinical parameter. In some embodiments, the subject has a disease. In some embodiments, the subject has unknown responsiveness to the treatment. In some embodiments, the parameter is of a subject with unknown responsiveness. In some embodiments, in the inference step, a trained machine learning algorithm is applied. In some embodiments, the applied is applied to input. In some embodiments, the input is received input. In some embodiments, the inference step is predicting responsiveness. In some embodiments, the responsiveness is responsiveness to the treatment of a subject with unknown responsiveness.
[0125] In some embodiments, the machine learning algorithm outputs a resistance score. In some embodiments, the output resistance score is scaled from 0 to 1. In some embodiments, 1 is completely similar to a non-responder and 0 is completely similar to a responder. In some embodiments, the machine learning algorithm calculates similarity to a responder. In some embodiments, the machine learning algorithm calculates similarity to a non-responder. In some embodiments, the machine learning algorithm outputs a numerical value of similarity to a responder and a non-responder. In some embodiments, a protein is considered to be a RAP if its resistance score exceeds a certain threshold. In some embodiments, the threshold for the resistance score is calculated on a scale of 0 to 1. In some embodiments, the threshold for the resistance score for a particular protein is 0.2 to 0.95. In some embodiments, the threshold for a resistance score for a particular protein is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold for a resistance score is 0.25. In some embodiments, the threshold for a resistance score is 0.42. In some embodiments, the threshold for a resistance score is 0.6. In some embodiments, the threshold for a resistance score as calculated by a machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.25. In some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.42. In some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.6.
[0126] In some embodiments, the response probability is determined by the calculation (1 - resistance score). In some embodiments, 1 - resistance score is 1 - overall resistance score. In some embodiments, the resistance score is the overall resistance score. In some embodiments, the response probability is the response score. In some embodiments, the machine learning algorithm outputs a response score. In some embodiments, the output response score is scaled from 0 to 1. In some embodiments, 1 is completely similar to a responder and 0 is completely similar to a non-responder. In some embodiments, the machine learning algorithm calculates similarity to responders. In some embodiments, the machine learning algorithm calculates similarity to non-responders. In some embodiments, the machine learning algorithm outputs numerical values of similarity to responders and non-responders. In some embodiments, a protein is considered to be a RAP if its response score exceeds a certain threshold. In some embodiments, a protein is considered to be an active RAP if its response score exceeds a certain threshold. In some embodiments, the threshold for the response score is calculated on a scale of 0 to 1. In some embodiments, the threshold for a response score for a particular protein is between 0.2 and 0.95. In some embodiments, the threshold for a response score for a particular protein is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold for a response score is 0.25. In some embodiments, the threshold for a response score is 0.276. In some embodiments, the threshold for a response score is 0.42. In some embodiments, the threshold for a response score is 0.5. In some embodiments, the threshold for a response score is 0.6.In some embodiments, the threshold for the response score when calculated by the machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold for the response score when calculated by the machine learning algorithm is 0.25. In some embodiments, the threshold for the response score is 0.276. In some embodiments, the threshold for the response score when calculated by the machine learning algorithm is 0.42. In some embodiments, the threshold for the response score when calculated by the machine learning algorithm is 0.5. In some embodiments, the threshold for the response score when calculated by the machine learning algorithm is 0.6. In some embodiments, the algorithm outputs a response probability, and the response probability is calculated on a scale of 0 to 1. In some embodiments, the algorithm outputs a response probability, where the response probability is calculated on a scale of 0 to 10. In some embodiments, the algorithm outputs a response probability, where the response probability is calculated on a scale of 0% to 100%, where 100% is a complete responder and 0% is a complete non-responder. In some embodiments, a response probability greater than 50% indicates that the subject is likely to respond. In some embodiments, a response probability less than 50% indicates that the subject is unlikely to respond. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.25. In some embodiments, proteins with a response score greater than 0.25 are active in the subject. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.5. In some embodiments, proteins with a response score greater than 0.5 are active in the subject. In some embodiments, the algorithm outputs a probability of clinical benefit. In some embodiments, the probability of clinical benefit is calculated on a scale of 0 to 1. In some embodiments, a clinical benefit probability of 0 indicates a 0% likelihood of clinical benefit for the subject.In some embodiments, a Clinical Benefit Probability of 1 indicates a likelihood of Clinical Benefit in 100% of subjects. In some embodiments, the algorithm outputs a Clinical Benefit Probability, where the Clinical Benefit Probability is calculated on a scale of 0 to 10. In some embodiments, a Clinical Benefit Probability of 10 indicates a likelihood of Clinical Benefit in 100% of subjects. In some embodiments, the algorithm outputs a Clinical Benefit Probability, where the Clinical Benefit Probability is calculated on a scale of 0% to 100%. In some embodiments, a 100% Clinical Benefit Probability indicates a likelihood of Clinical Benefit in 100% of subjects. In some embodiments, a 0% Clinical Benefit Probability indicates a likelihood of Clinical Benefit in 0% of subjects. In some embodiments, a Clinical Benefit Probability in a subject greater than 50% indicates that the subject should continue or discontinue treatment. In some embodiments, the treatment is a monotherapy. In some embodiments, the treatment is a combination therapy. In some embodiments, the threshold for the Clinical Benefit Probability is the median Clinical Benefit Probability in the development set. In some embodiments, the threshold for the probability of clinical benefit is the median probability of clinical benefit in the development set, where a probability of clinical benefit greater than the median probability of clinical benefit is a responder, and a probability of clinical benefit less than the median probability of clinical benefit is a non-responder. In other embodiments, a response probability or clinical benefit probability greater than 50% indicates that the subject will respond to the treatment. In other embodiments, a response probability or clinical benefit probability less than 50% indicates that the subject will not respond to the treatment. In some embodiments, the response probability or clinical benefit probability is between 0 and 10, where a response probability or clinical benefit probability greater than 5 indicates that the subject will respond to the treatment. In some embodiments, the response probability or clinical benefit probability is between 0 and 10, where a response probability or clinical benefit probability less than 5 indicates that the subject will not respond to the treatment.
[0127] In some embodiments, the score is between 0 and 1. In some embodiments, the activity is active in cancer. In some embodiments, the activity is active in the subject. In some embodiments, the activity is active in promoting resistance. In some embodiments, above a threshold is below a threshold. In some embodiments, above a threshold is above a threshold. In some embodiments, the predetermined threshold is 0.5, 0.4, 0.3, 0.25, 0.2, 0.15, 0.1, 0.05, 0.01, 0.005, 0.001, 0.0005, or 0.0001. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold is 0.05. In some embodiments, the threshold is 5%. In some embodiments, the number of active RAPs are combined to provide a total number of RAPs active in the subject. In some embodiments, the number of active RAPs is linearized to provide a total score between 0 and 1. In some embodiments, the linearized are linearly scaled. In some embodiments, the linearizing comprises linear regression. In some embodiments, the number of active RAPs is converted into an overall score of 0 to 1.
[0128] In some embodiments, the predetermined threshold is determined by performing cross-validation within a training set. In some embodiments, the predetermined threshold is the median score in the training set. In some embodiments, the predetermined threshold is the score that best distinguishes between responders and non-responders in the training set.
[0129] In some embodiments, the machine learning algorithm outputs a resistance score. In some embodiments, the resistance score is a RAP score. In some embodiments, the output resistance score is scaled from 0 to 1. In some embodiments, 1 is completely similar to a non-responder and 0 is completely similar to a responder. In some embodiments, for a response score, 1 is completely similar to a responder and 0 is completely similar to a non-responder. In some embodiments, the machine learning algorithm calculates similarity to a responder. In some embodiments, the machine learning algorithm calculates similarity to a non-responder. In some embodiments, the machine learning algorithm outputs a numerical value of similarity to a responder and a non-responder. In some embodiments, a protein is considered to be a RAP if its resistance score exceeds a certain threshold. In some embodiments, the threshold for the resistance score is calculated on a scale of 0 to 1. In some embodiments, the threshold for the resistance score for a particular protein is 0.2 to 0.95. In some embodiments, the threshold for a resistance score for a particular protein is about 0.01, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold for a resistance score is 0.25. In some embodiments, the threshold for a resistance score is 0.42. In some embodiments, the threshold for a resistance score is 0.6. In some embodiments, the threshold for the resistance score when calculated by a machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold for the resistance score when calculated by a machine learning algorithm is 0.25.In some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.42, hi some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.6.
[0130] In some embodiments, the probability of response is determined by the calculation (1 - resistance score). In some embodiments, 1 - resistance score is 1 - overall resistance score. In some embodiments, the resistance score is the overall resistance score. In some embodiments, the probability of response is the response score. In some embodiments, the machine learning algorithm outputs a response score. In some embodiments, the output response score is scaled from 0 to 1. In some embodiments, 1 is completely similar to a responder and 0 is completely similar to a non-responder. In some embodiments, the machine learning algorithm calculates similarity to responders. In some embodiments, the machine learning algorithm calculates similarity to non-responders. In some embodiments, the machine learning algorithm outputs numerical values of similarity to responders and non-responders. In some embodiments, a protein is considered to be a RAP if its response score exceeds a certain threshold. In some embodiments, beyond is above. In some embodiments, above is below. In some embodiments, the threshold for the response score is calculated on a scale of 0 to 1. In some embodiments, the threshold for a response score of a particular protein is between 0.2 and 0.95. In some embodiments, the threshold for a response score of a particular protein is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for a response score is 0.25. In some embodiments, the threshold for a response score is 0.42. In some embodiments, the threshold for a response score is 0.6. In some embodiments, the threshold for the response score as calculated by a machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95.Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.25. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.42. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.6.
[0131] In some embodiments, the calculated resistance scores are combined to generate an overall resistance score. In some embodiments, the calculated response scores are combined to generate an overall response score. Those skilled in the art will understand that response scores and resistance scores are always interchangeable, since they are simply one minus the other. The conversion of resistance to response can be performed at the individual factor level, or at the overall level after combining the scores. In some embodiments, the combination is a sum. In some embodiments, the resistance scores are summed to generate an overall resistance score. In some embodiments, the combination is an average. In some embodiments, the resistance scores are averaged to generate an overall resistance score. In some embodiments, the scores are weighted when combined.
[0132] In some embodiments, the method includes determining the number of factors of a plurality of factors that are active in the subject. In some embodiments, active factors are factors that have a resistance score above a predetermined threshold. In some embodiments, the threshold is 0.25. In some embodiments, factors with a resistance score above 0.25 are active factors in the subject. In some embodiments, the threshold is 0.276. In some embodiments, factors with a resistance score above 0.276 are active factors in the subject. In some embodiments, only active factors are combined. In some embodiments, combining the calculated resistance scores involves combining active resistance scores. In some embodiments, combining includes adding the number of factors that are active in the subject. In some embodiments, the number of factors that are active in the subject is converted to a score of 0 to 1. In some embodiments, the number of factors that are active in the subject is converted to a score of 0 to 10. In some embodiments, converting includes applying a linear regression model. In some embodiments, the number of active factors is linearized to provide a total score of 0 to 1. In some embodiments, the number of active factors is linearized to provide a total score of 0 to 10. In some embodiments, the linearizing comprises linear scaling. In some embodiments, the linearizing comprises linear regression. In some embodiments, the threshold is 5.
[0133] In some embodiments, the machine learning model is a machine learning algorithm. In some embodiments, the algorithm is a supervised learning algorithm. In some embodiments, the algorithm is an unsupervised learning algorithm. In some embodiments, the algorithm is a reinforcement learning algorithm. In some embodiments, the machine learning model is a convolutional neural network (CNN). In some embodiments, at least one hardware processor trains the machine learning model. In some embodiments, the model is based, at least in part, on a training set. In some embodiments, the model is based on a training set. In some embodiments, the model is trained on a training set. In some embodiments, the at least one hardware processor applies the machine learning model to expression levels of factors from the subject.
[0134] In some embodiments, calculating comprises calculating the mean expression for each protein in responders. In some embodiments, calculating comprises calculating the mean expression for each protein in non-responders. In some embodiments, calculating comprises calculating the mean expression for each protein in responders and the mean expression for each protein in non-responders. In some embodiments, calculating comprises calculating a distribution of expression for each protein in responders and non-responders. In some embodiments, calculating comprises calculating the standard deviation of expression for each protein in responders and non-responders. In some embodiments, �� in responders is �� in the responder population. In some embodiments, �� in non-responders is �� in the non-responder population. In some embodiments, the resistance score is based on the ratio of the deviation of the expression of the factor in the subject from the calculated mean in responders to the deviation of the expression of the factor in the subject from the calculated mean in non-responders. Calculation of the deviation is familiar to those skilled in the art. It will be understood that the further the expression in a subject is from the mean, the greater the deviation. Thus, factors that are very different from the mean in responders have a large numerator in the calculation of this ratio, and factors that are not very different from the mean in non-responders have a small denominator. Thus, the more dissimilar the expression of a factor in a subject is to responder expression and the more similar it is to non-responder expression, the higher the resistance score. In some embodiments, a resistance score above a predetermined threshold indicates that the factor is a factor associated with resistance. In some embodiments, the factor associated with resistance is a resistance-associated protein (RAP). In some embodiments, the factor associated with resistance is a RAP if its expression in responders is statistically different from its expression in non-responders.
[0135] In some embodiments, the calculating further comprises calculating a distribution for each factor in responders. In some embodiments, the calculating further comprises calculating a distribution for each factor in non-responders. In some embodiments, the calculating further comprises calculating a distribution for each factor in responders and a distribution for each factor in non-responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in non-responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in responders and a standard deviation for each factor in non-responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in a mix of responders and non-responders. In some embodiments, the deviation is measured as a multiple of the calculated standard deviation. One skilled in the art will appreciate that by scaling the deviation to the standard deviation for a group of expression values, the deviation is provided more absolutely, allowing for comparison of factors with very small and very large standard deviations (and which may have very low and very high expression levels).
[0136] In some embodiments, the resistance score is based on the Z-score of the expression level for each factor in the subject. In some embodiments, the resistance score is based on the Z-score for responders. In some embodiments, the resistance score is based on the Z-score for non-responders. In some embodiments, the resistance score is based on both the Z-score for responders and the Z-score for non-responders. In some embodiments, the resistance score is based on the ratio of the Z-score for responders to the Z-score for non-responders. It is well known to those skilled in the art that the Z-score takes into account the distance of an individual level from the population mean in units of population standard deviation. In some embodiments, the Z-score is calculated according to Formula 1.
[0137] In some embodiments, the resistance score is calculated using the formula (|Z R | / (|z NR In some embodiments, Z R is the deviation of expression of the factor of interest from the mean value calculated in the responders. In some embodiments, Z NR is the deviation of the expression of the factor in the subject from the calculated mean value in non-responders. In some embodiments, || is the Z score of the deviation. In some embodiments, || is the standardization of the deviation to a multiple of the standard deviation. In some embodiments, c is a constant. In some embodiments, the constant is the Z NR = 0. In some embodiments, the tolerance score is calculated according to Equation 2. In some embodiments, monotonoic is an ad-hoc function that prevents the tolerance score from decreasing for extreme values in the non-responder distribution. In some embodiments, the function is the function provided in Algorithm 1.
[0138] In some embodiments, a resistance score above a predetermined threshold indicates that the agent is RAP. In some embodiments, exceed is above. In some embodiments, the threshold is a predetermined threshold. In some embodiments, the threshold is a threshold value. In some embodiments, the threshold for the resistance score is about 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, 5.0, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 8.1, 8.1, 8.2, 8.3, 8.4, 8.5, 8.6, 8.7, 8.8, 8.9, 9.1, 9.2, 9.3, 9.4, 9.5, 9.6, 9.7, 9.8, 9.9, 10.1, 10. In some embodiments, the threshold is about 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.67, 0.7, 0.75, 0.8, 0.85, or 0.9. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold for the resistance score is about 2.9. In some embodiments, the threshold for the resistance score is 2.9. In some embodiments, the threshold for the resistance score is about 3.0. In some embodiments, the threshold for the resistance score is 3.0. In some embodiments, the threshold for the resistance score is calculated on an arbitrary unit scale. In some embodiments, the threshold for the resistance score as calculated by mathematical calculation is about 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, or 5.0. Each possibility represents a separate embodiment of the present invention.In some embodiments, the threshold for the resistance score when calculated in a mathematical calculation is about 2.9. In some embodiments, the threshold for the resistance score when calculated in a mathematical calculation is 2.9. In some embodiments, the threshold for the resistance score when calculated in a mathematical calculation is about 3.0. In some embodiments, the threshold for the resistance score when calculated in a mathematical calculation is 3.0. In some embodiments, the method wherein the mathematical calculation comprises calculating an average expression value for each protein.
[0139] In some embodiments, subjects with a number of resistance-associated factors (e.g., RAP) greater than a predetermined number are predicted to be resistant to the treatment. In some embodiments, subjects with a number of resistance-associated factors greater than a predetermined number are predicted to not respond to the treatment. In some embodiments, subjects with a number of resistance-associated factors greater than a predetermined number are predicted to be non-responders to the treatment. In some embodiments, subjects with a number of resistance-associated factors less than a predetermined number are predicted to be amenable to the treatment. In some embodiments, subjects with a number of resistance-associated factors less than a predetermined number are predicted to respond to the treatment. In some embodiments, subjects with a number of resistance-associated factors less than a predetermined number are predicted to be responders to the treatment. In some embodiments, subjects with a number of resistance-associated factors equal to or less than a predetermined number are predicted to be amenable to the treatment. In some embodiments, subjects with a number of resistance-associated factors equal to or less than a predetermined number are predicted to respond to the treatment. In some embodiments, subjects with a number of resistance-associated factors equal to or less than a predetermined number are predicted to be responders to the treatment.
[0140] In some embodiments, the predetermined number is a threshold number. In some embodiments, the predetermined number is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. Each possibility represents a separate embodiment of the present invention. In some embodiments, the predetermined number is 3. In some embodiments, the predetermined number is 4. In some embodiments, the predetermined number is 7. In some embodiments, the predetermined number is 13.
[0141] In some embodiments, the method further comprises classifying the resistance-associated factors into at least one pathway, process, or network. In some embodiments, the method further comprises performing an analysis of the resistance-associated factors to determine at least one pathway, process, or network in which the resistance-associated factors are involved. In some embodiments, the pathway, process, or network causes non-responsiveness to treatment. In some embodiments, the analysis is selected from pathway analysis, process analysis, and network analysis. In some embodiments, the method further comprises performing a pathway analysis on the RAP. In some embodiments, the method further comprises performing a process analysis on the RAP. In some embodiments, the method further comprises performing a network analysis on the RAP. In some embodiments, the at least one pathway, process, or network comprises at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 pathways, processes, or networks. Each possibility represents a separate embodiment of the present invention. In some embodiments, the at least one pathway, process, or network is all pathways, processes, or networks known to include resistance-associated factors. In some embodiments, the at least one pathway, process, or network is all pathways, processes, or networks enriched for resistance-associated factors. In some embodiments, enriched is most enriched. In some embodiments, enriched includes the most RAPs of either a pathway, process, or network.
[0142] In some embodiments, the method includes selecting a pathway, process, or network. In some embodiments, the selected pathway, process, or network is hypothesized to influence non-responsiveness to treatment. In some embodiments, the selected pathway, process, or network is hypothesized to cause non-responsiveness to treatment. In some embodiments, the selected pathway, process, or network is known to be druggable. In some embodiments, the known druggable includes a known therapeutic agent that modulates the pathway, process, or network. In some embodiments, the known therapeutic agent is in or has completed clinical trials. In some embodiments, the known therapeutic agent is approved for use in humans. In some embodiments, approved for use in humans means approved for use in treating a human disease. In some embodiments, the disease is cancer. In some embodiments, the method further includes administering to a subject who is or is predicted to be a non-responder an agent that modulates at least one pathway, process, or network comprising a factor associated with resistance. In some embodiments, the agent inhibits a target in the pathway, process, or network. In some embodiments, the target is a gene. In some embodiments, the target is a protein. In some embodiments, the protein is a regulatory RNA. In some embodiments, the target is a response-associated factor. In some embodiments, the target is not a response-associated factor. In some embodiments, the agent activates a target in a pathway, process, or network. In some embodiments, the agent modulates a pathway, process, or network. In some embodiments, pathway activity induces unresponsiveness and the agent inhibits the pathway. In some embodiments, pathway activity reduces unresponsiveness and the agent activates the pathway. One of skill in the art will understand that a response-associated factor is identified by its expression in a subject being more similar to its expression in a non-responder than in a responder.Thus, for example, if a factor is more highly expressed in non-responders and increases the activity of a pathway / process / network, the agent would inhibit the pathway. For example, if the factor is more highly expressed in non-responders but decreases the activity of a pathway / process / network, the agent would activate the pathway / process / network. Similarly, if the factor is less expressed in non-responders and decreases the activity of a pathway / process / network, the agent would inhibit the pathway / process / network. And finally, for example, if the factor is less expressed in non-responders but increases the activity of a pathway / process / network, the agent would activate the pathway / process / network. Essentially, the agent must induce the pathway / process / network to function better than it does in responders. In some embodiments, the agent targets a hub target in the pathway. In some embodiments, the agent targets a regulator target in the pathway. In some embodiments, activity of the process induces non-responsiveness, and the agent inhibits the process. In some embodiments, activity of the process reduces unresponsiveness and the agent activates the process. In some embodiments, the agent targets a hub target in the process. In some embodiments, the agent targets a regulator target in the process. In some embodiments, network activity induces unresponsiveness and the agent inhibits the network. In some embodiments, network activity reduces unresponsiveness and the agent activates the network. In some embodiments, the agent targets a hub factor in the network. In some embodiments, the agent targets a regulator factor in the network. In some embodiments, the regulator is a master regulator. Factors can be classified into pathways, protein interactions, or signals using any analytical tool known in the art.Examples include, but are not limited to, GO analysis, Ingenuity analysis, Metacore analysis (Clarivate Analytics), reactome pathway analysis, and functional analysis.
[0143] According to another aspect, there is provided a computer program product comprising a non-transitory computer-readable storage medium having program code embodied thereon, said program code being executable by at least one hardware processor to perform the method of the present invention.
[0144] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0145] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device having instructions recorded thereon, and any suitable combination of the above. As used herein, a computer-readable storage medium is not to be construed as a transitory signal itself, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted over a wire. Rather, the computer-readable storage medium is a non-transitory (ie, non-volatile) medium.
[0146] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.
[0147] Computer-readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry to perform aspects of the present invention.
[0148] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine whereby the instructions, executing via the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts identified in the flowchart and / or block diagram blocks. These computer-readable program instructions may also be stored on a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other apparatus to function in a particular manner, whereby a computer-readable storage medium having instructions stored therein includes an article of manufacture containing instructions that implement aspects of the functions / acts identified in the flowchart and / or block diagram blocks.
[0149] The computer-readable program instructions may also be loaded into a computer, other programmable data processing device, or other device that causes a series of operational steps to be performed on the computer, such that the instructions implemented on the computer, other programmable device, or other device may perform the functions / acts identified in the blocks of the flowcharts and / or block diagrams. As used herein, the term "about" when combined with a value represents ±10% of the reference value. For example, a length of about 1000 nanometers (nm) represents a length of 1000 nm ±100 nm.
[0150] It should be noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to a "polynucleotide" includes a plurality of such polynucleotides; a reference to a "polypeptide" includes a reference to one or more polypeptides and equivalents thereof known to those skilled in the art; and so forth. It should be further noted that the claims may be drafted to exclude any element. As such, this specification is intended to serve as a basis for the antecedent use of exclusive terminology, such as "solely," "only," and the like, or for the use of a "negative" limitation in connection with the recitation of claim elements.
[0151] When traditional language similar to "at least one of A, B, and C, etc." is used, such syntax is generally intended in the sense that one of ordinary skill in the art would understand the language (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Those of ordinary skill in the art will further understand that virtually any conjunction word and / or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" is understood to include the possibilities of "A" or "B," or "A and B."
[0152] It is understood that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination. All combinations of embodiments according to the present invention are specifically embraced by the present invention and are disclosed herein as if each and every combination were individually and explicitly disclosed. Furthermore, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein as if each and every such subcombination were individually and explicitly disclosed herein.
[0153] Additional objects, advantages, and novel features of the present invention will become apparent to those skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as described hereinabove and as claimed in the claims section below finds experimental support in the following examples.
[0154] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.
[0155] Example Generally, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular biological, biochemical, microbiological, and recombinant DNA techniques. Such techniques are fully explained in the literature. See, e.g., "Molecular Cloning: A Laboratory Manual" by Sambrook et al. (1989); "Current Protocols in Molecular Biology" Volumes I-III, Ausubel, R.M., ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology," John Wiley & Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning," John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA," Scientific American Books, New York; and Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series," Vols. 1-4, Cold Spring Harbor Laboratory Press, New York. (1998); methods described in U.S. Patent Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I-III, Cellis, JE, ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley-Liss, NY (1994), Third Edition; "Current Protocols in Immunology", Volumes I-III, Coligan, JE, ed. (1994); Stites et al.(eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual" CSHL Press (1996), all of which are incorporated by reference. Other general references are provided throughout this book.
[0156] material and method Plasma samples and clinical data were collected from 610 patients with advanced NSCLC undergoing ICI-based treatment at 20 participating medical centers. Patient cohort and specimen collection: Comprehensive clinical data were collected for each patient and verified by comparison with source references. All patients were treated with an ICI-based regimen, including ICI monotherapy (pembrolizumab, atezolizumab, or nivolumab), ICI and chemotherapy combination therapy (pembrolizumab / atezolizumab plus chemotherapy), or ICI combination therapy (ipilimumab plus nivolumab). Inclusion criteria were informed consent; age >18 years; stage IIIB-IV NSCLC; ECOG performance status 0-2; and normal hematologic, renal, and hepatic function. Additional exclusion criteria were any concurrent and / or other active malignancies requiring systemic treatment within 2 years prior to the first dose of ICI-based treatment. The overall cohort size was set once performance stabilized in the development set.
[0157] Specimens were collected before treatment initiation, immediately prior to the first treatment dose on the same day (n=244), or within 10 days (n=52), 11-30 days (n=38), or 31-58 days (n=5) before treatment initiation (this number represents the cohort after patient exclusion; see the Patient Exclusion section for further details). Specimen collection was performed as follows: blood samples were collected from each patient into EDTA clot tubes; plasma was isolated from whole blood by centrifugation at 1200 × g for 10-20 minutes at room temperature within 4 hours of venipuncture; plasma supernatants were collected, stored frozen at -80°C, and shipped frozen to the laboratory for analysis.
[0158] A separate retrospective cohort of 85 patients receiving chemotherapy was included for specific comparisons. In addition to the ICI-based cohort, a retrospective cohort of patients receiving chemotherapy as monotherapy was recruited. Samples were collected between September 2015 and October 2018 using the same protocol. Inclusion criteria: advanced-stage NSCLC receiving first-line chemotherapy without a change to ICI treatment or the addition of an ICI to the treatment regimen. For comparisons between the ICI-based treatment and chemotherapy cohorts, patient baseline characteristics were compared between the ICI-based development and chemotherapy sets using chi-square tests for categorical data and t-tests for continuous variables.
[0159] Assessment of therapeutic benefit: Clinical benefit data were retrieved from patients' medical records and verified by the investigator through review of radiological images, i.e., chest / abdominal CT and brain MRI, performed every 2-3 months, based on RECIST (Response Evaluation Criteria In Solid Tumors) 1.1. Clinical benefit (CB) was also assessed based on progression-free survival (PFS) at 12 months after the start of treatment. Therapeutic benefit was assessed based on progression events at 12 months. Because patients may reach the 12-month clinical assessment around 12 months, we chose to examine the range from 330 to 400 days after the start of treatment as follows: patients were assigned as having clinical benefit (CB) if a progression event was determined beyond 400 days or if there was no progression by 330 days; patients were assigned as not having clinical benefit (NCB) if there was a progression event by 400 days (inclusive) after the start of treatment; patients were considered "not a clinical benefit signifier" and excluded from the classifier development or validation process if there was no progression event by 330 days (inclusive).
[0160] Alternatively, treatment benefit was assessed at 3, 6, and 12 months after the start of treatment, and patients were classified as having clinical benefit (CB) or not having clinical benefit (NCB) at each time point. Patients who demonstrated a complete response, partial response, or stable disease at 3 and 6 months were classified as CB patients, and patients who developed progressive disease or died were classified as NCB patients. Sustained clinical benefit was assessed at 12 months after the start of treatment. Patients who were confirmed to be alive and free of progressive disease for at least 12 months after starting treatment were classified as CB patients. Patients who discontinued treatment before the 12-month criterion due to treatment-related adverse events (but who showed no signs of progression for at least 12 months) were also classified as CB patients. All other patients were classified as NCB patients. All patients were followed for at least 2 years. The time at which progressive disease and / or death occurred was recorded. If treatment was changed because of treatment-related adverse events or the patient's refusal to continue treatment, only patients who received two or more cycles of ICI remained in the study. Patients who discontinued chemotherapy but continued ICI therapy remained in the study.
[0161] Proteomic measurements, data normalization, and quality control: Proteomic profiling of plasma samples was performed using an assay that simultaneously measures approximately 7,000 protein targets. The assay is based on chemically modified single-stranded oligonucleotides that fold into molecular structures that can bind to proteins with high affinity and specificity. Measurements are performed using DNA microarray technology and provide a readout in relative fluorescence units (RFU). The assay simultaneously measures a total of 7,596 protein targets, of which 7,289 targets are human proteins.
[0162] Cohort samples were run in two running batches. Each sample was profiled once. Quality control and normalization were performed. Because the distribution of protein levels is approximately log-normal (i.e., the logarithm of the measurements is normally distributed), and given that many statistical methods assume normality, a log2 transformation was applied unless otherwise noted. There was no data imputation in model development and validation. If a patient had a "not available" (NA) data entry for a clinical parameter, the entry was treated as NA.
[0163] The proteomic dataset was narrowed down to a set of proteins with high analytical confidence by comparing the proteomic dataset of the current cohort with that of another cohort not participating in this study. For each assayed protein, the expression level distribution was compared between the two cohorts by applying the Kolmogorov-Smirnov test. Proteins with p-values less than 0.05 were excluded, resulting in 1578 proteins for model development.
[0164] Model development and validation was performed on patients receiving ICI-based therapy with clinical benefit assessment. Models were constructed in a development set (n=228) and tested in a blinded manner in an unrelated validation set (n=272).
[0165] Patient Exclusions:
[0166] Of the 610 patients enrolled, 65 were excluded for technical or clinical reasons (Supplementary Figure 1). Ten patients failed the SomaScan® quality check or had missing measurements. Samples from 13 patients were excluded because they were not collected within the defined time frame for blood collection (blood was collected more than 2 months before treatment or after ICI-based treatment). Thirty-two patients were excluded due to treatment-related issues (untreated; not naive to immunotherapy; received chemotherapy less than 60 days before ICI treatment; received ipilimumab in combination with nivolumab; the latter group in particular was excluded because this is a different treatment compared to anti-PD-L1 with or without chemotherapy, and 16 patients in this category were not enough for robust analysis. Future studies will include these patients if the group size is large enough. Ten patients were excluded due to eligibility issues (ECOG >2; psychiatric disorder, driver mutation, multiple cancer types). After patient exclusion, 545 patients remained in the dataset.
[0167] For each downstream analysis, a different exclusion of the ICI-based cohort was performed; for the development and validation of the PROphet model, the analysis required patients with clinical benefit assessment and only first-line ICI treatment in the validation set, leaving 500 patients for this analysis; for the analysis that included a combination of PD-L1 expression level and PROphet, patients without PD-L1 assessment and patients with advanced / unknown treatment preference were excluded, leaving 441 patients on ICI-based therapy in the analysis.
[0168] Resistance-associated protein (RAP) model development and validation: To avoid data leakage, the cohort was divided into a development set and a validation set. The model was constructed in the development set (n = 228). After the final model was constructed, blinded validation was performed in the validation set (n = 272). After model development was completed and the model configuration was fixed, proteomic and clinical data were obtained for an additional 272 patients comprising the validation set, and blinded validation was performed on this sample set. Notably, the data for the validation set were not available at the time of model development; although this method ensured that the validation was fully blinded, it was not possible to ensure that the distributions of various clinical parameters were similar between the development and validation sets. In fact, few clinical parameters (gender, ECOG, PD-L1 expression level, age) showed statistically significant differences between the development and validation sets. It is important to emphasize that, although similar distributions of clinical parameters are desirable, they cannot be achieved when constructing the validation set after model development is complete. Furthermore, models are expected to perform best when applied to similar populations; therefore, differences between the development and validation sets place additional stress on the model. The same division into development (n = 228) and validation (n = 272) sets described above was applied to the PD-L1-based predictive model and the predictive model. To improve performance, the PD-L1 model was based on PD-L1 numerical values rather than categorical values (i.e., PD-L1 ≥ 50%, PD-L1 1%–49%, PD-L1 < 1%); because not all samples had numerical values, 210 and 204 patients were included in the development and validation sets of this model, respectively.
[0169] The models were developed using a random sampling approach with multiple iterations. In each iteration, the development set was randomly divided into a training set and a test set (75% and 25% of the development set, respectively). In each iteration, the training set was used for feature selection and model training as follows: Proteins showing differential levels between CB and NCB patients were identified using the Kolmogorov-Smirnov test. Single-protein predictive models were built on the iterative training set for each of the 50 proteins with the lowest p-value (i.e., 50 independent models were built, each based on a single protein). The XGBoost algorithm was used to build models for each single protein using two features: protein expression level and patient gender. Gender was included as a feature in the model because it affects the protein's plasma expression level (thus, the model is not biased toward the majority of patients who are male). The output of each single-protein model is a probability between 0 and 1; the lower the probability, the more likely the patient is to show clinical benefit. The overall patient score for each iteration was extracted by summing the number of single-protein models that exhibited NCB, where a single-protein model exhibited NCB if its predicted probability exceeded the observed cohort CB rate (=0.276). This methodology resulted in the output of the iterative model being an integer between 0 and 50, with numbers closer to 0 corresponding to CB and numbers closer to 50 corresponding to NCB. The steps described above for one iteration were repeated for 80 iterations, where the overall patient outcome was the average of the outcomes from the 80 iterations. Finally, the model score was linearly scaled from 0 to 10, with values below 5 indicating a negative outcome and values above 5 indicating a positive outcome.
[0170] Model performance was evaluated blindly on an independent validation set using two metrics: (i) agreement between the predicted probability of CB and the observed CB rate in terms of goodness of fit (R2 of linear regression), where the observed CB rate for each CB value was defined as the proportion of CB patients in the group of patients within a window of CB probability ±0.05. (ii) by examining the hazard ratio (HR) for the positive versus negative population, calculated using a Cox proportional hazards model. Additional predictive models: To maintain consistency with the RAP model, all predictive models described in this study underwent a similar development pipeline. First, the same development and validation sets used for the RAP model were used for the other models. Next, the development set was randomly split 80 times into training and test sets (75% and 25% of the development set, respectively). In each iteration, a model was developed on the training set using the XGBoost algorithm, and predictions were inferred on the test set. Predictions from all iterations were averaged and returned as the probability of CB. For the PD-L1-based model, PD-L1 status was the only input (high, low, or negative). For the clinical model, four clinical parameters were used as inputs: (i) PD-L1 status (high, low, or negative); (ii) ECOG performance status; (iii) patient gender; and (iv) treatment choice (primary or advanced). The integrated model (i.e., the RAP model combined with another model) was developed in two steps. In the first step, the RAP model was developed as described above. In the second step, the output of the RAP model, along with the relevant clinical parameters, served as input features. The development set was again split into training and test sets 80 times, each time creating a new split into training and test sets, and the predicted values from all iterations were averaged. The model output was the probability of CB. Performance evaluation and comparison were performed using ROC curves and linear regression between the predicted probability of CB and the observed CB rate, as described above.
[0171] Data Analysis: All data analyses were performed using Python, the Perseus computing platform, and GraphPad Prism (San Diego, CA, USA; graphpad.com). Multivariate Cox proportional hazards regression with stepwise model reduction was used to obtain hazard ratios for treatment effect adjusted for all other factors and to assess interactions between treatment and predictor classes. Factors initially found to have an effect on the hazard ratio were also tested for interactions with treatment. Hazard ratios are reported with 95% confidence intervals and p-values. A level of 0.05 or less was considered significant. R statistical software was used for analysis with the Survival and MASS packages. For the analysis of overall survival and progression-free survival, 444 patients treated with first-line ICIs whose PD-L1 levels were determined were examined, along with the chemotherapy cohort (n = 85).
[0172] Associations between CBs and clinical parameters were assessed using chi-squared tests for categorical parameters and t-tests for numerical parameters. A network of RAPs was constructed based on the STRING database. Voronoi plots of proteins within each consensus cluster were plotted using Proteomaps. Enrichment analysis of CB probability values was performed using a 2D enrichment test (false positive rate <0.05)
[37] . Enrichment analysis of RAPs selected with at least 10 replicates was performed using Fisher's exact test (false positive rate <0.1) against a comprehensive background of 1578 tested proteins. Enrichment analysis of RAP functionality was performed using different replicate cutoffs with similar results. Protein categories were based on the Human Protein Atlas (proteinatlas.org), CHAT, the ECM maristome project, and UniProt (keywords).
[0173] Statistical Analysis: Log-rank tests and multivariate Cox proportional hazards regression tests were used to determine hazard ratios for treatment effects, taking into account predictor class and adjusting for the effects of other patient covariates.
[0174] Example 1: Response prediction based on resistance-associated proteins (RAPs) - proof of concept Data collection The proof-of-concept for response prediction was based on the analysis of blood samples from 108 non-small cell lung cancer (NSCLC) patients treated with immune checkpoint inhibitors (ICIs). The various treatments performed are summarized in Table 1.
[0175] [Table 1]
[0176] Plasma protein levels were measured in 108 patients, measuring approximately 1100 non-redundant protein targets. Samples were taken before the start of ICI treatment (T0) and after the first treatment run (T1), for a total of 156 samples in batches.
[0177] Building a Classifier Proteomic levels and response markers were incorporated using a supervised learning algorithm to predict response to treatment. Response markers were responder (R) and non-responder (NR), determined based on overall response rate (ORR) assessment at 3 months. Specifically, progressive disease (PD) or early death related to disease progression was classified as NR. Stable disease (SD), minimal response (MR), partial response (PR), and complete response (CR) were classified as R. ORR assessment was performed using RECIST 1.1 or other validated methods for ORR assessment, as described in the "Primary Outcome Measures" section of clinical trial NCT04056247 (clinicaltrials.gov / ct2 / show / NCT04056247, incorporated herein by reference in its entirety). Changes in blood levels of different proteins representing the host response [time frames: at baseline (pre-treatment, T0) and after the first treatment (post-treatment, T1)] were determined as described.
[0178] Samples were divided into training set and test set.All stages of algorithm development were carried out using training set, while test set was only used in the final stage to test the performance of the final algorithm.Training set included samples from n=78 patients (59 responders and 19 non-responders), and test set included samples analyzed from n=30 patients.
[0179] The response classifier takes features as input and predicts response based on the feature values. Features are protein levels measured in plasma at two time points: baseline (T0) and after the first treatment (T1). Measurements of the same protein at different time points are considered independent features. Furthermore, some proteins have more than one measurement in a single proteomic profile (e.g., the protein IL-6 is measured four times). Each repetition was treated as an independent feature.
[0180] Resistance-associated proteins Resistance-associated protein (RAP) refers to a specific protein whose expression in a given patient leads to resistance to treatment, i.e., RAP is patient-specific. A protein is considered to be a RAP if its expression level in each patient is more similar to its expression distribution in a non-responder population than in a responder population (see Figure 1A-1C for illustration). RAP can be determined in various ways. This document provides mathematical calculation of RAP, machine learning algorithms for classifying RAP, and a method that combines both. These methods are merely exemplary, and any method for calculating RAP can be used.
[0181] To quantitatively define the above concept, a RAP score (i.e., resistance score) was determined for each protein. A low RAP score value represents the expression level typical of a responder population, and a high RAP score represents the expression level typical of a non-responder population. A protein is considered to be a RAP if its RAP score exceeds a certain threshold (e.g., is greater than or less than, depending on the score configuration). The RAP score threshold optimization process is described herein below.
[0182] Calculating the RAP score requires knowing the expression level distribution of each protein in the responder and non-responder populations, and the protein level expression data of the tested patients.To allow comparison between multiple different proteins with different ranges of expression levels, it is important that the RAP score is not affected by or sensitive to the protein level expression scale.This is particularly important for plasma samples, where there is a large dynamic range of 11 orders of magnitude in protein expression levels.To achieve this, the RAP score is based on the Z score, which measures the distance of each level from the population mean value in units of the population standard deviation.In technical terms, the Z score is defined by Equation 1: Equation 1: Z = (x - μ) / σ where x is the protein level in the patient being tested, μ is the mean protein level in the population, and σ is the standard deviation of the population. The Z-score for a given patient is calculated separately for the responder and non-responder populations. Z R For calculation of the Z-score for the responder population, denoted by Z, the distribution measures, μ and σ, are calculated using the responder population. NR For calculation of the Z-score for the non-responder population, denoted by , the distribution measures, μ and σ, are calculated using the non-responder population. Finally, the RAP score is defined by 2. Formula 2: monotonic(|Z R | / (|z NR |+c)) In the formula, c is Z NR is a regularization constant that prevents score deviations at =0, and monotonous is an ad-hoc function designed to prevent RAP scores from decreasing for extreme values in the non-responder distribution. The function implementation is The pseudocode is provided in Algorithm 1. The RAP score values for representative responder and non-responder distributions are shown in Figure 2.
[0183] Algorithm 1: Monotonic functions used in Eq.
number
[0184] To determine the exact number of RAPs for a given patient, a threshold was determined for all proteins, and proteins with a RAP score above the determined threshold were considered RAPs. The threshold was determined using cross-validation applied to the training set. Specifically, a cross-validation dataset consisting of one-third of the training set and a non-cross-validation dataset consisting of a further one-third of the training set were sampled, while maintaining similar numbers of responders and non-responders between the cross-validation and non-cross-validation datasets. Calculations were performed for each patient in the non-cross-validation set and then in the cross-validation dataset. RAP scores were calculated for all features (i.e., all measured proteins at T0 and T1) using the expression level distributions of responders and non-responders. The number of RAPs was then used to predict response, and the area under the receiver operating characteristic (ROC) curve (AUC), which quantifies the performance of the prediction, was calculated for each threshold (Figures 3A-3B). To minimize noise associated with small datasets, 100 realizations were performed for each threshold (i.e., different samplings of the cross-validation set from the training set), and the average AUC across the 100 realizations was examined. Notably, the average ROC AUC curve in Figure 3A shows a single broad peak, suggesting that the predictive power of RAP counts is not very sensitive to the selected threshold. For features containing measurements at T0 and T1, the thresholds were set at 1.61 (Figure 3B) and 2.9 (Figure 3A), respectively.
[0185] Evaluating Machine Learning: Purely mathematical approaches, while powerful (both conceptually and practically), have some drawbacks that need to be addressed: 1. Because the RAP score function depends on the underlying distribution of protein expression levels, its validity may be platform-dependent (especially since different proteomics systems use different measurement methods and units that do not naturally scale). 2. The current implementation does not provide a natural way to include clinical parameters (patient condition, indication details, treatment details, etc.) in the predictors.
[0186] We have invented another approach that utilizes decision tree learning based on machine learning algorithms to classify proteins as RAPs for a given subject. For each measured protein, a predictive model was created based on the training set data using a machine learning algorithm (e.g., the XGBoost algorithm). Such data from the training set may include not only protein expression levels and responder / non-responder tags, but also other characteristics such as patient age, gender, condition, treatment type, treatment choice, and biomarker expression, such as PD-L1 expression. This approach makes no assumptions about protein distribution and provides a natural framework for utilizing clinical parameters.
[0187] To test this approach, samples from a cohort of 76 patients were screened using two different protein analysis platforms: approximately 1200 proteins (O) and approximately 7500 other proteins (S), with approximately 1000 proteins common to both platforms. The treatments given to these subjects are summarized in Table 2.
[0188] [Table 2]
[0189] The cohort of 76 patients was divided into a training set containing 51 subjects (38 responders and 13 non-responders) and a test set containing 25 subjects (19 responders and 6 non-responders). The XGBoost algorithm was chosen for this analysis due to the non-linear nature of the problem and the algorithm's efficiency in learning on small datasets. To avoid multiple comparisons in the test set, which would increase the risk of false discoveries, and because the purpose of the study was to verify the predictive feasibility (rather than to identify the optimal model configuration), the following predetermined configurations were used for the training model: The hyperparameters of the model were set as follows: a.Maximum tree depth = 4 b. Ridging factors: eta = 0.8, lambda = 5, alpha = 2 c.num_parallel_tree=100 d.purpose=binary:logistic e.eval_metric=logloss The parameters were chosen to handle small, noisy data sets.
[0190] For this evaluation, a machine learning algorithm was trained on protein expression levels alone, excluding other considerations. Patient expression results were evaluated for each protein individually, and a protein classifier was calculated for each single protein. The machine learning algorithm output a score between 0 and 1, with 1 being most similar to non-responders and 0 being most similar to responders.
[0191] Two input protein configurations were used to evaluate this approach. In the first configuration, all proteins were used as potential predictors. This is similar to that used in the mathematical approach; however, while this method is expected to be effective for large cohorts, false positives may hinder predictive power for small cohort sizes (compared to the number of features). In the second configuration, single-protein models were ranked according to their tendency to divide patients into responders and nonresponders (i.e., protein models with more balanced predicted classes were given higher ranks). As an extreme example, if a model predicted that all patients belonged to a single class (responder or nonresponder), this model received the lowest potential balanced rank. On the opposite scale, models that evenly divided the population between responders and nonresponders received the highest balanced rank. After ranking the different protein models, the machine learning approach was evaluated using the 200 proteins with the highest balanced ranks.
[0192] Both methods were used to evaluate subjects based on their "O" and "S" expression values at T0 and T1. The model's performance for "O" (measured by AUC) was above 0.8 for thresholds ranging from 0.4 to 0.8, with stable and smooth behavior (Figure 3C), peaking at AUC = 0.89 and a 95% confidence interval of [0.594, 0.995]. Therefore, for these samples, the threshold was set at approximately 0.6. This result represents a slight improvement over the AUC = 0.846 obtained for the same dataset using the mathematical RAP method. However, due to the large confidence intervals (which are a result of the small dataset size), the statistical significance of this difference is moderate.
[0193] The peak performance of the model for "O" when limiting the predictors to 200 proteins was AUC = 0.91 with a 95% confidence interval of [0.602, 0.996] (Figure 3C). The threshold is essentially the same in this case, and the AUC shows a slight improvement over the full protein set configuration.
[0194] Model performance (as measured by AUC) for "S" was above 0.75 for thresholds ranging from 0.4 to 0.9, with stable and smooth behavior (Figure 3C), peaking at an AUC of 0.81 and a 95% confidence interval of [0.587, 0.924]. Therefore, for these samples, the threshold can be set slightly lower at approximately 0.59, although this difference may be negligible. This result is inferior to the behavior observed from the same model configuration using the "O" data (approximately one standard deviation less), which is not unexpected due to the significantly larger number of proteins and small dataset size.
[0195] Model performance peaked for "S" when restricting predictors to 200 proteins at AUC = 0.87, with a 95% confidence interval of [0.597, 0.992] (Figure 3C). This threshold is therefore essentially the same as that obtained for the "S" analysis, representing a considerable improvement compared to the full protein set configuration, consistent with the reduced false discovery rate imposed by this configuration. Still, the performance of the 200 protein configuration in "S" is slightly lower than the same configuration using "O"; however, this difference is of low statistical significance (<0.3 standard deviations).
[0196] Response prediction by RAP number The RAP score described above allows for the identification of proteins specific to patients with expression levels corresponding to non-responsiveness, as reflected by the expression of responders and non-responders.Therefore, it was hypothesized that the number of RAPs a particular patient has will predict the response of that patient.Patients with a small number of RAPs or no RAPs at all are predicted to respond to treatment, since almost all measured proteins show expression levels consistent with the responder population. Patients with a number of RAPs are predicted to develop resistance because the expression level of some proteins is similar to that of non-responder population.This method does not consider the nature of RAPs, and each subject may have completely different RAPs.Rather, in some cases, it is the total number of RAPs that is important, not the identity of RAPs.
[0197] The predictive performance of the RAP score was tested using the test set. Specifically, for each patient in the test set (n = 30), a RAP score was calculated for all features using the protein level distributions of R and NR for all patients in the training set (n = 78). Combined with the threshold calculated using the training set as described above, it is possible to estimate the number and identity of RAPs for each patient in the test set. Figure 4A shows 30 subjects from the test set and the number of RAPs calculated for each subject using the TO and T1 data (using a mathematical method). The threshold was set at three RAPs, and subjects with more than three RAPs were predicted to be non-responders. The ROC curve showed an AUC of 0.88, indicating the high predictive power of this analysis (Figure 4B).
[0198] Targeting RAP Improved understanding of the molecular and immunological mechanisms of resistance to ICI therapy may not only identify novel predictive biomarkers but also suggest targets for ICI combination therapies, which aim to selectively block ICI resistance proteins to improve ICI outcomes in non-responding patients.
[0199] To find targets for combination therapy, we evaluated all RAPs with a score >2.9 (a defined threshold) found in the test set of patients. We then performed a search for clinical trials targeting RAPs from this list in combination with ICIs in patients with non-small cell lung cancer (NSCLC) or solid tumors. Mapping clinical trials with combination therapy yielded 1,300 clinical trials targeting 430 proteins in combination with ICIs or with 500 drugs in NSCLC or solid tumors. Comparing the 30 RAPs that passed the score threshold in the test set (RAPs with a score >2.9 that appeared in at least one of 30 patients) with the list of proteins found to be targeted in clinical trials in combination with ICIs revealed four RAPs that were also targeted in combination with ICIs in NSCLC clinical trials: KDR (VEGFR2), IL6, EPHA2, and TACSD2.
[0200] IL-6 is one of the targetable RAPs identified in a patient test set cohort. Recently, we demonstrated that the therapeutic efficacy of anti-CTLA-4 was significantly improved by coadministration of anti-IL-6 in tumor-bearing mice (Khononov, et al., 2021, "Host response to immune checkpoint inhibitors contributes to tumor aggressiveness," J. Immunother. Cancer, Mar. 9). These results are consistent with previous publications showing improved therapeutic outcomes when anti-IL-6 was combined with anti-PD1 or anti-PD-L1 treatment. Furthermore, Khononov et al.'s in vitro experiments demonstrated that inhibiting IL-6 reduced the invasive properties of tumor cells induced by anti-PD-1, further supporting the idea that blocking specific therapy-induced host factors is a strategy for overcoming therapy resistance.
[0201] An alternative approach for therapeutic targeting based on RAP is to associate proteins with key biological processes associated with cancer. To this end, each protein was assigned to a cancer signature that captures key tumorigenesis processes. An enrichment analysis was then performed for each patient using RAP as input (Fisher's exact test; Figure 5). Preliminary analysis of six patients revealed a total of four enriched processes. One patient showed significant enrichment in all four processes; four patients showed enrichment in one to three processes; and one patient showed no significant processes.
[0202] After performing enrichment analysis on a patient, the treating physician can select a treatment based on the enhanced biological processes. For example, if angiogenesis is significantly enhanced, the physician may choose to combine an approved drug that targets angiogenesis (e.g., Avastin) with an ICI. Another example is a patient with a high proliferation signal. In this case, the physician may choose to combine an ICI with chemotherapy against tumor cell proliferation.
[0203] To further examine the biological aspects of RAPs, 19 RAPs obtained in at least three patients in the test set cohort were examined. Most patients had 4-5 RAPs. The most common RAP among the patients examined was VEGFR2 (KDR; identified as RAP in 12 patients). Notably, the majority of RAPs were identified at T1, suggesting that resistance to treatment was primarily acquired and due to a host response. VEGFR2 was identified as a RAP in both T0 and T1, but in T1, it was defined as a RAP in more patients (12 compared with 8 at T0). VEGFR2 is one of the two receptors for vascular endothelial growth factor (VEGF), a major growth factor for endothelial cells, and its expression is higher in responders.
[0204] Network analysis revealed that most RAPs are functionally related to each other, with five of them being highly interconnected (Figure 6). Most proteins are associated with at least one hallmark of cancer, further suggesting that these RAPs are indeed associated with resistance to therapy. Several hallmarks of cancer were significantly enriched for the 19 RAPs (Figure 7), and multiple intracellular and membrane proteins were identified as RAPs (Figure 6). Therefore, putative cell-of-origin analysis was performed to further understand these results (Figure 8). Enrichment was observed for lung and bronchial as cell-of-origin types. Furthermore, various cancer types were examined for expression of the 19 RAPs, and enrichment for lung cancer was also observed (Figure 9).
[0205] Example 2: Combining RAP and Clinical Data A cohort of 184 NSCLC patients was obtained from whom blood samples were collected before (T0) and after (T1) the first dose of ICI. Protein levels were measured. Response assessment was based on ORR at 3 and 6 months after treatment initiation and on durable clinical benefit (DCB) at 1 year after treatment initiation. Progression-free survival (PFS) and overall survival (OS) were also monitored. For the 3- and 6-month evaluations, subjects with progressive disease or who died were considered non-responders, whereas subjects with stable disease, minimal remission, partial remission, and complete remission were considered responders. DCB was defined as 1-year PFS while continuing ICI treatment. Patients who discontinued ICI treatment due to adverse events (but without signs of progression) were considered responders. Additional clinical information collected throughout the study included treatment choice (first-line or advanced-line), PD-L1 immunostaining (<1%, 1–49%, >50%), age, and gender (see Figures 10A–10F). Analyses presented are based on T0 only. A summary of the ICIs / treatments used is shown in Table 3.
[0206] [Table 3]
[0207] The cohort was divided into a development set (60% of subjects) and a validation set (40% of subjects). The development set was further divided into a training set and a test set. The model was trained on the training set, and predictions were generated for a subset of patients not seen by the model during training (i.e., the test set). In order to generate stable predictions for all patients in the development set, the division of the development set into a training set and a test set was performed multiple times (each time to train the model on a different subset of the development set and perform predictions on the remaining patients, i.e., the training set and the test set were mixed and remixed, and dozens of iterations were performed to test the effectiveness of the model / classifier across the entire development set). Prediction quality was then quantified by calculating the ROC AUC for patients included in the development set. The validation set was used only at the very end of the analysis to verify the performance of the final classifier. This division was performed multiple times.
[0208] A model was created based on response assessments at three time points: 3 months, 6 months, and 1 year after initiation of treatment. A total of 184 patients were evaluated at 3 months, 177 at 6 months, and 146 at 1 year. Tolerance increased over time. 26% of subjects were non-responders at 3 months, 45% at 6 months, and 74% at 1 year. These rates were similar between the development and validation sets.
[0209] During model generation based on the development set, the development set was randomly divided into a training set and a test set 60 times. In each iteration, the top candidate proteins were selected using the Kolmogorov-Smirnov test, which defined for each protein how well it distinguished between responders and non-responders. For each selected protein, a single-protein XGBoost model (SP model) was generated based on the training set, and predictions were performed on the test set. A protein was defined as a RAP for a particular patient if its predicted resistance probability (i.e., resistance score) exceeded a predetermined threshold, and the average value across all iterations was used for each patient. A uniform threshold was assigned to all models to address class imbalance. Different thresholds were defined for each time point (e.g., 3-month threshold = 0.25, 6-month threshold = 0.42, 1-year threshold = 0.45). For each patient, the number of proteins whose model scores exceeded the defined threshold (i.e., the number of RAPs) was calculated.
[0210] The number of RAPs alone was predictive for this cohort. However, a predictive model was created that also integrated clinical data. The clinical classifier shown used as input the number of RAPs, treatment choice (ICI was first-line or advanced treatment), subject age, and percent PD-L1 staining within the tumor (<1% of cells, 1-49%, ≥50% positive). The classifier then generated an overall resistance score ranging from 0 to 1, with 0 being most similar to a responder and 1 being most similar to a non-responder. Subjects with a score above a predefined threshold were predicted to be non-responders. Similarly, a response score, 1 minus the resistance score, was also calculated. For the response score, subjects with a score above a predefined threshold were predicted to be responders.
[0211] To test the performance of the classification model, ROC AUC was calculated using the overall resistance score together with the actual response. ROC AUC was calculated separately for 3-month ORR, 6-month ORR, and 1-year DCB for both T0 and T1. The results are summarized in Figure 11A. The classifier was found to be predictive in both the development and validation sets at all time points. A similar analysis showed that the classifier was also predictive at all time points for the T1 data (Figure 11B).
[0212] To further confirm the performance of the classification model, we also tested the correlation between the predicted response probability (response score) assigned to each patient by the classification model and the observed response probability. To this end, for each value of the response score S, the observed response probability was given by the proportion of responders among patients assigned a response score within S ± 0.1. The choice of the ± 0.1 interval was arbitrary and reflected the size of the validation set; in larger validation sets, the interval could be further reduced. The agreement between the predicted response score and the actual response probability was quantified by the goodness of fit, R^2. The goodness of fit for all three time points (3-month ORR, 6-month ORR, 1-year DCB) was R^2 = 0.98 at time T (Figures 12A-12B).
[0213] The patients in the validation set were stratified into a long-term benefit group and a limited benefit group, and the stratification was based on the 3-month predicted response score. In survival analysis, the quality of stratification was measured by the hazard ratio (HR), which gives the ratio of the probability of an event per time unit in the two groups. For example, an HR of 4 for overall survival (OS) means that the probability of a death event per time unit in the limited benefit group is four times that per time unit in the long-term benefit group. The HRs in the validation set were 2.27 for PFS, p<0.004 (Figure 13A), and 4.50 for OS, p<0.0001 (Figure 13B).
[0214] This validation experiment demonstrates that the classifier incorporating clinical data and RAP number is highly predictive of patient response.
[0215] Functional network analysis of RAP The RAP-based analysis was further used as the basis for creating a resistance map (Figure 14A). The resistance map displays both the interactions between RAPs and RAP function. To this end, a RAP was defined as a protein selected in at least 10 model iterations in one or more patients (during RAP calculation, the model runs 60 iterations, and the number of times a given protein is selected for that model is recorded), resulting in a total of 73 RAPs in the current patient cohort. Each node represents a RAP, and the edges between nodes indicate functional relationships. Nodes with larger sizes indicate investigational new drugs (INDs) combined with immunotherapy. Nodes are color-coded based on protein function. The map shows multiple interactions between different RAPs, while RAPs are involved in different functional processes that may be related to resistance to therapy, such as splicing, immunomodulation, angiogenesis, and cell proliferation. Patient-specific maps can be created based on the patient's RAPs, which can help 1) map resistance mechanisms in individual patients and 2) identify targeted treatments that counteract resistance. Two examples of patients in the cohort are shown in Figure 14B. In these examples, the non-responder had 44 RAPs and a response probability score of 0.44 (which corresponds to a resistance score of 0.56, above the predetermined threshold of 0.2 for non-response). This patient had RAPs from multiple functional groups, but no DNA-associated RAPs were present in this patient. The second subject was a responder with 10 RAPs, below the predetermined threshold. These RAPs were primarily associated with the cytoskeleton. This patient had a response probability of 0.91 (which corresponds to a resistance score of 0.09, below the predetermined threshold of 0.2).
[0216] Further investigation of the patient RAPs revealed functional differences between RAPs, with higher occurrence in each response group (Figure 15). RAPs in non-responders are involved in splicing, signal transduction, and cytoskeleton-related processes, whereas RAPs in responders are primarily involved in proteolysis and cell adhesion. Interestingly, the elevated RAPs in the responder group contain two peptidases that are involved in antigen presentation, thereby potentially promoting response to therapy. To convert non-responders to responders, RAPs for which there are known therapeutic agents are selected. This agent is selected so that it modulates RAP and alters pathway function so that it more closely resembles that in responders. If therapeutic agents targeting RAP are unavailable or undesirable, therapeutic agents that modulate pathways involving RAP are selected. The selected agent must modulate a pathway involving RAP and alter pathway function so that it more closely resembles that in responders. Therapeutic agents are used to convert non-responders to responders or as combination therapy with ICIs.
[0217] Example 3: RAP-based models predict differential outcomes in patients based on PD-L1 tumor expression. To develop a blood-based model for predicting benefit from first-line PD-(L)1-based ICI therapy, plasma samples and clinical data were collected from ICI-treated patients with advanced NSCLC. Pretreatment plasma samples from 425 patients were profiled using a protein assay that measured approximately 7,000 proteins per plasma sample. After patient exclusion for technical or clinical reasons, the study cohort consisted of 339 remaining patients.
[0218] Clinical parameters of the patients are shown in Figure 16. The median age was 65 years, and one-third of the patients were female, with a male predominance. Consistent with expected proportions, the majority of patients (78.47%) had non-squamous cell carcinoma (mostly adenocarcinoma), and 21.24% had squamous cell carcinoma. Most patients had an ECOG performance status of 0-1 (94%). Patients were treated with either ICI-chemotherapy combination therapy (59.8860%) or ICI monotherapy (40.12%). Patients with PD-L1-negative, PD-L1-low, and PD-L1-high tumors were approximately evenly distributed, with negative, low, and high representing PD-L1 expression in less than 1%, 1-49%, and 50% or more of tumor cells, respectively. The PD-L1-high group was the most prevalent (36%). Clinical benefit (CB) was defined as described above.
[0219] Treatment benefit was assessed at 3, 6, and 12 months after treatment initiation. For each time point, patients were classified into clinical benefit (CB) and no clinical benefit (NCB) groups as follows: Patients who demonstrated a complete response, partial response, or stable disease at 3 and 6 months were classified as CB patients, while patients who demonstrated progressive disease or died were classified as NCB patients. At 12 months, patients who were alive and demonstrated a sustainable clinical benefit (defined as no progressive disease for at least 1 year after treatment initiation) were classified as CB patients, and all other patients were classified as NCB patients. Based on these criteria, 69.32%, 46.02%, and 24.78% of patients achieved CB at 3, 6, and 12 months, respectively (Figure 16). Cohort size varied between time points due to patient death or lack of clinical benefit data at each time point (Figure 17). Therefore, the data set included 339, 331, and 299 patients at 3, 6, and 12 months, respectively.
[0220] Various clinical parameters were found to be associated with CB (Figure 18). At 3 months, a higher proportion of CB patients was found in the ICI-chemotherapy treated group compared with the ICI monotherapy group (78% vs. 57%, respectively; p = 0.001), while no association between treatment type and CB was found at other time points. PD-L1 status correlated with CB at 6 and 12 months, with a higher proportion of CB patients found in the PD-L1 high group compared with the combined PD-L1 low and PD-L1 negative group (66% vs. 59%, respectively, at 6 months; p = 0.010; 40% vs. 22%, respectively, at 12 months; p = 0.010). Furthermore, at 12 months, the non-squamous lung cancer group had a higher CB rate compared with the squamous cell carcinoma group (31% vs. 18%, respectively; p = 0.039). ECOG performance status correlated with CB at 3 months, with a higher proportion of CB patients in the ECOG 0 and 1 groups compared to the ECOG 2 group (68% and 72% vs. 44%, respectively; p-value = 0.047). Finally, a higher CB rate was found in women compared to men at 12 months (36% vs. 24%, respectively; p-value 0.038).
[0221] Example 4: Prediction of benefit from ICI treatment based on clinical parameters Although PD-L1-based companion diagnostic tests recommend the use of ICI monotherapy for NSCLC patients with high PD-L1 levels, clinical evidence also shows a trend toward increased benefit with increasing tumor PD-L1 levels in patients treated with ICI-chemotherapy combination therapy. The predictive performance of the PD-L1 biomarker was evaluated across various expression levels (i.e., <1%, 1–49%, and ≥50%) in a mixed cohort of patients treated with either ICI monotherapy or ICI-chemotherapy combination therapy. Predictive models were generated for each CB assessment time point (3, 6, and 12 months), and the cohort was divided into a development set and a validation set. The development set consisted of 75% of the patients (n=254) and was used for model generation. After model development, overall performance was evaluated in a blinded manner in an unrelated validation set (n=85; Figure 19A) consisting of the remaining 25% of patients (n=85).
[0222] Although PD-L1 expression correlated with CB at 6 and 12 months (p-value=0.01; Figure 18), its prediction of CB at each of the three time points was poor, with the area under the curve (AUC) of the receiver operating characteristic (ROC) plot being 0.50 (p-value=5.13e-01), 0.60 (p-value=6.13e-02), and 0.55 (p-value=2.76e-01) at 3, 6, and 12 months, respectively (Figure 19A).
[0223] We next investigated whether incorporating additional clinical parameters would improve the predictive ability of the PD-L1 biomarker. We considered three clinical parameters known to correlate with treatment benefit: patient gender, ECOG performance status, and treatment choice. Accordingly, we developed a predictive model based on PD-L1, gender, ECOG, and treatment choice, referred to herein as the "clinical model." This clinical model showed only a slight improvement in response prediction ability compared with PD-L1 alone, with AUCs of 0.52, 0.60, and 0.62 at 3, 6, and 12 months, respectively (Figure 19B). Therefore, a more powerful predictive model is needed.
[0224] Example 5: Resistance-associated protein (RAP) prediction model Aiming to develop a more robust predictive model, we designed an additive model in which the output is based on the sum of predicted values from a large set of individual features related to therapeutic benefit. Because each feature itself has only a small effect on the final output, the effect of any false positives is minimized and the model remains stable. This approach likely mitigates the significant heterogeneity between patients and the effect of a large number of features in a relatively small cohort.
[0225] Briefly, this model is based on a set of proteins that exhibit differential plasma level distributions in CB and NCB populations, as determined by statistical tests. Such proteins, called resistance-associated proteins (RAPs), serve as potential indicators of treatment benefit depending on their plasma concentrations in individual patients (Figure 20A). Specifically, for a given patient, a machine learning (ML)-based model trained on the CB and NCB populations infers a CB or NCB prediction from the patient's plasma level of each RAP within the entire RAP set. In this way, patients are assigned a set of predictions based on their personal RAP profile, and the sum of all predictions reflects the patient's likelihood of benefiting from treatment. Patients with a high number of CB predictions are more likely to benefit, while patients with a high number of NCB predictions are less likely to benefit.
[0226] Three RAP-based models were developed, one for each of the three CB assessment time points. These models were developed following the same workflow, in which CB markers at 3, 6, or 12 months were used as inputs, along with protein expression data and patient gender (Figure 20A). First, to define the set of RAPs on which the final models were based, the development set (75% of the patient cohort; n = 254) was divided into a training set and a test set, consisting of 75% and 25% of the development set patients, respectively (Figure 20B; see also Materials and Methods). Proteins showing statistically significant differences between their plasma level distributions in the CB and NCB populations were identified in the training set, and the 50 proteins with the lowest p-values were selected as RAPs (Figure 21; see also Methods). Next, for each selected RAP, an ML algorithm was trained on two features, namely, RAP expression level and patient gender, to develop a binary classifier for the therapeutic utility of each RAP. Predictions were then inferred for each RAP for each patient in the test set, and a RAP score was computed based on the collective predictions from the 50 selected RAPs. This three-step process (i.e., RAP selection, model training, and RAP score computation) was repeated 80 times, each time randomly dividing patients into training and test sets (Figure 20B). RAP scores were averaged for each patient and linearly scaled to generate a model whose final output was the CB probability—a clinically oriented metric reflecting the patient's likelihood of benefiting from treatment.
[0227] Because RAP selection occurred via an iterative process during model development (50 RAPs were selected from the training set after randomly mixing patients between the training and test sets 80 times), it was possible that the same RAPs were selected several times overall (Figure 22A). Of the total 287, 330, and 371 RAPs selected at 3, 6, and 12 months, respectively, approximately 100 RAPs were selected at least 10 times per time point (Figure 22B). A total of 598 RAPs were selected across the three time points, of which 113 RAPs were common to all three time points (Figure 22C). Additionally, approximately 30 RAPs were selected more than 10 times across the three time points (Figure 22D).
[0228] To gain insight into the biological functions of the selected RAPs, we first classified them by their cellular location and origin based on the Human Protein Atlas database. In most cases, RAPs were found to be intracellular proteins, most likely derived from immune cells. Approximately 8–10% of RAPs per time point are known to be highly expressed in lung tumors (Figure 22E). Next, we performed functional analysis of the selected RAPs at each time point. At all time points, multiple RAPs were found to be involved in splicing or alternative splicing (Figure 22F), while splicing was significantly enhanced at the 3-month time point (Figure 22G). Fisher's exact test, false positive rate <0.1; this test was applied using an identification cutoff of at least 10 replicates, but similar results were obtained using different cutoffs. Furthermore, at all three time points, multiple RAPs were associated with the complement and coagulation cascades (Figure 22F). Finally, extracellular matrix (ECM)-related pathways, represented by proteins such as osteopontin (SPP1) and TIMP1, were significantly enhanced at 3 and 6 months, whereas two hallmarks of cancer (i.e., sustained growth signaling and invasion and metastasis) were significantly enhanced at 6 and 12 months (Figure 22G). Notably, multiple RAPs, such as VEGFA, IL-6, FLT4, CSF1R, and CA125 (MUC16), are known targets of approved and investigational therapeutic agents, some of which are being investigated in combination with ICIs in clinical trials. Overall, these findings demonstrate a link between RAPs and biological pathways related to tumor progression and treatment resistance.
[0229] Example 6: RAP model predicts benefit from ICI treatment After model development, the RAP model for each time point was locked and blindly tested in an unrelated validation set (25% of the patient cohort; n = 85). The validation set consisted of patients with advanced NSCLC treated with first-line PD-(L)1-based ICI therapy, either as monotherapy or in combination with chemotherapy. CB probability was determined for each patient in the validation set at each time point. The range of the CB probability distribution varied at each time point, with the median CB probability decreasing over time (Figure 23A). Furthermore, the CB probability for all patients decreased from one time point to all subsequent time points (Figures 25A-25C), consistent with the decrease in the actual CB rate over time (Figure 25D). Notably, actual NCB patients clustered in the low range of predicted CB probability at all three time points, indicating the high predictive power of the model (Figure 23A). This finding was further supported by enrichment analysis based on CB probability (2D enrichment test; false positive rate <0.05). Notably, at all three time points, the patient group with high CB probability was significantly more likely to have CB, females, non-squamous cell carcinoma, and no progressive disease or death events, whereas the patient group with low CB probability values was significantly more likely to have NCB, males, squamous cell carcinoma, and patients with progressive disease or death events (Figure 24).
[0230] Next, using the median CB probability as the threshold, we classified patients into high or low CB probability groups. Specifically, patients with predicted CB probabilities above or below the median were assigned to high or low CB probability groups, respectively (Figure 23B). Log-rank tests showed that patients in the high CB probability group achieved significantly longer overall survival (OS) than patients in the low CB probability group across the three time points (Figure 23B, hazard ratio, HR = 0.24-0.38). Similar results were obtained for progression-free survival (PFS; Figure 23C, HR = 0.32-0.41). These findings indicate that the RAP-based model effectively classifies the survival outcomes of NSCLC patients treated with ICIs.
[0231] To further test the accuracy of the test model, the predicted CB probability was compared with the observed CB rate. Here, the latter refers to the proportion of CB patients observed within a group of patients assigned a similar CB probability (i.e., CB probability ±0.15). Linear regression analysis demonstrated a high goodness of fit (R2 = 0.97) between the predicted CB probability and the observed CB rate (Figure 23D). Furthermore, the AUCs of the ROC plot were 0.71, 0.77, and 0.78 at 3, 6, and 12 months, respectively (Figure 23E), demonstrating the strong predictive ability of the RAP model over the first year of ICI-based treatment. Notably, the RAP model demonstrated superior predictive performance compared with the PD-L1-based model (AUC = 0.5-0.6 over the first year) and the clinical model (AUC = 0.52-0.62 over the first year) (Figures 19A-19B).
[0232] We next investigated whether integrating clinical parameters into the RAP model would improve its predictive performance. To this end, we integrated a PD-L1-based model (PD-L1) or a clinical model (CM) into the RAP model and compared their predictive performance. Interestingly, adding PD-L1 parameters to the RAP model slightly improved its predictive performance at 6 months, but integrating the RAP model and the clinical model reduced its predictive performance overall (Figure 26A). In survival analysis, the RAP model showed the best HR compared to the other four models, but the HRs were not significant for the PD-L1-based model and the clinical model (Figure 26B).
[0233] Finally, we investigated the performance of the RAP model in different patient subsets (Figure 27). The model demonstrated strong predictive performance in both the ICI monotherapy and ICI-chemotherapy subsets, similar to its performance in the overall population. Histological subset analysis, on the other hand, demonstrated improved prediction for the squamous cell carcinoma subset at 3 months compared to the overall population. At 6 and 12 months, the strongest prediction was observed in the PD-L1-negative subset, while prediction was slightly weaker in the PD-L1-high subset compared to the overall population.
[0234] Example 7: The RAP model predicts differential outcomes in patient subgroups classified by PD-L1 expression. Because PD-L1 expression is a major factor influencing treatment selection, we investigated the properties of a model predicting survival outcomes when PD-L1 classification is taken into account. In our cohort, patients with high PD-L1 levels (≥50%) showed the best outcomes, with up to two-fold differences in median OS and PFS compared with patients with low PD-L1 levels (1-49%) and PD-L1-negative patients (<1%) (Figure 28). Most patients with high PD-L1 levels (65.3%) were treated with a single ICI according to current guidelines.
[0235] Among patients with high PD-L1 levels, it is possible to distinguish between those who would benefit from ICI monotherapy and those who would do better with ICI-chemotherapy combination therapy. To investigate this, we examined the properties of the 12-month RAP model to predict survival outcomes for patients with high PD-L1 levels receiving ICI monotherapy or ICI-chemotherapy combination therapy. Patients were classified into high or low CB probability groups using the median CB probability of the cohort as the threshold, and OS and PFS curves were plotted for each group. In the high CB probability group, patients receiving ICI monotherapy or combination therapy fared similarly (Figure 29, left panel). Median OS was 32.13 months vs. 28.95 months, and median PFS was 7.85 months vs. 13.08 months (monotherapy vs. combination therapy). This suggests that such patients are suitable candidates for monotherapy and may be spared the more toxic ICI-chemotherapy combination therapy. In contrast, in the low CB probability group, OS and PFS were significantly longer in patients receiving ICI-chemotherapy compared with ICI monotherapy (Figure 29, right panel). The median OS was 10.71 months versus 4.14 months (combination therapy vs. monotherapy; HR = 0.17; p = 0.001), and the median PFS was 14.29 months versus 4.14 months (combination therapy vs. monotherapy; HR = 0.40; p = 0.016). This suggests that PD-L1-high patients with low CB probability should be treated with ICI-chemotherapy combination therapy, regardless of their high PD-L1 levels.
[0236] We also investigated whether this model could provide insights for managing patients with PD-L1 levels below 50%. To this end, we tested the properties of the RAP model to predict survival outcomes in a mixed group of PD-L1 low and PD-L1 negative patients receiving ICI monotherapy or ICI-chemotherapy combination therapy (overall, 47 PD-L1 low and negative patients received ICI monotherapy, 87% of whom (Those treated with ICI as an advanced treatment option were included.) In this analysis, patients in the high CB probability group showed an OS benefit when treated with ICI-chemotherapy combination therapy compared with patients receiving monotherapy, although the difference did not reach statistical significance (Figure 30, left panel). The median OS was 27.83 months for ICI-chemotherapy versus 12.72 months for ICI monotherapy. Notably, the median OS in the ICI-chemotherapy subset was comparable to the median OS of all PD-L1-high patients (median OS 28.96; Figure 28). This result is consistent with current guidelines, which recommend ICI-chemotherapy over ICI monotherapy for patients with PD-L1 levels below 50%. However, patients in the low CB probability group also showed poor outcomes when treated with either of the two treatment modalities, with median OS of 10.02 months and 9.69 months for monotherapy and ICI-chemotherapy, respectively (Figure 30, right panel). This suggests that treatment types other than commonly used ICI-chemotherapy combinations, including platinum-based chemotherapy, first-line clinical trials, and novel combination therapies, may be considered for this patient subgroup. Similar trends were observed when such comparisons were made in subgroups consisting exclusively of patients with low or negative PD-L1 levels. The results presented here were obtained using the RAP model based on the 12-month time point, although similar results were obtained with the 3- and 6-month time point models (data not shown).
[0237] These collective findings demonstrate the model's potential clinical utility for optimizing treatment selection: When combined with PD-L1 testing, the model could help determine whether patients should receive ICI monotherapy, ICI-chemotherapy combinations, or alternatives to commonly used therapies.
[0238] Example 8: Further confirmation of RAP(PROphet) model predictions Plasma samples and clinical data were collected from 610 patients with advanced NSCLC treated with ICI as monotherapy or in combination with chemotherapy within the framework of the PROPHETIC clinical trial (NCT04056247). A separate cohort of 85 patients treated with chemotherapy alone was used for specific comparisons. Samples analyzed in this study were analyzed by proteomic profiling of approximately 7,000 proteins. Of the 610 enrolled patients, 65 were excluded for technical or clinical reasons, leaving 545 in the analyzed cohort (Figure 31A). Clinical benefit (CB) was assessed 12 months after the start of treatment. Patients who demonstrated progression-free survival (PFS) for at least 12 months after the start of treatment were classified as CB patients. All other patients were classified as "no clinical benefit" (NCB) patients.
[0239] Clinical parameters of the patients are shown in Figure 31B. Focusing on the ICI-based treatment cohort, the median age was 66 years (range 33-89 years), and male patients predominated (61%). Most patients (80%) had non-squamous cell carcinoma and an ECOG performance status of 0-1 (91%). In terms of clinically significant metastatic sites, 30%, 16%, and 24% of patients had bone, liver, or brain metastases, respectively. Overall, 25%, 27%, and 41% of patients had PD-L1 levels of <1%, 1-49%, and ≥50%, respectively. Patients were treated with either ICI-chemotherapy combination therapy (59%) or ICI monotherapy (41%). Most patients were either former smokers (51%) or current smokers (39%). Overall, 25% of patients achieved CB at 12 months.
[0240] Proteomics-based models were developed and evaluated in patients undergoing ICI-based therapy and undergoing clinical benefit assessment. Models were developed in a development set (n = 228) and blindly tested in an unrelated validation set (n = 272; Figure 31C). A set of 388 proteins (Table 4) that showed differential plasma levels between CB and NCB populations was identified using the Kolmogorov-Smirnov statistical test in 80 replicates of randomly selected training and test sets (Figure 31C). These proteins are called resistance-associated proteins (RAPs) and serve as likely indicators of CB based on the XGBoost algorithm. The sum of the 388 predictions for a given patient is called the PROphet score (total response score) and reflects the patient's likelihood of benefiting from treatment.
[0241] Because PD-L1-based tests are currently used to guide treatment in NSCLC patients, we evaluated the predictive performance of the PD-L1 biomarker in the validation set. In this study, cancers with PD-L1 ≥ 50% showed a non-significant overall survival (OS) benefit compared with cancers with PD-L1 < 50% (p-value = 0.0655; hazard ratio (HR) between PD-L1 ≥ 50% and PD-L1 < 50% of 0.74, confidence interval (CI) of 0.53-1.02; Figure 32A). Furthermore, the PD-L1-based predictive model showed poor correlation between the predicted probability of clinical benefit and the observed benefit rate (R 2 =0.35; Figure 32B), demonstrating the limited predictive power of the PD-L1 biomarker.
[0242] A model was developed that outputs CB probability (a continuous metric) using proteomic and clinical data from patients undergoing ICI-based treatment. Patients with a predicted CB probability equal to or greater than the median in the development set versus patients with a predicted CB probability less than the median in the development set were classified into positive or negative groups, respectively. This proteomic-based model was called PROphet (Figures 31C-31D). This model demonstrated excellent predictive performance compared to PD-L1, with a hazard ratio (HR) of 0.51 between the positive and negative groups (CL = 0.37-0.70; p-value < 0.001; Figure 32C) and a median OS of 25.9 months and 10.8 months for the positive and negative groups, respectively. Furthermore, this model demonstrated a high goodness of fit (R ) between the predicted CB probability and the observed CB rate. 2 = 0.97; Figure 32D), demonstrating strong predictive performance overall. When testing the model's performance in a retrospective cohort of treatment-naive patients receiving chemotherapy alone, the PROphet subgroup showed no significant difference in OS (HR = 0.68; CI = 0.43-1.06; p-value = 0.0853) (Figure 33A). There was poor correlation between predicted CB probability and observed CB rate (R 2 =0.09, Figure 33B). Together, this suggests that the PROphet test is more predictive for ICI-based treatments than for chemotherapy.
[0243] Next, we evaluated the clinical utility of combining the model results with PD-L1 expression levels (patient stratification is shown in Figure 34). The subgroup of patients with PD-L1 ≥ 50% who had positive results performed similarly well in terms of OS and PFS when receiving ICI monotherapy or combination therapy (OS HR = 0.77; CI = 0.42-1.43; p-value = 0.4096; Figures 35A and 36A). This suggests that such patients may be good candidates for monotherapy and may be spared the more toxic ICI-chemotherapy combination therapy. In contrast, in patients with negative PD-L1 levels >50%, both OS and PFS were significantly longer in the ICI-chemotherapy group compared with ICI monotherapy (Figures 35D and 36D), with median OS not reached in the combination therapy group and 11.10 months in the monotherapy group (HR = 0.29, CI = 0.14-0.59, p < 0.001). Multivariate Cox proportional hazards regression analysis identified a significant interaction between the model results and treatment regimen (Figure 37; ECOG performance status scale was also significant), indicating that the effect of treatment depended on the model results. This suggests that patients with negative PD-L1 levels >50% should be considered for ICI-chemotherapy combination therapy despite their high PD-L1 levels, in contrast to patients with positive results. This is consistent with a comparison between negative and positive subgroups with PD-L1 ≥ 50%, where a significant difference between the two groups was observed for ICI monotherapy but not for ICI-chemotherapy combination therapy (Figures 38A-38B).
[0244] Next, the subgroup of patients with PD-L1 <50% was analyzed. The subgroup of patients with a positive PD-L1 <50% result showed a significant benefit in OS with the ICI-chemotherapy combination compared with chemotherapy alone (Figures 35B-35C), with HRs of 0.39 and 0.41 for patients with PD-L1 1-49% and <1%, respectively. The median OS was 27.9 months and 23.2 months for patients with PD-L1 <49% and <1%, respectively, receiving ICI-chemotherapy, compared with 8.6 months for chemotherapy. PFS benefited from ICI-chemotherapy only in patients with PD-L1 1-49% (Figure 36B), whereas no significant difference was observed in PFS for patients with PD-L1 <1% (HR = 0.67, CI = 0.43-1.03; p = 0.0675 (Figure 36C)). Of note, in the PD-L1 < 50% group, ICI monotherapy was only used in the comparison for the PD-L1 1-49% group because it is not standard treatment for these patients (Figures 35G-35H). Furthermore, patients receiving chemotherapy were not stratified based on PD-L1 expression level because values were not available for many patients treated with chemotherapy alone. Overall, our findings suggest that patients with PD-L1 < 50% positivity will benefit from guideline-based treatment (i.e., the combination of ICI and chemotherapy).
[0245] When examining the subgroup of patients with PD-L1 1-49% and negative results, a significant difference between ICI-chemotherapy and chemotherapy alone was observed, with a HR of 0.51 and a median OS of 11.5 months and 6.7 months for the combination therapy versus chemotherapy, respectively (Figure 35F), while no significant difference in PFS was observed between the two arms (Figure 36E). Multivariate analysis showed no interaction between treatment and PROphet results, indicating that in the subgroup of PD-L1 1-49% patients, there was no effect on treatment outcome, as both negative and positive patients benefited from the combination therapy. Because our results showed only a moderate benefit from ICI-chemotherapy for PD-L1 1-49% patients with negative PROphet results, these patients may wish to consider other approved therapies or first-line clinical trials, as also recommended by the NCCN guidelines.
[0246] Conversely, PD-L1-negative patients with <1% PD-L1 expression showed similarly poor outcomes for both treatment modalities, with median OS of 7.5 and 6.7 months for combination therapy and chemotherapy, respectively (Figure 35E), and a median PFS of 4.5 months for both treatment modalities (Figure 36F). These findings suggest that such patients are less likely to benefit from ICI-based combination therapy than chemotherapy and may opt for consideration of other approved therapies or first-line clinical trials. Consistently, patients with PD-L1 <1% PD-L1 expression showed significant differences between the negative and positive subgroups (Figure 38D). Guidelines for patients with PD-L1 >50% support ICI monotherapy or ICI in combination with chemotherapy, but there is no clear guidance as to which treatment modality is more beneficial for these patients. Our model successfully distinguishes between patients who will benefit from combination therapy and those for whom ICI monotherapy may be sufficient and who may avoid chemotherapy-related toxicity. This study could improve overall survival by directing patients with a negative response score of PD-L1 ≥ 50% to ICI-chemotherapy treatment.
[0247] Guidelines for patients with PD-L1 <50% recommend the combination of ICI and chemotherapy. Patients with a positive PROphet response score and either PD-L1 1-49% or PD-L1 <1% expression levels demonstrated prolonged OS when receiving ICI in combination with chemotherapy. Thus, this study successfully identified patients who could benefit from standard therapy. However, patients with a negative response score showed differential results for PD-L1 1-49% and PD-L1 <1% expression levels, whereas patients with PD-L1 1-49% demonstrated significant benefit from combination therapy, and patients with PD-L1 <1% did not.
[0248] PD-L1 biomarkers are currently used to guide treatment selection, but as mentioned above, they are not completely reliable. The model described in the present invention provides proteomic analysis of pre-treatment plasma samples in combination with PD-L1 testing for stratifying patients into subgroups, providing an additional solution to consider when selecting treatment regimens, thus providing a novel tool for treatment decision-making and prediction of clinical benefit in NSCLC patients receiving ICI-based therapy, thus addressing an unmet need.
[0249] [Table 4-1] [Table 4-2] [Table 4-3]
[0250] Example 9: Evaluation of response prediction using the PROphet model in melanoma and SCLC patients. It was hypothesized that immunotherapy response involves common mechanisms across cancer types, and therefore the NSCLC response prediction classifier was applied to protein measurements from blood samples from subjects with a variety of other cancers within the framework of the PROPHETIC clinical trial (NCT04056247).
[0251] T0 plasma samples and clinical data were collected from 68 patients with unresectable metastatic melanoma treated with anti-PD1 alone or in combination with anti-CTLA-4. The performance of the response prediction model was quantified based on the receiver operating characteristic curve (ROC) area under the curve for CB prediction at 1 year. Specifically, the goodness of fit between predicted and observed response probabilities was measured as the ROC area under the curve from the best fit line. 2 Distance was used to assess the risk based on 1-year CB. Hazard ratios (HRs) between positive and negative patients were also calculated.
[0252] For a classifier to be considered predictive, the following criteria must be met. First, the 1-year CB ROC AUC must be greater than 0.60 with a p-value less than 0.05. A threshold of 0.6 was chosen to ensure that the model response probability was better than random and relatively low. A more stringent threshold was not chosen because goodness of fit was the more important criterion. The second criterion was the linear fit between the predicted and observed response probabilities. For the 1-year CB, the fit should be greater than R^2 > 0.85 for the best-fit line. The slope must be greater than 0.9. Third, the predicted response probabilities for the 1-year CB should span a range of at least 0.25 (i.e., if the high response probability assigned to patients in the validation set is 0.6, the lowest response probability should be 0.35 or less). Finally, the hazard ratio between positive and negative patients must be less than 0.8. As seen in Figures 39A-39C, the ROC AUC for the model of 1-year sustainable clinical benefit was 0.69, p=0.004 (39A), and the linear fit between predicted and observed response probabilities based on the PROphet model was R 2= 0.93 (39B), and a Kaplan Meier plot of PROphet-positive and PROphet-negative patients showed a hazard ratio of 0.27 (39C), with a predicted probability range of 0.27. Because all acceptance criteria were met, this classifier is also considered to predict response in melanoma patients treated with anti-PD-1 therapy.
[0253] A similar analysis was performed on TO plasma samples from a cohort of 54 small cell lung cancer (SCLC) patients. Patients with at least 7 months of follow-up were included in the analysis, and response to treatment was defined as PFS of at least 7 months after treatment initiation. Patients were treated with a combination of ICI (atezolizumab, 48 cases; durvalumab, 6 cases) and chemotherapy (carboplatin and etoposide). As can be seen in Figure 39D, the Kaplan-Meier plot for PROphet-positive and PROphet-negative patients showed a hazard ratio of 0.60, p=0.18, indicating that this classifier also appears to predict the response of SCLC patients to treatment with PD-(L)1 inhibitors.
[0254] Example 10: Evaluation of response prediction using the PROphet model in HPV-associated malignancies.
[0255] Patients with HPV-associated malignancies were also evaluated using the PROphet classifier (Figure 39H). TO serum samples from a cohort of 43 patients with HPV-associated malignancies, including anogenital cancer (Figure 39E), cervical cancer (Figure 39F), and head and neck cancer (Figure 39G), who were treated with anti-PDL1 / TGFβ-trap fusion protein, were analyzed for their proteome expression profiles and their survival probability using the PROphet classifier. As can be seen from the Kaplan-Meier curves for PROphet-positive and PROphet-negative patients in Figures 39E-39H, the hazard ratios of the PROphet classifier were also predictive for these cancers. This finding indicates that the resistance mechanisms identified by this classifier are likely pan-cancer phenomena and that this classifier is therefore useful for all cancers.
[0256] Example 11: Evaluation of response prediction using the PROphet model in NSCLC patients with targetable mutations
[0257] NSCLC patients with EGFR, ALK, or ROS1 mutations usually respond poorly to immunotherapy and are therefore initially treated with tyrosine kinase inhibitors (TKIs). To date, no biomarkers exist for identifying NSCLC patients with EGFR, ALK, or ROS1 mutations who may benefit from treatment with PD-(L)1 inhibitors. A cohort of 35 highly selected NSCLC patients, either previously treated with or without TKIs before treatment with PD-(L)1 inhibitors, was analyzed using the PROphet model. As can be seen in Figure 40A, the linear fit between PROphet score and overall survival (days) was significantly improved by R 2 =0.41, p=0.0073. Kaplan Meier curves for PROphet-positive and PROphet-negative patients showed a hazard ratio of 0.36, p=0.07 (Figure 40B). These results demonstrate the ability of this model to predict response to treatment in this specific subpopulation and also to distinguish between NSCLC patients with targetable mutations who may benefit from PD-(L)1 treatment (PROphet-positive patients) and those who do not (PROphet-negative patients).
[0258] While the present invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, the present invention is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and broad scope of the appended claims.
Claims
1. 1. A method for predicting the response of a subject suffering from a PD-L1-high cancer to a monotherapy comprising anti-PD-1 / PD-L1 immunotherapy, the method comprising: a. i. In a population of subjects with cancer and known to respond to said monotherapy (responders); ii. In a population of subjects known to be suffering from cancer and not responding to the monotherapy (non-responders); and iii. In the subject receiving factor expression levels for a plurality of factors; b. Calculating a resistance score for a factor of the plurality of factors, comprising applying a machine learning algorithm trained with a training set including the expression levels of the received factors in responders and non-responders and the gender of each of the responders and non-responders to the expression levels of each received factor from the subject and the gender of the subject, and the machine learning algorithm outputs the resistance score; c. Combining the calculated resistance scores to generate an overall resistance score. Including, subjects with a composite resistance score above a predetermined threshold are predicted to not respond to the monotherapy and subjects with a composite resistance score within the predetermined threshold are predicted to respond to the monotherapy; thereby predicting the subject's response to monotherapy.
2. 2. The method of claim 1, wherein the overall resistance score is converted to an overall response score, and an overall response score above a predetermined threshold indicates that the subject will respond to the monotherapy and an overall response score below a predetermined threshold indicates that the subject will not respond to the monotherapy.
3. The method of claim 1 or 2, wherein the training set further comprises expression levels of factors received in subjects (combo-responders) who are afflicted with cancer and known to be responsive to a combination therapy comprising anti-PD-1 / PD-L1 immunotherapy and chemotherapy, expression levels of factors received in subjects (combo-non-responders) who are afflicted with cancer and known not to be responsive to the combination therapy, and the genders of the combo-responders and combo-non-responders, respectively.
4. 1. A method for predicting the response of a subject suffering from PD-L1 low or negative cancer to a combination therapy comprising anti-PD-1 / PD-L1 immunotherapy and chemotherapy, comprising: a. i. In a population of subjects with cancer and known to respond to said monotherapy (responders); ii. In a population of subjects known to be suffering from cancer and not responding to the monotherapy (non-responders); and iii. In the subject receiving factor expression levels for a plurality of factors; b. Calculating a resistance score for a factor of the plurality of factors, comprising applying a machine learning algorithm trained with a training set including the expression levels of the received factors in responders and non-responders and the gender of each of the responders and non-responders to the expression levels of each received factor from the subject and the gender of the subject, and the machine learning algorithm outputs the resistance score; c. Combining the calculated resistance scores to generate an overall resistance score. Including, subjects having a composite resistance score above a predetermined threshold are predicted to not respond to the combination therapy and subjects having a composite resistance score within the predetermined threshold are predicted to respond to the combination therapy; thereby predicting the subject's response to the combination therapy.
5. 5. The method of claim 4, wherein the overall resistance score is converted to an overall response score, and an overall response score above a predetermined threshold indicates that the subject will respond to the combination therapy, and an overall response score below a predetermined threshold indicates that the subject will not respond to the combination therapy.
6. The method of claim 4 or 5, wherein the training set further comprises expression levels of the received factors in subjects (mono-responders) who are afflicted with cancer and known to respond to monotherapy comprising anti-PD-1 / PD-L1 immunotherapy, expression levels of the received factors in subjects (mono-non-responders) who are afflicted with cancer and known not to respond to the monotherapy, and the genders of the mono-responders and the mono-non-responders.
7. 1. A method for predicting the response of a subject suffering from cancer to anti-PD-1 / PD-L1 immunotherapy, comprising: a. i. in a population of subjects who are afflicted with cancer and who are known to respond to said immunotherapy (responders); ii. In a population of subjects known to be suffering from cancer and not responding to said immunotherapy (non-responders); and iii. In the subject receiving factor expression levels for a plurality of factors; b. Calculating a resistance score for a factor of the plurality of factors, comprising applying a machine learning algorithm trained with a training set including the expression levels of the received factors in responders and non-responders and the gender of each of the responders and non-responders to the expression levels of each received factor from the subject and the gender of the subject, and the machine learning algorithm outputs the resistance score; c. combining the calculated resistance scores to generate an overall resistance score; Subjects with a comprehensive resistance score above a predetermined threshold are predicted to not respond to anti-PD-1 / PD-L1 immunotherapy; thereby predicting the subject's response to anti-PD-1 / PD-L1 immunotherapy.
8. 8. The method of any one of claims 1 to 7, wherein the plurality of factors comprises at least two factors selected from the factors provided in Table 4.
9. 9. The method of claim 8, wherein the plurality of factors consists of factors selected from Table 4.
10. 10. The method of any one of claims 1 to 9, wherein the responders and non-responders are determined based on progression-free survival (PFS) at one year after initiation of the monotherapy or combination therapy.
11. 11. The method of any one of claims 1 to 10, further comprising, prior to b), selecting a subset of the plurality of factors, wherein the subset comprises factors that best distinguish between the responders and non-responders, and wherein said calculating is for each factor of the subset.
12. 12. The method of claim 11, wherein said selecting comprises applying a statistical test to the expression levels of said received factors, optionally wherein said statistical test is a Kolmogorov-Smirnov test.
13. 13. The method of claim 11 or 12, wherein the subset consists of at least 50 factors.
14. The method of any one of claims 1 to 13, wherein the expression level of the factor is from a time point prior to administration of anti-PD-1 / PD-L1 immunotherapy to the subject.
15. The method of any one of claims 1 to 14, wherein said combining is averaging.
16. 15. The method of any one of claims 1 to 14, wherein said combining comprises determining a total number of factors having a resistance score above a predetermined threshold and creating an overall resistance score proportional to said total number.
17. The method of any one of claims 1 to 16, further comprising performing a dimensionality reduction step on said plurality of factors to reduce the number of said plurality of factors.
18. 18. The method of any one of claims 1 to 17, wherein the cancer is selected from hepatobiliary cancer, cervical cancer, genitourinary cancer, anogenital cancer, prostate cancer, thyroid cancer, ovarian cancer, nervous system cancer, eye cancer, lung cancer, soft tissue cancer, bone cancer, pancreatic cancer, bladder cancer, skin cancer, intestinal cancer, liver cancer, rectal cancer, colorectal cancer, esophageal cancer, stomach cancer, gastroesophageal cancer, breast cancer, kidney cancer, skin cancer, head and neck cancer, leukemia and lymphoma.
19. 19. The method of claim 18, wherein the cancer is selected from lung cancer, skin cancer, anogenital cancer, cervical cancer, and head and neck cancer.
20. 20. The method of claim 18 or 19, wherein the cancer is non-small cell lung cancer (NSCLC).
21. The method according to any one of claims 1 to 20, wherein the cancer is a tyrosine kinase inhibitor-resistant cancer.
22. The method of any one of claims 1 to 21, wherein the predetermined threshold is determined by cross-validation within the training set or is the median score of the training set.
23. The method of any one of claims 1 to 22, wherein the plurality of factors is at least 200 factors.
24. The method according to any one of claims 1 to 23, wherein the expression level of the factor is the expression level of the factor in a biological sample provided by the subject.
25. 25. The method of claim 24, wherein the biological sample is selected from plasma, whole blood, serum, or peripheral blood mononuclear cells.
26. 26. The method of claim 25, wherein the biological sample is plasma or serum.
27. 27. The method of any one of claims 1-3 and 8-26, further comprising administering the monotherapy to the subject predicted to respond to the monotherapy, or administering a combination therapy comprising the anti-PD-1 / PD-L1 immunotherapy and chemotherapy to the subject predicted not to respond to the monotherapy.
28. 27. The method of any one of claims 4 to 6 and 8 to 26, further comprising administering the combination therapy to the subject predicted to respond to the combination therapy or administering an alternative therapy to the subject predicted not to respond to the combination therapy.
29. The method of any one of claims 7 to 26, further comprising administering the anti-PD-1 / PD-L1 immunotherapy to the subject predicted to respond to the anti-PD-1 / PD-L1 immunotherapy or administering an alternative therapy to the subject predicted not to respond to the anti-PD-1 / PD-L1 immunotherapy.
30. 30. The method of any one of claims 1 to 29, wherein the anti-PD-1 / PD-L1 immunotherapy is selected from pembrolizumab, nivolumab, durvalumab, and atezolizumab.
31. The method of any one of claims 3 to 6 and 8 to 30, wherein the chemotherapy is selected from carboplatin, paclitaxel, nab-paclitaxel, pemetrexed, vinorelbine, and cisplatin.
32. the combination therapy comprising: Carboplatin, durvalumab, and paclitaxel; b. Atezolizumab, bevacizumab, carboplatin, and paclitaxel; c. Carboplatin, nab-paclitaxel, and pembrolizumab; d. Carboplatin, nivolumab, and paclitaxel; e. Carboplatin, nivolumab, pemetrexed; f. Carboplatin, paclitaxel, pembrolizumab; g. Carboplatin, paclitaxel, pembrolizumab, and radiation; h. carboplatin and pembrolizumab; i. Carboplatin, pembrolizumab, and pemetrexed; j. carboplatin, pembrolizumab, and vinorelbine; and k. Cisplatin, pembrolizumab, and pemetrexed 32. The method of claim 31 , wherein the compound is selected from the group consisting of:
33. The method of any one of claims 1 to 32, wherein predicting response comprises predicting overall survival.
34. 33. The method of any one of claims 1 to 32, wherein predicting response comprises predicting progression-free survival.
35. 35. The method of claim 34, wherein progression-free survival is one year from the start of said monotherapy or combination therapy.
36. The method of any one of claims 4 to 35, wherein the subject is suffering from a PD-L1 negative cancer.
37. 37. The method of any one of claims 1 to 6 and 8 to 36, wherein the PD-L1 high cancer comprises at least 50% of cancer cells that are positive for surface expression of PD-L1 and the PD-L1 low or negative cancer comprises less than 50% of cancer cells that are positive for surface expression of PD-L1.
38. 38. The method of any one of claims 4 to 6 and 8 to 37, wherein the PD-L1 low or negative cancer is a PD-L1 negative cancer comprising less than 1% of cells positive for surface expression of PD-L1.
39. the trained machine learning algorithm generating a trained machine learning algorithm; During the training stage, (i) factor expression levels of resistance-associated factors in samples from subjects who are suffering from cancer and who are known to be responsive to anti-PD-1 / PD-L1 immunotherapy, and factor expression levels of resistance-associated factors in samples from subjects who are suffering from said cancer and who are known not to be responsive to said anti-PD-1 / PD-L1 immunotherapy; (ii) at least one clinical parameter of the known responsive subject and the known non-responsive subject; and (iii) a label associated with the response of a subject suffering from said cancer. training a machine learning algorithm on a training set comprising:
39. The method of any one of claims 1 to 38, wherein the trained machine learning algorithm is trained to output the resistance score.
40. 40. The method of claim 39, wherein the expression level of the resistance-associated factor and the at least one clinical parameter are labeled with the label.
41. 41. The method of claim 39 or 40, wherein a predetermined threshold for the overall resistance score is 5 and a resistance score above 5 indicates that the subject is resistant to the treatment, or wherein the overall resistance score is converted to an overall response score by the formula (10 - overall resistance score) and an overall response score above a predetermined threshold indicates that the subject is responsive to the treatment, optionally wherein the overall response score predetermined threshold is 5.