Predicting Patient Response

JP2024532762A5Pending Publication Date: 2025-08-19ONCOHOST LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024508379
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-27
Filing Date
2022-08-11
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Current immunotherapy approaches for treating non-small cell lung cancer (NSCLC) exhibit limited efficacy, with only 20-30% objective responses, and the immune mechanisms underlying these responses remain poorly understood, necessitating improved methods for predicting patient-specific responses to therapy.

Method used

A method for predicting a subject's response to therapy by calculating a resistance score based on protein expression levels and classifying factors as resistance-associated proteins (RAPs) using machine learning algorithms, incorporating clinical parameters to determine therapy resistance or responsiveness.

Benefits of technology

The method accurately predicts therapy resistance or responsiveness in patients, enabling personalized treatment decisions and potentially improving treatment outcomes by guiding therapy adjustments or alternative therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for predicting the response of a subject suffering from a disease to a therapy is provided, comprising: calculating a resistance score for factors expressed by the subject; classifying the factors as resistance-associated factors based on the resistance score; and a number of resistance-associated factors exceeding a predetermined threshold indicates that the subject is predicted to be resistant to the therapy. A method for predicting response based on the number of resistance-associated factors and at least one clinical parameter is also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 231,770, filed August 11, 2021, and U.S. Provisional Patent Application No. 63 / 324,116, filed March 27, 2022, the contents of all of which are incorporated herein by reference in their entireties.

[0002] The present invention is in the field of patient-specific diagnostics. [Background technology]

[0003] One of the major complications in oncology is resistance to therapy. Many studies have focused on the involvement of mutations and epigenetic changes in tumor cells in conferring drug resistance. However, in recent years, studies have shown that in response to most types of anticancer therapy, the patient (i.e., the host) can produce pro-tumorigenic and pro-metastatic effects. This phenomenon, called the host response, is a physiological reaction of the patient to a cancer therapy that potentially counteracts the anti-tumor activity of the treatment.

[0004] Lung cancer has the highest mortality rate among various cancer types, with approximately 2.1 million lung cancer cases and 1.8 million deaths in 2018 worldwide. More than 85% of lung cancer cases are classified as non-small cell lung cancer (NSCLC), of which the two most common histological subtypes are lung adenocarcinoma and lung squamous cell carcinoma. The treatment of NSCLC is moving away from the use of primarily chemotherapy to a more personalized approach. Currently, patient-specific genetic alterations in tumor cells determine eligibility to receive targeted agents. In particular, the status of programmed death-ligand-1 (PD-L1) expression levels in tumors determines eligibility to receive this immunotherapy.

[0005] Immunotherapy is a type of treatment based on immune response modulation. Currently, one of the most common immunotherapy approaches is the use of immune checkpoint inhibitors (ICIs), which target regulators of the immune system to stimulate the immune system. Currently, there are several approved ICIs in the form of monoclonal antibodies that target the immune checkpoint proteins CTLA4, PD-1, and PD-L1. ICI therapy has been approved for the treatment of multiple cancer types, including melanoma and NSCLC.

[0006] When used as monotherapy, these therapeutic agents present several limitations, with objective responses observed in only 20-30% of patients. Furthermore, the immune mechanisms involved in the response to these therapeutic interventions remain poorly elucidated. Thus, advanced proteomic techniques that allow for facile, non-invasive means to discover blood-based protein biomarkers hold promise for identifying host and tumor alterations associated with immunotherapy response / non-response and uncovering biological mechanisms underlying host-associated primary resistance. There is a great need for methods to determine patient-specific response to therapy, particularly immunotherapy. Summary of the Invention [Problem to be solved by the invention]

[0007] The present invention provides a method for predicting the response of a subject to a therapy. A method for predicting the response of a subject suffering from a disease to a therapy is provided, which comprises: calculating a resistance score for the proteins expressed by the subject; classifying the proteins as resistance-associated proteins (RAPs) based on the resistance score; and a number of resistance-associated proteins that exceeds a predetermined threshold indicates that the subject is predicted to be resistant to the therapy. A method for predicting response based on the number of resistance-associated factors and at least one clinical parameter is also provided. [Means for solving the problem]

[0008] According to a first aspect, there is provided a method of predicting response to a therapy in a subject suffering from a disease, comprising: a. i. In a population of subjects known to be afflicted with a disease and responsive to therapy (responders), ii. in a population of subjects known to be afflicted with a disease and who are not responsive to therapy (non-responders); and iii. In the subject receiving protein expression levels of a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, the resistance score being based on the similarity of the factor expression level in the subject to the factor expression level in the responders and the similarity of the factor expression level in the subject to the factor expression level in the non-responders; and c. classifying factors of the plurality of factors having a resistance score above a predefined threshold as resistance-associated factors; Subjects having a number of resistance-associated factors greater than a predetermined number are predicted to be resistant to the therapy, and subjects having a number of resistance-associated factors equal to or less than a predetermined number are predicted to respond to the therapy; Thereby predicting a subject's response to therapy. A method is provided that includes:

[0009] According to another aspect, there is provided a method of predicting a response to a therapy in a subject suffering from a disease, comprising: a. i. In a population of subjects known to be afflicted with a disease and responsive to therapy (responders), ii. in a population of subjects known to be afflicted with a disease and who are not responsive to therapy (non-responders); and iii. In the subject receiving factor expression levels for a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, the resistance score being based on the similarity of the factor expression levels in the subject to the factor expression levels in the responders and the similarity of the factor expression levels in the subject to the factor expression levels in the non-responders; c. classifying factors of the plurality of factors having a resistance score above a predefined threshold as resistance-associated factors; d. Summarizing the number of resistance-associated factors present in the subject; and e. applying a trained machine learning algorithm to the number of resistance-associated factors and at least one clinical parameter of the subject, where the trained machine learning algorithm outputs a final resistance score and a final resistance score that exceeds a predetermined threshold, indicating that the subject is resistant to the therapy; Thereby predicting a subject's response to therapy. A method is provided that includes:

[0010] According to another aspect, there is provided a method of predicting a response to a therapy in a subject suffering from a disease, comprising: a. i. In a population of subjects known to be afflicted with a disease and responsive to therapy (responders), ii. in a population of subjects known to be afflicted with a disease and who are not responsive to therapy (non-responders); and iii. In the subject receiving factor expression levels for a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, the resistance score being based on the similarity of the factor expression levels in the subject to the factor expression levels in the responders and the similarity of the factor expression levels in the subject to the non-responders, the calculation comprising applying a machine learning algorithm trained on a training set including the received factor expression levels in the responders and non-responders and the genders of the responders and non-responders, respectively, to each received factor expression level from the subject and gender of the subject, the machine learning algorithm outputting a resistance score; and c. summing the calculated resistance scores to generate a total resistance score; Subjects with a total resistance score above a predefined threshold are predicted to be resistant to the therapy. Thereby predicting a subject's response to therapy. A method is provided that includes:

[0011] According to another aspect, there is provided a method, comprising: During the training phase, (i) the number of resistance-associated factors expressed in samples from subjects known to be responsive to the therapy and the number of resistance-associated factors expressed in samples from subjects known to be non-responsive to the therapy; (ii) at least one clinical parameter of subjects known to be responsive and subjects known to be non-responsive; and (iii) a marker associated with the responsiveness of a subject suffering from a disease Train a machine learning algorithm on a training set including generating a trained machine learning algorithm, the trained machine learning algorithm being trained to predict the responsiveness of a subject suffering from a disease to a therapy. A method is provided.

[0012] According to another aspect, there is provided a method, comprising: During the training phase, (i) factor expression levels of resistance-associated factors in samples from subjects known to be afflicted with a disease and responsive to a therapy, and factor expression levels of resistance-associated factors in samples from subjects known to be afflicted with a disease and unresponsive to a therapy; (ii) at least one clinical parameter of subjects known to be responsive and subjects known to be non-responsive; and (iii) a marker associated with the responsiveness of a subject suffering from a disease Train a machine learning algorithm on a training set including generating a trained machine learning algorithm, the trained machine learning algorithm being trained to predict activity of the resistance-associated factor in the subject; A method is provided.

[0013] According to some embodiments, the number of resistance-associated factors and the at least one clinical parameter are labeled with a label.

[0014] According to some embodiments, the predetermined threshold for the final resistance score is 0.2, and a resistance score greater than 0.2 indicates that the subject is resistant to the therapy, or the final resistance score is converted to a response score by the formula (1-final resistance score), and a response score greater than the predetermined threshold indicates that the subject is responding to the therapy, optionally, the predetermined threshold for the response score is 0.8.

[0015] According to some embodiments, the resistance-associated factors in each subject are: a. i. In a population of subjects known to be afflicted with a disease and responsive to therapy (responders), ii. in a population of subjects known to be afflicted with a disease and who are not responsive to therapy (non-responders); and iii. In the subject receiving factor expression levels for a plurality of factors; b. calculating a resistance score for each factor of the plurality of proteins, the resistance score being based on the similarity of the factor expression level in each subject to the factor expression level in the responders and the similarity of the factor expression level in the subject to the factor expression level in the non-responders; and c. classifying factors of the plurality of factors having a resistance score above a predefined threshold as resistance-associated factors. The method is determined by a method including:

[0016] According to some embodiments, the method includes, prior to (b), selecting a subset of the plurality of factors, the subset including factors that best distinguish between responders and non-responders, and the calculating is for each factor of the subset.

[0017] According to some embodiments, the selecting comprises applying a statistical test to the received factor expression levels, optionally, the statistical test is a Kolmogorov-Smirnov test.

[0018] According to some embodiments, the calculating comprises applying a machine learning algorithm trained on a training set including the received factor expression levels in responders and non-responders to each received factor expression level from the subject, where the machine learning algorithm outputs a resistance score.

[0019] According to some embodiments, the training set further comprises the gender of each responder and non-responder, and the machine learning algorithm is applied to the individual factor expression levels received from the subjects and gender of the subjects.

[0020] According to some embodiments, the method further comprises performing a dimensionality reduction step on the multiple factors to reduce the number of multiple factors.

[0021] According to some embodiments, the dimensionality reduction step identifies a subset of key factors, and the training set includes only expression levels of the subset of key factors, optionally the subset of key factors are those factors that most evenly balance the number of predicted responders and non-responders.

[0022] According to some embodiments, the predetermined threshold is determined by performing cross-validation within the training set.

[0023] According to some embodiments, the calculating comprises calculating a mean expression for each factor in the responders and non-responders, and the resistance score is based on a ratio of the deviation of the factor expression in the subject from the calculated mean in the responders to the deviation of the factor expression in the subject from the calculated mean in the non-responders.

[0024] According to some embodiments, the calculating further comprises calculating the distribution and standard deviation for each factor in responders and non-responders, where the deviation is measured as a multiple of the calculated standard deviation.

[0025] According to some embodiments, the resistance score is calculated based on a monotonic function of the formula

number

[0026] According to some embodiments, the pre-defined threshold for the resistance score is about 2.9, and a resistance score above 2.9 indicates that the factor is a resistance-associated factor.

[0027] According to some embodiments, the plurality of factors is at least 200 factors.

[0028] According to some embodiments, the predetermined number of resistance-associated factors is three, and subjects having more than three resistance-associated factors are predicted to be resistant to the therapy.

[0029] According to some embodiments, receiving factor expression levels for a plurality of factors comprises: a. receiving factor expression levels for a group of greater than a plurality of factors in a population of responders and a population of non-responders; b. for each factor of the group, applying a machine learning algorithm trained on the received factor expression levels in the responders and non-responders; c. The algorithm selects subgroups of factors that most evenly divide the subjects in the population into responders and non-responders; and d. Specifying subgroups of factors as multiple factors; Includes.

[0030] According to some embodiments, the factor expression level is the factor expression level of a biological sample provided by the subject.

[0031] According to some embodiments, the biological sample is selected from plasma, whole blood, serum, or peripheral blood mononuclear cells.

[0032] According to some embodiments, the biological sample is plasma.

[0033] According to some embodiments, the biological sample is provided by the subject prior to receiving the therapy.

[0034] According to some embodiments, prior is up to 24 hours prior.

[0035] According to some embodiments, the biological sample is provided by a subject after undergoing a therapy.

[0036] According to some embodiments, the post-receiving therapy is after a first treatment with the therapy.

[0037] According to some embodiments, later is at least 24 hours later.

[0038] According to some embodiments, the biological sample provided by each subject in the population is the same type of biological sample.

[0039] According to some embodiments, the responder population, the non-responder population, and the biological samples provided by the subjects are all the same type of biological sample.

[0040] According to some embodiments, the trained machine learning algorithm is trained by the methods of the present invention.

[0041] According to some embodiments, the disease is cancer.

[0042] According to some embodiments, the therapy is immune checkpoint inhibition, and optionally, the immune checkpoint inhibition inhibits the PD-1 / PD-L1 axis.

[0043] According to some embodiments, the at least one clinical parameter is selected from subject age, sex, line of treatment, and expression of a biomarker in a sample from the subject.

[0044] According to some embodiments, the disease is cancer, the therapy is an anti-PD-1 or anti-PD-L1 therapy, and the target expression in the sample is PD-L1 expression in a tumor sample.

[0045] According to some embodiments, the method further comprises, in the inference step, receiving as input the number of resistance-associated factors expressed in a sample from a subject suffering from a disease and with unknown responsiveness to the therapy and at least one clinical parameter of the subject with unknown responsiveness, and applying the trained machine learning algorithm to the received input to predict the responsiveness of the subject with unknown responsiveness to the therapy.

[0046] According to some embodiments, the method further comprises administering the therapy or continuing to administer the therapy to the subject predicted to respond to the therapy.

[0047] According to some embodiments, the method further comprises discontinuing the therapy or not administering the therapy to a subject predicted to be resistant to the therapy.

[0048] According to some embodiments, the method further comprises administering an alternative therapy to the subject predicted to be resistant to the therapy.

[0049] According to some embodiments, the method further comprises administering to a subject predicted to be resistant to the therapy, the therapy in combination with an agent that modulates at least one of the resistance-associated factors or a factor of a functional pathway that includes at least one resistance-associated factor; a. at least one resistance-associated factor is more highly expressed in non-responders than in responders, increasing the activity of a pathway, and the agent inhibits the resistance-associated factor or pathway, or b. at least one resistance-associated factor is more highly expressed in non-responders than in responders, reducing activity of the pathway, and the agent inhibits the resistance-associated factor or activates the pathway; c. at least one resistance-associated factor is expressed less in non-responders than in responders, increases the activity of a pathway, and the agent activates the resistance-associated factor or pathway, or d. At least one resistance-associated factor is expressed less in non-responders than in responders, reducing activity of a pathway, and the agent activates the resistance-associated factor or inhibits the pathway.

[0050] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description given hereinafter. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description. [Brief description of the drawings]

[0051] [Figure 1A] Graphical representation of protein expression distribution in responder and non-responder populations at the single protein level. (1A-1C) Computer-generated examples of distribution of protein expression for responder and non-responder populations. Shown are examples of protein expression levels that may be considered RAP (light grey dashed line) or not RAP (dark grey dashed line) based on population expression distribution data. [Figure 1B]Graphical representation of protein expression distribution in responder and non-responder populations at the single protein level. (1A-1C) Computer-generated examples of distribution of protein expression for responder and non-responder populations. Shown are examples of protein expression levels that may be considered RAP (light grey dashed line) or not RAP (dark grey dashed line) based on population expression distribution data. [Figure 1C] Graphical representation of protein expression distribution in responder and non-responder populations at the single protein level. (1A-1C) Computer-generated examples of distribution of protein expression for responder and non-responder populations. Shown are examples of protein expression levels that may be considered RAP (light grey dashed line) or not RAP (dark grey dashed line) based on population expression distribution data. [Diagram 2] Illustration of the RAP score of Equation 2 implemented in Algorithm 1. The RAP score was calculated using synthetic data, where the responder and non-responder populations were generated by sampling from normal distributions. Expression levels of the synthetic populations are shown in a histogram, with the responder population in dark grey and the non-responder population in light grey. Taking these distributions into account, the RAP score was calculated for each expression level. The resulting RAP scores are plotted on a blue curve and the values ​​are shown on the secondary Y-axis on the right. [Figure 3A] Determination of RAP score threshold based on AUC as a function of RAP score. The AUC at each RAP score was calculated and the peak of the resulting curve was determined as the threshold (dotted line) for determining whether a particular protein is a RAP or not. (3A-3B) Graphs describing the determination of the RAP score threshold using a mathematical approach for protein measurement at (3A) T1 and (3B) T0. (3C) Graph illustrating the determination of the RAP score threshold using a machine learning approach. [Figure 3B]Determination of RAP score threshold based on AUC as a function of RAP score. The AUC at each RAP score was calculated and the peak of the resulting curve was determined as the threshold (dotted line) for determining whether a particular protein is a RAP or not. (3A-3B) Graphs describing the determination of the RAP score threshold using a mathematical approach for protein measurement at (3A) T1 and (3B) T0. (3C) Graph illustrating the determination of the RAP score threshold using a machine learning approach. [Figure 3C] Determination of RAP score threshold based on AUC as a function of RAP score. The AUC at each RAP score was calculated and the peak of the resulting curve was determined as the threshold (dotted line) for determining whether a particular protein is a RAP or not. (3A-3B) Graphs describing the determination of the RAP score threshold using a mathematical approach for protein measurement at (3A) T1 and (3B) T0. (3C) Graph illustrating the determination of the RAP score threshold using a machine learning approach. [Figure 4A] (FIG. 4A) Bar graph showing the number of RAPs for each patient in the study cohort (n=30). Responders-light grey; non-responders-dark grey. (4B) ROC curve for RAP analysis. [Figure 4B] (FIG. 4A) Bar graph showing the number of RAPs for each patient in the study cohort (n=30). Responders-light grey; non-responders-dark grey. (4B) ROC curve for RAP analysis. [Diagram 5] Heatmap of cancer features significantly enriched in six patients. Enrichment analysis was based on Fisher's exact test (FDR<0.05). Next to each patient identifier, the number of RAPs for that patient is shown in parentheses. [Figure 6]Protein-protein network of significant RAPs in the current cohort. The network is based on the STRING database. Each node (protein) is colored based on the cancer feature with which it is associated. Black circular boxes indicate RAPs that can be targeted. The size of each node correlates with the number of patients with RAPs examined. The compartments of each node are shown in the center (based on the Human Protein Atlas). I, intracellular. M, membrane. S, soluble. Proteins can have more than one compartment. [Figure 7] Chart of significantly enriched features of cancer among the 19 RAPs. Analysis was performed using Fisher's exact test. An enrichment factor greater than 1 indicates enrichment. [Figure 8] Heatmap of protein expression levels of 19 RAPs in healthy tissues. Expression data are based on the Human Protein Atlas (HPA) database. [Figure 9] Heatmap of the percentage of moderate to high staining in patients with various cancer types, including NSCLC. Expression data is based on the Human Protein Atlas (HPA) database. [Figure 10A]Clinical description of the 184 patients included in the analysis. (10A) Heatmaps representing the clinical characteristics of the patients: response to treatment (ORR1, ORR2, 1-year DCB); percent of cells expressing PD-L1 in biopsy immunostaining, a prognostic marker of response to treatment; treatment type: ICI only or combination ICI and chemotherapy treatment; treatment line: first line indicates ICI treatment was given as the first systemic treatment for NSCLC, advanced line indicates previous non-ICI treatment was given before the current ICI treatment was administered. Gender indicates the patient's gender at birth. Histology indicates lung cancer histology type (ADC-adenocarcinoma, SCC-squamous cell carcinoma). (10B-10C) Violin plots of the correlation between patient age and response at each time point: (10B) ORR1 and (10C) ORR2. ORR1 and ORR2 are the overall response determined at 3 and 6 months after treatment initiation, respectively. (10D-10E) Graphical representation of response groups in (10D)ORR1 and (10E)ORR2. NR=non-responder. R=responder (partial or complete responder). SD=stable disease (included in the responder group in the model). (10F) Graphical representation of the division of the population into development and validation sets. [Figure 10B]Clinical description of the 184 patients included in the analysis. (10A) Heatmaps representing the clinical characteristics of the patients: response to treatment (ORR1, ORR2, 1-year DCB); percent of cells expressing PD-L1 in biopsy immunostaining, a prognostic marker of response to treatment; treatment type: ICI only or combination ICI and chemotherapy treatment; treatment line: first line indicates ICI treatment was given as the first systemic treatment for NSCLC, advanced line indicates previous non-ICI treatment was given before the current ICI treatment was administered. Gender indicates the patient's gender at birth. Histology indicates lung cancer histology type (ADC-adenocarcinoma, SCC-squamous cell carcinoma). (10B-10C) Violin plots of the correlation between patient age and response at each time point: (10B) ORR1 and (10C) ORR2. ORR1 and ORR2 are the overall response determined at 3 and 6 months after treatment initiation, respectively. (10D-10E) Graphical representation of response groups in (10D)ORR1 and (10E)ORR2. NR=non-responder. R=responder (partial or complete responder). SD=stable disease (included in the responder group in the model). (10F) Graphical representation of the division of the population into development and validation sets. [Figure 10C]Clinical description of the 184 patients included in the analysis. (10A) Heatmaps representing the clinical characteristics of the patients: response to treatment (ORR1, ORR2, 1-year DCB); percent of cells expressing PD-L1 in biopsy immunostaining, a prognostic marker of response to treatment; treatment type: ICI only or combination ICI and chemotherapy treatment; treatment line: first line indicates ICI treatment was given as the first systemic treatment for NSCLC, advanced line indicates previous non-ICI treatment was given before the current ICI treatment was administered. Gender indicates the patient's gender at birth. Histology indicates lung cancer histology type (ADC-adenocarcinoma, SCC-squamous cell carcinoma). (10B-10C) Violin plots of the correlation between patient age and response at each time point: (10B) ORR1 and (10C) ORR2. ORR1 and ORR2 are the overall response determined at 3 and 6 months after treatment initiation, respectively. (10D-10E) Graphical representation of response groups in (10D)ORR1 and (10E)ORR2. NR=non-responder. R=responder (partial or complete responder). SD=stable disease (included in the responder group in the model). (10F) Graphical representation of the division of the population into development and validation sets. [Figure 10D]Clinical description of the 184 patients included in the analysis. (10A) Heatmaps representing the clinical characteristics of the patients: response to treatment (ORR1, ORR2, 1-year DCB); percent of cells expressing PD-L1 in biopsy immunostaining, a prognostic marker of response to treatment; treatment type: ICI only or combination ICI and chemotherapy treatment; treatment line: first line indicates ICI treatment was given as the first systemic treatment for NSCLC, advanced line indicates previous non-ICI treatment was given before the current ICI treatment was administered. Gender indicates the patient's gender at birth. Histology indicates lung cancer histology type (ADC-adenocarcinoma, SCC-squamous cell carcinoma). (10B-10C) Violin plots of the correlation between patient age and response at each time point: (10B) ORR1 and (10C) ORR2. ORR1 and ORR2 are the overall response determined at 3 and 6 months after treatment initiation, respectively. (10D-10E) Graphical representation of response groups in (10D)ORR1 and (10E)ORR2. NR=non-responder. R=responder (partial or complete responder). SD=stable disease (included in the responder group in the model). (10F) Graphical representation of the division of the population into development and validation sets. [Figure 10E]Clinical description of the 184 patients included in the analysis. (10A) Heatmaps representing the clinical characteristics of the patients: response to treatment (ORR1, ORR2, 1-year DCB); percent of cells expressing PD-L1 in biopsy immunostaining, a prognostic marker of response to treatment; treatment type: ICI only or combination ICI and chemotherapy treatment; treatment line: first line indicates ICI treatment was given as the first systemic treatment for NSCLC, advanced line indicates previous non-ICI treatment was given before the current ICI treatment was administered. Gender indicates the patient's gender at birth. Histology indicates lung cancer histology type (ADC-adenocarcinoma, SCC-squamous cell carcinoma). (10B-10C) Violin plots of the correlation between patient age and response at each time point: (10B) ORR1 and (10C) ORR2. ORR1 and ORR2 are the overall response determined at 3 and 6 months after treatment initiation, respectively. (10D-10E) Graphical representation of response groups in (10D)ORR1 and (10E)ORR2. NR=non-responder. R=responder (partial or complete responder). SD=stable disease (included in the responder group in the model). (10F) Graphical representation of the division of the population into development and validation sets. [Figure 10F]Clinical description of the 184 patients included in the analysis. (10A) Heatmaps representing the clinical characteristics of the patients: response to treatment (ORR1, ORR2, 1-year DCB); percent of cells expressing PD-L1 in biopsy immunostaining, a prognostic marker of response to treatment; treatment type: ICI only or combination ICI and chemotherapy treatment; treatment line: first line indicates ICI treatment was given as the first systemic treatment for NSCLC, advanced line indicates previous non-ICI treatment was given before the current ICI treatment was administered. Gender indicates the patient's gender at birth. Histology indicates lung cancer histology type (ADC-adenocarcinoma, SCC-squamous cell carcinoma). (10B-10C) Violin plots of the correlation between patient age and response at each time point: (10B) ORR1 and (10C) ORR2. ORR1 and ORR2 are the overall response determined at 3 and 6 months after treatment initiation, respectively. (10D-10E) Graphical representation of response groups in (10D)ORR1 and (10E)ORR2. NR=non-responder. R=responder (partial or complete responder). SD=stable disease (included in the responder group in the model). (10F) Graphical representation of the division of the population into development and validation sets. [Figure 11A] Performance of classification models. ROC AUC was calculated using the final resistance score along with the actual overall response assessment at 3 months ORR, 6 months ORR, and 1-year sustained clinical benefit (DCB) for both T0 and T1. (11A, upper panel) Results at T0 for the development set and (11A, lower panel and 11B, upper panel) validation set are shown. (11B, lower panel) A similar classification model was generated based on T1. [Figure 11B] Performance of classification models. ROC AUC was calculated using the final resistance score along with the actual overall response assessment at 3 months ORR, 6 months ORR, and 1-year sustained clinical benefit (DCB) for both T0 and T1. (11A, upper panel) Results at T0 for the development set and (11A, lower panel and 11B, upper panel) validation set are shown. (11B, lower panel) A similar classification model was generated based on T1. [Figure 12A](12A) Patients classified with a response probability score (calculated as 1-resistance score) based on protein levels at T0. The actual observed response at 3 months ORR is indicated by color for each patient. (12B) Dot plot of the agreement between the predicted response probability based on T0 protein expression and the observed response probability at either 3 months, 6 months, or 1 year. Each point on the graph represents a specific patient, and different time points are indicated by different hues and marker types. The black diagonal line represents the y=x line, and the red diagonal line represents the regression line fitted for all points, indicating the goodness of fit of the regression (R2). The horizontal line represents the average observed response probability for the three time points (color coded) across the entire validation set. [Figure 12B] (12A) Patients classified with a response probability score (calculated as 1-resistance score) based on protein levels at T0. The actual observed response at 3 months ORR is indicated by color for each patient. (12B) Dot plot of the agreement between the predicted response probability based on T0 protein expression and the observed response probability at either 3 months, 6 months, or 1 year. Each point on the graph represents a specific patient, and different time points are indicated by different hues and marker types. The black diagonal line represents the y=x line, and the red diagonal line represents the regression line fitted for all points, indicating the goodness of fit of the regression (R2). The horizontal line represents the average observed response probability for the three time points (color coded) across the entire validation set. [Figure 13A] Survival analysis based on predicted outcomes for ORR at 3 months based on T0 protein measurements for (13A) PFS and (13B) OS. [Figure 13B] Survival analysis based on predicted outcomes for ORR at 3 months based on T0 protein measurements for (13A) PFS and (13B) OS. [Figure 14A](14A) Network of functions of all potential RAPs from this analysis. Each node represents a RAP and edges between nodes indicate functional relationships. Nodes with larger size and provided protein names indicate investigational new drugs (IND) in combination with immunotherapy. Nodes are colored based on protein function. (14B) Functional network of two patients, predicted non-responder (top) and predicted responder (bottom). RAPs detected for each patient are outlined in black. The non-responder patient had 44 RAPs detected and the responder had 10 RAPs detected. [Figure 14B] (14A) Network of functions of all potential RAPs from this analysis. Each node represents a RAP and edges between nodes indicate functional relationships. Nodes with larger size and provided protein names indicate investigational new drugs (IND) in combination with immunotherapy. Nodes are colored based on protein function. (14B) Functional network of two patients, predicted non-responder (top) and predicted responder (bottom). RAPs detected for each patient are outlined in black. The non-responder patient had 44 RAPs detected and the responder had 10 RAPs detected. [Figure 15] Functional differentiation between RAPs was higher in each response group. Each polygon in the Voronoi plot represents a RAP, and the size correlates with the difference between responders and non-responders. Non-responder RAPs were involved in splicing, signal transduction, and cytoskeleton-related processes, whereas responder RAPs were mainly involved in protein degradation and cell adhesion. Each color indicates a different overall function. [Figure 16] Table describing the clinical parameters of the 339 patients included in the analysis. [Figure 17] Line graphs of the number of patients at each time point are shown in total by response group (NR, non-responder; R, responder). Patient cohorts were divided into development and validation sets. [Figure 18A]Performance of clinical parameter-based prediction models. (18A) Receiver operating characteristic (ROC) plot of the PD-L1-based prediction model. (18B) ROC plot of the prediction model based on PD-L1, age and line of treatment. Area under the curve (AUC) values ​​are shown for each time point. [Figure 18B] Performance of clinical parameter-based prediction models. (18A) Receiver operating characteristic (ROC) plot of the PD-L1-based prediction model. (18B) ROC plot of the prediction model based on PD-L1, age and line of treatment. Area under the curve (AUC) values ​​are shown for each time point. [Figure 19A](19A) Development of RAP-based predictive models. A cohort of stage IV NSCLC patients receiving ICB-based therapy was recruited. Pretreatment blood samples were obtained and plasma proteomes were profiled. Clinical response to treatment was assessed at 3, 6, and 12 months after treatment initiation, while patients were followed up for up to 2 years. For each response assessment time point, a predictive model of ICB response was developed as follows: Proteins showing differential plasma levels in responders and non-responders (collectively referred to as response-associated proteins; RAPs) were selected for training the model using statistical tests. Predictive models of response were developed for each RAP using machine learning algorithms. Response predictions inferred from each RAP were summed to obtain a RAP score for each patient. RAP scores were linearly scaled to values ​​between 0 and 1, allowing the conversion of a given patient's RAP score into a probability of response. (19B) Development and validation of RAP-based models. The cohort was split into a development set and a validation set (75% and 25%, respectively). The development set was further randomly split into a training set and a test set (75% and 25%, respectively). The training set was used for RAP selection and subsequent model training, resulting in a predictive model for each RAP. A response prediction was then generated for each RAP for each patient in the test set. The response predictions from all selected RAPs were summed to obtain a RAP score for each patient in the test set. This process was repeated 80 times, each time randomly splitting the patients in the development set into a training set and a test set. The RAP scores were averaged for each patient in the development set and then linearly scaled, allowing the RAP score for a given patient to be converted into a response probability (a value between 0 and 1). The model was then locked and tested on an independent validation set. [Figure 19B](19A) Development of RAP-based predictive models. A cohort of stage IV NSCLC patients receiving ICB-based therapy was recruited. Pretreatment blood samples were obtained and plasma proteomes were profiled. Clinical response to treatment was assessed at 3, 6, and 12 months after treatment initiation, while patients were followed up for up to 2 years. For each response assessment time point, a predictive model of ICB response was developed as follows: Proteins showing differential plasma levels in responders and non-responders (collectively referred to as response-associated proteins; RAPs) were selected for training the model using statistical tests. Predictive models of response were developed for each RAP using machine learning algorithms. Response predictions inferred from each RAP were summed to obtain a RAP score for each patient. RAP scores were linearly scaled to values ​​between 0 and 1, allowing the conversion of a given patient's RAP score into a probability of response. (19B) Development and validation of RAP-based models. The cohort was split into a development set and a validation set (75% and 25%, respectively). The development set was further randomly split into a training set and a test set (75% and 25%, respectively). The training set was used for RAP selection and subsequent model training, resulting in a predictive model for each RAP. A response prediction was then generated for each RAP for each patient in the test set. The response predictions from all selected RAPs were summed to obtain a RAP score for each patient in the test set. This process was repeated 80 times, each time randomly splitting the patients in the development set into a training set and a test set. The RAP scores were averaged for each patient in the development set and then linearly scaled, allowing the RAP score for a given patient to be converted into a response probability (a value between 0 and 1). The model was then locked and tested on an independent validation set. [Figure 20A]RAP identification during model development. (20A) Histogram showing the number of identified RAPs grouped according to the number of times they were selected across 80 iterations. The top, middle and bottom histograms are for the 3, 6 and 12 month time points, respectively. (20B) Total number of RAPs identified per time point. Some proteins quantified by this method are redundant. The number of total and non-redundant RAPs are shown in light and medium grey bars, respectively). The number of non-redundant RAPs identified at least 40 times in a total of 80 iterations is shown in dark grey bars. (20C) Venn diagram showing the number of RAPs identified per time point. (20D) Hierarchical clustering-based heatmap showing the number of iterations in which a given protein was classified as a RAP. (20E) Voronoi plot displaying the main biological functions of RAPs per time point. Each polygon represents a RAP and the size correlates with the number of times the protein was selected as a RAP. [Figure 20B] RAP identification during model development. (20A) Histogram showing the number of identified RAPs grouped according to the number of times they were selected across 80 iterations. The top, middle and bottom histograms are for the 3, 6 and 12 month time points, respectively. (20B) Total number of RAPs identified per time point. Some proteins quantified by this method are redundant. The number of total and non-redundant RAPs are shown in light and medium grey bars, respectively). The number of non-redundant RAPs identified at least 40 times in a total of 80 iterations is shown in dark grey bars. (20C) Venn diagram showing the number of RAPs identified per time point. (20D) Hierarchical clustering-based heatmap showing the number of iterations in which a given protein was classified as a RAP. (20E) Voronoi plot displaying the main biological functions of RAPs per time point. Each polygon represents a RAP and the size correlates with the number of times the protein was selected as a RAP. [Figure 20C]RAP identification during model development. (20A) Histogram showing the number of identified RAPs grouped according to the number of times they were selected across 80 iterations. The top, middle and bottom histograms are for the 3, 6 and 12 month time points, respectively. (20B) Total number of RAPs identified per time point. Some proteins quantified by this method are redundant. The number of total and non-redundant RAPs are shown in light and medium grey bars, respectively). The number of non-redundant RAPs identified at least 40 times in a total of 80 iterations is shown in dark grey bars. (20C) Venn diagram showing the number of RAPs identified per time point. (20D) Hierarchical clustering-based heatmap showing the number of iterations in which a given protein was classified as a RAP. (20E) Voronoi plot displaying the main biological functions of RAPs per time point. Each polygon represents a RAP and the size correlates with the number of times the protein was selected as a RAP. [Figure 20D] RAP identification during model development. (20A) Histogram showing the number of identified RAPs grouped according to the number of times they were selected across 80 iterations. The top, middle and bottom histograms are for the 3, 6 and 12 month time points, respectively. (20B) Total number of RAPs identified per time point. Some proteins quantified by this method are redundant. The number of total and non-redundant RAPs are shown in light and medium grey bars, respectively). The number of non-redundant RAPs identified at least 40 times in a total of 80 iterations is shown in dark grey bars. (20C) Venn diagram showing the number of RAPs identified per time point. (20D) Hierarchical clustering-based heatmap showing the number of iterations in which a given protein was classified as a RAP. (20E) Voronoi plot displaying the main biological functions of RAPs per time point. Each polygon represents a RAP and the size correlates with the number of times the protein was selected as a RAP. [Figure 20E]RAP identification during model development. (20A) Histogram showing the number of identified RAPs grouped according to the number of times they were selected across 80 iterations. The top, middle and bottom histograms are for the 3, 6 and 12 month time points, respectively. (20B) Total number of RAPs identified per time point. Some proteins quantified by this method are redundant. The number of total and non-redundant RAPs are shown in light and medium grey bars, respectively). The number of non-redundant RAPs identified at least 40 times in a total of 80 iterations is shown in dark grey bars. (20C) Venn diagram showing the number of RAPs identified per time point. (20D) Hierarchical clustering-based heatmap showing the number of iterations in which a given protein was classified as a RAP. (20E) Voronoi plot displaying the main biological functions of RAPs per time point. Each polygon represents a RAP and the size correlates with the number of times the protein was selected as a RAP. [Figure 21] Performance of the RAP-based predictive model. Waterfall plot showing predicted response probabilities sorted from lowest to highest. Observed responders and non-responders are shown as light and dark grey bars, respectively. [Figure 22A] Comparison between response probabilities at successive time points. Each dot represents a patient in the cohort. The probability of response at one time point is plotted against the probability of response at the subsequent time point. Color indicates the patient response indicator at each time point and whether the response indicator changed between time points. R, responder; NR, non-responder. (22A). Comparison between 3 months and 6 months. (22B). Comparison between 3 months and 12 months. (22C). Comparison between 6 months and 12 months. (22D). Sankey plot displaying the flow of response indicators over time. R, responder; NR, non-responder; NA, not available. [Figure 22B]Comparison between response probabilities at successive time points. Each dot represents a patient in the cohort. The probability of response at one time point is plotted against the probability of response at the subsequent time point. Color indicates the patient response indicator at each time point and whether the response indicator changed between time points. R, responder; NR, non-responder. (22A). Comparison between 3 months and 6 months. (22B). Comparison between 3 months and 12 months. (22C). Comparison between 6 months and 12 months. (22D). Sankey plot displaying the flow of response indicators over time. R, responder; NR, non-responder; NA, not available. [Figure 22C] Comparison between response probabilities at successive time points. Each dot represents a patient in the cohort. The probability of response at one time point is plotted against the probability of response at the subsequent time point. Color indicates the patient response indicator at each time point and whether the response indicator changed between time points. R, responder; NR, non-responder. (22A). Comparison between 3 months and 6 months. (22B). Comparison between 3 months and 12 months. (22C). Comparison between 6 months and 12 months. (22D). Sankey plot displaying the flow of response indicators over time. R, responder; NR, non-responder; NA, not available. [Figure 22D] Comparison between response probabilities at successive time points. Each dot represents a patient in the cohort. The probability of response at one time point is plotted against the probability of response at the subsequent time point. Color indicates the patient response indicator at each time point and whether the response indicator changed between time points. R, responder; NR, non-responder. (22A). Comparison between 3 months and 6 months. (22B). Comparison between 3 months and 12 months. (22C). Comparison between 6 months and 12 months. (22D). Sankey plot displaying the flow of response indicators over time. R, responder; NR, non-responder; NA, not available. [Diagram 23]Enrichment analysis for response probability and observed proportion at each time point. Enrichment analysis was performed using 2D enrichment test. X-axis shows enrichment factor of predicted response probability. Y-axis shows enrichment factor of observed proportion (defined by the proportion of responders within ±0.15 response probability). Positive and negative values ​​indicate enrichment in high or low response probability or observed proportion, respectively. Solid line shows Y=X line. [Figure 24A] (24A) Overall survival analysis of patients stratified into high and low response probability groups. The median response probability at each time point was used as the stratification threshold. (24B) Progression-free survival analysis of patients stratified into high and low response probability groups. The median response probability at each time point was used as the stratification threshold. HR, hazard ratio. CI, confidence interval. [Figure 24B] (24A) Overall survival analysis of patients stratified into high and low response probability groups. The median response probability at each time point was used as the stratification threshold. (24B) Progression-free survival analysis of patients stratified into high and low response probability groups. The median response probability at each time point was used as the stratification threshold. HR, hazard ratio. CI, confidence interval. [Figure 25A] (25A) Predicted response probability as a function of observed response rate. Each dot represents a patient. The observed response rate for each predicted response probability data point refers to the proportion of observed responders among the patient group assigned a response probability ±0.15. Y=X is shown as a black line. Goodness of fit is indicated. (25B) Receiver operating characteristic (ROC) of the RAP-based model is plotted for each time point. Area under the curve (AUC) is shown. The red dashed line indicates AUC=0.5. [Figure 25B] (25A) Predicted response probability as a function of observed response rate. Each dot represents a patient. The observed response rate for each predicted response probability data point refers to the proportion of observed responders among the patient group assigned a response probability ±0.15. Y=X is shown as a black line. Goodness of fit is indicated. (25B) Receiver operating characteristic (ROC) of the RAP-based model is plotted for each time point. Area under the curve (AUC) is shown. The red dashed line indicates AUC=0.5. [Figure 26A]The RAP-based model outperforms the clinical parameter-based model. Predictive performance was compared across five models: RAP-based model (RAP); PD-L1-based model (PD-L1); clinical model (CM); integrated RAP+PD-L1; integrated RAP+CM. (26A) ROC curve plots of the five models at each time point. CM, clinical model. Dashed lines indicate AUC=0.5. (26B) Forest plots comparing the five models. Top, Cox regression analysis based on overall survival (OS) data. Bottom, Cox regression analysis based on progression-free survival (PFS) data. [Figure 26B] The RAP-based model outperforms the clinical parameter-based model. Predictive performance was compared across five models: RAP-based model (RAP); PD-L1-based model (PD-L1); clinical model (CM); integrated RAP+PD-L1; integrated RAP+CM. (26A) ROC curve plots of the five models at each time point. CM, clinical model. Dashed lines indicate AUC=0.5. (26B) Forest plots comparing the five models. Top, Cox regression analysis based on overall survival (OS) data. Bottom, Cox regression analysis based on progression-free survival (PFS) data. [Figure 27A] Performance of the RAP-based clinical prediction model at 3 months (27A), 6 months (27B), and 1 year (27C). (27D) Predicted response probability as a function of observed response rate. Each dot represents a patient. The observed response rate for each predicted response probability data point refers to the proportion of observed responders among the patient group assigned a response probability ±0.15. Y=X is shown as a black line. Goodness of fit is indicated. [Figure 27B] Performance of the RAP-based clinical prediction model at 3 months (27A), 6 months (27B), and 1 year (27C). (27D) Predicted response probability as a function of observed response rate. Each dot represents a patient. The observed response rate for each predicted response probability data point refers to the proportion of observed responders among the patient group assigned a response probability ±0.15. Y=X is shown as a black line. Goodness of fit is indicated. [Figure 27C]Performance of the RAP-based clinical prediction model at 3 months (27A), 6 months (27B), and 1 year (27C). (27D) Predicted response probability as a function of observed response rate. Each dot represents a patient. The observed response rate for each predicted response probability data point refers to the proportion of observed responders among the patient group assigned a response probability ±0.15. Y=X is shown as a black line. Goodness of fit is indicated. [Figure 27D] Performance of the RAP-based clinical prediction model at 3 months (27A), 6 months (27B), and 1 year (27C). (27D) Predicted response probability as a function of observed response rate. Each dot represents a patient. The observed response rate for each predicted response probability data point refers to the proportion of observed responders among the patient group assigned a response probability ±0.15. Y=X is shown as a black line. Goodness of fit is indicated. [Figure 28] Performance of the RAP-based model on different patient subsets. [Figure 29A] The RAP-based model identifies a PD-L1 high subset of patients who may benefit from combination therapy. (29A) Kaplan-Meier plots of the three PD-L1 groups. Left: OS; Right: PFS. (29B) Overall survival analysis of PD-L1 high and PD-L1 low negative subgroups under the treatment modalities of ICB (part I) or ICB chemotherapy (part II). [Figure 29B] The RAP-based model identifies a PD-L1 high subset of patients who may benefit from combination therapy. (29A) Kaplan-Meier plots of the three PD-L1 groups. Left: OS; Right: PFS. (29B) Overall survival analysis of PD-L1 high and PD-L1 low negative subgroups under the treatment modalities of ICB (part I) or ICB chemotherapy (part II). [Figure 30A] Progression-free survival analysis of PD-L1 high and PD-L1 low negative subgroups under the treatment modality of ICB (30A) or ICB chemotherapy (30B). [Figure 30B] Progression-free survival analysis of PD-L1 high and PD-L1 low negative subgroups under the treatment modality of ICB (30A) or ICB chemotherapy (30B). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0052] The present invention, in some embodiments, provides methods of predicting a subject's response to a therapy.

[0053] According to a first aspect, there is provided a method of predicting a subject's response to a therapy, comprising the steps of: a. i. In a population of subjects known to respond to therapy (responders), ii. in populations of subjects known to not respond to therapy (non-responders), and iii. In the subject receiving expression levels of a plurality of factors; b. calculating a resistance score for at least one factor of the plurality of factors; and c. classifying factors with resistance scores above a threshold as resistance-associated factors; Subjects having a number of resistance-associated factors greater than a predetermined number are predicted to be resistant to the therapy, thereby predicting the subject's response to the therapy. A method is provided, comprising:

[0054] According to another aspect, there is provided a method of predicting a subject's response to a therapy, comprising: a. i. In a population of subjects known to respond to therapy (responders), ii. in populations of subjects known to not respond to therapy (non-responders), and iii. In the subject receiving expression levels of a plurality of factors; b. calculating a resistance score for at least one factor of the plurality of factors; c. classifying factors with resistance scores above a threshold as resistance-associated factors; d. Sum the number of resistance-associated factors, and e. applying a trained machine learning algorithm to the number of resistance-associated factors and the at least one clinical parameter, wherein the trained machine learning algorithm outputs a final resistance score and a final resistance score that exceeds a predetermined threshold, indicating that the subject is resistant to the therapy; Thereby predicting a subject's response to therapy. A method is provided that includes:

[0055] According to another aspect, there is provided a method of predicting a subject's response to a therapy, comprising: a. i. In a population of subjects known to respond to therapy (responders), ii. in populations of subjects known to not respond to therapy (non-responders), and iii. In the subject receiving expression levels of a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, the resistance score being based on the similarity of the factor expression levels in the subject to the factor expression levels in the responders and the similarity of the factor expression levels in the subject to the factor expression levels in the non-responders, the calculation comprising applying a trained machine learning algorithm that outputs a resistance score; and c. summing the calculated resistance scores to generate a total resistance score; Subjects with a total resistance score above a predefined threshold are predicted to be resistant to the therapy. Thereby predicting a subject's response to therapy. A method is provided that includes:

[0056] According to another aspect, there is provided a method, comprising: During the training phase, (i) the number of resistance-associated factors expressed in samples from subjects known to be responsive to the therapy and subjects known to be non-responsive to the therapy; (ii) at least one clinical parameter of the subject; and (iii) a marker related to the responsiveness of the subject Train a machine learning algorithm on a training set including Generating trained machine learning algorithms A method is provided, comprising:

[0057] According to another aspect, there is provided a method, comprising: During the training phase, (i) factor expression levels of resistance-associated factors in samples from subjects known to be responsive to a therapy and subjects known to be non-responsive to a therapy; (ii) at least one clinical parameter of the subject; and (iii) a marker related to the responsiveness of the subject Train a machine learning algorithm on a training set including Generating trained machine learning algorithms A method is provided, comprising:

[0058] In some embodiments, the method is a diagnostic method. In some embodiments, the method is an in vitro method. In some embodiments, the method is an ex vivo method. In some embodiments, the method is a computer-implemented method. In some embodiments, the method is a statistical method. In some embodiments, the method is a method that cannot be performed in the human mind. In some embodiments, the method is a computerized method. In some embodiments, the processor is a computer processor. In some embodiments, the processor is a computer.

[0059] In some embodiments, the method is for predicting a response to a therapy. In some embodiments, the method is for determining a response to a therapy. In some embodiments, the method is for determining a response score. In some embodiments, the method is for determining a response probability. In some embodiments, the response probability is a response score. According to some embodiments, a resistance score is determined. According to other embodiments, a prediction of the resistance probability is determined. According to some other embodiments, a resistance probability of less than 20% indicates that the subject will be responsive to the therapy. According to some embodiments, a response score is determined. According to other embodiments, a prediction of the response probability is determined. According to some other embodiments, a response probability of more than 80% indicates that the subject will be responsive to the therapy. In some embodiments, "beyond" is "above". In some embodiments, "beyond" is "below". It will be understood by one of skill in the art that the scale can be designed to be measured in either direction, and thus up / down will depend on the construction of the scale.

[0060] In some embodiments, the method is for determining whether a subject is a responder to a therapy. In some embodiments, the method is for determining whether a subject is a non-responder to a therapy. In some embodiments, the method is for predicting a subject's response to a therapy. In some embodiments, the method is for monitoring the response to a therapy. In some embodiments, the method is for determining whether a therapy should be continued or adjusted (e.g., by further treating the subject with an additional therapy, including but not limited to an agent determined by the RAP analysis provided below). In some embodiments, the method is for determining a subject as a responder to a therapy or a non-responder to a therapy. In some embodiments, the method is for determining a subject as a responder to a therapy, a non-responder to a therapy, or as having a stable disease state. In some embodiments, the method is for predicting whether a subject will respond to a therapy or will not respond to a therapy.

[0061] In some embodiments, the non-response comprises progressive disease. In some embodiments, the non-response comprises progression of the cancer. In some embodiments, the non-response comprises stable disease. In some embodiments, the non-response comprises worsening of symptoms of the disease. In some embodiments, the non-response is not the onset of side effects. In some embodiments, the non-response comprises growth, metastasis and / or continued proliferation of the cancer. In some embodiments, the response is stable. In some embodiments, the response comprises remission. In some embodiments, the remission is minimal remission. In some embodiments, the remission is partial remission. In some embodiments, the remission is complete remission. In some embodiments, the response is measured using overall response rate (ORR). Trained physicians are familiar with methods to determine response, specifically ORR. In some embodiments, the response is measured using Response Evaluation Criteria in Solid Tumors (RECIST). In some embodiments, the response comprises survival. In some embodiments, the survival is overall survival. In some embodiments, the survival is progression-free survival. In some embodiments, the response comprises sustained clinical benefit (DCB).

[0062] In some embodiments, the subject is a mammal. In some embodiments, the subject is a human. In some embodiments, the subject is suffering from a disease. In some embodiments, the disease is treatable by therapy. In some embodiments, the disease is cancer. In some embodiments, the disease is treatable by immune checkpoint inhibitors (ICIs). In some embodiments, the cancer is a PD-L1 positive cancer. In some embodiments, the cancer is a PD-L1 negative cancer. In some embodiments, the cancer is a solid cancer. In some embodiments, the cancer is a tumor. In some embodiments, the cancer is selected from hepatic biliary cancer, cervical cancer, genitourinary cancer (e.g., urothelial cancer), testicular cancer, prostate cancer, thyroid cancer, ovarian cancer, nervous system cancer, eye cancer, lung cancer, soft tissue cancer, bone cancer, pancreatic cancer, bladder cancer, skin cancer, intestinal cancer, liver cancer, rectal cancer, colorectal cancer, esophageal cancer, gastric cancer, gastroesophageal cancer, breast cancer (e.g., triple negative breast cancer), renal cancer (e.g., renal carcinoma), skin cancer, head and neck cancer, leukemia, and lymphoma. In some embodiments, the cancer is selected from skin cancer and lung cancer. In some embodiments, the cancer is skin cancer. In some embodiments, the cancer is lung cancer. In some embodiments, the skin cancer is melanoma. In some embodiments, the lung cancer is small cell lung cancer. In some embodiments, the lung cancer is non-small cell lung cancer. In some embodiments, the subject is naive to the therapy prior to the first determination. In some embodiments, the subject has not received the therapy prior to the first determination. In some embodiments, the subject has previously received the therapy. In some embodiments, the subject has previously been treated with a therapy other than the therapy. In some embodiments, the subject is naive to any therapy. In some embodiments, the subject is naive to immunotherapy. In some embodiments, the therapy is a first line of treatment. In some embodiments, the therapy is an advanced line of treatment.

[0063] In some embodiments, the therapy is an anti-cancer therapy. In some embodiments, the anti-cancer therapy is radiation. In some embodiments, the anti-cancer therapy is chemotherapy. In some embodiments, the therapy is immunotherapy. In some embodiments, the anti-cancer therapy is immunotherapy. In some embodiments, the anti-cancer therapy is targeted therapy. In some embodiments, the anti-cancer therapy is selected from radiation, chemotherapy, immunotherapy, targeted therapy, hormonal therapy, anti-angiogenic therapy and photodynamic therapy, hyperthermia, surgery, and combinations thereof. In some embodiments, the immunotherapy is selected from immune checkpoint inhibition, immune checkpoint modulation, immune checkpoint blockade, adoptive cell transfer therapy, oncolytic virus therapy, vaccine therapy, immune system modulation, and therapy using monoclonal antibodies. In some embodiments, the immunotherapy is selected from immune checkpoint inhibitors, immune checkpoint modulating agents, immune checkpoint blockade, adoptive cell transfer therapy, oncolytic virus therapy, therapeutic vaccines, immune system modulating agents, and monoclonal antibodies. In some embodiments, the immunotherapy is an immune checkpoint inhibitor. In some embodiments, the immunotherapy is immune checkpoint blockade. In some embodiments, immunotherapy is administered in combination with one or more conventional cancer therapies, including chemotherapy, targeted therapy, steroids, and radiation therapy.The combination of ICI and chemotherapy / radiotherapy / targeted therapy has been studied in multiple clinical trials.It will be understood by those skilled in the art that the predictive proteins disclosed herein are predictive of immunotherapy as a single agent therapy, as well as part of combination therapy.

[0064] In some embodiments, the immunotherapy is multiple immunotherapies. In some embodiments, the immunotherapy is immune checkpoint blockade. In some embodiments, the immunotherapy is immune checkpoint protein inhibition. In some embodiments, the immunotherapy is immune checkpoint protein modulation. In some embodiments, the immunotherapy comprises immune checkpoint inhibition. In some embodiments, the immunotherapy comprises immune checkpoint modulation. In some embodiments, the immune checkpoint blockade and / or immune checkpoint inhibition comprises administering an immune checkpoint inhibitor to the subject. In some embodiments, the inhibition comprises administering an immune checkpoint inhibitor. In some embodiments, the inhibitor is a blocking antibody. In some embodiments, the immunotherapy comprises immune checkpoint blockade. In some embodiments, the modulation comprises administering an immune checkpoint modulating agent. In some embodiments, the immune checkpoint modulation comprises administering an immune checkpoint modulating agent to the subject.

[0065] As used herein, the term "immune checkpoint inhibitor (ICI)" refers to a single ICI, a combination of ICIs, and a combination of an ICI with another cancer therapy. The ICI can be a monoclonal antibody, a bispecific antibody, a humanized antibody, a fully human antibody, a fusion protein, or a combination thereof, directed to block, inhibit or modulate immune checkpoint proteins. In some embodiments, the immune checkpoint inhibitor is an immune checkpoint modulator. In some embodiments, the immune checkpoint inhibitor is an immune checkpoint blocker. In some embodiments, the immune checkpoint protein is PD-1 (Programmed Death-1), PD-L1, PD-L2, CTLA-4 (Cytotoxic T Lymphocyte-Associated Protein 4), A2AR (Adenosine A2A Receptor), also known as ADORA2A, B7-H3 also known as CD276, B7-H4 also known as VTCN1, B7-H5, BTLA (B and T Lymphocyte Associated Protein 4), also known as CD272, B7-H5, also known as CD273, B7-H6, also known as CD274, B7-H7, also known as CD275, B7-H8, also known as CD276, B7-H9, also known as CD275, B7-H1, also known as CD276, B7-H2, also known as CD275, B7-H1 ... Attenuator), IDO (indoleamine 2,3-dioxygenase), KIR (killer cell immunoglobulin-like receptor), LAG-3 (lymphocyte activation gene 3), TDO (tryptophan 2,3-dioxygenase), TIM-3 (T cell immunoglobulin and mucin domains 3), VISTA (V domain Ig inhibitor of T cell activation), NOX2 (nicotinamide adenine dinucleotide phosphate NADPH oxidase isoform 2), SIGLEC7 (sialic acid-binding immunoglobulin-type lectin 7), also known as CD328, SIGLEC9 (sialic acid-binding immunoglobulin-type lectin 9), also known as CD329, OX40 (tumor necrosis factor receptor superfamily, member 4), also known as CD134, and TIGIT. In some embodiments, the immune checkpoint protein is selected from PD-1, PD-L1, and PD-L2. In some embodiments, the immune checkpoint protein is selected from PD-1 and PD-L1. In some embodiments, the immune checkpoint protein is CTLA-4. In some embodiments, the immune checkpoint protein is PD-1.In some embodiments, immune checkpoint blockade comprises anti-PD-1 / PD-L1 / PD-L2 immunotherapy. In some embodiments, immune checkpoint blockade comprises anti-PD-1 immunotherapy. In some embodiments, immune checkpoint blockade comprises anti-PD-1 and / or anti-PD-L1 immunotherapy. In some embodiments, immune checkpoint blockade comprises anti-CTLA-4 immunotherapy. In some embodiments, immune checkpoint blockade comprises anti-PD-1 and / or anti-PD-L1 immunotherapy and anti-CTLA-4 immunotherapy.

[0066] In some embodiments, the resistance-associated factor is a. i. In a population of subjects known to respond to therapy (responders), ii. in populations of subjects known to not respond to therapy (non-responders), and iii. In the subject receiving expression levels of a plurality of factors; b. calculating a resistance score for at least one factor of the plurality of factors; and c. classifying factors with resistance scores above a threshold as resistance-associated factors The method is determined by the method comprising:

[0067] In some embodiments, the resistance associated factor is present in each subject. In some embodiments, the resistance associated factor is in a responder. In some embodiments, the resistance associated factor is in a non-responder. In some embodiments, the resistance associated factor is labeled with a label. In some embodiments, the resistance associated factor is a resistance associated protein.

[0068] In some embodiments, the immunotherapy is a blocking antibody. In some embodiments, the immunotherapy is the administration of a blocking antibody to a subject.

[0069] In some embodiments, the ICI is a monoclonal antibody (mAb) against PD-1 or PD-L1. In some embodiments, the ICI is a mAb that neutralizes / blocks / inhibits / modulates the PD-1 pathway. In some embodiments, the ICI is a mAb against PD-1. In some embodiments, the anti-PD-1 mAb is pembrolizumab (Penbromo moniliforme; formerly known as Lambrolizumab). In some embodiments, the anti-PD-1 mAb is Nivolumab (Opdivo). In some embodiments, the anti-PD-1 mAb is Pidilizumab (CT0011). In some embodiments, the anti-PD-1 mAb is Cemiplimab (Libtayo, REGN2810). In some embodiments, the anti-PD-1 mAb is any one of AMP-224, MEDI0680, or PDR001. In some embodiments, the ICI is a mAb against PD-L1. In some embodiments, the anti-PD-L1 mAb is selected from atezolizumab (Tecentriq), avelumab (Bavencio), and durvalumab (Imfinzi). In some embodiments, the anti-PD-L1 mAb is atezolizumab. In some embodiments, the anti-PD-L1 mAb is durvalumab. In some embodiments, the ICI is a mAb against CTLA-4. In some embodiments, the anti-CTLA-4 mAb is ipilimumab.

[0070] As used herein, the term "factor" refers to any measurable biological molecule produced by a subject. In some embodiments, the factor is a protein. In some embodiments, the factor is RNA. In some embodiments, the factor is a gene. In some embodiments, the factor is a secreted factor. In some embodiments, the secreted factor is selected from a cytokine, a chemokine, a growth factor, a soluble receptor, and an enzyme. In some embodiments, the factor is a soluble factor. In some embodiments, the factor is a cellular factor. In some embodiments, the factor is a membrane factor. In some embodiments, the factor is a cell adhesion molecule. In some embodiments, the factor is a factor found in the blood. In some embodiments, the factor is a host-generated factor. In some embodiments, the factor is a resistance factor.

[0071] In some embodiments, the expression is protein expression. In some embodiments, the expression is secreted protein expression. In some embodiments, the protein expression is soluble protein expression. In some embodiments, the expression is cellular protein expression. In some embodiments, the expression is membrane protein expression. In some embodiments, the expression is mRNA expression. In some embodiments, the expression is protein expression or mRNA expression. In some embodiments, the expression level is a concentration. In some embodiments, the concentration is a concentration level. It will be understood by those skilled in the art that when the presence of a factor is measured in a liquid sample, the expression can be provided as a concentration, such as mg / ml, or in any unit according to the method of determining the expression of the factor. The any unit can be selected from relative fluorescence units (RFU) and normalized protein expression (NPX), or any other unit used as a measurement of expression. The terms "expression" and "expression level" are used interchangeably herein and refer to the amount of a gene product present in a sample. In some embodiments, the gene product comprises a polynucleotide, such as tumor DNA, circulating tumor DNA, or circulating DNA. In some embodiments, the DNA is cell-free DNA. In some embodiments, determining comprises quantifying the expression level. In some embodiments, determining comprises normalizing the expression level. The expression level of the factor can be determined by any method known in the art. Methods for determining protein expression include, for example, antibody array, immunoblotting, immunohistochemistry, flow cytometry (FACS), ELISA, proximity extension assay (PEA), aptamer-based assay, proteomics array, proteome sequencing, flow cytometry (CyTOF), multiplex assay, mass spectrometry and chromatography. In some embodiments, determining the protein expression level comprises ELISA. In some embodiments, determining the protein expression level comprises protein array hybridization. In some embodiments, determining the protein expression level comprises mass spectrometry quantification.In some embodiments, determining protein expression level comprises PEA. In some embodiments, determining protein expression level comprises aptamer. Methods for determining mRNA expression include, for example, RT-PCR, quantitative PCR, real-time PCR, microarray, Northern blotting, in situ hybridization, next-generation sequencing, and massively parallel sequencing.

[0072] In some embodiments, the received factor expression level is indicative of a factor expression level. In some embodiments, the received factor expression level is determining a factor expression level. In some embodiments, determining is measuring. In some embodiments, measuring is performed on a sample. In some embodiments, the expression level is detected in a sample. In some embodiments, the sample is a biological sample. In some embodiments, the sample is provided by a subject. In some embodiments, the sample is provided by a subject. In some embodiments, the sample is provided by a responder. In some embodiments, the sample is provided by a non-responder. In some embodiments, each subject of a population of responders provided a sample. In some embodiments, each subject of a population of non-responders provided a sample. In some embodiments, the sample is provided by a subject prior to receiving a therapy. In some embodiments, the sample is provided by a subject after receiving a therapy. In some embodiments, determining is performed directly on the sample. In some embodiments, determining is on an unprocessed sample. In some embodiments, determining is on a processed sample. In some embodiments, the method further comprises processing the sample. In some embodiments, processing comprises isolating protein from the sample. In some embodiments, the processing comprises isolating nucleic acid from the sample. In some embodiments, the nucleic acid is RNA. In some embodiments, the RNA is mRNA. In some embodiments, the processing comprises lysing cells of the sample.

[0073] As used herein, the terms "peptide", "polypeptide" and "protein" are used interchangeably to refer to a polymer of amino acid residues. In another embodiment, the terms "peptide", "polypeptide" and "protein" as used herein encompass natural peptides, peptidomimetics (typically containing non-peptide bonds or other synthetic modifications) as well as peptide analogs, peptoids and semi-peptoids or any combination thereof. In another embodiment, the described peptides, polypeptides and proteins have modifications that make them more stable in the body or more capable of penetrating cells. In one embodiment, the terms "peptide", "polypeptide" and "protein" apply to naturally occurring amino acid polymers. In another embodiment, the terms "peptide", "polypeptide" and "protein" apply to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding natural amino acids.

[0074] In some embodiments, the sample is a biological sample. In some embodiments, the sample is a tissue. In some embodiments, the sample is a fluid. In some embodiments, the fluid is a biological fluid. In some embodiments, the sample is from a subject. In some embodiments, the sample is not a tumor sample. In some embodiments, the sample is a tumor sample. In some embodiments, the sample is not a hematopoietic cancer and the sample is a blood sample. In some embodiments, the sample is a sample that does not contain cancer cells. In some embodiments, the blood sample includes a peripheral blood sample, a serum sample, and a plasma sample. In some embodiments, the sample is a plasma sample. In some embodiments, the sample is a serum sample. In some embodiments, the processing comprises isolating plasma. In some embodiments, the processing comprises isolating serum. In some embodiments, the biological fluid is selected from blood, plasma, serum, lymphatic fluid, cerebrospinal fluid, urine, feces, semen, tumor fluid, and gastric fluid. In some embodiments, the samples obtained from the subject and the responder are the same type of sample. In some embodiments, the samples obtained from the subject and the responder are different types of samples. In some embodiments, the samples obtained from the subject and the non-responder are the same type of sample. In some embodiments, the samples obtained from the subject and the non-responder are different types of samples. In some embodiments, the samples obtained from the non-responder and the responder are the same type of sample. In some embodiments, the samples obtained from the non-responder and the responder are different types of samples. In some embodiments, the samples obtained from the subject, the non-responder and the responder are the same type of sample. In some embodiments, the samples obtained from the subject, the non-responder and the responder are blood samples. In some embodiments, the samples obtained from the subject, the non-responder and the responder are plasma samples. In some embodiments, the samples obtained from the subject, the non-responder and the responder are serum samples.In some embodiments, the samples obtained from the subject, non-responder, and responder are different types of samples.

[0075] In some embodiments, the factor is a factor of a plurality of factors. In some embodiments, expression levels of a plurality of factors are received. In some embodiments, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 11 In one embodiment, expression levels of 00, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 12000, 15000, 20000, 25000, 30000, 35000, or 40000 factors are received. Each possibility represents a separate embodiment of the present invention. In some embodiments, a plurality is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 90 0, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 12000, 15000, 20000, 25000, 30000, 35000, or 40000. Each possibility represents a separate embodiment of the present invention. In some embodiments, expression levels of at least 200 factors are received. In some embodiments, expression levels of at least 400 factors are received. In some embodiments, expression levels of at least 1000 factors are received. In some embodiments, expression levels of at least 5000 factors are received. In some embodiments, expression levels of at least 6000 factors are received. In some embodiments, expression levels of at least 7000 factors are received. In some embodiments, expression levels of at least 8000 factors are received.

[0076] In some embodiments, the population of responders suffers from a disease. In some embodiments, the responders all have the same disease. In some embodiments, the population of non-responders suffers from a disease. In some embodiments, the non-responders all suffer from the same disease. In some embodiments, the population of responders and the population of non-responders all suffer from the same disease. In some embodiments, the population of responders and the subject suffers from the same disease. In some embodiments, the population of non-responders and the subject suffers from the same disease. In some embodiments, the population of non-responders, the population of responders and the subject suffer from the same disease.

[0077] In some embodiments, the expression level is from the subject before undergoing therapy. In some embodiments, the expression level for the subject is determined before undergoing therapy. In some embodiments, the expression level is from time T0. In some embodiments, the sample is provided by the subject before undergoing therapy. In some embodiments, the expression level is from the subject before undergoing the first treatment of therapy. In some embodiments, the treatment is a medication. In some embodiments, the treatment is a regimen.

[0078] In some embodiments, before is at least 1 hour, 2 hours, 3 hours, 6 hours, 8 hours, 12 hours, 1 day, 2 days, 3 days, 5 days, 1 week, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, or 6 months before the therapy or administration of the therapy. Each possibility represents a separate embodiment of the invention. In some embodiments, before is at least 1 hour. In some embodiments, before is just before the therapy or just before administration of the therapy. In some embodiments, before is up to 1 hour, 2 hours, 3 hours, 4 hours, 6 hours, 9 hours, 12 hours, 18 hours, 24 hours, 2 days, 3 days, 5 days, 1 week, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, or 6 months before the therapy or administration of the therapy. Each possibility represents a separate embodiment of the present invention. In some embodiments, prior is up to 24 hours prior to or before administering the therapy. In some embodiments, administering the therapy is the first administered therapy. In some embodiments, administering the therapy is the any administered therapy.

[0079] In some embodiments, the expression level is from a subject after undergoing a therapy. In some embodiments, the expression level is from time T1. In some embodiments, the sample is provided by a subject after undergoing a therapy. In some embodiments, the expression level is from a subject after undergoing a first treatment of a therapy. In some embodiments, the expression level is from a subject after undergoing any treatment with a therapy.

[0080] In some embodiments, later is a time point after the initiation of therapy or after the administration of therapy, sufficient for the change in expression of at least one factor. In some embodiments, later is a time point after the initiation of therapy or after the administration of the first treatment of therapy. In some embodiments, later is at least 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 4 weeks, 6 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, or 1 year later. Each possibility represents a separate embodiment of the invention. In some embodiments, later is at least 24 hours later. In some embodiments, later is at least 2 weeks later. In some embodiments, later is at least 3 weeks later. In some embodiments, later is at least 6 weeks later. In some embodiments, later is up to 1 week, 2 weeks, 3 weeks, 4 weeks, 6 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, or 1 year later after the initiation of therapy or after the administration of therapy. Each possibility represents a separate embodiment of the present invention.

[0081] In some embodiments, receiving the expression levels includes receiving factor expression levels for a group of factors larger than a plurality of factors. In some embodiments, the expression levels received for the larger group are received for responders and non-responders. In some embodiments, a subgroup of proteins is selected from the group. In some embodiments, the subgroup is designated a plurality of factors. In some embodiments, the method includes designating. In some embodiments, receiving further includes for each factor of the group applying a machine learning algorithm. In some embodiments, the algorithm classifies the factors from responders and non-responders. In some embodiments, the algorithm outputs if the subject who provided the sample with the measured factor expression level is a responder or a non-responder. In some embodiments, receiving further includes selecting a subgroup of factors where the algorithm most evenly divides the subjects into responders and non-responders. In some embodiments, the subjects are all subjects in the population of responders and non-responders. In some embodiments, the factors are processed with an algorithm that most evenly divides all subjects, responders and non-responders into groups of responders and non-responders, and are selected as subgroups (even if the assignment is incorrect). In some embodiments, the algorithm is trained on the received factor expression levels in responders and non-responders. In some embodiments, the algorithm is trained on a training set. In some embodiments, the training is on expression levels and tags indicating whether the expression levels were from responders or non-responders. In some embodiments, the training is on expression levels, clinical information, and tags indicating whether the expression levels were from responders or non-responders.

[0082] In some embodiments, the subgroups include factors that the algorithm most evenly divides the subjects. In some embodiments, there is an even division into responders and non-responders. In some embodiments, the subgroups are the top 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 750, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, or 5000. Each possibility represents a separate embodiment of the invention. In some embodiments, the subgroups are the top 50. In some embodiments, the subgroups are the top 100. In some embodiments, the subgroups are the top 200. In some embodiments, the subgroups are the top 500.

[0083] In some embodiments, the method further comprises performing a dimension reduction step. In some embodiments, the reduction is for a plurality of factors. In some embodiments, the reduction is to reduce the number of a plurality of factors. In some embodiments, the dimension reduction step identifies a subgroup or subset of factors. In some embodiments, the factors are major factors. In some embodiments, the training set includes only expression levels of a subset / subgroup of factors. In some embodiments, the subgroup or subset of factors is the factor that most evenly balances the number of predicted responders and non-responders. In some embodiments, the prediction is predicted by a machine learning algorithm. In some embodiments, the machine learning algorithm is a trained machine learning algorithm. In some embodiments, the machine learning algorithm is a machine learning algorithm in training.

[0084] In some embodiments, a pre-processing step can be performed to pre-process the received expression levels. In some embodiments, the pre-processing step can include at least one of data cleaning and normalization, feature selection, feature extraction, dimensionality reduction, and / or any other suitable pre-processing method or technique. Feature selection can be performed by a statistical test such as the Kolmogorov-Smirnov (KS) test, or any other test known in the art.

[0085] In some embodiments, a factor selection and / or dimension reduction step can be performed to reduce the number of factors in each sample and / or to obtain key factors, e.g., a set of factors that may have significant predictive power. Thus, in some embodiments, the factor selection and / or dimension reduction step can result in a reduction in the number of factors in each sample and / or set of values. In some embodiments, the dimension reduction selects key factors, e.g., proteins, based on the level of response predictive power that the factors generate for the desired prediction. In certain embodiments, the dimension reduction includes considering all or some of the factors as vector components and calculating their norms.

[0086] In some embodiments, any suitable factor selection and / or dimensionality reduction method or technique may be used, including but not limited to: ANOVA with S0 parameter: Analysis of variance with an additional parameter (S0) that controls the relative importance of features based on the obtained test p-values ​​and the mean group differences (see for example Tusher, Tibshirani and Chu, PNAS 98, pp5116-21, 2001). Scalable EMpirical Bayes Model Selection (SEMMS): An empirical Bayes feature selection method that applies parsimonious mixture models to identify significant predictors (see, e.g., Bar, Booth, and Wells. A scalable empirical Bayes approach to variable selection in generalized linear models, 2019). · L2N: A method for differential expression analysis using a three-component mixture model, which consists of two log-normal components (L2) for differentially expressed features, one component for under-expressed features, one component for over-expressed features, and one normal component (N) for non-differentially expressed features (see, e.g., Bar and Schifano. Differential variation and expression analysis. Stat 8, e237, doi:10.1002 / sta4.237, 2019). Genetic Algorithms: A family of heuristic optimization algorithms that employ organic evolutionary techniques such as random mutation, recombination, and natural selection as a method to achieve an optimal configuration (see, for example, Popovic, Sifrim, Pavlopoulos, Moreau, and Bart De Moor. A Simple Genetic Algorithm for Biomarker Mining. 2012). Naive classifiers: Naive classifiers evaluate the response score by reducing the dimensionality to a single score. This is done by considering all features (e.g. specific profiles such as protein expression levels) as components of a vector and calculating its norm. The reduction in dimensionality reduces the possible risk of overfitting. In some embodiments, the vector components are normalized according to the typical component values ​​among patients belonging to the same response group (e.g. responders) such that the normalized norm quantifies the amount of deviation from the typical respective class value. In further embodiments, the naive classifier allows training using data of subjects belonging to only a portion of the response group.

[0087] As used herein, the terms "responder" or "known to respond" subject are used interchangeably and refer to a subject who, when administered a therapy, shows an improvement in at least one criterion of the disease being treated by the therapy or shows no increase in disease severity. In some embodiments, a responder is a subject who, when administered a therapy, shows an improvement in the disease being treated by the therapy. In some embodiments, a responder is a subject who, when administered a therapy, shows no increase in disease severity. In some embodiments, the increase is severity over time. In some embodiments, the increase in severity does not show a stable increase. In some embodiments, a responder is a subject who shows a mixed response when administered a therapy. In some embodiments, a responder is a subject who, when administered a therapy, shows a mixed response, where the mixed response is an improvement in at least one criterion of the disease but not an improvement in other criteria of the disease. In some embodiments, the mixed response is a shrinkage of some lesions combined with the growth of new or existing lesions. In some embodiments, a responder is a subject in whom a therapy results in an anti-disease response. In some embodiments, for a subject with cancer, a responder is a subject in which the therapy produces an anti-cancer response. In some embodiments, the response is not a reduction in side effects. In some embodiments, the response is a reduction in side effects. In some embodiments, the response is a response to the disease itself. In some embodiments, the anti-cancer response is an anti-tumor response. In some embodiments, the anti-tumor response comprises tumor regression. In some embodiments, the anti-tumor response comprises tumor shrinkage. In some embodiments, the anti-tumor response comprises a lack of tumor growth. In some embodiments, the anti-tumor response comprises a lack of tumor metastasis. In some embodiments, the anti-tumor response comprises a lack of tumor hyperproliferation. In some embodiments, the improvement is in at least one symptom of the disease. In some embodiments, the response is a complete response. In some embodiments, the response is a minimal response. In some embodiments, the response is a partial response.In some embodiments, the response includes stable disease. In some embodiments, a responder is a subject who has a favorable response to a therapy. In some embodiments, a non-responder is a subject who has an unfavorable response to a therapy. In some embodiments, the unfavorable response is an increase in tumor burden. An increase in tumor burden can include either an increase in tumor size or total number of cancer cells, such as an increase in tumor size, an increase in tumor spread, an increase in metastasis, an increase in tumor cell proliferation, or any other increase.

[0088] As used herein, a "favorable response" of a cancer patient refers to the "responsiveness" of the cancer patient to treatment with a therapy, i.e., treatment of a responsive cancer patient with a therapy results in a desired clinical outcome, such as tumor regression, tumor shrinkage or tumor necrosis, reduction in tumor burden, anti-tumor response by the immune system, preventing or delaying tumor recurrence, tumor growth or tumor metastasis. In some embodiments, the subject is a complete responder, or treatment with a cancer therapy results in plateau. In some embodiments, a complete responder is a subject who has no detectable cancer after treatment with a therapy. In this case, treatment of a responsive cancer patient with a therapy can be continued and is recommended, or treatment can be discontinued if the patient is no longer afflicted with cancer. In some embodiments, the method further comprises continuing to administer the therapy to a subject who is not a non-responder. In some embodiments, the subject is a non-responder, a minimal responder, a partial responder, or has plateaued, and the method further comprises continuing to administer the therapy to the subject, as well as treating the subject with an additional therapy (e.g., as determined using the resistance associated protein (RAP) analysis provided herein) to enhance responsiveness. In some embodiments, a subject that is not a non-responder is a responder.

[0089] As used herein, the terms "non-responder" and "known non-responsive" subjects are used interchangeably and refer to subjects who do not show improvement or stabilization of disease when administered a therapy. In some embodiments, a non-responder shows worsening of disease when administered a therapy. In some embodiments, a non-responder is not a subject who experiences side effects of a therapy. In some embodiments, a non-responder is a subject whose disease progresses. In some embodiments, a non-responder is a subject whose disease does not stabilize after a therapy. In some embodiments, a non-responder is a subject whose disease does not improve after a therapy. In some embodiments, a non-responder is a subject who is not a responder as defined above herein. In some embodiments, a non-responder is a subject who has an unfavorable response to a therapy. In some embodiments, a non-responder is a subject who is resistant to a therapy. In some embodiments, a non-responder is a subject who is refractory to a therapy.

[0090] As used herein, an "unfavorable response" of a cancer patient refers to a "non-responsiveness" of a cancer patient to treatment with a therapy, such that treatment of a non-responsive cancer patient with a therapy does not result in a desired clinical outcome, potentially resulting in undesirable outcomes such as tumor expansion, recurrence, or metastasis. In some embodiments, the method further comprises discontinuing administration of the therapy to a subject who is a non-responder. In some embodiments, the method further comprises continuing to administer the therapy to the subject in combination with an additional therapy. In some embodiments, the additional therapy enhances the responsiveness of the non-responsive patient.

[0091] In some embodiments, the methods are for determining whether a response is considered a durable response (eg, progression-free survival greater than 6 months).

[0092] In some embodiments, the method further comprises administering the therapy to the subject predicted to respond to the therapy. In some embodiments, the method further comprises continuing to administer the therapy to the subject predicted to respond to the therapy. In some embodiments, the method further comprises not administering the therapy to the subject predicted not to respond to the therapy. In some embodiments, the method further comprises discontinuing the therapy to the subject predicted not to respond to the therapy. In some embodiments, the method further comprises administering an alternative therapy to the subject predicted to be a non-responder. In some embodiments, the alternative therapy is an additional therapy. In some embodiments, the method further comprises administering or continuing to administer the therapy in combination with an agent or therapy that blocks or inhibits at least one of the resistance-associated factors in the subject predicted to be resistant to the therapy. In some embodiments, the agent or therapy that blocks or inhibits at least one of the resistance-associated factors is an additional therapy. In some embodiments, a combination therapy is administered to the subject predicted to be a non-responder.

[0093] In some embodiments, the method further comprises administering to the subject (e.g., the non-responder) an agent that modulates at least one factor. In some embodiments, modulating includes inhibiting, blocking, and modulating. In some embodiments, modulating is inhibiting. In some embodiments, the method further comprises administering to the subject (e.g., the non-responder) an agent that modulates a pathway that includes at least one factor. In some embodiments, modulating at least one factor is modulating a pathway that includes at least one factor. In some embodiments, modulating a pathway is modulating a driver protein / gene that controls the at least one factor. In some embodiments, modulating a pathway is modulating a driver protein / gene that controls the pathway. In some embodiments, modulating a pathway that includes at least one factor is modulating a receptor (e.g., using one or more receptor agonists), a ligand or factor, a paralog of the factor, or a combination thereof. In some embodiments, modulating is modulating multiple factors. In some embodiments, modulating is modulating multiple factors of a signature. In some embodiments, modulating is modulating each factor of a signature. In some embodiments, modulating achieves a better response to therapy. In some embodiments, the factor is a resistance-associated factor.

[0094] In some embodiments, the resistance score is a RAP score. In some embodiments, the resistance score is a response score. In some embodiments, the resistance score is a 1 response score. In some embodiments, the response score is a 1-resistance score. In some embodiments, the resistance score is a total resistance score. In some embodiments, the response score is a total response score. In some embodiments, the RAP score is a total RAP score. In some embodiments, the resistance score is based on the similarity of factor expression levels in the subject to factor expression levels in non-responders. In some embodiments, the resistance score is based on the similarity of factor expression levels in the subject to factor expression levels in responders. In some embodiments, based on something is calculated based on something. In some embodiments, the similarity is a state of lack of similarity. In some embodiments, the similarity with a responder is a state of lack of similarity with a non-responder. In some embodiments, the similarity with a non-responder is a state of lack of similarity with a responder. In some embodiments, the similarity is measured on a scale. In some embodiments, the scale is 0 to 1, with 1 being completely similar to a non-responder and 0 being completely similar to a responder. In some embodiments, the resistance score is 0 to 1, with 1 being completely similar to a non-responder and 0 being completely similar to a responder. In some embodiments, the resistance score is based on the similarity of factor expression levels in the subject to the factor expression levels in the non-responders and the factor expression levels in the responders.

[0095] In some embodiments, the method includes selecting a subset of factors prior to step (b). In some embodiments, prior to step (b) is prior to calculating. In some embodiments, the subset is a subject of a plurality of factors. In some embodiments, the subject includes factors that best distinguish between responders and non-responders. In some embodiments, the factors that provide the best discrimination are the top percentage. In some embodiments, the top percentage is the top 1, 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50% of factors. Each possibility represents a separate embodiment of the present invention. In some embodiments, the top percentage is the top 20%. In some embodiments, the top factors are the top 10, 20, 25, 30, 40, 50, 60, 70, 75, 80, 90, or 100 factors. Each possibility represents a separate embodiment of the present invention. In some embodiments, the top factors are the top 50 factors. In some embodiments, the selecting includes applying a Kolmogorov-Smirnov test. In some embodiments, a Kolmogorov-Smirnov test is applied to the received factor expression levels. In some embodiments, the Kolmogorov-Smirnov test determines how well a factor discriminates between responders and non-responders. In some embodiments, the Kolmogorov-Smirnov test outputs a measure of how well a factor discriminates, and the best factor is the factor with the highest score. In some embodiments, the selecting includes applying an XGBoost algorithm. In some embodiments, the calculating is for the subset. In some embodiments, the calculating is for each factor of the subset.

[0096] In some embodiments, the calculating comprises applying a machine learning algorithm. In some embodiments, the calculating comprises applying a machine learning model. In some embodiments, the machine learning model is a machine learning algorithm. In some embodiments, the machine learning model implements a machine learning algorithm. In some embodiments, the algorithm is a classifier. In some embodiments, the algorithm is a regression model. In some embodiments, the algorithm is supervised. In some embodiments, the algorithm is unsupervised. In some embodiments, the machine learning algorithm is trained on expression levels in responders. In some embodiments, the machine learning algorithm is trained on expression levels in non-responders. In some embodiments, the machine learning algorithm is trained on expression levels in responders and non-responders. In some embodiments, the machine learning algorithm is trained on a training set. In some embodiments, the machine learning algorithm is trained by a method of the invention. In some embodiments, the machine learning algorithm is applied to a factor of the plurality of factors. In some embodiments, the machine learning algorithm is applied to each factor of the plurality of factors. In some embodiments, the machine learning algorithm is applied to a subset. In some embodiments, the machine learning algorithm is applied to a subset of factors. In some embodiments, the machine learning algorithm is applied to each factor of the subset of factors. In some embodiments, each factor is analyzed and calculated separately, and the machine learning algorithm does not use the expression level of more than one factor as training set.In some embodiments, the trained machine learning algorithm is applied to the individual protein expression level from the subject.In some embodiments, the machine learning algorithm trained on the expression level of a particular factor in responders and non-responders is applied to the expression level of that particular factor in the subject.Those skilled in the art will understand that for each factor of the multiple factors, a different algorithm is trained and then applied to each expression level of the subject. Thus, if three algorithms are trained separately for expression in responders and non-responders for factor A, factor B and factor C, the algorithm trained for factor A expression level is applied to the expression level of the subject of factor A, the algorithm trained for factor B expression level is applied to the expression level of the subject of factor B, and the algorithm trained for factor C expression level is applied to the expression level of the subject of factor C. In some embodiments, during the training phase, a machine learning model is trained on a training set that includes the expression data of a single factor from responders and non-responders, using the corresponding annotation of "responder" or "non-responder" to predict or classify the factor expression data according to the class "responder" and "non-responder". In some embodiments, during the inference phase, the machine learning model is applied to the expression data of a single factor from the subject to predict the classification of the factor as responder or non-responder. In some embodiments, the classification is a resistance score. In some embodiments, the classification is a response score. In some embodiments, the classification is a measure of how similar the agent is to a non-responder and how dissimilar it is to a responder.

[0097] In some embodiments, the trained machine learning algorithm is trained to predict the responsiveness of a subject suffering from a disease to a therapy. In some embodiments, the trained machine learning algorithm is trained to output a resistance score. In some embodiments, the trained machine learning algorithm is trained to output a resistance probability. In some embodiments, the trained machine learning algorithm is trained to output an activity score. In some embodiments, the trained machine learning algorithm is trained to predict the activity of a resistance associated factor in a subject. In some embodiments, the trained machine learning algorithm is trained to predict whether a factor is a resistance associated factor in a subject. In some embodiments, the trained machine learning algorithm is trained to predict whether a factor in a subject is a resistance associated factor in a subject.

[0098] In some embodiments, the training set includes the received factor expression levels. In some embodiments, the training set includes the received factor expression levels in both responders and non-responders. In some embodiments, the training set includes the received factor expression levels for only one factor. In some embodiments, the training set includes the number of resistance-associated factors expressed in the samples. In some embodiments, the samples are from subjects suffering from a disease. In some embodiments, the samples are from responders. In some embodiments, the samples are from non-responders. In some embodiments, the training set includes at least one clinical parameter. In some embodiments, the clinical parameter is from a subject. In some embodiments, the subjects are responders and non-responders. In some embodiments, the training set includes a label. In some embodiments, the label is associated with the responsiveness of the subject. In some embodiments, the label is a responder or a non-responder. In some embodiments, the resistance-associated factor is labeled with a label. In some embodiments, the at least one clinical parameter is labeled with a label.

[0099] According to some embodiments, the training set further comprises at least one clinical parameter of each responder and non-responder, and the machine learning algorithm is applied to the individual factor expression levels received from the subjects and the at least one clinical parameter of the subjects. In some embodiments, the at least one clinical parameter is the sex of the subject. In some embodiments, the training set further comprises the sex of the subject. In some embodiments, the subject is each subject. In some embodiments, the sex is gender. In some embodiments, the at least one clinical parameter is sex. In some embodiments, the sex is the sex of the subject. In some embodiments, the sex is male or female. In some embodiments, the sex is sex at birth. In some embodiments, the clinical parameter is age. In some embodiments, the age is the age of the subject. In some embodiments, the clinical parameter is a line of treatment. In some embodiments, the parameter of the line of treatment is whether the therapy was a first line of treatment or an advanced treatment. In some embodiments, the line of treatment is a first line of treatment. In some embodiments, the line of treatment is a second line of treatment. In some embodiments, the second line of treatment is an advanced treatment. It will be understood by those skilled in the art that the advanced therapy can be any line of therapy after the first therapy, e.g., second therapy, third therapy, fourth therapy, fifth therapy, etc. In some embodiments, the clinical parameter is whether the therapy is a first therapy or an advanced therapy. In some embodiments, the clinical parameter is PD-L1 status. In some embodiments, the PD-L1 status is the PD-L1 status of the cancer. Methods for measuring PD-L1 levels of cancer cells (e.g., tumors) are well known in the art, and any such method can be used. In some embodiments, the PD-L1 status includes high PD-L1 or low PD-L1. In some embodiments, the PD-L1 status includes high PD-L1, low PD-L1, or no PD-L1. In some embodiments, the PD-L1 status includes high PD-L1, medium PD-L1, or low PD-L1.In some embodiments, PD-L1 status includes expression of PD-L1 in less than 1% of cancer cells, 1-49% of cancer cells, or 50% or more of cancer cells. In some embodiments, expression of PD-L1 in less than 1% of cancer cells is no PD-L1 expression. In some embodiments, expression of PD-L1 in less than 1% of cancer cells is low PD-L1 expression. In some embodiments, expression of PD-L1 in 1-49% of cancer cells is low PD-L1 expression. In some embodiments, expression of PD-L1 in 1-49% of cancer cells is moderate PD-L1 expression. In some embodiments, expression of PD-L1 in 50% or more of cancer cells is high PD-L1 expression.

[0100] In some embodiments, the clinical parameter is a known biomarker of the disease or a mutation in a known biomarker of the disease. In some embodiments, the biomarker is selected from MYC, NOTCH, EGFR, HER2, BRAF, KRAS, MAP2K1, MET, NRAS, NTRK1, NTRK2, NTRK3, PIK3CA, RET, ROS1, TP53, ALK, CDKN2A, KIT, NF1, BFAST, FGFR, LDH, PTEN, RB1, PD-L1, MSI (microsatellite instability), TMB (tumor mutation burden), or a combination thereof. In some embodiments, the clinical parameter is the expression of a biomarker. In some embodiments, the expression is a percentage of expression. In some embodiments, the expression is a mutation status.

[0101] In some embodiments, the training set further comprises the gender, age and PD-L1 status of each responder and non-responder. In some embodiments, the training set further comprises the gender of each responder and non-responder. In some embodiments, the training set further comprises the age and PD-L1 status of each responder and non-responder. In some embodiments, the machine learning algorithm is applied to the individual received factor expression levels from the subject and the gender of the subject. In some embodiments, the machine learning algorithm is applied to the individual received factor expression levels from the subject and the gender, age and PD-L1 status of the subject. In some embodiments, the calculating comprises applying a machine learning algorithm trained on a training set comprising the received factor expression levels in the responder and non-responder and at least one clinical parameter to the expression levels from the subject and at least one clinical parameter of the subject, and the machine learning algorithm outputs a resistance score. In some embodiments, the training includes the received factor expression levels in the responders and non-responders and clinical parameters of each of the responders and non-responders, and a machine learning algorithm is applied to the individual received factor expression levels from the subjects and the clinical parameters of the subjects, and the machine learning algorithm outputs a response prediction. In some embodiments, the training includes the received factor expression levels in the responders and non-responders, and clinical parameters selected from gender, age, and PD-L1 expression of each of the responders and non-responders, or any combination thereof, and a machine learning algorithm is applied to the individual received factor expression levels from the subjects and the clinical parameters of the subjects, and the machine learning algorithm outputs a response prediction. In some embodiments, the training set includes the number of resistance-associated factors in each of the responders and non-responders and at least one clinical parameter, and a machine learning algorithm is applied to the number of resistance-associated factors from the subjects and at least one clinical parameter of the subjects, and the machine learning algorithm outputs a response prediction.In some embodiments, the training set includes the number of resistance-associated factors in each responder and non-responder and the gender of each responder and non-responder, and a machine learning algorithm is applied to the number of resistance-associated factors from the subjects and the gender of the subjects, and the machine learning algorithm outputs a response prediction. In some embodiments, the training set includes the number of resistance-associated factors in each responder and non-responder, the age and PD-L1 status of each responder and non-responder, and a machine learning algorithm is applied to the number of resistance-associated factors from the subjects and the age and PD-L1 status of the subjects, and the machine learning algorithm outputs a response prediction.

[0102] In some embodiments, the clinical parameter is a type of therapy. In some embodiments, the clinical parameter is expression of a target of the therapy. In some embodiments, the clinical parameter is expression of a protein in a process that is a target of the therapy. In some embodiments, the process is a process that includes a target of the therapy. In some embodiments, the expression is expression in a subject. In some embodiments, the expression is expression in diseased tissue. In some embodiments, the expression is expression in a diseased tissue sample. In some embodiments, the expression is expression in a tumor. In some embodiments, the expression is expression in a tumor sample. In some embodiments, the tumor sample is a biopsy. In some embodiments, the expression is not expression in a tumor. In some embodiments, the expression is not expression in a tumor sample. In some embodiments, the expression is expression in a liquid biopsy. In some embodiments, the expression is a percentage of expression. In some embodiments, the percentage is a percentage of cells. In some embodiments, the therapy is an anti-PD-1 therapy and the protein in the process is PD-L1. In some embodiments, the therapy is an anti-PD-L1 therapy and the target protein is PD-L1. In some embodiments, the clinical parameter is expression of PD-L1. In some embodiments, the training set includes at least one clinical parameter selected from line of treatment, expression of PD-L1, gender, and age. In some embodiments, the training set includes protein expression level and gender. In some embodiments, the training set includes number of RAPs, age, and PD-L1 status.

[0103] In addition, clinical parameters may also be included.Those skilled in the art will be able to select relevant clinical parameters to include in training set.Examples of additional clinical parameters include, but are not limited to, the tissue type of sample (e.g., adenocarcinoma, squamous cell carcinoma, etc.), metastatic site, tumor site, cancer stage (e.g., tumor, lymph node and metastasis, TNM, stage, etc.), performance status (e.g., ECOG performance status), gene mutation, epigenetic status, general medical history, vital signs, blood measurements, renal and hepatic function, weight, height, pulse, blood pressure and smoking history.

[0104] In some embodiments, in the inference stage, a trained machine learning algorithm is applied. In some embodiments, the trained machine learning algorithm is applied to the individual received factor expression levels. In some embodiments, the trained machine learning algorithm is applied to the individual received factor expression levels and at least one clinical parameter. In some embodiments, the trained machine learning algorithm is applied to the individual received factor expression levels from the subject and the gender of the subject. In some embodiments, the trained machine learning algorithm is applied to several resistance-associated proteins. In some embodiments, the trained machine learning algorithm is applied to the number of resistance-associated factors. In some embodiments, the trained machine learning algorithm is applied to several resistance-associated factors and at least one clinical parameter.

[0105] In some embodiments, in the inference step, an input is received. In some embodiments, the input comprises a number of resistance-associated factors expressed in the sample. In some embodiments, the sample is from a subject. In some embodiments, the input comprises at least one clinical parameter. In some embodiments, the subject is afflicted with a disease. In some embodiments, the subject has an unknown responsiveness to a therapy. In some embodiments, the parameters are of a subject with an unknown responsiveness. In some embodiments, in the inference step, a trained machine learning algorithm is applied. In some embodiments, the application is applied to an input. In some embodiments, the input is a received input. In some embodiments, the inference step is predicting responsiveness. In some embodiments, the responsiveness is responsiveness to a therapy of a subject with an unknown responsiveness.

[0106] In some embodiments, the machine learning algorithm outputs a resistance score. In some embodiments, the output resistance score is scaled from 0 to 1. In some embodiments, 1 is completely similar to a non-responder and 0 is completely similar to a responder. In some embodiments, the machine learning algorithm calculates a similarity to a responder. In some embodiments, the machine learning algorithm calculates a similarity to a non-responder. In some embodiments, the machine learning algorithm outputs a similarity number to a responder and a non-responder. In some embodiments, a protein is considered to be a RAP if its resistance score exceeds a certain threshold. In some embodiments, the resistance score threshold is calculated on a scale of 0 to 1. In some embodiments, the threshold for a resistance score of a particular protein is between 0.2 and 0.95. In some embodiments, the resistance score threshold for a particular protein is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the resistance score threshold is 0.25. In some embodiments, the resistance score threshold is 0.42. In some embodiments, the resistance score threshold is 0.6. In some embodiments, the resistance score threshold when calculated by a machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention. In some embodiments, the resistance score threshold is 0.25 when calculated using a machine learning algorithm. In some embodiments, the resistance score threshold is 0.42 when calculated using a machine learning algorithm. In some embodiments, the resistance score threshold is 0.6 when calculated using a machine learning algorithm.

[0107] In some embodiments, the response probability is determined by the calculation (1-resistance score). In some embodiments, 1-resistance score is 1-final resistance score. In some embodiments, the resistance score is the final resistance score. In some embodiments, the response probability is the response score. In some embodiments, the machine learning algorithm outputs a response score. In some embodiments, the output response score is scaled from 0 to 1. In some embodiments, 1 is completely similar to a responder and 0 is completely similar to a non-responder. In some embodiments, the machine learning algorithm calculates a similarity to a responder. In some embodiments, the machine learning algorithm calculates a similarity to a non-responder. In some embodiments, the machine learning algorithm outputs a similarity value to a responder and a non-responder. In some embodiments, a protein is considered to be a RAP if its response score exceeds a certain threshold. In some embodiments, the response score threshold is calculated on a scale of 0 to 1. In some embodiments, the threshold for a response score of a particular protein is between 0.2 and 0.95. In some embodiments, the response score threshold for a particular protein is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the response score threshold is 0.25. In some embodiments, the response score threshold is 0.42. In some embodiments, the response score threshold is 0.6. In some embodiments, the response score threshold when calculated by a machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention.In some embodiments, the response score threshold when calculated using a machine learning algorithm is 0.25. In some embodiments, the response score threshold when calculated using a machine learning algorithm is 0.42. In some embodiments, the response score threshold when calculated using a machine learning algorithm is 0.6. In some embodiments, the algorithm outputs a response probability, the response probability being calculated on a scale of 0 to 1. In some embodiments, the algorithm outputs a response probability, the response probability being calculated on a scale of 0% to 100%, where 100% are responders and 0% are non-responders.

[0108] In some embodiments, the score is between 0 and 1. In some embodiments, the active agent is active in cancer. In some embodiments, the active agent is active in the subject. In some embodiments, the active agent is active in promoting resistance. In some embodiments, exceeding the threshold is below the threshold. In some embodiments, exceeding the threshold is above the threshold. In some embodiments, the predetermined threshold is 0.5, 0.4, 0.3, 0.25, 0.2, 0.15, 0.1, 0.05, 0.01, 0.005, 0.001, 0.0005, or 0.0001. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold is 0.05. In some embodiments, the threshold is 5%.

[0109] In some embodiments, the machine learning algorithm outputs a resistance score. In some embodiments, the resistance score is a RAP score. In some embodiments, the output resistance score is scaled from 0 to 1. In some embodiments, 1 is completely similar to a non-responder and 0 is completely similar to a responder. In some embodiments, the machine learning algorithm calculates a similarity to a responder. In some embodiments, the machine learning algorithm calculates a similarity to a non-responder. In some embodiments, the machine learning algorithm outputs a numerical value of similarity to a responder and a non-responder. In some embodiments, a protein is considered to be a RAP if its resistance score exceeds a certain threshold. In some embodiments, the resistance score threshold is calculated on a scale of 0 to 1. In some embodiments, the threshold for a resistance score of a particular protein is between 0.2 and 0.95. In some embodiments, the resistance score threshold for a particular protein is about 0.01, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention. In some embodiments, the resistance score threshold is 0.25. In some embodiments, the resistance score threshold is 0.42. In some embodiments, the resistance score threshold is 0.6. In some embodiments, the resistance score threshold calculated by machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention. In some embodiments, the resistance score threshold calculated by machine learning algorithm is 0.25. In some embodiments, the resistance score threshold calculated by machine learning algorithm is 0.42.In some embodiments, the resistance score threshold is 0.6 when calculated using a machine learning algorithm.

[0110] In some embodiments, the response probability is determined by the calculation (1-resistance score). In some embodiments, 1-resistance score is 1-final resistance score. In some embodiments, the resistance score is the final resistance score. In some embodiments, the response probability is the response score. In some embodiments, the machine learning algorithm outputs a response score. In some embodiments, the output response score is scaled from 0 to 1. In some embodiments, 1 is completely similar to a responder and 0 is completely similar to a non-responder. In some embodiments, the machine learning algorithm calculates similarity to responders. In some embodiments, the machine learning algorithm calculates similarity to non-responders. In some embodiments, the machine learning algorithm outputs similarity values ​​to responders and non-responders. In some embodiments, a protein is considered to be a RAP if its response score is above a certain threshold. In some embodiments, "beyond" is "above". In some embodiments, "beyond" is "below". In some embodiments, the response score threshold is calculated on a scale of 0 to 1. In some embodiments, the threshold for a response score of a particular protein is between 0.2 and 0.95. In some embodiments, the threshold for a response score of a particular protein is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for a response score is 0.25. In some embodiments, the threshold for a response score is 0.42. In some embodiments, the threshold for a response score is 0.6.In some embodiments, the response score threshold calculated by the machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the present invention. In some embodiments, the response score threshold calculated by the machine learning algorithm is 0.25. In some embodiments, the response score threshold calculated by the machine learning algorithm is 0.42. In some embodiments, the response score threshold calculated by the machine learning algorithm is 0.6.

[0111] In some embodiments, the machine learning model is a machine learning algorithm. In some embodiments, the algorithm is a supervised learning algorithm. In some embodiments, the algorithm is an unsupervised learning algorithm. In some embodiments, the algorithm is a reinforcement learning algorithm. In some embodiments, the machine learning model is a convolutional neural network (CNN). In some embodiments, the at least one hardware processor trains the machine learning model. In some embodiments, the model is based, at least in part, on a training set. In some embodiments, the model is based on a training set. In some embodiments, the model is trained on a training set. In some embodiments, the at least one hardware processor applies the machine learning model to factor expression levels from the subject.

[0112] In some embodiments, the calculating comprises calculating the average expression for each protein in the responders. In some embodiments, the calculating comprises calculating the average expression for each protein in the non-responders. In some embodiments, the calculating comprises calculating the average expression for each protein in the responders and the average expression for each protein in the non-responders. In some embodiments, the calculating comprises calculating the distribution of expression for each protein in the responders and the non-responders. In some embodiments, the calculating comprises calculating the standard deviation of expression for each protein in the responders and the non-responders. In some embodiments, the responders are in the responder population. In some embodiments, the non-responders are in the non-responder population. In some embodiments, the resistance score is based on the ratio of the deviation of the factor expression in the subject from the calculated average in the responders and the deviation of the factor expression in the subject from the calculated average in the non-responders. The calculation of the deviation is well known to those skilled in the art. It will be understood that the more dissimilar the expression in the subjects is from the average, the greater the deviation. Thus, factors that are less similar to the responder average will have a larger numerator in this ratio calculation, and factors that are less similar to the non-responder average will have a smaller denominator. Thus, the more similar the expression of a factor in a subject is to responder expression and the more similar it is to non-responder expression, the higher the resistance score. In some embodiments, a resistance score above a predefined threshold indicates that the factor is a resistance-associated factor. In some embodiments, the resistance-associated factor is a resistance-associated protein (RAP).

[0113] In some embodiments, the calculating further comprises calculating a distribution for each factor in the responders. In some embodiments, the calculating further comprises calculating a distribution for each factor in the non-responders. In some embodiments, the calculating further comprises calculating a distribution for each factor in the responders and a distribution for each factor in the non-responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in the responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in the non-responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in the responders and a standard deviation for each protein in the non-responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in the mix of responders and non-responders. In some embodiments, the deviation is measured as a multiple of the calculated standard deviation. It will be appreciated by those skilled in the art that by scaling the deviation to the standard deviation of the group of expression values, the deviation can be given in more absolute terms, allowing for the comparison of factors and populations with very small and very large stand deviations (which may also have very low and very high expression levels).

[0114] In some embodiments, the resistance score is based on the Z-score for each expression level of each factor in the subject. In some embodiments, the resistance score is based on the Z-score for responders. In some embodiments, the resistance score is based on the Z-score for non-responders. In some embodiments, the resistance score is based on both the Z-score for responders and the Z-score for non-responders. In some embodiments, the resistance score is based on the ratio of the Z-score for responders to the Z-score for non-responders. It is well known to those skilled in the art that the Z-score measures the distance of an individual level from the mean of a population in units of population standard deviation. In some embodiments, the Z-score is calculated by Formula 1.

[0115] In some embodiments, the resistance score is calculated by the formula

number

[0116] In some embodiments, a resistance score above a pre-determined threshold indicates that the agent is RAP. In some embodiments, "beyond" is "above." In some embodiments, the threshold is a pre-determined threshold. In some embodiments, the threshold is a threshold. In some embodiments, the resistance score threshold is about 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, 5.0, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 5.9, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 5.1, 5.2, 5.4, 5.5, 5.6, 5.8, 5.9, 5.1, 5.2, 5.5, 5.6, 5.7, 5.8, 5.9, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.8, 5.9, 5.1, 5.2, 5.4, 5.5, 5.6, 7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 5.0, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7, 5.8, 5.9, 6.0.6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, or 7.0. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold is about 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.67, 0.7, 0.75, 0.8, 0.85, or 0.9. Each possibility represents a separate embodiment of the present invention. In some embodiments, the resistance score threshold is about 2.9. In some embodiments, the resistance score threshold is 2.9. In some embodiments, the resistance score threshold is about 3.0. In some embodiments, the resistance score threshold is 3.0. In some embodiments, the resistance score threshold is calculated on an arbitrary unit scale. In some embodiments, the threshold resistance score calculated by mathematical calculation is about 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, or 5.0. Each possibility represents a separate embodiment of the present invention. In some embodiments, the threshold resistance score calculated by mathematical calculation is about 2.9.In some embodiments, the resistance score threshold when calculated using mathematical calculations is 2.9. In some embodiments, the resistance score threshold when calculated using mathematical calculations is about 3.0. In some embodiments, the resistance score threshold when calculated using mathematical calculations is 3.0. In some embodiments, the mathematical calculations are methods that include calculating the average expression of each protein.

[0117] In some embodiments, subjects with a number of resistance-associated factors (e.g., RAP) above a predetermined number are predicted to be resistant to the therapy. In some embodiments, subjects with a number of resistance-associated factors above a predetermined number are predicted to not respond to the therapy. In some embodiments, subjects with a number of resistance-associated factors above a predetermined number are predicted to be non-responders to the therapy. In some embodiments, subjects with a number of resistance-associated factors below a predetermined number are predicted to be suitable for the therapy. In some embodiments, subjects with a number of resistance-associated factors below a predetermined number are predicted to respond to the therapy. In some embodiments, subjects with a number of resistance-associated factors below a predetermined number are predicted to be responders to the therapy. In some embodiments, subjects with a number of resistance-associated factors at or below a predetermined number are predicted to be suitable for the therapy. In some embodiments, subjects with a number of resistance-associated factors at or below a predetermined number are predicted to respond to the therapy. In some embodiments, subjects with a number of resistance-associated factors at or below a predetermined number are predicted to be responders to the therapy.

[0118] In some embodiments, the predetermined number is a threshold number. In some embodiments, the predetermined number is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. Each possibility represents a separate embodiment of the invention. In some embodiments, the predetermined number is 3. In some embodiments, the predetermined number is 4. In some embodiments, the predetermined number is 7. In some embodiments, the predetermined number is 13.

[0119] In some embodiments, the method further comprises categorizing the resistance-associated factor into at least one pathway, process, or network. In some embodiments, the method further comprises performing an analysis on the resistance-associated factor to determine at least one pathway, process, or network in which the resistance-associated factor is involved. In some embodiments, the pathway, process, or network causes non-responsiveness to therapy. In some embodiments, the analysis is selected from pathway analysis, process analysis, and network analysis. In some embodiments, the method further comprises performing a pathway analysis on the RAP. In some embodiments, the method further comprises performing a process analysis on the RAP. In some embodiments, the method further comprises performing a network analysis on the RAP. In some embodiments, the at least one pathway, process, or network comprises at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 pathways, processes, or networks. Each possibility represents a separate embodiment of the present invention. In some embodiments, the at least one pathway, process, or network is all pathways, processes, or networks known to include resistance-associated factors. In some embodiments, the at least one pathway, process, or network is all pathways, processes, or networks that are enriched for resistance-associated factors. In some embodiments, the enriched ones are the most enriched. In some embodiments, the enriched ones contain the most RAPs of any or the pathways, processes, or networks.

[0120] In some embodiments, the method includes selecting a pathway, process, or network. In some embodiments, the selected pathway, process, or network is hypothesized to affect non-responsiveness to therapy. In some embodiments, the selected pathway, process, or network is hypothesized to cause non-responsiveness to therapy. In some embodiments, the selected pathway, process, or network is known to be druggable. In some embodiments, the known druggable includes a known therapeutic agent that modulates the pathway, process, or network. In some embodiments, the known therapeutic agent is in or has completed clinical trials. In some embodiments, the known therapeutic agent is approved for use in humans. In some embodiments, the approved for use in humans is approved for use in treating the disease in humans. In some embodiments, the disease is cancer. In some embodiments, the method further includes administering to the subject who is or is predicted to be a non-responder an agent that modulates at least one pathway, process, or network containing a resistance-associated factor. In some embodiments, the agent inhibits a target in the pathway, process, or network. In some embodiments, the target is a gene. In some embodiments, the target is a protein. In some embodiments, the protein is a regulatory RNA. In some embodiments, the target is a response-associated factor. In some embodiments, the target is not a response-associated factor. In some embodiments, the agent activates a target in a pathway, process, or network. In some embodiments, the agent modulates a pathway, process, or network. In some embodiments, activity of the pathway induces unresponsiveness and the agent inhibits the pathway. In some embodiments, activity of the pathway reduces unresponsiveness and the agent activates the pathway. It will be understood by those skilled in the art that a response-associated factor is identified by its expression in a subject that is more similar to its expression in non-responders than in responders.Thus, for example, if the factor is more highly expressed in non-responders and increases the activity of the pathway / process / network, the agent inhibits the pathway. For example, if the factor is more highly expressed in non-responders but decreases the activity of the pathway / process / network, the agent activates the pathway / process / network. Similarly, if the factor is, for example, less expressed in non-responders and decreases the activity of the pathway / process / network, the agent inhibits the pathway / process / network. And finally, for example, if the factor is less expressed in non-responders but increases the activity of the pathway / process / network, the agent activates the pathway / process / network. In essence, the agent should induce the pathway / process / network to function more in the responder. In some embodiments, the agent targets a hub target in the pathway. In some embodiments, the agent targets a regulator target in the pathway. In some embodiments, the process activity induces unresponsiveness and the agent inhibits the process. In some embodiments, the process activity reduces unresponsiveness and the agent activates the process. In some embodiments, the agent targets a hub target in the process. In some embodiments, the agent targets a target of a regulator of the process. In some embodiments, the network activity induces unresponsiveness and the agent inhibits the network. In some embodiments, the network activity reduces unresponsiveness and the agent activates the network. In some embodiments, the agent targets a hub factor of the network. In some embodiments, the agent targets a regulator of the network. In some embodiments, the regulator is a master regulator. The factors can be classified into pathways, protein interactions, or signals using any analytical tool known in the art. Examples include, but are not limited to, GO analysis, Ingenuity analysis, Metacore analysis (Clarivate Analytics), reactome pathway analysis, and functional analysis.

[0121] According to another aspect, a computer program product is provided that includes a non-transitory computer readable storage medium having program code embodied therein executable by at least one hardware processor for performing the method of the present invention.

[0122] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium(s) having computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0123] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), static random access memories (SRAMs), portable compact disk read-only memories (CD-ROMs), digital versatile disks (DVDs), memory sticks, floppy disks, mechanically encoded devices with instructions recorded thereon, and any suitable combinations of the above. A computer-readable storage medium as used herein should not be construed as a transitory signal itself, such as an electric signal transmitted over a wire, or an electromagnetic wave or other freely propagating electromagnetic wave, or an electromagnetic wave propagating through a wave guide or other transmission medium (e.g., a light pulse passing through a fiber optic cable). Rather, the computer-readable storage medium is a non-transitory (ie, non-volatile) medium.

[0124] The computer-readable program instructions described herein can be downloaded to each computing / processing device from a computer-readable storage medium or via an external computer or storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fiber, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium of the respective computing / processing device.

[0125] The computer readable program instructions for carrying out the operations of the present invention may be either source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, or object oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry to perform aspects of the invention.

[0126] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, and the instructions executing via the processor of the computer or other programmable data processing apparatus create means for performing the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams. These computer readable program instructions may also be stored on a computer readable storage medium capable of directing a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that a computer readable storage medium having instructions stored thereon includes a product including instructions implementing aspects of the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.

[0127] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause the computer, other programmable apparatus, or other device to execute a series of operational steps to generate a computer-implemented process, such that the instructions executed by the computer, other programmable apparatus, or other device perform the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams. As used herein, the term "about" when combined with a value refers to ±10% of a reference value. For example, a length of about 1000 nanometers (nm) refers to a length of 1000 nm +- 100 nm.

[0128] It should be noted that, as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "a polynucleotide" includes a plurality of such polynucleotides, a reference to "the polynucleotide" includes a reference to one or more polypeptides and equivalents thereof known to those of skill in the art, and so forth. It should be further noted that the claims may be drafted to exclude any optional element. Thus, this statement is intended to serve as a predicate for the use of exclusive language such as "solely," "only," and the like in connection with the recitation of claim elements or the use of a "negative" limitation.

[0129] When a convention similar to "at least one of A, B, and C, etc." is used, such a configuration is generally intended in the sense that one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B and C together, etc.). It will be further understood by one of ordinary skill in the art that substantially any disjunctive word and / or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to consider the possibility of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" is understood to include the possibilities of "A" or "B" or "A and B."

[0130] It is understood that certain features of the invention that are described for clarity in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, various features of the invention that are described for brevity in the context of a single embodiment may also be provided separately or in any suitable subcombination. All combinations of the embodiments related to the present invention are specifically embraced by the present invention and are disclosed herein as if each and every combination were individually and expressly disclosed. Moreover, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein as if each and every such subcombination were individually and expressly disclosed herein.

[0131] Additional objects, advantages, and novel features of the present invention will become apparent to those skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below finds experimental support in the following examples.

[0132] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples. EXAMPLES

[0133] In general, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are fully explained in the literature. See, for example, "Molecular Cloning: A laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, RM, ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley&Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York. New York (1998); methods described in U.S. Patent Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659; and 5,272,057; “Cell Biology: A Laboratory Handbook”, Volumes I-III Cellis, JE, ed. (1994); “Culture of Animal Cells-A Manual of Basic Technique” by Freshney, Wiley-Liss, NY (1994), Third Edition; “Current Protocols in Immunology” Volumes I-III Coligan JE, ed. (1994); Stites et al.(eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization-A Laboratory Course Manual" CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this document.

[0134] Example 1: Response prediction based on resistance associated proteins (RAPs) - proof of concept Data collection The response prediction proof-of-concept is based on the analysis of blood samples from 108 non-small cell lung cancer (NSCLC) patients under immune checkpoint inhibitor (ICI) treatment. The different treatments administered are summarized in Table 1.

[0135] [Table 1]

[0136] Plasma protein levels were measured in 108 patients, measuring approximately 1100 non-redundant protein targets. For a total of 156 samples in the batch, samples were taken before the start of ICI treatment (T0) and after the first treatment was administered (T1).

[0137] Classifier configuration Proteomic levels and response markers were incorporated by a supervised learning algorithm to predict response to treatment. Response markers were responder (R) and non-responder (NR), determined based on overall response rate (ORR) assessment at 3 months. Specifically, progressive disease (PD) or early death related to disease progression was classified as NR, and stable disease (SD), minimal response (MR), partial response (PR) and complete response (CR) were classified as R. ORR assessment was performed as described in clinical trial NCT04056247 (clinicaltrials.gov / ct2 / show / NCT04056247, incorporated herein by reference in its entirety) in the section "Primary Outcome Measures" by RECIST1.1 or other validated methods for ORR assessment. Changes in blood levels of various proteins [time frames: baseline (pre-therapy) and after the first drug administration (post-therapy)] representing host response were determined as described.

[0138] The samples were divided into a training set and a test set. All development stages of the algorithm were performed using the training set, and the test set was used only in the final stage to test the performance of the final algorithm. The training set contained samples from n=78 patients (59 responders and 19 non-responders), and the test set contained samples analyzed in n=30 patients.

[0139] The response classifier treats features as input and predicts the response based on the feature values. The features are protein levels measured in plasma at two time points: baseline (T0) and after the first treatment (T1). Measurements of the same protein at different time points are considered as independent features. Furthermore, some proteins have more than one measurement in a single proteomic profile (e.g., the protein IL-6 is measured four times). Each repetition was treated as an independent feature.

[0140] Resistance-associated proteins Resistance-associated proteins (RAPs) refer to specific proteins whose expression in a given patient confers resistance to therapy, i.e., RAPs are patient-specific. A protein is considered to be a RAP if its expression level in the respective patient is more similar to its expression distribution in the non-responder population than in the responder population (see Figures 1A-1C for illustration). RAPs can be determined in a variety of ways. Provided herein are mathematical calculations of RAPs, as well as machine learning algorithms for classifying RAPs, and methods for combining the two. These methods are merely exemplary, and any method of calculating RAPs can be used.

[0141] To quantitatively express the above concept, a RAP score (i.e., resistance score) was determined for each protein. A low RAP score value represents an expression level typical of a responder population, and a high RAP score indicates an expression level typical of a non-responder population. A protein is considered a RAP if its RAP score exceeds a certain threshold (e.g., above or below depending on the composition of the score). The RAP score threshold optimization process is described below.

[0142] Calculation of the RAP score requires knowing the expression level distribution of each protein in both the responder and non-responder populations, as well as data on the protein level expression of the patients tested. To allow comparison between several different proteins at different expression level ranges, it is important that the RAP score is insensitive to and not sensitive to the scale of protein level expression. This is particularly important for plasma samples, where there is a large dynamic range of 11 orders of magnitude in protein expression levels. To achieve this, the RAP score is based on the Z-score, which counts the distance of an individual level from the population mean in units of the population standard deviation. Technically, the Z-score is defined by Equation 1: Formula 1:

number

number

[0143] Algorithm 1: The monotonic function used in Eq. if|mean(R)-mean(NR)|>c·std(NR)then if mean(NR)>mean(R) then

number

number

[0144] To determine the exact number of RAPs for a given patient, a threshold was determined for all proteins, and proteins with a RAP score above the determined threshold were considered RAPs. The threshold was determined using cross-validation applied to the training set. Specifically, a cross-validation dataset consisting of 1 / 3 of the training set and a non-cross-validation dataset consisting of a further 1 / 3 of the training set were sampled while keeping the number of responders and non-responders similar between the cross-validation and non-cross-validation datasets. Calculations were performed on the non-cross-validation set, and then for each patient in the cross-validation dataset, the expression level distribution of responders and non-responders was used to calculate a RAP score for each feature (i.e., all proteins measured at T0 and T1). The number of RAPs was then used to predict response, and the area under the receiver operating characteristic (ROC) curve (AUC), which quantifies the predictive performance, was calculated for each threshold (Figure 3A-B). To minimize the noise associated with small datasets, we performed 100 realizations for each threshold (i.e., different sampling of the cross-validation set from the training set) and considered the average of the AUC over the 100 realizations. Notably, the average ROC AUC curve in Figure 3A shows a single broad peak, suggesting that the predictive power of RAP count is not very sensitive to the chosen threshold. For features including measurements at T0 and T1, the thresholds were set at 1.61 (Figure 3B) and 2.9 (Figure 3A), respectively.

[0145] Machine Learning Evaluation: Purely mathematical approaches are valid (both conceptually and practically), but they have some shortcomings that need to be addressed. 1. Because the RAP score function depends on the underlying distribution of protein expression levels, its effectiveness may be platform-dependent (especially since different proteomics systems use different units of measurement that do not scale naturally). 2. The current implementation does not provide a natural way to include clinical parameters (e.g., patient condition, indication details, treatment details, etc.) in the predictor.

[0146] An alternative approach was invented that utilizes decision tree learning based on machine learning algorithms to classify proteins as RAPs for a given subject. For each protein measured, a machine learning algorithm (e.g., XGBoost algorithm) was used to generate a predictive model based on the data in the training set. Such data from the training set may include not only protein expression levels and responder / non-responder tags, but also other features such as patient age, sex, condition, type of treatment, line of treatment, biomarker expression such as expression of PD-L1. This approach makes no assumptions on protein distribution and provides a natural framework for utilizing clinical parameters.

[0147] To test this approach, samples from a cohort of 76 patients were screened using two different protein analysis platforms: approximately 1200 proteins (O), and approximately 7500 proteins for other measurements (S), with approximately 1000 proteins common to both platforms. The treatments administered to these subjects are summarized in Table 2.

[0148] [Table 2]

[0149] The cohort of 76 patients was split into a training set containing 51 subjects (38 responders and 13 non-responders) and a test set containing 25 subjects (19 responders and 6 non-responders). The XGBoost algorithm was chosen for this analysis because of the non-linear nature of the problem and the algorithm's reputation for its efficiency in learning on small datasets. To avoid multiple comparisons in the test set, which would increase the risk of false discoveries, and because the goal of the study was to validate predictive potential (rather than identifying an optimal model configuration), the following predefined configurations were used for the training models: The model hyperparameters were set as follows: a. Max tree depth=4 b.Riding factors: eta=0.8,lambda=5,alpha=2 c.num_parallel_tree=100 d.objective=binary:logistic e.eval_metric=logloss The parameters were chosen to handle small noisy data sets.

[0150] For the purposes of this evaluation, the machine learning algorithm was trained on protein expression levels only, excluding other considerations. Patient expression results were evaluated for each protein separately, and a protein classifier was calculated for each single protein. The machine learning algorithm output a score between 0 and 1, with 1 being most similar to non-responders and 0 being most similar to responders.

[0151] To evaluate the approach, we used two configurations of proteins as input. In the first configuration, all proteins were used as potential predictors. This is similar to that used in the mathematical approach, but it is expected that the method will be effective in large cohorts, and in small cohort sizes (compared to the number of features), false positives may hinder the predictive power. In the second configuration, we used ranking single-protein models according to their tendency to split patients into responders and non-responders (i.e., giving higher ranks to protein models with more balanced predicted classes). As an extreme example, if the model predicted that all patients belonged to a single class (responders or non-responders), the model received the lowest possible balanced rank. At the opposite side of the scale, models that split the population evenly between responders and non-responders received the highest balanced rank. After ranking the different protein models, the 200 proteins with the highest balanced ranks were used to evaluate the machine learning approach.

[0152] Using both approaches, subjects were assessed based on the values ​​of "O" and "S" expression at T0 and T1. The model performance for "O" (measured by AUC) was above 0.8 in the range of thresholds from 0.4 to 0.8, with a stable smooth behavior (Figure 3C), peaking at AUC = 0.89 and 95% confidence interval [0.594, 0.995]. Therefore, for these samples, the threshold was set at approximately 0.6. This result is a slight improvement over the AUC = 0.846 obtained for the same dataset using the mathematical RAP approach. However, due to the large confidence interval (a consequence of the small dataset size), the statistical significance of this difference is moderate.

[0153] Peak model performance for "O" when restricting predictors to 200 proteins was AUC=0.91 and 95% confidence interval of [0.602,0.996] (Figure 3C). In this case, the threshold is essentially the same and the AUC represents a slight improvement over the construction of the full protein set.

[0154] Model performance (measured by AUC) for "S" was above 0.75 for a threshold range of 0.4-0.9, showing a stable and smooth behavior (Figure 3C), peaking at an AUC of 0.81, 95% confidence interval [0.587, 0.924]. Thus, for these samples, the threshold can be set slightly lower at about 0.59, although this difference may be negligible. This result is inferior (about one standard deviation lower) to the behavior observed from the same model configuration using the "O" data, which is not unexpected given the significantly higher number of proteins and smaller size of the dataset.

[0155] Peak model performance for "S" when restricting predictors to 200 proteins was AUC=0.87 and 95% confidence interval of [0.597,0.992] (Figure 3C). The threshold is therefore essentially the same as that found in the "S" analysis and represents a considerable improvement compared to the full protein set configuration, consistent with the reduced false discovery rate imposed by this configuration. Still, the performance of the 200 protein configuration in "S" is slightly lower than the same configuration using "O". However, the difference (<0.3 standard deviations) is of low statistical significance.

[0156] Response prediction by RAP number The above RAP score allows to identify patient-specific proteins whose expression levels correspond to non-responsiveness, as reflected by the expression of responders and non-responders. Therefore, it was hypothesized that the number of RAPs a particular patient has predicts the response of the patient. Since almost all of the measured proteins show expression levels that match the responder population, patients with few or no RAPs are expected to respond to treatment. Patients with a high number of RAPs are expected to develop resistance, since the expression levels of some proteins are similar to the non-responder population. This method does not take into account the nature of the RAPs, and each subject may have completely different RAPs. Rather, it is the total number of RAPs that is important, not the identity of the RAPs.

[0157] The predictive performance of the RAP score was tested using the test set. Specifically, for each patient in the test set (n=30), the RAP score was calculated for all features using the R and NR protein level distributions of all patients in the training set (n=78). Together with the threshold calculated using the training set as explained above, it is possible to infer the number and discriminatory information of RAPs for each patient in the test set. Figure 4A shows 30 subjects from the test set and the number of RAPs calculated (using mathematical methods) for each subject using T0+T1 data. The threshold was set at 3 RAPs, and subjects with more than 3 RAPs were predicted to be non-responders. The ROC curve shows an AUC of 0.88, indicating that the analysis is highly predictive (Figure 4B).

[0158] Targeted RAP Improved understanding of the molecular and immunological mechanisms of resistance to ICI therapy may not only identify novel predictive biomarkers but also suggest targets for combination ICI therapy, which aims to selectively block ICI resistance proteins to improve ICI outcomes in non-responder patients.

[0159] To find targets for combination therapy, we evaluated all RAPs with a score >2.9 (a defined threshold) seen in the test set patients. We then investigated the search for clinical trials targeting RAPs from this list in combination with ICIs in non-small cell lung cancer (NSCLC) patients or patients with solid tumors. Mapping of clinical trials with combination therapy yielded 1300 clinical trials targeting 430 proteins in combination with ICIs in NSCLC or solid tumors or by 500 different drugs. Comparing the 30 RAPs that met the score threshold in the test set (RAPs that appeared in at least one patient out of 30 patients and had a score higher than 2.9) with the list of proteins found to be targeted in clinical trials in combination with ICIs, we found four RAPs that were also targeted in combination with ICIs in NSCLC trials: KDR (VEGFR2), IL6, EPHA2, and TACSD2.

[0160] IL-6 is one of the targetable RAPs identified in a test set cohort of patients. Recently, we have shown that the therapeutic efficacy of anti-CTLA-4 is significantly improved by co-administration of anti-IL-6 in tumor-bearing mice (Khononov, et al., 2021, “Host response to immune checkpoint inhibitors contributes to tumor aggressiveness”, J. Immunother. Cancer, Mar;9)3_:e001996; incorporated herein by reference in its entirety). These results are consistent with previous publications demonstrating improved therapeutic outcomes when anti-IL-6 is combined with anti-PD1 or anti-PD-L1 treatment. Furthermore, in vitro experiments in Khononov et al. demonstrate that inhibition of IL-6 reduces anti-PD-1-induced tumor cell invasive properties, further supporting the idea that blocking specific therapy-induced host factors represents a strategy to overcome therapy resistance.

[0161] An alternative approach for RAP-based therapy targeting is by relating proteins to key biological processes related to cancer. To this end, each protein was assigned to a cancer hallmark capturing a key tumorigenesis process. An enrichment analysis was then performed for each patient with RAP as input (Fisher's exact test; Figure 5). Preliminary analysis of six patients revealed a total of four enriched processes. One patient had significant enrichment in all four processes; four patients showed enrichment in one to three processes; one patient did not have any significant process.

[0162] Once enrichment analysis is performed on a patient, the treating physician can select therapy based on the enriched biological process. For example, if angiogenesis is significantly enriched, the physician can choose to combine an approved drug that targets angiogenesis (e.g., Avastin) with ICI. Another example is a patient with high proliferation signal. In this case, the physician can choose to combine ICI with chemotherapy against tumor cell proliferation.

[0163] To further explore the biological aspects of RAP, we examined 19 RAPs obtained in at least three patients in the study set cohort. Most patients had 4–5 RAPs. The most common RAP among the patients tested was VEGFR2 (KDR; identified as a RAP in 12 patients). Of note, most of the RAPs were identified in T1, suggesting that resistance to therapy was primarily acquired and arises from a host response. VEGFR2 was identified as a RAP in both T0 and T1, but in T1 it was defined as a RAP in more patients (12 patients compared with 8 in T0). VEGFR2 is one of the two receptors for vascular endothelial growth factor (VEGF), a major growth factor for endothelial cells, whose expression was higher in responders.

[0164] Network analysis revealed that the majority of RAPs are functionally related to each other, with five of them being highly interconnected (Figure 6). Most of the proteins were associated with at least one hallmark of cancer, further implying that these RAPs are indeed associated with resistance to therapy. Several hallmarks of cancer were significantly enriched in the 19 RAPs (Figure 7). Multiple intracellular and membrane proteins were identified as RAPs (Figure 6). Therefore, an analysis of putative cell of origin was performed to further understand the results (Figure 8). Enrichment of lung and bronchus as cell types of origin was observed. Furthermore, expression of the 19 RAPs was examined in various cancer types, and enrichment for lung cancer was also observed (Figure 9).

[0165] Example 2: Fusion of RAP and clinical data A cohort of 184 NSCLC patients was obtained in whom blood samples were obtained before (T0) and after (T1) the first dose with ICI. Protein levels were measured. Response assessment was based on ORR at 3 and 6 months after treatment initiation and durable clinical response (DCB) at 1 year. Progression-free survival (PFS) and overall survival (OS) were also monitored. For the 3- and 6-month assessments, subjects with progressive disease or death were considered as non-responders, and subjects with stable disease, minimal remission, partial remission, and complete remission were considered as responders. DCB was defined as 1-year PFS with continued ICI treatment. Cases who stopped ICI treatment due to adverse events (but without signs of progression) were treated as responders. Additional clinical information collected throughout the study included line of treatment (first or progression), PD-L1 immunostaining (<1%, 1–49%, >50%), age, and sex (see Figures 10A–10F). The analyses presented are based on T0 only. The breakdown of ICIs / therapy used is shown in Table 3.

[0166] [Table 3]

[0167] The cohort was split into a development set (60% of subjects) and a validation set (40% of subjects). The development set was further split into a training set and a test set. The model was trained on the training set and predictions were generated for the subset of patients not seen by the model during training (i.e., the test set). In order to generate stable predictions for all patients in the development set, the split of the development set into training and test sets was performed multiple times (each time the model was trained on a different subset of the development set and predictions were performed for the remaining patients, i.e., the training and test sets were mixed and remixed, and dozens of iterations were performed to test that the model / classifier was effective across the entire development set). The quality of the predictions was then quantified by calculating the ROC AUC for patients included in the development set. The validation set was used only at the end of the analysis to verify the functionality of the final classifier. This split was performed multiple times.

[0168] Models were generated based on response assessments at three time points: 3 months, 6 months, and 1 year after initiation of treatment. All 184 patients were evaluated at 3 months, 177 at 6 months, and 146 at 1 year. Resistance increased over time. 26% of subjects were nonresponders at 3 months, 45% at 6 months, and 74% at 1 year. These proportions were similar between the development and validation sets.

[0169] During model generation based on the development set, the development set was randomly split into a training set and a test set 60 times. At each iteration, the top candidate proteins were selected using the Kolmogorov-Smirnov test, which defines how well each protein distinguishes between responders or non-responders. For each selected protein, a single-protein XGBoost model (SP model) was generated based on the training set to make predictions for the test set. A protein was defined as a RAP for a particular patient if its predicted probability of resistance (i.e., resistance score) exceeded a predefined threshold, and the average of all iterations was used for each patient. To handle class imbalance, a uniform threshold was assigned to all models. Different thresholds were defined for each time point (e.g., 3-month threshold = 0.25, 6-month threshold = 0.42, 1-year threshold = 0.45). For each patient, the number of proteins whose model scores exceeded the predefined threshold (i.e., number of RAPs) was calculated.

[0170] In this cohort, only ascertaining the number of RAPs was predictive. However, a predictive model was created that could also integrate clinical data. The presented clinical classifier used as input the number of RAPs, the line of treatment (ICI was the first or progression line), the subject's age and the percentage of PD-L1 staining in the tumor (<1%, 1-49%, or >50% of positive cells). The classifier then generated a final resistance score between 0 and 1, where 0 was most similar to a responder and 1 was most similar to a non-responder. Subjects with a score above a predefined threshold were predicted to be non-responders. Similarly, a response score was also calculated, which was 1-resistance score. For the response score, subjects with a score above a predefined threshold were predicted to be responders.

[0171] To test the performance of the classification model, the final resistance score along with the actual response was used to calculate the ROC AUC. The ROC AUC was calculated separately for 3-month ORR, 6-month ORR, and 1-year DCB for both T0 and T1. The results are summarized in Figure 11A. The classifier was found to be predictive for both the development and validation sets at all time points. A similar analysis showed that the classifier was also found to be predictive for the T1 data at all time points (Figure 11B).

[0172] In addition to checking the performance of the classification model, we also examined the correlation between the predicted response probability (response score) assigned to each patient by the classification model and the observed response probability. To this end, for each value of the response score S0, the observed response probability is given by the proportion of responders among patients assigned a response score within the range of S0 ± 0.1. The choice of the interval of ± 0.1 is arbitrary and reflects the size of the validation set. Within a larger validation set, the interval can be further reduced. The agreement between the predicted response score and the actual response probability was quantified by the construction of the goodness of fit R^2. The goodness of fit for all three time points (3-month ORR, 6-month ORR, and 1-year DCB) was R^2 = 0.98 for time point T0 (Figures 12A-12B).

[0173] Patients in the validation set were stratified into long-term benefit and limited benefit groups, and stratification was based on predicted 3-month response scores. In survival analysis, the quality of stratification was measured by hazard ratio (HR), which gives the ratio of the probability of an event per unit time in the two groups. For example, an HR of 4 for overall survival (OS) means that the probability of a death event per unit time in the limited benefit group is four times that per unit time in the long-term benefit group. The HR in the validation set was 2.27, p<0.004 for PFS (Figure 13A) and 4.50, p<0.0001 for OS (Figure 13B).

[0174] This validation experiment demonstrates that a classifier incorporating clinical data and RAP number is highly predictive of patient response.

[0175] Functional network analysis of RAP The RAP-based analysis is further used as the basis for the generation of a resistance map (Figure 14A). The resistance map displays both the interactions between RAPs and the RAP functions. For this purpose, a RAP is defined if a protein was selected in at least 10 model iterations in one or more patients (during RAP calculation, the model runs 60 iterations and the number of times a given protein was selected for the model is recorded), resulting in a total of 73 RAPs in the current patient cohort. Each node represents a RAP, and the edges between the nodes indicate functional relationships. Nodes of larger size indicate investigational new drugs (INDs) combined with immunotherapies. The nodes are colored based on the protein's function. Although the map shows multiple interactions between different RAPs, RAPs are involved in different functional processes that may be related to resistance to therapies such as splicing, immune regulation, angiogenesis, and cell proliferation. Patient-specific maps can be generated based on the patient's RAPs, which can help 1) map resistance mechanisms in individual patients, and 2) identify targeted therapies to counter resistance. Two examples of patients in the cohort are shown in Figure 14B. In these examples, the non-responder had 44 RAPs and a response probability score of 0.44 (which corresponds to a resistance score of 0.56, which is above the predefined threshold of 0.2 for non-response). This patient had RAPs from multiple functional groups, but no DNA-related RAPs were present in this patient. The second subject was a responder with 10 RAPs, which is below the predefined threshold. These RAPs were mainly related to the cytoskeleton. This patient had a high response probability of 0.91 (which corresponds to a resistance score of 0.09, which is below the predefined threshold of 0.2).

[0176] Further examination of patient RAPs shows functional differences between RAPs, with higher orders represented in each responder group (Figure 15). Non-responder RAPs are involved in splicing, signal transduction, and cytoskeleton-related processes, whereas responder RAPs are primarily involved in protein degradation and cell adhesion. Interestingly, the higher RAPs in the responder group contain two peptidases that may be involved in antigen presentation, thereby facilitating response to therapy. To convert non-responders to responders, RAPs for which known therapeutic agents exist are selected. Agents are selected to modulate RAPs to alter the function of the pathway, more closely approximating the function of the pathway in responders. When therapeutic agents targeting RAPs are unavailable or undesirable, therapeutic agents that modulate pathways that include RAPs are selected. The selected agent must modulate pathways that include RAPs to alter the function of the pathway, so that it more closely approximates the function of the pathway in responders. The therapeutic agents are used to convert non-responders to responders, or as combination therapy with ICIs.

[0177] While the present invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and broad scope of the appended claims.

[0178] Example 3: Validation of the RAP-based model To validate the RAP-based model described above, we used a larger cohort of NSCLC patients. Plasma samples and clinical data from 339 ICI-treated NSCLC patients were collected. Pre-treatment plasma samples were profiled by a protein assay that quantified approximately 7000 proteins in a single plasma sample.

[0179] The clinical parameters of the patients are shown in Figure 16A. The median age was 65. Approximately one-third of the patients were female. The majority of patients (78.47%) had non-squamous cell carcinoma (mostly adenocarcinoma), and 21.24% of patients had squamous cell carcinoma. Patients were treated with either immune checkpoint blockade (ICB)-chemotherapy combination (59.88%) or ICB monotherapy (40.12%). The distribution of patients with PD-L1 negativity (<1%), low expression (1–49%), and high PD-L1 expression (≥50%) was roughly equivalent. The high PD-L1 group was the largest (36%).

[0180] Clinical response to treatment was assessed at 3, 6, and 12 months after initiation of treatment, and at each time point, patients were classified as responders or non-responders. At 3 and 6 months, patients with complete remission, partial remission, or stable disease were classified as responders, and patients with progressive disease were classified as non-responders. At 12 months, patients with sustained clinical response (defined as no progressive disease for at least 1 year after initiation of treatment) were classified as responders, and all other patients were classified as non-responders. Based on these criteria, 69.32%, 46.02%, and 24.78% of patients were classified as responders at 3, 6, and 12 months, respectively (Figure 16). The size of the cohort varied between time points due to patient deaths (Figure 17). The dataset included 339, 331, and 299 patients for the 3, 6, and 12 month time points, respectively.

[0181] Even though PD-L1 expression correlated with response at 6 and 12 months, response prediction using this parameter alone at each of the three time points was poor, with areas under the curve (AUC) of the receiver operating characteristic (ROC) plots of 0.5, 0.6, and 0.55 at 3, 6, and 12 months, respectively (Figure 18A).

[0182] We next asked whether integrating additional clinical parameters would improve the predictive ability of the PD-L1 biomarker. Three clinical parameters, namely patient age, patient gender, and line of treatment, are known to correlate with response. Thus, we developed a predictive model based on PD-L1, age, gender, and line of treatment (the "clinical model"). The clinical model showed only a slight improvement in response prediction ability compared to PD-L1 alone, with AUCs of 0.52, 0.6, and 0.62 at 3, 6, and 12 months, respectively (Figure 18B). Further improvement in predictive performance is needed.

[0183] With the goal of developing a more robust predictive model, we designed an additional model whose output is the sum of individual features associated with therapy response. Each feature has a small effect on the final output by itself, thus minimizing the impact of false discoveries and maintaining the stability of the model, thereby potentially mitigating the impact of significant heterogeneity between patients and a large number of features in a relatively small cohort. The model is based on a set of proteins that show differential plasma levels in responder and non-responder populations, as determined by statistical tests (RAPs). Such proteins serve as potential indicators of treatment response, depending on their plasma levels in individual patients. Specifically, for a given patient, the ML-based model trained on the responder and non-responder populations infers a prediction of "active RAP" or "inactive RAP" from the plasma levels of each of the patient's potential RAPs in the total RAP set. In this way, an individualized RAP profile is assigned to the patient. The sum of the number of individual RAPs reflects the likelihood that the patient will respond to treatment. Patients who exhibit a large number of "active predictions" in the RAP set (and therefore many individual RAPs) are more likely to not respond to treatment, whereas patients with a large number of "inactive" predictions in the RAP set (and therefore few individual RAPs) are more likely to respond to treatment. Similarly, the ML-based model actually gives an activity score for each RAP. These scores can be combined (summed up) to generate a total RAP score, which in turn can be used to predict patient response (the more activity from the RAPs, the higher the likelihood of non-response). Three RAP-based models were developed, one for each of the three response assessment time points. The models were developed following the same workflow, using response labels for the 3-, 6- or 12-month time points, together with protein expression data and patient gender, as inputs to determine an individual's RAP and response probability (determined by the total RAP score) (Figure 19A). The patient cohort was split into a development set (75% of the patient cohort; n=254) and a validation set (25% of the patient cohort; n=85).The development set was further divided into a training set and a test set consisting of 75% and 25% of the patients in the development set, respectively (Figure 19B). Proteins showing statistically significant differences in plasma level distribution between the responder and non-responder populations were identified in the training set, and the 50 most significant proteins were identified as the common RAP set. For each RAP, an ML algorithm was trained on two features, namely RAP plasma expression level and patient gender, to develop a predictive model of RAP activity / expression based on the training set. Activity predictions (proteins of the RAP active / expressed in a given subject) were then generated for each RAP for each patient in the test set, resulting in a prediction of "active" or "inactive" for each single RAP. The three-step process (i.e., RAP selection, training the model, and predicting activity) was repeated 80 times, each time randomly splitting patients into training and test sets. In each iteration, the activity scores across the 50 selected RAPs were summed for each patient to obtain a total RAP score. After 80 iterations, the total RAP scores were averaged for each patient and linearly scaled to a value between 0 and 1. The final output of the model was converted to a response probability, a clinically oriented metric that reflects the likelihood that a patient will respond to treatment. (Similarly, the total number of active RAPs can be used to make response predictions, as described above.)

[0184] Using this method, patients were randomly mixed between the training and test sets 80 times, and 50 RAPs were selected from the training set at each time point, with the same RAPs being selected several times overall (Figure 20A). Of the total of 287, 330, and 371 RAPs selected for the 3, 6, and 12 month time points, respectively, approximately 30 RAPs were selected in at least 50% (≧40 times) of the repetitions for each time point (Figure 20B). Across the three time points, a total of 598 RAPs were selected, of which 113 RAPs were common to all three time points, while 97, 85, and 139 RAPs were unique to the 3, 6, and 12 month time points, respectively (Figure 20C). Notably, a large number of RAPs were selected multiple times across the three time points (Figure 20D). Biological processes associated with RAPs across all time points include splicing, complement and coagulation cascades, and peptidase activity, as well as multiple signaling cascades (Figure 20E).

[0185] After model development, the RAP-based models for each time point were tested on an independent validation set (25% of the patient cohort; n=85) that was not used during model training. At each time point, the response probability was determined for each patient in the validation set. The range of the response probability distribution differed at each time point, and the median response probability decreased over time (Figure 16A; Figure 6). Furthermore, the response probability for all patients decreased from one time point to any subsequent time point (Figure 22A-22C), consistent with the observed decrease in response rate over time (Figure 22D). Notably, the observed non-responders clustered in the lower range of predicted response probabilities for all three time points, indicating that the model has high predictive power (Figure 21). Furthermore, enrichment analysis based on response probability (2D enrichment test; false discovery rate <0.05) showed that high response probability was significantly enriched for responders, females, non-squamous cell carcinoma patients, and patients without progressive disease or death events at all three time points. On the other hand, low response probability was significantly enriched in non-responders, males, patients with squamous cell carcinoma and patients with progressive disease or death events (Figure 23).

[0186] The median response probability was used as a threshold to classify patients into high or low response probability groups, and patients with predicted response probabilities above or below the median were assigned to the high or low response probability group, respectively. Cox regression analysis demonstrated that across the three time points, patients in the high response probability group achieved significantly longer overall survival than patients in the low response probability group (Figure 24A, hazard ratio, HR = 0.24-0.38). Similar results were obtained for progression-free survival (PFS; Figure 24B, HR = 0.32-0.41). These findings demonstrate that the RAP-based model successfully classifies survival outcomes of ICB-treated NSCLC patients.

[0187] To further test the accuracy of the model, the predicted response probability was compared to the observed response rate. The latter refers to the proportion of observed responders among a group of patients assigned a similar probability of response (i.e., probability of response ±0.15). Linear regression analysis demonstrated a high goodness of fit (R 2 = 0.97) (Figure 25A). Moreover, the AUC of the ROC curve was 0.71, 0.77, and 0.78 at 3, 6, and 12 months, respectively (Figure 25B), demonstrating the strong predictive ability of the RAP-based model over the first year of ICB-based treatment. Notably, the RAP-based model showed superior predictive performance compared with the PD-L1-based model (AUC = 0.5-0.6 in the first year) and the clinical model (AUC = 0.52-0.62 in the first year) (Figure 18).

[0188] We next tested whether integrating clinical parameters into the RAP model would improve its predictive performance. To this end, we integrated the PD-L1-based model or the clinical model with the RAP-based model and compared the predictive performance. Interestingly, adding the PD-L1 parameter to the RAP-based model slightly increased the predictive performance at 6 months, whereas integrating the RAP-based and clinical models resulted in an overall worse predictive performance (Figure 26A). In survival analysis, the RAP model showed the best HR compared to the other four models, whereas the HR was not significant for the PD-L1-based and clinical models (Figure 26B).

[0189] In the RAP model, another model that integrated the number of active RAPs obtained according to PD-L1 expression and age, the AUC of the ROC curve was 0.66, 0.71, and 0.68 at 3, 6, and 12 months, respectively (Figures 27A-27C). Linear regression analysis demonstrated a high goodness of fit (R 2 = 0.94) (Figure 27D). Thus, the combination of RAP count and clinical parameters was found to be as good or superior to the combination of total RAP score (individual RAP activity level) and clinical parameters in its ability to accurately predict patient response.

[0190] Finally, we tested the performance of the RAP-based model in different patient subsets (Figure 28). The model showed strong predictive performance in both ICI monotherapy and ICI chemotherapy subsets, similar to its performance in the whole population. Meanwhile, histology subset analysis showed improved prediction in the squamous cell carcinoma subset at 3 months compared to the whole population. At 6 and 12 months, the strongest prediction was observed in the PD-L1 negative subset, while prediction was slightly weaker in the PD-L1 high subset compared to the whole population.

[0191] Example 4: RAP-based models predict different outcomes in PD-L1-high patients According to current guidelines for the early treatment of driver mutation-negative NSCLC, patients with PD-L1-high tumors are treated with ICI monotherapy or ICI in combination with chemotherapy, the latter therapy option being recommended in cases of aggressive disease. For patients with PD-L1-low or PD-L1-negative tumors, ICI in combination with chemotherapy is the only recommended option. In the cohort used, PD-L1-high patients showed a better prognosis and up to two-fold differences in median OS and PFS compared to PD-L1-low and PD-L1-negative patients (Figure 29A). Most PD-L1-high patients (65.3%) were treated with monotherapy ICI.

[0192] Among PD-L1 high patients receiving monotherapy ICI, it is possible to identify those who would have fared better with ICI-chemotherapy combinations. To investigate this, the ability of the RAP-based model to predict survival outcomes in subsets of patients (i.e., PD-L1 high patients receiving monotherapy ICI vs. PD-L1 high patients receiving ICI-chemotherapy combinations) was tested. In the monotherapy subset, patients in the high response probability group had significantly longer OS than patients in the low response probability group at all three time points (3 months HR=0.24, p<0.001; 6 months HR=0.36, p=0.004; 12 months HR=0.42, p=0.01; Figure 29B). Notably, in patients with high response probability, the median OS was not reached at 3 and 6 months, and was 32.6 months at 12 months, compared with a mean of 10.98 months in patients with low response probability. Notably, it exceeded the median OS of PD-L1 high patients overall (29.4 months). In contrast, in the combination subset, there was no significant difference in OS at any time point when comparing high and low response probability groups (Figure 29B). Notably, no significant difference in PFS was observed between high and low response probability groups in the monotherapy or combination subsets (Figures 30A, 30B, respectively). Taken together, the OS analysis demonstrates that the RAP-based model can identify PD-L1 high patients who may benefit less from ICI monotherapy and who may achieve better outcomes with the combination of ICI and chemotherapy.

[0193] For PD-L1-low / negative patients, both the high and low response probability groups showed poor prognosis in the monotherapy subset (Figure 29B). This indicates that the model did not identify a subgroup of PD-L1-low / negative patients who would benefit from ICI monotherapy. In the combination therapy subset, the high response probability group survived significantly longer than the low response probability group (3 months HR=0.34, p<0.0001; 6 months HR=0.32, p<0.0001; 12 months HR=0.38, p<0.0001), showing a similar outcome to PD-L1-high patients in this subset (Figure 29B). Thus, the RAP-based model distinguished between PD-L1-low / negative patients who would show a longer OS from combination therapy, where benefit is recommended, and those who would not. However, there are currently no therapeutic options to address the latter case.

Claims

1. 1. A method of predicting a response to a therapy in a subject suffering from a disease, comprising: a. i. in a population of subjects known to be afflicted with the disease and to respond to the therapy (responders), ii. in a population of subjects known to be suffering from the disease and who do not respond to the therapy (non-responders), and iii. In the subject receiving protein expression levels of a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, wherein the resistance score is based on the similarity of the factor expression level in the subject to the factor expression level in the responder and the similarity of the factor expression level in the subject to the factor expression level in the non-responder; c. classifying factors of said plurality of factors having a resistance score above a predetermined threshold as resistance-associated factors; Subjects with a number of resistance-associated factors greater than a predetermined number are predicted to be resistant to the therapy, and subjects with a number of resistance-associated factors equal to or less than a predetermined number are predicted to respond to the therapy; thereby predicting a subject's response to therapy. A method comprising:

2. 1. A method of predicting a response to a therapy in a subject suffering from a disease, comprising: a. i. in a population of subjects known to be afflicted with the disease and to respond to the therapy (responders), ii. in a population of subjects known to be suffering from the disease and who do not respond to the therapy (non-responders), and iii. In the subject receiving factor expression levels for a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, wherein the resistance score is based on the similarity of the factor expression level in the subject to the factor expression level in the responder and the similarity of the factor expression level in the subject to the factor expression level in the non-responder; c. classifying factors of said plurality of factors having a resistance score above a predetermined threshold as resistance-associated factors; d. Summarizing the number of resistance-associated factors present in the subject; and e. applying a trained machine learning algorithm to the number of resistance-associated factors and at least one clinical parameter of the subject, wherein the trained machine learning algorithm outputs a final resistance score and a final resistance score above a predetermined threshold, indicating the subject is resistant to the therapy; thereby predicting a subject's response to therapy. A method comprising:

3. 1. A method of predicting a response to a therapy in a subject suffering from a disease, comprising: a. i. in a population of subjects known to be afflicted with the disease and to respond to the therapy (responders), ii. in a population of subjects known to be suffering from the disease and who do not respond to the therapy (non-responders), and iii. In the subject receiving factor expression levels for a plurality of factors; b. calculating a resistance score for a factor of the plurality of factors, the resistance score being based on a similarity of the factor expression level in the subject to the factor expression level in the responder and a similarity of the factor expression level in the subject to the factor expression level in the non-responder, wherein the calculating comprises applying a machine learning algorithm trained with a training set including the received factor expression levels in responders and non-responders and the genders of the responders and non-responders, to each received factor expression level from the subject and gender of the subject, wherein the machine learning algorithm outputs the resistance score; c. summing the calculated resistance scores to generate a total resistance score; a subject having a total resistance score above a predetermined threshold is predicted to be resistant to said therapy; thereby predicting a subject's response to therapy. A method comprising:

4. 1. A method comprising: During the training phase, (i) the number of resistance-associated factors expressed in samples from subjects known to be afflicted with a disease and responsive to a therapy, and the number of resistance-associated factors expressed in samples from subjects known to be afflicted with said disease and unresponsive to said therapy; (ii) at least one clinical parameter of the subject known to be responsive and the subject known to be non-responsive; and (iii) a label associated with the responsiveness of the subject suffering from the disease. Train a machine learning algorithm on a training set including generating a trained machine learning algorithm, wherein the trained machine learning algorithm is trained to predict the responsiveness of a subject suffering from the disease to the therapy; method.

5. 1. A method comprising: During the training phase, (i) factor expression levels of resistance-associated factors in samples from subjects known to be afflicted with a disease and responsive to a therapy, and factor expression levels of resistance-associated factors in samples from subjects known to be afflicted with said disease and unresponsive to said therapy; (ii) at least one clinical parameter of the subject known to be responsive and the subject known to be non-responsive; and (iii) a label associated with the responsiveness of the subject suffering from the disease. Train a machine learning algorithm on a training set including generating a trained machine learning algorithm, wherein the trained machine learning algorithm is trained to predict activity of a resistance-associated factor in a subject; method.

6. The method of claim 4 or 5, wherein the number of resistance-associated factors and the at least one clinical parameter are labeled with the label.

7. 6. The method of claim 4 or 5, wherein a predetermined threshold for the final resistance score is 0.2, and a resistance score above 0.2 indicates that the subject is resistant to the therapy, or wherein the final resistance score is converted to a response score by the formula (1 - final resistance score), and a response score above a predetermined threshold indicates that the subject is responding to therapy, and optionally wherein the predetermined threshold for the response score is 0.

8.

8. The resistance-associated factors in each subject are a. i. in a population of subjects known to be afflicted with the disease and to respond to the therapy (responders), ii. in a population of subjects known to be suffering from the disease and who do not respond to the therapy (non-responders), and iii. In the subject receiving factor expression levels for a plurality of factors; b. calculating a resistance score for each factor of the plurality of proteins, the resistance score being based on the similarity of the factor expression level in each subject to the factor expression level in the responders and the similarity of the factor expression level in the subject to the factor expression level in the non-responders; c. classifying factors of said plurality of factors having a resistance score above a predetermined threshold as resistance-associated factors. The method of claim 4 , wherein the temperature is determined by a method comprising:

9. The resistance-associated factor in each subject is a. i. in a population of subjects known to be afflicted with the disease and to respond to the therapy (responders), ii. in a population of subjects known to be suffering from the disease and who do not respond to the therapy (non-responders), and iii. In the subject receiving factor expression levels for a plurality of factors; b. calculating a resistance score for each factor of the plurality of proteins, the resistance score being based on the similarity of the factor expression level in each subject to the factor expression level in the responders and the similarity of the factor expression level in the subject to the factor expression level in the non-responders; c. classifying factors of said plurality of factors having a resistance score above a predetermined threshold as resistance-associated factors. The method of claim 5 , wherein the temperature is determined by a method comprising:

10. 10. The method of any one of claims 1-3 and 8-9, further comprising, prior to (b), selecting a subset of the plurality of factors, the subset comprising factors that best distinguish between the responders and non-responders, and wherein said calculating is for each factor of the subset.

11. 11. The method of claim 10, wherein said selecting comprises applying a statistical test to said received factor expression levels, optionally wherein said statistical test is a Kolmogorov-Smirnov test.

12. 10. The method of any one of claims 1-3 and 8-9, wherein said calculating comprises applying a machine learning algorithm trained on a training set including the received factor expression levels in responders and non-responders to each received factor expression level from the subject, wherein the machine learning algorithm outputs the resistance score.

13. 13. The method of claim 12, wherein the training set further comprises the gender of each responder and non-responder, and the machine learning algorithm is applied to the individual factor expression levels received from the subjects and gender of the subjects.

14. The method of claim 12 , further comprising performing a dimensionality reduction step on the plurality of factors to reduce the number of the plurality of factors.

15. 15. The method of claim 14, wherein the dimensionality reduction step identifies a subset of key factors, and wherein the training set includes only expression levels of the subset of key factors, optionally the subset of key factors being those factors that most evenly balance the number of predicted responders and non-responders.

16. The method of claim 14 , wherein the predetermined threshold is determined by performing cross-validation within the training set.

17. 10. The method of any one of claims 1-3 and 8-9, wherein said calculating comprises calculating a mean expression for each factor in responders and non-responders, and wherein said resistance score is based on a ratio of deviation of said factor expression in said subject from the calculated mean in responders to deviation of said factor expression in said subject from the calculated mean in non-responders.

18. 18. The method of claim 17, wherein said calculating further comprises calculating a distribution and standard deviation for each factor in responders and non-responders, said deviations being measured as multiples of the calculated standard deviation.

19. The resistance score is a monotonic function of the formula [Equation 1] where Z R is the deviation of the factor expression in the subject from the calculated mean in responders, and Z NR 19. The method of claim 18, wherein x is the deviation of the factor expression in the subject from the calculated mean in non-responders, and c is a constant.

20. 20. The method of claim 19, wherein the predetermined threshold for the resistance score is about 2.9, and a resistance score above 2.9 indicates that the factor is a resistance-associated factor.

21. 10. The method of any one of claims 1 to 5 and 8 to 9, wherein the plurality of factors is at least a factor of 200.

22. 10. The method of any one of claims 1 to 3 and 8 to 9, wherein the predetermined number of resistance-associated factors is three, and a subject having more than three resistance-associated factors is predicted to be resistant to the therapy.

23. receiving factor expression levels for a plurality of factors; receiving factor expression levels for the greater than plurality of factors in the population of responders and the population of non-responders; b. for each factor in the group, applying a machine learning algorithm trained on the received factor expression levels in responders and non-responders; c. the algorithm selects a subgroup of factors that most evenly divides the subjects in the population into responders and non-responders; and d. designating a subgroup of said factors as said plurality of factors; The method according to any one of claims 1 to 3 and 8 to 9, comprising:

24. The method of any one of claims 1 to 3 and 8 to 9, wherein the factor expression level is the factor expression level in a biological sample provided by the subject.

25. 25. The method of claim 24, wherein the biological sample is selected from plasma, whole blood, serum, or peripheral blood mononuclear cells.

26. 26. The method of claim 25, wherein the biological sample is plasma.

27. 25. The method of claim 24, wherein the biological sample is provided by the subject prior to receiving the therapy.

28. 28. The method of claim 27, wherein before is up to 24 hours before.

29. 25. The method of claim 24, wherein the biological sample is provided by the subject after undergoing the therapy.

30. 30. The method of claim 29, wherein the after-treatment with the therapy is after a first treatment with the therapy.

31. 30. The method of claim 29, wherein after is at least 24 hours.

32. 25. The method of claim 24, wherein the biological sample provided by each subject in the population is the same type of biological sample.

33. 25. The method of claim 24, wherein the population of responders, the population of non-responders, and the biological samples provided by the subjects are all the same type of biological sample.

34. 10. The method of any one of claims 2 and 8-9, wherein the trained machine learning algorithm is trained by the method of claim 4.

35. The method of any one of claims 3 and 8-9, wherein the trained machine learning algorithm is trained by the method of claim 5.

36. The method according to any one of claims 1 to 5 and 8 to 9, wherein the disease is cancer.

37. 10. The method of any one of claims 1-5 and 8-9, wherein said therapy is immune checkpoint inhibition, and optionally said immune checkpoint inhibition inhibits the PD-1 / PD-L1 axis.

38. 10. The method of any one of claims 2 to 5 and 8 to 9, wherein the at least one clinical parameter is selected from subject age, sex, line of treatment, and expression of a biomarker in a sample from the subject.

39. 39. The method of claim 38, wherein the disease is cancer, the therapy is an anti-PD-1 or anti-PD-L1 therapy, and the target expression in the sample is PD-L1 expression in a tumor sample.

40. 6. The method of claim 4 or 5, further comprising, in the inference step, receiving as input the number of the resistance-associated factors expressed in samples from subjects suffering from the disease and with unknown responsiveness to the therapy and at least one clinical parameter of the subjects with unknown responsiveness, and applying the trained machine learning algorithm to the received input to predict the responsiveness of the subjects with unknown responsiveness to the therapy.