A system and method for predicting probability of clinical benefit of a treatment in a patient

Machine-learning models analyzing RAPs in biological samples predict treatment response variability in immunotherapy, enhancing personalized treatment decisions and clinical outcomes.

WO2026053217A1PCT designated stage Publication Date: 2026-03-12ONCOHOST LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-07
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing immunotherapy treatments for cancer exhibit wide variability in efficacy among patients, with many experiencing primary or subsequent resistance, necessitating a method to predict patient-specific response.

Method used

A method utilizing machine-learning based prediction models to analyze expression levels of Resistance-Associated Proteins (RAPs) in biological samples to determine clinical benefit scores, combining multiple models to calculate the probability of treatment response.

Benefits of technology

Enhances the prediction of treatment response by identifying differential protein expression patterns, enabling personalized treatment decisions and improving clinical outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050771_12032026_PF_FP_ABST
    Figure IL2025050771_12032026_PF_FP_ABST
Patent Text Reader

Abstract

A method and computer-based system for predicting probability of clinical benefit of a treatment in a target patient suffering from a first disease may process expression levels of proteins from a biological sample originating from the target patient. The system may select a first group of Resistance-Associated Proteins (RAPs) defined by differential expression between CB and Non-Clinical Benefit (NCB) patients of the first disease. The system may apply a first machine-learning based prediction model on the first group of RAPs to determine a first CB score. The system may identify a second group of RAPs having differential expression between CB and NCB patients of a second, different disease, and apply a second ML-based prediction model on the second group of RAPs to determine a second CB score. The system may calculate the CB probability based on the first CB score and the second CB score, as elaborated herein.
Need to check novelty before this filing date? Find Prior Art

Description

A SYSTEM AND METHOD FOR PREDICTING PROBABILITY OF CLINICAL BENEFIT OF A TREATMENT IN A PATIENTCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from Israeli Patent Application No. 315485, entitled “A SYSTEM AND METHOD FOR PREDICTING PROBABILITY OF CLINICAL BENEFIT OF A TREATMENT IN A PATIENT”, filed on September 5, 2024, the entire contents of which are incorporated herein by reference in their entirety.FIELD OF INVENTION

[0002] The present invention is in the field of cancer diagnostics.BACKGROUND OF THE INVENTION

[0003] Immunotherapy based on immune checkpoint inhibitors (ICIs) represents a significant breakthrough in clinical oncology. ICIs augment an anti-tumor immune response by targeting checkpoint proteins such as PD-1, PD-L1 and CTLA-4 expressed on tumor and immune cells. Although ICI therapies can achieve unprecedented long-term disease control across multiple tumor types, efficacy varies widely between patients, with the majority exhibiting primary or subsequent acquired resistance to therapy. Indeed, this patient variability is true for all anticancer therapies and truly for all therapeutic agents regardless of the disease. A new method of determining patient-specific response to therapy is greatly needed.SUMMARY OF THE INVENTION

[0004] The present invention provides methods of predicting probability of clinical benefit (CB) of a treatment in a target patient suffering from a first disease. Systems for performing method of the invention are also provided.

[0005] According to a first aspect, there is provided a method of predicting probability of clinical benefit (CB) of a treatment in a target patient of a first disease, the method comprising: receiving expression levels of proteins from a biological sample originating from the target patient; based on the received expression levels, determining expression levels of a first group of Resistance- Associated Proteins (RAPs) in the target patient, wherein the first group of RAPs are defined by differential levels of expression between CB and non-clinical benefit (NCB) patients of the first disease, in relation to the treatment; applying, or inferring a first, pretrained machine-learning (ML) based prediction model on the expression levels of the first group of RAPs in the target patient, to determine a first CB score; identifying at least one second group of RAPs, as having differential levels of expression between CB and NCB patients of at least one respective second disease, in relation to the treatment; based on the received expression levels, determining expression levels of a subset of the second group of RAPs in the target patient; applying, or inferring a second, pretrained ML-based prediction model on the expression levels of the subset of the second group of RAPs in the target patient, to determine at least one respective second CB score; and calculating the CB probability based on the first CB score and the at least one second CB score.

[0006] Embodiments of the method may include: obtaining a first dataset, comprising (i) expression levels of proteins in a cohort of patients of the first disease, and (ii) at least one annotation, labeling at least one patient of the cohort of patients as either a CB patient or a NCB patient; and applying a statistical test on the expression levels of proteins in the first dataset, to identify the first group of RAPs.

[0007] Additionally, or alternatively, embodiments of the method may include constructing the first prediction model as an ensemble of a plurality of decision trees such as an XGBoost model, a random forest, and the like. For each decision tree of the plurality of decision trees, embodiments may use at least one annotation as supervisory data, to train that decision tree so as to predict an interim CB score, based on expression of a respective, unique protein of the first group of RAPs. Embodiments of the invention may configure the ensemble of decision trees to calculate the first CB score based on the interim CB scores of the plurality of decision trees.

[0008] According to some embodiments, the method further comprises: obtaining a second dataset, comprising (i) expression levels of proteins in a second cohort of patients of the second disease, and (ii) at least one annotation, labeling at least one patient of the second cohort of patients as either a CB patient or an NCB patient; applying a statistical calculation on the expression levels of proteins in the second dataset, to identify the at least one second group of RAPs.

[0009] According to some embodiments, the method further comprises: calculating a correlation matrix, representing correlation of expression between RAPs of the second group; based on the correlation matrix, calculating a graph data element, representing a strength of correlation between pairs of RAPs of the second group; applying, or inferring agraph analysis algorithm on the graph data element, to collect the RAPs of the second group into a plurality of clusters; selecting a subset of the plurality of clusters, based on a number of RAPs within each cluster; and in each cluster of the subset of clusters, identifying a hub RAP as one whose expression is most correlated to other RAPs of that cluster, thereby selecting the subset of the second group of RAPs.

[0010] According to some embodiments, the subset of the at least one second group of RAPs does not comprise RAPs from the first group of RAPs.

[0011] According to some embodiments, the first disease and second disease are a first and a second type of cancer.

[0012] According to some embodiments, the first disease and second disease afflict the same tissue or cell type in the subject.

[0013] According to some embodiments, the first type of cancer and the second type of cancer are cancers of the same tissue.

[0014] According to some embodiments, the clinical benefit is overall survival.

[0015] According to some embodiments, the clinical benefit is progression free survival.

[0016] According to some embodiments, the treatment is an immunotherapy.

[0017] According to some embodiments, the statistical test is a Kolmogorov-Smirnov test.

[0018] According to some embodiments, the biological sample originating from the target patient is obtained before the target patient received the treatment.

[0019] According to some embodiments, the biological sample is selected from blood plasma, whole blood, blood serum or peripheral blood mononuclear cells.

[0020] According to some embodiments, the biological sample is blood plasma.

[0021] According to another aspect, there is provided a system for predicting probability of CB of a treatment in a target patient of a first disease, the system comprising: a non-transitory memory device, wherein modules of instruction code are stored, and at least one processor associated with the memory device, and configured to execute the modules of instruction code, whereupon execution of the modules of instruction code, the at least one processor is configured to: receiving expression levels of proteins from a biological sample originating from the target patient; based on the received expression levels, determine expression levels of a first group of Resistance- Associated Proteins (RAPs) in the target patient, wherein the first group of RAPs are defined by differential levels of expression between CB and non- clinical benefit (NCB) patients of the first disease, in relation to the treatment; infer a first, pretrained machine-learning (ML) based prediction model on the expression levels of the first group of RAPs in the target patient, to determine a first CB score; identify at least onesecond group of RAPs, as having differential levels of expression between CB and NCB patients of a second disease, in relation to the treatment; based on the received expression levels, determine expression levels of a subset of the second group of RAPs in the target patient; infer a second, pretrained ML-based prediction model on the expression levels of the subset of the second group of RAPs in the target patient, to determine a second CB score; and calculate the CB probability based on the first CB score and the second CB score.

[0022] According to some embodiments, the expression levels are received from a hardware or software module.

[0023] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description given hereinafter. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figs. 1A-1F: Main clinical parameters of the SCLC cohort. 1A. Kaplan Meier plot for the overall survival rate for patients who displayed clinical benefit (CB) or no clinical benefit (NCB). CB was defined based on progression-free survival (PFS) at 6 months. mOS, median overall survival. HR, hazard ration. CI, confidence interval. IB. Distribution of patient age in the cohort. 1C. Proportion of males and females in the cohort. ID. ECOG performance status (ECOG PS). IE. ICI-based treatment. IF. Proportion of CB and NCB patients based on 6 months PFS.

[0025] Fig. 2 : Illustration of the combined model. Model 1 is based on the Resistance- Associated Proteins approach followed by XGBoost machine learning model (Christopoulos et al., JCO PO 2024) applied on the SCLC cohort. Model 2 is based on generation of a correlation matrix between 372 proteins and the identification of hub proteins in the 4 largest clusters, followed by computational model (Cox regression) training on the NSCLC cohort applied on the SCLC cohort. Each model output, the clinical benefit score per patient, is then averaged, and a linear transformation is applied to reach the clinical probability (PROphet) score, a scale between 0 and 10, where 6.7 serves as the threshold between patients with PROphet-POSITIVE result and PROphet-NEGATIVE result.

[0026] Fig. 3 : Illustration of model 2 development.

[0027] Fig. 4 : Network of the 372 proteins. Each color indicates a different cluster. The network was generated using Fruchterman-Reingold force-directed algorithm. The intensity of the edge is correlated with the weight.

[0028] Figs. 5A-5C: Model performance of model 1 (the Resistance-Associated Protein model, or RAP model, on SCLC cohort; 5A), model 2 (the 4-protein model on SCLC cohort, which was developed on the NSCLC cohort; 5B), and the combined model (average of the two models’ outputs; 5C). The left figure in each panel is the Receiver Operating Characteristics (ROC) curve demonstrating the ability of the model to distinguish between CB vs NCB in the SCLC cohort. The area under the curve (AUC) is indicated in each plot. The middle figure in each panel describes the fit between the predicted clinical benefit probability (i.e., each model output) and the observed clinical benefit rate (i.e., the rate of patients who are defined as CB patients among a window of ±0.15 CB probability). The right figure in each panel is a Kaplan Meier plot describing the overall survival of patients with PROphet-NEGATIVE vs. PROphet-POSITIVE result. PROphet-POSITIVE and PROphet- NEGATIVE were defined based on division of the SCLC cohort into 1 / 3 of the cohort with the highest CB probability (PROphet POSITIVE) and 2 / 3 with the lowest CB probability (PROphet NEGATIVE).

[0029] Fig. 6 : The correlation between model 1 and model 2 output. Each dot represents a patient from the SCLC cohort. The colors indicate the observed clinical benefit based on Progression-Free Survival (PFS) at 6 months.

[0030] Figs. 7A-7B: Relevance of the PROphetNSCLC model to SCLC. 7A. Kaplan Meier analysis of the PROphetNSCLC model applied on the SCLC cohort. The PROphetNSCLC model was developed and validated on a cohort of 500 NSCLC patients as previously described (Christopulos et el., JCO PO 2024). 7B. Venn diagram showing NSCLC and SCLC RAPs. There is a minimal overlap between PROphetNSCLC RAPs and PROphetSCLC RAPs, suggesting that different mechanisms drive ICI resistance and clinical outcomes in the two cancer types.

[0031] Fig. 8 : Potential source of the SCLC RAPs. RAPs were categorized based on published data from CPTAC and HPA databases. The diagram presents the number of RAPs in each category (horizontal bars) and the number of RAPs unique to each category or common to more than one category (vertical bars).

[0032] Figs. 9A-9B: Functional analysis of the SCLC RAPs. 9A. Selected categories significantly enriched among the SCLC RAPs in enrichment analysis (Fisher exact test, FDR <0.1). The SCLC RAPs are significantly enriched with multiple processes related to cancerprogression or lung cancer. 9B. Protein-protein interaction map showing the significantly enriched categories.

[0033] Fig. 10: Proteomaps analysis of SCLC RAPs. The Voronoi plot displays the main biological processes in which RAPs may be involved. Each polygon represents a RAP. Polygon size correlates with RAP strength (the number of iterations in which the protein was selected as a RAP during model development). RAPs are categorized into 6 main categories based on the KEGG database.

[0034] Fig. 11: A block diagram depicting a computing device, which may be included within an embodiment of a system for predicting probability of CB, according to some embodiments of the invention.

[0035] Fig. 12: A block diagram depicting an exemplary implementation of a system for predicting probability of CB in a target patient.

[0036] Figs. 13A-13B: 13A. A block diagram, depicting an exemplary implementation of a training stage of a first analysis module that may be included in a system for predicting CB in a target patient, according to some embodiments of the invention. 13B. A block diagram, depicting an exemplary implementation of a training stage of a second analysis module that may be included in a system for predicting CB in a target patient, according to some embodiments of the invention.

[0037] Fig. 14: A flow diagram depicting a method of predicting probability of CB of treatment in a target patient by at least one processor, according to some embodiments of the invention.DETAILED DESCRIPTION OF THE INVENTION

[0038] The present invention, in some embodiments, provides methods and systems for predicting the probability of clinical benefit to a treatment in a subject.

[0039] By a first aspect, there is provided a method of predicting response of a subject suffering from a first disease to a treatment, the method comprising: receiving expression levels of proteins from a biological sample from the subject; based on the received expression levels, determining expression levels of a first group of resistance-associated proteins (RAPs) in the subject; applying, or inferring a first, predetermined machine-learning (ML) based prediction model on the expression levels of the first group of RAPs in the subject, to determine a first response score; identifying at least one second group of RAPs from a second disease; based on the received expression levels, determining expressionlevels of a subject of the at least one second group of RAPs in the subject; applying, or inferring a second, pretrained ML based prediction model on the expression levels of a subset of the second group of RAPs in the subject, to determine at least one respective second response score corresponding to the at least one second disease; and calculating the response probability based on the first response score and the second response score.

[0040] In some embodiments, the method is for determining if a subject is a responder to the therapy. In some embodiments, the method is for determining if a subject is a nonresponder to the therapy. In some embodiments, the method is for predicting a subject’s response to therapy. In some embodiments, the method is for monitoring response to the therapy. In some embodiments, the method is for determining if the therapy should continue or be adjusted (e.g., by further treating the subject with an additional therapy including but not limited to an agent determined by the RAP analysis provided hereinbelow). In some embodiments, the method is for determining a subject as being a responder to the therapy, or a non-responder to the therapy. In some embodiments, the method is for determining a subject as being a responder to the therapy, a non-responder to the therapy, or as having a stable diseased state. In some embodiments, the method is for predicting if a subject will respond to the therapy, or not respond to the therapy. In some embodiments, the method is a prognostic method. In some embodiments, the method is a predictive method. In some embodiments, the method is a diagnostic method. In some embodiments, the method is a prognostic method. In some embodiments, the method is an in vitro method.

[0041] In some embodiments, non-response comprises progressive disease. In some embodiments, non-response comprises cancer progression. In some embodiments, nonresponse comprises stable disease. In some embodiments, non-response comprises a worsening of symptoms of the disease. In some embodiments, non-response comprises no improvement in disease symptoms in response to the therapy. In some embodiments, nonresponse is not the development of side effects. In some embodiments, non-response comprises growth, metastasis and / or continued proliferation of a cancer. In some embodiments, response is stable disease. In some embodiments, response comprises remission. In some embodiments, remission is minimal remission. In some embodiments, remission is partial remission. In some embodiments, remission is complete remission. In some embodiments, response is measured using the overall response rate (ORR). A trained physician will be familiar with methods of determining response and specifically the ORR. In some embodiments, response is measured using Response Evaluation Criteria In Solid Tumors (RECIST). In some embodiments, response comprises survival. In someembodiments, survival is overall survival. In some embodiments, survival is progression free survival. In some embodiments, response comprises a durable clinical benefit (DCB). In some embodiments, progression is hyper-progression. In some embodiments, progression is pseudo-progression. In some embodiments, non-response is hyper-progression. In some embodiments, a non-responder is hyper-progressor. In some embodiments, response is pseudo-progression. In some embodiments, a responder is pseudo-progressor.

[0042] In some embodiments, response comprises improvement of symptoms of the disease. In some embodiments, response comprises improvement in disease symptoms in response to the therapy. In some embodiments, non-response comprises progressive disease. In some embodiments, non-response comprises cancer progression. In some embodiments, nonresponse comprises stable disease. In some embodiments, non-response comprises a worsening of symptoms of the disease. In some embodiments, non-response is not the development of side effects. In some embodiments, non-response comprises growth, metastasis and / or continued proliferation of a cancer. In some embodiments, non-response comprises no clinical benefit (NCB). In some embodiments, non-response is non-survival. In some embodiments, non-response is non-survival and / or cancer progression. In some embodiments, response is stable disease. In some embodiments, response comprises remission. In some embodiments, remission is minimal remission. In some embodiments, remission is partial remission. In some embodiments, remission is complete remission. In some embodiments, response is survival. In some embodiments, response is progression free survival. In some embodiments, response is long progression free survival. In some embodiments, response is measured using the overall response rate (ORR). A trained physician will be familiar with methods of determining response and specifically the ORR. In some embodiments, response is measured using Response Evaluation Criteria In Solid Tumors (RECIST). In some embodiments, response is measured using Immune Response Evaluation Criteria in Solid Tumors (iRECIST). In some embodiments, response comprises survival. In some embodiments, survival is overall survival. In some embodiments, survival is progression free survival. In some embodiments, survival is overall survival. In some embodiments, response comprises a clinical benefit (CB). In some embodiments, response comprises a durable clinical benefit (DCB). In some embodiments, response comprises a disease-free survival (DFS). In some embodiments, CB is DCB. In some embodiments, CB is PFS. In some embodiments, CB is PFS at 12 months after the commencement of treatment. In some embodiments, CB is PFS at 7 months after the commencement of treatment. In someembodiments, CB is PFS at 6 months after the commencement of treatment. In some embodiments, the population of subject known to respond and known not to respond are determined based on PFS and the predicted response comprises OS. In some embodiments, PFS is PFS at 12 months. In some embodiments, PFS is PFS at 7 months. In some embodiments, PFS is PFS at 6 months. In some embodiments, PFS is PFS at 3 months. In some embodiments, OS is OS at 12 months. In some embodiments, OS is OS at 7 months. In some embodiments, OS is OS at 6 months. In some embodiments, OS is OS at 3 months. In some embodiments, no clinical benefit or non-clinical benefit is the absence of a clinical benefit described herein.

[0043] In some embodiments, response is clinical benefit (CB). In some embodiments, a responder is a subject that would have clinical benefit (CB) from receiving the treatment. In some embodiments, a non-responder would not have a clinical benefit (NCB) from receiving the treatment. In some embodiments, the method is a method of predicting CB. In some embodiments, the first response score is a first CB score. In some embodiments, the second response score is a second CB score. In some embodiments, the response probability is CB probability. In some embodiments, the method is a method of predicting CB probability. In some embodiments, the response probability is a total response probability. In some embodiments, the response probability is a combined response probability. In some embodiments, the response probability is a response score. In some embodiments, the response probability is a CB score. In some embodiments, the response probability is a total CB score. In some embodiments, the response probability is a combined CB score. In some embodiments, the CB probability is a CB score. In some embodiments the CB probability is determined using survival analysis. In some embodiments the CB probability is determined using regression. In some embodiments the CB probability is determined using Cox regression analysis. In some embodiments the CB probability is determined using log-rank test. In some embodiments, the CB probability is determined based on progression free survival. In some embodiments the survival analysis is based on overall survival. In some embodiments the survival analysis is based on progression free survival. In some embodiments the survival analysis is based on durable clinical benefit. In some embodiments the survival analysis is based on disease free survival.

[0044] In some embodiments, the subject is a mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a patient. In some embodiments, the patient is a target patient. In some embodiments, the subject suffers from a disease. In someembodiments, the disease is treatable by the treatment. In some embodiments, the disease is cancer. In some embodiments, the disease is treatable by an immune checkpoint inhibitor (ICI). In some embodiments, the disease is treatable by a combination of ICI and additional agent. In some embodiments, the cancer is a PD-L1 positive cancer. In some embodiments, the cancer is a PD-L1 negative cancer. In some embodiments, the cancer is solid cancer. In some embodiments, the cancer is a tumor. In some embodiments, the cancer is selected from hepato-biliary cancer, cervical cancer, urogenital cancer (e.g., urothelial cancer), testicular cancer, prostate cancer, thyroid cancer, ovarian cancer, nervous system cancer, ocular cancer, lung cancer, soft tissue cancer, bone cancer, pancreatic cancer, bladder cancer, skin cancer, intestinal cancer, hepatic cancer, rectal cancer, colorectal cancer, esophageal cancer, gastric cancer, gastroesophageal cancer, breast cancer (e.g., triple negative breast cancer), renal cancer (e.g., renal carcinoma), skin cancer, head and neck cancer, leukemia and lymphoma. In some embodiments, the cancer is selected from skin cancer, and lung cancer. In some embodiments, the cancer is skin cancer. In some embodiments, the cancer is lung cancer. In some embodiments, the skin cancer is melanoma. In some embodiments, the lung cancer is small cell lung cancer. In some embodiments, the lung cancer is non-small cell lung cancer. In some embodiments, the subject is naive to therapy before the first determining. In some embodiments, the subject has not received the therapy before the first determining. In some embodiments, the subject has received the therapy previously. In some embodiments, the subject has previously been treated by a therapy other than the therapy. In some embodiments, the subject is naive to any therapy. In some embodiments, the subject is naive to immunotherapy. In some embodiments, the therapy is the first line of treatment. In some embodiments, the therapy is an advanced line of treatment. In some embodiments, determining the expression levels of proteins is before a first treatment with the treatment. In some embodiments, the expression levels are from a subject naive to the treatment. In some embodiments, the expression levels are from a subject before receiving the treatment. In some embodiments, determining the expression levels of proteins is after a first treatment with the treatment. In some embodiments, the expression levels are from a subject after receiving the treatment.

[0045] In some embodiments, the first disease and the second disease are different diseases. In some embodiments, the first disease and the second disease are both cancer. In some embodiments, the first disease and the second disease are a first type and a second type of cancer. In some embodiments, the first disease and the second disease afflict the same tissue. In some embodiments, the first disease and the second disease afflict the same cell type. Insome embodiments, the first type of cancer and the second type of cancer are cancers of the same tissue or cell type. In some embodiments, the first disease and the second disease are the same disease. In some embodiments, the first disease and the second disease are treated by the same treatment. In some embodiments, the first disease and the second disease are treated by similar treatments. In some embodiments, the first disease and the second disease are treated by different treatments.

[0046] As used herein, the terms “responder”, CB, or a subject “known to respond” are used interchangeably and refer to a subject that when administered a therapy displays an improvement in at least one criteria of the disease being treated by the therapy or does not show an increase in severity of the disease. In some embodiments, a responder is a subject that when administered a therapy displays an improvement in the disease that is being treated by the therapy. In some embodiments, a responder is a subject that when administered a therapy does not show an increase in severity of the disease. In some embodiments, an increase is in severity is over time. In some embodiments, does not show an increase in severity is stable disease. In some embodiments, a responder is a subject that when administered a therapy show mixed response. In some embodiments, a responder is a subject that when administered a therapy show mixed response, wherein mixed response is improvement in at least one criteria of the disease but does not show an improvement in other criteria of the disease. In some embodiments, mixed response is shrinkage of some lesions in combination with growth of new or existing lesions. In some embodiments, a responder is a subject for which the therapy produces an anti-disease response. In some embodiments, for a subject with cancer, a responder is a subject in which the therapy produces an anticancer response. In some embodiments, a response is not a reduction in side effects. In some embodiments, a response is a reduction in side effects. In some embodiments, a response is a response against the disease itself. In some embodiments, an anticancer response is an antitumor response. In some embodiments, an antitumor response comprises tumor regression. In some embodiments, an antitumor response comprises tumor shrinkage. In some embodiments, an antitumor response comprises a lack of tumor growth. In some embodiments, an antitumor response comprises a lack of tumor metastasis. In some embodiments, an antitumor response comprises a lack of tumor hyperproliferation. In some embodiments, an improvement is in at least one symptom of the disease. In some embodiments, response is complete response. In some embodiments, response is minimal response. In some embodiments, response is partial response. In some embodiments, response comprises stable disease. In some embodiments, responder is a subject with afavorable response to the therapy. In some embodiments, non-responder is a subject with a non-favorable response to the therapy. In some embodiments, a non-favorable response is an increase in tumor burden. Increases in tumor burden can encompass any increase in tumor size or total cancer cell number such as increase in tumor size, increase in tumor spread, increase in metastasis, increase in tumor cell proliferation or any other increase. In some embodiments, a responder is pseudo-progressor. In some embodiments, a pseudo-progressor is a subject that when treated with the therapy shows an apparent increase in tumor size or new lesions after starting the treatment followed by tumor shrinkage or stabilization. In some embodiments, a non-responder is hyper-progressor. In some embodiments, a hyper- progressor is a subject that when treated with the therapy shows an unexpected rapid acceleration of tumor growth.

[0047] As used herein, a “favorable response” or “clinical benefit” of the cancer patient indicates “responsiveness” of the cancer patient to the treatment with the therapy, namely, the treatment of the responsive cancer patient with the therapy will lead to the desired clinical outcome such as tumor regression, tumor shrinkage or tumor necrosis; reduction in tumor burden; an anti-tumor response by the immune system; preventing or delaying tumor recurrence, tumor growth or tumor metastasis; improvement in disease symptoms and clinical status. In some embodiments, the subject is complete responder or treatment with the cancer therapy leads to stable disease. In some embodiments, a complete responder is a subject in which there is an absence of detectable cancer after treatment with the therapy. In this case, it is possible and advised to continue the treatment of the responsive cancer patient with the therapy or if the patient is cancer free to discontinue treatment. In some embodiments, the method further comprises continuing to administer the therapy to a subject that is not a non-responder. In some embodiments, the subject is non-responder, and the method further comprises continuing to administer the therapy to the subject. In some embodiments, the subject is non-responder, a minimal responder, partial responder or has a stable disease, and the method further comprises continuing to administer the therapy to a subject, as well as treating the subject with an additional therapy (e.g., determined using the resistance associated protein (RAP) analysis provided herein) to increase responsiveness. In some embodiments, a subject that is not a non-responder is a responder. In some embodiments, CB is overall survival. In some embodiments, CB comprises overall survival. In some embodiments, CB is progression free survival. In some embodiments, CB comprises progression free survival. In some embodiments, a subject that is non-responder is an NCBsubject. In some embodiments, clinical benefit is or comprises progression-free survival for a predetermined amount of time. In some embodiments, the predetermined about of time is 6 months. In some embodiments, the predetermined about of time is 9 months. In some embodiments, the predetermined about of time is 12 months.

[0048] As used herein, the term “non-responder”, NCB, and a subject “known to not respond” are used interchangeably and refer to a subject that when administered a therapy displays no improvement in disease symptoms or stabilization in disease. In some embodiments, a non-responder is a subject that when administered a therapy displays no improvement in clinical status. In some embodiments, a non-responder displays a worsening of disease when administered a therapy. In some embodiments, a non-responder displays a worsening of disease symptoms when administered a therapy. In some embodiments, non- responder is not a subject that experiences a side effect of the therapy. In some embodiments, a non-responder is a subject in which the disease progresses. In some embodiments, a non- responder is a subject in which the disease does not stabilize after therapy. In some embodiments, a non-responder is a subject in which the disease does not improve after therapy. In some embodiments, a non-responder is a subject that is not a responder as defined hereinabove. In some embodiments, a non-responder is a subject with a non-favorable response to the therapy. In some embodiments, a non-responder is a subject resistant to the therapy. In some embodiments, a non-responder is a subject refractory to the therapy.

[0049] As used herein, the terms “non-favorable response” or NCB of the cancer patient may indicate “non-responsiveness” of the cancer patient to the treatment with the therapy and thus the treatment of the non-responsive cancer patient with the therapy will not lead to the desired clinical outcome, and potentially to a non-desired outcomes such as tumor expansion, recurrence, or metastases. In some embodiments, a non-favorable response is resistance to the therapy. In some embodiments, a non-favorable response is cancer progress, growth and / or metastasis. In some embodiments, the method further comprises discontinuing administration of the therapy to a subject that is a non-responder. In some embodiments the method further comprises continuing to administer the therapy to a subj ect, in combination with an additional therapy. In some embodiments, the additional therapy increases responsiveness of a non-responsive patient.

[0050] In some embodiments, the method is for determining whether the response is considered a durable response. In some embodiments, durable response is response to treatment that lasts for a prespecified duration (e.g., a progression-free survival of more than 6 months).

[0051] In some embodiments, the treatment is a therapy. In some embodiments, the method further comprises administering the therapy to the subject predicted to respond to the therapy. In some embodiments, the method further comprises continuing to administering the therapy to the subject predicted to respond to the therapy. In some embodiments, the method further comprises not administering the therapy to the subject predicted to not respond to the therapy. In some embodiments, the method further comprises discontinuing the therapy to the subject predicted to not respond to the therapy. In some embodiments, the method further comprises administering an alternative therapy to the subject predicted to be a non-responder. In some embodiments, the alternative therapy is an additional therapy. In some embodiments, alternative therapy is a different therapy. In some embodiments, alternative therapy is a therapy with a different mechanism of action. In some embodiments, the method further comprises administering the therapy or continuing to administer the therapy in combination with an agent or therapy that blocks or inhibits at least one of the resistance-associated factors in the subject predicted to be resistant to the therapy. In some embodiments, resistance-associated factors are resistance-associated proteins. In some embodiments, resistance associated proteins are resistance-associated factors. In some embodiments, an agent or therapy that blocks or inhibits at least one of the resistance- associated factors is an additional therapy. In some embodiments, the combination therapy is administered to a subject predicted to be a non-responder or NCB.

[0052] In some embodiments, the therapy is an anticancer therapy. In some embodiments, the anticancer therapy is radiation. In some embodiments, the anticancer therapy is chemotherapy. In some embodiments, the therapy is immunotherapy. In some embodiments, the anticancer therapy is immunotherapy. In some embodiments, the anticancer therapy is targeted therapy. In some embodiments, the anticancer therapy is selected from radiation, chemotherapy, immunotherapy, targeted therapy, hormonal therapy, anti-angiogenic therapy and photodynamic therapy, thermotherapy, surgery, T-cell redirecting therapies, antibodydrug conjugates, radiotherapy, cellular therapies, oncolytic viruses, bi-specific antibodies, epigenetic therapy, chemoradiotherapy, and a combination thereof. In some embodiments, the immunotherapy is selected from immune checkpoint inhibition, immune checkpoint modulation, immune checkpoint blockade, adoptive-cell transfer therapy, oncolytic virus therapy, vaccine therapy, immune system modulation and therapy using monoclonal antibodies. In some embodiments, an immunotherapy is selected from immune checkpoint inhibitors, immune checkpoint modulators, immune checkpoint blockers, adoptive-cell transfer therapy, oncolytic virus therapy, treatment vaccines, immune system modulatorsand monoclonal antibodies. In some embodiments, the immunotherapy is an immune checkpoint inhibitor. In some embodiments, the immunotherapy is immune checkpoint blockade. In some embodiments, an immunotherapy is administered in combination with one or more conventional cancer therapy including chemotherapy, targeted therapy, steroids, and radiotherapy. Combinations of ICI and chemotherapy / radiotherapy / targeted therapy have been studied in multiple clinical trials. It will be understood by a skilled artisan that the predictive proteins disclosed herein are predictive in immunotherapy as a monotherapy, as well as part of a combination therapy.

[0053] In some embodiments, the treatment is a chemotherapy. In some embodiments, the treatment is a monotherapy. In some embodiments, the monotherapy is chemotherapy. In some embodiments, the monotherapy is immunotherapy. In some embodiments, the monotherapy is targeted therapy. Examples of chemotherapies include, but are not limited to Carboplatin, etoposide, cisplatin, Gemcitabine, pemetrexed, paclitaxel, nab-paclitaxel, Altretamine, Bendamustine, Busulfan, Chlorambucil, Cyclophosphamide, Dacarbazine, Ifosfamide, Meehl orethamine, Melphalan, Oxaliplatin, Procarbazine, Temozolomide, Thiotepa, vinorelbine and Trabectedin. In some embodiments, the chemotherapy is an alkylating agent. In some embodiments, the chemotherapy is an antimetabolite. In some embodiments, the chemotherapy is a topoisomerase inhibitor. In some embodiments, the chemotherapy is a mitotic inhibitor. In some embodiments, the chemotherapy is a platinumbased chemotherapy.

[0054] In some embodiments, the treatment is targeted therapy. In some embodiments, the targeted therapy is a tyrosine kinase inhibitor (TKI). In some embodiments, the targeted therapy is Bevacizumab. In some embodiments, the targeted therapy is TKI, Bevacizumab or both. In some embodiments, the subject has been previously treated with a TKI. In some embodiments, the subject has not been previously treated with a TKI. In some embodiments, the TKI is selected from Osimertinib, Erlotinib, Afatinib, Gefitinib, Dacomitinib, dacomitinib, Amivantamab-vmjw, Mobocertinib, Sotorasib, Adagrasib, Alectinib, Brigatinib, Lorlatinib, Ceritinib, Crizotinib, entrectinib, Dabrafenib, ceritinib, trametinib, Vemurafenib, Tepotinib, Pertuzumab, Lapatinib, Palbociclib, Ribociclib, Capmatinib, Selpercatinib, Olaparib, Rucaparib, Niraparib, Pralsetinib, Fam-trastuzumab, deruxtecan- nxki, Trastuzumab Deruxtecan, Elahere, Ado-trastuzumab, emtansine, Cabozantinib, Ado- trastuzumab emtansine, Larotrectinib, alectinib, Cetuximab, cobimetinib, Encorafenib, binimetinib, Lenvatinib, imatinib, dasatinib, nilotinib, sunitinib, pazopanib, axitinib, sorafenib, Belzutifan, tivozinib, Regorafenib, Everolimus, Alpelisib, aflibercept,Ramucirumab, temsirolimus, Bortezomib, Carfilzomib, Ixazomib, Daratumumab, Elotuzumab, Panobinostat, Abiraterone, Enzalutamide, Vandetanib, Avapritinib, Vismodegib, Sonidegib, and Ripretinib. In some embodiments the targeted therapy is for blocking the activity of kinases, growth factor receptors, angiogenesis drivers, cell surface antigens, or DNA repair enzymes. In some embodiments, the targeted therapy is PARP inhibitor. In some embodiments, the targeted therapy is ATR inhibitors. In some embodiments, the targeted therapy is CHK1 and WEE1 inhibitors. In some embodiments, the targeted therapy is DNA-PK inhibitors. In some embodiments, the targeted therapy is EGFR inhibitor. In some embodiments, the targeted therapy is ALK inhibitor. In some embodiments, the targeted therapy is VEGF pathway inhibitor. In some embodiments, the targeted therapy is BRAF inhibitor. In some embodiments, the targeted therapy is MEK inhibitor. In some embodiments, the targeted therapy is mTOR inhibitor. In some embodiments, the targeted therapy is proteasome histone deacetylase inhibitor. In some embodiments, the targeted therapy is hormonal pathway inhibitor. In some embodiments, the targeted therapy is antibody-drug conjugate (ADC). In some embodiments, the ADC or bi-specific antibodies are DLL3 inhibitors. In some embodiments, the DLL3 inhibitor is tarlatamab. In some embodiments, that targeted therapy is BCL2 inhibitor. In some embodiments, the therapy is DNA repair therapy.

[0055] In some embodiments, the immunotherapy is a plurality of immunotherapies. In some embodiments, the immunotherapy is immune checkpoint blockade. In some embodiments, the immunotherapy is immune checkpoint protein inhibition. In some embodiments, the immunotherapy is immune checkpoint protein modulation. In some embodiments, the immunotherapy comprises immune checkpoint inhibition. In some embodiments, the immunotherapy comprises immune checkpoint modulation. In some embodiments, immune checkpoint blockade and / or immune checkpoint inhibition comprises administering to the subject an immune checkpoint inhibitor. In some embodiments, inhibition comprises administering an immune checkpoint inhibitor. In some embodiments, the inhibitor is a blocking antibody. In some embodiments, the immunotherapy comprises immune checkpoint blockade. In some embodiments, modulation comprises administering an immune checkpoint modulator. In some embodiments, immune checkpoint modulation comprises administering to the subject an immune checkpoint modulator.

[0056] As used herein, the term “immune checkpoint inhibitor (ICI)” refers to a single ICI, a combination of ICIs and a combination of an ICI with another cancer therapy. The ICI may be a monoclonal antibody, a dual-specific antibody, a humanized antibody, a fully humanantibody, a fusion protein, or a combination thereof directed to blocking, inhibition or modulation of immune checkpoint proteins. In some embodiments, an immune checkpoint inhibitor is an immune checkpoint modulator. In some embodiments, an immune checkpoint inhibitor is an immune checkpoint blocker. In some embodiments, the immune checkpoint protein is selected from PD-1 (Programmed Death-1); PD-L1 (Programmed Death-Ligand 1); PD-L2 (Programmed Death-Ligand 2); CTLA-4 (Cytotoxic T-Lymphocyte- Associated protein 4); A2AR (Adenosine A2A receptor), also known as AD0RA2A; B7-H3, also called CD276; B7-H4, also called VTCN1 (V-set domain-containing T cell activation inhibitor 1); B7-H5; BTLA (B and T Lymphocyte Attenuator), also called CD272; IDO (Indoleamine 2,3 -dioxygenase); KIR (Killer-cell Immunoglobulin-like Receptor); LAG-3 (Lymphocyte Activation Gene-3); TDO (Tryptophan 2,3 -dioxygenase); TIM-3 (T-cell Immunoglobulin domain and Mucin domain 3); VISTA (V-domain Ig suppressor of T cell activation); N0X2 (nicotinamide adenine dinucleotide phosphate NADPH oxidase isoform 2); SIGLEC7 (Sialic acid-binding immunoglobulin-type lectin 7), also called CD328; SIGLEC9 (Sialic acid-binding immunoglobulin-type lectin 9), also called CD329; 0X40 (Tumor necrosis factor receptor superfamily, member 4) also called CD 134; and TIGIT. In some embodiments, the immune checkpoint protein is selected from PD-1, PD-L1 and PD-L2. In some embodiments, the immune checkpoint protein is selected from PD-1 and PD-L1. In some embodiments, the immune checkpoint protein is CTLA-4. In some embodiments, the immune checkpoint protein is PD-1. In some embodiments, immune checkpoint blockade comprises an anti-PD-l / PD-Ll / PD-L2 immunotherapy. In some embodiments, immune checkpoint blockade comprises an anti-PD-1 immunotherapy. In some embodiments, immune checkpoint blockade comprises an anti-PD-1 and / or anti-PD-Ll immunotherapy. In some embodiments, immune checkpoint blockade comprises an anti CTLA-4 immunotherapy. In some embodiments, immune checkpoint blockade comprises an anti-PD- 1 and / or anti-PD-Ll immunotherapy and an anti CTLA-4 immunotherapy.

[0057] In some embodiments, the immunotherapy is a blocking antibody. In some embodiments, the immunotherapy is administration of a blocking antibody to the subject.

[0058] In some embodiments, the ICI is a monoclonal antibody (mAb) against PD-1 or PD- Ll. In some embodiments, the ICI is a mAb that neutralizes / blocks / inhibits / modulates the PD-1 pathway. In some embodiments, the ICI is a mAb against PD-1. In some embodiments, the anti-PD-1 mAb is Pembrolizumab (Keytruda; formerly called lambrolizumab). In some embodiments, the anti-PD-1 mAb is Nivolumab (Opdivo). In some embodiments, the anti- PD-1 mAb is Pidilizumab (CT0011). In some embodiments, the anti-PD-1 mAb isCemiplimab (Libtayo, REGN2810). In some embodiments, the anti-PD-1 mAb is Toripalimab (Loqtorzi). In some embodiments, the anti-PD-1 mAb is Retifanlimab (Zynyz). In some embodiments, the anti-PD-1 mAb is Dostarlimab (Jemperli). In some embodiments, the anti-PD-1 mAb is any one of AMP -224, MEDI0680, or PDR001. In some embodiments, the ICI is a mAb against PD-L1. In some embodiments, the anti-PD-Ll mAb is selected from Atezolizumab (Tecentriq), Avelumab (Bavencio), and Durvalumab (Imfinzi). In some embodiments, the anti-PD-Ll mAb is Atezolizumab. In some embodiments, the anti-PD-Ll mAb is Durvalumab. In some embodiments, the ICI is a mAb against CTLA-4. In some embodiments, the anti -CTLA-4 mAb is ipilimumab. In some embodiments, the anti-CTLA- 4 mAb is tremelimumab (Imjuno). In some embodiments, the ICI is a mAb against LAG-3. In some embodiments, the anti-LAG-3 mAb is Relatlimab.

[0059] As used herein, the term “factor” refers to any measurable biological molecule produced by the subject. In some embodiments, the factor is a protein. In some embodiments, the protein is a proteoform. In some embodiments, the factor is RNA (ribonucleic acid). In some embodiments, the factor is a gene. In some embodiments, the factor is a secreted factor. In some embodiments, the secreted factors are selected from cytokines, chemokines, growth factors, soluble receptors and enzymes. In some embodiments, the factor is a soluble factor. In some embodiment, the factor is cellular factor. In some embodiments, the factor is membranal factor. In some embodiments, the factor is a cell adhesion molecule. In some embodiments, the factor is a circulating factor. In some embodiments, the factor is a factor found in blood. In some embodiments, the factor is a host-generated factor. In some embodiments, the factor is a resistance factor. In some embodiments, the factor is a resistance-associated protein (RAP). As used herein, the term “resistance-associate factor” refers to a factor differentially expressed between responders and non-responders and which is therefore associated with resistance. In some embodiments, a RAP is upregulated in non- responders as compared to responders.

[0060] In some embodiments, a factor is a factor of the plurality of factors. In some embodiments, the RAP is selected from the NSCLC RAPs provided in Table 3. In some embodiments, the RAPs are all the NSCLC RAPs provided in Table 3. In some embodiments, the RAP is selected from the NSCLC RAPs provided in Table 4. In some embodiments, the RAPs are all the NSCLC RAPs provided in Table 4. In some embodiments, the RAP is selected from the SCLC RAPs provided in Table 3. In some embodiments, the RAPs are all the SCLC RAPs provided in Table 3. In some embodiments, the RAP is selected from the SCLC RAPs provided in Table 4. In some embodiments, theRAPs are all the SCLC RAPs provided in Table 4. In some embodiments, a plurality is at least 2, 3, 4, 5, 6 ,7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 12000, 15000, 20000, 25000, 30000, 35000, or 40000. Each possibility represents a separate embodiment of the invention. In some embodiments, a plurality is at least 50 factors. In some embodiments, a plurality is at least 100 factors. In some embodiments, a plurality is at least 200 factors. In some embodiments, a plurality is at least 388 factors.

[0061] In some embodiments, the plurality of factors comprises at least one factor selected from the group consisting of: KCNAB2, IL12B, IL23A, MCL1, KIR2DS2, AGA, RPN1, LAT, MFAP2, PUF60, MPZ, ACE, RNF122, TXNDC5, CDH15, FGFBP3, COL11A2, INPP5E, ADH7, MVK, RNF146, SOCS3, RBFOX2, ARFGAP1, SRSF6, RBM23, DDR1, APOF, TRA2B, MCTS1, TBCA, RGS7, PTPN9, CSNK1G2, ILF3, TPPP2, ARHGEF2, SRSF7, EWSR1, FSTL1, SPP1, FLRT2, FLRT3, VTN, ATP1B1, WFIKKN2, NRAC, PKD2, HSPA9, EMC4, ASAP2, NAP1L2, HTR7, DCUN1D3, RBL2, MAD1L1, GRB14, RBBP5, NAB2, CSF1R, CCN4, GPD1, KLK3, CXCL13, GZMA, C9, IL12B, RAP1GAP, IGFBP1, DHX58, COPS2, IL1RAP, CCL25, HPX, ADM, CD93, ISG15, MYL6B, HSPA1A, MBD1, TRAPPC3, AKT2, CRLF1, FTL, RBBP4, BMPER, SERPINB5, PMP2, OTC, OTOR, AOC1, FGFBP1, ATRN, NAGLU, SAA1, SAA4, CLSTN1, GSS, DLD, EPHB4, PRSS27, MUC16, CFHR2, HTRA1, KRT19, RBP4, SMOC2, BTD, TXLNA, MZB1, FADD, GSN, CDH17, LECT2, ADAMTSL1, RNASET2, SEMA4A, DDOST, BDH2, SNRPB2, GOLM1, RAB3A, CD46, SEPTIN6, WWOX, WDR5, HPCAL1, ALDH5A1, VAT1, SARS1, AFM, CDA, ITLN1, LRIG1, GREM1, PTGR2, UBE2L6, CLTA, GSR, PDCD6, SNCG, CRH, RGS21, UBE2R2, BASP1, GBP5, LMNB2, POP7, RAET1L, SEMA5B, CNTN3, UBL3, MMACHC, GTF2B, GCHFR, LRATD2, SGK1, TSEN15, SAR1B, CDK5RAP3, HAUS1, NKIRAS1, PHOSPHO2, PCDH17, TRIM5, ALDH7A1, TXNL4A, CEP20, PDE1B, ITGA4, ITGB1, LRFN3, ADGRB1, SGSH, MGAT5, B3GAT1, MGAT5, FBLN7, APBB1IP, PON2, PPP2R5D, RBFOX1, TIMP1, GEMIN7, CSNK1A1L, PHF11, BTN2A2, SKP2, SPATA46, LIN7A, BORCS5, ARRDC5, PCYT1A, PHYH, ANKRD63, VCX, NTAN1, STARD7, APOL2, FLT4, RCSD1, INIP, VMAC, XPNPEP3, IFNE, NELFA, KDM8, NCBP1, USF2, LRRC75A, APCS, PLCD1, ESPN, RFX5, RPS6KB2, N0M02, TCEAL2, CES3, DYRK1A, CYP2C19, CFI, IGFBP3, IL6, LEP, CRTC3, VEGFA, IL1RAP, HGF, PLA2G2A, CCL25, SERPINA7, POR, CCN3, HPX, IGFBP1, MMP3, FGA, FGB, FGG, BCAM, SPINT1, HAT1, GHR, CFP, CNTN1,SERPINF2, IL 19, MB, IGHM, LBP, NAAA, HAPLN1, IDS, NIDI, AC AN, TGFBI, DLL4, FCGR3B, ACY1, IBSP, SERPINA4, POSTN, SELE, B2M, HAMP, SERPINA1, AHSG, CKB, CKM, PROC, ANGPTL4, MBD4, PSMD7, IGHE, CXCL10, KLKB1, CFH, PFDN5, RBM39, DCTPP1, PRSS22, KYNU, IL6, SERPINA6, ITIH4, SFN, CCL7, LYZ, MMP13, STC1, CAPG, PI3, GPC5, HRG, SCGB2A1, SIRT2, TNFAIP6, CD300C, GPNMB, KRT18, TNFSF14, LEPR, PRKCG, FGL1, PGLYRP2, NPFF, MFAP4, TMX3, PRKCSH, DEFBI 12, SEMA4D, ACP6, AFP, NGF, FTH1, FTL, DMKN, EPHA10, CHRDL2, TP53, AOC1, IFNA8, CSH1, CSH2, TNC, PLTP, CCN1, CLSTN3, OIT3, GGT2, FMOD, C5orf 8, VWA1, INHBC, ADGRF5, C1QL2, PCYOX1, AOC2, CFHR4, LRRC15, UBE2J1, GFRAL, IGF2, LILRB5, LILRA6, APOA2, VWA2, DEPPI, C1QTNF3, SERPINA9, CFHR5, DLG3, GLTPD2, HBQ1, ENTPD1, AGGF1, NRG2, SPON2, FAM241B, JAML, BCHE, GPNMB, APOD, DLL1, PEAR1, RSPO4, ARL8B, PCDH10, MFAP3L, CD14, COL15A1, HAVCR1, ARHGEF10, MAN1A2, CRYZL1, TFPI2, PLXDC1, ACP2, MFAP2, ITIH2, EFCAB14, PLA1A, GZMK, YBX1, IDO1, NQO1, SPOCK3, and NXT1. In some embodiments, the plurality of factors comprises at least one factor selected from the group consisting of: LAT, COL11A2, FSTL1, PKD2, MAD1L1, RBBP5, NAB2, CCL25, RBBP4, BMPER, GSS, KRT19, BTD, SNRPB2, GSR, LMNB2, PPP2R5D, APOL2, VMAC, RFX5, IGFBP3, CRTC3, CCL25, HAT1, SERPINF2, NIDI, POSTN, ANGPTL4, MBD4, PSMD7, SFN, CCL7, STC1, KRT18, MFAP4, DEFBI 12, CCN1, GLTPD2, HAVCR1, TFPI2, PLA1A and SPOCK3. The amino acid sequences of these factors can be found in the Uniprot database, for example, and each factor’s Uniprot accession number is provided hereinbelow. Further, methods, reagents, and assays for measuring expression levels of these factors are well known in the art and are commercially available.

[0062] In some embodiments, the plurality of factors comprises at least one factor selected from the group consisting of: LAT, MSMB, SLITRK1, ITGB7, COL11 A2, PCBD1, KRT7, INA, LAMTOR3, ADSS2, RFXAP, PPP1R3B, FSTL1, HK2, CHAD, PKD2, MPP6, HTR2A, SNUB, MAD1L1, RBBP5, NAB2, C7, INHBA, S100A11, CCL25, RBBP4, APOCI, BMPER, GSTM1, NRG1, CD248, GSS, GUSB, KRT19, SORCS1, BTD, NANP, MAGOHB, SNRPB2, MORF4L1, ASRGL1, HDGFL3, EXOSC8, GSR, INHBA, DCPS, LMNB2, FAM50A, IZUMO4, EXOSC5, LRFN4, HLA-C, SH2B3, PPP2R5D, IP6K2, CCL22, APOL2, VMAC, DLST, KHSRP, CLUH, RFX5, HIP1R, SH3KBP1, TRIM28, IGFBP2, IGFBP3, CRTC3, EEF1A1, HM0X2, MRC1, CCL25, HAT1, YES1, SERPINF2, F10, RARRES2, DIABLO, CTSA, NIDI, LRP8, TBK1, BMX, CMA1, POSTN, CCL22,PRTN3, ARSA, ANGPTL4, CKB, FGF23, FGFR2, LMNB1, MBD4, PSMD7, NTF3, EIF4EBP2, P0N1, PLCG1, CSF2, APOD, SFN, F10, CCL7, STC1, FGFR4, IL22RA2, PDE7A, CKAP2, KRT18, CD5, MLN, MFAP4, DEFBI 12, CRISPLD2, FCRL1, PTH, ST6GAL1, CCN1, COL9A1, NTN1, TRAPPC4, CHST11, GLTPD2, MASP1, PTPRU, NTM, IGFBP2, LRP1, NEO1, PTK2B, PRG3, HAVCR1, IFNAR1, TFPI2, GGH, CAMP, DCBLD1, IL22RA2, PLA1A, GSTM3, SPOCK3, and TMEM59L.

[0063] In some embodiments, the plurality of factors comprises Voltage-gated potassium channel subunit beta-2 (KCNAB2). The human KCNAB2 gene can be found at Entrez gene #8514. The human KCNAB2 protein can be found at Uniprot ID Q13303. In some embodiments, the plurality of factors comprises Interleukin 12 subunit beta (IL12B). The human IL12B gene can be found at Entrez gene #3593. The human IL12B protein can be found at Uniprot ID P29460. In some embodiments, the plurality of factors comprises Interleukin 23 subunit alpha (IL23 A). The human IL23A gene can be found at Entrez gene #51561. The human IL23A protein can be found at Uniprot ID Q9NPF7. In some embodiments, the plurality of factors comprises Induced myeloid leukemia cell differentiation protein Mcl-1 (MCL1). The human MCLlgene can be found at Entrez gene #4170. The human MCLlprotein can be found at Uniprot ID Q07820. In some embodiments, the plurality of factors comprises Killer cell immunoglobulin-like receptor, two domains, short cytoplasmic tail, 1 (KIR2DS2). The human KIR2DS2gene can be found at Entrez gene #3806. The human KIR2DS2protein can be found at Uniprot ID Q14954. In some embodiments, the plurality of factors comprises N(4)-(beta-N-acetylglucosaminyl)-L- asparaginase (AGA). The human AGA gene can be found at Entrez gene #175. The human AGA protein can be found at Uniprot ID P20933. In some embodiments, the plurality of factors comprises Dolichyl-diphosphooligosaccharide — protein glycosyltransferase subunit 1 (RPN1). The human RPNlgene can be found at Entrez gene #6184. The human RPNlprotein can be found at Uniprot ID P04843. In some embodiments, the plurality of factors comprises Linker for activation of T cells (LAT). The human LAT gene can be found at Entrez gene #27040. The human LAT protein can be found at Uniprot ID 043561. In some embodiments, the plurality of factors comprises Microfibrillar-associated protein 2 (MFAP2). The human MFAP2 gene can be found at Entrez gene #4237. The human MFAP2protein can be found at Uniprot ID P55001. In some embodiments, the plurality of factors comprises Poly(U)-binding-splicing factor PUF60 (PUF60). The human PUF60 gene can be found at Entrez gene #22827. The human PUF60 protein can be found at Uniprot ID Q9UHX1. In some embodiments, the plurality of factors comprises Myelin protein zero(MPZ). The human MPZ gene can be found at Entrez gene #4359. The human MPZ protein can be found at Uniprot ID P25189. In some embodiments, the plurality of factors comprises Angiotensin-converting enzyme (ACE). The human ACE gene can be found at Entrez gene #1636. The human ACE protein can be found at Uniprot ID P12821. In some embodiments, the plurality of factors comprises RING finger protein 122 (RNF122). The human RNF122 gene can be found at Entrez gene #79845. The human RNF122 protein can be found at Uniprot ID Q9H9V4. In some embodiments, the plurality of factors comprises Thioredoxin domain-containing protein 5 (TXNDC5). The human TXNDC5gene can be found at Entrez gene #81567. The human TXNDC5protein can be found at Uniprot ID Q8NBS9. In some embodiments, the plurality of factors comprises Cadherin-15 (CDH15). The human CDH15 gene can be found at Entrez gene #1013. The human CDH15 protein can be found at Uniprot ID P55291. In some embodiments, the plurality of factors comprises Fibroblast growth factor-binding protein 1 (FGFBP3). The human FGFBP3 gene can be found at Entrez gene #9982. The human FGFBP3 protein can be found at Uniprot ID Q 14512. In some embodiments, the plurality of factors comprises Collagen alpha-2(XI) chain (COL11A2). The human COL11A2 gene can be found at Entrez gene #1302. The human COL11A2 protein can be found at Uniprot ID P13942. In some embodiments, the plurality of factors comprises 72 kDa inositol polyphosphate 5-phosphatase (INPP5E). The human INPP5E gene can be found at Entrez gene #56623. The human INPP5E protein can be found at Uniprot ID Q2YD81. In some embodiments, the plurality of factors comprises Alcohol dehydrogenase class 4 mu / sigma chain (ADH7). The human ADH7gene can be found at Entrez gene #131. The human ADH7protein can be found at Uniprot ID P40394. In some embodiments, the plurality of factors comprises Mevalonate kinase (MVK). The human MVK gene can be found at Entrez gene #4598. The human MVK protein can be found at Uniprot ID Q03426. In some embodiments, the plurality of factors comprises RING finger protein 146 (RNF146). The human RNF146 gene can be found at Entrez gene #81847. The human RNF146 protein can be found at Uniprot ID Q9NTX7. In some embodiments, the plurality of factors comprises Suppressor of cytokine signaling 3 (SOCS3). The human SOCS3 gene can be found at Entrez gene #9021. The human SOCS3 protein can be found at Uniprot ID 014543 or Q6FI39. In some embodiments, the plurality of factors comprises RNA binding motif protein 9 (RBFOX2). The human RBFOX2 gene can be found at Entrez gene #23543. The human RBFOX2 protein can be found at Uniprot ID 043251. In some embodiments, the plurality of factors comprises ADP-ribosylation factor GTPase-activating protein 1 (ARFGAP1). The human ARFGAP1 gene can be found at Entrez gene #55738.The human ARFGAP1 protein can be found at Uniprot ID Q8N6T3. In some embodiments, the plurality of factors comprises Splicing factor, arginine / serine-rich 6 (SRSF6). The human SRSF6 gene can be found at Entrez gene #6431. The human SRSF6 protein can be found at Uniprot ID QI 3247. In some embodiments, the plurality of factors comprises Probable RNA-binding protein 23 (RBM23). The human RBM23 gene can be found at Entrez gene #55147. The human RBM23 protein can be found at Uniprot ID Q88U06. In some embodiments, the plurality of factors comprises Discoidin domain receptor family, member 1 (DDR1). The human DDR1 gene can be found at Entrez gene #780. The human DDR1 protein can be found at Uniprot ID Q0345. In some embodiments, the plurality of factors comprises Apolipoprotein F (APOF). The human APOF gene can be found at Entrez gene #319. The human APOF protein can be found at Uniprot ID Q13790. In some embodiments, the plurality of factors comprises Transformer-2 protein homolog beta (TRA2B). The human TRA2B gene can be found at Entrez gene #6434. The human TRA2B protein can be found at Uniprot ID P62995. In some embodiments, the plurality of factors comprises MCTS1, reinitiation and release factor (MCTS1). The human MCTS1 gene can be found at Entrez gene #28985. The human MCTS1 protein can be found at Uniprot ID Q9ULC4. In some embodiments, the plurality of factors comprises Tubulin-specific chaperone A (TBCA). The human TBCA gene can be found at Entrez gene #6902. The human TBCA protein can be found at Uniprot ID 075347. In some embodiments, the plurality of factors comprises Regulator of G-protein signaling 7 (RGS7). The human RGS7 gene can be found at Entrez gene #6000. The human RGS7 protein can be found at Uniprot ID P49802. In some embodiments, the plurality of factors comprises Tyrosine-protein phosphatase non-receptor type (PTPN9). The human PTPN9 gene can be found at Entrez gene #5780. The human PTPN9 protein can be found at Uniprot ID P43378. In some embodiments, the plurality of factors comprises casein kinase 1 gamma 2 (CSNK1G2). The human CSNK1G2 gene can be found at Entrez gene #1455. The human CSNK1G2 protein can be found at Uniprot ID P78368. In some embodiments, the plurality of factors comprises Interleukin enhancerbinding factor 3 (ILF3). The human ILF3 gene can be found at Entrez gene #3609. The human ILF3 protein can be found at Uniprot ID Q 12906. In some embodiments, the plurality of factors comprises Tubulin polymerization-promoting protein (TPPP2). The human TPPP2 gene can be found at Entrez gene #11076. The human TPPP2 protein can be found at Uniprot ID 094811. In some embodiments, the plurality of factors comprises Rho guanine nucleotide exchange factor 2 (ARHGEF2). The human ARHGEF2 gene can be found at Entrez gene #9181. The human ARHGEF2 protein can be found at Uniprot ID Q92974. In someembodiments, the plurality of factors comprises Serine / arginine-rich splicing factor 7 (SRSF7). The human SRSF7 gene can be found at Entrez gene #6432. The human SRSF7 protein can be found at Uniprot ID QI 6629. In some embodiments, the plurality of factors comprises RNA-binding protein EWS (EWSR1). The human EWSR1 gene can be found at Entrez gene #2130. The human EWSR1 protein can be found at Uniprot ID Q01844. In some embodiments, the plurality of factors comprises Follistatin-related protein 1 (FSTL1). The human FSTL1 gene can be found at Entrez gene #11167. The human FSTL1 protein can be found at Uniprot ID Q12841. In some embodiments, the plurality of factors comprises Osteopontin also known as secreted phosphoprotein 1 (SPP1). The human SPP1 gene can be found at Entrez gene #6696. The human SPP1 protein can be found at Uniprot ID Pl 0451. In some embodiments, the plurality of factors comprises Fibronectin leucine rich transmembrane protein 2 (FLRT2). The human FLRT2 gene can be found at Entrez gene #23768. The human FLRT2 protein can be found at Uniprot ID 043155. In some embodiments, the plurality of factors comprises Fibronectin leucine rich transmembrane protein 3 (FLRT3). The human FLRT3 gene can be found at Entrez gene #23767. The human FLRT3 protein can be found at Uniprot ID Q9NZU0. In some embodiments, the plurality of factors comprises Vitronectin (VTN). The human VTN gene can be found at Entrez gene #7448. The human VTN protein can be found at Uniprot ID P04004. In some embodiments, the plurality of factors comprises Sodium / potassium -transporting ATPase subunit beta-1 (ATP1B1). The human ATP1B1 gene can be found at Entrez gene #481. The human ATP1B1 protein can be found at Uniprot ID P05026. In some embodiments, the plurality of factors comprises WAP, follistatin / kazal, immunoglobulin, kunitz and netrin domain containing 2 (WFIKKN2). The human WFIKKN2 gene can be found at Entrez gene #134857. The human WFIKKN2 protein can be found at Uniprot ID Q8TEU8. In some embodiments, the plurality of factors comprises Nutritionally-regulated adipose and cardiac enriched protein homolog (NRAC). The human NRAC gene can be found at Entrez gene #400258. The human NRAC protein can be found at Uniprot ID Q8N912. In some embodiments, the plurality of factors comprises Polycystin-2 (PKD2). The human PKD2 gene can be found at Entrez gene #5311. The human PKD2 protein can be found at Uniprot ID Q13563. In some embodiments, the plurality of factors comprises Mitochondrial 70kDa heat shock protein (HSPA9). The human HSPA9 gene can be found at Entrez gene #3313. The human HSPA9 protein can be found at Uniprot ID P38646. In some embodiments, the plurality of factors comprises ER membrane protein complex subunit 4 (EMC4). The human EMC4 gene can be found at Entrez gene #51234. The human EMC4 protein can be found atUniprot ID Q5J8M3. In some embodiments, the plurality of factors comprises Arf-GAP with SH3 domain, ANK repeat and PH domain-containing protein 2 (ASAP2). The human ASAP2 gene can be found at Entrez gene #8853. The human ASAP2 protein can be found at Uniprot ID 043150. In some embodiments, the plurality of factors comprises Nucleosome assembly protein 1 -like 2 (NAP1L2). The human NAP1L2 gene can be found at Entrez gene #4674. The human NAP1L2 protein can be found at Uniprot ID Q9ULW6. In some embodiments, the plurality of factors comprises 5-HT7 receptor (HTR7). The human HTR7 gene can be found at Entrez gene #3363. The human HTR7 protein can be found at Uniprot ID P34969. In some embodiments, the plurality of factors comprises DCNl-like protein 3 (DCUN1D3). The human DCUN1D3 gene can be found at Entrez gene #123879. The human DCUN1D3 protein can be found at Uniprot ID Q8IWE4. In some embodiments, the plurality of factors comprises Retinoblastoma-like protein 2 (RBL2). The human RBL2 gene can be found at Entrez gene #5934. The human RBL2 protein can be found at Uniprot ID Q08999. In some embodiments, the plurality of factors comprises Mitotic spindle assembly checkpoint protein MAD1 (MAD1L1). The human MAD1L1 gene can be found at Entrez gene #8379. The human MAD1L1 protein can be found at Uniprot ID Q9Y6D9. In some embodiments, the plurality of factors comprises Growth factor receptor-bound protein 14 (GRB14). The human GRB14 gene can be found at Entrez gene #2888. The human GRB14 protein can be found at Uniprot ID Q14449. In some embodiments, the plurality of factors comprises Retinoblastoma-binding protein 5 (RBBP5). The human RBBP5 gene can be found at Entrez gene #5929. The human RBBP5 protein can be found at Uniprot ID Q 15291. In some embodiments, the plurality of factors comprises NGFI-A-binding protein 2 (NAB2). The human NAB2 gene can be found at Entrez gene #4665. The human NAB2 protein can be found at Uniprot ID QI 5742. In some embodiments, the plurality of factors comprises Colony stimulating factor 1 receptor (CSF1R). The human CSF1R gene can be found at Entrez gene #1436. The human CSF1R protein can be found at Uniprot ID P07333. In some embodiments, the plurality of factors comprises CCN family member 4 (CCN4). The human CCN4 gene can be found at Entrez gene #8840. The human CCN4 protein can be found at Uniprot ID 095388. In some embodiments, the plurality of factors comprises Glycerol-3- phosphate dehydrogenase 1-like protein (GPD1). The human GPD1 gene can be found at Entrez gene #2819. The human GPD1 protein can be found at Uniprot ID Q8N335. In some embodiments, the plurality of factors comprises Prostate-specific antigen also known as kallikrein-3 (KLK3). The human KLK3 gene can be found at Entrez gene #354. The human KLK3 protein can be found at Uniprot ID P07288. In some embodiments, the plurality offactors comprises Chemokine (C-X-C motif) ligand 13 (CXCL13). The human CXCL13 gene can be found at Entrez gene #10563. The human CXCL13 protein can be found at Uniprot ID 043927. In some embodiments, the plurality of factors comprises Granzyme A (GZMA). The human GZMA gene can be found at Entrez gene #3001. The human GZMA protein can be found at Uniprot ID P12544. In some embodiments, the plurality of factors comprises Complement component C9 (C9). The human C9 gene can be found at Entrez gene #735. The human C9 protein can be found at Uniprot ID P06683. In some embodiments, the plurality of factors comprises Rapl GTPase-activating protein 1 (RAP 1 GAP). The human RAP 1 GAP gene can be found at Entrez gene #5909. The human RAP1GAP protein can be found at Uniprot ID P47736. In some embodiments, the plurality of factors comprises Insulin-like growth factor-binding protein 1 (IGFBP1). The human IGFBP1 gene can be found at Entrez gene #3484. The human IGFBP1 protein can be found at Uniprot ID P08833. In some embodiments, the plurality of factors comprises Probable ATP-dependent RNA helicase DHX58 (DHX58). The human DHX58 gene can be found at Entrez gene #79132. The human DHX58 protein can be found at Uniprot ID Q96C10. In some embodiments, the plurality of factors comprises COP9 signalosome complex subunit 2 (COPS2). The human COPS2 gene can be found at Entrez gene #9318. The human COPS2 protein can be found at Uniprot ID P61201. In some embodiments, the plurality of factors comprises Interleukin-1 receptor accessory protein (IL1RAP). The human IL1RAP gene can be found at Entrez gene #3556. The human IL1RAP protein can be found at Uniprot ID Q9NPH3. In some embodiments, the plurality of factors comprises Chemokine (C-C motif) ligand 25 (CCL25). The human CCL25 gene can be found at Entrez gene #6370. The human CCL25 protein can be found at Uniprot ID Q68A93. In some embodiments, the plurality of factors comprises Hemopexin (HPX). The human HPX gene can be found at Entrez gene #3263. The human HPX protein can be found at Uniprot ID P02790. In some embodiments, the plurality of factors comprises Adrenomedullin (ADM). The human ADM gene can be found at Entrez gene #133. The human ADM protein can be found at Uniprot ID P35318. In some embodiments, the plurality of factors comprises Cluster of Differentiation 93 (CD93). The human CD93 gene can be found at Entrez gene #22918. The human CD93 protein can be found at Uniprot ID Q9NPY3. In some embodiments, the plurality of factors comprises Interferon-stimulated gene 15 (ISG15). The human ISG15 gene can be found at Entrez gene #9636. The human ISG15 protein can be found at Uniprot ID P05161. In some embodiments, the plurality of factors comprises Myosin light chain 6B (MYL6B). The human MYL6B gene can be found at Entrez gene #140465. The human MYL6B protein can be found atUniprot ID P14649. In some embodiments, the plurality of factors comprises Heat shock 70 kDa protein 1 (HSPA1A). The human HSPA1A gene can be found at Entrez gene #3303. The human HSPA1 A protein can be found at Uniprot ID P0DMV8 or P0DMV9. In some embodiments, the plurality of factors comprises Methyl-CpG-binding domain protein 1 (MBD1). The human MBD1 gene can be found at Entrez gene #4152. The human MBD1 protein can be found at Uniprot ID Q9UIS9. In some embodiments, the plurality of factors comprises Trafficking protein particle complex subunit 3 (TRAPPC3). The human TRAPPC3 gene can be found at Entrez gene #27095. The human TRAPPC3 protein can be found at Uniprot ID 043617. In some embodiments, the plurality of factors comprises RAC- beta serine / threonine-protein kinase (AKT2). The human AKT2 gene can be found at Entrez gene #208. The human AKT2 protein can be found at Uniprot ID P31751. In some embodiments, the plurality of factors comprises Cytokine receptor-like factor 1 (CRLF1). The human CRLF1 gene can be found at Entrez gene #9244. The human CRLF1 protein can be found at Uniprot ID 075462. In some embodiments, the plurality of factors comprises Ferritin light chain (FTL). The human FTL gene can be found at Entrez gene #2512. The human FTL protein can be found at Uniprot ID P02792. In some embodiments, the plurality of factors comprises Histone-binding protein RBBP4 (RBBP4). The human RBBP4 gene can be found at Entrez gene #5928. The human RBBP4 protein can be found at Uniprot ID Q09028. In some embodiments, the plurality of factors comprises BMP binding endothelial regulator (BMPER). The human BMPER gene can be found at Entrez gene #168667. The human BMPER protein can be found at Uniprot ID Q8N8U9. In some embodiments, the plurality of factors comprises Maspin (SERPINB5). The human SERPINB5 gene can be found at Entrez gene #5268. The human SERPINB5 protein can be found at Uniprot ID P36952. In some embodiments, the plurality of factors comprises Myelin P2 protein (PMP2). The human PMP2 gene can be found at Entrez gene #5375. The human PMP2 protein can be found at Uniprot ID P02689. In some embodiments, the plurality of factors comprises Ornithine transcarbamylase (OTC). The human OTC gene can be found at Entrez gene #5009. The human OTC protein can be found at Uniprot ID P00480. In some embodiments, the plurality of factors comprises Otoraplin (OTOR). The human OTOR gene can be found at Entrez gene #56914. The human OTOR protein can be found at Uniprot ID Q9NRC9. In some embodiments, the plurality of factors comprises Diamine oxidase [copper-containing] (AOC1). The human AOC1 gene can be found at Entrez gene #26. The human AOC1 protein can be found at Uniprot ID Q8JZQ5. In some embodiments, the plurality of factors comprises Fibroblast growth factor-binding protein 1 (FGFBP1). The human FGFBP1 genecan be found at Entrez gene #9982. The human FGFBP1 protein can be found at Uniprot ID Q14512. In some embodiments, the plurality of factors comprises Attractin (ATRN). The human ATRN gene can be found at Entrez gene #8455. The human ATRN protein can be found at Uniprot ID 075882. In some embodiments, the plurality of factors comprises N- acetylglucosaminidase, alpha (NAGLU). The human NAGLU gene can be found at Entrez gene #4669. The human NAGLU protein can be found at Uniprot ID P54802. In some embodiments, the plurality of factors comprises Serum amyloid Al (SAA1). The human SAA1 gene can be found at Entrez gene #6288. The human SAA1 protein can be found at Uniprot ID P0DJI8. In some embodiments, the plurality of factors comprises Serum amyloid A4 (SAA4). The human SAA4 gene can be found at Entrez gene #6291. The human SAA4 protein can be found at Uniprot ID P35542. In some embodiments, the plurality of factors comprises Calsyntenin-1 (CLSTN1). The human CLSTN1 gene can be found at Entrez gene #22883. The human CLSTN1 protein can be found at Uniprot ID 094985. In some embodiments, the plurality of factors comprises Glutathione synthetase (GSS). The human GSS gene can be found at Entrez gene #2937. The human GSS protein can be found at Uniprot ID C9K4X8. In some embodiments, the plurality of factors comprises Dihydrolipoamide dehydrogenase (DLD). The human DLD gene can be found at Entrez gene #1738. The human DLD protein can be found at Uniprot ID P09622. In some embodiments, the plurality of factors comprises Ephrin type-B receptor 4 (EPHB4). The human EPHB4 gene can be found at Entrez gene #2050. The human EPHB4 protein can be found at Uniprot ID P54760. In some embodiments, the plurality of factors comprises Serine protease 27 (PRSS27). The human PRSS27 gene can be found at Entrez gene #83886. The human PRSS27 protein can be found at Uniprot ID Q9BQR3. In some embodiments, the plurality of factors comprises Mucin- 16 (MUC16). The human MUC16 gene can be found at Entrez gene #94025. The human MUC16 protein can be found at Uniprot ID Q8WXI7. In some embodiments, the plurality of factors comprises Complement factor H-related protein 2 (CFHR2). The human CFHR2 gene can be found at Entrez gene #3080. The human CFHR2 protein can be found at Uniprot ID P36980. In some embodiments, the plurality of factors comprises Serine protease HTRA1 (HTRA1). The human HTRA1 gene can be found at Entrez gene #5654. The human HTRA1 protein can be found at Uniprot ID Q92743. In some embodiments, the plurality of factors comprises Keratin, type I cytoskeletal 19 (KRT19). The human KRT19 gene can be found at Entrez gene #3880. The human KRT19 protein can be found at Uniprot ID P08727. In some embodiments, the plurality of factors comprises Retinol binding protein 4 (RBP4). Thehuman RBP4 gene can be found at Entrez gene #5950. The human RBP4 protein can be found at Uniprot ID P02753. In some embodiments, the plurality of factors comprises SPARC -related modular calcium-binding protein 2 (SMOC2). The human SMOC2 gene can be found at Entrez gene #64094. The human SMOC2 protein can be found at Uniprot ID Q9H3U7. In some embodiments, the plurality of factors comprises Biotinidase (BTD). The human BTD gene can be found at Entrez gene #686. The human BTD protein can be found at Uniprot ID P43251. In some embodiments, the plurality of factors comprises Alpha-taxili (TXLNA). The human TXLNA gene can be found at Entrez gene #200081. The human TXLNA protein can be found at Uniprot ID P40222. In some embodiments, the plurality of factors comprises Marginal zone B and Bl cell-specific protein (MZB1). The human MZB1 gene can be found at Entrez gene #51237. The human MZB 1 protein can be found at Uniprot ID Q8WU39. In some embodiments, the plurality of factors comprises FAS-associated death domain protein (FADD). The human FADD gene can be found at Entrez gene #8772. The human FADD protein can be found at Uniprot ID Q13158. In some embodiments, the plurality of factors comprises Gelsolin (GSN). The human GSN gene can be found at Entrez gene #2934. The human GSN protein can be found at Uniprot ID P06396. In some embodiments, the plurality of factors comprises Cadherin-17 (CDH17). The human CDH17 gene can be found at Entrez gene #1015. The human CDH17 protein can be found at Uniprot ID Q12864. In some embodiments, the plurality of factors comprises Leukocyte cell-derived chemotaxin-2 (LECT2). The human LECT2 gene can be found at Entrez gene #3950. The human LECT2 protein can be found at Uniprot ID 014960. In some embodiments, the plurality of factors comprises ADAMTS-like protein 1 (ADAMTSL1). The human ADAMTSL1 gene can be found at Entrez gene #92949. The human ADAMTSL1 protein can be found at Uniprot ID Q8N6G6. In some embodiments, the plurality of factors comprises Ribonuclease T2 (RNASET2). The human RNASET2 gene can be found at Entrez gene #8635. The human RNASET2 protein can be found at Uniprot ID 000584. In some embodiments, the plurality of factors comprises Semaphorin-4A (SEMA4A). The human SEMA4A gene can be found at Entrez gene #64218. The human SEMA4A protein can be found at Uniprot ID Q9H3S1. In some embodiments, the plurality of factors comprises Dolichyl-diphosphooligosaccharide — protein glycosyltransferase 48 kDa subunit (DDOST). The human DDOST gene can be found at Entrez gene #1650. The humanDDOST protein can be found at Uniprot ID P39656. In some embodiments, the plurality of factors comprises Dehydrogenase / reductase SDR family member 6 (BDH2 / DHRS6). The human BDH2 gene can be found at Entrez gene # 56898. The human BDH2 protein can be foundat Uniprot ID Q9BUT1. In some embodiments, the plurality of factors comprises U2 small nuclear ribonucleoprotein B (SNRPB2). The human SNRPB2 gene can be found at Entrez gene #6629. The human SNRPB2 protein can be found at Uniprot ID P08579. In some embodiments, the plurality of factors comprises Golgi membrane protein 1 (G0LM1). The human G0LM1 gene can be found at Entrez gene #51280. The human GOLM1 protein can be found at Uniprot ID Q8NBJ4. In some embodiments, the plurality of factors comprises Ras-related protein Rab-3 A (RAB3 A). The human RAB3 A gene can be found at Entrez gene #5864. The human RAB3A protein can be found at Uniprot ID P20336. In some embodiments, the plurality of factors comprises CD46 complement regulatory protein (CD46). The human CD46 gene can be found at Entrez gene #4179. The human CD46 protein can be found at Uniprot ID P15529. In some embodiments, the plurality of factors comprises Septin-6 (SEPTIN6). The human SEPTIN6 gene can be found at Entrez gene # 23157. The human SEPTIN6 protein can be found at Uniprot ID Q3SZN0. In some embodiments, the plurality of factors comprises WW domain-containing oxidoreductase (WWOX). The human WWOX gene can be found at Entrez gene #51741. The human WWOX protein can be found at Uniprot ID Q9NZC7. In some embodiments, the plurality of factors comprises WD repeat-containing protein 5 (WDR5). The human WDR5 gene can be found at Entrez gene #11091. The human WDR5 protein can be found at Uniprot ID P61964. In some embodiments, the plurality of factors comprises Hippocalcin-like protein 1 (HPCAL1). The human HPCAL1 gene can be found at Entrez gene #3241. The human HPCAL1 protein can be found at Uniprot ID P37235. In some embodiments, the plurality of factors comprises Aldehyde dehydrogenase 5 family, member Al (ALDH5A1). The human ALDH5A1 gene can be found at Entrez gene #7915. The human ALDH5A1 protein can be found at Uniprot ID P51649. In some embodiments, the plurality of factors comprises Synaptic vesicle membrane protein VAT-1 homolog (VAT1). The human VAT1 gene can be found at Entrez gene #10493. The human VAT1 protein can be found at Uniprot ID Q99536. In some embodiments, the plurality of factors comprises Cytoplasmic seryl-tRNA synthetase (SARS1). The human SARS1 gene can be found at Entrez gene #6301. The human SARS1 protein can be found at Uniprot ID P49591. In some embodiments, the plurality of factors comprises Afamin (AFM). The human AFM gene can be found at Entrez gene #173. The human AFM protein can be found at Uniprot ID P43652. In some embodiments, the plurality of factors comprises Cytidine deaminase (CDA). The human CDA gene can be found at Entrez gene #978. The human CDA protein can be found at Uniprot ID P32320. In some embodiments, the plurality of factors comprises Intelectin-1(ITLN1). The human ITLN1 gene can be found at Entrez gene #55600. The human ITLN1 protein can be found at Uniprot ID Q8WWA0. In some embodiments, the plurality of factors comprises Leucine-rich repeats and immunoglobulin-like domains protein 1 (LRIG1). The human LRIG1 gene can be found at Entrez gene #26018. The human LRIG1 protein can be found at Uniprot ID Q96JA1. In some embodiments, the plurality of factors comprises Gremlin (GREM1). The human GREM1 gene can be found at Entrez gene #26585. The human GREM1 protein can be found at Uniprot ID 060565. In some embodiments, the plurality of factors comprises Prostaglandin reductase 2 (PTGR2). The human PTGR2 gene can be found at Entrez gene # 145482. The human PTGR2 protein can be found at Uniprot ID Q8N8N7. In some embodiments, the plurality of factors comprises Gelsolin (GSN). The human GSN gene can be found at Entrez gene #2934. The human GSN protein can be found at Uniprot ID P06396. In some embodiments, the plurality of factors comprises Ubiquitin / ISG15-conjugating enzyme E2 L6 (UBE2L6). The human UBE2L6 gene can be found at Entrez gene #9246. The human UBE2L6 protein can be found at Uniprot ID 014933. In some embodiments, the plurality of factors comprises Clathrin light chain A (CLTA). The human CLTA gene can be found at Entrez gene # 1211. The human CLTA protein can be found at Uniprot ID P09496. In some embodiments, the plurality of factors comprises Glutathione reductase (GSR). The human GSR gene can be found at Entrez gene #2936. The human GSR protein can be found at Uniprot ID P00390. In some embodiments, the plurality of factors comprises Programmed cell death protein 6 (PDCD6). The human PDCD6 gene can be found at Entrez gene #10016. The human PDCD6 protein can be found at Uniprot ID 075340. In some embodiments, the plurality of factors comprises Gamma- synuclein (SNCG). The human SNCG gene can be found at Entrez gene #6623. The human SNCG protein can be found at Uniprot ID 076070. In some embodiments, the plurality of factors comprises Corticotropin-releasing hormone (CRH). The human CRH gene can be found at Entrez gene #1392. The human CRH protein can be found at Uniprot ID P06850. In some embodiments, the plurality of factors comprises Regulator of G-protein signaling 21 (RGS21). The human RGS21 gene can be found at Entrez gene # 431704. The human RGS21 protein can be found at Uniprot ID Q2M5E4. In some embodiments, the plurality of factors comprises Ubiquitin-conjugating enzyme E2 R2 (UBE2R2). The human UBE2R2 gene can be found at Entrez gene #54926. The human UBE2R2 protein can be found at Uniprot ID Q712K3. In some embodiments, the plurality of factors comprises Brain acid soluble protein 1 (BASP1). The human BASP1 gene can be found at Entrez gene #10409. The human BASP1 protein can be found at Uniprot ID P80723. In some embodiments, theplurality of factors comprises Guanylate binding protein 5 (GBP5). The human GBP5 gene can be found at Entrez gene #115362. The human GBP5 protein can be found at Uniprot ID Q96PP8. In some embodiments, the plurality of factors comprises Lamin B2 (LMNB2). The human LMNB2 gene can be found at Entrez gene #84823. The human LMNB2 protein can be found at Uniprot ID Q03252. In some embodiments, the plurality of factors comprises Ribonuclease P protein subunit p20 (POP7). The human POP7 gene can be found at Entrez gene #10248. The human POP7 protein can be found at Uniprot ID 075817. In some embodiments, the plurality of factors comprises Retinoic acid early transcript IL (RAET1L / ULBP6). The human RAET1L gene can be found at Entrez gene # 154064. The human RAET1L protein can be found at Uniprot ID Q5VY80. In some embodiments, the plurality of factors comprises Semaphorin-5B (SEMA5B). The human SEMA5B gene can be found at Entrez gene # 54437. The human SEMA5B protein can be found at Uniprot ID Q9P283. In some embodiments, the plurality of factors comprises Contactin-3 (CNTN3). The human CNTN3 gene can be found at Entrez gene #5067. The human CNTN3 protein can be found at Uniprot ID Q9P232. In some embodiments, the plurality of factors comprises Ubiquitin-like protein 3 (UBL3). The human UBL3 gene can be found at Entrez gene #5412. The human UBL3 protein can be found at Uniprot ID 095164. In some embodiments, the plurality of factors comprises Methylmalonic aciduria and homocystinuria type C protein (MMACHC). The human MMACHC gene can be found at Entrez gene #25974. The human MMACHC protein can be found at Uniprot ID Q9 Y4U 1. In some embodiments, the plurality of factors comprises Transcription factor II B (GTF2B). The human GTF2B gene can be found at Entrez gene #2959. The human GTF2B protein can be found at Uniprot ID Q00403. In some embodiments, the plurality of factors comprises GTP cyclohydrolase 1 feedback regulatory protein (GCHFR). The human GCHFR gene can be found at Entrez gene #2644. The human GCHFR protein can be found at Uniprot ID P30047. In some embodiments, the plurality of factors comprises Protein LRATD2 (LRATD2). The human LRATD2 gene can be found at Entrez gene # 157638. The human LRATD2 protein can be found at Uniprot ID Q96KN1. In some embodiments, the plurality of factors comprises Serine / threonine-protein kinase Sgkl (SGK1). The human SGK1 gene can be found at Entrez gene #6446. The human SGK1 protein can be found at Uniprot ID 000141. In some embodiments, the plurality of factors comprises tRNA-splicing endonuclease subunit Sent 5 (TSEN15). The human TSEN15 gene can be found at Entrez gene #116461. The human TSEN15 protein can be found at Uniprot ID Q8WW01. In some embodiments, the plurality of factors comprises GTP -binding protein SARlb (SAR1B). The human SAR1B gene can be found at Entrezgene # 51128. The human SAR1B protein can be found at Uniprot ID Q9Y6B6. In some embodiments, the plurality of factors comprises CDK5 regulatory subunit-associated protein 3 (CDK5RAP3). The human CDK5RAP3 gene can be found at Entrez gene #80279. The human CDK5RAP3 protein can be found at Uniprot ID Q96JB5. In some embodiments, the plurality of factors comprises HAUS augmin-like complex subunit 1 (HAUS1). The human HAUS1 gene can be found at Entrez gene #115106. The human HAUS1 protein canbe found at Uniprot ID Q96CS2. In some embodiments, the plurality of factors comprises NF-kappa- B inhibitor alpha (NKIRAS1). The human NKIRAS1 gene can be found at Entrez gene # 28512. The human NKIRAS1 protein can be found at Uniprot ID P25963. In some embodiments, the plurality of factors comprises Pyridoxal phosphate phosphatase PHOSPHO2 (PHOSPHO2). The human PHOSPHO2 gene can be found at Entrez gene # 493911. The human PHOSPHO2 protein can be found at Uniprot ID Q8TCD6. In some embodiments, the plurality of factors comprises Protocadherin-17 (PCDH17). The human PCDH17 gene can be found at Entrez gene #27253. The human PCDH17 protein can be found at Uniprot ID 014917. In some embodiments, the plurality of factors comprises Tripartite motif-containing protein 5 (TRIM5). The human TRIM5 gene can be found at Entrez gene #85363. The human TRIM5 protein can be found at Uniprot ID Q9C035. In some embodiments, the plurality of factors comprises Aldehyde dehydrogenase 7 family, member Al (ALDH7A1). The human ALDH7A1 gene can be found at Entrez gene #501. The human ALDH7A1 protein can be found at Uniprot ID P49419. In some embodiments, the plurality of factors comprises Thioredoxin-like protein 4A (TXNL4A). The human TXNL4A gene can be found at Entrez gene #10907. The human TXNL4A protein can be found at Uniprot ID P83876. In some embodiments, the plurality of factors comprises Centrosomal protein 20 (CEP20). The human CEP20 gene can be found at Entrez gene # 123811. The human CEP20 protein can be found at Uniprot ID Q96NB1. In some embodiments, the plurality of factors comprises Calcium / calmodulin-dependent 3',5'-cyclic nucleotide phosphodiesterase IB (PDE1B). The human PDE1B gene can be found at Entrez gene #5153. The human PDE1B protein can be found at Uniprot ID Q01064. In some embodiments, the plurality of factors comprises Integrin alpha-4 (ITGA4). The human ITGA4 gene can be found at Entrez gene #3676. The human ITGA4 protein can be found at Uniprot ID P13612. In some embodiments, the plurality of factors comprises Integrin beta- 1 (ITGB1). The human ITGB1 gene can be found at Entrez gene #3688. The human ITGB1 protein can be found at Uniprot ID P05556. In some embodiments, the plurality of factors comprises Leucine-rich repeat neuronal protein 3 (LRFN3). The human LRFN3 gene can befound at Entrez gene #54674. The human LRFN3 protein can be found at Uniprot ID Q9H3W5. In some embodiments, the plurality of factors comprises Adhesion G protein- coupled receptor Bl (ADGRB1). The human ADGRB1 gene can be found at Entrez gene # 575. The human ADGRB1 protein can be found at Uniprot ID 014514. In some embodiments, the plurality of factors comprises N-sulphoglucosamine sulphohydrolase (SGSH). The human SGSH gene can be found at Entrez gene #6448. The human SGSH protein can be found at Uniprot ID P51688. In some embodiments, the plurality of factors comprises Alpha-1, 6-mannosylglycoprotein 6-beta-N-acetylglucosaminyltransferase A (MGAT5). The human MGAT5 gene can be found at Entrez gene #4249. The human MGAT5 protein can be found at Uniprot ID Q09328. In some embodiments, the plurality of factors comprises 3-beta-glucuronosyltransferase 1 (B3GAT1). The human B3GAT1 gene can be found at Entrez gene #27087. The human B3GAT1 protein can be found at Uniprot ID Q9P2W7. In some embodiments, the plurality of factors comprises Alpha- 1,6- mannosylglycoprotein 6-beta-N-acetylglucosaminyltransferase A (MGAT5). The human MGAT5 gene can be found at Entrez gene #4249. The human MGAT5 protein can be found at Uniprot ID P06396. n can be found at Uniprot ID Q09328. In some embodiments, the plurality of factors comprises Fibulin-7 (FBLN7). The human FBLN7 gene can be found at Entrez gene # 129804. The human FBLN7 protein can be found at Uniprot ID Q501P1. In some embodiments, the plurality of factors comprises Amyloid beta A4 precursor proteinbinding family B member 1 -interacting protein (APBB1IP). The human APBB1IP gene can be found at Entrez gene #54518. The human APBB1IP protein can be found at Uniprot ID Q7Z5R6. In some embodiments, the plurality of factors comprises Serum paraoxonase / arylesterase 2 (PON2). The human PON2 gene can be found at Entrez gene #5445. The human PON2 protein can be found at Uniprot ID Q15165. In some embodiments, the plurality of factors comprises Serine / threonine-protein phosphatase 2A 56 kDa regulatory subunit delta isoform (PPP2R5D). The human PPP2R5D gene can be found at Entrez gene #5528. The human PPP2R5D protein can be found at Uniprot ID QI 4738. In some embodiments, the plurality of factors comprises Fox-1 homolog A (RBFOX1). The human RBFOX1 gene can be found at Entrez gene #54715. The human RBFOX1 protein can be found at Uniprot ID Q9NWB1. In some embodiments, the plurality of factors comprises TIMP metallopeptidase inhibitor 1 (TIMP1). The human TIMP1 gene can be found at Entrez gene #7076. The human TIMP1 protein can be found at Uniprot ID P01033 or Q6FGX5. In some embodiments, the plurality of factors comprises Gem-associated protein 7 (GEMIN7). The human GEMIN7 gene can be found at Entrez gene #79760. Thehuman GEMIN7 protein can be found at Uniprot ID Q9H840. In some embodiments, the plurality of factors comprises Casein kinase I isoform alpha (CSNK1A1L). The human CSNK1 AIL gene can be found at Entrez gene #1452. The human CSNK1 AIL protein can be found at Uniprot ID P48729. In some embodiments, the plurality of factors comprises PHD finger protein 11 (PHF11). The human PHF11 gene can be found at Entrez gene # 51131. The human PHF11 protein can be found at Uniprot ID Q9UIL8. In some embodiments, the plurality of factors comprises Butyrophilin subfamily 2 member A2 (BTN2A2). The human BTN2A2 gene can be found at Entrez gene #10385. The human BTN2A2 protein can be found at Uniprot ID Q8WVV5. In some embodiments, the plurality of factors comprises S-phase kinase-associated protein 2 (SKP2). The human SKP2 gene can be found at Entrez gene #6502. The human SKP2 protein can be found at Uniprot ID Q13309. In some embodiments, the plurality of factors comprises Spermatogenesis- associated protein 16 (SPATA46). The human SPATA46 gene can be found at Entrez gene # 284680. The human SPATA46 protein can be found at Uniprot ID Q5T0L3. In some embodiments, the plurality of factors comprises Lin-7 homolog A (LIN7A). The human LIN7A gene can be found at Entrez gene #8825. The human LIN7A protein can be found at Uniprot ID 014910. In some embodiments, the plurality of factors comprises BLOC-1- related complex subunit 5 (BORCS5). The human BORCS5 gene can be found at Entrez gene # 118426. The human BORCS5 protein can be found at Uniprot ID Q969J3. In some embodiments, the plurality of factors comprises Arrestin domain-containing protein 5 (ARRDC5). The human ARRDC5 gene can be found at Entrez gene # (>45432. The human ARRDC5 protein can be found at Uniprot ID A6NEK1. In some embodiments, the plurality of factors comprises Choline-phosphate cytidylyltransferase A (PCYT1A). The human PCYT1A gene can be found at Entrez gene #5130. The human PCYT1A protein can be found at Uniprot ID P49585. In some embodiments, the plurality of factors comprises Phytanoyl-CoA dioxygenase, peroxisomal (PHYH). The human PHYH gene can be found at Entrez gene # 5264. The human PHYH protein can be found at Uniprot ID 014832. In some embodiments, the plurality of factors comprises Ankyrin repeat domaincontaining protein 63 (ANKRD63). The human ANKRD63 gene can be found at Entrez gene #100131244. The human ANKRD63 protein can be found at Uniprot ID C9JTQ0. In some embodiments, the plurality of factors comprises Variable charge X-linked protein 1 (VCX). The human VCX gene can be found at Entrez gene #26609. The human VCX protein can be found at Uniprot ID Q9H320. In some embodiments, the plurality of factors comprises Protein N-terminal asparagine amidohydrolase (NTAN1). The human NTAN1gene can be found at Entrez gene #123803. The human NTAN1 protein can be found at Uniprot ID Q96AB6. In some embodiments, the plurality of factors comprises StAR-related lipid transfer domain protein 7 (STARD7). The human STARD7 gene can be found at Entrez gene #56910. The human STARD7 protein can be found at Uniprot ID Q9NQZ5. In some embodiments, the plurality of factors comprises Apolipoprotein L2 (APOL2). The human APOL2 gene can be found at Entrez gene #23780. The human APOL2 protein can be found at Uniprot ID Q9BQE5. In some embodiments, the plurality of factors comprises Fms- related tyrosine kinase 4 (FLT4). The human FLT4 gene can be found at Entrez gene #2324. The human FLT4 protein can be found at Uniprot ID P35916. In some embodiments, the plurality of factors comprises CapZ-interacting protein (RCSD1). The human RCSD1 gene can be found at Entrez gene # 92241. The human RCSD1 protein can be found at Uniprot ID Q6JBY9. In some embodiments, the plurality of factors comprises INTS3 and NABP interacting protein (INIP). The human INIP gene can be found at Entrez gene # 58493. The human INIP protein can be found at Uniprot ID Q9NRY2. In some embodiments, the plurality of factors comprises Vimentin-type intermediate filament-associated coiled-coil protein (VMAC). The human VMAC gene can be found at Entrez gene # 400673. The human VMAC protein can be found at Uniprot ID Q2NL98. In some embodiments, the plurality of factors comprises Xaa-Pro aminopeptidase 3 (XPNPEP3). The human XPNPEP3 gene can be found at Entrez gene # 63929. The human XPNPEP3 protein can be found at Uniprot ID Q9NQH7. In some embodiments, the plurality of factors comprises Interferon epsilon (IFNE). The human IFNE gene can be found at Entrez gene # 338376. The human IFNE protein can be found at Uniprot ID Q80ZF2. In some embodiments, the plurality of factors comprises Negative elongation factor A (NELFA). The human NELFA gene can be found at Entrez gene # 7469. The human NELFA protein can be found at Uniprot ID Q8BG30. In some embodiments, the plurality of factors comprises Lysine demethylase 8 (KDM8). The human KDM8 gene can be found at Entrez gene #79831. The human KDM8 protein can be found at Uniprot ID Q8N371. In some embodiments, the plurality of factors comprises Nuclear cap-binding protein complex (NCBP1). The human NCBP1 gene can be found at Entrez gene #4686. The human NCBP1 protein can be found at Uniprot ID Q56A27. In some embodiments, the plurality of factors comprises Upstream stimulatory factor 2 (USF2). The human USF2 gene can be found at Entrez gene #7392. The human USF2 protein can be found at Uniprot ID Q15853. In some embodiments, the plurality of factors comprises Leucine-rich repeat-containing protein 75A (LRRC75A). The humanLRRC75A gene can be found at Entrez gene # 388341. The human LRRC75A protein canbe found at Uniprot ID Q8NAA5. In some embodiments, the plurality of factors comprises Amyloid P component, serum (APCS). The human APCS gene can be found at Entrez gene #325. The human APCS protein can be found at Uniprot ID P02743. In some embodiments, the plurality of factors comprises l-Phosphatidylinositol-4,5-bisphosphate phosphodiesterase delta- 1 (PLCD1). The human PLCD1 gene can be found at Entrez gene #5333. The human PLCD1 protein can be found at Uniprot ID P51178. In some embodiments, the plurality of factors comprises Espin (ESPN). The human ESPN gene can be found at Entrez gene #83715. The human ESPN protein can be found at Uniprot ID Q5JYL1. In some embodiments, the plurality of factors comprises DNA-binding protein RFX5 (RFX5). The human RFX5 gene can be found at Entrez gene #5993. The human RFX5 protein can be found at Uniprot ID P48382. In some embodiments, the plurality of factors comprises Ribosomal protein S6 kinase beta-2 (RPS6KB2). The human RPS6KB2 gene can be found at Entrez gene # 6199. The human RPS6KB2 protein can be found at Uniprot ID Q9UBS0. In some embodiments, the plurality of factors comprises BOS complex subunit N0M02 (N0M02). The human N0M02 gene can be found at Entrez gene # 283820. The human N0M02 protein can be found at Uniprot ID Q5JPE7. In some embodiments, the plurality of factors comprises Transcription elongation factor A protein-like 2 (TCEAL2). The human TCEAL2 gene can be found at Entrez gene # 140597. The human TCEAL2 protein can be found at Uniprot ID Q9H3H9. In some embodiments, the plurality of factors comprises Carboxylesterase 3 (CES3). The human CES3 gene can be found at Entrez gene #23491. The human CES3 protein can be found at Uniprot ID Q6UWW8. In some embodiments, the plurality of factors comprises Dual specificity tyrosine-phosphorylation- regulated kinase 1 A (DYRK1 A). The human DYRK1 A gene can be found at Entrez gene #1859. The human DYRK1A protein can be found at Uniprot ID Q13627. In some embodiments, the plurality of factors comprises Cytochrome P450 2C19 (CYP2C19). The human CYP2C19 gene can be found at Entrez gene #1557. The human CYP2C19 protein can be found at Uniprot ID P33261. In some embodiments, the plurality of factors comprises Complement factor I (CFI). The human CFI gene can be found at Entrez gene #3426. The human CFI protein can be found at Uniprot ID P05156 or Q8WW88. In some embodiments, the plurality of factors comprises Insulin-like growth factor-binding protein 3 (IGFBP3). The human IGFBP3 gene can be found at Entrez gene #3486. The human IGFBP3 protein can be found at Uniprot ID P17936. In some embodiments, the plurality of factors comprises Interleukin 6 (IL6). The human IL6 gene can be found at Entrez gene #3569. The human IL6 protein can be found at Uniprot ID P05231. In some embodiments, the plurality of factorscomprises Leptin (LEP). The human LEP gene can be found at Entrez gene #3952. The human LEP protein can be found at Uniprot ID P41159. In some embodiments, the plurality of factors comprises CREB -regulated transcription coactivator 3 (CRTC3). The human CRTC3 gene can be found at Entrez gene #64784. The human CRTC3 protein can be found at Uniprot ID Q6UUV7. In some embodiments, the plurality of factors comprises Vascular endothelial growth factor A (VEGFA). The human VEGFA gene can be found at Entrez gene #7422. The human VEGFA protein can be found at Uniprot ID Pl 5692. In some embodiments, the plurality of factors comprises Interleukin- 1 receptor accessory protein (IL1RAP). The human IL1RAP gene can be found at Entrez gene #3556. The human IL1RAP protein can be found at Uniprot ID Q9NPH3. In some embodiments, the plurality of factors comprises Hepatocyte growth factor (HGF). The human HGF gene can be found at Entrez gene #3082. The human HGF protein can be found at Uniprot ID P14210. In some embodiments, the plurality of factors comprises Phospholipase A2, membrane associated (PLA2G2A). The human PLA2G2A gene can be found at Entrez gene #5320. The human PLA2G2A protein can be found at Uniprot ID P 14555. In some embodiments, the plurality of factors comprises Chemokine (C-C motif) ligand 25 (CCL25). The human CCL25 gene can be found at Entrez gene #6370. The human CCL25 protein can be found at Uniprot ID 015444. In some embodiments, the plurality of factors comprises serpin family A member 7 (SERPINA7). The human SERPINA7 gene can be found at Entrez gene #6906. The human SERPINA7 protein can be found at Uniprot ID P05543. In some embodiments, the plurality of factors comprises Cytochrome P450 reductase (POR). The human POR gene can be found at Entrez gene #5447. The human POR protein can be found at Uniprot ID Pl 6435. In some embodiments, the plurality of factors comprises CCN family member 3 (CCN3). The human CCN3 gene can be found at Entrez gene # 4856. The human CCN3 protein can be found at Uniprot ID P48745. In some embodiments, the plurality of factors comprises Hemopexin (HPX). The human HPX gene can be found at Entrez gene #3263. The human HPX protein can be found at Uniprot ID P02790. In some embodiments, the plurality of factors comprises Insulin-like growth factor-binding protein 1 (IGFBP1). The human IGFBP1 gene can be found at Entrez gene #3484. The human IGFBP1 protein can be found at Uniprot ID P08833. In some embodiments, the plurality of factors comprises matrix metalloproteinase-3 (MMP3). The human MMP3 gene can be found at Entrez gene #4314. The human MMP3 protein can be found at Uniprot ID P08254. In some embodiments, the plurality of factors comprises Fibrinogen alpha chain (FGA). The human FGA gene can be found at Entrez gene #2243. The human FGA protein can be found at Uniprot ID P02671. In some embodiments,the plurality of factors comprises Fibrinogen beta chain (FGB). The human FGB gene can be found at Entrez gene #2244. The human FGB protein can be found at Uniprot ID P02675. In some embodiments, the plurality of factors comprises Fibrinogen gamma chain (FGG). The human FGG gene can be found at Entrez gene #2266. The human FGG protein can be found at Uniprot ID P02679. In some embodiments, the plurality of factors comprises Basal cell adhesion molecule (BCAM). The human BCAM gene can be found at Entrez gene #4059. The human BCAM protein can be found at Uniprot ID P50895. In some embodiments, the plurality of factors comprises Kunitz-type protease inhibitor 1 (SPINT1). The human SPINT1 gene can be found at Entrez gene #6692. The human SPINT1 protein can be found at Uniprot ID 043278. In some embodiments, the plurality of factors comprises Histone acetyltransferase 1 (HAT1). The human HAT1 gene can be found at Entrez gene #8520. The human HAT1 protein can be found at Uniprot ID 014929. In some embodiments, the plurality of factors comprises Growth hormone receptor (GHR). The human GHR gene can be found at Entrez gene #2690. The human GHR protein can be found at Uniprot ID P10912. In some embodiments, the plurality of factors comprises Complement factor properdin (CFP). The human CFP gene can be found at Entrez gene #5199. The human CFP protein can be found at Uniprot ID P27918. In some embodiments, the plurality of factors comprises Contactin 1 (CNTN1). The human CNTN1 gene can be found at Entrez gene #1272. The human CNTN1 protein can be found at Uniprot ID Q12860. In some embodiments, the plurality of factors comprises Alpha 2-antiplasmin (SERPINF2). The human SERPINF2 gene can be found at Entrez gene #5345. The human SERPINF2 protein can be found at Uniprot ID P08697. In some embodiments, the plurality of factors comprises Interleukin- 19 (IL19). The human IL19 gene can be found at Entrez gene # 29949. The human IL19 protein can be found at Uniprot ID Q9UHD0. In some embodiments, the plurality of factors comprises Myoglobin (MB). The human MB gene can be found at Entrez gene #4151. The human MB protein can be found at Uniprot ID P02144. In some embodiments, the plurality of factors comprises Immunoglobulin heavy constant mu (IGHM). The human IGHM gene can be found at Entrez gene #3507. The human IGHM protein can be found at Uniprot ID PO 1871. In some embodiments, the plurality of factors comprises Lipopolysaccharide binding protein (LBP). The human LBP gene can be found at Entrez gene #3929. The human LBP protein can be found at Uniprot ID P18428. In some embodiments, the plurality of factors comprises N-acylethanolamine-hydrolyzing acid amidase (NAAA). The human NAAA gene can be found at Entrez gene # 27163. The human NAAA protein can be found at Uniprot ID Q02083. In some embodiments, the plurality offactors comprises Hyaluronan and proteoglycan link protein 1 (HAPLN1). The human HAPLN1 gene can be found at Entrez gene #1404. The human HAPLN1 protein can be found at Uniprot ID P10915. In some embodiments, the plurality of factors comprises Iduronate 2-sulfatase (IDS). The human IDS gene can be found at Entrez gene # 3423. The human IDS protein can be found at Uniprot ID P22304. In some embodiments, the plurality of factors comprises Nidogen-1 (NIDI). The human NIDI gene can be found at Entrez gene #4811. The human NIDI protein can be found at Uniprot ID Pl 4543. In some embodiments, the plurality of factors comprises Aggrecan (ACAN). The human ACAN gene can be found at Entrez gene #176. The human ACAN protein can be found at Uniprot ID P16112. In some embodiments, the plurality of factors comprises Transforming growth factor, beta-induced, 68kDa (TGFBI). The human TGFBI gene can be found at Entrez gene #7045. The human TGFBI protein can be found at Uniprot ID Q15582. In some embodiments, the plurality of factors comprises Delta-like 4 (DLL4). The human DLL4 gene can be found at Entrez gene #54567. The human DLL4 protein can be found at Uniprot ID Q9NR61. In some embodiments, the plurality of factors comprises Fc fragment of IgG, low affinity Illb, receptor (FCGR3B). The human FCGR3B gene can be found at Entrez gene #2215. The human FCGR3B protein can be found at Uniprot ID 075015. In some embodiments, the plurality of factors comprises Aminoacylase-1 (ACY1). The human ACY1 gene can be found at Entrez gene #95. The human ACY 1 protein can be found at Uniprot ID Q03154. In some embodiments, the plurality of factors comprises Integrin binding sialoprotein (IBSP). The human IBSP gene can be found at Entrez gene #3381. The human IBSP protein can be found at Uniprot ID P21815. In some embodiments, the plurality of factors comprises Kallistatin (SERPINA4). The human SERPINA4 gene can be found at Entrez gene #5267. The human SERPINA4 protein can be found at Uniprot ID P29622. In some embodiments, the plurality of factors comprises Periostin (POSTN). The human POSTN gene can be found at Entrez gene #10631. The human POSTN protein can be found at Uniprot ID Q15063. In some embodiments, the plurality of factors comprises E-selectin (SELE). The human SELE gene can be found at Entrez gene #6401. The human SELE protein can be found at Uniprot ID P16581. In some embodiments, the plurality of factors comprises P2 microglobulin (B2M). The human B2M gene can be found at Entrez gene #567. The human B2M protein can be found at Uniprot ID P61769. In some embodiments, the plurality of factors comprises Hepcidin (HAMP). The human HAMP gene can be found at Entrez gene #57817. The human HAMP protein can be found at Uniprot ID P81172. In some embodiments, the plurality of factors comprises Alpha-1 antitrypsin (SERPINA1). The human SERPINA1 gene can befound at Entrez gene #5265. The human SEREIN A l protein can be found at Uniprot ID P01009. In some embodiments, the plurality of factors comprises alpha-2 -HS-gly coprotein (AHSG). The human AHSG gene can be found at Entrez gene #197. The human AHSG protein can be found at Uniprot ID P02765. In some embodiments, the plurality of factors comprises Brain-type creatine kinase (CKB). The human CKB gene can be found at Entrez gene #1152. The human CKB protein can be found at Uniprot ID P12277. In some embodiments, the plurality of factors comprises Creatine kinase, muscle (CKM). The human CKM gene can be found at Entrez gene #1158. The human CKM protein can be found at Uniprot ID P06732. In some embodiments, the plurality of factors comprises Protein C (PROC). The human PROC gene can be found at Entrez gene #5624. The human PROC protein can be found at Uniprot ID P04070. In some embodiments, the plurality of factors comprises Angiopoietin-like 4 (ANGPTL4). The human ANGPTL4 gene can be found at Entrez gene #51129. The human ANGPTL4 protein can be found at Uniprot ID Q9BY76. In some embodiments, the plurality of factors comprises Methyl -CpG-binding domain protein 4 (MBD4). The human MBD4 gene can be found at Entrez gene #8930. The human MBD4 protein can be found at Uniprot ID 095243. In some embodiments, the plurality of factors comprises 26S proteasome non-ATPase regulatory subunit 7 (PSMD7). The human PSMD7 gene can be found at Entrez gene #5713. The human PSMD7 protein can be found at Uniprot ID P51665. In some embodiments, the plurality of factors comprises immunoglobulin heavy constant epsilon (IGHE). The human IGHE gene can be found at Entrez gene #3497. The human IGHE protein can be found at Uniprot ID P01854. In some embodiments, the plurality of factors comprises C-X-C motif chemokine ligand 10 (CXCL10). The human CXCL10 gene can be found at Entrez gene #3627. The human CXCL10 protein can be found at Uniprot ID P02778. In some embodiments, the plurality of factors comprises Plasma kallikrein (KLKB1). The human KLKB1 gene can be found at Entrez gene # 3818. The human KLKB1 protein can be found at Uniprot ID P03952. In some embodiments, the plurality of factors comprises Factor H (CFH). The human CFH gene can be found at Entrez gene #3075. The human CFH protein can be found at Uniprot ID P08603. In some embodiments, the plurality of factors comprises Prefoldin subunit 5 (PFDN5). The human PFDN5 gene can be found at Entrez gene #5204. The human PFDN5 protein can be found at Uniprot ID Q99471. In some embodiments, the plurality of factors comprises RNA- binding protein 39 (RBM39). The human RBM39 gene can be found at Entrez gene #9584 The human RBM39 protein can be found at Uniprot ID Q14498 or Q5QP23. In some embodiments, the plurality of factors comprises dCTP pyrophosphatase 1 (DCTPP1). Thehuman DCTPP1 gene can be found at Entrez gene #79077. The human DCTPP1 protein can be found at Uniprot ID Q9H773. In some embodiments, the plurality of factors comprises Brain-specific serine protease 4 (PRSS22). The human PRSS22 gene can be found at Entrez gene #64063. The human PRSS22 protein can be found at Uniprot ID Q9GZN4. In some embodiments, the plurality of factors compriseskynureninase (KYNU). The human KYNU gene can be found at Entrez gene # 8942. The human KYNU protein can be found at Uniprot ID QI 6719. In some embodiments, the plurality of factors comprises Interleukin 6 (IL6). The human IL6 gene can be found at Entrez gene #3569. The human IL6 protein can be found at Uniprot ID P05231. In some embodiments, the plurality of factors comprises Transcortin (SERPINA6). The human SERPINA6 gene can be found at Entrez gene #866. The human SERPINA6 protein can be found at Uniprot ID P08185. In some embodiments, the plurality of factors comprises Inter-alpha-trypsin inhibitor heavy chain H4 (ITH44). The human ITIH4 gene can be found at Entrez gene # 3700. The human ITH44 protein can be found at Uniprot ID Q 14624. In some embodiments, the plurality of factors comprises Stratifin (SFN). The human SFN gene can be found at Entrez gene # 2810. The human SFN protein can be found at Uniprot ID P31947. In some embodiments, the plurality of factors comprises Chemokine (C-C motif) ligand 7 (CCL7). The human CCL7 gene can be found at Entrez gene # 6354. The human CCL7 protein can be found at Uniprot ID P80098. In some embodiments, the plurality of factors comprises Lysozyme C (LYZ). The human LYZ gene can be found at Entrez gene # 4069. The human LYZ protein can be found at Uniprot ID P61626. In some embodiments, the plurality of factors comprises Collagenase 3 (MMP13). The human MMP13 gene can be found at Entrez gene # 4322. The human MMP13 protein can be found at Uniprot ID P45452. In some embodiments, the plurality of factors comprises Stanniocalcin-1 (STC1). The human STC1 gene can be found at Entrez gene # 6781. The human STC1 protein can be found at Uniprot ID P52823. In some embodiments, the plurality of factors comprises Macrophage-capping protein (CAPG). The human CAPG gene can be found at Entrez gene # 822. The human CAPG protein can be found at Uniprot ID P40121. In some embodiments, the plurality of factors comprises Peptidase inhibitor 3 (PI3). The human PI3 gene can be found at Entrez gene # 5266. The human PI3 protein can be found at Uniprot ID Pl 9957. In some embodiments, the plurality of factors comprises Glypican-5 (GPC5). The human GPC5 gene can be found at Entrez gene #2262. The human GPC5 protein can be found at Uniprot ID P78333. In some embodiments, the plurality of factors comprises Histidine-rich glycoprotein (HRG). The human HRG gene can be found at Entrez gene # 3273. The human HRG protein can be foundat Uniprot ID P04196. In some embodiments, the plurality of factors comprises Mammaglobin-B (SCGB2A1). The human SCGB2A1 gene can be found at Entrez gene # 4246. The human SCGB2A1 protein can be found at Uniprot ID 075556. In some embodiments, the plurality of factors comprises NAD-dependent deacetylase sirtuin 2 (SIRT2). The human SIRT2 gene can be found at Entrez gene # 22933. The human SIRT2 protein can be found at Uniprot ID Q8IXJ6. In some embodiments, the plurality of factors comprises Tumor necrosis factor-inducible gene 6 protein (TNFAIP6). The human TNFAIP6 gene can be found at Entrez gene #2934. The human TNFAIP6 protein can be found at Uniprot ID P06396. In some embodiments, the plurality of factors comprises CMRF35-like molecule 6 (CD300C). The human CD300C gene can be found at Entrez gene# 10871. The human CD300C protein can be found at Uniprot ID Q08708. In some embodiments, the plurality of factors comprises Transmembrane glycoprotein NMB (GPNMB). The human GPNMB gene can be found at Entrez gene # 10457. The human GPNMB protein can be found at Uniprot ID Q14956. In some embodiments, the plurality of factors comprises Keratin 18 (KRT18). The human KRT 18 gene can be found at Entrez gene# 3875. The human KRT18 protein can be found at Uniprot ID P05783. In some embodiments, the plurality of factors comprises tumor necrosis factor superfamily member 14 (TNFSF14). The human TNFSF14 gene can be found at Entrez gene # 8740. The human TNFSF14 protein can be found at Uniprot ID 043557. In some embodiments, the plurality of factors comprises Leptin receptor (LEPR). The human LEPR gene can be found at Entrez gene # 3953. The human LEPR protein can be found at Uniprot ID P48357. In some embodiments, the plurality of factors comprises Protein kinase C gamma type (PRKCG). The human PRKCG gene can be found at Entrez gene # 5582. The human PRKCG protein can be found at Uniprot ID P05129. In some embodiments, the plurality of factors comprises Fibrinogen-like protein 1 (FGL1). The human FGLlgene can be found at Entrez gene # 2267. The human FGL1 protein can be found at Uniprot ID Q08830. In some embodiments, the plurality of factors comprises Peptidoglycan recognition protein 2 (PGLYRP2). The human PGLYRP2 gene can be found at Entrez gene # 114770. The human PGLYRP2 protein can be found at Uniprot ID Q96PD5. In some embodiments, the plurality of factors comprises Neuropeptide FF (NPFF). The human NPFF gene can be found at Entrez gene # 8620. The human NPFF protein can be found at Uniprot ID 015130. In some embodiments, the plurality of factors comprises Microfibril-associated glycoprotein 4 (MFAP4). The human MFAP4 gene can be found at Entrez gene # 4239. The human MFAP4 protein can be found at Uniprot ID P55083. In some embodiments, the plurality of factors comprisesProtein disulfide-isomerase (TMX3). The human TMX3 gene can be found at Entrez gene # 54495. The human TMX3 protein can be found at Uniprot ID Q96JJ7. In some embodiments, the plurality of factors comprises Glucosidase 2 subunit beta (PRKCSH). The human PRKCSH gene can be found at Entrez gene # 5589. The human PRKCSH protein can be found at Uniprot ID P14314. In some embodiments, the plurality of factors comprises Beta- defensin 112 (DEFBI 12). The human DEFBI 12 gene can be found at Entrez gene # 245915. The human DEFBI 12 protein can be found at Uniprot ID Q30KQ8. In some embodiments, the plurality of factors comprises Semaphorin-4D (SEMA4D). The human SEMA4D gene can be found at Entrez gene # 10507. The human SEMA4D protein can be found at Uniprot ID Q92854. In some embodiments, the plurality of factors comprises Lysophosphatidic acid phosphatase type 6 (ACP6). The human ACP6 gene can be found at Entrez gene # 51205. The human ACP6 protein can be found at Uniprot ID Q9NPH0. In some embodiments, the plurality of factors comprises Alpha-fetoprotein (AFP). The human AFP gene can be found at Entrez gene # 174. The human AFP protein can be found at Uniprot ID P02771. In some embodiments, the plurality of factors comprises Nerve growth factor (NGF). The human NGF gene can be found at Entrez gene # 4803. The human NGF protein can be found at Uniprot ID P01138. In some embodiments, the plurality of factors comprises Ferritin heavy chain (FTH1). The human FTH1 gene can be found at Entrez gene # 2495. The human FTH1 protein can be found at Uniprot ID P02794. In some embodiments, the plurality of factors comprises Ferritin light chain (FTL). The human FTL gene can be found at Entrez gene # 2512. The human FTL protein can be found at Uniprot ID P02792. In some embodiments, the plurality of factors comprises Dermokine (DMKN). The human DMKN gene can be found at Entrez gene # 93099. The human DMKN protein can be found at Uniprot ID Q6E0U4. In some embodiments, the plurality of factors comprises EPH receptor A10 (EPHA10). The human EPHA10 gene can be found at Entrez gene # 284656. The human EPHA10 protein can be found at Uniprot ID Q5JZY3. In some embodiments, the plurality of factors comprises Chordin-like protein 2 (CHRDL2). The human CHRDL2 gene can be found at Entrez gene # 25884. The human CHRDL2 protein can be found at Uniprot ID Q6WN34. In some embodiments, the plurality of factors comprises Tumor protein P53 (TP53). The human TP53 gene can be found at Entrez gene # 7157. The human TP53 protein can be found at Uniprot ID P04637. In some embodiments, the plurality of factors comprises Diamine oxidase [copper-containing] (AOC1). The human AOC1 gene can be found at Entrez gene # 26. The human AOC1 protein can be found at Uniprot ID Pl 9801. In some embodiments, the plurality of factors comprises Interferon alpha-8 (IFNA8). The humanIFNA8 gene can be found at Entrez gene # 3445. The human IFNA8 protein can be found at Uniprot ID P32881. In some embodiments, the plurality of factors comprises Chorionic somatomammotropin hormone 1 (CSH1). The human CSH1 gene can be found at Entrez gene # 1442. The human CSH1 protein can be found at Uniprot ID P0DML2. In some embodiments, the plurality of factors comprises Chorionic somatomammotropin hormone 2 (CSH2). The human CSH2 gene can be found at Entrez gene # 1443. The human CSH2 protein can be found at Uniprot ID P0DML3. In some embodiments, the plurality of factors comprises Tenascin C (TNC). The human TNC gene can be found at Entrez gene # 3371. The human TNC protein can be found at Uniprot ID P24821. In some embodiments, the plurality of factors comprises Phospholipid transfer protein (PL TP). The human PL TP gene can be found at Entrez gene # 5360. The human PL TP protein can be found at Uniprot ID P55058. In some embodiments, the plurality of factors comprises cellular communication network factor 1 (CCN1). The human CCN1 gene can be found at Entrez gene # 3491. The human CCN1 protein can be found at Uniprot ID 000622. In some embodiments, the plurality of factors comprises Calsyntenin-3 (CLSTN3). The human CLSTN3 gene can be found at Entrez gene # 9746. The human CLSTN3 protein can be found at Uniprot ID Q9BQT9. In some embodiments, the plurality of factors comprises Oncoprotein-induced transcript 3 protein (OIT3). The human OIT3 gene can be found at Entrez gene # 170392. The human OIT3 protein can be found at Uniprot ID Q8WWZ8. In some embodiments, the plurality of factors comprises Inactive glutathione hydrolase 2 (GGT2). The human GGT2 gene can be found at Entrez gene # 102724197. The human GGT2 protein can be found at Uniprot ID P36268. In some embodiments, the plurality of factors comprises Fibromodulin (FMOD). The human FMOD gene can be found at Entrez gene # 2331. The human FMOD protein can be found at Uniprot ID Q06828. In some embodiments, the plurality of factors comprises Putative uncharacterized protein IRX2-DT (C5orf38). The human C5orf 8 gene can be found at Entrez gene # 153571. The human C5orf 8 protein can be found at Uniprot ID Q86SI9. In some embodiments, the plurality of factors comprises von Willebrand factor A domain-containing protein 1 (VWA1). The human VWA1 gene can be found at Entrez gene # 64856. The human VWA1 protein can be found at Uniprot ID Q6PCB0. In some embodiments, the plurality of factors comprises Inhibin beta C chain (INHBC). The human INHBC gene can be found at Entrez gene # 3626. The human INHBC protein can be found at Uniprot ID P55103. In some embodiments, the plurality of factors comprises Adhesion G protein-coupled receptor F5 (ADGRF5). The human ADGRF5 gene can be found at Entrez gene # 221395. The human ADGRF5 protein can be found at Uniprot ID Q8IZF2. In someembodiments, the plurality of factors comprises Complement Clq-like protein 2 (C1QL2). The human C1QL2 gene can be found at Entrez gene # 165257. The human C1QL2 protein can be found at Uniprot ID Q7Z5L3. In some embodiments, the plurality of factors comprises Prenylcysteine oxidase 1 (PCY0X1). The human PCY0X1 gene can be found at Entrez gene # 51449. The human PCY0X1 protein can be found at Uniprot ID Q9UHG3. In some embodiments, the plurality of factors comprises Amine oxidase, copper containing 2 (A0C2). The human A0C2 gene can be found at Entrez gene # 314. The human A0C2 protein can be found at Uniprot ID 075106. In some embodiments, the plurality of factors comprises Complement factor H-related protein 4 (CFHR4). The human CFHR4 gene can be found at Entrez gene # 10877. The human CFHR4 protein can be found at Uniprot ID Q92496. In some embodiments, the plurality of factors comprises Leucine rich repeat containing 15 (LRRC15). The human LRRC15 gene can be found at Entrez gene # 131578. The human LRRC15 protein can be found at Uniprot ID Q8TF66. In some embodiments, the plurality of factors comprises Periostin (POSTN). The human POSTN gene can be found at Entrez gene # 10631. The human POSTN protein can be found at Uniprot ID Q 15063. In some embodiments, the plurality of factors comprises Ubiquitin-conjugating enzyme E2 JI (UBE2J1). The human UBE2J1 gene can be found at Entrez gene # 51465. The human UBE2J1 protein can be found at Uniprot ID Q9Y385. In some embodiments, the plurality of factors comprises GDNF family receptor alpha like (GFRAL). The human GFRAL gene can be found at Entrez gene # 389400. The human GFRAL protein can be found at Uniprot ID Q6UXV0. In some embodiments, the plurality of factors comprises Insulin-like growth factor 2 (IGF2). The human IGF2 gene can be found at Entrez gene # 3481. The human IGF2 protein can be found at Uniprot ID P01344. In some embodiments, the plurality of factors comprises Leukocyte immunoglobulin-like receptor subfamily B member 5 (LILRB5). The human LILRB5 gene can be found at Entrez gene # 10990. The human LILRB5 protein can be found at Uniprot ID 075023. In some embodiments, the plurality of factors comprises Leukocyte immunoglobulin-like receptor subfamily A member 6 (LILRA6). The human LILRA6 gene can be found at Entrez gene # 79168. The human LILRA6 protein can be found at Uniprot ID Q6PI73. In some embodiments, the plurality of factors comprises Apolipoprotein A-II (APOA2). The human APOA2 gene can be found at Entrez gene # 336. The human APOA2 protein can be found at Uniprot ID P02652. In some embodiments, the plurality of factors comprises von Willebrand factor A domain-containing protein 2 (VWA2). The human VWA2 gene can be found at Entrez gene # 340706. The human VWA2 protein can be found at Uniprot ID Q5GFL6. In some embodiments, the plurality of factorscomprises Protein DEPPI (DEPPI). The human DEPPI gene can be found at Entrez gene # 11067. The human DEPPI protein can be found at Uniprot ID Q9NTK1. In some embodiments, the plurality of factors comprises Complement Clq tumor necrosis factor- related protein 3 (C1QTNF3). The human C1QTNF3 gene can be found at Entrez gene # 114899. The human C1QTNF3 protein can be found at Uniprot ID Q9BXJ4. In some embodiments, the plurality of factors comprises Serpin A9 (SERPINA9). The human SERPINA9 gene can be found at Entrez gene # 327657. The human SERPINA9 protein can be found at Uniprot ID Q86WD7. In some embodiments, the plurality of factors comprises Complement factor H-related protein 5 (CFHR5). The human CFHR5 gene can be found at Entrez gene # 81494. The human CFHR5 protein can be found at Uniprot ID Q9BXR6. In some embodiments, the plurality of factors comprises Disks large homolog 3 (DLG3). The human DLG3 gene can be found at Entrez gene # 1741. The human DLG3 protein can be found at Uniprot ID Q92796. In some embodiments, the plurality of factors comprises Glycolipid transfer protein domain-containing protein 2 (GLTPD2). The human GLTPD2 gene can be found at Entrez gene # 388323. The human GLTPD2 protein can be found at Uniprot ID A6NH11. In some embodiments, the plurality of factors comprises Hemoglobin subunit theta-1 (HBQ1). The human HBQ1 gene can be found at Entrez gene # 3049. The human HBQ1 protein can be found at Uniprot ID P09105. In some embodiments, the plurality of factors comprises Ectonucleoside triphosphate diphosphohydrolase- 1 (ENTPD1). The human ENTPD1 gene can be found at Entrez gene # 953. The human ENTPD1 protein can be found at Uniprot ID P49961. In some embodiments, the plurality of factors comprises Angiogenic factor with G patch and FHA domains 1 (AGGF1). The human AGGF1 gene can be found at Entrez gene # 55109. The human AGGF1 protein can be found at Uniprot ID Q8N302. In some embodiments, the plurality of factors comprises Neuregulin 2 (NRG2). The human NRG2 gene can be found at Entrez gene # 9542. The human NRG2 protein can be found at Uniprot ID 014511. In some embodiments, the plurality of factors comprises Spondin 2 (SPON2). The human SPON2 gene can be found at Entrez gene # 10417. The human SPON2 protein can be found at Uniprot ID Q9BUD6. In some embodiments, the plurality of factors comprises Protein FAM241B (FAM241B). The human FAM241B gene can be found at Entrez gene # 219738. The human FAM241B protein can be found at Uniprot ID Q96D05. In some embodiments, the plurality of factors comprises Junctional adhesion molecule-like (J AML). The human JAML gene can be found at Entrez gene # 120425. The human JAML protein can be found at Uniprot ID Q86YT9. In some embodiments, the plurality of factors comprises Butyrylcholinesterase (BCHE). Thehuman BCHE gene can be found at Entrez gene # 590. The human BCHE protein can be found at Uniprot ID P06276. In some embodiments, the plurality of factors comprises Transmembrane glycoprotein NMB (GPNMB). The human GPNMB gene can be found at Entrez gene # 10457. The human GPNMB protein can be found at Uniprot ID Q14956. In some embodiments, the plurality of factors comprises Apolipoprotein D (APOD). The human APOD gene can be found at Entrez gene # 347. The human APOD protein can be found at Uniprot ID P05090. In some embodiments, the plurality of factors comprises Deltalike protein 1 (DLL1). The human DLL1 gene can be found at Entrez gene # 28514. The human DLL1 protein can be found at Uniprot ID 000548. In some embodiments, the plurality of factors comprises Platelet endothelial aggregation receptor 1 (PEAR1). The human PEAR1 gene can be found at Entrez gene # 375033. The human PEAR1 protein can be found at Uniprot ID Q5VY43. In some embodiments, the plurality of factors comprises R-spondin-4 (RSPO4). The human RSPO4 gene can be found at Entrez gene # 343637. The human RSPO4 protein can be found at Uniprot ID Q2I0M5. In some embodiments, the plurality of factors comprises Leptin (LEP). The human LEP gene can be found at Entrez gene # 3952. The human LEP protein can be found at Uniprot ID P41159. In some embodiments, the plurality of factors comprises ADP-ribosylation factor-like protein 8B (ARL8B). The human ARL8B gene can be found at Entrez gene # 55207. The human ARL8B protein can be found at Uniprot ID Q9NVJ2. In some embodiments, the plurality of factors comprises Protocadherin-10 (PCDH10). The human PCDH10 gene can be found at Entrez gene # 57575. The human PCDH10 protein can be found at Uniprot ID Q9P2E7. In some embodiments, the plurality of factors comprises Microfibrillar-associated protein 3- like (MFAP3L). The human MFAP3L gene can be found at Entrez gene # 9848. The human MFAP3L protein can be found at Uniprot ID 075121. In some embodiments, the plurality of factors comprises Monocyte differentiation antigen CD14 (CD14). The human CD14 gene can be found at Entrez gene # 929. The human CD14 protein can be found at Uniprot ID P08571. In some embodiments, the plurality of factors comprises Collagen alpha-l(XV) chain (COL15A1). The human COL15A1 gene can be found at Entrez gene # 1306. The human COL15A1 protein can be found at Uniprot ID P39059. In some embodiments, the plurality of factors comprises Hepatitis A virus cellular receptor 1 (HAVCR1). The human HAVCR1 gene can be found at Entrez gene # 26762. The human HAVCR1 protein can be found at Uniprot ID Q96D42. In some embodiments, the plurality of factors comprises Rho guanine nucleotide exchange factor 10 (ARHGEF10). The human ARHGEF10 gene can be found at Entrez gene # 9639. The human ARHGEF10 protein can be found at Uniprot ID015013. In some embodiments, the plurality of factors comprises Mannosyl-oligosaccharide 1,2-alpha-mannosidase IB (MAN1A2). The human MAN1A2 gene can be found at Entrez gene # 10905. The human MAN1A2 protein can be found at Uniprot ID 060476. In some embodiments, the plurality of factors comprises Quinone oxidoreductase-like protein 1 (CRYZL1). The human CRYZL1 gene can be found at Entrez gene # 9946. The human CRYZL1 protein can be found at Uniprot ID 095825. In some embodiments, the plurality of factors comprises Tissue factor pathway inhibitor 2 (TFPI2). The human TFPI2 gene can be found at Entrez gene # 7980. The human TFPI2 protein can be found at Uniprot ID P48307. In some embodiments, the plurality of factors comprises Plexin domain-containing protein 1 (PLXDC1). The human PLXDC1 gene can be found at Entrez gene # 57125. The human PLXDC1 protein can be found at Uniprot ID Q8IUK5. In some embodiments, the plurality of factors comprises Lysosomal acid phosphatase (ACP2). The human ACP2 gene can be found at Entrez gene # 53. The human ACP2 protein can be found at Uniprot ID Pl 1117. In some embodiments, the plurality of factors comprises Biotinidase (BTD). The human BTD gene can be found at Entrez gene # 686. The human BTD protein can be found at Uniprot ID P43251. In some embodiments, the plurality of factors comprises Microfibrillar-associated protein 2 (MFAP2). The human MFAP2 gene can be found at Entrez gene # 4237. The human MFAP2 protein can be found at Uniprot ID P55001. In some embodiments, the plurality of factors comprises Inter-alpha-trypsin inhibitor heavy chain H2 (ITH42). The human ITIH2 gene can be found at Entrez gene # 3698. The human ITIH2 protein can be found at Uniprot ID Pl 9823. In some embodiments, the plurality of factors comprises EF-hand calcium-binding domain-containing protein 14 (EFCAB14). The human EFCAB14 gene can be found at Entrez gene # 9813. The human EFCAB14 protein can be found at Uniprot ID 075071. In some embodiments, the plurality of factors comprises Phospholipase Al member A (PLA1A). The human PLA1A gene can be found at Entrez gene # 51365. The human PLA1A protein can be found at Uniprot ID Q53H76. In some embodiments, the plurality of factors comprises Granzyme K (GZMK). The human GZMK gene can be found at Entrez gene # 3003. The human GZMK protein can be found at Uniprot ID P49863. In some embodiments, the plurality of factors comprises Y box binding protein 1 (YBX1). The human YBX1 gene can be found at Entrez gene # 4904. The human YBX1 protein can be found at Uniprot ID P67809. In some embodiments, the plurality of factors comprises Indoleamine-pyrrole 2,3-dioxygenase (IDO1). The human IDO1 gene can be found at Entrez gene # 3620. The human IDO1 protein can be found at Uniprot ID P14902. In some embodiments, the plurality of factors comprises NAD(P)H dehydrogenase[quinone] 1 (NQO1). The human NQO1 gene can be found at Entrez gene # 1728. The human NQO1 protein can be found at Uniprot ID P15559. In some embodiments, the plurality of factors comprises Testican-3 (SPOCK3). The human SPOCK3 gene can be found at Entrez gene # 50859. The human SPOCK3 protein can be found at Uniprot ID Q9BQ16. In some embodiments, the plurality of factors comprises NTF2-related export protein 1 (NXT1). The human NXT1 gene can be found at Entrez gene # 29107. The human NXT1 protein can be found at Uniprot ID Q9UKK6.

[0064] In some embodiments, the expression is protein expression. In some embodiments, the expression is secreted protein expression. In some embodiments, protein expression is soluble protein expression. In some embodiments, the expression is cellular protein expression. In some embodiments, the expression is membranal protein expression. In some embodiments, the expression is circulating protein expression. In some embodiments, the expression is relative expression. In some embodiments, the expression is absolute expression. In some embodiments, the expression is normalized expression. In some embodiments, expression level is concentration. In some embodiments, concentration is concentration level. It will be understood by a skilled artisan that when the presence of factor is measured in a liquid sample the expression can be provided as a concentration such as mg / ml or in arbitrary units according to the method of determining the factor’s expression. Arbitrary units can be selected from relative fluorescence unit (RFU) and Normalized Protein expression (NPX), or any other arbitrary units used as measurement of expression. The terms “expression” and “expression levels” are used herein interchangeably and refer to the amount of a gene product present in the sample. In some embodiments, the protein levels are determined. In some embodiments, the method comprises performing an assay on the biological sample to determine protein levels. In some embodiments, the assay is proteomic sequencing. In some embodiments, the assay is mass-spec. In some embodiments, the assay is a proteomics array. In some embodiments, determining comprises quantification of expression levels. In some embodiments, determining comprises normalization of expression levels. Determining of the expression level of the factor can be performed by any method known in the art. Methods of determining protein expression include, for example, antibody arrays, immunoblotting, immunohistochemistry, flow cytometry (FACS), Enzyme- Linked Immunosorbent Assay (ELISA), proximity extension assay (PEA), aptamer-based assays, proteomics arrays, proteome sequencing, flow cytometry (CyTOF), multiplex assays, mass spectrometry, reporter gene assays, and chromatography. In someembodiments, determining protein expression levels comprises ELISA. In some embodiments, the ELISA is next-generation ELISA (nELISA). In some embodiments, determining protein expression levels comprises protein array hybridization. In some embodiments, determining protein expression levels comprises mass-spectrometry quantification. In some embodiments, determining protein expression levels comprises PEA. In some embodiments, determining protein expression levels comprises aptamers. In some embodiments, determining protein expression levels is based on protein sequencing.

[0065] In some embodiments, the receiving factor expression levels is providing factor expression levels. In some embodiments, the receiving factor expression levels is determining factor expression levels. In some embodiments, determining is measuring. In some embodiments, the measuring is in a sample. In some embodiments, the expression levels were detected in a sample. In some embodiments, the sample is a biological sample. In some embodiments, the sample is provided by the subjects. In some embodiments, the sample is provided by the patients. In some embodiments, the sample is provided by the subject. In some embodiments, the sample is provided by the patient. In some embodiments, the sample is provided by a responder. In some embodiments, the sample is provided by a non-responder. In some embodiments, each subject of the population of responders provided a sample. In some embodiments, each subject of the population of non-responders provided a sample. In some embodiments, the sample is provided by a CB subject. In some embodiments, the sample is provided by a NCB subject. In some embodiments, each subject of the CB population provided a sample. In some embodiments, each subject of the NCB population provided a sample. In some embodiments, the sample is provided by a subject before receiving the therapy. In some embodiments, the sample is provided by a subject after receiving the therapy. In some embodiments, the sample is provided by the subject during receiving the therapy. In some embodiments, the method is a method of monitoring response and the sample is provided during receiving the therapy or after receiving the therapy. In some embodiments, the determining is directly in the sample. In some embodiments, the determining is in the unprocessed sample. In some embodiments, the determining is in a processed sample. In some embodiments, the method further comprises processing the sample. In some embodiments, processing comprises isolating proteins from the sample. In some embodiments, processing comprises isolating nucleic acids from the sample. In some embodiments, the nucleic acid is RNA. In some embodiments, the RNA is mRNA. In some embodiments, the processing comprises lysing cells in the sample. In some embodiments,processing comprises isolating DNA from the sample. In some embodiments, processing comprises isolating of cell free DNA (cfDNA).

[0066] As used herein, the terms “peptide”, "polypeptide", “proteoform” and "protein" are used interchangeably to refer to a polymer of amino acid residues. In another embodiment, the terms "peptide", "polypeptide", “proteoform” and "protein" as used herein encompass native peptides, peptidomimetics (typically including non-peptide bonds or other synthetic modifications) and the peptide analogues peptoids and semipeptoids or any combination thereof. In another embodiment, the peptides polypeptides and proteins described have modifications rendering them more stable while in the body or more capable of penetrating into cells. In one embodiment, the terms “peptide”, "polypeptide", “proteoform” and "protein" apply to naturally occurring amino acid polymers. In another embodiment, the terms “peptide”, "polypeptide", “proteoform” and "protein" apply to amino acid polymers in which one or more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid.

[0067] In some embodiments, the sample is a biological sample. In some embodiments, the sample is tissue. In some embodiments, the sample is fluid. In some embodiments, the fluid is a biological fluid. In some embodiments, the sample is from the subject. In some embodiments, the sample is not a tumor sample. In some embodiments, the sample is a tumor sample. In some embodiments, the sample is not a hematopoietic cancer and the sample is a blood sample. In some embodiments, the sample is a sample that does not comprise cancer cells. In some embodiments, the sample is a sample that comprises cancer cells. In some embodiments, a blood sample comprises a peripheral blood sample, serum sample and a plasma sample. In some embodiments, the sample is a plasma sample. In some embodiments, the sample is a serum sample. In some embodiments, processing comprises isolating plasma. In some embodiments, processing comprises isolating serum. In some embodiments, processing comprises isolating cells. In some embodiments, processing comprises isolating PBMCs. In some embodiments, the biological fluid is selected from blood, plasma, serum, lymph, cerebral spinal fluid, urine, feces, semen, tumor fluid and gastric fluid. In some embodiments, the sample obtained from the subject and the responders are the same type of sample. In some embodiments, the sample obtained from the subject and the responders are different types of samples. In some embodiments, the sample obtained from the subject and the non-responders are the same type of sample. In some embodiments, the sample obtained from the subject and the non-responders are different types of samples. In some embodiments, the sample obtained from the non-responders and the responders are the sametype of sample. In some embodiments, the sample obtained from the non-responders and the responders are different types of samples. In some embodiments, the sample obtained from the subject, the non-responders and the responders are the same type of sample. In some embodiments, the sample obtained from the subject, the non-responders and the responders are blood samples. In some embodiments, the sample obtained from the subject, the non- responders and the responders are plasma samples. In some embodiments, the sample obtained from the subject, the non-responders and the responders are serum samples. In some embodiments, the sample obtained from the subject, the non-responders and the responders are different types of samples.

[0068] In some embodiments, the RAPs are RAPs in the subject. In some embodiments, the first group of RAPs are in the subject. In some embodiments, the first group of RAPs are defined by differential levels of expression in CB and non-CB (NCB) subjects. In some embodiments, the RAPs are for a first disease, and the CB and NCB subjects have the first disease. In some embodiments, response is to the treatment. In some embodiments, the CB and NCB subjects have the first disease and their response to the treatment is known. In some embodiments, CB and NCB are with respect to the treatment. As defined here, RAPs are defined as factors or proteins conferring resistance to a therapy.

[0069] In some embodiments, the method further comprises obtaining a first dataset, comprising (i) expression levels of proteins in a cohort of patients of the first disease, and (ii) at least one annotation, labeling at least one patient of the cohort of patients as either a CB patient or a NCB patient. In some embodiments, the first dataset comprises expression level of CB patients with the first disease and with respect to response to the treatment. In some embodiments, the first dataset comprises expression level of NCB patients with the first disease and with respect to response to the treatment.

[0070] In some embodiments, a statistical test is applied to the expression levels of the first dataset to identify RAPs. In some embodiments, expression in the first dataset is protein expression. In some embodiments, the statistical test is a Kolmogorov- Smirnov test. Examples of other possible statistical tests include T-tests and Cox regression among many others known in the art. In some embodiments, T-test is Student’s T-test. In some embodiments, T-test is Welch’s t-test. In some embodiments, T-test is one-sample t-test. In some embodiments, T-test is two-sample t-test. In some embodiments, T-test is paired t-test. In some embodiments, T-test is one-tailed t-test. In some embodiments, T-test is Hotelling’s T2test. In some embodiments, T-test is Bayesian t-test. In some embodiments, T-test ispermutation / bootstrapped t-test. In some embodiments, Cox regression is Cox proportional hazard regression. In some embodiments, the statistical test is a proportional hazard model.

[0071] In some embodiments, the method further comprises constructing a first prediction model as an ensemble of a plurality of decision trees; for each decision tree of the plurality of decision trees, using at least one annotation as supervisory data, to train that decision tree so as to predict an interim response score, based on expression of a respective, unique protein of the first group of RAPs; and configuring the ensemble of decision trees to calculate the first response score based on the interim response scores of the plurality of decision trees. In some embodiments, the method further comprises constructing a first prediction model as an ensemble of a plurality of decision trees; for each decision tree of the plurality of decision trees, using at least one annotation as supervisory data, to train that decision tree so as to predict an interim response score, based on expression of a protein of the first group of RAPs; and configuring the ensemble of decision trees to calculate the first response score based on the interim response scores of the plurality of decision trees. In some embodiments, the decision tree is a prediction model for a single RAP. In some embodiments, the decision tree is a prediction model for a single RAP, and the first prediction model comprises a plurality of single RAPs prediction models. In some embodiments, the single RAP prediction model or decision tree is constructed using XGBoost algorithm. In some embodiments, a protein of the first group of RAPs is all proteins of the first group of RAPs. In some embodiments the first response score is CB score. In some embodiments, the interim response score is interim CB score. In some embodiments, the first response score is first CB score.

[0072] In some embodiments, the method further comprises, obtaining a second dataset, comprising (i) expression levels of proteins in a cohort of patients of the second disease, and (ii) at least one annotation, labeling at least one patient of the cohort of patients as either a CB patient or aNCB patient. In some embodiments, the second dataset comprises expression level of CB patients with the second disease and with respect to response to the treatment. In some embodiments, the second dataset comprises expression level of NCB patients with the second disease and with respect to response to the treatment. In some embodiments, a statistical test is applied to the expression levels of the second dataset to identify RAPs. In some embodiments, expression in the second dataset is protein expression. In some embodiments, the statistical test is a Kolmogorov- Smirnov test. Examples of other possible statistical tests include T-tests and Cox regression among many others known in the art. In some embodiments, T-test is Student’s T-test. In some embodiments, T-test is Welch’s t-test. In some embodiments, T-test is one-sample t-test. In some embodiments, T-test is two-sample t-test. In some embodiments, T-test is paired t-test. In some embodiments, T-test is one-tailed t-test. In some embodiments, T-test is Hotelling’s T2test. In some embodiments, T-test is Bayesian t-test. In some embodiments, T-test is permutation / bootstrapped t-test. In some embodiments, Cox regression is Cox proportional hazard regression. In some embodiments, the statistical test is a proportional hazard model. In some embodiments, the treatment for the first disease and the treatment for the second disease are the same treatment. In some embodiments, the treatment for the first disease and the treatment for the second disease are the same type of treatment. In some embodiments, the treatment for the first disease and the treatment for the second disease are both immunotherapy. In some embodiments, the treatment for the first disease and the treatment for the second disease are different treatments. In some embodiments, the treatment for the first disease and the treatment for the second disease are different treatments of the same type of treatment.

[0073] In some embodiments, the method further comprises calculating a correlation matrix, representing correlation of expression between RAPs of the second group. In some embodiments, the method comprises based on the correlation matrix, calculating a graph data element, representing a strength of correlation between pairs of RAPs of the second group. In some embodiments, based on the correlation matrix, calculating a strength of correlation between pairs of RAPs of the second group. In some embodiments, the method comprises applying, or inferring a graph analysis algorithm on the graph data element, to collect the RAPs of the second group into a plurality of clusters. In some embodiments, the method comprises clustering the RAPs of the second group based on similar expression. In some embodiments, the method comprises clustering the RAPs of the second group based on correlation of expression. In some embodiments, the absolute correlation is used. In some embodiments, a weighted correlation or mathematically modified correlation is used. In some embodiments, the method comprises selecting a subset of the plurality of clusters, based on a number of RAPs within each cluster. In some embodiments, the method comprises selecting a subset of the plurality of clusters based on a criterion. In some embodiments, the criterion is biological function or role. In some embodiments, the criterion is statistical, biological / clinical / functional significance, or topological. In some embodiments, criterion is strength of correlation, variance of correlations, cluster size, cluster diameter or radius, centrality, and significance of correlations.

[0074] In some embodiments, the method comprises in each cluster of the subset of clusters, identifying a hub RAP as one whose expression is most correlated to other RAPs of that cluster. In some embodiments, the method comprises in each cluster of the subset of clusters,identifying a hub RAP based on any one of degree, betweenness, closeness and eigenvector. In some embodiments, the method comprises in each cluster of the subset of clusters, identifying a hub RAP based on biological, clinical or functional significance. In some embodiments, the hub protein is based on CB and NCB. In some embodiments, the method comprises selecting a representative RAP from each cluster. In some embodiments, the representative RAP is a hub RAP. In some embodiments, the representative RAP is the one whose expression is most correlated to other RAPs of that cluster. In some embodiments, the selected RAPs are the subset of RAPs. In some embodiments, the hub RAPs are the subset of RAPs. In some embodiments, the subset of RAPs is the subset of the second group ofRAPs. In some embodiments, the hub RAPs are EPHB4, CLSTN3, BCHE, andDEFB132. In some embodiments, the hub proteins are EPHB4, CLSTN3, BCHE, and DEFB132. In some embodiments, the hub RAPs are selected from the RAPs provided in Table 3. In some embodiments, the hub RAPs are selected from the RAPs provided in Table 4.

[0075] In some embodiments, the subset of the second group of RAPs does not comprise RAPs from the first group of RAPs. In some embodiments, the subset of the second group of RAPs does comprise RAPs from the first group of RAPs. It will be understood that the first group ofRAPs are RAPs determined from responders and non-responders (e.g., CB and NCB patients) with the first disease to the treatment. While the second group of RAPs are RAPs determined from responder and non-responder (e.g., CB and NCB) with the second disease to the treatment. A subset is taken from the second group of RAPs and used to analyze a subject suffering from the first disease.

[0076] In some embodiments, the method further comprises constructing a second prediction model. In some embodiments, the method further comprises constructing a second prediction model using the hub RAPs. In some embodiments, the second prediction model is constructed using Cox regression model.

[0077] In some embodiments, the resistance-associated factor is determined by a method comprising: (a) receiving expression levels for a plurality of factors (i) in a population of subjects known to respond to the therapy (responders), (ii) in a population of subjects known to not respond to the therapy (non-responders), and (iii) in the subject; (b) calculate for at least one factor of the plurality of factors a resistance score; and (c) classify a factor with a resistance score beyond a threshold as a resistance-associated factor. In some embodiments, the resistance-associated factor is determined by a method comprising: (a) receiving expression levels for a plurality of factors (i) in a population of subjects known to clinical benefit from the therapy, (ii) in a population of subjects known to not have clinical benefitfrom the therapy, and (iii) in the subject; (b) calculate for at least one factor of the plurality of factors a resistance score; and (c) classify a factor with a resistance score beyond a threshold as a resistance-associated factor.

[0078] In some embodiments, the resistance-associated factor is a biological determinant (such as a genetic mutation, altered signaling pathway, protein expression, or influence from the tumor microenvironment) that enables cancer cells to evade or withstand the intended effects of a therapy or cancer treatment. In some embodiments, resistance-associated factor is a resistance-associated protein.

[0079] In some embodiments, a resistance score is a RAP score. In some embodiments, a resistance score is a response score. In some embodiments, a resistance score is 1 -response score. In some embodiments, response score is 1 -resistance score. In some embodiments, the response score is the RAP score. In some embodiments, the response score is CB score. In some embodiments, the response score is NCB score. In some embodiments, the score is interim CB score. In some embodiments, resistance score is total resistance score. In some embodiments, response score is total response score. In some embodiments, the response score is CB probability. In some embodiments, the response score is NCB probability. In some embodiments, the score is interim NCB score. In some embodiments, a RAP score is a total RAP score. In some embodiments, the resistance score is based on similarity of the factor expression level in the subject to the factor expression level in the non -responders. In some embodiments, the resistance score is based on similarity of the factor expression level in the subject to the factor expression level in the responders. In some embodiments, based on is calculated based on. In some embodiments, similarity is lack of similarity. In some embodiments, similarity to responders is lack of similarity to non-responders. In some embodiments, similarity to non-responders is lack of similarity to responders. In some embodiments, similarity is measured on a scale. In some embodiments, the scale is from 0 to 1, wherein 1 is perfectly similar to non-responders and 0 is perfectly similar to responders. In some embodiments, the resistance score is from 0 to 1, wherein 1 is perfectly similar to non-responders and 0 is perfectly similar to responders. In some embodiments, the resistance score is based on similarity of the factor expression level in the subject to the factor expression level in the non-responders and the factor expression level in the responders. In some embodiments, the scale is from 0 to 1, wherein 1 is perfectly similar to responders and 0 is perfectly similar to non-responders. In some embodiments, the response score is from 0 to 1, wherein 1 is perfectly similar to responders and 0 is perfectly similar to non-responders. In some embodiments, the scale is from 0 to 10, wherein 10 is perfectly similar to respondersand 0 is perfectly similar to non-responders. In some embodiments, the resistance score is from 0 to 10, wherein 10 is perfectly similar to non-responders and 0 is perfectly similar to responders. In some embodiments, the resistance score is based on similarity of the factor expression level in the subject to the factor expression level in the non-responders and the factor expression level in the responders. In some embodiments, the scale is from 0 to 10, wherein 10 is perfectly similar to responders and 0 is perfectly similar to non-responders. In some embodiments, the response score is from 0 to 10, wherein 10 is perfectly similar to responders and 0 is perfectly similar to non-responders. In some embodiments, the resistance score is from 0 to 10, wherein 10 is perfectly similar to non-responders and 0 is perfectly similar to responders In some embodiments, the response score is based on similarity of the factor expression level in the subject to the factor expression level in the non-responders and the factor expression level in the responders.

[0080] In some embodiments, resistance-associated factors are in each subject. In some embodiments, resistance-associated factors are in the responders. In some embodiments, resistance-associated factors are in the non-responders. In some embodiments, the resistance-associated factors are labeled with the labels. In some embodiments, the resistance-associated factors are resistance-associated proteins. In some embodiments, response-associated factors are in each subject. In some embodiments, response-associated factors are in the responders. In some embodiments, response-associated factors are in the non-responders. In some embodiments, the response-associated factors are labeled with the labels. In some embodiments, the response-associated factors are resistance-associated proteins.

[0081] In some embodiments, the population of responders or the CB patients suffers from the first disease. In some embodiments, the population of responders or the CB patients suffers from the second disease. In some embodiments, the responders or the CB patients all have the same disease. In some embodiments, the population of non-responders or NCB patients suffers from the first disease. In some embodiments, the population of non- responders or NCB patients suffers from the second disease. In some embodiments, the non- responders or the NCB patients all suffer from the same disease. In some embodiments, the population of responders and non-responders all suffer from the same disease. In some embodiments, the population of responders and the subject suffer from the same disease. In some embodiments, the population of CB and NCB all suffer from the same disease. In some embodiments, the population of CB and the subject suffer from the same disease. In some embodiments, the population of non-responders and the subject suffer from the same disease.In some embodiments, the population of NCB and the subject suffer from the same disease. In some embodiments, the population of non-responders, the population of responders and the subject suffer from the same disease. In some embodiments, the population of NCB patients, the population of CB patients and the subject suffer from the same disease.

[0082] In some embodiments, the first and second diseases are different diseases. In some embodiments, the first and second diseases are the same disease.

[0083] In some embodiments, calculating comprises applying a machine learning algorithm. In some embodiments, calculating comprises applying a machine learning model. In some embodiments, the machine learning model is a machine learning algorithm. In some embodiments, the machine learning model implements a machine learning algorithm. In some embodiments, the algorithm is a classifier. In some embodiments, the algorithm is a regression model. In some embodiments, the algorithm is supervised. In some embodiments, the algorithm is unsupervised. In some embodiments, the machine learning algorithm is trained on the expression levels in responders. In some embodiments, the machine learning algorithm is trained on the expression levels in non-responders. In some embodiments, the machine learning algorithm is trained on the expression levels in responders and non- responders. In some embodiments, the machine learning algorithm is trained on expression levels of a plurality of factors in responders and non-responders or CB and NCB patients In some embodiments, the machine learning algorithm is trained on expression levels of a subset of factors in responders and non-responders or CB and NCB patients. In some embodiments, the machine learning algorithm is trained on the expression levels of the RAPs in responders and non-responders. In some embodiments, the machine learning algorithm is trained on the expression levels of the hub RAPs in responders and non-responders. In some embodiments, the machine learning algorithm is trained on a training set. In some embodiments, the machine learning algorithm is trained by a method of the invention. In some embodiments, the trained machine learning algorithm is applied to a plurality of received protein expression levels. In some embodiments, a machine learning algorithm is applied to factors of the plurality of factors. In some embodiments, a machine learning algorithm is applied to each factor of the plurality of factors. In some embodiments, a machine learning algorithm is applied to the subset. In some embodiments, a machine learning algorithm is applied to the subset of factors. In some embodiments, a machine learning algorithm is applied to each factor of the subset of factors. In some embodiments, each factor is analyzed and calculated separately, and the machine learning algorithm does not use expression levels of more than one factor as the training set. In some embodiments,a trained machine learning algorithm is applied to individual protein expression levels from the subject. In some embodiments, a machine learning algorithm trained on expression levels of a specific factor in responders and non-responders or CB and NCB patients is applied to the expression level of that specific factor in the subject. It will be understood by a skilled artisan, that for each of the factors of the plurality of factors, a different algorithm will be trained and then applied to each expression level of the subject. Thus, if three algorithms are separately trained on expression in responders and non-responders or CB and NCB patients for Factor A, Factor B and Factor C, then the algorithm trained on Factor A expression levels will be applied to the subject’s expression level of Factor A, the algorithm trained on Factor B expression levels will be applied to the subject’s expression level of Factor B, and the algorithm trained on Factor C expression levels will be applied to the subject’s expression level of Factor C. In some embodiments, during a training phase, the machine learning model is trained on a training set comprising expression data for a single factor from responders and non-responders, using corresponding annotations of “responder” or “non-responder” to predict or classify factor expression data according to classes “responder” and “non- responder”. In some embodiments, during an inference stage, the machine learning model is applied to expression data of the single factor from a subject to predict classification of the factor as similar to a responder or non-responder. In some embodiments, the classification is a resistance score. In some embodiments, the classification is a response score. In some embodiments, the classification is a measure of how similar the factor is to non-responders and dissimilar to responders. In some embodiments, during a training phase, the machine learning model is trained on a training set comprising expression data for a single factor from CB and NCB population, using corresponding annotations of “CB” or “NCB” to predict or classify factor expression data according to classes “CB” and “NCB”. In some embodiments, during an inference stage, the machine learning model is applied to expression data of the single factor from a subject to predict classification of the factor as similar to a CB or NCB. In some embodiments, the classification is a resistance score. In some embodiments, the classification is a CB score. In some embodiments, the classification is a CB probability score. In some embodiments, the classification is a NCB score. In some embodiments, the classification is a NCB probability score. In some embodiments, the classification is a measure of how similar the factor is to non-responders and dissimilar to responders.

[0084] In some embodiments, the trained machine learning algorithm is trained to predict responsiveness of subjects suffering from the disease to the therapy. In some embodiments, the trained machine learning algorithm is trained to output a resistance score. In someembodiments, the trained machine learning algorithm is trained to output a resistance probability. In some embodiments, the trained machine learning algorithm is trained to output an activity score. In some embodiments, the trained machine learning algorithm is trained to predict activity of a resistance-associated factor in a subject. In some embodiments, the trained machine learning algorithm is trained to predict if a factor is a resistance-associated factor in the subject. In some embodiments, the trained machine learning algorithm is trained to predict if a factor of the subject is a resistance-associated factor in the subject. In some embodiments, the trained machine learning algorithm is trained to predict clinical benefit probability. In some embodiments, cross-validation is used during training.

[0085] In some embodiments, the training set comprises received factor expression levels. In some embodiments, the training set comprises received factor expression levels in both responders and non-responders. In some embodiments, the training set comprises received factor expression levels in both CB and NCB populations. In some embodiments, the training set comprises received factor expression levels for only one factor. In some embodiments, the training set comprises received factor expression levels for a plurality of factors. In some embodiments, the training set comprises received factor expression levels for a subset of factors. In some embodiments, the training set comprises the number of resistance-associated factors expressed in samples. In some embodiments, the samples are from subjects suffering from the disease. In some embodiments, the samples are from responders. In some embodiments, the samples are from CB population. In some embodiments, the samples are from non-responders. In some embodiments, the samples are from NCB population. In some embodiments, the training set comprises at least one clinical parameter. In some embodiments, the clinical parameter is from subjects. In some embodiments, subjects are responders and non-responders. In some embodiments, subjects are CB and NCB subjects. In some embodiments, the training set comprises labels. In some embodiments, the training set comprises annotations. In some embodiments, the labels or annotations are associated with the responsiveness or clinical benefit to the treatment of the subjects. In some embodiments, the labels are responder or non-responder. In some embodiments, the labels or annotations are responder or non-responder. In some embodiments, the labels or annotations are CB or NCB. In some embodiments, the labels or annotations are suffering from side effects or not suffering from side effects. In some embodiments, the resistance- associated factors are labeled with the labels. In some embodiments, the at least one clinical parameter is labeled with the label.

[0086] In some embodiments, at an inference stage the trained machine learning algorithm is applied. In some embodiments, the trained machine learning algorithm is applied to individual received factor expression levels. In some embodiments, the trained machine learning algorithm is applied to individual received factor expression levels and at least one clinical parameter. In some embodiments, the trained machine learning algorithm is applied to individual received factor expression levels from the subjects and the subject’s sex. In some embodiments, the trained machine learning algorithm is applied to the number of resistance-associated proteins. In some embodiments, the trained machine learning algorithm is applied to the number of resistance-associated factors. In some embodiments, the trained machine learning algorithm is applied to the number of resistance-associated factors and at least one clinical parameter.

[0087] In some embodiments, at the inference stage an input is received. In some embodiments, the input comprises the number of resistance-associated factors expressed in a sample. In some embodiments, the sample is from a subject. In some embodiments, the input comprises at least one clinical parameter. In some embodiments, the subject suffers from the disease. In some embodiments, the subject has unknown responsiveness to the therapy. In some embodiments, the parameter is of the subj ect with unknown responsiveness. In some embodiments, the subject has unknown clinical benefit to therapy. In some embodiments, the parameter is of the subject with unknown clinical benefit. In some embodiments, at the inference stage the trained machine learning algorithm is applied. In some embodiments, applied is applied to the input. In some embodiments, the input is the received input. In some embodiments, the inference stage is to predict responsiveness. In some embodiments, the inference stage is to predict clinical benefit. In some embodiments, responsiveness is responsiveness to the therapy of the subject with unknown responsiveness. In some embodiments, clinical benefit is clinical benefit to the therapy of the subject with unknown clinical benefit.

[0088] In some embodiments, the machine learning algorithm outputs the resistance score. In some embodiments, resistance score is CB score. In some embodiments, the outputted resistance score is scaled from 0 to 1. In some embodiments, 1 is perfectly similar to nonresponders and 0 is perfectly similar to responders. In some embodiments, the outputted resistance score is scaled from 0 to 10. In some embodiments, 10 is perfectly similar to nonresponders and 0 is perfectly similar to responders. In some embodiments, the machine learning algorithm calculates similarity to responders. In some embodiments, the machine learning algorithm calculates similarity to non-responders. In some embodiments, themachine learning algorithm outputs a numeric value of similarity to responders and nonresponders. In some embodiments, a protein is considered to be a RAP if its resistance score is beyond a certain threshold. In some embodiments, the threshold for the resistance score is calculated on a scale of 0 to 1. In some embodiments, the threshold for the resistance score of a certain protein is between 0.2 and 0.95. In some embodiments, the threshold for the resistance score of a certain protein is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the resistance score is 0.25. In some embodiments, the threshold for the resistance score is 0.42. In some embodiments, the threshold for the resistance score is 0.6. In some embodiments, the threshold for the resistance score when calculated by a machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.25. In some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.42. In some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.6.

[0089] In some embodiments, response probability is determined by the calculation (1- resistance score). In some embodiments, 1 -resistance score is 1 -final resistance score. In some embodiments, the resistance score is the final resistance score. In some embodiments, response probability is a response score. In some embodiments, the machine learning algorithm outputs the response score. In some embodiments, the outputted response score is scaled from 0 to 1. In some embodiments, 1 is perfectly similar to responders and 0 is perfectly similar to non-responders. In some embodiments, the outputted response score is scaled from 0 to 10. In some embodiments, 10 is perfectly similar to responders and 0 is perfectly similar to non-responders. In some embodiments, the machine learning algorithm calculates similarity to responders. In some embodiments, the machine learning algorithm calculates similarity to non-responders. In some embodiments, the machine learning algorithm outputs a numeric value of similarity to responders and non-responders. In some embodiments, a protein is considered to be a RAP if its response score is beyond a certain threshold. In some embodiments, the threshold for the response score is calculated on a scale of 0 to 1. In some embodiments, the threshold for the response score of a certain protein is between 0.2 and 0.95. In some embodiments, the threshold for the response score of a certain protein is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85,0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the response score is 0.25. In some embodiments, the threshold for the response score is 0.42. In some embodiments, the threshold for the response score is 0.6. In some embodiments, the threshold for the response score when calculated by a machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.25. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.42. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.6. In some embodiments, the algorithm outputs response probability, and the response probability is calculated on a scale of 0 to 1. In some embodiments, the algorithm outputs response probability, and the response probability is calculated on a scale of 0% to 100%, wherein 100% is responder and 0% is non-responder.

[0090] In some embodiments, the score is between zero and 1. In some embodiments, the score is between zero and 10. In some embodiments, active is active in the cancer. In some embodiments, active is active in the subject. In some embodiments, active is active in promoting resistance. In some embodiments, beyond a threshold is below a threshold. In some embodiments, beyond a threshold is above a threshold. In some embodiments, the predetermined threshold is 0.5, 0.4, 0.3, 0.25, 0.2, 0.15, 0.1, 0.05, 0.01, 0.005, 0.001, 0.0005 or 0.0001. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold is 0.05. In some embodiments, the threshold is 5%.

[0091] In some embodiments, the machine learning algorithm outputs the resistance score. In some embodiments, the resistance score is the RAP score. In some embodiments, the outputted resistance score is scaled from 0 to 1. In some embodiments, 1 is perfectly similar to non-responders and 0 is perfectly similar to responders. In some embodiments, the outputted resistance score is scaled from 0 to 10. In some embodiments, 10 is perfectly similar to non-responders and 0 is perfectly similar to responders. In some embodiments, the machine learning algorithm calculates similarity to responders. In some embodiments, the machine learning algorithm calculates similarity to non-responders. In some embodiments, the machine learning algorithm outputs a numeric value of similarity to responders and non- responders. In some embodiments, a protein is considered to be a RAP if its resistance score is beyond a certain threshold. In some embodiments, the threshold for the resistance score is calculated on a scale of 0 to 1. In some embodiments, the threshold for the resistance scoreof a certain protein is between 0.2 and 0.95. In some embodiments, the threshold for the resistance score of a certain protein is about 0.01, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the resistance score is 0.25. In some embodiments, the threshold for the resistance score is 0.42. In some embodiments, the threshold for the resistance score is 0.6. In some embodiments, the threshold for the resistance score when calculated by a machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.25. In some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.42. In some embodiments, the threshold for the resistance score when calculated with a machine learning algorithm is 0.6.

[0092] In some embodiments, response probability is determined by the calculation (1- resistance score). In some embodiments, 1 -resistance score is 1 -final resistance score. In some embodiments, the resistance score is the final resistance score. In some embodiments, response probability is a response score. In some embodiments, the machine learning algorithm outputs the response score. In some embodiments, the outputted response score is scaled from 0 to 1. In some embodiments, 1 is perfectly similar to responders and 0 is perfectly similar to non -responders. In some embodiments, the outputted response score is scaled from 0 to 10. In some embodiments, 10 is perfectly similar to responders and 0 is perfectly similar to non -responders. In some embodiments, the predetermined threshold is 6.7. In some embodiments, the predetermined threshold is 6.7 and a score higher than 6.7 is similar to responders an a score lower than 6.7 is similar to non-responders.

[0093] According to some embodiments, the threshold may be determined based on a percentage of CB or NCB patients or the number of responders or non-responders in the cohort of patients.

[0094] In some embodiments, the machine learning algorithm calculates similarity to responders. In some embodiments, the machine learning algorithm calculates similarity to non-responders. In some embodiments, the machine learning algorithm outputs a numeric value of similarity to responders and non-responders. In some embodiments, a protein is considered to be a RAP if its response score is beyond a certain threshold. In some embodiments, beyond is above. In some embodiments, beyond is below.

[0095] In some embodiments, the threshold for the response score is calculated on a scale of 0 to 1. In some embodiments, the threshold for the response score of a certain protein is between 0.2 and 0.95. In some embodiments, the threshold for the response score of a certain protein is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the response score is 0.25. In some embodiments, the threshold for the response score is 0.42. In some embodiments, the threshold for the response score is 0.6. In some embodiments, the threshold for the response score when calculated by a machine learning algorithm is about 0.2, 0.25, 0.3, 0.35, 0.4, 0.42, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.25. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.42. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 0.6.

[0096] In some embodiments, the threshold for the response score is calculated on a scale of 0 to 10. In some embodiments, the threshold for the response score of a certain protein is between 2 and 9.5. In some embodiments, the threshold for the response score of a certain protein is about 2, 2.5, 3, 3.5, 4, 4.2, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, or 9.5. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the response score is 2.5. In some embodiments, the threshold for the response score is 4.2. In some embodiments, the threshold for the response score is 6. In some embodiments, the threshold for the response score when calculated by a machine learning algorithm is about 2, 2.5, 3, 3.5, 4, 4.2, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, or 9.5. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 2.5. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 4.2. In some embodiments, the threshold for the response score when calculated with a machine learning algorithm is 6.

[0097] In some embodiments, the machine learning model is a machine learning algorithm. In some embodiments, the algorithm is a supervised learning algorithm. In some embodiments, the algorithm is an unsupervised learning algorithm. In some embodiments, the algorithm is a reinforcement learning algorithm. In some embodiments, the machine learning model is a Convolutional Neural Network (CNN). Other examples of acceptable models include, but are not limited to XGBoost, SVM, linear regression, COX regression,decision trees in general, naive-based, logistic regression, LLM, and K means clustering. In some embodiments, the at least one hardware processor trains a machine learning model. In some embodiments, the model is based, at least in part, on a training set. In some embodiments, the model is based on a training set. In some embodiments, the model is trained on a training set. In some embodiments, the at least one hardware processor applies the machine learning model to a factor expression level from a subject. In some embodiments, cross-validation is used to construct the training set. In some embodiments, the training set is cross-validated.

[0098] In some embodiments, the calculating comprises calculating a mean expression for each protein in responders. In some embodiments, the calculating comprises calculating a mean expression for each protein in non-responders. In some embodiments, the calculating comprises calculating a mean expression for each protein in responders and a mean expression for each protein in non-responders. In some embodiments, the calculating comprises calculating a distribution of the expression for each protein in responders and non- responders. In some embodiments, the calculating comprises calculating a standard deviation of expression for each protein in responders and non-responders. In some embodiments, in responders is in the responders population. In some embodiments, in non- responders is in the non-responders population. In some embodiments, the resistance score is based on the ratio of deviation of the factor expression in the subject from the calculated mean in responders to the deviation of the factor expression in the subject from the calculated mean in non-responders. Calculation of deviation is well known to one skilled in the art. It will be understood that the more dissimilar the expression in the subject is from a mean the larger the deviation will be. Thus, factors that are very dissimilar to the mean in responders will have a large numerator in the calculation of this ratio and factors that are lowly dissimilar to the mean in non-responders will have a small denominator. Thus, the more dissimilar to responder expression and the more similar to non-responder expression is expression of a factor in a subject the higher the resistance score will be. In some embodiments, a resistance score beyond a predetermined threshold indicates a factor is a resistance-associated factor. In some embodiments, a resistance-associated factor is a resistance-associated protein (RAP).

[0099] In some embodiments, the calculating further comprises calculating a distribution for each factor in responders. In some embodiments, the calculating further comprises calculating a distribution for each factor in non-responders. In some embodiments, the calculating further comprises calculating a distribution for each factor in responders and adistribution for each factor in non-responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in non-responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in responders and a standard deviation for each protein in non-responders. In some embodiments, the calculating further comprises calculating a standard deviation for each factor in a mix of responders and non-responders. In some embodiments, the deviation is measured as a multiple of the calculated standard deviation. It will be understood by a skilled artisan that by scaling the deviation to the standard deviation for a group of expression values the deviation can be given in more absolute terms allow for the comparison of factors and populations with very small and very large stand deviations (which may also have very low and very high expression levels).

[0100] In some embodiments, the resistance score is based on a Z-score for the expression level of each factor in the subject. In some embodiments, the resistance score is based on the Z-score relative to responders. In some embodiments, the resistance score is based on the Z- score relative to non-responders. In some embodiments, the resistance score is based on both the Z-score relative to responders and the Z-score relative to non-responders. In some embodiments, the resistance score is based on the ratio of the Z-score relative to responders to the Z-score relative to non-responders. It will be well known to a skilled artisan that a Z- score counts the distance of the individual level from the population mean in units of the population standard deviation. In some embodiments, the Z-score is calculated by Equation 1.some embodiments, ZR is the deviation of the factor expression in the subject from the calculated mean in responders. In some embodiments, ZN is the deviation of the factor expression in the subject from the calculated mean in non-responders. In some embodiments, | | is the Z-score of the deviation. In some embodiments, | | is the standardizing of the deviation to a multiple of the standard deviation. In some embodiments, c is a constant. In some embodiments, constant is a regulation constant that prevents the score from divergence for ZNR= 0. In some embodiments, the resistance score is calculated by Equation 2. In some embodiments, monotonoic is an ad-hoc function that prevents the resistance score from decreasing for extreme values within the non-responder distributions. In some embodiments, function is the function provided in Algorithm 1.

[0102] In some embodiments, a resistance score beyond a predetermined threshold indicates a factor is a RAP. In some embodiments, beyond is above. In some embodiments, the threshold is a predetermined threshold. In some embodiments, threshold is a threshold value. In some embodiments, the threshold for the resistance score is about 1.0, 1.1, 1.2, 1.3, 1.4,1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5,3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 5.0, 5.1, 5.2, 5.3, 5.4, 5.5, 5.6, 5.7,5.8, 5.9, 6.0. 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, or 7.0. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold is any number between 0-1. In some embodiments, the threshold is about 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.67, 0.7, 0.75, 0.8, 0.85 or 0.9. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the resistance score is about 2.9. In some embodiments, the threshold for the resistance score is 2.9. In some embodiments, the threshold for the resistance score is about 3.0. In some embodiments, the threshold for the resistance score is 3.0. In some embodiments, the threshold for the resistance score is calculated on a scale of arbitrary units. In some embodiments, the threshold for the resistance score when calculated by a mathematical calculation is about 1.0,1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1,3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, or 5.0. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the resistance score when calculated with a mathematical calculation is about2.9. In some embodiments, the threshold for the resistance score when calculated with a mathematical calculation is 2.9. In some embodiments, the threshold for the resistance score when calculated with a mathematical calculation is about 3.0. In some embodiments, the threshold for the resistance score when calculated with a mathematical calculation is 3.0. In some embodiments, a mathematical calculation is a method that comprises calculating a mean expression for each protein. In some embodiments the threshold is calculated based on the sum, median, average, of the protein in the population. In some embodiments, the threshold is any number between 0-10. In some embodiments, the threshold is about 5, 6, 7, 8, 9, or 10. Each possibility represents a separate embodiment of the invention. In some embodiments, the threshold for the resistance score is about 5. In some embodiments, the threshold for the resistance score is about 6.7. In some embodiments, the number of active RAPs is linearized to provide a total score between 0 and 10. In some embodiments, linearized is linearly scaled. In some embodiments, linearizing comprises a linear regression.In some embodiments, the number of active RAPs is converted to a total score between 0 and 10.

[0103] In some embodiments, the scale is from 0 to 10, wherein 10 is perfectly similar to responders and 0 is perfectly similar to non-responders. In some embodiments, the resistance score is from 0 to 10, wherein 10 is perfectly similar to non-responders and 0 is perfectly similar to responders. In some embodiments, the score is based on similarity of the factor expression level in the subject to the factor expression level in the non-responders and the factor expression level in the responders. In some embodiments, the score is from 0 to 10, wherein 10 is perfectly similar to responders and 0 is perfectly similar to non-responders. In some embodiments, the score is the PROphet score. In some embodiments, the score is the total score. In some embodiments, a prophet positive subject is a subject with a score above a predetermined threshold. In some embodiments, a positive subject is a subject with a score above a predetermined threshold. In some embodiments, a prophet negative subject is a subject with a score below a predetermined threshold. In some embodiments, a negative subject is a subject with a score below a predetermined threshold. In some embodiments, the score is based on similarity of the factor expression level in the subject to the factor expression level in the non-responders and the factor expression level in the responders. In some embodiments, a response score from 5 to 10 indicates the subject is a responder. In some embodiments, a response score above 5 indicates the subject is a responder. In some embodiments, a response score from 5 to 0 indicates the subject is a non-responder. In some embodiments, a response score below 5 indicates the subject is a non-responder. In some embodiments, a response score from 6 to 10 indicates the subject is a responder. In some embodiments, a response score above 6 indicates the subject is a responder. In some embodiments, a response score from 6 to 0 indicates the subject is a non-responder. In some embodiments, a response score below 6 indicates the subject is a non-responder. In some embodiments, a response score from 6.7 to 10 indicates the subject is a responder. In some embodiments, a response score above 6.7 indicates the subject is a responder. In some embodiments, a response score from 6.7 to 0 indicates the subject is a non-responder. In some embodiments, a response score below 6.7 indicates the subject is a non-responder.

[0104] In some embodiments, the method comprises selecting a subset of factors. In some embodiments, the subset is a subset of the plurality of factors. In some embodiments, the subset comprises the factors that best differentiate between the responders and non- responders. In some embodiments, the factors that best differentiate are the top percentage.In some embodiments, the top percentage is the top 1, 3, 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50% of factors. Each possibility represents a separate embodiment of the invention. In some embodiments, the top percentage is the top 20%. In some embodiments, the top factors are the top 10, 20, 25, 30, 40, 50, 60, 70, 75, 80, 90 or 100 factors. Each possibility represents a separate embodiment of the invention. In some embodiments, the top factors are the top 50 factors. In some embodiments, selection comprises applying a Kolmogorov- Smirnov test. In some embodiments, the Kolmogorov-Smirnov test is applied to the received factor expression levels. In some embodiments, the Kolmogorov- Smirnov test determines how well a factor differentiates between responders and non-responders. In some embodiments, the Kolmogorov- Smirnov test outputs a measure of how well a factor differentiates, and the best factors are the factors with the highest scores. In some embodiments, selection comprises applying an XGBoost algorithm. In some embodiments, the calculating is for the subset. In some embodiments, the calculating is for each factor of the subset.

[0105] In some embodiments, calculating comprises applying a machine learning algorithm. In some embodiments, calculating comprises applying a machine learning model. In some embodiments, the machine learning model is a machine learning algorithm. In some embodiments, the machine learning model implements a machine learning algorithm. In some embodiments, the algorithm is a classifier. In some embodiments, the algorithm is a regression model. In some embodiments, the algorithm is supervised. In some embodiments, the algorithm is unsupervised. In some embodiments, the machine learning algorithm is trained on the expression levels in responders. In some embodiments, the machine learning algorithm is trained on the expression levels in non-responders. In some embodiments, the machine learning algorithm is trained on the expression levels in responders and non- responders. In some embodiments, the machine learning algorithm is trained on a training set. In some embodiments, the machine learning algorithm is trained by a method of the invention. In some embodiments, a machine learning algorithm is applied to factors of the plurality of factors. In some embodiments, a machine learning algorithm is applied to each factor of the plurality of factors. In some embodiments, a machine learning algorithm is applied to the subset. In some embodiments, a machine learning algorithm is applied to the subset of factors. In some embodiments, a machine learning algorithm is applied to each factor of the subset of factors. In some embodiments, each factor is analyzed and calculated separately, and the machine learning algorithm does not use expression levels of more than one factor as the training set. In some embodiments, a trained machine learning algorithm isapplied to individual protein expression levels from the subject. In some embodiments, a machine learning algorithm trained on expression levels of a specific factor in responders and non-responders is applied to the expression level of that specific factor in the subject. It will be understood by a skilled artisan, that for each of the factors of the plurality of factors, a different algorithm will be trained and then applied to each expression level of the subject. Thus, if three algorithms are separately trained on expression in responders and non- responders for Factor A, Factor B and Factor C, then the algorithm trained on Factor A expression levels will be applied to the subject’s expression level of Factor A, the algorithm trained on Factor B expression levels will be applied to the subject’s expression level of Factor B, and the algorithm trained on Factor C expression levels will be applied to the subject’s expression level of Factor C. In some embodiments, during a training phase, the machine learning model is trained on a training set comprising expression data for a single factor from responders and non-responders, using corresponding annotations of “responder” or “non-responder” to predict or classify factor expression data according to classes “responder” and “non-responder”. In some embodiments, during an inference stage, the machine learning model is applied to expression data of the single factor from a subject to predict classification of the factor as similar to a responder or non-responder. In some embodiments, the classification is a resistance score. In some embodiments, the classification is a response score. In some embodiments, the classification is a measure of how similar the factor is to non-responders and dissimilar to responders.

[0106] In some embodiments, the trained machine learning algorithm is trained to predict responsiveness of subjects suffering from the disease to the therapy. In some embodiments, the trained machine learning algorithm is trained to output a resistance score. In some embodiments, the trained machine learning algorithm is trained to output a resistance probability. In some embodiments, the trained machine learning algorithm is trained to output clinical benefit probability. In some embodiments, the trained machine learning algorithm is trained to output an activity score. In some embodiments, the trained machine learning algorithm is trained to predict activity of a resistance-associated factor in a subject. In some embodiments, the trained machine learning algorithm is trained to predict if a factor is a resistance-associated factor in the subject. In some embodiments, the trained machine learning algorithm is trained to predict if a factor of the subject is a resistance-associated factor in the subject.

[0107] In some embodiments, the trained machine learning algorithm is trained to predict responsiveness of subjects suffering from the disease to the therapy. In some embodiments, the trained machine learning algorithm is trained to output a response score. In some embodiments, the trained machine learning algorithm is trained to output a response probability. In some embodiments, the trained machine learning algorithm is trained to output clinical benefit probability. In some embodiments, the trained machine learning algorithm is trained to output an activity score. In some embodiments, the trained machine learning algorithm is trained to predict activity of a response-associated factor in a subject. In some embodiments, the trained machine learning algorithm is trained to predict if a factor is a response-associated factor in the subject. In some embodiments, the trained machine learning algorithm is trained to predict if a factor of the subject is a response-associated factor in the subject.

[0108] In some embodiments, the training set comprises received factor expression levels. In some embodiments, the training set comprises received factor expression levels in both responders and non-responders. In some embodiments, the training set comprises received factor expression levels in both mono-responders and mono-non-responders. In some embodiments, the training set comprises received factor expression levels in both comboresponders and combo-non-responders. In some embodiments, the training set comprises received factor expression levels in mono-responders, mono-non-responders, comboresponders and combo-non-responders. In some embodiments, the training set comprises received factor expression levels for only one factor. In some embodiments, the training set comprises the number of resistance-associated factors or response-associated factors expressed in samples. In some embodiments, the sample are from subjects suffering from the disease. In some embodiments, the sample are from responders. In some embodiments, the sample are from non-responders. In some embodiments, the training set comprises at least one clinical parameter. In some embodiments, the clinical parameter is from subjects. In some embodiments, subjects are responders and non-responders. In some embodiments, the training set comprises labels. In some embodiments, the labels are associated with the responsiveness of the subjects. In some embodiments, the labels are responder or nonresponder. In some embodiments, the resistance-associated factors are labeled with the labels. In some embodiments, the expression levels of the resistance-associated factors are labeled with the labels. In some embodiments, the at least one clinical parameter is labeled with the label.

[0109] According to some embodiments, the training set further comprises at least one clinical parameter of each responder and non-responder and the machine learning algorithm is applied to individual received factor expression levels from the subject and the subject’s at least one clinical parameter. In some embodiments, the at least one clinical parameter is the sex of the subjects. In some embodiments, the training set further comprises the sex of the subjects. In some embodiments, the subjects are each subject. In some embodiments, sex is gender. In some embodiments, the at least one clinical parameter is sex. In some embodiments, sex is a subject’s sex. In some embodiments, sex is male or female. In some embodiments, sex is sex at birth. In some embodiments, the training set comprises the sex of each responder. In some embodiments, the training set comprise the sex of each non- responder. In some embodiments, the training set comprises the sex of each mono-responder. In some embodiments, the training set comprise the sex of each mono-non-responder. In some embodiments, the training set comprises the sex of each combo-responder. In some embodiments, the training set comprise the sex of each combo-non-responder. In some embodiments, the clinical parameter is age. In some embodiments, age is a subject’s age. In some embodiments, the clinical parameter is the line of treatment. In some embodiments, the line of treatment parameter is whether the therapy was a first line of treatment or an advanced treatment. In some embodiments, a line of treatment is first line treatment. In some embodiments, a line of treatment is a secondary treatment. In some embodiments, secondary treatment is an advanced treatment. It will be understood by a skilled artisan that advanced treatment may be any line of treatment after the first, e.g., second line, third line, fourth line, fifth line, etc. In some embodiments, the clinical parameter is whether the treatment is a first line treatment or an advanced treatment. In some embodiments, the clinical parameter is PD- L1 status. In some embodiments, PD-L1 status is PD-L1 status of the cancer. Methods of measuring PD-L1 levels in cancer cells (e.g., a tumor) are well known in the art and any such method may be employed. In some embodiments, PD-L1 status comprises high PD-L1 or low PD-L1. In some embodiments, PD-L1 status comprises high PD-L1, low PD-L1 or no PD-L1. In some embodiments, PD-L1 status comprises high PD-L1, medium PD-L1 or low PD-L1. In some embodiments, PD-L1 levels are numeric values between 0 to 100. In some embodiments, PD-L1 levels are percentages between 0 to 100. In some embodiments, PD- L1 status comprises PD-L1 expression in less than 1% of cancer cells, in 1-49% of cancer cells, or in 50% or more of cancer cells. In some embodiments, PD-L1 expression in less than 1% of cancer cells is no PD-L1 expression. In some embodiments, PD-L1 low or negative cancer comprises fewer than 50% of cancer cells being positive for PD-L1expression. In some embodiments, expression is surface expression. In some embodiments, PD-L1 negative cancer comprises fewer than 1% of cancer cells being positive for PD-L1 expression. In some embodiments, PD-L1 expression in less than 1% of cancer cells is low PD-L1 expression. In some embodiments, PD-L1 expression in 1-49% of cancer cells is low PD-L1 expression. In some embodiments, PD-L1 low cancer comprises fewer than 1-49% of cancer cells being positive for PD-L1 expression. In some embodiments, PD-L1 expression in 1-49% of cancer cells is medium PD-L1 expression. In some embodiments, PD-L1 expression in 50% or more of cancer cells is high PD-L1 expression. In some embodiments, a high PD-L1 cancer comprises expression in at least 50% of cells. In some embodiments, PD-L1 high cancer comprises at least 50% of cancer cells being positive for PD-L1 expression. In some embodiments, a low PD-L1 cancer comprises expression in 1- 49% of cells. In some embodiments, a no PD-L1 cancer comprises expression in 0% of cells. In some embodiments, a no PD-L1 cancer comprises expression in less than 1% of cells. In some embodiments, the PD-L1 low or negative cancer is PD-L1 low cancer. In some embodiments, the PD-L1 low or negative cancer is PD-L1 negative cancer. In some embodiments, a no PD-L1 cancer is a PD-L1 negative cancer.

[0110] In some embodiments, the clinical parameter is a known biomarker of the disease or mutations in known biomarkers of the disease. In some embodiments, the biomarker is selected from MYC, NOTCH, EGFR, HER2, BRAF, KRAS, MAP2K1, MET, NRAS, NTRK1, NTRK2, NTRK3, PIK3CA, RET, ROS1, TP53, ALK, CDKN2A, KIT, NF1, BFAST, FGFR, LDH, PTEN, RBI, PD-L1, MSI (Microsatelite Instability), TMB (Tumor Mutational Burden), or a combination thereof. In some embodiments, the clinical parameter is expression of the biomarker. In some embodiments, expression is percent expression. In some embodiments, expression is mutational status.

[0111] In some embodiments, the training set further comprises the sex, age and PD-L1 status of each responder and non-responder. In some embodiments, the training set further comprises the sex of each responder and non-responder. In some embodiments, the training set further comprises the age and PD-L1 status of each responder and non-responder. In some embodiments, the machine learning algorithm is applied to individual received factor expression levels from the subject and the subject’s sex. In some embodiments, the machine learning algorithm is applied to individual received factor expression levels from the subject and the subject’s sex, age andPD-Ll status. In some embodiments, the calculating comprises applying a machine learning algorithm trained on a training set comprising the receivedfactor expression levels in responders and non-responders and at least one clinical parameter, to the expression levels from the subject and the subject’s at least one clinical parameter and wherein the machine learning algorithm outputs the resistance score. In some embodiments, the training comprises the received factor expression levels in responders and non- responders and clinical parameters of each responder and non-responder and the machine learning algorithm is applied to individual received factor expression levels from the subject and the subject’s clinical parameters and wherein the machine learning algorithm outputs response score. In some embodiments, the training comprises the received factor expression levels in responders and non-responders and a clinical parameter selected from sex, age and PD-L1 expression, or any combination thereof, of each responder and non-responder and the machine learning algorithm is applied to individual received factor expression levels from the subject and the subject’s clinical parameters and wherein the machine learning algorithm outputs response prediction. In some embodiments, the training set comprises the number of resistance associated factors in each responder and non-responder and at least one clinical parameter and the machine learning algorithm is applied to the number of resistance associated factors from the subject and the subject’s at least one clinical parameters and wherein the machine learning algorithm outputs a response prediction. In some embodiments, the training set comprises the number of resistance associated factors in each responder and non-responder and sex of each responder and non-responder and the machine learning algorithm is applied to the number of resistance associated factors from the subject and the subject’s sex and wherein the machine learning algorithm outputs a response prediction. In some embodiments, the training set comprises the number of resistance associated factors in each responder and non-responder, age and PD-L1 status of each responder and non-responder and the machine learning algorithm is applied to the number of resistance associated factors from the subject and the subject’s age and PD-L1 status and wherein the machine learning algorithm outputs a response prediction.

[0112] In some embodiments, the training set comprises the received factor expression levels in responder and non-responders. In some embodiments, the training set comprises the received factor expression levels in responder and non-responders and a clinical parameter. In some embodiments, the training set comprises the received factor expression levels in responder and non-responders and sex of each of the responders and non- responders. In some embodiments, the trained machine learning algorithm is applied to individual received factor expression levels from the subject. In some embodiments, thetrained machine learning algorithm is applied to each received factor expression levels from the subject. In some embodiments, the trained machine learning algorithm is applied to individual received factor expression levels from the subject and a clinical parameter from the subject. In some embodiments, the trained machine learning algorithm is applied to individual received factor expression levels from the subject and the subject’s sex.

[0113] In some embodiments, the clinical parameter is the type of treatment. In some embodiments, the clinical parameter is expression of a target of the therapy. In some embodiments, the clinical parameter is expression of a protein within a process that is a target of the therapy. In some embodiments, the process is a process comprising the target of the therapy. In some embodiments, expression is expression in the subject. In some embodiments, expression is expression in a diseased tissue. In some embodiments, expression is expression in a diseased tissue sample. In some embodiments, expression is expression in the tumor. In some embodiments, expression is expression in a tumor sample. In some embodiments, a tumor sample is a biopsy. In some embodiments, expression is expression not in the tumor. In some embodiments, expression is expression not in a tumor sample. In some embodiments, expression is expression in a liquid biopsy. In some embodiments, expression is percent expression. In some embodiments, percent is percent of cells. In some embodiments, the therapy is anti-PD-1 therapy and the protein in the process is PD-L1. In some embodiments, the therapy is anti-PD-Ll therapy, and the target protein is PD-L1. In some embodiments, the clinical parameter is PD-L1 expression. In some embodiments the training set comprises at least one clinical parameter selected from line of treatment, PD-L1 expression, sex and age. In some embodiments the training set comprises protein expression levels and sex. In some embodiments the training set comprises number of RAPs, age and PD-L1 status.

[0114] Additionally clinical parameters may also be included. A skilled artisan will be able to select relevant clinical parameters for inclusion in the training set. Examples of additional clinical parameters include, but are not limited to, histological type of the sample (e.g., adenocarcinoma, squamous cell carcinoma, etc.), metastatic location, tumor location, cancer staging (such as tumor, nodes and metastases, TNM, staging for example), performance status (such as ECOG performance status), genetic mutations, epigenetic status, general medical history, vital signs, blood measurements, renal and liver function, weight, height, pulse, blood pressure and smoking history.

[0115] By another aspect there is provided, a computer program product comprising a non- transitory computer-readable storage medium having program code embodied thereon, the program code executable by at least one hardware processor to perform a method of the invention.

[0116] By another aspect, there is provided a system for performing a method of the invention.

[0117] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention.

[0118] By another aspect there is provided a system for predicting probability of CB of a treatment in a target patient of a first disease, the system comprising: a non-transitory memory device, wherein modules of instruction code are stored, and at least one processor associated with the memory device, and configured to execute the modules of instruction code, whereupon execution of said modules of instruction code, the at least one processor is configured to: receiving expression levels of proteins from a biological sample originating from the target patient; based on said received expression levels, determine expression levels of a first group of Resistance- Associated Proteins (RAPs) in the target patient, wherein said first group of RAPs are defined by differential levels of expression between CB and non- clinical benefit (NCB) patients of the first disease, in relation to the treatment; infer a first, pretrained machine-learning (ML) based prediction model on the expression levels of the first group of RAPs in the target patient, to determine a first CB score; identify at least one second group of RAPs, as having differential levels of expression between CB and NCB patients of at least one respective second disease, in relation to the treatment; based on the received expression levels, determine expression levels of a subset of the second group of RAPs in the target patient; infer a second, pretrained ML-based prediction model on the expression levels of the subset of the second group of RAPs in the target patient, to determine a second CB score; and calculate the CB probability based on the first CB score and the second CB score. For example, the CB probability may be calculated as a sum of the first CB score and the second CB score, a weighted sum of the first CB score and the second CB score, an average of the first CB score and the second CB score, a median between the first CB score and the second CB score, and the like.

[0119] In some embodiments, the method further comprises administering the treatment to a subject determined to be a responder. In some embodiments, the method further comprisesadministering the treatment to a subject probable to be a responder. In some embodiments, the method further comprises administering the treatment to a subject determined to have clinical benefit from the treatment. In some embodiments, the method further comprises administering the treatment to a subject determined to probably have clinical benefit from the treatment. In some embodiments, the method further comprises administering an alternative treatment to a subject determined to be a non-responder. In some embodiments, the method further comprises administering an alternative treatment to a subject probable to be a non-responder. In some embodiments, the method further comprises administering an alternative treatment to a subject determined not to have clinical benefit from the treatment. In some embodiments, the method further comprises administering an alternative treatment to a subject determined to probably not have clinical benefit from the treatment.

[0120] As used herein, the term "about" when combined with a value refers to plus and minus 10% of the reference value. For example, a length of about 1000 nanometers (nm) refers to a length of 1000 nm+- 100 nm.

[0121] It is noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polynucleotide" includes a plurality of such polynucleotides and reference to "the polypeptide" includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only" and the like in connection with the recitation of claim elements, or use of a "negative" limitation.

[0122] In those instances where a convention analogous to "at least one of A, B, and C, etc." is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."

[0123] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a singleembodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.

[0124] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents, unless the context clearly dictates otherwise. The terms “a” (or “an”) as well as the terms “one or more” and “at least one” can be used interchangeably.

[0125] Furthermore, “and / or” is to be taken as specific disclosure of each of the two specified features or components with or without the other. Thus, the term “and / or” as used in a phrase such as “A and / or B” is intended to include A and B, A or B, A (alone), and B (alone). Likewise, the term “and / or” as used in a phrase such as “A, B, and / or C” is intended to include A, B, and C; A, B, or C; A or B; A or C; B or C; A and B; A and C; B and C; A (alone); B (alone); and C (alone).

[0126] Wherever embodiments are described with the language “comprising,” otherwise analogous embodiments described in terms of “consisting of’ and / or “consisting essentially of’ are included.

[0127] Additional objects, advantages, and novel features of the present invention will become apparent to one ordinarily skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below finds experimental support in the following examples.

[0128] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.EXAMPLES

[0129] Generally, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are thoroughly explained in the literature. See, for example, "Molecular Cloning: A laboratory Manual" Sambrook et al., (1989); "Current Protocols inMolecular Biology" Volumes I-III Ausubel, R. M., ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York (1998); methodologies as set forth in U.S. Pat. Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I-III Cellis, J. E., ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley-Liss, N. Y. (1994), Third Edition; "Current Protocols in Immunology" Volumes I-III Coligan J. E., ed. (1994); Stites et al. (eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual" CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this document.Materials and Methods

[0130] Patient cohort and specimen collection:

[0131] Blood plasma samples and clinical data were collected from 79 small-cell lung cancer (SCLC) patients prior to commencement of ICI-based treatment. All patients were treated with ICI-based regimens including atezolizumab or durvalumab combined with chemotherapy or an ICI combination (ipilimumab plus nivolumab) as the first line treatment. In addition, plasma samples and clinical data were collected from 206 non-small cell lung cancer (NSCLC) patients prior to commencement of ICI-based treatment. All NSCLC patients were treated with ICI-based regimens including ICI alone or combined with chemotherapy.

[0132] Blood samples were collected from each patient into EDTA-anticoagulated tubes. Following plasma separation, the plasma samples were stored frozen at -80°C and were shipped frozen to the analysis laboratory.

[0133] Inclusion criteria: provision of informed consent; age older than 18 years; stage IIIB-IV NSCLC or SCLC; ECOG performance status 0-2; normal hematological, renal and liver functions. In addition, exclusion criterion was any concurrent and / or other active malignancy that required systemic treatment within 2 years prior to receiving the first dose of ICI-based treatment.

[0134] Information on Overall survival (OS) and progression-free survival (PFS) was evaluated in the patients. Clinical benefit data were retrieved from patient medical recordsand verified by the investigators through a review of radiologic images, i.e., CT chest / abdomen and brain MRI performed every 2-3 months, based on Response Evaluation Criteria In Solid Tumors (RECIST) 1.1 or radiological assessment. In the SCLC cohort, clinical benefit (CB) was assessed based on Progression Free Survival (PFS) at 6 months after the commencement of treatment. In the NSCLC cohort, clinical benefit was defined based on PFS at 12 months.

[0135] Proteomic measurements:

[0136] Proteomic profiling of plasma samples was performed using an assay (SomaScan) that simultaneously measures a total of 7596 protein targets. The assay utilizes chemically modified single-stranded oligonucleotides (modified DNA aptamers) that fold into specific molecular structures, allowing them to bind proteins with high affinity and specificity. This enables a broad and detailed analysis of the proteome. The measurement is performed using DNA microarray technology with a readout provided in relative fluorescence units (RFU).

[0137] Since protein level distributions are roughly log-normal (i.e., the logarithm of the measurement is normally distributed), and given that many statistical methods assume normality, log2 transformation was applied unless stated otherwise. There were no data imputations.

[0138] The proteomic dataset was narrowed down to a set of proteins with high analytical reliability.

[0139] Model development:

[0140] In order to predict CB probability in SCLC patients, two models were integrated by averaging the output from each model, which is the clinical benefit probability. The average outcome was transformed to deciles and scaled to values between 0 and 10, where values below 6.7 indicate a NEGATIVE result, while values equal or greater than 6.7 indicate a POSITIVE result.

[0141] Model 1 development: The model was developed using cross-validation by applying random sampling approach with multiple iterations. In each iteration the set was randomly divided into a train set and a test set (75% and 25% of the cohort, respectively) while maintaining the CB rate proportion. In each iteration, proteins displaying differential levels between CB and NCB patients were identified using Kolmogorov- Smirnov test, these proteins were termed Resistance- Associated Proteins (RAPs). A prediction model based on a single protein was constructed for each RAP using XGBoost algorithm. Based on the output of all RAPs, the CB probability was determined.

[0142] Model 2 development: Analyzing plasma proteins in cancer patients poses a challenge due to the intricate interplay between the tumor, the immune system, and other tissues. To decipher biological mechanisms associated with resistance to immunotherapy, an analysis was performed on a group of 388 pre-defined RAPs measured in plasma samples from 206 advanced-stage NSCLC patients prior to initiation of PD-1 / PD-L1 inhibitor-based treatment in the following manner. This group of 388 RAPs was determined based on similar methods as elaborated herein, e.g., in relation to model 1 above. It will be understood that within the group of 388 there were sometimes two different probes for the same protein. Thus, while 388 is referred to throughout, it will also be understood that this is actually 372 distinct proteins.

[0143] A correlation matrix was constructed for all 388 RAPs in all 206 patients under the basic assumption that proteins exhibiting similar expression levels may correlate with each other and might be involved in the same biological process and treatment resistance mechanisms. Selecting a representative protein from each cluster of proteins may therefore enable to reduce the number of features before applying machine learning algorithms to predict CB, while maintaining the biological input. Based on the correlation matrix, a weighted graph was constructed, in which every node corresponds to a protein, and the strength of the correlation (to power of 4) between each two proteins is depicted by the weight of the edge connecting their respective nodes. Using Louvain method, the proteins were grouped into clusters, and from each cluster, a hub protein was selected based on the degree centrality (total sum of the weights of all edges connected to the protein); a hub protein is the one with the highest degree of centrality. The hubs of the 4 largest clusters were selected for machine learning based model development. The model was developed on the NSCLC cohort based on Cox regression analysis using the progression-free survival (PFS) data, and the model output is the expected PFS. The model was developed for each of the four selected protein hubs - EPHB4, CLSTN3, BCHE, DEFB132. The expected PFS was normalized by dividing the result by the maximal expected PFS of the NSCLC train set. This was defined as the CB probability of model 2. In other words, processor 2 may train the at least one second model 130CM2 on the second cohort 20C2 based on Cox regression analysis, using progression-free survival (PFS) data in annotation 30CA2 for each of the selected protein hubs 120RS2, to predict the respective at least one second CB score 140CS2.

[0144] Model performance:

[0145] The performance of all models (model 1, model 2 and the hybrid model) was evaluated on the entire SCLC set using three metrics: (i) The area under the curve (AUC) ofthe receiver operating characteristics plot, (ii) Agreement between the predicted CB probability and the observed CB rate in terms of goodness of fit (R2of a linear regression), where the observed CB rate for each CB value was defined as the proportion of CB patients among a group of patients within the range of the CB probability ±0.15 window, (iii) By examining the hazard ratio (HR) for the positive population vs. the negative population, as calculated using Cox proportional hazard model.

[0146] Data analysis:

[0147] All data analyses were conducted using Python, Perseus computational platform and GraphPad Prism (San Diego, California, USA, graphpad.com).

[0148] Hazard ratios are reported with 95% confidence intervals and p-values. A level of 0.05 or lower was considered significant. The network of RAPs was generated based on STRING database. Enrichment analysis for the RAPs was done using Fisher exact test against the overall background of 1578 examined proteins (false discovery rate < 0.1). Potential source of the proteins was based on published databases such as the Human Protein Atlas (HP A) or Clinical Proteomic Tumor Analysis Consortium (CPTAC) analysis of lung cancer.Example 1: Response prediction based on hybrid model

[0149] The main clinical parameters of the SCLC cohort are described in Table 1 and Fig. 1. The SCLC cohort is balanced in terms of males and females (Fig. 1C). Most of the patients have performance status (ECOG PS) of 0-1 (Fig. ID), with lower ECOG in patients who experienced CB. Most of the patients (93.7%) received ICI as a first line of treatment, and most of them were treated using a combination of anti-PD-Ll and chemotherapy (Fig. IE). The median age was 65 (Fig. IB). Clinical benefit was defined based on PFS at 6 months; using this definition, 44 patients (56%) did not display clinical benefit (Fig. IF). CB patients had significantly longer OS compared to NCB patients (14.6 and 6.08 months, respectively.Fig. 1A).

[0150] Table 1 : clinical parameters of the SCLC cohort

[0151] The CB prediction was based on the integration of two predictive models, the SCLC RAP -based model (model 1) and the Cox -regression model (model 2), by averaging their output and scaling the result as a score between 0 and 10, that stratifies the patients into two groups, ‘POSITIVE’ (1 / 3 of the population, score 6.7-10) which is associated with high likelihood for clinical benefit or ‘NEGATIVE’ (2 / 3 of the population, score 0-6.6) which is associated with low likelihood for clinical benefit. The integrated model is termed the PROphetSCLC model. The development of the model as a hybrid model that integrates predictions from two distinct models enables the inclusion of potential SCLC-specific predictive biomarkers while mitigating bias from small sample size.

[0152] Model 1 involves identification of Resistance-Associated Proteins (RAPs) that are specific for SCLC population. This was done using a cross-validation approach, where a subset of 75% of the patients was randomly sampled for training and 25% for evaluation, while maintaining the CB proportion (stratification). The proteins that were differentially expressed between CB and NCB populations were identified using statistical test applied on the SCLC cohort. The statistical test applied here was Kolmogorov-Smirnov test. The 10 proteins with the lowest p-value were considered as Resistance-Associated Proteins (RAPs). This step (random selection of 75% of the SCLC population and applying Kolmogorov- Smirnov test to obtain the 10 proteins with the lowest p-value) was applied 80 times.Altogether, 146 RAPs were identified for the SCLC cohort (Table 3). Some of the proteins were selected in more than one iteration. Prediction model 130M1 may be trained based on the protein level of each RAP and the CB / NCB information in annotation 30CA1. In other words, for each patient, each RAP serves as an indicator for whether the patient will be determined as CB orNCB by applying a single-protein machine learning-based model using the XGBoost algorithm. The integration of all single-protein models provides the CB probability of the patient, which is the output of model 1 (also denoted CB score 140CS1). CB score 140CS1 is a value between 0 and 1 (Fig. 2 provides a general illustration of the model).

[0153] Model 2 was developed on the NSCLC cohort and applied on the SCLC cohort in the following manner (the process is illustrated in Fig. 2 and Fig. 3). In a cohort comprised of 206 NSCLC patients (Described in Table 2), the expression levels of 388 proteins were examined (this list of 388 proteins is based on a previous analysis of RAPs in a different NSCLC cohort that is disclosed in International Patent Applications W02024033930 and PCT / IL2025 / 050387, the contents of which are hereby incorporated by reference in their entirety, Table 3).

[0154] The correlation between the expression levels of each pair of proteins among the 388 RAPs was calculated. Based on the calculated correlations, a protein network was generated, where each node is a protein, and the correlation between nodes is defined by a line (edge), where the weight was determined by the correlation to the power of 4 (Fig. 4). Using Louvain method, the proteins were grouped into clusters. Overall, using this approach, 9 clusters of proteins were identified. In each cluster, the degree centrality (i.e., the sum of correlations it is involved in) was calculated for each protein and a hub protein was defined as the protein with the highest degree centrality. Focusing on the 4 largest clusters of proteins (cluster 8 with 134 proteins; Cluster 0 with 97 proteins; cluster 2 with 48 proteins; and cluster 5 with 27 proteins), four hub proteins were selected, one per each cluster, as features for machine learning model - EPHB4 (UniProt ID P54760), CLSTN3 (UniProt ID: Q9BQT9), BCHE (UniProt ID: P06276), and DEFB 132 (UniProt ID: Q7Z7B70). The train set was the NSCLC cohort. A Cox regression model was applied on the progression-free survival (PFS) data for the four proteins. The output of the model is the expected PFS, and it was normalized by dividing the result by the maximal expected PFS as measured in the NSCLC train set. The output of the model is the CB score of the patient, (also denoted CB score 140CS2). CB score 140CS2 is a value between 0 and 1 (Fig. 2 provides a general illustration of the model).

[0155] Table 2: clinical parameters of the NSCLC cohort

[0156] The performance of each model (model 1 - the 146-protein model developed based on the SCLC cohort, model 2 - the 4-protein model developed based on the NSCLC cohort, and the combined model) was examined on the SCLC cohort (Fig. 5). The AUCs of the ROC curves for model 1, model 2 and the combined model were 0.62, 0.63 and 0.63, respectively,with significant p-values (<0.05), as shown on the left-hand panels of Figs. 5 A, 5B and 5C respectively. When evaluating the correlation between the clinical benefit probability (the output of each model) and the clinical benefit rate, the three models showed different results, with goodness of fit (R2) values of 0.81, 0.85 and 0.93 for model 1, model 2 and the combined model, as shown on the middle panels of Figs. 5 A, 5B and 5C respectively. Altogether, all three models displayed a good fit, with the combined model showing the highest R2. The CB probability was further used to divide the patients into two subgroups- PROphet-POSITIVE and PROphet-NEGATIVE.

[0157] When evaluating the overall survival (OS) profile using Kaplan Meier plot, there was a significant difference between the two patient subgroups in two out of the three models, with hazard ration (HR) of 0.59 (p-value = 0.09), 0.48 (p-value = 0.02) and 0.47 (p-value = 0.02) for model 1, model 2 and the combined model, as shown on the right-hand panels of Figs. 5A, 5B and 5C respectively. Overall, when considering all three metrics for model performance evaluation, the combined model displayed the strongest performance. The fit between model 1 and model 2 is 0.37 (Pearson’s r) (Fig. 6).

[0158] Another RAP -based model was previously developed for NSCLC (see W02024033930). It was developed using a cohort comprised of 500 advanced stage NSCLC patients. The model was developed on one subset of samples (n=228) and then validated using another set of samples (n=272). The model is termed PROphetNSCLC. When applying the PROphetNSCLC model on the SCLC cohort, there was no significant difference between patients with PROphet POSITIVE and PROphet NEGATIVE result (Fig. 7A). When comparing the 388 RAPs of the PROphetNSCLC model and the 138 RAPs of the PROphetSCLC model (model 1 which is described here, Table 3), there was an overlap of overall 44 proteins between the two RAP sets (see Fig. 7B and Table 4), indicating that there are differences in the underlying biology between the two types of cancer which may explain why the PROphetNSCLC model was not predictive in the SCLC cohort. Notably, the proteins that overlap between the two models are mainly involved in extracellular matrix (ECM).

[0159] Table 3: List of NSCLC and SCLC RAPs

[0160] Table 4: 42 RAPs common to NSCLC and SCLC

[0161] Bioinformatic analysis of the SCLC RAPs showed that the RAPs potentially originate either from the host (normal tissues, immune cells) or the tumor (Fig. 8). Most of the proteins were highly expressed in all normal tissues, while 65 proteins were highly expressed in specific immune cells. Enrichment analysis of the SCLC RAPs (146 RAPs) showed that the RAPS are significantly enriched with multiple biological processes related to cancer progression or to lung cancer.

[0162] The significantly enriched categories are displayed in the protein-protein interaction map (Fig. 9A, Fig. 9B), where intermediate filament-related proteins, which may indicate involvement of invasion and metastasis, shows the largest enrichment factor. Additional significantly enriched processes include poor prognosis in lung cancer (involving 22 RAPs), extra cellular matrix (ECM), proteins related to lung cancer, replicative immortality and proteins that are highly expressed in all normal tissues.

[0163] Further exploration of the biological processes reveals the involvement of multiple cell adhesion-related proteins, different signaling cascades and metabolism, including glycan metabolism (Fig. 10). Among the RAPs related to metabolism that were identified are GSTM1, GSTM3, HK2, GSS, GSR, ADSS2, and DLST. RAPs associated with fibroblasts were also detected, including FGFR2, FGFR4, FGF23, CHAD, and CCN1. In addition, seven lung cancer poor prognosis-related proteins were found, such as STC1, GGH, POSTN, LRP8, INHBA, KRT18 and ANGPTL4.

[0164] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.

[0165] Reference is now made to Fig. 11, which is a block diagram depicting a computing device, which may be included within an embodiment of a system 10 for predicting probability of CB, according to some embodiments.

[0166] Computing device 1 may include a processor or controller 2 that may be, for example, a central processing unit (CPU) processor, a chip or any suitable computing or computationaldevice, an operating system 3, a memory 4, executable code 5, a storage system 6, input devices 7 and output devices 8. Processor 2 (or one or more controllers or processors, possibly across multiple units or devices) may be configured to carry out methods described herein, and / or to execute or act as the various modules, units, etc. More than one computing device 1 may be included in, and one or more computing devices 1 may act as the components of, a system 10 according to embodiments of the invention.

[0167] Operating system 3 may be or may include any code segment (e.g., one similar to executable code 5 described herein) designed and / or configured to perform tasks involving coordination, scheduling, arbitration, supervising, controlling or otherwise managing operation of computing device 1, for example, scheduling execution of software programs or tasks or enabling software programs or other modules or units to communicate. Operating system 3 may be a commercial operating system. It will be noted that an operating system 3 may be an optional component, e.g., in some embodiments, a system may include a computing device that does not require or include an operating system 3.

[0168] Memory 4 may be or may include, for example, a Random-Access Memory (RAM), a read only memory (ROM), a Dynamic RAM (DRAM), a Synchronous DRAM (SD-RAM), a double data rate (DDR) memory chip, a Flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short term memory unit, a long term memory unit, or other suitable memory units or storage units. Memory 4 may be or may include a plurality of possibly different memory units. Memory 4 may be a computer or processor non-transitory readable medium, or a computer non-transitory storage medium, e.g., a RAM. In one embodiment, a non- transitory storage medium such as memory 4, a hard disk drive, another storage device, etc. may store instructions or code which when executed by a processor may cause the processor to carry out methods as described herein.

[0169] Executable code 5 may be any executable code, e.g., an application, a program, a process, task, or script. Executable code 5 may be executed by processor or controller 2 possibly under control of operating system 3. For example, executable code 5 may be an application that may predict probability of CB as further described herein. Although, for the sake of clarity, a single item of executable code 5 is shown in Fig. 11, a system according to some embodiments of the invention may include a plurality of executable code segments similar to executable code 5 that may be loaded into memory 4 and cause processor 2 to carry out methods described herein.

[0170] Storage system 6 may be or may include, for example, a flash memory as known in the art, a memory that is internal to, or embedded in, a micro controller or chip as known in the art, a hard disk drive, a CD-Recordable (CD-R) drive, a Blu-ray disk (BD), a universal serial bus(USB) device or other suitable removable and / or fixed storage unit. Data pertaining to protein expression may be stored in storage system 6 and may be loaded from storage system 6 into memory 4 where it may be processed by processor or controller 2. In some embodiments, some of the components shown in Fig. 11 may be omitted. For example, memory 4 may be a nonvolatile memory having the storage capacity of storage system 6. Accordingly, although shown as a separate component, storage system 6 may be embedded or included in memory 4.

[0171] Input devices 7 may be or may include any suitable input devices, components, or systems, e.g., a detachable keyboard or keypad, a mouse and the like. Output devices 8 may include one or more (possibly detachable) displays or monitors, speakers and / or any other suitable output devices. Any applicable input / output (VO) devices may be connected to Computing device 1 as shown by blocks 7 and 8. For example, a wired or wireless network interface card (NIC), a universal serial bus (USB) device or external hard drive may be included in input devices 7 and / or output devices 8. It will be recognized that any suitable number of input devices 7 and output device 8 may be operatively connected to Computing device 1 as shown by blocks 7 and 8.

[0172] A system according to some embodiments of the invention may include components such as, but not limited to, a plurality of central processing units (CPU) or any other suitable multi-purpose or specific processors or controllers (e.g., similar to element 2), a plurality of input units, a plurality of output units, a plurality of memory units, and a plurality of storage units.

[0173] The term neural network (NN) or artificial neural network (ANN), e.g., a neural network implementing a machine learning (ML) or artificial intelligence (Al) function, may be used herein to refer to an information processing paradigm that may include nodes, referred to as neurons, organized into layers, with links between the neurons. The links may transfer signals between neurons and may be associated with weights. A NN may be configured or trained for a specific task, e.g., pattern recognition or classification. Training a NN for the specific task may involve adjusting these weights based on examples. Each neuron of an intermediate or last layer may receive an input signal, e.g., a weighted sum of output signals from other neurons, and may process the input signal using a linear or nonlinear function (e.g., an activation function). The results of the input and intermediate layers may be transferred to other neurons and the results of the output layer may be provided as the output of the NN. Typically, the neurons and links within a NN are represented by mathematical constructs, such as activation functions and matrices of data elements and weights. At least one processor (e.g., processor 2 of Fig. 11) such as one or more CPUs or graphics processing units (GPUs), or a dedicated hardware device may perform the relevant calculations.

[0174] Reference is now made to Fig. 12, which is a block diagram depicting an exemplary implementation of a system 10 for predicting probability of CB in a target patient.

[0175] Fig. 12 emphasizes aspects of inference, or application of system 10 on data pertaining to one or more specific patients (also referred to herein as “target” patients), suffering from a first disease (e.g., SCLC), to predict CB in relation to specific treatments of these patients, prior to actual application of these treatments. The preparation, or training of system 10 for that purpose is also discussed herein, e.g., in relation to Fig. 13A and Fig.13B

[0176] According to some embodiments of the invention, system 10 may be implemented as a software module, a hardware module, or any combination thereof. For example, system 10 may be, or may include a computing device such as element 1 of Fig. 11, and may be adapted to execute one or more modules of executable code (e.g., element 5 of Fig. 11) for predicting probability of CB, as described herein. As shown in Fig. 12, arrows may represent flow of one or more data elements to, and / or from system 10, and / or among modules or elements of system 10. Some arrows have been omitted in Fig. 12 for the purpose of clarity.

[0177] As shown in Fig. 12, system 10 may receive (e.g., via input 7 of Fig. 11) expression levels of a set of proteins 20 from a biological sample originating from the target patient.

[0178] The set of proteins 20 and their respective expression levels 20 may be used herein interchangeably.

[0179] The set of proteins 20 may include expression of a wide variety of proteins, from which system 10 may select specific groups of proteins to be analyzed. For example, expression levels data 20 may include pretreatment protein expression data 20CT1 of a first group of RAPs (also denoted RAP1), defined by differential levels of expression between CB and Non-Clinical Benefit (NCB) patients of a first disease (e.g., SCLC), in relation to a specific treatment. Additionally, or alternatively, the set of proteins 20 may include pretreatment protein expression data 20CT2 of at least one second group of RAPs (also denoted RAP2), defined by differential levels of expression between CB and NCB patients corresponding to at least one second, different disease (e.g., NSCLC), in relation to that same specific treatment.

[0180] Pretreatment protein expression data 20 may originate from a biological sample collected from the target patient including for example a tissue, a tumor sample, blood, plasma, serum, lymph, cerebral spinal fluid, urine, feces, semen, tumor fluid, gastric fluid, whole blood, peripheral blood mononuclear cells, and the like.

[0181] Protein expression data 20 may be obtained using various analytical techniques, including for example, mass spectrometry, proteomics arrays, proximity extension assay (PEA), aptamer-based assays, immunoblotting, immunohistochemistry, flow cytometry (FACS), or proteome sequencing, and the like.

[0182] As shown in Fig. 12, system 10 may include two parallel processing pathways that operate on different groups of RAPs (denoted RAP1, RAP2).

[0183] A first analysis module 100M1 (also referred to herein as “model 1”) may be adapted to select the first group (RAP1) of target patient RAPs from protein set 20. Analysis module 100M1 may subsequently process the expression data 20CT1 of the first group RAP1, thereby focusing on RAPs displaying differential levels of expression between CB and NCB patients of the first disease, in relation to a specific treatment.

[0184] A second analysis module 100M2 (also referred to herein as “model 2”) may operate in parallel to analysis module 100M1. Analysis module 100M2 may be adapted to select the at least one second group (RAP2) of target patient RAPs from protein set 20. Analysis module 100M2 may subsequently process the target patient RAP2 expression data 20CT2, focusing on RAPs displaying differential levels of expression between CB and NCB patients of the second, different disease, in relation to the specific treatment.

[0185] According to some embodiments, analysis module (“model 1”) 100M1 may incorporate a prediction model 130CM1 that may be trained to apply machine-learning based algorithms, to determine a first CB score 140CS1 based on RAP1 expression data 20CT1. The first CB score 140CS1 may represent a likelihood that the target patient will experience clinical benefit from the treatment based on the first group of RAPs. In some embodiments, the protein expression data 20 (and thus 20CT1) may be provided in absolute concentrations (e.g., ml, pl) or relative fluorescence units (RFU) or Normalized Protein expression (NPX) or other arbitrary units, which may be processed by the prediction model 130CM1 to generate CB score 140CS1 as a quantitative assessment of treatment response probability.

[0186] Additionally, or alternatively, prediction model 130CM2 may apply machinelearning based algorithms to the expression levels 20CT2 of a group of the at least one second group RAP2 of RAPs in protein data set 20, to determine a second CB score 140CS2. The second CB score 140CS2 may represent a likelihood that the target patient will experience clinical benefit from the treatment based on the at least one second group of RAPs (RAP2).

[0187] According to some embodiments, prediction model 130CM2 may be implemented as, or may include, a Cox regression model that processes the expression levels of the secondgroup of RAPs (RAP2) to determine the second CB score 140CS2. The Cox regression model 130CM2 may provide significant advantages for clinical benefit prediction by incorporating time-to-event data such as progression-free survival or overall survival, allowing the analysis module (model 2) 100M2 to capture temporal aspects of treatment response and resistance development. This survival-based approach may enable the prediction model 130CM2 to assess not only whether a target patient will experience clinical benefit, but also the expected duration and timing of that benefit, providing more comprehensive prognostic information than binary classification models.

[0188] Embodiments of the invention may include training the Cox regression model 130CM2 by utilizing a dataset comprising expression levels of proteins in a cohort of patients suffering from the second disease, along with corresponding survival outcomes and clinical benefit annotations. The training process may involve estimating hazard ratios for each RAP in the second group RAP2, where the model learns to associate specific protein expression patterns with different risks of treatment failure or disease progression overtime. The trained Cox regression model may output expected progression-free survival values that are subsequently normalized to generate the second CB score 140CS2, representing the likelihood that the target patient will experience clinical benefit from the treatment based on the second group of RAPs and their associated survival predictions.

[0189] The outputs from both analysis modules (100M1, 100M2) may be processed by a combination module 180, which may calculate a clinical benefit probability 40 based on the first CB score 140CS1 and the second CB score 140CS2.

[0190] Combination module 180 may implement mathematical algorithms that integrate the outputs from both prediction models to generate a unified assessment of treatment response likelihood. For example, combination module 180 may calculate clinical benefit probability 40 by averaging, or summing the first CB score 140CS1 and second CB score 140CS2, and applying a linear transformation on the result, to scale the result on a linear scale (e.g., in the range between 0 and 10). Combination module 180 may subsequently apply a predetermined threshold (e.g., 6.7) on the resulting clinical benefit probability 40 to obtain positive / negative classification of CB.

[0191] System 10 may then present a recommendation for treatment (e.g., via output device 8 of Fig. 11) of the target patient based on the calculated CB probability value 40. For example, when CB probability value 40 surpasses the predetermined threshold value, system 10 may output a message via a computing device of a physician, including recommendation for applying the relevant treatment to the target patient.

[0192] Additionally, or alternatively, system 10 may repeat the process of CB probability calculation, to obtain a plurality of CB probability values, each corresponding to respective treatment of a plurality of treatments. Based on the calculated CB probability values, system 10 may select a specific treatment, or indicate a treatment that is likelier to benefit the specific treatment (e.g., the treatment having the highest CB probability value). System 10 may then emit, or present a recommendation message (e.g., via output 8 of Fig. 11) for treatment of the target patient based on that selection.

[0193] In other words, clinical benefit probability 40 may provide clinicians with a qualitative measure (e.g., negative / positive) and / or a quantitative measure to guide treatment decisions for the target patient, where values above a precalculated threshold may indicate a higher likelihood of clinical benefit from the treatment, while values below the threshold may suggest a lower probability of favorable treatment response.

[0194] Reference is now made to Fig. 13A which is a block diagram, depicting an exemplary implementation of a training stage of a first analysis module 100M1 (also referred to as “model 1”), which may be included in a system 10 for predicting CB in a target patient, according to some embodiments of the invention. Analysis module 100M1 may be the same as analysis module 100M1 (“model 1”) of Fig. 2 and Fig. 12.

[0195] As explained herein, analysis module 100M1 may include a prediction model 130CM1 for determining CB probability based on resistance-associated proteins from a first disease.

[0196] System 10 may receive protein expression data 20C1 representing pretreatment protein expression levels from a first cohort of patients suffering from the first disease, along with cohort annotation data 30CA1 of the first disease.

[0197] Cohort annotation data 30CA1 may label one or more (e.g., each) patient of the first cohort as either a CB patient or a NCB patient in relation to a specific treatment.

[0198] In the example of Fig. 2, this first cohort includes patients suffering from small-cell lung cancer, and cohort annotation data 30CA1 may label specific patients as CB or NCB in relation to a specific type of treatment such as chemotherapy, a combination of Immune Checkpoint Inhibitor (ICI) therapy and chemotherapy, a combination of ICI therapies, and the like.

[0199] In some embodiments, protein expression data 20C1 may undergo a non-linear transformation that may enable statistical methods that assume normality to operate more effectively on the biological data. Such transformation may include, for example a logistic function or z-score standardization, where outliers may be replaced by threshold values.

[0200] In another example, protein expression data 20C1 may undergo a logarithmic (e.g., Iog2) transformation since protein level distributions are roughly log-normal.

[0201] Additionally, or alternatively, system 10 may include additional clinical parameters such as patient sex, age, ECOG performance status, treatment line, specific mutations, biomarkers expression, tumor location, metastasis, concomitant medications, adverse events, and treatment type as input features that may enhance the predictive capability of the analysis module 100M1.

[0202] Analysis module 100M1 may incorporate a RAP selection module 110R1 that may process the protein expression data 20C1 to identify proteins with differential expression patterns between CB and NCB patient populations. RAP selection module 110R1 may employ a statistical test module 120 that may analyze the expression levels of proteins in the first dataset to identify a first group 120R1 of RAPs, also referred to as RAPE

[0203] For example, statistical test module 120 may utilize statistical tests such as Z-score calculations, based on deviation from population means in responders and non-responders, with resistance scores calculated using mathematical formulas that may include regulation constants to prevent score divergence under certain conditions.

[0204] Additionally, or alternatively, statistical test module 120 may apply statistical tests that implement various analytical approaches such as a Mann -Whitney U test, T-tests, Cox regression, Cox proportional hazard regression, proportional hazard models in addition to Kolmogorov- Smirnov test, each offering distinct advantages for identifying resistance- associated proteins based on different statistical assumptions and data characteristics.

[0205] For example, statistics module 120 may implement a Kolmogorov- Smirnov test, which may be beneficial for comparing protein expression distributions between CB and NCB patient groups without requiring assumptions about the underlying distribution shape or normality. In other words, the Kolmogorov- Smirnov test may provide robust detection of differential expression patterns even when protein levels exhibit non-normal distributions commonly observed in biological samples.

[0206] In another example, the statistical tests module 120 may apply T-tests, which may offer computational efficiency and interpretability when protein expression data approximates normal distributions, and may be useful for identifying proteins with clear mean expression differences between patient groups.

[0207] In another example, statistical tests module 120 may employ a Cox regression model, a proportional hazard model or other equivalent models for selecting the first group 120R1 of RAPs. As known in the art, Cox regression and proportional hazard models may providethe advantage of incorporating time-to-event data such as progression-free survival or overall survival, allowing for identification of RAPs that may be associated with clinical benefit duration rather than limited binary outcomes. Such survival -based approaches may be beneficial for capturing the temporal aspects of treatment response and resistance development.

[0208] In another example, statistical tests module 120 may apply a Mann-Whitney U test to provide robust, non-parametric analysis of protein expression differences between CB and NCB patient groups without requiring assumptions about data normality or equal variances. As known in the art, the Mann-Whitney U test may be advantageous when protein expression data contains outliers or exhibits skewed distributions, as it compares median values and rank orders rather than means, making it less sensitive to extreme values that could skew results in parametric tests.

[0209] Other types of statistical tests may also be implemented by statistical tests module 120, as required by specific characteristics of the protein expression data 20C1 of the first cohort of patients.

[0210] With continued reference to Fig. 13A, analysis module 100M1 may include a training module, configured to train prediction model 130CM1 to generate predictions CB for specific target patients, based on the selected RAPs 120R1, using the cohort annotation data 30CA1 as supervisory data.

[0211] For example, training module 150T1 may receive inputs from both the RAP1 group 120R1 and the cohort annotation data (first disease) 30CA1 to construct the prediction model 130CM1 as an ensemble of one or more (e.g., a plurality of) decision trees. Training module 150T1 may, for example, use the XGBoost algorithm to train prediction model 130CM1. Training module 150T1 may thus configure each decision tree of the plurality of decision trees to use at least one annotation from the cohort annotation data 30CA1 as supervisory data, training that decision tree to predict an interim CB score, based on expression of a respective, unique protein of the RAP1 group 120R1. The ensemble approach may enable the prediction model 130CM1 to leverage the predictive power of multiple individual protein-based models while reducing overfitting and improving generalization to new patient data.

[0212] Training module 150T1 may subsequently configure the ensemble of decision trees to calculate the CB score 140CS1 based on the interim CB scores of the plurality of decision trees, creating a comprehensive assessment that may integrate information from multiple resistance-associated proteins. This training process may result in a fully trained predictionmodel 130CM1 that may output the CB score 140CS1 representing the likelihood that a target patient will experience clinical benefit from the treatment based on the first group 120R1 of RAPs associated with the first disease. Additionally, or alternatively, CB score 140CS1 may be calculated as an average of predictions from one or more (e.g., each) of the single-protein, interim CB scores. Additionally, or alternatively, CB score 140CS1 may be based on the number of RAPs predicting / indicating that the patient will be a CB or NCB.

[0213] It may be appreciated that the above-described architecture of prediction model 130CM1, which involves an aggregation of decision trees may be beneficial for training when dealing with small datasets.

[0214] Additionally, or alternatively, prediction model 130CM1 may be implemented using various machine learning architectures, e.g., beyond the ensemble of decision trees, each offering distinct advantages depending on the characteristics of the protein expression data 20C1 and 30CA1, and clinical requirements.

[0215] For example, Support Vector Machines (SVM) may provide robust performance when dealing with high-dimensional protein expression data, effectively finding optimal decision boundaries in complex feature spaces even with limited training data.

[0216] In another example, linear regression models may offer interpretability advantages, allowing clinicians to understand the direct contribution of each resistance-associated protein to the clinical benefit prediction, which may be valuable for clinical decision-making and regulatory approval processes.

[0217] In another example, Neural Network architectures, such as deep learning models, may capture complex non-linear relationships between protein expression patterns and clinical outcomes that simpler models might miss.

[0218] In another example, Logistic regression models may provide probabilistic outputs that directly correspond to clinical benefit likelihood, offering clear statistical interpretation of results with well-defined confidence intervals and p-values for each resistance-associated protein coefficient. The logistic regression architecture may be advantageous when binary clinical benefit outcomes are desired, as it naturally produces probability scores between 0 and 1 that can be easily interpreted by clinicians and integrated into existing clinical decision-making frameworks.

[0219] In another example, Bayesian models such as Naive Bayes may offer robust performance in scenarios with limited training data by incorporating prior knowledge about protein expression distributions and clinical outcomes. These Bayesian architectures may provide uncertainty quantification in their predictions, allowing clinicians to assess theconfidence level of clinical benefit predictions and make more informed treatment decisions. Naive Bayes models may be effective when dealing with high-dimensional protein expression data where feature independence assumptions are reasonable, and they may demonstrate computational efficiency that enables real-time clinical benefit assessment in resource-constrained environments.

[0220] Other types of architectures may also be used, to accommodate specific requirements and characteristics of the available cohort dataset.

[0221] Reference is now made to Fig. 13B which is a block diagram, depicting an exemplary implementation of a training stage of a second analysis module 100M2 (also referred to as “model 2”), which may be included in a system 10 for predicting CB in a target patient, according to some embodiments of the invention. Analysis module 100M2 may be the same as analysis module 100M2 (“model 2”) of Fig. 2 and Fig. 12.

[0222] As explained herein, analysis module 100M2 may include a second prediction model 130CM2 for determining clinical benefit probability based on resistance-associated proteins 120R2 (also referred to herein as RAP2), pertaining to a second disease.

[0223] System 10 may receive protein expression data 20C2 representing pretreatment protein expression levels from a second cohort of patients suffering from the second disease, along with cohort annotation data (second disease) 30CA2 that may label one or more (e.g., each) patient of the second cohort as either a CB patient or a NCB patient in relation to the same treatment.

[0224] The first disease and second disease may be different diseases, affecting the same tissue or cell type, allowing analysis module 100M2 (“model 2”) to leverage cross-disease patterns of resistance-associated protein expression that may provide complementary predictive information to analysis module 100M1 (“model 1”), in relation to the same treatment.

[0225] In the example of Fig. 2, the second cohort includes patients suffering from nonsmall cell lung cancer, and cohort annotation data 30CA2 may label specific patients of that cohort as CB or NCB in relation to the same specific treatment. In both cases subjects received ICI+chemotherapy, although the specific drugs were different. Thus, the treatment maybe the same or different between the two cohorts.

[0226] Analysis module 100M2 may incorporate a RAP selection module 110R2 that may process the protein expression data 20C2 to identify proteins with differential expression patterns between CB and NCB patient populations in the second disease context. RAP selection module 110R2 may employ a statistical test module 120 that may be the same asstatistical test module 120 of Fig. 13A. As explained herein, statistical test module 120 may apply statistical calculations on the expression levels of proteins in the second dataset to identify a second group of RAPs 120R2, also referred to as RAP2. Statistical test module 120 may analyze the protein expression data 20C2 using various analytical approaches, as explained herein in relation to Fig. 13A, to determine which proteins of the second cohort demonstrate differential expression between patients of the second disease who experience clinical benefit and those who do not. The function of statistical test module 120 will not be repeated herein for the sake of brevity.

[0227] With continued reference to Fig- 13B, a correlation module 160 may receive group 120R2 as input from the statistical test module 120 and may calculate a correlation matrix 160CM representing correlation of expression among RAPs of the RAP2 group 120R2.

[0228] In other words, correlation module 160 may analyze expression patterns across the second cohort of patients to identify proteins 20C2 that exhibit similar expression behaviors, operating under the assumption that proteins with correlated expression levels may be involved in similar biological processes and / or treatment resistance mechanisms.

[0229] According to some embodiments, RAP selection module 110R2 may include a clustering model 170, configured to group highly correlated RAPs 120R2 into clusters 170C, as depicted in the example of Fig. 2 (Clusters 1-4).

[0230] For example, correlation module 160 may generate a graph data element 160G that may include a plurality of nodes, interconnected by arcs, such as presented in the example of Fig. 4. In the graph data element 160G, each node may represent a specific RAP from the RAP2 group 120R2, and each arc may represent a value of correlation between expression of RAPs of the interconnected nodes.

[0231] Additionally, or alternatively, correlation module 160 may implement a hierarchical clustering, or consensus clustering algorithm which, as known in the art, may not require generating a graph data element 160G.

[0232] According to some embodiments, correlation module 160 may further apply mathematical transformations on the correlation values (e.g., raise their value to a predetermined power) as edge weights, to emphasize strong correlations, while diminishing weaker relationships. Correlation module 160 may thereby create a weighted graph 160G, or network structure that may facilitate more effective clustering of functionally related proteins.

[0233] The correlation module 160 may provide input (e.g., correlation matrix 160CM / graph 160G) to clustering model 170, that may group highly correlated RAPs 120R2 intoclusters 170C, For example, clustering model 170 may apply a graph analysis algorithm on graph data element 160G to collect the RAPs of the RAP2 group 120R2 into a plurality of clusters 170C. In such embodiments, clustering model 170 may implement the Louvain method for clustering proteins 20C2 (now RAPs 120R2) into groups, or clusters 170C based on the correlation matrix 160CM.

[0234] Clustering model 170 may thereby identify communities of proteins that exhibit strong internal correlations while maintaining weaker connections to proteins in other clusters. Clusters 170C may thus reflect correlation between expression levels of pairs or groups of RAPs of the RAP2 group 120R2, enabling the identification of functionally related protein groups that contribute to treatment resistance through coordinated biological pathways. The clustering approach may reduce the dimensionality of the protein expression data 20C2 while preserving biological relevance by grouping proteins that may participate in similar cellular processes or resistance mechanisms.

[0235] As further shown in Fig. 13B, RAP selection module 110R2 may output a RAP2 subset 120RS2, that may contain a subset of the RAP2 group 120R2 selected based on correlation patterns and cluster characteristics.

[0236] For example, RAP selection module 110R2 may select the subset of the plurality of clusters based on at least one topological criterion such as a number of RAPs within each cluster, a cluster diameter, a cluster centrality, a cluster morphology, and the like.

[0237] The number of RAPs within each cluster may serve as a primary selection criterion, where the RAP selection module 110R2 may prioritize larger clusters that contain sufficient protein members to provide robust predictive signals while maintaining statistical power for machine learning model development. Clusters with fewer than a predetermined threshold of RAPs may be excluded to ensure adequate representation of biological processes and reduce noise from sparsely populated protein groups.

[0238] Cluster diameter may represent the maximum distance between RAPs within a cluster, measured in terms of correlation coefficients or other similarity metrics. The RAP selection module 110R2 may select clusters with smaller diameters that indicate tighter functional relationships among constituent proteins, suggesting more coherent biological processes. Cluster centrality may measure the relative importance or connectivity of each cluster within the overall protein network, where the RAP selection module 110R2 may prioritize clusters that occupy central positions and demonstrate high interconnectivity with other protein groups.

[0239] Cluster morphology may encompass the geometric or structural characteristics of each cluster, including measures of compactness, elongation, or branching patterns within the correlation network. The RAP selection module 110R2 may select clusters with specific morphological properties that indicate distinct biological functions or regulatory mechanisms.

[0240] Additionally, or alternatively, RAP selection module 110R2 may select the subset of clusters 170C based on at least one criterion of significance, selected from a list consisting of: a statistical significance metric, a biological significance metric, a clinical significance metric, a functional significance metric, and the like.

[0241] Statistical significance metrics may include, for example p-values from differential expression analysis, false discovery rates, or confidence intervals that indicate the reliability of observed expression differences between CB and NCB patient populations. These statistical measures may ensure that RAP selection module 110R2 selects clusters containing proteins with robust and reproducible expression patterns that are unlikely to result from random variation.

[0242] Biological significance metrics may encompass pathway enrichment scores, proteinprotein interaction strengths, or functional annotation clustering that reflects the underlying biological processes involved in treatment resistance. For example, RAP selection module 110R2 may prioritize clusters enriched for extracellular matrix proteins, immune signaling pathways, or metabolic enzymes that demonstrate higher biological significance due to their established roles in cancer progression and therapeutic response.

[0243] Clinical significance metrics may include, for example, hazard ratios from survival analysis, correlation coefficients with clinical outcomes such as progression-free survival or overall survival duration, and the like.

[0244] Functional significance metrics may evaluate the molecular functions, cellular components, or biological processes represented within each cluster based on gene ontology annotations or pathway databases. RAP selection module 110R2 may select clusters containing proteins involved in drug metabolism, DNA repair mechanisms, or immune checkpoint regulation that demonstrate higher functional significance for predicting treatment response.

[0245] Additionally, or alternatively, RAP selection module 110R2 may select the subset of the plurality of clusters 170C based on at least one correlation parameter such as a strength of correlation among RAPs of the cluster 170, a variance of correlations among RAPs of the cluster 170C, and the like.

[0246] The strength of correlation among RAPs of the cluster may, for example, be measured as the average correlation coefficient between all pairs of proteins within each cluster, where RAP selection module 110R2 may prioritize clusters exhibiting higher mean correlation values that indicate stronger functional relationships among constituent proteins. Clusters with high correlation strength may represent proteins that are co-regulated through common biological pathways or participate in coordinated cellular responses to treatment, making them more likely to provide coherent predictive signals for clinical benefit assessment.

[0247] The variance of correlations among RAPs of the cluster may quantify the heterogeneity of correlation strengths within each cluster, where lower variance values indicate more uniform correlation patterns among cluster members. RAP selection module 110R2 may, for example, select clusters with low correlation variance that demonstrate consistent inter-protein relationships, suggesting stable functional associations that may be more reliable for predictive modeling.

[0248] RAP selection module 110R2 may apply these correlation parameters individually or in combination with other selection criteria to identify clusters that exhibit both strong and consistent correlation patterns. This approach may enhance the biological relevance and predictive accuracy of the selected protein groups by ensuring that chosen clusters represent well-defined functional units with stable inter-protein relationships across different patient populations.

[0249] Within each cluster 170C of the subset of clusters, RAP selection module 110R2 may identify a hub RAP as one whose expression is highly correlated (e.g., beyond a predetermined threshold), or the most correlated to other RAPs of that cluster. RAP selection module 110R2 may thereby select the hub RAPs of each of the selected clusters 170C, to obtain subset 120RS2 of the at least one second group 120R2 of RAPs.

[0250] The hub RAP identification may be based, for example, on degree centrality, betweenness, closeness, or eigenvector measures within each cluster, where degree centrality may calculate the total sum of correlations that a protein is involved in, providing a quantitative measure of each protein's connectivity and potential influence within the cluster network.

[0251] Analysis module 100M2 may provide the RAP2 subset 120RS2 as input to a training module 150T2. Training module 150T2 may be configured to apply machine-learning algorithms to the expression levels of the RAPs of subset 120RS2, to train a machinelearning based prediction model 130CM2. Prediction model 130CM2 may thus be trained topredict clinical benefit score 140CS2 based on the representative proteins from one or more (e.g., each) of the selected clusters 170C, using the cohort annotation data 30CA2 as supervisory information.

[0252] Reference is now made to Fig. 14 which is a flow diagram depicting a method of predicting probability of CB (e.g., element 40 of Fig. 12) of treatment in a target patient by at least one processor (e.g., processor 2 of Fig. 11).

[0253] As shown in step SI 005, the at least one processor 2 may receive (e.g., via input device 7 of Fig. 11) expression levels of a set of proteins (e.g., element 20 of Fig. 12) from a biological sample originating from the target patient.

[0254] The at least one processor 2 may select (step S1010) a first group of RAPs (e.g., RAP1 expression level 20CT1 of Fig. 12) from the set of proteins 20. The first group of RAPs 20CT1 may be defined by differential levels of expression 20CT1 between CB and NCB patients of the first disease, in relation to the treatment.

[0255] As shown in step S 1015, the at least one processor 2 may apply a first ML-based, prediction model (e.g., model 1, 100M1 of Fig. 12) on the expression levels of the first group ofRAPs 20CTl, to determine a first CB score (e.g., 140CS1 of Fig. 13A). The first CB score 140CS1 may represent a likelihood that the target patient will experience clinical benefit from the treatment based on expression of the first group of RAPs 20CT1.

[0256] The at least one processor 2 may further identify (step SI 020) at least one second group of RAPs (e.g., 20CT2 of Fig. 12) in the set of proteins 20, as having differential levels of expression between CB and NCB patients of at least one respective, second, different disease, in relation to the treatment.

[0257] As shown in step SI 025, the at least one processor 2 may apply at least one second, ML-based, prediction model (e.g., model 2, 100M2 of Fig. 12) on the expression levels of the at least one respective second group of RAPs 20CT2, to determine at least one second CB score (e.g., 140CS2 of Fig. 13B). Each CB score of the at least one second CB score 140CS2 may represent a likelihood that the target patient will experience clinical benefit from the treatment based on expression of a respective group of the at least one second group of RAPs 20CT2.

[0258] As explained herein, the at least one processor 2 may employ a combination module (e.g., 180 of Fig. 12), to calculate the CB probability 40 (step S1030) based on the first CB score 140CS1 and the at least one second CB score 140CS2.

[0259] As shown in step S1035, the at least one processor 2 may repeat steps S1015-S1030, to calculate a plurality of CB probability values 40. At least one (e.g., each) CB probability value 40 may respectively correspond to a specific treatment for the target patient.

[0260] The at least one processor 2 may subsequently select (step SI 040) a specific treatment based on its respective CB probability value 40. For example, the at least one processor 2 may select an optimal treatment for at least one target patient as one having a maximal CB probability value 40. Additionally, or alternatively, processor 2 may recommend on treatments for which there is high probability that the patient will respond to, or conversely indicate treatments would be less beneficial for the patient.

[0261] The at least one processor 2 may further present a recommendation (e.g., 40REC of Fig. 12) for treatment of the target patient based on that selection (step SI 045). For example, the at least one processor 2 may send an electronic message (e.g., an email message) or present a notification via output device 8 of Fig. 11, to a computing device of a person of interest, such as a physician or caregiver. The at least one processor 2 may thus notify the person of interest, advising them of the recommended, optimal treatment for the target patient’s disease.

[0262] As elaborated herein, the output of the at least one first model 100M1 is referred to as a first “CB score” 140CS1, and the output of the at least one second model 100M2 is referred to as at least one second “CB score” 140CS2, the value of which may be analyzed, or combined to generate “CB probability” 40CB . It may be appreciated that the terms “score” and “probability” may be used interchangeably, e.g., where the outcome of models 100M1, 100M2 are referred to as “probabilities” (e.g., values between 0 and 1), whereas the outcome 40CB may be referred to as a “score”, e.g., having values between 0 and 10, or -5 and 5.

Claims

CLAIMS1. A method of predicting probability of Clinical Benefit (CB) of a treatment in a target patient suffering from a first disease by at least one processor, the method comprising: receiving expression levels of a set of proteins from a biological sample originating from the target patient; selecting a first group of Resistance-Associated Proteins (RAPs) from the set of proteins, wherein the first group of RAPs are defined by differential levels of expression between CB and Non-Clinical Benefit (NCB) patients of the first disease, in relation to the treatment; applying a first, machine-learning (ML) based, prediction model on the expression levels of the first group of RAPs, to determine a first CB score, representing a likelihood that the target patient will experience clinical benefit from the treatment based on the first group of RAPs; identifying at least one second group of RAPs in the set of proteins, as having differential levels of expression between CB and NCB patients of at least one respective second, different disease, in relation to the treatment; applying a second, ML-based, prediction model on the expression levels of the at least one second group of RAPs, to determine at least one second CB score, representing a likelihood that the target patient will experience clinical benefit from the treatment based on the at least one second group of RAPs; and calculating the CB probability based on the first CB score and the at least one second CB score.

2. The method of claim 1, further comprising: identifying a subset of the second group of RAPs based on correlation between expression levels of the second group of RAPs in a cohort of patients of a specific disease of the at least one second disease; and applying the second ML-based prediction model on the expression levels of the subset of the second group of RAPs, to determine the at least one second CB score.

3. The method of claim 1, further comprising: obtaining a first dataset, comprising: (i) expression levels of proteins in a first cohort of patients, suffering from the first disease, and (ii) at least one annotation, labeling at least one patient of the first cohort as either a CB patient or a NCB patient in relation to the treatment; andapplying a statistical test on the expression levels of proteins in the first dataset, to identify the first group of RAPs.

4. The method of claim 3, further comprising: constructing the first prediction model as an ensemble of a plurality of decision trees; for each decision tree of the plurality of decision trees, using at least one annotation as supervisory data, to train that decision tree so as to predict an interim CB score, based on expression of a respective, unique protein of the first group of RAPs; and configuring the ensemble of decision trees to calculate the first CB score based on the interim CB scores of the plurality of decision trees.

5. The method of claim 1, further comprising: obtaining a second dataset, comprising (i) expression levels of proteins in a second cohort of patients, suffering from a specific disease of the at least one second disease, and (ii) at least one annotation, labeling at least one patient of the second cohort of patients as either a CB patient or an NCB patient in relation to the treatment; and applying a statistical calculation on the expression levels of proteins in the second dataset, to identify the second group of RAPs.

6. The method of claim 2, further comprising: calculating a correlation matrix, representing correlation of expression among RAPs of a specific group of RAPs of the at least one second group; based on the correlation matrix, generating a clustering model comprising a plurality of clusters, each reflecting correlation between expression levels of pairs of RAPs of the specific group; selecting a subset of the plurality of clusters; and in each cluster of the subset of clusters, identifying a hub RAP as one whose expression is most correlated to other RAPs of that cluster, thereby selecting the subset of the specific group of RAPs of the at least one second group of RAPs.

7. The method of claim 6, wherein selecting the subset of the plurality of clusters is based on at least one criterion of significance, selected from a list consisting of: a statistical significance metric, a biological significance metric, a clinical significance metric and a functional significance metric.

8. The method of any one of claims 6-7, wherein selecting the subset of the plurality of clusters is based on at least one topological criterion selected from a list consisting of: a number of RAPs within each cluster, a cluster diameter, a cluster centrality, and a cluster morphology.

9. The method of any one of claims 6-8, wherein selecting the subset of the plurality of clusters is based on at least one correlation parameter selected from a list consisting of: a strength of correlation among RAPs of the cluster, a variance of correlations among RAPs of the cluster.

10. The method of any one of claims 6-9, wherein generating the clustering model comprises: based on the correlation matrix, calculating a graph data element comprising a plurality of nodes interconnected by arcs, wherein each node represents a specific RAP, and each arc represents a value of correlation between RAPs of the interconnected nodes; and applying a graph analysis algorithm on the graph data element, to collect the RAPs of the specific group of RAPs of the at least one second group of RAPs into the plurality of clusters.

11. The method of any one of claims 6-10 further comprising training the at least one second model on the second cohort based on Cox regression analysis, using progression-free survival (PFS) data for each of the selected protein hubs, to predict the respective at least one second CB score.

12. The method of any one of claims 1 to 11, wherein said subset of the at least one second group of RAPs does not comprise RAPs from said first group of RAPs.

13. The method of any one of claims 1 to 12, wherein said first disease and at least one second disease are a first and at least one second type of cancer.

14. The method of any one of claims 1 to 13, wherein said first disease and at least one second disease afflict the same tissue or cell type in said subject.

15. The method of claim 13 or 14, wherein said first type of cancer and said at least one second type of cancer are cancers of the same tissue.

16. The method of any one of claims 1 to 15, wherein said clinical benefit is overall survival.

17. The method of any one of claims 1 to 15, wherein said clinical benefit is progression free survival.

18. The method of any one of claims 1 to 17, wherein said treatment is an immunotherapy.

19. The method of any one of claims 1 to 18, wherein said statistical test is a Kolmogorov- Smirnov test.

20. The method of any one of claims 1 to 19, wherein the biological sample originating from the target patient is obtained before said target patient received said treatment.

21. The method of any one of claims 1 to 20, wherein said biological sample is selected from blood plasma, whole blood, blood serum or peripheral blood mononuclear cells.

22. A system for predicting Clinical Benefit (CB) of a treatment in a target patient of a first disease, the system comprising: a non-transitory memory device, wherein modules of instruction code are stored, and at least one processor associated with the memory device, and configured to execute the modules of instruction code, whereupon execution of said modules of instruction code, the at least one processor is configured to: receive expression levels of a set of proteins from a biological sample originating from the target patient; select a first group of Resistance-Associated Proteins (RAPs) from the set of proteins, wherein the first group of RAPs are defined by differential levels of expression between CB and Non-Clinical Benefit (NCB) patients of the first disease, in relation to the treatment; apply a first, machine-learning (ML) based, prediction model on the expression levels of the first group of RAPs, to determine a first CB score, representing a likelihood that the target patient will experience clinical benefit from the treatment based on the first group of RAPs; identify at least one second group of RAPs in the set of proteins, as having differential levels of expression between CB and NCB patients of at least one respective, second, different disease, in relation to the treatment; apply a second, ML-based, prediction model on the expression levels of at least one specific group of the at least one second group of RAPs to determine at least one second CB score, representing a likelihood that the target patient will experience clinical benefit from the treatment based on the at least one specific group of RAPs; and calculate the CB probability based on the first CB score and the at least one second CB score.

23. The system of claim 22, wherein the at least one processor is further configured to present a recommendation for treatment of the target patient based on the calculated CB probability score.

24. The system according to any one of claims 22-23, wherein the at least one processor is further configured to: calculate a plurality of CB probability scores, respectively corresponding to a plurality of treatments; select a specific treatment based on its respective CB probability score; and present a recommendation for treatment of the target patient based on said selection.

Citation Information

Patent Citations

  • Systems, Devices and Methods for Constructing and Using a Biomarker

    US20170218456A1

  • Methods and systems for determining responders to treatment

    US20210295952A1

  • Machine learning prediction of therapy response

    US20230049979A1