Plasma proteomic biomarkers in lung cancer
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE FRANCIS CRICK INST LTD
- Filing Date
- 2025-12-11
- Publication Date
- 2026-07-23
AI Technical Summary
Current lung cancer screening methods, such as low-dose CT and risk prediction models, are limited by cost, accessibility, and low positive predictive value, missing opportunities for early detection in individuals without traditional risk factors like smoking, particularly in lower-resource settings.
A proteomic signature comprising the expression levels of proteins CXCL17, WFDC2, CEACAM5, and additional proteins (e.g., ALPP, GDF15, MMP12) is used to predict and detect lung cancer risk, applicable in primary care settings, independent of demographic factors.
The proteomic signature effectively predicts lung cancer risk up to 5 years in advance and facilitates early detection, improving survival rates by identifying individuals previously undetected by current methods.
Smart Images

Figure EP2025086701_23072026_PF_FP_ABST
Abstract
Description
[0001] PLASMA PROTEOMIC BIOMARKERS IN LUNG CANCER
[0002] FIELD OF THE INVENTION
[0003] The present invention relates to the use of proteins as biomarkers for identifying individuals with, or at risk of developing, lung cancer.
[0004] BACKGROUND
[0005] Lung cancer (LC) remains the leading cause of global cancer mortality (accounting for 1.8 million deaths in 2022 alone). Despite significant advancements in therapies over the past several decades,145.6% of patients in England are diagnosed with stage IV disease which has a dire 5- year survival of 4.3%. In comparison, 5-year survival at stage I was 62.7%, underscoring the need for early detection strategies, particularly in high-risk groups.2
[0006] There is clear unmet need to identify individuals who have never smoked but are nevertheless at risk of LC, since current LC screening guidelines are based exclusively on age and smoking history.3Consequently, individuals who have never smoked are explicitly excluded, despite accounting for up to 20% of LC cases in the UK and LC in never smokers (LCINS) being the fifth leading cause of cancer-related deaths.4This results in missed opportunities for early detection when tumours are at lower stages, less invasive, and more responsive to therapy.4Furthermore, whilst low-dose CT screening has proven effective in reducing LC mortality, its implementation is impeded by cost, accessibility, and scalability. Consequently, it is predominantly accessible in high-income countries, rendering it infeasible for widespread use in lower-resource settings.3This underscores the urgent need for more inclusive, affordable, and scalable risk prediction and early detection strategies.
[0007] In addition, whilst LC risk prediction models integrating patient and demographic factors (such as age, smoking history and BMI) exist, e.g.: (the Liverpool Lung Project, PLCOm2012 model), they suffer from low positive predictive values. This limits their utility for secondary prevention strategies and uses as inclusion criteria for prevention trials due to the large number of patients required in a trial setting to demonstrate efficacy.8
[0008] Thus, in view of the above, there remains a need to identify further individuals (particularly those lacking demographic risk factors, such as smoking status, that have previously been associated with LC) with, or at risk of developing, LC. SUMMARY OF THE INVENTION
[0009] The inventors surprisingly discovered a proteomic signature that can be used as an effective tool for identifying individuals with, or at risk of developing, lung cancer. The invention is minimally invasive, more practical and cost-effective than current screening methods, and can be implemented in primary care settings reducing the reliance on hospital-based, resource-intensive procedures. The invention provides a scalable and biologically validated foundation for early detection and prevention strategies. The proteomic signature can be used for risk prediction and / or early detection of lung cancer. In particular, it was unexpected that the proteomic signature of the invention was able to effectively predict the risk of an individual developing lung cancer in advance (e.g., 1-5 years) from being clinically diagnosed with lung cancer and despite the individual also lacking known demographic risk factors (e.g., smoking status). Further, it was also unexpected that the same proteomic signature would be effective in both early detection and risk prediction of lung cancer since the latter would typically rely on the protein signature being reflective of changes in the microenvironment of the individual, rather than the overt malignancy. Given the dwell time of lung cancer is thought to be around 1.5-2 years, it appears to be that this signature reflects the risk of an individual developing lung cancer rather than that of overt malignancy14.
[0010] Generally, the methods described herein relate to methods of predicting lung cancer, for example determining that a subject is of risk of developing lung cancer.
[0011] In a first aspect of the invention, there is provided a method of predicting, or determining the risk of, a subject developing lung cancer, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. predicting, or determining the risk of, the subject developing lung cancer if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0012] In a second aspect of the invention, there is provided a method of diagnosing lung cancer in a subject, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. diagnosing lung cancer in the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0013] In a third aspect of the invention, there is provided a method of classifying a subject, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. classifying the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0014] In a fourth aspect of the invention, there is provided a method of treating a subject: i) predicted to develop, or be at risk of developing, lung cancer according to the method described herein; ii) diagnosed with lung cancer according to the method described herein; or iii) classified according to the method described herein; wherein the method comprises administering a treatment for lung cancer to the subject.
[0015] In a fifth aspect of the invention, there is provided a method of treating lung cancer in a subject, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. administering a treatment for lung cancer to the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0016] In a sixth aspect of the invention, there is provided a method of processing a sample obtained from a subject, the method comprising: a. providing the sample obtained from the subject; b. processing the sample for determination of the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 c. determining the level of expression of each of the proteins in part b.; d. comparing the level of expression of each of the proteins determined in part c. to a reference value for each protein.
[0017] In a seventh aspect of the invention, there is provided a method of predicting, or determining the risk of, a subject developing lung cancer, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample obtained from the subject; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. predicting, or determining the risk of, the subject developing lung cancer if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0018] In an eighth aspect of the invention, there is provided a method of diagnosing lung cancer in a subject, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample from the subject; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. diagnosing lung cancer in the subject if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins. In a ninth aspect of the invention, there is provided a method of classifying a subject, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample obtained from the subject; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. classifying the subject if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0019] In a tenth aspect of the invention, there is provided a method of treating a subject: i) predicted to develop, or be at risk of developing lung cancer according to the method described herein; ii) diagnosed with lung cancer according to the method described herein; or iii) classified according to the method described herein; wherein the method comprises administering a treatment for lung cancer to the subject.
[0020] In an eleventh aspect of the invention, there is provided a method of treating lung cancer in a subject, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample from the subject; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. administering a treatment for lung cancer to the subject if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0021] In a twelfth aspect of the invention, there is provided a kit for predicting the presence or absence of lung cancer in a subject, wherein the kit comprises means for determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 and a sample collection apparatus.
[0022] BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 : Overview of the machine learning framework used to predict lung cancer incidence
[0024] Overview of the machine learning framework used to predict lung cancer from 48,099 individuals from the UK Biobank. A stratified train-test split of 75% of the data was used to train the data using a machine learning algorithm, which was then tested on 25% held-out data. Figure 2: Modelling framework utilising the proposed 15 proteins outperforms existing lung cancer risk prediction benchmark models
[0025] Figure 2A: On unseen data comprising of 12,025 patients, the model outperformed previous lung cancer risk models (p< 0.001 for comparison against the LCRAT and LLPv3 model5). These are demographic models which use age- and smoking- based cut-offs in order to predict the probability of an individual developing lung cancer and thus have been proposed for stratifying individuals who might benefit from screening.
[0026] Figure 2B: The model outperformed previous lung cancer risk models in the time bins 0-2 years and 2-4 years prior to diagnosis.
[0027] Figure 3: Proposed proteins associate with lung cancer incidence across multiple cohorts Meta analysis across five external proteomics cohorts demonstrating that the proteins of interest are associated with lung cancer incidence, after adjusting for the effects of age and sex.
[0028] Figure 4: Longitudinal proteomics sampling demonstrates that signature proteins increase prior to diagnosis of lung cancers in smokers and never-smokers
[0029] 3 / 15 (CEACAM5, CXCL17, WFDC2) proteins were measured in 100 women, five times, 1 year apart, prior to diagnosis of lung cancer. This was compared with 150 age-matched controls taken from a longitudinal ovarian cancer screening trial. The difference in protein levels between women who develop LC and those who do not is significant 2 years prior to lung cancer diagnosis. Data taken from a sub-study analysis of the UKCTOCS trial.6
[0030] Figure 5. Proteins of interest are enriched in healthy lung tissue
[0031] Using bulk-RNAseq analysis from the healthy tissues of individuals following death, it was found that the proteins of interest were highest expressed in the lung epithelial.
[0032] Figure 6. Proteins of interested are enriched in individuals from a pre-dominantly never smoker population
[0033] 4 / 10 proteins of interest (CEACAM5, CXCL17, WFDC2 and ALPP) were measured in 251 individuals who later developed a diagnosis of lung cancer and were compared with 502 age and sex matched controls. There was a significant difference with these proteins in cases compared with controls (p<0.05 by Wald test).
[0034] DETAILED DESCRIPTION
[0035] The present invention uses a proteomic signature comprising the level of expression of a plurality of proteins in order to identify a new population of patients at risk of developing or having lung cancer. The surprising identification of such patients enables earlier detection and access to new treatment options.
[0036] Definitions
[0037] Below are provided certain definitions of terms, technical means, and embodiments used herein.
[0038] The proteins may be selected from those shown in Tables 1-4. The proteins may be selected from those shown in Tables 1 or 2. The proteins may be selected from those shown in Tables 3 or 4.
[0039] The proteins may comprise of CXCL17, WFDC2 and CEACAM5 and at least 1 other protein selected from the group consisting of ALPP, GDF15, MMP12, TNFSF13B, LAMP3 and SFTPA1. The proteins may comprise of CXCL17, WFDC2, ALPP, GDF15, MMP12, TNFSF13B, LAMP3 and SFTPA1.
[0040] The proteins may comprise of CXCL17, WFDC2 and CEACAM5 and at least 1 other protein selected from the group consisting of ALPP, GDF15, MMP12, CDCP1 , LAMP3 and PLAUR. The proteins may comprise of CXCL17, WFDC2, ALPP, GDF15, MMP12, CDCP1 , LAMP3 and PLAUR.
[0041] The proteins may comprise of CXCL17, WFDC2 and CEACAM5 and at least 1 other protein selected from the group consisting of ALPP, GDF15, MMP12, CDCP1 , LAMP3, SFTPA1 , TNSF13B and PLAUR. The proteins may comprise of CXCL17, WFDC2, ALPP, GDF15, MMP12, CDCP1 , LAMP3, SFTPA1 , TNSF13B and PLAUR.
[0042] The proteins may consist of CXCL17, WFDC2 and CEACAM5 and at least 1 other protein selected from the group consisting of ALPP, GDF15, MMP12, TNFSF13B, LAMP3 and SFTPA1. The proteins may consist of CXCL17, WFDC2, ALPP, GDF15, MMP12, TNFSF13B, LAMP3 and SFTPA1.
[0043] The proteins may consist of CXCL17, WFDC2 and CEACAM5 and at least 1 other protein selected from the group consisting of ALPP, GDF15, MMP12, CDCP1 , LAMP3 and PLAUR. The proteins may consist of CXCL17, WFDC2, ALPP, GDF15, MMP12, TNFSF13B, LAMP3 and PLAUR.
[0044] The proteins may consist of CXCL17, WFDC2 and CEACAM5 and at least 1 other protein selected from the group consisting of ALPP, GDF15, MMP12, CDCP1 , LAMP3 and PLAUR. The proteins may consist of CXCL17, WFDC2, ALPP, GDF15, MMP12, TNFSF13B, LAMP3 and PLAUR.
[0045] In some embodiments, the proteins may be CXCL17, WFDC2 and CEACAM5 and at least 1, at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. BisGDF15;
[0046] Hi. C is MMP12; iv. DisTNFSF13B; v. E is LAMP3; and vi. FisSFTPAI.
[0047] In some embodiments, the proteins may be CXCL17, WFDC2 and CEACAM5 and at least 1, at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. BisGDF15;
[0048] Hi. C is MMP12; iv. DisCDCPI; v. E is LAMP3; and vi. FisPLAUR.
[0049] In some embodiments, the proteins may be CXCL17, WFDC2 and CEACAM5 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is ALPP; ii. BisGDF15;
[0050] Hi. C is MMP12; iv. DisCDCPI; v. E is LAMP3; vi. FisPLAUR; vii G is SFTPA1; and vii HisTNSF13B.
[0051] In some embodiments, the proteins may consist of CXCL17, WFDC2 and CEACAM5 and at least 1, at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. AisALPP; ii. BisGDF15; iii. C is MMP12; iv. DisTNFSF13B; v. E is LAMP3; and vi. FisSFTPAI.
[0052] In some embodiments, the proteins may consist of CXCL17, WFDC2 and CEACAM5 and at least 1, at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. AisALPP; ii. BisGDF15; iii. C is MMP12; iv. DisCDCPI; v. E is LAMP3; and vi. FisPLAUR.
[0053] In some embodiments, the proteins may consist of CXCL17, WFDC2 and CEACAM5 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. AisALPP; ii. BisGDF15; iii. C is MMP12; iv. DisCDCPI; v. E is LAMP3; vi. FisPLAUR; vii G is SFTPA1; and vii HisTNSF13B.
[0054] One or more (additional) proteins may be selected from those shown in Tables 1-4. The proteins may comprise the proteins shown in Table 1. The proteins may consist of the proteins shown in Table 1. The proteins may comprise the proteins shown in Table 2. The proteins may consist of the proteins shown in Table 2. The proteins may comprise the proteins shown in Table 3. The proteins may consist of the proteins shown in Table 3. The proteins may comprise the proteins shown in Table 4. The proteins may consist of the proteins shown in Table 4.
[0055] In preferred embodiments, the proteins may comprise or consist of the proteins shown in Table 4. The proteins may be used in different combinations. For example, the proteins may be selected from the group consisting of: (A), (B), (C), (D), (E), (F), (G), (H), (I), (A, B), (A, C), (A, D), (A, E), (A, F), (A, G), (A, H), (A, I), (B, C), (B, D), (B, E), (B, F), (B, G), (B, H), (B, I), (C, D), (C, E), (C, F), (C, G), (C, H), (C, I), (D, E), (D, F), (D, G), (D, H), (D, I), (E, F), (E, G), (E, H), (E, I), (F, G), (F, H), (F, I), (G, H), (G, I), (H, I), (A, B, C), (A, B, D), (A, B, E), (A, B, F), (A, B, G), (A, B, H), (A, B, I), (A, C, D), (A, C, E), (A, C, F), (A, C, G), (A, C, H), (A, C, I), (A, D, E), (A, D, F), (A, D, G), (A, D, H), (A, D, I), (A, E, F), (A, E, G), (A, E, H), (A, E, I), (A, F, G), (A, F, H), (A, F, I), (A, G, H), (A, G, I),
[0056] (A, H, I), (B, C, D), (B, C, E), (B, C, F), (B, C, G), (B, C, H), (B, C, I), (B, D, E), (B, D, F), (B, D, G),
[0057] (B, D, H), (B, D, I), (B, E, F), (B, E, G), (B, E, H), (B, E, I), (B, F, G), (B, F, H), (B, F, I), (B, G, H), (B, G, I), (B, H, I), (C, D, E), (C, D, F), (C, D, G), (C, D, H), (C, D, I), (C, E, F), (C, E, G), (C, E, H),
[0058] (C, E, I), (C, F, G), (C, F, H), (C, F, I), (C, G, H), (C, G, I), (C, H, I), (D, E, F), (D, E, G), (D, E, H),
[0059] (D, E, I), (D, F, G), (D, F, H), (D, F, I), (D, G, H), (D, G, I), (D, H, I), (E, F, G), (E, F, H), (E, F, I),
[0060] (E, G, H), (E, G, I), (E, H, I), (F, G, H), (F, G, I), (F, H, I), (G, H, I), (A, B, C, D), (A, B, C, E), (A, B,
[0061] C, F), (A, B, C, G), (A, B, C, H), (A, B, C, I), (A, B, D, E), (A, B, D, F), (A, B, D, G), (A, B, D, H),
[0062] (A, B, D, I), (A, B, E, F), (A, B, E, G), (A, B, E, H), (A, B, E, I), (A, B, F, G), (A, B, F, H), (A, B, F,
[0063] I), (A, B, G, H), (A, B, G, I), (A, B, H, I), (A, C, D, E), (A, C, D, F), (A, C, D, G), (A, C, D, H), (A, C,
[0064] D, I), (A, C, E, F), (A, C, E, G), (A, C, E, H), (A, C, E, I), (A, C, F, G), (A, C, F, H), (A, C, F, I), (A,
[0065] C, G, H), (A, C, G, I), (A, C, H, I), (A, D, E, F), (A, D, E, G), (A, D, E, H), (A, D, E, I), (A, D, F, G),
[0066] (A, D, F, H), (A, D, F, I), (A, D, G, H), (A, D, G, I), (A, D, H, I), (A, E, F, G), (A, E, F, H), (A, E, F,
[0067] I), (A, E, G, H), (A, E, G, I), (A, E, H, I), (A, F, G, H), (A, F, G, I), (A, F, H, I), (A, G, H, I), (B, C, D, E), (B, C, D, F), (B, C, D, G), (B, C, D, H), (B, C, D, I), (B, C, E, F), (B, C, E, G), (B, C, E, H), (B, C, E, I), (B, C, F, G), (B, C, F, H), (B, C, F, I), (B, C, G, H), (B, C, G, I), (B, C, H, I), (B, D, E, F), (B, D, E, G), (B, D, E, H), (B, D, E, I), (B, D, F, G), (B, D, F, H), (B, D, F, I), (B, D, G, H), (B, D, G, I), (B, D, H, I), (B, E, F, G), (B, E, F, H), (B, E, F, I), (B, E, G, H), (B, E, G, I), (B, E, H, I), (B, F, G,
[0068] H), (B, F, G, I), (B, F, H, I), (B, G, H, I), (C, D, E, F), (C, D, E, G), (C, D, E, H), (C, D, E, I), (C, D, F, G), (C, D, F, H), (C, D, F, I), (C, D, G, H), (C, D, G, I), (C, D, H, I), (C, E, F, G), (C, E, F, H), (C,
[0069] E, F, I), (C, E, G, H), (C, E, G, I), (C, E, H, I), (C, F, G, H), (C, F, G, I), (C, F, H, I), (C, G, H, I), (D,
[0070] E, F, G), (D, E, F, H), (D, E, F, I), (D, E, G, H), (D, E, G, I), (D, E, H, I), (D, F, G, H), (D, F, G, I),
[0071] (D, F, H, I), (D, G, H, I), (E, F, G, H), (E, F, G, I), (E, F, H, I), (E, G, H, I), (F, G, H, I), (A, B, C, D, E), (A, B, C, D, F), (A, B, C, D, G), (A, B, C, D, H), (A, B, C, D, I), (A, B, C, E, F), (A, B, C, E, G),
[0072] (A, B, C, E, H), (A, B, C, E, I), (A, B, C, F, G), (A, B, C, F, H), (A, B, C, F, I), (A, B, C, G, H), (A,
[0073] B, C, G, I), (A, B, C, H, I), (A, B, D, E, F), (A, B, D, E, G), (A, B, D, E, H), (A, B, D, E, I), (A, B, D,
[0074] F, G), (A, B, D, F, H), (A, B, D, F, I), (A, B, D, G, H), (A, B, D, G, I), (A, B, D, H, I), (A, B, E, F, G),
[0075] (A, B, E, F, H), (A, B, E, F, I), (A, B, E, G, H), (A, B, E, G, I), (A, B, E, H, I), (A, B, F, G, H), (A, B,
[0076] F, G, I), (A, B, F, H, I), (A, B, G, H, I), (A, C, D, E, F), (A, C, D, E, G), (A, C, D, E, H), (A, C, D, E,
[0077] I), (A, C, D, F, G), (A, C, D, F, H), (A, C, D, F, I), (A, C, D, G, H), (A, C, D, G, I), (A, C, D, H, I), (A,
[0078] C, E, F, G), (A, C, E, F, H), (A, C, E, F, I), (A, C, E, G, H), (A, C, E, G, I), (A, C, E, H, I), (A, C, F, G, H), (A, C, F, G, I), (A, C, F, H, I), (A, C, G, H, I), (A, D, E, F, G), (A, D, E, F, H), (A, D, E, F, I),
[0079] (A, D, E, G, H), (A, D, E, G, I), (A, D, E, H, I), (A, D, F, G, H), (A, D, F, G, I), (A, D, F, H, I), (A, D,
[0080] G, H, I), (A, E, F, G, H), (A, E, F, G, I), (A, E, F, H, I), (A, E, G, H, I), (A, F, G, H, I), (B, C, D, E,
[0081] F), (B, C, D, E, G), (B, C, D, E, H), (B, C, D, E, I), (B, C, D, F, G), (B, C, D, F, H), (B, C, D, F, I),
[0082] (B, C, D, G, H), (B, C, D, G, I), (B, C, D, H, I), (B, C, E, F, G), (B, C, E, F, H), (B, C, E, F, I), (B, C,
[0083] E, G, H), (B, C, E, G, I), (B, C, E, H, I), (B, C, F, G, H), (B, C, F, G, I), (B, C, F, H, I), (B, C, G, H,
[0084] I), (B, D, E, F, G), (B, D, E, F, H), (B, D, E, F, I), (B, D, E, G, H), (B, D, E, G, I), (B, D, E, H, I), (B,
[0085] D, F, G, H), (B, D, F, G, I), (B, D, F, H, I), (B, D, G, H, I), (B, E, F, G, H), (B, E, F, G, I), (B, E, F,
[0086] H, I), (B, E, G, H, I), (B, F, G, H, I), (C, D, E, F, G), (C, D, E, F, H), (C, D, E, F, I), (C, D, E, G, H),
[0087] (C, D, E, G, I), (C, D, E, H, I), (C, D, F, G, H), (C, D, F, G, I), (C, D, F, H, I), (C, D, G, H, I), (C, E,
[0088] F, G, H), (C, E, F, G, I), (C, E, F, H, I), (C, E, G, H, I), (C, F, G, H, I), (D, E, F, G, H), (D, E, F, G,
[0089] I), (D, E, F, H, I), (D, E, G, H, I), (D, F, G, H, I), (E, F, G, H, I), (A, B, C, D, E, F), (A, B, C, D, E,
[0090] G), (A, B, C, D, E, H), (A, B, C, D, E, I), (A, B, C, D, F, G), (A, B, C, D, F, H), (A, B, C, D, F, I), (A, B, C, D, G, H), (A, B, C, D, G, I), (A, B, C, D, H, I), (A, B, C, E, F, G), (A, B, C, E, F, H), (A, B, C,
[0091] E, F, I), (A, B, C, E, G, H), (A, B, C, E, G, I), (A, B, C, E, H, I), (A, B, C, F, G, H), (A, B, C, F, G, I), (A, B, C, F, H, I), (A, B, C, G, H, I), (A, B, D, E, F, G), (A, B, D, E, F, H), (A, B, D, E, F, I), (A, B,
[0092] D, E, G, H), (A, B, D, E, G, I), (A, B, D, E, H, I), (A, B, D, F, G, H), (A, B, D, F, G, I), (A, B, D, F, H, I), (A, B, D, G, H, I), (A, B, E, F, G, H), (A, B, E, F, G, I), (A, B, E, F, H, I), (A, B, E, G, H, I), (A,
[0093] B, F, G, H, I), (A, C, D, E, F, G), (A, C, D, E, F, H), (A, C, D, E, F, I), (A, C, D, E, G, H), (A, C, D,
[0094] E, G, I), (A, C, D, E, H, I), (A, C, D, F, G, H), (A, C, D, F, G, I), (A, C, D, F, H, I), (A, C, D, G, H, I), (A, C, E, F, G, H), (A, C, E, F, G, I), (A, C, E, F, H, I), (A, C, E, G, H, I), (A, C, F, G, H, I), (A, D,
[0095] E, F, G, H), (A, D, E, F, G, I), (A, D, E, F, H, I), (A, D, E, G, H, I), (A, D, F, G, H, I), (A, E, F, G, H, I), (B, C, D, E, F, G), (B, C, D, E, F, H), (B, C, D, E, F, I), (B, C, D, E, G, H), (B, C, D, E, G, I), (B,
[0096] C, D, E, H, I), (B, C, D, F, G, H), (B, C, D, F, G, I), (B, C, D, F, H, I), (B, C, D, G, H, I), (B, C, E, F,
[0097] G, H), (B, C, E, F, G, I), (B, C, E, F, H, I), (B, C, E, G, H, I), (B, C, F, G, H, I), (B, D, E, F, G, H),
[0098] (B, D, E, F, G, I), (B, D, E, F, H, I), (B, D, E, G, H, I), (B, D, F, G, H, I), (B, E, F, G, H, I), (C, D, E,
[0099] F, G, H), (C, D, E, F, G, I), (C, D, E, F, H, I), (C, D, E, G, H, I), (C, D, F, G, H, I), (C, E, F, G, H, I),
[0100] (D, E, F, G, H, I), (A, B, C, D, E, F, G), (A, B, C, D, E, F, H), (A, B, C, D, E, F, I), (A, B, C, D, E,
[0101] G, H), (A, B, C, D, E, G, I), (A, B, C, D, E, H, I), (A, B, C, D, F, G, H), (A, B, C, D, F, G, I), (A, B,
[0102] C, D, F, H, I), (A, B, C, D, G, H, I), (A, B, C, E, F, G, H), (A, B, C, E, F, G, I), (A, B, C, E, F, H, I), (A, B, C, E, G, H, I), (A, B, C, F, G, H, I), (A, B, D, E, F, G, H), (A, B, D, E, F, G, I), (A, B, D, E, F,
[0103] H, I), (A, B, D, E, G, H, I), (A, B, D, F, G, H, I), (A, B, E, F, G, H, I), (A, C, D, E, F, G, H), (A, C, D,
[0104] E, F, G, I), (A, C, D, E, F, H, I), (A, C, D, E, G, H, I), (A, C, D, F, G, H, I), (A, C, E, F, G, H, I), (A,
[0105] D, E, F, G, H, I), (B, C, D, E, F, G, H), (B, C, D, E, F, G, I), (B, C, D, E, F, H, I), (B, C, D, E, G, H,
[0106] I), (B, C, D, F, G, H, I), (B, C, E, F, G, H, I), (B, D, E, F, G, H, I), (C, D, E, F, G, H, I), (A, B, C, D,
[0107] E, F, G, H), (A, B, C, D, E, F, G, I), (A, B, C, D, E, F, H, I), (A, B, C, D, E, G, H, I), (A, B, C, D, F,
[0108] G, H, I), (A, B, C, E, F, G, H, I), (A, B, D, E, F, G, H, I), (A, C, D, E, F, G, H, I), (B, C, D, E, F, G,
[0109] H, I) and (A, B, C, D, E, F, G, H, I); wherein A is ALPP; B is GDF15; C is MMP12; D is TNFSF13B; E is LAMP3; F is SFTPA1 ; G is CXCL17; H is WFDC2 and I is CEACAM5. In the preceding list, the proteins are listed between parentheses. For example, “(A, B, C, D, H, I)” indicates that the combination of proteins is ALPP, GDF15, MMP12, TNFSF13B, WFDC2 and CEACAM5. The one or more additional proteins may further comprise PLALIR in any combination above. The one or more additional proteins may further comprise CDCP1 in any combination above. The one or more additional proteins may further comprise PLALIR and CDCP1 in any combination above.
[0110] The proteins may also be selected from the group consisting of: (A), (B), (C), (D), (E), (F), (G), (H), (I), (A, B), (A, C), (A, D), (A, E), (A, F), (A, G), (A, H), (A, I), (B, C), (B, D), (B, E), (B, F), (B,
[0111] G), (B, H), (B, I), (C, D), (C, E), (C, F), (C, G), (C, H), (C, I), (D, E), (D, F), (D, G), (D, H), (D, I),
[0112] (E, F), (E, G), (E, H), (E, I), (F, G), (F, H), (F, I), (G, H), (G, I), (H, I), (A, B, C), (A, B, D), (A, B, E), (A, B, F), (A, B, G), (A, B, H), (A, B, I), (A, C, D), (A, C, E), (A, C, F), (A, C, G), (A, C, H), (A, C, I),
[0113] (A, D, E), (A, D, F), (A, D, G), (A, D, H), (A, D, I), (A, E, F), (A, E, G), (A, E, H), (A, E, I), (A, F, G),
[0114] (A, F, H), (A, F, I), (A, G, H), (A, G, I), (A, H, I), (B, C, D), (B, C, E), (B, C, F), (B, C, G), (B, C, H), (B, C, I), (B, D, E), (B, D, F), (B, D, G), (B, D, H), (B, D, I), (B, E, F), (B, E, G), (B, E, H), (B, E, I), (B, F, G), (B, F, H), (B, F, I), (B, G, H), (B, G, I), (B, H, I), (C, D, E), (C, D, F), (C, D, G), (C, D, H), (C, D, I), (C, E, F), (C, E, G), (C, E, H), (C, E, I), (C, F, G), (C, F, H), (C, F, I), (C, G, H), (C, G, I),
[0115] (C, H, I), (D, E, F), (D, E, G), (D, E, H), (D, E, I), (D, F, G), (D, F, H), (D, F, I), (D, G, H), (D, G, I),
[0116] (D, H, I), (E, F, G), (E, F, H), (E, F, I), (E, G, H), (E, G, I), (E, H, I), (F, G, H), (F, G, I), (F, H, I), (G, H, I), (A, B, C, D), (A, B, C, E), (A, B, C, F), (A, B, C, G), (A, B, C, H), (A, B, C, I), (A, B, D, E), (A,
[0117] B, D, F), (A, B, D, G), (A, B, D, H), (A, B, D, I), (A, B, E, F), (A, B, E, G), (A, B, E, H), (A, B, E, I),
[0118] (A, B, F, G), (A, B, F, H), (A, B, F, I), (A, B, G, H), (A, B, G, I), (A, B, H, I), (A, C, D, E), (A, C, D,
[0119] F), (A, C, D, G), (A, C, D, H), (A, C, D, I), (A, C, E, F), (A, C, E, G), (A, C, E, H), (A, C, E, I), (A,
[0120] C, F, G), (A, C, F, H), (A, C, F, I), (A, C, G, H), (A, C, G, I), (A, C, H, I), (A, D, E, F), (A, D, E, G),
[0121] (A, D, E, H), (A, D, E, I), (A, D, F, G), (A, D, F, H), (A, D, F, I), (A, D, G, H), (A, D, G, I), (A, D, H,
[0122] I), (A, E, F, G), (A, E, F, H), (A, E, F, I), (A, E, G, H), (A, E, G, I), (A, E, H, I), (A, F, G, H), (A, F,
[0123] G, I), (A, F, H, I), (A, G, H, I), (B, C, D, E), (B, C, D, F), (B, C, D, G), (B, C, D, H), (B, C, D, I), (B,
[0124] C, E, F), (B, C, E, G), (B, C, E, H), (B, C, E, I), (B, C, F, G), (B, C, F, H), (B, C, F, I), (B, C, G, H),
[0125] (B, C, G, I), (B, C, H, I), (B, D, E, F), (B, D, E, G), (B, D, E, H), (B, D, E, I), (B, D, F, G), (B, D, F,
[0126] H), (B, D, F, I), (B, D, G, H), (B, D, G, I), (B, D, H, I), (B, E, F, G), (B, E, F, H), (B, E, F, I), (B, E,
[0127] G, H), (B, E, G, I), (B, E, H, I), (B, F, G, H), (B, F, G, I), (B, F, H, I), (B, G, H, I), (C, D, E, F), (C,
[0128] D, E, G), (C, D, E, H), (C, D, E, I), (C, D, F, G), (C, D, F, H), (C, D, F, I), (C, D, G, H), (C, D, G, I),
[0129] (C, D, H, I), (C, E, F, G), (C, E, F, H), (C, E, F, I), (C, E, G, H), (C, E, G, I), (C, E, H, I), (C, F, G,
[0130] H), (C, F, G, I), (C, F, H, I), (C, G, H, I), (D, E, F, G), (D, E, F, H), (D, E, F, I), (D, E, G, H), (D, E,
[0131] G, I), (D, E, H, I), (D, F, G, H), (D, F, G, I), (D, F, H, I), (D, G, H, I), (E, F, G, H), (E, F, G, I), (E, F,
[0132] H, I), (E, G, H, I), (F, G, H, I), (A, B, C, D, E), (A, B, C, D, F), (A, B, C, D, G), (A, B, C, D, H), (A,
[0133] B, C, D, I), (A, B, C, E, F), (A, B, C, E, G), (A, B, C, E, H), (A, B, C, E, I), (A, B, C, F, G), (A, B, C,
[0134] F, H), (A, B, C, F, I), (A, B, C, G, H), (A, B, C, G, I), (A, B, C, H, I), (A, B, D, E, F), (A, B, D, E, G), (A, B, D, E, H), (A, B, D, E, I), (A, B, D, F, G), (A, B, D, F, H), (A, B, D, F, I), (A, B, D, G, H), (A,
[0135] B, D, G, I), (A, B, D, H, I), (A, B, E, F, G), (A, B, E, F, H), (A, B, E, F, I), (A, B, E, G, H), (A, B, E, G, I), (A, B, E, H, I), (A, B, F, G, H), (A, B, F, G, I), (A, B, F, H, I), (A, B, G, H, I), (A, C, D, E, F), (A, C, D, E, G), (A, C, D, E, H), (A, C, D, E, I), (A, C, D, F, G), (A, C, D, F, H), (A, C, D, F, I), (A,
[0136] C, D, G, H), (A, C, D, G, I), (A, C, D, H, I), (A, C, E, F, G), (A, C, E, F, H), (A, C, E, F, I), (A, C, E,
[0137] G, H), (A, C, E, G, I), (A, C, E, H, I), (A, C, F, G, H), (A, C, F, G, I), (A, C, F, H, I), (A, C, G, H, I),
[0138] (A, D, E, F, G), (A, D, E, F, H), (A, D, E, F, I), (A, D, E, G, H), (A, D, E, G, I), (A, D, E, H, I), (A, D,
[0139] F, G, H), (A, D, F, G, I), (A, D, F, H, I), (A, D, G, H, I), (A, E, F, G, H), (A, E, F, G, I), (A, E, F, H,
[0140] I), (A, E, G, H, I), (A, F, G, H, I), (B, C, D, E, F), (B, C, D, E, G), (B, C, D, E, H), (B, C, D, E, I), (B,
[0141] C, D, F, G), (B, C, D, F, H), (B, C, D, F, I), (B, C, D, G, H), (B, C, D, G, I), (B, C, D, H, I), (B, C, E,
[0142] F, G), (B, C, E, F, H), (B, C, E, F, I), (B, C, E, G, H), (B, C, E, G, I), (B, C, E, H, I), (B, C, F, G, H),
[0143] (B, C, F, G, I), (B, C, F, H, I), (B, C, G, H, I), (B, D, E, F, G), (B, D, E, F, H), (B, D, E, F, I), (B, D,
[0144] E, G, H), (B, D, E, G, I), (B, D, E, H, I), (B, D, F, G, H), (B, D, F, G, I), (B, D, F, H, I), (B, D, G, H,
[0145] I), (B, E, F, G, H), (B, E, F, G, I), (B, E, F, H, I), (B, E, G, H, I), (B, F, G, H, I), (C, D, E, F, G), (C,
[0146] D, E, F, H), (C, D, E, F, I), (C, D, E, G, H), (C, D, E, G, I), (C, D, E, H, I), (C, D, F, G, H), (C, D, F,
[0147] G, I), (C, D, F, H, I), (C, D, G, H, I), (C, E, F, G, H), (C, E, F, G, I), (C, E, F, H, I), (C, E, G, H, I),
[0148] (C, F, G, H, I), (D, E, F, G, H), (D, E, F, G, I), (D, E, F, H, I), (D, E, G, H, I), (D, F, G, H, I), (E, F,
[0149] G, H, I), (A, B, C, D, E, F), (A, B, C, D, E, G), (A, B, C, D, E, H), (A, B, C, D, E, I), (A, B, C, D, F,
[0150] G), (A, B, C, D, F, H), (A, B, C, D, F, I), (A, B, C, D, G, H), (A, B, C, D, G, I), (A, B, C, D, H, I), (A,
[0151] B, C, E, F, G), (A, B, C, E, F, H), (A, B, C, E, F, I), (A, B, C, E, G, H), (A, B, C, E, G, I), (A, B, C,
[0152] E, H, I), (A, B, C, F, G, H), (A, B, C, F, G, I), (A, B, C, F, H, I), (A, B, C, G, H, I), (A, B, D, E, F, G),
[0153] (A, B, D, E, F, H), (A, B, D, E, F, I), (A, B, D, E, G, H), (A, B, D, E, G, I), (A, B, D, E, H, I), (A, B,
[0154] D, F, G, H), (A, B, D, F, G, I), (A, B, D, F, H, I), (A, B, D, G, H, I), (A, B, E, F, G, H), (A, B, E, F, G,
[0155] I), (A, B, E, F, H, I), (A, B, E, G, H, I), (A, B, F, G, H, I), (A, C, D, E, F, G), (A, C, D, E, F, H), (A,
[0156] C, D, E, F, I), (A, C, D, E, G, H), (A, C, D, E, G, I), (A, C, D, E, H, I), (A, C, D, F, G, H), (A, C, D,
[0157] F, G, I), (A, C, D, F, H, I), (A, C, D, G, H, I), (A, C, E, F, G, H), (A, C, E, F, G, I), (A, C, E, F, H, I),
[0158] (A, C, E, G, H, I), (A, C, F, G, H, I), (A, D, E, F, G, H), (A, D, E, F, G, I), (A, D, E, F, H, I), (A, D,
[0159] E, G, H, I), (A, D, F, G, H, I), (A, E, F, G, H, I), (B, C, D, E, F, G), (B, C, D, E, F, H), (B, C, D, E,
[0160] F, I), (B, C, D, E, G, H), (B, C, D, E, G, I), (B, C, D, E, H, I), (B, C, D, F, G, H), (B, C, D, F, G, I),
[0161] (B, C, D, F, H, I), (B, C, D, G, H, I), (B, C, E, F, G, H), (B, C, E, F, G, I), (B, C, E, F, H, I), (B, C,
[0162] E, G, H, I), (B, C, F, G, H, I), (B, D, E, F, G, H), (B, D, E, F, G, I), (B, D, E, F, H, I), (B, D, E, G, H,
[0163] I), (B, D, F, G, H, I), (B, E, F, G, H, I), (C, D, E, F, G, H), (C, D, E, F, G, I), (C, D, E, F, H, I), (C,
[0164] D, E, G, H, I), (C, D, F, G, H, I), (C, E, F, G, H, I), (D, E, F, G, H, I), (A, B, C, D, E, F, G), (A, B, C,
[0165] D, E, F, H), (A, B, C, D, E, F, I), (A, B, C, D, E, G, H), (A, B, C, D, E, G, I), (A, B, C, D, E, H, I), (A,
[0166] B, C, D, F, G, H), (A, B, C, D, F, G, I), (A, B, C, D, F, H, I), (A, B, C, D, G, H, I), (A, B, C, E, F, G,
[0167] H), (A, B, C, E, F, G, I), (A, B, C, E, F, H, I), (A, B, C, E, G, H, I), (A, B, C, F, G, H, I), (A, B, D, E,
[0168] F, G, H), (A, B, D, E, F, G, I), (A, B, D, E, F, H, I), (A, B, D, E, G, H, I), (A, B, D, F, G, H, I), (A, B,
[0169] E, F, G, H, I), (A, C, D, E, F, G, H), (A, C, D, E, F, G, I), (A, C, D, E, F, H, I), (A, C, D, E, G, H, I), (A, C, D, F, G, H, I), (A, C, E, F, G, H, I), (A, D, E, F, G, H, I), (B, C, D, E, F, G, H), (B, C, D, E, F,
[0170] G, I), (B, C, D, E, F, H, I), (B, C, D, E, G, H, I), (B, C, D, F, G, H, I), (B, C, E, F, G, H, I), (B, D, E,
[0171] F, G, H, I), (C, D, E, F, G, H, I), (A, B, C, D, E, F, G, H), (A, B, C, D, E, F, G, I), (A, B, C, D, E, F,
[0172] H, I), (A, B, C, D, E, G, H, I), (A, B, C, D, F, G, H, I), (A, B, C, E, F, G, H, I), (A, B, D, E, F, G, H,
[0173] I), (A, C, D, E, F, G, H, I), (B, C, D, E, F, G, H, I) and (A, B, C, D, E, F, G, H, I); wherein A is ALPP;
[0174] B is GDF15; C is MMP12; D is CDCP1 ; E is LAMP3; F is PLAUR; G is CXCL17; H is WFDC2 and I is CEACAM5. In the preceding list, the proteins are listed between parentheses. For example, “(A, B, C, D, H, I)” indicates that the combination of proteins is ALPP, GDF15, MMP12, CDCP1 , WFDC2 and CEACAM5. The one or more additional proteins may further comprise SFTPA1 in any combination above. The one or more additional proteins may further comprise TNSF13B in any combination above. The one or more additional proteins may further comprise SFTPA1 and TNSF13B in any combination above.
[0175] The one or more additional proteins (used in combination with CXCL17, WFDC2 and CEACAM5) may be used in different combinations. For example, the one or more additional proteins may be selected from the group consisting of: (A), (B), (C), (D), (E), (F), (A, B), (A, C), (A, D), (A, E), (A, F), (B, C), (B, D), (B, E), (B, F), (C, D), (C, E), (C, F), (D, E), (D, F), (E, F), (A, B, C), (A, B, D), (A, B, E), (A, B, F), (A, C, D), (A, C, E), (A, C, F), (A, D, E), (A, D, F), (A, E, F), (B, C, D), (B, C, E), (B, C, F), (B, D, E), (B, D, F), (B, E, F), (C, D, E), (C, D, F), (C, E, F), (D, E, F), (A, B, C, D), (A, B, C, E), (A, B, C, F), (A, B, D, E), (A, B, D, F), (A, B, E, F), (A, C, D, E), (A, C, D, F), (A, C, E, F), (A, D, E, F), (B, C, D, E), (B, C, D, F), (B, C, E, F), (B, D, E, F), (C, D, E, F), (A, B, C, D, E), (A, B, C, D, F), (A, B, C, E, F), (A, B, D, E, F), (A, C, D, E, F), (B, C, D, E, F), (A, B, C, D, E, F); wherein A is ALPP; B is GDF15; C is MMP12; D is TNFSF13B; E is LAMP3; F is SFTPA1 ; G is CXCL17; H is WFDC2 and I is CEACAM5. In the preceding list, the one or more additional proteins used in combination with CXCL17, WFDC2 and CEACAM5 are the proteins listed between parentheses. For example, “(A, B, C, D)” indicates that the combination of proteins is CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12 and TNFSF13B. The one or more additional proteins may further comprise PLAUR in any combination above. The one or more additional proteins may further comprise CDCP1 in any combination above. The one or more additional proteins may further comprise PLAUR and CDCP1 in any combination above.
[0176] The one or more additional proteins (used in combination with CXCL17, WFDC2 and CEACAM5) may be used in different combinations. For example, the one or more additional proteins may be selected from the group consisting of: (A), (B), (C), (D), (E), (F), (A, B), (A, C), (A, D), (A, E), (A, F), (B, C), (B, D), (B, E), (B, F), (C, D), (C, E), (C, F), (D, E), (D, F), (E, F), (A, B, C), (A, B, D), (A, B, E), (A, B, F), (A, C, D), (A, C, E), (A, C, F), (A, D, E), (A, D, F), (A, E, F), (B, C, D), (B, C, E), (B, C, F), (B, D, E), (B, D, F), (B, E, F), (C, D, E), (C, D, F), (C, E, F), (D, E, F), (A, B, C, D), (A, B, C, E), (A, B, C, F), (A, B, D, E), (A, B, D, F), (A, B, E, F), (A, C, D, E), (A, C, D, F), (A, C, E, F), (A, D, E, F), (B, C, D, E), (B, C, D, F), (B, C, E, F), (B, D, E, F), (C, D, E, F), (A, B, C, D, E), (A, B, C, D, F), (A, B, C, E, F), (A, B, D, E, F), (A, C, D, E, F), (B, C, D, E, F), (A, B, C, D, E, F); wherein A is ALPP; B is GDF15; C is MMP12; D is CDCP1 ; E is LAMP3; F is PLAUR; G is CXCL17; H is WFDC2 and I is CEACAM5. In the preceding list, the one or more additional proteins used in combination with CXCL17, WFDC2 and CEACAM5 are the proteins listed between parentheses. For example, “(A, B, C, D)” indicates that the combination of proteins is CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12 and CDCP1. The one or more additional proteins may further comprise SFTPA1 in any combination above. The one or more additional proteins may further comprise TNSF13B in any combination above. The one or more additional proteins may further comprise SFTPA1 and TNSF13B in any combination above.
[0177] The one or more additional proteins (used in combination with CXCL17, WFDC2 and CEACAM5, optionally in the combinations set out above) may be used in different combinations. For example, the one or more additional proteins may be selected from the group consisting of: (G), (H), (I), (J), (K), (L), (G, H), (G, I), (G, J), (G, K), (G, L), (H, I), (H, J), (H, K), (H, L), (I, J), (I, K), (I, L), (J, K), (J, L), (K, L), (G, H, I), (G, H, J), (G, H, K), (G, H, L), (G, I, J), (G, I, K), (G, I, L), (G, J, K), (G, J, L), (G, K, L), (H, I, J), (H, I, K), (H, I, L), (H, J, K), (H, J, L), (H, K, L), (I, J, K), (I, J, L), (I, K, L), (J, K, L), (G, H, I, J), (G, H, I, K), (G, H, I, L), (G, H, J, K), (G, H, J, L), (G, H, K, L), (G, I, J, K), (G, I, J, L), (G, I, K, L), (G, J, K, L), (H, I, J, K), (H, I, J, L), (H, I, K, L), (H, J, K, L), (I, J, K, L), (G, H, I,
[0178] J, K), (G, H, I, J, L), (G, H, I, K, L), (G, H, J, K, L), (G, I, J, K, L), (H, I, J, K, L) and (G, H, I, J, K, L); wherein G is CDCP1 ; H is PIGR; I is PRSS8; J is SFTPD; K is AGER; and L is PLAUR. In the preceding list, the one or more additional proteins used in combination with CXCL17, WFDC2 and CEACAM5 are the proteins listed between parentheses. For example, “(G, H, I, J)” indicates that the combination of proteins is CXCL17, WFDC2, CEACAM5, CDCP1 , PIGR, PRSS8 and SFTPD. The one or more additional proteins may further comprise SFTPA1 in any combination above. The one or more additional proteins may further comprise TNSF13B in any combination above. The one or more additional proteins may further comprise SFTPA1 and TNSF13B in any combination above.
[0179] The one or more additional proteins (used in combination with CXCL17, WFDC2 and CEACAM5, optionally in the combinations set out above) may be used in different combinations. For example, the one or more additional proteins may be selected from the group consisting of: (G), (H), (I), (J), (K), (L), (G, H), (G, I), (G, J), (G, K), (G, L), (H, I), (H, J), (H, K), (H, L), (I, J), (I, K), (I, L), (J, K), (J, L), (K, L), (G, H, I), (G, H, J), (G, H, K), (G, H, L), (G, I, J), (G, I, K), (G, I, L), (G, J, K), (G, J, L), (G, K, L), (H, I, J), (H, I, K), (H, I, L), (H, J, K), (H, J, L), (H, K, L), (I, J, K), (I, J, L), (I, K, L), (J,
[0180] K, L), (G, H, I, J), (G, H, I, K), (G, H, I, L), (G, H, J, K), (G, H, J, L), (G, H, K, L), (G, I, J, K), (G, I, J, L), (G, I, K, L), (G, J, K, L), (H, I, J, K), (H, I, J, L), (H, I, K, L), (H, J, K, L), (I, J, K, L), (G, H, I, J, K), (G, H, I, J, L), (G, H, I, K, L), (G, H, J, K, L), (G, I, J, K, L), (H, I, J, K, L) and (G, H, I, J, K, L); wherein G is TNSF13B; H is PIGR; I is PRSS8; J is SFTPD; K is AGER; and L is SFTPA1. In the preceding list, the one or more additional proteins used in combination with CXCL17, WFDC2 and CEACAM5 are the proteins listed between parentheses. For example, “(G, H, I, J)” indicates that the combination of proteins is CXCL17, WFDC2, CEACAM5, TNSF13B, PIGR, PRSS8 and SFTPD.
[0181] The proteins may comprise of CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, TNFSF13B, LAMP3, SFTPA1 , CDCP1 , PIGR, PRSS8, SFTPD, AGER and PLAUR. The proteins may consist of CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, TNFSF13B, LAMP3, SFTPA1 , CDCP1 , PIGR, PRSS8, SFTPD, AGER and PLAUR.
[0182] The proteins may comprise of CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, LAMP3, CDCP1 , PIGR, PRSS8, SFTPD, AGER and PLAUR. The proteins may consist of CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, LAMP3, CDCP1 , PIGR, PRSS8, SFTPD, AGER and PLAUR.
[0183] The proteins may comprise of CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, LAMP3, CDCP1 , PIGR, PRSS8, SFTPD, AGER, SFTPA1 , TNSF13B and PLAUR. The proteins may consist of CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, LAMP3, CDCP1 , PIGR, PRSS8, SFTPD, AGER, SFTPA1 , TNSF13B and PLAUR.
[0184] The methods described herein may include applying a predictive model to the level of expression of the proteins described herein. The performance of the predictive model may be characterised by an area under the curve (AUG) of at least 0.60, at least 0.61 , at least 0.62, at least 0.63, at least 0.64, at least 0.65, at least 0.66, at least 0.67, at least 0.68, at least 0.69, at least 0.70, at least 0.71 , at least 0.72, at least 0.73, at least 0.74, at least 0.75, at least 0.76, at least 0.77, at least 0.78, at least 0.79, at least 0.80, at least 0.81 , at least 0.82, at least 0.83, at least 0.84, at least 0.85, at least 0.86, at least 0.87, at least 0.88, at least 0.89, or at least 0.90. Preferably the performance of the predictive model may be characterised by an area under the curve (AUG) of at least 0.88.
[0185] The lung cancer may be selected from the group consisting of: an adenocarcinoma, an adenosquamous cell cancer, a large cell cancer, a neuroendocrine cancer, a non-small cell lung cancer (NSCLC), a small cell cancer, or asquamous cell cancer. The lung cancer may be an early stage cancer. The cancer may be stage I, stage II, stage III and / or stage IV lung cancer. In some embodiments, the lung cancer may be stage I or stage II lung cancer. In some embodiments, the lung cancer may be stage I lung cancer. In some embodiments, the lung cancer may be stage II lung cancer. In some embodiments, the lung cancer may be a lung cancer characterised by or associated with a mutation in one or more genes. The genes may be selected from the group consisting of EGFR, KRAS, ALK and BRAF.
[0186] The expression levels of the proteins may be determined from a sample obtained from the subject (e.g., a biopsy). The sample may be whole blood, serum, plasma, or any combination thereof. Generally, the sample has been previously obtained from a subject. Preferably, the sample is a blood sample. More preferably, the sample is a plasma sample.
[0187] The invention may include a step of obtaining or providing the biological sample, or alternatively the sample may have already been obtained from a subject, for example in ex vivo methods.
[0188] Biological samples obtained from a subject can be stored until needed. Suitable storage methods include freezing within two hours of collection. Maintenance at -80°C can be used for long-term storage.
[0189] The sample may be processed prior to determining the level of expression of the biomarkers. The sample may be subject to enrichment (for example to increase the concentration of the biomarkers being quantified), centrifugation or dilution. Expression products of the genes (such as protein or nucleic acids, but in particular RNA) may be extracted from the sample prior to analysis.
[0190] In some embodiments of the invention, the biological sample may be enriched for gene expression products prior to detection and quantification (i.e. measurement). The step of enrichment can be any suitable pre-processing method step to increase the concentration of gene expression products in the sample. For example, the step of enrichment may comprise centrifugation and filtration to remove cells or unwanted analytes from the sample. For RNA, methods of the invention may include a step of amplification to increase the amount of RNA that is detected and quantified. Methods of amplification include PCR amplification. Such methods may be used to enrich the sample for any biomarkers of interest.
[0191] Generally speaking, the gene expression products will need to be extracted from the biological sample. This can be achieved by a number of suitable methods. For example, extraction may involve separating the gene expression products from the biological sample. Methods include chemical extraction (comprising the use of, for example, guanidium thiocyante) and solid-phase extraction (for example on silica columns). Preferred methods include chromatographic methods (for example spin column chromatography), in particular chromatographic methods comprising the use of a silica column. Chromatographic methods comprise lysing cells (if required), addition of a binding solution, centrifugation in a spin column to force the binding solution through a silica gel membrane, optional washing to remove further impurities, and elution of the nucleic acid. Commercial kits are available for such methods, for example Norgen’s urine microRNA purification kit (other kits available, for example from Qiagen or Exigon).
[0192] If gene expression products such as RNA are extracted from a sample, the extracted solution may require enrichment to increase the relative abundance of RNA in the sample.
[0193] In one embodiment of the invention, the method the sample is processed prior to analysis, wherein processing of the sample comprises:
[0194] (i) removal of cells and / or debris from the sample;
[0195] (ii) optional purification of the sample to obtain a purified sample comprising expression products (for example protein or nucleic acid molecules) corresponding to the proteins being measured; and / or
[0196] (iii) extraction or isolation expression products (for example protein or nucleic acid molecules) corresponding to the proteins being measured.
[0197] The methods of the invention may be carried out on one test sample from a subject. Alternatively, a plurality of test samples may be taken from a subject, for example 2, 3, 4 or 5 samples. Each sample may be subjected to a single assay to quantify one of the biomarker panel members, or alternatively a sample may be tested for a plurality of or all of the biomarkers being quantified.
[0198] In one embodiment, there is provided a method comprising: a) measuring proteins of the biomarker panels of the invention described herein in a biological sample obtained from a subject that has previously received therapy for lung cancer; b) comparing the measurement determined in step a) with a previously determined level of expression of the same biomarker or biomarkers; and c) maintaining, changing or withdrawing the therapy for lung cancer.
[0199] Methods of diagnosis
[0200] The present invention also provides methods of diagnosis of lung cancer in a subject. The present invention also provides one of proteins described herein for use in the diagnosis of lung cancer.
[0201] Methods of treatment The present invention also provides methods of treatment of lung cancer in a subject. The present invention also provides a substance or composition described herein for use in the treatment of lung cancer in a subject. A sample from the subject may have undergone a method of diagnosis or prognosis of the invention to determine the subject’s suitability for treatment. In some embodiments, the methods of treatment include the steps of diagnosis or prognosis according to a method of the invention.
[0202] In some embodiments, the methods comprise only recommending the subject for, or assigning a treatment to, the subject. In other embodiments, the methods include the steps of treatment administration.
[0203] In one embodiment of the invention, the method comprises:
[0204] (i) providing or obtaining a sample from a subject;
[0205] (ii) measuring the level of expression of proteins from the biomarker panels of the invention described herein in the subject sample;
[0206] (iii) determining the presence or absence of lung cancer based on the measurement in step (ii); and
[0207] (iv) administering a lung cancer therapy or initiating a therapeutic regimen for lung cancer if lung cancer is diagnosed or suspected.
[0208] In another embodiment of the invention, the method comprises:
[0209] (i) providing or obtaining a sample from a subject;
[0210] (ii) optionally enriching the sample for protein or RNA and / or extracting protein or RNA from the sample;
[0211] (iii) diagnosing or prognosing lung cancer according to a method of diagnosis or prognosis of the invention; and
[0212] (iv) selecting a treatment regimen for the subject according to the presence or absence of lung cancer as determined in step (iii).
[0213] In another embodiment of the invention, there is provided a method of predicting a subject’s responsiveness to a lung cancer treatment, comprising
[0214] (i) providing or obtaining a sample from a subject;
[0215] (ii) optionally enriching the sample for protein or RNA and / or extracting protein or RNA from the sample;
[0216] (iii) diagnosing or prognosing lung cancer according to a method of diagnosis or prognosis of the invention;
[0217] (iv) predicting a subject’s responsiveness to a lung cancer treatment according to the presence or absence of lung cancer as determined in step (iii). The method may comprise a prior step of administering the therapy for lung cancer to the subject. In another embodiment, the method may also comprise a pre-step of measuring proteins of the biomarker panels of the invention described herein in a biological sample obtained from the same subject prior to administration of the therapy. In step c), the therapy for lung cancer may be maintained if an appropriate adjustment in the level(s) of expression of the biomarker or biomarkers is determined. If the levels of expression are unchanged or have worsened, this may be indicative of a worsening of the subject’s condition, and hence an alternative therapy for lung cancer. In this way, drug candidates useful in the treatment of lung cancer can be screened.
[0218] The subject may be human.
[0219] The subject may lack one or more demographic risk factor(s) previously associated with lung cancer, such as age, race, geographic location, genetic history, personal and / or medical history (e.g., smoking status, alcohol, drugs, carcinogenic agents, diet, obesity, BMI, COPD, diabetes, physical activity, sun exposure, radiation exposure, exposure to asbestos, exposure to radon gas, exposure to arsenic, exposure to diesel exhaust fumes, exposure to infectious agents such as viruses, family history of lung cancer and / or occupational hazard). The subject may therefore not have previously been considered to be at risk at developing lung cancer based on the one or more of the demographic risk factors. In particularly, the subject may be a non-smoker. More particularly, the subject may never have smoked. The subject may never have been exposed to second-hand smoke. The subject may be suspected of having an early stage lung cancer. The subject may not be suspected of having an early stage lung cancer. Certain characteristics of non-smokers / never-smokers include: increased incidence of lung adenocarcinoma, highest incidence in East Asian and South Asian women, mutant EGFR-driven, better response to EGFR tyrosine kinase inhibitors and generally lower mutational burden in tumours. Certain characteristics of smokers include mutant KRAS or TP53 driven.
[0220] Alternatively, the subject may be a smoker and / or have been exposed to second-hand smoke.
[0221] The subject may not have previously been diagnosed with lung cancer and / or been treated for lung cancer. The subject may not have an imminent lung cancer diagnosis (e.g., a diagnosis within 3 years). The subject may be greater than 1 , 2, 3, 4, 5, 6, 7, 8, 9 or 10 years from lung cancer diagnosis. The subject may be great than 4 years from lung cancer diagnosis. The subject may be great than 5 years from lung cancer diagnosis.
[0222] The subject may not have lung cancer. The subject may have been diagnosed not to have lung cancer. The subject may not have detectable lung cancer. The subject may not have detectable circulating tumour DNA (ctDNA), preferably ctDNA that is specific for lung cancer. The subject may not have detectable tumours (e.g., by any means known in the art).
[0223] In another embodiment of the invention, there is provided a method of preventing lung cancer, or preventing lung cancer progression, in a subject the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15;
[0224] Hi. C is MMP12; iv. D is TNFSF13B v. E is LAMP3; and vi. F is SFTPAI ; c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. administering treatment to the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0225] In another embodiment of the invention, there is provided a method of preventing lung cancer, or preventing lung cancer progression, in a subject the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. administering treatment to the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0226] In another embodiment of the invention, there is provided a method of preventing lung cancer, or preventing lung cancer progression, in a subject the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. administering treatment to the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0227] The present invention also provides a drug, wherein the drug is a lung cancer therapy, for use in the treatment of lung cancer in a subject, wherein the subject has been predicted as having or being at risk of developing lung cancer by a method described herein. In another embodiment of the invention, there is provided a method identifying a drug useful for the treatment of lung cancer, comprising:
[0228] (a) measuring the level of expression of proteins described herein in a biological sample obtained from a subject;
[0229] (b) administering a candidate drug to the subject;
[0230] (c) measuring in a biological sample obtained from the same subject at a point in time after administration of the candidate drug; and
[0231] (d) comparing the value determined in step (a) with the value determined in step (c), to determine the suitability of the drug candidate as a treatment for lung cancer.
[0232] Treatment may comprise any suitable treatment of lung cancer, including those described herein. The treatment may be surgery, chemotherapy, radiation therapy, radiosurgery, targeted (drug) therapy, immunotherapy, cellular therapies (e.g., CAR T therapy), or any combination thereof.
[0233] Typical chemotherapeutic agents include alkylating agents (for example nitrogen mustards (such as mechlorethamine, cyclophosphamide, melphalan, chlorambucil, ifosfamide and busulfan), nitrosoureas (such as N-Nitroso-N-methylurea (MNU), carmustine (BCNll), lomustine (CCNll) and semustine (MeCCNll), fotemustine and streptozotocin), tetrazines (such as dacarbazine, mitozolomide and temozolomide), aziridines (such as thiotepa, mytomycin and diaziquone), cisplatins and derivatives thereof (such as carboplatin and oxaliplatin), and non-classical alkylating agents (such as procarbazine and hexamethylmelamine)), antimetabolites (for example anti-folates (such as methotrexate and pemetrexed), fluoropyrimidines (such as fluorouracil and capecitabine), deoxynucleoside analogues (such as cytarabine, gemcitabine, decitabine, Vidaza, fludarabine, nelarabine, cladribine, clofarabine and pentostatin) and thiopurines (such as thioguanine and mercaptopurine)), anti-microtubule agents ( for example Vinca alkaloids (such as vincristine, vinblastine, vinorelbine, vindesine, and vinflunine) and taxanes (such as paclitaxel and docetaxel)), platins (such as cisplatin and carboplatin), topoisomerase inhibitors (for example irinotecan, topotecan, camptothecin, etoposide, doxorubicin, mitoxantrone, teniposide, novobiocin, merbarone, and aclarubicin), and cytotoxic antibiotics (for example anthracyclines (such as doxorubicin, daunorubicin apirubicin, idarubicin, pirarubicin, aclarubicin, mitoxantrone), bleomycins, mitomycin C, mitoxantrone, and actinomycin), and combinations thereof.
[0234] Measuring protein expression
[0235] An assay may be performed to determine the expression levels of one or more of the proteins. The assay may be a Proximity Extension Assay (PEA), a xMAP Multiplex Assay, a single molecule array (SIMOA) assay, mass spectrometry-based protein or peptide assay, or an aptamer-based assay. Performing the assay may comprise contacting a sample with a plurality of reagents comprising antibodies. The antibodies may comprise one of monoclonal and polyclonal antibodies. The antibodies may comprise both monoclonal and polyclonal antibodies. The assay may comprise Olink®’s Proximity Extension Assay (PEA). The assay may comprise an absolute quantification assay (e.g., such as Olink® Flex15). The assay may comprise an ELISA (e.g., such as from Abeam16).
[0236] The level of expression of the one or more proteins may represent or be provided as an aggregate score of the expression levels of all proteins being analysed. It may not be necessary to know how the expression level of any individual protein (relative to healthy control(s)) to classify the subject as having a presence or absence of, or is at risk of developing, lung cancer. Rather, it may be the aggregate combination of how the expression level of all proteins has changed relative to healthy control(s) that are determinative of whether the subject has a presence or absence of, or is at risk of developing, lung cancer. One or more parameters (e.g., coefficients) may be assigned to each protein and the parameter may represent the importance of the particular protein associated with the parameter in determining the lung cancer prediction. Thus, the prediction may heavily consider the expression level of certain proteins (e.g., those associated with parameters of higher values) in comparison to other proteins (e.g., those associated with parameters of lower values) when determining the lung cancer prediction.
[0237] The methods may involve comparing the level of expression each of the one or more proteins to one or more reference values (or cut-off or threshold value). As used herein, "reference" refer to a previously determined value, such as a "healthy reference value” corresponding to one or more healthy subjects or a "cancer reference value" corresponding to one or more cancerous subjects. For example, a healthy reference value may correspond to healthy subjects, a subject's own baseline at a prior timepoint when the subject did not exhibit cancer activity (e.g., longitudinal analysis), subjects clinically diagnosed with cancer but not exhibiting cancer activity (e.g., cancer remission), or a healthy reference threshold (e.g., a cutoff). As another example, a "cancer reference value" may correspond to subjects previously diagnosed with cancer, subjects exhibiting cancer activity, or a cancer reference threshold (e.g., a cutoff). The threshold can be derived from a cancer case I non-cancer control ROC curve analysis. The ROC curve can be derived using a logistic regression probability, or any other predictive method that can calculate a score that may be used for classification (e.g., for instance, a neural network). The reference value may therefore be the (aggregated) level of expression each of the one or more proteins prior to diagnosis of the subject with lung cancer (for example, up to 0.25, 0.5, 0.75, 1 , 1.25, 1 .5, 1.75 or 2 years before diagnosis). The reference value may also be the (aggregated) level of expression of each of the one or more proteins in individuals without lung cancer. Accordingly, the reference value for each protein described herein may correspond to the level of expression of each of the same proteins in a sample from a subject who does not have lung cancer.
[0238] In some embodiments where expression of the proteins is positively correlated with lung cancer, either i) the level of expression of each of the proteins determined is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part is higher than the total of the reference values for all proteins. In other embodiments where expression of the proteins is negatively correlated with lung cancer either i) the level of expression of each of the proteins determined is lower than the reference value for each protein; ii) the level of expression of one or more of the proteins determined is lower than the reference value for each protein; or iii) the total level of expression of all proteins determined in part is lower than the total of the reference values for all proteins. In other embodiments where expression of some proteins is positively correlated with lung cancer and where expression of some proteins is negatively correlated with lung cancer, a ratio of the total level of expression of all proteins positively correlated with lung cancer to the total level of expression of all proteins negatively correlated with lung cancer can be determined and compared to an appropriate reference value as described herein.
[0239] The methods may involve determining whether the level of expression of each of the one or more proteins is above or below a threshold cutoff score. If the predicted score is above the threshold cutoff score, the subject is determined to have a presence of, or a risk of developing, lung cancer. If the predicted score is below the threshold cutoff score, the subject is determined to have an absence of, or a risk of developing, lung cancer. If the predicted score is above the threshold cutoff score, the subject may be determined to have an absence, or a risk of developing, of lung cancer. If the predicted score is below the threshold cutoff score, the subject may be determined to have a presence of, or a risk of developing, lung cancer.
[0240] References herein to “higher than” also generally include “higher than or equal to”. For example, if the reference value is 19%, a level of protein expression of 19% or higher would indicate a presence of, or a risk of developing, lung cancer, and a level of protein expression of 18% would indicate an absence of, or a risk of developing, lung cancer.
[0241] The level of expression of each of the one or more proteins may be determined by determining a percentage of cells in the sample that express the protein. The reference value may therefore correspond to a percentage (%) of cells in the sample that express the protein. The reference value may be 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31 %, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41 %, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, or 50%.
[0242] The level of expression of each of the one or more proteins may be determined by determining a percentage of the sample that is positive for the protein. The reference value may therefore correspond to a percentage (%) of the sample that is positive for the protein. The reference value may be 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21 %, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31 %, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, or 50%.
[0243] Kits
[0244] The present invention also relates to a kit for predicting the presence or absence of lung cancer in a subject, wherein the kit comprises means for determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15;
[0245] Hi. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPAI ; and a sample collection apparatus. Other proteins may be used, as discussed above. The kit may comprise instructions for use.
[0246] The present invention also relates to a kit for predicting the presence or absence of lung cancer in a subject, wherein the kit comprises means for determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15;
[0247] Hi. C is MMP12; iv. D is CDCPI ; v. E is LAMP3; and vi. F is PLAUR; and a sample collection apparatus. Other proteins may be used, as discussed above. The kit may comprise instructions for use.
[0248] The present invention also relates to a kit for predicting the presence or absence of lung cancer in a subject, wherein the kit comprises means for determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is ALPP; ii. B is GDF15;
[0249] Hi. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; vi. F is SFTPAI ; vii. G is SFTPA1 ; and vii. H is TNSF13B; and a sample collection apparatus. Other proteins may be used, as discussed above. The kit may comprise instructions for use.
[0250] The present invention provides methods comprising the following steps: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample obtained from the subject; wherein i. A is ALPP; ii. B is GDF15;
[0251] Hi. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPA1 ; and b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein.
[0252] The present invention provides methods comprising the following steps: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample obtained from the subject; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is CDCP1 ; v. E is LAMP3; and vi. F is PLAUR; b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein.
[0253] The present invention provides methods comprising the following steps: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; vi. F is SFTPA1 ; vii. G is SFTPA1 ; and vii. H is TNSF13B; b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein.
[0254] In some embodiments, the methods comprise a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B v. E is LAMP3; and vi. F is SFTPA1 ; and c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein.
[0255] In some embodiments, the methods comprise a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is CDCP1 ; v. E is LAMP3; and vi. F is PLALIR; and c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein.
[0256] In some embodiments, the methods comprise a. obtaining a sample from the subject; determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; vi. F is SFTPA1 ; vii. G is SFTPA1 ; and vii. H is TNSF13B; and b. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein.
[0257] The methods may further comprise determining if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0258] In some embodiments, the methods comprise a. providing or determining an aggregate level of expression of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample obtained from the subject; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPA1; and b. comparing the aggregate level of expression of the proteins in part a. to a reference value.
[0259] In some embodiments, the methods comprise a. providing or determining an aggregate level of expression of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1, at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample obtained from the subject; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is CDCPI; v. E is LAMP3; and vi. F is PLALIR; and b. comparing the aggregate level of expression of the proteins in part a. to a reference value.
[0260] In some embodiments, the methods comprise a. providing or determining an aggregate level of expression of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; vi. F is SFTPA1 ; vii. G is SFTPA1 ; and vii. H is TNSF13B; and b. comparing the aggregate level of expression of the proteins in part a. to a reference value.
[0261] The methods may further comprise determining if the aggregate level of expression is higher than a reference value, wherein the reference value is an aggregate level of expression of the same proteins used to provide or determine the aggregate level of expression in step (a).
[0262] The methods provided herein include a method of predicting the risk of a subject developing lung cancer, a method of diagnosing lung cancer in a subject or a method of classifying a subject. In some embodiments, the methods are methods of treatment, which further comprise a step of administering a therapy (for example a lung cancer therapy) to the subject. In some embodiments, the methods of treatment are carried out on a subject who has been determined as being at risk of developing lung cancer, who has been diagnosed as having lung, or who has been classified according to a method of the present invention.
[0263] The invention also relates to kits for predicting the presence or absence of lung cancer in a subject or otherwise related to the methods described herein.
[0264] In one embodiment, the kit of the invention may comprise biosensor. A biosensor incorporates a biological sensing element and provides information on a biological sample, for example the presence (or absence) or concentration of an analyte. Specifically, they combine a biorecognition component (a bioreceptor) with a physiochemical detector for detection and / or quantification of an analyte (such as an RNA, a cDNA or a protein).
[0265] The bioreceptor specifically interacts with or binds to the analyte of interest and may be, for example, an antibody or antibody fragment, an enzyme, a nucleic acid, an organelle, a cell, a biological tissue, imprinted molecule or a small molecule. The bioreceptor may be immobilised on a support, for example a metal, glass or polymer support, or a 3-dimensional lattice support, such as a hydrogel support.
[0266] Biosensors are often classified according to the type of biotransducer present. For example, the biosensor may be an electrochemical (such as a potentiometric), electronic, piezoelectric, gravimetric, pyroelectric biosensor or ion channel switch biosensor. The transducer translates the interaction between the analyte of interest and the bioreceptor into a quantifiable signal such that the amount of analyte present can be determined accurately. Optical biosensors may rely on the surface plasmon resonance resulting from the interaction between the bioreceptor and the analyte of interest. The SPR can hence be used to quantify the amount of analyte in a test sample. Other types of biosensor include evanescent wave biosensors, nanobiosensors and biological biosensors (for example enzymatic, nucleic acid (such as DNA), antibody, epigenetic, organelle, cell, tissue or microbial biosensors).
[0267] The invention also provides microarrays (RNA, DNA or protein) comprising capture molecules (such as RNA or DNA oligonucleotides) specific for each of the biomarkers or biomarker panels being quantified, wherein the capture molecules are immobilised on a solid support. The microarrays are useful in the methods of the invention.
[0268] In particular, the present invention provides a combination of binding molecules, wherein each binding molecule specifically binds a different target analyte.
[0269] The binding molecules may be present on a solid substrate, such an array (for example an RNA microarray, in which case the binding molecules are RNAs that hybridise to the target miRNA). The binding molecules may all be present on the same solid substrate. Alternatively, the binding molecules may be present on different substrates. In some embodiments of the invention, the binding molecules are present in solution.
[0270] These kits may further comprise additional components, such as a buffer solution. Other components may include a labelling molecule for the detection of the bound miRNA and so the necessary reagents (i.e. enzyme, buffer, etc) to perform the labelling; binding buffer; washing solution to remove all the unbound or non-specifically bound miRNAs. Hybridisation will be dependent on the size of the putative binder, and the method use may be to be determined experimentally, as is standard in the art. As an example, hybridisation can be performed at ~20°C below the melting temperature (Tm), over-night. (Hybridisation buffer: 50% deionised formamide, 0.3 M NaCI, 20 mM Tris-HCI, pH 8.0, 5 mM EDTA, 10 mM phosphate buffer, pH 8.0, 10% dextran sulfate, 1 x Denhardt’s solution, and 0.5 mg / mL yeast tRNA). Washes can be performed at 4-6°C higher than hybridization temperature with 50% Formamide / 2x SSC (20x Standard Saline Citrate (SSC), pH 7.5: 3 M NaCI, 0.3 M sodium citrate, the pH is adjusted to 7.5 with 1 M HCI). A second wash can be performed with 1xPBS / 0.1% Tween 20.
[0271] Binding or hybridisation of the binding molecules to the target analyte may occur under standard or experimentally determined conditions. The skilled person would appreciate what stringent conditions are required, depending on the biomarkers being measured. The stringent conditions may include a hybridisation buffer that is be high in salt concentration, and a temperature of hybridisation high enough to reduce non-specific binding.
[0272] As used herein, “stringent conditions for hybridization” are known to those skilled in the art.23Stringent conditions may be defined as equivalent to hybridization in 6X sodium chloride / sodium citrate (SSC) at 45°C, followed by a wash in 0.2 X SSC, 0.1 % SDS at 65°C. Alternatively, stringent conditions may be defined as equivalent to hybridization in 50 % v / v formamide, 10 % w / v Dextran sulphate, 2X SSC at 37°C, followed by a wash in 50 % formamide 12x SSC at 42°C.
[0273] In one embodiment of the invention, the kit is able to simultaneously measure both miRNA biomarkers and protein biomarkers.
[0274] The present invention also provides a microarray, comprising specific binding molecules that hybridize to an expression product comprising biomarker panels of the invention described herein. The microarray can be a DNA or RNA microarray. The microarray may comprise a sample from a subject. In some embodiments, the specific binding molecules are oligonucleotides. When in use, the expression products may be hybridized to the corresponding specific binding molecules.
[0275] The present invention also provides use of the biomarker panels of the invention (or subset selection thereof) in a method of diagnosis or prognosis of lung cancer. Such uses are generally in vitro or ex vivo uses. The present invention also provides the use of the biomarker panels of the invention (or subset selection thereof) in the manufacture of a biosensor, such as a microarray, suitable for detection and / or quantification or each of the biomarkers.
[0276] When the invention uses one or more biomarkers, the biomarkers may all be measured in a single sample obtained from a subject. Alternatively, multiple samples may be taken from the subject. If multiple samples are available, or if a sample is divided into separate samples, different samples can be used for each protein being measured.
[0277] In some embodiments of the invention, the method may comprise providing an expression profile comprising the expression level of each of the proteins being measured. A measurement of expression, such as an expression profile, may be provided by quantifying one or more expression products. The expression products may be proteins or nucleic acids. In some preferred embodiments, the methods comprise quantifying RNA corresponding to the proteins being measured. In other preferred embodiments, the methods comprise quantifying proteins, for example using immunohistochemical methods. Thus, in one embodiment of the invention there is provided a method comprising:
[0278] (i) providing or obtaining a subject sample;
[0279] (ii) determining the protein expression profile of the sample, wherein the protein expression profile is based on the expression the proteins described herein being measured;
[0280] (iii) optionally correlating the protein expression profile of the sample to a reference; and
[0281] (iv) diagnosing or prognosing lung cancer in the subject.
[0282] In some embodiments of the invention, the method comprises contacting the sample with a binding molecule or binding molecules specific for the protein being measured. The binding molecule can be any suitable binding molecule, for example a nucleic acid, an antibody, an antibody fragment, a protein or an aptamer, depending on the method being used.
[0283] Measurement of the proteins / biomarkers in the sample generally comprises a measurement of the level of expression of the proteins. This may be carried out using any suitable means, for example a measurement or analysis of expression products, such as proteins or nucleic acids. Analysis of RNA may be preferred. The RNA may be converted to cDNA prior to analysis. In other embodiments, immunohistochemical analysis, or other methods of quantification of proteins, may be preferred.
[0284] Levels of expression may be determined by, for example, quantifying the expression products (such as nucleic acids (e.g. RNA) or proteins) of the biomarkers in the sample (such as a tissue sample). Methods include real-time quantitative PCR, microarray analysis, RNA sequencing, Northern blot analysis and in situ hybridisation. There is also an nCounter Analysis system from NanoString and 'Integrated Comprehensive Droplet Digital Detection' (IC 3D) that has been developed for the digital quantification of RNA directly in plasma.17In this system the plasma sample containing target RNAs is encapsulated into microdroplets, enzymatically amplified and digitally counted using a novel, high-throughput 3D particle counter.
[0285] Methods of real-time qPCR can use stem-loop primers or a poly(A)tailing technique, to reverse transcribe RNA into complementary DNA (cDNA) for the amplification step. Generally using predesigned assays that target specific RNAs of interest, microarray analysis may comprise the steps of fluorescently labelling the RNAs, hybridization of the labelled RNAs to DNA (or RNA or LNA) probes on a solid-substrate array, washing the array, and scanning the array. RNA enrichment techniques may be particularly useful in methods involving microarrays. RNA sequencing is another method that can benefit from RNA enrichment, although this is not always necessary. RNA sequencing techniques generally use next generation sequencing methods (also known as high-throughput or massively parallel sequencing). These methods use a sequencing-by-synthesis approach and allow relative quantification and precise identification of RNA sequences. In situ hybridisation techniques can be used on tissue samples, both in vivo and ex vivo.
[0286] In some methods of the invention, detection and quantification of cDNA-binding molecule complexes may be used to determine RNA expression. For example, RNA transcripts in a sample may be converted to cDNA by reverse-transcription, after which the sample is contacted with binding molecules specific for the RNAs being quantified, detecting the presence of a of cDNA- specific binding molecule complex, and quantifying the expression of the corresponding gene. There is therefore provided the use of cDNA transcripts corresponding to one or more of the RNAs of interest, or combinations thereof, for use in methods of detecting, diagnosing or prognosis on lung cancer. In some embodiments of the invention, the method may therefore comprise a step of conversion of the RNAs to cDNA to allow a particular analysis to be undertaken and to achieve RNA quantification.
[0287] Methods for detecting the levels of protein expression include any methods known in the art. For example, protein levels can be measured indirectly using DNA or mRNA arrays. Alternatively, protein levels can be measured directly by measuring the level of protein synthesis or measuring protein concentration.
[0288] DNA and RNA arrays (microarrays) for use in quantification of the RNAs of interest comprise a series of microscopic spots of DNA or RNA oligonucleotides, each with a unique sequence of nucleotides that are able to bind complementary nucleic acid molecules. In this way the oligonucleotides are used as probes to which only the correct target sequence will hybridise under high-stringency conditions. In the present invention, the target sequence can be the coding DNA sequence or unique section thereof, corresponding to the RNA whose expression is being detected. Most commonly the target sequence is the RNA biomarker of interest itself.
[0289] Protein microarrays can also be used to directly detect protein expression. These are similar to DNA and RNA microarrays in that they comprise capture molecules fixed to a solid surface.
[0290] Capture molecules include antibodies, proteins, aptamers, nucleic acids, receptors and enzymes, which might be preferable if commercial antibodies are not available for the analyte being detected. Capture molecules for use on the arrays can be externally synthesised, purified and attached to the array. Alternatively, they can be synthesised in-situ and be directly attached to the array. The capture molecules can be synthesised through biosynthesis, cell-free DNA expression or chemical synthesis. In-situ synthesis is possible with the latter two. The appropriate capture molecule will depend on the nature of the target (e.g. mRNA, protein or cDNA).
[0291] Once captured on a microarray, detection methods can be any of those known in the art. For example, fluorescence detection can be employed. It is safe, sensitive and can have a high resolution. Other detection methods include other optical methods (for example colorimetric analysis, chemiluminescence, label free Surface Plasmon Resonance analysis, microscopy, reflectance etc.), mass spectrometry, electrochemical methods (for example voltametry and amperometry methods) and radio frequency methods (for example multipolar resonance spectroscopy).
[0292] With respect to protein biomarkers, direct measurement of protein expression and identification of the proteins being expressed in a given sample can be done by any one of a number of methods known in the art. For example, 2-dimensional polyacrylamide gel electrophoresis (2D-PAGE) has traditionally been the tool of choice to resolve complex protein mixtures and to detect differences in protein expression patterns between normal and diseased tissue. Differentially expressed proteins observed between normal and tumour samples are separate by 2D-PAGE and detected by protein staining and differential pattern analysis. Alternatively, 2-dimensional difference gel electrophoresis (2D-DIGE) can be used, in which different protein samples are labelled with fluorescent dyes prior to 2D electrophoresis. After the electrophoresis has taken place, the gel is scanned with the excitation wavelength of each dye one after the other. This technique is particularly useful in detecting changes in protein abundance, for example when comparing a sample from a healthy subject and a sample form a diseased subject.
[0293] Commonly, proteins subjected to electrophoresis are also further characterised by mass spectrometry methods. Such mass spectrometry methods can include matrix-assisted laser desorption / ionisation time-of-f light (MALDI-TOF).
[0294] MALDI-TOF is an ionisation technique that allows the analysis of biomolecules (such as proteins, peptides and sugars), which tend to be fragile and fragment when ionised by more conventional ionisation methods. Ionisation is triggered by a laser beam (for example, a nitrogen laser) and a matrix is used to protect the biomolecule from being destroyed by direct laser beam exposure and to facilitate vaporisation and ionisation. The sample is mixed with the matrix molecule in solution and small amounts of the mixture are deposited on a surface and allowed to dry. The sample and matrix co-crystallise as the solvent evaporates. Protein microarrays can also be used to directly detect protein expression. These are similar to DNA and mRNA microarrays in that they comprise capture molecules fixed to a solid surface. Capture molecules are most commonly antibodies specific to the proteins being detected, although antigens can be used where antibodies are being detected in serum. Further capture molecules include proteins, aptamers, nucleic acids, receptors and enzymes, which might be preferable if commercial antibodies are not available for the protein being detected. Capture molecules for use on the protein arrays can be externally synthesised, purified and attached to the array. Alternatively, they can be synthesised in-situ and be directly attached to the array. The capture molecules can be synthesised through biosynthesis, cell-free DNA expression or chemical synthesis. In-situ synthesis is possible with the latter two. There is therefore provided a protein microarray comprising capture molecules (such as antibodies) specific for each of the biomarkers being quantified immobilised on a solid support.
[0295] Once captured on a microarray, detection methods can be any of those known in the art. For example, fluorescence detection can be employed. It is safe, sensitive and can have a high resolution. Other detection methods include other optical methods (for example colorimetric analysis, chemiluminescence, label free Surface Plasmon Resonance analysis, microscopy, reflectance etc.), mass spectrometry, electrochemical methods (for example voltametry and amperometry methods) and radio frequency methods (for example multipolar resonance spectroscopy).
[0296] Additional methods of determine protein concentration include mass spectrometry and / or liquid chromatography, such as LC-MS, LIPLC, or a tandem UPLC-MS / MS system.
[0297] Methods of the invention involving quantitative analysis, such as quantitative microarray analysis, may be preferred.
[0298] Immunohistochemical methods are useful in the present invention for quantification of protein expression. Such methods are known to the person of skill in the art, for example those discussed in Cregger et al., 2006, Arch Pathol Lab Med, 130(7):1026-1030.18. An example of a suitable technique is paraffin-embedded Q-IHC.
[0299] Once the level of expression or concentration has been determined, the level can be compared to a previously measured level of expression or concentration (either in a sample from the same subject but obtained at a different point in time, or in a sample from a different subject, for example a healthy subject, i.e. a control or reference sample) to determine whether the level of expression or concentration is higher or lower in the sample being analysed. Hence, the methods of the invention may further comprise a step of correlating said detection or quantification with a control or reference to determine if lung cancer is present (or suspected) or not, or to determine the lung cancer prognosis. Said correlation step may also detect the presence of particular types of lung cancer and to distinguish these subjects from healthy subjects, in which no lung cancer is present. In particular, the invention is particularly useful for predicting lung cancer metastasis.
[0300] Said step of correlation may include comparing the amount (expression or concentration) of the biomarkers with the amount of the corresponding biomarker(s) in a reference sample, for example in a biological sample taken from a healthy subject. Generally, the methods of the invention do not include the steps of determining the amount of the corresponding biomarker in a reference sample, and instead such values will have been previously determined. However, in some embodiments the methods of the invention may include carrying out the method steps from a healthy subject who is used as a control. Alternatively, the method may use reference data obtained from samples from the same subject at a previous point in time. In this way, the effectiveness of any treatment can be assessed and a prognosis for the subject determined.
[0301] Internal controls can be also used, for example quantification of one or more different RNAs or proteins not part of the biomarker panel. This may provide useful information regarding the relative amounts of the biomarkers in the sample, allowing the results to be adjusted for any variances according to different populations or changes introduced according to the method of sample collection, processing or storage. Therefore, in some embodiments of the invention, the method may comprise the step of comparing the measured level of expression with one or more housekeeping proteins. Suitable housekeeping proteins are known to the skilled person.
[0302] As would be apparent to a person of skill in the art, any measurements of analyte concentration or expression may need to be normalised to take in account the type of test sample being used and / or and processing of the test sample that has occurred prior to analysis. Data normalisation also assists in identifying biologically relevant results. Invariant RNAs may be used to determine appropriate processing of the sample. Differential expression calculations may also be conducted between different samples to determine statistical significance.
[0303] In some embodiments of the invention, the methods comprise determining a ratio of the average expression level of the proteins positively correlated with disease score to that of the remaining negatively correlated proteins. This ratio is termed the matrix index and is indicative of metastasis and can be used to calculate the hazard ratio, which is indicative of the probability of subject survival.
[0304] In general, the methods of the present invention may comprise the steps of: a) providing or obtaining a biological sample, such as a tissue sample or bodily fluid sample (such as a blood or urine sample); b) optionally processing the sample, for example to extract the gene expression products (for example RNA or protein) from the sample; c) quantification of the gene expression products (such as RNA or protein) in the sample.
[0305] The methods may further comprise the step of: d) comparison of the level of gene expression product from step c) with a control or reference sample or value.
[0306] Alternatively, the method may comprise the step of: a) determining the average level of proteins expression of the proteins positively correlated with disease; b) determining the average level of proteins expression of the proteins negatively correlated with disease; c) determining a ratio of expression of the value determined in step (a) and the value determining in step (b); and optionally d) determining a hazard ratio by associating matrix index with subject survival.
[0307] The above methods provide a hazard ratio and gives an indication of the prognosis of the diseases (such as the risk of metastasis and / or an indication of the probability of long-term survival of the subject). The average level of protein expression may be normalised prior to determining the ration of expression.
[0308] A hazard ratio, for example a multivariate hazard ratio, may be determined by any suitable method known to the skilled person. For example, a hazard ratio may be derived from a Cox proportional hazards regression model. Such an analysis allows easier comparison across lung cancer types and / or datasets using the matrix index.
[0309] In embodiments where only one protein that is positively correlated with disease is measured, then no average needs to be determined. Similarly, in embodiments where only one protein that is negatively correlated with disease is measured, then no average needs to be determined. Instead, the expression level of the positively and / or negatively correlated protein can be used to determine the ratio of expression.
[0310] Where the hazard ratio is greater than 1 (preferably with a confidence internal of at least 95%), a poor prognosis is indicated and the probability of metastasis is increased. Where the hazard ratio is less than 1 (preferably with a confidence interval of at least 95%), a better prognosis is indicated and the probability of metastasis is decreased.
[0311] Certain aspect of methods of the invention may be carried out by a computer. The present invention therefore provides a computer programmed to carry out the methods of the invention, for example to determine average levels of expression of proteins in the protein panel, determine a ratio of positively to negatively correlated proteins, determine a matrix index and / or determine a hazard ratio as described herein. The computer may be further programmed to generate a report providing the results of the calculations, for example the matrix index and / or the hazard ratio.
[0312] In some embodiments of the invention, the step of quantification of protein expression may comprise the following steps: a) contacting the sample or extracted RNA or protein with a binding partner that specifically binds to the RNA(s) or protein(s) of interest b) quantifying the amount of RNA-binding partners or protein-binding partners to determine the amount of the RNA(s) or protein(s) present in the original sample.
[0313] The present invention therefore provides a reaction mixture, comprising either the RNAs or proteins of interest, or a biological sample (such as a tissue sample) containing the RNAs or proteins of interest, wherein the RNAs or proteins of interest are bound to a binding partner specific to the RNA or protein. The binding partner may be, for example, an oligonucleotide that hybridises to the RNA, or an antibody or antigen binding fragment thereof that specifically binds to the protein.
[0314] Alternatively, the reaction mixture may comprise cDNA molecules corresponding to the RNAs of interest, and it is the cDNAs that are bound to a specific binding partner. The RNAs of interest correlate to the proteins of the biomarkers being analysed.
[0315] The method of the invention can be carried out using a binding molecules or reagents specific for the expression products or cDNAs being detected. Binding molecules and reagents are those molecules that have an affinity for the target such that they can form binding molecule / reagent- biomarker complexes that can be detected using any method known in the art. The binding molecule of the invention can be an antibody, an antibody fragment, a nucleic acid, an oligonucleotide, a protein or an aptamer or molecularly imprinted polymeric structure, depending on the nature of the target (for example RNA or, in some embodiments, cDNA or protein). Methods of the invention may comprise contacting the biological sample with an appropriate binding molecule or molecules. Said binding molecules may form part of a kit of the invention, in particular they may form part of the biosensors of in the present invention.
[0316] Antibodies can include both monoclonal and polyclonal antibodies and can be produced by any means known in the art. Techniques for producing monoclonal and polyclonal antibodies which bind to a particular protein are now well developed in the art. They are discussed in standard immunology textbooks.19Polyclonal antibodies can be raised by stimulating their production in a suitable animal host (e.g. a mouse, rat, guinea pig, rabbit, sheep, chicken, goat or monkey) when the antigen is injected into the animal. If necessary, an adjuvant may be administered together with the antigen. The antibodies can then be purified by virtue of their binding to antigen or as described further below. Monoclonal antibodies can be produced from hybridomas. These can be formed by fusing myeloma cells and B-lymphocyte cells which produce the desired antibody in order to form an immortal cell line.20The antibodies may be human or humanised, or may be from other species.
[0317] The present invention includes antibody derivatives which are capable of binding to antigen. Thus, the present invention includes antibody fragments and synthetic constructs. Examples of antibody fragments and synthetic constructs are given in Dougall et al. (1994) Trends Biotechnol, 12:372- 379.21
[0318] Antibody fragments or derivatives, such as Fab, F(ab')2 or Fv may be used, as may single-chain antibodies (scAb) such as described by Huston et al. (1993) Int Rev Immunol, 10:195-21722, domain antibodies (dAbs), for example a single domain antibody, or antibody-like single domain antigen-binding receptors. In addition, antibody fragments and immunoglobulin-like molecules, peptidomimetics or non-peptide mimetics can be designed to mimic the binding activity of antibodies. Fv fragments can be modified to produce a synthetic construct known as a single chain Fv (scFv) molecule. This includes a peptide linker covalently joining VH and VL regions which contribute to the stability of the molecule. The present invention therefore also extends to single chain antibodies or scAbs.
[0319] Other synthetic constructs include CDR peptides. These are synthetic peptides comprising antigen binding determinants. These molecules are usually conformationally restricted organic rings which mimic the structure of a CDR loop and which include antigen-interactive side chains. Synthetic constructs also include chimeric molecules. Thus, for example, humanised (or primatised) antibodies or derivatives thereof are within the scope of the present invention. An example of a humanised antibody is an antibody having human framework regions, but rodent hypervariable regions. Synthetic constructs also include molecules comprising a covalently linked moiety which provides the molecule with some desirable property in addition to antigen binding. For example the moiety may be a label (e.g. a detectable label, such as a fluorescent or radioactive label) or a pharmaceutically active agent.
[0320] In those embodiments of the invention in which the binding molecule is an antibody or antibody fragment, the method of the invention can be performed using any immunological technique known in the art. For example, ELISA, radio immunoassays, bead-based, or similar techniques may be utilised. In general, an appropriate autoantibody is immobilised on a solid surface and the sample to be tested is brought into contact with the autoantibody. If the lung cancer biomarker recognised by the autoantibody is present in the sample, an antibody-marker complex is formed. The complex can then be directed or quantitatively measured using, for example, a labelled secondary antibody which specifically recognises an epitope of the biomarker. The secondary antibody may be labelled with biochemical markers such as, for example, horseradish peroxidase (HRP) or alkaline phosphatase (AP), and detection of the complex can be achieved by the addition of a substrate for the enzyme which generates a colorimetric, chemiluminescent or fluorescent product. Alternatively, the presence of the complex may be determined by addition of a protein labelled with a detectable label, for example an appropriate enzyme. In this case, the amount of enzymatic activity measured is inversely proportional to the quantity of complex formed and a negative control is needed as a reference to determining the presence of antigen in the sample. Another method for detecting the complex may utilise antibodies or antigens that have been labelled with radioisotopes followed by a measure of radioactivity. Examples of radioactive labels for antigens include 3H, 14C and 1251.
[0321] Aptamers are oligonucleotides or peptide molecules that bind a specific target molecule. Oligonucleotide aptamers include DNA aptamers and RNA aptamers. Aptamers can be created by an in vitro selection process from pools of random sequence oligonucleotides or peptides. Aptamers can be optionally combined with ribozymes to self-cleave in the presence of their target molecule.
[0322] Aptamers can be made by any process known in the art. For example, a process through which aptamers may be identified is systematic evolution of ligands by exponential enrichment (SELEX). This involves repetitively reducing the complexity of a library of molecules by partitioning on the basis of selective binding to the target molecule, followed by re-amplification. A library of potential aptamers is incubated with the target biomarker before the unbound members are partitioned from the bound members. The bound members are recovered and amplified (for example, by polymerase chain reaction) in order to produce a library of reduced complexity (an enriched pool). The enriched pool is used to initiate a second cycle of SELEX. The binding of subsequent enriched pools to the target biomarker is monitored cycle by cycle. An enriched pool is cloned once it is judged that the proportion of binding molecules has risen to an adequate level. The binding molecules are then analysed individually. SELEX is reviewed in Fitzwater & Polisky (1996) Methods Enzymol, 267:275-30124.
[0323] Thus, in one embodiment of the invention, there is provided a method of analysing a biological sample from a subject, comprising contacting the sample with reagents or binding molecules specific for the biomarker(s) being quantified, and measuring the abundance of biomarkerreagent or biomarker-binding molecule complexes, and correlating the abundance of biomarkerreagent or biomarker-binding molecule complexes with the concentration of the relevant biomarker in the biological sample. For example, in one embodiment of the invention, the method comprises the steps of: a) contacting a biological sample with reagents or binding molecules specific for the proteins in a biomarker panel of the invention described herein; b) quantifying the abundance of biomarker-reagent or biomarker-binding molecule complexes for proteins in a biomarker panel of the invention described herein; and c) correlating the abundance of biomarker-reagent or biomarker-binding molecule complexes with the concentration or expression of proteins in a biomarker panel of the invention described herein in the biological sample.
[0324] The method may further comprise the step of d) comparing the concentration or expression of the biomarkers in step c) with a reference to diagnose or prognose lung cancer. The subject can then be treated accordingly. Alternatively, a ratio between the proteins positively correlated with disease to the proteins negatively associated with disease may be determined. As discussed elsewhere, suitable reagents or binding molecules may include an antibody or antibody fragment, an enzyme, a nucleic acid, an organelle, a cell, a biological tissue, imprinted molecule or a small molecule. Such methods may be carried out using kits or biosensors of the invention.
[0325] Determining the level of expression of the proteins in the sample may comprise measuring the concentration of the proteins on cells present in the sample by a method known in the art. Measuring the concentration of the proteins on cells present in the sample may comprise the use of antibodies, or binding fragments thereof, which specifically bind to each protein. The antibodies, or binding fragments thereof, may comprise an oligonucleotide. Measuring the concentration of the proteins on cells present in the sample may comprise amplifying and detecting the oligonucleotide in the sample. The oligonucleotide may be amplified and detected when a pair of antibodies, or binding fragments thereof, specific for the same protein bind to a cell, the oligonucleotides of each antibody hybridise and the hybridisation product is amplified and detected. The hybridisation product may be amplified and detected by qPCR. The step of determining the level of expression of the proteins in the sample may comprise quantitative PCR (qPCR). Alternatively, measuring the concentration of the proteins on cells present in the sample may comprise flow cytometry (e.g., FACS), immunofluorescence, Western blot, surface plasmon resonance, labelling with fluorescent or biotinylated antibodies, lateral flow assay, cell surface biotinylation and pull-down assay, cytokine bead array, immunoprecipitation, CRISPR / Cas9- based surface protein tagging or single-cell RNA sequencing (scRNA-seq).
[0326] Alternatively, determining the level of expression of the proteins in the sample may comprise measuring the concentration of the proteins in solution (e.g., in a plasma sample from the subject) by a method known in the art. Such methods include Bradford assay, BCA (bicinchoninic acid) assay, lowry assay, UV absorption, NanoDrop (spectrophotometry), ELISA (Enzyme-Linked Immunosorbent Assay), protein chip (microarray), colorimetric and fluorometric methods.
[0327] Aspects and embodiments described herein with the term “comprising” may include other features or steps within the scope. It is also understood that aspects and embodiments described as “comprising” also describes aspect and embodiments wherein the term “comprising” is replaced by the term “consisting essentially of’ or “consisting of”.
[0328] The phrase "selected from the group comprising" may be substituted with the phrase "selected from the group consisting of" and vice versa, wherever they occur herein.
[0329] It is also understood that the application discloses all combinations of any of the above aspects and embodiments described above with each other, unless the context demands otherwise. Similarly, the application discloses all combinations of the preferred and / or optional features either singly or together with any of the other aspects, unless the context demands otherwise.
[0330] Preferred features for the second and subsequent aspects are as provided for the first aspect, mutatis mutandis.
[0331] The invention will now be further described by way of the following Examples, which are meant to serve to assist one of ordinary skill in the art in carrying out the invention and are not intended in any way to limit the scope of the invention, with reference to the Figures.
[0332] EXAMPLES
[0333] Table 1 : 15 protein biomarkers
[0334] Table 2: 9 protein biomarkers Table 3: 13 protein biomarkers
[0335] Table 4: 9 protein biomarkers
[0336] OTHER EMBODIMENTS
[0337] Other embodiments of the invention are set out below. 1. A method of predicting, or determining the risk of, a subject developing lung cancer, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1, at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B v. E is LAMP3; and vi. F is SFTPAI ; c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. predicting, or determining the risk of, the subject developing lung cancer if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0338] 2. A method of diagnosing lung cancer in a subject, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPAI ; c. comparing the level of expression of each of the proteins determined in part b. to a reference values for each protein; and d. diagnosing lung cancer in the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0339] 3. A method of classifying a subject, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15;
[0340] Hi. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPAI ; c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. classifying the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0341] 4. A method of treating a subject: i) predicted to develop, or be at risk of developing, lung cancer according to the method of embodiment 1 ; ii) diagnosed with lung cancer according to the method of embodiment 2; or iii) classified according to the method of embodiment 3; wherein the method comprises administering a treatment for lung cancer to the subject.
[0342] 5. A method of treating lung cancer in a subject, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPAI ; c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. administering a treatment for lung cancer to the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0343] 6. The method according to any one of the preceding embodiments, wherein the lung cancer is selected from the group consisting of: adenocarcinoma, adenosquamous cell cancer, large cell cancer, neuroendocrine cancer, a non-small cell lung cancer (NSCLC), a small cell cancer or a squamous cell cancer.
[0344] 7. The method according to any one of the preceding embodiments wherein the lung cancer is early stage lung cancer.
[0345] 8. The method according to embodiment 7, wherein the lung cancer is stage I, stage II, stage III and / or stage IV lung cancer.
[0346] 9. A method of processing a sample obtained from a subject, the method comprising: a. providing the sample obtained from the subject; b. processing the sample for determination of the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPAI ; c. determining the level of expression of each of the proteins in part b.; d. comparing the level of expression of each of the proteins determined in part c. to a reference value for each protein.
[0347] 10. A method of predicting, or determining the risk of, a subject developing lung cancer, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample obtained from the subject; wherein i. A is ALPP; ii. B is GDF15;
[0348] Hi. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPAI ; b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. predicting, or determining the risk of, the subject developing lung cancer if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins. ethod of diagnosing lung cancer in a subject, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample from the subject; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPAI ; b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. diagnosing lung cancer in the subject if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins ethod of classifying a subject, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample obtained from the subject; wherein i. A is ALPP; ii. B is GDF15;
[0349] Hi. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPAI ; b. comparing the level of expression of each of the proteins in part a. to a reference values for each protein; and c. classifying the subject if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins. A method of treating a subject: i) predicted to develop, or be at risk of developing lung cancer according to the method of embodiment 10; ii) diagnosed with lung cancer according to the method of embodiment 11 ; or iii) classified according to the method of embodiment 12; wherein the method comprises administering a treatment for lung cancer to the subject. A method of treating lung cancer in a subject, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in a sample from the subject; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPAI ; b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. administering a treatment for lung cancer to the subject if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
[0350] 15. The method according to any one of embodiments 10 to 14 as further defined by any one of embodiments 1 to 8.
[0351] 16. The method according to any one of the preceding embodiments wherein, wherein the subject has not previously been diagnosed with lung cancer.
[0352] 17. The method according to any one of the preceding embodiments wherein, wherein the subject is a non-smoker.
[0353] 18. The method according to any one of the preceding embodiments wherein, wherein the subject is human.
[0354] 19. The method according to any one of the preceding embodiments wherein the sample is a blood sample.
[0355] 20. The method according to embodiment 19 wherein the sample is a plasma sample.
[0356] 21. The method according to any one of the preceding embodiments wherein the reference value for each protein corresponds to the level of expression of each of the same proteins in a sample from a subject who does not have lung cancer.
[0357] 22. A kit for predicting the presence or absence of lung cancer in a subject, wherein the kit comprises means for determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5 or all 6 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is ALPP; ii. B is GDF15; iii. C is MMP12; iv. D is TNFSF13B; v. E is LAMP3; and vi. F is SFTPAI; and a sample collection apparatus. The method or kit according to any one of the preceding embodiments, wherein the additional proteins are selected from the group consisting of: a) A; b) B; c) C; d) D; e) E; f) F; g) A and B; h) A and C; i) A and D; j) A and E; k) A and F; l) B and C; m) B and D; n) B and E; o) B and F; p) C and D; q) C and E; r) C and F; s) D and E; t) D and F; u) E and F; v) A, B and C; w) A, B and D; x) A, B and E; y) A, B and F; z) A, C and D; aa) A, C and E; ab) A, C and F; ac) A, D and E; ad) A, D and F; ae) A, E and F; af) B, C and D; ag) B, C and E; ah) B, C and F; ai) B, D and E; aj) B, D and F; ak) B, E and F; al) C, D and E; am) C, D and F; an) C, E and F; ao) D, E and F; ap) A, B, C and D; aq) A, B, C and E; ar) A, B, C and F; as) A, B, D and E; at) A, B, D and F; au) A, B, E and F; av) A, C, D and E; aw) A, C, D and F; ax) A, C, E and F; ay) A, D, E and F; az) B, C, D and E; ba) B, C, D and F; bb) B, C, E and F; be) B, D, E and F; bd) C, D, E and F; be) A, B, C, D and E; bf) A, B, C, D and F; bg) A, B, C, E and F; bh) A, B, D, E and F; bi) A, C, D, E and F; bj) B, C, D, E and F; and bk) A, B, C, D, E and F. The method or kit according to any one of the preceding embodiments, wherein the proteins further comprise those selected from the group consisting of: a) G; b) H; c) l; d) J; e) K; f) L; g) G and H; h) G and I; i) G and J; j) G and K; k) G and L; l) H and I; m) H and J; n) H and K; o) H and L; p) I and J; q) I and K; r) I and L; s) J and K; t) J and L; u) K and L; v) G, H and I; w) G, H and J; x) G, H and K; y) G, H and L; z) G, I and J; aa) G, I and K; ab) G, I and L; ac) G, J and K; ad) G, J and L; ae) G, K and L; af) H, I and J; ag) H, I and K; ah) H, I and L; ai) H, J and K; aj) H, J and L; ak) H, K and L; al) I, J and K; am) I, J and L; an) I, K and L; ao) J, K and L; ap) G, H, I and J; aq) G, H, I and K; ar) G, H, I and L; as) G, H, J and K; at) G, H, J and L; au) G, H, K and L; av) G, I, J and K; aw) G, I, J and L; ax) G, I, K and L; ay) G, J, K and L; az) H, I, J and K; ba) H, I, J and L; bb) H, I, K and L; be) H, J, K and L; bd) I, J, K and L; be) G, H, I, J and K; bf) G, H, I, J and L; bg) G, H, I, K and L; bh) G, H, J, K and L; bi) G, I, J, K and L; bj) H, I, J, K and L; and bk) G, H, I, J, K and L; wherein: i. G is CDCPI ; ii. H is PIGR; iii. I is PRSS8; iv. J is SFTPD; v. K is AGER; and vi. L is PLAUR. The method or kit according to any one of the preceding embodiments, wherein the proteins comprise of CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, TNFSF13B, LAMP3, SFTPA1 , CDCP1 , PIGR, PRSS8, SFTPD, AGER and PLAUR. The method or kit according to any one of the preceding embodiments, wherein the proteins consist of CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, TNFSF13B, LAMP3, SFTPA1 , CDCP1 , PIGR, PRSS8, SFTPD, AGER and PLAUR. 27. The method or kit according to any one of the preceding embodiments, wherein determining the level of expression of the proteins in the sample comprises measuring the concentration of the proteins on cells present in the sample.
[0358] 28. The method or kit according to embodiment 29, wherein measuring the concentration of the proteins on cells present in the sample comprises the use of antibodies, or binding fragments thereof, which specifically bind to each protein.
[0359] 29. The method or kit according to embodiment 30, wherein the antibodies, or binding fragments thereof, comprise an oligonucleotide.
[0360] 30. The method or kit according to embodiment 31 , wherein measuring the concentration of the proteins on cells present in the sample comprises amplifying and detecting the oligonucleotide in the sample.
[0361] 31. The method or kit according to embodiment 32, wherein the oligonucleotide is amplified and detected when a pair of antibodies, or binding fragments thereof, specific for the same protein bind to a cell, the oligonucleotides of each antibody hybridise and the hybridisation product is amplified and detected.
[0362] 32. The method or kit according to embodiment 33, when the hybridisation product is amplified and detected by qPCR.
[0363] 33. The method or kit according to embodiment 29, wherein the step of determining the level of expression of the proteins in the sample comprises quantitative PCR (qPCR).
[0364] EXAMPLE 1 - DEVELOPING A MACHINE LEARNING FRAMEWORK TO ELUCIDATE BIOMARKERS OF LUNG CANCER AND IDENTIFYING PROTEINS INDICATIVE OF FUTURE LUNG CANCER RISK
[0365] The inventors interrogated the association between circulating protein signatures and future lung cancer diagnoses using the UK Biobank (Figure 1). The UK Biobank (UKBB) is a prospective cohort study that recruited n=502,401 participants, aged 37-73, from 2006-2010, with a subset of patients (n=54,219) having Olink plasma proteomics measured (for 2923 proteins) from baseline blood samples7. The inventors determined cancer incidence through the UKBB’s linkage with the national cancer registries with diagnoses recorded using the tenth revision of the International Classification of Diseases (ICD10) codes. The inventors excluded participants in the model according to the following criteria: any cancer diagnosed pre-recruitment, a cancer diagnosis date entry but no corresponding cancer annotation, and this subsequently subset to patients with available baseline Olink plasma proteomics data (n=48,099). Of these individuals, n=375 developed lung cancer (LC) 0-6 years post-sample acquisition.
[0366] Next, the dataset was split into a train (75%) and held-out test (25%) set, stratifying by smoking status, sex, household income, educational attainment, LC diagnosis, age at baseline, body-mass index (BMI), and pack years of smoking. The latter three continuous variables were first categorized into quartiles. Missing values were treated as a separate category for the purposes of splitting, to allow the distribution of missingness across these variables to be factored into the selection of train and held-out samples. Splitting was performed using the MultilabelStratifiedShuffleSplit() function from the iterative-stratification (vO.1.7) Python package. Missing covariate data were imputed separately in the train and held out sets to minimise data leakage.
[0367] Missing covariate data were handled using multiple imputation with chained equations (MICE). Imputed covariables were smoking status (categorized into never, previous, and current; <1% missing), passive smoking (weekly hours of home tobacco exposure; 10.0% missing), pack-years of smoking (15.4% missing), body-mass index (BMI) (<1% missing), household income (dichotomized by <£31 ,000 annually; 14.6% missing) and educational attainment (split by degree or professional qualification status; 1.31% missing). To predict values for missing data points, MICE imputation models incorporated these variables, lung cancer diagnosis and follow-up duration in addition to: PM2.5, age at baseline, and sex. Continuous variables were imputed with predictive mean matching, whilst random forest and logistic regression were used to impute higher order categorical and binary variables, respectively. This yielded 15 variant imputed datasets, where missing covariate data were replaced with imputed values (thus yielding complete datasets). Due to variation in the modelling process, these 15 complete datasets naturally contain small variations across imputed variables. Each imputed dataset was independently used in the same analysis protocol.
[0368] The LIKBB proteomics dataset comprised 2923 protein markers, and the inventors utilised feature selection to identify which proteins should be included in downstream modelling. Within the training data, the dataset was stratified by the number of incident lung cancer diagnoses and split into 5 folds. Each hyper-parameter optimization trial consisted of both demographic and proteomic features with 5 repeats of 5-fold cross-validation and 100 total trials. Following each trial, the model suggested a set of optimised hyperparameters and each feature was given a SHAP (SHapley Additive exPlanations) score, which measures the contribution of a feature to the model prediction. After each trial, the model eliminated the features with the lowest 5% of SHAP scores. Features that achieved the peak performance in the validation set were selected. For hyperparameter optimisation, model training, and performance evaluation, the inventors used the ROC-AUC as the objective function for the Extreme gradient boosting classification model (XGBoost, v2.0.3). Bayesian hyperparameter optimisation was performed using the Tree- structured Parzen (TPE) sampler from the Optuna python package (v3.5.0;) with 100 trials. Due to data imbalance (n=375 participants with a LC diagnosis compared to n=47,720 without), the inventors randomly undersampled the majority class to generate a 1 :1 ratio between cases and controls (using imbalanced-learn v0.12.0;). The inventors utilised the mean predicted probability of each the 100 different models that were created as the predicted probability per individual (Figure 1).
[0369] This modelling framework achieved a ROC-AUC of 0.928 (95% Confidence Interval 0.876 - 0.947) on the CV test and a ROC-AUC of 0.878 (0.841-0.911) on the held-out test set, confirming its ability to accurate distinguish between individuals who do and do not develop a future LC diagnosis. Furthermore, through using recursive feature selection, the inventors were able to capture a maximally informative protein subset for LC early detection.
[0370] EXAMPLE 2 - IDENTIFIED PROTEINS OUTPERFORM EXISTING LUNG CANCER RISK PREDICTION BENCHMARK MODELS
[0371] To benchmark the final model, consisting of 15 proteins (Table 1) alongside age, smoking status, BMI, and pack years, its performance was compared against the top two probabilistic lung cancer (LC) risk prediction models (LCRAT and LLPv3) in the UK Biobank dataset5. These are demographic models which use age- and smoking- based cut-offs in order to predict the probability of an individual developing lung cancer and thus have been proposed for stratifying individuals who might benefit from screening. This comparison was conducted on the held-out test 25% of the dataset. The area under the receiver operating characteristic curve (AUC) was calculated using the Icmodels package, and the statistical significance of differences in AUC values was assessed using DeLong's test (Figure 2A). For sensitivity analysis, the held-out test set was further stratified by two-year intervals prior to lung cancer diagnosis to evaluate model performance (Figure 2B). Therefore, the proteins identified by the inventors confer superior predictive performance for identifying individuals at risk of LC, when compared to existing alternatives.
[0372] EXAMPLE 3 - CANDIDATE PROTEINS ASSOCIATE WITH LUNG CANCER INCIDENCE IN MULTIPLE COHORTS
[0373] To investigate the association between specific proteins and lung cancer incidence, the inventors analysed proteomic data from individuals collected prior to their lung cancer diagnosis across five validation cohorts6’10'13(Figure 3). Within each cohort, protein levels were measured, and the association with lung cancer incidence was quantified using hazard ratios derived from Cox proportional hazards models (where time-to-diagnosis data were available). To account for heterogeneity in cohort design and characteristics, a random-effects meta-analysis was employed to integrate findings across cohorts, ensuring robust estimation of the protein-lung cancer associations. This confirms that all 15 identified proteins are associated with LC risk across multiple cohorts, after adjusting for the effects of age and sex, and thus likely reflect processes underlying lung tumour development.
[0374] EXAMPLE 4- LONGITUDINAL PROTEOMICS SAMPLING CONFIRMS THAT LUNG CANCER RISK PROTEINS INCREASE PRIOR TO DIAGNOSIS
[0375] The temporal patterns of plasma protein levels were analysed longitudinally in a cohort of 250 women (100 lung cancer cases and 150 controls) over five years preceding clinical lung cancer diagnosis6. Plasma samples from all were analysed annually using the Olink Oncology II proteomics panel. Statistical analyses were conducted using locally estimated scatterplot smoothing (LOESS) curves to assess trends in protein levels (which were Z-scored prior to analysis) over time relative to lung cancer diagnosis (Figure 4). The difference in protein levels between women who develop LC and those who do not is significant 2 years prior to LC diagnosis with these proteins being expressed more highly in individuals who develop a future LC.
[0376] EXAMPLE 5 - IDENTIFIED PROTEINS ARE ENRICHED WITHIN HEALTHY LUNG TISSUE AND THUS DO NOT MERELY REFLECT PRE-EXISTING LUNG TUMOURS
[0377] To further elucidate the potential tissue origins of these proteins within the human body, the inventors analysed bulk RNA sequencing (RNA-seq) data from normal tissues spanning a diverse range of human organs. This analysis aimed to determine the relative expression levels of the proteins across different tissue types. Notably, these proteins exhibited the highest expression in lung tissue compared to other organs, suggesting the lung as their primary site of production or activity within the body (Figure 5). Since these proteins were enriched within normal lung tissues, they are unlikely to reflect pre-existing lung tumours and instead are more likely to capture processes occurring within local microenvironmental niches that could predispose towards lung cancer development.
[0378] In conclusion, the inventors have used a robust modelling framework to identify a set of circulating proteins that allow maximal LC predictive performance, validate across five independent datasets, increase prior to LC diagnosis, and are enriched within lung tissue. This highlights their potential as minimally invasive early biomarkers for lung cancer detection. EXAMPLE 6 - IDENTIFICATION OF A PREFERRED PANEL OF 9 PROTEIN BIOMARKERS
[0379] From the list of 15 proteins recited in Table 1 , 9 preferred proteins were selected (see Table 2) that were validated across the most datasets.
[0380] EXAMPLE 7 - FURTHER VALIDATION OF THE PREDICTIVE POWER OF THE PROTEIN BIOMARKERS IN A NEVER-SMOKING COHORT
[0381] Further validation that the proteomic signature of the biomarkers described herein increases prior to diagnosis in never smoker lung cancer cases will be conducted by taking plasma samples from lung cancer cases and controls from the TALENT study in Taiwan. This is the only prospective trial to date which looks at screening lung cancer in never smokers, comprising of screening 12,011 individuals for lung cancer using low-dose CT every two years, resulting in a 2% diagnosis rate of invasive lung cancer.9The analysis will also further validate using the proteomic signature as part of a minimally invasive test use to predict lung cancer.
[0382] EXAMPLE 8 - RESULTS OF TALENT STUDY IN TAIWAN (FROM EXAMPLE 7)
[0383] To understand how the proteomic signature of the biomarkers described herein changed in never- smokers who later developed lung cancer, the inventors further interrogated the ability of the proteomic signature to predict future lung cancer diagnosis from a subset of never-smoking individuals enrolled in the multi-centre, prospective TALENT (Taiwan Lung Cancer Screening in Never-Smoker Trial) clinical trial. This was a multi-centre prospective cohort study conducted across 17 tertiary medical centres in Taiwan between 2015 and 20199and is one of the largest screening trials of its kind. All patients screened negative for lung cancer via a chest X-ray and had plasma samples collected at study entry. Subsequent incident diagnoses of invasive lung adenocarcinoma were identified through linkage with the cancer registry. For each case, two controls were selected by 1 :2 matching on age, sex, and smoking status. Quality control followed an established protocol7: samples with QC warnings were excluded, as were samples whose median NPX exceeded ±5 standard deviations from the median NPX across all samples. Proteins with >50% of measurements below the plate-specific limit of detection were removed. The inventors performed Olink® proteomics using four Target 96 panels which covered 10 / 14 (71.4%) proteins, on baseline plasma samples taken from a subset of the trial in TALENT (n = 251 cases, n = 501 age, sex and baseline smoking status matched controls, 81.3% female, median 144 days to diagnosis, 87.6% stage 1A disease and 62.1% adenocarcinomas). Four proteins (WFDC2, CXCL17, CEACAM5 and ALPP) were significantly associated with future lung cancer diagnoses. The above Example demonstrates that the proteomic signature of the biomarkers described herein can predict incident lung cancer in non-smokers / never-smokers.
[0384] In summary, the above data therefore demonstrates that the proteins of the invention can be used to identify individuals with, or at risk of developing, lung cancer that may not previously have been identified on basis of known demographic risk factors (e.g., smoking status).
[0385] REFERENCES
[0386] 1. https: / / gco.iarc.who.int / media / globocan / factsheets / cancers / 39-all-cancers-fact-sheet.pdf
[0387] 2. Santucci, C., Carioli, G., Bertuccio, P., Malvezzi, M., Pastorino, II., Boffetta, P., Negri, E., Bosetti, C. and La Vecchia, C., 2020. Progress in cancer mortality, incidence, and survival: a global overview. European Journal of Cancer Prevention, 29(5), pp.367-381.
[0388] 3. Oudkerk, M., Liu, S., Heuvelmans, M.A., Walter, J.E. and Field, J.K., 2021. Lung cancer LDCT screening and mortality reduction — evidence, pitfalls and future perspectives. Nature reviews Clinical oncology, 18(3), pp.135-151.
[0389] 4. Khan, S., Hatton, N., Tough, D., Rintoul, R.C., Pepper, C., Caiman, L., McDonald, F., Harris, C., Randle, A., Turner, M.C. and Haley, R.A., 2023. Lung cancer in never smokers (LCINS): development of a UK national research strategy. BJC Reports, 1(1), p.21.
[0390] 5. Feng, X., Goodley, P., Alcala, K., Guida, F., Kaaks, R., Vermeulen, R., Downward, G.S., Bonet, C., Colorado-Yohar, S.M., Albanes, D. and Weinstein, S.J., 2024. Evaluation of risk prediction models to select lung cancer screening participants in Europe: a prospective cohort consortium analysis. The Lancet Digital Health, 6(9), pp.e614-e624.
[0391] 6. Jacobs, I. J., Menon, U., Ryan, A., Gentry-Maharaj, A., Burnell, M., Kalsi, J.K., Amso, N.N., Apostolidou, S., Benjamin, E., Cruickshank, D. and Crump, D.N., 2016. Ovarian cancer screening and mortality in the UK Collaborative Trial of Ovarian Cancer Screening (UKCTOCS): a randomised controlled trial. The Lancet, 387(10022), pp.945-956.
[0392] 7. Sun, B.B., Chiou, J., Traylor, M., Benner, C., Hsu, Y.H., Richardson, T.G., Surendran, P., Mahajan, A., Robins, C., Vasquez-Grinnell, S.G. and Hou, L., 2023. Plasma proteomic associations with genetics and health in the UK Biobank. Nature, 622(7982), pp.329-338.
[0393] 8. The Genotype-Tissue Expression (GTEx) Project was supported by the Common Fund of the Office of the Director of the National Institutes of Health, and by NCI, NHGRI, NHLBI, NIDA, NIMH, and NINDS. The data used for the analyses described in this manuscript were obtained from: bulk RNAseq the GTEx Portal on 09 / 12 / 2024.
[0394] 9. Chang, G.C., Chiu, C.H., Yu, C.J., Chang, Y.C., Chang, Y.H., Hsu, K.H., Wu, Y.C., Chen, C.Y., Hsu, H.H., Wu, M.T. and Yang, C.T., 2024. Low-dose CT screening among never- smokers with or without a family history of lung cancer in Taiwan: a prospective cohort study. The Lancet Respiratory Medicine, 12(2), pp.141-152. 10. Carrasco-Zanini, J., Pietzner, M., Koprulu, M., Wheeler, E., Kerrison, N.D., Wareham, N.J. and Langenberg, C., 2024. Proteomic prediction of diverse incident diseases: a machine learning-guided biomarker discovery study using data from a prospective cohort study. The Lancet Digital Health, 6(7), pp.e470-e479.
[0395] 11. Eldjarn, G.H., Ferkingstad, E., Lund, S.H., Helgason, H., Magnusson, O.T., Gunnarsdottir, K., Olafsdottir, T.A., Halldorsson, B.V., Olason, P.I., Zink, F. and Gudjonsson, S.A., 2023. Large-scale plasma proteomics comparisons through genetics and disease associations. Nature, 622(7982), pp.348-358.
[0396] 12. Lind, L., Mazidi, M., Clarke, R., Bennett, D.A. and Zheng, R., 2024. Measured and genetically predicted protein levels and cardiovascular diseases in UK Biobank and China Kadoorie Biobank. Nature Cardiovascular Research, pp.1-10.
[0397] 13. "The blood proteome of imminent lung cancer diagnosis." Nature communications 14, no. 1 (2023): 3042.
[0398] 14. Ten Haaf K, van Rosmalen J, de Koning HJ. Lung cancer detectability by test, histology, stage, and gender: estimates from the NLST and the PLCO trials. Cancer Epidemiol Biomarkers Prev. 2015 Jan;24(1):154-61. doi: 10.1158 / 1055-9965. EPI- 14-0745. Epub 2014 Oct 13.
[0399] 15. https: / / olink.com / products / olink-flex.
[0400] 16 us / search?facets catcscoryT yp6=Ass3y'^ K ts&kayyvor s=alis ^ ss y^sortCriit(Bri '=r€H(Bv
[0401] 17. Zhang K, Kang DK, Ali MM, Liu L, Labanieh L, Lu M, Riazifar H, Nguyen TN, Zell JA, Digman MA, Gratton E, Li J, Zhao W. Digital quantification of miRNA directly in plasma using integrated comprehensive droplet digital detection. Lab Chip. 2015 Nov 7;15(21):4217-26.
[0402] 18. Cregger M, Berger AJ, Rimm DL. Immunohistochemistry and quantitative analysis of protein expression. Arch Pathol Lab Med. 2006 Jul; 130(7): 1026-30.
[0403] 19. Roitt, I, Brostoff, J, Male, D. Immunology, 2nded. London: Churchill Livingstone, 1984.
[0404] 20. Kohler, G., Milstein, C. Continuous cultures of fused cells secreting antibody of predefined specificity. Nature 256, 495-497 (1975).
[0405] 21. Dougall WC, Peterson NC, Greene Ml. Antibody-structure-based design of pharmacological agents. Trends Biotechnol. 1994 Sep; 12(9):372-379.
[0406] 22. Huston JS, McCartney J, Tai MS, Mottola-Hartshorn C, Jin D, Warren F, Keck P, Oppermann H. Medical applications of single-chain antibodies. Int Rev Immunol. 1993; 10(2-3): 195-217.
[0407] 23. Current Protocols in Molecular Biology, John Wiley & Sons, N.Y., 6.3.1-6.3.6, 1991.
[0408] 24. Fitzwater T, Polisky B. A SELEX primer. Methods Enzymol. 1996;267:275-301. EQUIVALENTS AND SCOPE
[0409] Those skilled in the art will appreciate that the present invention is defined by the appended claims and not by the Examples or other description of certain embodiments included herein.
[0410] Similarly, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise.
[0411] Unless defined otherwise above, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention. Generally, nomenclatures used in connection with, and techniques described herein are those well-known and commonly used in the art, or according to manufacturer’s specifications.
[0412] All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.
Claims
CLAIMS1. A method of predicting, or determining the risk of, a subject developing lung cancer, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. predicting, or determining the risk of, the subject developing lung cancer if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
2. A method of diagnosing lung cancer in a subject, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3;68vii. G is TNFSF13B; and viii. H is SFTPA1 c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. diagnosing lung cancer in the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
3. A method of classifying a subject, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. classifying the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
4. A method of treating a subject: i) predicted to develop, or be at risk of developing, lung cancer according to the method of claim 1 ; ii) diagnosed with lung cancer according to the method of claim 2; or69iii) classified according to the method of claim 3; wherein the method comprises administering a treatment for lung cancer to the subject.
5. A method of treating lung cancer in a subject, the method comprising: a. obtaining a sample from the subject; b. determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 c. comparing the level of expression of each of the proteins determined in part b. to a reference value for each protein; and d. administering a treatment for lung cancer to the subject if i) the level of expression of each of the proteins determined in part b. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
6. The method according to any one of the preceding claims, wherein the lung cancer is selected from the group consisting of: adenocarcinoma, adenosquamous cell cancer, large cell cancer, neuroendocrine cancer, a non-small cell lung cancer (NSCLC), a small cell cancer or a squamous cell cancer.
7. The method according to any one of the preceding claims wherein the lung cancer is early stage lung cancer.
8. The method according to claim 7, wherein the lung cancer is stage I, stage II, stage III and / or stage IV lung cancer.
9. A method of processing a sample obtained from a subject, the method comprising:70a. providing the sample obtained from the subject; b. processing the sample for determination of the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 c. determining the level of expression of each of the proteins in part b.; d. comparing the level of expression of each of the proteins determined in part c. to a reference value for each protein.
10. A method of predicting, or determining the risk of, a subject developing lung cancer, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in a sample obtained from the subject; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. predicting, or determining the risk of, the subject developing lung cancer if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the71total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
11. A method of diagnosing lung cancer in a subject, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in a sample from the subject; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. diagnosing lung cancer in the subject if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins12. A method of classifying a subject, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in a sample obtained from the subject; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. classifying the subject if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
13. A method of treating a subject: i) predicted to develop, or be at risk of developing lung cancer according to the method of claim 10; ii) diagnosed with lung cancer according to the method of claim 11 ; or iii) classified according to the method of claim 12; wherein the method comprises administering a treatment for lung cancer to the subject.
14. A method of treating lung cancer in a subject, the method comprising: a. providing the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E, F, G and H in a sample from the subject; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 b. comparing the level of expression of each of the proteins in part a. to a reference value for each protein; and c. administering a treatment for lung cancer to the subject if i) the level of expression of each of the proteins in part a. is higher than the reference value for each protein; ii) the level of expression of one or more of the proteins determined in part b. is higher than the reference value for each protein; or iii) the total level of expression of all proteins determined in part b. is higher than the total of the reference values for all proteins.
15. The method according to any one of claims 10 to 14 as further defined by any one of claims 1 to 8.
16. The method according to any one of the preceding claims wherein, wherein the subject has not previously been diagnosed with lung cancer and / or has not previously received treatment for lung cancer.
17. The method according to any one of the preceding claims wherein, wherein the subject is a non-smoker and / or has never smoked.
18. The method according to any one of the preceding claims wherein, wherein the subject is human.
19. The method according to any one of the preceding claims wherein the sample is a blood sample.
20. The method according to claim 19 wherein the sample is a plasma sample.21 . The method according to any one of the preceding claims wherein the reference value for each protein corresponds to the level of expression of each of the same proteins in a sample from a subject who does not have lung cancer.
22. A kit for predicting the presence or absence of lung cancer in a subject, wherein the kit comprises means for determining the level of expression of each of the proteins CXCL17, WFDC2 and CEACAM5 and at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or all 8 additional proteins selected from the group consisting of A, B, C, D, E and F in the sample; wherein i. A is PLAUR; ii. B is ALPP; iii. C is CDCPI ; iv. D is GDF15 v. E is MMP12; vi. F is LAMP3; vii. G is TNFSF13B; and viii. H is SFTPA1 and a sample collection apparatus.
23. The method or kit according to any one of the preceding claims, wherein the additional proteins are selected from the group consisting of. a) A; b) B; c) C; d) D; e) E; f) F; g) G; h) H; i) A, B; j) A, C; k) A, D; l) A, E; m) A, F; n) A, G; o) A, H;P) B, C; q) B, D; r) B, E; s) B, F; t) B, G; u) B, H; v) C, D; w) C, E; x) C, F; y) C, G; z) C, H; aa) D, E; ab) D, F; ac) D, G; ad) D, H; ae) E, F; af) E, G; ag) E, H; ah) F, G; ai) F, H; aj) G, H;ak) A, B, C; al) A, B, D; am) A, B, E; an) A, B, F; ao) A, B, G; ap) A, B, H; aq) A, C, D; ar) A, C, E; as) A, C, F; at) A, C, G; au) A, C, H; av) A, D, E; aw) A, D, F; ax) A, D, G; ay) A, D, H; az) A, E, F; ba) A, E, G; bb) A, E, H; be) A, F, G; bd) A, F, H; be) A, G, H; bf) B, C, D; bg) B, C, E; bh) B, C, F; bi) B, C, G; bj) B, C, H; bk) B, D, E; bl) B, D, F; bm) B, D, G; bn) B, D, H; bo) B, E, F; bp) B, E, G; bq) B, E, H; br) B, F, G; bs) B, F, H; bt) B, G, H; bu) C, D, E; bv) C, D, F;bw) C, D, G; bx) C, D, H; by) C, E, F; bz) C, E, G; ca) C, E, H; cb) C, F, G; cc) C, F, H; cd) C, G, H; ce) D, E, F; cf) D, E, G; eg) D, E, H; ch) D, F, G; ci) D, F, H; cj) D, G, H; ck) E, F, G; cl) E, F, H; cm) E, G, H; cn) F, G, H; co) A, B, C, D; cp) A, B, C, E; eq) A, B, C, F; er) A, B, C, G; cs) A, B, C, H; ct) A, B, D, E; cu) A, B, D, F; cv) A, B, D, G; cw) A, B, D, H; ex) A, B, E, F; cy) A, B, E, G; cz) A, B, E, H; da) A, B, F, G; db) A, B, F, H; de) A, B, G, H; dd) A, C, D, E; de) A, C, D, F; df) A, C, D, G; dg) A, C, D, H; dh) A, C, E, F;di) A, C, E, G; dj) A, C, E, H; dk) A, C, F, G; dl) A, C, F, H; dm) A, C, G, H; dn) A, D, E, F; do) A, D, E, G; dp) A, D, E, H; dq) A, D, F, G; dr) A, D, F, H; ds) A, D, G, H; dt) A, E, F, G; du) A, E, F, H; dv) A, E, G, H; dw) A, F, G, H; dx) B, C, D, E; dy) B, C, D, F; dz) B, C, D, G; ea) B, C, D, H; eb) B, C, E, F; ec) B, C, E, G; ed) B, C, E, H; ee) B, C, F, G; ef) B, C, F, H; eg) B, C, G, H; eh) B, D, E, F; ei) B, D, E, G; ej) B, D, E, H; ek) B, D, F, G; el) B, D, F, H; em) B, D, G, H; en) B, E, F, G; eo) B, E, F, H; ep) B, E, G, H; eq) B, F, G, H; er) C, D, E, F; es) C, D, E, G; et) C, D, E, H;eu) C, D, F, G; ev) C, D, F, H; ew) C, D, G, H; ex) C, E, F, G; ey) C, E, F, H; ez) C, E, G, H; fa) C, F, G, H; fb) D, E, F, G; fc) D, E, F, H; fd) D, E, G, H; fe) D, F, G, H; ff) E, F, G, H; fg) A, B, C, D, E; fh) A, B, C, D, F; fi) A, B, C, D, G; fj) A, B, C, D, H; fk) A, B, C, E, F; fl) A, B, C, E, G; fm) A, B, C, E, H; fn) A, B, C, F, G; fo) A, B, C, F, H; fp) A, B, C, G, H; fq) A, B, D, E, F; fr) A, B, D, E, G; fs) A, B, D, E, H; ft) A, B, D, F, G; fu) A, B, D, F, H; fv) A, B, D, G, H; fw) A, B, E, F, G; fx) A, B, E, F, H; fy) A, B, E, G, H; fz) A, B, F, G, H; ga) A, C, D, E, F; gb) A, C, D, E, G; gc) A, C, D, E, H; gd) A, C, D, F, G; ge) A, C, D, F, H; gf) A, C, D, G, H;gg) A, C, E, F, G; gh) A, C, E, F, H; gi) A, C, E, G, H; gj) A, C, F, G, H; gk) A, D, E, F, G; gl) A, D, E, F, H; gm) A, D, E, G, H; gn) A, D, F, G, H; go) A, E, F, G, H; gp) B, C, D, E, F; gq) B, C, D, E, G; gr) B, C, D, E, H; gs) B, C, D, F, G; gt) B, C, D, F, H; gu) B, C, D, G, H; gv) B, C, E, F, G; gw) B, C, E, F, H; gx) B, C, E, G, H; gy) B, C, F, G, H; gz) B, D, E, F, G; ha) B, D, E, F, H; hb) B, D, E, G, H; he) B, D, F, G, H; hd) B, E, F, G, H; he) C, D, E, F, G; hf) C, D, E, F, H; hg) C, D, E, G, H; hh) C, D, F, G, H; hi) C, E, F, G, H; hj) D, E, F, G, H; hk) A, B, C, D, E, F; hl) A, B, C, D, E, G; hm) A, B, C, D, E, H; hn) A, B, C, D, F, G; ho) A, B, C, D, F, H; hp) A, B, C, D, G, H; hq) A, B, C, E, F, G; hr) A, B, C, E, F, H;80hs) A, B, C, E, G, H; ht) A, B, C, F, G, H; hu) A, B, D, E, F, G; hv) A, B, D, E, F, H; hw) A, B, D, E, G, H; hx) A, B, D, F, G, H; hy) A, B, E, F, G, H; hz) A, C, D, E, F, G; ia) A, C, D, E, F, H; ib) A, C, D, E, G, H; ic) A, C, D, F, G, H; id) A, C, E, F, G, H; ie) A, D, E, F, G, H; if) B, C, D, E, F, G; ig) B, C, D, E, F, H; ih) B, C, D, E, G, H; ii) B, C, D, F, G, H; ij) B, C, E, F, G, H; ik) B, D, E, F, G, H; il) C, D, E, F, G, H; im) A, B, C, D, E, F, G; in) A, B, C, D, E, F, H; io) A, B, C, D, E, G, H; ip) A, B, C, D, F, G, H; iq) A, B, C, E, F, G, H; ir) A, B, D, E, F, G, H; is) A, C, D, E, F, G, H; it) B, C, D, E, F, G, H; and iu) A, B, C, D, E, F, G, H.
24. The method or kit according to any one of the preceding claims, wherein the proteins further comprise those selected from the group consisting of: a) I; b) J; c) K; d) L; e) I, J; f) I, K;81g) I, L; h) J, K; i) J, L; j) K, L; k) I, J, K; l) I, J, L; m) I, K, L; n) J, K, L; and o) I, J, K, L; wherein: i. I is PIGR; ii. J is PRSS8; iii. K is SFTPD; and iv. L is AGER.
25. The method or kit according to any one of the preceding claims, wherein the proteins comprise CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, LAMP3, CDCP1 , PIGR, PRSS8, SFTPD, AGER and PLAUR.
26. The method or kit according to any one of the preceding claims, wherein the proteins consist of CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, LAMP3, CDCP1 , PIGR, PRSS8, SFTPD, AGER and PLAUR.
27. The method or kit according to any one of the preceding claims, wherein the proteins comprise CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, TNFSF13B, LAMP3, SFTPA1 , CDCP1 , PIGR, PRSS8, SFTPD, AGER and PLAUR.
28. The method or kit according to any one of the preceding claims, wherein the proteins consist of CXCL17, WFDC2, CEACAM5, ALPP, GDF15, MMP12, TNFSF13B, LAMP3, SFTPA1 , CDCP1 , PIGR, PRSS8, SFTPD, AGER and PLAUR.
29. The method or kit according to any one of the preceding claims, wherein determining the level of expression of the proteins in the sample comprises measuring the concentration of the proteins on cells present in the sample.
30. The method or kit according to claim 29, wherein measuring the concentration of the proteins on cells present in the sample comprises the use of antibodies, or binding fragments thereof, which specifically bind to each protein.8231. The method or kit according to claim 30, wherein the antibodies, or binding fragments thereof, comprise an oligonucleotide.
32. The method or kit according to claim 31 , wherein measuring the concentration of the proteins on cells present in the sample comprises amplifying and detecting the oligonucleotide in the sample.
33. The method or kit according to claim 32, wherein the oligonucleotide is amplified and detected when a pair of antibodies, or binding fragments thereof, specific for the same protein bind to a cell, the oligonucleotides of each antibody hybridise and the hybridisation product is amplified and detected.
34. The method or kit according to claim 33, when the hybridisation product is amplified and detected by qPCR.
35. The method or kit according to claim 29, wherein the step of determining the level of expression of the proteins in the sample comprises quantitative PCR (qPCR).
36. The method or kit according to any one of the preceding claims, wherein the proteins do not comprise TNSF13B.
37. The method or kit according to any one of the preceding claims, wherein the proteins do not comprise SFTPA1 .
38. The method or kit according to any one of the preceding claims, wherein the proteins do not comprise TNSF13B or SFTPA1.