Lung cancer prediction and its applications
Patent Information
- Application Number
- JP2024520616
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-07
- Filing Date
- 2022-10-07
- Publication Date
- 2025-10-03
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 253,509, filed October 7, 2021, which is incorporated by reference herein in its entirety for all purposes.
[0002] The present application relates generally to methods for detecting biomarkers and assessing risk of lung cancer in an individual, and more specifically to one or more biomarkers, methods, devices, reagents, systems, and kits used to assess an individual for predicting risk of developing lung cancer within a specified time frame. [Background technology]
[0003] The following discussion provides a brief summary of information related to the present application and is not an admission that any of the information provided herein or publications referenced are prior art to the present application.
[0004] Lung cancer is the second most common type of cancer and the leading cause of cancer death for both men and women in the United States (Siegel et al. “Cancer Statistics, 2021.” CA Cancer J Clin 2021;71:7-33).
[0005] There are two major categories of lung cancer classified according to cell type, immunohistochemical, and molecular features: 1) non-small cell lung cancer, which includes squamous cell carcinoma, large cell carcinoma, and adenocarcinoma, accounting for approximately 85% of lung cancers; and 2) small cell lung cancer, which includes small cell carcinoma and small cell combined carcinoma. Small cell lung cancer grows quickly, and approximately 70% of patients with this type of cancer have already had the disease spread by the time of diagnosis. ("About Lung Cancer", American Cancer Society. Available online at https: / / www.cancer.org / cancer / lung-cancer / about / and "Cancer Stat Facts: Lung and Bronchus Cancer" National Cancer Institute. Surveillance, Epidemiology, and End Results Program. Available online at https: / / seer.cancer.gov / statfacts / html / lungb.html).
[0006] Lung cancer patients may present with various stages of disease, with early symptoms commonly observed as persistent cough, shortness of breath, and bloody sputum. The diagnosis of lung cancer is based on the initial presence of pulmonary nodules detected by chest imaging (low-dose computed tomography is the gold standard, but some clinical settings may also include chest x-ray or MRI as alternatives), followed by a biopsy. The stage of lung cancer is determined by the size, invasiveness, and spread of the tumor to lymph nodes. The prognosis of lung cancer is poor and worsens with each stage of the disease. The 5-year survival rate for patients with localized lung cancer stage is 59%, while the 5-year survival rate for patients with lung cancer stage with local spread drops to 32%, and the 5-year survival rate for patients with lung cancer stage with distant metastasis drops to 6% (https: / / seer.cancer.gov / statfacts / html / lungb.html).
[0007] Patients with lung cancer have different types of treatment options, including surgery, radiation therapy, chemotherapy, targeted therapy, and immunotherapy, depending on the stage of the cancer and their overall health. Patient Version” (August 2021) National Cancer Institute. Available online at https: / / www.cancer.gov / types / lung / patient / non-small-cell-lung-treatment-pdq#_118.
[0008] Approximately 1 in 15 men and 1 in 17 women will develop lung cancer during their lifetime (https: / / seer.cancer.gov / statfacts / html / lungb.html). However, the risk increases dramatically in smokers, with more than 80% of lung cancer cases in the United States occurring in smokers. “Lung Cancer Among People Who Never Smoked” (November 2020) Center for Disease Control and Prevention. Available online at https: / / www.cdc.gov / cancer / lung / nonsmokers / index.htm). Lung cancer risk in both male and female smokers increases with cumulative smoking amount and duration (defined as “pack-years”) and decreases in former smokers as time since quitting increases (Bruder et al. “Estimating lifetime and 10-year risk of lung cancer.” Prev Med Rep 2018;11:125-30 and Samet JM“Health benefits of smoking cessation.” Clin Chest Med 1991;12:669-79 and (Figure 1).
[0009] Although smoking is well-established to be the most penetrant risk factor for lung cancer, age is also a significant risk factor for lung cancer, with the median age at which lung cancer is diagnosed being 71 years, and lung cancer most frequently diagnosed in those aged 65-74 years. https: / / seer.cancer.gov / statfacts / html / lungb.html.
[0010] Additionally, other risk factors such as exposure to smoke, workplace chemicals (e.g., asbestos), and radiation, as well as clinical factors such as a family history of lung cancer, have also been associated with an increased risk of lung cancer or pulmonary disease in individuals. https: / / www.cancer.gov / types / lung / patient / non-small-cell-lung-treatment-pdq#_118.
[0011] The United States Preventive Services Task Force (USPSTF) recommends with moderate certainty (Grade B rating) that annual screening with low-dose computed tomography (LDCT) has a moderate net benefit for individuals considered at high risk for lung cancer. High-risk individuals are defined as those aged 50-80 years, with a smoking history of at least 20 pack-years, who are currently smoking or have quit smoking within the past 15 years. (Force USPST, et al. “Screening for Lung Cancer: US Preventive Services Task Force Recommendation Statement.” JAMA 2021;325:962-70).
[0012] These screening guidelines are restricted to high-risk individuals (based on age and smoking status) because there is ample evidence that annual LDCT screening reduces lung cancer mortality in high-risk individuals. For example, the National Lung Cancer Screening Trial (NLST) compared the effectiveness of LDCT scanning with chest radiography in individuals at high risk for lung cancer based on previous USPSTF guidelines (ages 55-80, with a smoking history of at least 30 pack-years, and either currently smoking or have quit within the past 15 years). The trial reported a 20% reduction in lung cancer mortality with LDCT screening, which was significantly better than chest radiography. (N National Lung Screening Trial Research Team. “Reduced lung-cancer mortality with low-dose computed tomographic screening.” N Engl J Med 2011;365:395-409). If an abnormality is found on LDCT, subsequent lung cancer screening is often changed to more frequent LDCT according to the Lung-RADS assessment category. (“Lung CT Screening Reporting & Data System (Lung-RADS)” American College of Radiology, available online at https: / / www.acr.org / Clinical-Resources / Reporting-and-Data-Systems / Lung-Rads).
[0013] The USPSTF does not recommend lung cancer screening for low-risk individuals (non-smokers) because there is insufficient evidence of net benefit in this population and the risks of harm from screening (including false-positive results leading to unnecessary testing, invasive procedures, overdiagnosis, radiation-induced cancers, incidental findings, and increased distress or anxiety) outweigh the benefits in low-risk populations. However, some health systems, under the supervision of a physician, may recommend periodic lung cancer screening for ineligible individuals who are lung cancer survivors, have a strong family history of lung cancer, or have occupational asbestos exposure. ("Lung Cancer Screening" (March 2021) Mayo Clinic. Available online at https: / / www.mayoclinic.org / tests-procedures / lung-cancer-screening / about / pac-20385024). Reimbursement for screening for these individuals is not guaranteed.
[0014] In addition to LDCT, lung cancer screening can also be performed by chest radiography, sputum cytology, and biomarker measurements, but there is insufficient evidence that these screening methods provide a mortality benefit, and the sensitivity of these techniques is lower than that of LDCT. (Force USPST, et al. “Screening for Lung Cancer: US Preventive Services Task Force Recommendation Statement.” JAMA 2021;325:962-70). Furthermore, if abnormal findings are found on any of these alternative methods of lung cancer screening, follow-up LDCT screening would be recommended to achieve a mortality benefit.
[0015] Lung cancer screening is a shared decision-making process between the patient and the healthcare provider and should be performed in addition to smoking cessation counseling (for current smokers). The risks, benefits, and level of evidence for each screening modality should be discussed, along with advice on where screening should be performed (at a high-quality lung cancer and treatment center that uses the standard Lung-RADS classification). If the risk of lung cancer is particularly high (e.g., a screening-eligible current heavy smoker, with a history of COPD, and a strong family history of lung cancer), physicians may change their guidance to strongly recommend LDCT as the screening modality (as it is the gold standard) and that screening be performed at a Screening Center of Excellence (to ensure that the highest levels of sensitivity and specificity are achieved).
[0016] Although it is well established that the benefit of lung cancer screening with LDCT is in reducing lung cancer mortality, available data indicate that uptake of lung cancer screening is low, with studies showing that only 14% of individuals eligible for lung cancer screening had been screened for lung cancer in the past year. (Zahnd et al. “Lung Cancer Screening Utilization: A Behavioral Risk Factor Surveillance System Analysis.” Am J Prev Med 2019;57:250-5). With an estimated 130,000+ deaths from lung cancer in the United States in 2021, (https: / / seer.cancer.gov / statfacts / html / lungb.html) there is a need for additional clinical courses of care to advise patients of their risk of lung cancer. There is currently no clinically accepted standard of care to assess a patient's risk for a future lung cancer diagnosis.
[0017] As explained above, physicians may recommend different lung cancer screening methods or referral to a lung cancer screening center based on an individual patient's risk level, however, screening tools are limited to detecting current lung cancer and cannot detect future risk. Various clinical risk calculators for future lung risk have been developed to predict an individual's lung cancer risk from a combination of demographics, personal and family health history, lifestyle, and carcinogen exposure levels, however, these calculators have not been routinely validated / replicated in independent cohorts, contain patient self-reported information, and many require non-standard clinical outcomes. Currently, no clinical risk calculators are widely used in clinical practice as standard of care. Thus, there is a need for biomarkers, methods, devices, reagents, systems, and kits to assess an individual's lung cancer risk. Summary of the Invention
[0018] This application discloses biomarkers, methods, devices, reagents, systems, and kits for assessing an individual's risk of lung cancer diagnosis within a specific time frame. In one aspect, the purpose of the lung cancer risk test disclosed herein is to create a model that predicts a current or former smoker's risk for lung cancer diagnosis within 5 years of blood draw.
[0019] Benefits of the lung cancer risk test disclosed herein include a convenient way to obtain personalized knowledge of the degree of risk for a future lung cancer diagnosis without relying on self-reported demographics or genetic background, the test results may impact compliance with lung cancer screening guidelines allowing for possible early detection of lung cancer and improving lung cancer survival, the test may impact positive behavioral change in modifiable risk-related behaviors (e.g., smoking cessation, dietary changes, weight loss), and the test may assist in the decision / recommendation of lung cancer screening by healthcare providers or the patient's desire for lung screening based on the test results (e.g., when a patient in a high-risk category first undergoes LDCT screening method, which is considered the gold standard, instead of a less sensitive alternative such as chest x-ray). Lung cancer risk testing may include identifying subjects who have lung cancer at the time of sampling.
[0020] The following numbered paragraphs
[0021] ~
[0116] includes a description of the broad combinations of technical features of the invention disclosed in this specification:
[0021] 1. a) measuring the level of PSP-94 protein and the levels of at least one, two, three, four, five, or six proteins selected from the group consisting of MMP-12, SP-D, HE4, PH, FUT5, and CRLF1 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on the level of PSP-94 and the levels of at least one, two, three, four, five, or six proteins. The method includes:
[0022] 2. a) Levels of PH protein, as well as MMP-12, SP- measuring the levels of at least one, two, three, four, five, or six proteins selected from the group consisting of D, HE4, PSP-94, FUT5, and CRLF1; b) identifying said human subject as at risk for developing lung cancer based on the level of said PH and the levels of said at least one, two, three, four, five, or six proteins.
[0023] 3. a) measuring the level of FUT5 protein and the level of at least one, two, three, four, five, or six proteins selected from the group consisting of MMP-12, SP-D, HE4, PSP-94, PH, and CRLF1 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on said level of FUT5 and said levels of at least one, two, three, four, five, or six proteins. The method includes:
[0024] 4. a) measuring the level of CRLF1 protein and the levels of at least one, two, three, four, five, or six proteins selected from the group consisting of MMP-12, SP-D, HE4, PSP-94, PH, and FUT5 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on the level of said CRLF1 and the levels of said at least one, two, three, four, five, or six proteins. The method includes:
[0025] 5. a) contacting a sample from a human subject with a set of capture reagents, each capture reagent having affinity for a different protein of a set of proteins comprising PSP-94 protein and at least 1, 2, 3, 4, 5, or 6 proteins selected from the group consisting of MMP-12, SP-D, HE4, PH, FUT5, and CRLF1; b) measuring the level of each protein of said set of proteins using said set of capture reagents; The method includes:
[0026] 6. a) contacting a sample from a human subject with a set of capture reagents, each capture reagent having affinity for a different protein of a set of proteins comprising a PH protein and at least 1, 2, 3, 4, 5, or 6 proteins selected from the group consisting of MMP-12, SP-D, HE4, PSP-94, FUT5, and CRLF1; b) measuring the level of each protein of said set of proteins using said set of capture reagents; The method includes:
[0027] 7. a) contacting a sample from a human subject with a set of capture reagents, each capture reagent having affinity for a different protein of a set of proteins comprising FUT5 protein and at least 1, 2, 3, 4, 5, or 6 proteins selected from the group consisting of MMP-12, SP-D, HE4, PSP-94, PH, and CRLF1; b) measuring the level of each protein of said set of proteins using said set of capture reagents; The method includes:
[0028] 8. a) contacting a sample from a human subject with a set of capture reagents, each capture reagent having affinity for a different protein of a set of proteins comprising CRLF1 protein and at least 1, 2, 3, 4, 5, or 6 proteins selected from the group consisting of MMP-12, SP-D, HE4, PSP-94, PH, and FUT5; b) measuring the level of each protein of said set of proteins using said set of capture reagents; The method includes:
[0029] 9. The method of embodiment 1 or embodiment 5, wherein said method comprises measuring PSP-94 and MMP-12; PSP-94 and SP-D; PSP-94 and HE4; PSP-94 and PH; PSP-94 and FUT5; or PSP-94 and CRLF1.
[0030] 10. The method of embodiment 1 or embodiment 5, wherein the method comprises measuring PSP-94, MMP-12, and SP-D; PSP-94, MMP-12, and HE4; PSP-94, MMP-12, and PH; PSP-94, MMP-12, and FUT5; PSP-94, MMP-12, and CRLF1; PSP-94, SP-D, and HE4; PSP-94, SP-D, and PH; PSP-94, SP-D, and FUT5; PSP-94, SP-D, and CRLF1; PSP-94, HE4, and PH; PSP-94, HE4, and FUT5; PSP-94, HE4, and CRLF1; PSP-94, PH, and FUT5; PSP-94, PH, and CRLF1; or PSP-94, FUT5, and CRLF1.
[0031] 11. The method of embodiment 2 or embodiment 6, wherein the method comprises measuring PH and MMP-12; PH and SP-D; PH and HE4; PH and PSP-94; PH and FUT5; or PH and CRLF1.
[0032] 12. The method of embodiment 2 or embodiment 6, wherein the method comprises measuring PH, MMP-12, and SP-D; PH, MMP-12, HE4; PH, MMP-12, and PSP-94; PH, MMP-12, and FUT5; PH, MMP-12, and CRLF1; PH, SP-D, and HE4; PH, SP-D, and PSP-94; PH, SP-D, and FUT5; PH, SP-D, and CRLF1; PH, HE4, and PSP-94; PH, HE4, and FUT5; PH, HE4, and CRLF1; PH, PSP-94, and FUT5; PH, PSP-94, and CRLF1; or PH, FUT5, and CRLF1.
[0033] 13. The method of embodiment 3 or embodiment 7, wherein said method comprises measuring FUT5 and MMP-12; FUT5 and SP-D; FUT5 and HE4; FUT5 and PSP-94; FUT5 and PH; or FUT5 and CRLF1.
[0034] 14. The method comprises the steps of: FUT5, MMP-12, and SP-D; FUT5, MMP-12, and HE4; FUT5, MMP-12, and PSP-94; FUT5, MMP-12, and PH; FUT5, MMP-12, and CRLF1; FUT5, SP-D, and HE4; FUT5, SP-D, and PSP-94; FUT5, SP-D, and PH; FUT5, SP-D, and CRLF1; FUT5, HE4, and PSP-94; FUT5, HE4, and PH; FUT5, HE4, and CRLF1; FUT5, PSP-94, and PH; FUT5, PSP-94, and CRLF1; or FUT5, PH, and CRLF1.
[0035] 15. The method of embodiment 4 or embodiment 8, wherein said method comprises measuring CRLF1 and MMP-12; CRLF1 and SP-D; CRLF1 and HE4; CRLF1 and PSP-94; CRLF1 and PH; or CRLF1 and FUT5.
[0036] 16. The method of embodiment 4 or embodiment 8, wherein the method comprises measuring CRLF1, MMP-12, and SP-D; CRLF1, MMP-12, and HE4; CRLF1, MMP-12, and PSP-94; CRLF1, MMP-12, and PH; CRLF1, MMP-12, and FUT5; CRLF1, SP-D, and HE4; CRLF1, SP-D, and PSP-94; CRLF1, SP-D, and PH; CRLF1, SP-D, and FUT5; CRLF1, HE4, and PSP-94; CRLF1, HE4, and PH; CRLF1, HE4, and FUT5; CRLF1, PSP-94, and PH; CRLF1, PSP-94, and FUT5; or CRLF1, PH, and FUT5.
[0037] 17. The method of embodiment 1 or embodiment 5, wherein said method comprises measuring PSP-94 and PH, and at least one of the following proteins selected from MMP-12, SP-D, HE4, FUT5, and CRLF1.
[0038] 18. The method according to embodiment 1 or embodiment 5, comprising measuring PSP-94 and FUT5, and at least one of the following proteins selected from MMP-12, SP-D, HE4, PH, and CRLF1.
[0039] 19. The method according to embodiment 1 or embodiment 5, comprising measuring PSP-94 and CRLF1, and at least one of the following proteins selected from MMP-12, SP-D, HE4, PH, and FUT5.
[0040] 20. The method according to embodiment 2 or embodiment 6, comprising measuring PH and FUT5, and at least one protein selected from the following: MMP-12, SP-D, HE4, PSP-94, and CRLF1.
[0041] 21. The method according to embodiment 2 or embodiment 6, comprising measuring PH and CRLF1, and at least one of the following proteins selected from MMP-12, SP-D, HE4, PSP-94, and FUT5.
[0042] 22. A method according to embodiment 3 or embodiment 7, comprising measuring FUT5 and CRLF1, and at least one of the following proteins selected from MMP-12, SP-D, HE4, PSP-94 and PH.
[0043] twenty three. a) contacting a sample from a human subject with two capture reagents, one capture reagent having affinity for PSP-94 protein and a second capture reagent having affinity for PH protein; b) measuring the levels of each protein using the two capture reagents; The method includes:
[0044] twenty four. a) A sample from a human subject is treated with two capture reagents, one of which is PSP-9 a first capture reagent having affinity for the FUT4 protein and a second capture reagent having affinity for the FUT5 protein; b) measuring the levels of each protein using the two capture reagents; The method includes:
[0045] twenty five. a) contacting a sample from a human subject with two capture reagents, one capture reagent having affinity for PSP-94 protein and a second capture reagent having affinity for CRLF1 protein; b) measuring the levels of each protein using the two capture reagents; The method includes:
[0046] 26. a) contacting a sample from a human subject with two capture reagents, one capture reagent having affinity for the PH protein and a second capture reagent having affinity for the FUT5 protein; b) measuring the levels of each protein using the two capture reagents; The method includes:
[0047] 27. a) contacting a sample from a human subject with two capture reagents, one capture reagent having affinity for the PH protein and a second capture reagent having affinity for the CRLF1 protein; b) measuring the levels of each protein using the two capture reagents; The method includes:
[0048] 28. a) contacting a sample from a human subject with two capture reagents, one capture reagent having affinity for FUT5 protein and a second capture reagent having affinity for CRLF1 protein; b) measuring the levels of each protein using the two capture reagents; The method includes:
[0049] 29. a) measuring the levels of PSP-94 and PH in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on said levels of PSP-94 and PH. The method includes:
[0050] 30. a) measuring the levels of PSP-94 and FUT5 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on said levels of PSP-94 and FUT5. The method includes:
[0051] 31. a) measuring the levels of PSP-94 and CRLF1 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on said levels of PSP-94 and CRLF1. The method includes:
[0052] 32. a) measuring the levels of PH and FUT5 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on said levels of PH and FUT5. The method includes:
[0053] 33. a) measuring the levels of PH and CRLF1 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on said levels of PH and CRLF1. The method includes:
[0054] 34. a) measuring the levels of FUT5 and CRLF1 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on said levels of FUT5 and CRLF1. The method includes:
[0055] 35. a) contacting a sample from a human subject with three capture reagents, each of the three capture reagents having an affinity for a protein selected from PSP-94, PH, and FUT5; b) measuring the levels of each protein using the three capture reagents; The method includes:
[0056] 36. a) contacting a sample from a human subject with three capture reagents, each of the three capture reagents having an affinity for a protein selected from PSP-94, PH, and CRLF1; b) measuring the levels of each protein using the three capture reagents; The method includes:
[0057] 37. a) contacting a sample from a human subject with three capture reagents, each of the three capture reagents having an affinity for a protein selected from PH, FUT5, and CRLF1; b) measuring the levels of each protein using the three capture reagents; The method includes:
[0058] 38. a) contacting a sample from a human subject with three capture reagents, each of the three capture reagents having an affinity for a protein selected from FUT5, CRLF1, and PSP-94; b) measuring the levels of each protein using the three capture reagents; The method includes:
[0059] 39. a) measuring the levels of PSP-94, PH, and FUT5 in a sample from a human subject; b) determining whether the human subject has a pulmonary Identifying those at risk for developing cancer The method includes:
[0060] 40. a) measuring the levels of PSP-94, PH, and CRLF1 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on said levels of PSP-94, PH, and CRLF1. The method includes:
[0061] 41. a) measuring the levels of PH, FUT5, and CRLF1 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on said levels of PH, FUT5, and CRLF1. The method includes:
[0062] 42. a) measuring the levels of FUT5, CRLF1, and PSP-94 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on said levels of FUT5, CRLF1, and PSP-94. The method includes:
[0063] 43. A method according to any one of aspects 23 to 42, further comprising measuring the level of MMP-12 protein.
[0064] 44. A method according to any one of aspects 23 to 43, further comprising measuring the level of SP-D protein.
[0065] 45. A method according to any one of aspects 23 to 44, further comprising measuring the level of HE4 protein.
[0066] 46. a) measuring the levels of at least three, four, five, six, or seven proteins selected from the group consisting of MMP-12, SP-D, HE4, PSP-94, PH, FUT5, and CRLF1 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on the levels of said at least three, four, five, six, or seven proteins. The method includes:
[0067] 47. The method comprises the steps of: MMP-12, SP-D, and HE4; MMP-12, SP-D, and PSP-94; MMP-12, SP-D, and PH; MMP-12, SP-D, and FUT5; MMP-12, SP-D, and CRLF1; MMP-12, HE4, and PSP-94; MMP-12, HE4, and PH; MMP-12, HE4, and FUT5; MP-12, HE4, and CRLF1; MMP-12, PSP-94, and PH; MMP-12, PSP-94, and FUT5; MMP -12, PSP-94, and CRLF1;MMP-12, PH, and FUT5;MMP-12, PH, and CRLF1;MMP-12, FUT5, and CRLF1;SP-D, HE4, and PSP-94;SP-D, HE4, and PH;SP-D, HE4, and FUT5;SP-D, HE4, and CRLF1;SP-D, PSP-94, and PH;SP-D, PSP-94, and FUT5;SP-D, PSP-94, and CRLF1;SP-D, PH, and FUT5;SP-D, PH, and 47. The method of embodiment 46, comprising measuring CRLF1; SP-D, FUT5, and CRLF1; HE4, PSP-94, and PH; HE4, PSP-94, and FUT5; HE4, PSP-94, and CRLF1; HE4, PH, and FUT5; HE4, PH, and CRLF1; HE4, FUT5, and CRLF1; PSP-94, PH, and FUT5; PSP-94, PH, and CRLF1; PSP-94, FUT5, and CRLF1; or PH, FUT5, and CRLF1.
[0068] 48. The method of embodiment 46 or 47, further comprising measuring one or more of PSP-94, PH, FUT5, and CRLF1.
[0069] 49. a) contacting a sample from a human subject with a set of capture reagents, each capture reagent having affinity for a different protein of a set of proteins that includes at least 3, 4, 5, 6, or 7 proteins selected from the group consisting of MMP-12, SP-D, HE4, PSP-94, FUT5, and CRLF1 in the sample from the human subject; b) measuring the level of each protein of said set of proteins using said set of capture reagents; The method includes:
[0070] 50. The method comprises the steps of: MMP-12, SP-D, and HE4; MMP-12, SP-D, and PSP-94; MMP-12, SP-D, and PH; MMP-12, SP-D, and FUT5; MMP-12, SP-D, and CRLF1; MMP-12, HE4, and PSP-94; MMP-12, HE4, and PH; MMP-12, HE4, and FUT5; MP-12; HE4, and CRLF1; MMP-12, PSP-94, and PH; MMP-12, PSP-94, and FUT5; MMP-12, PSP-94, and CRLF1; MMP-12, PH, and FUT5; MMP-12, PH, and CRLF1; MMP-12, FUT5, and CRLF1; SP-D, HE4, and PSP-94; SP-D, HE4, and PH; SP- 50. The method of embodiment 49, comprising measuring: D, HE4, and FUT5; SP-D, HE4, and CRLF1; SP-D, PSP-94, and PH; SP-D, PSP-94, and FUT5; SP-D, PSP-94, and CRLF1; SP-D, PH, and FUT5; SP-D, PH, and CRLF1; SP-D, FUT5, and CRLF1; HE4, PSP-94, and PH; HE4, PSP-94, and FUT5; HE4, PSP-94, and CRLF1; HE4, PH, and FUT5; HE4, PH, and CRLF1; HE4, FUT5, and CRLF1; PSP-94, PH, and FUT5; PSP-94, PH, and CRLF1; PSP-94, FUT5, and CRLF1; or PH, FUT5, and CRLF1.
[0071] 51. The method of embodiment 49 or 50, further comprising measuring one or more of PSP-94, PH, FUT5, and CRLF1.
[0072] 52. a) measuring the level of MMP-12 protein and the level of at least one, two, three, four, five, or six proteins selected from the group consisting of SP-D, HE4, PSP-94, PH, FUT5, and CRLF1 in a sample from a human subject; b) identifying the human subject as at risk for developing lung cancer based on the level of MMP-12 and the levels of at least one, two, three, four, five, or six proteins. The method includes:
[0073] 53. a) measuring the level of SP-D protein and the levels of at least one, two, three, four, five, or six proteins selected from the group consisting of MMP-12, HE4, PSP-94, PH, FUT5, and CRLF1 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on the level of SP-D and the level of at least one, two, three, four, five, six, seven, eight, or nine proteins. The method includes:
[0074] 54. a) measuring the level of HE4 protein and the levels of at least one, two, three, four, five, or six proteins selected from the group consisting of MMP-12, SP-D, PSP-94, PH, FUT5, and CRLF1 in a sample from a human subject; b) identifying the human subject as at risk for developing lung cancer based on the level of HE4 and the levels of at least one, two, three, four, five, six, seven, eight, or nine proteins. The method includes:
[0075] 55. a) contacting a sample from a human subject with a set of capture reagents, each capture reagent having affinity for a different protein of a set of proteins comprising MMP-12 protein and at least 1, 2, 3, 4, 5, or 6 proteins selected from the group consisting of SP-D, HE4, PSP-94, PH, FUT5, and CRLF1; b) measuring the level of each protein of said set of proteins using said set of capture reagents; The method includes:
[0076] 56. a) contacting a sample from a human subject with a set of capture reagents, each capture reagent having affinity for a different protein of a set of proteins comprising SP-D protein and at least 1, 2, 3, 4, 5, or 6 proteins selected from the group consisting of MMP-12, HE4, PSP-94, PH, FUT5, and CRLF1; b) measuring the level of each protein of said set of proteins using said set of capture reagents; The method includes:
[0077] 57. a) contacting a sample from a human subject with a set of capture reagents, each capture reagent having affinity for a different protein of a set of proteins comprising HE4 protein and at least 1, 2, 3, 4, 5, or 6 proteins selected from the group consisting of MMP-12, SP-D, PSP-94, PH, FUT5, and CRLF1; b) measuring the level of each protein of said set of proteins using said set of capture reagents; The method includes:
[0078] 58. Any of aspects 5 to 28, 36 to 38, 49 to 51, and 55 to 57, wherein the set of capture reagents is selected from an aptamer, an antibody, and a combination of an aptamer and an antibody. 2. The method according to claim 1.
[0079] 59. The method of any one of the preceding aspects, wherein said measuring is performed using mass spectrometry, an aptamer-based assay, and / or an antibody-based assay.
[0080] 60. The method of embodiment 58 or claim 59, wherein the level of each measured biomarker protein is determined from relative fluorescence units (RFU) or protein concentration.
[0081] 61. The method of any one of the preceding aspects, wherein the sample is selected from blood, plasma, serum, or urine.
[0082] 62. The method of any one of the preceding aspects, wherein said protein levels are used to identify human subjects as at risk for developing lung cancer.
[0083] 63. The method of embodiment 62, wherein the risk of developing lung cancer is within 5 years.
[0084] 64. The method of embodiment 62 or embodiment 63, wherein the risk of developing lung cancer is within 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, or 10 years.
[0085] 65. The method of any one of the preceding aspects, wherein the human subject is a current smoker or a former smoker.
[0086] 66. The method of any one of the preceding aspects, wherein the subject has lung cancer.
[0087] 67. The method of any one of the preceding aspects, wherein the method provides an area under the curve (AUC) of 0.62, 0.67, 0.68, 0.68, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76 or greater.
[0088] 68. The method of any one of the preceding aspects, wherein the method provides an area under the curve (AUC) of about 0.6 to about 0.8, about 0.61 to about 0.78, about 0.62 to about 0.76, about 0.62 to about 0.68, about 0.67 to about 0.72, about 0.69 to about 0.74, about 0.71 to about 0.74, about 0.73 to about 0.76, or about 0.74 to about 0.76.
[0089] 69. The method of any one of the preceding aspects, wherein the prediction of the risk of developing lung cancer is based on input of the levels of said measured protein in a statistical model.
[0090] 70. The method of embodiment 69, wherein the prediction comprises analyzing the measured protein levels using an accelerated failure time (AFT) Weibull survival model.
[0091] 71. The method of any one of the preceding aspects, further comprising performing a diagnostic screening.
[0092] 72. The method of aspect 71, wherein the diagnostic screening is selected from low-dose computed tomography (LDCT), chest x-ray, and sputum cytology.
[0093] 73. The method of any one of the preceding aspects, comprising predicting the risk of developing lung cancer for purposes of determining medical or life insurance premiums.
[0094] 74. The method of embodiment 73, wherein the method further comprises determining health or life insurance coverage or premiums.
[0095] 75. The method of any one of aspects 1-72, wherein the method further comprises using information obtained by the method to predict and / or manage utilization of medical resources.
[0096] 76. The method of any one of aspects 1-72, wherein the method further comprises using information obtained by the method to enable a decision to acquire or purchase a medical business, hospital, or company.
[0097] 77. A kit comprising N protein capture reagents, where N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and at least one of the N protein capture reagents specifically binds to a protein selected from PSP-94, MMP-12, SP-D, HE4, PH, FUT5, and CRLF1.
[0098] 78. The kit of aspect 77, wherein N is at least 2, and at least one of said two N protein capture reagents specifically binds to said protein selected from PSP-94, MMP-12, SP-D, HE4, PH, FUT5, and CRLF1.
[0099] 79. The kit according to embodiment 77 or 78, wherein N is 2-7, or N is 3-7, or N is 4-7, or N is 5-7, or N is 6-7.
[0100] 80. The kit of any one of aspects 77-79, wherein N is 2, N is 3, N is 4, N is 5, N is 6, or N is 7.
[0101] 81. The kit of any one of aspects 77-80, wherein each of the N protein capture reagents specifically binds to a different biomarker protein.
[0102] 82. The kit of any one of aspects 77 to 81, wherein each of the N protein capture reagents specifically binds to a protein selected from PSP-94, MMP-12, SP-D, HE4, PH, FUT5, and CRLF1.
[0103] 83. The kit of any one of aspects 77-81, wherein two of the N protein capture reagents specifically bind to PSP-94 and MMP-12, or two of the N protein capture reagents specifically bind to PSP-94 and SP-D, or two of the N protein capture reagents specifically bind to PSP-94 and HE4, or two of the N protein capture reagents specifically bind to PSP-94 and PH, or two of the N protein capture reagents specifically bind to PSP-94 and FUT5, or two of the N protein capture reagents specifically bind to PSP-94 and CRLF1.
[0104] 84. Three of the N protein capture reagents specifically bind to PSP-94, MMP-12, and SP-D; or three of the N protein capture reagents specifically bind to PSP-94, MMP-12, and HE4; or three of the N protein capture reagents specifically bind to PSP-94, MMP-12, and PH; or three of the N protein capture reagents specifically bind to PSP-94, MMP-12, and FUT5; or three of the N protein capture reagents specifically bind to PSP-94, MMP-12, and CRLF1; or three of the N protein capture reagents specifically bind to PSP-94, SP-D, and HE4; or three of the N protein capture reagents specifically bind to PSP or three of the N protein capture reagents specifically bind to PSP-94, SP-D, and PH; or three of the N protein capture reagents specifically bind to PSP-94, SP-D, and CRLF1; or three of the N protein capture reagents specifically bind to PSP-94, HE4, and PH; or PSP-94, HE4, and FUT5; or 82. The kit of any one of aspects 77 to 81, wherein three of the drugs specifically bind to PSP-94, HE4, and CRLF1, or three of the N protein capture reagents specifically bind to PSP-94, PH, and FUT5, or three of the N protein capture reagents specifically bind to PSP-94, PH, and CRLF1, or three of the N protein capture reagents specifically bind to PSP-94, FUT5, and CRLF1.
[0105] 85. The kit of any one of aspects 77 to 81, wherein two of the N protein capture reagents specifically bind to PH and MMP-12, or two of the N protein capture reagents specifically bind to PH and SP-D, or two of the N protein capture reagents specifically bind to PH and HE4, or two of the N protein capture reagents specifically bind to PH and PSP-94; PH and FUT5, or two of the N protein capture reagents specifically bind to PH and CRLF1.
[0106] 86. Three of the N protein capture reagents specifically bind to PH, MMP-12, and SP-D, or three of the N protein capture reagents specifically bind to PH, MMP-12, and HE4, or three of the N protein capture reagents specifically bind to PH, MMP-12, and PSP-94, or three of the N protein capture reagents specifically bind to PH, MMP-12, and FUT5, or three of the N protein capture reagents specifically bind to PH, MMP-12, and CRLF1, or three of the N protein capture reagents specifically bind to PH, SP-D, and HE4, or three of the N protein capture reagents specifically bind to PH, SP-D, and PSP-94, or three of the N protein capture reagents specifically bind to PH, MMP-12, and FUT5. 82. The kit of any one of aspects 77-81, wherein three of the reagents specifically bind to PH, SP-D, and FUT5, or three of the N protein capture reagents specifically bind to PH, SP-D, and CRLF1, or three of the N protein capture reagents specifically bind to PH, HE4, and PSP-94, or three of the N protein capture reagents specifically bind to PH, HE4, and FUT5, or three of the N protein capture reagents specifically bind to PH, HE4, and CRLF1, or three of the N protein capture reagents specifically bind to PH, PSP-94, and FUT5; PH, PSP-94, and CRLF1, or three of the N protein capture reagents specifically bind to PH, FUT5, and CRLF1.
[0107] 87. The kit of any one of aspects 77-81, wherein two of the N protein capture reagents specifically bind to FUT5 and MMP-12, or two of the N protein capture reagents specifically bind to FUT5 and SP-D, or two of the N protein capture reagents specifically bind to FUT5 and HE4, or two of the N protein capture reagents specifically bind to FUT5 and PSP-94, or two of the N protein capture reagents specifically bind to FUT5 and PH, or two of the N protein capture reagents specifically bind to FUT5 and CRLF1.
[0108] 88. Three of the N protein capture reagents are FUT5, MMP-12, and or three of the N protein capture reagents specifically bind to FUT5, MMP-12, and HE4, or three of the N protein capture reagents specifically bind to FUT5, MMP-12, and PSP-94, or three of the N protein capture reagents specifically bind to FUT5, MMP-12, and FUT5, or three of the N protein capture reagents specifically bind to FUT5, MMP-12, and CRLF1, or three of the N protein capture reagents specifically bind to FUT5, SP-D, and HE4, or three of the N protein capture reagents specifically bind to FUT5, SP-D, and PSP-94, or three of the N protein capture reagents specifically bind to FUT5, SP-D, and PH. or three of the N protein capture reagents specifically bind to FUT5, SP-D, and CRLF1, or three of the N protein capture reagents specifically bind to FUT5, HE4, and PSP-94, or three of the N protein capture reagents specifically bind to FUT5, HE4, and PH, or three of the N protein capture reagents specifically bind to FUT5, HE4, and CRLF1, or three of the N protein capture reagents specifically bind to FUT5, PSP-94, and PH, or three of the N protein capture reagents specifically bind to FUT5, PSP-94, and CRLF1, or three of the N protein capture reagents specifically bind to FUT5, PH, and CRLF1.
[0109] 89. The kit of any one of aspects 77-81, wherein two of the N protein capture reagents specifically bind to CRLF1 and MMP-12, or two of the N protein capture reagents specifically bind to CRLF1 and SP-D, or two of the N protein capture reagents specifically bind to CRLF1 and HE4; or CRLF1 and PSP-94, or two of the N protein capture reagents specifically bind to CRLF1 and PH, or two of the N protein capture reagents specifically bind to CRLF1 and FUT5.
[0110] 90. Three of the N protein capture reagents specifically bind to CRLF1, MMP-12, and SP-D, or three of the N protein capture reagents specifically bind to CRLF1, MMP-12, and HE4, or three of the N protein capture reagents specifically bind to CRLF1, MMP-12, and PSP-94, or three of the N protein capture reagents specifically bind to CRLF1, MMP-12, and PH, or three of the N protein capture reagents specifically bind to CRLF1, MMP-12, and FUT5, or three of the N protein capture reagents specifically bind to CRLF1, SP-D, and HE4, or three of the N protein capture reagents specifically bind to CRLF1, SP-D, and PSP-94, or three of the N protein capture reagents specifically bind to CRLF 82. The kit of any one of aspects 77-81, wherein three of said N protein capture reagents specifically bind to CRLF1, SP-D, and FUT5, or three of said N protein capture reagents specifically bind to CRLF1, HE4, and PSP-94, or three of said N protein capture reagents specifically bind to CRLF1, HE4, and PH, or three of said N protein capture reagents specifically bind to CRLF1, HE4, and FUT5, or three of said N protein capture reagents specifically bind to CRLF1, PSP-94, and PH, or three of said N protein capture reagents specifically bind to CRLF1, PSP-94, and FUT5, or three of said N protein capture reagents specifically bind to CRLF1, PH, and FUT5.
[0111] 91. A kit comprising N protein capture reagents, the kit comprising protein capture reagents for carrying out a method according to any one of claims 1 to 76.
[0112] 92. The kit of any one of aspects 77-91, wherein each of the N biomarker protein capture reagents is an antibody or an aptamer.
[0113] 93. The kit of embodiment 92, wherein each biomarker protein capture reagent is an aptamer.
[0114] 94. A kit according to any one of aspects 77 to 93 for use in detecting the N biomarker proteins in a sample from a subject.
[0115] 95. The kit according to embodiment 94, for use in predicting the risk of developing lung cancer in a subject.
[0116] 96. The kit of aspect 95, wherein the subject is suffering from lung cancer. [Brief description of the drawings]
[0117] [Figure 1] Shows the effect of smoking and smoking cessation on lifetime risk of lung cancer in men and women since 1995. Image from Bruder et al.“Estimating lifetime and 10-year risk of lung cancer.” Prev Med Rep 2018;11:125-30. [Diagram 2] Kaplan-Meier plots of training and validation data using relative risk bins are shown. [Diagram 3] 1 illustrates an exemplary computer system for use with various computer-implemented methods described herein. [Figure 4] 1 is a flowchart of a method for assessing risk of lung cancer, according to one embodiment. [Diagram 5] FIG. 1 shows a line graph of lung cancer risk prediction from Visit 2 to Visit 3, stratified by change in individual smoking behavior over time. [Figure 6]Box plots of lung cancer predictions in individuals with (Y) and without (N) prevalent lung cancer at ARIC visit 3 are shown. [Figure 7] Box plots of lung cancer predictions in individuals with (Y) and without (N) prevalent lung cancer at ARIC visit 5 are shown. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0118] Reference will now be made in detail to exemplary embodiments of the present invention. While the invention will be described in conjunction with the enumerated embodiments, it will be understood that it is not intended that the invention be limited to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents which may be included within the scope of the present invention as defined by the claims.
[0119] One skilled in the art will recognize many methods and materials similar or equivalent to those described herein, which could be used in the practice of the present invention and are within the scope of the present invention, and is in no way limited to the methods and materials described.
[0120] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods, devices, and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods, devices, and materials are now described. Reveal.
[0121] All publications, published patent documents, and patent applications cited in this application are indicative of the level of skill in the technical field(s) to which this application pertains. All publications, published patent documents, and patent applications cited herein are incorporated by reference to the same extent as if each individual publication, published patent document, or patent application was specifically and individually indicated to be incorporated by reference herein.
[0122] As used in this application, including the appended claims, the singular forms "a," "an," and "the" include plural references and are used synonymously with "at least one" and "one or more," unless the context clearly dictates otherwise. Thus, reference to a "SOMAmer" includes mixtures of SOMAmers, reference to a "probe" includes mixtures of probes, and so forth.
[0123] As used herein, the term "about" refers to a small modification or variation of a numerical value that does not change the basic function of the item to which the numerical value is associated.
[0124] As used herein, "backward selection" refers to a method for feature selection and pruning. In certain aspects, backward selection is a form of stepwise regression that starts with all features included in the model. For example, in an iterative process, features are considered for pruning using AUC as a selection criterion.
[0125] As used herein, the words "comprises," "comprising," "includes," "including," "contains," "containing," and any variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product of a process, or composition that comprises, includes, or contains an element or list of elements may include not only those elements, but also other elements that are not expressly listed or that are inherent in such process, method, product of a process, or composition.
[0126] The terms "biological sample", "sample" and "test sample" are used interchangeably herein and refer to any material, biological fluid, tissue or cell obtained from or otherwise derived from an individual. This includes blood (including whole blood, white blood cells, peripheral blood mononuclear cells, buffy coat, plasma and serum), dried blood spots (e.g., from infants), sputum, tears, mucus, nasal washings, nasal aspirates, exhaled breath, urine, semen, saliva, peritoneal washings, ascites, cyst fluid, cerebrospinal fluid, glandular fluid, pancreatic juice, lymphatic fluid, pleural fluid, nipple aspirate, bronchial aspirate, bronchial scraping, synovial fluid, joint aspirate, organ secretions, cells, cell extracts and cerebrospinal fluid. This also includes fractions separated for all of the above experiments. For example, blood samples can be fractionated into serum, plasma or fractions containing specific types of blood cells, e.g., red blood cells or leukocytes (white blood cells). Optionally, the sample may be a combination of samples from an individual, such as a combination of tissue and liquid samples. The term "biological sample" also includes materials containing homogenized solid material, such as from a stool sample, tissue sample, or tissue biopsy. The term "biological sample" also includes materials from tissue culture or cell culture. Any suitable method for obtaining a biological sample may be used, and exemplary methods include, for example, venisection, swab (e.g., buccal swab), and fine needle aspiration biopsy. Exemplary tissues that can be aspirated include lymph node, lung, lung lavage, BAL (bronchoalveolar lavage), thyroid, breast, pancreas, and liver. Samples may also be collected, for example, by microdissection (e.g., laser capture microdissection (LCM) or laser microdissection (LMD)), bladder washing, smear (e.g., PAP smear), or ductal lavage. A "biological sample" obtained from or derived from an individual includes any suitable method that can be used after being obtained from the individual. The present invention includes any such sample that has been treated in a suitable manner.
[0127] It should further be understood that the biological sample may be obtained by taking and pooling biological samples from multiple individuals or by pooling aliquots of the biological sample from each individual.
[0128] As mentioned above, the biological sample can be urine. Urine samples have certain advantages over blood and serum samples. Obtaining blood or plasma samples by venipuncture is more complicated than desired, can be variable in deliverable amounts, can be worrisome for the patient, and carries some (small) risk of infection. Phlebotomy also requires skilled personnel. Due to the simplicity of urine sample collection, the method may be more widely applicable.
[0129] As used herein, "calculation" refers to any kind of mathematical calculation, including arithmetic and non-arithmetic steps.
[0130] For purposes of this specification, the phrase "data resulting from a biological sample from an individual" is intended to mean data in any form derived from or generated using a biological sample from an individual. The data may have been reformatted, modified, or to some extent numerically altered after generation, such as by conversion from units in one measurement system to units in another, but the data is understood to have been obtained from or generated using a biological sample.
[0131] "Target", "target molecule", and "analyte" are used interchangeably herein and refer to any molecule of interest that may be present in a biological sample. "Molecule of interest" includes any minor change in a particular molecule, e.g., in the case of a protein, minor changes in amino acid sequence, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, e.g., conjugation with a labeling moiety, that does not substantially change the identity of the molecule. "Target molecule", "target", or "analyte" is a type or set of copies of a molecule or multi-molecular structure. "Target molecule", "target", and "analyte" refer to a set of such molecules. Exemplary target molecules include proteins, polypeptides, nucleic acids, carbohydrates, lipids, polysaccharides, glycoproteins, hormones, receptors, antigens, antibodies, affibodies, antibody mimetics, viruses, pathogens, toxic substances, substrates, metabolites, transition state analogs, cofactors, inhibitors, drugs, dyes, nutrients, growth factors, cells, tissues, and any fragments or portions of any of the foregoing. In certain embodiments, the "analyte" is a protein target of the capture reagent, e.g., an aptamer. In certain further embodiments, the capture reagent is a SOMAmer.
[0132] As used herein, the terms "polypeptide", "peptide" and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The polymers may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acids. These terms also encompass amino acid polymers that are modified naturally or by intervention, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component. For example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids, etc.), and other modifications known in the art, are also included in this definition. Polypeptides may be single chains or associated chains. This definition also includes preproteins and intact mature proteins, peptides or polypeptides derived from mature proteins, fragments of proteins, splice variants, recombinant forms of proteins, variants of proteins with amino acid modifications, deletions, or substitutions, digests, and post-translational modifications such as glycosylation, acetylation, phosphorylation.
[0133] As used herein, "marker" and "biomarker" and "feature" are used interchangeably and refer to a target molecule that is indicative of or indicative of a normal or abnormal process in an individual, or a disease or other condition in an individual. More specifically, a "marker" or "biomarker" or "feature" is an anatomical, physiological, biochemical, or molecular parameter associated with the presence of a particular physiological state or process, whether normal or abnormal, and if abnormal, whether chronic or acute. Biomarkers can be detected and measured by a variety of methods, including laboratory assays and medical imaging. When a biomarker is a protein, the expression of the corresponding gene can also be used as a surrogate measure of the amount or presence or absence of the corresponding protein biomarker in a biological sample, or the methylation status of the gene that codes for the biomarker or the protein that controls the expression of the biomarker. In certain aspects, the feature is an analyte / SOMAmer reagent of other predictors in a statistical model.
[0134] As used herein, "biomarker value," "value," "biomarker level," "characteristic level," and "level" are used interchangeably and refer to a measurement obtained using any analytical method to detect a biomarker in a biological sample and indicating the presence, absence, absolute amount or concentration, relative amount or concentration, titer, level, expression level, ratio of measured levels, etc. of a biomarker in a biological sample, relating to, or corresponding to, a biomarker in a biological sample. The exact nature of a "value" or "level" will depend on the specific design and components of the particular analytical method used to detect the biomarker.
[0135] When a biomarker indicates or is a sign of an abnormal process, disease, or other condition in an individual, the biomarker is generally described as being either overexpressed or underexpressed compared to the expression level or value of the biomarker that indicates or is a sign of a normal process or the absence of a disease or other condition in an individual. "Upregulation," "upregulated," "overexpression," "overexpressed," and any variations thereof are used interchangeably and refer to a value or level of a biomarker in a biological sample that is higher than the value or level (or range of values or levels) of the biomarker that is normally detected in a similar biological sample from a healthy or normal individual. The term may also refer to a value or level of a biomarker in a biological sample that is higher than the value or level (or range of values or levels) of the biomarker that can be detected at different stages of a particular disease.
[0136] "Downregulation," "downregulated," "underexpression," "underexpressed," and any variations thereof, are used interchangeably and refer to a value or level of a biomarker in a biological sample that is below the value or level (or range of values or levels) of the biomarker that is normally detected in a similar biological sample from a healthy or normal individual. The terms may also refer to a value or level of a biomarker in a biological sample that is below the value or level (or range of values or levels) of the biomarker that can be detected at different stages of a particular disease.
[0137] Furthermore, a biomarker that is either overexpressed or underexpressed may also be referred to as being "differentially expressed" or having a "differential level" or "differential value" compared to a "normal" expression level or value of the biomarker that is indicative of or indicative of a normal process or the absence of a disease or other condition in an individual. Thus, "differential expression" of a biomarker is also referred to as a variation from the "normal" expression level of the biomarker.
[0138] The terms "differential gene expression" and "differential expression" are used interchangeably and refer to the expression of a gene in a subject suffering from a particular disease or condition compared to its expression in normal or control subjects. It refers to a gene (or its corresponding protein expression product) whose expression is activated to a higher or lower level. The term also includes genes (or their corresponding protein expression products) whose expression is activated to a higher or lower level at different stages of the same disease or condition. It is also understood that differentially expressed genes can be activated or inhibited at the nucleic acid level or protein level, or can undergo alternative splicing to produce different polypeptide products. Such differences can be evidenced by a variety of changes, including mRNA levels, surface expression of polypeptides, secretion, or other distribution. Differential gene expression can include a comparison of expression between two or more genes or their gene products, or a comparison of the expression ratio between two or more genes or their gene products, or a comparison of two differently processed products of the same gene that differ between normal subjects and subjects suffering from a disease, or a comparison between different stages of the same disease. Differential expression includes, for example, quantitative and qualitative differences in the temporal or cellular expression patterns of genes or their expression products between normal and diseased cells, or between cells that have undergone different disease events or disease stages.
[0139] As used herein, "individual" refers to a subject or patient. An individual may be a mammal or a non-mammal. In various embodiments, an individual is a mammal. A mammalian individual may be a human or a non-human. In various embodiments, an individual is a human. A healthy or normal individual is an individual in which a disease or condition of interest (including, for example, lung cancer) is not detected by conventional diagnostic methods.
[0140] "Diagnosing", "diagnosis", "diagnosis" and variations thereof refer to detecting, quantifying, or recognizing the health or condition of an individual based on one or more signs, symptoms, data, or other information associated with the individual. An individual's health condition may be diagnosed as healthy / normal (i.e., a diagnosis of the absence of a disease or condition) or diseased / abnormal (i.e., a diagnosis of the presence or characterization of a disease or condition). The terms "diagnosing", "diagnosis", "diagnosis" and the like, with respect to a particular disease or condition, encompass the early detection of the disease, the characterization or classification of the disease, the detection of the progression, remission, or recurrence of the disease, and the detection of disease response following administration of a treatment or therapy to an individual.
[0141] As used herein, "elastic net logistic regression" refers to a machine learning method that utilizes penalized regression techniques to select features that best predict an endpoint while allowing correlated features to be grouped.
[0142] As used herein, a "feature" refers to an analyte or other predictor in a statistical model.
[0143] As used herein, "forward selection" refers to a method for feature selection and reduction. In certain aspects, it is a form of stepwise regression that starts with zero features included in the model. For example, in an iterative process, features are considered for additional consideration using AUC as a selection criterion.
[0144] As used herein, "Mean Absolute Error" or "MAE" refers to the average of the absolute value of the prediction errors across all instances in a dataset.
[0145] As used herein, "normalized root mean square error" or "NRMSE" refers to the standard deviation of the prediction errors (residuals) divided by the mean of the outcome.
[0146] As used herein, "parent study" refers to the external research source for a pilot study.
[0147] As used herein, "prediction error curve" or "Brier score" refers to the difference between the predicted and observed survival times for each individual, with higher values thus representing a worse model.
[0148] As used herein, "population adaptive median normalization" refers to the process of normalizing analytes to reduce site bias and sample handling issues.
[0149] As used herein, "principal component analysis" refers to a method for assessing and identifying significant sources of variability in data.
[0150] As used herein, "root mean square error" or "RMSE" refers to the standard deviation of the prediction errors (residuals).
[0151] As used herein, the term "predict" refers to a prediction about a current or future state or condition. In one embodiment, predicting or making a prediction refers to a prediction about the risk of lung cancer within a specified time period. In one embodiment, the time period is 5 years. In one embodiment, the subject is suffering from lung cancer.
[0152] "Prognose," "prognosing," "prognosis," and variations thereof, refer to the prediction of the future course of a disease or condition in an individual who has the disease or condition (e.g., predicting patient survival), and such terms encompass assessing the response of a disease or condition after administering a treatment or therapy to an individual.
[0153] As used herein, the term "R 2 ” refers to the proportion of variance in the outcome that can be explained by the model.
[0154] "Evaluate," "assessing," "evaluation," and variations thereof encompass both "diagnosis" and "prognosis," and also encompass the determination or prediction of the current or future course of a disease or condition in individuals who may or may not have the disease, and the determination or prediction of the risk of recurrence of a disease or condition in individuals who have apparently been cured of a disease or in remission of a condition. The term "evaluating" also encompasses evaluating an individual's response to a treatment, e.g., determining whether an individual is likely to respond well or unlikely to respond to a therapeutic agent (or, for example, experience toxic or other undesirable side effects), selecting a therapeutic agent to administer to the individual, or monitoring or measuring an individual's response to a treatment administered to the individual.
[0155] As used herein, "additional biomedical information" refers to one or more assessments of an individual other than using any of the biomarkers described herein that relate to the current state of lung health. "Additional biomedical information" includes physical descriptors of the individual, including the individual's height and / or weight, the individual's age, the individual's sex, weight change, the individual's ethnicity, occupational history, family history of lung cancer, the presence of genetic marker(s) that correlate with high risk of lung cancer in the individual, clinical symptoms such as chest pain, weight gain or loss gene expression values, physical descriptors of the individual, including physical descriptors observed by radiological imaging, smoking status, alcohol use history, occupational history, dietary habits - salt, saturated fat, and cholesterol intake, caffeine consumption, and imaging information. Testing of biomarker levels in combination with evaluation of additional biomedical information, including other clinical tests, may improve the sensitivity, specificity, and / or AUC in estimating or determining the current state of lung health, for example, compared to biomarker testing alone or to the sole evaluation of a particular item of additional biomedical information. Additional biomedical information may be used by those skilled in the art. The biomarker level may be obtained from the individual using routine techniques known in the art, such as from the individual themselves using a routine patient or health history questionnaire, or may be obtained from a medical professional, etc. Combining testing of biomarker levels with evaluation of any additional biomedical information may improve the sensitivity, specificity, and / or threshold for estimating or determining current pulmonary health status, as compared, for example, to biomarker testing alone or evaluation of any particular item of additional biomedical information alone (e.g., CT imaging alone).
[0156] As used herein, "detecting" or "determining" with respect to a biomarker value includes the use of both the instrumentation required to observe and record a signal corresponding to the biomarker value as well as the substance / substances required to generate that signal. In various embodiments, the biomarker value is detected using any suitable method, including fluorescence, chemiluminescence, surface plasmon resonance, surface acoustic waves, mass spectrometry, infrared spectroscopy, Raman spectroscopy, atomic force microscopy, scanning tunneling microscopy, electrochemical detection methods, nuclear magnetic resonance, quantum dots, and the like.
[0157] As used herein, a "solid support" refers to any substrate having a surface to which molecules can be directly or indirectly attached, either covalently or non-covalently. A "solid support" can have a variety of physical forms, including, for example, membranes, chips (e.g., protein chips), slides (e.g., glass slides or cover slips), columns, hollow, solid, semi-solid, particles with holes or cavities, such as beads, gels, fibers, including fiber optic materials, matrices, and sample containers. Exemplary sample containers include sample wells, tubes, capillaries, vials, and any other container, groove, or depression that can hold a sample. Sample containers can be mounted on multi-sample platforms, such as microtiter plates, glass slides, microfluidic devices, and the like. Supports can be composed of natural or synthetic materials, organic or inorganic materials. The composition of the solid support to which the capture reagent is attached generally depends on the method of attachment (e.g., covalent attachment). Other exemplary containers include microdroplets, microfluidic controlled, or bulk oil-in-water emulsions in which assays and related operations can be performed. Suitable solid supports include, for example, plastics, resins, polysaccharides, silica or silica-based materials, functionalized glass, modified silicon, carbon, metals, inorganic glass, membranes, nylon, natural fibers (e.g., silk, wool, and cotton), polymers, etc. The material constituting the solid support may contain reactive groups, such as, for example, carboxy, amino, or hydroxyl groups, which are used to bind the capture reagent. Polymeric solid supports include, for example, polystyrene, polyethylene glycol tetraphthalate, polyvinyl acetate, polyvinyl chloride, polyvinylpyrrolidone, polyacrylonitrile, polymethylmethacrylate, polytetrafluoroethylene, butyl rubber, styrene-butadiene rubber, natural rubber, polyethylene, polypropylene, (poly)tetrafluoroethylene, (poly)vinylidene fluoride, polycarbonate, and polymethylpentene. Suitable solid support particles that may be used include, for example, coded particles, such as Luminex® type coded particles, magnetic particles, and glass particles.
[0158] As used herein, "stability selection" refers to methods for feature selection and reduction that use regularization techniques and subsampling approaches such that the Type I error rate is controlled throughout the feature selection process.
[0159] As used herein, "maximum likelihood adaptive normalization" refers to the process of normalizing samples to reduce site bias.
[0160] As used herein, "Lin concordance correlation coefficient" or "Lin's CCC" refers to the concordance correlation coefficient that measures the agreement between a new test and an existing test that is considered the gold standard. means.
[0161] As used herein, a "study" refers to a set of samples and clinical data that are analyzed to derive a test.
[0162] As used herein, "test dataset" refers to the final subset of data used to evaluate the performance of the final model developed in the validation dataset.
[0163] As used herein, "training dataset" means a subset of the data from a study that is used to fit a model.
[0164] As used herein, "validation dataset" refers to the final subset of data used to evaluate the performance of the final model developed in the validation dataset.
[0165] As used herein, a "validation dataset" means a separate subset of data used to provide an unbiased evaluation of a model that fits the training dataset while adjusting model parameters.
[0166] As used herein, the terms "necessary" or "required" refer to a judgment made by a health care provider regarding the treatment of a patient that the health care provider believes to be beneficial to the patient's health status.
[0167] In one embodiment, a lung cancer risk test is disclosed that provides a model to predict a current or former smoker's risk for lung cancer diagnosis within a specific time period, for example, within 5 years after blood draw.
[0168] In certain aspects, the endpoint used for model development is lung cancer diagnosis as determined by electronic health record and cancer registry review.
[0169] In one aspect, the lung cancer risk test was developed using the Atherosclerosis Risk in Communities (ARIC) Visit 3 cohort, which was divided into training (70%), validation (15%), and validation (15%) datasets. The ARIC study was initially intended to longitudinally investigate the contribution of genetic, environmental, and demographic risk factors to atherosclerosis and related cardiovascular disease, but the study's objectives were expanded to also investigate cancer-related outcomes. (Joshu et al. "Enhancing the Infrastructure of the Atherosclerosis Risk in Communities (ARIC) Study for Cancer Epidemiology Research: ARIC Cancer." Cancer Epidemiol Biomarkers Prev 2018;27:295-305).
[0170] In a particular embodiment, the target population for use in this test are adults 50 years or older who are current or former smokers and are eligible for lung cancer screening based on current guidelines. The final model is a 7-feature, protein-only accelerated failure time (AFT) Weibull model. The model output can be reported as absolute risk probability of lung cancer diagnosis within 5 years, or relative risk probability of lung cancer diagnosis within 5 years, compared to the average risk of the "ever smoker" cohort used to develop the model. Relative risks range from 0.010 to 25.
[0171] In one aspect, the minimum performance requirements for this test are to assess future lung cancer risk in current and former smokers. The area under the curve (AUC) was at least equivalent to the published performance of the National Lung Cancer Screening Trial (NLST) for predicting risk of lung cancer (AUC=0.689) (Tammemagi,et al.“Selection criteria for lung-cancer screening.” N Engl J Med 2013;368:728-36). Training and validation results exceeded the performance index of AUC ≥ 0.689 (Table 1). Additional exploratory analyses were also performed to evaluate the performance of the lung cancer risk model in nonsmokers.
[0172] [Table 1]
[0173] In certain embodiments, the intended use of the lung cancer risk test disclosed herein is to predict the risk probability of an individual diagnosed with lung cancer within 5 years from a blood sample. In further embodiments, the test is intended for cancer-free adults 50 years or older who are current or former smokers and are eligible for lung cancer screening based on current guidelines. In certain embodiments, the test is not intended for use in individuals with currently known cancer. The benefits and risks of using RUO are relevant to decision-making in research studies for participant monitoring, stratification, and enrichment. A benefit / risk analysis of clinical LDT use is described below.
[0174] Benefits of the lung cancer risk test disclosed herein include a convenient way to obtain personalized knowledge of the degree of risk for a future lung cancer diagnosis without relying on self-reported demographics or genetic background; the test results may impact compliance with lung cancer screening guidelines allowing for possible early detection of lung cancer and improving lung cancer survival; the test may impact positive behavioral change in modifiable risk-related behaviors (e.g., smoking cessation, dietary changes, weight loss); and the test may assist in the decision / recommendation of lung cancer screening by healthcare providers or the patient's desire for lung screening based on the test results (e.g., when a patient in a high-risk category first undergoes LDCT screening, which is considered the gold standard, instead of an alternative less sensitive method such as chest x-ray).
[0175] The tests disclosed herein may be combined with additional evaluations, including health status assessments, including but not limited to assessment of comorbid conditions such as diabetes, additional laboratory tests, including but not limited to measurements of serum creatinine, urinary albumin, clinical pathology, lung imaging, and histology. It can be used.
[0176] In one aspect, one or more biomarkers are provided for use alone or in various combinations to predict lung cancer risk. As described in more detail below, exemplary embodiments include the biomarkers provided in Table 6, which were identified using multiplex SOMAmer-based assays.
[0177] In a preferred embodiment, the model has seven features (Table 6) for predicting lung cancer risk.
[0178] In one embodiment, the number of biomarkers useful in a biomarker subset or panel is based on the selection of biomarkers that have a non-zero coefficient as a measure of their predictive power for lung cancer risk.
[0179] The risk analysis is shown in Table 2. [Table 2]
[0180] The test of the present disclosure provides a novel and simple method for health care providers to assess and monitor the risk of lung cancer.
[0181] The impact of a false-negative result from this test is moderate based on the risk of potential harm, but these results are comparable to standard treatment. The risk of progression should not preclude or preclude standard treatment (e.g., the method and frequency of lung cancer screening recommended according to current guidelines) and should not be interpreted as a reason to discontinue or reduce standard treatment.
[0182] The impact of a false-positive result from this test is low. A false-positive result may lead health care providers to recommend a more invasive screening method (e.g., LDCT vs. sputum cytology) or lifestyle changes to reduce known risk factors for lung cancer. Although there is a small risk of radiation exposure associated with LDCT screening (although the radiation exposure is several times lower than that of a chest radiograph), risk-benefit analysis studies have concluded that the large reduction in mortality achieved with screening makes this risk acceptable. This test does not need to be the sole source of information for screening decisions.
[0183] In one aspect, the performance of the model was compared with the ability of risk factors that account for lung cancer screening eligibility to accurately predict future risk of lung cancer. The risk factors that determine screening criteria are those used in clinical practice, so the NLST clinical model was selected to reflect the best comparator. The NLST is the largest lung cancer screening clinical trial ever completed and is the basis for the USPSTF lung cancer screening eligibility guidelines. The NLST study included individuals with age and smoking risk factors in line with the existing USPSTF guidelines (age 55 years or older, current or former smokers with a cumulative total of 30 pack years or more (if former smokers have quit smoking within 15 years)). (National Lung Screening Trial Research Team. "Reduced lung-cancer mortality with low-dose computed tomographic screening." N Engl J Med 2011;365:395-409). The NLST risk factor-based baseline model showed AUC performance of 0.689 and 0.670 in the two study arms. (Tammemagi, et al. "Selection criteria for lung-cancer screening." N Engl J Med 2013;368:728-36). Therefore, in developing models for lung cancer risk testing, AUC ≥ 0.689 was used as the performance threshold.
[0184] Thus, the lung cancer risk test disclosed herein met the minimum performance requirements set forth in Table 3. [Table 3]
[0185] In one embodiment, the number of biomarkers useful in a biomarker subset or panel is based on the sensitivity and specificity values for a particular combination of biomarker values. The terms "sensitivity" and "specificity" are used herein with respect to the ability to accurately classify an individual as having an increased risk of lung cancer within 5 years or not having an increased relative risk of lung cancer within the same time period based on one or more biomarker values detected in a biological sample. "Sensitivity" refers to the performance of a biomarker(s) with respect to accurately classifying individuals with an increased risk of lung cancer. "Specificity" refers to the performance of a biomarker(s) with respect to accurately classifying individuals with not an increased relative risk of lung cancer.
[0186] Alternatively, the score may be reported on a continuous range with thresholds of high, intermediate, or low risk of lung cancer, the thresholds being determined based on clinical findings.
[0187] In some embodiments, the overall performance of a panel of one or more biomarkers is represented by an area under the curve (AUC) value. The AUC value is derived from a receiver operating characteristic (ROC) curve. The ROC curve plots the true positive rate (sensitivity) of a test against the false positive rate (specificity) of the test. The term "area under the curve" or "AUC" refers to the area under the curve of a receiver operating characteristic (ROC) curve, both of which are well known in the art. The AUC measurement is useful for comparing the accuracy of classifiers across a range of data. A classifier with a larger AUC is more capable of accurately classifying unknown individuals between two groups of interest (e.g., normal individuals and individuals at risk for lung cancer). The ROC curve is useful for plotting the performance of a particular feature (e.g., any of the biomarkers described herein and / or any item of additional biomedical information) in distinguishing between two populations. Typically, the feature data for the entire population is sorted in ascending order based on the value of a single feature. The true positive rate and false positive rate of the data are then calculated for each value of that feature. The true positive rate is determined by counting the number of cases above the value of the feature and dividing by the total number of cases. The false positive rate is determined by counting the number of controls above the value of the feature and dividing by the total number of controls. Although this definition refers to the scenario where the feature is high compared to the control, this definition also applies to the scenario where the feature is low compared to the control (in such a scenario, one would count samples below the control value for the feature). ROC curves can be generated for single features as well as other single outputs, for example, a combination of two or more features can be mathematically combined (e.g., added, subtracted, multiplied, etc.) to provide a single sum value, which can be plotted on the ROC curve. Additionally, any combination of multiple features that derives a single output value can be plotted on the ROC curve.
[0188] Another factor that may influence the number of biomarkers used in a biomarker subset or panel is the procedure used to collect the biological specimens from the individuals being evaluated for lung cancer risk. In a carefully controlled sample procurement environment, the number of biomarkers required to meet the desired sensitivity and specificity and / or thresholds will be fewer than in situations where there may be greater variability in sample collection, handling, and storage.
[0189] Exemplary Uses of Biomarkers In various exemplary embodiments, methods are provided for estimating or determining lung cancer risk by detecting one or more biomarker values corresponding to one or more biomarkers present in an individual's circulation, such as in serum or plasma, by any number of analytical methods, including any of the analytical methods described herein.
[0190] In addition to testing biomarker levels as an independent diagnostic test, biomarker levels may be tested in conjunction with determining SNPs or other genetic lesions or genetic variability that indicate an increased risk of susceptibility to a disease or condition (see, e.g., Amos et al., Nature Genetics 40, 616-622 (2009)).
[0191] In addition to examining biomarker levels as an independent diagnostic test, biomarker levels can also be used in conjunction with screening methods such as lung imaging techniques, more specifically radiological screening. Biomarker levels may also be used in conjunction with associated symptoms or genetic testing. Detection of any of the biomarkers described herein may be useful in assessing and / or directing appropriate clinical care for an individual, regardless of whether the individual has healthy or unhealthy lung function. Biomarker levels in conjunction with associated symptoms or risk factors can also be used in conjunction with associated symptoms or risk factors. In addition to examining the bell, information regarding the biomarkers may be evaluated along with other types of data, particularly data indicative of the individual's current pulmonary health (e.g., the patient's medical history, symptoms, family history, smoking or drinking history, risk factors, such as the presence of genetic marker(s) and / or the status of other biomarkers, etc.). These various data may be evaluated by automated methods, such as computer programs / software that may be embodied in a computer or other apparatus / device.
[0192] In addition to testing biomarker levels in conjunction with radiological screening in high-risk individuals (e.g., evaluating biomarker levels in conjunction with detection of abnormal lung function), information regarding the biomarkers may be evaluated along with other types of data, particularly data indicative of the individual's pulmonary health (e.g., the patient's medical history, symptoms, family history of lung disease, risk factors such as whether the individual is a smoker, heavy alcohol drinker, and / or other biomarker status, etc.). These various data may be evaluated by automated methods, such as computer programs / software that may be embodied in a computer or other apparatus / device.
[0193] Any of the described biomarkers may also be used in imaging studies. For example, imaging agents can be combined with any of the described biomarkers, which can be used, among other uses, to help determine lung health and the presence or absence of abnormal lung function, to monitor responses to therapeutic interventions, and to select target populations in clinical trials.
[0194] DETECTION AND DETERMINATION OF BIOMARKERS AND BIOMARKER VALUES The biomarker values of the biomarkers described herein can be detected using any of a variety of known analytical methods. In one embodiment, the biomarker values are detected using a capture reagent. As used herein, "capture agent" or "capture reagent" refers to a molecule that can specifically bind to a biomarker. In various embodiments, the capture reagent can be exposed to the biomarker in solution or with the capture reagent immobilized to a solid support. In other embodiments, the capture reagent includes a feature that is reactive to a secondary feature on the solid support. In these embodiments, the capture reagent can be exposed to the biomarker in solution, and then the feature on the capture reagent can be used in combination with the secondary feature on the solid support to immobilize the biomarker on the solid support. The capture reagent is selected based on the type of analysis to be performed. Capture reagents include, but are not limited to, SOMAmers, antibodies, adnectins, ankyrins, other antibody mimetics and other protein scaffolds, autoantibodies, chimeras, small molecules, F(ab')2 fragments, single chain antibody fragments, Fv fragments, single chain Fv fragments, nucleic acids, lectins, ligand binding receptors, affibodies, nanobodies, imprinted polymers, avimers, peptidomimetics, hormone receptors, cytokine receptors, and synthetic receptors, as well as modified versions and fragments thereof.
[0195] In some embodiments, the biomarker level is detected using a biomarker / capture reagent complex.
[0196] In other embodiments, the biomarker value is obtained from the biomarker / capture reagent complex and is indirectly detected, e.g., as a result of a reaction following biomarker / capture reagent interaction, but is dependent on the formation of the biomarker / capture reagent complex.
[0197] In some embodiments, the biomarker value is detected directly from the biomarker in a biological sample.
[0198] In one embodiment, biomarkers are detected using a multiplexing format that allows for simultaneous detection of two or more biomarkers in a biological sample. In one embodiment of the multiplexing format, capture reagents are immobilized directly or indirectly, by covalent or non-covalent attachment, at separate locations on a solid support. In another embodiment, the multiplexing format uses separate solid supports, where each solid support has a unique capture reagent attached to its solid support, e.g., quantum dots. In another embodiment, separate devices are used to detect each of the multiple biomarkers to be detected in a biological sample. The separate devices can be configured to allow each biomarker in a biological sample to be processed simultaneously. For example, a microtiter plate can be used, whereby each well in the plate is used to uniquely analyze one of the multiple biomarkers to be detected in a biological sample.
[0199] In one or more of the foregoing embodiments, a fluorescent tag can be used to label a component of the biomarker / capture complex to detect the biomarker value. In various embodiments, a fluorescent label can be conjugated to a capture reagent specific for any of the biomarkers described herein using known techniques, and the fluorescent label can then be used to detect the corresponding biomarker value. Suitable fluorescent labels include rare earth chelates, fluorescein and its derivatives, rhodamine and its derivatives, dansyl, allophycocyanin, PBXL-3, Qdot 605, Lissamine, phycoerythrin, Texas Red, and other similar compounds.
[0200] In one embodiment, the fluorescent label is a fluorescent dye molecule. In some embodiments, the fluorescent dye molecule comprises at least one substituted indolium ring system, in which the substituent at the carbon at the 3-position of the indolium ring contains a chemically reactive group or a conjugated substance. In some embodiments, the dye molecule comprises an AlexaFluor molecule, such as, for example, AlexaFluor488, AlexaFluor532, AlexaFluor647, AlexaFluor680, or AlexaFluor700. In other embodiments, the dye molecule comprises a first type of dye molecule and a second type of dye molecule, for example, two different Alexafluor molecules. In other embodiments, the dye molecule comprises a first type of dye molecule and a second type of dye molecule, in which the two types of dye molecules have different emission spectra.
[0201] Fluorescence can be measured by a variety of instrumentation adaptable to a wide range of assay formats. For example, spectrofluorometers are designed to analyze microtiter plates, microscope slides, printed arrays, cuvettes, etc. See Principles of Fluorescence Spectroscopy by J.R. Lakowicz, Springer Science + Business Media, Inc., 2004. See Bioluminescence & Chemiluminescence: Progress & Current Applications; Philip E. Stanley Larry J. Kricka editors, World Scientific Publishing Company, January 2002.
[0202] In one or more of the foregoing embodiments, a chemiluminescent tag can optionally be used to label a component of the biomarker / capture complex to detect the biomarker value. Suitable chemiluminescent materials include any of oxalyl chloride, rhodamine 6G, Ru(bipy)32+, TMAE (tetrakis(dimethylamino)ethylene), pyrogallol (1,2,3-trihydroxybenzene), lucigenin, peroxyoxalates, aryloxalates, acridinium esters, dioxetanes, and the like.
[0203] In yet other embodiments, the detection method comprises generating a detectable signal corresponding to the biomarker value. The enzymes include enzyme / substrate combinations that produce chromogenic dyes. Generally, the enzyme catalyzes a chemical change in a chromogenic substrate, which can be measured using a variety of techniques, including spectrophotometry, fluorescence, and chemiluminescence. Suitable enzymes include, for example, luciferase, luciferin, malate dehydrogenase, urease, horseradish peroxidase (HRPO), alkaline phosphatase, β-galactosidase, glucoamylase, lysozyme, glucose oxidase, galactose oxidase, and glucose-6-phosphate dehydrogenase, uricase, xanthine oxidase, lactoperoxidase, microperoxidase, and the like.
[0204] In yet other embodiments, the detection method may be a combination of fluorescent, chemiluminescent, radionuclide, or enzyme / substrate combinations that generate a measurable signal. Multiplexed signal generation may have unique and advantageous features in biomarker assay formats.
[0205] More specifically, the biomarker values of the biomarkers described herein may be detected using known analytical methods, including singleplex SOMAmer assays, multiplex SOMAmer assays, singleplex or multiplex immunoassays, mRNA expression profiling, miRNA expression profiling, mass spectrometry, histological / cytological methods, etc., as detailed below.
[0206] Measuring Biomarker Levels Using SOMAmer-Based Assays Assays aimed at the detection and quantification of physiologically significant molecules in biological and other samples are important tools in scientific research and the health care field. One class of such assays involves the use of microarrays that contain one or more aptamers immobilized on a solid support. Each aptamer can bind to a target molecule in a highly specific manner and with extremely high affinity. See, for example, U.S. Pat. No. 5,475,096, entitled "Nucleic Acid Ligands." See also, for example, U.S. Pat. Nos. 6,242,246, 6,458,543, and 6,503,715, each entitled "Nucleic Acid Ligand Diagnostic Biochip." When the microarray is contacted with a sample, the aptamers bind to the respective target molecules present in the sample, thereby allowing the biomarker value corresponding to the biomarker to be measured.
[0207] As used herein, "aptamer" refers to a nucleic acid that has a specific binding affinity to a target molecule. It is recognized that affinity interactions are a matter of degree, however, in this context, the "specific binding affinity" of an aptamer to its target generally means that the aptamer binds to its target with a much higher degree of binding affinity than it binds to other components in the test sample. An "aptamer" is a set of copies of one type or type of nucleic acid molecule with a specific nucleotide sequence. An aptamer can contain any suitable number of nucleotides, including any number of chemically modified nucleotides. An "aptamer" refers to a set of two or more such molecules. Different aptamers can have either the same or different number of nucleotides. Aptamers can be DNA or RNA or chemically modified nucleic acids, and can be single-stranded, double-stranded, or contain double-stranded regions and can include higher order structures. The aptamer may be a photoaptamer, in which case a photoreactive or chemically reactive functional group is included in the aptamer so that the aptamer can be covalently linked to its corresponding target. Any of the aptamer methods disclosed herein may include the use of two or more aptamers that specifically bind to the same target molecule. As further described below, the aptamer may include a tag. If the aptamer includes a tag, it is not necessary that all copies of the aptamer have the same tag. Furthermore, if different aptamers each include a tag, these different aptamers may be combined to produce a single aptamer. The tamers can have either the same tag or different tags.
[0208] Aptamers can be identified using any known method, including the SELEX process. Once identified, aptamers can be prepared or synthesized according to any known method, including chemical and enzymatic synthesis.
[0209] As used herein, "SOMAmer" or Slow Off-Rate Modified Aptamers refer to aptamers with improved off-rate properties. SOMAmers can be generated using the modified SELEX method described in U.S. Publication No. 2009 / 0004667, entitled "Method for Generating Aptamers with Improved Off-Rates."
[0210] The terms "SELEX" and "SELEX process" are used interchangeably herein and generally refer to the combination of (1) the selection of aptamers that interact with a target molecule in a desired manner, e.g., bind to a protein with high affinity, and (2) the amplification of those selected nucleic acids. The SELEX process can be used to identify aptamers with high affinity for a specific target or biomarker.
[0211] SELEX generally involves preparing a mixture of candidate nucleic acids, binding the candidate mixture to a desired target molecule to form an affinity complex, separating the affinity complex from unbound candidate nucleic acids, separating and isolating the nucleic acids from the affinity complex, purifying the nucleic acids, and identifying a specific aptamer sequence. The process may include multiple rounds to further increase the affinity of the selected aptamers. The process may include an amplification step at one or more points in the process. See, for example, U.S. Patent No. 5,475,096, entitled "Nucleic Acid Ligands." The SELEX process may be used to generate aptamers that bind covalently to a target as well as aptamers that bind non-covalently to a target. See, for example, "Systematic Evolution of Nucleic Acid Ligands by See U.S. Patent No. 5,705,337 entitled "Exponential Enrichment: Chemi-SELEX."
[0212] The SELEX process can be used to identify high affinity aptamers containing modified nucleotides that confer improved properties to the aptamer, such as improved in vivo stability or improved delivery properties. Examples of such modifications include chemical substitutions at the ribose and / or phosphate and / or base positions. Aptamers containing modified nucleotides identified by the SELEX process are described in U.S. Patent No. 5,660,985, entitled "High Affinity Nucleic Acid Ligands Containing Modified Nucleotides," which describes oligonucleotides containing nucleotide derivatives chemically modified at the 5' and 2' positions of the pyrimidines. U.S. Patent No. 5,580,737 (see above) describes highly specific aptamers containing one or more nucleotides modified with 2'-amino (2'-NH2), 2'-fluoro (2'-F), and / or 2'-O-methyl (2'-OMe). See also U.S. Patent Application Publication No. 20090098549, entitled "SELEX and PHOTOSELEX," which describes nucleic acid libraries with enhanced physical and chemical properties and their use in SELEX and photoSELEX.
[0213] SELEX can also be used to identify aptamers with desirable off-rate properties. See U.S. Patent Application Publication No. 20090004667, entitled "Slow Off-Rate Aptamers with Improved Off-Rates." As mentioned above, these slow off-rate aptamers are known as "SOMAmers." A method is described for making aptamers or SOMAmers and photoaptamer SOMAmers with slower off-rates from their respective target molecules. The method includes contacting a candidate mixture with a target molecule, forming a nucleic acid-target complex, and performing a process to enrich for aptamers with slow off-rates, where the nucleic acid-target complexes with fast off-rates dissociate and do not reform, while the complexes with slow off-rates remain intact. Additionally, the method includes using modified nucleotides in generating the candidate nucleic acid mixture to produce aptamers or SOMAmers with improved off-rate performance.
[0214] A variation of this assay uses aptamers that contain photoreactive functional groups that allow the aptamer to covalently bind or photocrosslink to its target molecule. See, e.g., U.S. Patent No. 6,544,776, entitled "Nucleic Acid Ligand Diagnostic Biochip." These photoreactive aptamers are also referred to as photoaptamers. See, e.g., U.S. Patent Nos. 5,763,177, 6,001,577, and 6,291,184, entitled "Systematic Evolution of Nucleic Acid Ligands by Exponential Enrichment: Photoselection of Nucleic Acid Ligands and Solution SELEX," respectively. See also, e.g., U.S. Patent No. 6,458,539, entitled "Photoselection of Nucleic Acid Ligands." After the microarray is contacted with the sample and the photoaptamers are given a chance to bind to their target molecules, the photoaptamers are photoactivated and the solid support is washed to remove any non-specifically bound molecules. Stringent washing conditions may be used since the target molecules bound to the photoaptamers are not usually removed due to the covalent bond generated by the photoactivated functional group(s) on the photoaptamers. In this way, the assay can detect biomarker values corresponding to the biomarkers in the test sample.
[0215] In both of these assay formats, the aptamer or SOMAmer is immobilized on a solid support before contacting with the sample. However, under certain circumstances, immobilizing the aptamer or SOMAmer before contacting with the sample may not provide an optimal assay. For example, pre-immobilizing the aptamer or SOMAmer may result in inefficient mixing of the aptamer or SOMAmer with the target molecule on the solid support surface, possibly prolonging the reaction time and thus extending the incubation time that allows the aptamer or SOMAmer to efficiently bind with its target molecule. Furthermore, when photoaptamers or photoSOMAmers are used in the assay, and depending on the material utilized as the solid support, the solid support may tend to scatter or absorb the light used to affect the formation of covalent bonds between the photoaptamer or photoSOMAmer and its target molecule. Furthermore, depending on the method used, the surface of the solid support may also be exposed to and affected by any labeling agent used, making the detection of the target molecule bound to the aptamer or photoSOMAmer prone to inaccuracies. Finally, immobilization of an aptamer or SOMAmer on a solid support generally involves an aptamer or SOMAmer preparation step (i.e., immobilization) prior to exposure of the aptamer or SOMAmer to a sample, which preparation step may affect the activity or functionality of the aptamer or SOMAmer.
[0216] This allows the SOMAmer to capture its target in solution, and then uses a separation step designed to remove specific components of the SOMAmer-target mixture prior to detection. SOMAmer assays that utilize multiplexed nucleic acid sequences have also been described (see U.S. Patent Application Publication No. 20090042206, entitled "Multiplexed Analyses of Test Samples"). The described SOMAmer assays can detect and quantify non-nucleic acid targets (e.g., protein targets) in a test sample by detecting and quantifying nucleic acids (i.e., SOMAmers). The described methods create nucleic acid surrogates (i.e., SOMAmers) for detecting and quantifying non-nucleic acid targets, thus allowing a wide variety of nucleic acid techniques, including amplification, to be applied to a wider range of desired targets, including protein targets.
[0217] SOMAmers can be constructed to facilitate separation of assay components from the SOMAmer-biomarker complex (or photoSOMAmer-biomarker covalent complex) and allow isolation of the SOMAmer for detection and / or quantification. In one embodiment, these constructs can include a cleavable or releasable element within the SOMAmer sequence. In other embodiments, additional functionality can be introduced into the SOMAmer, such as a label or detectable moiety, a spacer moiety, or a specific binding tag or immobilization element. For example, a SOMAmer can include a tag linked to the SOMAmer via a cleavable moiety, a label, a spacer component separating the labels, and a cleavable moiety. In one embodiment, the cleavable element is a photocleavable linker. The photocleavable linker can be attached to a biotin moiety and a spacer moiety and can include an NHS group for derivatization of an amine and can be used to introduce a biotin group into the SOMAmer, thereby allowing release of the SOMAmer later in the assay method.
[0218] Homogeneous assays, performed with all assay components in solution, do not require separation of samples and reagents prior to signal detection. These methods are rapid and easy to use. These methods generate signals based on molecular capture or binding reagents that react with specific targets. In the case of estimating or determining the risk of lung cancer, the molecular capture reagents can be aptamers (e.g., modified aptamer or SOMAmer reagents) or antibodies, and the specific targets can be biomarkers such as those in Table 6.
[0219] In one embodiment, the signal generation method utilizes the anisotropic signal change resulting from the interaction of a fluorophore-labeled capture reagent with its specific biomarker target. When the labeled capture reagent reacts with its target, the increased molecular weight causes the rotational motion of the fluorophore bound to the complex to become very slow, resulting in a change in anisotropy value. By monitoring the anisotropy change, the binding event can be used to quantitatively measure the biomarker in solution. Other methods include fluorescence polarization assays, molecular beacon methods, time-resolved fluorescence quenching, chemiluminescence, fluorescence resonance energy transfer, etc.
[0220] An exemplary solution-based SOMAmer assay that may be used to detect a biomarker value corresponding to a biomarker in a biological sample includes: (a) preparing a mixture by contacting the biological sample with a SOMAmer that includes a first tag and has a specific affinity for the biomarker (if the biomarker is present in the sample, a SOMAmer affinity complex is formed); (b) exposing the mixture to a first solid support that includes a first capture element, allowing the first tag to associate with the first capture element; (c) removing any components of the mixture that are not associated with the first solid support; and (d) removing the second tag from the first solid support. (e) releasing the SOMAmer affinity complex from the first solid support, (f) exposing the released SOMAmer affinity complex to a second solid support comprising a second capture element, allowing the second tag to associate with the second capture element, (g) removing any uncomplexed SOMAmers from the mixture by separating any uncomplexed SOMAmers from the SOMAmer affinity complex, (h) eluting the SOMAmers from the solid support, and (i) detecting the SOMAmer component of the SOMAmer affinity complex. This includes detecting more biomarkers.
[0221] Any means known in the art can be used to detect the SOMAmer components of the SOMAmer affinity complex to detect the biomarker value. Many different detection methods for detecting the SOMAmer components of the affinity complex can be used, such as hybridization assays, mass spectrometry, or QPCR. In some embodiments, nucleic acid sequencing can be used to detect the SOMAmer components of the SOMAmer affinity complex to detect the biomarker value. In summary, a test sample can be subjected to any type of nucleic acid sequencing method to identify and quantify one or more SOMAmer sequences or sequences present in the test sample. In some embodiments, the sequence includes the entire SOMAmer molecule, or any portion of the molecule that can be used to uniquely identify the molecule. In other embodiments, the identifying sequence is a specific sequence added to the SOMAmer, and such sequences are often referred to as "tags," "barcodes," or "zip codes." In some embodiments, the sequencing method includes an enzymatic step to amplify the SOMAmer sequence or to convert any type of nucleic acid, including RNA and DNA, containing chemical modifications at any position, into any other type of nucleic acid suitable for sequencing.
[0222] In some embodiments, the sequencing method comprises one or more cloning steps, while in other embodiments, the sequencing method comprises a direct sequencing method without cloning.
[0223] In some embodiments, the sequencing method comprises a directed approach using specific primers that target one or more SOMAmers in the test sample. In other embodiments, the sequencing method comprises a shotgun approach that targets all SOMAmers in the test sample.
[0224] In some embodiments, the sequencing method includes an enzymatic step to amplify the molecule targeted for sequencing. In other embodiments, the sequencing method directly sequences a single molecule. An exemplary nucleic acid sequencing-based method that can be used to detect a biomarker value corresponding to a biomarker in a biological sample includes (a) converting a mixture of SOMAmers containing chemically modified nucleotides using an enzymatic step to unmodified nucleic acid, (b) shotgun sequencing the resulting unmodified nucleic acid using a massively parallel sequencing platform, such as 454 Sequencing System (454 Life Sciences / Roche), Illumina Sequencing System (Illumina), ABI SOLiD Sequencing System (Applied Biosystems), HeliScope 1 Molecule Sequencer (Helicos Biosciences) or Pacific BioSciences Real-Time 1 Molecule Sequencing System (Pacific BioSciences), or Polonator G Sequencing Systems (Dover Systems), and (c) identifying and quantifying the SOMAmers present in the mixture by specific sequences and sequence counts.
[0225] Measuring biomarker levels using immunoassays Immunoassay methods are based on the reaction of antibodies to corresponding targets or analytes, and can detect the analytes in a sample depending on the specific assay format. To improve the specificity and sensitivity of immunoreactivity-based assay methods, monoclonal antibodies are frequently used due to their specific epitope recognition. Polyclonal antibodies have also been successfully used in various immunoassays due to their higher affinity for targets compared to monoclonal antibodies. Immunoassays are designed for use with a wide range of biological sample matrices. Immunoassay formats are designed to provide qualitative, semi-quantitative, and quantitative results.
[0226] Quantitative results are obtained by using a standard curve prepared with known concentrations of the specific analyte to be detected. The response or signal from an unknown sample is plotted against the standard curve to determine the amount or value corresponding to the target in the unknown sample.
[0227] Numerous immunoassay formats have been designed. ELISA or EIA can be quantitative for the detection of the analyte. The method is based on the binding of a label to either the analyte or the antibody, the label component comprising an enzyme, either directly or indirectly. ELISA tests can be in a format for direct, indirect, competitive or sandwich detection of the analyte. Other methods are based on labels such as, for example, radioisotopes (I125) or fluorescence. Additional techniques include, for example, agglutination, nephelometry, turbidimetry, Western blot, immunoprecipitation, immunocytostaining, immunohistostaining, flow cytometry, Luminex assays, etc. (ImmunoAssay: A Practical (See the Law Guide, edited by Brian Law, published by Taylor & Francis, Ltd., 2005 edition).
[0228] Exemplary assay formats include enzyme-linked immunosorbent assays (ELISAs), radioimmunoassays, fluorescence, chemiluminescence, and fluorescence resonance energy transfer (FRET) or time-resolved FRET (TR-FRET) immunoassays. Exemplary techniques for detecting biomarkers include biomarker immunoprecipitation followed by quantitative methods that allow size and peptide level discrimination, such as gel electrophoresis, capillary electrophoresis, planar electrochromatography, and the like.
[0229] The method of detecting and / or quantifying the detectable label or signal generating substance depends on the nature of the label. The products of the reaction catalyzed by the appropriate enzyme (wherein the detectable label is an enzyme; see above) can be, but are not limited to, fluorescent, luminescent, or radioactive, or they can absorb visible or ultraviolet light. Examples of detectors suitable for detecting such detectable labels include, but are not limited to, X-ray film, radioactivity counters, scintillation counters, spectrophotometers, colorimeters, fluorometers, luminometers, and densitometers.
[0230] Any detection method can be carried out in any format that allows any suitable preparation, processing, and analysis of the reaction. Detection methods can be carried out, for example, in multi-well assay plates (e.g., 96-well or 384-well) or using any suitable array or microarray. Stock solutions of various agents can be made manually or robotically, and all subsequent pipetting, dilution, mixing, dispensing, washing, incubation, sample reading, data collection, and analysis can be carried out robotically using commercially available analysis software, robots, and detection equipment that can detect detectable labels.
[0231] Measuring Biomarker Levels Using Gene Expression Profiling Measurement of mRNA in a biological sample may be used as a proxy to detect the level of the corresponding protein in the biological sample. Thus, any of the biomarkers or biomarker panels described herein can be detected by detecting the appropriate RNA.
[0232] The mRNA expression levels are measured by reverse transcription quantitative polymerase chain reaction (RT-PCR followed by qPCR). RT-PCR is used to generate cDNA from the mRNA. The cDNA may be used in a qPCR assay to generate fluorescence as the DNA amplification process proceeds. qPCR determines the amount of mRNA per cell by comparison to a standard curve. Absolute measurements such as copy number of a gene can be obtained. Northern blots, microarrays, Invader assays, and RT-PCR in combination with capillary electrophoresis have all been used to measure mRNA expression levels in a sample. See Gene Expression Profiling: Methods and Protocols, Richard A. Shimkets, editor, Humana Press, 2004.
[0233] miRNA molecules are small RNAs that are non-coding but can regulate gene expression. Any method suitable for measuring mRNA expression levels can be used for the corresponding miRNA. Recently, many laboratories have been investigating the use of miRNAs as biomarkers for disease. Many diseases involve extensive transcriptional regulation, so it is not surprising that miRNAs may serve as biomarkers. The relationship between miRNA concentrations and disease is often not as clear as the relationship between protein levels and disease, but there is still considerable value in miRNA biomarkers. Of course, as with other RNAs that are differentially expressed during disease, the challenges facing the development of in vitro diagnostic products include requirements such as miRNAs surviving in diseased cells and being easily extractable for analysis, or released into blood or other matrices where they must persist for a sufficient period of time to be measured. Protein biomarkers have similar requirements, although many potential protein biomarkers are purposely secreted at the site of pathology and function in a paracrine manner during disease. Many potential protein biomarkers are designed to function outside of the cells in which their proteins are synthesized.
[0234] Detection of biomarkers using in vivo molecular imaging techniques Any of the described biomarkers (see Table 6) may also be used in imaging studies. For example, imaging agents can be combined with any of the described biomarkers and used to help estimate the risk of lung cancer, monitor responses to therapeutic interventions, and select populations for particular clinical trials.
[0235] In vivo imaging techniques provide a non-invasive method for determining the state of a particular disease or condition within an individual's body. For example, all parts of the body or the entire body may be displayed as a three-dimensional image, thereby providing useful information regarding the morphology and structure of the body. Such techniques may be combined with the detection of biomarkers as described herein to provide information regarding an individual's risk of lung cancer.
[0236] Various technological advances have led to the development of in vivo molecular imaging techniques. These advances include the development of new contrast agents or labels, such as radiolabels and / or fluorescent labels, that can generate strong signals within the body, as well as powerful new imaging techniques that can detect and analyze these signals from outside the body with sufficient sensitivity and accuracy to provide useful information. The contrast agents can be visualized with a suitable imaging system, thereby providing an image of the part or parts of the body in which they are present. The contrast agents can be bound to or associated with, for example, capture agents such as SOMAmers or antibodies, and / or peptides or proteins or oligonucleotides (e.g., for detecting gene expression), or complexes that include any of these with one or more macromolecules and / or other particulate forms.
[0237] Contrast agents may feature radioactive atoms useful in imaging. Radioactive atoms suitable for scintigraphic studies include technetium 99m or iodine 123. Other readily detectable moieties are, for example, spin labels for magnetic resonance imaging (MRI), such as iodine 123, iodine 131, indium 111, fluorine 19, carbon 13, nitrogen 15, oxygen 17, gadolinium, manganese, or iodine 123. or iron. Such labels are well known in the art and can be readily selected by one of ordinary skill in the art.
[0238] Standard imaging techniques include, but are not limited to, magnetic resonance imaging, computed tomography (coronary calcium score), positron emission tomography (PET), single photon emission computed tomography (SPECT), computed tomography angiography, etc. For in vivo diagnostic imaging, the type of detection instrument available is an important factor in the selection of a given contrast agent, e.g., a given radionuclide and the particular biomarker (protein, mRNA, etc.) that is to be targeted with it. The radionuclide typically chosen exhibits some kind of decay that is detectable by a given type of instrument. Also, in When selecting a radionuclide for in vivo diagnosis, its half-life should be long enough to permit detection upon maximal uptake by the target tissue, yet short enough to minimize harmful radiation to the host.
[0239] Exemplary imaging techniques include, but are not limited to, PET and SPECT, which are imaging techniques in which radionuclides are administered synthetically or locally to an individual. Subsequent radioactive tracer uptake is measured over time and used to obtain information about the target tissue and biomarkers. Depending on the high-energy (gamma-ray) emission of the particular isotope used, and the sensitivity and sophistication of the instrument used to detect it, the two-dimensional distribution of radioactivity can be estimated from outside the body.
[0240] Commonly used positron-emitting nuclides in PET include, for example, carbon-11, nitrogen-13, oxygen-15, and fluorine-18. SPECT uses isotopes that decay by electron capture and / or gamma emission, including, for example, iodine-123 and technetium-99m. An exemplary method for labeling amino acids with technetium-99m is the reduction of pertechnetate ions in the presence of a chelating precursor to form an unstable technetium-99m-precursor complex, which then reacts with the metal binding group of a bifunctionally modified chemotactic peptide to form a technetium-99m-chemotactic peptide conjugate.
[0241] In such in vivo imaging diagnostic methods, antibodies are frequently used. Preparation and use of antibodies for in vivo diagnosis is well known in the art. For the purpose of diagnosing or assessing the disease risk of an individual, a labeled antibody that specifically binds to any of the biomarkers in Table 6 can be injected into the individual to be assessed for lung cancer risk detectable by the particular biomarker used. The label used is selected according to the imaging modality used, as described above. The localization of the label can determine tissue damage or other indications related to lung cancer. The amount of label in an organ or tissue can also determine the involvement of lung cancer biomarkers in that organ or tissue.
[0242] Similarly, SOMAmers may be used in such in vivo imaging methods. For example, SOMAmers used to identify (and therefore specifically bind to) a particular biomarker listed in Table 6 may be appropriately labeled and injected into an individual to be evaluated for lung cancer risk detectable by the particular biomarker, for the purpose of diagnosing or assessing the level of tissue damage, components of the inflammatory response, and other factors associated with lung cancer risk in an individual. The label used is selected according to the imaging modality to be used, as described above. The localization of the label allows the determination of the site of the process leading to an increased risk. The amount of label in an organ or tissue may also determine the infiltration of the pathological process in that organ or tissue. SOMAmer-directed imaging agents may have unique and advantageous properties compared to other imaging agents in terms of tissue penetration, biodistribution, kinetics, clearance, potency, and selectivity.
[0243] Such techniques may be carried out using optionally labeled oligonucleotides, for example, to detect gene expression by imaging with antisense oligonucleotides.These methods are used, for example, in situ hybridization, using fluorescent molecules or radionuclides as labels.Other methods for detecting gene expression include, for example, detecting the activity of reporter genes.
[0244] Another common type of imaging technique is optical imaging, in which fluorescent signals within a subject are detected by optical devices external to the subject. These signals can be due to actual fluorescence and / or bioluminescence. Improvements in the sensitivity of optical detection devices have increased the usefulness of optical imaging for in vivo diagnostic assays.
[0245] For example, in vivo molecular biomarker imaging is increasingly being used, including in clinical trials, to measure clinical efficacy more rapidly in clinical trials for the treatment of new diseases or conditions and / or to avoid long-term placebo treatment for these diseases, e.g., multiple sclerosis, which may be considered ethically questionable.
[0246] For a review of other techniques, see N. Blow, Nature Methods, 6, 465-469, 2009.
[0247] Measuring Biomarker Levels Using Mass Spectrometry Mass spectrometers of various configurations can be used to detect biomarker values. Several types of mass spectrometers are available or can be manufactured in various configurations. Generally, mass spectrometers have the following main components: sample inlet, ion source, mass analyzer, detector, vacuum system, and instrument control system, and data system. The differences in the sample inlet, ion source, and mass analyzer generally define the type of instrument and its capabilities. For example, the inlet can be a capillary column liquid chromatography source, or a direct probe or stage such as used in matrix-assisted laser desorption. Common ion sources are, for example, electrospray, including nanospray and microspray, or matrix-assisted laser desorption. Common mass analyzers include quadrupole mass filters, ion trap mass analyzers, and time-of-flight mass analyzers. Additional mass spectrometry methods are well known in the art (see Burlingame et al., Anal. Chem. 70:647R-716R (1998); Kinter and Sherman, New York (2000)).
[0248] Protein biomarkers and biomarker values may be detected and measured by any of the following: electrospray ionization mass spectrometry (ESI-MS), ESI-MS / MS, ESI-MS / (MS)n, matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF-MS), surface-enhanced laser desorption / ionization time-of-flight mass spectrometry (SELDI-TOF-MS), silicon-assisted desorption / ionization (DIOS), secondary ion mass spectrometry (SIMS), quadrupole time-of-flight (Q-TOF), a tandem time-of-flight (TOF / TOF) technology called UltraFlex III TOF / TOF, atmospheric pressure chemical ionization mass spectrometry (APCI-MS), APCI-MS / MS, APCI-(MS)N, atmospheric pressure photoionization mass spectrometry (APPI-MS), APPI-MS / MS, and APPI-(MS)N, quadrupole mass spectrometry, Fourier transform mass spectrometry (FTMS), quantitative mass spectrometry, and ion trap mass spectrometry.
[0249] Sample preparation strategies are used to label and enrich samples before characterizing protein biomarkers and measuring biomarker values by mass spectrometry. Labeling methods include, but are not limited to, iso-mass tags for relative or absolute quantification (iTRAQ) and stable isotope labeling with amino acids in cell culture (SILAC). Capture reagents used to selectively enrich potential biomarker proteins in samples prior to mass spectrometry analysis include, but are not limited to, SOMAmers, antibodies, nucleic acid probes, chimeras, small molecules, F(ab')2 fragments, single-chain antibody fragments, Fv fragments, single-chain Fv fragments, nucleic acids, lectins, ligand-binding receptors, affibodies, nanobodies, ankyrin, domain antibodies, alternative antibody scaffolds (e.g., diabodies, etc.), imprinted polymers, avimers, peptidomimetics, peptoids, peptide nucleic acids, threose nucleic acids, hormone receptors, cytokine receptors, and synthetic receptors, as well as modified versions and fragments thereof.
[0250] Measuring biomarker levels using proximity ligation assays Proximity ligation assays can be used to measure biomarker values. In summary, a test sample is contacted with a pair of affinity probes, which can be a pair of antibodies or a pair of SOMAmers, with each member of the pair extended with an oligonucleotide. The targets of a pair of affinity probes can be two different determinants on one protein, or one determinant each on two different proteins that can exist as homo- or heteromultimeric complexes. When the probes bind to the target determinants, the free ends of the oligonucleotide extensions are brought close enough together to hybridize together. Hybridization of the oligonucleotide extensions is facilitated by a common connector oligonucleotide that serves to bridge the oligonucleotide extensions together if they are positioned close enough together. Once the oligonucleotide extensions of the probes are hybridized, the ends of the extensions are joined together by enzymatic DNA ligation.
[0251] Each oligonucleotide extension contains a primer site for PCR amplification. When the oligonucleotide extensions are ligated together, the oligonucleotides form a continuous DNA sequence, and through PCR amplification, information about the identity and amount of the target protein, as well as information about protein-protein interactions when the target determinants are on two different proteins, is revealed. Proximity ligation can provide a highly sensitive and specific assay for real-time protein concentration and interaction information by using real-time PCR. Probes that do not bind to the determinants of interest will not bring the corresponding oligonucleotide extensions into proximity, and ligation or PCR amplification cannot proceed, resulting in no signal generation.
[0252] The aforementioned assays can detect biomarker values useful in methods for determining or estimating lung cancer risk, comprising detecting biomarker values in a biological sample from an individual, each corresponding to a biomarker selected from the group of biomarkers shown in Table 6, and using the biomarker values to assess the risk of lung cancer in the individual, as described in detail below. Although some of the described lung cancer risk biomarkers are useful alone for estimating or determining lung cancer risk, methods are also described herein for grouping multiple subsets of lung cancer risk biomarkers, each of which is useful as a panel of three or more biomarkers. According to any of the methods described herein, biomarker values can be detected and evaluated individually, or can be detected and evaluated together, such as in a multiplex assay format.
[0253] A biomarker "signature" of a given diagnostic or predictive test comprises a set of markers, each of which has different levels in a population of interest. In this regard, different levels can refer to different means of marker levels for individuals in two or more groups, different levels of markers for individuals in two or more groups, or different levels of markers for individuals in two or more groups. It may refer to a different variance in, or a combination of both. In the simplest form of diagnostic testing, markers can be used to assign unknown samples from individuals to one of two groups: with or without lung cancer risk. Assigning a sample to one of two or more groups is known as classification, and the techniques used to achieve this assignment are known as classifiers or classification methods. Classification methods may also be referred to as scoring methods. There are numerous classification methods that can be used to build diagnostic classifiers from a set of biomarker values. In general, classification methods are most easily performed using supervised learning techniques, where a data set is collected using samples taken from individuals in the two (or more in the case of multiple classification situations) separate groups that one wishes to distinguish. Since the class (group or population) to which each sample belongs is known in advance for each sample, the classifier can be trained to produce the desired classification response. It is also possible to use unsupervised learning techniques to generate diagnostic classifiers.
[0254] Common techniques for developing diagnostic classifiers include decision trees; bagging, boosting, forests, and random forests; inference rule-based learning; Parzen windows; linear models; logistic curves; neural network methods; unsupervised clustering; K-means; hierarchical ascending / descending classification; semi-supervised learning; prototype methods; nearest neighbor methods; kernel density estimation; support vector machines; hidden Markov models; Boltzmann learning, and classifiers may be combined simply or in a manner that minimizes a particular objective function. For a general discussion, see, e.g., Pattern Classification, R.O.Duda, et al., editors, John Wiley & Sons, 2nd edition, 2001. See also The Elements of Statistical Learning - Data Mining, Inference, and Prediction, T.Hastie, et al., editors, Springer Science+Business Media, LLC, 2nd edition, 2009 (each of which is incorporated by reference in its entirety).
[0255] To generate a classifier using supervised learning techniques, a set of samples, called training data, is obtained. For diagnostic testing, the training data includes samples from different groups (classes) to which unknown samples will later be assigned. For example, samples collected from individuals in a control population and samples collected from individuals in a population with a particular disease, condition, or event may constitute training data for developing a classifier that can classify unknown samples (or, more specifically, the individuals from whom the samples were taken) as either diseased or disease-free. The development of a classifier from training data is known as training the classifier. The specific details of classifier training depend on the nature of the supervised learning approach (see, for example, Pattern Classification, RODuda, et al., "Pattern Classification," in J. Med. Soc. 1999, 143:131-135, 2001). al., editors, John Wiley & Sons, 2nd edition, 2001; see also The Elements of Statistical Learning - Data Mining, Inference, and Prediction, T. Hastie, et al., editors, Springer Science+Business Media, LLC, 2nd edition, 2009).
[0256] Usually, there may be many higher biomarker values compared to the samples of training set, so care must be taken to avoid overfitting.Overfitting occurs when statistical model represents random error or noise instead of the underlying relationship.Overfitting can be avoided in various ways, including, for example, limiting the number of markers used during classifier development, assuming that the responses of markers are independent of each other, limiting the complexity of the underlying statistical model used, and ensuring that the underlying statistical model fits data.
[0257] To identify a set of biomarkers associated with the occurrence of an event, the combined set of control and early event samples was analyzed using principal component analysis (PCA). PCA presents samples against the axis defined by the strongest variation among all samples, regardless of case or control outcome, thus reducing the risk of overfitting the case-control distinction. Because the development of lung cancer has a strong chance component, one would not expect to see a clear separation between the control and event sample sets. Although the observed separation between cases and controls is not large, it occurs in the second principal component and represents approximately 10% of the total variation in this sample set, indicating that the underlying biological variation is relatively easy to quantify.
[0258] In the next series of analyses, the biomarkers can be analyzed for components of inter-sample differences specific to the separation between control and early event samples. One method that can be used is to use DSGA to predict patient survival from gene expression data (Bair, E. and Tibshirani, R. (2004) Semi-supervised methods to predict patient survival from gene expression data. PLOS Biol., 2, 511-522) to remove (shrink) the first three principal component directions of variation between samples in the control set. Dimensionality reduction is performed on the control set to detect, but through PCA on both samples in the controls and samples from the early event samples. Separation of cases from the early event can be observed along the horizontal axis. Cross-validation of a selection of proteins related to lung cancer risk estimation or determination
[0259] To avoid overfitting the predictive power of proteins to the idiosyncratic features of a particular sample selection, cross-validation and dimensionality reduction approaches can be employed. Cross-validation involves the combination of multiple selections of sample sets to determine the association of risk with proteins, and the use of unselected samples to monitor the ability of the method to be applied to samples that were not used to create the risk model (The Elements of Statistical Learning - Data Mining, Inference, and Prediction, T. Hastie, et al., editors, Springer Science+Business Media, LLC, 2nd edition, 2009). The supervised PCA method of Tibshirani et al. (Bair, E. and Tibshirani, R. (2004) Semi-supervised methods to predict patient survival from gene expression data. PLOS Biol.,2,511-522.), can be applied to high-dimensional datasets in modeling lung cancer risk. The Supervised PCA (SPCA) method involves the univariate selection of a set of proteins that are statistically associated with the event hazard observed in the data and the determination of correlated components that combine information from all these proteins. This correlation component determination is a dimensionality reduction step that not only combines information across proteins, but also mitigates the possibility of overfitting by reducing the number of independent variables from a full protein menu of over 1000 proteins down to a small number of principal components (in this study, only the first principal component was examined).
[0260] Univariate and multivariate analyses of the relationship between individual proteins and time to event The Cox proportional hazard model (Cox, David R (1972). “Regression Models and Life-Tables”. Journal of the Royal Statistical Society. Series B(Methodological)34(2):187-220.) is widely used in medical statistics. Cox regression avoids fitting a specific time function to the cumulative survival rate and instead uses a baseline hazard. We employ relative risk models that refer to a baseline hazard function (which can change over time). The baseline hazard function describes the common shape of the survival time distribution for all individuals, while the relative risk indicates the level of hazard for a set of covariate values (such as a single individual or group) as a multiple of the baseline hazard. In the Cox model, the relative risk is constant over time.
[0261] Accelerated failure time (AFT) models are a subclass of survival models. Survival models predict time-to-event data based on partial information. For example, in the lung cancer model data, the event is a diagnosis of lung cancer, but time-to-diagnosis event data is available for a portion of the subjects in the study. The information available for the remaining subjects is that they were not diagnosed with lung cancer from the time of blood collection to the end of the study. This second category is called "censored" because of the partial information and uncertainty of whether or when lung cancer will be diagnosed.
[0262] Survival models take into account censoring and can still use data from censored subjects, whereas other longitudinal models that attempt to predict when an event will occur can only use information from subjects diagnosed with lung cancer. Also, because survival models take into account time to event, they can generate predicted probabilities of an event occurring within any time window, which differs from most classification models (logistic regression, random forests).
[0263] In particular, the AFT survival model is a regression model that specifies / assumes a linear relationship between the model covariates and the log (time to event). Thus, a subject whose covariate (protein RFU count) is 2-fold higher than baseline may be predicted to "survive" twice as long from a lung cancer diagnosis as baseline.
[0264] The two most common survival models are the AFT model and the proportional hazards model, with the AFT-Weibull model being both. The definition of the proportional hazards model is a bit more complicated than that of the AFT model, in which a subject with a covariate twice as high as baseline will have a twice as high hazard at any given time point, and the hazard is the negative derivative of the time course of the survival curve.
[0265] Other common proportional hazards models are the exponential and Cox models. The exponential model is a subtype of the Weibull model. The Cox model is more limited in use. The Cox model does not provide predicted probabilities of time to event, only relative risks. The AFT model can give both absolute and relative risks.
[0266] kit For example, a kit suitable for use in practicing the methods disclosed herein can be used to detect any combination of the biomarkers in Table 6. Additionally, any kit can include one or more detectable labels as described herein, such as, for example, fluorescent moieties.
[0267] In one embodiment, the kit comprises (a) one or more capture reagents (e.g., at least one SOMAmer or antibody, etc.) for detecting one or more biomarkers in a biological sample, the biomarkers including any of the biomarkers listed in Table 6, and optionally (b) one or more software or computer program products for calculating lung cancer risk. Alternatively, rather than one or more computer program products, one or more instructions for a person to manually perform the above steps may be provided.
[0268] The combination of a solid support and a corresponding capture reagent having a signal generating substance is described herein. In the present specification, the biological sample is referred to as a "detection device" or "kit." The kit may also include instructions for use of the device and reagents, handling of samples, and analysis of data. Additionally, the kit may be used with a computer system or software for analyzing biological samples and reporting the results of the analysis.
[0269] The kit may also include one or more reagents for processing the biological sample (e.g., a solubilization buffer, a detergent, a washing solution, or a buffer). Any of the kits described herein may also include, for example, buffers, blocking agents, mass spectrometry matrix materials, antibody capture agents, positive control samples, negative control samples, software, and information, such as protocols, guidelines, and reference data.
[0270] In one aspect, the present invention provides a kit for analyzing the risk of lung cancer. The kit includes one or more SOMAmer PCR primers specific to a biomarker selected from Table 6. The kit may further include instructions for use and instructions for correlating the biomarker with estimating or determining lung cancer risk. The kit may also include a DNA array containing one or more aptamer or SOMAmer reagent complements specific to a biomarker selected from Table 6, reagents, and / or enzymes for amplifying or isolating sample DNA. The kit may include reagents for real-time PCR, such as TaqMan probes and / or primers, and enzymes.
[0271] For example, the kit may include (a) reagents including at least a capture reagent for quantifying one or more biomarkers in a test sample, the biomarkers including those listed in Table 6, or any other biomarker or panel of biomarkers described herein, and, optionally, (b) one or more algorithms or computer programs for performing the steps of: comparing the amount of each quantified biomarker in the test sample to one or more predefined cutoffs, assigning a score for each quantified biomarker based on the comparison, combining the scores assigned to each quantified biomarker to obtain a total score, comparing the total score to a predefined score, and using the comparison to determine whether the individual is at risk for lung cancer. Alternatively, rather than one or more algorithms or computer programs, one or more instructions for use may be provided for a human to manually perform the above steps.
[0272] Computer Methods and Software Once a biomarker or panel of biomarkers has been selected, a method for diagnosing an individual may include: 1) collecting or obtaining a biological sample; 2) performing an analytical method to detect and measure one or more biomarkers in the panel in the biological sample; 3) performing data normalization or standardization required for the method used to collect biomarker values; 4) calculating marker scores; 5) combining the marker scores to obtain a total diagnostic or predictive score; and 6) reporting the individual's diagnostic or predictive score. In this approach, the diagnostic or predictive score may be a single number determined from the sum of all marker calculations compared to a pre-set threshold indicating the presence or absence of disease. Alternatively, the diagnostic or predictive score may be a series of bars, each representing a biomarker value, and the response pattern may be compared to a pre-set pattern for determining the presence or absence of elevated (or not elevated) risk of a disease, condition, or event.
[0273] At least some embodiments of the methods described herein may be implemented using a computer. An example of a computer system 100 is shown in FIG. 3. Referring to FIG. 3, the system 100 includes a baud rate controller including a processor 101, an input device 102, an output device 103, a storage device 104, a computer-readable storage medium reader 105a, a communication system 106, an accelerated processing device (e.g., a DSP or special purpose processor) 107, and a memory 109. The system 100 is shown comprised of hardware elements electrically connected via a computer readable storage medium reader 105a, which is further connected to a computer readable storage medium 105b, which combination collectively represents a storage medium, memory, etc., in addition to remote, local, fixed, and / or removable storage devices for containing computer readable information on a temporary and / or more persistent basis, which combination includes storage device 104, memory 109, and / or any other such accessible system 100 resources. The system 100 also includes software elements (shown here as residing in working memory 191) including an operating system 192, and other code 193, e.g., programs, data, and the like.
[0274] Referring to FIG. 4, the system 100 has a wide range of flexibility and configurability. Thus, for example, a single architecture may be utilized to implement one or more servers, which may further be configured according to generally desired protocols, protocol modifications, extensions, and the like. However, it will be apparent to one skilled in the art that embodiments may be utilized according to more specific application requirements. For example, one or more system elements may be implemented as sub-elements within the components of the system 100 (e.g., within the communications system 106). Custom hardware may also be utilized, and / or specific elements may be implemented in hardware, software, or both. Additionally, connections to other computing devices, such as network input / output devices (not shown), may be utilized, although it should be understood that wired, wireless, modem, and / or other connections or multiple connections to other computing devices may be utilized.
[0275] In one embodiment, the system may include a database that includes biomarker features characteristic for estimating or determining the risk of lung cancer. Biomarker data (or biomarker information) may be utilized as input to a computer for use as part of a computer-implemented method. Biomarker data may include data described herein.
[0276] In one aspect, the system further comprises one or more devices for providing input data to the one or more processors.
[0277] The system further comprises a memory for storing the dataset of ranked data elements.
[0278] In another embodiment, the device for providing input data includes a detector for detecting characteristics of the data elements, such as, for example, a mass spectrometer or a gene chip reader.
[0279] The system may additionally include a database management system. A user's request or query may be formatted in an appropriate language that can be understood by the database management system, which processes the query and extracts relevant information from the training set database.
[0280] The system may be connectable to a network to which a network server and one or more clients are connected. The network may be a local area network (LAN) or a wide area network (WAN), as is known in the art. Preferably, the server includes the necessary hardware to execute a computer program product (e.g., software) that accesses database data to process user requests.
[0281] The system may include an operating system (e.g., UNIX or Linux) for executing instructions from a database management system. The operating system may operate over a global communication network, such as the Internet, and may be connected to such network using global communication network servers.
[0282] The system may include one or more devices with a graphical display interface that includes interface elements such as buttons, pull-down menus, scroll bars, fields for entering text, etc., routinely found in graphical user interfaces known in the art. Requests entered into the user interface are transmitted to application programs in the system and formatted to search for relevant information in one or more system databases. User-entered requests or queries may be formulated in a suitable database language.
[0283] A graphical user interface may be generated by graphical user interface code as part of an operating system and may be used to input data and / or display input data. The results of processed data may be displayed in the interface, printed on a printer in communication with the system, stored on a storage device, and / or transmitted over a network, or provided in the form of a computer readable medium.
[0284] The system can be in communication with an input device for providing data regarding the data elements (e.g., values of expressions) to the system. In one aspect, the input device can include a gene expression profiling system, including, for example, a mass spectrometer, a gene chip, or an array reader.
[0285] According to various embodiments, the method and device for analyzing lung cancer risk using biomarker information can be implemented in any suitable manner, for example, using a computer program running on a computer system. Conventional computer systems including a processor and random access memory can be used, such as a remotely accessible application server, a network server, a personal computer, or a workstation. Additional computer system elements can include a storage or information storage system, such as a mass storage system, and a user interface, such as a conventional monitor, keyboard, and tracking device. The computer system can be a stand-alone system or part of a network of computers, including a server and one or more databases.
[0286] Lung cancer risk assessment using a biomarker analysis system may provide functions and operations to complete data analysis, such as data collection, processing, analysis, reporting, and / or diagnosis. For example, in one embodiment, a computer system may execute a computer program that may receive, store, retrieve, analyze, and report information about lung cancer risk biomarkers. The computer program may include multiple modules that perform various functions or operations, such as a processing module for processing raw data and generating supplemental data, and an analysis module for analyzing the raw data and supplemental data to generate an assessment of lung cancer risk. Calculating lung cancer risk may optionally include generating or retrieving any other information, including additional biomedical information regarding the individual's status associated with a disease, condition, or event, ascertaining whether further testing may be desirable, or assessing the individual's health status.
[0287] Referring now to FIG. 4, an example of a computer-implemented method in accordance with the principles of the disclosed embodiments is shown. FIG. 4 shows a flow chart 3000. At block 3004, biomarker information for an individual can be obtained. The biomarker information can be, for example, After performing the test on the body biological sample, it can be obtained from the computer database. The biomarker information can include biomarker values corresponding to one or more of the biomarkers in Table 6. At block 3008, a calculation can be performed for each of the biomarker values using a computer. Then, at block 3012, an estimate or determination can be made regarding the risk of lung cancer. The indication can be output to a display or other display device for human viewing. Thus, for example, it can be displayed on a display screen of a computer or other output device.
[0288] Some embodiments described herein may be implemented to include a computer program product, which may include a computer readable medium having computer readable program code embodied therein for executing an application program on a computer with a database.
[0289] As used herein, a "computer program product" refers to a set of instructions organized in the form of natural language statements or programming language statements that can be contained in a physical medium of any nature (e.g., written, electronic, magnetic, optical, etc.) and used by a computer or other automated data processing system. Such programming language statements, when executed by a computer or data processing system, cause the computer or data processing system to operate according to the specific content of the statements. Computer program products include, but are not limited to, source and object code embedded in a computer readable medium and / or programs in test or data libraries. Furthermore, computer program products that enable a computer system or data processing device to operate in a preselected manner can be provided in a number of forms, including, but not limited to, original source code, assembly code, object code, machine language, encrypted or compressed versions of the foregoing, and any equivalents.
[0290] In one aspect, a computer program product for estimating lung cancer risk is provided, the computer program product including a computer readable medium having program code embodied therein executable by a processor of a computing device or system, the program code including code for obtaining data originating from a biological sample from an individual, the data including biomarker values each corresponding to one or more biomarkers in Table 6, and code for performing a calculation method that indicates the individual's lung cancer risk as a function of the biomarker values.
[0291] In yet another aspect, a computer program product is provided that takes into account lung cancer risk, the computer program product including a computer readable medium having program code embodied therein executable by a processor of a computing device or system, the program code including code for obtaining data originating from a biological sample from an individual, the data including biomarker values corresponding to one or more biomarkers in Table 6, and code for performing a calculation method that indicates the individual's lung cancer risk as a function of the biomarker values.
[0292] While various embodiments have been described as methods or apparatus, it should be understood that the embodiments may be implemented via code in conjunction with a computer, e.g., code resident on or accessible by a computer. For example, software and databases may be utilized to implement many of the methods described above. Thus, in addition to hardware-implemented embodiments, it should also be noted that these embodiments may be implemented through the use of an article of manufacture consisting of a computer usable medium having computer readable program code embodied therein that enables the functionality disclosed herein. Thus, the embodiments are preferably considered to be protected by this patent in their program code means as well. Additionally, the embodiments may be implemented in a variety of formats, including RAM, ROM, magnetic media, optical media, and the like. The embodiments may be embodied as code stored in virtually any type of computer readable memory, including but not limited to, a programmable logic (PLC) or magneto-optical medium. Still more generally, the embodiments may be implemented in software, or in hardware, or any combination thereof, including but not limited to software running on a general purpose processor, microcode, PLA, or ASIC.
[0293] It is further contemplated that embodiments may be achieved as a computer signal embodied in a carrier wave and a signal propagated through a transmission medium (e.g., electrical and optical). Thus, the various types of information described above may be formatted in structures, such as data structures, and transmitted as an electrical signal over a transmission medium or stored on a computer-readable medium.
[0294] It should also be noted that many of the structures, materials, and acts recited herein may be recited as a means for performing a function or a step for performing a function, and such language should therefore be understood to be entitled to cover all structures, materials, or acts disclosed within this specification, including those incorporated by reference, and equivalents thereof.
[0295] The biomarker identification process, uses of the biomarkers, and various methods for determining biomarker values disclosed herein have been detailed above with respect to assessing lung cancer risk. However, the application of the process, the use of the identified biomarkers, and the methods for determining biomarker values are fully applicable to identifying other specific types of diseases or medical conditions, or individuals who may or may not benefit from adjunctive medical treatment.
[0296] Other methods In some embodiments, the biomarkers and methods described herein are used to determine health insurance premiums or coverage decisions and / or life insurance premiums or coverage decisions. In some embodiments, the results of the methods described herein are used to determine health insurance premiums and / or life insurance premiums. In some such instances, an organization providing health or life insurance requests or otherwise obtains information regarding a subject's tobacco use status and uses that information to determine appropriate health or life insurance premiums for the subject. In some embodiments, the tests are requested by and paid for by the organization providing health or life insurance. In some embodiments, the tests are used by potential acquirers of a business or insurance plan or company to forecast future liabilities or costs and determine whether the acquisition should proceed.
[0297] In some embodiments, the biomarkers and methods described herein are used to predict and / or manage utilization of medical resources. In some such embodiments, the methods are not performed for the purpose of such prediction, but information obtained by the methods is used in predicting and / or managing utilization of such medical resources. For example, a laboratory or hospital may collect information on a large number of subjects by the methods in order to predict and / or manage utilization of medical resources in a particular facility or in a particular geographic area. EXAMPLES
[0298] The following examples are provided for illustrative purposes only and are not intended to limit the scope of this application, which is defined by the appended claims. All examples described herein were performed using standard techniques well known and routine to those skilled in the art. Routine molecular biology techniques described in the following examples are described in detail in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd Edition. d., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, (2001).
[0299] Example 1. Model specifications 1.1 Description of Model Results. In certain embodiments, the endpoint of these analyses is the outcome of time to onset of lung cancer, which has two components. 1) The number of days from blood draw to diagnosis of lung cancer (as primary cancer) or to end / completion of the study. 2) a binary variable indicating whether a pulmonary diagnosis was observed during the study period;
[0300] 1.2 Final Model Information. In one embodiment, the final model is a 7-feature (see Table 6), accelerated failure time (AFT) Weibull survival model. The model was trained over the entire study period, with maximum performance at 5 years.
[0301] The model provides two predictions: 1) Absolute Risk: This output is the absolute probability of not being diagnosed with lung cancer within 5 years (traditionally Pr(LC-free), a value between 0 and 1). When delivered as a result and evaluated by business rules, the resulting predicted probability is subtracted from 1 to provide the absolute 5-year probability of being diagnosed with lung cancer. 2) Relative Risk: This is a continuous variable calculated using the absolute risk probability generated by the model (above) and divided by the "baseline" absolute probability (defined below) of the training cohort. This approach allows the model's risk probability predictions to be interpreted such that higher values indicate a higher likelihood of being diagnosed with lung cancer within the next five years.
[0302] The baseline risk probability score represents the "average" person in the training cohort based on the model algorithm. A "baseline" individual is defined as an individual with model feature values set to zero. All features in the model are global mean centered, meaning that for any given feature value, a value of 0 is equal to the mean (i.e., the average value). The baseline value is calculated by setting all features to zero and generating absolute risk probabilities for those "zeroed" features. Thus, a score less than 1 represents a lower than average risk and a score greater than 1 represents a higher than average risk. Lung cancer diagnosis rates in the "ever smoker" ARIC dataset are consistent with lung cancer diagnosis event rates in the US population of the same age as the intended population.
[0303] The baseline risk score for the training data is 0.0095 or 0.95%.
[0304] The stratification of absolute risk scores is based on quartiles, with the first two quartiles collapsed into a single risk bin. (Preliminary analysis of the training data did not show a strong separation between the first two quartiles.) Thus, as shown in the table below (Table 4), the groupings represent Q1+Q2, Q3, and Q4, with the baseline risk (0.0095) close to the upper limit of Q2.
[0305] To support reporting lung cancer risk as a relative risk for the purposes of the LDT, a Kaplan-Meier plot stratifying subjects into the three relative risk groups is shown in Figure 2, and a summary of the absolute probability and relative risk stratifications and corresponding event rates is provided in Table 4. [Table 4]
[0306] Based on this stratification, the scoring rules for the relative risk of lung cancer risk tests are shown in Table 5. [Table 5] [Table 6] [Table 7] [Table 8] [Table 9] [Table 10] [Table 11] [Table 12] [Table 13]
[0307] 1.3 In one embodiment, the output of the model is Pr(lung cancer free) at 5 years. The output is reported as the probability of being diagnosed with lung cancer, which is (1-Pr(lung cancer free)). The 5-year event probability is reported as a continuous variable. Since the output of this model is a probability, values outside the range [0,1] are malfunctions and will not be reported.
[0308] Example: Based on the proteomic model, the absolute predicted probability of a hypothetical patient is 0.0100. The relative risk for this patient is 1.06, placing him in the medium risk bin.
number
[0309] Example 2. Data set for test development and validation 2.1 Development and Validation Cohort(s). The Atherosclerosis Risk in Communities (ARIC) study is a prospective epidemiologic study conducted in four communities in the United States: Forsyth County, NC; Jackson, MS; the northwest suburbs of Minneapolis, MN; and Washington County, MD. The ARIC study enrolled 15,792 participants aged 45-64 years. Enrollment occurred from 1987 to 1989, with 30 years of follow-up currently ongoing through the sixth study visit in 2016-2017. Although the ARIC study was originally designed to investigate the etiology and natural history of atherosclerosis, etiology of clinical atherosclerotic disease, cardiovascular risk factors, and variations in medical care and disease by race, sex, location, and date, the study has expanded to implement initiatives to advance cancer epidemiology research. Cancer incidence and mortality were determined through review of medical records, abstract records, and cancer registries in combination with self-reported information obtained during follow-up. (Joshu et al. “Enhancing the Infrastructure of the Atherosclerosis Risk in Communities (ARIC) Study for Cancer Epidemiology Research: ARIC Cancer.” Cancer Epidemiol Biomarkers Prev 2018;27:295-305). A lung cancer risk model was developed based on documented lung cancer diagnoses among ever smokers (defined as current or former smokers at visit 3) from visit 3 conducted in 1993-1995 through a 20-year follow-up period, although the model is intended to include lung cancer diagnoses for up to 5 years of follow-up.
[0310] For the recent (2014-2018) US population, the 5-year risk of lung cancer for men and women aged 40-74 years is 0.715%, regardless of smoking status. (SEER *Explorer” National Cancer Institute. Surveillance, Epidemiology, and End Results Program. (Available online at https: / / seer.cancer.gov / explorer.) This is consistent with a 5-year incidence rate of lung cancer of 0.759% when considering both ever and never smokers, and an increased 5-year incidence rate of 1.23% among ever smokers, in the ARIC Visit 3 dataset for men and women aged 50-73 years.
[0311] 2.2 Dataset stratification. In this study, the data was split independently into training (70%), validation (15%), and validation (15%) sets, stratified with lung cancer diagnosis as the endpoint, to identify robust models while mitigating the overfitting problem. The validation dataset was not used in the POC or refinement phases.
[0312] 2.2.1 Model training data [Table 14] [Table 15]
[0313] 2.2.2 Model Validation Data [Table 16] [Table 17]
[0314] 2.2.3 Model Validation Data [Table 18] [Table 19]
[0315] Example 3. Development results 3.1 Data QC and Pre-analysis Results
[0316] The original clinical dataset contained 11,288 samples. The following number of samples were removed based on various flags. The removals are detailed in Table 14. [Table 20]
[0317] In addition, 363 analytes did not pass the target confirmation specificity test and were therefore removed before the analysis began, leaving 4921 analytes available for analysis.
[0318] No other issues were identified during data QC or pre-analysis.
[0319] Samples for this study were run from May 13, 2019 to July 8, 2019 with assay version 4.0 master mix lot 1.
[0320] 3.2 Refinements Approach and Results. The final model for lung cancer risk testing included seven features and was developed using the AFT survival model with a Weibull distribution. The model was trained on 70% of the ARIC Visit 3 dataset using individuals with no prevalent cancer and current or past smoking. Validation metrics were calculated on another 15% of the dataset, and an additional 15% of the dataset was retained for use in validation.
[0321] Observations across the entire study period were used to optimize model performance over a 5-year period. The model was built using a reduced feature list of the top 100 univariate features and further refined by stability selection with upsampling of event classes. Features with high CV (CV>10%) were removed from the remaining analytes. A backward feature selection process using an AFT Weibull model was used to further refine the features to find the best performing model.
[0322] The improved model was further evaluated using a model enrichment tool and refined to ensure concordance between assay versions V4.0 and V4.1.
[0323] The final model was evaluated for predicting lung cancer diagnosis in individuals with a smoking history using the 5-year (1825-day) AUC. Additional metrics such as C-index, PEC, sensitivity, and specificity were reported. Results for the training and validation datasets are shown in Table 15. [Table 21]
[0324] For informational purposes only, the final model was further used to predict the risk of lung cancer diagnosis 10 and 15 years after blood draw for ever-smoking individuals, and performance metrics were calculated. The performance metrics are detailed in Table 16. [Table 22]
[0325] For research purposes only, the final model was further used to predict the risk of lung cancer diagnosis in non-smokers from ARIC Visit 3 at 5, 10, and 15 years after blood draw, and performance metrics were calculated. Lung cancer event rates and summary demographics for the non-smoker dataset are detailed in Tables 17a and 17b, respectively. Performance metrics are detailed in Table 18. [Table 23] [Table 24] [Table 25]
[0326] Example 4. Longitudinal analysis of lung cancer risk prediction Data from ARIC Visits 2, 3, and 5 were used to analyze changes in lung cancer risk over time according to specific parameters: changes in lung cancer risk associated with changes or concordance in smoking status between ARIC Visits 2 and 3 (Objective 1), changes in lung cancer risk associated with changes or concordance in smoking status between ARIC Visits 2 and 3 in subjects diagnosed with lung cancer at different proxies after Visit 3 (Objective 2), and changes in lung cancer risk associated with changes or concordance in smoking status between ARIC Visits 2 and 3 in subjects diagnosed with lung cancer at different proxies after Visit 3 (Objective 3). We assessed changes in cancer risk over time (Objective 2) and differences in lung cancer risk predictions between individuals with and without a diagnosis of prevalent lung cancer at blood draw at Visit 3 or Visit 5 (Objective 3). Demographics of subjects at ARIC Visits 2, 3, and 5 are shown in Tables 19-22. [Table 26] [Table 27] [Table 28] [Table 29]
[0327] Data quality control (QC) pre-analysis was performed on ARIC Visits 2, 3, and 5. There were 11,779 samples from Visit 2, 11,360 samples from Visit 3, and 5,281 samples from Visit 5 with clinical and RFU data available for analysis.
[0328] Data QC identified 28 (0.238%) samples from Visit 2, 41 (0.361%) samples from Visit 3, and 27 (0.511%) samples from Visit 5 as outlier samples, indicating that 5% or more of the analytes were greater than 6 times the absolute deviation from the median. Data QC also indicated that 17 (0.144%) samples from Visit 2, 36 (0.317%) samples from Visit 3, and 0 (0.0%) samples from Visit 5 failed the row check, meaning that either the hybridization or at least one of the three scale factor medians was outside the range of 0.4 to 2.5. A failed row check indicates that that particular sample had a technical issue (e.g., clogging) that was not corrected by running the sample again. Table 23 summarizes the data QC and the samples that were removed from each ARIC dataset prior to analysis. [Table 30]
[0329] Samples from Visits 2 and 3 were used for Aim 1 and Aim 2. There were 10,048 Visit 2 and Visit 3 samples available for these purposes. There were 19 prevalent lung cancer cases at both time points, which were excluded from the analysis, leaving 10,029 (99.811%) samples for analysis. Smoking behavior variables were constructed based on self-reported smoking exposure. Table 24 summarizes the smoking behavior variables.
[0330] We excluded 274 missing or misclassified samples (10 individuals were missing smoking status at Visit 2, 23 were missing smoking status at Visit 3, and 241 individuals were misclassified and reported as a type of smoker at Visit 2 and a non-smoker at Visit 3) from the analysis, leaving 9,755 (97.084%) for analyses in Objective 1. For analyses of smoking exposure variables, we combined “new current smokers” and “new ex-smokers” into a “new smoker” variable because the sample size of the “new current smoker” exposure group was small.
[0331] Of the 10,029 samples available for analysis, there were 624 individuals with missing lung cancer diagnosis information, leaving 9,405 (93.8%) samples for analysis in objective 2. [Table 31]
[0332] Results of Objective 1
[0333] Objective 1 examined changes in lung cancer incidence due to changes or consistency in smoking status between Visit 2 and Visit 3. All individuals (N=9,755) who were not prevalent in lung cancer in the Visit 2 and Visit 3 samples were used in this analysis. The mean difference in lung cancer prediction (Figure 5) was calculated as the difference between the mean lung cancer prediction at Visit 3 (0.0111) and the mean lung cancer prediction at Visit 2 (0.0092) (mean difference=0.0020; p<0.001). An ANOVA test was performed to determine whether there were differences in lung cancer predictions between smoking status groups (p<0.001). Post-hoc t-tests were performed to determine which smoking groups were statistically different from each other and are summarized in Table 25. [Table 32] [Table 33]
[0334] The mean change in prediction between visits 2 and 3 for nonsmokers was 0.0019, for persistent smokers 0.003, for new smokers 0.0025, for quitters -0.0023, for relapsers 0.0048, and for persistent quitters 0.0019 (summarised in Table 25).
[0335] Post-hoc t-tests showed no significant differences between continued quitters and continued smokers (difference = -0.0011; p < 0.001), between continued smokers and non-smokers (difference = 0.0011; p < 0.001), between continued quitters and intermediate quitters (difference = 0.0042; p < 0.001), between continued smokers and intermediate quitters (difference = 0.0053; p < 0.001), and between non-smokers and intermediate quitters (difference = 0.0 Significant differences were noted between continuers and non-smokers, between continuers and new smokers, between continuers and new smokers, and between non-smokers and new smokers (difference = 0.042; p < 0.001), between new smokers and late quitters (difference = 0.0048; p < 0.001), between continuers and relapsers (difference = -0.0029; p = 0.0054), between non-smokers and relapsers (difference = -0.0029; p = 0.0054), and between late quitters and relapsers (difference = -0.0071; p < 0.001). There were no significant differences between continuers and non-smokers, between continuers and new smokers, between continuers and new smokers, and between non-smokers and new smokers (summarized in Table 26).
[0336] Results of Objective 2
[0337] Of the Visit 2 and Visit 3 individuals (N = 9,755) used in the analyses for Objective 1, 624 individuals (6.6%) were missing lung cancer diagnosis information, leaving 9,405 (93.4%) individuals for analysis with samples from Visits 2 and 3. There were 313 individuals with an incidental diagnosis of lung cancer after Visit 3 (by the end of follow-up) and 63 individuals with an incidental diagnosis of lung cancer within 5 years of Visit 3.
[0338] The change in lung cancer risk score was calculated between Visit 2 and Visit 3. A t-test was used to determine whether there was a significant difference in the change in lung cancer risk score between Visit 2 and Visit 3 for individuals diagnosed with cancer compared to individuals never diagnosed with lung cancer (summarized in Table 27). [Table 34]
[0339] Lung cancer visit 3 predictions were examined in more detail in individuals who developed lung cancer within 5 years compared to those who developed lung cancer after 5 years (these two groups are mutually exclusive). Of the 313 individuals who had lung cancer at visit 3, 63 developed lung cancer within 5 years and 250 developed lung cancer after 5 years. These lung cancer predictions were compared using t-tests and are summarized in Table 28. [Table 35]
[0340] Concordance was measured to determine whether changes in lung cancer risk predictions were predictive of time to cancer diagnosis. There were 2 (0.02%) individuals (lung cancer diagnosis = "N") with 0 days of follow-up and were excluded from this analysis (number of analyses = 9,403). Table 29 summarizes the concordance analysis and shows that changes in lung cancer predictions were not predictive of time to cancer diagnosis (lung cancer total concordance = 0.550; lung cancer 5-year concordance = 0.556). [Table 36]
[0341] Results of Objective 3
[0342] Of the individuals with visit 5 samples (5,254), there were 300 (5.7%) individuals with missing lung cancer diagnosis information who were excluded from this analysis. There were 9,454 (94.3%) samples available for analysis. The purpose 2 samples were used for purpose 3 analyses.
[0343] A t-test was performed to determine whether there was a significant difference in lung cancer risk prediction between individuals with and without prevalent lung cancer. Tables 30 and 31 summarize this analysis at Visit 3 and Visit 5, respectively, and show a significant difference between individuals with prevalent lung cancer (Visit 3 mean predicted value=0.24; Visit 5 mean predicted value=0.33) and individuals without prevalent lung cancer (Visit 3 mean predicted value=0.011; Visit 5 mean predicted value=0.021) (p=0.003 and p=0.005, respectively). One individual at Visit 3 had follow-up data at Visit 5. [Table 37]
[0344] [Table 38] References All references listed below, or elsewhere throughout the specification, are hereby incorporated by reference in their entirety. 1.Joshu et al. “Enhancing the Infrastructure of the Atherosclerosis Risk in Communities (ARIC) Study for Cancer Epidemiology Research: ARIC Cancer.” Cancer Epidemiol Biomarkers Prev 2018;27:295-305. 2.Tammemagi, et al. “Selection criteria for lung-cancer screening.” N Engl J Med 2013;368:728-36. 3.Siegel et al. “Cancer Statistics, 2021.” CA Cancer J Clin 2021;71:7-33. 4.“About Lung Cancer”, American Cancer Society. Available online at https: / / www.cancer.org / cancer / lung-cancer / about / . 5. “Cancer Stat Facts: Lung and Bronchus Cancer.” National Cancer Institute. Surveillance, Epidemiology, and End Results Program. Available online at: https: / / seer.cancer.gov / statfacts / html / lungb.html. 6.“Non-Small Cell Lung Cancer Treatment (PDQ®)-Patient Version” (August 2021). National Cancer Institute. Available online at https: / / www.cancer.gov / types / lung / patient / non-small-cell-lung-treatment-pdq#_118. 7.“Lung Cancer Among People Who Never Smoked”(November 2020). Centers for Disease Control and Prevention. Available online at https: / / www.cdc.gov / cancer / lung / nonsmokers / index.htm. 8. Bruder et al. “Estimating lifetime and 10-year risk of lung cancer.” Prev Med Rep 2018;11:125-30. 9. Samet JM “Health benefits of smoking cessation.” Clin Chest Med 1991;12:669-79. 10.Force USPST, et al. “Screening for Lung Cancer: US Preventive Services Task Force Recommendation Statement.” JAMA 2021;325:962-70. 11.National Lung Screening Trial Research Team. “Reduced lung-cancer mortality with low-dose computed tomographic screening.” N Engl J Med 2011;365:395-409. 12. “Lung CT Screening Reporting & Data System (Lung-RADS).” American College of Radiology, Available online at: https: / / www.acr.org / Clinical-Resources / Reporting-and-Data-Systems / Lung-Rads. 13.“Lung Cancer Screening”(March 2021)Ma Mayo Clinic. Available online at https: / / www.mayoclinic.org / tests-procedures / lung-cancer-screening / about / pac-20385024. 14. Zahnd et al. “Lung Cancer Screening Utilization: A Behavioral Risk Factor Surveillance System Analysis.” Am J Prev Med 2019;57:250-5. 15.Tammemagi MC “Application of risk prediction models to lung cancer screening: a review.” J Thorac Imaging 2015;30:88-100. 16. “SEER *Explorer” National Cancer Institute. Surveillance, Epidemiology, and End Results Program. Available online at https: / / seer.cancer.gov / explorer. 17. Fuentes et al. “Comprehension of Top 200 Prescribed Drugs in the US as a Resource for Pharmacy Teaching, Training and Practice.” Pharmacy (Basel) 2018;6. 18. Roman J. “On the ‘TRAIL’ of a Killer: MMP12 in Lung Cancer.” Am J Respir Crit Care Med. 2017 Aug 1;196(3):262 - 264. 19. Mehan et al. “Validation of a blood protein signature for non - small cell lung cancer.” Clin Proteomics. 2014 Aug 1;11(1):32. doi: 10.1186 / 1559 - 0275 - 11 - 32. 20. Umeda et al. “Surfactant protein D inhibits activation of non - small cell lung cancer - associated mutant EGFR and affects clinical outcomes of patients.” Oncogene. 2017 Nov 16;36(46):6432 - 6445. 21.Fu et al. “A meta-analysis of influence of MSMB promoter rs10993994 polymorphisms on prostate cancer risk.” Eur Rev Med Pharmacol Sci. 2019 Nov;23(21):9295-9303. 22.Lv et al. “Diagnostic value of human epididymis protein 4 in malignant pleural effusion in lung cancer.” Cancer Biomark. 2019;26(4):523-528 23.Moore,et al.“HE4 (WFDC2) gene overexpression promotes ovarian tumor growth.” Sci Rep 4,3574(2014)
Claims
1. a) measuring the level of PH protein and the levels of at least one, two, three, four, five, or six proteins selected from MMP-12, SP-D, HE4, PSP-94, FUT5, and CRLF1 in a sample from a human subject; b) identifying said human subject as at risk for developing lung cancer based on the level of said PH and the levels of said at least 1, 2, 3, 4, 5, or 6 proteins.
2. a) measuring the level of FUT5 protein and the levels of at least one, two, three, four, five, or six proteins selected from MMP-12, SP-D, HE4, PSP-94, PH, and CRLF1 in a sample from a human subject; b) identifying the human subject as being at risk for developing lung cancer based on the level of FUT5 and the levels of at least one, two, three, four, five, or six proteins. A method comprising:
3. a) measuring the level of CRLF1 protein and the levels of at least one, two, three, four, five, or six proteins selected from MMP-12, SP-D, HE4, PSP-94, PH, and FUT5 in a sample from a human subject; b) identifying said human subject as being at risk for developing lung cancer based on said level of CRLF1 and said levels of at least one, two, three, four, five, or six proteins. A method comprising:
4. a) measuring the level of PSP-94 protein and the levels of at least one, two, three, four, five, or six proteins selected from MMP-12, SP-D, HE4, PH, FUT5, and CRLF1 in a sample derived from a human subject; b) identifying said human subject as being at risk for developing lung cancer based on the level of said PSP-94 and the levels of said at least one, two, three, four, five, or six proteins. A method comprising:
5. The method comprises: PH and MMP-12; PH and SP-D; PH and HE4; PH and PSP-94; PH and FUT5; PH and CRLF1; PH, MMP-12, and SP-D; PH, MMP-12, and HE4; PH, MMP-12, and PSP-94; PH, MMP-12, and FUT5; PH, MMP-12, and CRLF1; PH, SP-D, and HE4; 2. The method of claim 1, comprising measuring PH, SP-D, and PSP-94; PH, SP-D, and FUT5; PH, SP-D, and CRLF1; PH, HE4, and PSP-94; PH, HE4, and FUT5; PH, HE4, and CRLF1; PH, PSP-94, and FUT5; PH, PSP-94, and CRLF1; or PH, FUT5, and CRLF1.
6. The method comprises the steps of: FUT5 and MMP-12; FUT5 and SP-D; FUT5 and HE4; FUT5 and PSP-94; FUT5 and PH; FUT5 and CRLF1; FUT5, MMP-12, and SP-D; FUT5, MMP-12, and HE4; FUT5, MMP-12, and PSP-94; FUT5, MMP-12, and PH; FUT5, MMP-12, and CRLF1; FUT5, SP-D, and The method of claim 2, comprising measuring HE4; FUT5, SP-D, and PSP-94; FUT5, SP-D, and PH; FUT5, SP-D, and CRLF1; FUT5, HE4, and PSP-94; FUT5, HE4, and PH; FUT5, HE4, and CRLF1; FUT5, PSP-94, and PH; FUT5, PSP-94, and CRLF1; or FUT5, PH, and CRLF1.
7. The method comprises: CRLF1 and MMP-12; CRLF1 and SP-D; CRLF1 and HE4; CRLF1 and PSP-94; CRLF1 and PH; CRLF1 and FUT5; CRLF1, MMP-12, and SP-D; CRLF1, MMP-12, and HE4; CRLF1, MMP-12, and PSP-94; CRLF1, MMP-12, and PH; CRLF1, MMP-12, and FUT5; CRLF1, SP-D; and HE4; CRLF1, SP-D, and PSP-94; CRLF1, SP-D, and PH; CRLF1, SP-D, and FUT5; CRLF1, HE4, and PSP-94; CRLF1, HE4, and PH; CRLF1, HE4, and FUT5; CRLF1, PSP-94, and PH; CRLF1, PSP-94, and FUT5; or CRLF1, PH, and FUT5.
8. The method according to claim 7, wherein the method comprises the step of: PSP-94 and MMP-12; PSP-94 and SP-D; PSP-94 and HE4; PSP-94 and PH; PSP-94 and FUT5; PSP-94 and CRLF1; PSP-94, MMP-12, and SP-D; PSP-94, MMP-12, and HE4; PSP-94, MMP-12, and PH; PSP-94, MMP-12, and FUT5; PSP-94, MMP-12, and CRLF1; PSP-94, SP 5. The method of claim 4, comprising measuring PSP-94, SP-D, and HE4; PSP-94, SP-D, and PH; PSP-94, SP-D, and FUT5; PSP-94, SP-D, and CRLF1; PSP-94, HE4, and PH; PSP-94, HE4, and FUT5; PSP-94, HE4, and CRLF1; PSP-94, PH, and FUT5; PSP-94, PH, and CRLF1; or PSP-94, FUT5, and CRLF1.
9. The method comprises: a) PH and at least one protein selected from the following: MMP-12, SP-D, HE4, FUT5, and CRLF1; b) FUT5 and at least one protein selected from the following: MMP-12, SP-D, HE4, PH, and CRLF1; or c) CRLF1 and at least one protein selected from the following: MMP-12, SP-D, HE4, PH, and FUT5. The method of claim 4, comprising measuring:
10. The method comprises: a) FUT5 and at least one protein selected from the following: MMP-12, SP-D, HE4, PSP-94, and CRLF1; or b) CRLF1 and at least one protein selected from the following: MMP-12, SP-D, HE4, PSP-94, and FUT5. The method of claim 1 , comprising measuring:
11. 3. The method of claim 2, wherein the method comprises measuring FUT5 and CRLF1, and at least one protein selected from the following: MMP-12, SP-D, HE4, PSP-94, and PH.
12. a) measuring the levels of at least 3, 4, 5, 6, or 7 proteins selected from PH, MMP-12, SP-D, HE4, PSP-94, FUT5, and CRLF1 in a sample from a human subject; b) identifying said human subject as being at risk for developing lung cancer based on the levels of said at least three, four, five, six, or seven proteins. A method comprising:
13. The method comprises the steps of: MMP-12, SP-D, and HE4; MMP-12, SP-D, and PSP-94; MMP-12, SP-D, and PH; MMP-12, SP-D, and FUT5; MMP-12, SP-D, and CRLF1; MMP-12, HE4, and PSP-94; MMP-12, HE4, and PH; MMP-12, HE4, and FUT5; MMP-12, HE 4, and CRLF1; MMP-12, PSP-94, and PH; MMP-12, PSP-94, and FUT5; MMP-12, PSP-94, and CRLF1; MMP-12, PH, and FUT5; MMP-12, PH, and CRLF1; MMP-12, FUT5, and CRLF1; SP-D, HE4, and PSP-94; SP-D, HE4, and PH; SP-D; HE4, and FUT5; SP-D, HE4, and CRLF1; SP-D, PSP-94, and PH; SP-D, PSP-94, and FUT5; SP-D, PSP-94, and CRLF1; SP-D, PH, and FUT5; SP-D, PH, and CRLF1; SP-D, FUT5, and CRLF1; HE4, PSP-94, and PH; HE4, PSP-94, and FU 13. The method of claim 12, comprising measuring T5; HE4, PSP-94, and CRLF1; HE4, PH, and FUT5; HE4, PH, and CRLF1; HE4, FUT5, and CRLF1; PSP-94, PH, and FUT5; PSP-94, PH, and CRLF1; PSP-94, FUT5, and CRLF1; or PH, FUT5, and CRLF1.
14. 14. The method of claim 13, further comprising measuring one or more of PSP-94, PH, FUT5, and CRLF1.
15. The method according to any one of claims 1 to 14, wherein the measurement is carried out using an aptamer, an antibody, or a combination of an aptamer and an antibody.
16. The method of any one of claims 1 to 14, wherein the measuring is performed using mass spectrometry, an aptamer-based assay, and / or an antibody-based assay.
17. 17. The method of claim 16, wherein the level of each measured biomarker protein is determined from relative fluorescence units (RFU) or protein concentration.
18. The method of any one of claims 1 to 14, wherein the sample is selected from blood, plasma, serum, or urine.
19. 15. The method of any one of claims 1 to 14, wherein the risk of developing lung cancer is within 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, or 10 years.
20. The method of any one of claims 1 to 14, wherein the human subject is a current or ex-smoker, and optionally, the subject has lung cancer.
21. 15. The method of any one of claims 1-14, wherein the method provides an area under the curve (AUC) of 0.62, 0.67, 0.68, 0.68, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, or greater; or an area under the curve (AUC) of about 0.6 to about 0.8, about 0.61 to about 0.78, about 0.62 to about 0.76, about 0.62 to about 0.68, about 0.67 to about 0.72, about 0.69 to about 0.74, about 0.71 to about 0.74, about 0.73 to about 0.76, or about 0.74 to about 0.
76.
22. 15. The method of any one of claims 1 to 14, wherein the prediction of the risk of developing lung cancer is based on input of the measured protein levels in a statistical model, optionally wherein the prediction comprises analyzing the measured protein levels using an accelerated failure time (AFT) Weibull survival model.
23. 1. A kit comprising N protein capture reagents, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and at least one of the N protein capture reagents specifically binds to a protein selected from PH, PSP-94, MMP-12, SP-D, HE4, FUT5, and CRLF1.
24. 24. The kit of claim 23, wherein N is at least 2 and at least one of the two N protein capture reagents specifically binds to the protein selected from PH, PSP-94, MMP-12, SP-D, HE4, FUT5, and CRLF1.
25. 24. The kit of claim 23, wherein N is 2 to 7, or N is 3 to 7, or N is 4 to 7, or N is 5 to 7, or N is 6 to 7; or N is 2, N is 3, N is 4, N is 5, N is 6, or N is 7.
26. 26. The kit of any one of claims 23 to 25, wherein each of the N protein capture reagents specifically binds to a different biomarker protein, and optionally, each of the N protein capture reagents specifically binds to a protein selected from PH, PSP-94, MMP-12, SP-D, HE4, FUT5, and CRLF1.
27. Two of the N protein capture reagents specifically bind to PSP-94 and MMP-12, or two of the N protein capture reagents specifically bind to PSP-94 and SP-D, or two of the N protein capture reagents specifically bind to PSP-94 and HE4, or two of the N protein capture reagents specifically bind to PSP-94 and PH, or two of the N protein capture reagents specifically bind to PSP-94 and FUT5, or two of the N protein capture reagents specifically bind to PSP-94 and CRLF1, or two of the N protein capture reagents specifically bind to PH and MMP-12, or two of the N protein capture reagents specifically bind to PH and SP-D, or two of the N protein capture reagents specifically bind to PH and HE4, or 26. The kit of any one of claims 23 to 25, wherein two of the protein capture reagents specifically bind to PH and FUT5, or two of the N protein capture reagents specifically bind to PH and CRLF1, or two of the N protein capture reagents specifically bind to FUT5 and MMP-12, or two of the N protein capture reagents specifically bind to FUT5 and SP-D, or two of the N protein capture reagents specifically bind to FUT5 and HE4, or two of the N protein capture reagents specifically bind to FUT5 and CRLF1, or two of the N protein capture reagents specifically bind to CRLF1 and MMP-12, or two of the N protein capture reagents specifically bind to CRLF1 and SP-D, or two of the N protein capture reagents specifically bind to CRLF1 and HE4.
28. Three of the N protein capture reagents specifically bind to PSP-94, MMP-12, and SP-D, or three of the N protein capture reagents specifically bind to PSP-94, MMP-12, and HE4, or three of the N protein capture reagents specifically bind to PSP-94, MMP-12, and PH, or three of the N protein capture reagents specifically bind to PSP-94, MMP-12, and FUT5, or three of the N protein capture reagents specifically bind to PSP three of the N protein capture reagents specifically bind to PSP-94, SP-D, and HE4; three of the N protein capture reagents specifically bind to PSP-94, SP-D, and PH; three of the N protein capture reagents specifically bind to PSP-94, SP-D, and FUT5; three of the N protein capture reagents specifically bind to PSP-94, SP-D, and CRLF1; Three of the N protein capture reagents specifically bind to PSP-94, HE4, and PH, or three of the N protein capture reagents specifically bind to PSP-94, HE4, and FUT5, or three of the N protein capture reagents specifically bind to PSP-94, HE4, and CRLF1, or three of the N protein capture reagents specifically bind to PSP-94, PH, and FUT5, or three of the N protein capture reagents specifically bind to PSP-94, PH, and CRLF1. or three of the N protein capture reagents specifically bind to PSP-94, FUT5, and CRLF1, or three of the N protein capture reagents specifically bind to PH, MMP-12, and SP-D, or three of the N protein capture reagents specifically bind to PH, MMP-12, and HE4, or three of the N protein capture reagents specifically bind to PH, MMP-12, and FUT5, or three of the N protein capture reagents specifically bind to PH, MMP-12,and CRLF1, or three of the N protein capture reagents specifically bind to PH, SP-D, and HE4, or three of the N protein capture reagents specifically bind to PH, SP-D, and FUT5, or three of the N protein capture reagents specifically bind to PH, SP-D, and CRLF1, or three of the N protein capture reagents specifically bind to PH, HE4, and FUT5, or three of the N protein capture reagents specifically bind to PH, HE4 and CRLF1, or three of the N protein capture reagents specifically bind to PH, PSP-94, and FUT5, or three of the N protein capture reagents specifically bind to PH, PSP-94, and CRLF1, or three of the N protein capture reagents specifically bind to PH, FUT5, and CRLF1, or three of the N protein capture reagents specifically bind to FUT5, MMP-12, and SP-D, or three of the N protein capture reagents specifically bind to FUT5, MMP-12, and HE4, or three of the N protein capture reagents specifically bind to FUT5, MMP-12, and CRLF1, or three of the N protein capture reagents specifically bind to FUT5, SP-D, and HE4, or three of the N protein capture reagents specifically bind to FUT5, SP-D, and CRLF1, or three of the N protein capture reagents specifically bind to FUT5, HE4, and CRLF1, or three of the N protein capture reagents specifically bind to FUT5, PSP-94, and CRLF1, or three of the N protein capture reagents specifically bind to FUT5, PH, and CRLF1, or three of the N protein capture reagents specifically bind to CRLF1, MMP-12, and SP-D, or three of the N protein capture reagents specifically bind to CRLF1, MMP-12, and HE4, or three of the N protein capture reagents specifically bind to CRLF1, SP-D,and the kit according to any one of claims 23 to 25, which specifically binds to HE4.
29. The kit of any one of claims 23 to 25, wherein each of the N biomarker protein capture reagents is an antibody or an aptamer.
30. 26. The kit of any one of claims 23 to 25, for use in detecting the N biomarker proteins in a sample from a subject, optionally wherein detecting the N biomarker proteins is predictive of the subject's risk of developing lung cancer.