Methods for assessing tobacco use status
A method using multiple biomarker proteins and statistical models addresses the limitations of self-reported tobacco use by accurately predicting former smokers, improving the identification of at-risk individuals for smoking-related health issues.
Patent Information
- Application Number
- JP2025514516
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-23
- Filing Date
- 2023-09-22
- Publication Date
- 2025-10-15
AI Technical Summary
Current methods for assessing tobacco use status, particularly distinguishing between current and former smokers, rely heavily on subjective self-reporting and lack reliable biomarkers, limiting the accuracy of identifying clinically relevant populations for smoking-related health risks.
Development of a method involving the detection of multiple biomarker proteins, such as EPHA6, placental alkaline phosphatase, and others, using mass spectrometry or antibody-based assays to predict the probability of an individual being a former tobacco user, employing statistical models for risk prediction.
Provides a more accurate assessment of tobacco use history by identifying former smokers through objective biomarker analysis, enhancing the identification of at-risk populations for smoking-related diseases.
Smart Images

Figure 2025534223000125 
Figure 2025534223000126 
Figure 2025534223000127
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 409,338, filed September 23, 2022, which is incorporated herein by reference in its entirety for all purposes.
[0002] This application relates generally to detection and methods of biomarkers that predict tobacco use status in an individual, and more specifically to one or more biomarkers, methods, devices, reagents, systems, and kits used to predict tobacco use status in an individual. [Background technology]
[0003] The following provides a summary of information that may be relevant to the present application and is not an admission that any of the information provided or publications referenced herein is prior art to the present application.
[0004] According to the Centers for Disease Control and Prevention (CDC), tobacco smoking is the direct cause of more than 480,000 deaths in the United States each year. Smoking is strongly correlated with numerous chronic health conditions (including cardiovascular disease, pulmonary disease, and lung cancer) that reduce quality of life, and the amount and frequency of smoking are directly proportional to the increased risk of disease (Banks E, et al., BMC Med, 2019;17(1):128). Quitting smoking before age 40 can reduce the risk of dying from smoking-related diseases by up to 90%, and therefore the USDaPt Department of Health and Services has been encouraging cessation of heavy smoking for nearly 60 years (Brawley OW, et al. CA Cancer J Clin, 2014;64(1):5-8). Former smokers are also at higher risk of developing comorbidities compared to never smokers (Kramarow EA. Health of former cigarette smokers aged 65 and over: United States, 2018. National Center for Health Statistics. July 2020. Available at: https: / / www.cdc.gov / nchs / data / nhsr / nhsr145-508.pdf).
[0005] The current standard of care guidelines, as outlined by the Clinical Practice Guideline Treating Tobacco Use and Dependence 2008 Update Panel, Liaisons, and Staff, recommend the 5 As for tobacco intervention by primary care providers: "(1) ask the patient if they are using tobacco, (2) advise them to quit, (3) assess their intention to quit, (4) assist those who are willing to quit, and (5) arrange for follow-up contact to prevent relapse" (The Clinical Practice Guideline Treating Tobacco Use and Dependence 2008 Update Panel, Liaisons, and Staff. A Clinical Practice Guideline for Treating Tobacco Use and Dependence: 2008 Update, A US Public Health Service Report. Am J Prev Med, 2008;35(2):158-176). In practice, clinicians often rely on patients' self-reported smoking habits to assess their risk for smoking-related diseases and their eligibility for screening. Secondarily, biochemical measures (such as cotinine, carbon monoxide, or thiocyanate) can be used to assess the accuracy of self-reported smoking habits, but all biochemical measures of smoking are limited to assessing "current" smoking habits (Jarvis MJ, et al. Am J Public Health, 1987;77:1435-1438. Patrick DL, et al. Am J Public Health, 1994;84(7):1086-1093). The development of surrogates for assessing tobacco use status, for example, via blood-based measurements, has the potential to identify clinically relevant populations that encompass "former" and "current" tobacco users.
[0006] Recently, the United States Preventative Services Task Force (USPSTF) concluded that behavioral and pharmacological interventions significantly increase smoking cessation in non-pregnant individuals (Patnode CD, et al., JAMA, 2021;325(3):280-298). For patients who intend to quit smoking, interventions include counseling or medication; for patients who do not intend to quit, clinicians are encouraged to implement interventions (e.g., providing information and advice) that motivate patients to consider quitting. For patients who have recently quit smoking, clinicians are advised to support patients in preventing relapse, although prevention of relapse is not required in patients who have quit smoking for several years.
[0007] One of several first-line treatments may be prescribed to assist patients in smoking cessation: bupropion SR, nicotine gum, nicotine inhalant, nicotine lozenge, nicotine nasal spray, nicotine patch, or varenicline. Second-line treatments (including clonidine and nortriptyline) are not recommended unless a contraindication to the first-line treatment (such as pregnancy) is identified or the patient has not responded to the first-line treatment (The Clinical Practice Guideline Treating Tobacco Use and Dependence 2008 Update Panel, Liaisons, and Staff. A Clinical Practice Guideline for Treating Tobacco Use and Dependence: 2008 Update, A US Public Health Service Report. Am J Prev Med, 2008;35(2):158-176).
[0008] Cotinine (a nicotine metabolite) is often used to identify "current" smoking habits, but there are currently no biomarkers used in clinical practice that can identify "former" smokers (Centers for Disease Control and Prevention. Biomonitoring Summary: Cotinine. April 2017. Available at www.cdc.gov / biomonitoring / Cotinine_BiomonitoringSummary.html). Epigenetic tests that measure DNA methylation at specific gene locations are available, such as the Smoke Signature® test by Behavioral Diagnostics, LLC, but as commercially available, this test claims only to stratify "current" smokers from "never" smokers (Philibert R, et al., Front Genet, 2018;9:137; Dawes K, et al. Sci Rep, 2021;11(1):21627).
[0009] The development of proteomic models for predicting the probability that an individual is an ever tobacco user would be highly desirable. Such proteomic models would provide insight into current or former (i.e., "ever") tobacco use habits without relying on subjective, misleading, or potentially falsified self-reporting. A need exists for biomarkers, methods, devices, reagents, systems, and kits that allow for the prediction of the probability that an individual is an ever tobacco user. Summary of the Invention
[0010] The present application encompasses biomarkers, methods, reagents, devices, systems, and kits for predicting the probability that an individual is a former tobacco user. In some embodiments, methods are provided for predicting the probability that an individual is a former tobacco user. In some embodiments, methods are provided for detecting the levels of N biomarker proteins in a sample.
[0011] In some embodiments, there is provided a method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of EPHA6 biomarker protein and a level of each of N biomarker proteins in a sample from the subject, where N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 2, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is selected from placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0012] In some embodiments, there is provided a method of detecting levels of N biomarker proteins in a sample, comprising obtaining a sample from a subject and detecting the level of each of N biomarker proteins in the sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and wherein at least one of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0013] In some embodiments, there is provided a method of predicting the probability that a subject is a ever tobacco user, comprising detecting in a sample from the subject a level of placental alkaline phosphatase biomarker protein and a level of each of N biomarker proteins, where N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and wherein at least one of the N biomarker proteins is selected from EPHA6, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0014] In some embodiments, there is provided a method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of an SCF biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, The method is provided wherein the number of biomarker proteins is at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0015] In some embodiments, there is provided a method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of MASP3:heavy chain biomarker protein and a level of each of N biomarker proteins in a sample from the subject, where N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0016] In some embodiments, there is provided a method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of DUS10 biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 2, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0017] In some embodiments, there is provided a method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of an IgG4-κ biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and wherein at least one of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0018] In some embodiments, there is provided a method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of a DCC biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, The method is provided wherein the number of biomarker proteins is at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0019] In some embodiments, a method of predicting the probability that a subject is a former tobacco user comprises detecting a level of RCAN3 biomarker protein and a level of each of N biomarker proteins in a sample from the subject, where N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 2, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0020] In some embodiments, there is provided a method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of CD248 biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 2, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0021] In some embodiments, N is 2 to 39, or N is 3 to 39, or N is 4 to 39, or N is 5 to 39, or N is 6 to 39, or N is 7 to 39, or N is 8 to 39, or N is 9 to 39, or N is 10 to 39, or N is 11 to 39, or N is 12 to 39, or N is 13 to 39, or N is 14 to 39, or N is 15 to 39, or N is 16 to 39, or N is 17 to 39, or N is 18 to 39. or N is 19-39, or N is 20-39, or N is 21-39, or N is 22-39, or N is 23-39, or N is 24-39, or N is 25-39, or N is 26-39, or N is 27-39, or N is 28-39, or N is 29-39, or N is 30-39, or N is 31-39, or N is 32-39, or N is 33-39, or N is 34-39, or N is 35-39, or N is 36-39, or N is 37-39, or N is 38-39. In some embodiments, N is 2, or N is 3, or N is 4, or N is 5, or N is 6, or N is 7, or N is 8, or N is 9, or N is 10, or N is 11, or N is 12, or N is 13, or N is 14, or N is 15, or N is 16, or N is 17, or N is 18, or N is 19, or N is 20, or N is 21, or N is 22, or N is 23, or N is 24, or N is 25, or N is 26, or N is 27, or N is 28, or N is 29, or N is 30, or N is 31, or N is 32, or N is 33, or N is 34, or N is 35, or N is 36, or N is 37, or N is 38, or N is 39.
[0022] In some embodiments, each of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B. In some embodiments, at least one of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, and CD248. In some embodiments, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, or at least 9 of the N protein biomarkers are selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, and CD248.
[0023] In some embodiments, two of the N biomarker proteins are EPHA6 and placental alkaline phosphatase, or two of the N biomarker proteins are EPHA6 and SCF, or two of the N biomarker proteins are EPHA6 and MASP3:heavy chain, or two of the N biomarker proteins are EPHA6 and DUS10, or two of the N biomarker proteins are EPHA6 and IgG4-κ, or two of the N biomarker proteins are EPHA6 and DCC, or two of the N biomarker proteins are EPHA6 and RCAN3, or two of the N biomarker proteins are EPHA6 and CD248. In some embodiments, two of the N biomarker proteins are placental alkaline phosphatase and SCF, or two of the N biomarker proteins are placental alkaline phosphatase and MASP3:heavy chain, or two of the N biomarker proteins are placental alkaline phosphatase and DUS10, or two of the N biomarker proteins are placental alkaline phosphatase and IgG4-κ, or two of the N biomarker proteins are placental alkaline phosphatase and DCC, or two of the N biomarker proteins are placental alkaline phosphatase and RCAN3, or two of the N biomarker proteins are placental alkaline phosphatase and CD248. In some embodiments, two of the N biomarker proteins are SCF and MASP3:heavy chain, or two of the N biomarker proteins are SCF and DUS10, or two of the N biomarker proteins are SCF and IgG4-κ, or two of the N biomarker proteins are SCF and DCC, or two of the N biomarker proteins are SCF and RCAN3, or two of the N biomarker proteins are SCF and CD248.In some embodiments, two of the N biomarker proteins are MASP3:heavy chain and DUS10, or two of the N biomarker proteins are MASP3:heavy chain and IgG4-κ, or two of the N biomarker proteins are MASP3:heavy chain and RCAN3, or two of the N biomarker proteins are MASP3:heavy chain and CD248, or two of the N biomarker proteins are MASP3:heavy chain and DCC. In some embodiments, two of the N biomarker proteins are DUS10 and IgG4-κ, or two of the N biomarker proteins are DUS10 and DCC, or two of the N biomarker proteins are DUS10 and CD248, or two of the N biomarker proteins are DUS10 and RCAN3. In some embodiments, two of the N biomarker proteins are IgG4-κ and DCC, or two of the N biomarker proteins are IgG4-κ and RCAN3, or two of the N biomarker proteins are IgG4-κ and CD248. In some embodiments, two of the N biomarker proteins are DCC and RCAN3, or two of the N biomarker proteins are DCC and CD248. In some embodiments, two of the N biomarker proteins are RCAN3 and CD248.
[0024] In some embodiments, methods are provided for predicting the probability that a subject is a former tobacco user, comprising detecting a level of T11L1 biomarker protein. In some embodiments, methods are provided for predicting the probability that a subject is a former tobacco user, comprising detecting a level of MXRA8 biomarker protein. In some embodiments, methods are provided for predicting the probability that a subject is a former tobacco user, comprising detecting a level of XTP3A biomarker protein. In some embodiments, methods are provided for predicting the probability that a subject is a former tobacco user, comprising detecting a level of CRAC1 biomarker protein. In some embodiments, the method further comprises detecting at least one biomarker protein selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit.
[0025] In some embodiments, the sample is a blood sample, a plasma sample, or a serum sample. In some embodiments, detection is accomplished using mass spectrometry, an aptamer-based assay, and / or an antibody-based assay. In some embodiments, the method comprises contacting biomarker proteins of the sample(s) with a set of biomarker capture reagents, wherein each biomarker capture reagent of the set of biomarker capture reagents specifically binds to a different biomarker protein being detected. In some embodiments, each biomarker capture reagent is an antibody or an aptamer. In some embodiments, each biomarker capture reagent is an aptamer. In some embodiments, at least one aptamer is a slow-off aptamer. In some embodiments, the at least one slow-off aptamer comprises nucleotides with at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten modifications. In some embodiments, each slow off aptamer has a dissociation rate (t) of ≥ 20 minutes, ≥ 30 minutes, ≥ 60 minutes, ≥ 90 minutes, ≥ 120 minutes, ≥ 150 minutes, ≥ 180 minutes, ≥ 210 minutes, or ≥ 240 minutes. 1 / 2 ) and binds to its target protein. In some embodiments, the level of each measured biomarker protein is determined from relative fluorescence units (RFU) or protein concentration.
[0026] In some embodiments, predicting the probability that a subject is a ever tobacco user is based on inputting the measured levels of N biomarker proteins into a statistical model. In some embodiments, the determining includes analyzing the levels of the N biomarker proteins using an elastic net logistic regression model. In some embodiments, the model has an area under the curve (AUC) selected from at least 0.65, at least 0.66, at least 0.67, at least 0.68, at least 0.69, at least 0.7, at least 0.75, at least 0.8, at least 0.85, at least 0.9, or at least 0.95. In some embodiments, the model provides an absolute risk probability of ever being a tobacco user. In some embodiments, the model provides a value between 0 and 1, where a value >0.460 is predictive of ever being a tobacco user. In some embodiments, the model provides an absolute probability of ever being a tobacco user based on the levels of each of the biomarker proteins selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, and CD248. In some embodiments, the model provides an absolute probability of being an ever tobacco user based on the level of each of the biomarker proteins selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0027] In some embodiments, a method is provided for predicting the probability that a subject is a former tobacco user for purposes of determining medical or life insurance premiums. In some embodiments, the method further comprises determining medical or life insurance coverage or premiums. In some embodiments, a method is provided further comprising using information resulting from the method to predict and / or manage healthcare resource utilization. In some embodiments, the method further comprises using information resulting from the method to facilitate decisions to acquire or purchase healthcare businesses, hospitals, or companies. In any of the embodiments, the tobacco user is a smoker.
[0028] In some embodiments, a kit comprises N biomarker protein capture reagents, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38 or at least 39, wherein at least one of the N biomarker protein capture reagents specifically binds to a biomarker protein selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
[0029] In some embodiments, a kit is provided that includes N biomarker protein capture reagents for performing any of the methods described herein. In some embodiments, each of the N biomarker protein capture reagents is an antibody or an aptamer. In some embodiments, each biomarker protein capture reagent is an aptamer. In some embodiments, at least one aptamer is a slow-off aptamer. In some embodiments, at least one slow-off aptamer comprises nucleotides with at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 modifications. In some embodiments, each slow-off aptamer has a dissociation rate (t) of ≥ 20 minutes, ≥ 30 minutes, ≥ 60 minutes, ≥ 90 minutes, ≥ 120 minutes, ≥ 150 minutes, ≥ 180 minutes, ≥ 210 minutes, or ≥ 240 minutes. 1 / 2 ) binds to its target protein. In some embodiments, the kit is for use in detecting N biomarker proteins in a sample from a subject. In some embodiments, the kit is for use in predicting the probability that a subject is a former tobacco user. [Brief explanation of the drawings]
[0030] [Figure 1] Receiver operating characteristic (ROC) curves for the tobacco use status reference model for the training, validation, and verification datasets are shown. [Figure 2] Boxplots of predicted probabilities stratified by dataset (training, validation, or validation) and classification of tobacco use as "ever" vs. "never." The horizontal black line corresponds to the decision cutoff of 0.46. [Figure 3-1]
[0023] Illustrated are specific exemplary modifications in aptamers that may occur at the 5-position of uridine. The chemical structures of C-5 modifications include exemplary amide linkages connecting the modifications to the 5-position of uridine. Illustrated 5-position moieties include benzyl moieties (e.g., Bn, PE, and PP), naphthyl moieties (e.g., Nap, 2Nap, NE), butyl moieties (e.g., iBu), fluorobenzyl moieties (e.g., FBn), tyrosyl moieties (e.g., Tyr), 3,4-methylenedioxybenzyl (e.g., MBn), morpholino moieties (e.g., MOE), benzofuranyl moieties (e.g., BF), indole moieties (e.g., Trp), and hydroxypropyl moieties (e.g., Thr). [Figure 3-2]
[0023] Illustrated are specific exemplary modifications in aptamers that may occur at the 5-position of uridine. The chemical structures of C-5 modifications include exemplary amide linkages connecting the modifications to the 5-position of uridine. Illustrated 5-position moieties include benzyl moieties (e.g., Bn, PE, and PP), naphthyl moieties (e.g., Nap, 2Nap, NE), butyl moieties (e.g., iBu), fluorobenzyl moieties (e.g., FBn), tyrosyl moieties (e.g., Tyr), 3,4-methylenedioxybenzyl (e.g., MBn), morpholino moieties (e.g., MOE), benzofuranyl moieties (e.g., BF), indole moieties (e.g., Trp), and hydroxypropyl moieties (e.g., Thr). [Figure 3-3]
[0023] Illustrated are specific exemplary modifications in aptamers that may occur at the 5-position of uridine. The chemical structures of C-5 modifications include exemplary amide linkages connecting the modifications to the 5-position of uridine. Illustrated 5-position moieties include benzyl moieties (e.g., Bn, PE, and PP), naphthyl moieties (e.g., Nap, 2Nap, NE), butyl moieties (e.g., iBu), fluorobenzyl moieties (e.g., FBn), tyrosyl moieties (e.g., Tyr), 3,4-methylenedioxybenzyl (e.g., MBn), morpholino moieties (e.g., MOE), benzofuranyl moieties (e.g., BF), indole moieties (e.g., Trp), and hydroxypropyl moieties (e.g., Thr). [Figure 4-1]Specific exemplary modifications in aptamers that may occur at the 5-position of cytidine are shown. The chemical structures of C-5 modifications include exemplary amide linkages connecting the modifications to the 5-position of cytidine. The 5-position moieties shown include benzyl moieties (e.g., Bn, PE, and PP), naphthyl moieties (e.g., Nap, 2Nap, NE, and 2NE), and tyrosyl moieties (e.g., Tyr). [Figure 4-2] Specific exemplary modifications in aptamers that may occur at the 5-position of cytidine are shown. The chemical structures of C-5 modifications include exemplary amide linkages connecting the modifications to the 5-position of cytidine. The 5-position moieties shown include benzyl moieties (e.g., Bn, PE, and PP), naphthyl moieties (e.g., Nap, 2Nap, NE, and 2NE), and tyrosyl moieties (e.g., Tyr). [Figure 5] 1 illustrates an exemplary computer system for use with various computer-implemented methods described herein. [Figure 6] 1 is a flowchart for a method of assessing the probability that an individual is an ever tobacco user, according to one or more embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0031] Reference will now be made in detail to exemplary embodiments of the present invention. While the invention will be described in conjunction with enumerated specific embodiments, it will be understood that it is not intended that the invention be limited to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents which may be included within the scope of the present invention as defined by the claims.
[0032] One skilled in the art will recognize many methods and materials similar or equivalent to those described herein, which could be used in and are within the scope of the present invention. The present invention is in no way limited to the methods and materials described.
[0033] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods, devices, and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, particular methods, devices, and materials are described herein.
[0034] All publications, published patent documents, and patent applications cited in this application are indicative of the level of skill in the technical field(s) to which this application pertains. All publications, published patent documents, and patent applications cited herein are incorporated by reference to the same extent as if each individual publication, published patent document, or patent application was specifically and individually indicated to be incorporated by reference.
[0035] As used in this application, including the appended claims, the singular forms "a," "an," and "the" include plural references unless the content clearly dictates otherwise, and are used interchangeably with "at least one" and "one or more." Thus, reference to a "SOMAmer" includes a mixture of SOMAmers, reference to a "probe" includes a mixture of probes, and so forth.
[0036] As used herein, the term "about" refers to an insignificant modification or variation of a numerical value such that the basic function of the item to which the numerical value is associated remains unchanged.
[0037] As used herein, the terms "comprise," "including," "include," "including," "containing," "containing," and any variations thereof are intended to cover a non-exclusive inclusion, so that a process, method, product-by-process, or composition of matter that includes, includes, or contains elements, or a list of elements, does not include only those elements, but may also include other elements not expressly listed or inherent in such process, method, product-by-process, or composition of matter.
[0038] The present application encompasses biomarkers, methods, devices, reagents, systems, and kits for determining the probability that an individual is a former tobacco user.
[0039] "Tobacco use" as used herein includes combustible tobacco and chewing tobacco. "Tobacco use" as used herein does not include e-cigarettes.
[0040] "Biological sample," "sample," and "test sample" are used interchangeably herein to refer to any material, biological fluid, tissue, or cell obtained from or otherwise derived from an individual. The terms encompass blood (including whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, and serum), dried blood spots (e.g., from an infant), sputum, tears, mucus, nasal washes, nasal aspirates, exhaled breath, urine, semen, saliva, peritoneal washings, ascites, cyst fluid, cerebrospinal fluid, glandular fluid, pancreatic juice, lymphatic fluid, pleural effusion, nipple aspirate, bronchial aspirate, bronchial brush, synovial fluid, joint aspirate, organ secretions, cells, cell extracts, and cerebrospinal fluid. The terms also encompass experimentally isolated fractions of all of the foregoing. For example, a blood sample can be fractionated into serum, plasma, or fractions containing specific types of blood cells, such as red blood cells or white blood cells (leukocytes). If desired, a sample can be a combination of samples from an individual, such as a combination of a tissue sample and a fluid sample. The term "biological sample" also encompasses materials containing homogenized solid material (e.g., from a stool sample, tissue sample, or tissue biopsy). The term "biological sample" also encompasses materials derived from tissue or cell culture. Any suitable method for obtaining a biological sample can be used. Exemplary methods include, for example, phlebotomy, swabs (e.g., buccal swabs), and fine needle aspiration biopsy procedures. Exemplary tissues amenable to fine needle aspiration include lymph nodes, lung, lung lavage fluid, BAL (bronchoalveolar lavage fluid), thyroid, breast, pancreas, and liver. Samples may also be collected, for example, by microdissection (e.g., laser capture microdissection (LCM) or laser microdissection (LMD)), bladder washing, smear (e.g., PAP smear), or breast ductal lavage. A "biological sample" obtained from or derived from an individual includes any such sample that has been processed in any suitable manner after being obtained from the individual.
[0041] For purposes of this specification, the phrase "data attributable to a biological sample from an individual" is intended to mean data in some form that is derived from or generated using an individual's biological sample. Although the data may be reformatted, modified, or mathematically altered to some extent after it is generated, such as by converting units from one measurement system to another, the data is understood to be derived from or generated using a biological sample.
[0042] "Target," "target molecule," and "analyte" are used interchangeably herein to refer to any molecule of interest that may be present in a biological sample. "Molecule of interest" encompasses any minor variation of a particular molecule (e.g., in the case of proteins, slight variations in amino acid sequence, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, etc.) or other manipulation or modification (e.g., conjugation with a labeling component) that does not substantially alter the identity of the molecule. A "target molecule," "target," or "analyte" is a set of copies of one type or species of molecule or multimolecular structure. "Multiple target molecules," "multiple targets," and "multiple analytes" refer to a set of more than one such molecule. Exemplary target molecules include proteins, polypeptides, nucleic acids, carbohydrates, lipids, polysaccharides, glycoproteins, hormones, receptors, antigens, antibodies, affibodies, antibody mimics, viruses, pathogens, toxicants, substrates, metabolites, transition-state analogs, cofactors, inhibitors, drugs, dyes, nutrients, growth factors, cells, tissues, and fragments or portions of any of the above. In some embodiments, the target molecule is a protein, in which case the target molecule may be referred to as a "target protein."
[0043] As used herein, "capture agent" or "capture reagent" refers to a molecule capable of specifically binding to a biomarker. "Target protein capture reagent" refers to a molecule capable of specifically binding to a target protein. Non-limiting exemplary capture reagents include aptamers, antibodies, adnectins, ankyrins, other antibody mimetics and other protein scaffolds, autoantibodies, chimeras, small molecules, nucleic acids, lectins, ligand-binding receptors, imprinted polymers, avimers, peptidomimetics, hormone receptors, cytokine receptors, synthetic receptors, and modifications and fragments of any of the foregoing capture reagents. In some embodiments, the capture reagent is selected from an aptamer and an antibody.
[0044] As used herein, the terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. Polymers can be linear or branched, can contain modified amino acids, and can be interrupted by non-amino acids. The terms also encompass amino acid polymers that are modified naturally or by intervention, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or other manipulation or modification (such as conjugation with a labeling component). Also included within the definition are polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids) and other modifications known in the art. Polypeptides can be single chains or associated chains. Also included within the definition are preproteins and intact mature proteins; peptides or polypeptides derived from mature proteins; fragments of proteins; splice variants; recombinant forms of proteins; protein variants with amino acid modifications, deletions, or substitutions; digests; and post-translational modifications (such as glycosylation, acetylation, phosphorylation, and the like).
[0045] The term "antibody" refers to full-length antibodies of any species, as well as fragments and derivatives of such antibodies, including Fab and F(ab')2 fragments, single-chain antibodies, Fv fragments, and single-chain Fv fragments. The term "antibody" also refers to synthetically derived antibodies, such as phage-display-derived antibodies and fragments, affibodies, nanobodies, etc.
[0046] As used herein, "marker," "biomarker," and "feature" are used interchangeably to refer to a target molecule that indicates or is symptomatic of a normal or abnormal process in an individual, or a disease or other condition in an individual. More specifically, a "marker" or "biomarker" or "feature" is an anatomical, physiological, biochemical, or molecular parameter that correlates with the presence of a particular physiological state or process, whether normal or abnormal, and if abnormal, whether inert or acute. Biomarkers are detectable and measurable by a variety of methods, including laboratory assays and medical imaging. When a biomarker is a protein, it is also possible to use the expression of the corresponding gene or the methylation status of the gene encoding the biomarker or the protein that controls the expression of the biomarker as a surrogate measure of the amount or presence or absence of the corresponding protein biomarker in a biological sample. In certain embodiments, a feature is an analyte / SOMAmer reagent that is another predictor in a statistical model.
[0047] As used herein, "biomarker value," "value," "biomarker level," "feature level," and "level" are used interchangeably to refer to a measurement made using any analytical method for the detection of a biomarker in a biological sample and indicating the presence, absence, absolute amount or concentration, relative amount or concentration, titer, level, expression level, ratio of measured levels, or the like, of, or corresponding to, a biomarker in a biological sample. The exact nature of a "value" or "level" will depend on the specific design and components of the particular analytical method used to detect the biomarker.
[0048] When a biomarker indicates or is symptomatic of an abnormal process or disease or other condition in an individual, the biomarker is generally described as being either overexpressed or underexpressed compared to an expression level or value of the biomarker that indicates or is symptomatic of the absence of a normal process or disease or other condition in the individual. "Upregulation," "upregulated," "overexpression," "overexpressed," and any variations thereof are used interchangeably to refer to a value or level of a biomarker in a biological sample that exceeds the value or level (or range of values or levels) of the biomarker typically detected in a similar biological sample from a healthy or normal individual. The term can also refer to a value or level of a biomarker in a biological sample that exceeds the value or level (or range of values or levels) of the biomarker that can be detected at different stages of a particular disease.
[0049] "Downregulation," "downregulated," "underexpression," "underexpressed," and any variations thereof, are used interchangeably to refer to a value or level of a biomarker in a biological sample that is less than the value or level (or range of values or levels) of the biomarker typically detected in a similar biological sample from a healthy or normal individual. The terms may also refer to a value or level of a biomarker in a biological sample that is less than the value or level (or range of values or levels) of the biomarker that can be detected at different stages of a particular disease.
[0050] Furthermore, a biomarker that is either over- or under-expressed may also be referred to as being "differentially expressed" or having a "differential level" or "differential value" compared to a "normal" expression level or value of the biomarker that is indicative of, or symptomatic of, a normal process or the absence of a disease or other pathological condition in an individual. Thus, "differential expression" of a biomarker may also be referred to as a variation from the "normal" expression level of the biomarker.
[0051] The terms "differential gene expression" and "differential expression" are used interchangeably to refer to a gene (or its corresponding protein expression product) whose expression is activated to a higher or lower level in a subject suffering from a particular disease or condition compared to its expression in a normal or control subject. The term also encompasses genes (or their corresponding protein expression products) whose expression is activated to a higher or lower level at different stages of the same disease or condition. It is also understood that a differentially expressed gene may be activated or inhibited at the nucleic acid or protein level, or may undergo alternative splicing to result in a different polypeptide product. Such differences may be evidenced by various changes, including mRNA levels, surface expression, secretion, or other distribution of polypeptides. Differential gene expression may encompass a comparison of expression between two or more genes or their gene products; or a comparison of the ratio of expression between two or more genes or their gene products; or even a comparison of two differentially processed products of the same gene that differ between normal and diseased subjects; or between various stages of the same disease. Differential expression encompasses both quantitative and qualitative differences in the temporal or cellular expression patterns of a gene or its expression products, for example, among normal and diseased cells, or among cells that have undergone different disease events or stages.
[0052] A "control level" of a target molecule refers to the level of the target molecule in properly handled samples of the same sample type. A control level may refer to the average level of the target molecule in properly handled samples from a population of individuals.
[0053] As used herein, "individual" refers to a test subject or patient. An individual can be a mammal or a non-mammal. In various embodiments, the individual is a mammal. A mammalian individual can be a human or a non-human. In various embodiments, the individual is a human.
[0054] "Diagnosing," "diagnosis," and variations thereof refer to detecting, determining, or recognizing an individual's health or condition based on one or more signs, symptoms, data, or other information pertaining to that individual. An individual's health may be diagnosed as healthy / normal (i.e., diagnosing the absence of a disease or condition) or as diseased / abnormal (i.e., diagnosing the presence of a disease or condition or assessing the characteristics of a disease or condition). The terms "diagnosing," "diagnosing," "diagnosis," and the like, with respect to a particular disease or condition, encompass the initial detection of disease; the characterization or classification of disease; the detection of disease progression, remission, or recurrence; and the detection of disease response after administration of a treatment or therapy to an individual. The probability of ever being a tobacco user encompasses distinguishing individuals who were ever tobacco users (former or current) from individuals who were never tobacco users.
[0055] "Prognosticate," "prognosing," "prognosis," and variations thereof refer to predicting the future course of a disease or condition in an individual who has the disease or condition (e.g., predicting patient survival), and such terms encompass the assessment of the response of a disease or condition following administration of a treatment or therapy to an individual.
[0056] "Evaluate," "assessing," "evaluation," and variations thereof encompass both "diagnosing" and "prognosing," and also encompass the determination or prediction of the future course of a disease or condition in disease-free individuals, and the determination or prediction of the risk of the disease or condition recurring in individuals who have apparently been cured of the disease or have resolved the condition. The term "evaluating" also encompasses assessing an individual's response to therapy (e.g., predicting whether an individual is likely to respond favorably or unlikely to respond to a therapeutic agent (or, for example, will experience toxic effects or other undesirable side effects), selecting a therapeutic agent for administration to an individual, or monitoring or determining an individual's response to a therapy administered to an individual, etc.).
[0057] As used herein, "additional biomedical information" refers to one or more assessments of an individual other than using any of the biomarkers described herein associated with tobacco use. "Additional biomedical information" includes any of the following: physical descriptors of the individual, including the individual's height and / or weight; the individual's age; the individual's sex; weight change; the individual's ethnicity; occupational history; family history of tobacco use; the presence of genetic marker(s); clinical symptoms (abdominal pain, weight gain or loss, gene expression values, etc.); physical descriptors of the individual (including physical descriptors observed by radiological imaging); tobacco use status; history of alcohol use; occupational history; dietary habits (salt, saturated fat, and cholesterol intake); caffeine consumption; and imaging information. Additional biomedical information can be obtained from the individual (such as from the individual themselves, using routine patient questionnaires or health history questionnaires, or from a healthcare professional) using routine techniques known in the art.
[0058] As used herein, "detecting" or "determining" with respect to a biomarker value encompasses both the use of the equipment required to observe and record a signal corresponding to the biomarker value as well as the material(s) required to generate that signal. In various embodiments, the biomarker value is detected using any suitable method, including fluorescence, chemiluminescence, surface plasmon resonance, surface acoustic waves, mass spectrometry, infrared spectroscopy, Raman spectroscopy, atomic force microscopy, scanning tunneling microscopy, electrochemical detection, nuclear magnetic resonance, quantum dots, and the like.
[0059] "Solid support" herein refers to any substrate having a surface to which molecules can be directly or indirectly attached via either covalent or non-covalent bonds. A "solid support" can have a variety of physical formats, including, for example, membranes; chips (e.g., protein chips); slides (e.g., glass slides or coverslips); columns; hollow, solid, semi-solid, hole-containing, or cavity-containing particles (e.g., beads); gels; fibers (including fiber optic materials); matrices; and sample vessels. Exemplary sample vessels include sample wells, tubes, capillaries, vials, and other vessels, grooves, or depressions that can hold a sample. Sample vessels can be contained in multi-sample platforms (e.g., microtiter plates, slides, microfluidic devices, and the like). Supports can be composed of natural or synthetic, organic, or inorganic materials. The composition of the solid support to which the capture reagent is attached generally depends on the method of attachment (e.g., covalent bonding). Other exemplary vessels include microdroplets, microfluidically controlled emulsions, or bulk oil / water emulsions in which assays and associated manipulations can occur. Suitable solid supports include, for example, plastics, resins, polysaccharides, silica or silica-based materials, functionalized glass, modified silicon, carbon, metals, inorganic glass, membranes, nylon, natural fibers (e.g., silk, wool, cotton, etc.), polymers, and the like. The material comprising the solid support may contain reactive groups (e.g., carboxyl, amino, or hydroxyl groups) used for binding capture reagents. Polymeric solid supports may include, for example, polystyrene, polyethylene glycol tetraphthalate, polyvinyl acetate, polyvinyl chloride, polyvinylpyrrolidone, polyacrylonitrile, polymethyl methacrylate, polytetrafluoroethylene, butyl rubber, styrene butadiene rubber, natural rubber, polyethylene, polypropylene, (poly)tetrafluoroethylene, (poly)vinylidene fluoride, polycarbonate, and polymethylpentene.Suitable solid support particles that can be used include, for example, encoded particles (such as Luminex® type encoded particles), magnetic particles, and glass microparticles.
[0060] As used herein, an "analyte" is a protein target of a capture reagent. In certain embodiments, the capture reagent is an aptamer. In certain further embodiments, the capture reagent is a SOMAmer.
[0061] As used herein, "Lin's CCC" refers to the concordance correlation coefficient that measures the agreement between a new test and an existing test that is considered the gold standard.
[0062] As used herein, a "test" refers to a set of samples and clinical data that are analyzed to derive a test.
[0063] As used herein, "training data set" means a subset of data from a study that is used to fit a model.
[0064] As used herein, "validation dataset" means the final subset of data used to assess the performance of the selected model developed on the validation dataset.
[0065] As used herein, a "validation dataset" means a separate subset of data used to provide an unbiased evaluation of a model fitted to a training dataset while adjusting model parameters.
[0066] As used herein, the term "need" or "required" refers to a judgment made by a healthcare provider regarding treatment of a patient that is determined by the healthcare provider to be beneficial to the patient's health condition.
[0067] Risk Analysis In some embodiments, disclosed herein are objective tests for predicting the probability that an individual is a former tobacco user.
[0068] The risk analysis profile can be described as in Table 1. [Table 1]
[0069] The testing methods disclosed herein may provide insight into current or former (i.e., "ever") tobacco use habits without relying on subjective, misleading, or falsified self-reporting. The testing methods disclosed herein may influence positive behavioral changes (i.e., tobacco use cessation) that may prevent or delay tobacco use-related pathologies. The testing methods disclosed herein may enable health care providers and patients to monitor tobacco use habits, adherence to tobacco use cessation, and the effects of tobacco use on the proteome over time. The testing methods disclosed herein may be considered for individual risk assessment for disease screening eligibility. [Table 2]
[0070] The performance threshold was established based on the performance of DNA methylation of the AHRR gene, the leading commercially available epigenetic prediction model for predicting smoking habits (Langdon RJ, et al., Clin Epigenetics, 2021;13(1):206; Philibert R, et al., Front Genet, 2018;9:137).
[0071] There are no commercially available tests that predict "former" and "current" tobacco use status. Cotinine (a nicotine metabolite) is commonly used in clinical settings to predict recent (i.e., <3 days) tobacco use (Centers for Disease Control and Prevention. Biomonitoring Summary: Cotinine. April 2017. Available at https: / / www.cdc.gov / biomonitoring / Cotinine_BiomonitoringSummary.html). However, cotinine is not a comparable reference tool because it does not indicate the long-term biological effects of tobacco smoke exposure. The test methods disclosed herein at least rival the performance of current epigenetic prediction models; for example, the Smoke Signature® test offered by Behavioral Diagnostics, LLC claims to stratify "current" smokers vs. non-smokers only using a cutoff for DNA methylation of the AHRR gene (Philibert R, et al., Front Genet, 2018;9:137; Dawes K, et al., Sci Rep, 2021;11(1):21627). Recently, AHRR methylation was applied to a two-stage approach to assess its ability to classify three-component smoking status, with AUCs ranging from 0.717 to 0.902 (Langdon RJ, et al., Clin Epigenetics, 2021;13(1):206). While epigenetic changes can be quantitative (Philibert R, et al., Front Psychiatry, 2016;7:55), these changes require significant alterations in response to smoke exposure and take extended periods (i.e., years) to show measurable differences in smoking status classes (McCartney DL, et al., EBioMedicine, 2018;37:214-220), and therefore, dynamic proteomic assessment may be more predictive of more rapid changes in smoking habits.Embodiments of the methods disclosed herein perform with an AUC of 0.65 or greater, 0.66, 0.67, 0.68, 0.69, 0.70, and greater than 0.717, with 0.717 being the lower limit of AHRR methylation performance reported for ternary classification.
[0072] In some embodiments, the number of biomarkers useful in a biomarker subset or panel is based on the sensitivity and specificity values for a particular combination of biomarker levels. The terms "sensitivity" and "specificity" are used herein to refer to the ability to accurately classify an individual as an ever smoker or a never smoker based on one or more biomarker levels detected in a biological sample. "Sensitivity" refers to the performance of a biomarker(s) in accurately classifying an individual as an ever (current and previous) smoker or a never smoker. "Specificity" refers to the performance of a biomarker(s) in accurately classifying individuals who do not smoke.
[0073] In some embodiments, the overall performance of a panel of one or more biomarkers is represented by an area under the curve (AUC) value. AUC values are derived from a receiver operating characteristic (ROC) curve. An ROC curve is a plot of the true positive rate (sensitivity) of a test against the false positive rate (1-specificity) of the test. The terms "area under the curve" or "AUC" refer to the area under a receiver operating characteristic (ROC) curve, both of which are well known in the art. AUC measurements are useful for comparing the accuracy of classifiers across the complete data range. A classifier with a larger AUC has a higher ability to accurately classify unknowns between two groups of interest (e.g., ever tobacco users (former and current) and never tobacco users). ROC curves are useful for plotting the performance of a particular feature (e.g., any of the biomarkers described herein and / or any item of additional biomedical information) in distinguishing two populations. Typically, feature data across populations is sorted in ascending order based on the value of a single feature. Then, for each value of the feature, the true positive rate and false positive rate of the data are calculated. The true positive rate is determined by counting the number of cases that exceed the value for the feature and then dividing by the total number of cases. The false positive rate is determined by counting the number of controls that exceed the value for the feature and then dividing by the total number of controls. While this definition refers to a scenario in which the feature is elevated compared to the control, this definition also applies to scenarios in which the feature is lower compared to the control (in such a scenario, samples below the value for the feature would be counted). ROC curves can be generated for single features and other single outputs; for example, a combination of two or more features can be mathematically combined (e.g., added, subtracted, multiplied, etc.) to provide a single total value, which can be plotted in the ROC curve. Additionally, any combination of multiple features whose combination leads to a single output value can be plotted in the ROC curve.
[0074] Exemplary Uses of Biomarkers In various exemplary embodiments, methods are provided for predicting the probability that an individual is a former tobacco user by detecting one or more biomarker values corresponding to one or more biomarkers present in the individual's circulation (such as in blood, serum, or plasma) by a number of analytical methods (including any of the analytical methods described herein). These biomarkers are differentially expressed in individuals who are former (former and current) tobacco users compared to, for example, individuals who have never smoked. Detection of differential expression of biomarkers in an individual can be used, for example, to enable prediction of the probability of being a former tobacco user.
[0075] In addition to examining biomarker levels as a stand-alone diagnostic test, biomarker levels can also be performed in conjunction with determining SNPs or other genetic lesions or variations that represent an increased risk of susceptibility to a disease or condition (see, e.g., Amos et al., Nature Genetics 40, 616-622 (2009)).
[0076] Any of the described biomarkers can be used in imaging studies. For example, imaging agents can be coupled to any of the described biomarkers, which can be used to aid in predicting the probability that an individual is a former tobacco user, to monitor response to therapeutic interventions, to select target populations in clinical trials, among other uses.
[0077] Detection and Determination of Biomarkers and Biomarker Levels Biomarker levels for the biomarkers described herein can be detected using any of a variety of known analytical methods. In one embodiment, biomarker levels are detected using a capture reagent. As used herein, "capture agent" or "capture reagent" refers to a molecule capable of specifically binding to a biomarker. In various embodiments, the capture reagent can be exposed to the biomarker in solution, or the capture reagent can be exposed to the biomarker while immobilized on a solid support. In other embodiments, the capture reagent contains features reactive with secondary features on the solid support. In these embodiments, the capture reagent can be exposed to the biomarker in solution, and then the features on the capture reagent can be used in conjunction with the secondary features on the solid support to immobilize the biomarker on the solid support. The capture reagent is selected based on the type of analysis to be performed. Capture reagents include, but are not limited to, SOMAmers, antibodies, adnectins, ankyrins, other antibody mimetics and other protein scaffolds, autoantibodies, chimeras, small molecules, F(ab')2 fragments, single chain antibody fragments, Fv fragments, single chain Fv fragments, nucleic acids, lectins, ligand binding receptors, affibodies, nanobodies, imprinted polymers, avimers, peptidomimetics, hormone receptors, cytokine receptors and synthetic receptors, and modifications and fragments thereof.
[0078] In some embodiments, biomarker levels are detected using a biomarker / capture reagent complex.
[0079] In other embodiments, the biomarker level is derived from the biomarker / capture reagent complex and is detected indirectly, such as as a result of a reaction subsequent to the biomarker / capture reagent interaction, but dependent on the formation of the biomarker / capture reagent complex.
[0080] In some embodiments, the biomarker level is detected directly from the biomarker in the biological sample.
[0081] In one embodiment, biomarkers are detected using a multiplex format that allows for simultaneous detection of two or more biomarkers in a biological sample. In one embodiment of a multiplex format, capture reagents are immobilized directly or indirectly by covalent or non-covalent attachment to discrete locations on a solid support. In another embodiment, the multiplex format uses discrete solid supports, each with a unique capture reagent (e.g., quantum dots, etc.) associated with it. In another embodiment, a separate device is used for the detection of each of the multiple biomarkers to be detected in a biological sample. The separate device can be configured to allow each biomarker in a biological sample to be processed simultaneously. For example, a microtiter plate can be used such that each well in the plate is used to uniquely analyze one of the multiple biomarkers to be detected in a biological sample.
[0082] In one or more of the foregoing embodiments, a component of the biomarker / capture complex may be labeled using a fluorescent tag to enable detection of the biomarker value. In various embodiments, a fluorescent label may be conjugated to a capture reagent specific for any of the biomarkers described herein using known techniques, and the fluorescent label may then be used to detect the corresponding biomarker value. Suitable fluorescent labels include rare earth chelates, fluorescein and its derivatives, rhodamine and its derivatives, dansyl, allophycocyanin, PBXL-3, Qdot 605, Lissamine, phycoerythrin, Texas Red, and other such compounds.
[0083] In one embodiment, the fluorescent label is a fluorescent dye molecule. In some embodiments, the fluorescent dye molecule comprises at least one substituted indolium ring system in which a substituent on the 3-carbon of the indolium ring contains a chemically reactive group or a conjugated substance. In some embodiments, the dye molecule comprises an AlexFluor molecule, such as AlexaFluor 488, AlexaFluor 532, AlexaFluor 647, AlexaFluor 680, or AlexaFluor 700. In other embodiments, the dye molecule comprises a first type and a second type of dye molecule (e.g., two different AlexaFluor molecules). In other embodiments, the dye molecule comprises a first type and a second type of dye molecule, wherein the two dye molecules have different emission spectra.
[0084] Fluorescence can be measured by a variety of instruments compatible with a wide range of assay formats. For example, spectrofluorometers are designed to analyze microtiter plates, microscope slides, printed arrays, cuvettes, etc. See, Principles of Fluorescence Spectroscopy, by J.R. Lakowicz, Springer Science + Business Media, Inc., 2004; and Bioluminescence & Chemiluminescence: Progress & Current Applications; Philip E. Stanley and Larry J. Kricka editors, World Scientific Publishing Company, January 2002.
[0085] In one or more of the foregoing embodiments, a chemiluminescent tag may optionally be used to label a component of the biomarker / capture complex to enable detection of the biomarker value. Suitable chemiluminescent materials include any of oxalyl chloride, rhodamine 6G, Ru(bipy)32+, TMAE (tetrakis(dimethylamino)ethylene), pyrogallol (1,2,3-trihydroxibenzene), lucigenin, peroxyoxalates, aryloxalates, acridinium esters, dioxetanes, and others.
[0086] In yet other embodiments, the detection method involves an enzyme / substrate combination that generates a detectable signal corresponding to the biomarker value. Generally, the enzyme catalyzes a chemical alteration of a chromogenic substrate that can be measured using a variety of techniques, including spectrophotometry, fluorescence, and chemiluminescence. Suitable enzymes include, for example, luciferase, luciferin, malate dehydrogenase, urease, horseradish peroxidase (HRPO), alkaline phosphatase, β-galactosidase, glucoamylase, lysozyme, glucose oxidase, galactose oxidase and glucose-6-phosphate dehydrogenase, uricase, xanthine oxidase, lactoperoxidase, microperoxidase, and the like.
[0087] In yet other embodiments, the detection method may be a combination of fluorescent, chemiluminescent, or radionuclide, or enzyme / substrate combinations that generate a measurable signal. Multimodal signal generation can be a unique and advantageous feature in biomarker assay formats.
[0088] More specifically, biomarker levels for the biomarkers described herein can be detected using known analytical methods, including singleplex SOMAmer assays, multiplex SOMAmer assays, singleplex or multiplex immunoassays, mRNA expression profiling, miRNA expression profiling, mass spectrometry, histological / cytological methods, etc., as described in more detail below.
[0089] Determining biomarker levels using aptamer-based assays Assays aimed at detecting and quantifying physiologically significant molecules in biological and other samples are important tools in scientific research and healthcare. One class of such assays involves the use of microarrays containing one or more aptamers immobilized on a solid support. Each aptamer is capable of binding to a target molecule in a highly specific manner and with very high affinity. See, e.g., U.S. Pat. No. 5,475,096, entitled "Nucleic Acid Ligands." See also, e.g., U.S. Pat. Nos. 6,242,246, 6,458,543, and 6,503,715, each entitled "Nucleic Acid Ligand Diagnostic Biochip." Once the microarray is contacted with a sample, the aptamers bind to their respective target molecules present in the sample, thereby enabling the determination of biomarker values corresponding to the biomarkers.
[0090] As used herein, "aptamer" refers to a nucleic acid that has specific binding affinity to a target molecule. While it is recognized that affinity interactions are a matter of degree, in this context, the "specific binding affinity" of an aptamer for its target generally means that the aptamer binds to its target with a degree of affinity that is much higher than that with which it binds to other components in a test sample. An "aptamer" is a set of copies of one type or species of nucleic acid molecule having a specific nucleotide sequence. An aptamer can contain any suitable number of nucleotides (including any number of chemically modified nucleotides). "Aptamers" refers to a set of more than one such molecule. Different aptamers can have either the same or different numbers of nucleotides. Aptamers can be DNA or RNA or chemically modified nucleic acids, and can be single-stranded, double-stranded, or contain double-stranded regions and higher-order structures. Aptamers can also be photoaptamers, where photoreactive or chemically reactive functional groups are included in the aptamer, allowing the aptamer to be covalently linked to its corresponding target.Any of the aptamer methods disclosed herein can include the use of two or more aptamers that specifically bind the same target molecule.As will be further described below, aptamers can include tags.If an aptamer includes a tag, all copies of the aptamer do not need to have the same tag.Furthermore, if different aptamers each include a tag, these different aptamers can have either the same tag or different tags.
[0091] Aptamers can be identified using any known method, including the SELEX process. Once identified, aptamers can be prepared or synthesized according to any known method, including chemical and enzymatic synthesis.
[0092] As used herein, "SOMAmer" or Slow Off-Rate Modified Aptamer refers to an aptamer with improved off-rate characteristics. SOMAmers can be generated using the improved SELEX method described in U.S. Patent Application Publication No. 2009 / 0004667, entitled "Method for Generating Aptamers with Improved Off-Rates."
[0093] The terms "SELEX" and "SELEX process" are generally used interchangeably herein to refer to the combination of (1) the selection of aptamers that interact with a target molecule in a desired manner (e.g., bind to a protein with high affinity) and (2) the amplification of these selected nucleic acids. The SELEX process can be used to identify aptamers with high affinity for a specific target or biomarker.
[0094] SELEX generally involves preparing a candidate mixture of nucleic acids, binding the candidate mixture to a desired target molecule to form an affinity complex, separating the affinity complex from unbound candidate nucleic acids, separating and isolating the nucleic acid from the affinity complex, purifying the nucleic acid, and identifying a specific aptamer sequence. The process may include multiple rounds to further fine-tune the affinity of the selected aptamer. The process may include an amplification step at one or more points in the process. See, for example, U.S. Pat. No. 5,475,096, entitled "Nucleic Acid Ligands." The SELEX process can be used to generate aptamers that bind covalently to a target, as well as aptamers that bind non-covalently to a target. See, for example, U.S. Pat. No. 5,705,337, entitled "Systematic Evolution of Nucleic Acid Ligands by Exponential Enrichment: Chemi-SELEX."
[0095] The SELEX process can be used to identify high-affinity aptamers containing modified nucleotides that confer improved characteristics on the aptamer, such as improved in vivo stability or improved delivery characteristics. Examples of such modifications include chemical substitutions at the ribose and / or phosphate and / or base positions. SELEX process-identified aptamers containing modified nucleotides are described in U.S. Pat. No. 5,660,985, entitled "High Affinity Nucleic Acid Ligands Containing Modified Nucleotides," which describes oligonucleotides containing nucleotide derivatives chemically modified at the 5' and 2' positions of the pyrimidine. U.S. Pat. No. 5,580,737 (see above) describes highly specific aptamers containing one or more nucleotides modified with 2'-amino (2'-NH2), 2'-fluoro (2'-F), and / or 2'-Omethyl (2'-OMe). See also U.S. Patent Application Publication No. 20090098549 entitled "SELEX and PHOTOSELEX," which describes physical and chemical property amplified nucleic acid libraries and their use in SELEX and photoSELEX.
[0096] SELEX can be used to identify aptamers with desired off-rate characteristics. See U.S. Patent Application Publication No. 20090004667, entitled "Method for Generating Aptamers with Improved Off-Rates," which describes an improved SELEX method for generating aptamers capable of binding to target molecules. As mentioned above, these slow-off-rate aptamers are known as "SOMAmers." A method for producing aptamers or SOMAmers and photoaptamers or photoSOMAmers with slower off-rates from their respective target molecules is described. The method includes contacting a candidate mixture with the target molecule, allowing nucleic acid-target complexes to form, and performing an enrichment process for those with slow off-rates, such that fast off-rate nucleic acid-target complexes dissociate and do not reform, while slow off-rate complexes remain intact. Additionally, the methods include the use of modified nucleotides in the production of candidate nucleic acid mixtures to generate aptamers or SOMAmers with improved off-rate performance. Non-limiting exemplary modified nucleotides include, for example, modified pyrimidines shown in Figures 3 and 4.
[0097] A variation of this assay uses aptamers containing photoreactive functional groups that allow the aptamer to covalently bind or "photocrosslink" to its target molecule. See, e.g., U.S. Patent No. 6,544,776, entitled "Nucleic Acid Ligand Diagnostic Biochip." These photoreactive aptamers are also referred to as photoaptamers. See, e.g., U.S. Patent Nos. 5,763,177, 6,001,577, and 6,291,184, each entitled "Systematic Evolution of Nucleic Acid Ligands by Exponential Enrichment: Photoselection of Nucleic Acid Ligands and Solution SELEX." See also, e.g., U.S. Patent No. 6,458,539, entitled "Photoselection of Nucleic Acid Ligands." After the microarray is contacted with the sample and the photoaptamers have had a chance to bind to their target molecules, the photoaptamers are photoactivated and the solid support is washed to remove any non-specifically bound molecules. Generally, harsh washing conditions can be used because the target molecules bound to the photoaptamers are not removed due to the covalent bond generated by the photoactivated functional group(s) on the photoaptamers. In this manner, the assay allows for the detection of biomarker values corresponding to the biomarkers in the test sample.
[0098] In both of these assay formats, the aptamer or SOMAmer is immobilized on a solid support before contacting with the sample. However, under certain circumstances, immobilizing the aptamer or SOMAmer before contacting with the sample may not provide an optimal assay. For example, pre-immobilization of the aptamer or SOMAmer may result in inefficient mixing of the aptamer or SOMAmer with the target molecule on the surface of the solid support, possibly leading to a long reaction time; therefore, the incubation period may be extended to allow the aptamer or SOMAmer to efficiently bind to their target molecule. Furthermore, when photoaptamers or photoSOMAmers are used in the assay, and depending on the material used as the solid support, the solid support may tend to scatter or absorb the light used to achieve covalent bond formation between the photoaptamer or photoSOMAmer and their target molecule. Furthermore, depending on the method used, detection of target molecules bound to aptamers or photoSOMAmers may be prone to inaccuracies, since the surface of the solid support may also be exposed to and affected by any labeling agent used. Finally, immobilization of aptamers or SOMAmers on a solid support generally involves a preparation step (i.e., immobilization) of the aptamer or SOMAmer prior to exposing the aptamer or SOMAmer to a sample, which preparation step may affect the activity or functionality of the aptamer or SOMAmer.
[0099] SOMAmer assays have also been described that allow a SOMAmer to capture its target in solution, followed by a separation step designed to remove specific components of the SOMAmer-target mixture prior to detection (see U.S. Patent Application Publication No. 20090042206, entitled "Multiplexed Analyses of Test Samples"). The described SOMAmer assay methods allow for the detection and quantification of non-nucleic acid targets (e.g., protein targets) in a test sample by detecting and quantitating nucleic acids (i.e., SOMAmers). The described methods generate nucleic acid surrogates (i.e., SOMAmers) for the detection and quantitation of non-nucleic acid targets, thereby allowing various nucleic acid technologies (including amplification) to be applied to a wide range of desired targets (including protein targets).
[0100] SOMAmers can be constructed to facilitate separation of assay components from the SOMAmer-biomarker complex (or photo-SOMAmer-biomarker covalent complex), allowing for isolation of the SOMAmer for detection and / or quantification. In some embodiments, these constructs can include cleavable or releasable elements within the SOMAmer sequence. In other embodiments, additional functionality can be introduced into the SOMAmer, such as a label or detectable component, a spacer component, or a specific binding tag or immobilization element. For example, a SOMAmer can include a cleavable moiety, a label, a spacer component that separates the label, and a tag connected to the SOMAmer via the cleavable moiety. In one embodiment, the cleavable element is a photocleavable linker. The photocleavable linker can be attached to a biotin moiety and a spacer section and can include an NHS group for amine derivatization, which can be used to introduce a biotin group into the aptamer, thereby allowing for release of the aptamer later in the assay method.
[0101] Homogeneous assays performed with all assay components in solution do not require separation of sample and reagents prior to signal detection. These methods are rapid and easy to use. These methods generate signals based on molecular capture or binding reagents that react with specific targets. For prediction of tobacco use status, the molecular capture reagent would be an aptamer or antibody or the like, and the specific target would be a tobacco use biomarker such as those in Table 4.
[0102] In some embodiments, the method for signal generation utilizes the anisotropic signal change resulting from the interaction of a fluorophore-labeled capture reagent with its specific biomarker target. When the labeled capture reagent reacts with its target, the increase in molecular weight significantly slows the rotational motion of the fluorophore bound to the complex, changing the anisotropy value. By monitoring the anisotropy change, the binding event can be used to quantitatively measure the biomarker in solution. Other methods include fluorescence polarization assays, molecular beacon methods, time-resolved fluorescence quenching, chemiluminescence, fluorescence resonance energy transfer, and the like.
[0103] An exemplary solution-based aptamer assay that can be used to detect a biomarker value corresponding to a biomarker in a biological sample includes the following steps: (a) preparing a mixture by contacting the biological sample with an aptamer that includes a first tag and has specific affinity for the biomarker, such that if the biomarker is present in the sample, an aptamer affinity complex is formed; (b) exposing the mixture to a first solid support that includes a first capture element, allowing the first tag to associate with the first capture element; (c) removing any components of the mixture that do not associate with the first solid support; (d) removing a second solid support that does not associate with the first capture element; (e) releasing the aptamer affinity complex from the first solid support; (f) exposing the released aptamer affinity complex to a second solid support comprising a second capture element, allowing the second tag to associate with the second capture element; (g) removing any uncomplexed aptamer from the mixture by partitioning the uncomplexed aptamer from the aptamer affinity complex; (h) eluting the aptamer from the solid support; and (i) detecting the biomarker by detecting the aptamer component of the aptamer affinity complex.
[0104] Any means known in the art can be used to detect biomarker values by detecting the aptamer component of an aptamer affinity complex. Many different detection methods can be used to detect the aptamer component of an affinity complex, such as hybridization assays, mass spectroscopy, or QPCR. In some embodiments, nucleic acid sequencing methods can be used to detect the aptamer component of an aptamer affinity complex and thereby detect biomarker values. Briefly, a test sample can be subjected to any type of nucleic acid sequencing method to identify and quantify one or more aptamer sequence(s) present in the test sample. In some embodiments, the sequence comprises the entire aptamer molecule or any portion of the molecule that can be used to uniquely identify the molecule. In other embodiments, an identification sequence is a specific sequence added to the aptamer; such sequences are often referred to as "tags," "barcodes," or "zip codes." In some embodiments, the sequencing method includes an enzymatic step to amplify the aptamer sequence or to convert any type of nucleic acid (including RNA and DNA containing chemical modifications at any position) into other types of nucleic acid suitable for sequencing.
[0105] In some embodiments, the sequencing method comprises one or more cloning steps. In other embodiments, the sequencing method comprises a direct sequencing method.
[0106] In some embodiments, the sequencing method involves a directed approach using specific primers that target one or more aptamers in the test sample. In other embodiments, the sequencing method involves a shotgun approach that targets all aptamers in the test sample.
[0107] In some embodiments, the sequencing method includes an enzymatic step to amplify the targeted molecules for sequencing. In other embodiments, the sequencing method directly sequences single molecules. An exemplary nucleic acid sequencing-based method that can be used to detect biomarker values corresponding to biomarkers in biological samples includes: (a) converting a mixture of aptamers containing chemically modified nucleotides into unmodified nucleic acids through an enzymatic step; (b) shotgun sequencing the resulting unmodified nucleic acids using a massively parallel sequencing platform (e.g., 54 Sequencing System (454 Life Sciences / Roche), Illumina Sequencing System (Illumina), ABI SOLiD Sequencing System (Applied Biosystems), HeliScope Single Molecule Sequencer (Helicos Biosciences), or Pacific Biosciences Real-Time Single Molecule Sequencing System (Pacific BioSciences), or Polonator G Sequencing System (Dover Systems)); and (c) identifying and quantifying the SOMAmers present in the mixture by specific sequences and sequence counts.
[0108] Determining Biomarker Values Using Immunoassays Immunoassays are based on the reaction of antibodies with their corresponding targets or analytes and can detect the analyte in a sample depending on the specific assay format. To improve the specificity and sensitivity of immunoreactivity-based assay methods, monoclonal antibodies are often used due to their specific epitope recognition. Polyclonal antibodies have also been successfully used in various immunoassays due to their increased affinity to the target compared to monoclonal antibodies. Immunoassays are designed for use with a wide range of biological sample matrices. Immunoassay formats are designed to provide qualitative, semi-quantitative, and quantitative results.
[0109] Quantitative results are generated through the use of a standard curve generated with known concentrations of the specific analyte to be detected. The response or signal from an unknown sample is plotted onto the standard curve to establish the amount or value corresponding to the target in the unknown sample.
[0110] Numerous immunoassay formats have been designed. ELISA or EIA can be quantitative for the detection of analytes. The methods rely on the attachment of a label to either the analyte or the antibody, with the label component comprising an enzyme, either directly or indirectly. ELISA tests can be formatted for direct, indirect, competitive, or sandwich detection of analytes. Other methods rely on labels, such as radioisotopes (I125) or fluorescence. Additional techniques include, for example, agglutination, nephelometry, turbidimetry, Western blot, immunoprecipitation, immunocytochemistry, immunohistochemistry, flow cytometry, Luminex assays, and others (see ImmunoAssay: A Practical Guide, edited by Brian Law, published by Taylor & Francis, Ltd., 2005 edition).
[0111] Exemplary assay formats include enzyme-linked immunosorbent assays (ELISAs), radioimmunoassays, fluorescent immunoassays, chemiluminescent immunoassays, and fluorescence resonance energy transfer (FRET) or time-resolved FRET (TR-FRET) immunoassays. Exemplary procedures for detecting biomarkers include biomarker immunoprecipitation followed by quantitative methods that allow for size- and peptide-level discrimination, such as gel electrophoresis, capillary electrophoresis, planar electrochromatography, and the like.
[0112] Methods for detecting and / or quantifying a detectable label or signal-producing substance depend on the nature of the label. The product of the reaction catalyzed by an appropriate enzyme (when the detectable label is an enzyme; see above) may be, without limitation, fluorescent, luminescent, or radioactive, or may absorb visible or ultraviolet light. Examples of detectors suitable for detecting such detectable labels include, without limitation, X-ray film, radioactivity counters, scintillation counters, spectrophotometers, colorimeters, fluorometers, luminometers, and densitometers.
[0113] Any of the detection methods can be performed in any format that allows for any suitable preparation, processing, and analysis of the reactants. This can be done, for example, in multi-well assay plates (e.g., 96-well or 384-well) or using any suitable array or microarray. Stock solutions for various agents can be made manually or robotically, and all subsequent pipetting, dilution, mixing, distribution, washing, incubation, sample reading, data collection, and analysis can be performed robotically using commercially available analysis software, robotics, and detection equipment capable of detecting the detectable label.
[0114] Determining Biomarker Values Using Gene Expression Profiling Measurement of mRNA in a biological sample can be used as a surrogate for detecting the level of the corresponding protein in the biological sample. Thus, any of the biomarkers or biomarker panels described herein can also be detected by detecting the appropriate RNA.
[0115] mRNA expression levels are measured by reverse transcription quantitative polymerase chain reaction (RT-PCR followed by qPCR). RT-PCR is used to generate cDNA from mRNA. The cDNA can be used in a qPCR assay to produce fluorescence as the DNA amplification process progresses. By comparison to a standard curve, qPCR can generate absolute measurements (such as the number of mRNA copies per cell). Northern blots, microarrays, Invader assays, and RT-PCR combined with capillary electrophoresis have all been used to measure mRNA expression levels in samples. See Gene Expression Profiling: Methods and Protocols, Richard A. Shimkets, editor, Humana Press, 2004.
[0116] miRNA molecules are small, non-coding RNAs that can regulate gene expression. Any method suitable for measuring mRNA expression levels can also be used for the corresponding miRNA. Recently, many laboratories have been investigating the use of miRNAs as biomarkers for disease. Given that many diseases involve widespread transcriptional regulation, it is not surprising that miRNAs could find a role as biomarkers. While the relationship between miRNA concentrations and disease is often less clear than that between protein levels and disease, the value of miRNA biomarkers can be substantial. Naturally, as with any RNA differentially expressed during disease, challenges facing the development of in vitro diagnostic products include the requirement that miRNAs persist in diseased cells and be easily extracted for analysis, or that they be released into blood or other matrices, where they must remain long enough to be measured. Protein biomarkers have similar requirements, but many potential protein biomarkers are intentionally secreted in a paracrine manner at sites of pathology and function during disease. Many potential protein biomarkers are designed to function outside the cells in which the proteins are synthesized.
[0117] Biomarker detection using in vivo molecular imaging technologies Any of the described biomarkers (see, e.g., Table 4) can be used in molecular imaging studies. For example, imaging agents can be coupled to any of the described biomarkers and used to aid in the assessment of tobacco use status, to monitor response to therapeutic interventions, and to select populations for clinical trials, among other uses.
[0118] In vivo imaging techniques provide a non-invasive method for determining the status of a particular disease or condition in an individual's body. For example, entire body parts or even the entire body can be viewed as three-dimensional images, thereby providing useful information about the morphology and structure of the body. Such techniques can be combined with the detection of biomarkers described herein to provide information about an individual's tobacco use status.
[0119] The use of in vivo molecular imaging techniques has expanded due to various technological advances. These advances include the development of new contrast agents or labels (such as radioisotope and / or fluorescent labels) that can provide strong signals within the body; and the development of powerful new imaging technologies that can detect and analyze these signals from outside the body with sufficient sensitivity and precision to provide useful information. Contrast agents can be visualized in an appropriate imaging system, thereby providing an image of the body part(s) in which they are located. Contrast agents can be bound to or associated with capture agents (such as aptamers or antibodies), and / or peptides or proteins, or oligonucleotides (e.g., for detecting gene expression), or complexes containing any of these together with one or more macromolecules and / or other particulate forms.
[0120] Contrast agents may also feature radioactive atoms useful in imaging. Suitable radioactive atoms include technetium-99m or iodine-123 for scintigraphic studies. Other readily detectable moieties include spin labels for magnetic resonance imaging (MRI) (e.g., again, iodine-123, iodine-131, indium-111, fluorine-19, carbon-13, nitrogen-15, oxygen-17, gadolinium, manganese, or iron, etc.). Such labels are well known in the art and can be readily selected by those skilled in the art.
[0121] Standard imaging techniques include, but are not limited to, magnetic resonance imaging, computed tomography scans (coronary artery calcium score), positron emission tomography (PET), single-photon emission computed tomography (SPECT), computed tomography angiography, and the like. For diagnostic in vivo imaging, the type of detection device available is a major factor in the selection of a given contrast agent (such as a given radionuclide) and the specific biomarker (protein, mRNA, and the like) to be targeted using it. The radionuclide chosen typically has a type of decay that is detectable by a given type of device. Also, when selecting a radionuclide for in vivo diagnosis, its half-life should be long enough to allow detection at the time of maximum uptake by the target tissue, yet short enough so that harmful radiation to the host is minimized.
[0122] Exemplary imaging techniques include, but are not limited to, PET and SPECT, which are imaging techniques in which radionuclides are administered synthetically or locally to an individual. The subsequent uptake of the radiotracer is measured over time and used to obtain information about the targeted tissue and biomarkers. Because the specific isotopes used emit with high energy (gamma rays) and the equipment used to detect them is sensitive and sophisticated, the two-dimensional distribution of radioactivity can be inferred from outside the body.
[0123] Positron-emitting isotopes commonly used in PET include, for example, carbon-11, nitrogen-13, oxygen-15, and fluorine-18. Isotopes that decay by electron capture and / or gamma emission are used in SPECT, including, for example, iodine-123 and technetium-99m. An exemplary method for labeling amino acids with technetium-99m involves reducing pertechnetate ions in the presence of a chelate precursor to form an unstable technetium-99m precursor complex, which in turn reacts with the metal-binding group of a bifunctionally modified chemotactic peptide to form a technetium-99m chemotactic peptide conjugate.
[0124] Antibodies are frequently used for such in vivo imaging diagnostic methods. The preparation and use of antibodies for in vivo diagnosis is well known in the art. Labeled antibodies that specifically bind to any of the biomarkers in Table 4 can be injected into an individual to be detected according to the particular biomarker used for the purpose of diagnosing or evaluating the individual's disease state or condition. The label used is selected according to the imaging modality used, as previously described. Localization of the label allows for the determination of tissue damage or other indicators related to the probability that the individual is a former tobacco user. The amount of label in an organ or tissue also allows for the determination of the contribution of biomarkers attributable to "ever" tobacco users in that organ or tissue.
[0125] Similarly, aptamers can be used for such in vivo imaging diagnostic methods. For example, aptamers used to identify (and therefore specifically bind to) specific biomarkers listed in Table 4 can be appropriately labeled and injected into an individual being evaluated for the determination of detectable tobacco use status according to the specific biomarker for the purpose of diagnosing or assessing the levels of tissue damage, atherosclerotic plaques, components of the inflammatory response, and other factors associated with tobacco use status in the individual. The label used is selected according to the imaging modality used, as previously described. Localization of the label allows for the determination of the site of the process leading to increased risk. The amount of label within an organ or tissue also allows for the determination of the infiltration of pathological processes in that organ or tissue. Aptamer-directed imaging agents can have unique and advantageous characteristics related to tissue permeability, tissue distribution, kinetics, excretion, efficacy, and selectivity compared to other imaging agents.
[0126] Such techniques can also be optionally carried out with labeled oligonucleotides, for example, for detecting gene expression through imaging with antisense oligonucleotides. These methods are used, for example, for in situ hybridization with fluorescent molecules or radionuclides as labels. Other methods for detecting gene expression include, for example, detecting the activity of reporter genes.
[0127] Another common type of imaging technique is optical imaging, in which fluorescent signals within a subject are detected by an optical device outside the subject. These signals can result from actual fluorescence and / or bioluminescence. Improvements in the sensitivity of optical detection devices have increased the usefulness of optical imaging for in vivo diagnostic assays.
[0128] The use of in vivo molecular biomarker imaging is increasing, including for clinical trials, to more rapidly measure clinical efficacy in testing, for example, new treatments for diseases or conditions, and / or to avoid long-term placebo treatment for diseases where such treatment may be deemed ethically questionable (such as multiple sclerosis).
[0129] For a review of other techniques, see N. Blow, Nature Methods, 6, 465-469, 2009.
[0130] Determination of Biomarker Values Using Mass Spectrometric Methods Mass spectrometers of various configurations can be used to detect biomarker values. Multiple types of mass spectrometers are available or can be manufactured in various configurations. Generally, a mass spectrometer has the following major components: a sample inlet, an ion source, a mass analyzer, a detector, a vacuum system, and an instrument control system and a data system. The differences in the sample inlet, ion source, and mass analyzer generally define the type of instrument and its capabilities. For example, the inlet can be a capillary column liquid chromatography source or a direct probe or stage used in matrix-assisted laser desorption, etc. Common ion sources are, for example, electrospray (including nanospray and microspray) or matrix-assisted laser desorption. Common mass analyzers include quadrupole mass filters, ion trap mass analyzers, and time-of-flight mass analyzers. Additional mass spectrometry methods are well known in the art (see Burlingame et al. Anal. Chem. 70:647 R-716R (1998); Kinter and Sherman, New York (2000)).
[0131] Protein biomarkers and biomarker values may be detected and measured by any of the following: electrospray ionization mass spectrometry (ESI-MS), ESI-MS / MS, ESI-MS / (MS)n, matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF-MS), surface-enhanced laser desorption / ionization time-of-flight mass spectrometry (SELDI-TOF-MS), desorption / ionization on silicon (DIOS), secondary ion mass spectrometry (SIMS), quadrupole-time-of-flight (Q-TOF), tandem time-of-flight (TOF / TOF) technology called ultraflex III TOF / TOF, atmospheric pressure chemical ionization mass spectrometry (APCI-MS), APCI-MS / MS, APCI-(MS)n, atmospheric pressure photoionization mass spectrometry (APPI-MS), APPI-MS / MS, and APPI-(MS)n, quadrupole mass spectrometry, Fourier transform mass spectrometry (FTMS), quantitative mass spectrometry, and ion trap mass spectrometry.
[0132] Sample preparation strategies are used to label and enrich samples prior to mass spectrometric characterization of protein biomarkers and determination of biomarker values. Labeling methods include, but are not limited to, isobaric tagging (iTRAQ) for relative and absolute quantification and stable isotope labeling of amino acids in cell culture (SILAC). Capture reagents used to selectively enrich samples for candidate biomarker proteins prior to mass spectrometric analysis include, but are not limited to, aptamers, antibodies, nucleic acid probes, chimeras, small molecules, F(ab')2 fragments, single-chain antibody fragments, Fv fragments, single-chain Fv fragments, nucleic acids, lectins, ligand-binding receptors, affibodies, nanobodies, ankyrins, domain antibodies, surrogate antibody scaffolds (e.g., diabodies), imprinted polymers, avimers, peptidomimetics, peptoids, peptide nucleic acids, threose nucleic acids, hormone receptors, cytokine receptors, and synthetic receptors, as well as modifications and fragments thereof.
[0133] Determination of biomarker values using proximity ligation assays Proximity ligation assays can be used to determine biomarker values. Briefly, a test sample is contacted with a pair of affinity probes, which can be a pair of antibodies or a pair of aptamers, and each member of the pair is extended with an oligonucleotide. The targets for the pair of affinity probes can be two distinct determinants on one protein, or one determinant on each of two different proteins that can exist as a homo- or hetero-multimeric complex. When the probes bind to the target determinants, the free ends of the oligonucleotide extensions are brought into sufficient proximity to hybridize together. Hybridization of the oligonucleotide extensions is facilitated by a common connector oligonucleotide, which serves to bridge the oligonucleotide extensions together when they are positioned sufficiently close. Once the oligonucleotide extensions of the probes are hybridized, the ends of the extensions are joined together by enzymatic DNA ligation.
[0134] Each oligonucleotide extension contains a primer site for PCR amplification. Once the oligonucleotide extensions are ligated together, the oligonucleotides form a continuous DNA sequence that, through PCR amplification, reveals information about the identity and quantity of the target protein, and, if the target determinants are on two different proteins, information about protein-protein interactions. Proximity ligation can provide a highly sensitive and specific assay for real-time protein concentration and interaction information through the use of real-time PCR. Probes that do not bind the determinants of interest will not bring the corresponding oligonucleotide extensions into proximity, and ligation or PCR amplification cannot proceed, resulting in no signal being generated.
[0135] The foregoing assays allow for the detection of biomarker levels useful in methods for predicting the probability of being a former tobacco user, the methods comprising detecting, in a biological sample from an individual, biomarker levels each corresponding to a biomarker selected from the group consisting of the biomarkers provided in Table 4, wherein classification using the biomarker levels, as described in detail below, indicates whether the individual has a probability of being a former tobacco user. While certain of the described biomarkers are useful alone in determining the probability that an individual is a former tobacco user, methods are also described herein for grouping multiple subsets of biomarkers, each useful as a panel of two or more biomarkers. According to any of the methods described herein, biomarker levels can be detected and classified individually or collectively, for example, in a multiplex assay format.
[0136] Biomarker classification and disease score calculation In some embodiments, a biomarker "signature" contains a set of markers for a given diagnostic or predictive test, each marker having different levels in a population of interest. Different levels, in this context, can refer to different means of marker levels for individuals in two or more groups, different variances in two or more groups, or a combination of both. For the simplest form of a diagnostic test, these markers can be used to assign an unknown sample from an individual into one of two groups (such as with or without a history of tobacco use). Assigning a sample into one of two or more groups is known as classification, and the procedure used to achieve this assignment is known as a classifier or classification method. Classification methods can also be referred to as scoring methods. There are many classification methods that can be used to construct a diagnostic classifier from a set of biomarker values. Generally, classification methods are most easily accomplished using supervised learning techniques when a dataset is collected using samples obtained from individuals in two (or more for multiple classification situations) distinct groups that one wishes to distinguish. Since the class (group or population) to which each sample belongs is known in advance for each sample, classification methods can be trained to give the desired classification response. It is also possible to use unsupervised learning techniques to generate classifiers for diagnostic purposes.
[0137] Common approaches for developing diagnostic classifiers include decision trees; bagging, boosting, forests, and random forests; rule-based learning; Parzen windows; linear models; logistic regression; neural network methods; unsupervised clustering; k-means; hierarchical ascending / descending; semi-supervised learning; prototype methods; nearest neighbors; kernel density estimation; support vector machines; hidden Markov models; and Boltzmann learning, and classifiers can be combined simply or in a manner that minimizes a specific objective function. For reviews, see, e.g., *Pattern Classification*, *RODuda*, et al., editors, *John Wiley & Sons*, 2nd edition, 2001. See also *The Elements of Statistical Learning—Data Mining, Inference, and Prediction*, *T. Hastie*, et al., editors, *Springer Science+Business Media*, LLC, 2nd edition, 2009. Each of these documents is incorporated by reference in its entirety.
[0138] To generate a classifier using supervised learning techniques, a set of samples called training data is obtained. In the context of diagnostic testing, training data includes samples from distinct groups (classes) to which unknown samples will subsequently be assigned. For example, samples collected from individuals in a control population and individuals in a particular disease, condition, or event population (such as "ever" tobacco users (former or current tobacco users)) may constitute training data for developing a classifier that can classify unknown samples (or more specifically, the individuals from whom the samples were obtained) as either "ever" tobacco users or "never tobacco users." Developing a classifier from training data is known as training the classifier. The specific details of classifier training depend on the nature of the supervised learning technique (see, e.g., Pattern Classification, R.O.Duda, et al., editors, John Wiley & Sons, 2nd edition, 2001; see also, The Elements of Statistical Learning - Data Mining, Inference, and Prediction, T. Hastie, et al., editors, Springer Science+Business Media, LLC, 2nd edition, 2009).
[0139] Typically, there are more possible biomarker values than samples in the training set, so care must be taken to avoid overfitting. Overfitting occurs when the statistical model represents random error or noise instead of background associations. Overfitting can be avoided in various ways, including, for example, by limiting the number of markers used in developing the classifier, by assuming that marker responses are independent of each other, by limiting the complexity of the background statistical model used, and by ensuring that the background statistical model fits the data.
[0140] To identify a set of biomarkers associated with the occurrence of an event, the combined set of control and early event samples was analyzed using principal component analysis (PCA). PCA presents samples relative to an axis defined by the greatest variation among all samples, without considering the outcome of the cases or controls, thus reducing the risk of overfitting the distinction between cases and controls. Because the occurrence of a serious thrombotic event involves a strong chance component and requires reporting of unstable plaque in a life-threatening blood vessel that is about to rupture, we did not expect to observe a clear separation between the control and event sample sets. While the separation observed between cases and controls was not large, it occurred in the second principal component, accounting for approximately 10% of the total variation in this set of samples, indicating a relatively simple approach to quantifying background biological variation.
[0141] In the next set of analyses, biomarkers can be analyzed for components of differences between samples that were specific to the separation between control samples and early event samples. One method that can be used is to use DSGA (Bair, E. and Tibshirani, R. (2004) Semi-supervised methods to predict patient survival from gene expression data. PLOS Biol., 2, 511-522) to remove (deflate) the top three principal component directions of variation between samples in the control set. While dimensionality reduction is performed on the control set to be explored, both samples in the control and samples from the early event are run through PCA. Separation of cases from the early event can be observed along the horizontal axis.
[0142] Cross-validated selection of proteins associated with tobacco use status To avoid overfitting protein prediction power to the unique features of a particular sample selection, cross-validation and dimensionality reduction approaches can be employed. Cross-validation involves multiple selections of a set of samples to determine the association of risk with proteins, combined with the use of unselected samples to monitor the ability of the method to be applied to samples not used to generate the risk model (The Elements of Statistical Learning - Data Mining, Inference, and Prediction, T. Hastie, et al., editors, Springer Science+Business Media, LLC, 2nd edition, 2009). We applied the supervised PCA method of Tibshirani et al. (Bair, E. and Tibshirani, R. (2004) Semi-supervised methods to predict patient survival from gene expression data. PLOS Biol., 2, 511-522), which can be applied to high-dimensional datasets in modeling tobacco use status. The supervised PCA (SPCA) method involves the univariate selection of a set of proteins that are statistically associated with the event risk observed in the data, and the determination of correlated components that combine information from all of these proteins. This determination of correlated components is a dimensionality reduction step that not only combines information across proteins, but also mitigates the possibility of overfitting by reducing the number of independent variables from the full protein menu of over 1000 proteins to a small number of principal components (in this work, we considered only the first principal component).
[0143] Univariate and multivariate analyses of individual protein-to-event relationships The Cox proportional hazards model (Cox, David R (1972). "Regression Models and Life-Tables". Journal of the Royal Statistical Society. Series B (Methodological) 34(2):187-220)) is widely used in medical statistics. Cox regression avoids fitting a specific time function to cumulative survival rates and instead uses a model of relative risk, called the baseline hazard function, which can vary over time. The baseline hazard function describes the general shape of the survival time distribution for all individuals, while the relative risk gives the level of risk for a set of covariate values (such as a single individual or a group) as a multiple of the baseline risk. The relative risk is constant over time in the Cox model.
[0144] Accelerated mortality time (AFT) models are a subclass of survival models. Survival models predict time-to-event data with partial information. Because survival models account for censoring, they can still use data from these censored subjects, whereas other longitudinal models that attempt to predict when an event will occur may only use information from subjects with a determined outcome. And because survival models account for time-to-event, they can generate predicted probabilities of an event occurring within any time frame, which differs from most classification models (logistic regression, random forests).
[0145] The AFT survival model is a regression model that, among other things, defines / assumes a linear relationship between the model's covariates and log(time to event). Thus, a subject with a covariate (protein RFU count) twice as high as baseline may be predicted to "survive" twice as long from diagnosis as baseline.
[0146] The two most common survival models are the AFT model and the proportional hazards model, and the AFT Weibull model is both. The proportional hazards model is slightly more complex to define than the AFT model; in a proportional hazards model, a subject with a covariate that is twice as high as baseline will have a twice as high risk at any time point, where risk is the negative derivative of the survival curve over time.
[0147] Other common proportional hazards models are the exponential and Cox models. The exponential model is a subtype of the Weibull model. The Cox model is more limited in use; predicted probabilities of time to event are not available from the Cox model, only relative risks. The AFT model can provide both absolute and relative risks.
[0148] kit Any combination of the biomarkers in Table 4 can be detected using a kit suitable for use in performing the methods disclosed herein, etc. Additionally, any kit can contain one or more detectable labels (such as fluorescent moieties) described herein.
[0149] In one embodiment, the kit includes (a) one or more capture reagents (e.g., at least one aptamer or antibody) for detection of one or more biomarkers in a biological sample, where the biomarkers include any of the biomarkers set forth in Table 4, and optionally (b) one or more software or computer program products for classifying an individual from whom the biological sample was obtained as either having or not having a probability of being a "ever" tobacco user, as further described herein. Alternatively, rather than one or more computer program products, one or more instructions for manually performing the above steps by a human may be provided.
[0150] The combination of a solid support and a corresponding capture reagent with a signal-generating agent is referred to herein as a "detection device" or "kit." The kit may also include instructions for use of the device and reagents, sample handling, and data analysis. Additionally, the kit may be used with a computer system or software to analyze biological samples and report the results of the analysis.
[0151] The kits may also contain one or more reagents for processing a biological sample (e.g., solubilization buffer, detergent, cleaning agent, or buffer). Any of the kits described herein may also include, for example, buffers, blocking agents, mass spectrometry matrix materials, antibody capture agents, positive control samples, negative control samples, software, and information (such as protocols, guidance, and reference data).
[0152] In one aspect, the present invention provides a kit for assessing tobacco use status. The kit includes PCR primers for one or more aptamers specific to a biomarker selected from Table 4. The kit may further include instructions for use and instructions for correlating the biomarker with a prediction of the probability of being an "ever" tobacco user. The kit may also include a DNA array containing complements of one or more of the aptamers specific to a biomarker selected from Table 4, reagents, and / or enzymes for amplifying or isolating sample DNA. The kit may include reagents for real-time PCR (e.g., TaqMan probes and / or primers, and enzymes).
[0153] For example, a kit may include (a) reagents including at least a capture reagent for quantifying one or more biomarkers in a test sample, wherein the biomarkers include the biomarkers set forth in Table 4 or other biomarkers or biomarker panels described herein, and optionally (b) one or more algorithms or computer programs for performing the steps of comparing the amount of each quantified biomarker in the test sample to one or more predetermined cutoffs and assigning a score for each quantified biomarker based on the comparison, combining the assigned scores for each quantified biomarker to obtain a total score, comparing the total score to a predetermined score, and using the comparison to determine the probability that the individual was an "ever" tobacco user. Alternatively, one or more instructions for manually performing the above steps by a human, rather than one or more algorithms or computer programs, may be provided.
[0154] Computer methods and software Once a biomarker or panel of biomarkers has been selected, a method for diagnosing an individual may include: 1) collecting or otherwise obtaining a biological sample; 2) performing an analytical method to detect and measure the biomarker or panel of biomarkers in the biological sample; 3) performing any data normalization or standardization required by the method used to collect biomarker levels; 4) calculating marker scores; 5) combining the marker scores to obtain a total diagnostic or predictive score; and 6) reporting the individual's diagnostic or predictive score. In this approach, the diagnostic or predictive score may be a single number determined from the sum of all marker calculations, which is compared to a preset threshold that indicates the presence or absence of disease, or that indicates "ever" or "never" tobacco use. Alternatively, the diagnostic or predictive score may be a series of bars, each representing a biomarker level, and the pattern of response may be compared to a preset pattern to determine the presence or absence of increased (or non-increased) risk of a disease, condition, or event.
[0155] At least some embodiments of the methods described herein may be implemented using a computer. An example of a computer system 100 is shown in FIG. 5. Referring to FIG. 5, system 100 is shown to be comprised of hardware elements electrically coupled via a bus 108, including a processor 101, input devices 102, output devices 103, storage devices 104, computer-readable storage media reader 105a, a communication system 106, processing acceleration (e.g., a DSP or dedicated processor) 107, and memory 109. Computer-readable storage media reader 105a is further coupled to computer-readable storage media 105b, which combination collectively represents remote, local, fixed, and / or removable storage devices, plus storage media, memory, etc., for temporarily and / or more permanently containing computer-readable information, which may include storage device 104, memory 109, and / or any other such accessible system 100 resource. System 100 also includes software elements (generally shown as residing in working memory 191), including an operating system 192 and other code 193 (such as programs, data, and the like).
[0156] With reference to FIG. 5 , system 100 has a wide range of flexibility and configurability. Thus, for example, a single architecture may be utilized to implement one or more servers, which may be further configured according to commonly desired protocols, protocol variations, extensions, and the like. However, it will be apparent to those skilled in the art that embodiments may be better utilized according to more specific application requirements. For example, one or more system elements may be implemented as sub-elements within a system 100 component (e.g., within communications system 106). Customized hardware may also be utilized, and / or particular elements may be implemented in hardware, software, or both. Additionally, connections to other computing devices, such as network input / output devices (not shown), may be utilized, although it should be understood that wired, wireless, modem, and / or other connection(s) to other computing devices may also be utilized.
[0157] In one embodiment, the system may include a database containing features of biomarker characteristics predictive of the probability that an individual is an "ever" tobacco user. The biomarker data (or biomarker information) may be utilized as input to a computer for use as part of a computer-implemented method. The biomarker data may include the data described herein.
[0158] In one embodiment, the system further includes one or more devices for providing input data to the one or more processors.
[0159] The system further includes a memory for storing the dataset of ranked data elements.
[0160] In another embodiment, the device for providing input data includes a detector (such as a mass spectrometer or gene chip reader) for detecting characteristics of the data elements.
[0161] The system may additionally include a database management system. User requests or queries may be formatted in an appropriate language understood by the database management system, which processes the queries to extract relevant information from the training set database.
[0162] The system may be connectable to a network to which a network server and one or more clients are connected. The network may be a local area network (LAN) or a wide area network (WAN), as known in the art. Preferably, the server includes the necessary hardware to execute computer program products (e.g., software) and access database data for processing user requests.
[0163] The system may include an operating system (e.g., UNIX or Linux) for execution of instructions from the database management system. In one embodiment, the operating system may operate over a global communications network (such as the Internet) and utilize a global communications network server for connecting to such a network.
[0164] The system may include one or more devices that include a graphical display interface that includes interface elements (such as buttons, pull-down menus, scroll bars, text entry fields, and the like) conventionally found in graphical user interfaces known in the art. Requests registered on the user interface may be sent to application programs in the system for formatting and searching for relevant information in one or more of the system databases. Requests or queries registered by users may be constructed in any suitable database language.
[0165] A graphical user interface may be generated by graphical user interface code as part of the operating system and may be used to input data and / or display input data. The results of processed data may be displayed in the interface, printed on a printer in communication with the system, stored in a memory device, and / or transmitted over a network, or provided in the form of a computer-readable medium.
[0166] The system may be in communication with an input device to provide data (e.g., expression values) regarding the data elements to the system. In one embodiment, the input device may include a gene expression profiling system, including, for example, a mass spectrometer, gene chip, or array reader, and the like.
[0167] According to various embodiments, the methods and apparatus for analyzing tobacco use status biomarker information can be implemented in any suitable manner, for example, using a computer program running on a computer system. Conventional computer systems including a processor and random access memory (such as a remotely accessible application server, network server, personal computer, or workstation) can be used. Additional computer system components can include memory devices or information storage systems (such as a mass storage system) and user interfaces (e.g., conventional monitors, keyboards, and tracking devices). The computer system can be a stand-alone system or part of a network of computers including a server and one or more databases.
[0168] A tobacco use assessment biomarker analysis system may provide functions and operations for completing data analysis (such as data collection, processing, analysis, reporting, and / or diagnosis). For example, in one embodiment, a computer system may execute a computer program that may receive, store, retrieve, analyze, and report information related to predictive biomarkers for tobacco use assessment. The computer program may include multiple modules that perform various functions or operations (such as a processing module for processing raw data and generating supplemental data and an analysis module for analyzing the raw data and supplemental data to generate a predicted tobacco use status). Determining the probability that an individual is a "ever" tobacco user may optionally include generating or collecting other information (including additional biomedical information) regarding the individual's status related to a disease, condition, or event, identifying whether further testing may be desired, or otherwise assessing the individual's health status.
[0169] Referring now to FIG. 6, an example of a computer-implemented method according to the principles of the disclosed embodiments can be seen. In FIG. 6, a flowchart 3000 is shown. In block 3004, biomarker information can be retrieved for an individual. The biomarker information can be retrieved from a computer database, for example, after testing of the individual's biological sample has been performed. The biomarker information can include biomarker levels, each corresponding to one or more of the biomarkers in Table 4. In block 3008, a computer can be utilized to classify each of the biomarker levels. Then, in block 3012, a determination can be made on a plurality of classifications as to the probability that the individual is an "ever" tobacco user. The display can be output to a display or other display device for human viewing. Thus, for example, it can be displayed on a display screen of a computer or other output device.
[0170] Some embodiments described herein may be implemented to include a computer program product, which may include a computer-readable medium having computer-readable program code embodied in the medium for causing an application program to run on a computer with a database.
[0171] As used herein, a "computer program product" refers to a set of instructions, organized in the form of natural language statements or programming language statements, contained on a physical medium of any nature (e.g., written, electronic, magnetic, optical, or other) and usable by a computer or other automated data processing system. Such programming language statements, when executed by a computer or data processing system, cause the computer or data processing system to operate according to the specific content of the instructions. Computer program products include, but are not limited to, programs in source code and object code and / or test or data libraries embodied in a computer-readable medium. Furthermore, computer program products that enable a computer system or data processing equipment device to operate in a preselected manner may be provided in many forms, including, but not limited to, original source code, assembly code, object code, machine language, encrypted or compressed versions of the above, and any and all equivalents.
[0172] In one embodiment, a computer program product is provided for assessing tobacco use status. The computer program product includes a computer-readable medium having program code embodied thereon executable by a processor of a computing device or system, the program code including: code for retrieving data belonging to a biological sample from an individual, the data including biomarker values each corresponding to one or more of the biomarkers in Table 4; and code for implementing a classification method that indicates the probability of the individual's "ever" tobacco user status as a function of the biomarker values.
[0173] While various embodiments have been described as methods or apparatus, it should be understood that embodiments can be implemented via code coupled to a computer (e.g., code resident on or accessible by a computer). For example, software and databases can be utilized to implement many of the methods discussed above. Accordingly, in addition to embodiments achieved through hardware, it should also be noted that these embodiments can be achieved through the use of an article of manufacture consisting of a computer-usable medium having computer-readable program code embodied therein that causes the performance of the functions disclosed herein. Accordingly, it is desired that embodiments be deemed protected by this patent in their program code means as well. Furthermore, embodiments can be embodied as code stored in virtually any type of computer-readable memory, including, but not limited to, RAM, ROM, magnetic media, optical media, or magneto-optical media. Even more generally, embodiments can be implemented in software or hardware, or any combination thereof, including, but not limited to, software running on a general-purpose processor, microcode, PLA, or ASIC.
[0174] It is also contemplated that embodiments may be achieved as a computer signal embodied in a carrier wave and as a signal propagated over a transmission medium (e.g., electrical and optical). Thus, various types of information discussed above may be formatted in structures (such as data structures) and transmitted as an electrical signal over a transmission medium or stored on a computer-readable medium.
[0175] It should also be noted that many of the structures, materials, and acts recited herein may be recited as a means for performing a function or a step for performing a function, and such language should therefore be understood to be entitled to cover all such structures, materials, or acts disclosed herein, including those incorporated by reference, and their equivalents.
[0176] The biomarker identification process, uses of the biomarkers, and various methods for determining biomarker values disclosed herein are described in detail above with respect to assessing the probability that an individual is an "ever" tobacco user. However, the application of the process, uses of the identified biomarkers, and methods for determining biomarker values are fully applicable to identifying other specific types of diseases or medical conditions, or individuals who would or would not benefit from adjunctive medical treatment.
[0177] Other methods In some embodiments, the biomarkers and methods described herein are used to determine health insurance premium or coverage determinations and / or life insurance premium or coverage determinations. In some embodiments, the results of the methods described herein are used to determine health insurance premium and / or life insurance premium determinations. In some such instances, an organization providing health insurance or life insurance requests or otherwise obtains information regarding a subject's tobacco use status and uses that information to determine appropriate health insurance or life insurance premiums for the subject. In some embodiments, the test is requested by and paid for by an organization providing health insurance or life insurance. In some embodiments, the test is used by a potential acquirer of the practice or health system or company to predict future liabilities or costs if the acquisition proceeds.
[0178] In some embodiments, the biomarkers and methods described herein are used to predict and / or manage medical resource utilization. In some such embodiments, the methods are not performed for the purpose of such prediction, but information obtained from the methods is used in such prediction and / or management of medical resource utilization. For example, a laboratory or hospital may collect information about a large number of subjects from the methods to predict and / or manage medical resource utilization in a particular facility or in a particular geographic area. [Example]
[0179] The following examples are provided for illustrative purposes only and are not intended to limit the scope of the present application, which is defined by the appended claims. All examples described herein were carried out using standard techniques well known and routine to those skilled in the art. The routine molecular biology techniques described in the following examples can be carried out as described in standard laboratory manuals (e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, (2001)).
[0180] Example 1: Multiplexed aptamer assay and statistical approach for biomarker identification A multiplex aptamer assay was used to analyze test and control samples to identify predictive biomarkers for the probability that an individual is a "ever" tobacco user. The multiplex assay used in this experiment included aptamers for detecting approximately 5,000 proteins in blood from a small sample volume (approximately 65 μl of serum or plasma) with a low detection limit (median 1 pM), a dynamic range of approximately 7 logs, and a median coefficient of variation of approximately 5%. Multiplex aptamer assays are generally described, for example, in Gold et al. (2010) Aptamer-Based Multiplexed Proteomic Technology for Biomarker Discovery. PLoS ONE 5(12):e15004; and U.S. Patent Application Publication Nos. 2012 / 0101002 and 2012 / 0077695.
[0181] Example 2: Model specifications The primary outcome was tobacco use status, an assessment of self-reported categories of tobacco use habits, with three levels: "never tobacco user," "past" tobacco user, or "current" tobacco user. "Past" (e.g., "former") tobacco users were combined with "current" tobacco users, labeled "ever," and compared to the "never tobacco user" group for model development.
[0182] The model selected was a 39-feature, protein-only (see Table 4) elastic net logistic regression model trained on a binary classification endpoint of "ever" ("current" and "former") tobacco users or "never" tobacco users.
[0183] The model provides a single prediction output, predicted probability. This output is the absolute probability (i.e., absolute risk) of being an "ever" tobacco user at the time of blood draw, and ranges from 0 to 1. Based on the absolute probability and rounded to thirds, the predictions are categorized into two classes: "never" tobacco users and "ever" tobacco users. Table 3 below shows the absolute probability ranges and their associated category labels. The receiver operating characteristic (ROC) curve for the final model is shown in Figure 1, and boxplots of the predicted probabilities for the "never" and "ever" tobacco user classifications are shown in Figure 2.
[0184] [Table 3] [Table 4-1] [Table 4-2]
[0185] Table 5 shows the performance metrics for the model. [Table 5]
[0186] The model output is a binary classification of "ever" tobacco user or "never" tobacco user at the time of blood draw. Predicted probabilities less than 0.460 were labeled as "never" tobacco users, and otherwise as "ever" tobacco users. For RUO purposes, the probability of being an ever tobacco user will be reported as a categorical class ("ever" vs. "never") along with the predicted probability score. Because the output of this model is based on probability, values outside the range [0, 1] are insufficient and will not be reported. [Table 6]
[0187] Example 3: Dataset for test development and validation Development and Validation Cohort(s). The Covance study, also known as the collection of blood samples and baseline information from healthy individuals for biomarker testing, is an observational study conducted across six regions of the United States: Austin, TX; Boise, ID; Dallas, TX; Honolulu, HI; Portland, OR; and San Diego, CA. Between 2008 and 2009, the Covance study enrolled 1,029 individuals aged 19-90 years who matched U.S. demographics for self-reported medical history and ethnic / racial and gender demographics, and blood samples were collected at a single time point. For test development, plasma samples were available for 1,020 individuals aged 19-89 years.
[0188] The exclusion criteria for this study are as follows: 1. Uncontrolled hypertension (i.e., two readings 10 minutes apart >160 / 95); 2. Self-reported treatment for malignancies other than squamous cell carcinoma or basal cell carcinoma of the skin in the past 2 years; 3. Self-reported pregnancy; 4. Self-reported chronic infectious condition(s) (e.g., hepatitis B, hepatitis C, HIV), autoimmune condition(s), or other inflammatory condition(s) (e.g., SLE, scleroderma, MS, Crohn's disease, or ulcerative colitis); 5. Self-reported chronic kidney or liver disease, 6. Self-reported chronic heart failure or diagnosed myocardial infarction in the past 3 months; 7. Self-reported uncontrolled diabetes (HbA1c>8 if known); 8. Self-reported acute viral or bacterial infection or temperature >38°C within 24 hours of enrollment; 9. Self-reported participation in any therapeutic trial within 14 days prior to blood sampling; Self-reported intake of prednisone or related medications >10.20 mg / day was included.
[0189] The selected model was developed based on self-reported tobacco use status at the time of blood draw. The Covance dataset is representative of tobacco use in the United States, with an overall prevalence of "current" tobacco users of 14.6% (13-15% between the model development and validation datasets). A 2020 report from the Surgeon General found that the prevalence of "current" smokers in the United States fell to 13.8% in 2018, the lowest population to date (assessments began in 1965). The overall prevalence of “former” smokers in the Covance dataset was 33.3% (33–35% between the model development and validation datasets), higher than the 2018 prevalence of 20.9%, which may be due to bias in sampling location or demographics (e.g., race / ethnicity) (United States Public Health Service Office of the Surgeon General; National Center for Chronic Disease Prevention and Health Promotion (US) Office on Smoking and Health. Smoking Cessation: A Report of the Surgeon General. Washington (DC): US Department of Health and Human Services; 2020).
[0190] Dataset stratification. For this study, the data were independently split into three sets (70% training / 15% validation / 15%) and stratified into ever tobacco users vs. never tobacco users to allow for the identification of robust models while mitigating overfitting issues. The validation dataset was not used in the proof-of-concept or refinement stages.
[0191] Model development data [Table 7-1] [Table 7-2]
[0192] Model Validation Data [Table 8-1] [Table 8-2]
[0193] Model Validation Data [Table 9-1] [Table 9-2]
[0194] Example 4: Results from development Data QC and pre-analysis results. The original dataset contained 1,020 samples, but some samples were removed based on the data QC analysis results. The removals are detailed in Table 10. [Table 10]
[0195] An additional 363 analytes failed target confirmation specificity testing and were removed prior to analysis, leaving 7,233 analytes available for analysis. No other issues were identified during data QC or pre-analysis.
[0196] POC Approach and Results: Model performance requirements were met by an AUC ≥ 0.717.
[0197] Refinement Approach and Results. The final model for the tobacco use status test is a 39-analyte, protein-only elastic net logistic regression model. The model was trained on the 70% Covance dataset. Validation metrics were calculated on a separate 15% dataset, with an additional 15% dataset held out for use in validation. The primary model output is the predicted probability of ever being a tobacco user at the time of blood draw.
[0198] Feature selection was based on significant features for univariate analysis during POC. In particular, the top univariate features from the comparisons of "ever" vs. "never," "previous" vs. "current," "never" vs. "current," and "never" vs. "previous" were considered for inclusion in the "ever" vs. "never" model.
[0199] Candidate models were fitted using cross-validation with 10 folds and 5 replicates. Upsampling was used to reduce class imbalance. Cross-validation AUC for the "ever" vs. "never" endpoint was the primary metric used to evaluate model performance. Models with cross-validation AUC less than 90% of the AUC calculated on the full training set were overfitted and therefore eliminated from consideration.
[0200] We compared logistic regression models built with 27 different feature sets. Of these, 14 models met the AUC threshold and showed no signs of overfitting. The top three models were selected based on AUC on the validation dataset.
[0201] The selected models were chosen based on their robustness as assessed through various model robustness measures. In particular, the final model had the smallest tolerance interval and passed QC replicate testing.
[0202] The final model was an elastic net linear regression model with 39 non-zero features and 4 zero-coefficient features, α=0.100, and λ=0.0254. In development, the decision cutoff (for labeling "ever" vs. "never") was set as 0.46, the average of the optimal decision cutoffs to two decimal points, where "optimal" is the Youden's J-squared for the "ever" vs. "never" test computed on both the training and validation datasets. 21 The final model will report predicted probabilities to three decimal places using a decision cutoff of 0.460.
[0203] The performance metrics for the training and validation data are shown in Table 11. The AUC values for training and validation meet the requirement of an AUC of 0.717 or greater. [Table 11]
[0204] Performance of stratifying "current" and "former" tobacco users within the "ever" tobacco user class. Predicted probabilities were stratified across the three tobacco use classes: "current," "former," and "never."
[0205] Example 5: Model validation plan and results Clinical Validation Plan. Validation will be assessed on a holdout dataset of 15% of the unused Covance data. The final model from refinement will be a 39-feature elastic net logistic regression model. Predicted probabilities for being an "ever" tobacco user will be calculated using this final model, and these predictions will be compared to the true norm of "ever" vs. "never." Models will be required to have an AUC of at least 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.717, or better on the validation dataset.
[0206] Clinical Results for Validation Data. The required passing criterion for validation was that the tobacco use status model have an AUC of at least 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.717, or greater on the validation data set. The model passed validation with an AUC equal to 0.729, as shown in Table 12. [Table 12]
[0207] The model performance metrics are slightly lower on the validation data than on the training or validation data sets, but the AUC is still high enough to pass the performance criteria.
[0208] The results from the simulations in Table 13 below show that the best approach to dealing with outlying aptamers is analyte deletion.
[0209] [Table 13]
[0210] The results from the margin of error simulation account for the variance in the model resulting from assay noise. The 95% and 99% upper and lower bounds in linear space and the maximum bound in probability space are listed in Table 14. [Table 14]
[0211] In some embodiments, one or more biomarkers may be included in the analysis and / or determination of univariate performance. By non-limiting example, biomarkers may include dCTP pyrophosphatase 1 (target name XTP3A, Entrez gene symbol DCTPP1, Uniprot ID Q9H773); matrix remodeling-associated protein 8 (target name MXRA8, Entrez gene symbol MXRA8, Uniprot ID Q9BRK3); T-complex protein 11-like protein 1 (target name T11L1, Entrez gene symbol TCP11L1, Uniprot ID Q9NUJ3), and / or cartilage acidic protein 1 (target name CRAC1, Entrez gene symbol CRTAC1, Uniprot ID Q9NQ79). Univariate performance is shown in Table 25.
[0212] Example 6: Analysis of a tobacco use model biomarker panel Model biomarker panels comprising various combinations of the biomarkers listed in Table 4 were analyzed to determine area under the curve (AUC) values for the various combinations. The model biomarker panel may be based on a panel of N biomarker proteins having an AUC value of at least 0.65, at least 0.66, at least 0.67, at least 0.68, at least 0.69, at least 0.7, at least 0.75, at least 0.8, at least 0.85, at least 0.9, or at least 0.95, where N is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, and / or 39 of the biomarker proteins listed in Table 4. The table below shows the results of an exemplary model when various combinations including biomarker proteins 1-39 are measured. [Table 15-1] [Table 15-2]
Table 15-3
Table 15-4
Table 15-5
Table 15-6
Table 15-7
Table 15-8
Table 15-9
Table 15-10
Table 15-11
Table 16-1
Table 16-2
Table 16-3
Table 16-4
Table 16-5
Table 16-6
Table 16-7
Table 16-8
Table 16-9
Table 16-10
Table 16-11
Table 16-12
Table 16-13
Table 16-14
Table 17-1
Table 17-2
Table 17-3
Table 17-4
Table 17-5
Table 17-6
Table 17-7
Table 17-8
Table 17-9
Table 17-10
Table 17-11
Table 17-12
Table 17-13
Table 18-1
Table 18-2
Table 18-3
Table 18-4
Table 18-5
Table 18-6
Table 18-7
Table 18-8
Table 18-9
Table 18-10
Table 18-11
Table 19-1
Table 19-2
Table 19-3
Table 19-4
Table 19-5
Table 19-6
Table 19-7
Table 19-8
Table 19-9
Table 19-10
Table 19-11
Table 20-1
Table 20-2
Table 20-3
Table 20-4
Table 20-5
Table 20-6
Table 20-7
Table 20-8
Table 20-9
Table 21-1
Table 21-2
Table 21-3
Table 21-4
Table 21-5
Table 21-6
Table 21-7
Table 21-8
Table 21-9
Table 21-10
Table 21-11
Table 22-1
Table 22-2
Table 22-3
Table 22-4
Table 22-5
Table 22-6
Table 22-7
Table 22-8
Table 22-9
Table 22-10
Table 22-11
Table 23-1
Table 23-2
Table 23-3
Table 23-4
Table 23-5
Table 23-6
Table 23-7
Table 23-8
Table 23-9
Table 23-10
Table 23-11
Table 24-1
Table 24-2
Table 25-1
Table 25-2
Claims
1. 1. A method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of EPHA6 biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin , secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
2. 1. A method for detecting levels of N biomarker proteins in a sample, comprising obtaining the sample from a subject and detecting the level of each of the N biomarker proteins in the sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, At least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is EPHA6, placental alkaline phosphatase, SCF, MASP3: heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agu phospholipid, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
3. 1. A method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of placental alkaline phosphatase biomarker protein and levels of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is EPHA6, SCF, MASP3: heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin , secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
4. 1. A method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of an SCF biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is EPHA6, placental alkaline phosphatase, MASP3: heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin , secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
5. 1. A method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of a MASP3:heavy chain biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32. , at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is EPHA6, placental alkaline phosphatase, SCF, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, or agrin. , secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
6. 1. A method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of DUS10 biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is EPHA6, placental alkaline phosphatase, SCF, MASP3: heavy chain, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin , secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
7. 1. A method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of an IgG4-κ biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin , secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
8. 1. A method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of a DCC biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is EPHA6, placental alkaline phosphatase, SCF, MASP3: heavy chain, DUS10, IgG4-κ, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin , secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
9. 1. A method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of RCAN3 biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is EPHA6, placental alkaline phosphatase, SCF, MASP3: heavy chain, DUS10, IgG4-κ, DCC, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin , secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
10. 1. A method of predicting the probability that a subject is a former tobacco user, comprising detecting a level of CD248 biomarker protein and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, and at least one of the N biomarker proteins is EPHA6, placental alkaline phosphatase, SCF, MASP3: heavy chain, DUS10, IgG4-κ, DCC, RCAN3, EDIL3, renin, PSP-94, PIGR, URB, agrin , secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
11. N is 2 to 39, or N is 3 to 39, or N is 4 to 39, or N is 5 to 39, or N is 6 to 39, or N is 7 to 39, or N is 8 to 39, or N is 9 to 39, or N is 10 to 39, or N is 11 to 39, or N is 12 to 39, or N is 13 to 39, or N is 14 to 39, or N is 15 to 39, or N is 16 to 39, or N is 17 to 39, or N is 18 to 39, or N is 19 to 39, or N is 20-39, or N is 21-39, or N is 22-39, or N is 23-39, or N is 24-39, or N is 25-39, or N is 26-39, or N is 27-39, or N is 28-39, or N is 29-39, or N is 30-39, or N is 31-39, or N is 32-39, or N is 33-39, or N is 34-39, or N is 35-39, or N is 36-39, or N is 37-39, or N is 38-39.
12. N is 2, or N is 3, or N is 4, or N is 5, or N is 6, or N is 7, or N is 8, or N is 9, or N is 10, or N is 11, or N is 12, or N is 13, or N is 14, or N is 15, or N is 16, or N is 17, or N is 18, or N is 19, or N is 20 10. A method according to any one of the preceding claims, wherein N is 21, or N is 22, or N is 23, or N is 24, or N is 25, or N is 26, or N is 27, or N is 28, or N is 29, or N is 30, or N is 31, or N is 32, or N is 33, or N is 34, or N is 35, or N is 36, or N is 37, or N is 38, or N is 39.
13. 10. The method of any one of the preceding claims, wherein each of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
14. 10. The method of any one of the preceding claims, wherein at least one of the N biomarker proteins is selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, and CD248.
15. 10. The method of any one of the preceding claims, wherein at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, or at least 9 of the N protein biomarkers are selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, and CD248.
16. 10. The method of any one of the preceding claims, wherein two of said N biomarker proteins are EPHA6 and placental alkaline phosphatase, or two of said N biomarker proteins are EPHA6 and SCF, or two of said N biomarker proteins are EPHA6 and MASP3: heavy chain, or two of said N biomarker proteins are EPHA6 and DUS10, or two of said N biomarker proteins are EPHA6 and IgG4-κ, or two of said N biomarker proteins are EPHA6 and DCC, or two of said N biomarker proteins are EPHA6 and RCAN3, or two of said N biomarker proteins are EPHA6 and CD248.
17. 10. The method of any one of the preceding claims, wherein two of said N biomarker proteins are placental alkaline phosphatase and SCF, or two of said N biomarker proteins are placental alkaline phosphatase and MASP3:heavy chain, or two of said N biomarker proteins are placental alkaline phosphatase and DUS10, or two of said N biomarker proteins are placental alkaline phosphatase and IgG4-κ, or two of said N biomarker proteins are placental alkaline phosphatase and DCC, or two of said N biomarker proteins are placental alkaline phosphatase and RCAN3, or two of said N biomarker proteins are placental alkaline phosphatase and CD248.
18. 10. The method of any one of the preceding claims, wherein two of the N biomarker proteins are SCF and MASP3:heavy chain, or two of the N biomarker proteins are SCF and DUS10, or two of the N biomarker proteins are SCF and IgG4-κ, or two of the N biomarker proteins are SCF and DCC, or two of the N biomarker proteins are SCF and RCAN3, or two of the N biomarker proteins are SCF and CD248.
19. 10. The method of any one of the preceding claims, wherein two of the N biomarker proteins are MASP3:heavy chain and DUS10, or two of the N biomarker proteins are MASP3:heavy chain and IgG4-κ, or two of the N biomarker proteins are MASP3:heavy chain and RCAN3, or two of the N biomarker proteins are MASP3:heavy chain and CD248, or two of the N biomarker proteins are MASP3:heavy chain and DCC.
20. 10. The method of any one of the preceding claims, wherein two of the N biomarker proteins are DUS10 and IgG4-κ, or two of the N biomarker proteins are DUS10 and DCC, or two of the N biomarker proteins are DUS10 and CD248, or two of the N biomarker proteins are DUS10 and RCAN3.
21. 10. The method of any one of the preceding claims, wherein two of said N biomarker proteins are IgG4-κ and DCC, or two of said N biomarker proteins are IgG4-κ and RCAN3, or two of said N biomarker proteins are IgG4-κ and CD248.
22. 10. The method of any one of the preceding claims, wherein two of the N biomarker proteins are DCC and RCAN3, or two of the N biomarker proteins are DCC and CD248.
23. 10. The method of any one of the preceding claims, wherein two of the N biomarker proteins are RCAN3 and CD248.
24. 10. A method for predicting the probability that a subject is a former tobacco user, the method comprising detecting the level of T11L1 biomarker protein.
25. 10. A method for predicting the probability that a subject is a former tobacco user, the method comprising detecting the level of MXRA8 biomarker protein.
26. 10. A method for predicting the probability that a subject is a former tobacco user, the method comprising detecting the level of XTP3A biomarker protein.
27. 10. A method for predicting the probability that a subject is a former tobacco user, the method comprising detecting the level of a CRAC1 biomarker protein.
28. 28. The method of any one of claims 24 to 27, further comprising detecting at least one biomarker protein selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit.
29. 10. The method of any one of the preceding claims, wherein the sample is a blood sample, a plasma sample, or a serum sample.
30. 10. The method of any one of the preceding claims, wherein said detection is performed using mass spectrometry, an aptamer-based assay, and / or an antibody-based assay.
31. 10. The method of any one of the preceding claims, wherein the method comprises contacting biomarker proteins of the sample(s) with a set of biomarker capture reagents, each biomarker capture reagent of the set specifically binding to a different biomarker protein being detected.
32. 32. The method of claim 31 , wherein each biomarker capture reagent is an antibody or an aptamer.
33. 33. The method of claim 32, wherein each biomarker capture reagent is an aptamer.
34. 34. The method of claim 33, wherein at least one aptamer is a slow dissociation aptamer.
35. 35. The method of claim 34, wherein at least one slow-dissociating aptamer comprises nucleotides with at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten modifications.
36. Each slow dissociation aptamer exhibits a dissociation rate (t 1/2 36. The method of claim 34 or claim 35, wherein the antibody binds to its target protein by
37. 37. The method of any one of claims 30 to 36, wherein the level of each measured biomarker protein is determined from relative fluorescence units (RFU) or protein concentration.
38. 10. The method of any one of the preceding claims, wherein the prediction of the probability that the subject is a former tobacco user is based on inputting the measured levels of the N biomarker proteins into a statistical model.
39. 39. The method of claim 38, wherein said determining comprises analyzing the levels of said N biomarker proteins using an elastic net logistic regression model.
40. 40. The method of claim 38 or 39, wherein the model has an area under the curve (AUC) selected from at least 0.65, at least 0.66, at least 0.67, at least 0.68, at least 0.69, at least 0.7, at least 0.75, at least 0.8, at least 0.85, at least 0.9, or at least 0.
95.
41. 36. The method of any one of claims 33 to 35, wherein the model provides absolute risk probabilities of ever being a tobacco user.
42. 42. The method of any one of claims 38 to 41, wherein the model provides a value between 0 and 1, where a value of >0.460 is predictive of ever tobacco use.
43. 43. The method of any one of claims 38-42, wherein the model provides an absolute probability of being an ever tobacco user based on the level of each of the biomarker proteins selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, and CD248.
44. 44. The method of any one of claims 38-43, wherein the model provides an absolute probability of being an ever tobacco user based on the level of each of the biomarker proteins selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
45. 10. A method according to any one of the preceding claims, wherein the method comprises predicting the probability that a subject is an ever tobacco user for the purposes of determining health or life insurance premiums.
46. 46. The method of claim 45, wherein the method further comprises determining medical or life insurance coverage or premiums.
47. 47. The method of any one of claims 1 to 46, wherein the method further comprises using information resulting from the method to predict and / or manage healthcare resource utilization.
48. 48. The method of any one of claims 1 to 47, wherein the method further comprises using information resulting from the method to enable a decision to acquire or purchase a medical business, hospital, or company.
49. The method of any one of claims 1 to 40, wherein the tobacco user is a smoker.
50. A kit comprising N biomarker protein capture reagents, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, or at least 39, wherein at least one of the N biomarker protein capture reagents specifically binds to a biomarker protein selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
51. N is at least 2, and at least one of the two N biomarker protein capture reagents specifically binds to the biomarker protein selected from EPHA6, placental alkaline phosphatase, SCF, MASP3: heavy chain, DUS10, IgG4-κ, DCC, RCAN3, and CD248, and the other biomarker proteins among the two N biomarker proteins are EPHA6, placental alkaline phosphatase, SCF, MASP3: heavy chain, DUS10, IgG4-κ, DCC, RCAN3, and CD248. AN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
52. N is 2 to 39, or N is 3 to 39, or N is 4 to 39, or N is 5 to 39, or N is 6 to 39, or N is 7 to 39, or N is 8 to 39, or N is 9 to 39, or N is 10 to 39, or N is 11 to 39, or N is 12 to 39, or N is 13 to 39, or N is 14 to 39, or N is 15 to 39, or N is 16 to 39, or N is 17 to 39, or N is 18 to 39, or N is 19 to 39, or N is 52. The kit of claim 50 or 51, wherein N is 20-39, N is 21-39, N is 22-39, N is 23-39, N is 24-39, N is 25-39, N is 26-39, N is 27-39, N is 28-39, N is 29-39, N is 30-39, N is 31-39, N is 32-39, N is 33-39, N is 34-39, N is 35-39, N is 36-39, N is 37-39, or N is 38-39.
53. N is 2, or N is 3, or N is 4, or N is 5, or N is 6, or N is 7, or N is 8, or N is 9, or N is 10, or N is 11, or N is 12, or N is 13, or N is 14, or N is 15, or N is 16, or N is 17, or N is 18, or N is 19, or N is 20 or N is 21, or N is 22, or N is 23, or N is 24, or N is 25, or N is 26, or N is 27, or N is 28, or N is 29, or N is 30, or N is 31, or N is 32, or N is 33, or N is 34, or N is 35, or N is 36, or N is 37, or N is 38, or N is 39.
54. 54. The kit of any one of claims 50 to 53, wherein each of the N biomarker protein capture reagents specifically binds to a different biomarker protein.
55. 55. The kit of any one of claims 50 to 54, wherein each of the N biomarker protein capture reagents specifically binds to a biomarker protein selected from EPHA6, placental alkaline phosphatase, SCF, MASP3:heavy chain, DUS10, IgG4-κ, DCC, RCAN3, CD248, EDIL3, renin, PSP-94, PIGR, URB, agrin, secretoglobin family 3A member 1, SOX2, SREC-II, IGFALS, SLPI, WFDC1, REG4, MMP-8, LPLC1, NET4, PSP, DLDH, TCP10, MUC18, TAGL, TIMP-4, FAM3B, fibulin 1, CRAC1, PPBN, 6Ckine, cathepsin V, HE4, and PP2A subunit B.
56. 56. The kit of any one of claims 50 to 55, wherein at least one of the N biomarker protein capture reagents specifically binds to a biomarker protein selected from EPHA6, placental alkaline phosphatase, SCF, MASP3: heavy chain, DUS10, IgG4-κ, DCC, RCAN3, and CD248.
57. 57. The kit of any one of claims 50 to 56, wherein at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, or at least 9 of the N biomarker protein capture reagents specifically bind to a biomarker protein selected from EPHA6, placental alkaline phosphatase, SCF, MASP3: heavy chain, DUS10, IgG4-κ, DCC, RCAN3, and CD248.
58. 58. The kit of any one of claims 50 to 57, wherein two of the N biomarker protein capture reagents specifically bind to EPHA6 and placental alkaline phosphatase, or two of the N biomarker protein capture reagents specifically bind to EPHA6 and SCF, or two of the N biomarker protein capture reagents specifically bind to EPHA6 and MASP3: heavy chain, or two of the N biomarker protein capture reagents specifically bind to EPHA6 and DUS10, or two of the N biomarker protein capture reagents specifically bind to EPHA6 and IgG4-κ, or two of the N biomarker protein capture reagents specifically bind to EPHA6 and DCC, or two of the N biomarker protein capture reagents specifically bind to EPHA6 and RCAN3, or two of the N biomarker protein capture reagents specifically bind to EPHA6 and CD248.
59. 58. The kit of any one of claims 50 to 57, wherein two of the N biomarker protein capture reagents specifically bind to placental alkaline phosphatase and SCF, or two of the N biomarker protein capture reagents specifically bind to placental alkaline phosphatase and MASP3: heavy chain, or two of the N biomarker protein capture reagents specifically bind to placental alkaline phosphatase and DUS10, or two of the N biomarker protein capture reagents specifically bind to placental alkaline phosphatase and IgG4κ, or two of the N biomarker protein capture reagents specifically bind to placental alkaline phosphatase and DCC, or two of the N biomarker protein capture reagents specifically bind to placental alkaline phosphatase and RCAN3, or two of the N biomarker protein capture reagents specifically bind to placental alkaline phosphatase and CD248.
60. 58. The kit of any one of claims 50-57, wherein two of the N biomarker protein capture reagents specifically bind SCF and MASP3: heavy chain, or two of the N biomarker protein capture reagents specifically bind SCF and DUS10, or two of the N biomarker protein capture reagents specifically bind SCF and IgG4-κ, or two of the N biomarker protein capture reagents specifically bind SCF and DCC, or two of the N biomarker protein capture reagents specifically bind SCF and RCAN3, or two of the N biomarker protein capture reagents specifically bind SCF and CD248.
61. 58. The method of any one of claims 50-57, wherein two of the N biomarker protein capture reagents specifically bind to MASP3:heavy chain and DUS10, or two of the N biomarker protein capture reagents specifically bind to MASP3:heavy chain and IgG4-κ, and two of the N biomarker protein capture reagents specifically bind to MASP3:heavy chain and DCC, and two of the N biomarker protein capture reagents specifically bind to MASP3:heavy chain and RCAN3, and two of the N biomarker protein capture reagents specifically bind to MASP3:heavy chain and CD248.
62. The kit of any one of claims 50 to 57, wherein two of the N biomarker protein capture reagents specifically bind to DUS10 and IgG4-κ, or two of the N biomarker protein capture reagents specifically bind to DUS10 and DCC, or two of the N biomarker protein capture reagents specifically bind to DUS10 and RCAN3, or two of the N biomarker protein capture reagents specifically bind to DUS10 and CD248.
63. 58. The kit of any one of claims 50 to 57, wherein two of the N biomarker protein capture reagents specifically bind to IgG4-κ and DCC, or two of the N biomarker protein capture reagents specifically bind to IgG4-κ and RCAN3, or two of the N biomarker protein capture reagents specifically bind to IgG4-κ and CD248.
64. 58. The kit of any one of claims 50 to 57, wherein two of the N biomarker protein capture reagents specifically bind to DCC and RCAN3, or two of the N biomarker protein capture reagents specifically bind to DCC and CD248.
65. 58. The kit of any one of claims 50 to 57, wherein two of the N biomarker protein capture reagents specifically bind to RCAN3 and CD248.
66. A kit comprising N biomarker protein capture reagents, the kit comprising biomarker protein capture reagents for carrying out the method of any one of claims 1 to 49.
67. 67. The kit of any one of claims 50 to 66, wherein each of the N biomarker protein capture reagents is an antibody or an aptamer.
68. 68. The kit of claim 67, wherein each biomarker protein capture reagent is an aptamer.
69. 69. The kit of claim 68, wherein at least one aptamer is a slow dissociation aptamer.
70. The kit of claim 69, wherein at least one slow-dissociating aptamer comprises nucleotides with at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least 10 modifications.
71. Each slow dissociation aptamer exhibits a dissociation rate (t 1/2 71. The kit of claim 69 or claim 70, wherein the antibody binds to its target protein via a nucleotide sequence selected from the group consisting of nucleotides (A, B, C, D, E ...
72. 72. A kit according to any one of claims 50 to 71 for use in detecting the N biomarker proteins in a sample from a subject.
73. 73. The kit of claim 72 for use in predicting the probability that a subject is an ever tobacco user.
74. 74. The kit of claim 73, wherein the tobacco user is a smoker.