Pulmonary biomarkers and their uses
A method using physiochemically distinct particles to form biomolecular coronas with plasma samples for NSCLC detection addresses the limitations of current plasma protein detection, achieving high sensitivity and specificity for early NSCLC identification.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-03-10
AI Technical Summary
Current methods for early detection of non-small cell lung cancer (NSCLC) using plasma proteins are limited by wide concentration ranges and impractical complex biochemical workflows, leading to insufficient validation and replication in clinical trials.
A novel method involving physiochemically distinct particles to form biomolecular coronas with plasma samples, followed by mass spectrometry and classifier analysis to identify NSCLC with high sensitivity and specificity, using proteins like ANGL6, HTRA1, and others.
Enables high-throughput detection of NSCLC with sensitivity and specificity of 80% or higher, matching genome-based studies and identifying new biomarkers for improved early disease detection.
Smart Images

Figure 2026041851000009 
Figure 2026041851000010 
Figure 2026041851000011
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 62 / 967,995, filed January 30, 2020, which is incorporated herein by reference in its entirety. [Background technology]
[0002] background Early detection of NSCLC is key to favorable prognosis, yet little progress has been made in developing useful clinical trials. Plasma proteins should be a valuable biomarker discovery matrix, given plasma's contact with nearly every tissue in the body. However, plasma proteins can be problematic due to several factors, including a wide range of concentrations (e.g., 10 orders of magnitude). While complex biochemical workflows have attempted to circumvent these challenges, such workflows may be impractical for discovery studies of sufficient size to warrant validation and replication. Alternatively, biomarker research is limited to assessing or reassessing known markers without substantial improvement in clinical outcomes. Summary of the Invention [Means for solving the problem]
[0003] overview Disclosed herein are systems and methods for analyzing protein-particle and protein-protein interactions. Interactions between biomolecules and particles, and protein-protein interactions on particles, can provide insight into protein-protein interactions throughout a biological sample.
[0004] In various aspects, the present disclosure provides methods that include obtaining a dataset including protein information from biomolecular coronas corresponding to physiochemically distinct particles incubated with a biofluid sample from a subject; and using a classifier to identify biofluid samples indicative of a healthy condition, a cancerous condition, or a comorbidity thereof in the subject based on the dataset.
[0005] In some embodiments, the cancer is non-small cell lung cancer (NSCLC) and the comorbidity is a pulmonary comorbidity. In some embodiments, the pulmonary comorbidity is a chronic lung disease other than non-small cell lung cancer. In some embodiments, the pulmonary comorbidity is selected from the group consisting of chronic obstructive pulmonary disease (COPD), emphysema, cardiovascular disease, hypertension, pulmonary fibrosis, asthma, chronic lung disease, and any combination thereof. In some embodiments, the cancer condition is identified with a sensitivity or specificity of about 80% or higher.
[0006] In some embodiments, the protein information comprises expression information for a protein selected from the group consisting of angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), and platelet glycoprotein Ib beta chain (GP1BB). In some embodiments, the protein information comprises expression information for a protein selected from the group consisting of ANGL6, HTRA1, PXDN, ANTR2, CSPG2, ANTR1, NOTUM, CILP1, CAN2, and GP1BB.
[0007] In some embodiments, obtaining the dataset includes contacting the biological fluid sample with physiochemically distinct particles to form a biomolecular corona. In some embodiments, the physiochemically distinct particles include lipid particles, metal particles, silica particles, or polymer particles. In some embodiments, the physiochemically distinct particles include carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles.
[0008] In some aspects, obtaining the dataset comprises detecting biomolecular corona proteins by mass spectrometry, chromatography, liquid chromatography, high performance liquid chromatography, solid phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blot, dot blot, or immunostaining, or a combination thereof. In some aspects, obtaining the dataset comprises detecting biomolecular corona proteins by mass spectrometry. In some aspects, obtaining the dataset comprises measuring a readout indicative of the presence, absence, or amount of biomolecular corona proteins.
[0009] In some embodiments, the classifier is generated by removing or filtering out biomolecules associated with acute phase responses. In some embodiments, the NSCLC comprises early stage NSCLC (stage 1, stage 2, or stage 3). In some embodiments, the NSCLC comprises late stage NSCLC (stage 4). In some embodiments, the method further comprises administering an NSCLC treatment to the subject based on the disease state. In some embodiments, the classifier has increased protein detection consistency compared to a second classifier generated using proteome data from depleted plasma samples.
[0010] In some embodiments, the biological fluid comprises a blood sample from which red blood cells have been removed. In some embodiments, the biological fluid comprises plasma.
[0011] In various aspects, the present disclosure provides methods of assessing cancer status, comprising measuring biomarkers in a biological fluid sample from a subject suspected of having cancer to obtain biomarker measurements, wherein the biomarkers comprise one or more biomarkers selected from the group consisting of angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), anthrax toxin receptor 2 (ANTR2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), platelet glycoprotein Ib beta chain (GP1BB).
[0012] In some embodiments, measuring one or more biomarkers involves using a detection reagent that binds to the protein and produces a detectable signal. In some embodiments, measuring one or more biomarkers involves measuring a readout that indicates the presence, absence, or amount of one or more biomarkers. In some embodiments, measuring the biomarkers involves performing mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blot, dot blot, or immunostaining, or a combination thereof. In some embodiments, measuring the biomarkers involves performing mass spectrometry. In some embodiments, measuring the biomarkers involves performing an immunoassay. In some embodiments, measuring the biomarkers involves contacting the biological fluid sample with a plurality of physiochemically distinct nanoparticles.
[0013] In some aspects, the cancer comprises lung cancer. In some aspects, the method further comprises applying a classifier to the biomarker measurements. In some aspects, the classifier distinguishes between cancer and chronic lung disorder, chronic obstructive pulmonary disease, emphysema, cardiovascular disease, hypertension, pulmonary fibrosis, or asthma.
[0014] In some embodiments, the cancer comprises non-small cell lung cancer (NSCLC). In some aspects, the NSCLC comprises early stage NSCLC (stage 1, stage 2, or stage 3). In some aspects, the method further comprises applying a classifier to the biomarker measurements, wherein the classifier comprises features for distinguishing between early stage and late stage NSCLC. In some embodiments, the features comprise one of more biomarkers selected from the group consisting of SDC1, OC085, KV401, MYL6, JIP2, HV459, HV461, HV169, HNRPC, ROA1, STON2, LV301, KVD20, SAE1, PDE5A, RTN3, HV373, LV325, H2B1C, H2B1D, H2B1H, H2B1K, H2B1L, H2B1M, H2B1N, H2B2F, H2BFS, and NMT1. In some embodiments, the method further comprises applying a classifier to the biomarker measurements, wherein the classifier comprises features for distinguishing between the presence and absence of NSCLC. In some embodiments, the features include one of more biomarkers selected from the group consisting of SDC1, ANGL6, PXDN, ANTR1, OC085, SAA2, HTRA1, KPCB, KV401, OCL18, MYL6, ANTR2, GTPB2, HDGF, TBA1A, CSRP1, TCO2, CSPG2, PTPRZ, ILF2, SIAT1, ITA2B, DOK2, H31, H31T, H32, H33, H3C, RAC2, ARRB1, DHB4, HV102, RHG18, GDF15, PCSK6, FHOD1, or ITLN2.
[0015] In some embodiments, the method further comprises identifying the subject as having cancer based on the biomarker measurements. In some embodiments, the method further comprises administering a cancer treatment to the subject.
[0016] In some embodiments, the biological fluid comprises a blood sample that has had red blood cells removed or comprises plasma. In some embodiments, the subject is a human.
[0017] In various aspects, the present disclosure provides a method for detecting and classifying non-small cell lung cancer using a classifier (e.g., a trained classifier) comprising: (a) assaying a biological sample from a subject to identify a biomolecule; (b) using a classifier (e.g., a trained classifier) to identify the sample or the subject as positive or negative for non-small cell lung cancer based on the biomolecule identified in (a), wherein the trained classifier is trained using data from training samples that include known healthy samples and known non-small cell lung cancer samples, and the training samples are physicochemically modified to obtain the data. Assayed using a plurality of particles having distinctly different properties.
[0018] In some embodiments, the biomolecule comprises a protein. In some embodiments, the biomolecule is a protein. In some embodiments, the data comprises proteomic data identifying the presence or absence of a protein in the training sample. In some embodiments, the trained classifier is configured to remove acute phase response bias or stress protein bias. In some embodiments, the trained classifier comprises protein-related features, the features selected to remove acute phase response and / or stress protein bias in the biological sample.
[0019] In some embodiments, the features of the classifier exclude a protein selected from the group consisting of C-reactive protein (CRP), haptoglobin, and S10a8 / 9. In some embodiments, the features of the classifier exclude CRP, haptoglobin, and S10a8 / 9. In some embodiments, the features of the classifier exclude a protein listed in Table 5. In some embodiments, the features of the classifier include a plurality of proteins listed in Table 7. In some embodiments, the features of the classifier are listed in Table 7. In some embodiments, the features include tubulin alpha-1A chain (TBA1A) and syndecan-1 (SDC1).
[0020] In some embodiments, the method further comprises obtaining a biological sample from the subject. In some embodiments, the biological sample is a complex biological sample. The complex biological sample is a plasma sample or a serum sample. In some embodiments, the plurality of particles with distinct physicochemical properties comprises two or more particles listed in Table 4. In some embodiments, the plurality of particles with distinct physicochemical properties are listed in Table 4. In some embodiments, the trained classifier has improved performance based on area under the curve (AUC) compared to a classifier trained using proteomic data from a depleted plasma sample from the same subject as the known healthy sample and the known non-small cell lung cancer sample.
[0021] In some embodiments, the method further includes outputting a report indicating that the sample or the subject is positive or negative for the non-small cell lung cancer. In some embodiments, the assaying includes performing mass spectrometry or ELISA, and the biomolecule includes a protein. In some embodiments, the assaying includes targeted mass spectrometry. In some embodiments, the trained classifier is a trained algorithm. In some embodiments, the known non-small cell lung cancer sample includes an early stage (stages 1-3) non-small cell lung cancer sample. In some embodiments, the trained classifier identifies the sample or the subject as positive for the non-small cell lung cancer and identifies the stage of the non-small cell lung cancer.
[0022] In some embodiments, the method further comprises identifying the sample or the subject as positive or negative for non-small cell lung cancer with a sensitivity of greater than about 80%. In some embodiments, the sensitivity is greater than about 85%, about 90%, about 95%, or about 99%. In some embodiments, the method further comprises identifying the sample or the subject as positive or negative for non-small cell lung cancer with a specificity of greater than about 80%. In some embodiments, the specificity is greater than about 85%, about 90%, about 95%, or about 99%.
[0023] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0024] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the U.S. Patent and Trademark Office upon request and payment of the necessary fee. The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings, in which: [Brief explanation of the drawings]
[0025] [Figure 1] Figure 1 shows the age and sex breakout for the 268 subjects in the NSCLC biomarker discovery study.
[0026] [Figure 2] Figure 2 shows protein counts by each study group, including healthy, comorbid, NSCLC stage 1 "NSCLC_1", NSCLC stage 2 "NSCLC_2", NSCLC stage 3 "NSCLC_3", and NSCLC stage 4 "NSCLC_4".
[0027] [Figure 3] Figure 3 shows protein counts for the depleted plasma DP and particle panel.
[0028] [Figure 4] FIG. 4 shows the resulting summary of fractional detection of proteins across subjects versus their mean abundance for all 10 particle types of the particle panel and for depleted plasma (DP).
[0029] [Figure 5] FIG. 5 shows the performance of the cross-validated particle panel classifier, where the x-axis shows the fraction of classifications that are false positives and the y-axis shows the fraction of classifications that are true positives.
[0030] [Figure 6] Figure 6 shows graphs of the random forest model for NSCLC (stages 1, 2 and 3) in healthy individuals for depleted plasma (left) and the 10 particle panel (right), depicting the false positive rate on the x-axis and the true positive rate on the y-axis.
[0031] [Figure 7-1] FIG. 7 shows the performance of the classifier features across the study samples. [Figure 7-2] FIG. 7 shows the performance of the classifier features across the study samples. [Figure 7-3] FIG. 7 shows the performance of the classifier features across the study samples.
[0032] [Figure 8] FIG. 8 shows the results from 10 replicates of 10 rounds of 10-fold cross-validation with randomized target class assignment, with false positive rate on the x-axis and true positive rate on the y-axis.
[0033] [Figure 9] FIG. 9 shows the ROC plots for 13 peptides by MRM-MS and 2 proteins by ELISA after removal of proteins found in depleted plasma.
[0034] [Figure 10] Figure 10 shows the random forest model for the overall study group comparison.
[0035] [Figure 11-1] Figure 11 shows the differentiation of important characteristics in the study group comparisons. [Figure 11-2] Figure 11 shows the differentiation of important characteristics in the study group comparisons.
[0036] [Figure 12] FIG. 12 shows protein counts (e.g., number of proteins identified from corona analysis) for panel sizes ranging from 1 particle type to 12 particle types.
[0037] [Figure 13] FIG. 13 illustrates a computer system that is programmed or otherwise configured to implement the methods provided herein.
[0038] [Figure 14] FIG. 14 shows some example biomarkers for use in the classifiers described herein.
[0039] [Figure 15] FIG. 15 shows examples of biomarkers. DETAILED DESCRIPTION OF THE INVENTION
[0040] Detailed Description Disclosed herein are compositions and methods for identifying high-performance protein-based classifiers for comparing healthy versus non-small cell lung cancer (NSCLC) based on deep plasma protein profiling on a novel multi-particle-type panel platform. The compositions and methods disclosed herein facilitate the discovery of superior protein-based NSCLC biomarkers using the novel particle-type panel proteomic profiling platform disclosed herein for plasma protein quantification. Particles (e.g., nanoparticles) can specifically and reproducibly interrogate a subset of proteins from biofluids and can have high proteomic profiling efficiency and effectiveness. The low-complexity particle panel workflow disclosed herein enables the study of sizes, such as those described herein and larger. By profiling NSCLC subjects against healthy and pulmonary comorbidity control subjects, a multi-protein classifier panel was identified using the particle-based platform. These classifiers include previously unknown proteins that contribute to NSCLC. Therefore, the particle panel disclosed herein could identify new markers for improved early disease detection.
[0041] In one study, the particle panel disclosed herein identified 1,779 proteins in 288 subjects in 7 weeks, a throughput made possible by the simplicity and robustness of the particle platform. The performance of the classifier for early-stage NSCLC (stages 1, 2, and 3) in healthy subjects (AUC 0.90) included known and unknown proteins involved in NSCLC. Therefore, the proteins disclosed herein enable a better assay for early disease detection. This marks the first time that a deep plasma protein biomarker profiling study has achieved throughput that matches genome-based studies and enables complementary studies involving proteins and nucleic acids.
[0042] The present disclosure also provides methods for using a classifier trained by training a classifier with biomarkers found to be associated with NSCLC to classify samples as healthy, comorbid, or NSCLC, which sensitively and specifically distinguishes NSCLC from healthy and comorbid states.
[0043] The biomarkers disclosed herein were developed using particle panels having one or more distinct particle types, which were then incubated with a sample to form a biomolecular corona on the particle surface and assayed for proteins in the biomolecular corona. The particle panel can have multiple distinct particle types, and proteins from the sample are enriched on the distinct biomolecular coronas formed on the surface of these distinct particle types. The particle types included in the particle panels disclosed herein are particularly well suited to enriching multiple proteins over a wide dynamic range in an unbiased manner. The combinations of particle types selected for inclusion in the particle panels of the present disclosure vary in their physicochemical properties (e.g., size, surface charge, core material, shell material, surface chemistry, porosity, morphology, and other properties). However, particle types may share some of the physicochemical properties. As used herein, the term "biomolecular corona" may be used interchangeably with the term "protein corona" and refers to the formation of a layer of protein on the surface of a particle after the particle is contacted with a sample (e.g., plasma). This method is sometimes referred to interchangeably as corona analysis, or in some instances, as "proteographic" analysis, which combines a multi-particle-type protein corona strategy with mass spectrometry (MS). The particle types included in the particle panels disclosed herein can be superparamagnetic and thus rapidly separated or isolated from unbound proteins (proteins not adsorbed to the particle surface to form a corona) after incubation of the particles in a sample. Biomolecules can include proteins. Methods involving biomolecule or protein coronas can involve either biomolecules or proteins, or vice versa.
[0044] Disclosed herein is a method comprising obtaining a dataset comprising protein information from biomolecular coronas corresponding to physiochemically distinct particles incubated with a biological fluid sample from a subject; and using a classifier to identify a biological fluid sample indicative of a healthy state, a cancerous state, or a comorbidity thereof in the subject based on the dataset.
[0045] Disclosed herein is a method that includes obtaining a dataset including proteins detected in a biomolecular corona corresponding to physiochemically distinct particles incubated with a biological sample, including a biological fluid sample or a blood sample from which red blood cells have been removed (e.g., an acellular biological sample such as plasma); and identifying a disease state based on the dataset using a classifier.
[0046] Disclosed herein is a method for assessing cancer status, comprising measuring biomarkers in a biological sample from a subject suspected of having cancer to obtain a biomarker measurement value from the biomarkers described herein. The biomarkers may include one or more biomarkers selected from the group consisting of angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), anthrax toxin receptor 2 (ANTR2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), and platelet glycoprotein Ib beta chain (GP1BB).
[0047] Disclosed herein is a method for assaying one or more biomarkers in a sample from a subject suspected of having lung cancer, the method comprising measuring one or more biomarkers in the sample, for example, a biomarker selected from the group consisting of angiopoietin-related protein 6 (ANGL6), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), 60S acidic ribosomal protein P2 (RLA2), and platelet glycoprotein Ib beta chain (GP1BB), or peptide fragments thereof, to detect the presence, absence, or amount of the one or more biomarkers.
[0048] 1. A method for assaying one or more biomarkers in a sample from a subject suspected of having non-small cell lung cancer (NSCLC), comprising: detecting one or more biomarkers in the sample, such as angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 2 (ANTR2), and / or vasoconstrictor protein 1 (VSPG1). Disclosed herein are methods that include measuring a biomarker selected from the group consisting of anticoagulant receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), or platelet glycoprotein Ib beta chain (GP1BB), or a peptide fragment thereof, to detect the presence, absence, or amount of one or more biomarkers.
[0049] 1. A method of treatment, comprising: (a) detecting one or more biomarkers in a sample from a subject suspected of having lung cancer, such as angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate zona protein (CARDI). Disclosed herein are methods comprising: (a) obtaining or receiving a measurement of a biomarker selected from the group consisting of calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), and platelet glycoprotein Ib beta chain (GP1BB), or a peptide fragment thereof; and (b) administering to the subject a lung cancer treatment based on the presence of the one or more biomarkers measured in (a), and monitoring the subject without administering the lung cancer treatment based on the absence of the one or more biomarkers in (a).
[0050] Disclosed herein are methods comprising: (a) assaying a biological sample from a subject to identify a biomolecule; and (b) using a classifier to identify the sample as positive or negative for non-small cell lung cancer (NSCLC) based on the biomolecule identified in (a), wherein the classifier is generated using data from the sample assayed using a plurality of particles having distinct physicochemical properties to obtain the data.
[0051] Disclosed herein is a system including: (a) a communications interface that receives, via a communications network, biomolecular data from a plurality of particles having distinct physicochemical properties that have been exposed to a sample from a subject that contains biomolecules; and (b) a computer in communication with the communications interface, the computer including a computer processor and a computer-readable medium, the computer-readable medium including machine-executable code that, when executed by the computer processor, implements a method that includes: (i) receiving the biomolecular data via the communications network; (ii) combining the biomolecular data to generate a biomolecular fingerprint for the sample; and (iii) assigning a label to the biomolecular fingerprint that corresponds to the presence or absence of non-small cell lung cancer (NSCLC) in the subject.
[0052] 1. A system including a communications interface that receives, via a communications network, biomarker data from a sample from a subject suspected of having non-small cell lung cancer (NSCLC), the sample comprising biomarkers, the biomarkers being selected from the group consisting of angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), and ribosomal protein 1 (RIP1). Disclosed herein is a system comprising one or more biomarkers selected from the group consisting of versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), or platelet glycoprotein Ib beta chain (GP1BB), or a peptide fragment thereof.
[0053] Biomarkers Biomarkers disclosed herein (e.g., for a disease state, e.g., NSCLC, a comorbidity, or a healthy state) can include at least one of the following: protein S100-A9 (P06702; S10A9_HUMAN), C-reactive protein (P02741; CRP_HUMAN), inter-alpha trypsin inhibitor heavy chain H2 (P19823; ITIH2_HUMAN), protein S100-A8 (P05109; S10A8_HUMAN), serine protease HTRA1 (Q92743; HTRA1_HUMAN), α-glucanase (Q92743; α-glucanase) ... Angiopoietin-related protein 6 (Q8NI99; ANGL6_HUMAN), haptoglobin-related protein (P00739; HPTR_HUMAN), CC motif chemokine 18 (P55774; CCL18_HUMAN), actin, cytoplasmic 1 (P60709; ACTB_HUMAN), actin, cytoplasmic 2 (P63261; ACTG_HUMAN), serum amyloid A-1 protein (P0DJI8; SAA1_HUMAN), immunoglobulin kappa constant (P01834; IGKC_HUMAN), angiopoietin-related protein 6 (Q8NI99; ANGL6_HUMAN), peroxidasin homolog (Q92743; PXDN_HUMAN), anthrax toxin receptor 2 (P58335; ANTR2_HUMAN), tubulin alpha-1A chain (Q71U36; TBA1A_HUMAN), syndecan-1 (P18827; SDC1_HUMAN), serum amyloid A-2 protein (P0DJI9; SAA2_HUMAN), versican core protein (P13611; CSPG2_HUMAN), anthrax toxin receptor 1 (Q9H6X2; ANTR1_HUMAN), palmitoyltransferase (P18827; SDC1_HUMAN), and ribosomal protein (P18827; SDC1_HUMAN). leoyl-protein carboxylesterase NOTUM (Q6P988; NOTUM_HUMAN), cartilage intermediate lamina protein 1 (O75339; CILP1_HUMAN), calpain-2 catalytic subunit (P17655; CAN2_HUMAN), 60S acidic ribosomal protein P2 (P05387; RLA2_HUMAN), beta-galactoside alpha-2,6-sialyltransferase 1 (P15907; SIAT1_HUMAN), and platelet glycoprotein Ib beta chain (P13224; GP1BB_HUMAN).The biomarkers may include any biomarker(s) in Figure 11. Any one or more of the above biomarkers in various combinations can be used to train a classifier to distinguish whether a subject has lung cancer (e.g., NSCLC), or whether they have co-morbidities or are healthy. In some embodiments, at least one of the biomarkers, at least two of the biomarkers, at least three of the biomarkers, at least four of the biomarkers, at least five of the biomarkers, at least six of the biomarkers, at least seven of the biomarkers, at least eight of the biomarkers, at least nine of the biomarkers, at least ten of the biomarkers, at least fifteen of the biomarkers, at least twenty of the biomarkers, at least twenty-five of the biomarkers, or all of the biomarkers can be used together to train a classifier to distinguish whether a subject has lung cancer (e.g., NSCLC), or whether they have co-morbidities or are healthy. In some embodiments, at least one of the biomarkers, at least two of the biomarkers, at least three of the biomarkers, at least four of the biomarkers, at least five of the biomarkers, at least six of the biomarkers, at least seven of the biomarkers, at least eight of the biomarkers, at least nine of the biomarkers, at least ten of the biomarkers, at least fifteen of the biomarkers, at least twenty of the biomarkers, at least twenty-five of the biomarkers, or all of the biomarkers together can be used in a diagnostic assay to determine whether a subject has lung cancer. Diagnostic assays can be performed using the trained classifiers disclosed herein.
[0054] The present disclosure provides a method for detecting low-abundance peptides in complex biological samples. Many of the diagnostic peptides disclosed herein are inaccessible by traditional blood analysis methods due to the high concentrations of albumin, immunoglobulins, and other high-abundance blood proteins. Diagnostic peptides may exist at concentrations three, four, five, six, seven, eight, nine, ten, eleven, twelve, or more orders of magnitude lower than the highest-abundance proteins in blood samples, and therefore would not be detectable by many traditional proteomic methods. The present disclosure not only provides a method for enriching low-abundance biomolecules (e.g., proteins) from complex biological samples such as plasma, but also a method for quantifying the enriched biomolecules.
[0055] Examples of lung cancer diagnostic peptides are provided in Table 1. The methods of the present disclosure may include assaying a sample from a subject to detect the presence, absence, or abundance of one or more peptides or fragments of peptides from among the peptides listed in Table 1. In some cases, the methods include identifying a ratio between the abundance of two peptides or fragments of peptides from among the peptides listed in Table 1. In some cases, the methods include identifying a ratio between the abundance of a peptide or fragment of a peptide to the abundance of another peptide from among the peptides listed in Table 1 from the same biological sample. For example, a method may include identifying the ratio of the relative abundance of APOC1 and ceruloplasmin in a plasma sample from a subject suspected of having lung cancer. In some cases, the method includes assaying the sample to detect the presence, absence, or abundance of one or more peptides or fragments of peptides from the group consisting of angiopoietin-related protein 6 (ANGL6), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), 60S acidic ribosomal protein P2 (RLA2), and platelet glycoprotein Ib beta chain (GP1BB). In some cases, the method includes assaying the sample to detect the presence, absence, or abundance of at least two, at least three, at least four, at least five, at least six, at least eight, at least ten, at least twelve, at least fifteen, at least twenty, at least twenty-five, at least thirty, or at least thirty-five peptides or fragments of peptides from the peptides listed in Table 1.
[0056] The disclosed methods allow for quantification of distinct biomarkers across a wide concentration range. In some cases, cancer (e.g., NSCLC) is evidenced by the relative concentrations of two or more proteins from a patient sample. In some cases, the disclosed methods include identifying an abundance (e.g., concentration) ratio between at least two peptides from among the peptides listed in Table 1. In some cases, the disclosed methods include identifying an abundance ratio between at least three peptides from among the peptides listed in Table 1. In some cases, the disclosed methods include identifying an abundance ratio between at least four peptides from among the peptides listed in Table 1. In some cases, the disclosed methods include identifying an abundance ratio between at least five peptides from among the peptides listed in Table 1. In some cases, the disclosed methods include identifying an abundance ratio between at least six peptides from among the peptides listed in Table 1. In some cases, the disclosed methods include identifying an abundance ratio between at least seven peptides from among the peptides listed in Table 1. In some cases, the methods of the present disclosure include identifying abundance ratios between at least 8 peptides from among the peptides listed in Table 1. In some cases, the methods of the present disclosure include identifying abundance ratios between at least 9 peptides from among the peptides listed in Table 1. In some cases, the methods of the present disclosure include identifying abundance ratios between at least 10 peptides from among the peptides listed in Table 1. In some cases, the methods of the present disclosure include identifying abundance ratios between at least 12 peptides from among the peptides listed in Table 1. In some cases, the methods of the present disclosure include identifying abundance ratios between at least 15 peptides from among the peptides listed in Table 1. In some cases, the methods of the present disclosure include identifying abundance ratios between at least 20 peptides from among the peptides listed in Table 1. In some cases, the methods of the present disclosure include identifying abundance ratios between at least 25 peptides from among the peptides listed in Table 1. In some cases, the sample is a blood sample (e.g., plasma).
[0057] In some cases, the method includes assaying the sample to detect the presence, absence, or abundance of at least two, at least three, at least four, or all five of ANGL6, NOTUM, CILP1, RLA2, or GP1BB. In some cases, the one or more peptides or fragments of peptides from among the peptides listed in Table 1 are selected from the group consisting of actin (e.g., beta-actin), anthrax toxin receptor 2, cartilage intermediate lamina protein 1, collectin 11, and kallistatin. In some cases, one or more peptides or peptide fragments from among the peptides listed in Table 1 are identified as a specific protein or fragment of angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versicolor protein (VAS), or vasopressin (VAS). and selected from the group consisting of can core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), and platelet glycoprotein Ib beta chain (GP1BB).In some cases, the one or more peptides or fragments of peptides from among those listed in Table 1 are selected from the group consisting of __, and the one or more biomarkers are leucine-rich alpha-2-glycoprotein (A2GL); actin, cytoplasmic 1 (ACTB); actin, cytoplasmic 2 (ACTG); apolipoprotein CI (APOC1); apolipoprotein M (APOM); voltage-gated calcium channel subunit alpha-2 / delta-1 (CA2D1); cadherin-13 (CAD13); beta-Ala-His dipeptidase (CNDP1); ciliary body Further included are neurotrophin receptor subunit alpha (CNTFR); collectin-11 (COL11); C-reactive protein (CRP); hemoglobin subunit alpha (HBA); haptoglobin-related protein (HPT); haptoglobin-related protein (HPTR); inter-alpha trypsin inhibitor heavy chain H2 (ITIH2); kallistatin (KAIN); plasma kallikrein (KLKB1); neural cell adhesion molecule 1 (NCAM1); protein S100-A8 (S10A8); protein S100-A9 (S10A9); and maintenance of chromosome structure protein 4 (SMC4). In some cases, one or more peptides or fragments of peptides from among the peptides listed in Table 1 are selected from the group consisting of A2GL, ACTB, ACTG, APOC1, APOM, CA2D1, CAD13, CNDP1, CNTFR, COL11, CRP, HBA, HPT, HPTR, ITIH2, KAIN, KLKB1, NCAM1, S10A8, S10A9, or SMC4.In some cases, one or more peptides or fragments of peptides from among the peptides listed in Table 1 include at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of A2GL, ACTB, ACTG, APOC1, APOM, CA2D1, CAD13, CNDP1, CNTFR, COL11, CRP, HBA, HPT, HPTR, ITIH2, KAIN, KLKB1, NCAM1, S10A8, S10A9, or SMC4. [Table 1]
[0058] In some cases, the methods include targeting angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitate 1 (PA1), and ribosomal protein 1 (RI). The method includes detecting the presence, absence, or abundance of one or more peptides selected from the group consisting of oil-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), and platelet glycoprotein Ib beta chain (GP1BB). In some cases, the method includes identifying angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), and / or anthrax toxin receptor 2 (ANTR2). The method includes identifying the ratio between the abundances of two peptides selected from the group consisting of palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), and platelet glycoprotein Ib beta chain (GP1BB).In some cases, the method includes identifying angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM). , cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), and platelet glycoprotein Ib beta chain (GP1BB).
[0059] Biomarkers (e.g., proteins) may include angiopoietin-related proteins, serine proteases, peroxidasin homologs, CC motif chemokines, anthrax toxin receptors, tubulin proteins, syndecan proteins, serine amyloid A proteins, versican proteins, anthrax toxin receptor proteins, palmitoleoyl-protein carboxylesterase proteins, cartilage intermediate lamina proteins, calpain proteins or subunits, 60S acidic ribosomal proteins, beta-galactoside alpha-2,6-sialyltransferase proteins, or platelet glycoproteins, or subunits or fragments of any of the foregoing proteins. Biomarkers may include angiopoietin-related proteins. Biomarkers may include serine proteases. Biomarkers may include peroxidasin homologs. Biomarkers may include CC motif chemokines. Biomarkers may include anthrax toxin receptors. Biomarkers may include tubulin proteins. Biomarkers may include syndecan proteins. The biomarker may include serum amyloid A protein. The biomarker may include versican protein. The biomarker may include anthrax toxin receptor protein. The biomarker may include palmitoleoyl-protein carboxylesterase protein. The biomarker may include cartilage intermediate lamina protein. The biomarker may include calpain protein or subunit. The biomarker may include 60S acidic ribosomal protein. The biomarker may include beta-galactoside alpha-2,6-sialyltransferase protein. The biomarker may include platelet glycoprotein. The biomarker may be secreted.
[0060] Biomarkers (e.g., proteins) can include angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), or platelet glycoprotein Ib beta chain (GP1BB). Biomarkers (e.g., proteins) can include angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), and platelet glycoprotein Ib beta chain (GP1BB). In some cases, the biomarker is a secreted protein.
[0061] The biomarkers may include ANGL6, HTRA1, PXDN, ANTR2, CSPG2, ANTR1, NOTUM, CILP1, CAN2, or GP1BB.The biomarkers may include ANGL6, HTRA1, PXDN, ANTR2, CSPG2, ANTR1, NOTUM, CILP1, CAN2, and GP1BB.
[0062] In some cases, the method includes assaying a plasma sample to detect the presence, absence, or abundance of one or more peptides or fragments of peptides from among the peptides listed in Table 1. In some cases, the method includes assaying a buffy coat sample to detect the presence, absence, or abundance of one or more peptides or fragments of peptides from among the peptides listed in Table 1. In some cases, the method includes assaying a granulocyte sample to detect the presence, absence, or abundance of one or more peptides or fragments of peptides from among the peptides listed in Table 1. In some cases, the method includes assaying a homogenized tissue (e.g., a homogenized lung biopsy tissue sample) to detect the presence, absence, or abundance of one or more peptides or fragments of peptides from among the peptides listed in Table 1.
[0063] The present methods enable rapid and deep biomolecular profiling from complex biological samples. In many cases, the methods detect and identify hundreds or thousands of distinct biomolecules. Such extensive analysis not only enables deeper profiling of complex samples but also increases the diagnostic utility of individual peptides. The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 50 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1. The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 100 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1. The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 200 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1. The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 400 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1. The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 600 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1. The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 800 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1. The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 1000 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1.The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 1200 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1. The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 1400 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1. The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 1600 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1. The disclosed methods may include assaying a sample from a subject to detect the presence, absence, or abundance of at least 1800 peptides from the biological sample in addition to one or more additional peptides or peptide fragments from among the peptides listed in Table 1. Methods of the disclosure may include identifying an abundance or signal intensity (e.g., mass spectrometric signal intensity) ratio of at least a subset of at least 50, at least 100, at least 200, at least 400, at least 600, at least 800, at least 1000, at least 1200, at least 1400, at least 1600, or at least 1800 peptides to one or more additional peptides or fragments of peptides from among the peptides listed in Table 1.
[0064] The methods of the present disclosure may include monitoring lung cancer progression over time. The method may include collecting two samples from a patient at two different time points and detecting at least two peptides from among the peptides listed in Table 1 in each of the samples. The method may include collecting two samples from a patient at two different time points and detecting at least three peptides from among the peptides listed in Table 1 in each of the samples. The method may include collecting two samples from a patient at two different time points and detecting at least four peptides from among the peptides listed in Table 1 in each of the samples. The method may include collecting two samples from a patient at two different time points and detecting at least five peptides from among the peptides listed in Table 1 in each of the samples. The method may include collecting two samples from a patient at two different time points and detecting at least six peptides from among the peptides listed in Table 1 in each of the samples. The method may include collecting two samples from a patient at two different time points and detecting at least seven peptides from among the peptides listed in Table 1 in each of the samples. The method may include collecting two samples from a patient at two different time points and detecting at least eight peptides from among the peptides listed in Table 1 in each of the samples. The method may include collecting two samples from a patient at two different time points and detecting at least nine peptides from among the peptides listed in Table 1 in each of the samples. The method may include collecting two samples from a patient at two different time points and detecting at least ten peptides from among the peptides listed in Table 1 in each of the samples. The method may include collecting two samples from a patient at two different time points and detecting at least 12 peptides from among the peptides listed in Table 1 in each of the samples. The method may include collecting two samples from a patient at two different time points and detecting at least 15 peptides from among the peptides listed in Table 1 in each of the samples.The method may include collecting two samples from a patient at two different time points and detecting at least 20 peptides from the peptides listed in Table 1 in each sample. The second of the two samples may be collected at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 5 weeks, at least 6 weeks, at least 8 weeks, at least 12 weeks, at least 15 weeks, at least 18 weeks, at least 24 weeks, at least 36 weeks, at least 52 weeks, at least 78 weeks, at least 104 weeks, at least 130 weeks, at least 156 weeks, at least 208 weeks, or at least 260 weeks apart. One or both samples may be collected during cancer treatment, such as chemotherapy, to determine the effectiveness of the treatment. A sample may be collected during cancer remission to detect recurrence, dormancy, or progression to complete remission.
[0065] Disclosed herein are methods involving biomarkers. The biomarkers may include angiopoietin-related protein 6 (ANGL6), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), 60S acidic ribosomal protein P2 (RLA2), and platelet glycoprotein Ib beta chain (GP1BB), or peptide fragments thereof. The biomarkers may include at least one, at least two, at least three, or at least four of ANGL6, NOTUM, CILP1, RLA2, or GP1BB. The biomarkers may include ANGL6, NOTUM, CILP1, RLA2, and GP1BB. In some cases, any of these biomarkers is useful for identifying lung cancer. The biomarkers can be included in a classifier for distinguishing between lung cancers.
[0066] Disclosed herein are methods involving biomarkers. Biomarkers may include angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), or platelet glycoprotein Ib beta chain (GP1BB), or a peptide fragment thereof. The biomarkers may include at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 11, at least 12, at least 13, at least 14, or at least 15 of ANGL6, HTRA1, PXDN, CCL18, ANTR2, TBA1A, SDC1, SAA2, CSPG2, ANTR1, NOTUM, CILP1, CAN2, RLA2, SIAT1, or GP1BB. The biomarkers may include ANGL6, HTRA1, PXDN, CCL18, ANTR2, TBA1A, SDC1, SAA2, CSPG2, ANTR1, NOTUM, CILP1, CAN2, RLA2, SIAT1, and GP1BB. The biomarkers may be included in a classifier.
[0067] Disclosed herein are methods including biomarkers, such as leucine-rich alpha-2-glycoprotein (A2GL); actin, cytoplasmic 1 (ACTB); actin, cytoplasmic 2 (ACTG); apolipoprotein CI (APOC1); apolipoprotein M (APOM); voltage-gated calcium channel subunit alpha-2 / delta-1 (CA2D1); cadherin-13 (CAD13); beta-Ala-His dipeptidase (CNDP1); ciliary neurotrophic factor receptor subunit alpha (CNTFR); collectin-11 (C OL11), C-reactive protein (CRP); hemoglobin subunit alpha (HBA); haptoglobin-associated protein (HPT); haptoglobin-related protein (HPTR); inter-alpha trypsin inhibitor heavy chain H2 (ITIH2); kallistatin (KAIN); plasma kallikrein (KLKB1); neural cell adhesion molecule 1 (NCAM1); protein S100-A8 (S10A8); protein S100-A9 (S10A9); or maintenance of chromosome structure protein 4 (SMC4). The biomarkers may include at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of A2GL, ACTB, ACTG, APOC1, APOM, CA2D1, CAD13, CNDP1, CNTFR, COL11, CRP, HBA, HPT, HPTR, ITIH2, KAIN, KLKB1, NCAM1, S10A8, S10A9, or SMC4. Biomarkers may include A2GL, ACTB, ACTG, APOC1, APOM, CA2D1, CAD13, CNDP1, CNTFR, COL11, CRP, HBA, HPT, HPTR, ITIH2, KAIN, KLKB1, NCAM1, S10A8, S10A9, and SMC4. Biomarkers can be included in a classifier.
[0068] Disclosed herein are methods or classifiers comprising a biomarker (or multiple biomarkers). The biomarker may include ANGL6. The biomarker may include HTRA1. The biomarker may include PXDN. The biomarker may include CCL18. The biomarker may include ANTR2. The biomarker may include TBA1A. The biomarker may include SDC1. The biomarker may include SAA2. The biomarker may include CSPG2. The biomarker may include ANTR1. The biomarker may include NOTUM. The biomarker may include CILP1. The biomarker may include CAN2. The biomarker may include RLA2. The biomarker may include SIAT1. The biomarker may include GP1BB. The biomarker may include A2GL. The biomarker may include ACTB. The biomarker may include ACTG. The biomarker may include APOC1. The biomarker may include APOM. The biomarker may include CA2D1. The biomarker may include CAD13. The biomarker may include CNDP1. The biomarker may include CNTFR. The biomarker may include COL11. The biomarker may include CRP. The biomarker may include HBA. The biomarker may include HPT. The biomarker may include HPTR. The biomarker may include ITIH2. The biomarker may include KAIN. The biomarker may include KLKB1. The biomarker may include NCAM1. The biomarker may include S10A8. The biomarker may include S10A9. The biomarker may include SMC4.
[0069] Disclosed herein are methods and classifiers comprising biomarkers. The biomarkers may exclude ANGL6. The biomarkers may exclude HTRA1. The biomarkers may exclude PXDN. The biomarkers may exclude CCL18. The biomarkers may exclude ANTR2. The biomarkers may exclude TBA1A. The biomarkers may exclude SDC1. The biomarkers may exclude SAA2. The biomarkers may exclude CSPG2. The biomarkers may exclude ANTR1. The biomarkers may exclude NOTUM. The biomarkers may exclude CILP1. The biomarkers may exclude CAN2. The biomarkers may exclude RLA2. The biomarkers may exclude SIAT1. The biomarkers may exclude GP1BB. The biomarkers may exclude A2GL. The biomarkers may exclude ACTB. The biomarkers may exclude ACTG. The biomarker may exclude APOC1. The biomarker may exclude APOM. The biomarker may exclude CA2D1. The biomarker may exclude CAD13. The biomarker may exclude CNDP1. The biomarker may exclude CNTFR. The biomarker may exclude COL11. The biomarker may exclude CRP. The biomarker may exclude HBA. The biomarker may exclude HPT. The biomarker may exclude HPTR. The biomarker may exclude ITIH2. The biomarker may exclude KAIN. The biomarker may exclude KLKB1. The biomarker may exclude NCAM1. The biomarker may exclude S10A8. The biomarker may exclude S10A9. The biomarker may exclude SMC4.
[0070] Classification index (classifier) The methods described herein may include the use of a classifier. The methods described herein may include generating a classifier. The methods described herein may include using the classifier to identify a disease state based on a dataset. The methods described herein may include applying the classifier to biomarker measurements.
[0071] A method for determining a set of proteins associated with a disease or disorder and / or disease state involves analyzing biomarkers (e.g., corona or proteins) of at least one or two samples. This determination, analysis, and statistical classification can be performed by methods known in the art, including, but not limited to, a wide range of supervised and unsupervised data analysis, machine learning, deep learning, and clustering techniques, including, for example, hierarchical cluster analysis (HCA), principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA), random forests, logistic regression, decision trees, support vector machines (SVM), k-nearest neighbors, naive Bayes, linear regression, polynomial regression, SVM for regression, k-means clustering, and hidden Markov models, among others. In other words, the proteins of each sample (e.g., in the corona) are compared / analyzed with each other using statistical significance to determine what patterns are common among the proteins of interest and determine a set of proteins associated with a disease or disorder or disease state. Any of these methods can be used to generate a classifier for use herein.
[0072] A model can be trained with one or more biomarkers using deep learning, hierarchical cluster analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k-nearest neighbor analysis, naive Bayes analysis, k-means clustering analysis, or hidden Markov analysis. A model can be trained with one or more biomarkers using deep learning. A model can be trained with one or more biomarkers using hierarchical cluster analysis. A model can be trained with one or more biomarkers using principal component analysis. A model can be trained with one or more biomarkers using partial least squares discriminant analysis. A model can be trained with one or more biomarkers using random forest classification analysis. A model can be trained with one or more biomarkers using support vector machine analysis. A model can be trained with one or more biomarkers using k-nearest neighbor analysis. A model can be trained with one or more biomarkers using naive Bayes analysis. A model can be trained with one or more biomarkers using K-means cluster analysis. A model can be trained with one or more biomarkers using hidden Markov analysis. The methods described herein may include the use of a model. The method may include generating a model.
[0073] The model can be trained with the measured values of biomarkers (for example, any of those described herein) in control samples from control subjects.In some cases, the one or more biomarkers that the model is trained on do not include depleted plasma proteins.The control subjects can have a specific stage of NSCLC.
[0074] Generally, machine learning algorithms are used to build models that accurately assign class labels to examples based on input features describing the examples (e.g., healthy, comorbid, or NSCLC stage 1, 2, or 3). In some cases, it may be advantageous to utilize machine learning and / or deep learning techniques in the methods described herein. For example, machine learning can be used to correlate a set of biomarkers (ser) with various disease states (e.g., The proteome data can be correlated with a prognosis (e.g., disease-free, disease precursor, early or late stage of disease, etc.). For example, in some cases, one or more machine learning algorithms are utilized in connection with the methods of the present invention to analyze data detected or obtained by the protein corona and protein set obtained from the methods of the present invention. For example, in one embodiment, machine learning is coupled with the particle panels described herein to determine whether a subject is in a pre-cancer stage, has cancer, or has or has not developed cancer, as well as to differentiate between cancer types, e.g., lung cancers such as NSCLC. A classifier can have increased protein detection consistency compared to a second classifier generated using proteome data from a depleted plasma sample. For example, a classifier can be generated by contacting a sample with particles and have increased protein detection consistency compared to a second classifier generated using proteome data from a depleted plasma sample that has not been contacted with particles.
[0075] The determination, analysis, and statistical classification can be performed using methods known in the art, including, for example, a wide range of supervised and unsupervised data analysis and clustering techniques, such as hierarchical cluster analysis (HCA), principal component analysis (PCA), partial least squares discriminant analysis (PLSDA), machine learning (also known as random forests), logistic regression, decision trees, support vector machines (SVM), k-nearest neighbors, naive Bayes, linear regression, polynomial regression, SVM for regression, k-means clustering, and hidden Markov models, among others. The system or method can analyze biomarkers, such as protein sets or protein coronas, of the present disclosure. The analysis can include comparing / analyzing biomarkers from one or more (e.g., several) samples to determine what patterns are common among the biomarkers with statistical significance to determine biomarkers (e.g., protein sets) associated with a biological state. The system or method can be used to develop classifiers for detecting and distinguishing between different protein sets or protein coronas (e.g., compositional features of protein coronas). Data collected from a method or system described herein (e.g., a system including a sensor array) can be used to train a machine learning algorithm, e.g., an algorithm that receives array measurements from patients and outputs the specific biomolecular corona composition from each patient.
[0076] Machine learning can be generalized as the ability of a learning machine to accurately perform new, unseen examples / tasks after experiencing a training dataset. Machine learning can include the following concepts and methods: Supervised learning concepts include AODE; artificial neural networks, e.g., backpropagation, autoencoders, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, and spiking neural networks; Bayesian statistics, e.g., Bayesian networks and Bayesian knowledge bases; case-based reasoning; Gaussian process regression; gene expression programming; group data processing methods (GMDH); inductive logic programming; instance-based learning; lazy learning; learning automata; learning vector quantization; logistic model trees; minimum message length (decision trees, decision graphs, etc.), e.g., nearest neighbor algorithms and analogical modeling; probabilistic and approximately correct learning (PAC) learning; ripple-down rules, knowledge bases, and the like. These may include: knowledge acquisition methodologies; symbolic machine learning algorithms; support vector machines; random forests; ensembles of classifiers, e.g., bootstrap aggregating (bagging) and boosting (meta-algorithms); ordinal classification; information fuzzy networks (IFNs); conditional random fields; ANOVA; linear classifiers, e.g., Fisher linear discriminant, linear regression, logistic regression, multinomial logistic regression, naive Bayes classifier, perceptron, support vector machines; quadratic classifiers; k-nearest neighbors; boosting; decision trees, e.g., C4.5, random forest, ID3, CART, SLIQ, SPRINT; Bayesian networks, e.g., naive Bayes; and hidden Markov models. Unsupervised learning concepts may include expectation-maximization algorithms; vector quantization; generative phase maps; information bottom-neck methods; artificial neural networks, e.g., self-organizing maps; association rule learning, e.g., the Apriori algorithm, the Eclat algorithm, and the FPgrowth algorithm; hierarchical clustering, e.g., shortest distance clustering and concept clustering; cluster analysis, e.g., the K-means algorithm, fuzzy clustering, DBSCAN, and the OPTICS algorithm; and outlier detection, e.g., the local outlier factor method.Semi-supervised learning concepts may include generative models, sparse separation, graph-based methods, and co-training. Reinforcement learning concepts may include temporal difference learning, Q-learning, learning automata, and SARSA. Deep learning concepts may include deep belief networks, deep Boltzmann machines, deep convolutional neural networks, deep recurrent neural networks, and hierarchical temporal memories.
[0077] The methods described herein can include the use of a classifier to identify or distinguish between disease states, such as cancer (e.g., lung cancer or NSCLC). The classifier can distinguish between disease states and co-morbidities, such as chronic lung disorder, chronic obstructive pulmonary disease, emphysema, cardiovascular disease, hypertension, pulmonary fibrosis, or asthma.
[0078] A classifier can be generated by removing or filtering out biomolecules associated with an acute phase response. In some embodiments, the classifier is configured to remove acute phase response bias or stress protein bias. In some embodiments, the classifier comprises protein-related features. The features can be selected to remove acute phase response and / or stress protein bias in the biological sample.
[0079] The classifier may include features (e.g., biomarker information) for distinguishing between the disease states in Figure 11 and other states (e.g., healthy states or comorbidity states). Any of the features or biomarkers in Figure 11 may be used in methods for distinguishing between disease states and other states. The biomarker information may include information including the expression level or amount of the biomarker.
[0080] Classifier can comprise the feature for distinguishing NSCLC presence or absence.For example, feature can comprise information about one of many biomarkers, including SDC1, ANGL6, PXDN, ANTR1, CC085, SAA2, HTRA1, KPCB, KV401, CCL18, MYL6, ANTR2, GTPB2, HDGF, TBA1A, CSRP1, TCO2, CSPG2, PTPRZ, ILF2, SIAT1, ITA2B, DOK2, H31, H31T, H32, H33, H3C, RAC2, ARRB1, DHB4, HV102, RHG18, GDF15, PCSK6, FHOD1 or ITLN2, or any combination thereof.Any of these features or biomarkers can be included in the method for distinguishing NSCLC presence or absence.
[0081] The classifier may include features for distinguishing between healthy and early stage NSCLC (e.g., NSCLC stage 1, 2, and / or 3). Such features may include information about one of many biomarkers, including SDC1, ANGL6, PXDN, ANTR1, SAA2, HTRA1, CCL18, MYL6, ANTR2, TBA1A, TCO2, CSPG2, SIAT1, H31, H31T, H32, H33, H3C, or HV102, or any combination thereof. Any of these features or biomarkers can be included in a method for distinguishing between healthy and early stage NSCLC.
[0082] Classifier can comprise features for distinguishing between healthy state and late stage NSCLC (for example, NSCLC stage 4). Such features can comprise information about one of many biomarkers, including SDC1, ANGL6, PXDN, ANTR1, CC085, HTRA1, CCL18, MYL6, HDGF, TBA1A, ILF2, SIAT1, H31, H31T, H32, H33, H3C, GDF15 or PCSK6, or any combination thereof. Any of these features or biomarkers can be included in the method for distinguishing between healthy state and late stage NSCLC.
[0083] Classifier can comprise the features for distinguishing between healthy state and comorbidity.Such features can comprise information about one of many biomarkers, including SAA2, HTRA1, SYWC, RAB14, CSPG2, CTHR1, ITA6, FA8, ITA2B, DOK2, CILP1, CD9, CD36, INF2, CYFP1, ACTA or ACTH, or any combination thereof.Any of these features or biomarkers can be included in the method for distinguishing between healthy state and comorbidity.
[0084] Classifier can comprise the features for distinguishing between early stage NSCLC and late stage NSCLC.For example, feature can comprise information about one of many biomarkers, including SDC1, CC085, KV401, MYL6, JIP2, HV459, HV461, HV169, HNRPC, ROA1, STON2, LV301, KVD20, SAE1, PDE5A, RTN3, HV373, LV325, H2B1C, H2B1D, H2B1H, H2B1K, H2B1L, H2B1M, H2B1N, H2B2F, H2BFS or NMT1, or any combination thereof.Any of these features or biomarkers can be included in the method for distinguishing between early stage NSCLC and late stage NSCLC.
[0085] Classification index can comprise the feature for distinguishing early stage NSCLC from comorbidity.For example, feature can comprise information about one of many biomarkers, including ANGL6, ANTR1, CC085, SAA2, KPCB, GTPB2, HDGF, CSRP1, TCO2, PTPRZ, DOK2, RAC2, ARRB1 or DHB4, or any combination thereof.Any of these features or biomarkers can be included in the method for distinguishing early stage NSCLC from comorbidity.
[0086] Classification index can comprise the features for distinguishing between late stage NSCLC and comorbidity.For example, feature can comprise information about one of many biomarkers, including SDC1, ANGL6, PXDN, ANTR1, CC085, CCL18, HNRPC, HDGF, CSRP1, PTPRZ, ILF2, ITA2B, RHG18, FHOD1 or ITLN2, or any combination thereof.Any of these features or biomarkers can be included in the method for distinguishing between late stage NSCLC and comorbidity.
[0087] Disease Detection One or more of the biomarkers disclosed herein can be used in assays for detecting cancer in a sample from a subject. For example, in some embodiments, the biomarkers disclosed herein can be used to detect lung cancer in a sample from a subject. The lung cancer can be non-small cell lung cancer (NSCLC). The lung cancer can be adenosquamous carcinoma of the lung. The lung cancer can include lung nodules. The lung cancer can be or include metastatic lung cancer. The lung cancer can be large cell neuroendocrine carcinoma. The lung cancer can be salivary gland-type lung cancer. The lung cancer can be mesothelioma. In some cases, the present disclosure provides methods for identifying a lung cancer biomarker disclosed herein in a sample from a patient (e.g., by mass spectrometry or ELISA). In some cases, the present disclosure provides methods for obtaining a sample from a subject, incubating the sample with a particle panel disclosed herein, and performing targeted mass spectrometry analysis of biomolecular coronas formed on various particle types of the particle panel to assess the presence or absence of one or more of the biomarkers disclosed herein associated with NSCLC. The protein data obtained using the above methods can be further processed using the classifiers disclosed herein to classify samples as healthy, comorbid, or NSCLC.
[0088] The biomarkers disclosed herein can not only be used to detect the presence of lung cancer, but also to identify the type and stage of lung cancer in patients. Determining the stage, type, and grade of lung cancer is often beyond the scope of the present method, as little is known about the genetic and molecular factors that mediate lung cancer progression. Successful treatment is highly dependent on accurate lung cancer characterization, but current methods for determining information about the state of lung cancer in patients are often slow, invasive, expensive, and very time-consuming. There has long been an unmet need for a rapid, non-invasive method that can accurately diagnose the stage and type of lung cancer. The present compositions and methods overcome this deficiency by enabling the identification and characterization of lung cancer from small patient samples.
[0089] In many cases, the compositions or methods of the present disclosure can identify lung cancer from less than 100 mL, less than 50 mL, less than 30 mL, less than 25 mL, less than 20 mL, less than 15 mL, less than 10 mL, less than 8 mL, less than 6 mL, less than 5 mL, less than 3 mL, less than 2 mL, or less than 1 mL of blood (plasma) from a patient. Additionally, some compositions or methods of the present disclosure can determine the type of lung cancer from less than 100 mL, less than 50 mL, less than 30 mL, less than 25 mL, less than 20 mL, less than 15 mL, less than 10 mL, less than 8 mL, less than 6 mL, less than 5 mL, less than 3 mL, less than 2 mL, or less than 1 mL of blood (plasma) from a patient. The methods or compositions of the disclosure can also determine the stage of lung cancer from less than 100 mL, less than 50 mL, less than 30 mL, less than 25 mL, less than 20 mL, less than 15 mL, less than 10 mL, less than 8 mL, less than 6 mL, less than 5 mL, less than 3 mL, less than 2 mL, or less than 1 mL of blood (plasma) from a patient.
[0090] The disclosed methods may include monitoring the progression of cancer in a patient. Various methods can distinguish between healthy, early-stage, and late-stage cancers. The disclosed methods can also determine whether a patient is in complete or partial remission. Thus, the methods may involve analyzing samples collected from a patient at separate time points. Such methods can identify and subsequently track a patient's health or cancer progression without invasive or costly procedures. Because small, localized cancers often require biopsies or lengthy imaging sessions for detection, tracking early-stage cancers can be particularly challenging and time-consuming for patients. Conversely, the present disclosure provides various methods for tracking small, localized cancers using blood analysis alone. For example, patients with stage 0 or stage 1 lung cancer can undergo bimonthly plasma analysis consistent with the disclosed methods to monitor cancer metastasis or progression. The patient can be subjected to the diagnostic analysis of the present disclosure daily, twice weekly, weekly, biweekly, monthly, bimonthly, quarterly (every three months), twice yearly, annually, or biennially. The patient can be monitored periodically to track remission, early cancer status, late cancer status, or maintenance of a healthy or precancerous status. In some cases, lung cancer can be diagnosed using the particles and methods of the present disclosure up to one year, two years, three years, four years, five years, six years, seven years, eight years, nine years, ten years, fifteen years, twenty years, or twenty-five years prior to the onset of symptoms of the lung cancer.
[0091] In some cases, the total assay time from obtaining the sample, sample preparation, incubation of the sample with the particle panel, and LC-MS (e.g., targeted mass spectrometry) to identify the protein or proteins can be about 8 hours. In some embodiments, the total assay time from a single pooled sample, including sample preparation and LC-MS, can be about at least 1 hour, at least 2 hours, at least 3 hours, at least 4 hours, at least 5 hours, at least 6 hours, at least 7 hours, at least 8 hours, at least 9 hours, at least 10 hours, less than 20 hours, less than 19 hours, less than 18 hours, less than 17 hours, less than 16 hours, less than 15 hours, less than 14 hours, less than 13 hours, less than 12 hours, less than 11 hours, less than 10 hours, less than 9 hours, less than 8 hours, less than 7 hours, less than 6 hours, less than 5 hours, less than 4 hours, less than 3 hours, less than 2 hours, less than 1 hour, at least 5-10 minutes, at least 10-20 minutes, at least 20-30 minutes, or at least The incubation time may be 30 to 40 minutes, at least 40 to 50 minutes, at least 50 to 60 minutes, at least 1 to 1.5 hours, at least 1.5 to 2 hours, at least 2 to 2.5 hours, at least 2.5 to 3 hours, at least 3 to 3.5 hours, at least 3.5 to 4 hours, at least 4 to 4.5 hours, at least 4.5 to 5 hours, at least 5 to 5.5 hours, at least 5.5 to 6 hours, at least 6 to 6.5 hours, at least 6.5 to 7 hours, at least 7 to 7.5 hours, at least 7.5 to 8 hours, at least 8 to 8.5 hours, at least 8.5 to 9 hours, at least 9 to 9.5 hours, or at least 9.5 to 10 hours.
[0092] The disease state can be identified with a sensitivity or specificity of about 80% or greater. The disease state can be identified with a sensitivity or specificity of about 85% or greater. The disease state can be identified with a sensitivity or specificity of about 90% or greater. The disease state can be identified with a sensitivity or specificity of about 95% or greater.
[0093] In some embodiments, any of the classifiers disclosed herein can be constructed using any of the biomarkers disclosed herein to determine whether a sample from a subject has a disease state selected from a healthy state, a comorbidity state, NSCLC stage 1, NSCLC stage 2, NSCLC stage 3, NSCLC stage 4, or NSCLC stage 1, 2, or 3. In some embodiments, the classifier can distinguish samples with high sensitivity and specificity as healthy versus NSCLC stage 1, 2, or 3. In some embodiments, the classifier can distinguish samples with high sensitivity and specificity as comorbid versus NSCLC stage 1, 2, or 3.
[0094] The present disclosure provides several peptides that can be useful in diagnosing various cancers, including lung cancer. In some cases, the absence, presence, or abundance of a single peptide can indicate a particular cancer. However, in many cases, analysis of a collection of peptides as disclosed herein can provide significantly more accurate diagnoses. The methods of the present disclosure can not only identify cancer in a patient, but also identify the stage (e.g., stage I vs. stage II, stage I vs. stage III, early vs. late), degree of metastasis, and tissue or site of origin. Furthermore, the methods of the present disclosure can complement other forms of analysis. For example, immunohistological analysis of tissue biopsies can be combined with plasma proteomic analysis to improve the accuracy of cancer diagnosis. Alternatively, a single method of the present disclosure may be sufficient for accurate cancer diagnosis.
[0095] An advantage of many of the disclosed methods can be their low invasiveness and minimal patient participation. In many cases, the disclosed diagnostic peptides can be identified in blood (e.g., whole blood, granulocyte, buffy coat, or plasma) samples, and can provide insights equal to or greater than those of intensive tissue biopsies or lengthy and expensive imaging procedures.
[0096] The methods described herein may include the detection or discrimination of a disease state. The disease state may include cancer. The disease state may include lung cancer. The disease state may include non-small cell lung cancer (NSCLC). Lung cancer may include NSCLC. NSCLC may include early stage NSCLC (e.g., stage 1 NSCLC, stage 2 NSCLC, or stage 3 NSCLC). NSCLC may include late stage NSCLC (e.g., stage 4 NSCLC).
[0097] The methods described herein include identifying a subject as having a disease state, such as cancer, based on biomarker measurements. Methods for assessing cancer status are disclosed herein. The methods include measuring biomarkers in a biological sample. The sample can be from a subject suspected of having cancer. The measuring can include obtaining biomarker measurements. The method can include obtaining biomarker measurements. The biomarkers can include biomarkers described herein.
[0098] The methods described herein may include identifying a biological sample from a subject as indicating a healthy state, a cancerous state, or a comorbidity thereof in the subject based on biomarker measurements obtained in the subject. The cancer may be lung cancer, such as NSCLC. The method may include using a classifier such as a classifier described herein. The method may distinguish between a comorbidity and a cancerous state. The method may distinguish between a healthy state and a cancerous state. The comorbidity of the lung may include diseases other than cancer.
[0099] The methods described herein can identify or distinguish between comorbidities. The comorbidity can be a pulmonary comorbidity. The pulmonary comorbidity can include a pulmonary disease other than cancer. The pulmonary comorbidity can be selected from the group consisting of chronic obstructive pulmonary disease (COPD), emphysema, cardiovascular disease, hypertension, pulmonary fibrosis, asthma, chronic lung disease, and any combination thereof. The pulmonary comorbidity can include COPD. The pulmonary comorbidity can include emphysema. The pulmonary comorbidity can include cardiovascular disease. The pulmonary comorbidity can include hypertension. The pulmonary comorbidity can include pulmonary fibrosis. The pulmonary comorbidity can include asthma. The pulmonary comorbidity can include chronic lung disease.
[0100] Disclosed herein are methods for assaying one or more biomarkers in a sample from a subject suspected of having lung cancer. The method may include measuring one or more biomarkers in the sample. The measurement may include detecting the presence of one or more biomarkers. The measurement may include detecting the absence of one or more biomarkers. The measurement may include detecting the amount of one or more biomarkers. The biomarkers may include a biomarker selected from the group consisting of angiopoietin-related protein 6 (ANGL6), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), 60S acidic ribosomal protein P2 (RLA2), and platelet glycoprotein Ib beta chain (GP1BB), or a peptide fragment thereof.
[0101] Disclosed herein are methods for assaying one or more biomarkers in a sample from a subject suspected of having lung cancer, including non-small cell lung cancer (NSCLC). Measuring can include detecting the presence of one or more biomarkers. Measuring can include detecting the absence of one or more biomarkers. Measuring can include detecting the amount of one or more biomarkers. The biomarkers may include a biomarker selected from the group consisting of angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), or platelet glycoprotein Ib beta chain (GP1BB), or a peptide fragment thereof.
[0102] The method can include comparing the amount of the biomarker to a control. The control can include an index. The control can include a threshold. The control can include a control sample from a control subject. In some cases, the control sample includes a blood sample, a plasma sample, or a serum sample. In some cases, the control subject does not have lung cancer.
[0103] In some cases, the lung cancer comprises stage 1-4 NSCLC. In some cases, the subject has lung cancer. In some cases, the control subject has stage 1-4 NSCLC. In some cases, the subject's NSCLC comprises a different stage than the control subject's NSCLC.
[0104] The control subject may have a chronic lung disorder, chronic obstructive pulmonary disease, emphysema, cardiovascular disease, hypertension, pulmonary fibrosis, or asthma. The control subject may have a lung disorder. The control subject may have a chronic lung disorder. The control subject may have chronic obstructive pulmonary disease. The control subject may have emphysema. The control subject may have cardiovascular disease. The control subject may have hypertension. The control subject may have fibrosis. The control subject may have pulmonary fibrosis. The control subject may have asthma.
[0105] The method may include identifying a subject as having or not having lung cancer based on the measurement of one or more biomarkers. The method may include identifying the presence or absence of lung cancer cells or components thereof in a sample based on the measurement of one or more biomarkers. The presence of one or more biomarkers may indicate the presence of NSCLC cells or components thereof in the sample. The method may include determining the likelihood of the subject having lung cancer based on the measurement of one or more biomarkers. The method may include identifying a subject as having lung cancer based on the measurement of one or more biomarkers. The method may include identifying a stage of the cancer based on the measurement.
[0106] The method may include assaying a biological sample from a subject to identify a biomolecule. The method may include using a classifier to identify the sample as positive for non-small cell lung cancer (NSCLC) based on the identified biomolecule. The method may include using the classifier to identify the sample as negative for non-small cell lung cancer (NSCLC) based on the identified biomolecule. The classifier can be generated using data from a sample assayed using multiple particles with distinct physicochemical properties to obtain data. The classifier can be trained using data from samples, including known healthy samples and known NSCLC samples. The biomolecule may include a protein or biomarker described herein. The data may include proteomic data that identifies the presence or absence of a protein in the sample.
[0107] Detection Method The present disclosure provides various methods for detecting biomolecules (e.g., protein biomarkers) from biological samples. Several different analytical techniques can be used to identify, measure, and quantify biomolecular (e.g., proteomic) data from biological samples. For example, SDS-PAGE or any gel-based separation technique can be used to analyze proteomic data. Alternatively, mass spectrometry, high-performance liquid chromatography, LC-MS / MS, Edman degradation, immunoaffinity techniques, binding reagent analysis (e.g., immunostaining or aptamer binding assays), enzyme-linked immunosorbent assay (ELISA), chromatography, Western blot analysis, mass spectrometry, or any combination thereof can be used to identify, measure, and quantify proteomic data. Biomolecules can be enriched using particles or particle panels prior to analysis. Particles can be used to collect a subset of biomolecules from a biological sample, optionally eluted into solution, optionally treated (e.g., digested or chemically reduced), and then analyzed. Particle-based biomolecule collection can enrich biomolecules from biological samples, thereby enabling rapid detection and quantification of low-abundance biomolecules.
[0108] Various methods of the present disclosure for detecting biomolecules include binding reagent assays. A biological sample, or a collection of biomolecules from a biological sample, can be contacted with a target-specific binding reagent, such as an antibody, affibody, affimer, alphabody, avimer, DARPin, chimeric antigen receptor, T cell receptor, aptamer, or fragment thereof. The binding reagent can be detectable. The binding reagent can include a barcode sequence that allows detection and quantification of the binding reagent by nucleic acid sequencing analysis. The binding reagent can include an optically detectable label or moiety (e.g., a fluorescent protein, such as GFP or YFP, or a fluorescent dye). The binding reagent assay can include multiple binding reagents targeting multiple biomolecules and containing different detectable signals (e.g., nucleic acid barcode sequences or optically detectable moieties), thereby enabling multiplexed detection and quantification of selected biomarkers from a sample. For example, a sample can be contacted with multiple antibodies containing distinct detectable labels that target different proteins from those listed in Table 1. In some cases, the binding reagent can be contacted with a biomolecule that is covalently or non-covalently immobilized on a support (e.g., a membrane, surface, resin, or slide). In some cases, the binding reagent can be contacted with a biomolecule that is adsorbed to a particle (e.g., disposed within the biomolecular corona of the particle).
[0109] Various methods of the present disclosure for detecting biomolecules include ELISA. The method may include a sandwich ELISA assay in which a biomolecule (e.g., a peptide from among the peptides listed in Table 1) is contacted with a first antibody immobilized on a solid phase and a second antibody linked to a detectable moiety (e.g., an optically detectable moiety), where the first antibody comprises a first paratope of a first epitope on the biomolecule and the second antibody comprises a second paratope of a second epitope on the biomolecule. The ELISA assay may include immobilizing the biomolecule of interest on a support (e.g., a glass slide or the bottom of a well of a multiwell plate) and contacting the biomolecule with a first molecule that has binding affinity for the biomolecule. The first antibody may be linked to a detectable moiety or may be contacted with a second antibody, which is linked to a detectable moiety and binds to the first antibody. ELISA assays can have low detection limits (eg, >1 pg / ml) for target detection and quantification and therefore may be suitable for analysis of the cancer biomarkers disclosed herein.
[0110] The method of the present disclosure can include mass spectrometric analysis of biomolecules, such as proteins, peptides, or portions thereof.Mass spectrometric analysis can be performed in tandem with chromatographic separation techniques, such as liquid chromatography, so that biomolecules or biomolecule fragments are subjected to mass spectrometric analysis at different times.Mass spectrometric analysis can include two or more mass spectrometric steps (for example, tandem mass spectrometry), in which ions are fragmented and then subjected to further analysis.
[0111] The methods described herein may include measuring a biomarker (e.g., one or more biomarkers) in a sample from a subject. Measuring a biomarker may include performing an assay method. Measuring a biomarker may include performing mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blot, dot blot, or immunostaining, or a combination thereof. Measuring a biomarker may include performing mass spectrometry. Measuring a biomarker may include performing chromatography. Measuring a biomarker may include performing liquid chromatography. Measuring a biomarker may include performing high-performance liquid chromatography. Measuring a biomarker may include performing solid-phase chromatography. Measuring a biomarker may include performing a lateral flow assay. Measuring a biomarker may include performing an immunoassay. Measuring a biomarker may include performing an enzyme-linked immunosorbent assay. Measuring a biomarker may include performing a blot, such as a Western blot. Measuring a biomarker may include performing a dot blot. Measuring a biomarker may include performing immunostaining. Measuring biomarkers can include contacting a biological sample with multiple physiochemically distinct nanoparticles. Measuring biomarkers can include performing a combination of assay methods. For example, the methods described herein can include the use of particles followed by an immunoassay, such as an ELISA, to assess proteins or biomolecules in the biomolecule or protein corona. The methods described herein can include detecting proteins in the biomolecule corona by mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blot, dot blot, or immunostaining, or a combination thereof.The methods described herein can include detecting proteins of the biomolecular corona by mass spectrometry.
[0112] Measuring biomarkers can include using a detection reagent that binds to protein and produces a detectable signal.Methods described herein that can include detecting protein can include measuring a reading that indicates the presence, absence or amount of protein.Measuring biomarkers can include measuring a reading that indicates the presence, absence or amount of one or more biomarkers.
[0113] The method may include concentrating the biomarkers in the sample prior to measuring the biomarkers. Measuring the biomarkers may include concentrating the sample. Measuring the biomarkers may include filtering the sample. Measuring the biomarkers may include centrifuging the sample.
[0114] Measuring the biomarker may include contacting the sample with an assay reagent. The assay reagent may include a particle. The assay reagent may include an antibody. The assay reagent may include a biomolecule-binding molecule.
[0115] Particles and Types Disease detection methods can include the use of particles. The methods described herein can include contacting a biological sample with physiochemically distinct particles to form a biomolecular corona. The particles can adsorb biomolecules from the biological sample, thereby forming a biomolecular corona on the particle's surface. Upon contact with the biological sample, the particles can adsorb multiple peptides, proteins, nucleic acids, lipids, sugars, small molecules (e.g., metabolites (natural or foreign), terpenes, polyketides, and cyclic peptides), or any combination thereof. Thus, methods can include collecting a subset of biomolecules from a biological sample (e.g., a complex biological sample such as human plasma) onto particles and analyzing the biomolecules collected on the particles, analyzing biomolecules remaining in the biological sample, or analyzing the biomolecules collected on the particles and biomolecules remaining in the biological sample. The biomolecules, biomolecular corona, or portions thereof can be eluted from the particles into solution prior to analysis.
[0116] The relationship between particle properties and biomolecular corona composition can be exploited to manipulate biomolecule collection from samples. In some cases, a set of particle properties can favor the binding of specific biomolecule types, families, or superfamilies. For example, humans express over 100 proteins from the Ras superfamily, which share a conserved GTP-binding motif within a 20-kilodalton (kDa) N-terminal domain. Particles, or particle ensembles (e.g., mixtures containing five types of particles), can be functionalized to favor Ras protein adsorption and thus tailored to preferentially adsorb Ras proteins from complex biological samples, thereby enabling their enrichment for further analysis.
[0117] Particles, or mixtures of different particles, can be tailored for broad profiling of samples. In many biological samples, a small number of biomolecules constitute the majority of biological material. For example, only 20 of the approximately 3,500 human plasma proteins account for more than 99% of the protein mass in human plasma. Analysis of such samples can be extremely challenging, as a small number of highly abundant biomolecules can saturate detection or enrichment schemes. Particles, or collections of multiple particle types, can be tailored to broadly profile complex biological samples so that low-abundance biomolecules are preferentially enriched over or along with highly abundant biomolecules from complex biological samples. Particles, or collections of multiple particle types, can have similar binding affinities for multiple biomolecules, thus favoring the adsorption of multiple biomolecules from a sample. Particles can have low affinity for a highly abundant protein or set of highly abundant proteins in a sample, thereby preferentially adsorbing and enriching the low-abundance biomolecules. A particle collection can include particle types that have affinities for different types or classes of biomolecules, such that the particle collection can adsorb a wide range of biomolecules from a sample. Thus, the present disclosure provides a wide range of particle types with distinct physicochemical properties.
[0118] Particle types consistent with the methods disclosed herein can be made from a variety of materials. For example, particle materials consistent with the present disclosure include metals, polymers, magnetic materials, and lipids. Magnetic particles can be iron oxide particles. Examples of metallic materials include gold, silver, copper, nickel, cobalt, palladium, platinum, iridium, osmium, rhodium, ruthenium, rhenium, vanadium, chromium, manganese, niobium, molybdenum, tungsten, tantalum, iron, and cadmium, or any other materials described in U.S. Patent No. 7,749,299 (the contents of which are incorporated herein by reference in their entirety). Particles can be magnetic (e.g., ferromagnetic or ferrimagnetic). For example, particles can include superparamagnetic iron oxide nanoparticles (SPIONs).
[0119] The particles may be a plurality of physiochemically distinct particles (e.g., two or more sets of physiochemically distinct particles, in this case One set of particles may comprise particles that are physiochemically distinct from another set of particles. The physiochemically distinct particles may comprise lipid particles, metal particles, silica particles, or polymer particles. The physiochemically distinct particles may comprise carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles.
[0120] The particles may comprise a polymer, examples of which include polyethylene, polycarbonate, polyanhydrides, polyhydroxy acids, and polypropylfumerates. , polycaprolactone, polyamide, polyacetal, polyether, polyester, poly(orthoester), polycyanoacrylate, polyvinyl alcohol, polyurethane, polyphosphazene, polyacrylate, polymethacrylate, polycyanoacrylate, polyurea, polystyrene, or polyamine, polyalkylene glycol (e.g., polyethylene glycol (PEG)), polyester (e.g., poly(lactide-co-glycolide) (PLGA), polylactic acid, or polycaprolactone), or copolymers of two or more polymers, for example, copolymers of polyalkylene glycol (e.g., PEG) and polyester (e.g., PLGA). The polymer can be a lipid-terminated polyalkylene glycol, and polyester, or any other material disclosed in U.S. Pat. No. 9,549,901, the contents of which are incorporated herein by reference in their entirety.
[0121] The particles may comprise lipids. Examples of lipids that may be used to form the particles of the present disclosure include cationic lipids, anionic lipids, and lipids with a neutral charge. For example, the particles may be composed of dioleylphosphatidylglycerol (DOPG), diacylphosphatidylcholine, diacylphosphatidylethanolamine, ceramide, sphingomyelin, cephalin, cholesterol, cerebrosides and diacylglycerol, dioleylphosphatidylcholine (DOPC), dimyristoylphosphatidylcholine (DMPC), and dioleylphosphatidylserine (DOPS), phosphatidylglycerol, cardiolipin, diacylphosphatidylserine, diacylphosphatidic acid, N-dodecanoylphosphatidylethanolamine, N-succinylphosphatidylethanolamine, N-glutarylphosphatidylethanolamine, lysylphosphatidylglycerol, palmitoyloleoylphosphatidylglycerol (POPG), lecithin, lysolecithin, phosphatidylethanolamine, lysophosphatidyl Dioleoylethanolamine, dioleoylphosphatidylethanolamine (DOPE), dipalmitoylphosphatidylethanolamine (DPPE), dimyristoylphosphoethanolamine (DMPE), distearoyl-phosphatidyl-ethanolamine (DSPE), palmitoyloleoyl-phosphatidylethanolamine (POPE), palmitoyloleoylphosphatidylcholine (POPC), egg phosphatidylcholine (EPC), distearoylphosphatidylcholine (DSPC), dioleoylphosphatidylcholine (DOPC), dipalmitoylphosphatidylcholine (DPPC), dioleoylphosphatidylglycerol (DOPG), dipalmitoylphosphatidylglycerol (DPPG), palmitoyloleoylphosphatidylglycerol (POPG), 16-O-monomethyl PE, 16-O-dimethyl PE, 18-1-transIt may be made of any one or any combination of PE, palmitoyloleoyl-phosphatidylethanolamine (POPE), 1-stearoyl-2-oleoyl-phosphatidylethanolamine (SOPE), phosphatidylserine, phosphatidylinositol, sphingomyelin, cephalin, cardiolipin, phosphatidic acid, cerebrosides, diacetylphosphate, and cholesterol, or any other materials disclosed in U.S. Pat. No. 9,445,994, the entirety of which is incorporated herein by reference.
[0122] Examples of particles of the present disclosure are provided in Table 2. [Table 2]
[0123] Examples of particle types of the present disclosure include carboxylate (citrate) superparamagnetic iron oxide nanoparticles (SPIONs), phenol-formaldehyde coated SPIONs, silica coated SPIONs, polystyrene coated SPIONs, carboxylated poly(styrene-co-methacrylic acid) coated SPIONs, N-(3-trimethoxysilylpropyl)diethylenetriamine coated SPIONs, poly(N-(3-(dimethylamino)propyl)methacrylamide) (PDMAPMA) coated SPIONs, 1,2,4,5-benzenetetracarboxylic acid coated SPIONs, poly(vinylbenzyltrimethylammonium chloride) (PVBTMAC) coated SPIONs, carboxylate, PAA coated SPIONs, poly(oligo(ethylene glycol)-coated SPIONs), and poly(ethylene glycol)-coated SPIONs. The particles can be (ethylene glycol) methyl ether methacrylate (POEGMA) coated SPIONs, carboxylate microparticles, carboxyl-functionalized polystyrene particles, carboxylic acid coated particles, silica particles, carboxylic acid particles with a diameter of approximately 150 nm, amino-surfaced microparticles with a diameter of approximately 0.4-0.6 μm, amino-functionalized silica microparticles with a diameter of approximately 0.1-0.39 μm, Jeffamine-surfaced particles with a diameter of approximately 0.1-0.39 μm, polystyrene microparticles with a diameter of approximately 2.0-2.9 μm, silica particles, carboxylated particles with a proprietary coating with a diameter of approximately 50 nm, particles with a diameter of approximately 0.13 μm coated with a dextran-based coating, or silanol-coated silica particles with low acidity.
[0124] Particles consistent with the present disclosure can be produced and used in a method to form a protein corona of a wide range of sizes after incubation in a biological fluid. The particles of the present disclosure can be nanoparticles. The nanoparticles of the present disclosure can be from about 10 nm to about 1000 nm in diameter. For example, the nanoparticles disclosed herein can be at least 10 nm, at least 100 nm, at least 200 nm, at least 300 nm, at least 400 nm, at least 500 nm, at least 600 nm, at least 700 nm, at least 800 nm, at least 900 nm, 10 nm to 50 nm, 50 nm to 100 nm, 100 nm to 150 nm, 150 nm to 200 nm, 200 nm to 250 nm, 250 nm to 300 nm, 300 nm to 350 nm, 350 nm to 400 nm, 400 nm to 450 nm, 450 nm to 500 nm, 500 nm to 550 nm, The nanoparticles may be 550nm to 600nm, 600nm to 650nm, 650nm to 700nm, 700nm to 750nm, 750nm to 800nm, 800nm to 850nm, 850nm to 900nm, 100nm to 300nm, 150nm to 350nm, 200nm to 400nm, 250nm to 450nm, 300nm to 500nm, 350nm to 550nm, 400nm to 600nm, 450nm to 650nm, 500nm to 700nm, 550nm to 750nm, 600nm to 800nm, 650nm to 850nm, 700nm to 900nm, or 10nm to 900nm.
[0125] The particles of the present disclosure may be microparticles. Microparticles may be particles having a diameter of about 1 μm to about 1000 μm. For example, the nanoparticles disclosed herein may have a diameter of at least 1 μm, at least 10 μm, at least 100 μm, at least 200 μm, at least 300 μm, at least 400 μm, at least 500 μm, at least 600 μm, at least 700 μm, at least 800 μm, at least 900 μm, 10 μm to 50 μm, 50 μm to 100 μm, 100 μm to 150 μm, 150 μm to 200 μm, 200 μm to 250 μm, 250 μm to 300 μm, 300 μm to 350 μm, 350 μm to 400 μm, 400 μm to 450 μm, 450 μm to 500 μm, 500 μm to 550 μm, or 500 μm to 550 μm. 0μm, 550μm~600μm, 600μm~650μm, 650μm~700μm, 700μm~750μm, 750μm~800μm, 8 00μm~850μm, 850μm~900μm, 100μm~300μm, 150μm~350μm, 200μm~400μm, 250μm~4 The microparticles may be 50 μm, 300 μm to 500 μm, 350 μm to 550 μm, 400 μm to 600 μm, 450 μm to 650 μm, 500 μm to 700 μm, 550 μm to 750 μm, 600 μm to 800 μm, 650 μm to 850 μm, 700 μm to 900 μm, or 10 μm to 900 μm. The microparticles may be less than 1000 μm in diameter.
[0126] The surface area to mass ratio can be a determining factor in particle properties. For example, the number and type of bioparticles that a particle will adsorb from a solution can vary depending on the surface area to mass ratio of the particle. The particles disclosed herein have a surface area of 3-30 cm 2 / mg, 5-50cm 2 / mg, 10-60cm 2 / mg, 15-70cm 2 / mg, 20-80cm 2 / mg, 30-100cm 2 / mg, 35-120cm 2 / mg, 40-130cm 2 / mg, 45-150cm 2 / mg, 50-160cm 2 / mg, 60-180cm 2 / mg, 70-200cm 2 / mg, 80-220cm 2 / mg, 90-240cm 2 / mg, 100-270cm 2 / mg, 120-300cm 2 / mg, 200-500cm 2 / mg, 10-300cm 2 / mg, 1-3000cm 2 / mg, 20-150cm 2 / mg, 25-120cm 2 / mg, or 40-85cm 2 / mg. Small particles (e.g., having a diameter of 50 nm or less) can have significantly higher surface area to mass ratios, due in part to the higher order dependence of surface area on mass on diameter. In some cases (e.g., for small particles), particles may have a surface area to mass ratio of 200-1000 cm 2 / mg, 500-2000cm 2 / mg, 1000-4000cm 2 / mg, 2000-8000cm 2 / mg, or 4000-10000cm 2 In some cases (e.g., for large particles), the particles may have a surface area to mass ratio of 1-3 cm 2 / mg, 0.5-2cm 2 / mg, 0.25-1.5cm 2 / mg, or 0.1-1cm 2 / mg surface area to mass ratio.
[0127] In some cases, a plurality of particles (e.g., of a particle panel) used in the methods described herein can have a range of surface area to mass ratios. In some cases, the range of surface area to mass ratios for a plurality of particles can be greater than 100 cm. 2 / mg, 80cm 2 / mg, 60cm 2 / mg, 40cm 2 / mg, 20cm 2 / mg, 10cm 2 / mg, 5cm 2 / mg, or 2cm2 In some cases, the surface area to mass ratio for the plurality of particles varies from particle to particle by up to 40%, 30%, 20%, 10%, 5%, 3%, 2%, or 1% in most cases. In some cases, the plurality of particles can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or more different types of particles.
[0128] In some cases, the plurality of particles (e.g., in a particle panel) may have a wider range of surface area to mass ratios. In some cases, the range of surface area to mass ratios for the plurality of particles may be greater than 100 cm. 2 / mg, 150cm 2 / mg, 200cm 2 / mg, 250cm 2 / mg, 300cm 2 / mg, 400cm 2 / mg, 500cm 2 / mg, 800cm 2 / mg, 1000cm 2 / mg, 1200cm 2 / mg, 1500cm 2 / mg, 2000cm 2 / mg, 3000cm 2 / mg, 5000cm 2 / mg, 7500cm 2 / mg, 10000cm 2 / mg or greater. In some cases, the surface area to mass ratio for a plurality of particles (e.g., in a panel) can vary by more than 100%, 200%, 300%, 400%, 500%, 1000%, 10000% or more. In some cases, a plurality of particles having a wide range of surface area to mass ratios includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or more different types of particles.
[0129] The surface functional groups may include polymerizable functional groups, positively or negatively charged functional groups, zwitterionic functional groups, acidic or basic functional groups, polar functional groups, or any combination thereof. The surface functional groups may include carboxyl groups, hydroxyl groups, thiol groups, cyano groups, nitro groups, ammonium groups, alkyl groups, imidazolium groups, sulfonium groups, pyridinium groups, pyrrolidinium groups, phosphonium groups, aminopropyl groups, amine groups, boronic acid groups, N-succinimidyl ester groups, PEG groups, streptavidin, methyl ether groups, triethoxypropylaminosilane groups, PCP groups, citrate groups, lipoic acid groups, BPEI groups, or any combination thereof. The particles from the plurality of particles may be micelles, liposomes, iron oxide particles, silver particles, gold particles, palladium particles, quantum dots, platinum particles, titanium particles, silica particles, metal or inorganic oxide particles, synthetic polymer particles, copolymer particles, terpolymer particles, polymer particles with a metal core, polymer particles with a metal oxide core, polystyrene sulfonate particles, polyethylene oxide particles, polyoxyethylene glycol particles, polyethyleneimine particles, polylactic acid particles, polycaprolactone particles, polyglycolic acid particles, poly(lactide-co-glycolide) polymer particles, cellulose ether polymer particles, polyvinylpyrrolidone particles, polyvinyl acetate particles, polyvinylpyrrolidone-vinyl acetate copolymer particles, polyvinyl alcohol particles, acrylate particles, polyacrylic acid particles, crotonic acid copolymer particles, polyethylene phosphonate particles, polyalkylene particles, Carboxyvinyl polymer particles, sodium alginate particles, carrageenan particles, xanthan gum particles, acacia gum particles, gum arabic particles, guar gum particles, pullulan particles, agar particles, chitin particles, chitosan particles, pectin particles, karaya gum particles, locust bean gum particles, maltodextrin particles, amylose particles, corn starch particles, potato starch particles, rice starch particles, tapioca starch particles, pea starch particles, sweet potato starch particles, barley starch particles, wheat starch particles, hydroxypropylated high amylase starch particles, dextrin particles, levan particles, elsinan particles, gluten particles, collagen particles, whey protein isolate particles, casein particles, milk protein particles, soy protein particles, keratin particles, polyethylene particles, polycarbonate particles, polyanhydride particles, polyhydroxy acid particles, polypropyl fumarate e The polymer may be selected from the group consisting of poly(lactic acid-co-glycolic acid) particles, polyanhydride particles, bioreducible polymer particles, and 2-(3-aminopropylamido)ethanol particles, and any combination thereof.
[0130] The multiple particles (e.g., particles with distinct physicochemical characteristics) include carboxylate (citrate) superparamagnetic iron oxide nanoparticles (SPIONs), phenol-formaldehyde-coated SPIONs, silica-coated SPIONs, polystyrene-coated SPIONs, carboxylated poly(styrene-co-methacrylic acid)-coated SPIONs, N-(3-trimethoxysilylpropyl)diethylenetriamine-coated SPIONs, poly(N-(3-(dimethylamino)propyl)methacrylamide) (PDMAPMA)-coated SPIONs, 1,2,4,5-benzenetetracarboxylic acid-coated SPIONs, poly(vinylbenzyltrimethylammonium bromide)-coated SPIONs, and poly(vinylbenzyltrimethylammonium bromide). The particles may comprise one or more particle types selected from the group consisting of poly(trimethylsilyl methyl ether methacrylate) (PVB TMAC) coated SPIONs, carboxylates, PAA coated SPIONs, poly(oligo(ethylene glycol) methyl ether methacrylate) (POEGMA) coated SPIONs, carboxylate microparticles, carboxyl functionalized polystyrene particles, carboxylic acid coated particles, silica particles, carboxylic acid particles, amino surface particles, amino functionalized silica particles, Jeffamine surface particles, polystyrene particles, particles about 0.13 μm in diameter coated with a dextran based coating, or silanol coated silica particles.
[0131] The multiple particles (e.g., particles with distinct physicochemical characteristics) include carboxylate (citrate) superparamagnetic iron oxide nanoparticles (SPIONs), phenol-formaldehyde-coated SPIONs, silica-coated SPIONs, polystyrene-coated SPIONs, carboxylated poly(styrene-co-methacrylic acid)-coated SPIONs, N-(3-trimethoxysilylpropyl)diethylenetriamine-coated SPIONs, poly(N-(3-(dimethylamino)propyl)methacrylamide) (PDMAPMA)-coated SPIONs, 1,2,4,5-benzenetetracarboxylic acid-coated SPIONs, poly(vinylbenzyltrimethylammonium bromide)-coated SPIONs, and poly(vinylbenzyltrimethylammonium bromide). The particles may comprise one or more particle types selected from the group consisting of poly(trimethylsilyl methyl ether methacrylate) (PVB TMAC) coated SPIONs, carboxylates, PAA coated SPIONs, poly(oligo(ethylene glycol) methyl ether methacrylate) (POEGMA) coated SPIONs, carboxylate microparticles, carboxyl functionalized polystyrene particles, carboxylic acid coated particles, silica particles, carboxylic acid particles, amino surface particles, amino functionalized silica particles, Jeffamine surface particles, polystyrene particles, particles about 0.13 μm in diameter coated with a dextran based coating, or silanol coated silica particles.
[0132] The plurality of particles (e.g., physicochemically distinct particles) may comprise one or more particle types selected from the group consisting of silica particles, poly(acrylamide) particles, polyethylene glycol particles, or combinations thereof. One or more of the particles may comprise a paramagnetic or superparamagnetic core material. The particles may comprise silica particles. The particles may comprise poly(acrylamide) particles. The particles may comprise polyethylene glycol particles.
[0133] The plurality of particles can include multiple particle types. In some cases, the plurality of particles includes at least two types of particles. In some cases, the plurality of particles includes at least three types of particles. In some cases, the plurality of particles includes at least five types of particles. In some cases, the plurality of particles includes at least six types of particles. In some cases, the plurality of particles includes at least eight types of particles. In some cases, the plurality of particles includes at least ten types of particles. In some cases, the plurality of particles includes at least twelve types of particles. In some cases, the plurality of particles includes at least fifteen types of particles. In some cases, the plurality of particles includes at least eighteen types of particles. In some cases, the plurality of particles includes at least twenty types of particles.
[0134] A particle may include layers with distinct properties. A particle may include a core with a first set of properties and a shell with a second set of properties. A particle may include multiple shells with distinct properties (e.g., a core including a first material, an inner shell including a second material, and an outer shell including a third material). A layer of a particle may include multiple materials. For example, a layer of a particle may include multiple polymers. The polymers may be uniformly dispersed within the layer, phase-separated, or unevenly coated.
[0135] In some cases, the one or more physicochemical properties are selected from the group consisting of composition, size, surface charge, hydrophobicity, hydrophilicity, surface functional groups, surface topography, surface curvature, shape, and any combination thereof. In some embodiments, the surface functional groups further comprise chemical functionalization. In some embodiments, the small molecule functionalization is selected from the group consisting of amine functionalization, carboxylate functionalization, monosaccharide functionalization, oligosaccharide functionalization, phosphate sugar functionalization, sulfate functionalization, alcohol functionalization, ether functionalization, ester functionalization, amide functionalization, carbonate functionalization, carbamate functionalization, urea functionalization, benzyl functionalization, phenyl functionalization, phenol functionalization, aniline functionalization, imidazole functionalization, indole functionalization, fluoride functionalization, chloride functionalization, bromide functionalization, and the like. Functionalized particles include: functionalized with hydroxyl groups, sulfide groups, nitro groups, thiol groups, nitrogen-containing base groups, aminopropyl groups, boronic acid groups, N-succinimidyl ester groups, PEG groups, methyl ether groups, triethoxypropylaminosilane groups, silicon alkoxide groups, phenol-formaldehyde groups, organosilane groups, ethylene glycol groups, PCP groups, citrate groups, lipoic acid groups, or any combination thereof. In some embodiments, the small molecule functionalized particles include silica-functionalized particles, amine-functionalized particles, silicon alkoxide-functionalized particles, polystyrene-functionalized particles, and saccharide-functionalized particles. In some embodiments, the small molecule functionalized particles include amine-functionalized particles, phosphate-sugar groups, carboxylate groups, silica groups, organosilane groups, or any combination thereof. In some embodiments, the small molecule functionalization comprises silica functionalization, ethylene glycol functionalization, and amine functionalization, or any combination thereof.
[0136] The particles of the present disclosure can be synthesized, or the particles of the present disclosure can be purchased from commercial suppliers.For example, the particles consistent with the present disclosure can be purchased from commercial suppliers, including Sigma-Aldrich, Life Technologies, Fisher Biosciences, nanoComposix, Nanopartz, Spherotech and other commercial suppliers.The suitable particles of the present disclosure can be purchased from commercial suppliers, and then further modified, coated or functionalized.
[0137] The present disclosure includes compositions and methods comprising two or more particles from among those differing in at least one physicochemical property. Such compositions and methods may comprise at least two to at least 20 particles from among a plurality of particles differing in at least one physicochemical property. Such compositions and methods may comprise at least three to at least six particles from among a plurality of particles differing in at least one physicochemical property. Such compositions and methods may comprise at least four to at least eight particles from among a plurality of particles differing in at least one physicochemical property. Such compositions and methods may comprise at least four to at least 10 particles from among a plurality of particles differing in at least one physicochemical property. Such compositions and methods may comprise at least five to at least 12 particles from among a plurality of particles differing in at least one physicochemical property. Such compositions and methods may comprise at least six to at least 14 particles from among a plurality of particles differing in at least one physicochemical property. Such compositions and methods may comprise at least eight to at least 15 particles from among a plurality of particles differing in at least one physicochemical property. Such compositions and methods can include at least 10 to at least 20 particles from a plurality of particles that differ in at least one physicochemical property. Such compositions and methods can include at least two distinct particle types, at least three distinct particle types, at least four distinct particle types, at least five distinct particle types, at least six distinct particle types, at least seven distinct particle types, at least eight distinct particle types, at least nine distinct particle types, at least 10 distinct particle types, at least 11 distinct particle types, at least 12 distinct particle types, at least 13 distinct particle types, at least 14 distinct particle types, at least 15 distinct particle types, at least 20 distinct particle types, at least 25 particle types, or at least 30 distinct particle types.
[0138] The particles of the present disclosure can be contacted with a biological sample (e.g., a biological fluid) to form a biomolecular corona. When contacted with a complex biological sample, one or more types of particles of a plurality of particles can adsorb 100 or more types of proteins (e.g., approximately 10 of a given type in a 100 μl aliquot of biological sample containing 100 pM of particles of one type). 10 (A total of 100 or more types of particles may adsorb proteins.) The particles and biomolecular corona can be separated from the biological sample by, for example, centrifugation, magnetic separation, filtration, or gravity separation. The particle types and biomolecular corona can be separated from the biological sample using several separation techniques. Non-limiting examples of separation techniques include magnetic separation, column-based separation, filtration, spin column-based separation, centrifugation, ultracentrifugation, density or gradient-based centrifugation, gravity separation, or any combination thereof. Protein corona analysis can be performed on the separated particles and biomolecular corona. Protein corona analysis can include identifying one or more proteins in the biomolecular corona, for example, by mass spectrometry. The method can include contacting a single particle type (e.g., a particle type listed in Table 2) with the biological sample. The method can also include contacting multiple particle types (e.g., multiple particle types provided in Table 2) with the biological sample. Multiple particle types can be combined and contacted with a single sample volume of the biological sample. Multiple particle types can be contacted sequentially with a biological sample, with subsequent particle types being separated from the biological sample before contacting the biological sample. Protein corona analysis of biomolecular coronas can compress the dynamic range of analysis compared to whole protein analysis methods.
[0139] Contacting a particle or particles with a biological sample may include adding a defined concentration of particles to the biological sample. Contacting a particle or particles with a biological sample may include adding 1 pM to 100 nM particles to the biological sample. Contacting a particle or particles with a biological sample may include adding 1 pM to 500 pM particles to the biological sample. Contacting a particle or particles with a biological sample may include adding 10 pM to 1 nM particles to the biological sample. Contacting a particle or particles with a biological sample may include adding 100 pM to 10 nM particles to the biological sample. Contacting a particle or particles with a biological sample may include adding 500 pM to 100 nM particles to the biological sample. Contacting a particle or particles with a biological sample may include adding 50 μg / ml to 300 μg / ml (particle mass to biological sample volume) of particles to the biological sample. Contacting the biological sample with the particle or particles may include adding 100 μg / ml to 500 μg / ml of particles to the biological sample. Contacting the biological sample with the particle or particles may include adding 250 μg / ml to 750 μg / ml of particles to the biological sample. Contacting the biological sample with the particle or particles may include adding 400 μg / ml to 1 mg / ml of particles to the biological sample. Contacting the biological sample with the particle or particles may include adding 600 μg / ml to 1.5 mg / ml of particles to the biological sample. Contacting the biological sample with the particle or particles may include adding 800 μg / ml to 2 mg / ml of particles to the biological sample. Contacting the biological sample with the particle or particles may include adding 1 mg / ml to 3 mg / ml of particles to the biological sample. Contacting the particle or particles with the biological sample can include adding between 2 mg / ml and 5 mg / ml of particles to the biological sample. Contacting the particle or particles with the biological sample can include adding greater than 5 mg / ml of particles to the biological sample.
[0140] The particles in the plurality of particles may have varying degrees of size and shape uniformity. The standard deviation of diameter for a collection of particles of a particular type may be less than 20%, 10%, 5%, or 2% of the average diameter of that particle type (e.g., less than 2 nm for particles with an average diameter of 100 nm). This may correspond to a low polydispersity index for a sample containing a plurality of particles of less than 2, less than 1, less than 0.8, less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2, less than 0.1, or less than 0.05. Conversely, the plurality of particles may have a high degree of dispersion in average size and shape. The polydispersity index for a sample containing a plurality of particles may be greater than 3, greater than 4, greater than 5, greater than 8, greater than 10, greater than 12, greater than 15, or greater than 20. The size and shape uniformity among the plurality of particles may affect the number and type of biomolecules adsorbed to the particles. For some methods, particle-to-particle size uniformity (e.g., a low polydispersity index) allows for greater enrichment of specific biomolecules and a stronger correspondence between enriched biomolecule abundance and particle type. For some methods, low size uniformity allows for the collection of more types of biomolecules.
[0141] The method disclosed herein includes obtaining a dataset including proteins detected in biomolecular coronas corresponding to physiochemically distinct particles incubated with a biological sample. The biological sample may include a blood sample (e.g., an acellular sample) from which red blood cells have been removed. Physiochemically distinct types of particles give rise to different biomolecular coronas. Physiochemically distinct types of particles give rise to different biomarkers. Physiochemically distinct types of particles give rise to different mass spectral patterns.
[0142] Particles Panel The present disclosure provides compositions and methods for assaying samples for proteins. The compositions described herein include particle panels containing one or more distinct particle types. The particle panels described herein can vary in the number and diversity of particle types within a single panel. For example, particles within a panel can vary by size, polydispersity, shape and morphology, surface charge, surface chemistry and functionalization, and base material. The panel can be incubated with a sample and analyzed for proteins and protein concentrations. Proteins in the sample adsorb to the surfaces of different particle types within the particle panel to form protein coronas. The exact proteins and protein concentrations adsorbed to a particular particle type within the particle panel can depend on the composition, size, and surface charge of the particle type. Thus, each particle type within a panel can have a different protein corona due to adsorption to a different set of proteins, different concentrations of specific proteins, or a combination of these. Each particle type within a panel can have mutually exclusive or overlapping protein coronas. Overlapping protein coronas may overlap in protein identity, may overlap in protein concentration, or may overlap in both.
[0143] The present disclosure also provides methods for selecting particle types for inclusion in a panel depending on the sample type. The particle types included in the panel can be a combination of particles optimized for removal of high abundance proteins. Particle types also compatible for inclusion in a panel are those selected for adsorbing a specific protein of interest. The particles can be nanoparticles. The particles can be microparticles. The particles can be a combination of nanoparticles and microparticles.
[0144] The particle panels disclosed herein can be used to identify the number of distinct proteins disclosed herein and / or the number of any of the specific proteins disclosed herein across a wide dynamic range. For example, a particle panel disclosed herein comprising distinct particle types can enrich proteins in a sample across the entire dynamic range over which these proteins are present in the sample (e.g., a plasma sample). In some cases, a particle panel comprising any number of distinct particle types disclosed herein enriches proteins across a dynamic range of at least two orders of magnitude. In some cases, a particle panel comprising any number of distinct particle types disclosed herein enriches proteins across a dynamic range of at least three orders of magnitude. In some cases, a particle panel comprising any number of distinct particle types disclosed herein enriches proteins across a dynamic range of at least four orders of magnitude. In some cases, a particle panel comprising any number of distinct particle types disclosed herein enriches proteins across a dynamic range of at least five orders of magnitude. In some cases, a particle panel comprising any number of distinct particle types disclosed herein enriches proteins across a dynamic range of at least six orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of at least 7 orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of at least 8 orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of at least 9 orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of at least 10 orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of at least 11 orders of magnitude.In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of at least 12 orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of 3-5 orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of 3-6 orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of 4-8 orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of 5-8 orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of 6-10 orders of magnitude. In some cases, particle panels comprising any number of distinct particle types disclosed herein enrich proteins over a dynamic range of 8-12 orders of magnitude. For example, particle panels can collect mM and fM concentrations of proteins in a sample, thereby enriching proteins over a 12 order of magnitude range.
[0145] A particle panel comprising any number of distinct particle types disclosed herein enriches a single protein or a group of proteins. In some cases, the single protein or group of proteins may include proteins with different post-translational modifications. For example, a first particle type within the particle panel may enrich a protein or group of proteins with a first post-translational modification, a second particle type within the particle panel may enrich the same protein or group of proteins with a second post-translational modification, and a third particle type within the particle panel may enrich the same protein or group of proteins lacking the post-translational modification. In some cases, a particle panel comprising any number of distinct particle types disclosed herein enriches a single protein or group of proteins by binding to different domains, sequences, or epitopes of the single protein or group of proteins. For example, a first particle type within the particle panel may enrich a protein or group of proteins by binding to a first domain of the protein or group of proteins, and a second particle type within the particle panel may enrich the same protein or group of proteins by binding to a second domain of the protein or group of proteins.
[0146] The particle panel may include a combination of particles having silica and polymer surfaces. For example, the particle panel may include SPIONs coated with a thin layer of silica, SPIONs coated with poly(dimethylaminopropylmethacrylamide) (PDMAPMA), and SPIONs coated with poly(ethylene glycol) (PEG). A particle panel consistent with the present disclosure may include two or more particles selected from the group consisting of silica-coated SPIONs, N-(3-trimethoxysilylpropyl)diethylenetriamine-coated SPIONs, PDMAPMA-coated SPIONs, carboxyl-functionalized polyacrylic acid-coated SPIONs, amino-surface-functionalized SPIONs, carboxyl-functionalized polystyrene SPIONs, silica particles, and dextran-coated SPIONs. Particle panels consistent with the present disclosure include surfactant-free carboxylate microparticles, carboxyl-functionalized polystyrene particles, silica-coated particles, silica particles, dextran-coated particles, oleic acid-coated particles, boronated nanopowder-coated particles, PDMAPMA-coated particles, poly(glycidyl methacrylate-benzylamine)-coated particles, and poly(N-[3-(dimethylamino)propyl]methacrylamide-co-[2-(methacryloyloxy)ethyl]dimethyl-(3-sulfopropyl)ammonium hydroxide, P(DMAPMA-co-S BMA)-coated particles. Particle panels consistent with the present disclosure may include silica-coated particles, N-(3-trimethoxysilylpropyl)diethylenetriamine-coated particles, poly(N-(3-(dimethylamino)propyl)methacrylamide) (PDMAPMA)-coated particles, phosphate-sugar functionalized polystyrene particles, amine-functionalized polystyrene particles, carboxyl-functionalized polystyrene particles, ubiquitin-functionalized polystyrene particles, dextran-coated particles, or any combination thereof.
[0147] Using the particle panels disclosed herein, several proteins, peptides, or protein groups can be identified using the workflow described herein (collectively referred to as the "Proteograph" workflow, which involves MS analysis of distinct biomolecular coronas corresponding to distinct particle types within the particle panel). Feature intensities, as disclosed herein, are derived from the intensities of individual spikes ("features") found in a plot of mass-to-charge ratio versus intensity from a mass spectrometry run of a sample. These features may correspond to variably ionized fragments of peptides and / or proteins. Using the data analysis methods described herein, feature intensities can be sorted into protein groups. A protein group refers to two or more proteins identified by a shared peptide sequence. Alternatively, a protein group can refer to a single protein identified using a unique identifying sequence. For example, if a peptide sequence shared between two proteins (protein 1: XYZZX and protein 2: XYZYZ) in a sample is assayed, the protein group can be an "XYZ protein group" with two members (protein 1 and protein 2). Alternatively, if the peptide sequence is unique to a single protein (protein 1), the protein group may be a "ZZX" protein group with one member (protein 1). Each protein group may be supported by more than one peptide sequence. A protein detected or identified according to the present disclosure may refer to a distinct protein detected in a sample (e.g., distinct from other proteins detected using mass spectrometry). Thus, analysis of proteins present in distinct coronas corresponding to distinct particle types within a particle panel will yield a large number of feature intensities. This number decreases as the feature intensities are processed into distinct peptides, which in turn decrease as distinct peptides are processed into distinct proteins, which in turn decrease as peptides are grouped into protein groups (two or more proteins sharing distinct peptide sequences).
[0148] Disclosed herein particle panels for assessing the presence or absence of one or more biomarkers associated with lung cancer (e.g., NSCLC) include at least 1 distinct distinct particle type, at least 2 distinct distinct particle types, at least 3 distinct distinct particle types, at least 4 distinct distinct particle types, at least 5 distinct distinct particle types, at least 6 distinct distinct particle types, at least 7 distinct distinct particle types, at least 8 distinct distinct particle types, at least 9 distinct distinct particle types, at least 10 distinct distinct particle types, at least 11 distinct distinct particle types, at least 12 distinct distinct particle types, at least 13 distinct distinct particle types, at least 14 distinct distinct particle types, at least 15 distinct distinct particle types, at least 16 distinct distinct particle types, at least 17 distinct distinct particle types, at least 18 distinct particle types, at least 19 distinct distinct particle types, at least 20 distinct distinct particle types, at least 25 distinct particle types, at least at least 30 distinct distinct particle types, at least 35 distinct distinct particle types, at least 40 distinct distinct particle types, at least 45 distinct distinct particle types, at least 50 distinct distinct particle types, at least 55 distinct distinct particle types, at least 60 distinct distinct particle types, at least 65 distinct distinct particle types, at least 70 distinct distinct particle types, at least 75 distinct distinct particle types, at least 80 distinct distinct particle types, at least 85 distinct distinct particle types, at least 90 distinct distinct particle types distinct particle types, at least 95 distinctly distinct particle types, at least 100 distinctly distinct particle types, 1-5 distinctly distinct particle types, 5-10 distinctly distinct particle types, 10-15 distinctly distinct particle types, 15-20 distinctly distinct particle types, 20-25 distinctly distinct particle types, 25-30 distinctly distinct particle types, 30-35 distinctly distinct particle types, 35-40 distinctly distinct particle types, 40-45 distinctly distinct particle types, 45-50 distinctly distinct particle types, 50-55 distinctly distinct particle types,The panel may have 55-60 distinct particle types, 60-65 distinct particle types, 65-70 distinct particle types, 70-75 distinct particle types, 75-80 distinct particle types, 80-85 distinct particle types, 85-90 distinct particle types, 90-95 distinct particle types, 95-100 distinct particle types, 1-100 distinct particle types, 20-40 distinct particle types, 5-10 distinct particle types, 3-7 distinct particle types, 2-10 distinct particle types, 6-15 distinct particle types, or 10-20 distinct particle types. In certain embodiments, the present disclosure provides a panel size of 3-10 particle types. In certain embodiments, the present disclosure provides a panel size of 4-11 distinct particle types. In certain embodiments, the present disclosure provides a panel size of 5-15 distinct particle types. In certain embodiments, the present disclosure provides panel sizes of 5 to 15 distinct particle types. In certain embodiments, the present disclosure provides panel sizes of 8 to 12 distinct particle types. In certain embodiments, the present disclosure provides panel sizes of 9 to 13 distinct particle types. In certain embodiments, the present disclosure provides panel sizes of 10 distinct particle types. The particle types may include nanoparticle types.
[0149] Particle panels can be designed to broadly profile proteomes, such as the human plasma proteome. A major challenge in analyzing the human proteome is that only 20 proteins account for more than 99% of the mass of the approximately 3,500 proteins in human plasma. Plasma analysis methods are often saturated with these 20 proteins, providing minimal profiling depth for the remaining proteins. The particle panels of the present disclosure can include particle combinations that facilitate the collection of at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 1,100, at least 1,200, at least 1,300, at least 1,400, at least 1,500, at least 1,600, at least 1,700, at least 1,800, at least 1,900, at least 2,000, at least 2,100, or at least 2,200 distinct proteins from a single biological sample. A particle panel of the present disclosure may include a combination of particles that facilitates collection of at least 4%, at least 5%, at least 6%, at least 8%, at least 10%, at least 12%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70% of a protein type from a complex biological sample, such as human plasma. This can be achieved by providing (e.g., as a particle panel) multiple particles with distinct protein binding profiles. A particle panel may include two particles that, upon contact with a biological sample, jointly form a protein corona comprising less than 80%, 70%, 60%, 50%, 40%, 30%, 25%, 20%, 15%, or 10% of the proteins. In some cases, the biological sample is human plasma.
[0150] By increasing the number of particle types in the panel, the number of proteins that can be identified in a given sample can be increased. An example of how increasing the panel size can increase the number of identified proteins is shown in Figure 12, where a panel size of 1 particle type identified 419 different proteins, a panel size of 2 particle types identified 588 different proteins, a panel size of 3 particle types identified 727 different proteins, a panel size of 4 particle types identified 844 proteins, a panel size of 5 particle types identified 934 different proteins, a panel size of 6 particle types identified 1008 different proteins, a panel size of 7 particle types identified 1075 different proteins, a panel size of 8 particle types identified 1133 different proteins, a panel size of 9 particle types identified 1184 different proteins, a panel size of 10 particle types identified 1230 different proteins, a panel size of 11 particle types identified 1275 different proteins, and a panel size of 12 particle types identified 1318 different proteins.
[0151] Biomarker analysis in biological samples The compositions disclosed herein and methods for their use can identify numerous unique proteins in biological samples (e.g., biological fluids). Non-limiting examples of biological samples that can be analyzed using the methods described herein (e.g., protein corona analysis) include biological fluid samples (e.g., cerebrospinal fluid (CSF), synovial fluid (SF), urine, plasma, serum, tears, semen, whole blood, milk, nipple aspirate, ductal lavage, vaginal fluid, nasal secretions, ear fluid, gastric juice, pancreatic juice, trabecular fluid, lung lavage, prostatic fluid, sputum, fecal material, bronchial washings, swabs, bronchial aspirates, sweat, or saliva), fluid solids (e.g., tissue homogenates), or samples derived from cell culture. For example, the particles disclosed herein can be incubated with any of the biological samples disclosed herein to generate at least 100 unique proteins, at least 120 unique proteins, at least 140 unique proteins, at least 160 unique proteins, at least 180 unique proteins, at least 200 unique proteins, at least 220 unique proteins, at least 240 unique proteins, at least 260 unique proteins, at least 280 unique proteins, at least 300 unique proteins, at least 320 unique proteins, at least 340 unique proteins, at least 360 unique proteins, at least 380 unique proteins, at least 400 unique proteins, at least 420 unique proteins, at least 440 unique proteins, at least 460 unique proteins, at least 480 unique proteins, at least 500 unique proteins, at least 520 unique proteins, at least 540 unique proteins, at least 560 unique proteins, at least 580 unique proteins, at least 590 unique proteins, at least 600 unique proteins, at least 610 unique proteins, at least 620 unique proteins, at least 630 unique proteins, at least 640 unique proteins, at least 650 unique proteins, at least 660 unique proteins, at least 670 unique proteins, at least 680 unique proteins, at least 690 unique proteins, at least 700 unique proteins, at least 710 unique proteins, at least 720 unique proteins, at least 730 unique proteins, at least 740 unique proteins, at least 750 unique proteins, at least 760 unique proteins, at least 770 unique proteins, at least 780 unique proteins, at least 790 unique proteins, at least 800 unique proteins, at least 810 unique proteins, at least 820 unique proteins, at least 830 unique proteins, at least 840 unique proteins, at least 850 unique proteins, at least 860 unique proteins, at least 870 unique proteins, at least 88 unique proteins, at least 440 unique proteins, at least 460 unique proteins, at least 480 unique proteins, at least 500 unique proteins, at least 520 unique proteins, at least 540 unique proteins, at least 560 unique proteins, at least 580 unique proteins, at least 600 unique proteins, at least 620 unique proteins, at least 640 unique proteins, at least 660 unique proteins, at least 680 unique proteins, at least 700 unique proteins, at least 720 unique proteins, at least 740 unique proteins, at least 760 unique proteins, at least 780 unique proteins, at least 800 unique proteins,at least 820 unique proteins, at least 840 unique proteins, at least 860 unique proteins, at least 880 unique proteins, at least 900 unique proteins, at least 920 unique proteins, at least 940 unique proteins, at least 960 unique proteins, at least 980 unique proteins, at least 1000 unique proteins, 100-1000 unique proteins, 150-950 unique proteins, 200-900 unique proteins, 250-850 unique proteins, 300-800 unique proteins, 350-750 unique proteins, 400-700 unique proteins, 450-650 unique proteins Protein coronas can be formed that include proteins, 500-600 unique proteins, 200-250 unique proteins, 250-300 unique proteins, 300-350 unique proteins, 350-400 unique proteins, 400-450 unique proteins, 450-500 unique proteins, 500-550 unique proteins, 550-600 unique proteins, 600-650 unique proteins, 650-700 unique proteins, 700-750 unique proteins, 750-800 unique proteins, 800-850 unique proteins, 850-900 unique proteins, 900-950 unique proteins, and 950-1000 unique proteins. Similar numbers of proteins can be assessed, in some cases, without the use of particles or using the assay methods described herein. In some embodiments, several different types of particles can be used separately or in combination to identify multiple proteins in a particular biological sample. In other words, particles can be multiplexed to bind to and identify multiple proteins in a biological sample.
[0152] The compositions and methods disclosed herein can be used to identify various biological states in a particular biological sample. For example, a biological state can refer to elevated or decreased levels of a particular protein or set of proteins, or can be evidenced by the ratio between the abundances of two or more biomolecules. In another example, a biological state can refer to the identification of a disease, such as cancer. One or more particle types can be incubated with a biological sample, such as human plasma, to allow the formation of a protein corona. The protein corona can then be analyzed to identify a protein pattern. Analysis can include gel electrophoresis, mass spectrometry, chromatography, ELISA, immunohistochemistry, or any combination thereof. Analysis of the protein corona (e.g., by mass spectrometry or gel electrophoresis) is sometimes referred to as corona analysis. The protein pattern can be compared with the same method performed on a control sample. By comparing the protein patterns, it can be identified that a first sample has elevated levels of a marker corresponding to a particular type of lung cancer. Thus, the particles and their methods of use can be used to diagnose specific disease states.
[0153] Assays can include protein collection of particles, protein digestion, and mass spectrometry (e.g., MS, LC-MS, LC-MS / MS). Digestion can include chemical digestion, such as digestion with cyanogen bromide or 2-nitro-5-thiocyanatobenzoic acid (NTCB). Digestion can include enzymatic digestion, such as digestion with trypsin or pepsin. Digestion can include enzymatic digestion with multiple proteases. Digestion can include trypsin, chymotrypsin, Glu C, Lys C, esterase, subtilisin, proteinase K, thrombin, factor X, Arg C, papain, Asp N, thermolysin, pepsin, aspartyl protease, cathepsin D, zinc metalloprotease (ZMP), and the like. mealloprotease), glycoprotein endopeptidase, proline, aminopeptidase The digestion method may include a protease selected from the group consisting of a prenyl protease, a caspase, a kex2 endoprotease, or any combination thereof. The digestion method may randomly cleave peptides or may cleave peptides at a specific position or set of positions. The assay may utilize multiple digestion methods (e.g., two or more proteases). The assay may include dividing a sample into multiple portions and subjecting the portions to different digestion methods and separate analyses (e.g., separate mass spectrometry analyses). The digestion may cleave peptides at specific positions (e.g., at methionine) or sequences (e.g., glutamate-histidine-glutamate). The digestion may allow for the differentiation of similar proteins. For example, an assay may separate eight distinct proteins into a single protein group in a first digestion method and eight separate proteins with distinct signals in a second digestion method. The digestion may produce an average peptide fragment length of 8 to 15 amino acids. The digestion may produce an average peptide fragment length of 12 to 18 amino acids. The digestion may produce an average peptide fragment length of 15 to 25 amino acids. The digestion may produce an average peptide fragment length of 20 to 30 amino acids. The digestion may produce an average peptide fragment length of 30 to 50 amino acids.
[0154] The various methods disclosed herein enable measurements over a wide concentration range. Biomolecular analysis methods are often limited to a narrow concentration range. For example, mass spectroscopic proteomic analysis is often limited to a concentration range of three, four, or five orders of magnitude. As a result, the presence of relatively high concentrations of biomolecules (e.g., present at mg / ml concentrations) can mask the detection of biomolecules at lower concentrations and further limit the accuracy of quantification of low-concentration biomolecules. The methods disclosed herein may enable the detection of molecules with concentrations spanning at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, or at least twelve orders of magnitude. As a result, the methods disclosed herein can detect and quantify both relatively high and relatively low concentrations of biomolecules from a single sample without first depleting the biomolecules from the sample. For example, a plasma assay consistent with the present disclosure can simultaneously quantify albumin (present at approximately 40 mg / ml) and interleukin-10 (present at approximately 6 pg / ml) from a single, non-depleted plasma sample, thereby simultaneously detecting two species that differ in concentration by approximately 10 orders of magnitude.
[0155] Dynamic Range Some methods described herein (e.g., biomolecular corona analysis) may involve assaying biomolecules in a sample of the present disclosure over a wide dynamic range. The dynamic range of a biomolecule assayed in a sample may be the range of biomolecule abundance as measured by an assay method (e.g., mass spectrometry, chromatography, gel electrophoresis, spectrometry, or immunoassay) for the biomolecule contained in the sample. For example, an assay capable of detecting proteins over a wide dynamic range may be capable of detecting proteins of very low abundance to very high abundance. The dynamic range of an assay can be directly related to the slope of the assay signal intensity as a function of biomolecule abundance. For example, an assay with a small dynamic range may have a small (but positive) slope of the assay signal intensity as a function of biomolecule abundance; e.g., the ratio of the signal detected for a high-abundance biomolecule to the signal detected for a low-abundance biomolecule may be smaller for an assay with a small dynamic range than for an assay with a large dynamic range. In certain cases, dynamic range may refer to the dynamic range of proteins within a sample or assay method.
[0156] The methods described herein can compress the dynamic range of an assay. The dynamic range of an assay can be compressed compared to another assay if the slope of the assay signal intensity as a function of biomolecule abundance is smaller than that of the other assay. For example, a plasma sample assayed using protein corona analysis with mass spectrometry can have a compressed dynamic range compared to a plasma sample assayed directly using mass spectrometry alone, or compared to the abundance values provided for plasma proteins in a database (e.g., the database provided in Keshishiane et al., Mol. Cell Proteomics 14, 2375-2393 (2015)), also referred to herein as the "Carr database"). A compressed dynamic range can enable the detection of less abundant biomolecules in biological samples by using biomolecular corona analysis with mass spectrometry compared to the use of mass spectrometry alone.
[0157] The dynamic range of the analysis can be compressed by collecting biomolecules on particles prior to analysis (e.g., mass spectrometry or ELISA analysis). 6 Two proteins present in a ratio of 1:1 were differentially adsorbed onto the particles, and their new ratio was 10 4 The adsorption can be performed in a solution such that the adsorption ratio is 1:1. Such differential adsorption can enable the simultaneous detection of two biomolecules with a concentration difference greater than the dynamic range of the analytical technique. For example, mass spectrometry is often limited to measuring species within a concentration range of 4-6 orders of magnitude, resulting in a 10 8It may be impossible to simultaneously detect two biomolecules present at a 2-fold concentration difference. Biomolecular corona-based enrichment of a sample can enable the simultaneous detection of two biomolecules in a single analytical method by concentrating a biomolecule (e.g., a first protein) that is dilute relative to a second biomolecule (e.g., a second protein). Similarly, particle-based enrichment can enable the quantification of low-concentration biomolecules in a sample. The dynamic range over which an analyte can be quantified is often narrower than the dynamic range over which the analyte can be detected. For example, ELISAs often cover a dynamic range spanning two to three orders of magnitude, but provide accurate concentration quantification for less than two orders of magnitude. By increasing the number of biomolecular targets within the desired concentration range, particle-based enrichment can enable the simultaneous quantification of two or more biomolecules present in a biological sample at concentrations outside the dynamic range of the analytical technique.
[0158] Thus, various methods of the present disclosure involve detecting two biomolecules present in a biological sample with a concentration difference greater than the dynamic range of the detection method. Many of the biomarker pairs disclosed herein span a concentration range that exceeds the detection limit of a biomolecular analysis technique (e.g., immunostaining or LC-MS / MS) and therefore may be unidentifiable or quantifiable without the enrichment-based methods of the present disclosure. In some cases, the disclosed methods involve detecting two biomolecules (e.g., two proteins) at concentrations that differ by at least three orders of magnitude (e.g., 1 mg / ml and 1 μg / ml, or 50 μM and 50 nM) in a biological sample. In some cases, the disclosed methods involve detecting two biomolecules (e.g., two proteins) at concentrations that differ by at least four orders of magnitude (e.g., 1 mg / ml and 100 ng / ml, or 50 μM and 5 nM) in a biological sample. In some cases, the disclosed methods involve detecting two biomolecules (e.g., two proteins) at concentrations that differ by at least five orders of magnitude (e.g., detecting HBA and NOTUM in human plasma). In some cases, the disclosed methods involve detecting two biomolecules (e.g., two proteins) at concentrations that differ by at least five orders of magnitude in a biological sample (e.g., detecting ITIH2 and ANGL6 in human plasma). In some cases, the disclosed methods involve detecting two biomolecules (e.g., two proteins) at concentrations that differ by at least six orders of magnitude in a biological sample (e.g., detecting HBA and NOTUM in human plasma). In some cases, the disclosed methods involve detecting two biomolecules (e.g., two proteins) at concentrations that differ by at least seven orders of magnitude in a biological sample (e.g., detecting ceruloplasmin and RLA2 in human plasma). In some cases, the disclosed methods involve detecting two biomolecules (e.g., two proteins) at concentrations that differ by at least seven orders of magnitude in a biological sample (e.g., detecting human serum albumin and CAN2 in human plasma). In some cases, the disclosed methods involve detecting two biomolecules (e.g., two proteins) at concentrations that differ by at least seven orders of magnitude in a biological sample (e.g., detecting human serum albumin and interleukin-6 in human plasma).
[0159] The dynamic range of a proteomic analysis assay can be the ratio of the signal generated by the most abundant protein (e.g., the top 10% of proteins by abundance) to the signal generated by the least abundant protein (e.g., the bottom 10% of proteins by abundance). Compressing the dynamic range of a proteomic analysis can include decreasing the ratio of the signal generated by the most abundant protein to the signal generated by the least abundant protein for a first proteomic analysis assay compared to that for a second proteomic analysis assay. The protein corona analysis assays disclosed herein can compress the dynamic range compared to the dynamic range of all protein analysis methods (e.g., mass spectrometry, gel electrophoresis, or liquid chromatography).
[0160] Several methods are provided herein for compressing the dynamic range of biomolecular analytical assays to facilitate the detection of low-abundance biomolecules relative to high-abundance biomolecules. For example, the particle types disclosed herein can be used to sequentially interrogate samples. Upon incubation of the particle type in the sample, the particle type forms a biomolecular corona on the surface of the particle type. If the biomolecules are detected directly in the sample without the particle type, for example by direct mass spectrometric analysis of the sample, the dynamic range can span a wider concentration range or orders of magnitude greater than when the biomolecules are oriented on the surface of the particle type. Thus, the particle types disclosed herein can be used to compress the dynamic range of biomolecules in a sample. Without being limited by theory, this effect may be observed due to greater capture of higher-affinity, lower-abundance biomolecules in the biomolecular corona of the particle type and less capture of lower-affinity, higher-abundance biomolecules in the biomolecular corona of the particle type.
[0161] The dynamic range of a proteomic analysis assay can be the slope of a plot of the protein signal measured by the proteomic analysis assay as a function of the total abundance of proteins in a sample. Compressing the dynamic range can include reducing the slope of a plot of the protein signal measured by the proteomic analysis assay as a function of the total abundance of proteins in a sample compared to the slope of a plot of the protein signal measured by a second proteomic analysis assay as a function of the total abundance of proteins in a sample. The protein corona analysis assay disclosed herein can compress the dynamic range compared to the dynamic range of all protein analysis methods (e.g., mass spectrometry, gel electrophoresis, or liquid chromatography).
[0162] Samples and subjects The methods described herein may include the use of a sample, such as a biological sample. For example, the method may include determining one or more biomarker measurements in the sample. The biological sample may be from a subject. The biological sample may include a blood sample from which red blood cells have been removed. For example, the biological sample may include a plasma sample. The biological sample may include a serum sample. The biological sample may include blood or blood components.
[0163] Samples consistent with the methods disclosed herein for assessing the presence or absence of one or more biomarkers associated with a disease state, such as lung cancer (e.g., NSCLC), include biological samples from subjects. The subjects can be human or non-human animals. The biological sample can be a biological fluid. For example, the biological fluid can be plasma, serum, CSF, urine, tears, cell lysate, tissue lysate, cell homogenate, tissue homogenate, nipple aspirate, fecal sample, synovial fluid, whole blood, or saliva. The sample can also be a non-biological sample, such as water, milk, solvent, or anything homogenized into a fluid state. The biological sample can contain multiple proteins or proteome data, which can be analyzed after adsorption of proteins to the surfaces of various particle types in a panel and subsequent digestion of the protein corona. Proteome data can include nucleic acids, peptides, or proteins. Any of the samples herein can contain several different analytes, which can be analyzed using the compositions and methods disclosed herein. The analyte can be a protein, peptide, small molecule, nucleic acid, metabolite, lipid, or any molecule that may have the potential to bind to or interact with a particle-type surface.
[0164] Biological samples can include biological fluid samples, such as cerebrospinal fluid (CSF), synovial fluid (SF), urine, plasma, serum, tears, interstitial fluid, semen, whole blood, milk, nipple aspirate, ductal lavage, vaginal fluid, nasal discharge, ear fluid, gastric juice, pancreatic juice, trabecular fluid, pulmonary lavage, prostatic fluid, sputum, fecal material, bronchial lavage, swabbing fluid, bronchial aspirate, sweat, or saliva. Biological fluids can be flowing solids, such as tissue homogenates, or fluids extracted from biological samples. Biological samples can be, for example, tissue samples or fine needle aspirate (FNA) samples. Biological samples can be cell culture samples. For example, samples that can be used in the methods disclosed herein can include cells grown in cell culture or acellular material harvested from cell culture. Biological fluids can be flowing biological samples. For example, biological fluids can be flowing cell culture extracts. Samples can be extracted from fluid samples, or samples can be extracted from solid samples. For example, the sample may include gas molecules (eg, volatile organic compounds) extracted from a fluidized solid.
[0165] Methods consistent with the present disclosure can include collecting (e.g., isolating, enriching, or purifying) species from a biological sample. The species can be a biomolecule (e.g., a protein), a biopolymer structure (e.g., a peptide aggregate or a ribosome), a cell, or a tissue. Species can be selectively collected from a biological sample. For example, a method can include isolating cancer cells from a tissue (e.g., as a tissue biopsy) or from a biological fluid such as whole blood, plasma, or buffy coat (e.g., as a liquid biopsy). The species can be treated prior to analysis. For example, proteins can be reduced and degraded, nucleic acids can be separated from histones, or cells can be lysed.
[0166] The method may include collecting tissue or cells from a biological sample. The tissue or cells can be collected from a tissue or liquid biological sample. The tissue or cells can be taken directly from a patient. The tissue or cells can be collected from tissue suspected of being cancerous or precancerous. In some cases, the tissue or cells are selected from a biological sample isolated from a patient. The method may include identifying a cell or tissue subsection of interest from the biological sample. For example, the method may include isolating lung tissue by transthoracic lung biopsy, identifying potentially cancerous cells by immunohistological staining, and isolating the potentially cancerous cells for further analysis.
[0167] The method may involve the parallel analysis of two or more species. The species can be compared to determine the disease state of the sample (e.g., the type and stage of the disease). The species may be derived from a single subject (e.g., a single patient suspected of early-stage non-small cell lung cancer) or from different subjects (e.g., a healthy patient and a lung cancer patient). The species may be a healthy species, a diseased species, or a potentially diseased species. The species may be collected from the same biological sample, for example, from a single tissue section, or may be collected from different biological samples, for example, from separate blood and tissue samples.
[0168] Parallel analysis of two or more species can improve the accuracy of diagnosis. In some cases, multi-species analysis includes a known healthy species and a species suspected or known to be diseased (e.g., cells from healthy tissue and cells from cancerous tissue). Analysis of the healthy species and the diseased species can identify the stage of disease in the diseased species. In some cases, the first species may be suspected of having a disease, and the second species (e.g., a portion of a plasma sample) may contain a potential biomarker for the disease. In certain cases, the first species may be suspected of having a disease, and the second species may include blood or a portion of a blood sample (e.g., plasma or buffy coat). For example, squamous cells can be identified as cancerous by DNA sequencing and then identified as early-stage cancer cells based on the patient's plasma proteome profile.
[0169] Compositions and methods for multi-ome analysis are disclosed herein. "Multi-ome (multi-omics)" or "multi-ome (multi-omics)" can refer to analytical techniques for analyzing large-scale biomolecules in which the data set is multiple omes, such as the proteome, genome, transcriptome, lipidome, and metabolome. Non-limiting examples of multi-ome data include proteome data, genome data, lipidome data, glycomic data, transcriptome data, or metabolomic data. The "biomolecule" in the "biomolecular corona" can refer to any molecule or biological component that can be produced by or present in a biological organism. Non-limiting examples of biomolecules include proteins (protein coronas), polypeptides, polysaccharides, sugars, lipids, lipoproteins, metabolites, oligonucleotides, nucleic acids (DNA, RNA, microRNA, plasmids, single-stranded nucleic acids, double-stranded nucleic acids), metabolomes, and small molecules, such as primary metabolites, secondary metabolites, and other natural products, or any combination thereof. In some embodiments, the biomolecule is selected from the group of proteins, nucleic acids, lipids, and metabolites.
[0170] In some cases, samples can be depleted prior to biomarker analysis. Commercially available kits can be used to deplete samples. For example, kits that can be used to deplete samples can be spin-column-based depletion kits, albumin depletion kits, immune depletion kits, or kits that deplete high-abundance proteins. Non-limiting examples of kits that can be used for sample depletion include PureProteome™ Human Examples of suitable depletion kits include Albumin / Immunoglobulin depletion kit (EMD Millipore Sigma), ProteoPrep® Immunoaffinity Albumin & IgG Depletion Kit (Millipore Sigma), Seppro® Protein Depletion kit (Millipore Sigma), Top 12 Abundant Protein Depletion Spin Columns (Pierce), or Proteome Purify™ Immunopletion Kit (R&D Systems). Depletion can remove highly abundant biomolecules from a sample. For example, a method can include removing albumin from a plasma sample prior to analysis of low-abundance biomarkers. The sample can include depleted plasma.
[0171] The samples described or used herein may be from a subject. The subject may be a vertebrate. The subject may be a mammal. The subject may be a human. The subject may be at least 18 years old.
[0172] treatment Disclosed herein are methods that include administering treatment or therapy to a subject in need thereof.Various methods of the present disclosure include treating a disease state, such as cancer, in a patient in need thereof, where a biomarker, such as a peptide from the peptides listed in Table 1, is identified in the patient's sample.Treatment or therapy can be administered according to or based on the biomarker measurement described herein.Biomarkers can be measured using the methods described herein.
[0173] The methods described herein may include administering a cancer treatment to a subject. The methods described herein may include administering a pulmonary disease treatment to a subject. The methods described herein may include administering a lung cancer treatment to a subject. The methods described herein may include administering a pulmonary disease treatment other than a cancer treatment to a subject. The methods described herein may include administering an NSCLC treatment to a subject. The methods described herein may include administering a cancer treatment to a subject based on the subject's disease state. The methods described herein may include administering a pulmonary treatment to a subject based on the subject's disease state. The methods described herein may include administering an NSCLC treatment to a subject based on the subject's disease state.
[0174] Disclosed herein are methods of treatment, including: (a) obtaining or receiving a measurement of one or more biomarkers described herein; (b) administering a lung cancer treatment to the subject based on the presence of one or more biomarkers; (c) monitoring the subject without administering a lung cancer treatment to the subject based on the absence of one or more biomarkers; and (d) identifying the subject as having lung cancer and administering treatment to the subject.
[0175] The biomarkers may include peptides. In some cases, at least two peptides, at least three peptides, four peptides, five peptides, eight peptides, ten peptides, fifteen peptides, or twenty peptides from among the peptides listed in Table 1 are identified in a patient sample. In some cases, the type, duration, dosage, or frequency of treatment is determined by the combination or relative abundance of peptides from among the peptides listed in Table 1 identified in a patient sample. In some cases, the efficacy of treatment is determined by the combination or relative abundance of peptides from among the peptides listed in Table 1 identified in a patient sample. In some cases, the combination or relative abundance of peptides from among the peptides listed in Table 1 diagnoses a patient as having or not having cancer. In some cases, the combination or relative abundance of peptides from among the peptides listed in Table 1 diagnoses a type of cancer. In some cases, the combination or relative abundance of peptides from among the peptides listed in Table 1 indicates whether a cancer treatment should or should not be administered to a patient. In some cases, the sample is a plasma sample. In some cases, the cancer is lung cancer, such as NSCLC.
[0176] Various methods of the present disclosure include tracking the progress of cancer treatment. The methods may include detecting biomarkers over time in multiple samples taken from a patient. In some cases, the methods include measuring changes in the level of at least one peptide from among the peptides listed in Table 1 in samples from the patient over time to determine whether to discontinue or modify treatment (e.g., adjust administration frequency or dose). For example, the methods may include measuring the concentrations of at least two proteins selected from the group consisting of ANGL6, NOTUM, CILP1, RLA2, and GP1BB in plasma samples taken from the patient at biweekly intervals, and determining when to discontinue treatment or initiate secondary treatment based on changes in the concentrations of the at least two proteins.
[0177] In some cases, the treatment includes chemotherapy, which may include, but is not limited to, adriamycin, amsacrine, azathioprine, bleomycin, busulfan, capecitabine, carboplatin, chlorambucil, cisplatin, cyclophosphamide, cytarabine, daunorubicin, docetaxel, doxorubicin, epirubicin, etoposide, floxuridine, fludarabine, gemcitabine, ifosfamide, iproplatin, irinotecan, leucovorin, mechlorethamine, melphalan, mercaptopurine, methotrexate, mitomycin, mitoxantrone, nitrosoureas, oxaliplatin, paclitaxel, plicamycin, podophyllotoxin, satraplatin, spiropsin, thiazolinone ... cyclophosphamide, ifosfamide, chlorambucil, busulfan, melphalan, mechlorethamine, uramustine, thiotepa, nitrosoureas, 5-fluorouracil, azathioprine, 6-mercaptopurine, methotrexate, leucovorin, capecitabine, cytarabine, floxuridine, fludarabine, gemcitabine, vincristine, vinblastine, vinorelbine, oxaliplatin, cisplatin, carboplatin, spiroplatin, iproplatin, satraplatin, cyclophosphamide, ifosfamide, chlorambucil ... busulfan, melphalan, mechlorethamine, uramustine, thiotepa, nitrosoureas, 5-fluorouracil, azathioprine, 6-mercaptopurine, methotrexate, leucovorin, capecitabine, cytarabine, floxuridine, fludarabine, gemcitabine, vincristine, vinblastine, vin Treatments include norelbine, vindesine, podophyllotoxin, paclita docetaxel, irinotecan, topotecan, amsacrine, etoposide, teniposide, doxorubicin, adriamycin, daunorubicin, epirubicin, actinomycin, bleomycin, mitomycin, mitoxantrone, plicamycin, or any combination thereof. In some cases, the treatment includes immunotherapy. In some cases, the treatment includes hormone therapy. In some cases, the treatment includes monoclonal antibody treatment. In some cases, the treatment includes an mTOR inhibitor. In some cases, the treatment includes stem cell transplantation. In some cases, the treatment includes radiation therapy. In some cases, the treatment includes gene therapy. In some cases, the treatment includes chimeric antigen receptor (CAR)-T cell or transgenic T cell administration. In some cases, the treatment includes resective surgery. For example, a CT scan may identify an adenocarcinoma tumor in a patient, and analysis of a protein selected from the group consisting of ANGL6, NOTUM, CILP1, RLA2, and GP1BB from a blood sample from the patient may determine that the tumor is malignant and therefore that removal of the tumor is likely to lead to a favorable outcome.
[0178] In some cases, the treatment includes a cancer treatment. In some cases, the treatment includes multiple cancer treatments. Examples of cancer treatments include: anticancer drugs such as abemaciclib, abiraterone acetate, Abraxane (paclitaxel-albumin-stabilized nanoparticles), ABVD, ABVE, ABVE-PC, AC, acalabrutinib, AC-T, Actemra (tocilizumab), Adcetris (brentuximab vedotin), ADE, Ado-trastuzumab emtansine, Adriamycin (doxorubicin hydrochloride), afatinib dimaleate, Afinitor (everolimus), Aquinzeo (netupitant and palonosetron hydrochloride), Aldara (imiquimod), aldesleukin, Alecensa (alectinib), Alemtuzumab, Alimta (pemetrexed disodium), Alicopa (copanlisib hydrochloride), Alkeran for injection (melphalan hydrochloride), Alkeran tablets (melphalan), Aloxi (palonosetron hydrochloride), alpelisib, Alumbrig (brigutinib), Amels (aminolevulinic acid hydrochloride), amifostine, aminolevulinic acid hydrochloride, anastrozole, apalutamide, aprepitant, Aranesp (darbepoetin alfa), Aredia (pamidronate disodium), Arimidex (anastrozole), Aromasin (exemestane), Alanon (nelarabine), arsenic trioxide, Arzera (ofatumumab), asparaginase Erwinia chrysanthemi, Asparagus (calaspargase pegol-mknl), atezolizumab, avapritinib, Avastin (bevacizumab), avelumab, axicarbagene ciloreucel, axitinib, Ivakit (avapritinib), azacitidine, azedra (iobenguane I) 131), Barversa (erdafitinib), Bavencio (avelumab), BEACOPP, belantamab mafodotin-blmf, Beleodac (belinostat), belinostat, bendamustine hydrochloride, Bendeca (bendamustine hydrochloride), BEP, Besponsa (inotuzumab ozogamicin), bevacizumab, bexarotene, bicalutamide, BiCNU (carmustine), binimetinib, Brenrep (belantamab mafodotin-blmf), bleomycin sulfate, blinatumomab,Blincyto (blinatumomab), bortezomib, Bosulif (bosutinib), bosutinib, Braftobi (encorafenib), brentuximab vedotin, brexcabtagene outrousel, brigutinib, Brukinsa (zanubrutinib), BuMel, busulfan, Busulfex (busulfan), cabazitaxel, Cabryvi (caplacizumab-yhdp), Cabometyx (cabozantinib-S-malate), cabozantinib-S-malate, CAF, calaspargase pegol-mknl, Calquence (acalabrutinib), Campagnolo (alemtuzumab), Camptosar (irinotecan hydrochloride), capecitabine, caplacizumab-yhdp, capmatinib hydrochloride, CAPOX, Carac (topical fluorouracil), carboplatin, CARBOPLATIN-TAXOL, carfilzomib, carmustine, carmustine implant, Casodex (bicalutamide), CEM, cemiplimab-rwlc, ceritinib, Cerbidine (daunorubicin hydrochloride), Cervarix (recombinant HPV bivalent vaccine), cetuximab, CEV, chlorambucil, CHLORAMBUCIL Prednisone, CHOP, cisplatin, cladribine, clofarabine, chloral (clofarabine), CMF, cobimetinib fumarate, Cometriq (cabozantinib-S-malate), copanlisib hydrochloride, COPDAC, Copictra (duvelisib), COPP, COPP-ABV, Cosmegen (dactinomycin), Cotellic (cobimetinib fumarate), crizotinib, CVP, cyclophosphamide, Cyramza (ramucirumab), cytarabine, dabrafenib mesylate, dacarbazine, Dacogen (decitabine), dacomitinib dactinomycin, daratumumab, daratumumab and hyaluronidase-fihj, darbepoetin alfa, darolutamide, Darzalex (daratumumab), Darzalex Faspro (daratumumab and hyaluronidase-fihj), dasatinib, daunorubicin hydrochloride, daunorubicin hydrochloride and cytarabine liposomal, Daurismo (glasdegib maleate), decitabine, decitabine and cedazuridine, defibrotide sodium, Defitelio (defibrotide sodium), degarelix, denileukin diftitox,Denosumab, dexamethasone, dexrazoxane hydrochloride, dinutuximab, docetaxel, Doxil (doxorubicin hydrochloride liposomal), doxorubicin hydrochloride, doxorubicin hydrochloride liposomal, durvalumab, duvelisib, Efudex (fluorouracil - topical), Eligard (leuprolide acetate), ERYTECH (rasburicase), Elence (epirubicin hydrochloride), elotuzumab, Eloxatin (oxaliplatin), eltrombopag olamine, Erzonris (taglaxofusp - erzs), emapalumab - lzsg, ibuprofen Mend (aprepitant), Empliti (elotuzumab), enasidenib mesylate, encorafenib, enfortumab vedotin-ejfv, Enhertz (Fam-trastuzumab deruxtecan-nxki), entrectinib, enzalutamide, epirubicin hydrochloride, EPOCH, epoetin alfa, Epogen (epoetin alfa), Erbitux (cetuximab), erdafitinib, eribulin mesylate, Erivedge (vismodegib), Erleada (apalutamide), erlotinib hydrochloride, Erwinase (Erwinia chrysanthemi-derived asparaginase), Ethiol (amifostine), Etopofos (etoposide phosphate), etoposide, etoposide phosphate, everolimus, Evista (raloxifene hydrochloride), Evomela (melphalan hydrochloride), exemestane, 5-FU (fluorouracil injection), 5-FU (fluorouracil - topical), Fam-trastuzumab deruxtecan-nxki, Fairston (toremifene), Farydak (panobinostat), Faslodex (fulvestrant), FEC, fedratinib hydrochloride, Femara La (letrozole), filgrastim, Filmagon (degarelix), fludarabine phosphate, Fluoroplex (topical fluorouracil), fluorouracil injection, topical fluorouracil, flutamide, FOLFIRI, FOLFIRI-BEVACIZUMAB, FOLFIRI-CETUXIMAB, FOLFIRINOX, FOLFOX, Folotin (pralatrexate), fostamatinib disodium, Fulfira (pegfilgrastim), FU-LV, fulvestrant, Gamifant (emapalumab-lzsg),Gardasil (recombinant HPV quadrivalent vaccine), Gardasil 9 (recombinant HPV nonavalent vaccine), Gabret (pralsetinib), Gadiva (obinutuzumab), gefitinib, gemcitabine hydrochloride, gemcitabine-cisplatin, gemcitabine-oxaliplatin, gemtuzumab ozogamicin, Gemzar (gemcitabine hydrochloride), Gilotrif (afatinib dimaleate), Gilteri Tinib fumarate, glasdegib maleate, Glivec (imatinib mesylate), Gliadelwehr (carmustine implant), glucarpidase, goserelin acetate, granisetron, granisetron hydrochloride, Granix (filgrastim), Halaven (eribulin mesylate), Hemangeol (propranolol hydrochloride), Herceptin, Hylecta (trastuzumab, and hyaluronidase-oysk), Herceptin (trastuzumab), HPV bivalent vaccine, recombinant, HPV nonavalent vaccine, recombinant, HPV quadrivalent vaccine, recombinant, Hycamtin (topotecan hydrochloride), Hydrea (hydroxyurea), hydroxyurea, Hyper CVAD, Ibrance (palbociclib), ibritumomab tiuxetan, ibrutinib, ICE, Iclusig (ponatinib hydrochloride), idarubicin PFS (idarubicin hydrochloride), idarubicin hydrochloride, idelalisib, Idhifa (enasidenib mesylate), Ifex (ifosfamide), Ifosfamide, IL-2 (aldesleukin), imatinib mesylate, Imbruvica (ibrutinib), Imfinzi (durvalumab), imiquimod, Imlizic (talimogene laherparepvec), Infugem (gemcitabine hydrochloride), Inlyta (axitinib), inotuzumab ozogamicin, Incovi (decitabine and cedazuridine), Inlevic (fedratinib hydrochloride), interferon alfa-2b, recombinant, interleukin-2 (aldesleukin), Intron A (recombinant interferon alfa-2b), iobenguane I 131, ipilimumab, Iressa (gefitinib), irinotecan hydrochloride, irinotecan hydrochloride liposomal, isatuximab-irfc, Istodax (romidepsin), ivosidenib, ixabepilone, ixazomib citrate, Ixempra (ixabepilone), Jakafi (ruxolitinib), JEB, Jelmit (mitomycin), Jevtana (cabazitaxel), Kadcyla (Ado-trastuzumab emtansine), Kepivance (palifermin), Keytruda (pembrolizumab), Kisqali (ribociclib), Cosergo (selumetinib sulfate), Kymriah (tisagenlecleucel), Kyprolis (carfilzomib), lanreotide acetate, lapatinib ditosylate, larotrectinib sulfate, lenvatinib mesylate, Lenvima (lenvatinib mesylate), letrozole, leucovorin calcium, Leukeran (chlorambucil), leuprolide acetate Loride, LevulanKerastik (aminolevulinic acid hydrochloride), Libtayo (cemiplimab-rwlc), lomustine, Lonsurf (trifluridine and tipiracil hydrochloride), Lobrena (lorlatinib), lorlatinib, Lumoxity (moxetumomab pasudotox-tdfk), Lupron Depot (leuprolide acetate), lurbinectedin, luspatercept-aamt, Lutatera (lutetium 177-dotatate), lutetium 177-dotatate, ...177-dotatate), Lynparza (olaparib), Marquibo (vincristine sulfate liposomal), Matulane (procarbazine hydrochloride), mechlorethamine hydrochloride, megestrol acetate, Mekinist (trametinib), Mektovi (binimetinib), melphalan, melphalan hydrochloride, mercaptopurine, mesna, Mesnex (mesna), methotrexate sodium, methylnaltrexone bromide, midostaurin, mitomycin, mitoxantrone hydrochloride, mogamulizumab-kpkc, Monjuvi (tafasitamab-cxix), moxetumomab-pasudotox-tdfk, Mozovir (plerixafor), MVAC, Mubashi (bevacizumab), Myleran (busulfan), Ma Ilotarg (gemtuzumab ozogamicin), nanoparticle paclitaxel (paclitaxel-albumin-stabilized nanoparticle formulation), necitumumab, nelarabine, neratinib maleate, Nerlynx (neratinib maleate), netupitant and palonosetron hydrochloride, Neulasta (pegfilgrastim), Neupogen (filgrastim), Nexavar (sorafenib tosylate), Nilandrone (nilutamide), nilotinib, nilutamide, Ninlaro (ixazomib citrate), niraparibut tosylate monohydrate, nivolumab, Nplate (romiplostim), Nubequo (darolutamide), Nibepria (pegfilgrastim), obinutuzumab, Odomzo (sonidegib) , OEPA, ofatumumab, OFF, olaparib, omacetaxine mepesuccinate, Oncaspar (pegaspargase), ondansetron hydrochloride, Onivyde (irinotecan hydrochloride liposomal), Ontak (denileukin diftitox), Onureg (azacitidine), Opdivo (nivolumab), OPPA, osimertinib mesylate, oxaliplatin, paclitaxel, paclitaxel-albumin-stabilized nanoparticle formulation, PAD, Padsev (enfortumab vedotin-ejfv), palbociclib, palifermin, palonosetron salt acid salt, palonosetron hydrochloride and netupitant, pamidronate disodium, panitumumab, panobinostat, pazopanib hydrochloride, PCV, PEB, pegaspargase, pegfilgrastim, peginterferon alfa-2b, PEG-Intron (peginterferon alfa-2b), Pemasil (pemigatinib), pembrolizumab, pemetrexed disodium, pemigatinib, Perjeta (pertuzumab), pertuzumab, pertuzumab, trastuzumab, and hyaluronidase-zzxf, pexidartinib hydrochloride, fes Go (pertuzumab, trastuzumab, and hyaluronidase-zzxf), Piclair (alpelisib), plerixafor, polatuzumab vedotin-piiq, Porivee (polatuzumab vedotin-piiq), ponatinib hydrochloride, Portraza (necitumumab), Potelizio (mogamulizumab-kpkc), pralatrexate, pralsetinib, prednisone, procarbazine hydrochloride, Procrit (epoetin alfa), Proleukin (aldesleukin), Prolia (denosumab), Promacta (eltrombopag olamine), propranolol ol hydrochloride, Provenzi (sipuleucel-T), Prinetol (mercaptopurine), Plixan (mercaptopurine), Kinloch (ripretinib), radium-223 dichloride, raloxifene hydrochloride, ramucirumab, rasburicase, ravulizumab-cwvz, Revlozil (raspatercept-aamt), R-CHOP, R-CVP, recombinant human papillomavirus (HPV) bivalent vaccine, recombinant human papillomavirus (HPV) nonavalent vaccine, recombinant human papillomavirus (HPV) quadrivalent vaccine, recombinant interferon alpha-2b,Regorafenib, Relistol (methylnaltrexone bromide), R-EPOCH, Retacrit (epoetin alfa), Letvimo (selpercatinib), ribociclib, R-ICE, ripretinib, Rituxan (rituximab), Rituxan Hycera (rituximab and hyaluronidase (human origin)), rituximab, rituximab and hyaluronidase (human origin), Rolapita methicillin hydrochloride, romidepsin, romiplostim, Rozlytrek (entrectinib), rubidomycin (daunorubicin hydrochloride), Rubraca (rucaparib camsylate), rucaparib camsylate, ruxolitinib phosphate, Ridapt (midostaurin), sacituzumab govitecan-hziy, Thankuso (granisetron), Circrisa (isatuximab-irfc), Sclerosol Intrapleural Aerosol (talc), selinexor, selpercatinib, selumetinib sulfate, siltuximab, sipuleucel-T, Somatuline Depot (lanreotide acetate), sonidegib, sorafenib tosylate, Sprycel (dasatinib), STANFORD, V, sterilized talc powder (talc), Steritalc (talc), Stivarga (regorafenib), sunitinib malate, Sustol (granisetron), Sutent (sunitinib malate), Silatron (peginterferon-alpha-2b), Silvant (siltuximab), Synribo (omacetaxine mepesuccinate), Tabloid (thioguanine), Tabrecta (capmatinib hydrochloride), TAC, tafasitamab-cxix, Tafinlar (dabrafenib mesylate), tagraxofusp-erzs, Tagrisso (osimertinib mesylate) ), talazoparibtosilate, talc, Talimogene laherparepvec, Tarzenna (talazoparibtosilate), tamoxifen citrate, Tarceva (erlotinib hydrochloride), Targretin (bexarotene), Tasigna (nilotinib), Tavarisse (fostamatinib disodium), Taxotere (docetaxel), tazemetostat hydrobromide, Tazverik (tazemetostat hydrobromide), Tecartas (brexcabutadiene outrusel), Tecentriq (atezolizumab), Temodar (temozolomide), temozolomide, temsiroli Mus, thioguanine, thiotepa, Tivsovo (ivosidenib), tisagenlecleucel, tocilizumab, Tolak (topical fluorouracil), topotecan hydrochloride, toremifene, Torisel (temsirolimus), Totect (dexrazoxane hydrochloride), TPF, trabectedin, trametinib, trastuzumab, trastuzumab and hyaluronidase-oysk, Treanda (bendamustine hydrochloride), Trexall (methotrexate sodium), trifluridine and tipiracil hydrochloride, Trisenox (diarsenic trioxide), Trodelvy (sacituzumab govitecan-hziy), Truxima (rituximab), tucatinib, Tukisa (tucatinib), Tulario (pexidartinib hydrochloride), Tykerb (lapatinib ditosylate), Ultomiris (ravulizumab-cwvz), Undencyca (pegfilgrastim), Unituxin (dinutuximab), uridine triacetate, VAC, valrubicin, Valstar (valrubicin), vandetanib, VAMP, Barbi (rolapitant hydrochloride), Vectibix (panitumumab), VeIP, Velcade (bortezomib), vemurafenib,Venclexta (venetoclax), venetoclax, Verzenio (abemaciclib), Vidaza (azacitidine), vinblastine sulfate, vincristine sulfate, vincristine sulfate liposomal, vinorelbine tartrate, VIP, vismodegib, Vistogard (uridine triacetate), Vitrakvi (larotrectinib sulfate), Visinpro (dacomitinib), Boraxase (glucarpidase), vorinostat, Votrien (pazopanib hydrochloride), Vyxeos (daunorubicin hydrochloride and cytarabine liposomal), Xalkori (crizotinib), Zatomep (methotrexate sodium), Xeloda (capecitabine), Xeliri, Xelox, Exziva (denosumab), Xofigo (radium-223 dichloride), Xospata (gilteritinib fumarate), Expovio (selinexor), Xtandi (enzalutamide), Yervoy (ipib Limumab), Yescarta (axicabtagene ciloreucel), Yondelis (trabectedin), Yonsa (abiraterone acetate), Zaltrap (Ziv-aflibercept), zanubrutinib, Zalxio (filgrastim), Zedula (niraparibut tosilate monohydrate), Zelboraf (vemurafenib), Zepzelca (lurbinectedin), Zevalin (ibritumomab tiuxetan), Ziextenzo ( pegfilgrastim), Zaincard (dexrazoxane hydrochloride), Zirabef (bevacizumab), Ziv-aflibercept, Zofran (ondansetron hydrochloride), Zoladex (goserelin acetate), zoledronic acid, Zolinza (vorinostat), Zometa (zoledronic acid), Zyclara (imiquimod), Zidelig (idelalisib), Zykadia (ceritinib), or Zytiga (abiraterone acetate).
[0179] kit Various aspects of the present disclosure provide kits for detecting (e.g., quantifying) the biomarkers disclosed herein. The kits may include reagents for detecting peptides from Table 1, such as anti-SAA2 antibodies. The kits may include multiple reagents for detecting multiple peptides from Table 1. The kits may include reagents for ELISA assays. The kits may also include reagents for detecting biomolecules that are not useful as biomarkers for a particular cancer. For example, the kits may include reagents for quantifying ANTR1 and ANTR2 in a biological sample, as well as reagents for quantifying ceruloplasmin, such that the ANTR1- and ANTR2-specific reagents generate cancer-specific information from the sample, and the ceruloplasmin-specific reagent serves as a calibration standard or control. The kit may include reagents for detecting at least one peptide from Table 1, at least two peptides from Table 1, at least three peptides from Table 1, at least four peptides from Table 1, at least five peptides from Table 1, at least six peptides from Table 1, at least eight peptides from Table 1, at least ten peptides from Table 1, at least twelve peptides from Table 1, at least fifteen peptides from Table 1, at least twenty peptides from Table 1, at least twenty-five peptides from Table 1, at least thirty peptides from Table 1, or at least forty peptides from Table 1, optionally together with a reagent or reagents for detecting at least one peptide not listed in Table 1. For example, the kit may include ELISA reagents for detecting at least one, at least two, at least three, at least four, at least five, at least six, at least eight, at least ten, at least twelve, at least fifteen, at least twenty, at least twenty-five, at least thirty, or at least forty peptides from Table 1, and optionally for at least one peptide not listed in Table 1.
[0180] The kit may include multiple antibodies targeting at least one, at least two, at least three, at least four, at least five, at least six, at least eight, at least ten, at least twelve, at least fifteen, at least twenty, at least twenty-five, at least thirty, or at least forty peptides from Table 1, and optionally at least one peptide not listed in Table 1.
[0181] The kit may include particles or a particle panel. Particles from a particle panel can be provided together (e.g., as a mixture) or separately. For example, a kit may include a particle panel with eight particle types, each provided in a separate well in a 96-well plate. The kit may include a particle panel containing at least one, at least two, at least three, at least four, at least five, at least six, at least eight, at least ten, at least twelve, or at least fifteen particles from the particles in Table 2. The kit may include multiple compositions containing the same particle or particles in different states (e.g., mixed with or suspended in different buffers or solutions) or in different amounts. For example, a well plate may include one set of wells with 20 μg of particles, one set of wells with 40 μg of particles, and one set of wells with 80 μg of particles. The kit may include buffers for suspending the particles, eluting biomolecules from the particles, or washing the particles. The kit may include a reagent for chemically modifying proteins (e.g., a reducing agent) or digesting proteins (e.g., a protease). The kit may include multiple reagents for enriching a subset of proteins from a sample (e.g., a particle panel) and multiple reagents for preparing the subset of proteins for mass spectrometric analysis (e.g., trypsin, a buffer, an alkylating reagent, and a reducing agent). The kit may include a reagent for lysing viruses or cells (e.g., a lysis buffer).
[0182] The kit can be configured for multiplex analysis. The kit can include multiple reagents and can be configured to collate multiple portions of a biological sample under different conditions or with different reagents. The kit can include multiple compartments, such as multiple wells in a well plate or multiple Eppendorf tubes. The compartments may be prepackaged with reagents. For example, the kit can include a well plate with multiple wells containing different affinity reagents specific for different peptides from Table 1.
[0183] The kit may be compatible for use with commercially available instruments, for example, the kit may include well plates configured for measurement of fluorescence in a microplate reader, or may include sample vials compatible with commercially available mass spectrometers.
[0184] system The present disclosure provides a system capable of implementing the methods described herein. The system may include a computer control system programmed to implement the methods of the present disclosure. Figure 13 shows a computer system programmed or otherwise configured to implement the methods provided herein. The computer system 1401 can regulate various aspects of the assays disclosed herein that are amenable to automation (e.g., the movement of any of the reagents disclosed herein on a support). The computer system 1401 can be a user's electronic device or a computer system remotely located relative to the electronic device. The electronic device can be a mobile electronic device.
[0185] The computer system 1401 includes a central processing unit (CPU, also referred to herein as "processor" or "computer processor") 1405, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 1401 may also include memory or memory locations 1410 (e.g., random access memory, read-only memory, flash memory), electronic storage 1415 (e.g., hard disk), communication interface 1420 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1425, such as cache, other memory, data storage, and / or electronic display adapters. The memory 1410, storage 1415, interface 1420, and peripheral devices 1425 are in communication with the CPU 1405 via a communication bus (solid lines), such as a motherboard. The storage 1415 may be a data storage device (or data repository) for storing data. Computer system 1401 can be operably coupled to a computer network (“network”) 1430 utilizing communication interface 1420. Network 1430 can be the Internet, an internet and / or extranet, or an intranet and / or extranet in communication with the Internet. Network 1430, in some cases, is a telecommunications and / or data network. Network 1430 can include one or more computer servers, which can enable distributed computing, such as cloud computing. Network 1430 can, in some cases, utilize computer system 1401 to implement a peer-to-peer network, which can enable devices coupled to computer system 1401 to operate as clients or servers.
[0186] The CPU 1405 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 1410. The instructions may be directed to the CPU 1405, which may then be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU 1405 may include fetch, decode, execute, and writeback.
[0187] The CPU 1405 may be part of a circuit, such as an integrated circuit. One or more other components of the system 1401 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0188] The storage device 1415 can store files, such as drivers, libraries, and saved programs. The storage device 1415 can store user data, such as user preferences and user programs. In some cases, the computer system 1401 can include one or more additional data storage devices external to the computer system 1401, such as those located on remote servers in communication with the computer system 1401 via an intranet or the Internet.
[0189] Computer server 1401 can communicate with one or more remote computer systems via network 1430. For example, computer system 1401 can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android®-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 1401 via network 1430.
[0190] The methods described herein may be implemented by machine (e.g., a computer processor) executable code stored in an electronic storage location of the computer system 1401, such as on memory 1410 or electronic storage 1415. The machine-executable or machine-readable code may be provided in the form of software. During use, the code may be executed by the processor 1405. In some cases, the code may be read from storage 1415 and stored in memory 1410 for easy access by the processor 1405. In some situations, the electronic storage 1415 may be eliminated, and the machine-executable instructions are stored in memory 1410.
[0191] The code may be pre-compiled and configured for use on a machine having a processor configured to execute the code, or may be compiled at run time. The code may be provided in a programming language that may be selected to enable execution of the code in a pre-compiled or contemporaneously compiled form.
[0192] Aspects of the systems and methods provided herein, such as computer system 1401, can be embodied in programming. Various aspects of the technology can be generally considered as a "product" or "article of manufacture" in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code can be stored in electronic storage, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. A "storage" type medium can include any or all of the tangible memory of a computer, processor, or the like, or its associated modules, such as various semiconductor memories, tape drives, disk drives, and the like, capable of providing non-transitory storage for software programming at any time. All or portions of the software can sometimes be communicated via the Internet or various other telecommunications networks. Such communication can enable, for example, loading of the software from one computer or processor to another, for example, from a management server or host computer to an application server computer platform. Thus, other types of media that may carry software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, over wired or optical telephone networks, and through various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, or the like, may also be considered media that carry software. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0193] Thus, machine-readable media such as computer-executable code may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices in any computer, such as those that may be used to implement the databases shown in the figures, or the like. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards, paper tape, any other physical storage medium with a pattern of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chip or cartridge, a carrier wave transmitting data or instructions, a cable or link transmitting such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0194] The computer system 1401 may include or be in communication with an electronic display 1435 that includes a user interface (UI) 1440, for example, to provide a readout of proteins identified using the methods disclosed herein. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0195] The methods and systems of the present disclosure may be implemented by one or more algorithms, which may be implemented by software when executed by the central processing unit 1405.
[0196] A system such as a computer system can be configured to implement the methods described herein. For example, the systems described herein can implement the statistical methods, classification methods or machine learning methods described herein. The data collected from the sensor array can be used to train a machine learning algorithm, for example, an algorithm that receives assay measurements from subjects and outputs specific assay results from each subject. Before training the algorithm, the raw data from the array can first be denoised to reduce the variability of individual variables.
[0197] The system may include a central computer server programmed to implement the methods described herein. The server may include a central processing unit (CPU, also "processor"), which may be a single-core processor, a multi-core processor, or multiple processors for parallel processing. The server may also include memory (e.g., random access memory, read-only memory, flash memory); electronic storage (e.g., hard disk); communication interfaces (e.g., network adapters) for communicating with one or more other systems; and peripheral devices, which may include cache, other memory, data storage, and / or electronic display adapters. The memory, storage, interfaces, and peripheral devices may be in communication with the processor via a communication bus (solid line), such as a motherboard. The storage may be a data storage device for storing data. The server is operably coupled to a computer network ("network") utilizing a communication interface. The network may be the Internet, an intranet and / or an extranet, an intranet and / or an extranet in communication with the Internet, a telecommunications or data network. In some cases, a network may utilize a server to implement a peer-to-peer network, which may allow devices coupled to the server to act as either clients or servers.
[0198] The storage device can store files, such as subject reports, and / or communication data about an individual, or any aspect of data relevant to the present disclosure.
[0199] The computer server can communicate over a network with one or more remote computer systems, which can be, for example, personal computers, laptops, tablets, phones, smartphones, or personal digital assistants.
[0200] In some applications, the computer system includes a single server. In other situations, the system includes multiple servers communicating with each other via an intranet, an extranet, and / or the Internet.
[0201] The server can be configured to store measurement data or databases as provided herein, patient information from the subject, such as medical history, family history, demographic data, etc., and / or other clinical or personal information of potential relevance to a particular application. Such information can be stored on a storage device or server, and such data can be transmitted over a network.
[0202] The methods described herein can be implemented by machine (or computer processor) executable code (or software) stored in an electronic storage location of a server, for example, on a memory or electronic storage device. During use, the code can be executed by the processor. In some cases, the code can be read from storage and stored in memory for easy access by the processor. In some situations, electronic storage can be omitted, and the machine-executable instructions can be stored in memory. Alternatively, the code can be executed on a second computer system.
[0203] Aspects of the systems and methods provided herein, such as a server, can be embodied in programming. Various aspects of the technology can be generally considered "products" or "articles of manufacture" in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code can be stored in electronic storage, such memory (e.g., read-only memory, random-access memory, flash memory), or a hard disk. "Storage" type media can include any or all of the tangible memory of a computer, processor, or the like, or its associated modules, such as various semiconductor memories, tape drives, disk drives, and the like, capable of providing non-transitory storage for software programming at any time. All or portions of the software can sometimes be communicated via the Internet or various other telecommunications networks. Such communication can enable, for example, loading of the software from one computer or processor to another, for example, from a management server or host computer to an application server computer platform. Thus, other types of media that may carry software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, over wired or optical telephone networks, and through various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, or the like, may also be considered media that carry software. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" may refer to any medium that participates in providing instructions to a processor for execution.
[0204] The computer systems described herein may include computer executable code for performing any of the algorithms or algorithm-based methods described herein. In some applications, the algorithms described herein will use storage consisting of at least one database.
[0205] Data related to the present disclosure can be transmitted over a network or connection for receipt and / or review by a recipient. The recipient can be, but is not limited to, the subject to whom the report pertains; or their caregiver, such as a healthcare provider, administrator, other healthcare professional, or other caregiver; or the person or company that performs and / or orders the analysis. The recipient can also be a local or remote system (e.g., a server or other system in a "cloud computing" architecture) for storing such reports. In one embodiment, the computer-readable medium comprises a medium suitable for transmitting the results of an analysis of a biological sample using the methods described herein.
[0206] Aspects of the systems and methods provided herein can be embodied in programming. Various aspects of the technology can be generally considered as "products" or "articles of manufacture" in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code can be stored in electronic storage, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media can include any or all of the tangible memory of a computer, processor, or the like, or its associated modules, such as various semiconductor memories, tape drives, disk drives, and the like, capable of providing non-transitory storage for software programming at any time. All or portions of the software can sometimes be communicated via the Internet or various other telecommunications networks. Such communication can enable, for example, loading of the software from one computer or processor to another, for example, from a management server or host computer to an application server computer platform. Thus, other types of media that may carry software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, over wired or optical telephone networks, and through various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, or the like, may also be considered media that carry software. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0207] Thus, machine-readable media such as computer-executable code may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices in any computer, such as those that may be used to implement the databases shown in the figures, or the like. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards, paper tape, any other physical storage medium with a pattern of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chip or cartridge, a carrier wave transmitting data or instructions, a cable or link transmitting such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0208] The system can include a communications interface for receiving biomolecular data from a sample from a subject containing biomolecules. The sample can be exposed to a plurality of particles having distinct physicochemical properties. The biomolecular data can be received over a communications network.
[0209] The system may include a communication interface that receives biomarker data from a sample from a subject suspected of having non-small cell lung cancer (NSCLC). The biomarkers may include one or more biomarkers described herein.
[0210] The system may include a computer in communication with the communications interface. The computer may include a computer processor. The computer may include a computer-readable medium containing machine-executable code that, when executed by the computer processor, implements a method. The method may include (i) receiving biomolecular data over a communications network, (ii) combining the biomolecular data to generate a biomolecular fingerprint for the sample, and / or (iii) assigning a label to the biomolecular fingerprint. The label may correspond to a disease state described herein. The label may correspond to the presence or absence of small cell lung cancer (NSCLC) in the subject. The system may include an output device configured to output information regarding the label.
[0211] The system may include a communication interface that receives biomarker data from a sample from a subject suspected of having cancer or lung cancer, such as non-small cell lung cancer (NSCLC). The sample may include one or more biomarkers described herein. The biomolecular data may be received through a communication network.
[0212] The system may include an assay device that generates biomolecular data. The assay device may generate the biomolecular data by performing at least one aspect of an assay, including mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blot, dot blot, or immunostaining. In some cases, the assay device transmits the biomolecular data over a communication network. The assay device may include a mass spectrometer. The biomolecular data may include mass spectra. The assay device may include a chromatography device (e.g., a liquid chromatography device or a high-performance liquid chromatography device). The biomolecular data may include chromatography data. The assay device may include a lateral flow assay device. The biomolecular data may include lateral flow assay data. The assay device may include an immunoassay device (e.g., an enzyme-linked immunosorbent assay device, a Western blot device, a dot blot device, or an immunostaining device). The biomolecular data may include immunoassay data, e.g., an image or blot. [Example]
[0213] The following examples are illustrative and not limiting of the scope of the devices, systems, fluidic devices, kits and methods described herein.
[0214] Example 1 Non-small cell lung cancer (NSCLC) research This example describes a non-small cell lung cancer (NSCLC) study.
[0215] Sample Design and Collection, Data Collection. Data for three arms were collected at multiple sites: NSCLC (all stages), pulmonary comorbidities, and healthy controls. Inclusion and exclusion criteria for sample selection were as follows: 1) informed consent to donate 50 mL if 18 years of age or older; 2) no history of any cancer; 3) for NSCLC subjects, a pathology-confirmed diagnosis and no previous treatment for newly diagnosed cancer; 4) for pulmonary comorbidity controls, subjects had one of more of the following: COPD, emphysema, cardiovascular disease, hypertension, pulmonary fibrosis, asthma, or any other chronic lung disease; 5) for healthy controls, callback from the collection site stating that the subject did not have NSCLC or a pulmonary disease (they could have other diseases). For NSCLC subjects with known previous diagnostic procedures and results, the median time from the diagnostic procedure was 26 days, and samples were collected either during a post-diagnosis informational visit or immediately prior to treatment. Data collected included: 1) nanoparticle-panel data: 10 particle types were incubated in depleted plasma ("DP") and samples were randomized across 4 plates per particle type / DP. Data collected included assay process and mass spectrometry (MS) injection controls; 2) targeted MS data: assays were developed and implemented for 51 peptides from 31 proteins based on a known panel; and 3) ELISA data: assays were implemented for two candidate proteins, including CA-125 and CK19. 288 subjects were included in the study over a 9-week period.
[0216] Twenty-four centers collected subject samples into NSCLC stages 1, 2, and 3 (early stage), NSCLC stage 4 (late stage), or healthy and pulmonary comorbidity control arms. Samples included plasma and serum tubes, PAXgene RNA tubes, and Streck blood cell collection tubes. A randomly selected cohort of 288 age- and sex-matched subjects was used for NP protein profiling. Peptides from proteins bound by NPs were evaluated by data-independent acquisition mass spectrometry (DIA-MS). Depleted plasma was also prepared for analysis. 268 subject samples provided complete data sets for all 10 particle types in the panel and depleted plasma (80 healthy, 80 comorbidity control, 61 early stage NSCLC (stages 1, 2, and 3), and 47 late stage NSCLC (stage 4)). MS data acquisition for all 288 samples took 7 weeks. Historically, analysis of depleted plasma alone has not been productive. The depth of protein profiling with the particle panel allowed for in silico removal of all proteins associated with depleted plasma prior to classifier analysis. This focused the analysis on novel proteins that would not otherwise be observable in a study of this size. Classification analysis was performed for each pairwise comparison of study arms using 10 rounds of 10-fold cross-validation with a random forest model.
[0217] Subjects were age- and sex-matched, and data from multiple centers were included in each class (those with comorbidities, healthy subjects, NSCLC stage 1 "NSCLC_1," NSCLC stage 2 "NSCLC_2," NSCLC stage 3 "NSCLC_3," and NSCLC stage 4 "NSCLC_4") to avoid bias. Figure 1 shows the age and sex breakout for 268 subjects in the NSCLC biomarker discovery study. NSCLC stages 1, 2, and 3 were thoroughly explored as "early NSCLC" to increase the power of classifier generation. This study had no age or sex bias by class for the 141 subjects used in the healthy (80 subjects) vs. NSCLC (61 subjects) classification study, as shown in Table 3. [Table 3]
[0218] A summary of the particle types in the 10 particle type panel is shown below in Table 4. All of these are superparamagnetic. [Table 4]
[0219] Initial observations from the NSCLC study quantified the number of proteins observed using a 10-particle type panel. The average protein count observed in samples using the 10-particle type panel was 1,797 ± 337. Figure 2 shows the protein counts by each study group, including healthy controls, those with comorbidities, NSCLC stage 1 (NSCLC_1), NSCLC stage 2 (NSCLC_2), NSCLC stage 3 (NSCLC_3), and NSCLC stage 4 (NSCLC_4). Figure 3 shows the protein counts for the depleted plasma DP and particle panel.
[0220] Particles were observed to achieve superior protein detection consistency compared to depleted plasma with similar intensity bias. Variation in protein group detection was assessed as a function of intensity. Proteins detected in healthy subjects (n=82) from the NSCLC study were scored for each particle type, including the number of subjects in which a given protein was detected and the average signal intensity for that protein. Figure 4 shows the resulting summary of protein trace detection across subjects versus the average abundance of that protein for all 10 particle types in the particle panel and depleted plasma (DP). The curves are smooth fits of the data. As shown in Figure 4, particles outperformed depleted plasma in terms of detection consistency. At a given intensity, depleted plasma showed the lowest trace protein detection across all samples.
[0221] An average of 1,779 proteins were detected in each of the 268 subject samples with the multiparticulate type panel, compared with only 413 in depleted plasma.
[0222] Classification of healthy versus early-stage NSCLC (stages 1, 2, and 3). The initial classifier build demonstrated comparable high performance between depleted plasma ("DP") and the 10-particle panel ("panel"). Examination of key features for both methods revealed acute phase response (APR) or stress-related proteins as potential drivers of initial classification. The diagnostic procedure itself and subject perception of the diagnostic result may induce APR and other stress-related proteins as (artifactual) classifier signals. Removal of any particle panel features for proteins was also observed in depleted plasma, which removed potential bias. This option cannot be used to "shallow" the profiling effort. The final cross-validated classifier leveraged the deep profiling available in the particle panel. Figure 5 shows the performance of the cross-validated particle panel classifier, with the x-axis representing the fraction of false-positive classifications and the y-axis representing the fraction of true-positive classifications. APR and stress protein bias was observed in depleted plasma and the 10-particle panel ("panel"). As shown in Tables 5 and 6 below, the top features were identified to be associated with APR and associated proteins, which were the main drivers of initial classification. The importance scores indicate that APR proteins, particularly CRP, drove the initial performance of the classifier. Figure 6 shows graphs of the random forest model for healthy versus NSCLC (stages 1, 2, and 3) for depleted plasma (left) and the 10 particle type panel (right), depicting the false positive rate on the x-axis and the true positive rate on the y-axis. [Table 5] [Table 6]
[0223] The final classifier included features that emphasize the importance of unbiased proteomics. This final classifier used proteins known to have high and low importance in NSCLC, as well as proteins not previously important in NSCLC. Table 7 shows the proteins in the final classifier. The OT score is the OpenTargets database score for the protein. An OT score of 0 indicates that the protein is not registered in OpenTargets for lung cancer. These proteins are newly discovered features from the studies described above. Higher OT scores provide valid validation for building the classifier using proteins associated with lung cancer. For example, TBA1A and SDC1 are drug targets for lung cancer and are part of the classifier. [Table 7]
[0224] Comparison of the top features including the NSCLC classifier with the top features including the comorbidity classifier showed significant differences that may allow for clinical differentiation. Furthermore, examination of the top 20 NSCLC classifier features highlights proteins both known and unknown to be involved in NSCLC, as judged by OpenTargets (OT) annotation.
[0225] Figure 7 shows the performance of classifier features across study samples. Each graph shows the difference in protein levels for the top 20 features across all subject data for various particle types. A difference of 0.3 on the y-axis represents approximately a 2-fold change in protein levels. Data were suitable for validation by ELISA.
[0226] Figure 8 shows the results from 10 replicates of 10 rounds of 10-fold cross-validation with randomized subject class assignment, with the false positive rate on the x-axis and the true positive rate on the y-axis. Because performing measurements on a small number of samples can result in some features overfitting, separating the two groups by chance, 10 rounds of 10-fold cross-validation were performed to avoid overfitting. The subject class ("healthy" or "NSCLC") was randomized 10 times. Each time, a new 10-round 10-fold cross-validation was performed. The data shown in Figure 8 are features present in the 10-particle type panel protein dataset after removing proteins found in depleted plasma. The mean area under the curve (AUC) for the class-randomized classifier was 0.52 ± 0.04 (maximum: 0.58). No overfitting was observed with this random forest classifier build.
[0227] The performance of candidate markers was assessed by targeted mass spectrometry (MS) and ELISA. Targeted MS and ELISA were used to evaluate candidate markers identified from a public NSCLC classifier panel. 51 peptides were targeted by MS, and two proteins were detected by ELISA. Proteins detected in depleted plasma were excluded from consideration for the particle panel data described above. Figure 9 shows the ROC plot for 13 peptides by MRM-MS and two proteins by ELISA after removing proteins found in depleted plasma. The x-axis shows the false positive rate, and the y-axis shows the true positive rate. Table 8 shows the proteins detected by targeted MS and ELISA. [Table 8]
[0228] Figure 10 shows the random forest model for the overall study group comparison. The classifier for the overall study group comparison included 10 rounds of 10-fold cross-validation after removal of depleted plasma-related features in the overall classifier build. Random classification of healthy versus early-stage NSCLC after removal of depleted plasma-related proteins achieved an average AUC of 0.90. Comparison of the same healthy subjects with late-stage NSCLC and comorbid disease subjects achieved average AUCs of 0.98 and 0.84, respectively.
[0229] Figure 11 shows the differentiation of important features in the study group comparisons. A comparison of proteins associated with the top 20 features for each of the six pairwise groups is depicted.
[0230] In one analysis shown in Figure 14, 13 of the top 17 proteins in the classifier (76%) were secreted proteins. In that analysis, for plasma, approximately 28% of the proteins picked up by the proteograph from the reference plasma were secreted proteins. Secreted proteins may play an important role in the mechanism of cancer disease and treatment. Some cancer driver mutations are intracellular (e.g., BRAF, KRAS, PIK3CA, TP53) or receptor proteins (e.g., EGFR). Figure 15 includes some necessary details for some biomarkers.
[0231] Example 2 Lung cancer detection This example describes the detection of lung cancer by using a classifier trained to distinguish between various biological conditions using the biomarkers disclosed herein. Drugs are engineered to target any one of the biomarkers listed in Table 7, including ANGL6_HUMAN, HTRA1_HUMAN, PXDN_HUMAN, CCL18_HUMAN, ANTR2_HUMAN, TBA1A_HUMAN, SDC1_HUMAN, SAA2_HUMAN, CSPG2_HUMAN, ANTR1_HUMAN, NOTUM_HUMAN, CILP1_HUMAN, CAN2_HUMAN, RLA2_HUMAN, SIAT1_HUMAN, or GP1BB_HUMAN. Optionally, the agent targets more than one of ANGL6_HUMAN, HTRA1_HUMAN, PXDN_HUMAN, CCL18_HUMAN, ANTR2_HUMAN, TBA1A_HUMAN, SDC1_HUMAN, SAA2_HUMAN, CSPG2_HUMAN, ANTR1_HUMAN, NOTUM_HUMAN, CILP1_HUMAN, CAN2_HUMAN, RLA2_HUMAN, SIAT1_HUMAN, or GP1BB_HUMAN.
[0232] A sample is obtained from a subject and incubated with a particle panel disclosed herein (e.g., the 10-particle panel in Table 4). The particles are separated from the sample to remove unbound proteins, and the biomolecular corona on the particles is analyzed by mass spectrometry for one or more of the biomarkers described above. The biological state of the sample is determined using a trained classifier trained to distinguish between healthy, comorbid, and NSCLC stage 1, 2, and 3 biological states based on one or more of the biomarkers described above.
[0233] Example 3 Lung Cancer Treatment This example describes the treatment of lung cancer with drugs that target the biomarkers disclosed herein.The drugs are engineered to target any one of the biomarkers listed in Table 7, including ANGL6_HUMAN, HTRA1_HUMAN, PXDN_HUMAN, CCL18_HUMAN, ANTR2_HUMAN, TBA1A_HUMAN, SDC1_HUMAN, SAA2_HUMAN, CSPG2_HUMAN, ANTR1_HUMAN, NOTUM_HUMAN, CILP1_HUMAN, CAN2_HUMAN, RLA2_HUMAN, SIAT1_HUMAN, or GP1BB_HUMAN. Optionally, the drug targets one or more of ANGL6_HUMAN, HTRA1_HUMAN, PXDN_HUMAN, CCL18_HUMAN, ANTR2_HUMAN, TBA1A_HUMAN, SDC1_HUMAN, SAA2_HUMAN, CSPG2_HUMAN, ANTR1_HUMAN, NOTUM_HUMAN, CILP1_HUMAN, CAN2_HUMAN, RLA2_HUMAN, SIAT1_HUMAN, or GP1BB_HUMAN. The drug is produced by chemical synthesis or recombinant expression. The drug is administered to a subject in need thereof. The subject has lung cancer. Administration to the subject alleviates lung cancer symptoms and / or targets and eliminates lung cancer cells.
[0234] While preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be utilized in practicing the disclosure. It is intended that the following claims define the scope of the disclosure, and that methods and structures within the scope of these claims and their equivalents be covered thereby. In certain embodiments, for example, the following are provided: (Item 1) obtaining a dataset containing protein information from biomolecular coronas corresponding to physiochemically distinct particles incubated with a biological fluid sample from a subject; and using a classifier to identify, based on the dataset, the biological fluid sample as indicative of a healthy state, a cancerous state, or a comorbidity thereof in the subject. A method comprising: (Item 2) 2. The method of item 1, wherein the cancer is non-small cell lung cancer (NSCLC) and the comorbidity is a pulmonary comorbidity. (Item 3) 3. The method of item 2, wherein the pulmonary comorbidity is a chronic lung disease other than non-small cell lung cancer. (Item 4) the pulmonary comorbidity is selected from the group consisting of chronic obstructive pulmonary disease (COPD), emphysema, cardiovascular disease, hypertension, pulmonary fibrosis, asthma, chronic lung disease, and any combination thereof; The method described in item 2. (Item 5) 5. The method of any one of items 1 to 4, wherein the cancer condition is identified with a sensitivity or specificity of about 80% or higher. (Item 6) The protein information includes angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), CC motif chemokine 18 (CCL18), anthrax toxin receptor 2 (ANTR2), tubulin alpha-1A chain (TBA1A), syndecan-1 (SDC1), serum amyloid A-2 protein (SAA2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitreole 6. The method of any one of items 1 to 5, wherein the expression information for a protein selected from the group consisting of yl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), 60S acidic ribosomal protein P2 (RLA2), beta-galactoside alpha-2,6-sialyltransferase 1 (SIAT1), and platelet glycoprotein Ib beta chain (GP1BB). (Item 7) 7. The method of any one of items 1 to 6, wherein the protein information comprises expression information for a protein selected from the group consisting of ANGL6, HTRA1, PXDN, ANTR2, CSPG2, ANTR1, NOTUM, CILP1, CAN2, and GP1BB. (Item 8) 8. The method of any one of items 1 to 7, wherein the protein information comprises expression information for secreted proteins. (Item 9) 9. The method of any one of items 1 to 8, wherein obtaining a dataset comprises contacting the biological fluid sample with the physiochemically distinct particles to form the biomolecular corona. (Item 10) 10. The method of any one of items 1 to 9, wherein the physiochemically distinct particles comprise lipid particles, metal particles, silica particles, or polymer particles. (Item 11) 11. The method of any one of items 1 to 10, wherein the physiochemically distinct particles comprise carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles. (Item 12) 12. The method of any one of items 1 to 11, wherein obtaining the dataset comprises detecting proteins of the biomolecular corona by mass spectrometry, chromatography, liquid chromatography, high performance liquid chromatography, solid phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blot, dot blot, or immunostaining, or a combination thereof. (Item 13) 13. The method of any one of items 1 to 12, wherein obtaining a dataset comprises detecting the proteins of the biomolecular corona by mass spectrometry. (Item 14) 14. The method of any one of items 1 to 13, wherein obtaining the dataset comprises measuring a readout indicative of the presence, absence or amount of a protein in the biomolecular corona. (Item 15) The classifier removes or filters biomolecules associated with an acute phase response. 15. The method of any one of items 1 to 14, wherein the nucleotide sequence is generated by multiplying and subtracting. (Item 16) 16. The method of any one of items 2 to 15, wherein the NSCLC comprises early stage NSCLC (stage 1, stage 2, or stage 3). (Item 17) 16. The method of any one of items 2 to 15, wherein the NSCLC comprises late stage NSCLC (stage 4). (Item 18) 18. The method of any one of items 1 to 17, further comprising administering to the subject an NSCLC treatment based on the disease status. (Item 19) 19. The method of any one of items 1 to 18, wherein the classifier has increased protein detection consistency compared to a second classifier generated using proteomic data from a depleted plasma sample. (Item 20) 20. The method of any one of items 1 to 19, wherein the biological fluid comprises a blood sample from which red blood cells have been removed. (Item 21) 21. The method of any one of items 1 to 20, wherein the biological fluid comprises plasma. (Item 22) 1. A method for assessing cancer status, comprising: 10. The method of claim 1, further comprising measuring biomarkers in a biological fluid sample from a subject suspected of having cancer to obtain a biomarker measurement value, wherein the biomarkers comprise one or more biomarkers selected from the group consisting of angiopoietin-related protein 6 (ANGL6), serine protease HTRA1 (HTRA1), peroxidasin homolog (PXDN), anthrax toxin receptor 2 (ANTR2), versican core protein (CSPG2), anthrax toxin receptor 1 (ANTR1), palmitoleoyl-protein carboxylesterase NOTUM (NOTUM), cartilage intermediate lamina protein 1 (CILP1), calpain-2 catalytic subunit (CAN2), and platelet glycoprotein Ib beta chain (GP1BB). (Item 23) 23. The method of claim 22, wherein measuring the one or more biomarkers comprises using a detection reagent that binds to the protein and produces a detectable signal. (Item 24) 24. The method of claim 22 or 23, wherein measuring the one or more biomarkers comprises measuring a readout indicative of the presence, absence, or amount of the one or more biomarkers. (Item 25) 25. The method of any one of items 22 to 24, wherein measuring the biomarkers comprises performing mass spectrometry, chromatography, liquid chromatography, high performance liquid chromatography, solid phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blot, dot blot, or immunostaining, or a combination thereof. (Item 26) 26. The method of any one of items 22 to 25, wherein measuring the biomarkers comprises performing mass spectrometry. (Item 27) 27. The method of any one of items 22 to 26, wherein measuring the biomarker comprises performing an immunoassay. (Item 28) 28. The method of any one of items 22 to 27, wherein measuring the biomarkers comprises contacting the biological fluid sample with a plurality of physiochemically distinct nanoparticles. (Item 29) 29. The method of any one of items 22 to 28, wherein the cancer comprises lung cancer. (Item 30) 30. The method of any one of items 22 to 29, further comprising applying a classifier to the biomarker measurements. (Item 31) 31. The method of claim 30, wherein the classifier distinguishes between the cancer and chronic lung disorder, chronic obstructive pulmonary disease, emphysema, cardiovascular disease, hypertension, pulmonary fibrosis, or asthma. (Item 32) 32. The method of any one of items 22 to 31, wherein the cancer comprises non-small cell lung cancer (NSCLC). (Item 33) 33. The method of item 32, wherein the NSCLC comprises early stage NSCLC (stage 1, stage 2, or stage 3). (Item 34) 33. The method of claim 32, further comprising applying a classifier to the biomarker measurements, wherein the classifier comprises features for distinguishing between early stage and late stage NSCLC. (Item 35) 35. The method of claim 34, wherein the characteristic comprises one of more biomarkers selected from the group consisting of SDC1, OC085, KV401, MYL6, JIP2, HV459, HV461, HV169, HNRPC, ROA1, STON2, LV301, KVD20, SAE1, PDE5A, RTN3, HV373, LV325, H2B1C, H2B1D, H2B1H, H2B1K, H2B1L, H2B1M, H2B1N, H2B2F, H2BFS, and NMT1. (Item 36) 33. The method of claim 32, further comprising applying a classifier to the biomarker measurements, wherein the classifier comprises features for distinguishing between the presence and absence of NSCLC. (Item 37) 37. The method of claim 36, wherein the characteristic comprises one of more biomarkers selected from the group consisting of SDC1, ANGL6, PXDN, ANTR1, OC085, SAA2, HTRA1, KPCB, KV401, OCL18, MYL6, ANTR2, GTPB2, HDGF, TBA1A, CSRP1, TCO2, CSPG2, PTPRZ, ILF2, SIAT1, ITA2B, DOK2, H31, H31T, H32, H33, H3C, RAC2, ARRB1, DHB4, HV102, RHG18, GDF15, PCSK6, FHOD1, or ITLN2. (Item 38) 38. The method of any one of items 22 to 37, further comprising identifying the subject as having said cancer based on said biomarker measurements. (Item 39) 39. The method of any one of items 22 to 38, further comprising administering to the subject a cancer treatment. (Item 40) 40. The method of any one of items 22 to 39, wherein the biological fluid comprises a blood sample from which red blood cells have been removed or comprises plasma. (Item 41) 41. The method of any one of items 22 to 40, wherein the subject is a human. (Item 42) (a) assaying a biological sample from a subject to identify a biomolecule; (b) using a trained classifier to identify the sample or the subject as positive or negative for non-small cell lung cancer based on the biomolecules identified in (a), wherein the trained classifier was trained using data from training samples including known healthy samples and known non-small cell lung cancer samples, and the training samples were assayed using a plurality of particles having distinct physicochemical properties to obtain the data. A method comprising: (Item 43) 43. The method of claim 42, wherein the biomolecule comprises a protein. (Item 44) Item 43. The method of item 42, wherein the biomolecule is a protein. (Item 45) 45. The method of any one of items 42 to 44, wherein the data comprises proteomic data identifying the presence or absence of proteins in the training samples. (Item 46) 46. The method of any one of items 42 to 45, wherein the trained classifier is configured to remove acute phase response bias or stress protein bias. (Item 47) 47. The method of any one of items 42 to 46, wherein the trained classifier comprises protein-related features, wherein the features are selected to exclude a bias of acute phase response and / or stress proteins in the biological sample. (Item 48) 48. The method of claim 47, wherein the features of the classifier exclude proteins selected from the group consisting of C-reactive protein (CRP), haptoglobin, and S10a8 / 9. (Item 49) 48. The method of claim 47, wherein the features of the classifier exclude CRP, haptoglobin, and S10a8 / 9. (Item 50) 48. The method of claim 47, wherein the features of the classifier exclude proteins listed in Table 3. (Item 51) 51. The method of any one of items 47 to 50, wherein the features of the classifier comprise secreted proteins. (Item 52) 52. The method of any one of items 47 to 51, wherein the features of the classifier comprise a plurality of proteins listed in Table 5. (Item 53) 53. The method of claim 52, wherein the features of the classifier are listed in Table 5. (Item 54) 54. The method of any one of items 47 to 53, wherein the features include tubulin alpha-1A chain (TBA1A) and syndecan-1 (SDC1). (Item 55) 55. The method of any one of items 42 to 54, further comprising obtaining a biological sample from the subject. (Item 56) 56. The method of item 55, wherein the biological sample is a complex biological sample. (Item 57) 57. The method of claim 56, wherein the complex biological sample is a plasma or serum sample. (Item 58) 58. The method of any one of items 42 to 57, wherein the plurality of particles having distinct physicochemical properties comprises two or more particles listed in Table 2 or Table 4. (Item 59) 59. The method of any one of items 42 to 58, wherein the plurality of particles having distinct physicochemical properties are listed in Table 2. (Item 60) 60. The method of any one of items 42 to 59, wherein the trained classifier has improved performance based on AUC compared to a classifier trained using proteomic data from depleted plasma samples from the same subject as the known healthy sample and the known non-small cell lung cancer sample. (Item 61) 61. The method of any one of items 42 to 60, further comprising outputting a report indicating that the sample or the subject is positive or negative for the non-small cell lung cancer. (Item 62) 62. The method of any one of items 42 to 61, wherein the assaying comprises performing mass spectrometry or ELISA, and the biomolecule comprises a protein. (Item 63) 63. The method of any one of items 42 to 62, wherein the assaying comprises performing targeted mass spectrometry. (Item 64) 64. The method of any one of items 42 to 63, wherein the trained classifier is a trained algorithm. (Item 65) 65. The method of any one of items 42 to 64, wherein the known non-small cell lung cancer sample comprises an early stage (stage 1 to 3) non-small cell lung cancer sample. (Item 66) 61. The method of any one of paragraphs 42 to 60, wherein the trained classifier identifies the sample or the subject as being positive for the non-small cell lung cancer and the stage of the non-small cell lung cancer. (Item 67) 67. The method of any one of items 42 to 66, further comprising identifying the sample or the subject as positive or negative for non-small cell lung cancer with a sensitivity greater than about 80%. (Item 68) 68. The method of claim 67, wherein the sensitivity is greater than about 85%, about 90%, about 95%, or about 99%. (Item 69) 69. The method of any one of items 42 to 68, further comprising identifying the sample or the subject as positive or negative for non-small cell lung cancer with a specificity greater than about 80%. (Item 70) 70. The method of claim 69, wherein the specificity is greater than about 85%, about 90%, about 95%, or about 99%.
Claims
[Claim 1] The invention described in the specification.