Multi-omics assessment
A multi-omics approach using intronic RNA and protein information from biofluid samples with a trained classifier effectively detects cancers early and accurately, reducing the need for invasive tests.
Patent Information
- Application Number
- PCT/US2024/062171
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-18
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-03
AI Technical Summary
There is a need for accurate and early detection of diseases such as cancer to improve treatment and prognosis, with existing methods often being invasive, costly, and prone to high false positive rates.
A multi-omics approach combining intronic RNA and protein information from biofluid samples, using a classifier trained on these data types to distinguish between the presence and absence of cancer, with feature selection and machine learning techniques to enhance accuracy.
This method enables non-invasive, early detection of cancers like pancreatic, liver, ovarian, and colon cancer with high accuracy, reducing the need for invasive procedures and improving patient outcomes.
Smart Images

Figure US2024062171_03072025_PF_FP_ABST
Abstract
Description
MULTI-OMICS ASSESSMENTCROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 615,253, filed December 27, 2023; and U.S. Provisional Application No. 63 / 661,415, filed June 18, 2024; each of which is entirely incorporated herein by reference.BACKGROUND
[0002] There is a need for methods of accurately detecting a disease state such as cancer at an early stage. Accurate and early disease detection can improve treatment and prognosis for subjects with the disease.SUMMARY
[0003] Aspects disclosed in methods comprise obtaining intronic RNA information and protein information from a biofluid sample from a subject suspected of having cancer or at risk of having cancer; inputting the intronic RNA information and protein information into a classifier having been trained using intronic RNA training information and protein training information to distinguish between the presence and absence of cancer; and generating an output using the classifier, wherein the output is indicative of a likelihood of cancer. In some embodiments, the intronic RNA information or protein information undergoes feature selection prior to applying the classifier. In some embodiments, the intronic RNA information comprise alternative splicing events. In some embodiments, the intronic RNA information comprise immature transcripts. In some embodiments, the immature transcripts are pre-mRNA. In some embodiments, the immature transcripts have not undergone splicing. In some embodiments, the classifier comprises a performance characteristic comprising an average or median area under the curve (AUC) of a receiver operating characteristic (ROC) curve of at least 0.80. In some embodiments, the cancer is selected from the group consisting of lung cancer, pancreatic cancer, colon cancer, liver cancer, breast cancer, and ovarian cancer. In some embodiments, the biofluid sample comprises a blood, serum, or plasma sample.
[0004] In some embodiments, said intronic RNA information is generated by sequencing, microarray analysis, hybridization, polymerase chain reaction, electrophoresis, or a combination thereof. In some embodiments, the method further comprises monitoring the subject when the subject does not have the cancer. In some embodiments, the method further comprises identifying, recommending, or administering a disease treatment based on a use of the biomarker or set of biomarkers. In some embodiments, the classifier is trained-using deep learning, a hierarchical cluster analysis, a principal component analysis, a partial least squares discriminant analysis, a random forest classification analysis, a gradient boosted analysis, a gradient boosted tree ensemble analysis, a support vector machine analysis, a k-nearest neighbors analysis, a naive Bayes analysis, a K-means clustering analysis, or a hidden Markov analysis. In some embodiments, the AUC is determined by application of the classifier to a data set derived from a trial of at least 20 subjects having the cancer, and over 20 control subjects not having the cancer, and wherein the ROC curve comprises a plot of true positive and false positive rates obtained when the classifier is applied to the held-out data set.In some embodiments, said protein information is generated by mass spectrometry. In some embodiments, said protein information is generated from biomolecules adsorbed to a plurality of nanoparticles. In some embodiments, the plurality of nanoparticles comprises physiochemically distinct groups of nanoparticles. In some embodiments, the physiochemically distinct groups of nanoparticles comprise lipid nanoparticles, metal nanoparticles, silica nanoparticles, or polymer nanoparticles. In some embodiments, the intronic RNA information comprises one or more of the following intronic RNAs: ENST00000216044.10+chr22_38721866_38724296, ENST00000260702.4-chrl0_98261128_98262034, ENST00000262067.5+chr7_16754031_16776210, ENST00000264870.8-chr6_6195886_6197222, ENST00000285873.8-chr5_154931578_154932098, ENST00000303212.3+chrl9_5824335_5827748, ENST00000309415.8+chr2_109419643_109432500, ENST00000309765.4+chr3_44270811_44281967, ENST00000314888.10-chr9_35714082_35714238, ENST00000316707.10+chr4_86745129_86758259, ENST00000322313.9-chr2_127708893_127709489, ENST00000327473.9+chrl9_4639630_4651866, ENST00000329335.3-chr8_10623221_10711956, ENST00000343815.10+chrl_154220245_154225083, ENST00000355754.7-chrI_89I87I03_89I8858I, ENST00000359236.I0-chrI2_I I8073993_118079345, ENST00000359314.5+chr6_47554767_47574063, ENST00000366922.3+chrl_220110938_220114313, ENST00000367059.3+chrl_207475217_207477884, ENST00000367211.6-chrl_219612268_219612624, ENST00000369780.8+chrl0_103585226_103589513, ENST00000369816.5-chr6_79926382_79947179, ENST00000370192.8-chrl_97306298_97373560, ENST00000372330.3+chr20_46013797_46014123, ENST00000372409.8+chr20_45943766_45944867, ENST00000374026.7-chr6_34606905_34654624, ENST00000374272.4-chrl_26051884_26053892, ENST00000375448.4+chrl_17342403_17346027, ENST00000376630.5+chr6_30492584_30492748, ENST00000377047.9+chrl3_93227617_93545262, ENST00000379936.3+chrl 1_6241781_6243948, ENST00000380760.4+chr7_73300639_73301066, ENST00000381340.8-chrl2_26658131_26659112, ENST00000381624.4+chrl9_5744511 5745912, ENST00000382353.6+chrl3_21672190_21681041, ENST00000389629.8+chrl5_41471350_41474619, ENST00000393409.3+chr3_127016685_127016943, ENST00000393980.8+chr5_160187455_160199048, ENST00000400405.4-chrl3_45784058_45851494, ENST00000403045.6-chr2_70648781_70663243, ENST00000403045.6-chr2_70648781_70663243, ENST00000404190.3+chrl0_88726913_88730982, ENST00000411427.3-chr21_41442747_41445489, ENST00000411433.1-chr2_218351227_218357913, ENST00000417390.1+chr6_128067597_128083670, ENST00000421239.7+chrl9_52456959_52491323, ENST00000423023.2+chrl3_19674733_19675617, ENST00000424662.1-chr20_59352322_59357308, ENST00000426173.6-chr 12_122948709_122949787, ENST00000428948.1+chr9_35646434_35646689, ENST00000429945.1-chr6_27286149_27311096, ENST00000431627.1+chrl_158880897_158883353, ENST00000435287. l+chr6_131951255_132077207, ENST00000444740.2+chrl9_41879176_41879534, ENST00000445019.5-chrl_26890852_26892669, ENST00000449581.1-chr20_7188568_7241828, ENST00000454681.2-chrl3_33383680_33439690, ENST00000468244.2+chr9_125254672_125256963, ENST00000469154.5-chr7_151243702_151275113, ENST00000469989.1+chr3_52523577_52523870, ENST00000470809.1+chrl_45787133_45824432, ENST00000486199.5+chr8_132814890_132817338, ENST00000490486.2-chr9_137008596_137008730, ENST00000500450.6-chr3_129171509_129171654,ENST00000504592.5+chr4_101721434_101829807, ENST00000505846.5+chr4_82448022_82454721, ENST00000505916.6-chr4_185378208_l 85394778, ENST00000510496.5-chr5_74805285_74834425, ENST0000051105L5-chr4_22346824_22389087, ENST00000511443.1-chr5_15192450_15243648, ENST00000512905.6-chr3_98521443_98562323, ENST00000513206.5+chr5_14502658_14531743, ENST00000518649.5+chrl4_91167379_91181185, ENST00000518836.5-chr5_159099464_159099582, ENST00000521442.1+chr8_20225407_20225688, ENST00000526372.1-chrl l_88149956_88175237, ENST00000526638.1+chrl l_75572400_75572550, ENST00000527263.1+chrl_77518585_77535917, ENST00000528313. l+chrl l_60462486_60466997, ENST00000534065.1+chrl l_66312993_66318833, ENST00000534068. l+chrl l_23731659_23794822, ENST00000535199.5-chr21_36098260_36126486, ENST00000535681.1-chrl7_80468918_80470526, ENST00000537256.5+chrl6_87601931_87644257, ENST00000538357. l-chrl2_l 18079607_l 18095532, ENST00000542280.5-chrl2_7479997_7482965, ENST00000543030.5+chrl6_57055520_57059466, ENST00000546412.2-chrl4_25664976_25835849, ENST00000551286. l-chrl2_56314967_56315820, ENST00000551516. l+chrl2_69251142_69262468, ENST00000551900.5-chrl2_53725259_53727403, ENST00000554119.5-chrl4_95710882_95712219, ENST00000554804.1+chrl4_100274759_100277417, ENST00000557595.1+chrl4_22556843_22556969, ENST00000559100.1-chrl5_50301369_50303965, ENST00000559494.1-chrl5_77054280_77070891, ENST00000563175.1-chrl6_68259822_68260309, ENST00000563237.3-chr7_102354922_102355863, ENST00000563730.1-chrl6_47459176_47463951, ENST00000563907.5-chrl5_72752896_72760406, ENST00000567873.1-chrl3_33167623_33350289, ENST00000568293.1+chrl6_56625880_56626879, ENST00000568759.1-chrl6_14735084_14740891, ENST00000576437.5-chrl6_4362132_4395288, ENST00000576479.4+chrl l_14254730_14255646, ENST00000576965.1+chrl7_4942942_4944714, ENST00000578379.5-chrl7_64006672_64020522, ENST00000581278.1+chrl8_24450815_24453149, ENST00000583510.1-chrl7_64855890_64859441, ENST00000583535.6-chrl7_10632017_10632475, ENST00000584348.5+chrl7_19495482_19542392, ENST00000586744.1-chrl9_45076877_45090068, ENST00000588212.1-chrl9_44340540_44397198, ENST00000588397.1-chrl8_13645198_13645530, ENST00000589143.5-chrl9_56376122_56390000, ENST00000592226.5-chrl7_44384586_44385163, ENST00000592234.5+chrl9_1271037_1271550, ENST00000595816.1-chrl9_17351643_17365972, ENST00000597430.2-chrl9_6592576_6603940, ENST00000598213.5-chrl9_57922630_57935079, ENST00000602162.5-chrl9_52706865_52735000, ENST00000602271.1+chrl9_51120384_51120509, ENST00000602644.5-chrl6_67663557_67663936, ENST00000608254.1-chrX_23064122_23270050, ENST00000614278.1-chr9_138179514_138179564, ENST00000616279.4-chr2_l 1217174 1218957, ENST00000616721.6-chrl9_39924733_39927058, ENST00000616721.6-chrl9_39928307_39934563, ENST00000617106.1-chrl9_39726221_39726702, ENST00000617305.4+chrl9_41264694_41305675, ENST00000619343.1-chr20_3921870_3923322, ENST00000630715.2-chr6_109454240_109465663, ENST00000635641.1-chr6_26196247_26196719, ENST00000635687.1-chrl_9151836_9182006, ENST00000640432.1-chrl4_94051070_94051339, ENST00000641481.1-chrl_38756606_38784539, ENST00000642991.1-chrl3_99409060_99417056, ENST00000643759.2-chrl_158652654_158653273, ENST00000648060. l-chrl3_76729511_76835980, ENST00000652524.1+chr4_83076353_83082729,ENST00000652631. 1-chr 14_40286951_40347879, ENST00000652635. l+chr2_239539446_239544463, ENST00000655309. l+chr6_5072735_5079555, ENST00000656176. l+chrl3_25169032_25186523, ENST00000660939.1+chr6_132147845_132160766, ENST00000663029.1-chr6_21369811_21382738, ENST00000665235.1+chrl4_75296565_75297208, ENST00000665970.1+chr3_40719916_40773867, ENST00000667372. 1 +chr3_53196917_53197230, ENST00000675287. 1- chrl5_34066812_34070587,ENST00000677787.1-chr8_100712458_100713086, ENST00000677919.1- chr8_63017631_63024079, ENST00000680973.1+chr9_64031415_64054475, ENST00000681659.1- chrl0_49493253_49505883, or ENST00000681827.1+chrl_227728285_227728689.INCORPORATION BY REFERENCE
[0005] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. International Publication No. W02023 / 240046 is incorporated herein in full. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Fig. 1A illustrates a multi-omics approach.
[0007] Fig. IB illustrates combining data sets in a multi-omics approach.
[0008] Fig. 1C shows various aspects that may be used in some methods described herein.
[0009] Fig. 2A shows examples of methods for generating and applying the classifiers described herein.
[0010] Fig. 2B is a flowchart showing some aspects that may be used in methods herein.
[0011] Fig. 3A shows examples of stages in screening and treatment of a patient having or suspected of having a disease state.
[0012] Fig. 3B shows examples of stages in pancreatic cancer patient screening and treatment.
[0013] Fig. 3C shows examples of stages in liver cancer patient screening and treatment.
[0014] Fig. 3D shows examples of stages in colon cancer patient screening and treatment.
[0015] Fig. 4A shows a non-limiting example of a computing device; in this case, a device with one or more processors, memory, storage, and a network interface.
[0016] Fig. 4B shows a non-limiting example of a web / mobile application provision system; in this case, a system providing browser-based and / or native mobile user interfaces; and
[0017] Fig. 4C shows a non-limiting example of a cloud-based web / mobile application provision system; in this case, a system comprising an elastically load balanced, auto-scaling web server and application server resources as well synchronously replicated databases.
[0018] Fig. 5 shows a diagram of classifier and feature information, in accordance with some aspects described herein.
[0019] Fig. 6 shows a non-limiting workflow for processing a sample and generating RNA data.
[0020] Fig. 7 shows a non-limiting workflow for identifying features from RNA data.
[0021] Fig. 8 shows a non-limiting workflow for processing and analyzing RNA data for the identification of subjects with or without cancer, or a risk of cancer.
[0022] Fig. 9 shows some differences recognized between pre-mRNA and mRNA.
[0023] Fig. 10 shows non-limiting examples of two products from pre-mRNA splicing including a fully spliced isoform and an intron retaining isoform.
[0024] Fig. 11 includes a receiver operating characteristic (ROC) curve and shows an area under the curve (AUC) using intronic RNA sequencing data.
[0025] Fig. 12 illustrates some extrapolated mRNA data showing differentially expressed proteins in metabolic pathways.
[0026] Fig. 13 includes a graphical depiction of protein concentrations for data obtained in a study described herein.
[0027] Fig. 14A shows some aspects that may be used in integrated models classification.
[0028] Fig. 14B shows some aspects that may be used in transformation-based classification.
[0029] Fig. 15 shows aspects of a 2-stage machine learning framework for analyzing and training multiple datatypes.
[0030] Fig. 16 shows a non-limiting example of a flowchart of machine training algorithm for improving the sensitivity and specificity of the classifier for predicating a disease described herein.
[0031] Fig. 17 includes some aspects such as subjects or test outcomes that may be included in a method described herein.
[0032] Fig. 18 depicts an example workflow showing sample preparation, data acquisition, and data analysis.
[0033] Fig. 19 depicts an example workflow including use of timsTOF Pro 2 and HT.
[0034] Fig. 20 depicts an example workflow to build spectral libraries using timsTOF.
[0035] Fig. 21 depicts an example of a computational pipeline.
[0036] Fig. 22 depicts graphs and charts showing details of a 3,044 subject study.
[0037] Fig. 23 depicts a graph shows coefficient of variation (CV%) values.
[0038] Fig. 24 depicts a methodology for running 180 plates on four machines and the high reproducibility achieved between the machines.
[0039] Fig. 25 depicts intra-batch reproducibility over 180 batches across four machines.
[0040] Fig. 26 depicts inter-batch reproducibility over 180 batches across four machines.
[0041] Fig. 27 depicts MARLE distributions of injections on the four machines.
[0042] Fig. 28 depicts a graph illustrating spectral libraries of different samples.
[0043] Fig. 29 depicts a graph illustrating improvement of two libraries.
[0044] Fig. 30 depicts a graph illustrating the number of peptides identified and the percentage of subjects the peptides were detected in.
[0045] Fig. 31 depicts a graph illustrating the number of protein groups identified and the percentage of subjects the proteins were detected in.
[0046] Fig. 32 depicts an image and a bar graph evaluating a first deep plasma spectral library content.
[0047] Fig. 33 depicts a graph illustrating HPPP proteins detected and the estimated plasma concentration of the detected proteins.
[0048] Fig. 34 depicts proteins that were not detected from a current spectral library.
[0049] Fig. 35 depicts the distribution of protein counts by subject and nanoparticle.
[0050] Fig. 36 depicts indices for platelet and erythrocyte contamination.
[0051] Fig. 37 depicts shows platelet and erythrocyte contamination indices for the various collection sites.
[0052] Fig. 38 depicts platelet and erythrocyte contamination indices for cancer and non-cancer samples.
[0053] Fig. 39 depicts proteins groups across five nanoparticles which were significantly different.
[0054] Fig. 40 depicts the number of proteins detected in various studies in relation to the number of samples analyzed.
[0055] Fig. 41 depicts a breakdown of the molecular features identified across the various groups.
[0056] Fig. 42 depicts erythrocyte and platelet contamination indices by nanoparticle.
[0057] Fig. 43 depicts the workflow for the two independent classifiers developed.
[0058] Fig. 44 depicts ROC graphs between the two classifiers for all-stage cancers.
[0059] Fig. 45 depicts ROC graphs between the two classifiers for Stage-1 cancers.
[0060] Fig. 46 depicts graphs with covariate distribution for sex, smoking status, and age.
[0061] Fig. 47 depicts a cofounder analysis on the XGBoost classifier.
[0062] Fig. 48 depicts top features identified by the XGBoost classifier.
[0063] Fig. 49 depicts a comparison of the abundance of 4 top features identified between cancer and noncancer subjects.
[0064] Fig. 50 depicts a comparison of the abundance of 4 top features identified between subjects with different cancer stages and non-cancer subjects.
[0065] Fig. 51 depicts peptide features and range.
[0066] Fig. 52 depicts an overview of the validation process for the XGBoost Classifier.
[0067] Fig. 53 depicts the composition of the protocol, cancer stage, and study sub-cohort used in the training and validation datasets.
[0068] Fig. 54 depicts ROC curves for all cancer stages for the training and validation datasets.
[0069] Fig. 55 depicts ROC curves for Stage-I and Stage-IA cancer subjects from validation datasets
[0070] Fig. 56 depicts sensitivity of validation data at 90% specificity.
[0071] Fig. 57 depicts analysis of covariates for roles in mis-classifications and the false negative rate and false positive rate based on age group.
[0072] Fig. 58 depicts the distribution and number of cancer and non-cancer samples across the multitude of clinical sites.
[0073] Fig. 59A depicts the breakdown of disease cohorts across the samples.
[0074] Fig. 59B depicts the breakdown of cancer stages across the cancer samples.
[0075] Fig. 60 depicts the distribution of clinical covariates across training and validation data sets.
[0076] Fig. 61 depicts the distribution of samples used in training and validation across the multitude of clinical sites.
[0077] Fig. 62 depicts the AUC generated with reduced features on the XGBOOST classifier.
[0078] Fig. 63A depicts the AUC generated with only proteomic and metabolomic data on the XGBOOST classifier.
[0079] Fig. 63B depicts the AUC generated with reduced proteomic and metabolomic features on the XGBOOST classifier.
[0080] Fig. 64A depicts the AUC generated with only proteomic, ortho, and metabolomic data on the XGBOOST classifier.
[0081] Fig. 64B depicts the AUC generated with reduced proteomic, ortho, and metabolomic features on the XGBOOST classifier.
[0082] Fig. 65A depicts the AUC generated with only proteomic and RNA data on the XGBOOST classifier.
[0083] Fig. 65B depicts the AUC generated with reduced proteomic and RNA sequencing features on the XGBOOST classifier.
[0084] Fig. 66 depicts the predicted probability of cancer with age in the XGBOOST classifier for cancer and non-cancer samples.
[0085] Fig. 67 compares the precursors, peptides, and protein groups identified on the timsTOF HT and Orbitrap Astral.
[0086] Fig. 68 compares the precursors, peptides, and protein groups identified on the timsTOF HT and Orbitrap Astral for each nanoparticle.
[0087] Fig. 69 compares the proteins in the library generated from timsTOF HT and Orbitrap Astral as well as a comparison between general coverage between the two platforms.
[0088] Fig. 70 compares peptide yield obtained from the Proteograph XT and Proteograph V1.2 sample processing protocols.
[0089] Fig. 71 compares the amount of peptides and proteins in libraries constructed using an original method, Spectronaut method, and PaSER method.
[0090] Fig. 72A compares average AUC and AUC distributions for different numbers of features chosen from the group of features chosen and used in the lung cancer classifier as compared to randomly chosen features.
[0091] Fig. 72B illustrates that subsets of multi -omic features retain substantially equivalent performance with AUC of 0.95 (53 proteomic features and 25 RNA transcriptomic features) and AUC of 0.94 (9 proteomic features and 10 RNA transcriptomic features) respectively
[0092] Fig. 73 shows an overview of proteomics, transcriptomics, and metabolomics assays.
[0093] Fig. 74 shows ROC curve for multiomics model (top) and unbiased proteomics model (bottom) comparing all-stage cancers versus non cancers.
[0094] Fig. 75A shows charts relating to the performance of a validated multi-omics classifier Stage-wise sensitivity and specificity of the multi-omics classifier for subjects from training (left) and validation (right) datasets. Error bars indicate 95% Clopper-Pearson confidence intervals. The number of subjects in each subgroup is denoted in parenthesis.
[0095] Fig. 75B shows mRNA features help preserve specificity and / or sensitivity of the classifier disclosed herein as the proteomic features are reduced.
[0096] Fig. 75C shows intronic RNA features help preserve specificity and / or sensitivity of the classifier disclosed herein as the proteomic features are reduced.
[0097] Fig. 76 shows charts relating to features of the multi-omics model ranked by importance score as determined by information gain.
[0098] Fig. 77 shows the association of analyte feature abundance and lung cancer stage by a plot of the slope and statistical significance of all 682 features from the validated multi-omics classifier with respect to association with cancer stage.
[0099] Fig. 78 shows a schematic for an example study data: subjects with and without lung cancer(N=2513) were enrolled in the MOSAIC study across 77 clinical studies; three blood samples were collected per subject and used for proteomics, RNA-seq, metabolomics, and targeted immunoassays; data from the omics assay were then divided into training and validation sets for the development of a machine learningbased lung cancer classification model.DETAILED DESCRIPTION
[0100] This disclosure provides non-invasive methods for diagnosing or ruling out the presence of a disease in a subject, or the risk of developing the disease in a subject. The disease may include a cancer such as pancreatic cancer, breast cancer, liver cancer, ovarian cancer, or colon cancer. Identifying an early-stage disease in a subject can save the subject from further development of the disease if treatment is provided early on. Non-invasive tests can also be used to rule out the presence of a disease, thereby saving subjects from having to undergo invasive testing such as a biopsy, which can be painful and stressful, or may risk damaging the subject.
[0101] This disclosure also provides non-invasive methods for detecting presence of a cancer such as pancreatic cancer, or risk of developing the cancer in a subject. Identifying cancer in a subject at an early stage can save the subject from further development of the cancer if treatment is provided early on. Non- invasive tests can also be used to rule out the presence of a cancer, thereby saving subjects from having to undergo invasive testing such as a biopsy, which can be painful and stressful, or may risk damaging the subject.
[0102] A multi-omics approach may unlock the ability to detect a disease at an early stage of development of the disease and may improve accuracy of detection of the disease. Fig. 1A shows some aspects of a multi-omics approach to early disease detection that may combine genomic DNA or DNA methylation information (an example of what may be a generally static indicator of risk) with molecular phenotype information coming from proteomics or metabolomics, which may be more dynamic indicators of function. Fig. 1C also shows some aspects that may be included in a multi -omic method and includes some examples of disease states that may be detected or assessed. Fig. IB shows an example of integration of multiple omic datatypes. Any aspect of these figures may be used in a method described herein.
[0103] Fig. 2A illustrates a non-limiting example of a method for predicting whether a subject has a disease such as cancer or is at risk of developing the disease. Analysis may include obtaining a biofluid sample from a subject (200). The sample may be assayed or analyzed. The biofluid sample can be any one of or any combination of the biofluids described herein. The sample can be either: directly analyzed to generate data(202) such as proteomic data; or contacted with particle described herein to obtain adsorbed biomolecules(203) prior to the analysis of 202. After obtaining the data from the analysis of 202, additional analysis (203) can be performed from the sample obtained from 200 or 201 to obtain additional data sets such as transcriptomic data, genomic data, metabolomic data, or a combination thereof. The data or data sets obtained from the analysis of 202 or 203 can be used to generate a classifier (205). The classifier can be applied to identify a likelihood of the subject having or at risk of having the disease. The generation or application of the classifier can be further repeated or refined to improve the analysis. Fig. 2B further illustrates some details that may be used in the methods described herein. Any of the aspects of Fig. 2A or Fig. 2B may be used in a method described herein such as a classification method.
[0104] Furthermore, an analysis as illustrated in Fig. 2A or Fig. 2B can be applied before or during a procedure at any step included in Fig. 3A. For example, an evaluation or analysis may be completed early on in a diseased patient’s journey before, shortly after, or as part of an invasive workup. It is useful to screen high-risk patients before performing an invasive procedure such as a biopsy or invasive treatment. Generally, an opportunity where a method described herein may be useful, may be in screening high risk patients for early detection of the disease. The methods described herein may be used for such detection with greater accuracy and convenience than other methods. In Fig. 3A, the non-invasive work-up may include medical imaging, or the invasive work-up may include obtaining a biopsy. The biopsy may be of a suspected tumor. Similar patient journeys are shown for pancreatic cancer, liver cancer, and colon cancer in Fig. 3B, Fig. 3C and Fig. 3D. An evaluation or analysis may be completed at or before any point in Fig. 3B, Fig. 3C, or Fig. 3D.
[0105] In some aspects, the cancer to be detected by the methods described herein can be pancreatic cancer. The pancreatic cancer may be early-stage pancreatic cancer. In other aspects, the pancreatic cancer may be late-stage pancreatic cancer. Non-invasively obtained samples can be used for cancer diagnosis by generating data and identifying patterns in the data that associate with the cancer such as pancreatic cancer. Diagnosis of cancer may be improved by obtaining proteomic data. Diagnosis of cancer may be improved by combining multiple types of data (e.g., multiple data sets) into the analysis. For example, combining multiple data types comprising proteomic, transcriptomic, genomic, metabolomic, or a combination thereof may improve the accuracy of prediction of whether a subject has the cancer. In some aspects, the methodsdescribed herein include generating or obtaining data and using the data to predict whether a subject has or does not have a cancer. Various ways of combining or analyzing the data are described, and the uses of the data for cancer assessment are further elaborated.
[0106] In certain aspects, the method of detecting a cancer may comprise additional screening or diagnosing methods such as a computed tomography (CT) scan indicative of pancreatic cancer, a magnetic resonance imaging (MRI) scan indicative of pancreatic cancer, a positron emission tomography (PET) scan indicative of pancreatic cancer, an ultrasound indicative of pancreatic cancer, a cholangiopancreatography indicative of pancreatic cancer, an angiography indicative of pancreatic cancer, a liver function test (LFT) indicative of pancreatic cancer, an elevated carcinoembryonic antigen (CEA) level relative to a control or baseline measurement, an elevated carbohydrate antigen (CA) 19-9 level relative to a control or baseline measurement, or a combination thereof. In some aspects, the method of detecting pancreatic cancer may comprise identifying a symptom of a subject such as jaundice, abdominal pain, gallbladder or liver enlargement, a blood clot, digestion problems, or depression, or a combination thereof.
[0107] Non-invasively obtained samples can be used for disease diagnosis by generating omic data and identifying patterns in the omic data that associate with a disease. Diagnosis of diseases may be improved by combining multiple types of data (e.g., multiple data sets such as omic data sets) into the analysis. For example, combining multiple data types may improve the accuracy of prediction of whether a subject has or does not have a particular disease. Combined data may be more accurate than individual data sets if the individual data sets err independently or do not overlap completely. The methods described herein include generating or obtaining multi -omics data, and using the multi -omics data to make a prediction about whether a subject has or does not have a disease. Various ways of combining or analyzing multi -omics data are described. Uses of the multi-omics data and disease assessment are further elaborated.
[0108] Some methods may be used to classify a lung nodule. Lung nodules can be either benign or malignant. Malignant lung nodules can rapidly progress into lung cancer, a common and deadly cancer. Improved identification of malignant and benign lung nodules is needed. On one hand, early diagnosis of a malignant lung nodule can lead to early treatment regimen and a more favorable prognosis for a subject having the malignant lung nodule. On the other hand, non-invasive diagnosis of a benign or non-malignant lung nodule can help in the avoidance of obtaining a lung biopsy, which can be costly and invasive, and thus also be more favorable for a subject having a lung nodule that is not malignant.
[0109] However, there has been little progress in the development of useful clinical tests for diagnosing and deciphering lung nodules as benign or malignant. Imaging methods often lead to high degree of misdiagnose (e.g., false positive) rates. Smaller nodules are usually not detected by these imaging methods. Other non- invasive methods such as screening for biomarkers also have limitations. Proteins in plasma may be a useful biomarker discovery matrix given plasma’s contact with many tissues in the body. However, plasma proteins can be problematic due to several factors including a wide range of concentration (e.g., 10-orders of magnitude). Complex biochemical workflows have attempted to circumvent these challenges but may not be practical for discovery studies of sufficient size to ensure validation and replication. Alternatively, biomarker studies have been limited to evaluating or re-evaluating existing markers without substantiveimprovement in clinical performance. Accordingly, there remains a need for methods for diagnosing or screening for the presence of benign or malignant lung nodule based on the analysis of biomarkers in a biofluid sample. The methods described herein may address this need.
[0110] Disclosed herein are methods that include obtaining biomolecule data. The biomolecule data may include multi-omics data. The biomolecule data may comprise RNA data. The method may include generating or receiving the data, and then using a classifier to make an evaluation. The evaluation may include applying a classifier, identifying a disease, ruling out a presence of a disease, predicting a likelihood of a disease, or selecting a treatment for the disease.[oni] Disclosed herein are methods that include assessing a biological state using multi -omic data. Disclosed herein are methods that include assessing a biological state comprising using a combination of protein makers, genetics, and metabolic markers. The biological state may include a disease such as cancer. The biological state may include a healthy state. The biological state may include a state free of the disease.
[0112] Disclosed herein are methods that include obtaining a multi-omics database comprising multi-omics data generated from biofluid samples. The samples may be of a population having varying disease states and patient characteristics. Some aspects include querying the multi-omics database. The querying may be to identify a biomarker or set of biomarkers capable of distinguishing individuals of the population as having a first disease state or patient characteristic from other individuals of the population as having a second disease state or patient characteristic. The multi-omics data may include a combination of comprises proteomics, metabolomics, lipidomics, transcriptomics, fragmentomics, methylomics, or genomics.
[0113] Disclosed herein are methods that include obtaining multi-omics data from one or more biofluid samples of a subject identified as having a lung nodule; and applying a classifier to the multi-omics data to evaluate the lung nodule. The evaluation may be to determine whether the lung nodule is cancerous or non- cancerous. The evaluation may be to rule out lung cancer.
[0114] Disclosed herein are methods that include obtaining multi-omics data from one or more biofluid samples of a subject suspected of having pancreatic cancer; and applying a classifier to the multi-omics data to evaluate the subject. The evaluation may include determining or indicating a likelihood of the subject having the pancreatic cancer or not.
[0115] Some aspects relate to sample preparation. Some aspects include preparing a sample for a method disclosed herein. Some methods include preparing multiple samples.Diseases
[0116] The methods described herein may be used to evaluate a disease state. The methods described herein may be used to predict or identify a disease state. A disease state may include a disease or disorder such as cancer. Examples of cancer include lung cancer, colon cancer, pancreatic cancer, liver cancer, ovarian cancer, breast cancer, prostate cancer, melanoma, bladder cancer, lymphoma, leukemia, renal cancer, or uterine cancer. In some aspects, the cancer is breast cancer. A disease may include a disorder. A disease state may include having a comorbidity related to a disease or disorder. A reference to whether a subject hasa disease state or not may include the subject being healthy. A healthy state may exclude a disease state. For example, a healthy state may exclude having cancer. A disease state may exclude being healthy.
[0117] The methods may be useful for cancer diagnosis. The methods may be useful for cancer screening. The method may be useful for cancer treatment. The method may include assaying proteins in a biofluid sample obtained from a subject having or suspected of having a nodule such as a lung nodule to obtain protein measurements. The method may include applying a classifier to the protein measurements, thereby identifying the protein measurements as indicative of the lung nodule being cancerous or non-cancerous. In some cases, the classifier is generated using proteomic data obtained by contacting training samples with particles such that the particles adsorb proteins in the training samples, and assaying the proteins adsorbed to the particles. Some aspects include obtaining of receiving the biofluid sample of the subject. The methods may include assaying RNA in a biofluid sample obtained from a subject having or suspected of having a cancer. The methods may include applying a classifier to the RNA, thereby identifying the RNA as indicative of having cancer or not. In some cases, the classifier is generated using RNA data obtained from training samples by assaying the RNA from training samples. Some aspects include obtaining or receiving the biofluid sample of the subject.
[0118] In some aspects, the cancer to be detected by the methods described herein can be pancreatic cancer, liver cancer, ovarian cancer, or colon cancer. Diagnosis of cancer may be improved by obtaining proteomic data or other omic data (such as lipidomic data). Diagnosis of cancer may be improved by combining multiple types of data (e.g., multiple data sets) into the analysis. For example, combining multiple data types comprising proteomic, transcriptomic, genomic, metabolomic, or a combination thereof may improve the accuracy of prediction of whether a subject has the cancer. In some aspects, the methods described herein include generating or obtaining data and using the data to predict whether a subject has or does not have a cancer. The method may include discriminating between cancer types (e.g., liver cancer vs. ovarian cancer). Various ways of combining or analyzing the data are described, and the uses of the data for cancer assessment are further elaborated.
[0119] The cancer may be at an early stage or a late stage. An example of an early stage of cancer may include stage I. An early stage may include stage I or II. An early stage may include stage I, II, or III. An example of late-stage cancer may include stage 4.
[0120] The cancer may include pancreatic cancer. The pancreatic cancer may be early-stage pancreatic cancer. In other aspects, the pancreatic cancer may be late-stage pancreatic cancer. Non-invasively obtained samples can be used for cancer diagnosis by generating data and identifying patterns in the data that associate with the cancer such as pancreatic cancer. In certain aspects, the method of detecting a cancer may comprise additional screening or diagnosing methods such as a computed tomography (CT) scan indicative of pancreatic cancer, a magnetic resonance imaging (MRI) scan indicative of pancreatic cancer, a positron emission tomography (PET) scan indicative of pancreatic cancer, an ultrasound indicative of pancreatic cancer, a cholangiopancreatography indicative of pancreatic cancer, an angiography indicative of pancreatic cancer, a liver function test (LFT) indicative of pancreatic cancer, an elevated carcinoembryonic antigen (CEA) level relative to a control or baseline measurement, an elevated carbohydrate antigen (CA) 19-9 levelrelative to a control or baseline measurement, or a combination thereof. In some aspects, the method of detecting pancreatic cancer may comprise identifying a symptom of a subject such as jaundice, abdominal pain, gallbladder or liver enlargement, a blood clot, digestion problems, or depression, or a combination thereof. Any of these aspects may be used in identifying a subject at risk of having pancreatic cancer.
[0121] The cancer may include liver cancer. In some aspects, the cancer to be detected by the methods described herein can be liver cancer. The liver cancer may be early-stage liver cancer. In other aspects, the liver cancer may be late-stage liver cancer. In some cases, the liver cancer can be stage I, II, III, or IV liver cancer. In some instances, the stage of the liver cancer is unknown. Non-invasively obtained samples can be used for cancer diagnosis by generating data and identifying patterns in the data that associate with the cancer such as liver cancer. In certain aspects, the method of detecting a cancer may comprise additional screening or diagnosing methods such as a dynamic contrast computed tomography (CT) scan indicative of liver cancer, having a magnetic resonance imaging (MRI) scan indicative of liver cancer, having a liver function test (LFT) indicative of liver cancer, having an elevated bilirubin level relative to a control or baseline measurement, having an elevated aminotransferase level relative to a control or baseline measurement, having an elevated alkaline phosphatase level relative to a control or baseline measurement, having hypoalbuminemia, having an elevated prothrombin time relative to a control or baseline measurement, having an elevated alpha-fetoprotein level relative to a control or baseline measurement, or having a liver nodule, or a combination thereof. In some aspects, the method of detecting a cancer may comprise identifying symptoms of a subject such as abdominal discomfort, pain, and tenderness, jaundice, white, chalky stools, nausea, vomiting, bruising, or bleeding easily, weakness, or fatigue, or a combination thereof. Any of these aspects may be used in identifying a subject at risk of having liver cancer.
[0122] The cancer may include ovarian cancer. In some aspects, the cancer to be detected by the methods described herein can be ovarian cancer. The ovarian cancer may be early-stage ovarian cancer. In other aspects, the ovarian cancer may be late-stage ovarian cancer. In some cases, the stage of the ovarian cancer may be unknown. In some aspects, the stage of the ovarian cancer may be stage I, II, III, or IV. Non- invasively obtained samples can be used for cancer diagnosis by generating data and identifying patterns in the data that associate with the cancer such as ovarian cancer. In certain aspects, the method of detecting a cancer may comprise additional screening or diagnosing methods such as a computed tomography (CT) scan indicative of ovarian cancer, having a magnetic resonance imaging (MRI) scan indicative of ovarian cancer, having a positron emission tomography (PET) scan indicative of ovarian cancer, having a transvaginal ultrasound indicative of ovarian cancer, having an elevated cancer antigen (CA)-125 level relative to a control or baseline measurement, or having an ovarian cyst, or a combination thereof. In some aspects, the method of detecting cancer may comprise identifying a symptom in a subject such as a heavy feeling in the pelvis, pain in the lower abdomen, bleeding from the vagina, weight gain, weight loss, abnormal periods, unexplained back pain that worsens over time, an increase in urination, gas, nausea, vomiting, or loss of appetite, or a combination thereof. Any of these aspects may be used in identifying a subject at risk of having ovarian cancer.
[0123] The cancer may include colon cancer or colorectal cancer (CRC). In some aspects, the cancer to be detected by the methods described herein can be colon cancer. The colon cancer may be early-stage colon cancer. In other aspects, the colon cancer may be late-stage colon cancer. Non-invasively obtained samples can be used for cancer diagnosis by generating data and identifying patterns in the data that associate with the cancer such as colon cancer. Diagnosis of cancer may be improved by obtaining proteomic data. In certain aspects, the method of detecting a cancer may comprise additional screening or diagnosing methods such as computed tomography (CT) scan for indication of colon cancer, a liver function test (LFT) for indication of colon cancer, measuring carcinoembryonic antigen (CEA) level relative to a control or baseline measurement, determining blood in a stool, performing a fecal immunochemical test (FIT), or a combination thereof. Any of these aspects may be used in identifying a subject at risk of having a colon cancer. For example, a subject identified as at risk of having colon cancer may be identified as at risk by one of these methods. The non-invasive methods described herein may save a patient who does not have colon cancer from undergoing further invasive testing or treatment procedures such as having a colonoscopy or cancer biopsy taken, or from undergoing a colon cancer treatment procedure. On the other hand, the non-invasive methods described herein may be used to identify a person who likely has colon cancer and confirm that the patient should undergo further testing (e.g., invasive testing) or treatment procedures. Colon cancer may be an example of colorectal cancer (CRC). References or teachings herein related to colon cancer may be applied to CRC, or vice versa.
[0124] The cancer may include lung cancer. An example of lung cancer is non-small cell lung cancer (NSCLC). An example of lung cancer is small cell lung cancer. Disclosed are lung nodule diagnosis methods. The method may be useful for diagnosing, treating, or screening a patient with an identified lung nodule from a computed tomography (CT) scan who has not had a lung biopsy. The method may be useful for informing a medical practitioner regarding a probability of the lung nodule being benign or malignant. With test results from such a method, a medical practitioner may avoid unnecessarily biopsying the patient. For example, the method may be used as a rule-out test. With test results from such a method, a medical practitioner may identify a subject who should be biopsied. For example, the method may be used as a rulein test.
[0125] Disclosed are diagnosis methods for identifying CT imaging candidates. The method may be useful for diagnosing, treating, or screening a patient who may be a CT imaging candidate. The method may be useful for a higher-risk patient (e.g., as defined by USPSTF or another body) who is a candidate for but has not received a CT scan for lung cancer screening. The method may inform a medical practitioner of a probability of the patient having a lung cancer. The method may therefore inform the medical practitioner of an urgency or need to obtain a CT scan of the patient’s lungs. Such a method may be useful for high-risk patients such as patients who are non-compliant to other CT screening methods. The method may improve selection or compliance of a patient for CT imaging. The method may improve selection or compliance of a patient for biopsy.
[0126] Disclosed are methods for recurrent monitoring. The method may be useful for monitoring a patient with a potentially resectable lung cancer. The method may be useful for monitoring a patient that has a post-surgical therapy intervention. The method may be useful for monitoring a patient that has an adjuvant chemotherapy or radiotherapy intervention. The method may be useful for detecting cancer recurrence before a CT scan or other medical imaging. The method may be useful for surveillance testing for recurrence. The method may be tailored or developed in partnership with a patient treatment method.
[0127] Described herein is a method, comprising: assaying proteins in a biofluid sample obtained from a subject having or suspected of having a lung nodule to obtain protein measurements; and applying a classifier to the protein measurements, thereby identifying the protein measurements as indicative of the lung nodule being cancerous or non-cancerous, wherein the classifier is generated using proteomic data obtained by contacting training samples with particles such that the particles adsorb proteins in the training samples, and assaying the proteins adsorbed to the particles. The method may be useful for cancer diagnosis or screening.
[0128] Described herein is a method, comprising obtaining a biofluid sample of a subject having a lung nodule; contacting the biofluid sample with particles such that the particles adsorb biomolecules comprising proteins to the particles; assaying the biomolecules adsorbed to the particles to generate proteomic data; and classifying the proteomic data as indicative of the lung nodule being cancerous or non-cancerous. The method may be useful for cancer diagnosis or screening.
[0129] Described herein are methods for determining lung nodule-related state in a sample obtained from a subject. In some embodiments, the lung nodule-related state includes the presence or absence of a lung nodule in the subject. In some embodiments, the lung nodule-related state includes determining whether the lung nodule is benign or malignant. In some embodiments, the method comprises screening for lung nodule- related state by assaying biomarkers in the sample obtained from the subject. In some embodiments, the biomarkers comprise at least one protein in the sample. In some embodiments, the sample is a biofluid sample. In some embodiments, the biofluid sample is contacted with a particle described herein to adsorb proteins in the biofluid sample. In some embodiments, the method comprises obtaining proteins measurements of the proteins in the sample. In some embodiments, the method comprises applying a classifier to the protein measurements, thereby identifying the protein measurements as indicative of the lung nodule being cancerous or non-cancerous. In some embodiments, the classifier is generated using proteomic data obtained by contacting training samples with particles such that the particles adsorb proteins in the training samples. The adsorbed proteins can then be assayed by the methods described herein. In some embodiments, the subject is suspected of having a lung nodule or is identified as having the lung nodule by imaging methods described herein. In some embodiments, a report is generated based on the identification of the protein measurements as indicative of the lung nodule being cancerous or non-cancerous. In some embodiments, the report indicates the likelihood or an indication that the lung nodule is cancerous or non- cancerous. In some embodiments, the report indicates that the lung nodule is cancerous. In some embodiments, the report indicates that the lung nodule comprises non-small-cell lung carcinoma (NSCLC). In some embodiments, the method described herein generates a classifier comprising features to indicate the protein measurements as indicative of the lung nodule being cancerous or non-cancerous. In some embodiments, the features comprise control protein measurements, mass spectra, m / z ratios,chromatography results, immunoassay results, or light or fluorescence intensities. In some embodiments, the classifier is trained using any one of the computation or machine-leaning methods described herein. In some embodiments, the method comprises obtaining RNA data and / or RNA measurements of the RNA in the sample. In some embodiments, the methods comprise applying a classifier to the RNA data and / or RNA measurements, thereby identifying the RNA data and / or RNA measurements as indicative of a mass being cancerous or non-cancerous. In some embodiments, the classifier is generated using RNA data obtained from training samples. In some embodiments, the RNA can be assayed by the methods described herein. In some embodiments, the subject is suspected of having a mass or is identified as having a mass by imaging methods described herein. In some embodiments, a report is generated based on the identification of the RNA data and / or RNA measurements as indicative of a mass being cancerous or non-cancerous. In some embodiments, the report indicates the likelihood or an indication that a mass is cancerous or non-cancerous. In some embodiments, the report indicates that a mass is cancerous. In some embodiments, the report indicates that a mass is non-cancerous. In some embodiments, the report indicates that a mass comprises non-small-cell lung carcinoma (NSCLC). In some embodiments, the methods described herein generates a classifier comprising features to indicate the RNA data and / or RNA measurements as indicative of a mass being cancerous or non-cancerous. In some embodiments, the RNA features comprise control RNA data and / or measurements. In some embodiments, the classifier is trained using any one of the computation or machine -leaning methods described herein.
[0130] Described herein, in some embodiments, are methods for recommending a lung cancer treatment for the subject when the subject is determined to have malignant lung nodule based on the analysis of the protein or RNA measurements described herein. In some embodiments, the protein or RNA measurements are classified as indicative of the lung nodule being cancerous.
[0131] Disclosed herein, in some aspects, are methods useful for diagnosing, screening, or treating a subject. Some aspects include assaying proteins in a biofluid sample obtained from a subject suspected of having a lung nodule to obtain protein measurements. Some aspects include applying a classifier to the protein measurements. Some aspects include identifying the protein measurements as indicative of the subject having the lung nodule. In some aspects, the classifier is generated using proteomic data obtained by contacting training samples with particles such that the particles adsorb proteins in the training samples and assaying the proteins adsorbed to the particles. Some aspects include assaying RNA in a biofluid sample obtained from a subject suspected of having a mass to obtain RNA measurements. Some aspects include applying a classifier to the RNA data and / or RNA measurements. Some aspects include identifying the RNA data and / or RNA measurements as indicative of the subject having a mass. In some aspects, the classifier is generated using RNA data obtained by contacting training samples and assaying the RNA.
[0132] Disclosed herein, in some aspects, are methods useful for diagnosing, screening, or treating a subject. Some aspects include assaying proteins in a biofluid sample obtained from a subject suspected of having a lung cancer to obtain protein measurements. Some aspects include applying a classifier to the protein measurements. Some aspects include identifying the protein measurements as indicative of the subject having the lung cancer. In some aspects, the classifier is generated using proteomic data obtained bycontacting training samples with particles such that the particles adsorb proteins in the training samples and assaying the proteins adsorbed to the particles. Some aspects include assaying RNA in a biofluid sample obtained from a subject suspected of having a lung cancer to obtain RNA data and / or RNA measurements. Some aspects include applying a classifier to the RNA data and / or RNA measurements. Some aspects include identifying the RNA data and / or RNA measurements as indicative of the subject having the lung cancer. In some aspects, the classifier is generated using RNA data and / or RNA measurements obtained by contacting training samples and assaying the RNA.
[0133] Disclosed herein, in some aspects, are methods useful for diagnosing, screening, or treating a subject. Some aspects include obtaining a biofluid sample of a subject suspected of having a lung cancer. Some aspects include contacting the biofluid sample with particles such that the particles adsorb biomolecules comprising proteins to the particles. Some aspects include assaying the biomolecules adsorbed to the particles to generate proteomic data. Some aspects include, based on the proteomic data, classifying the proteomic data as indicative of the subject having the lung cancer or as not indicative of the subject having the lung cancer.
[0134] Disclosed herein, in some aspects, are methods useful for monitoring a subject. Some aspects include obtaining a biofluid sample of a subject at risk of a lung cancer recurrence. Some aspects include contacting the biofluid sample with particles such that the particles adsorb biomolecules comprising proteins to the particles. Some aspects include assaying the biomolecules adsorbed to the particles to generate proteomic data. Some aspects include, based on the proteomic data, classifying the proteomic data as indicative of the subject having the lung cancer recurrence or as not indicative of the subject having the lung cancer recurrence. In some aspects, the subject has received a lung cancer treatment such as chemotherapy, radiotherapy, or surgery. In some aspects, the cancer may be resectable. In some aspects, the lung cancer comprises NSCLC.
[0135] In some cases, a lung nodule is described as malignant or cancerous. The terms, malignant and cancerous may be used interchangeably. A malignant or cancerous lung nodule may be referred to as a lung cancer, or vice versa. In some cases, a lung nodule is described as benign or non-cancerous. The terms, benign and non-cancerous may be used interchangeably.Samples & Subjects
[0136] Some aspects relate to a subject. For example, a subject may be evaluated, or a sample from a subject may be evaluated using methods described herein. Multi -omics data may be generated from a sample of a subject.
[0137] The methods described herein may be used to identify a subject as likely or at risk to have a disease such as cancer. The subject may have lung cancer, pancreatic cancer, liver cancer, ovarian cancer, or colon cancer. The cancer may include adenocarcinoma, for example pancreatic adenocarcinoma. The subject may have the cancer. The subject may not have the cancer. The subject may have the pancreatic cancer, liver cancer, ovarian cancer, or colon cancer. The subject may not have the pancreatic cancer, liver cancer, ovarian cancer, or colon cancer. The subject may be at risk of having pancreatic cancer, liver cancer, ovariancancer, or colon cancer. The subject may have a mass (e.g., nodule or cyst) in the pancreas. The subject may have a mass (e.g., nodule) in the liver. The liver cancer may include a hepatocellular carcinoma (HCC). The liver cancer may include stage I, stage II, stage III, or stage IV liver cancer. The subject may have a mass (e.g., nodule or cyst) in one or both ovaries. The ovarian cancer may include stage I, stage II, stage III, or stage IV ovarian cancer. The ovarian cancer may include stage III ovarian cancer. The ovarian cancer may include stage IV ovarian cancer. The subject may have a mass (e.g., nodule) in the colon. The subject may have a lung nodule, cancer. The subject may be at risk of having breast cancer. The subject may have a mass (e.g., nodule or cyst) in the breast.
[0138] A sample may be obtained from the subject for purposes of identifying a cancer in the subject. The subject may be suspected of having the cancer or as not having the cancer. The method may be used to confirm or refute the suspected cancer.
[0139] Data described herein may be generated from a sample of a subject. The sample may be a biofluid sample or a mass sample (e.g., an abnormal growth biopsied from the subject). Examples of biofluids include blood, serum, or plasma. The sample may include a blood sample. The sample may include a serum sample. The sample may include a plasma sample. One or more biofluid samples may comprise a blood, serum, or plasma sample. Other examples of biofluids include urine, tears, semen, milk, vaginal fluid, mucus, saliva, sweat, or cell homogenate.
[0140] A sample may be obtained from the subject for purposes of identifying a disease state in the subject. The subject may be suspected of having the disease state or as not having the disease state. The method may be used to confirm or refute the suspected disease state. In some aspects, a sample from the subject is used in determining whether a mass, nodule (e.g. a lung nodule), or cyst is cancerous or non-cancerous.
[0141] A biofluid sample may be obtained from a subject. For example, a blood, serum, or plasma sample may be obtained from a subject by a blood draw. Other ways of obtaining biofluid samples include aspiration or swabbing.
[0142] The biofluid sample may be cell-free or substantially cell-free. To obtain a cell-free or substantially cell-free biofluid sample, a biofluid may undergo a sample preparation method such as centrifugation and pellet removal.
[0143] A non-biofluid sample may be obtained from a subject or patient. For example, a sample may include a tissue sample. Some examples of organs or tissues that may be sampled include lung, colon, pancreatic, liver, breast, or ovarian tissue. The sample may include a mass taken from the organ or tissue of the subject. The mass may be suspected of being cancerous. The mass may include a nodule (e.g., a colon nodule or liver nodule). The mass may include a cyst (e.g., an ovarian cyst). The nodule or cyst may be identified by a physician as at a high risk or low risk of being cancerous prior to performing the methods described herein. The mass may be biopsied, for example by a needle biopsy procedure. A needle biopsy procedure may include insertion of a thin needle through the subject’s abdomen and into the liver to obtain a tissue sample, which may then be examined under a microscope for signs of cancer. The sample may include a cell sample. The sample may include a homogenate of a cell or tissue. The sample may include a supernatant of a centrifuged homogenate of a cell or tissue.
[0144] The sample may include lung tissue. The sample may include colon tissue. The sample may include pancreatic tissue. The sample may include liver tissue. The sample may include breast tissue. The sample may include ovarian tissue. The tissue may be cancerous. The tissue may be non-cancerous. The tissue may be suspected of being cancerous. The tissue may be malignant. The tissue may be non-malignant. The tissue may be suspected of being malignant.
[0145] The sample (e.g., biofluid or tissue sample) can be obtained from the subject during any phase of a screening procedure, such as before, during, or after a stage shown in Fig. 3A. The sample can be obtained before or during a stage where the subject is a candidate for a biopsy, pancreatoscopy, or colonoscopy, for early detection of a disease. The sample can be obtained before or during a non-invasive work-up, an invasive work-up, treatment, a monitoring stage.
[0146] Data may be generated from a single sample, or from multiple samples. Data from multiple samples may be obtained from the same subject. In some cases, different data types are obtained from samples collected differently or in separate containers. A sample may be collected in a container that includes one or more reagents such as a preservation reagent or a biomolecule isolation reagent. Some examples of reagents include heparin, ethylenediaminetetraacetic acid (EDTA), citrate, an anti-lysis agent, or a combination of reagents. Samples from a subject may be collected in multiple containers that include different reagents, such as for preserving or isolating separate types of biomolecules. A sample may be collected in a container that does not include any reagent in the container. The samples may be collected at the same time (e.g., same hour or day), or at different times. A sample may be frozen, refrigerated, heated, or kept at room temperature.
[0147] The methods described herein may be used to identify a subject as likely to have a disease state or not. A disease state may include cancer, including pancreatic cancer, liver cancer, ovarian cancer, or colon cancer. Some aspects of the present disclosure include identifying whether a lung nodule of a subject is cancerous or non-cancerous. The lung nodule may be in the subject’s lung. The subject may be identified as having the lung nodule. In some aspects, the subject has multiple lung nodules. The subject may have a lung cancer. The subject may be at risk of a lung cancer. The subject may have a lung complication. The subject may have a comorbidity described herein. The subject may have trouble breathing. The subject may have fluid in the lungs.
[0148] In some cases, the subject is monitored. For example, information about a likelihood of the subject having a disease state may be used to determine to monitor a subject without providing a treatment to the subject. In other circumstances, the subject may be monitored while receiving treatment to see if a disease state in the subject improves. In some aspects, a subject having a lung nodule may be monitored to determine progression of the lung nodule. A lung nodule in a subject may be monitored. A subject may be treated as described herein.
[0149] The subject may be a vertebrate. The subject may be a mammal. The mammal may include a rat, mouse, gerbil, guinea pig, or hamster. The mammal may include a fox, bear, dog, monkey, cow, pig, or sheep. The subject may be a primate. The primate may include an ape or monkey. The primate may include a chimpanzee, a lemur, a bonobo, an orangutan, or a baboon. The subject may be a human. The subject maybe an adult (e.g. at least 18-years-old). The subject may be male. The subject may be female. The subject may have a disease state. For example, the subject may have a disease or disorder, a comorbidity of a disease or disorder, or may be healthy.
[0150] The methods described herein may include use of a sample such as a biological sample. For example, a method may include determining one or more biomarker measurements in the sample. The biological sample may be from a subject such as a subject with a lung nodule. The biological sample may include a blood sample that has had red blood cells removed. For example, the biological sample may comprise a plasma sample. The biological sample may comprise a serum sample. The biological sample may comprise blood or a blood constituent. The biological sample may comprise a blood sample. A sample described or used herein may be from a subject described herein, such as a subject with an identified lung nodule.
[0151] Samples consistent with the methods disclosed herein of assessing for the presence or absence of one or more biomarkers associated with presence or malignancy state of lung nodule. The subject may be a human or a non-human animal. Biological samples may be a biofluid. For example, the biofluid may be plasma, serum, CSF, urine, tear, cell lysates, tissue lysates, cell homogenates, tissue homogenates, nipple aspirates, fecal samples, synovial fluid and whole blood, or saliva. Samples can also be non-biological samples, such as water, milk, solvents, or anything homogenized into a fluidic state. Said biological samples can contain a plurality of proteins or proteomic data, which may be analyzed after adsorption of proteins to the surface of the various particle types in a panel and subsequent digestion of protein coronas. Proteomic data can comprise nucleic acids, peptides, or proteins. Any of the samples herein can contain a number of different analytes, which can be analyzed using the methods disclosed herein. The analytes can be proteins, peptides, small molecules, nucleic acids, metabolites, lipids, or any molecule that could potentially bind or interact with the surface of a particle type.
[0152] The sample may be a biofluid. A biological sample may comprise a biofluid sample such as cerebrospinal fluid (CSF), synovial fluid (SF), urine, plasma, serum, tear, crevicular fluid, semen, whole blood, milk, nipple aspirate, ductal lavage, vaginal fluid, nasal fluid, ear fluid, gastric fluid, pancreatic fluid, trabecular fluid, lung lavage, prostatic fluid, sputum, fecal matter, bronchial lavage, fluid from swabbing, bronchial aspirant, sweat, or saliva. A biofluid may be a fluidized solid, for example a tissue homogenate, or a fluid extracted from a biological sample. A biological sample may be, for example, a tissue sample or a fine needle aspiration (FNA) sample. A biological sample may be a cell culture sample. For example, a sample that may be used in the methods disclosed herein can either include cells grow in cell culture or can include acellular material taken from cell cultures. A biofluid may be a fluidized biological sample. For example, a biofluid may be a fluidized cell culture extract. A sample may be extracted from a fluid sample, or a sample may be extracted from a solid sample. For example, a sample may comprise gaseous molecules extracted from a fluidized solid (e.g., a volatile organic compound). In some aspects, the biofluid comprises blood, plasma, or serum.
[0153] A method consistent with the present disclosure may comprise collecting (e.g., isolating, enriching, or purifying) a species from biological sample. The species may be a biomolecule (e.g., a protein), abiomacromolecular structure (e.g., a peptide aggregate or a ribosome), a cell, or tissue. The species may be selectively collected from the biological sample. For example, a method may comprise isolating cancer cells from tissue (e.g., as a tissue biopsy) or from a biofluid (e.g., as a liquid biopsy) such as whole blood, plasma, or a buffy coat. The method may include a sample without cancer cells. The species may be treated prior to analysis. For example, a protein may be reduced and degraded, a nucleic acid may be separated from histones, or a cell may be lysed.
[0154] The biological samples may be obtained or derived from a human subject. The biological samples may be stored in a variety of storage conditions before processing, such as different temperatures (e.g., at room temperature, under refrigeration or freezer conditions, at 25°C, at 4°C, at -18°C, -20°C, or at -80°C) or different suspensions (e.g., EDTA collection tubes, cell -free RNA collection tubes, or cell-free DNA collection tubes).
[0155] In some cases, a sample may be depleted prior to biomarker analysis. A sample may be depleted using a commercially available kit. For example, a kit that may be used to deplete a sample may be a spin column-based depletion kit, an albumin depletion kit, an immunodepletion kit, or an abundant protein depletion kit. Non-limiting examples of kits that may be used for sample depletion include a PureProteome™ Human Albumin / Immunoglobulin depletion kit (EMD Millipore Sigma), a ProteoPrep® Immunoaffinity Albumin & IgG Depletion Kit (Millipore Sigma), a Seppro® Protein Depletion kit (Millipore Sigma), Top 12 Abundant Protein Depletion Spin Columns (Pierce), or a Proteome Purify™ Immunodepletion Kit (R&D Systems). Depletion may remove a high concentration biomolecule from a sample. For example, a method may comprise removing albumin from a plasma sample prior to low concentration biomarker analysis. The sample may include depleted plasma.Data Generation and Use
[0156] The methods disclosed herein may include obtaining data such as multi -omics data generated from one or more biofluid samples collected from a subject. The data may include biomolecule measurements such as protein measurements, transcript measurements, genetic material measurements, or metabolite measurements. Omic data may include any of the following: proteomic data, genomic data, transcriptomic data, or metabolomic data. This section includes some ways of generating each of these types of omic data. Methods of generating or analyzing omic data may also be applied to methods of generating or analyzing individual biomolecules or sets of biomolecules. Other types of omic data may also be generated. Descriptions of generating or analyzing omic data may be applied to methods of generating or analyzing individual biomolecules or sets of biomolecules that do not necessarily include omic data. Aspects described in relation to biomolecule data may be relevant to biomolecule measurements, or vice versa. The data may be labeled or identified as indicative of a disease or as not indicative of a disease. The data may be labeled or identified as indicative of pancreatic cancer, liver cancer, ovarian cancer, or colon cancer or as not indicative of pancreatic cancer, liver cancer, ovarian cancer, or colon cancer. The methods described herein may include obtaining the multi-omics measurements such as by performing an assay.
[0157] The methods described herein may include generating or using omic data. Omic data may include data on all biomolecules of a certain type such as proteins, transcripts, genetic material, or metabolites. Omic data may include data on a subset of the biomolecules. For example, omic data may include data on 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 750 or more, 1000 or more, 2500 or more, 5000 or more, 10,000 or more, 25,000 or more, biomolecules of a certain type. The methods described herein may include obtaining measurements of over 10, over 20, over 30, over 40, over 50, over 75, over 100, over 250, over 500, over 750, over 1000, over 1250, over 2500, over 5000, over 7500, over 10,000, over 12,500, over 15,000, over 17,500, over 20,000, over 22,500, or over 25,000 biomolecules of a certain type. The methods described herein may include obtaining measurements of less than 10, less than 20, less than 30, less than 40, less than 50, less than 75, less than 100, less than 250, less than 500, less than 750, less than 1000, less than 1250, less than 2500, less than 5000, less than 7500, less than 10,000, less than 12,500, less than 15,000, less than 17,500, less than 20,000, less than 22,500, or less than 25,000 biomolecules of a certain type. Any of the aforementioned numbers of biomolecules may be measured for each of multiple data types. Multi -omics comprises at least 10 measurements of each of the at least two types of omic data. Multi -omics comprises at least 20 measurements of each of the at least two types of omic data. Multi -omics comprises at least 30 measurements of each of the at least two types of omic data. Multi -omics comprises at least 40 measurements of each of the at least two types of omic data. Multi -omics comprises at least 50 measurements of each of the at least two types of omic data. Multi -omics comprises at least 100 measurements of each of the at least two types of omic data. Multi -omics comprises at least 500 measurements of each of the at least two types of omic data. Multi -omics comprises at least 1000 measurements of each of the at least two types of omic data. Multi -omics comprises at least 10 measurements of each of the at least three types of omic data. Multi -omics comprises at least 20 measurements of each of the at least three types of omic data. Multi -omics comprises at least 30 measurements of each of the at least three types of omic data. Multi -omics comprises at least 40 measurements of each of the at least three types of omic data. Multi -omics comprises at least 50 measurements of each of the at least three types of omic data. Multi -omics comprises at least 100 measurements of each of the at least three types of omic data. Multi -omics comprises at least 500 measurements of each of the at least three types of omic data. Multi -omics comprises at least 1000 measurements of each of the at least three types of omic data. Multi -omics comprises at least 10 measurements of each of the at least four types of omic data. Multi -omics comprises at least 20 measurements of each of the at least four types of omic data. Multi -omics comprises at least 30 measurements of each of the at least four types of omic data. Multi -omics comprises at least 40 measurements of each of the at least four types of omic data. Multi -omics comprises at least 50 measurements of each of the at least four types of omic data. Multi -omics comprises at least 100 measurements of each of the at least four types of omic data. Multi -omics comprises at least 500 measurements of each of the at least four types of omic data. Multi -omics comprises at least 1000 measurements of each of the at least four types of omic data. Multi -omics comprises at least 10measurements of each of the at least five types of omic data. Multi -omics comprises at least 20 measurements of each of the at least five types of omic data. Multi -omics comprises at least 30 measurements of each of the at least five types of omic data. Multi -omics comprises at least 40 measurements of each of the at least five types of omic data. Multi -omics comprises at least 50 measurements of each of the at least five types of omic data. Multi -omics comprises at least 100 measurements of each of the at least five types of omic data. Multi -omics comprises at least 500 measurements of each of the at least five types of omic data. Multi -omics comprises at least 1000 measurements of each of the at least five types of omic data. The data may relate to a presence, absence, or amount of a given biomolecule. Examples of data types may include lipid, protein, peptide, transcript, mRNA, miRNA, intronic RNA sequences, DNA sequence, methylation, or metabolite data.
[0158] Deep proteome coverage is advantageous to a multi-omics approach. New technologies and sample availability address historical challenges to scale proteomics. Some challenges include access to large well- collected, annotated sample cohorts for specific clinical questions, technical challenges associated with plasma proteomics such as reproducibility, throughput and depth of coverage that may limit translation to the clinic, and reproducible measurement and integration of multi-omics datasets providing novel insights into cancer biology.
[0159] The concepts described herein may help address some of these challenges. For example, the use of particles or the inclusion of additional omic types may address these concerns.
[0160] Disclosed herein are methods for multi-omics analysis, “multi-omics(s)” or “multiomic(s)” may include an analytical approach for analyzing biomolecules at a large scale, wherein the data sets are multiple omes, such as proteome, genome, transcriptome, lipidome, and metabolome. Non-limiting examples of multi-omics data may include proteomic data, genomic data, lipidomic data, glycomic data, transcriptomic data, or metabolomics data. “Biomolecule” in “biomolecule corona” can refer to any molecule or biological component that can be produced by, or is present in, a biological organism. Non-limiting examples of biomolecules include proteins (protein corona), polypeptides, polysaccharides, a sugar, a lipid, a lipoprotein, a metabolite, an oligonucleotide, a nucleic acid (DNA, RNA, micro-RNA, plasmid, single stranded nucleic acid, double stranded nucleic acid), metabolome, as well as small molecules such as primary metabolites, secondary metabolites, and other natural products, or any combination thereof. In some embodiments, the biomolecule is selected from the group of proteins, nucleic acids, lipids, and metabolites.
[0161] Some aspects that may be included in a multi-omics strategy include a well-defined disease biobank with multiple sample types optimized for the multi-omics measurements, development and optimization of novel proteomics technologies to increase proteome coverage and throughput without compromising reproducibility, or an unbiased multi-omics platform deploying state-of-the-art instrumentation and advanced machine learning analysis to transform complex early disease detection.Proteomic Data
[0162] The data such as multi-omics data described herein may include protein data or proteomic data.Proteomic data may involve data about proteins, peptides, or proteoforms. This data may include justpeptides or proteins, or a combination of both. An example of a peptide is an amino acid chain. An example of a protein is a peptide or a combination of peptides. For example, a protein may include one, two or more peptides bound together. A protein may be a secreted protein. Proteomic data may include data about various proteoforms. Proteoforms can include different forms of a protein produced from a genome with any variety of sequence variations, splice isoforms, or post-translational modifications. The proteomic data may be generated using an unbiased, non-targeted approach, or may include a specific set of proteins. Aspects described in relation to proteomic data may be relevant to protein data, or vice versa. In an aspect, the disclosed systems and methods provide a classifier generated based on feature information derived from protein data or proteomic data obtained from biological samples. The classifier forms part of a predictive engine for distinguishing groups in a population based on peptide features, which a peptide feature may correspond to a proteomic measure of a peptide (quantification, identification, etc.). The peptide features may include any of the peptides / proteins disclosed herein, such as in Tables 5 and 15 and figures.
[0163] Proteomic data may include information on the presence, absence, or amount of various proteins, peptides. For example, proteomic data may include amounts of proteins. A protein amount may be indicated as a concentration or quantity of proteins, for example a concentration of a protein in a biofluid. A protein amount may be relative to another protein or to another biomolecule. Proteomic data may include information on the presence of proteins or peptides. Proteomic data may include information on the absence of proteins or peptides. Proteomic data may be distinguished by subtype, where each subtype includes a different type of protein, peptide, or proteoform.
[0164] Proteomic data generally includes data on a number of proteins or peptides. For example, proteomic data may include information on the presence, absence, or amount of 1000 or more proteins or peptides. In some cases, proteomic data may include information on the presence, absence, or amount of 5000, 10,000, 20,000, or more peptides, proteins, or proteoforms. Proteomic data may even include up to about 1 million proteoforms. Proteomic data may include a range of proteins, peptides, or proteoforms defined by any of the aforementioned numbers of proteins, peptides, or proteoforms.
[0165] Proteomic data may include protein information such as protein measurements in a biofluid. Some examples of protein biomarkers that may be useful in the methods disclosed herein. The protein measurements may be obtained with the use of particles such as those described herein. Any combination or number of such biomarkers may be included.
[0166] A fragment of any of the proteins may be used. Any of the biomarkers may be useful alone or in combination to assess a lung nodule (for example, to determine a likelihood of the lung nodule being cancerous or not). The protein measurements may be obtained with the use of internal standards.
[0167] Proteomic data may include peptide information such as peptide measurements in a biofluid. Some examples of peptide biomarkers that may be useful in the methods disclosed herein, such as evaluating a cancer such as lung cancer (for example, non-small cell lung cancer). The protein measurements may be obtained with the use of particles such as those described herein. Any combination or number of such biomarkers may be included. In some cases, a biomarker is useful when its feature importance score is above 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, or 0.10. The peptide or protein biomarkers (orcombination of said biomarkers) may be useful for identifying a presence, absence, or likelihood of a cancer described herein.
[0168] Proteomic data may include peptide information such as peptide measurements in a biofluid. Peptide measurements were obtained following the nanoparticle / Proteograph methods as described herein and following data generation on LC-MS, the data fdes were analyzed using data-independent acquisition-neural network (DIA-NN) (version 1.8.1) to determine identification and quantification (ng) of peptides and proteins. Some examples of peptide biomarkers or features that may be useful in the methods disclosed herein, such as using the peptide features disclosed herein to build a classifier, evaluating a cancer such as lung cancer, are included in Table 4 and Table 15 listed below. The protein measurements may be obtained with the use of particles (difference types of nanoparticles (NP1, NP2, NP3, NP4, NP5, etc.)) such as those described herein. Any combination or number of such biomarkers may be included. In some cases, a biomarker is useful when its feature importance score is above 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, or 0. 10. The features may include any of the following peptides (as indicated using a 1-letter amino acid code): HGEYWLGNK (SEQ ID NO: 1), AEPSAATQSHSISSSSFGAEPSAPGGGGSPGACPALGTK (SEQ ID NO: 2), MVGAGISTPSGIPDFR (SEQ ID NO: 3), MNPIVVVHGGGAGPISK (SEQ ID NO: 4), AAAATGTIFTFR (SEQ ID NO: 5), AAAEDVNVTFEDQQK (SEQ ID NO: 6), AAATCFAR (SEQ ID NO: 7), AADYLHVALDLLER (SEQ ID NO: 8), AALPEGLPEASR (SEQ ID NO: 9), AAMPPQIIQFPEDQK (SEQ ID NO: 10), AATVGSLAGQPLQERAQAWGER (SEQ ID NO: 11), AAVEWFDGK (SEQ ID NO: 12), AAVYHHFISDGVR (SEQ ID NO: 13), ACKKDACPINGGWGPWSPWDICSVTCGGGVQK (SEQ ID NO: 14), ACSMPQELPQSPR (SEQ ID NO: 15), ADHAHCCVEMGVDMIEAISLVR (SEQ ID NO: 16), ADMVIEAVFEDLSLK (SEQ ID NO: 17), ADRDQYELLCLDNTR (SEQ ID NO: 18), ADSPMDDFFQCVNGK (SEQ ID NO: 19), AECLNPSQPSR (SEQ ID NO: 20), AEEYILGDFCFLR (SEQ ID NO: 21), AEILLSSSKPVPK (SEQ ID NO: 22), AFCLACPFYGTTPFAGSR (SEQ ID NO: 23), AFDLIVDRPVTLVR (SEQ ID NO: 24), AFLHVPAK (SEQ ID NO: 25), AFQLWSNVTPLTFTK (SEQ ID NO: 26), AFVHWYVGEGMEEGEFSEAREDLAALEK (SEQ ID NO: 27), AGAAPYVQAFDSLLAGPVAEYLK (SEQ ID NO: 28), AGAFCLSEDAGLGISSTASLR (SEQ ID NO: 29), AGHIAWTSSGK (SEQ ID NO: 30), AGLFSQEQYER (SEQ ID NO: 31), AGLYGLPR (SEQ ID NO: 32), AIEMLGGELGSK (SEQ ID NO: 33), AIGPSQTHTIR (SEQ ID NO: 34), CDSSPDSAEDVRK (SEQ ID NO: 35), CLAALASLR (SEQ ID NO: 36), EDITQSAQHALR (SEQ ID NO: 37), EHAVEGDCDFQLLK (SEQ ID NO: 38), ETLLQDFR (SEQ ID NO: 39), GAEVSFGCGVLASGK (SEQ ID NO: 40), GDLDLELVLLCK (SEQ ID NO: 41), GECVPGEQEPEPILIPR (SEQ ID NO: 42), GECWCVNPNTGK (SEQ ID NO: 43), GLFLPEDENLR (SEQ ID NO: 44), GVDCMEVYEYPGYR (SEQ ID NO: 45), GVPWNFTFNVK (SEQ ID NO: 46), IPLDMVAGFNTPLVK (SEQ ID NO: 47), IQPSGGTNINEALLR (SEQ ID NO: 48), LAYVAPTIPR (SEQ ID NO: 49), LFIGGLNTETNEK (SEQ ID NO: 50), NGLTPLHVAVHHNNLDIVK (SEQ ID NO: 51), PYSLHAHGLSYEK (SEQ ID NO: 52), QGALELIK (SEQ ID NO: 53), SQNPVQPIGPQTPK (SEQ ID NO: 54), SVTHANALTVMGK (SEQ ID NO: 55), TEAPSATGQASSLLGGR (SEQ ID NO: 56), TVSSLSEDLESTR (SEQ ID NO: 57), VFDEFKPLVEEPQNLIK (SEQ ID NO: 58),VLLEAGEGLVTITPTTGSDGRPDAR (SEQ ID NO: 59), VVEESELAR (SEQ ID NO: 60), VVLAYEPVWAIGTGK (SEQ ID NO: 61), WFCHVDDDNYVNAR (SEQ ID NO: 62), YEMHELLR (SEQ ID NO: 63), YQSSPAKPDSSFYK (SEQ ID NO: 64), ALILGELEK (SEQ ID NO: 65), CDSSPDSAEDVRK (SEQ ID NO: 66), CTNLEGSFR (SEQ ID NO: 67), DFDFVPPVVR (SEQ ID NO: 68), DTPVLSELPEPVVAR (SEQ ID NO: 69), EGEAVVLPEVEPGLTAR (SEQ ID NO: 70), ELCCLVYTSWQIPQK (SEQ ID NO: 71), EQLGEFYEALDCLR (SEQ ID NO: 72), FGLLDEDGKK (SEQ ID NO: 73), GEALEDFTGPDCR (SEQ ID NO: 74), HQLYIDETVNSNIPTNLR (SEQ ID NO: 75), IDLADFEK (SEQ ID NO: 76), IDQYQGADAVGLEEK (SEQ ID NO: 77), IKPLQSPAEFSVYCDMSDGGGWTVIQR (SEQ ID NO: 78), ILLNPQDK (SEQ ID NO: 79), IRPNDFIPNVI (SEQ ID NO: 80), LGVLTVTDTTPDSMR (SEQ ID NO: 81), LQEAAELEAVELPVPIR (SEQ ID NO: 82), LSGNVLSYTFQVK (SEQ ID NO: 83), LVHVEEPHTETVRK (SEQ ID NO: 84), LYHSEAFTVNFGDTEEAK (SEQ ID NO: 85), NFEDVAFDEK (SEQ ID NO: 86), NIETIINTFHQYSVK (SEQ ID NO: 87), NRDWLTTTFVDDIK (SEQ ID NO: 88), NWGLSVYADKPETTK (SEQ ID NO: 89), PGQAPVLVIYK (SEQ ID NO: 90), QNLLAPQTLPSK (SEQ ID NO: 91), QSSGENCDVVVNTLGK (SEQ ID NO: 92), SSTGPGEQLR (SEQ ID NO: 93), STSESTAALGCLVK (SEQ ID NO: 94), TKLEEHLEGIVNIFHQYSVRK (SEQ ID NO: 95), TLMFGSYLDDEK (SEQ ID NO: 96), TVTAMDVVYALK (SEQ ID NO: 97), VHFLPVVISDNGMPSR (SEQ ID NO: 98), VLDFHNLPDGITK (SEQ ID NO: 99), YDVENCLANK (SEQ ID NO: 100), YNWRENLDR (SEQ ID NO: 101), AALPEGLPEASR (SEQ ID NO: 102), ALNSIIDVYHK (SEQ ID NO: 103), DTHFPICIFCCGCCHR (SEQ ID NO: 104), EDFLEQSEQLFGAK (SEQ ID NO: 105), EFTRPEEIIFLR (SEQ ID NO: 106), FCTALLPVNDR (SEQ ID NO: 107), GDTFSCMVGHEALPLAFTQK (SEQ ID NO: 108), GLEVTITAR (SEQ ID NO: 109), GPGGVWAAK (SEQ ID NO: 110), HLNQGTDEDIYLLGK (SEQ ID NO: 111), LMHSFCAFK (SEQ ID NO: 112), LSCAASGFTFSSYSMNWVR (SEQ ID NO: 113), TSVFHDVDGSVSEYPGSYLTK (SEQ ID NO: 114), VHQYFNVELIQPGAVK (SEQ ID NO: 115), VLTSLAWR (SEQ ID NO: 116), VTIDSSYDIAK (SEQ ID NO: 117), YEINSLIR (SEQ ID NO: 118), AALPEGLPEASR (SEQ ID NO: 119), AAPSAEFSVDR (SEQ ID NO: 120), AMGIMNSFVNDIFER (SEQ ID NO: 121), DGSFSVVITGLR (SEQ ID NO: 122), EIGELYLPK (SEQ ID NO: 290), EMPSVFGK (SEQ ID NO: 291), GDNQILQHHVLTR (SEQ ID NO: 292), GFLLWYSGR (SEQ ID NO: 123), GMCTSPPLIK (SEQ ID NO: 124), GNPTVEVDLHTAK (SEQ ID NO: 125), IDLADFEK (SEQ ID NO: 126), IMPNSFIMMFK (SEQ ID NO: 127), IYLSDSLTGK (SEQ ID NO: 128), MSEQLNDLTYDMEILQPLLEQGASLR (SEQ ID NO: 129), NEQVEIR (SEQ ID NO: 130), NNLELSTPLK (SEQ ID NO: 131), QWSEPPR (SEQ ID NO: 302), SDACQGDSGGPLACEK (SEQ ID NO: 132), SQYLLTAIHK (SEQ ID NO: 133), SSQSLLHSNGYNYLDWYLQK (SEQ ID NO: 134), SWWGDYWEPFR (SEQ ID NO: 135), TQFTCECSIGFR (SEQ ID NO: 307), TTCLVACDEGYR (SEQ ID NO: 136), VITANILQLQVK (SEQ ID NO: 137), VYTVDLGR (SEQ ID NO: 138), YQEALAK (SEQ ID NO: 139), ADPSLNPEQLK (SEQ ID NO: 140), APSDLYQIILK (SEQ ID NO: 141), ATAQMLEVMFK (SEQ ID NO: 142), AVLDVFEEGTEASAATAVK (SEQ ID NO: 143), AVSDWIDEQEK (SEQ ID NO: 144), CCHCCLLGR (SEQ ID NO: 145), CSVFYGAPSK (SEQ ID NO:146), DFHINLFR (SEQ ID NO: 147), DGIHNVEGVAVDWMGDNLYWTDDGPK (SEQ ID NO: 148), DTPVLSELPEPVVAR (SEQ ID NO: 149), DYENGFGNFVQK (SEQ ID NO: 150), DYTSGAMLTGELK (SEQ ID NO: 151), EAACLCPPGWVGER (SEQ ID NO: 152), EHAVEGDCDFQLLK (SEQ ID NO: 153), EIQVQHPAAK (SEQ ID NO: 154), EITEAAVLLFYR (SEQ ID NO: 155), EWVAIESDSVQPVPR (SEQ ID NO: 156), FLLFGIQDGK (SEQ ID NO: 157), FLLYGLHEGK (SEQ ID NO: 158), FMIELDGTENK (SEQ ID NO: 159), GSFRPIWVTLDTEDHK (SEQ ID NO: 160), HTLNQIDEVK (SEQ ID NO: 161), HVSSPLASYFLSFPYAR (SEQ ID NO: 162), IIYSPTVGDPIDEYTTVPGRR (SEQ ID NO: 163), ILEGFQPSGR (SEQ ID NO: 164), IREEYPDRIMNTFSVVPSPK (SEQ ID NO: 165), ITCQGDSLR (SEQ ID NO: 166), IVEGSDAEIGMSPWQVMLFRK (SEQ ID NO: 167), LGVLTVTDTTPDSMR (SEQ ID NO: 168), LLIYSNNQRPSGVPDR (SEQ ID NO: 169), LLQEDTPVRK (SEQ ID NO: 170), LMQVWCDQR (SEQ ID NO: 171), LSLLCIDFNK (SEQ ID NO: 172), LYLDRNLIAAVAPGAFLGLK (SEQ ID NO: 173), MRPSTDTITVMVENSHGLR (SEQ ID NO: 174), MCVDVNECQR (SEQ ID NO: 175), MGPTELLIEMEDWKGDK (SEQ ID NO: 176), MLTELEK (SEQ ID NO: 177), MRPSTDTITVMVENSHGLR (SEQ ID NO: 178), NSLFVGESGNVGTEMMDNR (SEQ ID NO: 179), PFTEAQLLCTQAGGQLASPR (SEQ ID NO: 180), PLQSPAEFSVYCDMSDGGGWTVIQR (SEQ ID NO: 181), QNLLAPQTLPSK (SEQ ID NO: 182), QSPVDIDTHTAK (SEQ ID NO: 183), QYADCSEIFNDGYK (SEQ ID NO: 184), RFNTFIHEDIWNIR (SEQ ID NO: 185), RIEESAIDEVVVTNTIPHEVQK (SEQ ID NO: 186), SIPLVQVLR (SEQ ID NO: 187), SPVTLLAAVMSLPEEHNK (SEQ ID NO: 189), SSEVYAQLCNVAR (SEQ ID NO: 190), SVGEVMAIGR (SEQ ID NO: 191), TCQSLHINEMCQER (SEQ ID NO: 192), TLEIPGNSDPNMIPDGDFNSYVR (SEQ ID NO: 193), TLMFGSYLDDEK (SEQ ID NO: 194), VFCNMDVNGGGWTVIQHR (SEQ ID NO: 195), VGLSGMAIADVTLLSGFHALRADLEK (SEQ ID NO: 196), VKPAPDETSFSEALLK (SEQ ID NO: 197), VSPHSGVVALTKPVPEPR (SEQ ID NO: 198), VSVNPSYLVPESDYTNNVVR (SEQ ID NO: 199), VVEPPEK(SEQ ID NO: 200), WEMPFDPQDTHQSR (SEQ ID NO: 201), WFLEWDAK (SEQ ID NO: 202), or YVAVMPPHIGDQPLTGAYTVTLDGR (SEQ ID NO: 203). A fragment or modification of any of these peptides may be used. Any of these biomarkers may be useful alone or in combination to assess a lung cancer. In some cases, any of these peptides may be useful as biomarkers when measured after being adsorbed from a biofluid sample to a particle. A biomarker may include HGEYWLGNK (SEQ ID NO: 1). A biomarker may include A biomarker may include AEPSAATQSHSISSSSFGAEPSAPGGGGSPGACPALGTK (SEQ ID NO: 2). A biomarker may include A biomarker may include MVGAGISTPSGIPDFR (SEQ ID NO: 3). A biomarker may include A biomarker may include MNPIWVHGGGAGPISK (SEQ ID NO: 4). A biomarker may include AAAATGTIFTFR (SEQ ID NO: 5). A biomarker may include AAAEDVNVTFEDQQK (SEQ ID NO: 6). A biomarker may include AAATCFAR (SEQ ID NO: 7). A biomarker may include AADYLHVALDLLER (SEQ ID NO: 8). A biomarker may include AALPEGLPEASR (SEQ ID NO: 9). A biomarker may include AAMPPQIIQFPEDQK (SEQ ID NO: 10). A biomarker may include AATVGSLAGQPLQERAQAWGER (SEQ ID NO: 11). A biomarker may include AAVEWFDGK (SEQ ID NO: 12). A biomarker may includeAAVYHHFISDGVR (SEQ ID NO: 13). A biomarker may include ACKKDACPINGGWGPWSPWDICSVTCGGGVQK (SEQ ID NO: 14). A biomarker may include ACSMPQELPQSPR (SEQ ID NO: 15). A biomarker may include ADHAHCCVEMGVDMIEAISLVR (SEQ ID NO: 16). A biomarker may include ADMVIEAVFEDLSLK (SEQ ID NO: 184). A biomarker may include ADRDQYELLCLDNTR (SEQ ID NO: 17). A biomarker may include ADSPMDDFFQCVNGK (SEQ ID NO: 19). A biomarker may include AECLNPSQPSR (SEQ ID NO: 20). A biomarker may include AEEYILGDFCFLR (SEQ ID NO: 21). A biomarker may include AEILLSSSKPVPK (SEQ ID NO: 22). A biomarker may include AFCLACPFYGTTPFAGSR (SEQ ID NO: 23). A biomarker may include AFDLIVDRPVTLVR (SEQ ID NO: 24). A biomarker may include AFLHVPAK (SEQ ID NO: 25). A biomarker may include AFQLWSNVTPLTFTK (SEQ ID NO: 26). A biomarker may include AFVHWYVGEGMEEGEFSEAREDLAALEK (SEQ ID NO: 27). A biomarker may include AGAAPYVQAFDSLLAGPVAEYLK (SEQ ID NO: 28). A biomarker may include AGAFCLSEDAGLGISSTASLR (SEQ ID NO: 29). A biomarker may include AGHIAWTSSGK (SEQ ID NO: 30). A biomarker may include AGLFSQEQYER (SEQ ID NO: 31). A biomarker may include AGLYGLPR (SEQ ID NO: 32). A biomarker may include AIEMLGGELGSK (SEQ ID NO: 33). A biomarker may include AIGPSQTHTIR (SEQ ID NO: 34). A biomarker may include CDSSPDSAEDVRK (SEQ ID NO: 35). A biomarker may include CLAALASLR (SEQ ID NO: 36). A biomarker may include EDITQSAQHALR (SEQ ID NO: 37). A biomarker may include EHAVEGDCDFQLLK (SEQ ID NO: 38). A biomarker may include ETLLQDFR (SEQ ID NO: 39). A biomarker may include GAEVSFGCGVLASGK (SEQ ID NO: 40). A biomarker may include GDLDLELVLLCK (SEQ ID NO: 41). A biomarker may include GECVPGEQEPEPILIPR (SEQ ID NO: 42). A biomarker may include GECWCVNPNTGK (SEQ ID NO: 43). A biomarker may include GLFLPEDENLR (SEQ ID NO: 44). A biomarker may include GVDCMEVYEYPGYR (SEQ ID NO: 45). A biomarker may include GVPWNFTFNVK (SEQ ID NO: 46). A biomarker may include IPLDMVAGFNTPLVK (SEQ ID NO: 47). A biomarker may include IQPSGGTNINEALLR (SEQ ID NO: 48). A biomarker may include LAYVAPTIPR (SEQ ID NO: 49). A biomarker may include LFIGGLNTETNEK (SEQ ID NO: 50). A biomarker may include NGLTPLHVAVHHNNLDIVK (SEQ ID NO: 51). A biomarker may include PYSLHAHGLSYEK (SEQ ID NO: 52). A biomarker may include QGALELIK (SEQ ID NO: 53). A biomarker may include SQNPVQPIGPQTPK (SEQ ID NO: 54). A biomarker may include SVTHANALTVMGK (SEQ ID NO: 55). A biomarker may include TEAPSATGQASSLLGGR (SEQ ID NO: 56). A biomarker may include TVSSLSEDLESTR (SEQ ID NO: 56). A biomarker may include VFDEFKPLVEEPQNLIK (SEQ ID NO: 57). A biomarker may include VLLEAGEGLVTITPTTGSDGRPDAR (SEQ ID NO: 58). A biomarker may include VVEESELAR (SEQ ID NO: 59). A biomarker may include VVLAYEPVWAIGTGK (SEQ ID NO: 60). A biomarker may include WFCHVDDDNYVNAR (SEQ ID NO: 61). A biomarker may include YEMHELLR (SEQ ID NO: 62). A biomarker may include YQSSPAKPDSSFYK (SEQ ID NO: 63). A biomarker may include ALILGELEK (SEQ ID NO: 64). A biomarker may include CDSSPDSAEDVRK (SEQ ID NO: 65). A biomarker may include CTNLEGSFR (SEQ ID NO: 66). A biomarker may include DFDFVPPVVR (SEQID NO: 67). A biomarker may include DTPVLSELPEPVVAR (SEQ ID NO: 236). A biomarker may include EGEAVVLPEVEPGLTAR (SEQ ID NO: 68). A biomarker may include ELCCLVYTSWQIPQK (SEQ ID NO: 69). A biomarker may include EQLGEFYEALDCLR (SEQ ID NO: 70). A biomarker may include FGLLDEDGKK (SEQ ID NO: 71). A biomarker may include GEALEDFTGPDCR (SEQ ID NO: 72). A biomarker may include HQLYIDETVNSNIPTNLR (SEQ ID NO: 73). A biomarker may include IDLADFEK (SEQ ID NO: 74). A biomarker may include IDQYQGADAVGLEEK (SEQ ID NO: 75). A biomarker may include IKPLQSPAEFSVYCDMSDGGGWTVIQR (SEQ ID NO: 76). A biomarker may include ILLNPQDK (SEQ ID NO: 77). A biomarker may include IRPNDFIPNVI (SEQ ID NO: 78). A biomarker may include LGVLTVTDTTPDSMR (SEQ ID NO: 79). A biomarker may include LQEAAELEAVELPVPIR (SEQ ID NO: 80). A biomarker may include LSGNVLSYTFQVK (SEQ ID NO: 81). A biomarker may include LVHVEEPHTETVRK (SEQ ID NO: 82). A biomarker may include LYHSEAFTVNFGDTEEAK (SEQ ID NO: 83). A biomarker may include NFEDVAFDEK (SEQ ID NO: 84). A biomarker may include NIETIINTFHQYSVK (SEQ ID NO: 85). A biomarker may include NRDVVLTTTFVDDIK (SEQ ID NO: 86). A biomarker may include NWGLSVYADKPETTK (SEQ ID NO: 87). A biomarker may include PGQAPVLVIYK (SEQ ID NO: 88). A biomarker may include QNLLAPQTLPSK (SEQ ID NO: 89). A biomarker may include QSSGENCDVVVNTLGK (SEQ ID NO: 90). A biomarker may include SSTGPGEQLR (SEQ ID NO: 91). A biomarker may include STSESTAALGCLVK (SEQ ID NO: 92). A biomarker may include TKLEEHLEGIVNIFHQYSVRK (SEQ ID NO: 93). A biomarker may include TLMFGSYLDDEK (SEQ ID NO: 94). A biomarker may include TVTAMDVVYALK (SEQ ID NO: 95). A biomarker may include VHFLPVVISDNGMPSR (SEQ ID NO: 96). A biomarker may include VLDFHNLPDGITK (SEQ ID NO: 97). A biomarker may include YD VENCLANK (SEQ ID NO: 98). A biomarker may include YNWRENLDR (SEQ ID NO: 99). A biomarker may include AALPEGLPEASR (SEQ ID NO: 100). A biomarker may include ALNSIIDVYHK (SEQ ID NO: 101). A biomarker may include DTHFPICIFCCGCCHR (SEQ ID NO: 102). A biomarker may include EDFLEQSEQLFGAK (SEQ ID NO: 103). A biomarker may include EFTRPEEIIFLR (SEQ ID NO: 104). A biomarker may include FCTALLPVNDR (SEQ ID NO: 105). A biomarker may include GDTFSCMVGHEALPLAFTQK (SEQ ID NO: 106). A biomarker may include GLEVTITAR (SEQ ID NO: 107). A biomarker may include GPGGVWAAK (SEQ ID NO: 108). A biomarker may include HLNQGTDEDIYLLGK (SEQ ID NO: 109). A biomarker may include LMHSFCAFK (SEQ ID NO: 110). A biomarker may include LSCAASGFTFSSYSMNWVR (SEQ ID NO: 111). A biomarker may include TSVFHDVDGSVSEYPGSYLTK (SEQ ID NO: 112). A biomarker may include VHQYFNVELIQPGAVK (SEQ ID NO: 113). A biomarker may include VLTSLAWR (SEQ ID NO: 114). A biomarker may include VTIDSSYDIAK (SEQ ID NO: 115). A biomarker may include YEINSLIR (SEQ ID NO: 116). A biomarker may include AALPEGLPEASR (SEQ ID NO: 117). A biomarker may include AAPSAEFSVDR (SEQ ID NO: 118). A biomarker may include AMGIMNSFVNDIFER (SEQ ID NO: 119). A biomarker may include DGSFSWITGLR (SEQ ID NO: 120). A biomarker may include EIGELYLPK (SEQ ID NO: 121). A biomarker may include EMPSVFGK (SEQ ID NO: 122). A biomarker may include GDNQILQHHVLTR (SEQ ID NO: 123). A biomarker may include GFLLWYSGR (SEQ ID NO: 124). A biomarker may includeGMCTSPPLIK (SEQ ID NO: 125). A biomarker may include GNPTVEVDLHTAK (SEQ ID NO: 126). A biomarker may include IDLADFEK (SEQ ID NO: 127). A biomarker may include IMPNSFIMMFK (SEQ ID NO: 128). A biomarker may include IYLSDSLTGK (SEQ ID NO: 129). A biomarker may include MSEQLNDLTYDMEILQPLLEQGASLR (SEQ ID NO: 130). A biomarker may include NEQVEIR (SEQ ID NO: 131). A biomarker may include NNLELSTPLK (SEQ ID NO: 132). A biomarker may include QWSEPPR (SEQ ID NO: 133). A biomarker may include SDACQGDSGGPLACEK (SEQ ID NO: 134). A biomarker may include SQYLLTAIHK (SEQ ID NO: 135). A biomarker may include SSQSLLHSNGYNYLDWYLQK (SEQ ID NO: 136). A biomarker may include SWWGDYWEPFR (SEQ ID NO: 137). A biomarker may include TQFTCECSIGFR (SEQ ID NO: 138). A biomarker may include TTCLVACDEGYR (SEQ ID NO: 139). A biomarker may include VITANILQLQVK (SEQ ID NO: 140). A biomarker may include VYTVDLGR (SEQ ID NO: 141). A biomarker may include Y QEALAK (SEQ ID NO: 142). A biomarker may include ADPSLNPEQLK (SEQ ID NO: 143). A biomarker may include APSDLYQIILK (SEQ ID NO: 144). A biomarker may include ATAQMLEVMFK (SEQ ID NO: 145). A biomarker may include AVLDVFEEGTEASAATAVK (SEQ ID NO: 146). A biomarker may include AVSDWIDEQEK (SEQ ID NO: 147). A biomarker may include CCHCCLLGR (SEQ ID NO: 148). A biomarker may include CSVFYGAPSK (SEQ ID NO: 149). A biomarker may include DFHINLFR (SEQ ID NO: 150). A biomarker may include DGIHNVEGVAVDWMGDNLYWTDDGPK (SEQ ID NO: 151). A biomarker may include DTPVLSELPEPVVAR (SEQ ID NO: 152). A biomarker may include DYENGFGNFVQK (SEQ ID NO: 153). A biomarker may include DYTSGAMLTGELK (SEQ ID NO: 154). A biomarker may include EAACLCPPGWVGER (SEQ ID NO: 155). A biomarker may include EHAVEGDCDFQLLK (SEQ ID NO: 156). A biomarker may include EIQVQHPAAK (SEQ ID NO: 157). A biomarker may include EITEAAVLLFYR (SEQ ID NO: 158). A biomarker may include EWVAIESDSVQPVPR (SEQ ID NO: 159). A biomarker may include FLLFGIQDGK (SEQ ID NO: 160). A biomarker may include FLLYGLHEGK (SEQ ID NO: 161). A biomarker may include FMIELDGTENK (SEQ ID NO: 162). A biomarker may include GSFRPIWVTLDTEDHK (SEQ ID NO: 163). A biomarker may include HTLNQIDEVK (SEQ ID NO: 164). A biomarker may include HVSSPLASYFLSFPYAR (SEQ ID NO: 165). A biomarker may include IIYSPTVGDPIDEYTTVPGRR (SEQ ID NO: 166). A biomarker may include ILEGFQPSGR (SEQ ID NO: 167). A biomarker may include IREEYPDRIMNTFSVVPSPK (SEQ ID NO: 168). A biomarker may include ITCQGDSLR (SEQ ID NO: 169). A biomarker may include IVEGSDAEIGMSPWQVMLFRK (SEQ ID NO: 170). A biomarker may include LGVLTVTDTTPDSMR (SEQ ID NO: 171). A biomarker may include LLIYSNNQRPSGVPDR (SEQ ID NO: 172). A biomarker may include LLQEDTPVRK (SEQ ID NO: 173). A biomarker may include LMQVWCDQR (SEQ ID NO: 174). A biomarker may include LSLLCIDFNK (SEQ ID NO: 175). A biomarker may include LYLDRNLIAAVAPGAFLGLK (SEQ ID NO: 176). A biomarker may include MRPSTDTITVMVENSHGLR (SEQ ID NO: 178). A biomarker may include MCVDVNECQR (SEQ ID NO: 179). A biomarker may include MGPTELLIEMEDWKGDK (SEQ ID NO: 180). A biomarker may include MLTELEK (SEQ ID NO: 181). A biomarker may include MRPSTDTITVMVENSHGLR (SEQ ID NO: 182). A biomarker may include NSLFVGESGNVGTEMMDNR (SEQ ID NO: 183). A biomarker mayinclude PFTEAQLLCTQAGGQLASPR (SEQ ID NO: 184). A biomarker may include PLQSPAEFSVYCDMSDGGGWTVIQR (SEQ ID NO: 185). A biomarker may include QNLLAPQTLPSK (SEQ ID NO: 186). A biomarker may include QSPVDIDTHTAK (SEQ ID NO: 187). A biomarker may include QYADCSEIFNDGYK (SEQ ID NO: 188). A biomarker may include RFNTFIHEDIWNIR (SEQ ID NO: 189). A biomarker may include RIEESAIDEVVVTNTIPHEVQK (SEQ ID NO: 190). A biomarker may include SIPLVQVLR (SEQ ID NO: 191). A biomarker may include SPVTLLAAVMSLPEEHNK (SEQ ID NO: 192). A biomarker may include SSEVYAQLCNVAR (SEQ ID NO: 193). A biomarker may include SVGEVMAIGR (SEQ ID NO: 194). A biomarker may include TCQSLHINEMCQER (SEQ ID NO: 195). A biomarker may include TLEIPGNSDPNMIPDGDFNSYVR (SEQ ID NO: 196). A biomarker may include TLMFGSYLDDEK (SEQ ID NO: 197). A biomarker may include VFCNMDVNGGGWTVIQHR (SEQ ID NO: 198). A biomarker may include VGLSGMAIADVTLLSGFHALRADLEK (SEQ ID NO: 199). A biomarker may include VKPAPDETSFSEALLK (SEQ ID NO: 200). A biomarker may include VSPHSGVVALTKPVPEPR (SEQ ID NO: 201). A biomarker may include VSVNPSYLVPESDYTNNVVR (SEQ ID NO: 202). A biomarker may include VVEPPEK (SEQ ID NO: 203). A biomarker may include WEMPFDPQDTHQSR (SEQ ID NO: 204). A biomarker may include WFLEWDAK (SEQ ID NO: 205). A biomarker may include or YVAVMPPHIGDQPLTGAYTVTLDGR (SEQ ID NO: 206). Any of the forementioned peptide or protein biomarkers (or combination of said biomarkers) may be useful for identifying a presence, absence, or likelihood of a cancer described herein. Additional peptide features listed in Table 15 below:Table 15 Additional peptides for use as features in building classifiers for indicating a presence or susceptibility of cancer (e.g., lung cancer)
[0169] A peptide may have a modification. A peptide may have one or more modifications. A modification can be any type of modification. A modification may be identified using a UNIMOD accession number. Non-limiting examples of a modification can include, but is not limited to a post-translational modification, an artifact, a chemical derivative. A modification may be an acetylation. A modification may be a carbamidomethylation. A modification may be an oxidation. A modification may be a hydroxylation. The modified peptides may include any of the following peptides(UniMod: l)AEPSAATQSHSISSSSFGAEPSAPGGGGSPGAC(UniMod:4)PALGTK (SEQ ID NO: 443),(UniMod: l)M(UniMod:35)VGAGISTPSGIPDFR (SEQ ID NO: 444), (UniMod: 1)MNPIVWHGGGAGPISK (SEQ ID NO: 445), AAATC(UniMod:4)FAR (SEQ ID NO: 446), AAM(UniMod:35)PPQIIQFPEDQK (SEQ ID NO: 447), AC(UniMod:4)KKDAC(UniMod:4)PINGGWGPWSPWDIC(UniMod:4)SVTC(UniMod:4)GGGVQK (SEQ ID NO: 448), AC(UniMod:4)SM(UniMod:35)PQELPQSPR (SEQ ID NO: 449), ADHAHC(UniMod:4)C(UniMod:4)VEMGVDMIEAISLVR (SEQ ID NO: 450), AEC(UniMod:4)LNPSQPSR (SEQ ID NO: 451), AEEYILGDFC(UniMod:4)FLR (SEQ ID NO: 452), AFC(UniMod:4)LAC(UniMod:4)PFYGTTPFAGSR (SEQ ID NO: 453), AGAFC(UniMod:4)LSEDAGLGISSTASLR (SEQ ID NO: 454), C(UniMod:4)DSSPDSAEDVRK (SEQ ID NO: 455), C(UniMod:4)LAALASLR (SEQ ID NO: 456), EHAVEGDC(UniMod:4)DFQLLK (SEQ ID NO: 457), GAEVSFGC(UniMod:4)GVLASGK (SEQ ID NO: 458), GDLDLELVLLC(UniMod:4)K (SEQ ID NO: 459), GEC(UniMod:4)VPGEQEPEPILIPR (SEQ ID NO: 460), GEC(UniMod:4)WC(UniMod:4)VNPNTGK (SEQ ID NO: 461), GVDC(UniMod:4)M(UniMod:35)EVYEYPGYR (SEQ ID NO: 462), WFC(UniMod:4)HVDDDNYVNAR (SEQ ID NO: 463), C(UniMod:4)DSSPDSAEDVRK (SEQ ID NO: 464), C(UniMod:4)TNLEGSFR (SEQ ID NO: 465), ELC(UniMod:4)C(UniMod:4)LVYTSWQIPQK (SEQ ID NO: 466), EQLGEFYEALDC(UniMod:4)LR (SEQ ID NO: 467), GEALEDFTGPDC(UniMod:4)R (SEQ ID NO: 400), IKPLQSPAEFSVYC(UniMod:4)DMSDGGGWTVIQR (SEQ ID NO: 468), QSSGENC(UniMod:4)DWVNTLGK (SEQ ID NO: 469), STSESTAALGC(UniMod:4)LVK (SEQ ID NO: 470), YDVENC(UniMod:4)LANK (SEQ ID NO: 471), DTHFPIC(UniMod:4)IFC(UniMod:4)C(UniMod:4)GC(UniMod:4)C(UniMod:4)HR (SEQ ID NO: 472), FC(UniMod:4)TALLPVNDR (SEQ ID NO: 473), GDTFSC(UniMod:4)M(UniMod:35)VGHEALPLAFTQK (SEQ ID NO: 474), LMHSFC(UniMod:4)AFK (SEQ ID NO: 475), LSC(UniMod:4)AASGFTFSSYSMNWVR (SEQ ID NO: 476), GMC(UniMod:4)TSPPLIK (SEQ ID NO: 477), M(UniMod:35)SEQLNDLTYDM(UniMod:35)EILQPLLEQGASLR (SEQ ID NO: 478), SDAC(UniMod:4)QGDSGGPLAC(UniMod:4)EK (SEQ ID NO: 479), TQFTC(UniMod:4)EC(UniMod:4)SIGFR (SEQ ID NO: 480), TTC(UniMod:4)LVAC(UniMod:4)DEGYR (SEQ ID NO: 481), C(UniMod:4)C(UniMod:4)HC(UniMod:4)C(UniMod:4)LLGR (SEQ ID NO: 482), C(UniMod:4)SVFYGAPSK (SEQ ID NO: 483), EAAC(UniMod:4)LC(UniMod:4)PPGWVGER (SEQ ID NO: 484), EHAVEGDC(UniMod:4)DFQLLK (SEQ ID NO: 485), ITC(UniMod:4)QGDSLR (SEQ ID NO: 486), IVEGSDAEIGMSPWQVM(UniMod:35)LFRK (SEQ ID NO: 487), LMQVWC(UniMod:4)DQR (SEQ ID NO: 488), LSLLC(UniMod:4)IDFNK (SEQ ID NO: 489),M(UniMod:35)RPSTDTITVMVENSHGLR (SEQ ID NO: 490), MC(UniMod:4)VDVNEC(UniMod:4)QR (SEQ ID NO: 491), MRPSTDTITVM(UniMod:35)VENSHGLR (SEQ ID NO: 492), PFTEAQLLC(UniMod:4)TQAGGQLASPR (SEQ ID NO: 493), PLQSPAEFSVYC(UniMod:4)DMSDGGGWTVIQR (SEQ ID NO: 494), QYADC(UniMod:4)SEIFNDGYK (SEQ ID NO: 495), SPVTLLAAVM(UniMod:35)SLPEEHNK (SEQ IDNO: 496), SSEVYAQLC(UniMod:4)NVAR (SEQ ID NO: 497), TC(UniMod:4)QSLHINEM(UniMod:35)C(UniMod:4)QER (SEQ ID NO: 498), VFC(UniMod:4)NMDVNGGGWTVIQHR (SEQ ID NO: 499), or VGLSGM(UniMod:35)AIADVTLLSGFHALRADLEK (SEQ ID NO: 500). A fragment of any of these peptides may be used. Any of these biomarkers may be useful alone or in combination to assess a lung cancer. In some cases, any of these peptides may be useful as biomarkers when measured after being adsorbed from a biofluid sample to a particle. A biomarker may include (UniMod: l)AEPSAATQSHSISSSSFGAEPSAPGGGGSPGAC(UniMod:4)PALGTK (SEQ ID NO: 501). A biomarker may include A biomarker may include (UniMod: l)M(UniMod:35)VGAGISTPSGIPDFR (SEQ ID NO: 502). A biomarker may include A biomarker may include (UniMod: 1)MNPIVVVHGGGAGPISK (SEQ ID NO: 503). A biomarker may include A biomarker may include AAATC(UniMod:4)FAR (SEQ ID NO: 504). A biomarker may include AAM(UniMod:35)PPQIIQFPEDQK (SEQ ID NO: 505). A biomarker may include AC(UniMod:4)KKDAC(UniMod:4)PINGGWGPWSPWDIC(UniMod:4)SVTC(UniMod:4)GGGVQK (SEQ ID NO: 506). A biomarker may include AC(UniMod:4)SM(UniMod:35)PQELPQSPR (SEQ ID NO: 507). A biomarker may include ADHAHC(UniMod:4)C(UniMod:4)VEMGVDMIEAISLVR (SEQ ID NO: 508). A biomarker may include AEC(UniMod:4)LNPSQPSR (SEQ ID NO: 509). A biomarker may include AEEYILGDFC(UniMod:4)FLR (SEQ ID NO: 510). A biomarker may include AFC(UniMod:4)LAC(UniMod:4)PFYGTTPFAGSR (SEQ ID NO: 511). A biomarker may include AGAFC(UniMod:4)LSEDAGLGISSTASLR (SEQ ID NO: 512). A biomarker may include C(UniMod:4)DSSPDSAEDVRK (SEQ ID NO: 513). A biomarker may include C(UniMod:4)LAALASLR (SEQ ID NO: 514). A biomarker may include EHAVEGDC(UniMod:4)DFQLLK (SEQ ID NO: 515). A biomarker may include GAEVSFGC(UniMod:4)GVLASGK (SEQ ID NO: 516). A biomarker may include GDLDLELVLLC(UniMod:4)K (SEQ ID NO: 517). A biomarker may include GEC(UniMod:4)VPGEQEPEPILIPR (SEQ ID NO: 518). A biomarker may include GEC(UniMod:4)WC(UniMod:4)VNPNTGK (SEQ ID NO: 519). A biomarker may include GVDC(UniMod:4)M(UniMod:35)EVYEYPGYR (SEQ ID NO: 520). A biomarker may include WFC(UniMod:4)HVDDDNYVNAR (SEQ ID NO: 521). A biomarker may include C(UniMod:4)DSSPDSAEDVRK (SEQ ID NO: 522). A biomarker may include C(UniMod:4)TNLEGSFR (SEQ ID NO: 523). A biomarker may include ELC(UniMod:4)C(UniMod:4)LVYTSWQIPQK (SEQ ID NO: 524). A biomarker may include EQLGEFYEALDC(UniMod:4)LR (SEQ ID NO: 525). A biomarker may include GEALEDFTGPDC(UniMod:4)R (SEQ ID NO: 526). A biomarker may include IKPLQSPAEFSVYC(UniMod:4)DMSDGGGWTVIQR (SEQ ID NO: 527). A biomarker may include QSSGENC(UniMod:4)DWVNTLGK (SEQ ID NO: 528). A biomarker may include STSESTAALGC(UniMod:4)LVK (SEQ ID NO: 529). A biomarker may include YDVENC(UniMod:4)LANK (SEQ ID NO: 530). A biomarker may include DTHFPIC(UniMod:4)IFC(UniMod:4)C(UniMod:4)GC(UniMod:4)C(UniMod:4)HR (SEQ ID NO: 531). A biomarker may include FC(UniMod:4)TALLPVNDR (SEQ ID NO: 532). A biomarker may includeGDTFSC(UniMod:4)M(UniMod:35)VGHEALPLAFTQK (SEQ ID NO: 533). A biomarker may include LMHSFC(UniMod:4)AFK (SEQ ID NO: 534). A biomarker may include LSC(UniMod:4)AASGFTFSSYSMNWVR (SEQ ID NO: 535). A biomarker may include GMC(UniMod:4)TSPPLIK (SEQ ID NO: 536). A biomarker may include M(UniMod:35)SEQLNDLTYDM(UniMod:35)EILQPLLEQGASLR (SEQ ID NO: 537). A biomarker may include SDAC(UniMod:4)QGDSGGPLAC(UniMod:4)EK (SEQ ID NO: 538). A biomarker may include TQFTC(UniMod:4)EC(UniMod:4)SIGFR (SEQ ID NO: 539). A biomarker may include TTC(UniMod:4)LVAC(UniMod:4)DEGYR (SEQ ID NO: 540). A biomarker may include C(UniMod:4)C(UniMod:4)HC(UniMod:4)C(UniMod:4)LLGR (SEQ ID NO: 541). A biomarker may include C(UniMod:4)SVFYGAPSK (SEQ ID NO: 542). A biomarker may include EAAC(UniMod:4)LC(UniMod:4)PPGWVGER (SEQ ID NO: 543). A biomarker may include EHAVEGDC(UniMod:4)DFQLLK (SEQ ID NO: 544). A biomarker may include ITC(UniMod:4)QGDSLR (SEQ ID NO: 545). A biomarker may include IVEGSDAEIGMSPWQVM(UniMod:35)LFRK (SEQ ID NO: 546). A biomarker may include LMQVWC(UniMod:4)DQR (SEQ ID NO: 547). A biomarker may include LSLLC(UniMod:4)IDFNK (SEQ ID NO: 548). A biomarker may include M(UniMod:35)RPSTDTITVMVENSHGLR (SEQ ID NO: 549). A biomarker may include MC(UniMod:4)VDVNEC(UniMod:4)QR (SEQ ID NO: 550). A biomarker may include MRPSTDTITVM(UniMod:35)VENSHGLR (SEQ ID NO: 551). A biomarker may include PFTEAQLLC(UniMod:4)TQAGGQLASPR (SEQ ID NO: 552). A biomarker may include PLQSPAEFSVYC(UniMod:4)DMSDGGGWTVIQR (SEQ ID NO: 553). A biomarker may include QYADC(UniMod:4)SEIFNDGYK (SEQ ID NO: 554). A biomarker may include SPVTLLAAVM(UniMod:35)SLPEEHNK (SEQ ID NO: 555). A biomarker may include SSEVYAQLC(UniMod:4)NVAR (SEQ ID NO: 556). A biomarker may include TC(UniMod:4)QSLHINEM(UniMod:35)C(UniMod:4)QER (SEQ ID NO: 557). A biomarker may include VFC(UniMod:4)NMDVNGGGWTVIQHR (SEQ ID NO: 558). A biomarker may include VGLSGM(UniMod:35)AIADVTLLSGFHALRADLEK (SEQ ID NO: 559). A cancer antigen may be used to assess a cancer status. Any one of the following cancer antigens may be used including C125a, CA153, CA199, or CEA. A biomarker may be C125a. A biomarker may be CA153. A biomarker may be CA199. A biomarker may be CEA. Any of the forementioned peptide or protein biomarkers (or combination of said biomarkers) may be useful for identifying a presence, absence, or likelihood of a cancer described herein.
[0170] The protein measurements may be obtained with the use of internal standards. Any combination or number of such biomarkers may be included. In some cases, a biomarker is useful when its feature importance score is above 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.10, 0.12, or 0.14. The features may include any of the following proteins: Immunoglobulin kappa variable 2-28(KV228 HUMAN; UniProt ID A0A075B6P5), Immunoglobulin heavy variable 3-21 (HV321_HUMAN; UniProt ID A0A0B4J1V1), Protein phosphatase 1 regulatory subunit 12A (MYPT1 HUMAN; UniProt ID 014974), Adenylate cyclase type 6 (ADCY6_HUMAN; UniProt ID 043306), Thioredoxin-like protein 1 (TXNL1_HUMAN; UniProt ID 043396), Phosphoribosyl pyrophosphate synthase-associated protein 2(KPRB HUMAN; UniProt ID 060256), Reticulon-3 (RTN3 HUMAN; UniProt ID 095197), NADH dehydrogenase (NDUBA HUMAN; UniProt ID 096000), Prothrombin (THRB HUMAN; UniProt ID P00734), Carbonic anhydrase 2 (CAH2_HUMAN; UniProt ID P00918), Alpha- 1 -antitrypsin(A1AT HUMAN; UniProt ID P01009), Alpha- 1 -antichymotrypsin (AACT HUMAN; UniProt ID P01011), Alpha-2-macroglobulin (A2MG_HUMAN; UniProt ID P01023), Complement C3 (C03_HUMAN; UniProt ID P01024), Immunoglobulin lambda variable 1-47 (LV147 HUMAN; UniProt ID P01700), Immunoglobulin lambda variable 3-19 (LV319 HUMAN; UniProt ID P01714), Immunoglobulin lambda variable 3-25 (UV325_HUMAN; UniProt ID P01717), Polymeric immunoglobulin receptor (PIGR HUMAN; UniProt ID PO 1833), Immunoglobulin heavy constant gamma 2 (IGHG2 HUMAN; UniProt ID P01859), Immunoglobulin heavy constant alpha 1 (IGHA1 HUMAN; UniProt ID P01876), Apolipoprotein E (APOE HUMAN; P02649), Fibrinogen beta chain (FIBB HUMAN; UniProt ID P02675), Band 3 anion transport protein (B3AT HUMAN; UniProt ID P02730), Complement component C9 (C09 HUMAN; UniProt ID P02748), Protein AMBP (AMBP HUMAN; UniProt ID P02760), Alpha- 1- acid glycoprotein 1 (A1AG1_HUMAN; UniProt ID P02763), Alpha-2 -HS-glycoprotein (FETUA_HUMAN; UniProt ID P02765), Albumin (ALBU HUMAN; UniProt ID P02768), Serotransferrin (TRFE HUMAN; UniProt ID P02787), Interstitial collagenase (MMP1 HUMAN; UniProt ID P03956), Apolipoprotein B-100 (APOB HUMAN; UniProt ID P04114), Superoxide dismutase (SODM HUMAN; UniProt ID P04179), von Willebrand factor (VWF HUMAN; UniProt ID P04275), Protein S100-A8 (S10A8_HUMAN; UniProt ID P05109), Plasma serine protease inhibitor (IPSP HUMAN; UniProt ID P05154), Complement factor I (CFAI HUMAN; UniProt ID P05156), Coagulation factor XIII B chain (F13B HUMAN; UniProt ID P05160), Myeloperoxidase (PERM HUMAN; UniProt ID P05164), Complement C2 (C02_HUMAN; UniProt ID P06681), Protein S100-A9 (S10A9_HUMAN; UniProt ID P06702), Vitamin K-dependent protein S (PROS HUMAN; UniProt ID P07225), Protein disulfide-isomerase (PDIA1 HUMAN; UniProt ID P07237), Calpain-1 catalytic subunit (CAN1_HUMAN; UniProt ID P07384), Fumarate hydratase, mitochondrial (FUMH HUMAN; UniProt ID P07954), Thrombospondin- 1 (TSP1 HUMAN; UniProt ID P07996), Alpha-2-antiplasmin (A2AP_HUMAN; UniProt ID P08697), Complement C4-A(C04A HUMAN; UniProt ID P0C0L4), Osteopontin (OSTP HUMAN; UniProt ID P10451), Complement component C7 (C07 HUMAN; UniProt ID P10643), Tissue factor pathway inhibitor (TFPI1 HUMAN; UniProt ID P10646), Protein 4.1 (EPB41 HUMAN; UniProt ID Pl 1171), Coagulation factor V(FA5 HUMAN; UniProt ID P12259), Versican core protein (CSPG2_HUMAN; UniProt ID P13611), Betaenolase (ENOB HUMAN; UniProt ID P13929), Nidogen-1 (NID1 HUMAN; UniProt ID P14543), Phospholipase A2, membrane associated (PA2GA HUMAN; UniProt ID P14555), Beta-galactoside alpha- 2,6-sialyltransferase 1 (SIAT1 HUMAN; UniProt ID P15907), Ankyrin-1 (ANK1 HUMAN; UniProt ID P16157), Insulin-like growth factor-binding protein 2 (IBP2 HUMAN; UniProt ID P18065), Lipopolysaccharide -binding protein (LBP HUMAN; UniProt ID P18428), Syndecan-1 (SDC1 HUMAN; UniProt ID P18827), Alpha-l-acid glycoprotein 2 (A1AG2 HUMAN; UniProt ID P19652), Inter-alphatrypsin inhibitor heavy chain H2 (ITIH2 HUMAN; UniProt ID P19823), Pregnancy zone protein(PZP HUMAN; UniProt ID P20742), Collagen alpha- 1(V) chain (C05A1 HUMAN; UniProt IDP20908), Tenascin-X (TENX HUMAN; UniProt ID P22105), Fibulin-1 (FBLN1 HUMAN; UniProt ID P23142), Tryptophan— tRNA ligase, cytoplasmic (SYWC_HUMAN; UniProt ID P23381), Protein-lysine 6- oxidase (LYOX_HUMAN; UniProt ID P28300), Proprotein convertase subtilisin / kexin type 6 (PCSK6 HUMAN; UniProt ID P29122), Flavin reductase (BLVRB HUMAN; UniProt ID P30043), Carbamoyl-phosphate synthase (CPSM_HUMAN; UniProt ID P31327), Cadherin-5 (CADH5_HUMAN; UniProt ID P33151), Ribonuclease 4 (RNAS4 HUMAN; UniProt ID P34096), Pulmonary surfactant- associated protein D (SFTPD HUMAN; UniProt ID P35247), Serum amyloid A-4 protein (SAA4_HUMAN; UniProt ID P35542), Insulin-like growth factor-binding protein complex acid labile subunit (ALS HUMAN; UniProt ID P35858), Chitinase-3 -like protein 1 (CH3L1 HUMAN; UniProt ID P36222), RNA-binding motif protein, X chromosome (RBMX_HUMAN; UniProt ID P38159), Collagen alpha- 1 (XVIII) chain (C0IA1 HUMAN; UniProt ID P39060), Trifunctional enzyme subunit alpha, mitochondrial (ECHA HUMAN; UniProt ID P40939), Fatty acid synthase (FAS HUMAN; UniProt ID P49327), T-complex protein 1 subunit gamma (TCPG HUMAN; UniProt ID P49368), Thrombospondin-3 (TSP3 HUMAN; UniProt ID P49746), Cartilage oligomeric matrix protein (COMP HUMAN; UniProt ID P49747), Thimet oligopeptidase (TH0P1 HUMAN; UniProt ID P52888), C-C motif chemokine 18 (CCL18 HUMAN; UniProt ID P55774), Triosephosphate isomerase (TPIS HUMAN; UniProt ID P60174), Histone H4 (H4 HUMAN; UniProt ID P62805), Histone H2B type 1-C / E / F / G / I (H2B1C HUMAN;UniProt ID P62807), Tubulin beta-4B chain (TBB4B HUMAN; UniProt ID P68371), Reelin (RELN HUMAN; UniProt ID P78509), Phosphatidylinositol-glycan-specific phospholipase D (PHLD HUMAN; UniProt ID P80108), Protein S100-A12 (S10AC HUMAN; UniProt ID P80511), Hepcidin (HEPC HUMAN; UniProt ID P81172), Adenylyl cyclase-associated protein 1 (CAP1 HUMAN; UniProt ID Q01518), RNA-binding protein EWS (EWS HUMAN; UniProt ID Q01844), Hepatocyte growth factor activator (HGFA HUMAN; UniProt ID Q04756), Prolow-density lipoprotein receptor-related protein 1 (LRP1 HUMAN; Q07954), Fibrinogen-like protein 1 (FGL1 HUMAN; Q08830), Nexilin (NEXN_HUMAN; UniProt ID Q0ZGT2), Aspartyl / asparaginyl beta-hydroxylase (ASPH_HUMAN;UniProt ID Q 12797), Interleukin enhancer-binding factor 2 (ILF2 HUMAN; UniProt ID Q 12905), Interleukin enhancer-binding factor 3 (ILF3 HUMAN; UniProt ID Q 12906), Multimerin-1 (MMRN1 HUMAN; UniProt ID Q13201), Protocadherin Fat 1 (FAT1 HUMAN; UniProt ID Q14517), Latent-transforming growth factor beta-binding protein 2 (LTBP2 HUMAN; UniProt ID Q14767), Periostin (POSTN_HUMAN; UniProt ID Q15063), Procollagen C-endopeptidase enhancer 1 (PCOC1_HUMAN; UniProt ID Q15113), Angiopoietin-1 (ANGP1_HUMAN; UniProt ID Q15389), Myosin light chain kinase, smooth muscle (MYLK HUMAN; UniProt ID Q15746), Sushi, von Willebrand factor type A, EGF and pentraxin domain-containing protein 1 (SVEP1 HUMAN; UniProt ID Q4LDE5), Phospholipase Al member A (PLA1A HUMAN; UniProt ID Q53H76), Transport and Golgi organization protein 1 homolog (TGO1 HUMAN; UniProt ID Q5JRA6), Isoaspartyl peptidase / L-asparaginase (ASGL1 HUMAN; UniProt ID Q7L266), Target of Nesh-SH3 (TARSH HUMAN; UniProt ID Q7Z7G0), Out at first protein homolog (OAF HUMAN; UniProt ID Q86UD1), Adipocyte enhancer-binding protein 1 (AEBP1 HUMAN; UniProt ID Q8IUX7), Secreted frizzled-related protein 1 (SFRP1 HUMAN; UniProt ID Q8N474), Ubiquitincarboxyl-terminal hydrolase MINDY-1 (MINY1 HUMAN; UniProt ID Q8N5J2), BPI fold-containing family B member 1 (BPIB1 HUMAN; UniProt ID Q8TDL5), Fas-binding factor 1 (FBF1 HUMAN; UniProt ID Q8TES7), Cell migration-inducing and hyaluronan-binding protein (CEMIP HUMAN; UniProt ID Q8WUJ3), Intelectin-1 (ITLN1 HUMAN; UniProt ID Q8WWA0), Transmembrane protein 40 (TMM40 HUMAN; UniProt ID Q8WWA1), Stonin-2 (STON2 HUMAN; UniProt ID Q8WXE9), Proteoglycan 4 (PRG4 HUMAN; UniProt ID Q92954), Phospholipid phosphatase-related protein type 2 (PLPR2 HUMAN; UniProt ID Q96GM1), Beta-Ala-His dipeptidase (CNDP1 HUMAN; UniProt ID Q96KN2), Proline-rich acidic protein 1 (PRAP1 HUMAN; UniProt ID Q96NZ9), Histone H4-like protein type G (H4G HUMAN; UniProt ID Q99525), Collagen alpha- 1 (XII) chain (COCA 1 HUMAN; UniProt ID Q99715), Growth / differentiation factor 15 (GDF15 HUMAN; UniProt ID Q99988), Sphingosine- 1- phosphate phosphatase 1 (SGPP1 HUMAN; UniProt ID Q9BX95), Complement factor H-related protein 5 (FHR5 HUMAN; UniProt ID Q9BXR6), Multimerin-2 (MMRN2 HUMAN; UniProt ID Q9H8L6), Spondin-1 (SPON1 HUMAN; UniProt ID Q9HCB6), Prefoldin subunit 4 (PFD4 HUMAN; UniProt ID Q9NQP4), Matrix-remodeling-associated protein 5 (MXRA5_HUMAN; UniProt ID Q9NR99), NAD- dependent protein deacetylase sirtuin-3, mitochondrial (SIR3 HUMAN; UniProt ID Q9NTG7), Dipeptidyl peptidase 3 (DPP3 HUMAN; UniProt ID Q9NY33), Tubulin alpha-8 chain (TBA8 HUMAN; UniProt ID Q9NY65), EH domain-containing protein 3 (EHD3_HUMAN; UniProt ID Q9NZN3), Protein Z-dependent protease inhibitor (ZPI HUMAN; UniProt ID Q9UK55), Angiopoietin-related protein 2 (ANGL2_HUMAN; UniProt ID Q9UKU9), Procollagen C-endopeptidase enhancer 2 (PCOC2_HUMAN; UniProt ID Q9UKZ9), Neurogenic locus notch homolog protein 3 (NOTC3_HUMAN; UniProt ID Q9UM47), Protein kinase C and casein kinase substrate in neurons protein 2 (PACN2_HUMAN; UniProt ID Q9UNF0), or Beta-l,3-N-acetylglucosaminyltransferase radical fringe (RFNG_HUMAN; UniProt ID Q9Y644).
[0171] A biomarker may include Immunoglobulin kappa variable 2-28 (KV228_HUMAN; UniProt ID A0A075B6P5). A biomarker may include A biomarker may include Immunoglobulin heavy variable 3-21 (HV321 HUMAN; UniProt ID A0A0B4J1V1). A biomarker may include A biomarker may include Protein phosphatase 1 regulatory subunit 12A (MYPT1 HUMAN; UniProt ID 014974). A biomarker may include A biomarker may include Adenylate cyclase type 6 (ADCY6_HUMAN; UniProt ID 043306). A biomarker may include Thioredoxin-like protein 1 (TXNL1 HUMAN; UniProt ID 043396). A biomarker may include Phosphoribosyl pyrophosphate synthase-associated protein 2 (KPRB HUMAN; UniProt ID 060256). A biomarker may include Reticulon-3 (RTN3 HUMAN; UniProt ID 095197). A biomarker may include NADH dehydrogenase (NDUBA_HUMAN; UniProt ID 096000). A biomarker may include Prothrombin (THRB HUMAN; UniProt ID P00734). A biomarker may include Carbonic anhydrase 2 (CAH2 HUMAN; UniProt ID P00918). A biomarker may include Alpha- 1 -antitrypsin (A1AT HUMAN; UniProt ID P01009). A biomarker may include Alpha- 1 -antichymotrypsin (AACT HUMAN; UniProt ID PO 1011). A biomarker may include Alpha-2 -macroglobulin (A2MG_HUMAN; UniProt ID P01023). A biomarker may include Complement C3 (C03 HUMAN; UniProt ID P01024). A biomarker may include Immunoglobulin lambda variable 1-47 (LV147 HUMAN; UniProt ID P01700). A biomarker may include Immunoglobulin lambdavariable 3-19 (LV319 HUMAN; UniProt ID P01714). A biomarker may include Immunoglobulin lambda variable 3-25 (LV325_HUMAN; UniProt ID P01717). A biomarker may include Polymeric immunoglobulin receptor (PIGR HUMAN; UniProt ID P01833). A biomarker may include Immunoglobulin heavy constant gamma 2 (IGHG2 HUMAN; UniProt ID P01859). A biomarker may include Immunoglobulin heavy constant alpha 1 (IGHA1 HUMAN; UniProt ID P01876). A biomarker may include Apolipoprotein E (APOE HUMAN; UniProt ID P02649). A biomarker may include Fibrinogen beta chain (FIBB HUMAN; UniProt ID P02675). A biomarker may include Band 3 anion transport protein (B3AT_HUMAN; UniProt ID P02730). A biomarker may include Complement component C9 (C09 HUMAN; UniProt ID P02748). A biomarker may include Protein AMBP (AMBP HUMAN; UniProt ID P02760). A biomarker may include Alpha-l-acid glycoprotein 1 (A1AG1 HUMAN; UniProt ID P02763). A biomarker may include Alpha-2 -HS-glycoprotein (FETUA_HUMAN; UniProt ID P02765). A biomarker may include Albumin (ALBU HUMAN; UniProt ID P02768). A biomarker may include Serotransferrin (TRFE HUMAN; UniProt ID P02787). A biomarker may include Interstitial collagenase (MMP1 HUMAN; UniProt ID P03956). A biomarker may include Apolipoprotein B-100 (APOB HUMAN; UniProt ID P04114). A biomarker may include Superoxide dismutase (SODM_HUMAN; UniProt ID P04179). A biomarker may include von Willebrand factor (VWF_HUMAN; UniProt ID P04275). A biomarker may include Protein S100-A8 (S10A8_HUMAN; UniProt ID P05109). A biomarker may include Plasma serine protease inhibitor (IPSP HUMAN; UniProt ID P05154). A biomarker may include Complement factor I (CFAI HUMAN; UniProt ID P05156). A biomarker may include Coagulation factor XIII B chain (F13B HUMAN; UniProt ID P05160). A biomarker may include Myeloperoxidase (PERM HUMAN; UniProt ID P05164). A biomarker may include Complement C2 (CO2_HUMAN; UniProt ID P06681). A biomarker may include Protein S100-A9 (S10A9_HUMAN; UniProt ID P06702). A biomarker may include Vitamin K-dependent protein S (PROS HUMAN; UniProt ID P07225). A biomarker may include Protein disulfide-isomerase (PDIA1 HUMAN; UniProt ID P07237). A biomarker may include Calpain-1 catalytic subunit (CAN1_HUMAN; UniProt ID P07384). A biomarker may include Fumarate hydratase. A biomarker may include mitochondrial (FUMH HUMAN; UniProt ID UniProt ID P07954). A biomarker may include Thrombospondin- 1 (TSP1 HUMAN; UniProt ID UniProt ID P07996). A biomarker may include Alpha-2 -antiplasmin (A2AP HUMAN; UniProt ID UniProt ID P08697). A biomarker may include Complement C4-A (CO4A HUMAN; UniProt ID UniProt ID P0C0L4). A biomarker may include Osteopontin (OSTP HUMAN; UniProt ID UniProt ID P10451). A biomarker may include Complement component C7 (CO7 HUMAN; UniProt ID UniProt ID P10643). A biomarker may include Tissue factor pathway inhibitor (TFPI1 HUMAN; UniProt ID UniProt ID P10646). A biomarker may include Protein 4.1 (EPB41 HUMAN ; UniProt ID UniProt ID P11171). A biomarker may include Coagulation factor V (FA5 HUMAN; UniProt ID UniProt ID P12259). A biomarker may include Versican core protein (CSPG2 HUMAN; UniProt ID UniProt ID P13611). A biomarker may include Beta-enolase (ENOB HUMAN; UniProt ID UniProt ID P13929). A biomarker may include Nidogen-1 (NID1 HUMAN; UniProt ID UniProt ID P14543). A biomarker may include Phospholipase A2. A biomarker may include membrane associated ( PA2GA HUMAN; UniProt ID UniProt ID P14555). A biomarker may includeBeta-galactoside alpha-2.6-sialyltransferase 1 (SIAT1 HUMAN; UniProt ID UniProt ID P15907). A biomarker may include Ankyrin-1 (ANK1 HUMAN; UniProt ID UniProt ID P16157). A biomarker may include Insulin-like growth factor-binding protein 2 (IBP2 HUMAN; UniProt ID UniProt ID P18065). A biomarker may include Uipopoly saccharide -binding protein (LBP HUMAN; UniProt ID UniProt ID P18428). A biomarker may include Syndecan-1 (SDC1 HUMAN ; UniProt ID UniProt ID P18827). A biomarker may include Alpha-l-acid glycoprotein 2 (A1AG2 HUMAN; UniProt ID UniProt ID P19652). A biomarker may include Inter-alpha-trypsin inhibitor heavy chain H2 (ITIH2 HUMAN; UniProt ID UniProt ID P19823). A biomarker may include Pregnancy zone protein (PZP HUMAN; UniProt ID UniProt ID P20742). A biomarker may include Collagen alpha-l(V) chain (C05A1 HUMAN; UniProt ID UniProt ID P20908). A biomarker may include Tenascin-X (TENX_HUMAN; UniProt ID P22105). A biomarker may include Fibulin-1 (FBUN 1 HUMAN; UniProt ID P23142). A biomarker may include Tryptophan— tRNA ligase. A biomarker may include cytoplasmic (SYWC_HUMAN; UniProt ID P23381). A biomarker may include Protein-lysine 6-oxidase (LYOX HUMAN; UniProt ID P28300). A biomarker may include Proprotein convertase subtilisin / kexin type 6 (PCSK6_HUMAN; UniProt ID P29122). A biomarker may include Flavin reductase (BLVRB HUMAN; UniProt ID P30043). A biomarker may include Carbamoyl-phosphate synthase (CPSM HUMAN; UniProt ID P31327). A biomarker may include Cadherin-5 (CADH5 HUMAN; UniProt ID P33151). A biomarker may include Ribonuclease 4 (RNAS4_HUMAN; UniProt ID P34096). A biomarker may include Pulmonary surfactant-associated protein D (SFTPD HUMAN; UniProt ID P35247). A biomarker may include Serum amyloid A-4 protein (SAA4_HUMAN; UniProt ID P35542). A biomarker may include Insulin-like growth factor-binding protein complex acid labile subunit (ALS HUMAN; UniProt ID P35858). A biomarker may include Chitinase-3- like protein 1 (CH3U1_HUMAN; UniProt ID P36222). A biomarker may include RNA -binding motif protein. A biomarker may include X chromosome (RBMX HUMAN; UniProt ID P38159). A biomarker may include Collagen alpha- 1 (XVIII) chain (C0IA1 HUMAN; UniProt ID P39060). A biomarker may include Trifunctional enzyme subunit alpha. A biomarker may include mitochondrial (ECHA HUMAN; UniProt ID P40939). A biomarker may include Fatty acid synthase (FAS HUMAN; UniProt ID P49327). A biomarker may include T-complex protein 1 subunit gamma (TCPG HUMAN; UniProt ID P49368). A biomarker may include Thrombospondin-3 (TSP3 HUMAN; UniProt ID P49746). A biomarker may include Cartilage oligomeric matrix protein (COMP HUMAN; UniProt ID P49747). A biomarker may include Thimet oligopeptidase (TH0P1 HUMAN; UniProt ID P52888). A biomarker may include C-C motif chemokine 18 (CCL18 HUMAN; UniProt ID P55774). A biomarker may include Triosephosphate isomerase (TPIS HUMAN; UniProt ID P60174). A biomarker may include Histone H4 (H4 HUMAN; UniProt ID P62805). A biomarker may include Histone H2B type 1-C / E / F / G / I (H2B1C HUMAN; UniProt ID P62807). A biomarker may include Tubulin beta-4B chain (TBB4B HUMAN; UniProt ID P68371). A biomarker may include Reelin (RELN HUMAN; UniProt ID P78509). A biomarker may include Phosphatidylinositol-glycan-specific phospholipase D (PHLD HUMAN; UniProt ID P80108). A biomarker may include Protein S 100-A12 (S 10AC HUMAN; UniProt ID P80511). A biomarker may include Hepcidin (HEPC HUMAN; UniProt ID P81172). A biomarker may include Adenylyl cyclase-associated protein 1(CAP1_HUMAN; UniProt ID Q01518). A biomarker may include RNA-binding protein EWS(EWS HUMAN; UniProt ID Q01844). A biomarker may include Hepatocyte growth factor activator (HGFA HUMAN; UniProt ID Q04756). A biomarker may include Prolow -density lipoprotein receptor- related protein 1 (LRP1 HUMAN; UniProt ID Q07954). A biomarker may include Fibrinogen-like protein 1 (FGL1_HUMAN; UniProt ID Q08830). A biomarker may include Nexilin (NEXN_HUMAN; UniProt ID Q0ZGT2). A biomarker may include Aspartyl / asparaginyl beta-hydroxylase (ASPH HUMAN; UniProt ID QI 2797). A biomarker may include Interleukin enhancer-binding factor 2 (ILF2 HUMAN; UniProt ID Q12905). A biomarker may include Interleukin enhancer-binding factor 3 (ILF3 HUMAN; UniProt ID Q12906). A biomarker may include Multimerin-1 (MMRN1_HUMAN; UniProt ID Q13201). A biomarker may include Protocadherin Fat 1 (FAT1 HUMAN; UniProt ID Q14517). A biomarker may include Latent- transforming growth factor beta-binding protein 2 (LTBP2 HUMAN; UniProt ID QI 4767). A biomarker may include Periostin (POSTN HUMAN; UniProt ID Q15063). A biomarker may include Procollagen C- endopeptidase enhancer 1 (PC0C1 HUMAN; UniProt ID Q15113). A biomarker may include Angiopoietin-1 (ANGP1_HUMAN; UniProt ID Q15389). A biomarker may include Myosin light chain kinase. A biomarker may include smooth muscle (MYLK HUMAN; UniProt ID Q15746). A biomarker may include Sushi. A biomarker may include von Willebrand factor type A. A biomarker may include EGF and pentraxin domain-containing protein 1 (SVEP1 HUMAN; UniProt ID Q4LDE5). A biomarker may include Phospholipase Al member A (PLA1A HUMAN; UniProt ID Q53H76). A biomarker may include Transport and Golgi organization protein 1 homolog (TG01 HUMAN; UniProt ID Q5JRA6). A biomarker may include Isoaspartyl peptidase / L-asparaginase (ASGL1 HUMAN; UniProt ID Q7L266). A biomarker may include Target of Nesh-SH3 (TARSH HUMAN; UniProt ID Q7Z7G0). A biomarker may include Out at first protein homolog (OAF HUMAN; UniProt ID Q86UD1). A biomarker may include Adipocyte enhancer-binding protein 1 (AEBP1 HUMAN; UniProt ID Q8IUX7). A biomarker may include Secreted frizzled-related protein 1 (SFRP1 HUMAN; UniProt ID Q8N474). A biomarker may include Ubiquitin carboxyl-terminal hydrolase MINDY-1 (MINY1_HUMAN; UniProt ID Q8N5J2). A biomarker may include BPI fold-containing family B member 1 (BPIB1 HUMAN; UniProt ID Q8TDL5). A biomarker may include Fas-binding factor 1 (FBF1 HUMAN; UniProt ID Q8TES7). A biomarker may include Cell migrationinducing and hyaluronan-binding protein (CEMIP HUMAN; UniProt ID Q8WUJ3). A biomarker may include Intelectin-1 (ITLN1 HUMAN; UniProt ID Q8WWA0). A biomarker may include Transmembrane protein 40 (TMM40_HUMAN; UniProt ID Q8WWA1). A biomarker may include Stonin-2 (STON2_HUMAN; UniProt ID Q8WXE9). A biomarker may include Proteoglycan 4 (PRG4_HUMAN; UniProt ID Q92954). A biomarker may include Phospholipid phosphatase-related protein type 2 (PLPR2 HUMAN; UniProt ID Q96GM1). A biomarker may include Beta-Ala-His dipeptidase (CNDP1_HUMAN; UniProt ID Q96KN2). A biomarker may include Proline-rich acidic protein 1 (PRAP1 HUMAN; UniProt ID Q96NZ9). A biomarker may include Histone H4-like protein type G (H4G HUMAN; UniProt ID Q99525). A biomarker may include Collagen alpha-l(XII) chain (COCA1 HUMAN; UniProt ID Q99715). A biomarker may include Growth / differentiation factor 15 (GDF15 HUMAN; UniProt ID Q99988). A biomarker may include Sphingosine-l-phosphate phosphatase 1(SGPP1 HUMAN; UniProt ID Q9BX95). A biomarker may include Complement factor H-related protein 5 (FHR5 HUMAN; UniProt ID Q9BXR6). A biomarker may include Multimerin-2 (MMRN2 HUMAN; UniProt ID Q9H8U6). A biomarker may include Spondin-1 (SP0N1 HUMAN; UniProt ID Q9HCB6). A biomarker may include Prefoldin subunit 4 (PFD4 HUMAN; UniProt ID Q9NQP4). A biomarker may include Matrix-remodeling-associated protein 5 (MXRA5_HUMAN; UniProt ID Q9NR99). A biomarker may include NAD-dependent protein deacetylase sirtuin-3. A biomarker may include mitochondrial(SIR3 HUMAN; UniProt ID Q9NTG7). A biomarker may include Dipeptidyl peptidase 3 (DPP3 HUMAN; UniProt ID Q9NY33). A biomarker may include Tubulin alpha-8 chain (TBA8 HUMAN; UniProt ID Q9NY65). A biomarker may include EH domain-containing protein 3 (EHD3 HUMAN; UniProt ID Q9NZN3). A biomarker may include Protein Z-dependent protease inhibitor (ZPI_HUMAN; UniProt ID Q9UK55). A biomarker may include Angiopoietin-related protein 2 (ANGL2_HUMAN; UniProt ID Q9UKU9). A biomarker may include Procollagen C-endopeptidase enhancer 2 (PC0C2_HUMAN; UniProt ID Q9UKZ9). A biomarker may include Neurogenic locus notch homolog protein 3 (N0TC3_HUMAN; UniProt ID Q9UM47). A biomarker may include Protein kinase C and casein kinase substrate in neurons protein 2 (PACN2_HUMAN; UniProt ID Q9UNF0). A biomarker may include Beta-1.3-N- acetylglucosaminyltransferase radical fringe (RFNG_HUMAN; UniProt ID Q9Y644). A fragment of any of these peptides may be used. Any of the forementioned peptide or protein biomarkers (or combination of said biomarkers) may be useful for identifying a presence, absence, or likelihood of a cancer described herein.
[0172] Any of the following protein biomarkers may be useful in detecting, identifying, or evaluating a presence, absence, or likelihood of cancer A2M, ABI3BP, ADCY6, AEBP1, AHSG, ALB, AMBP, ANGPT1, ANGPTL2, ANK1, APOB, APOE, ASPH, ASRGL1, BLVRB, BPIFB1, C2, C3, C4A, C7, C9, CA2, CAP1, CAPN1, CCL18, CCT3, CDH5, CEMIP, CFHR5, CFI, CHI3L1, CNDP1, COL12A1, COL18A1, COL5A1, COMP, CPS1, DPP3, EHD3, ENO3, EPB41, EWSR1, F13B, F2, F5, FASN, FAT1, FBLN1, FGB, FGL1, FH, GDF15, GPLD1, H2BC10, H2BC4, H2BC6, H2BC7, H2BC8, H4C1, H4C11, H4C12, H4C13, H4C14, H4C15, H4C16, H4C2, H4C3, H4C4, H4C5, H4C6, H4C7, H4C8, H4C9, HADHA, HAMP, HGFAC, IGFALS, IGFBP2, IGHA1, IGHG2, IGHV3-21, IGKV2-28, IGLV1-47, IGLV3-19, IGLV3-25, ILF2, ILF3, ITIH2, ITLN1, LBP, LOX, LRP1, LTBP2, MIA3, MINDY 1, MMP1, MMRN1, MMRN2, MPO, MXRA5, MYLK, NDUFB10, NEXN, NIDI, NOTCH3, OAF, 0RM1, 0RM2, P4HB, PACSIN2, PCOLCE, PCOLCE2, PCSK6, PFDN4, PIGR, PLA1A, PLA2G2A, POSTN, PPP1R12A, PRAP1, PRG4, PROS1, PRPSAP2, PZP, RBMX, RELN, RFNG, RNASE4, RTN3, S100A12, S100A8, S100A9, SAA4, SDC1, SERPINA1, SERPINA10, SERPINA3, SERPINA5, SERPINF2, SFRP1, SFTPD, SIRT3, SLC4A1, SOD2, SPON1, SPP1, ST6GAL1, STON2, SVEP1, TF, TFPI, THBS1, THBS3, THOP1, TMEM40, TNXB, TPI1, TUBA8, TUBB4B, TXNL1, VCAN, VWF, or WARSI. Any of the following protein biomarkers may be useful in detecting, identifying, or evaluating a presence, absence, or likelihood of cancer: KV228 HUMAN, HV321 HUMAN, MYPT1 HUMAN, ADCY6 HUMAN, TXNL1 HUMAN, KPRB HUMAN, RTN3 HUMAN, NDUBA HUMAN, THRB HUMAN, CAH2 HUMAN, Al AT HUMAN, AACT HUMAN, A2MG HUMAN, CO3 HUMAN, LV 147 HUMAN, LV319 HUMAN, LV325 HUMAN, PIGR HUMAN, IGHG2 HUMAN, IGHA1 HUMAN,APOE HUMAN, FIBB HUMAN, B3AT HUMAN, C09 HUMAN, AMBP HUMAN, A 1 AG 1 HUMAN, FETUA HUMAN, ALBU HUMAN, TRFE HUMAN, MMP1 HUMAN, APOB HUMAN,SODM HUMAN, VWF HUMAN, S10A8 HUMAN, IPSP HUMAN, CFAI HUMAN, F13B HUMAN, PERM HUMAN, C02 HUMAN, S10A9 HUMAN, PROS HUMAN, PDIA1 HUMAN,CAN 1 HUMAN, FUMH HUMAN, TSP 1 HUMAN, A2AP HUMAN, CO4A HUMAN,OSTP HUMAN, C07 HUMAN, TFPI1 HUMAN, EPB41 HUMAN, FA5 HUMAN, CSPG2 HUMAN, ENOB HUMAN, NIDI HUMAN, PA2GA HUMAN, SIAT1 HUMAN, ANK1 HUMAN,IBP2 HUMAN, LBP HUMAN, SDC1 HUMAN, A1AG2 HUMAN, ITIH2 HUMAN, PZP HUMAN, CO5A1 HUMAN, TENX HUMAN, FBLN1 HUMAN, SYWC HUMAN, LYOX HUMAN,PCSK6 HUMAN, BLVRB HUMAN, CPSM HUMAN, CADH5 HUMAN, RNAS4 HUMAN,SFTPD HUMAN, SAA4 HUMAN, ALS HUMAN, CH3 LI HUMAN, RBMX HUMAN,COIA1 HUMAN, ECHA HUMAN, FAS HUMAN, TCPG HUMAN, TSP3 HUMAN, COMP HUMAN, THOP1 HUMAN, CCL 18 HUMAN, TPIS HUMAN, H4 HUMAN, H2B1C HUMAN,TBB4B HUMAN, RELN HUMAN, PHLD HUMAN, S10AC HUMAN, EPC HUMAN,CAP 1 HUMAN, EWS HUMAN, HGFA HUMAN, LRP1 HUMAN, FGL1 HUMAN, NEXN HUMAN, ASPH HUMAN, ILF2 HUMAN, ILF3 HUMAN, MMRN1 HUMAN, FAT 1 HUMAN,LTBP2 HUMAN, POSTN HUMAN, PCOC1 HUMAN, ANGP1 HUMAN, MYLK HUMAN, SVEP1 HUMAN, PLA1A HUMAN, TGO1 HUMAN, ASGL1 HUMAN, TARSH HUMAN, OAF HUMAN, AEBP1 HUMAN, SFRP1 HUMAN, MINY1 HUMAN, BPIB1 HUMAN, FBF1 HUMAN, CEMIP HUMAN, ITLN1 HUMAN, TMM40 HUMAN, STON2 HUMAN, PRG4 HUMAN, PLPR2 HUMAN, CNDP1 HUMAN, PRAP1 HUMAN, H4G HUMAN, COCA 1 HUMAN, GDF 15 HUMAN, SGPP1 HUMAN, FHR5 HUMAN, MMRN2 HUMAN, SPON1 HUMAN, PFD4 HUMAN, MXRA5 HUMAN, SIR3 HUMAN, DPP3 HUMAN, TBA8 HUMAN, EHD3 HUMAN, ZPI HUMAN, ANGL2 HUMAN, PCOC2 HUMAN,N0TC3_HUMAN, PACN2_HUMAN, or RFNG_HUMAN. The cancer may include a lung cancer such as NSCLC. Any number or combination of the aforementioned biomarkers may be used. For example, 1, 2 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more of said proteins may be used.
[0173] Some aspects include a peptide transition. Some aspects include the use of multiple peptide transitions. For example, measurements of multiple peptide transitions from a biofluid sample may be useful in a diagnostic method, or in any multi -omic method.
[0174] In some aspects, the multi -omics data comprises measurements of at least about 10 peptides or protein groups, at least about 15 peptides or protein groups, at least about 20 peptides or protein groups, at least about 25 peptides or protein groups, at least about 30 peptides or protein groups, at least about 35 peptides or protein groups, at least about 40 peptides or protein groups, at least about 45 peptides or protein groups, at least about 50 peptides or protein groups, at least about 75 peptides or protein groups, at least about 100 peptides or protein groups, at least about 250 peptides or protein groups, at least about 500 peptides or protein groups, at least about 1,000 peptides or protein groups, at least about 2,500 peptides orprotein groups, at least about 5,000 peptides or protein groups, at least about 10,000 peptides or protein groups, at least about 15,000 peptides or protein groups, or at least about 20,000 peptides or protein groups. In some aspects, the protein data comprises measurements of no greater than 10 peptides or protein groups, no greater than 15 peptides or protein groups, no greater than 20 peptides or protein groups, no greater than 25 peptides or protein groups, no greater than 30 peptides or protein groups, no greater than 35 peptides or protein groups, no greater than 40 peptides or protein groups, no greater than 45 peptides or protein groups, no greater than 50 peptides or protein groups, no greater than 75 peptides or protein groups, no greater than 100 peptides or protein groups, no greater than 250 peptides or protein groups, no greater than 500 peptides or protein groups, no greater than 1,000 peptides or protein groups, no greater than 2,500 peptides or protein groups, no greater than 5,000 peptides or protein groups, no greater than 10,000 peptides or protein groups, no greater than 15,000 peptides or protein groups, or no greater than 20,000 peptides or protein groups.
[0175] A protein may also include a post-translational modification (PTM). An example of a PTM may include glycosylation. Proteins or peptides may include glycoproteins or glycopeptides. A protein may include a glycoprotein. A peptide may include a glycopeptide. An example of a PTM may include phosphorylation. Proteins or peptides may include phosphoproteins or phosphopeptides. A protein may include a phosphoprotein. A peptide may include a phosphopeptide. An example of a PTM may include carboxyamidomethylation. Proteins or peptides may include carbamidomethyl proteins or carbamidomethyl peptides. A protein may include a carbamidomethyl protein. A peptide may include a carbamidomethyl peptide. An example of a PTM may include oxidation or hydroxylation. Proteins or peptides may include oxidated or hydroxylated proteins or oxidated or hydroxylated peptides. A protein may include an oxidated protein. A protein may include a hydroxylated protein. A peptide may include an oxidated peptide, peptide may include a hydroxylated peptide.
[0176] Proteomic data may be generated by any of a variety of methods. Generating proteomic data may include using a detection reagent that binds to a peptide or protein and yields a detectable signal. After use of a detection reagent that binds to a peptide or protein and yields a detectable signal, a readout may be obtained that is indicative of the presence, absence or amount of the protein or peptide. Generating proteomic data may include concentrating, filtering, or centrifuging a sample.
[0177] Proteomic data may be generated using mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, a lateral flow assay, an immunoassay, an enzyme-linked immunosorbent assay, a western blot, a dot blot, or immunostaining, or a combination thereof. Some examples of methods for generating proteomic data include using mass spectrometry, a protein chip, or a reverse-phased protein microarray. Proteomic data may also be generated using an immunoassay such as an enzyme-linked immunosorbent assay, western blot, dot blot, or immunohistochemistry assay. Generating proteomic data may involve use of an immunoassay panel.
[0178] One way of obtaining proteomic data includes use of mass spectrometry. An example of a mass spectrometry method includes use of high resolution, two-dimensional electrophoresis to separate proteins from different samples in parallel, followed by selection or staining of differentially expressed proteins to be identified by mass spectrometry. Another method uses stable isotope tags to differentially label proteinsfrom two different complex mixtures. The proteins within a complex mixture may be labeled isotopically and then digested to yield labeled peptides. Then the labeled mixtures may be combined, and the peptides may be separated by multidimensional liquid chromatography and analyzed by tandem mass spectrometry. A mass spectrometry method may include use of liquid chromatography-mass spectrometry (LC-MS), a technique that may combine physical separation capabilities of liquid chromatography (e.g., HPLC) with mass spectrometry.
[0179] Proteins may be enriched prior to assaying or measuring them. The enrichment may enrich one set of proteins and not another set, or may enrich a single protein and not another protein. Enrichment may be obtained through the use of an affinity reagent, for example by incubating the affinity reagent with a sample prior to measuring proteins in the sample. The affinity reagent may include an antibody. The affinity reagent may include a particle such as a nanoparticle. Proteins may be adsorbed to the affinity reagent, separated from the rest of the sample, and then assayed by using a proteomic assay described herein.
[0180] Generating proteomic data may include contacting a sample with particles such that the particles adsorb biomolecules comprising proteins. The adsorbed proteins may be part of a biomolecule corona. The adsorbed proteins may be measured or identified in generating the proteomic data.
[0181] Generating proteomic data may include the use of known amounts internal reference proteins. The reference proteins may be labeled. The label may include an isotopic label. Generating proteomic data may include the use of known amounts of isotopically labeled internal reference proteins (referred to as “PiQuant”). The internal reference proteins may be spiked into a sample. The internal reference proteins may be used to identify mass spectra of individual endogenous proteins. The internal reference proteins may be used as standards for determining amounts of the individual endogenous proteins. Proteomic measurements may be generated based on amounts of proteins added into a sample of the one or more biofluid samples. Proteomic measurements may be generated based on amounts of labeled proteins added into a sample of the one or more biofluid samples. In some aspects, the proteomic data can include spatial proteomic data, where the spatial proteomic data can include detecting and quantifying protein subcellular localization in a cell. Spatial proteomic data can be obtained via microscopy, mass spectrometry and machine learning applications for data analysis. In some aspects, the spatial proteomic data is in situ proteomic data.RNA or Transcriptomic Data
[0182] The data such as multi -omics data described herein may include RNA data such as transcript data or transcriptomic data. Transcriptomic data may involve data about nucleotide transcripts such as RNA. RNA data may comprise one or more isoform features. RNA data may comprise one or more intron sequencing features. Intron sequencing features may comprise immature transcripts, intron retention, or any combination thereof. In some examples, RNA intron sequencing features are used as input datasets into trained algorithms (e.g., machine learning models or classifiers) to find correlations between sequence composition and subject groups (e.g., patient groups). Examples of such patient groups include presence or absence of diseases or conditions, elevated or non-elevated risk of diseases or conditions, stages of diseases orconditions, subtypes of diseases or conditions, responders to treatment vs. non-responders to treatment, and progressors versus non-progressors. In some examples, feature matrices are generated to compare samples obtained from subjects with known conditions or characteristics. In some embodiments, samples are obtained from healthy subjects, or subjects who do not have any of the known indications and samples from patients known to have cancer. In an aspect, the disclosed systems and methods provide a classifier generated based on feature information derived from RNA intron sequence analysis from biological samples. The classifier forms part of a predictive engine for distinguishing groups in a population based on RNA intron sequence features identified in biological samples.
[0183] Examples of RNA include messenger RNA (mRNA), ribosomal RNA (rRNA), signal recognition particle (SRP) RNA, transfer RNA (tRNA), small nuclear RNA (snRNA), small nucleoar RNA (snoRNA), long noncoding RNA (IncRNA), microRNA (miRNA), noncoding RNA (ncRNA), or piwi-interacting RNA (piRNA), or a combination thereof. The RNA may include mRNA. The RNA may include miRNA. Transcriptomic data may be distinguished by subtype, where each subtype includes a different type of RNA or transcript. For example, mRNA data may be included in one subtype, and data for one or more types of small non-coding RNAs such as miRNAs or piRNAs may be included in another subtype. A miRNA may include a 5p miRNA or a 3p miRNA. In some embodiments, the RNA comprises fragmented RNA. In some embodiments, the RNA comprises partially degraded RNA. In some embodiments, the RNA comprises a microRNA or portion thereof. In some embodiments, the RNA comprises an RNA molecule or a fragmented RNA molecule (RNA fragments) selected from: a microRNA (miRNA or miR), a pre-miRNA, a pri- miRNA, a mRNA, a pre-mRNA, a viral RNA, a viroid RNA, a virusoid RNA, circular RNA (circRNA), a ribosomal RNA (rRNA), a transfer RNA (tRNA), a pre-tRNA, a long non-coding RNA (IncRNA), a small nuclear RNA (snRNA), a circulating RNA, a cell-free RNA, an exosomal RNA, a vector-expressed RNA, an RNA transcript, a synthetic RNA, and combinations thereof.
[0184] RNA data may include information on the presence, absence, or amount of various RNAs. For example, transcriptomic data may include amounts of RNAs. An RNA amount may be indicated as a concentration or number or RNA molecules, for example a concentration of an RNA in a biofluid. An RNA amount may be relative to another RNA or to another biomolecule. RNA data may include information on the presence of RNA. RNA data may include information on the absence of RNA. Aspects described in relation to RNA data may be relevant to transcript or transcriptome data, or vice versa.
[0185] RNA data generally includes data on a number of RNAs. For example, RNA data may include information on the presence, absence, or amount of 1,000 or more RNAs. In some cases, RNA data may include information on the presence, absence, or amount of 5,000 or more, 10,000 or more, or 20,000 or more RNAs. A transcript amount may include a copy number. RNA data may even include up to about 200,000 transcripts. RNA data may include a range of transcripts defined by any of the aforementioned numbers of RNAs or transcripts.
[0186] Transcriptomic data may include information on the presence, absence, or amount of various RNAs. For example, transcriptomic data may include amounts of RNAs. An RNA amount may be indicated as a concentration or number or RNA molecules, for example a concentration of an RNA in a biofluid. An RNAamount may be relative to another RNA or to another biomolecule. Transcriptomic data may include information on the presence of RNAs. Transcriptomic data may include information on the absence of RNA. Aspects described in relation to transcriptomic data may be relevant to transcript or RNA data, or vice versa.
[0187] Transcriptomic data generally includes data on a number of RNAs. For example, transcriptomic data may include information on the presence, absence, or amount of 1000 or more RNAs. In some cases, transcriptomic data may include information on the presence, absence, or amount of 5000, 10,000, 20,000, or more RNAs. A transcript amount may include a copy number. Transcriptomic data may even include up to about 200,000 transcripts. Transcriptomic data may include a range of transcripts defined by any of the aforementioned numbers of RNAs or transcripts. Some examples of mRNAs that may be included in transcriptomic data are shown herein. Some examples of microRNAs that may be included in transcriptomic data are shown herein.
[0188] Some examples of RNAs that may be used as biomarkers include mRNAs. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 of the RNAs may be useful as biomarkers, for example, in cancer assessment. A cancer may be a lung cancer, such as NSCLC. Any of the following RNAs may be useful as such (listed with gene name): C12orf57 (Ensembl ID ENST00000229281.6),ZNF391 (Ensembl ID ENST00000244576.9), HIP1R (Ensembl ID ENST00000253083.9), CRISP3 (Ensembl ID ENST00000263045.9), CD80 (Ensembl ID ENST00000264246.8), ILDR2 (Ensembl ID ENST00000271417.8), PRCC (Ensembl ID ENST00000271526.9), N4BP3 (Ensembl ID ENST00000274605.6), NARS2 (Ensembl ID ENST00000281038.10), JAZF1 (Ensembl ID ENST00000283928.10), ZNF276 (Ensembl ID ENST00000289816.9), ANKRD11 (Ensembl ID ENST00000301030.10), LDB2 (Ensembl ID ENST00000304523.10), TNK2 (Ensembl ID ENST00000333602.14), FOXD2 (Ensembl ID ENST00000334793.6), RPL14 (Ensembl ID ENST00000338970.10), SAXO2 (Ensembl ID ENST00000339465.5), FUT8 (Ensembl ID ENST00000342677.10), PPL (Ensembl ID ENST00000345988.7), PFKL (Ensembl ID ENST00000349048.9), IFT122 (Ensembl ID ENST00000349441.6), KAT5 (Ensembl ID ENST00000352980.8), C6orf89 (Ensembl ID ENST00000355190.7), SYT17 (Ensembl ID ENST00000355377.7), ELK4 (Ensembl ID ENST00000357992.9), PLXNB1 (Ensembl ID ENST00000358536.8), LINGO4 (Ensembl ID ENST00000368820.4), VPS72 (Ensembl ID ENST00000368892.9), CDK19 (Ensembl ID ENST00000368911.8), RUNX2 (Ensembl ID ENST00000371432.7), ERI3 (Ensembl ID ENST00000372257.7), SPOCK2 (Ensembl ID ENST00000373109.7), MAGT1 (Ensembl ID ENST00000373336.3), ITPRIPL2 (Ensembl ID ENST00000381440.5), ITGAD (Ensembl ID ENST00000389202.3), TRAV13-1 (Ensembl ID ENST00000390436.2), SNORA53 (Ensembl ID ENST00000391141.1), WDR45B (Ensembl ID ENST00000392325.9), M0XD1 (Ensembl ID ENST00000392401.3), SEPTIN6 (Ensembl ID ENST00000394610.7), TMEM106B (Ensembl ID ENST00000396668.8), HBA1 (Ensembl ID ENST00000397797.1), SFR1P1 (Ensembl ID ENST00000403720.2), FAM168B (Ensembl ID ENST00000409185.5), AFF3 (Ensembl ID ENST00000409236.6), NAB1 (Ensembl ID ENST00000409581.5), RPL23AP18 (Ensembl ID ENST00000416139.1), SCYL1 (Ensembl IDENST00000420247.6), RPL37A (Ensembl ID ENST00000420712.1), TBCEL (Ensembl ID ENST00000422003.6), CTD-2666L21.2 (Ensembl ID ENST00000428058.1), OPTN (Ensembl ID ENST00000430081.5), CSNK1E (Ensembl ID ENST00000431632.5), BACH1-IT2 (Ensembl ID ENST00000436529.2), OGA (Ensembl ID ENST00000439817.5), SATB2 (Ensembl ID ENST00000440919.1), BRPF3 (Ensembl ID ENST00000441730.2), TPT1 (Ensembl ID ENST00000442760.2), TRGC1 (Ensembl ID ENST00000443402.6), N4BP2L2 (Ensembl ID ENST00000446957.6), CXCR2 (Ensembl ID ENST00000449014.5), EDRF1-AS1 (Ensembl ID ENST00000449436.5), KIF28P (Ensembl ID ENST00000451123.5), IKZF2 (Ensembl ID ENST00000452786.2), CEP85 (Ensembl ID ENST00000453146.1), MFSD2B (Ensembl ID ENST00000453731.1), ZBP1 (Ensembl ID ENST00000453793.1), AGFG1 (Ensembl ID ENST00000458212.1), XIAP-AS1 (Ensembl ID ENST00000458331.1), RNF13 (Ensembl ID ENST00000459632.5), CAMSAP1 (Ensembl ID ENST00000460094.1), AMY2B (Ensembl ID ENST00000462971.1), TFRC (Ensembl ID ENST00000465288.5), SELP (Ensembl ID ENST00000466167.2), GOLGA2 (Ensembl ID ENST00000468488.1), RAB5A (Ensembl ID ENST00000469122.5), MBD5 (Ensembl ID ENST00000469438.1), GON4L (Ensembl ID ENST00000471341.5), ASB1 (Ensembl ID ENST00000473306.1), TTLL3 (Ensembl ID ENST00000473661.5), ABHD16A (Ensembl ID ENST00000474007.5), SLAIN1 (Ensembl ID ENST00000474663.5), HBB (Ensembl ID ENST00000475226.1), GTPBP1 (Ensembl ID ENST00000475959.1), MIDEAS (Ensembl ID ENST00000476562.5), SRRT (Ensembl ID ENST00000477529.1), SF1 (Ensembl ID ENST00000477596.1), RP11-397E7.1 (Ensembl ID ENST00000480094.3), TDP2 (Ensembl ID ENST00000480495.1), RP11-296A18.5 (Ensembl ID ENST00000480705.1), ETV5 (Ensembl ID ENST00000480706.1), ARL5A (Ensembl ID ENST00000487723.1), PDCD10 (Ensembl ID ENST00000487947.6), INTS6 (Ensembl ID ENST00000488009.1), CREM (Ensembl ID ENST00000490460.5), ATP6AP1 (Ensembl ID ENST00000491569.5), PTPN7 (Ensembl ID ENST00000492451.1), THOC5 (Ensembl ID ENST00000492707.5), PGD (Ensembl ID ENST00000493288.1), CLEC16A (Ensembl ID ENST00000494853.1), DGKD (Ensembl ID ENST00000495901.1), CDKN2A (Ensembl ID ENST00000497750.1), ZNF518B (Ensembl ID ENST00000500268.6), GNPDA1 (Ensembl ID ENST00000500692.6), NSUN2 (Ensembl ID ENST00000502932.1), CTC-338M12.4 (Ensembl ID ENST00000506340.2), OCIAD1 (Ensembl ID ENST00000508293.5), CPLANE1 (Ensembl ID ENST00000509849.5), PPIP5K2 (Ensembl ID ENST00000510672.1), SH3RF1 (Ensembl ID ENST00000510806.1), UBE2D2 (Ensembl ID ENST00000511691.1), MACROH2A1 (Ensembl ID ENST00000513268.1), RACK1 (Ensembl ID ENST00000514455.5), CHMP7 (Ensembl ID ENST00000517325.5), HNRNPH1 (Ensembl ID ENST00000519033.5), PDE7A (Ensembl ID ENST00000519626.1), LYN (Ensembl ID ENST00000520050.1), CLK4 (Ensembl ID ENST00000520878.5), HNRNPH1 (Ensembl ID ENST00000521790.5), GFRA2 (Ensembl ID ENST00000522071.2), PPP1R2B (Ensembl ID ENST00000522232.3), NSMAF (Ensembl ID ENST00000523106.5), ARHGAP32 (Ensembl ID ENST00000524655.5), DGKZ (Ensembl IDENST00000524869.1), TTC17 (Ensembl ID ENST00000525135.5), RPS3 (Ensembl IDENST00000526608.5), KNL1 (Ensembl ID ENST00000526913.5), RAB30 (Ensembl IDENST00000527633.6), SBF2 (Ensembl ID ENST00000528478.1), CYHR1 (Ensembl IDENST00000528558.1), PLPP5 (Ensembl ID ENST00000528814.1), KCNJ5 (Ensembl IDENST00000529694.6), HSPA8 (Ensembl ID ENST00000530391.1), ARHGAP27 (Ensembl IDENST00000532038.5), S0RL1 (Ensembl ID ENST00000532451.1), TAF6L (Ensembl IDENST00000532915.1), FUZ (Ensembl ID ENST00000533418.5), NUMA1 (Ensembl IDENST00000535111.5), AP000462.1 (Ensembl ID ENST00000540545.1), KXD1 (Ensembl IDENST00000540691.5), ARHGDIA (Ensembl ID ENST00000541078.6), ZBP1 (Ensembl IDENST00000541799.1), PSMD9 (Ensembl ID ENST00000544254.1), GPR19 (Ensembl IDENST00000545876.5), TAF1D (Ensembl ID ENST00000546088.5), COROIC (Ensembl ID ENST00000547170.1), RP11-386G11.10 (Ensembl ID ENST00000547387.1), ZNF606 (Ensembl IDENST00000547828.5), DTX1 (Ensembl ID ENST00000548759.1), DIP2B (Ensembl IDENST00000549620.5), ITGB7 (Ensembl ID ENST00000550743.6), CFAP54 (Ensembl IDENST00000554108.6), FOXN3 (Ensembl ID ENST00000555353.5), ADAM10 (Ensembl IDENST00000558733.5), SMAD3 (Ensembl ID ENST00000560402.1), ANKRD9 (Ensembl IDENST00000560748.5), RBMX (Ensembl ID ENST00000561733.5), GALNS (Ensembl IDENST00000562593.5), GABARAPL2 (Ensembl ID ENST00000563744.1), MAPK8IP3 (Ensembl IDENST00000566064.1), MVP (Ensembl ID ENST00000566554.1), ARID3B (Ensembl IDENST00000569680.1), RP11-314A20.5 (Ensembl ID ENST00000570493.2), FOXK2 (Ensembl IDENST00000570585.1), MRPS21P9 (Ensembl ID ENST00000571150.1), CORO7 (Ensembl IDENST00000572666.5), SRRM2 (Ensembl ID ENST00000574866.1), NOL11 (Ensembl IDENST00000577687.1), CDK5RAP3 (Ensembl ID ENST00000578663.1), NBPF15 (Ensembl IDENST00000578953.1), CBX3P2 (Ensembl ID ENST00000579647.1), PITPNC1 (Ensembl IDENST00000580974.5), TVP23B (Ensembl ID ENST00000581733.1), KLF14 (Ensembl IDENST00000583337.4), NFE2L1 (Ensembl ID ENST00000585291.5), GRN (Ensembl IDENST00000585348.1), MLX (Ensembl ID ENST00000585403.5), OAZ1 (Ensembl IDENST00000589361.2), RP11-686D22.7 (Ensembl ID ENST00000590824.1), NOSIP (Ensembl IDENST00000594932.5), RPS11 (Ensembl ID ENST00000600027.5), HCG18 (Ensembl IDENST00000602290.2), RP11-214K3.19 (Ensembl ID ENST00000602741.1), CTD-2589H19.6 (Ensembl IDENST00000607068.1), RP5-1098D14.1 (Ensembl ID ENST00000608190.1), SATB1-AS1 (Ensembl IDENST00000609327.2), CTA-217C2.2 (Ensembl ID ENST00000609432.1), NPRL3 (Ensembl IDENST00000610509.1), NBPF9 (Ensembl ID ENST00000611593.1), BIRC3 (Ensembl IDENST00000615299.4), MRTFA (Ensembl ID ENST00000618417.1), TRPM2 (Ensembl IDENST00000621064.1), LINC01620 (Ensembl ID ENST00000623295.1), RP11-44F14.10 (Ensembl IDENST00000623877.1), FAM223A (Ensembl ID ENST00000625031.1), SPDYE5 (Ensembl IDENST00000625065.3), RP11-227G15.12 (Ensembl ID ENST00000625127.3), CFLAR-AS1 (Ensembl IDENST00000625224.2), XPC-AS1 (Ensembl ID ENST00000626655.2), EP300 (Ensembl IDENST00000635083.1), MIR34AHG (Ensembl ID ENST00000635405.2), MIR34AHG (Ensembl ID ENST00000635687.1), CLCN2 (Ensembl ID ENST00000637538.1), STAT6 (Ensembl ID ENST00000640254.2), RP11-582H21.1 (Ensembl ID ENST00000643055.1), ETAA1 (Ensembl ID ENST00000644028.1), AUTS2 (Ensembl ID ENST00000644506.1), CSNK2A1 (Ensembl ID ENST00000645234.1), LRPAP1 (Ensembl ID ENST00000650182.1), NFKB1 (Ensembl ID ENST00000652569.1), FOXCUT (Ensembl ID ENST00000652712.1), FLJ40194 (Ensembl ID ENST00000659323.1), RP4-644F6.4 (Ensembl ID ENST00000662176.1), RP11-702N8.3 (Ensembl ID ENST00000663064.1), LINC00910 (Ensembl ID ENST00000663186.1), LINC01727 (Ensembl ID ENST00000663538.1), AC009404.2 (Ensembl ID ENST00000664381.1), EGLN1 (Ensembl ID ENST00000667629.1), ZFAS1 (Ensembl ID ENST00000667889.1), SNHG14 (Ensembl ID ENST00000667925.1), SNHG14 (Ensembl ID ENST00000669064.1), JAK1 (Ensembl ID ENST00000672179.1), DCTN2 (Ensembl ID ENST00000676646.1), SPECC1 (Ensembl ID ENST00000676737.1), DNMT1 (Ensembl ID ENST00000676868.1), NDUFAF1 (Ensembl ID ENST00000677477.1), HNRNPA1 (Ensembl ID ENST00000678077.1), ADAR (Ensembl ID ENST00000680472.1), SCARB1 (Ensembl ID ENST00000680926.1), OAS3 (Ensembl ID ENST00000681085.1), AHI1 (Ensembl ID ENST00000681477.1), P4HB (Ensembl ID ENST00000681693.1), HEXA (Ensembl ID ENST00000683228.1), or PTPN23 (Ensembl ID ENST00000683708.1). RS2 (Ensembl ID ENST00000281038.10). A biomarker may include JAZF1 (Ensembl ID ENST00000283928. 10). A biomarker may include ZNF276 (Ensembl IDENST00000289816.9). A biomarker may include ANKRD11 (Ensembl ID ENST00000301030.10). A biomarker may include LDB2 (Ensembl ID ENST00000304523.10). A biomarker may include TNK2 (Ensembl ID ENST00000333602. 14). A biomarker may include FOXD2 (Ensembl ID ENST00000334793.6). A biomarker may include RPL14 (Ensembl ID ENST00000338970. 10). A biomarker may include SAXO2 (Ensembl ID ENST00000339465.5). A biomarker may include FUT8 (Ensembl ID ENST00000342677. 10). A biomarker may include PPL (Ensembl ID ENST00000345988.7). A biomarker may include PFKL (Ensembl ID ENST00000349048.9). A biomarker may include IFT122 (Ensembl ID ENST00000349441.6). A biomarker may include KAT5 (Ensembl ID ENST00000352980.8). A biomarker may include C6orf89 (Ensembl ID ENST00000355190.7). A biomarker may include SYT17 (Ensembl ID ENST00000355377.7). A biomarker may include ELK4 (Ensembl ID ENST00000357992.9). A biomarker may include PLXNB1 (Ensembl ID ENST00000358536.8). A biomarker may include LINGO4 (Ensembl ID ENST00000368820.4). A biomarker may include VPS72 (Ensembl ID ENST00000368892.9). A biomarker may include CDK19 (Ensembl ID ENST00000368911.8). A biomarker may include RUNX2 (Ensembl ID ENST00000371432.7). A biomarker may include ERI3 (Ensembl ID ENST00000372257.7). A biomarker may include SPOCK2 (Ensembl ID ENST00000373109.7). A biomarker may include MAGT1 (Ensembl ID ENST00000373336.3). A biomarker may include ITPRIPL2 (Ensembl IDENST00000381440.5). A biomarker may include ITGAD (Ensembl ID ENST00000389202.3). A biomarker may include TRAV13-1 (Ensembl ID ENST00000390436.2). A biomarker may include SNORA53 (Ensembl ID ENST00000391141.1). A biomarker may include WDR45B (Ensembl IDENST00000392325.9). A biomarker may include M0XD1 (Ensembl ID ENST00000392401.3). A biomarker may include SEPTIN6 (Ensembl ID ENST00000394610.7). A biomarker may include TMEM106B (Ensembl ID ENST00000396668.8). A biomarker may include HBA1 (Ensembl ID ENST00000397797. 1). A biomarker may include SFR1P1 (Ensembl ID ENST00000403720.2). A biomarker may include FAM168B (Ensembl ID ENST00000409185.5). A biomarker may include AFF3 (Ensembl ID ENST00000409236.6). A biomarker may include NAB1 (Ensembl ID ENST00000409581.5). A biomarker may include RPL23AP18 (Ensembl ID ENST00000416139.1). A biomarker may include SCYL1 (Ensembl ID ENST00000420247.6). A biomarker may include RPL37A (Ensembl ID ENST00000420712. 1). A biomarker may include TBCEL (Ensembl ID ENST00000422003.6). A biomarker may include CTD-2666L21.2 (Ensembl ID ENST00000428058. 1). A biomarker may include OPTN (Ensembl ID ENST00000430081.5). A biomarker may include CSNK1E (Ensembl IDENST00000431632.5). A biomarker may include BACH1-IT2 (Ensembl ID ENST00000436529.2). A biomarker may include OGA (Ensembl ID ENST00000439817.5). A biomarker may include SATB2 (Ensembl ID ENST00000440919. 1). A biomarker may include BRPF3 (Ensembl ID ENST00000441730.2). A biomarker may include TPT1 (Ensembl ID ENST00000442760.2). A biomarker may include TRGC1 (Ensembl ID ENST00000443402.6). A biomarker may include N4BP2L2 (Ensembl ID ENST00000446957.6). A biomarker may include CXCR2 (Ensembl ID ENST00000449014.5). A biomarker may include EDRF1-AS1 (Ensembl ID ENST00000449436.5). A biomarker may include KIF28P (Ensembl ID ENST00000451123.5). A biomarker may include IKZF2 (Ensembl ID ENST00000452786.2). A biomarker may include CEP85 (Ensembl ID ENST00000453146.1). A biomarker may include MFSD2B (Ensembl ID ENST00000453731.1). A biomarker may include ZBP1 (Ensembl ID ENST00000453793.1). A biomarker may include AGFG1 (Ensembl ID ENST00000458212. 1). A biomarker may include XIAP- AS1 (Ensembl ID ENST00000458331.1). A biomarker may include RNF13 (Ensembl ID ENST00000459632.5). A biomarker may include CAMSAP1 (Ensembl ID ENST00000460094.1). A biomarker may include AMY2B (Ensembl ID ENST00000462971.1). A biomarker may include TFRC (Ensembl ID ENST00000465288.5). A biomarker may include SELP (Ensembl ID ENST00000466167.2). A biomarker may include G0LGA2 (Ensembl ID ENST00000468488.1). A biomarker may include RAB5A (Ensembl ID ENST00000469122.5). A biomarker may include MBD5 (Ensembl ID ENST00000469438. 1). A biomarker may include G0N4L (Ensembl ID ENST00000471341.5). A biomarker may include ASB1 (Ensembl ID ENST00000473306. 1). A biomarker may include TTLL3 (Ensembl ID ENST00000473661.5). A biomarker may include ABHD16A (Ensembl ID ENST00000474007.5). A biomarker may include SLAIN1 (Ensembl ID ENST00000474663.5). A biomarker may include HBB (Ensembl ID ENST00000475226. 1). A biomarker may include GTPBP1 (Ensembl ID ENST00000475959. 1). A biomarker may include MIDEAS (Ensembl ID ENST00000476562.5). A biomarker may include SRRT (Ensembl ID ENST00000477529.1). A biomarker may include SF1 (Ensembl ID ENST00000477596.1). A biomarker may include RP11-397E7.1 (Ensembl ID ENST00000480094.3). A biomarker may include TDP2 (Ensembl ID ENST00000480495. 1). A biomarker may include RP11-296A18.5 (Ensembl ID ENST00000480705. 1). A biomarker may include ETV5 (Ensembl ID ENST00000480706. 1). A biomarkermay include ARL5A (Ensembl ID ENST00000487723.1). A biomarker may include PDCD10 (Ensembl ID ENST00000487947.6). A biomarker may include INTS6 (Ensembl ID ENST00000488009. 1). A biomarker may include CREM (Ensembl ID ENST00000490460.5). A biomarker may include ATP6AP1 (Ensembl ID ENST00000491569.5). A biomarker may include PTPN7 (Ensembl ID ENST00000492451.1). A biomarker may include THOC5 (Ensembl ID ENST00000492707.5). A biomarker may include PGD (Ensembl ID ENST00000493288. 1). A biomarker may include CLEC16A (Ensembl ID ENST00000494853. 1). A biomarker may include DGKD (Ensembl ID ENST00000495901.1). A biomarker may include CDKN2A (Ensembl ID ENST00000497750. 1). A biomarker may include ZNF518B (Ensembl ID ENST00000500268.6). A biomarker may include GNPDA1 (Ensembl ID ENST00000500692.6). A biomarker may include NSUN2 (Ensembl ID ENST00000502932. 1). A biomarker may include CTC- 338M12.4 (Ensembl ID ENST00000506340.2). A biomarker may include OCIAD1 (Ensembl ID ENST00000508293.5). A biomarker may include CPLANE1 (Ensembl ID ENST00000509849.5). A biomarker may include PPIP5K2 (Ensembl ID ENST00000510672. 1). A biomarker may include SH3RF1 (Ensembl ID ENST00000510806. 1). A biomarker may include UBE2D2 (Ensembl ID ENST00000511691.1). A biomarker may include MACROH2A1 (Ensembl ID ENST00000513268.1). A biomarker may include RACK1 (Ensembl ID ENST00000514455.5). A biomarker may include CHMP7 (Ensembl ID ENST00000517325.5). A biomarker may include HNRNPH1 (Ensembl IDENST00000519033.5). A biomarker may include PDE7A (Ensembl ID ENST00000519626.1). A biomarker may include LYN (Ensembl ID ENST00000520050.1). A biomarker may include CLK4 (Ensembl ID ENST00000520878.5). A biomarker may include HNRNPH1 (Ensembl ID ENST00000521790.5). A biomarker may include GFRA2 (Ensembl ID ENST00000522071.2). A biomarker may include PPP1R2B (Ensembl ID ENST00000522232.3). A biomarker may include NSMAF (Ensembl IDENST00000523106.5). A biomarker may include ARHGAP32 (Ensembl ID ENST00000524655.5). A biomarker may include DGKZ (Ensembl ID ENST00000524869.1). A biomarker may include TTC17 (Ensembl ID ENST00000525135.5). A biomarker may include RPS3 (Ensembl ID ENST00000526608.5). A biomarker may include KNL1 (Ensembl ID ENST00000526913.5). A biomarker may include RAB30 (Ensembl ID ENST00000527633.6). A biomarker may include SBF2 (Ensembl ID ENST00000528478.1). A biomarker may include CYHR1 (Ensembl ID ENST00000528558.1). A biomarker may include PLPP5 (Ensembl ID ENST00000528814.1). A biomarker may include KCNJ5 (Ensembl ID ENST00000529694.6). A biomarker may include HSPA8 (Ensembl ID ENST00000530391.1). A biomarker may include ARHGAP27 (Ensembl ID ENST00000532038.5). A biomarker may include SORL1 (Ensembl ID ENST00000532451.1). A biomarker may include TAF6L (Ensembl ID ENST00000532915.1). A biomarker may include FUZ (Ensembl ID ENST00000533418.5). A biomarker may include NUMA1 (Ensembl ID ENST00000535111.5). A biomarker may include AP000462.1 (Ensembl ID ENST00000540545.1). A biomarker may include KXD1 (Ensembl ID ENST00000540691.5). A biomarker may include ARHGDIA (Ensembl ID ENST00000541078.6). A biomarker may include ZBP1 (Ensembl ID ENST00000541799.1). A biomarker may include PSMD9 (Ensembl ID ENST00000544254.1). A biomarker may include GPR19 (Ensembl ID ENST00000545876.5). A biomarker may include TAF1D (Ensembl ID ENST00000546088.5).A biomarker may include COROIC (Ensembl ID ENST00000547170.1). A biomarker may include RP11- 386G11.10 (Ensembl ID ENST00000547387.1). A biomarker may include ZNF606 (Ensembl ID ENST00000547828.5). A biomarker may include DTX1 (Ensembl ID ENST00000548759.1). A biomarker may include DIP2B (Ensembl ID ENST00000549620.5). A biomarker may include ITGB7 (Ensembl ID ENST00000550743.6). A biomarker may include CFAP54 (Ensembl ID ENST00000554108.6). A biomarker may include FOXN3 (Ensembl ID ENST00000555353.5). A biomarker may include ADAM10 (Ensembl ID ENST00000558733.5). A biomarker may include SMAD3 (Ensembl ID ENST00000560402. 1). A biomarker may include ANKRD9 (Ensembl ID ENST00000560748.5). A biomarker may include RBMX (Ensembl ID ENST00000561733.5). A biomarker may include GALNS (Ensembl ID ENST00000562593.5). A biomarker may include GABARAPL2 (Ensembl ID ENST00000563744. 1). A biomarker may include MAPK8IP3 (Ensembl ID ENST00000566064.1). A biomarker may include MVP (Ensembl ID ENST00000566554.1). A biomarker may include ARID3B (Ensembl ID ENST00000569680. 1). A biomarker may include RP11-314A20.5 (Ensembl ID ENST00000570493.2). A biomarker may include FOXK2 (Ensembl ID ENST00000570585.1). A biomarker may include MRPS21P9 (Ensembl ID ENST00000571150. 1). A biomarker may include CORO7 (Ensembl ID ENST00000572666.5). A biomarker may include SRRM2 (Ensembl ID ENST00000574866.1). A biomarker may include NOLI 1 (Ensembl ID ENST00000577687. 1). A biomarker may include CDK5RAP3 (Ensembl ID ENST00000578663. 1). A biomarker may include NBPF15 (Ensembl ID ENST00000578953. 1). A biomarker may include CBX3P2 (Ensembl ID ENST00000579647. 1). A biomarker may include PITPNC1 (Ensembl ID ENST00000580974.5). A biomarker may include TVP23B (Ensembl ID ENST00000581733.1). A biomarker may include KLF14 (Ensembl ID ENST00000583337.4). A biomarker may include NFE2L1 (Ensembl ID ENST00000585291.5). A biomarker may include GRN (Ensembl ID ENST00000585348. 1). A biomarker may include MLX (Ensembl ID ENST00000585403.5). A biomarker may include OAZ1 (Ensembl ID ENST00000589361.2). A biomarker may include RP11- 686D22.7 (Ensembl ID ENST00000590824. 1). A biomarker may include NOSIP (Ensembl ID ENST00000594932.5). A biomarker may include RPS11 (Ensembl ID ENST00000600027.5). A biomarker may include HCG18 (Ensembl ID ENST00000602290.2). A biomarker may include RP11-214K3.19 (Ensembl ID ENST00000602741.1). A biomarker may include CTD-2589H19.6 (Ensembl ID ENST00000607068.1). A biomarker may include RP5-1098D14.1 (Ensembl ID ENST00000608190.1). A biomarker may include SATB1-AS1 (Ensembl ID ENST00000609327.2). A biomarker may include CTA- 217C2.2 (Ensembl ID ENST00000609432. 1). A biomarker may include NPRL3 (Ensembl ID ENST00000610509.1). A biomarker may include NBPF9 (Ensembl ID ENST00000611593.1). A biomarker may include BIRC3 (Ensembl ID ENST00000615299.4). A biomarker may include MRTFA (Ensembl ID ENST00000618417.1). A biomarker may include TRPM2 (Ensembl ID ENST00000621064.1). A biomarker may include LINC01620 (Ensembl ID ENST00000623295.1). A biomarker may include RP11-44F14.10 (Ensembl ID ENST00000623877. 1). A biomarker may include FAM223A (Ensembl ID ENST00000625031.1). A biomarker may include SPDYE5 (Ensembl ID ENST00000625065.3). A biomarker may include RP11-227G15.12 (Ensembl ID ENST00000625127.3). A biomarker may includeCFLAR-AS1 (Ensembl ID ENST00000625224.2). A biomarker may include XPC-AS1 (Ensembl ID ENST00000626655.2). A biomarker may include EP300 (Ensembl ID ENST00000635083. 1). A biomarker may include MIR34AHG (Ensembl ID ENST00000635405.2). A biomarker may include MIR34AHG (Ensembl ID ENST00000635687. 1). A biomarker may include CLCN2 (Ensembl ID ENST00000637538. 1). A biomarker may include STAT6 (Ensembl ID ENST00000640254.2). A biomarker may include RP11- 582H21.1 (Ensembl ID ENST00000643055.1). A biomarker may include ETAA1 (Ensembl ID ENST00000644028.1). A biomarker may include AUTS2 (Ensembl ID ENST00000644506.1). A biomarker may include CSNK2A1 (Ensembl ID ENST00000645234.1). A biomarker may include LRPAP1 (Ensembl ID ENST00000650182.1). A biomarker may include NFKB1 (Ensembl ID ENST00000652569.1). A biomarker may include FOXCUT (Ensembl ID ENST00000652712.1). A biomarker may include FLJ40194 (Ensembl ID ENST00000659323. 1). A biomarker may include RP4-644F6.4 (Ensembl ID ENST00000662176.1). A biomarker may include RP11-702N8.3 (Ensembl ID ENST00000663064.1). A biomarker may include LINC00910 (Ensembl ID ENST00000663186. 1). A biomarker may include LINC01727 (Ensembl ID ENST00000663538.1). A biomarker may include AC009404.2 (Ensembl ID ENST00000664381.1). A biomarker may include EGLN1 (Ensembl ID ENST00000667629.1). A biomarker may include ZFAS1 (Ensembl ID ENST00000667889.1). A biomarker may include SNHG14 (Ensembl ID ENST00000667925. 1). A biomarker may include SNHG14 (Ensembl ID ENST00000669064.1). A biomarker may include JAK1 (Ensembl ID ENST00000672179.1). A biomarker may include DCTN2 (Ensembl ID ENST00000676646. 1). A biomarker may include SPECC1 (Ensembl ID ENST00000676737.1). A biomarker may include DNMT1 (Ensembl ID ENST00000676868.1). A biomarker may include NDUFAF1 (Ensembl ID ENST00000677477. 1). A biomarker may include HNRNPA1 (Ensembl ID ENST00000678077.1). A biomarker may include ADAR (Ensembl ID ENST00000680472. 1). A biomarker may include SCARB 1 (Ensembl ID ENST00000680926. 1). A biomarker may include OAS3 (Ensembl ID ENST00000681085.1). A biomarker may include AHI1 (Ensembl ID ENST00000681477.1). A biomarker may include P4HB (Ensembl ID ENST00000681693.1). A biomarker may include HEXA (Ensembl ID ENST00000683228. 1). A biomarker may include PTPN23 (Ensembl ID ENST00000683708. 1). Any of these biomarkers may be useful alone or in combination to assess a cancer (for example, to determine a likelihood of having a cancer or not, such as lung cancer).
[0189] In some embodiments, an intron region from RNA sequencing data may be used to assess a cancer, such as lung cancer. In some embodiments, a long non-coding RNA (IncRNA) transcript sequence data may be used to assess a cancer, such as lung cancer. In some embodiments, IncRNAs can originate from various genomic locations, including intergenic regions and RNA transcripts derived solely from the intron regions within protein-coding genes. In some embodiments, an intronic RNA is a type of IncRNA that is specifically located within an intron of a protein-coding gene. Any of the following RNAs may be useful as such (listed with gene name followed by Ensembl transcript ID, strand, and location in the chromosome): GCNT7 (Ensembl ID ENST00000243913.8-chr20_56504295_56513795), C20orfl94 (Ensembl ID ENST00000252032.10-chr20_3318401_3322216), TRIM21 (Ensembl ID ENST00000254436.8- chrl 1_4386991_4388299), DTX1 (Ensembl ID ENST00000257600.3+chrl2_l 13078106_l 13093161),LOXL4 (Ensembl ID ENST00000260702.4-chrl0_98261128_98262034), GEMIN5 (Ensembl ID ENST00000285873.8-chr5_154931578_154932098), NRTN (Ensembl ID ENST00000303212.3+chrl9_5824335_5827748), SH3RF3 (Ensembl ID ENST00000309415.8+chr2_109419643_109432500), TOPAZ1 (Ensembl ID ENST00000309765.4+chr3_44270811_44281967), PTPN13 (Ensembl IDENST00000316707. 10+chr4_86745129_86758259), WDR33 (Ensembl ID ENST00000322313.9- chr2_127708893_127709489), TNFAIP8L1 (Ensembl ID ENST00000327473.9+chrl9_4639630_4651866), RP1L1 (Ensembl ID ENST00000329335.3-chr8_10623221_10711956), UBAP2L (Ensembl ID ENST00000343815.10+chrl_154220245_154225083), GBP4 (Ensembl ID ENST00000355754.7- chrl_89187103_89188581), CD2AP (Ensembl ID ENST00000359314.5+chr6_47554767_47574063), IARS2 (Ensembl ID ENST00000366922.3+chrl_220110938_220114313), CR2 (Ensembl ID ENST00000367059.3+chrl_207475217_207477884), ZC3H11B (Ensembl ID ENST00000367211.6- chrl_219612268_219612624), NEURL1 (Ensembl ID ENST00000369780.8+chrl0_103585226_103589513), ILRUN (Ensembl ID ENST00000374026.7- chr6_34606905_34654624), TRIM63 (Ensembl ID ENST00000374272.4-chrl_26051884_26053892), PADI4 (Ensembl ID ENST00000375448.4+chrl_17342403_17346027), GPC6 (Ensembl ID ENST00000377047.9+chrl3_93227617_93545262), CNGA4 (Ensembl ID ENST00000379936.3+chrl 1_6241781_6243948), POM121B (Ensembl ID ENST00000380760.4+chr7_73300639_73301066), ITPR2 (Ensembl ID ENST00000381340.8- chrl2_26658131_26659112), CATSPERD (Ensembl ID ENST00000381624.4+chrl9_5744511_5745912), PLXNA1 (Ensembl ID ENST00000393409.3+chr3_127016685_127016943), FABP6 (Ensembl ID ENST00000393980.8+chr5_160187455_160199048), SIAH3 (Ensembl ID ENST00000400405.4- chrl3_45784058_45851494), ADD2 (Ensembl ID ENST00000403045.6-chr2_70648781_70663243), AP001610.5 (Ensembl ID ENST00000411427.3-chr21_41442747_41445489), CATIP-AS2 (Ensembl ID ENST00000411433.1-chr2_218351227_218357913), ZNF578 (Ensembl ID ENST00000421239.7+chrl9_52456959_52491323), PSPC1-AS2 (Ensembl ID ENST00000423023.2+chrl3_19674733_19675617), RP4-614C15.2 (Ensembl ID ENST00000424662.1- chr20_59352322_59357308), ABCB9 (Ensembl ID ENST00000426173.6-chrl2_122948709_122949787), RP11-331F9.4 (Ensembl ID ENST00000428948.1+chr9_35646434_35646689), POM121L2 (Ensembl ID ENST00000429945.1-chr6_27286149_27311096), PYHIN5P (Ensembl ID ENST00000431627.1+chrl_158880897_158883353), GPATCH3 (Ensembl ID ENST00000445019.5- chrl_26890852_26892669), LINC01428 (Ensembl ID ENST00000449581.1-chr20_7188568_7241828), RP11-141M1.3 (Ensembl ID ENST00000454681.2-chrl3_33383680_33439690), RP11-65N13.8 (Ensembl ID ENST00000468244.2+chr9_125254672_125256963), SMARCD3 (Ensembl ID ENST00000469154.5- chr7_151243702_151275113), MAST2 (Ensembl ID ENST00000470809.1+chrl_45787133_45824432), PHF20L1 (Ensembl ID ENST00000486199.5+chr8_132814890_132817338), CNBP (Ensembl ID ENST00000500450.6-chr3_129171509_129171654), ENOPH1 (Ensembl ID ENST00000505846.5+chr4_82448022_82454721), LRP2BP (Ensembl ID ENST00000505916.6-chr4_185378208_185394778), FAM169A (Ensembl ID ENST00000510496.5-chr5_74805285_74834425), ADGRA3 (Ensembl ID ENST00000511051.5-chr4_22346824_22389087), LINC02149 (Ensembl ID ENST00000511443.1-chr5_15192450_15243648), RP11-227H4.8 (Ensembl ID ENST00000512905.6- chr3_98521443_98562323), DGLUCY (Ensembl ID ENST00000518649.5+chrl4_91167379_91181185), ATP6V1B2 (Ensembl ID ENST00000521442.1+chr8_20225407_20225688), RAB38 (Ensembl ID ENST00000526372.1-chrl l_88149956_88175237), SERPINH1 (Ensembl ID ENST00000526638.1+chrl l_75572400_75572550), RP11-867G23.13 (Ensembl ID ENST00000534065.1+chrl 1_66312993_66318833), AP000688.14 (Ensembl ID ENST00000535199.5- chr21_36098260_36126486), NPTX1 (Ensembl ID ENST00000535681.1-chrl7_80468918_80470526), CD163 (Ensembl ID ENST00000542280.5-chrl2_7479997_7482965), NLRC5 (Ensembl ID ENST00000543030.5+chrl6_57055520_57059466), LINC02306 (Ensembl ID ENST00000546412.2- chrl4_25664976_25835849), CNPY2 (Ensembl ID ENST00000551286. l-chrl2_56314967_56315820), CALCOCO1 (Ensembl ID ENST00000551900.5-chrl2_53725259_53727403), YY1 (Ensembl ID ENST00000554804.1+chrl4_100274759_100277417), AE000662.92 (Ensembl ID ENST00000557595.1+chrl4_22556843_22556969), GABPB1 (Ensembl ID ENST00000559100.1- chrl5_50301369_50303965), RP11-96D1.7 (Ensembl ID ENST00000563175.1- chrl6_68259822_68260309), SPDYE6 (Ensembl ID ENST00000563237.3-chr7_102354922_102355863), ITFG1 (Ensembl ID ENST00000563730.1-chrl6_47459176_47463951), STARD13 (Ensembl ID ENST00000567873.1-chrl3_33167623_33350289), MT1E (Ensembl ID ENST00000568293.1+chrl6_56625880_56626879), RP11-719K4.6 (Ensembl ID ENST00000568759.1- chrl6_14735084_14740891), CORO7 (Ensembl ID ENST00000576437.5-chrl6_4362132_4395288), SPON1 (Ensembl ID ENST00000576479.4+chrl l_14254730_14255646), RNF167 (Ensembl ID ENST00000576965.1+chrl7_4942942_4944714), ICAM2 (Ensembl ID ENST00000578379.5- chrl7_64006672_64020522), IMPACT (Ensembl ID ENST00000581278. l+chrl8_24450815_24453149), LRRC37A3 (Ensembl ID ENST00000583510.1-chrl7_64855890_64859441), MYH3 (Ensembl ID ENST00000583535.6-chrl7_10632017_10632475), SLC47A1 (Ensembl ID ENST00000584348.5+chrl7_19495482_19542392), RP11-701H16.4 (Ensembl ID ENST00000588397.1- chrl8_13645198_13645530), ZNF582 (Ensembl ID ENST00000589143.5-chrl9_56376122_56390000), CIRBP (Ensembl ID ENST00000592234.5+chrl9_1271037_1271550), CD70 (Ensembl ID ENST00000597430.2-chrl9_6592576_6603940), ZNF418 (Ensembl ID ENST00000598213.5- chrl9_57922630_57935079), ZNF611 (Ensembl ID ENST00000602162.5-chrl9_52706865_52735000), SIGLEC18P (Ensembl ID ENST00000602271.1+chrl9_51120384_51120509), ENKD1 (Ensembl ID ENST00000602644.5-chrl6_67663557_67663936), RP11-40F8.2 (Ensembl ID ENST00000608254.1- chrX_23064122_23270050), RP11-424E7.6 (Ensembl ID ENST00000614278.1- chr9_138179514_138179564), ROCK2 (Ensembl ID ENST00000616279.4-chr2_l 1217174 1218957), FCGBP (Ensembl ID ENST00000616721.6-chrl9_39924733_39927058), FCGBP (Ensembl ID ENST00000616721.6-chrl9_39928307_39934563), AC006133.7 (Ensembl ID ENST00000617106.1- chrl9_39726221_39726702), HNRNPUL1 (Ensembl IDENST00000617305.4+chrl9_41264694_41305675), RP4-681N20.5 (Ensembl ID ENST00000619343.1 - chr20_3921870_3923322), MICAL1 (Ensembl ID ENST00000630715.2-chr6_109454240_109465663), MIR34AHG (Ensembl ID ENST00000635687.1-chrl_9151836_9182006), DDX24 (Ensembl ID ENST00000640432.1-chrl4_94051070_94051339), SPTA1 (Ensembl ID ENST00000643759.2- chrl_158652654_158653273), RP11-19D22.1 (Ensembl ID ENST00000648060.1- chrl3_7672951 l_76835980), RP11-429022.1 (Ensembl ID ENST00000652524.1+chr4_83076353_83082729), RP11-1074012.2 (Ensembl ID ENST00000652631.1- chrl4_40286951_40347879), RP11-546M8.1 (Ensembl ID ENST00000652635.1+chr2_239539446_239544463), RP11-428J1.6 (Ensembl ID ENST00000655309.1+chr6_5072735_5079555), LINC01053 (Ensembl ID ENST00000656176.1+chrl3_25169032_25186523), LINC01013 (Ensembl ID ENST00000660939.1+chr6_132147845_132160766), RP1-135L22.2 (Ensembl ID ENST00000663029.1- chr6_21369811_21382738), LINC01220 (Ensembl ID ENST00000665235.1+chrl4_75296565_75297208), RP11-761N21.1 (Ensembl ID ENST00000665970.1+chr3_40719916_40773867), AVEN (Ensembl ID ENST00000675287.1-chrl5_34066812_34070587), PABPC1 (Ensembl ID ENST00000677787.1- chr8_100712458_100713086), GGH (Ensembl ID ENST00000677919.1-chr8_63017631_63024079), FLJ43315 (Ensembl ID ENST00000680973.1+chr9_64031415_64054475), MIR34AHG-204 (ENST00000635687.1-chrl_9151836_9182006), LINC01053-202 (ENST00000656176. l+chrl3_25169032_25186523), MMP9-201 (ENST00000372330.3+chr20_46014475_46016249), ENST00000635641.1-chr6_26196247_26196719, BANK1-206 (ENST00000504592.5+chr4_101721434_101829807), PADI4-201 (ENST00000375448.4+chrl_17342122_17342298), VSIG10-202 (ENST00000359236.10- chrl2_l 18073993_l 18079345), PLVAP-202 (ENST00000595816.1-chrl9_17351643_17365972), PTPRK- AS1-201 (ENST00000417390.1+chr6_128067597_128083670), FGF9-201 (ENST00000382353.6+chrl3_21672190_21681041), EBF1-205 (ENST00000518836.5- chr5_159099464_159099582), ELOVL4-201 (ENST00000369816.5-chr6_79926382_79947179), TSPAN 13-201 (ENST00000262067.5+chr7_l 6754031 16776210), ENST00000642991.1- chrl3_99409060_99417056, ENST00000667372.1+chr3_53196917_53197230, LINC02726-201 (ENST00000534068. 1+chrl 1 23731659_23794822), ITGA2B-206 (ENST00000592226.5- chrl7_44384586_44385163), MMP9-201 (ENST00000372330.3+chr20_46013797_46014123), CD79A- 202 (ENST00000444740.2+chrl9_41879176_41879534), F13A1-201 (ENST00000264870.8- chr6_6195886_6197222), MS4A1-204 (ENST00000528313.1+chrl l_60462486_60466997), ENST00000512905.6-chr3_98521443_98562323, ENST00000435287.1+chr6_131951255_132077207, VSIG10-204 (ENST00000538357. l-chrl2_l 18079607_l 18095532), ENST00000588212.1- chrl9_44340540_44397198, SNAP47 (Ensembl ID ENST00000681827.1+chrl_227728285_227728689), ENST00000372330.3+chr20_46014475_46016249, ENST00000635641.1-chr6_26196247_26196719, ENST00000504592.5+chr4_101721434_101829807, ENST00000375448.4+chrl_17342122_17342298,ENST00000359236. 10-chrl2_l 18073993 118079345, ENST00000595816.1-chrl9_17351643_17365972, ENST00000417390.1+chr6_128067597_128083670, ENST00000382353.6+chrl3_21672190_21681041, ENST00000518836.5-chr5_159099464_159099582, ENST00000369816.5-chr6_79926382_79947179, ENST00000262067.5+chr7_16754031_16776210, ENST00000642991.1-chrl3_99409060_99417056, ENST00000667372. 1 +chr3_53196917_53197230, ENST00000534068. 1+chrl 1 23731659_23794822, ENST00000592226.5-chrl7_44384586_44385163, ENST00000372330.3+chr20_46013797_46014123, ENST00000444740.2+chrl9_41879176_41879534, ENST00000264870.8-chr6_6195886_6197222, ENST00000528313. 1+chrl l_60462486_60466997, ENST00000435287. l+chr6_131951255J32077207, ENST00000538357. l-chrl2_118079607_l 18095532, ENST00000588212.1-chrl9_44340540_44397198. A biomarker may include GCNT7 (Ensembl ID ENST00000243913.8-chr20_56504295_56513795). A biomarker may include A biomarker may include C20orfl94 (Ensembl ID ENST00000252032.10- chr20_3318401 3322216). A biomarker may include A biomarker may include TRIM21 (Ensembl ID ENST00000254436.8-chrl 1_4386991_4388299). A biomarker may include A biomarker may include DTX1 (Ensembl ID ENST00000257600.3+chrl2_113078106_113093161). A biomarker may include A biomarker may include LOXL4 (Ensembl ID ENST00000260702.4-chrl0_98261128_98262034). A biomarker may include GEMIN5 (Ensembl ID ENST00000285873.8-chr5_154931578_154932098). A biomarker may include NRTN (Ensembl ID ENST00000303212.3+chrl9_5824335_5827748). A biomarker may include SH3RF3 (Ensembl ID ENST00000309415.8+chr2_109419643_109432500). A biomarker may include TOPAZ1 (Ensembl ID ENST00000309765.4+chr3_44270811_44281967). A biomarker may include PTPN13 (Ensembl ID ENST00000316707. 10+chr4_86745129_86758259). A biomarker may include WDR33 (Ensembl ID ENST00000322313.9-chr2_127708893_127709489). A biomarker may include TNFAIP8L1 (Ensembl ID ENST00000327473.9+chrl9_4639630_4651866). A biomarker may include RP1L1 (Ensembl ID ENST00000329335.3-chr8_10623221_10711956). A biomarker may include UBAP2L (Ensembl ID ENST00000343815.10+chrl_154220245_154225083). A biomarker may include GBP4 (Ensembl ID ENST00000355754.7-chrl_89187103_89188581). A biomarker may include CD2AP (Ensembl ID ENST00000359314.5+chr6_47554767_47574063). A biomarker may include IARS2 (Ensembl ID ENST00000366922.3+chrl_220110938_220114313). A biomarker may include CR2 (Ensembl ID ENST00000367059.3+chrl_207475217_207477884). A biomarker may include ZC3H1 IB (Ensembl ID ENST00000367211.6-chrl_219612268_219612624). A biomarker may include NEURL1 (Ensembl ID ENST00000369780.8+chrl0_103585226_103589513). A biomarker may include ILRUN (Ensembl ID ENST00000374026.7-chr6_34606905_34654624). A biomarker may include TRIM63 (Ensembl ID ENST00000374272.4-chrl_26051884_26053892). A biomarker may include PADI4 (Ensembl ID ENST00000375448.4+chrl_17342403_17346027). A biomarker may include GPC6 (Ensembl ID ENST00000377047.9+chrl3_93227617_93545262). A biomarker may include CNGA4 (Ensembl ID ENST00000379936.3+chrl 1_6241781_6243948). A biomarker may include POM121B (Ensembl ID ENST00000380760.4+chr7_73300639_73301066). A biomarker may include ITPR2 (Ensembl ID ENST00000381340.8-chrl2_26658131_26659112). A biomarker may include CATSPERD (Ensembl ID ENST00000381624.4+chrl9_5744511_5745912). A biomarker may include PLXNA1 (Ensembl IDENST00000393409.3+chr3_127016685_127016943). A biomarker may include FABP6 (Ensembl ID ENST00000393980.8+chr5_160187455_160199048). A biomarker may include SIAH3 (Ensembl ID ENST00000400405.4-chrl3_45784058_45851494). A biomarker may include ADD2 (Ensembl ID ENST00000403045.6-chr2_70648781_70663243). A biomarker may include AP001610.5 (Ensembl ID ENST00000411427.3-chr21_41442747_41445489). A biomarker may include CATIP-AS2 (Ensembl ID ENST00000411433. l-chr2_218351227_218357913). A biomarker may include ZNF578 (Ensembl ID ENST00000421239.7+chrl9_52456959_52491323). A biomarker may include PSPC1-AS2 (Ensembl ID ENST00000423023.2+chrl3_19674733_19675617). A biomarker may include RP4-614C15.2 (Ensembl ID ENST00000424662.1-chr20_59352322_59357308). A biomarker may include ABCB9 (Ensembl ID ENST00000426173.6-chrl2_122948709_122949787). A biomarker may include RP11-331F9.4 (Ensembl ID ENST00000428948.1+chr9_35646434_35646689). A biomarker may include POM121L2 (Ensembl ID ENST00000429945.1-chr6_27286149_27311096). A biomarker may include PYHIN5P (Ensembl ID ENST00000431627.1+chrl_158880897_158883353). A biomarker may include GPATCH3 (Ensembl ID ENST00000445019.5-chrl_26890852_26892669). A biomarker may include LINC01428 (Ensembl ID ENST00000449581. 1 -chr20_7188568_7241828) . A biomarker may include RP 11 - 141 M 1.3 (Ensembl ID ENST00000454681.2-chrl3_33383680_33439690). A biomarker may include RP11-65N13.8 (Ensembl ID ENST00000468244.2+chr9_125254672_125256963). A biomarker may include SMARCD3 (Ensembl ID ENST00000469154.5-chr7_151243702J51275113). A biomarker may include MAST2 (Ensembl ID ENST00000470809.1+chrl_45787133_45824432). A biomarker may include PHF20L1 (Ensembl ID ENST00000486199.5+chr8_132814890_132817338). A biomarker may include CNBP (Ensembl ID ENST00000500450.6-chr3_129171509_129171654). A biomarker may include ENOPH1 (Ensembl ID ENST00000505846.5+chr4_82448022_82454721). A biomarker may include LRP2BP (Ensembl ID ENST00000505916.6-chr4_185378208_185394778). A biomarker may include FAM169A (Ensembl ID ENST00000510496.5-chr5_74805285_74834425). A biomarker may include ADGRA3 (Ensembl ID ENST00000511051.5-chr4_22346824_22389087). A biomarker may include LINC02149 (Ensembl ID ENST00000511443.1-chr5_15192450_15243648). A biomarker may include RP11-227H4.8 (Ensembl ID ENST00000512905.6-chr3_98521443_98562323). A biomarker may include DGLUCY (Ensembl ID ENST00000518649.5+chrl4_91167379 91181185). A biomarker may include ATP6V1B2 (Ensembl ID ENST00000521442. l+chr8_20225407_20225688). A biomarker may include RAB38 (Ensembl ID ENST00000526372.1-chrl l_88149956_88175237). A biomarker may include SERPINH1 (Ensembl ID ENST00000526638.1+chrl l_75572400_75572550). A biomarker may include RP11-867G23.13 (Ensembl ID ENST00000534065.1+chrl l_66312993_66318833). A biomarker may include AP000688.14 (Ensembl ID ENST00000535199.5-chr21_36098260_36126486). A biomarker may include NPTX1 (Ensembl ID ENST00000535681.1-chrl7_80468918_80470526). A biomarker may include CD163 (Ensembl ID ENST00000542280.5-chrl2_7479997_7482965). A biomarker may include NLRC5 (Ensembl ID ENST00000543030.5+chrl6_57055520_57059466). A biomarker may include LINC02306 (Ensembl ID ENST00000546412.2-chrl4_25664976_25835849). A biomarker may include CNPY2 (Ensembl ID ENST00000551286. l-chr!2_56314967_56315820). A biomarker may include CALCOCO1 (Ensembl IDENST00000551900.5-chrl2_53725259_53727403). A biomarker may include YY1 (Ensembl ID ENST00000554804.1+chrl4_100274759_100277417). A biomarker may include AE000662.92 (Ensembl ID ENST00000557595.1+chrl4_22556843_22556969). A biomarker may include GABPB1 (Ensembl ID ENST00000559100.1-chrl5_50301369_50303965). A biomarker may include RP11-96D1.7 (Ensembl ID ENST00000563175. l-chrl6_68259822_68260309). A biomarker may include SPDYE6 (Ensembl ID ENST00000563237.3-chr7_102354922_102355863). A biomarker may include ITFG1 (Ensembl ID ENST00000563730.1-chrl6_47459176_47463951). A biomarker may include STARD13 (Ensembl ID ENST00000567873.1-chrl3_33167623_33350289). A biomarker may include MT1E (Ensembl ID ENST00000568293.1+chrl6_56625880_56626879). A biomarker may include RP11-719K4.6 (Ensembl ID ENST00000568759.1-chrl6_14735084_14740891). A biomarker may include CORO7 (Ensembl ID ENST00000576437.5-chrl6_4362132_4395288). A biomarker may include SPON1 (Ensembl ID ENST00000576479.4+chrl l_14254730_14255646). A biomarker may include RNF167 (Ensembl ID ENST00000576965.1+chrl7_4942942_4944714). A biomarker may include ICAM2 (Ensembl ID ENST00000578379.5-chrl7_64006672_64020522). A biomarker may include IMPACT (Ensembl ID ENST00000581278.1+chrl8_24450815_24453149). A biomarker may include LRRC37A3 (Ensembl ID ENST00000583510. l-chrl7_64855890_64859441). A biomarker may include MYH3 (Ensembl ID ENST00000583535.6-chrl7_10632017_10632475). A biomarker may include SLC47A1 (Ensembl ID ENST00000584348.5+chrl7_19495482_19542392). A biomarker may include RP11-701H16.4 (Ensembl ID ENST00000588397.1-chrl8_13645198_13645530). A biomarker may include ZNF582 (Ensembl ID ENST00000589143.5-chrl9_56376122_56390000). A biomarker may include CIRBP (Ensembl ID ENST00000592234.5+chrl9_1271037_1271550). A biomarker may include CD70 (Ensembl ID ENST00000597430.2-chrl9_6592576_6603940). A biomarker may include ZNF418 (Ensembl ID ENST00000598213.5-chrl9_57922630_57935079). A biomarker may include ZNF611 (Ensembl ID ENST00000602162.5-chrl9_52706865_52735000). A biomarker may include SIGLEC18P (Ensembl ID ENST00000602271. l+chrl9_51120384_51120509). A biomarker may include ENKD1 (Ensembl ID ENST00000602644.5-chrl6_67663557_67663936). A biomarker may include RP11-40F8.2 (Ensembl ID ENST00000608254.1-chrX_23064122_23270050). A biomarker may include RP11-424E7.6 (Ensembl ID ENST00000614278.1-chr9_138179514 138179564). A biomarker may include ROCK2 (Ensembl ID ENST00000616279.4-chr2_l 1217174 11218957). A biomarker may include FCGBP (Ensembl ID ENST00000616721.6-chrl9_39924733_39927058). A biomarker may include FCGBP (Ensembl ID ENST00000616721.6-chrl9_39928307_39934563). A biomarker may include AC006133.7 (Ensembl ID ENST00000617106.1-chrl9_39726221_39726702). A biomarker may include HNRNPUL1 (Ensembl ID ENST00000617305.4+chrl9_41264694_41305675). A biomarker may include RP4-681N20.5 (Ensembl ID ENST00000619343. l-chr20_3921870_3923322). A biomarker may include MICAL1 (Ensembl ID ENST00000630715.2-chr6_109454240_109465663). A biomarker may include MIR34AHG (Ensembl ID ENST00000635687.1-chrl_9151836_9182006). A biomarker may include DDX24 (Ensembl ID ENST00000640432.1-chrl4_94051070_94051339). A biomarker may include SPTA1 (Ensembl ID ENST00000643759.2-chrl_158652654_158653273). A biomarker may include RP11-19D22.1 (Ensembl IDENST00000648060.1-chrl3_76729511_76835980). A biomarker may include RP11-429022.1 (Ensembl ID ENST00000652524.1+chr4_83076353_83082729). A biomarker may include RP11-1074012.2 (Ensembl ID ENST00000652631.1-chrl4_40286951_40347879). A biomarker may include RP11-546M8.1 (Ensembl ID ENST00000652635.1+chr2_239539446_239544463). A biomarker may include RP11-428J1.6 (Ensembl ID ENST00000655309. l+chr6_5072735_5079555). A biomarker may include LINC01053 (Ensembl ID ENST00000656176.1+chrl3_25169032_25186523). A biomarker may include LINC01013 (Ensembl ID ENST00000660939.1+chr6_132147845_132160766). A biomarker may include RP1-135L22.2 (Ensembl ID ENST00000663029.1-chr6_21369811_21382738). A biomarker may include LINC01220 (Ensembl ID ENST00000665235.1+chrl4_75296565_75297208). A biomarker may include RP11-761N21.1 (Ensembl ID ENST00000665970.1+chr3_40719916_40773867). A biomarker may include AVEN (Ensembl ID ENST00000675287.1-chrl5_34066812_34070587). A biomarker may include PABPC1 (Ensembl ID ENST00000677787.1-chr8_100712458_100713086). A biomarker may include GGH (Ensembl ID ENST00000677919.1-chr8_63017631_63024079). A biomarker may include FLJ43315 (Ensembl ID ENST00000680973.1+chr9_64031415_64054475). A biomarker may include MIR34AHG-204 (ENST00000635687.1-chrl_9151836_9182006). A biomarker may include LINC01053-202 (ENST00000656176.1+chrl3_25169032_25186523). A biomarker may include MMP9-201 (ENST00000372330.3+chr20_46014475_46016249). A biomarker may include ENST00000635641.1- chr6_26196247_26196719. A biomarker may include BANK 1-206 (ENST00000504592.5+chr4_101721434_101829807). A biomarker may include PADI4-201 (ENST00000375448.4+chrl_17342122_17342298). A biomarker may include VSIG10-202 (ENST00000359236.10-chrl2_118073993_118079345). A biomarker may include PLVAP-202 (ENST00000595816. l-chrl9_17351643_17365972). A biomarker may include PTPRK-AS1-201 (ENST00000417390.1+chr6_128067597_128083670). A biomarker may include FGF9-201 (ENST00000382353.6+chrl3_21672190_21681041). A biomarker may include EBF1-205 (ENST00000518836.5-chr5_159099464_159099582). A biomarker may include ELOVL4-201 (ENST00000369816.5-chr6_79926382_79947179). A biomarker may include TSPAN13-201 (ENST00000262067.5+chr7_16754031_16776210). A biomarker may include ENST00000642991.1- chrl3_99409060_99417056. A biomarker may include ENST00000667372.1+chr3_53196917_53197230. A biomarker may include LINC02726-201 (ENST00000534068.1+chrl l_23731659_23794822). A biomarker may include ITGA2B-206 (ENST00000592226.5-chrl7_44384586_44385163). A biomarker may include MMP9-201 (ENST00000372330.3+chr20_46013797_46014123). A biomarker may include CD79A-202 (ENST00000444740.2+chrl9_41879176_41879534). A biomarker may include F13A1-201 (ENST00000264870.8-chr6_6195886_6197222). A biomarker may include MS4A1-204 (ENST00000528313. 1+chr 11-60462486-60466997). A biomarker may include ENST00000512905.6- chr3_98521443_98562323. A biomarker may include ENST00000435287.1+chr6_131951255_132077207. A biomarker may include VSIG10-204 (ENST00000538357.1-chrl2_l 18079607_l 18095532). A biomarker may include ENST00000588212. l-chrl9_44340540_44397198. A biomarker may include or SNAP47 (Ensembl ID ENST00000681827.1+chrl_227728285_227728689). Any of the forementioned nucleic acidbiomarkers (or combination of said biomarkers) may be useful for identifying a presence, absence, or likelihood of a cancer described herein. Any of these biomarkers may be useful alone or in combination to assess lung cancer (for example, non-small cell lung cancer).
[0190] Transcriptomic data may be generated by any of a variety of methods. Generating transcriptomic data may include using a detection reagent that binds to an RNA and yields a detectable signal. After use of a detection reagent that binds to an RNA and yields a detectable signal, a readout may be obtained that is indicative of the presence, absence, or amount of the RNA. Generating transcriptomic data may include concentrating, filtering, or centrifuging a sample.
[0191] Transcriptomic data may include RNA sequence data. Some examples of methods for generating RNA sequence data include use of sequencing, microarray analysis, hybridization, polymerase chain reaction (PCR), or electrophoresis, or a combination thereof. A microarray may be used for generating transcriptomic data. PCR may be used for generating transcriptomic data. PCR may include quantitative PCR (qPCR). Such methods may include use of a detectable probe (e.g., a fluorescent probe) that intercalates with double-stranded nucleotides, or that binds to a target nucleotide sequence. PCR may include reverse transcriptase quantitative PCR (RT-qPCR). Generating transcriptomic data may involve use of a PCR panel.
[0192] RNA sequence data may be generated by sequencing a subject’s RNA or by converting the subject’s RNA into DNA (e.g., complementary DNA (cDNA)) first and sequencing the DNA. Sequencing may include massive parallel sequencing. Examples of massive parallel sequencing techniques include pyrosequencing, sequencing by reversible terminator chemistry, sequencing-by-ligation mediated by ligase enzymes, or phospholinked fluorescent nucleotides or real-time sequencing. Generating transcriptomic data may include preparing a sample or template for sequencing. A reverse transcriptase may be used to convert RNA into cDNA. Some template preparation methods include use of amplified templates originating from single RNA or cDNA molecules, or single RNA or cDNA molecule templates. Examples of amplification methods include emulsion PCR, rolling circle, or solid-phase amplification.
[0193] In addition to any of the above methods, generating transcriptomic data may include contacting a sample with particles such that the particles adsorb biomolecules comprising RNA. The adsorbed RNA may be part of a biomolecule corona. The adsorbed RNA may be measured or identified in generating the transcriptomic data.
[0194] In some embodiments, the transcriptomic data can include spatial transcriptomics. For example, biopsy or tissue sample can be permeabilized and contacted with probes to determine in situ transcriptome, which includes transcriptomic data and its positional context of in a cell of the tissue. In some embodiments, the spatial transcriptomics can be obtained by microdissection (e.g., laser capture microdissection, RNA sequencing of individual cryosections, TIVA, tomo-seq, LCM-seq, Geo-seq, NICHE-seq, or ProximID); fluorescent in situ hybridization (e.g., smFISH, RNAscope, seqFISH, MERFISH, smHCR, osmFISH, seqFISH+, or DNA microscopy); in situ sequencing (e.g., ISS using padlock probes, FISSEQ, Barista-seq, or STARmap); in situ capture (e.g., GeoMx, Slide-seq, APEX-seq, HDST, or 10X Visium); in silico construction (e.g., Reconstruction using ISH or DistMap).
[0195] As shown in FIG. 1, the workflow starts at obtaining a sample from a subject. In this depicted embodiment, the sample obtained from the subject is a blood sample. In some embodiments, the sample is a biofluid sample or a tissue sample of a subject. Non-limiting examples of a biofluid sample include blood, serum, plasma, urine, tears, semen, milk, vaginal fluid, mucus, saliva, sweat, or cell homogenate. The biofluid samples in this workflow were full blood samples collected in PAXgene RNA tubes (although use of other biofluid sample types such as plasma or serum are also envisioned). PAXgene RNA tubes include an RNA stabilization reagent. The workflow continues to the extraction of RNA from the sample obtained from the subject. The extracted RNA undergoes one or more quality control steps. In some embodiments, the quality control of the RNA may be based on concentration, purity, absorbance, absorbance ratios, RNA integrity number (RIN), or any combination thereof. RNA that makes it past the initial quality control step of the workflow may be prepped for analysis. In this depicted embodiment, the RNA undergoes library preparation for downstream sequencing analysis. Library preparation of RNA may be done using a variety of methods, kits, or techniques common in the art. The library prepared RNA undergoes another quality control step. In some embodiments, the quality control of the RNA may be based on concentration, purity, absorbance, absorbance ratios, RNA integrity number (RIN), length, or any combination thereof. The library prepared RNA that passes quality control is sequenced on a sequencer. Sequencing can be performed with any appropriate sequencing technology, including but not limited to single-molecule real-time (SMRT) sequencing, Polony sequencing, sequencing by ligation, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis. Sequencing methods also include next-generation sequencing, e.g., modem sequencing technologies such as Illumina sequencing (e.g., Solexa), Roche 454 sequencing, Ion Torrent sequencing, PacBio sequencing, and SOLiD sequencing. In some cases, next-generation sequencing involves high- throughput sequencing methods. In some embodiments, RNA can be analyzed by quantitative polymerase chain reaction (qPCR), gel electrophoresis (including for e.g., Northern blot), immunochemistry, in situ hybridization such as fluorescent in situ hybridization (FISH), cytochemistry, or sequencing. In some embodiments, the sequencing technique comprises next generation sequencing. In some embodiments, the methods involve a hybridization assay such as Anorogenic qPCR (e.g., TaqManTM or SYBR green), which involves a nucleic acid amplification reaction with a specific primer pair, and hybridization of the amplified nucleic acid probes comprising a detectable moiety or molecule that is specific to a target nucleic acid sequence. An additional exemplary nucleic acid-based detection assay comprises the use of nucleic acid probes conjugated or otherwise immobilized on a bead, multi-well plate, array, or other substrate, wherein the nucleic acid probes are configured to hybridize with a target nucleic acid sequence. The sequencer may generate one or more FastQ files. In this depicted embodiment, the FastQ files are generated per lane. The FastQ files may be aligned to a reference sequence. In this depicted embodiment, the reference sequence is human. Data and samples associated with the data may undergo additional quality control steps. The features of the RNA data may be generated and may be used in a classifier to determine if a subject associated with a sample has cancer or not.
[0196] FIG. 7 depicts a non-limiting example of the analysis on the FastQ reads performed in FIG. 7. In this depicted embodiment, the sequencing reads of the FastQ fdes are aligned using STAR (Spliced Transcripts Alignment to a Reference). Other alignment tools for RNA data may be used. Non-limiting examples of RNA alignment tools include Bowtie2, segemehl, BWA (Burrows-Wheeler Aligner), BWA- MEM (Maximal Exact Match), GEM (Genome Multitool), BBMap, Manta, Isaac variant caller, and RNAReadCounter. Following the alignment of the RNA data, some of the data undergoes analysis for expression levels. In this depicted embodiment, RS EM (RNA-Seq by Expectation-Maximization) is used for determining expression levels. RSEM may estimate gene and isoform expression level. Other programs for expression analysis may be used such as samtools, salmon, or CuffDiff2. Following analysis of the RNA data with RSEM, the isoform features may be identified from the mature transcripts. Further depicted in FIG. 7 is the analysis of intron annotations with feature counts and feature quantification that can be used for developing intron features from the RNA data. In this depicted embodiment, there are two kinds of intron features: immature transcripts and intron retention.
[0197] FIG. 8 depicts a non-limiting example of the use of the intron features for determining between samples associated with subjects with and without cancer. In this depicted embodiment, intron features that have no variance across samples are filtered out. The features are further reduced by removing the features that are absent in 50% of the sample. The feature expression may also be normalized as depicted in FIG. 8. The normalization may be performed by DeSeq2, edgeR, Limma + Voom, EBSeq, NOISeq, Wilcoxon ranksum test, or dearseq. Finally, the features are run through the classifier. In this depicted embodiment, the classifier is XGBoost. Other classifiers may be used such as elastic net, support vector machines, sparse neural networks, or random forests.
[0198] The RNA data for use in identifying subjects with or without cancer may comprise intronic data. The intronic data may be from alternative splicing events. The intronic data may comprise immature transcripts that have not yet undergone splicing. In some embodiments, immature transcripts may be referred to a pre-mRNA. Pre-mRNA may not have the structures and / or features of mRNA, such as a 3' poly(A) tail or a 5' GTP cap (as shown in FIG. 9). FIG. 9 depicts a schematic for processing pre-mRNA to mRNA. In this depicted embodiment, all of the introns are removed and only the exons from the pre-mRNA are present in the mRNA. FIG. 10 depicts additional embodiments for the processing of pre-mRNA to mRNA. In this depicted embodiment, the fully spiced isoform comprising only exons, similar to the mRNA in FIG. 9, is labeled “fully spiced isoform” and an isoform that has retained an intron is labeled “intron retaining isoform.”
[0199] The methods and systems disclosed herein for determining if a subject associated with a sample has cancer or not may be determined by combining RNA data with one or more of an omics data type. The omics data may include, but is not limited to, transcriptomics, proteomics, genomics, metabolomics, lipidomics, fragmentomics, methylomics, epigenomics, orthology, or any combination thereof.Genomic Data
[0200] The data such as multi -omics data described herein may include data on genetic material or genomic data. Genomic data may include data about genetic material such as nucleic acids or histones. The nucleic acids may include DNA. Genomic data may include information on the presence, absence, or amount of the genetic material. An amount of genetic material may be indicated as a concentration, absolute number, or may be relative. Aspects described in relation to genomic data may be relevant to nucleic acid or DNA data, or vice versa. Nucleic acid data may include RNA data, or genomic data may include transcriptomic data. In some examples, genomic features are used as input datasets into trained algorithms (e.g., machine learning models or classifiers) to find correlations between sequence composition and subject groups (e.g., patient groups). Examples of such patient groups include presence or absence of diseases or conditions, elevated or non-elevated risk of diseases or conditions, stages of diseases or conditions, subtypes of diseases or conditions, responders to treatment vs. non-responders to treatment, and progressors versus non-progressors. In some examples, feature matrices are generated to compare samples obtained from subjects with known conditions or characteristics. In some embodiments, samples are obtained from healthy subjects, or subjects who do not have any of the known indications and samples from patients known to have cancer. In an aspect, the disclosed systems and methods provide a classifier generated based on feature information derived from genomic data analysis from biological samples. The classifier forms part of a predictive engine for distinguishing groups in a population based on RNA intron sequence features identified in biological samples.
[0201] Genomic data may include DNA sequence data. The sequence data may include gene sequences. For example, the genomic data may include sequence data for up to about 20,000 genes. The genomic data may also include sequence data for non-coding DNA regions. DNA sequence data may include information on the presence, absence, or amount of DNA sequences. The DNA sequence data may include information on the presence or absence of a mutation such as a single nucleotide polymorphism. The DNA sequence data may include DNA measurement of an amount of mutated DNA, for example a measurement of mutated DNA from cancer cells.
[0202] A DNA amount may include a copy number. Likewise, genomic data may include copy numbers of various sequences. Copy number variation may be determined for circulating cell free DNA (cfDNA). Copy number variation may be indicated for a genomic or chromosomal region. For example, cfDNA sequences found within a genomic region may be quantified as part of a copy number variation analysis. Copy number variation may be indicated as a gain or loss relative to a control, standard, or baseline copy number variation measurement.
[0203] Genomic data may include epigenetic data. Examples of epigenetic data include DNA methylation data, DNA hydroxymethylation data, or histone modification data. Epigenetic data may include DNA methylation or hydroxymethylation. DNA methylation or hydroxymethylation may be measured in whole or at regions within the DNA. Methylated DNA may include methylated cytosine (e.g., 5 -methylcytosine). Cytosine is often methylated at CpG sites and may be indicative of gene activation.
[0204] Epigenetic data may include histone modification data. Histone modification data may include the presence, absence, or amount of a histone modification. Examples of histone modifications include serotonylation, methylation, citrullination, acetylation, or phosphorylation. Some specific examples of histone modifications may include lysine methylation, glutamine serotonylation, arginine methylation, arginine citrullination, lysine acetylation, serine phosphorylation, threonine phosphorylation, or tyrosine phosphorylation. Histone modifications may be indicative of gene activation.
[0205] Genomic data may be distinguished by subtype, where each subtype includes a different type of genomic data. For example, DNA sequence data may be included in another subtype, and epigenetic data may be included in one subtype, or different types of epigenetic data may be included in different subtypes.
[0206] Genomic data may be generated by any of a variety of methods. Generating genomic data may include using a detection reagent that binds to a genetic material such as DNA or histones and yields a detectable signal. After use of a detection reagent that binds to genetic material and yields a detectable signal, a readout may be obtained that is indicative of the presence, absence, or amount of the genetic material. Generating genomic data may include concentrating, filtering, or centrifuging a sample.
[0207] Some examples of methods for generating DNA sequence data include use of sequencing, microarray analysis (e.g., a SNP microarray), hybridization, polymerase chain reaction, or electrophoresis, or a combination thereof. DNA sequence data may be generated by sequencing a subject’s DNA.Sequencing may include massive parallel sequencing. Examples of massive parallel sequencing techniques include pyrosequencing, sequencing by reversible terminator chemistry, sequencing-by-ligation mediated by ligase enzymes, or phospholinked fluorescent nucleotides or real-time sequencing. Generating genomic data may include preparing a sample or template for sequencing. Some template preparation methods include use of amplified templates originating from single DNA molecules, or single DNA molecule templates.Examples of amplification methods include emulsion PCR, rolling circle, or solid-phase amplification.
[0208] DNA methylation can be detected by use of mass spectrometry, methylation-specific PCR, bisulfite sequencing, a Hpall tiny fragment enrichment by ligation-mediated PCR assay, a Glal hydrolysis and ligation adapter dependent PCR assay, a chromatin immunoprecipitation (ChIP) assay combined with a DNA microarray (a ChIP -on-chip assay), restriction landmark genomic scanning, methylated DNA immunoprecipitation, pyrosequencing of bisulfite treated DNA, a molecular break light assay for DNA adenine methyltransferase activity, methyl sensitive Southern blotting, methylCpG binding proteins, high resolution melt analysis, a methylation sensitive single nucleotide primer extension assay, another methylation assay, or a combination thereof.
[0209] Histone modifications may be detected by using mass spectrometry or an immunoassay, an enzyme- linked immunosorbent assay, a western blot, a dot blot, or immunostaining, or a combination thereof.
[0210] In some aspects, the multi -omics data described herein may include ffagmentomics data. In some embodiments, the fragmentomics data comprises sequencing and measuring the abundance of cell-free DNA (cfDNA) bound by transcription factors (e.g., DNA that is not bound by histone). Fragment size is deduced by sequencing in single-base resolution. Information from the fragment can include preferred ends of the cfDNA (e.g., jagged ends are single -stranded ends carried by these double-stranded DNA molecules); DNAtopology (e.g., circular or linear forms); nucleosome footprints; or end motifs (e.g., characteristic sequences at the 5’ end of a fragment). The cfDNA fragmentation can comprise size between 10 to 200 base pairs.
[0211] In addition to any of the above methods, generating genomic data may include contacting a sample with particles such that the particles adsorb biomolecules comprising genetic material. The adsorbed genetic material may be part of a biomolecule corona. The adsorbed genetic material may be measured or identified in generating the genomic data.
[0212] Data may include circulating free DNA (cfDNA) methylation, mRNA, miRNA, circulating free miRNA (cf-miRNA), or whole exome sequencing data. Any sample type, isolation method, quality control (QC) aspect, or sequencing depth provided in the figure may be included. Any aspect shown in the figure may be included in, or used to generate, data such as multi-omics data.Lipidomic Data
[0213] The data such as multi-omics data described herein may include lipid data or lipidomic data. Lipidomic data may include information on the presence, absence, or amount of various lipids. For example, lipidomic data may include amounts of lipids. A lipid amount may be indicated as a concentration or quantity of lipids, for example a concentration of a lipid in a biofluid. A lipid amount may be relative to another lipid or to another biomolecule. Lipidomic data may include information on the presence of lipids. Lipidomic data may include information on the absence of lipids. Lipid or lipidomic data may be included in metabolite or metabolomic data. Aspects described in relation to lipidomic data may be relevant to lipid data, or vice versa. In some examples, lipidomic features are used as input datasets into trained algorithms (e.g., machine learning models or classifiers) to find correlations between sequence composition and subject groups (e.g., patient groups). Examples of such patient groups include presence or absence of diseases or conditions, elevated or non-elevated risk of diseases or conditions, stages of diseases or conditions, subtypes of diseases or conditions, responders to treatment vs. non-responders to treatment, and progressors versus non-progressors. In some examples, feature matrices are generated to compare samples obtained from subjects with known conditions or characteristics. In some embodiments, samples are obtained from healthy subjects, or subjects who do not have any of the known indications and samples from patients known to have cancer. In an aspect, the disclosed systems and methods provide a classifier generated based on feature information derived from lipidomic data analysis from biological samples. The classifier forms part of a predictive engine for distinguishing groups in a population based on RNA intron sequence features identified in biological samples.
[0214] Many organisms contain complex arrays of lipids (for example, humans express over 600 types of lipids), whose relative expression can serve as a powerful marker for biological state and health determinations. Lipids are a diverse class of biomolecules which include fatty acids (e.g., long carbohydrates with carboxylate tail groups), di-, tri-, and poly-glycerides, phospholipids, prenols, sterols (e.g., cholesterol), and ladderanes, among many other types. While lipids are primarily found in membranes, free, protein- complexed, and nucleic acid-complexed lipids are typically present in a range of biofluids, and in some cases may be differentially fractionated from membrane bound lipids. For example, lipid-binding proteins(e.g., albumin) may be collected from a sample by immunohistochemical precipitation, and then chemically induced to release bound lipids for subsequent collection and detection.
[0215] Lipids may be an integral component in the development of diseases such as cancer. For example, lipids may be key players in cancer biology, as they may affect or be involved in feeding membrane and cell proliferation, lipotoxicity (where lipid content balance may aid in protection from lipotoxicity), empowering cellular processes, membrane biophysics, oncogenic signaling and metastasis, protection from oxidative stress, signaling in the microenvironment, or immune modulation. Some lipid classes may be relevant to cancers, such as glycerophospholipids in hepatocellular carcinomas, glycerophospholipids and acylcamitines in prostate cancer, choline containing lipids and phospholipids increase during metastasis, or sphingolipid regulation of cancer cell survival and death.
[0216] Lipid data may be generated from a sample after the sample has been treated to isolate or enrich lipids in the sample. Generating lipid data may include concentrating, filtering, or centrifuging a sample. Lipid analysis can comprise lipid fractionation. In many cases, lipids may be readily separated from other biomolecule types for lipid-specific analysis. As many lipids are strongly hydrophobic, organic solvent extractions and gradient chromatography methods can cleanly separate lipids from other biomolecule-types present within a sample. Lipid data may be generated using mass spectrometry. Lipid analysis may then distinguish lipids by class (e.g., distinguish sphingolipids from chlorolipids) or by individual type.
[0217] Lipidomic data may be generated by any of a variety of methods. Generating lipidomic data may include using a detection reagent that binds to a lipid and yields a detectable signal. After use of a detection reagent that binds to a lipid and yields a detectable signal, a readout may be obtained that is indicative of the presence, absence or amount of the lipid. Generating lipidomic data may include concentrating, filtering, or centrifuging a sample.
[0218] Lipidomic data may be generated using mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, a lateral flow assay, an immunoassay, an enzyme-linked immunosorbent assay, a western blot, a dot blot, or immunostaining, or a combination thereof. An example of a method for generating lipidomic data includes using mass spectrometry. Mass spectrometry may include a separation method step such as liquid chromatography (e.g., HPLC). Mass spectrometry may include an ionization method such as electron ionization, atmospheric -pressure chemical ionization, electrospray ionization, or secondary electrospray ionization. Mass spectrometry may include surface-based mass spectrometry or secondary ion mass spectrometry. Another example of a method for generating lipidomic data includes nuclear magnetic resonance (NMR). Other examples of methods for generating lipidomic data include Fourier-transform ion cyclotron resonance, ion-mobility spectrometry, electrochemical detection (e.g., coupled to HPLC), or Raman spectroscopy and radiolabel (e.g., when combined with thin-layer chromatography). Some mass spectrometry methods described for generating lipidomic data may be used for generating proteomic data, or vice versa. Lipidomic data may also be generated using an immunoassay such as an enzyme-linked immunosorbent assay, western blot, dot blot, or immunohistochemistry. Generating lipidomic data may involve use of a lipid panel.
[0219] In addition to any of the above methods, generating lipidomic data may include contacting a sample with particles such that the particles adsorb biomolecules comprising lipids. The adsorbed lipids may be part of a biomolecule corona. The adsorbed lipids may be measured or identified in generating the lipidomic data.
[0220] Generating lipidomic data may include the use of known amounts internal reference lipids. The reference lipids may be labeled. The label may include an isotopic label. Generating lipidomic data may include the use of known amounts of isotopically labeled internal reference lipids. The internal reference lipids may be spiked into a sample. The internal reference lipids may be used to identify mass spectra of individual endogenous lipids. The internal reference lipids may be used as standards for determining amounts of the individual endogenous lipids. Lipidomic measurements may be generated based on amounts of lipids added into a sample of the one or more biofluid samples. Lipidomic measurements may be generated based on amounts of labeled lipids added into a sample of the one or more biofluid samples.
[0221] Lipids may have associations with biology of a disease such as cancer. Lipids may include phospholipids. Examples of phospholipids include phosphatidylethanolamine (PE), phosphatidylcholine (PC), phosphatidylinositol (PI), or phosphatidylglycerol (PG). Some phospholipids are components of cellular membrane and may play roles in cells such as chemical-energy storage, cellular signaling, cell membrane, or cellular interactions within tissue. A lipid may include a ceramide (CER). Ceramides may act as tumor suppressors, and may be a therapeutic pathway to target. For example, the efficacy of some chemotherapeutics and targeted therapies may be dictated by ceramide levels. A lipid may include a diacylglyceride (DAG). A lipid may include a triacylglyceride (TAG). A lipid may include a fatty acid (FA).
[0222] Examples of lipids may include any lipids in Table 8. Lipid data may include a measurement of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 of the lipids, or a range of any of the aforementioned numbers of lipids from these figures. Any number of the aforementioned lipids may be used. Any of the lipids may be used in a classifier. A combination of lipids may be included.
[0223] A lipid measurement may be affected (e.g., decreased) in a sample from a subject having liver cancer relative to a lipid measurement from a control sample, or relative to a baseline measurement. The lipid measurement may include a phospholipid measurement. The lipid measurement may be useful for evaluating liver cancer. The lipid measurement may include a measurement of a lipid or phospholipid, or a combination of lipids or phospholipids. The lipid measurement may be useful for evaluating ovarian cancer. The lipid measurement may include a measurement of a lipid or phospholipid, or a combination of lipids or phospholipids. The combination of lipids or phospholipids may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 of the lipids, or a range of lipids defined by any two of the aforementioned integers. The combination of lipids or phospholipids may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, or at least 45 of the lipids. The combination of lipids or phospholipids may include less than 3, less than 4, less than 5, less than 6, less than 7, less than 8, less than 9, less than 10, less than 15, less than 20, less than 25, less than 30, less than 35, less than 40, less than 45, or less than 50, of the lipids. In someaspects, the combination of lipids does not include any one or more lipids. In some aspects, the combination of lipids does not include any one or more lipids.
[0224] Any of the biomarkers may be useful alone or in combination to assess a lung nodule (for example, to determine a likelihood of the lung nodule being cancerous or not).
[0225] Lipidomic data may include lipid information such as lipid measurements in a biofluid. Lipid measurements were obtained using LC-MS. Some examples of lipid biomarkers that may be useful in the methods disclosed herein, such as evaluating a cancer such as lung cancer, such as NSCLC. Any combination or number of such biomarkers may be included. In some cases, a biomarker is useful when its feature importance score is above 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.10, 0.12, or 0.14. The features may include any of the following: 10-heptadecenoate (17: ln7), myristate (14:0), caproate (6:0), pentadecanoate (15:0), 1 -stearoyl -GPC (18:0), 1-arachidoyl-GPC (20:0), 1 -docosahexaenoyl -GPC (22:6), 1- arachidonoyl-GPE (20:4n6), 1 -stearoyl -2 -arachidonoyl -GPC (18:0 / 20:4), 1 -docosahexaenoyl -GPE (22:6), stearamide (18:0), sphingomyelin (dl8:2 / 14:0, dl8: 1 / 14: 1), behenoyl sphingomyelin (dl8: 1 / 22:0), lignoceroyl sphingomyelin (dl8: 1 / 24:0), 1,2-dilinoleoyl-GPC (18:2 / 18:2), 1 -stearoyl -2 -arachidonoyl -GPE (18:0 / 20:4), 1 -palmitoyl -2 -arachidonoyl -GPE (16:0 / 20:4), 1-linoleoyl -2 -arachidonoyl -GPC (18:2 / 20:4n6), l-palmitoleoyl-2-linolenoyl-GPC (16: 1 / 18:3), 5-dodecenoylcamitine (C12: l), 1-palmityl -2 -oleoyl -GPE (O- 16:0 / 18: 1), myristamide (14:0), myristoleamide (14: 1), margaramide (17:0), heptadecenamide (17: 1), linolenamide (18:3), oleoyl ethanolamide, linoleamide (18:2n6), or N-stearoyl-sphingosine (dl8: 1 / 18:0). A biomarker may include 10-heptadecenoate (17: ln7). A biomarker may include myristate (14:0). A biomarker may include caproate (6:0). A biomarker may include pentadecanoate (15:0). A biomarker may include 1 -stearoyl -GPC (18:0). A biomarker may include 1-arachidoyl-GPC (20:0). A biomarker may include 1 -docosahexaenoyl -GPC (22:6). A biomarker may include 1 -arachidonoyl -GPE (20:4n6). A biomarker may include 1 -stearoyl -2 -arachidonoyl -GPC (18:0 / 20:4). A biomarker may include 1- docosahexaenoyl-GPE (22:6). A biomarker may include stearamide (18:0). A biomarker may include sphingomyelin (d 18:2 / 14:0. dl 8: 1 / 14: 1). A biomarker may include behenoyl sphingomyelin (dl 8: 1 / 22:0). A biomarker may include lignoceroyl sphingomyelin (d 18: 1 / 24:0). A biomarker may include 1 ,2-dilinoleoyl- GPC (18:2 / 18:2). A biomarker may include 1 -stearoyl -2-arachidonoyl -GPE (18:0 / 20:4). A biomarker may include 1 -palmitoyl -2 -arachidonoyl -GPE (16:0 / 20:4). A biomarker may include 1-linoleoyl -2 -arachidonoyl - GPC (18:2 / 20:4n6). A biomarker may include l-palmitoleoyl-2-linolenoyl-GPC (16: 1 / 18:3). A biomarker may include 5-dodecenoylcamitine (C12: 1). A biomarker may include 1-palmityl -2 -oleoyl -GPE (O- 16:0 / 18: 1). A biomarker may include myristamide (14:0). A biomarker may include myristoleamide (14: 1). A biomarker may include margaramide (17:0). A biomarker may include heptadecenamide (17: 1). A biomarker may include linolenamide (18:3). A biomarker may include oleoyl ethanolamide, linoleamide (18:2n6). A biomarker may include N-stearoyl-sphingosine (dl 8: 1 / 18:0). Any of the forementioned nucleic acid biomarkers (or combination of said biomarkers) may be useful for identifying a presence, absence, or likelihood of a cancer described herein. Any of these biomarkers may be useful alone or in combination to assess lung cancer (for example, non-small cell lung cancer).Metabolomic Data
[0226] The data such as multi -omics data described herein may include metabolite data or metabolomic data. Metabolomic data may include information on small-molecule (e.g., less than 1.5 kDa) metabolites (such as metabolic intermediates, hormones or other signaling molecules, or secondary metabolites). Metabolomic data may involve data about metabolites. Metabolites may include are substrates, intermediates or products of metabolism. A metabolite may include a small molecule. A metabolite may be any molecule less than 1.5 kDa in size. Examples of metabolites may include sugars, lipids, amino acids, fatty acids, phenolic compounds, or alkaloids. Metabolomic data may be distinguished by subtype, where each subtype includes a different type of metabolite. Metabolomic data may include some lipid data. Metabolomic data may comprise lipidomic data. Aspects described in relation to metabolomic data may be relevant to metabolite data, or vice versa. In some examples, metabolomic data features are used as input datasets into trained algorithms (e.g., machine learning models or classifiers) to find correlations between sequence composition and subject groups (e.g., patient groups). Examples of such patient groups include presence or absence of diseases or conditions, elevated or non-elevated risk of diseases or conditions, stages of diseases or conditions, subtypes of diseases or conditions, responders to treatment vs. non-responders to treatment, and progressors versus non-progressors. In some examples, feature matrices are generated to compare samples obtained from subjects with known conditions or characteristics. In some embodiments, samples are obtained from healthy subjects, or subjects who do not have any of the known indications and samples from patients known to have cancer. In an aspect, the disclosed systems and methods provide a classifier generated based on feature information derived from metabolomic data analysis from biological samples. The classifier forms part of a predictive engine for distinguishing groups in a population based on RNA intron sequence features identified in biological samples. Metabolomic data may include metabolite measurements. Metabolite measurements may include measurements of lipids such as phospholipids.
[0227] Metabolomic data may include information on the presence, absence, or amount of various metabolites. For example, metabolomic data may include amounts of metabolites. A metabolite amount may be indicated as a concentration or quantity of metabolites, for example a concentration of a metabolite in a biofluid. A metabolite amount may be relative to another metabolite or to another biomolecule.Metabolomic data may include information on the presence of metabolites. Metabolomic data may include information on the absence of metabolites.
[0228] Metabolomic data generally includes data on a number of metabolites. For example, metabolomic data may include information on the presence, absence, or amount of 1000 or more metabolites. In some cases, metabolomic data may include information on the presence, absence, or amount of 5000, 10,000, 20,000, 50,000, 100,000, 500,000, 1 million, 1.5 million, 2 million, or more metabolites, or a range of metabolites defined by any two of the aforementioned numbers of metabolites.
[0229] Metabolomic data may be generated by any of a variety of methods. Generating metabolomic data may include using a detection reagent that binds to a metabolite and yields a detectable signal. After use of a detection reagent that binds to a metabolite and yields a detectable signal, a readout may be obtained that isindicative of the presence, absence, or amount of the metabolite. Generating metabolomic data may include concentrating, filtering, or centrifuging a sample.
[0230] Metabolomic data may be generated using mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, a lateral flow assay, an immunoassay, an enzyme-linked immunosorbent assay, a western blot, a dot blot, or immunostaining, or a combination thereof. An example of a method for generating metabolomic data includes using mass spectrometry. Mass spectrometry may include a separation method step such as liquid chromatography (e.g., HPLC). Mass spectrometry may include an ionization method such as electron ionization, atmospheric -pressure chemical ionization, electrospray ionization, or secondary electrospray ionization. Mass spectrometry may include surface-based mass spectrometry or secondary ion mass spectrometry. Another example of a method for generating metabolomic data includes nuclear magnetic resonance (NMR). Other examples of methods for generating metabolomic data include Fourier-transform ion cyclotron resonance, ion-mobility spectrometry, electrochemical detection (e.g., coupled to HPLC), or Raman spectroscopy and radiolabel (e.g., when combined with thin-layer chromatography). Some mass spectrometry methods described for generating metabolomic data may be used for generating proteomic data, or vice versa. Metabolomic data may also be generated using an immunoassay such as an enzyme- linked immunosorbent assay, western blot, dot blot, or immunohistochemistry. Generating metabolomic data may involve use of a lipid panel.
[0231] In addition to any of the above methods, generating metabolomic data may include contacting a sample with particles such that the particles adsorb biomolecules comprising metabolites. The adsorbed metabolites may be part of a biomolecule corona. The adsorbed metabolites may be measured or identified in generating the metabolomic data.
[0232] Generating metabolomic data may include the use of known amounts internal reference metabolites. The reference metabolites may be labeled. The label may include an isotopic label. Generating metabolomic data may include the use of known amounts of isotopically labeled internal reference metabolites. The internal reference metabolites may be spiked into a sample. The internal reference metabolites may be used to identify mass spectra of individual endogenous metabolites. The internal reference metabolites may be used as standards for determining amounts of the individual endogenous metabolites. Metabolomic measurements may be generated based on amounts of metabolites added into a sample of the one or more biofluid samples. Metabolomic measurements may be generated based on amounts of labeled metabolites added into a sample of the one or more biofluid samples.
[0233] A metabolite to be detected in a method described herein may include 5-Aminoimidazole-4- carboxamide ribonucleotide (AICAR). The metabolite may include a nucleotide such as a monophosphate nucleotide. Some examples of metabolites are shown in Fig. 36. A metabolite to be detected in a method described herein may include cytidine monophosphate (CMP). The metabolite may include AICAR or CMP. Metabolites to be detected may include AICAR and CMP. Any number of the aforementioned metabolites may be used. Any of the metabolites may be used in a method disclosed herein.-n-
[0234] Some aspects include use of a carbohydrate. An example of a carbohydrate that may be used as a biomarker may include CA-19-9. CA-19-9 levels may be indicative of a cancer such as pancreatic cancer. High levels of CA-19-9 may, in some instances, indicate other types of cancer or a noncancerous disorder such as cirrhosis or gallstones. Because of this, CA 19-9 may be more useful when combined with other biomarkers than when used by itself.
[0235] Data generated directed to metabolites may be used as biomarkers. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more of the metabolites may be useful as biomarkers, for example, in nonsmall cell lung cancer. Any of the following metabolites may be useful as such: pyroglutamylvaline, 3- carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), HWESASXX, 1,7-dimethylurate, 4- hydroxyhippurate, 2-hydroxy-3-methylvalerate, androstenediol (3beta,17beta) disulfate (2), androstenediol (3alpha, 17alpha) monosulfate (2), gamma-CEHC, beta-citrylglutamate, glycohyocholate, 2,8-quinolinediol sulfate, ferulic acid 4-sulfate, 3-(3-hydroxyphenyl)propionate sulfate, methyl glucopyranoside (alpha + beta), alpha-CMBHC glucuronide, 4-vinylguaiacol sulfate, 3 -methoxy catechol sulfate (2), N- acetylkynurenine (2), 2-hydroxylaurate, 7-hydroxyindole sulfate, perfluorooctanoate (PFOA), 3- formylindole, glutamine conjugate of C6H10O2 (1), 2-hydroxyfluorene sulfate, glycoursodeoxycholic acid sulfate (2), 2,6-dihydroxybenzoic acid, 4-vinylcatechol sulfate, bilirubin degradation product, C16H18N2O5 (3), citrate, 3-(4-hydroxyphenyl)lactate, X-11299, X-11381, X-11852, X-11880, X-11979, X-12411, X- 13846, X-15503, X-17146, X-17343, X-17654, X-18779, X-21467, X-21803, X-21821, X-21839, X-23429, X-23787, X-24306, X-24328, X-24338, allo-threonine, N-acetylaspartate (NAA), adipate (C6-DC), 4- hydroxy-2 -oxoglutaric acid, ascorbic acid 3 -sulfate, guanidinoacetate, mannose, arabinose, isobutyrylcamitine (C4), 1 -methyl -4-imidazoleacetate, tryptophanetaine, 5 -methylthioribose, gammacarboxyglutamate, N6-methyllysine, imidazole propionate, lanthionine, arginine, aspartate, S-l-pyrroline-5- carboxylate, Nicotinamide, urea, X-12462, X-12689, X-13507, X-23648, X-23662, X-23666, X-24728, X- 24970, X-24971, X-25172, X-25267, X-25948, X-25957, or X-24027. A biomarker may include pyroglutamyl valine. A biomarker may include 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF). A biomarker may include HWESASXX. A biomarker may include 1.7-dimethylurate. A biomarker may include 4-hydroxyhippurate. A biomarker may include 2 -hydroxy-3 -methylvalerate. A biomarker may include androstenediol (3beta,17beta) disulfate (2). A biomarker may include androstenediol(3alpha, 17alpha) monosulfate (2). A biomarker may include gamma-CEHC. A biomarker may include beta- citrylglutamate. A biomarker may include glycohyocholate. A biomarker may include 2,8-quinolinediol sulfate. A biomarker may include ferulic acid 4-sulfate. A biomarker may include 3-(3- hydroxyphenyl)propionate sulfate. A biomarker may include methyl glucopyranoside (alpha + beta). A biomarker may include alpha-CMBHC glucuronide. A biomarker may include 4-vinylguaiacol sulfate. A biomarker may include 3 -methoxy catechol sulfate (2). A biomarker may include N-acetylkynurenine (2). A biomarker may include 2-hydroxylaurate. A biomarker may include 7-hydroxyindole sulfate. A biomarker may include perfluorooctanoate (PFOA). A biomarker may include 3-formylindole. A biomarker may include glutamine conjugate of C6H10O2 (1). A biomarker may include 2-hydroxyfluorene sulfate. A biomarker may include glycoursodeoxycholic acid sulfate (2). A biomarker may include 2,6-dihydroxybenzoic acid. A biomarker may include 4-vinylcatechol sulfate. A biomarker may include bilirubin degradation product. A biomarker may include C16H18N2O5 (3). citrate. A biomarker may include 3 -(4-hydroxyphenyl)lactate . A biomarker may include X- 11299. A biomarker may include X- 11381. A biomarker may include X-l 1852. A biomarker may include X-l 1880. A biomarker may include X-l 1979. A biomarker may include X-l 2411. A biomarker may include X-l 3846. A biomarker may include X-l 5503. A biomarker may include X-17146. A biomarker may include X-17343. A biomarker may include X-17654. A biomarker may include X-18779. A biomarker may include X-21467. A biomarker may include X-21803. A biomarker may include X-21821. A biomarker may include X-21839. A biomarker may include X-23429. A biomarker may include X-23787. A biomarker may include X-24306. A biomarker may include X-24328. A biomarker may include X-24338. A biomarker may include allo-threonine. A biomarker may include N- acetylaspartate (NAA). A biomarker may include adipate (C6-DC). A biomarker may include 4-hydroxy-2- oxoglutaric acid. A biomarker may include ascorbic acid 3-sulfate. A biomarker may include guanidinoacetate. A biomarker may include mannose. A biomarker may include arabinose. A biomarker may include isobutyrylcamitine (C4). A biomarker may include 1-m ethyl -4-imidazoleacetate. A biomarker may include tryptophanetaine. A biomarker may include 5 -methylthioribose. A biomarker may include gamma-carboxyglutamate. A biomarker may include N6-methyllysine. A biomarker may include imidazole propionate. A biomarker may include lanthionine. A biomarker may include arginine. A biomarker may include aspartate. A biomarker may include S-l-pyrroline-5-carboxylate. A biomarker may include Nicotinamide. A biomarker may include urea. A biomarker may include X-l 2462. A biomarker may include X-12689. A biomarker may include X-13507. A biomarker may include X-23648. A biomarker may include X-23662. A biomarker may include X-23666. A biomarker may include X-24728. A biomarker may include X-24970. A biomarker may include X-24971 . A biomarker may include X-25172. A biomarker may include X-25267. A biomarker may include X-25948. A biomarker may include X-25957. A biomarker may include X-24027. Any of the forementioned nucleic acid biomarkers (or combination of said biomarkers) may be useful for identifying a presence, absence, or likelihood of a cancer described herein. Any of these biomarkers may be useful alone or in combination to assess lung cancer (for example, non-small cell lung cancer).Use of Particles
[0236] Samples may be contacted with particles, for example prior to generating data. The data described herein may generated using particles. For example, a method may include contacting a sample with particles such that the particles adsorb biomolecules. The particles may attract different sets of biomolecules than would normally be measured accurately by performing an omics measurement directly on a sample. For example, a dominant biomolecule may make up a large percentage of certain type of biomolecules (e.g., proteins, transcripts, genetic material, or metabolites) in a sample. For example, one protein may make up a large portion of proteins in circulation that is collected by blood sampling. By adhering biomolecules to particles prior to analyzing the biomolecules, a subset of biomolecules may be obtained that does not includethe dominant biomolecule. Removing dominant biomolecules in this way may increase the accuracy of biomolecule measurements and sensitivity of an analysis using those measurements.
[0237] Examples of biomolecules that may be adsorbed to particles include proteins, transcripts, genetic material, or metabolites. The adsorbed biomolecules may make up a biomolecule corona around the particle. The adsorbed biomolecules may be measured or identified in generating data such as omic data (e.g., proteomic data). In some aspects, the proteomic measurements are generated from proteins adsorbed to nanoparticles. The nanoparticles may enrich the proteins, or may enrich other biomolecule types.
[0238] Particles can be made from various materials. Such materials may include metals, magnetic particles, polymers, or lipids. A particle may be made from a combination of materials. A particle may comprise layers of different materials. The different materials may have different properties. A particle may include a core comprising one material, and be coated with another material. The core and the coating may have different properties.
[0239] A particle may include a metal. A particle may be magnetic (e.g., ferromagnetic or ferrimagnetic). A particle comprising iron oxide may be magnetic. A particle may include a superparamagnetic iron oxide nanoparticle (SPION).
[0240] A particle may include a polymer. A particle may be made from a combination of polymers.
[0241] A particle may include a lipid. A particle may be made from a combination of lipids.
[0242] Particles of various sizes may be used. The particles may include nanoparticles. Nanoparticles may be from about 10 nm to about 1000 nm in diameter.
[0243] The particles may include microparticles. A microparticle may be a particle that is from about 1 pm to about 1000 pm in diameter.
[0244] The particles may include physiochemically distinct sets of particles (for example, 2 or more sets of physiochemically particles where 1 set of particles is physiochemically distinct from another set of particles. Examples of physiochemical properties include charge (e.g., positive, negative, or neutral) or hydrophobicity (e.g., hydrophobic or hydrophilic). The particles may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more sets of particles, or a range of sets of particles including any of said numbers of sets of particles.Biomarker Analysis in Biological Samples
[0245] The methods of use thereof disclosed herein can identify a large number of biomarkers in a biological sample (e.g., a biofluid). Non-limiting examples of biological samples that may be analyzed using the methods (e.g. protein corona analysis) described herein include biofluid samples (e.g., cerebral spinal fluid (CSF), synovial fluid (SF), urine, plasma, serum, tears, semen, whole blood, milk, nipple aspirate, ductal lavage, vaginal fluid, nasal fluid, ear fluid, gastric fluid, pancreatic fluid, trabecular fluid, lung lavage, prostatic fluid, sputum, fecal matter, bronchial lavage, fluid from swabbings, bronchial aspirants, sweat or saliva), fluidized solids (e.g., a tissue homogenate), or samples derived from cell culture. For example, a particle disclosed herein can be incubated with any biological sample disclosed herein to form a protein corona comprising at least 100 unique proteins, or from 100 to 1000 unique proteins. Similar numbers ofproteins may be assessed in some cases without the use of particles, or with an assay method described herein. In some embodiments, several different types of particles can be used, separately or in combination, to identify large numbers of proteins in a particular biological sample. In other words, particles can be multiplexed in order to bind and identify large numbers of proteins in a biological sample.
[0246] The methods disclosed herein can be used to identify various biological states in a particular biological sample. For example, a biological state can refer to an elevated or low level of a particular protein or a set of proteins or may be evidenced by a ratio between the abundances of two or more biomolecules. In other examples, a biological state can refer to identification of a disease, such as cancer. The biological state may include a cancerous lung nodule. The biological state may include a non-cancerous lung nodule. One or more particle types can be incubated with a biological sample, such as human plasma, allowing for formation of a protein corona. Said protein corona can then be analyzed in order to identify a pattern of proteins. The analysis may comprise gel electrophoresis, mass spectrometry, chromatography, ELISA, immunohistology, or any combination thereof. Analysis of protein corona (e.g., by mass spectrometry or gel electrophoresis) may be referred to as corona analysis. The pattern of proteins can be compared to the same methods carried out on a control sample. Upon comparison of the patterns of proteins, it may be identified that the first sample comprises an elevated level of markers corresponding to a particular type of lung cancer. The particles and methods of use thereof, can thus be used to diagnose a particular disease state.
[0247] An assay may comprise protein collection of particles, protein digestion, and mass spectrometric analysis (e.g., MS, LC-MS, LC-MS / MS). The digestion may comprise chemical digestion, such as by cyanogen bromide or 2-Nitro-5-thiocyanatobenzoic acid (NTCB). The digestion may comprise enzymatic digestion, such as by trypsin or pepsin. The digestion may comprise enzymatic digestion by a plurality of proteases. The digestion may comprise a protease selected from among the group consisting of trypsin, chymotrypsin, Glu C, Lys C, elastase, subtilisin, proteinase K, thrombin, factor X, Arg C, papaine, Asp N, thermolysine, pepsin, aspartyl protease, cathepsin D, zinc mealloprotease, glycoprotein endopeptidase, proline, aminopeptidase, prenyl protease, caspase, kex2 endoprotease, or any combination thereof. A digestion method may randomly cleave peptides or may cleave peptides at a specific position or set of positions. An assay may utilize a plurality of digestion methods (e.g., two or more proteases). An assay may comprise splitting a sample into multiple portions and subjecting the portions to different digestion methods and separate analyses (e.g., separate mass spectrometric analyses). The digestion may cleave peptides at a specific position (e.g., at methionines) or sequence (e.g., glutamate-histidine-glutamate). The digestion may enable similar proteins to be distinguished. For example, an assay may resolve 8 distinct proteins as a single protein group with a first digestion method, and as 8 separate proteins with distinct signals with a second digestion method. The digestion may generate an average peptide fragment length of 8 to 15 amino acids. The digestion may generate an average peptide fragment length of 12 to 18 amino acids. The digestion may generate an average peptide fragment length of 15 to 25 amino acids. The digestion may generate an average peptide fragment length of 20 to 30 amino acids. The digestion may generate an average peptide fragment length of 30 to 50 amino acids.
[0248] Various methods of the present disclosure enable measurement over a broad concentration range. Biomolecule analysis methods are often limited to narrow concentration ranges. For example, mass spectrometric proteomic analyses are often limited to 3, 4, or 5 orders of magnitude in concentration. Thus, the presence of relatively high concentration biomolecules (e.g., present at mg / ml concentrations) may mask detection of lower concentration biomolecules, and furthermore may limit the accuracy of low concentration biomolecule quantitation. Methods of the present disclosure may enable detection of molecules spanning at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, or at least 12 orders of magnitude in concentration. Thus, a method of the present disclosure may detect and quantitate a relatively high concentration biomolecule and a relatively low concentration biomolecule from a single sample without first depleting biomolecules from the sample. For example, a plasma assay consistent with the present disclosure may simultaneously quantitate albumin (present at around 40 mg / ml) and interleukin 10 (present at around 6 pg / ml) from a single, non-depleted plasma sample, thereby simultaneously detecting two species who concentrations differ by about 10 orders of magnitude.
[0249] The present disclosure provides methods for detecting low abundance peptides in complex biological samples. Many of the diagnostic peptides of the present disclosure are inaccessible through traditional blood analysis methods due to the high concentrations of albumin, immunoglobulins, and other high abundance blood proteins. A diagnostic peptide may be present at 3-, 4-, 5-, 6-, 7-, 8-, 9-, 10-, 11-, 12- or more orders of magnitude lower concentration than the highest abundance proteins in a blood sample, and accordingly will cannot be detected by many traditional proteomic methods. The present disclosure provides methods for enriching low abundance biomolecules (e.g., proteins) from complex biological samples such as plasma, and also for quantifying the enriched biomolecules.
[0250] The methods of the present disclosure enable quantification of disparate biomarkers spanning wide concentration ranges. In some cases, a lung cancer (e.g., NSCLC) is evidenced by the relative concentrations of two or more proteins from a sample from a patient. In some cases, a method of the present disclosure comprises identifying abundance (e.g., concentration) ratios between at least 2 peptides from among the peptides listed in another table or figure provided herein. In some cases, the sample is a blood sample (e.g., plasma).
[0251] Any combination of biomarkers may be used. In some cases, the biomarker is a secreted protein. In some aspects, the biomarker includes a protein involved in a metabolic pathway. In some aspects, the biomarker includes a protein involved in oxidative phosphorylation. In some cases, the biomarker includes a cell-free RNA. In some cases, the biomarker is an RNA encoding a secreted protein. In some aspects, the biomarker includes an mRNA encoding a protein involved in a metabolic pathway. In some aspects, the biomarker includes an mRNA encoding a protein involved in oxidative phosphorylation.
[0252] In some cases, a method comprises assaying a plasma sample to detect a presence, absence, or abundance of one or more peptides or fragments of peptides from among the peptides listed in table or figure provided herein. In some cases, a method comprises assaying a buffy coat sample to detect a presence, absence, or abundance of one or more peptides or fragments of peptides from among the peptides listed in table or figure provided herein. In some cases, a method comprises assaying a granulocyte sample to detect apresence, absence, or abundance of one or more peptides or fragments of peptides from among the peptides listed in table or figure provided herein. In some cases, a method comprises assaying homogenized tissue (e.g. a homogenized lung biopsy tissue sample) to detect a presence, absence, or abundance of one or more peptides or fragments of peptides from among the peptides listed in table or figure provided herein.
[0253] The present methods enable rapid and deep biomolecule profiling from complex biological samples. In many cases, a method detects and identifies hundreds or thousands of distinct biomolecules. Such broad analysis enables deeper profiling of complex samples, and increases the diagnostic utility of individual peptides. A method of the present disclosure may comprise assaying a sample from a subject to detect a presence, absence, or abundance of at least 50 peptides from a biological sample along with one or more additional peptides or fragments of peptides from among the peptides listed in table or figure provided herein. A method of the present disclosure may comprise identifying abundance or signal intensity (e.g., mass spectrometric signal intensity) ratios between at least a subset of the at least 50, at least 100, at least 200, at least 400, at least 600, at least 800, at least 1000, at least 1200, at least 1400, at least 1600, or at least 1800 peptides and one or more additional peptides or fragments of peptides from among the peptides disclosed herein or another table or figure provided herein.
[0254] A method of the present disclosure may comprise monitoring a lung cancer progression over time. A method of the present disclosure may comprise monitoring a lung nodule over time. A method may comprise collecting two samples from a patient at two different points in time and detecting at least two peptides from among the peptides disclosed herein or another table or figure provided herein in each of the samples. A method may comprise collecting two samples from a patient at two different points in time and detecting at least three peptides from among the peptides disclosed herein or another table or figure provided herein in each of the samples. Any combination of biomarkers may be used. The second of the two samples may be collected at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 5 weeks, at least 6 weeks, at least 8 weeks, at least 12 weeks, at least 15 weeks, at least 18 weeks, at least 24 weeks, at least 36 weeks, at least 52 weeks, at least 78 weeks, at least 104 weeks, at least 130 weeks, at least 156 weeks, at least 208 weeks, or at least 260 weeks apart. A sample or both samples may be collected during the course of a cancer treatment, such as chemotherapy, to determine the efficacy of the treatment. A sample may be collected during a cancer remission stage in order to detect the reemergence, dormancy, or progression to complete remission.
[0255] Disclosed herein are methods that include biomarkers. In some cases, any of the biomarkers are useful for identifying a lung nodule as being cancerous or not. The biomarkers may be included in a classifier for distinguishing the lung nodule as being cancerous or not. Any combination of biomarkers may be used.
[0256] In some embodiments, the biomarkers include any protein in Fig. 13. The biomarkers may include all of the proteins in Fig. 13. The biomarkers may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 49 of the proteins in Fig. 13, or a range of proteins defined by any two of the aforementioned integers. The biomarkers may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, or at least 35, of the proteins in Fig.13. In some aspects, the biomarkers include less than 3, less than 4, less than 5, less than 6, less than 7, less than 8, less than 9, less than 10, less than 15, less than 20, less than 25, less than 30, less than 35, less than 40, less than 45, or less than 49, of the proteins in Fig. 13. In some aspects, the biomarkers exclude a protein in Fig. 13. Any combination of biomarkers may be used.
[0257] Although several examples of protein biomarkers have been included, other types of biomolecules may be useful for biomarkers. For example, biomolecules such as genetic material, transcripts, or metabolites may be used as biomarkers in the methods described herein.Further Detection Methods
[0258] The present disclosure provides a variety of methods for detecting biomolecules (e.g. biomarkers such as protein biomarkers) from a biological sample. Some embodiments include obtaining a biomarker measurement. The biomarker measurement may include a protein measurement such as a protein concentration or amount. Some embodiments include measuring a biomarker. The biological sample may be from a subject with a lung nodule. Biomolecular (e.g., proteomic) data of the biological sample can be identified, measured, and quantified using a number of different analytical techniques. For example, proteomic data can be analyzed using SDS-PAGE or any gel-based separation technique. Alternatively, proteomic data can be identified, measured, and quantified using mass spectrometry, high performance liquid chromatography, LC-MS / MS, Edman Degradation, an immunoaffinity technique, binding reagent analysis (e.g., immunostaining or an aptamer binding assay), an enzyme linked immunosorbent assay (ELISA), chromatography, western blot analysis, mass spectrometric analysis, or any combination thereof. The biomolecules may be enriched on a particle or particle panel prior to analysis. A subset of biomolecules from a biological sample may be collected on a particle, optionally eluted into a solution, optionally treated (e.g., digested or chemically reduced), and analyzed. Particle -based biomolecule collection may enrich a biomolecule from a biological sample, thereby enabling rapid detection and quantification of a low abundance biomolecule.
[0259] Various methods of the present disclosure for detecting a biomolecule comprise binding reagent analysis. A biological sample or collection of biomolecules from a biological sample may be contacted with a target-specific binding reagent, such as an antibody, an affibody, an affimer, an alphabody, an avimer, a DARPin, a chimeric antigen receptor, a T-cell receptor, an aptamer, or a fragment thereof. A binding reagent may be detectable. A binding reagent may comprise a barcode sequence that enables detection and quantification of the binding reagent by nucleic acid sequencing analysis. A binding reagent may comprise an optically detectable label or moiety (e.g., a fluorescent protein such as GFP or YFP or a fluorescent dye). Binding reagent analysis may comprise a plurality of binding reagents targeting a plurality of biomolecules and comprising different detectable signals (e.g., nucleic acid barcode sequences or optically detectable moieties), thereby enabling multiplexed detection and quantification of selected biomarkers from the sample. For example, a sample may be contacted with a plurality of antibodies comprising distinct detectable labels and targeting different proteins from among the proteins listed herein, another table or figure, or a classifier feature disclosed herein. In some cases, a binding reagent may contact a biomolecule covalently or non-covalently immobilized to a substrate (e.g., a membrane, a surface, a resin, or a slide). Insome cases, a binding reagent may contact a biomolecule adsorbed to a particle (e.g., disposed in a biomolecule corona of a particle).
[0260] In some aspects, assaying the proteins comprises measuring a readout indicative of the presence, absence or amount of the biomolecules. In some aspects, assaying the proteins comprises performing mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solidphase chromatography, a lateral flow assay, an immunoassay, an enzyme-linked immunosorbent assay, a western blot, a dot blot, or immunostaining, or a combination thereof. In some aspects, assaying the proteins comprises performing mass spectrometry.
[0261] Various methods of the present disclosure for detecting a biomolecule comprise ELISA. A method may comprise sandwich ELISA analysis, in which a biomolecule (e.g., a peptide from among the peptides listed in table or figure, or a classifier feature disclosed herein) is contacted to a first antibody immobilized to a solid phase and a second antibody coupled to a detectable moiety (e.g., an optically detectable dye molecule), wherein the first antibody comprises a first paratope for a first epitope on the biomolecule and the second antibody comprises a second paratope for a second epitope on the biomolecule. An ELISA assay may comprise immobilizing a biomolecule of interest to a substrate (e.g., a glass slide or the bottom of a well of a multiwell plate) and contacting the biomolecule with a first antibody comprising a binding affinity for the biomolecule. The first antibody may be coupled to a detectable moiety or may be contacted to a second antibody that is coupled to a detectable moiety and which binds to the first antibody. ELISA assays can comprise low detection limits (e.g., > 1 pg / ml) for target detection and quantitation and may thus be suitable for analyzing a cancer biomarker disclosed herein.
[0262] A method of the present disclosure may comprise mass spectrometric analysis of a biomolecule such as a protein, a peptide, or a portion thereof. The mass spectrometric analysis can be performed in tandem with a chromatographic separation technique, such as liquid chromatography, such that biomolecules or biomolecule fragments are subjected to mass spectrometric analysis at different points in time. Mass spectrometric analysis may comprise two or more mass analysis steps (e.g., tandem mass spectrometry), such that an ion is fragmented and then subjected to further analysis.
[0263] The methods described herein may include measuring a biomarker (e.g., one or more biomarkers) in a sample from a subject. Measuring a biomarker may include performing an assay method. Measuring a biomarker may include performing mass spectrometry, chromatography, liquid chromatography, high- performance liquid chromatography, solid-phase chromatography, a lateral flow assay, an immunoassay, an enzyme-linked immunosorbent assay, a western blot, a dot blot, or immunostaining, or a combination thereof. Measuring a biomarker may include performing mass spectrometry. Measuring a biomarker may include performing chromatography. Measuring a biomarker may include performing liquid chromatography. Measuring a biomarker may include performing high-performance liquid chromatography. Measuring a biomarker may include performing solid-phase chromatography. Measuring a biomarker may include performing a lateral flow assay. Measuring a biomarker may include performing an immunoassay. Measuring a biomarker may include performing an enzyme-linked immunosorbent assay. Measuring a biomarker may include performing a blot such as a western blot. Measuring a biomarker may includeperforming dot blot. Measuring a biomarker may include performing immunostaining. Measuring a biomarker may include contacting a biological sample with a plurality of physiochemically distinct nanoparticles. Measuring a biomarker may include performing a combination of assay methods. For example, a method described herein may include use of particles followed by an immunoassay such as an ELISA to assess proteins or biomolecules of biomolecule or protein coronas. The methods described herein may include detecting the proteins of the biomolecule coronas by mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, a lateral flow assay, an immunoassay, an enzyme-linked immunosorbent assay, a western blot, a dot blot, or immunostaining, or a combination thereof. The methods described herein may include detecting the proteins of the biomolecule coronas by mass spectrometry.
[0264] Measuring a biomarker may include using a detection reagent that binds to a protein and yields a detectable signal. The methods described herein may include detecting the proteins comprises measuring a readout indicative of the presence, absence or amounts of the proteins. Measuring a biomarker may include measuring a readout indicative of the presence, absence or amounts of the one or more biomarkers.
[0265] A method may include concentrating biomarkers in a sample prior to measuring the biomarkers. Measuring a biomarker may include concentrating a sample. Measuring a biomarker may include filtering a sample. Measuring a biomarker may include centrifuging a sample.
[0266] Measuring a biomarker may include contacting the sample with an assay reagent. The assay reagent may include a particle. The assay reagent may include an antibody. The assay reagent may include a biomolecule binding molecule.
[0267] The biological sample may contain one or more analytes capable of being assayed, such as cell-free ribonucleic acid (cfRNA) molecules suitable for assaying to generate transcriptomic data, cell-free deoxyribonucleic acid (cfDNA) molecules suitable for assaying to generate genomic data, proteins suitable for assaying to generate proteomic data, metabolites suitable for assaying to generate metabolomic data, or a mixture or combination thereof. One or more such analytes (e.g., cfRNA molecules, cfDNA molecules, proteins, or metabolites) may be isolated or extracted from one or more biological samples of a subject for downstream assaying using one or more suitable assays.
[0268] After obtaining a biological sample from the subject, the biological sample may be processed to generate datasets indicative of a lung nodule-related state of the subject. For example, a presence, absence, or quantitative assessment of nucleic acid molecules of the biological sample at a panel of lung nodule- related state-associated genomic loci (e.g., quantitative measures of RNA transcripts including intronic RNA regions or DNA at the lung nodule-related state-associated genomic loci), proteomic data comprising quantitative measures of proteins of the dataset at a panel of lung nodule-related state-associated proteins, and / or metabolome data comprising quantitative measures of a panel of lung nodule-related state-associated metabolites may be indicative of a lung nodule-related state. Processing the biological sample obtained from the subject may comprise (i) subjecting the biological sample to conditions that are sufficient to isolate, enrich, or extract a plurality of nucleic acid molecules, proteins, and / or metabolites, and (ii) assaying the plurality of nucleic acid molecules, proteins, and / or metabolites to generate the dataset.
[0269] In some embodiments, a plurality of nucleic acid molecules is extracted from the biological sample and subjected to sequencing to generate a plurality of sequencing reads. The nucleic acid molecules may comprise ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). The nucleic acid molecules (e.g., RNA or DNA) may be extracted from the biological sample by a variety of methods, such as a FastDNA Kit protocol from MP Biomedicals, a QIAamp DNA cell-free biological mini kit from Qiagen, or a cell-free biological DNA isolation kit protocol from Norgen Biotek. The extraction method may extract all RNA or DNA molecules from a sample. Alternatively, the extract method may selectively extract a portion of RNA or DNA molecules from a sample. Extracted RNA molecules from a sample may be converted to DNA molecules by reverse transcription (RT).
[0270] The sequencing may be performed by any suitable sequencing methods. The sequencing may comprise nucleic acid amplification (e.g., of RNA or DNA molecules). In some embodiments, the nucleic acid amplification is polymerase chain reaction (PCR). A suitable number of rounds of PCR (e.g., PCR, qPCR, reverse-transcriptase PCR, digital PCR, etc.) may be performed to sufficiently amplify an initial amount of nucleic acid (e.g., RNA or DNA) to a desired input quantity for subsequent sequencing. In some cases, the PCR may be used for global amplification of target nucleic acids. This may comprise using adapter sequences that may be first ligated to different molecules followed by PCR amplification using universal primers. In other cases, only certain target nucleic acids within a population of nucleic acids may be amplified. Specific primers, possibly in conjunction with adapter ligation, may be used to selectively amplify certain targets for downstream sequencing. The PCR may comprise targeted amplification of one or more genomic loci, such as genomic loci associated with lung nodule-related states. The sequencing may comprise use of simultaneous reverse transcription (RT) and polymerase chain reaction (PCR), such as a OneStep RT-PCR kit protocol by Qiagen, NEB, Thermo Fisher Scientific, or Bio-Rad.
[0271] RNA or DNA molecules isolated or extracted from a biological sample may be tagged, e.g., with identifiable tags, to allow for multiplexing of a plurality of samples. Any number of RNA or DNA samples may be multiplexed. For example, a plurality of biological samples may be tagged with sample barcodes such that each DNA molecule may be traced back to the sample (and the subject) from which the DNA molecule originated. Such tags may be attached to RNA or DNA molecules by ligation or by PCR amplification with primers.
[0272] After subjecting the nucleic acid molecules to sequencing, suitable bioinformatics processes may be performed on the sequence reads to generate the data indicative of the presence, absence, or relative assessment of the lung nodule-related state. For example, the sequence reads may be aligned to one or more reference genomes (e.g., a genome of one or more species such as a human genome). The aligned sequence reads may be quantified at one or more genomic loci to generate the datasets indicative of the lung nodule- related state. For example, quantification of sequences corresponding to a plurality of genomic loci associated with lung nodule-related states may generate the datasets indicative of the lung nodule-related state.
[0273] The biological sample may be processed without any nucleic acid extraction. For example, the lung nodule-related state may be identified or monitored in the subject by using probes configured to selectivelyenrich nucleic acid (e.g., RNA or DNA) molecules corresponding to the plurality of lung nodule-related state-associated genomic loci. The genomic loci may correspond to nucleic acids encoding the biomarkers described herein. The probes may be nucleic acid primers. The probes may have sequence complementarity with nucleic acid sequences from one or more of the plurality of lung nodule -related state-associated genomic loci or genomic regions. The plurality of lung nodule-related state-associated genomic loci or genomic regions may comprise at least 2, at least about 100, or more distinct lung nodule-related state- associated genomic loci or genomic regions. Aspects disclosed in this section related to a lung nodule or to lung cancer may be relevant to detecting another disease state or cancer. The plurality of lung nodule-related state-associated genomic loci or genomic regions may comprise one or more members (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, or more) encoding any one of the biomarkers in table or figure.
[0274] The probes may be nucleic acid molecules (e.g., RNA or DNA) having sequence complementarity with nucleic acid sequences (e.g., RNA or DNA) of the one or more genomic loci (e.g., lung nodule-related state-associated genomic loci). These nucleic acid molecules may be primers or enrichment sequences. The assaying of the biological sample using probes that are selective for the one or more genomic loci (e.g., lung nodule-related state-associated genomic loci) may comprise use of array hybridization (e.g., microarraybased), polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing). In some embodiments, DNA or RNA may be assayed by one or more of: isothermal DNA / RNA amplification methods (e.g., loop-mediated isothermal amplification (LAMP), helicase dependent amplification (HD A), rolling circle amplification (RCA), recombinase polymerase amplification (RPA)), immunoassays, electrochemical assays, surface-enhanced Raman spectroscopy (SERS), quantum dot (QD)-based assays, molecular inversion probes, droplet digital PCR (ddPCR), CRISPR / Cas-based detection (e.g., CRISPR-typing PCR (ctPCR), specific high-sensitivity enzymatic reporter un-locking (SHERLOCK), DNA endonuclease targeted CRISPR trans reporter (DETECTR), and CRISPR-mediated analog multi -event recording apparatus (CAMERA)), and laser transmission spectroscopy (LTS).
[0275] The assay readouts may be quantified at one or more genomic loci (e.g., lung nodule-related state- associated genomic loci) to generate the data indicative of the lung nodule-related state. For example, quantification of array hybridization or polymerase chain reaction (PCR) corresponding to a plurality of genomic loci (e.g., lung nodule-related state-associated genomic loci) may generate data indicative of the lung nodule-related state. Assay readouts may comprise quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized values thereof. The assay may be a home use test configured to be performed in a home setting.
[0276] In some embodiments, multiple assays are used to process biological samples of a subject. For example, a first assay may be used to process a first biological sample obtained or derived from the subject to generate a first dataset; and based at least in part on the first dataset, a second assay different from said first assay may be used to process a second biological sample obtained or derived from the subject to generate a second dataset indicative of said lung nodule-related state. The first assay may be used to screenor process biological samples of a set of subjects, while the second or subsequent assays may be used to screen or process biological samples of a smaller subset of the set of subjects. The first assay may have a low cost and / or a high sensitivity of detecting one or more lung nodule-related states (e.g., lung nodule- related complication), that is amenable to screening or processing biological samples of a relatively large set of subjects. The second assay may have a higher cost and / or a higher specificity of detecting one or more lung nodule-related states (e.g., lung nodule-related complication), that is amenable to screening or processing biological samples of a relatively small set of subjects (e.g., a subset of the subjects screened using the first assay). The second assay may generate a second dataset having a specificity (e.g., for one or more lung nodule-related states such as lung nodule -related complications) greater than the first dataset generated using the first assay. As an example, one or more biological samples may be processed using a cfRNA assay on a large set of subjects and subsequently a metabolomics assay on a smaller subset of subjects, or vice versa. The smaller subset of subjects may be selected based at least in part on the results of the first assay.
[0277] Alternatively, multiple assays may be used to simultaneously process biological samples of a subject. For example, a first assay may be used to process a first biological sample obtained or derived from the subject to generate a first dataset indicative of the lung nodule-related state; and a second assay different from the first assay may be used to process a second biological sample obtained or derived from the subject to generate a second dataset indicative of the lung nodule-related state. Any or all of the first dataset and the second dataset may then be analyzed to assess the lung nodule-related state of the subject. For example, a single diagnostic index or diagnosis score can be generated based on a combination of the first dataset and the second dataset. As another example, separate diagnostic indexes or diagnosis scores can be generated based on the first dataset and the second dataset.
[0278] The biological samples may be processed using a metabolomics assay. For example, a metabolomics assay can be used to identify a quantitative measure (e.g., indicative of a presence, absence, or relative amount) of each of a plurality of lung nodule-related state-associated metabolites in a biological sample of the subject. The metabolomics assay may be configured to process biological samples such as a blood sample or a urine sample (or derivatives thereof) of the subject. A quantitative measure (e.g., indicative of a presence, absence, or relative amount) of lung nodule-related state-associated metabolites in the biological sample may be indicative of one or more lung nodule-related states. The metabolites in the biological sample may be produced (e.g., as an end product or a byproduct) as a result of one or more metabolic pathways corresponding to lung nodule-related state-associated genes. Assaying one or more metabolites of the biological sample may comprise isolating or extracting the metabolites from the biological sample. The metabolomics assay may be used to generate datasets indicative of the quantitative measure (e.g., indicative of a presence, absence, or relative amount) of each of a plurality of lung nodule-related state-associated metabolites in the biological sample of the subject.
[0279] The biological samples may be processed using a methylation-specific assay. For example, a methylation-specific assay can be used to identify a quantitative measure (e.g., indicative of a presence, absence, or relative amount) of methylation each of a plurality of lung nodule-related state-associatedgenomic loci in a biological sample of the subject. The methylation-specific assay may be configured to process biological samples such as a blood sample or a urine sample (or derivatives thereof) of the subject. A quantitative measure (e.g., indicative of a presence, absence, or relative amount) of methylation of lung nodule-related state-associated genomic loci in the biological sample may be indicative of one or more lung nodule-related states. The methylation-specific assay may be used to generate datasets indicative of the quantitative measure (e.g., indicative of a presence, absence, or relative amount) of methylation of each of a plurality of lung nodule-related state-associated genomic loci in the biological sample of the subject.
[0280] The biological samples may be processed using a proteomics assay. For example, a proteomics assay can be used to identify a quantitative measure (e.g., indicative of a presence, absence, or relative amount) of each of a plurality of lung nodule-related state-associated proteins or polypeptides in a biological sample of the subject. The proteomics assay may be configured to process biological samples such as a blood sample or a urine sample (or derivatives thereof) of the subject. A quantitative measure (e.g., indicative of a presence, absence, or relative amount) of lung nodule -related state-associated proteins or polypeptides in the biological sample may be indicative of one or more lung nodule-related states. The proteins or polypeptides in the biological sample may be produced (e.g., as an end product or a byproduct) as a result of one or more biochemical pathways corresponding to lung nodule-related state-associated genes. Assaying one or more proteins or polypeptides of the biological sample may comprise isolating or extracting the proteins or polypeptides from the biological sample. The proteomics assay may be used to generate datasets indicative of the quantitative measure (e.g., indicative of a presence, absence, or relative amount) of each of a plurality of lung nodule-related state-associated proteins or polypeptides in the biological sample of the subject.
[0281] The proteomics assay may analyze a variety of proteins or polypeptides in the biological sample, such as proteins made under different cellular conditions (e.g., development, cellular differentiation, or cell cycle). The proteomics assay may comprise, for example, one or more of: an antibody-based immunoassay, an Edman degradation assay, a mass spectrometry-based assay (e.g., matrix-assisted laser desorption / ionization (MALDI) and electrospray ionization (ESI)), a top-down proteomics assay, a bottom- up proteomics assay, a mass spectrometric immunoassay (MSIA), a stable isotope standard capture with anti-peptide antibodies (SISCAP A) assay, a fluorescence two-dimensional differential gel electrophoresis (2-D DIGE) assay, a quantitative proteomics assay, a protein microarray assay, or a reverse -phased protein microarray assay. The proteomics assay may detect post-translational modifications of proteins or polypeptides (e.g., phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, and nitrosylation). The proteomics assay may identify or quantify one or more proteins or polypeptides from a database (e.g., Human Protein Atlas, PeptideAtlas, and UniProt).
[0282] Such descriptive labels may provide an identification of secondary clinical tests that may be appropriate to perform on the subject, and may comprise, for example, an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X- ray, a positron emission tomography (PET) scan, a PET-CT scan, a cell-free biological cytology, anamniocentesis, or any combination thereof. For example, such descriptive labels may provide a prognosis of the lung nodule-related state of the subject.
[0283] A method may comprise collecting tissue or a cell from a biological sample. The tissue or cell may be collected from a tissue or liquid biological sample. The tissue or cell may be collected directly from a patient. The tissue or cell may be collected from tissue suspected to be cancerous or premalignant. In some cases, the tissue or cell is selected from a biological sample isolated from a patient. The method may comprise identifying a cell or tissue subsection of interest from the biological sample. For example, a method may comprise isolating lung tissue in a transthoracic lung biopsy, identifying potentially cancerous cells through immunohistological staining, and isolating a potentially cancerous cell for further analysis.
[0284] A method may comprise parallel analysis of two or more species. The species may be compared to determine a disease state (e.g., the type and stage of a disease) of a sample. The species may originate from a single subject (e.g., a single patient suspected of having early-stage non-small cell lung cancer), or from different subjects (e.g., a health patient and a lung cancer patient). The species may comprise a healthy species and a diseased or potentially diseased species. The species may be collected from the same biological sample, for example from a single tissue section, or from different biological samples, for example from separate blood and tissue samples.
[0285] Parallel analysis of two or more species may increase the accuracy of a diagnosis. In some cases, multi-species analysis comprises a known healthy species and a suspected or known diseased species (e.g., a cell from healthy tissue and a cell from cancerous tissue). Analysis of the healthy and diseased species may identify the stage of disease of the diseased species. In some cases, the first species may be suspected of comprising a disease and the second species (e.g., a portion of a plasma sample) may comprise potential biomarkers for that disease. In particular cases, the first species may be suspected of comprising a disease and the second species may comprise blood or a portion of a blood sample (e.g., plasma or a buffy coat). For example, a squamous cell may be identified as cancerous through DNA sequencing, and then identified as an early-stage cancer cell based on a plasma proteomic profile of the patient.Multi-omics Databases
[0286] Described herein, in some aspects, are multi-omics databases. The database may be included in methods of obtaining or querying the multi-omics database. The database may be useful for querying. The multi-omics database can be generated by collecting multi-omics data from two or more populations that differ in some characteristic between the two or more populations. For example, the first population may be a population of subjects with a disease state such as cancer, and the second population may include a matched healthy control population. The differences can be differences among non-disease state, a first disease state, a second disease state, any additional disease states, or progression of any one of the disease states (e.g., early stage or late stage of a disease state. In some embodiments, multi-omics data can be comparison between one disease state to another disease state (e.g., between lung cancer and COPD). By querying multi-omics database and comparing the multi-omics database among the non-disease state or one or more disease states, candidate biomarkers (e.g., biomarkers indicative of a disease state or treatment for the disease state) can be identified. In some embodiments, the candidate biomarkers can validate biomarkersobtained from other methodologies or other sources. In some embodiments, the multi -omics database being queried can include proteomics, metabolomics, lipidomics, transcriptomics including intron RNA sequencing information, fragmentomics, methylomics, or genomics, or any combination thereof. Such diverse sets of multi -omics databases present an improvement in bother specificity and sensitivity in identifying or validating biomarkers associated with a disease state described herein compared to other methodologies that do not utilize the multi -omics database.
[0287] Described herein, in some embodiments, is a method, comprising: obtaining a multi-omics database comprising multi-omics data generated from biofluid samples of a population having varying disease states and patient characteristics, wherein the multi-omics data comprises proteomics, metabolomics, lipidomics, transcriptomics including intron RNA sequencing information, fragmentomics, methylomics, and genomics; and querying the multi-omics database to identify a biomarker or set of biomarkers capable of distinguishing individuals of the population as having a first disease state or patient characteristic from other individuals of the population as having a second disease state or patient characteristic. In some embodiments, the querying comprises identifying a biomarker or set of biomarkers as useful for identifying a third disease state or patient characteristic and determining that the biomarker or set of biomarkers is also useful for identifying the first or second first disease state or patient characteristic. In some embodiments, the querying comprises identifying another biomarker or set of biomarkers as useful for distinguishing individuals of the population as having the first disease state or patient characteristic from other individuals of the population as having the second disease state or patient characteristic and determining that the biomarker or set of biomarkers correlates with the other biomarker or set of biomarkers among individuals of the population. In some embodiments, the querying comprises comparing or correlating measurements values of the multi-omics data. In some embodiments, the querying the multi-omics database comprises correlating values of the multi-omics data with the first or second disease state or patient characteristic. In some embodiments, the querying comprises the use of machine learning.
[0288] In some embodiments, the multi-omics data are generated using untargeted omic measurement methods. In some embodiments, at least some of the multi-omics data are generated after using nanoparticle enrichment. In some embodiments, the biomarker or set of biomarkers comprises a secreted biomarker. In some embodiments, the biomarker or set of biomarkers comprises a protein, a lipid, a nucleic acid, an intron RNA sequencing information, a metabolite, or a combination thereof. In some embodiments, the set of biomarkers corresponds to a metabolic pathway. In some embodiments, the first disease state or patient characteristic comprises a cancer state (e.g., lung nodule, lung cancer, or pancreatic cancer). In some embodiments, the first or second disease state or patient characteristic comprises a comorbid state. In some embodiments, the comorbid state comprises a state associated with lung nodule, lung cancer, or pancreatic cancer. In some embodiments, the second disease state or patient characteristic comprises a healthy state. In some embodiments, the healthy state can be indicative of a likelihood of a subject subsequently developing a disease. For example, the multi-omics assessment method can assess the likelihood of a lung nodule developing into lung cancer in an otherwise healthy subject. In some embodiments, the first or second patient characteristic comprises age, sex, race, weight, height, dietary consumption, exercise habits, anactivity level, or smoking status. In some embodiments, the method further comprises using the biomarker or set of biomarkers to classify a subject as having the first disease state or patient characteristic or as having the second disease state or patient characteristic. In some embodiments, the method further comprises identifying, recommending, or administering a disease treatment based on an use of the biomarker or set of biomarkers. In some embodiments, the biofluid samples comprise blood, serum, or plasma samples. In some embodiments, the population comprises human subjects. In some embodiments, the multi-omics assessment method suggests treatment for a disease state described herein. In some embodiments, the treatment suggested by the multi-omics assessment method increases therapeutic effectiveness compared to other convention treatment for the disease state.
[0289] Described herein, in some aspects, is a method, comprising: obtaining multi-omics data from one or more biofluid samples of a subject identified as having a lung nodule; and applying a classifier to the multi- omics data to evaluate whether the lung nodule is cancerous or non-cancerous. In some embodiments, the multi-omics data comprise metabolomic, lipidomic, proteomic, or transcriptomic data including intron RNA sequencing information. In some embodiments, the proteomic data comprise targeted proteomic data. In some embodiments, the proteomic data comprise untargeted proteomic data. In some embodiments, the transcriptomic data comprise mRNA data and intron RNA sequencing information. In some embodiments, the transcriptomic data comprise microRNA data. In some embodiments, the classifier performs with an area under the curve of at least about 0.6, as determined in a receiver operating characteristic curve, when distinguishing biofluid samples as indicative of lung nodules being cancerous or not.
[0290] Described herein, in some aspects, is a method, comprising: obtaining multi-omics data from one or more biofluid samples of a subject suspected of having pancreatic cancer; and applying a classifier to the multi-omics data to evaluate a likelihood of the subject having the pancreatic cancer or not. In some embodiments, the classifier performs with an area under the curve of at least 0.85, at least 0.86, at least 0.87, at least 0.88, at least 0.89, at least 0.90, at least 0.91, at least 0.92, at least 0.93, at least 0.94, at least 0.95, at least 0.96, at least 0.97, or at least 0.98, as determined in a receiver operating characteristic curve, when distinguishing biofluid samples as indicative of the pancreatic cancer or not. In some embodiments, the pancreatic cancer comprises stage 1 or 2 pancreatic cancer. In some embodiments, the pancreatic cancer comprises stage 3 or 4 pancreatic cancer. In some embodiments, the multi-omics data comprise data on copy -number variation, fragmentomics, mRNA, intron RNA sequencing information, proteins, metabolites, or lipids. In some embodiments, the multi-omics data comprise copy-number variation data, fragmentomic data, transcriptomic data, proteomic data, metabolic data, and lipidomic data.
[0291] Described herein, in some aspects, is a method, comprising: obtaining a multi-omics database comprising multi-omics data generated from biofluid samples of a population having varying disease states and patient characteristics, wherein the multi-omics data comprises proteomics, metabolomics, lipidomics, transcriptomics including intron RNA sequencing information, fragmentomics, methylomics, and genomics; and querying the multi-omics database to identify whether a biomarker used to classify subjects as having a first disease state or patient characteristic is capable of classifying individuals of the population as having a second disease state or patient characteristic. Also described herein is a method, comprising: obtaining amulti-omics database generated from biofluid samples comprising a biomarker; and querying the database to identify a set of multi -omics biomarkers corresponding to the biomarker in a first population of subjects within the database and not to a second population of subjects within the database; wherein the first and second populations differ in the presence of a disease state or other clinical characteristic. In some aspects, described herein is a method, comprising: obtaining a multi-omics database from biofluid samples of subjects, wherein the database comprises a set of multi-omics biomarkers corresponding to a first population of the subjects; and querying the database to identify a second population of the subjects corresponding to the set of multi-omics biomarkers, wherein the first and second populations differ in the presence of a disease state or other clinical characteristic. In some aspects, described herein is a method, comprising: obtaining a multi-omics database from biofluid samples of a first population and a second population; and querying the database to identify a set of multi-omics biomarkers corresponding to the first population and a not corresponding to the second population; and generating a classifier based on the set of multi-omics biomarkers. In some aspects, described herein is a method, comprising: obtaining a multi-omics database comprising multiple types of biomarker measurements from biofluid samples of a first population and a second population; querying the database to identify a first set of multi-omics biomarkers corresponding to the first population, and a second set of multi-omics biomarkers corresponding to the second population; and generating a classifier bas...
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A multi -omic method, comprising: obtaining intronic RNA information from a biofluid sample from a subject; processing proteins from the biofluid sample to obtain proteomic data; generating a combined dataset comprising at least a portion of the intronic RNA information and at least a portion of the proteomic data; determining a presence or absence of cancer in said subject at least in part by inputting the combined dataset into a trained machine learning classifier, wherein the trained machine learning classifier is trained on a first sample of cancer samples and non-cancer samples; and generating an output using the classifier, wherein the output is indicative of the presence or absence of cancer.2 The method of claim 1, wherein the intronic RNA information or proteomic data undergoes feature selection prior to applying the classifier.3 The method of claim 1, wherein the intronic RNA information comprise alternative splicing events.4 The method of claim 1, wherein the intronic RNA information comprise immature transcripts.5 The method of claim 4, wherein the immature transcripts are pre-mRNA.6 The method of claim 4, wherein the immature transcripts have not undergone splicing.7 The method of claim 1, wherein the classifier comprises a performance characteristic comprising an average or median area under the curve (AUC) of a receiver operating characteristic (ROC) curve of at least 0.80, at least 0.85, at least 0.88, at least 0.90, at least 0.91, at least 0.92, at least 0.93, at least 0.94, at least 0.95, or at least 0.96.8 The method of claim 1, wherein the cancer is selected from the group consisting of: lung cancer, pancreatic cancer, colon cancer, liver cancer, breast cancer, and ovarian cancer.9 The method of claim 1, wherein the biofluid sample comprises a blood, serum, or plasma sample.10 The method of claim 1, wherein said intronic RNA information is generated by sequencing, microarray analysis, hybridization assay, polymerase chain reaction assay, electrophoresis assay, or a combination thereof.11 The method of claim 1, further comprising monitoring the subject when the subject does not have the cancer.
12. The method of claim 1, further comprising identifying, recommending, or administering a disease treatment based on a use of the biomarker or set of biomarkers.
13. The method of claim 1, wherein the classifier is trained-using deep learning, a hierarchical cluster analysis, a principal component analysis, a partial least squares discriminant analysis, a random forest classification analysis, a gradient boosted analysis, a gradient boosted tree ensemble analysis, a support vector machine analysis, a k-nearest neighbors analysis, a naive Bayes analysis, a K-means clustering analysis, or a hidden Markov analysis.
14. The method of claim 7, wherein the AUC is determined by application of the classifier to a data set derived from a trial of at least 20 subjects having the cancer, and over 20 control subjects not having the cancer, and wherein the ROC curve comprises a plot of true positive and false positive rates obtained when the classifier is applied to the held-out data set.
15. The method of claim 1, wherein said proteomic data is generated by mass spectrometry.
16. The method of claim 1, wherein said proteomic data is generated from biomolecules adsorbed to a plurality of nanoparticles.
17. The method of claim 16, wherein the plurality of nanoparticles comprises physiochemically distinct groups of nanoparticles.
18. The method of claim 17, wherein the physiochemically distinct groups of nanoparticles comprise lipid nanoparticles, metal nanoparticles, silica nanoparticles, or polymer nanoparticles.
19. The method of claim 1, wherein the intronic RNA information comprises one or more of the following intronic RNAs: ENST00000216044.10+chr22_38721866_38724296, ENST00000260702.4- chrl0_98261128_98262034, ENST00000262067.5+chr7_16754031_16776210, ENST00000264870.8- chr6_6195886_6197222, ENST00000285873.8-chr5_154931578_154932098, ENST00000303212.3+chrl9_5824335_5827748, ENST00000309415.8+chr2_109419643_109432500, ENST00000309765.4+chr3_44270811_44281967, ENST00000314888. 10-chr9_35714082_35714238, ENST00000316707. 10+chr4_86745129_86758259, ENST00000322313.9- chr2_127708893_127709489, ENST00000327473.9+chrl9_4639630_4651866, ENST00000329335.3- chr8_10623221_10711956, ENST00000343815.10+chrl_154220245_154225083, ENST00000355754.7-chrl_89187103_89188581, ENST00000359236.10- chrl2_l 18073993 118079345, ENST00000359314.5+chr6_47554767_47574063, ENST00000366922.3+chrl_220110938_220114313, ENST00000367059.3+chrl_207475217_207477884, ENST00000367211.6- chrl_219612268_219612624, ENST00000369780.8+chrl0_103585226_103589513, ENST00000369816.5-chr6_79926382_79947179, ENST00000370192.8-chrl_97306298_97373560,ENST00000372330.3+chr20_46013797_46014123, ENST00000372409.8+chr20_45943766_45944867, ENST00000374026.7-chr6_34606905_34654624, ENST00000374272.4-chrl_26051884_26053892, ENST00000375448.4+chrl_17342403_17346027, ENST00000376630.5+chr6_30492584_30492748, ENST00000377047.9+chrl3_93227617_93545262, ENST00000379936.3+chrl 1_6241781_6243948, ENST00000380760.4+chr7_73300639_73301066, ENST00000381340.8-chrl2_26658131_26659112, ENST00000381624.4+chrl9_5744511_5745912, ENST00000382353.6+chrl3_21672190_21681041, ENST00000389629.8+chrl5_41471350_41474619, ENST00000393409.3+chr3_127016685_127016943,ENST00000393980.8+chr5_160187455_160199048, ENST00000400405.4- chrl3_45784058_45851494, ENST00000403045.6-chr2_70648781_70663243, ENST00000403045.6- chr2_70648781_70663243, ENST00000404190.3+chrl0_88726913_88730982, ENST00000411427.3- chr21_41442747_41445489, ENST00000411433. l-chr2_218351227_218357913, ENST00000417390.1+chr6_128067597_128083670, ENST00000421239.7+chrl9_52456959_52491323, ENST00000423023.2+chrl3_19674733_19675617, ENST00000424662.1-chr20_59352322_59357308, ENST00000426173.6- chrl2_122948709_122949787, ENST00000428948.1+chr9_35646434_35646689, ENST00000429945.1-chr6_27286149_27311096, ENST00000431627.1+chrl_158880897_158883353, ENST00000435287. l+chr6_131951255J32077207,ENST00000444740.2+chrl9_41879176_41879534, ENST00000445019.5-chrl_26890852_26892669, ENST00000449581.1-chr20_7188568_7241828, ENST00000454681.2-chrl3_33383680_33439690, ENST00000468244.2+chr9_125254672_125256963, ENST00000469154.5- chr7_151243702_151275113, ENST00000469989. l+chr3_52523577_52523870, ENST00000470809.1+chrl_45787133_45824432, ENST00000486199.5+chr8_132814890_132817338, ENST00000490486.2-chr9_137008596_137008730, ENST00000500450.6- chr3_129171509_129171654, ENST00000504592.5+chr4_101721434_101829807, ENST00000505846.5+chr4_82448022_82454721, ENST00000505916.6-chr4_185378208_185394778, ENST00000510496.5-chr5_74805285_74834425, ENST00000511051.5-chr4_22346824_22389087, ENST00000511443.1-chr5_15192450_15243648, ENST00000512905.6-chr3_98521443_98562323, ENST00000513206.5+chr5_14502658_14531743, ENST00000518649.5+chrl4_91167379_91181185, ENST00000518836.5-chr5_159099464_159099582, ENST00000521442.1+chr8_20225407_20225688, ENST00000526372. 1-chrl 1_88149956_88175237, ENST00000526638.1+chrl l_75572400_75572550, ENST00000527263.1+chrl_77518585_77535917, ENST00000528313.1+chrl l_60462486_60466997, ENST00000534065. 1+chrl 1 66312993_66318833, ENST00000534068. 1+chrl 1 23731659_23794822, ENST00000535199.5-chr21_36098260_36126486, ENST00000535681.1-chrl7_80468918_80470526, ENST00000537256.5+chrl6_87601931_87644257, ENST00000538357.1- chr 12_118079607_l 18095532, ENST00000542280.5-chrl2_7479997_7482965,ENST00000543030.5+chrl6_57055520_57059466, ENST00000546412.2-chrl4_25664976_25835849, ENST00000551286. l-chrl2_56314967_56315820, ENST00000551516. l+chrl2_69251142_69262468,ENST00000551900.5-chrl2_53725259_53727403, ENST00000554119.5-chrl4_95710882_95712219, ENST00000554804.1+chrl4_100274759_100277417,ENST00000557595.1+chrl4_22556843_22556969, ENST00000559100.1-chrl5_50301369_50303965, ENST00000559494.1-chrl5_77054280_77070891, ENST00000563175.1-chrl6_68259822_68260309, ENST00000563237.3-chr7_102354922_102355863, ENST00000563730.1-chrl6_47459176_47463951, ENST00000563907.5-chrl5_72752896_72760406, ENST00000567873.1-chrl3_33167623_33350289, ENST00000568293.1+chrl6_56625880_56626879, ENST00000568759.1-chrl6_14735084_14740891,ENST00000576437.5-chrl6_4362132_4395288, ENST00000576479.4+chrl l_14254730_14255646, ENST00000576965.1+chrl7_4942942_4944714, ENST00000578379.5-chrl7_64006672_64020522, ENST00000581278. l+chrl8_24450815_24453149, ENST00000583510. l-chrl7_64855890_64859441, ENST00000583535.6-chrl7_10632017_10632475, ENST00000584348.5+chrl7_19495482_19542392, ENST00000586744.1-chrl9_45076877_45090068, ENST00000588212.1-chrl9_44340540_44397198, ENST00000588397.1-chrl8_13645198_13645530, ENST00000589143.5-chrl9_56376122_56390000,ENST00000592226.5-chrl7_44384586_44385163, ENST00000592234.5+chrl9_1271037_1271550, ENST00000595816. l-chrl9_17351643_17365972, ENST00000597430.2-chrl9_6592576_6603940, ENST00000598213.5-chrl9_57922630_57935079, ENST00000602162.5-chrl9_52706865_52735000, ENST00000602271. l+chrl9_51120384_51120509, ENST00000602644.5-chrl6_67663557_67663936, ENST00000608254.1-chrX_23064122_23270050, ENST00000614278.1-chr9_138179514_138179564,ENST00000616279.4-chr2_l 1217174 11218957, ENST00000616721.6-chrl9_39924733_39927058,ENST00000616721.6-chrl9_39928307_39934563, ENST00000617106.1-chrl9_39726221_39726702, ENST00000617305.4+chrl9_41264694_41305675, ENST00000619343.1-chr20_3921870_3923322, ENST00000630715.2-chr6_109454240_109465663, ENST00000635641.1-chr6_26196247_26196719, ENST00000635687.1-chrl_9151836_9182006, ENST00000640432.1-chrl4_94051070_94051339, ENST00000641481.1-chrl_38756606_38784539, ENST00000642991.1-chrl3_99409060_99417056,ENST00000643759.2-chrl_158652654_158653273, ENST00000648060.1-chrl3_76729511 76835980, ENST00000652524. l+chr4_83076353_83082729, ENST00000652631. l-chrl4_40286951_40347879, ENST00000652635.1+chr2_239539446_239544463, ENST00000655309.1+chr6_5072735_5079555, ENST00000656176.1+chrl3_25169032_25186523, ENST00000660939.1+chr6_132147845_132160766, ENST00000663029.1-chr6_21369811_21382738,ENST00000665235.1+chrl4_75296565_75297208, ENST00000665970.1+chr3_40719916_40773867, ENST00000667372. 1 +chr3_53196917_53197230, ENST00000675287. 1- chrl5_34066812_34070587,ENST00000677787.1-chr8_100712458_100713086, ENST00000677919.1- chr8_63017631_63024079, ENST00000680973.1+chr9_64031415_64054475, ENST00000681659.1- chrl0_49493253_49505883, or ENST00000681827.1+chrl_227728285_227728689.
20. The method of claim 19, wherein the intronic RNA information comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, or at least 120 of the following intronic RNAs: ENST00000216044.10+chr22_38721866_38724296, ENST00000260702.4-chrl0_98261128_98262034, ENST00000262067.5+chr7_16754031_16776210, ENST00000264870.8-chr6_6195886_6197222, ENST00000285873.8-chr5_154931578_154932098, ENST00000303212.3+chrl9_5824335_5827748, ENST00000309415.8+chr2_109419643_109432500, ENST00000309765.4+chr3_44270811_44281967, ENST00000314888. 10-chr9_35714082_35714238, ENST00000316707. 10+chr4_86745129_86758259, ENST00000322313.9- chr2_127708893_127709489, ENST00000327473.9+chrl9_4639630_4651866, ENST00000329335.3- chr8_10623221_10711956, ENST00000343815.10+chrl_154220245_154225083, ENST00000355754.7-chrl_89187103_89188581, ENST00000359236.10- chrl2_l 18073993 118079345, ENST00000359314.5+chr6_47554767_47574063, ENST00000366922.3+chrl_220110938_220114313, ENST00000367059.3+chrl_207475217_207477884, ENST00000367211.6- chrl_219612268_219612624, ENST00000369780.8+chrl0_103585226_103589513, ENST00000369816.5-chr6_79926382_79947179, ENST00000370192.8-chrl_97306298_97373560, ENST00000372330.3+chr20_46013797_46014123, ENST00000372409.8+chr20_45943766_45944867, ENST00000374026.7-chr6_34606905_34654624, ENST00000374272.4-chrl_26051884_26053892, ENST00000375448.4+chrl_17342403_17346027, ENST00000376630.5+chr6_30492584_30492748, ENST00000377047.9+chrl3_93227617_93545262, ENST00000379936.3+chrl 1_6241781_6243948, ENST00000380760.4+chr7_73300639_73301066, ENST00000381340.8-chrl2_26658131_26659112, ENST00000381624.4+chrl9_5744511_5745912, ENST00000382353.6+chrl3_21672190_21681041, ENST00000389629.8+chrl5_41471350_41474619, ENST00000393409.3+chr3_127016685_127016943, ENST00000393980.8+chr5_160187455_160199048, ENST00000400405.4- chrl3_45784058_45851494, ENST00000403045.6-chr2_70648781_70663243, ENST00000403045.6- chr2_70648781_70663243, ENST00000404190.3+chrl0_88726913_88730982, ENST00000411427.3- chr21_41442747_41445489, ENST00000411433. l-chr2_218351227 218357913, ENST00000417390.1+chr6_128067597_128083670, ENST00000421239.7+chrl9_52456959_52491323, ENST00000423023.2+chrl3_19674733_19675617, ENST00000424662.1-chr20_59352322_59357308, ENST00000426173.6- chrl2_122948709_122949787, ENST00000428948.1+chr9_35646434_35646689, ENST00000429945.1-chr6_27286149_27311096, ENST00000431627.1+chrl_158880897_158883353, ENST00000435287. l+chr6_131951255_132077207, ENST00000444740.2+chrl9_41879176_41879534, ENST00000445019.5-chrl_26890852_26892669, ENST00000449581.1-chr20_7188568_7241828, ENST00000454681.2-chrl3_33383680_33439690, ENST00000468244.2+chr9_125254672_125256963, ENST00000469154.5- chr7_151243702_151275113, ENST00000469989. l+chr3_52523577_52523870,ENST00000470809.1+chrl_45787133_45824432, ENST00000486199.5+chr8_132814890_132817338, ENST00000490486.2-chr9_137008596_137008730, ENST00000500450.6- chr3_129171509_129171654, ENST00000504592.5+chr4_101721434_101829807,ENST00000505846.5+chr4_82448022_82454721, ENST00000505916.6-chr4_185378208_185394778, ENST00000510496.5-chr5_74805285_74834425, ENST00000511051.5-chr4_22346824_22389087, ENST00000511443.1-chr5_15192450_15243648, ENST00000512905.6-chr3_98521443_98562323, ENST00000513206.5+chr5_14502658_14531743, ENST00000518649.5+chrl4_91167379_91181185, ENST00000518836.5-chr5_159099464_159099582, ENST00000521442.1+chr8_20225407_20225688, ENST00000526372. l-chrl l_88149956_88175237, ENST00000526638.1+chrl l_75572400_75572550, ENST00000527263.1+chrl_77518585_77535917, ENST00000528313.1+chrl l_60462486_60466997, ENST00000534065. 1+chrl 1 66312993_66318833, ENST00000534068. 1+chrl 1 23731659_23794822, ENST00000535199.5-chr21_36098260_36126486, ENST00000535681.1-chrl7_80468918_80470526, ENST00000537256.5+chrl6_87601931_87644257, ENST00000538357.1- chrl2_l 18079607_l 18095532, ENST00000542280.5-chrl2_7479997_7482965, ENST00000543030.5+chrl6_57055520_57059466, ENST00000546412.2-chrl4_25664976_25835849, ENST00000551286. l-chrl2_56314967_56315820, ENST00000551516. l+chrl2_69251142_69262468, ENST00000551900.5-chrl2_53725259_53727403, ENST00000554119.5-chrl4_95710882_95712219, ENST00000554804.1+chrl4_100274759_100277417, ENST00000557595.1+chrl4_22556843_22556969, ENST00000559100.1-chrl5_50301369_50303965, ENST00000559494.1-chrl5_77054280_77070891, ENST00000563175.1-chrl6_68259822_68260309, ENST00000563237.3-chr 7_102354922_102355863, ENST00000563730.1-chrl6_47459176_47463951, ENST00000563907.5-chrl5_72752896_72760406, ENST00000567873.1-chrl3_33167623_33350289, ENST00000568293.1+chrl6_56625880_56626879, ENST00000568759.1-chrl6_14735084_14740891, ENST00000576437.5-chrl6_4362132_4395288, ENST00000576479.4+chrl l_14254730_14255646, ENST00000576965.1+chrl7_4942942_4944714, ENST00000578379.5-chrl7_64006672_64020522, ENST00000581278. l+chrl8_24450815_24453149, ENST00000583510. l-chrl7_64855890_64859441, ENST00000583535.6-chrl7_10632017_10632475, ENST00000584348.5+chrl7_19495482_19542392, ENST00000586744.1-chrl9_45076877_45090068, ENST00000588212.1-chrl9_44340540_44397198, ENST00000588397. l-chrl8_13645198 13645530, ENST00000589143.5-chrl9_56376122_56390000, ENST00000592226.5-chrl7_44384586_44385163, ENST00000592234.5+chrl9_1271037_1271550, ENST00000595816. l-chrl9_17351643_17365972, ENST00000597430.2-chrl9_6592576_6603940, ENST00000598213.5-chr 19_57922630_57935079, ENST00000602162.5-chrl9_52706865_52735000, ENST00000602271. l+chrl9_51120384_51120509, ENST00000602644.5-chrl6_67663557_67663936, ENST00000608254.1-chrX_23064122_23270050, ENST00000614278.1-chr9_138179514_138179564, ENST00000616279.4-chr2_l 1217174_11218957, ENST00000616721.6-chrl9_39924733_39927058,ENST00000616721.6-chrl9_39928307_39934563, ENST00000617106.1-chrl9_39726221_39726702, ENST00000617305.4+chrl9_41264694_41305675, ENST00000619343.1-chr20_3921870_3923322, ENST00000630715.2-chr6_109454240_109465663, ENST00000635641.1-chr6_26196247_26196719,ENST00000635687.1-chrl_9151836_9182006, ENST00000640432.1-chrl4_94051070_94051339,ENST00000641481.1-chrl_38756606_38784539, ENST00000642991.1-chrl3_99409060_99417056,ENST00000643759.2-chrl_158652654_158653273, ENST00000648060.1-chrl3_76729511_76835980, ENST00000652524. l+chr4_83076353_83082729, ENST00000652631. l-chrl4_40286951_40347879, ENST00000652635.1+chr2_239539446_239544463, ENST00000655309.1+chr6_5072735_5079555, ENST00000656176.1+chrl3_25169032_25186523,ENST00000660939.1+chr6_132147845_132160766, ENST00000663029.1-chr6_21369811_21382738, ENST00000665235.1+chrl4_75296565_75297208, ENST00000665970.1+chr3_40719916_40773867, ENST00000667372. 1 +chr3_53196917_53197230, ENST00000675287. 1- chrl5_34066812_34070587,ENST00000677787.1-chr8_100712458_100713086, ENST00000677919.1- chr8_63017631_63024079, ENST00000680973.1+chr9_64031415_64054475, ENST00000681659.1- chrl0_49493253_49505883, or ENST00000681827.1+chrl_227728285_227728689.
21. The method of claim 1, wherein said proteomic data is generated by immunoassay.
22. The method of claim 21, wherein said proteomic data is generated without using a plurality of nanoparticles.
23. The method of claim 21, wherein said proteomic data is generated from biomolecules adsorbed to a plurality of nanoparticles.
24. A method, comprising: obtaining intronic RNA information from a biofluid sample from a subject; generating a dataset comprising at least a portion of the intronic RNA information; determining a presence or absence of cancer in said subject at least in part by inputting the dataset into a trained machine learning classifier, wherein the trained machine learning classifier has a performance characteristic comprising an area under the curve (AUC) of a receiver operating characteristic (ROC) curve of at least 0.85 and wherein the trained machine learning classifier is trained on a first sample of cancer samples and non cancer samples; and generating an output using the classifier, wherein the output is indicative of the presence or absence of cancer.
25. The method of claim 24, wherein the trained machine learning classifier has a performance characteristic comprising an area under the curve (AUC) of a receiver operating characteristic (ROC) curve of at least 0.90.
26. A multi -omic method, comprising: obtaining intronic RNA information from a biofluid sample from a subject;combining the biofluid sample with internal standard proteins; processing the internal standard proteins; processing endogenous proteins of the biofluid sample based on the internal standard proteins to obtain proteomic data generating a combined dataset comprising at least a portion of the intronic RNA information and at least a portion of the proteomic data; determining a presence or absence of cancer in said subject at least in part by inputting the combined dataset into a trained machine learning classifier, wherein the trained machine learning classifier is trained on a first sample of cancer samples and non-cancer samples; and generating an output using the classifier, wherein the output is indicative of the presence or absence of cancer.
27. A multi -omic method, comprising: obtaining intronic RNA information from a biofluid sample from a subject; processing proteins from the biofluid sample to obtain proteomic data; generating a combined dataset comprising at least a portion of the intronic RNA information and at least a portion of the proteomic data; determining a presence or absence of cancer in said subject at least in part by inputting the combined dataset into a trained machine learning classifier, wherein the trained machine learning classifier has a performance characteristic comprising an area under the curve (AUC) of a receiver operating characteristic (ROC) curve of at least 0.90 and wherein the trained machine learning classifier is trained on a first sample of cancer samples and non-cancer samples; and generating an output using the classifier, wherein the output is indicative of the presence or absence of cancer.
28. The method of any of the claims above, wherein the combined dataset further comprises genomic data.
29. The method of any of the claims above, wherein the combined dataset further comprises RNA transcriptomic data comprising mRNA or microRNA expression data.
30. The method of any of the claims above, wherein the combined dataset further comprises metabolomic data.
31. The method of any of the claims above, wherein the combined dataset further comprises lipid data.
Citation Information
Patent Citations
Multi-omic assessment
US20230223111A1