Multi-omics evaluation

A multi-omics method using proteomics and nucleic acid sequencing with high-performance classifiers addresses the challenge of inaccurate cancer detection, improving accuracy and enabling early intervention.

JP2026136233APending Publication Date: 2026-08-25PROGNOMIQ INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026085642
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-21
Filing Date
2026-05-21
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Current methods for detecting diseases such as cancer lack accuracy and early detection capabilities, which are crucial for improving treatment and prognosis.

Method used

A multi-omics approach involving proteomics and nucleic acid sequencing measurements from biological fluid samples, combined with classifiers achieving an area under the ROC curve of at least 0.9, to assess disease states and inform treatment decisions.

Benefits of technology

Enhances disease detection performance by at least 4% compared to single-omics methods, with a mean AUC of 0.9, enabling accurate early detection and personalized treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026136233000001_ABST
    Figure 2026136233000001_ABST
Patent Text Reader

Abstract

An accurate and early disease detection method to improve the treatment and prognosis of individuals with diseases such as cancer. [Solution] The present invention provides methods such as multi-omics methods for evaluating diseases such as cancer. Multi-omics methods can integrate proteomics data, transcriptomics data, genomics data, lipidomics data, or metabolomics data. Methods for screening diseases or conditions are also described herein. Methods for screening diseases or conditions from biological samples are also described herein. These methods may include evaluating whether nodules, tumors, or cysts are cancerous.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application is a U.S. provisional application No. 63 / 168,594 filed on March 31, 2021; U.S. provisional application No. 63 / 168,634 filed on March 31, 2021; U.S. provisional application No. 63 / 183,816 filed on May 4, 2021; U.S. provisional application No. 63 / 183,829 filed on May 4, 2021; and U.S. provisional application No. 63 / 183,829 filed on May 4, 2021. Provisional application No. 63 / 183,844, US provisional application No. 63 / 183,852 filed on May 4, 2021, US provisional application No. 63 / 184,498 filed on May 5, 2021, US provisional application No. 63 / 228,533 filed on August 2, 2021, US provisional application No. 63 / 228,543 filed on August 2, 2021, August Claiming the interests of U.S. Provisional Application No. 63 / 229,232 filed on 4th of the month, U.S. Provisional Application No. 63 / 229,242 filed on 4 August 2021, U.S. Provisional Application No. 63 / 256,482 filed on 15 October 2021, U.S. Provisional Application No. 63 / 278,637 filed on 12 November 2021, U.S. Provisional Application No. 63 / 288,825 filed on 13 December 2021, U.S. Provisional Application No. 63 / 288,827 filed on 13 December 2021, U.S. Provisional Application No. 63 / 312,455 filed on 22 February 2022, and U.S. Provisional Application No. 63 / 322,149 filed on 21 March 2022, each of which is incorporated herein by reference in whole. [Background technology]

[0002] There is a need for methods to accurately detect diseases such as cancer at an early stage. Accurate and early disease detection can improve the treatment and prognosis of individuals suffering from the disease. [Overview of the Initiative] [Means for solving the problem]

[0003] This specification discloses multi-omics methods in several embodiments. These methods include obtaining multi-omics data generated from one or more biological fluid samples collected from subjects suspected of having a disease, the multi-omics data including proteomics measurements and nucleic acid sequencing measurements; applying a classifier to the multi-omics data to evaluate the disease state; and any one of (i) to (iv): (i) proteomics measurements are generated after the samples of one or more biological fluid samples have undergone an enrichment protocol that enriches proteins or peptides without enriching other proteins or peptides; and (ii) proteomics measurements (iii) The classifier is generated based on the amount of protein or peptide added to a sample of one or more biological fluid samples; or (iv) the classifier includes performance characteristics including an area under the mean or median area under the receiver operating characteristic (ROC) curve of at least 0.9, determined in a dataset derived from a randomized controlled trial of at least 20 subjects with the disease and more than 20 control subjects without the disease; or (iv) the assessment includes selecting a cancer treatment based on multi-omics data, wherein the proteomics measurement is generated using mass spectrometry. In some embodiments, the proteomics measurement is generated after a sample of one or more biological fluid samples has undergone an enrichment protocol that enriches some proteins without enriching other proteins. In some embodiments, the proteomics measurement is generated from proteins adsorbed to nanoparticles. In some embodiments, the proteomics measurement is generated based on the amount of protein added to a sample of one or more biological fluid samples. In some embodiments, the protein added to the sample is labeled. In some embodiments, the nucleic acid sequencing measurement includes an mRNA sequencing measurement. In some embodiments, nucleic acid sequencing measurements include mRNA sequencing measurements and miRNA sequencing measurements. In some embodiments, multi-omics data include measurements of more than 45 peptides or proteins.In some embodiments, the evaluation performance is improved by at least 4% compared to when the classifier is applied to only one type of omics data, and the performance includes sensitivity at a given specificity, determined in a dataset derived from randomized controlled trials involving more than 25 subjects with the disease and more than 25 control subjects without the disease. In some embodiments, the classifier features a mean area under the curve (AUC) of the receiver operating characteristic (ROC) curve of at least 0.9, determined in a dataset derived from randomized controlled trials involving at least 20 subjects with the disease and more than 20 control subjects without the disease. In some embodiments, applying classifiers to multi-omics data to assess disease states includes applying a first classifier to proteomics measurements to generate a first label corresponding to the presence, absence, or possibility of a disease state; applying a second classifier to nucleic acid sequencing measurements to generate a second label corresponding to the presence, absence, or possibility of a disease state; and assessing disease states based on (a), (b), or (c): (a) an unweighted average of the first and second labels, (b) a weighted average of the first and second labels, or (c) a majority vote score based on the first and second labels. In some embodiments, assessing disease states based on a weighted average of the first and second labels is performed, where the weighted average is generated by assigning weights to the results of the first and second classifiers based on area under the ROC curve, area under the precision-recall curve, precision, precision, recall, sensitivity, F1-score, specificity, or a combination thereof. In some embodiments, applying a classifier to multi-omics data to assess a pathological condition includes: obtaining a subset of features from proteomics measurements; obtaining at least one subset of features from nucleic acid sequencing measurements; pooling a subset of features from a first omics data set and at least one subset of features from a second omics data set into a pooled set of features; and assessing the pathological condition based on the pooled set of features. In some embodiments, obtaining a subset of features from the first or second omics data includes obtaining the top-tier features based on univariate data.In some embodiments, the classifier is trained using deep learning, hierarchical cluster analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k nearest neighbor analysis, simple Bayesian analysis, K-means clustering analysis, or hidden Markov analysis. In some embodiments, the multi-omics data further includes metabolomics data. In some embodiments, the disease state includes cancer. In some embodiments, cancer is selected from the group consisting of lung cancer, pancreatic cancer, breast cancer, colorectal cancer, liver cancer, and ovarian cancer. In some embodiments, the assessment includes selecting a cancer treatment based on the multi-omics data. In some embodiments, based on the assessment, the subject is subjected to chemotherapy, pharmaceuticals, radiation, or surgical cancer treatment. In some embodiments, one or more biological fluid samples include blood, serum, or plasma samples. In some embodiments, the subject is human. This specification provides for obtaining multi-omics data generated from one or more blood, serum, or plasma samples taken from human subjects suspected of having cancer, wherein the multi-omics data includes proteomics measurements and RNA sequencing measurements; applying a classifier to the multi-omics data to assess cancer; selecting or administering cancer treatments to the subjects based on the assessment; and any one of (i) to (iii): (i) the proteomics measurements are generated after the samples of one or more blood, serum, or plasma samples are concentrated with an affinity reagent for proteins or peptides; (ii) the proteomics measurements are generated based on the amount of labeled proteins or peptides added to the samples of one or more blood, serum, or plasma samples; or (iii) the classifier is based on holdout data derived from randomized controlled trials of at least 25 subjects with the disease and more than 25 control subjects without the disease. A multi-omics method is disclosed that includes performance characteristics, which include a mean area under the receiver operating characteristic (ROC) curve (AUC) of at least 0.9, determined in the data set.In some embodiments, the proteomics measurement is generated after one or more blood, serum, or plasma samples have been concentrated with an affinity reagent. In some embodiments, the proteomics measurement is generated based on the amount of labeled protein added to one or more blood, serum, or plasma samples. In some embodiments, the classifier is characterized by a mean area under the curve (AUC) of at least 0.9 of the receiver operating characteristic (ROC) curve, determined in a dataset derived from a randomized controlled trial of at least 25 subjects with the disease and more than 25 control subjects without the disease.

[0004] This specification discloses a multi-omics disease detection method comprising: obtaining multi-omics data generated from one or more biological fluid samples collected from a subject, wherein the multi-omics data comprises first omics data including proteomics data, metabolomics data, transcriptomics data, or genomics data, and second omics data including proteomics data, metabolomics data, transcriptomics data, or genomics data different from the first omics data; and using a first classifier to assign first labels to the first omics data, including the presence, absence, or possibility of a disease state; using a second classifier to assign second labels to the second omics data, including the presence, absence, or possibility of a disease state, based on the first and second labels; and identifying the multi-omics data as indicating or not indicating a disease state. In some embodiments, the first omics data includes proteomics data, and the second omics data includes metabolomics data, transcriptomics data, or genomics data. In some embodiments, the proteomics data is generated by contacting one of several biofluid samples with particles so that the particles adsorb biomolecules, including proteins. In some embodiments, the particles include carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles. In some embodiments, the particles include a group of physiologically and chemically distinct nanoparticles. In some embodiments, the proteomics data is generated using mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blotting, dot blotting, or immunostaining, or a combination thereof.In some embodiments, genomics data or transcriptomics data are generated by sequencing, microarray analysis, hybridization, polymerase chain reaction, electrophoresis, or a combination thereof. In some embodiments, the second omics data includes transcriptomics data. In some embodiments, the transcriptomics data includes mRNA or microRNA expression data. In some embodiments, the second omics data includes genomics data. In some embodiments, the genomics data includes DNA sequence data or epigenetic data. In some embodiments, identifying multi-omics data as disease-indicating or non-disease-indicating includes identifying multi-omics data as disease-indicating or non-disease-indicating based on either a first label or a second label. In some embodiments, identifying multi-omics data as disease-indicating or non-disease-indicating includes generating or obtaining a majority vote score based on the first and second labels. In some embodiments, identifying multi-omics data as indicating or not indicating a pathological condition involves generating or obtaining a weighted average of first and second labels. In some embodiments, this involves assigning weights to the first and second classifiers based on the area under the receiver operating characteristic (ROC) curve, the area under the precision-recall curve, precision, precision, recall, sensitivity, F1 score, specificity, or a combination thereof, thereby obtaining a weighted average. In some embodiments, the first omics data is generated from a first biological fluid sample from a biological fluid sample, and the second omics data is generated from a second biological fluid sample from a biological fluid sample. In some embodiments, the first biological fluid sample is collected in a first container containing a first collection component comprising heparin, ethylenediaminetetraacetic acid (EDTA), citrate, or an antisolubilant, and the second biological fluid sample is collected in a second container containing a second collection component different from the first collection component, comprising heparin, EDTA, citrate, or an antisolubilant. In some embodiments, the multi-omics data further includes a third omics data that includes a third omics data type.The third omics data may include omics data types or subtypes different from the first and second omics data. Some embodiments include using a third classifier to assign a third label to the third omics data corresponding to the presence, absence, or possibility of a pathological condition. In some embodiments, identifying multi-omics data as indicating or not indicating a pathological condition includes identifying multi-omics data as indicating or not indicating a pathological condition based on a combination of the first, second, and third labels. Some embodiments include using a third classifier to assign a third label, including the presence, absence, or possibility of a pathological condition, to third omics data different from the first and second omics data, and identifying multi-omics data as indicating or not indicating a pathological condition based on the first and second labels includes identifying multi-omics data as indicating or not indicating a pathological condition based on the first, second, and third labels. In some embodiments, the first omics data type includes proteomics data, the second omics data type includes mRNA transcriptomics data, and the third omics data type includes microRNA transcriptomics data. In some embodiments, the system includes transmitting or outputting information related to identification. In some embodiments, the system includes recommending treatment for a disease condition.

[0005] This specification discloses, in several embodiments, methods for obtaining combined data comprising two, three, or four proteomics, metabolomics, transcriptomics, or genomics data generated from one or more biological fluid samples from a subject; and identifying the combined data as indicating or not indicating one or more pathological conditions using a classifier. In some embodiments, one or more biological fluid samples include 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more biological fluid samples. In some embodiments, the combined data are generated simultaneously. In some embodiments, simultaneous data generation comprises simultaneously assaying two, three, or four proteomics, metabolomics, transcriptomics, or genomics data. In some embodiments, simultaneous data generation comprises assaying two, three, or four proteomics, metabolomics, transcriptomics, or genomics data at separate locations on an assay substrate. In some embodiments, distinct locations include distinct wells, and the assay substrate includes an assay plate. In some embodiments, one or more biological fluid samples include two or more of whole blood samples, plasma samples, serum samples, or urine samples. In some embodiments, proteomics data are generated from one of the one or more biological fluid samples. In some embodiments, metabolomics data are generated from the biological fluid sample or from additional biological fluid samples, and the proteomics data and metabolomics data are combined to obtain combined data. In some embodiments, a classifier identifies the combined data as indicating or not indicating one or more pathological conditions with higher sensitivity or specificity than the proteomics data, metabolomics data, transcriptomics data, or genomics data alone. In some embodiments, the classifier includes features selected from the proteomics data, metabolomics data, genomics data, or transcriptomics data.In some embodiments, the classifier includes features selected from a combination of proteomics data, metabolomics data, genomics data, or transcriptomics data. In some embodiments, the classifier includes multiple classifiers. In some embodiments, the multiple classifiers include two, three, four, or more classifiers. In some embodiments, the multiple classifiers each include features selected from proteomics data, metabolomics data, genomics data, transcriptomics data, or a combination thereof. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more pathologies includes using multiple classifiers to identify combined data as indicating or not indicating one or more pathologies. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more pathologies includes selecting one of the outputs of multiple classifiers. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more pathologies includes majority voting across multiple classifiers. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more conditions involves a majority vote across a subset of multiple classifiers. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more conditions involves a weighted average of multiple classifiers. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more conditions involves a weighted average of a subset of multiple classifiers. In some embodiments, the weights of the weighted average are assigned based on the area under the receiver operating characteristic (ROC) curve. In some embodiments, the weights of the weighted average are assigned based on the area under the precision-recall curve. In some embodiments, the weights of the weighted average are assigned based on precision. In some embodiments, the weights of the weighted average are assigned based on precision. In some embodiments, the weights of the weighted average are assigned based on recall. In some embodiments, the weights of the weighted average are assigned based on sensitivity.In some embodiments, the weights of the weighted average are assigned based on the F1 score. In some embodiments, the weights of the weighted average are assigned based on the specificity.

[0006] This specification discloses methods, in several embodiments, for obtaining proteomics data generated from a biofluid sample from a subject; obtaining metabolomics, transcriptomics, or genomics data generated from a biofluid sample from a subject or from additional biofluid samples, and for combining the proteomics data with the metabolomics, transcriptomics, or genomics data to obtain combined data; and for using a classifier to identify whether the combined data indicates or does not indicate one or more pathological conditions. In some embodiments, proteomics data is generated by contacting a biofluid sample from a subject with particles so that the particles adsorb biomolecules, including proteins. In some embodiments, this includes contacting a biofluid sample from a subject with particles so that the particles adsorb biomolecules. In some embodiments, this includes analyzing the biomolecules adsorbed on the particles to generate proteomics data. In some embodiments, this includes analyzing the biofluid sample or additional biofluid samples to generate metabolomics data. In some embodiments, a classifier is used to identify combined data as indicating or not indicating one or more pathological conditions. In some embodiments, proteomics data is generated by measuring readings indicating the presence, absence, or quantity of biomolecules. In some embodiments, proteomics data is generated using mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blotting, dot blotting, or immunostaining, or a combination thereof. In some embodiments, proteomics data is generated using mass spectrometry. In some embodiments, proteins include secreted proteins. In some embodiments, particles include nanoparticles. In some embodiments, particles include lipid particles, metal particles, silica particles, or polymer particles.In some embodiments, the particles include carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles. In some embodiments, the particles include a group of physiologically and chemically distinct nanoparticles. In some embodiments, the metabolomics data is generated from a different biological fluid sample than the proteomics data. In some embodiments, the metabolomics data is generated using mass spectrometry, electrophoresis, colorimetric assays, fluorescence assays, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assays, immunoassays, or a combination thereof. In some embodiments, the metabolomics data is generated using mass spectrometry. In some embodiments, the metabolomics data is generated from the same biological fluid sample as the proteomics data. In some embodiments, the metabolomics data is generated by analyzing analytes adsorbed onto the particles. In some embodiments, the metabolomics data includes lipid metabolite data, carbohydrate metabolite data, vitamin metabolite data, cofactor metabolite data, or a combination thereof. In some embodiments, the biological fluid sample includes a blood sample, plasma sample, or serum sample. In some embodiments, additional biological fluid samples are collected from the subject in a separate container from the biological fluid sample. In some embodiments, combined data is generated from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more samples. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more samples are collected separately in 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more containers. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more containers contain multiple components in addition to the sample. In some embodiments, the biological fluid sample and additional biological fluid samples are collected in separate containers, and these separate containers contain different components. In some embodiments, a first of the separate containers contains a first component that is different from a second component in a second of the separate containers.In some embodiments, the biological fluid sample includes serum; is collected in a container containing ethylenediaminetetraacetic acid (EDTA), citrate, or heparin; or contains a preservative to prevent cell lysis. In some embodiments, the biological fluid sample is collected in a container containing ethylenediaminetetraacetic acid (EDTA). In some embodiments, the additional biological fluid sample includes a blood sample, plasma sample, or serum sample. In some embodiments, the additional biological fluid sample is processed to obtain cell-free DNA or RNA. In some embodiments, the process involves obtaining genomics or transcriptomics data generated from the biological fluid sample from the subject, from the additional biological fluid sample, or from a third biological fluid sample. In some embodiments, the combined data further includes genomics or transcriptomics data. In some embodiments, the process involves analyzing the biological fluid sample, the additional biological fluid sample, or the third biological fluid sample to generate genomics or transcriptomics data. In some embodiments, the third biological fluid sample includes a blood sample, plasma sample, or serum sample. In some embodiments, a third biological fluid sample is processed to obtain cell-free DNA or RNA. In some embodiments, a classifier is used to identify combined data as indicating or not indicating one or more pathological conditions. In some embodiments, genomics or transcriptomics data is generated by measuring readings indicating the presence, absence, or quantity of nucleic acids. In some embodiments, genomics or transcriptomics data is generated by sequencing, microarray analysis, hybridization, polymerase chain reaction, electrophoresis, or a combination thereof. In some embodiments, genomics or transcriptomics data is generated from a different biological fluid sample than the metabolomics data. In some embodiments, genomics or transcriptomics data is generated from the same biological fluid sample as the metabolomics data.In some embodiments, genomics data or transcriptomics data are generated from a different biological fluid sample than the p-data. In some embodiments, genomics data or transcriptomics data are generated from the same biological fluid sample as the proteomics data. In some embodiments, genomics data or transcriptomics data are generated by analyzing nucleic acids adsorbed to particles. In some embodiments, genomics data or transcriptomics data includes genomics data. In some embodiments, genomics data includes DNA sequence data. In some embodiments, genomics data includes DNA polymorphism data. In some embodiments, genomics data includes epigenetic data. In some embodiments, genomics data includes DNA methylation data. In some embodiments, epigenetic data includes histone modification data. In some embodiments, histone modification data includes acetylation data, methylation data, ubiquitination data, phosphorylation data, SUMOylation data, ribosylation data, or citrullination data. In some embodiments, genomics data or transcriptomics data includes transcriptomics data. In some embodiments, transcriptomics data includes RNA sequence data. In some embodiments, transcriptomics data includes RNA expression data. In some embodiments, transcriptomics data includes mRNA, tRNA, rRNA, microRNA, snRNA, snoRNA, or lncRNA expression data. In some embodiments, transcriptomics data includes mRNA expression data. In some embodiments, transcriptomics data includes microRNA expression data. In some embodiments, the classifier includes features for identifying combined data as representing one or more disease states. In some embodiments, features include control protein measurements, control metabolite measurements, control nucleic acid measurements, mass spectra, m / z ratios, chromatography results, immunoassay results, light or fluorescence intensity, or sequence information.In some embodiments, the classifier is trained using deep learning, hierarchical cluster analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k nearest neighbor analysis, simple Bayesian analysis, K-means clustering analysis, or hidden Markov analysis. In some embodiments, one or more disease states include one or more cancers. In some embodiments, one or more cancers include lung cancer, breast cancer, prostate cancer, colorectal cancer, colon cancer, melanoma, bladder cancer, lymphoma, leukemia, kidney cancer, uterine cancer, pancreatic cancer, or a combination thereof. In some embodiments, the classifier distinguishes one or more disease states. In some embodiments, the classifier distinguishes lung cancer, colon cancer, and pancreatic cancer. In some embodiments, the classifier distinguishes lung cancer, colon cancer, and pancreatic cancer. In some embodiments, lung cancer includes non-small cell lung cancer (NSCLC). Some embodiments include generating a report based on the use of a classifier to identify combined data as indicating or not indicating one or more conditions. In some embodiments, the report includes the likelihood or indicator that a biological fluid or subject has one or more conditions. Some embodiments include outputting or transmitting the report. In some embodiments, the report is used by a healthcare professional when making a diagnosis, providing medical advice, or providing treatment for at least one of the one or more conditions. Some embodiments include identifying combined data as indicating one or more conditions. In some embodiments, if one or more conditions include cancer and the combined data is identified as indicating cancer, it further includes recommending cancer treatment for the subject. In some embodiments, if one or more conditions include cancer and the combined data is identified as indicating cancer, it further includes administering cancer treatment to the subject. In some embodiments, cancer treatment includes chemotherapy, radiotherapy, ablation therapy, embolization, or surgery. Some embodiments include using a classifier to identify combined data as indicating a first of the one or more conditions and not indicating a second of the one or more conditions. Some aspects include administering or recommending treatment for the first condition rather than the second condition.Some aspects include identifying combinatorial data as not indicative of one or more medical conditions. Some aspects include observing a subject without providing treatment to the subject when the combinatorial data is identified as not indicative of one or more medical conditions. In some aspects, observing the subject without providing treatment includes analyzing biomolecules in a biological fluid sample later obtained from the subject. In some aspects, the subject is a mammal. In some aspects, the subject is a human. In some aspects, the classifier includes features selected from proteomic data, metabolomic data, genomic data, or transcriptomic data. In some aspects, the classifier includes proteomic data, metabolomic data, genomic data, or transcriptomic data. The classifier includes features selected from a combination of data. In some embodiments, the classifier includes multiple classifiers. In some embodiments, the multiple classifiers include two, three, four, or more classifiers. In some embodiments, the multiple classifiers each include features selected from proteomics data, metabolomics data, genomics data, transcriptomics data, or a combination thereof. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more conditions includes using multiple classifiers to identify combined data as indicating or not indicating one or more conditions. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more conditions includes selecting one output from any of the multiple classifiers. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more conditions includes a majority vote across multiple classifiers. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more conditions includes a majority vote across a subset of multiple classifiers. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more conditions includes a weighted average of multiple classifiers. In some embodiments, using a classifier to identify combined data as indicating or not indicating one or more conditions includes a weighted average of a subset of multiple classifiers. In some embodiments, the weights of the weighted average are assigned based on the area under the receiver operating characteristic (ROC) curve. In some embodiments, the weights of the weighted average are assigned based on the area under the precision-recall curve. In some embodiments, the weights of the weighted average are assigned based on precision. In some embodiments, the weights of the weighted average are assigned based on precision. In some embodiments, the weights of the weighted average are assigned based on recall. In some embodiments, the weights of the weighted average are assigned based on sensitivity. In some embodiments, the weights of the weighted average are assigned based on the F1 score. In some embodiments, the weights of the weighted average are assigned based on specificity.

[0007] This specification discloses a method, in several embodiments, for obtaining multi-omics data generated from one or more biological fluid samples collected from a subject, wherein the multi-omics data comprises first omics data and second omics data, wherein the first omics data comprises a first omics data type including proteomics data, metabolomics data, transcriptomics data, or genomics data, and the second omics data comprises a second omics data type different from the first omics data type, and includes proteomics data, metabolomics data, transcriptomics data, or genomics data; identifying a subset of first features from the first omics data; identifying a subset of second features from the second omics data; pooling the subsets of first and second features; and identifying the multi-omics data as indicative of or not indicative of a pathological condition based on the pooled subsets of features. In some embodiments, identifying a subset of first or second features from first or second omics data includes obtaining univariate data for the features of first or second omics data, and identifying the first or second subset based on the univariate data. In some embodiments, the subset of first or second features is identified from the features of a classifier for first or second omics data. In some embodiments, identifying a subset of first or second features from first or second omics data includes obtaining a classifier for first or second omics data, and identifying the first or second subset as the top-level features of the classifier. In some embodiments, identifying a subset of first or second features from first or second omics data includes obtaining a classifier for first or second omics data, removing one or more features from the classifier at a time, and identifying which features, when removed from the classifier, degrade the classifier's performance.

[0008] In some embodiments, the disease or disorder includes pancreatic cancer. In some aspects, multi-omics cancer detection methods for detecting pancreatic cancer are disclosed herein. In some aspects, methods for detecting pancreatic cancer in a subject are disclosed, comprising: identifying a subject at risk of having pancreatic cancer; obtaining a biofluid sample from the subject; contacting the biofluid sample with particles so that the particles adsorb biomolecules, including proteins, onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data as indicating or not indicating pancreatic cancer. In some aspects, methods are disclosed herein, comprising: assaying proteins in a biofluid sample obtained from a subject identified as at risk of pancreatic cancer to obtain protein measurements; and applying a classifier to the protein measurements to identify the protein measurements as indicating that the subject has pancreatic cancer, wherein the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and assaying the proteins adsorbed onto the particles. This specification discloses a therapeutic method that, in several embodiments, includes: identifying a tumor within the target pancreas; obtaining a biofluid sample from the target; contacting the biofluid sample with particles such that the particles adsorb protein-containing biomolecules onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data as indicating a tumor containing pancreatic cancer or as not indicating a tumor containing pancreatic cancer.In some aspects, a method for evaluating a subject suspected of having pancreatic cancer, the method comprising measuring a biomarker in a biological fluid sample from the subject, the biomarker comprising A2GL, AKR1B1, ANPEP, ANTXR1, ANTXR2, BTK, CALR, CDH1, CDH11, CDH2, CDHR2, CILP2, CLEC3B, COL18A1, CRP, EXT1, F13A1, FAT1, FGL1, FLT4, ICAM1, IDH2, LCN2, LPP, MAPK1, MAP2K1, MYH9, NOTCH1, NOTCH2, PIGR, PPP2R1A, PRKAR1A, PXDN, RELN, RHOA, S100A8, S100A9, S100A12, SAA1, SAA2, SERPINA3, SLAIN2, SND1, SVEP1, TSP2, TUBB, TUBB1, or VCAN, is disclosed. In some aspects, the present disclosure provides a method comprising assaying a biomolecule in a biological fluid sample obtained from a subject suspected of having pancreatic cancer to obtain a biomolecule measurement value; and applying a classifier to the biomolecule measurement value to identify a protein measurement value as indicating that the subject has pancreatic cancer or as indicating that the subject does not have pancreatic cancer, the classifier being characterized by a receiver operating characteristic (ROC) curve having an area under the curve (AUC) greater than 0.7, greater than 0.75, greater than 0.8, greater than 0.85, greater than 0.9, greater than 0.91, greater than 0.92, greater than 0.93, or greater than 0.94 based on the biomolecule measurement features. In some aspects, the AUC is less than or equal to 0.75, less than or equal to 0.8, less than or equal to 0.85, less than or equal to 0.9, less than or equal to 0.91, less than or equal to 0.92, less than or equal to 0.93, less than or equal to 0.94, less than or equal to 0.95, or less than or equal to 0.96. In some embodiments, the biomolecule comprises a protein, a lipid, or a metabolite, or a combination thereof.

[0009] In some embodiments, the disease or disorder includes liver cancer. In some aspects, multi-omics cancer detection methods for detecting liver cancer are disclosed herein. In some aspects, methods for detecting liver cancer in a subject are disclosed, comprising: identifying the subject as being at risk of having liver cancer; obtaining a biofluid sample from the subject; contacting the biofluid sample with particles so that the particles adsorb biomolecules, including proteins, onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data as indicating liver cancer or not indicating liver cancer. In some aspects, methods are disclosed herein, comprising: assaying proteins in a biofluid sample obtained from a subject identified as being at risk of liver cancer to obtain protein measurements; and applying a classifier to the protein measurements to identify the protein measurements as indicating that the subject has liver cancer, wherein the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and assaying the proteins adsorbed onto the particles. This specification discloses a therapeutic method, in several embodiments, comprising: identifying a tumor in the liver of a subject; obtaining a biofluid sample from the subject; contacting the biofluid sample with particles such that the particles adsorb protein-containing biomolecules onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data as indicating a tumor containing liver cancer or as not indicating liver cancer. This specification also discloses a method for detecting liver cancer in a subject, in several embodiments, comprising: identifying the subject as being at risk of having liver cancer; obtaining a biofluid sample from the subject; assaying lipids in the biofluid sample to obtain lipid data; and classifying the lipid data as indicating liver cancer or as not indicating liver cancer.

[0010] In some embodiments, the disease or disorder includes ovarian cancer. In several embodiments, multi-omics cancer detection methods for detecting ovarian cancer are disclosed herein. In several embodiments, methods for detecting ovarian cancer in a subject are disclosed, comprising: identifying the subject as being at risk of having ovarian cancer; obtaining a biofluid sample from the subject; contacting the biofluid sample with particles such that the particles adsorb protein-containing biomolecules onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data as indicating or not indicating ovarian cancer. In some embodiments, identifying a subject as being at risk of having ovarian cancer includes identifying a subject that has undergone a computed tomography (CT) scan indicating ovarian cancer, undergone a magnetic resonance imaging (MRI) scan indicating ovarian cancer, undergone a positron emission tomography (PET) scan indicating ovarian cancer, undergone a transvaginal ultrasound examination indicating ovarian cancer, has elevated cancer antigen (CA)-125 levels compared to a control or baseline measurement, or has an ovarian cyst, or a combination thereof. This specification discloses a method comprising, in several embodiments, assaying proteins in a biofluid sample obtained from a subject identified as being at risk of ovarian cancer to obtain a protein measurement; and applying a classifier to the protein measurement to identify the protein measurement as indicating that the subject has ovarian cancer, wherein the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and assaying the proteins adsorbed to the particles. In some embodiments, the proteins include ANTXR2, BMP1, CILP, EIF2AK2, ENO3, F13B, FGL1, or PEBP4.This specification discloses a therapeutic method comprising, in several embodiments, identifying a tumor in the ovary of a subject; obtaining a biofluid sample from the subject; contacting the biofluid sample with particles such that the particles adsorb protein-containing biomolecules onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data as indicating a tumor containing ovarian cancer or as not indicating ovarian cancer. This specification also discloses a method for detecting ovarian cancer in a subject comprising, in several embodiments, identifying the subject as being at risk of having ovarian cancer; obtaining a biofluid sample from the subject; assaying lipids in the biofluid sample to obtain lipid data; and classifying the lipid data as indicating ovarian cancer or as not indicating ovarian cancer. In some embodiments, the lipids include one or more phospholipids.

[0011] In some embodiments, the disease or disorder includes colorectal cancer. In some aspects, multi-omics cancer detection methods for detecting colorectal cancer are disclosed herein. In some aspects, methods for detecting colorectal cancer in a subject are disclosed, comprising: identifying the subject as being at risk of having colorectal cancer; obtaining a biofluid sample from the subject; contacting the biofluid sample with particles so that the particles adsorb biomolecules, including proteins, onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data as indicating colorectal cancer or not indicating colorectal cancer. In some aspects, methods are disclosed herein, comprising: assaying proteins in a biofluid sample obtained from a subject identified as being at risk of colorectal cancer to obtain protein measurements; and applying a classifier to the protein measurements to identify the protein measurements as indicating that the subject has colorectal cancer, wherein the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and assaying the proteins adsorbed onto the particles. In some embodiments, subjects are identified as being at risk of having colorectal cancer by identifying subjects who have undergone computed tomography (CT) scans indicating colorectal cancer, liver function tests (LFT) indicating colorectal cancer, elevated carcinoembryonic antigen (CEA) levels compared to control or baseline measurements, blood in the stool, fecal immunochemistry (FIT) indicating colorectal cancer, or the presence of colorectal nodules, or a combination thereof. Disclosed herein are therapeutic methods, in some embodiments, that include identifying a mass in the colon of a subject; obtaining a biofluid sample from the subject; contacting the biofluid sample with particles such that the particles adsorb protein-containing biomolecules onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data as indicating a mass containing colorectal cancer or as not indicating colorectal cancer.

[0012] This specification describes, in several aspects, the assay of proteins in biofluid samples obtained from subjects identified as having pulmonary nodules to obtain protein measurements; and the application of a classifier to the protein measurements to evaluate the pulmonary nodules; and (i), (ii), or (iii): (i) the classifier includes protein features of the assayed protein, and the classifier includes performance characteristics for distinguishing pulmonary nodules as malignant or non-malignant, determined in a dataset derived from a randomized controlled trial of more than 20 subjects with malignant pulmonary nodules and more than 20 control subjects with non-malignant pulmonary nodules, without including clinical features in the classifier. Disclosed is a method that includes performance characteristics including a mean or median area under the curve (AUC) of a receiver operating characteristic (ROC) curve greater than 0.65 (e.g., greater than 0.7) as determined in a dataset; (ii) the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and assaying the proteins adsorbed to the particles; or (iii) assaying a protein includes contacting a biological fluid sample with particles to adsorb proteins to the particles, and obtaining a protein measurement value from the adsorbed protein. In some embodiments, the classifier includes protein features of the assayed protein and features a mean ROC curve with a median AUC greater than 0.7 when distinguishing pulmonary nodules as malignant or non-malignant, where the AUC greater than 0.7 is determined without including non-protein features in a dataset derived from a randomized controlled trial of more than 20 subjects with malignant pulmonary nodules and more than 20 control subjects with non-malignant pulmonary nodules. In some embodiments, the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and then assaying the proteins adsorbed to the particles. In some embodiments, assaying proteins includes contacting a biological fluid sample with particles to adsorb proteins to the particles, and obtaining protein measurements from the adsorbed proteins.In some embodiments, the classifier is trained using deep learning, hierarchical clustering analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k nearest neighbor analysis, simple Bayes analysis, K-means clustering analysis, or hidden Markov analysis. In some embodiments, evaluating lung nodules includes identifying protein measurements that indicate the lung nodules are cancerous. In some embodiments, the subject is subjected to lung cancer treatment based on the evaluation. In some embodiments, lung cancer treatment includes chemotherapy, radiotherapy, percutaneous ablation, radiofrequency ablation, cryoablation, microwave ablation, chemoembolization, or surgery. In some embodiments, the subject is identified as having lung nodules through the use of medical imaging devices. In some embodiments, the classifier identifies lung cancer with sensitivity and specificity greater than 60%. In some embodiments, the particles include nanoparticles. In some embodiments, the particles include lipid particles, metal particles, silica particles, or polymer particles. In some embodiments, the particles include physiologically and chemically distinct groups of nanoparticles. In some embodiments, the biological fluid sample includes a blood, serum, or plasma sample. In some embodiments, the subject is human. In some embodiments, the protein measurement includes a measurement of a protein selected from the group consisting of APP, IGGH2, SERPING1, SAA2, SERPINF2, GC, IGHA1, HPR, SERPINA3, IGHA1, LTF, SERPINA1, PCSK6, PROS1, BPIF1, C6, CP, A2M, and IGFBP2.This specification describes, in several aspects, a method for obtaining protein measurements by assaying proteins in blood, serum, or plasma samples by mass spectrometry, wherein the samples are obtained from human subjects identified as having pulmonary nodules using a medical imaging device; evaluating the pulmonary nodules by applying a classifier to the protein measurements; and selecting or administering a lung cancer treatment to the subjects based on the evaluation; and (i), (ii), or (iii): (i) the classifier includes protein features of the assayed protein, and the classifier includes performance characteristics for distinguishing pulmonary nodules as cancerous or non-cancerous, with more than 25 subjects having cancerous pulmonary nodules and more than 25 controls having non-cancerous pulmonary nodules. Disclosed is a method comprising: (ii) including performance characteristics including an area under the curve (AUC) of the median receiver operating characteristic (ROC) curve greater than 0.7, determined in a holdout dataset derived from a randomized controlled trial and using only the protein features of the classifier; (ii) the classifier is generated using proteomics data obtained by contacting a training sample with nanoparticles so that the nanoparticles adsorb proteins in the training sample, and assaying the proteins adsorbed to the nanoparticles; or (iii) assaying proteins includes contacting a blood, serum, or plasma sample with nanoparticles to adsorb proteins to the nanoparticles, and obtaining protein measurements from the adsorbed proteins.

[0013] In some embodiments, the classifier includes protein features of the assayed protein and features a mean ROC curve with a median AUC greater than 0.7 when distinguishing pulmonary nodules as malignant or non-malignant, where an AUC greater than 0.7 is determined without including non-protein features in a holdout dataset derived from a randomized controlled trial of more than 25 subjects with malignant pulmonary nodules and more than 25 control subjects with non-malignant pulmonary nodules. In some embodiments, the classifier is generated using proteomics data obtained by contacting a training sample with nanoparticles so that the nanoparticles adsorb proteins in the training sample, and then assaying the proteins adsorbed to the nanoparticles. In some embodiments, assaying proteins includes contacting a blood, serum, or plasma sample with nanoparticles to adsorb proteins to the nanoparticles, and obtaining protein measurements from the adsorbed proteins.

[0014] This specification discloses a method comprising, in several embodiments, assaying proteins in a biofluid sample obtained from a subject identified as having a pulmonary nodule to obtain a protein measurement; and applying a classifier to the protein measurement to identify whether the pulmonary nodule is malignant or non-malignant, wherein the classifier features a receiver operating characteristic (ROC) curve having an area under the curve (AUC) greater than 0.7 based on the protein measurement features. In some embodiments, the AUC greater than 0.7 is generated without including non-protein clinical features. In some embodiments, the non-protein clinical features include clinical indicators of lung cancer. In some embodiments, the proteins include APP, IGHG2, SERPING1, SAA2, SERPINF2, GC, IGHA1, HPR, SERPINA3, IGHA1, LTF, SERPINA1, PCSK6, PROS1, BPIF1, C6, CP, A2M, or IGFBP2.

[0015] This specification discloses a method comprising, in several embodiments, assaying proteins in a biofluid sample obtained from a subject having or suspected of having pulmonary nodules to obtain protein measurements; and applying a classifier to the protein measurements to evaluate the pulmonary nodules, wherein the classifier is generated using proteomics data obtained by concentrating proteins with affinity reagents. This specification discloses a method comprising, in several embodiments, assaying proteins in a biofluid sample obtained from a subject having or suspected of having pulmonary nodules to obtain protein measurements; and applying a classifier to the protein measurements to identify the protein measurements as indicating whether the pulmonary nodules are cancerous or non-cancerous, wherein the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and assaying the proteins adsorbed to the particles. Some embodiments include obtaining or receiving a biofluid sample of the subject. In some embodiments, the subject is identified as having pulmonary nodules by medical imaging. In some embodiments, medical imaging includes computed tomography (CT) scans. In some embodiments, medical imaging is performed. In some embodiments, medical imaging is performed to identify pulmonary nodules. In some embodiments, a report is generated based on the identification of protein measurements indicating whether the pulmonary nodules are cancerous or non-cancerous. In some embodiments, the report includes the likelihood or indicator that the pulmonary nodules are cancerous or non-cancerous. In some embodiments, the report is printed or transmitted. In some embodiments, the report is used by healthcare professionals when making a diagnosis about the pulmonary nodules, providing medical advice, or providing treatment. In some embodiments, a biopsy is performed on the pulmonary nodules if the protein measurements are classified as indicating that the pulmonary nodules are cancerous. In some embodiments, the biopsy confirms the likelihood that the pulmonary nodules are cancerous or non-cancerous. In some embodiments, the pulmonary nodules are cancerous.In some embodiments, lung nodules include non-small cell lung cancer (NSCLC). In some embodiments, the classifier includes features that indicate whether a lung nodule is cancerous or non-cancerous, based on protein measurements. In some embodiments, the features include control protein measurements, mass spectra, m / z ratios, chromatography results, immunoassay results, or light or fluorescence intensity. In some embodiments, the classifier is trained using deep learning, hierarchical clustering analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k nearest neighbor analysis, simple Bayesian analysis, K-means clustering analysis, or hidden Markov analysis. In some embodiments, the classifier can distinguish lung cancer with sensitivity of 50%, 60%, 70%, 80%, or 90% or higher. In some embodiments, the classifier can distinguish lung cancer with specificity of 50%, 60%, 70%, 80%, or 90% or higher. In some embodiments, the method includes recommending lung cancer treatment to a subject when protein measurements classify the lung nodule as malignant. In some embodiments, the method includes administering lung cancer treatment to a subject when protein measurements classify the lung nodule as malignant. In some embodiments, lung cancer treatment includes chemotherapy, radiotherapy, percutaneous ablation, radiofrequency ablation, cryoablation, microwave ablation, chemoembolization, or surgery. In some embodiments, the lung nodule is non-malignant. In some embodiments, the method includes observing the subject without performing a biopsy when protein measurements classify the lung nodule as non-malignant. In some embodiments, observing the subject without performing a biopsy includes assaying proteins in a second biofluid sample obtained later from the subject. In some embodiments, the particles include nanoparticles. In some embodiments, the particles include lipid particles, metal particles, silica particles, or polymer particles.In some embodiments, the particles include carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles. In some embodiments, the particles include a group of physiologically different nanoparticles. In some embodiments, assaying a protein includes contacting a biological fluid sample with the particles so that the particles adsorb proteins onto the particles. In some embodiments, assaying a protein includes measuring a reading indicating the presence, absence, or amount of a biomolecule. In some embodiments, assaying a protein includes performing mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blotting, dot blotting, or immunostaining, or a combination thereof. In some embodiments, assaying a protein includes performing mass spectrometry. In some embodiments, the protein includes secretory proteins. In some embodiments, the biological fluid includes blood, plasma, or serum. In some embodiments, the pulmonary nodule is less than 3 cm in diameter. In some embodiments, the subject has multiple pulmonary nodules. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.

[0016] This specification discloses a method, in several embodiments, that includes: obtaining a biofluid sample from a subject having a pulmonary nodule; contacting the biofluid sample with particles such that the particles adsorb protein-containing biomolecules onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data as indicating whether the pulmonary nodule is cancerous or non-cancerous. In some embodiments, the subject is identified as having a pulmonary nodule by medical imaging. In some embodiments, the medical imaging includes computed tomography (CT) scans. In some embodiments, the method includes performing medical imaging. In some embodiments, the method includes identifying the pulmonary nodule by medical imaging. In some embodiments, the method includes performing a biopsy on the pulmonary nodule if the proteomics data is classified as indicating that the pulmonary nodule is cancerous. In some embodiments, the biopsy confirms whether the pulmonary nodule is cancerous or non-cancerous. In some embodiments, the pulmonary nodule is cancerous and includes a tumor. In some embodiments, the pulmonary nodule includes non-small cell lung cancer (NSCLC). In some embodiments, classifying proteomics data as indicating whether a lung nodule is malignant or non-malignant involves applying a classifier to the proteomics data. In some embodiments, the classifier includes features that indicate whether the lung cancer is malignant or non-malignant. In some embodiments, the classifier is trained using deep learning, hierarchical clustering analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k nearest neighbor analysis, simple Bayesian analysis, K-means clustering analysis, or hidden Markov analysis. In some embodiments, the proteomics data indicates whether a lung nodule is malignant or non-malignant with a sensitivity or specificity of approximately 80% or higher. In some embodiments, it involves recommending lung cancer treatment to a subject when the proteomics data classifies the lung nodule as malignant. In some embodiments, it involves administering lung cancer treatment to a subject when the proteomics data classifies the lung nodule as malignant.In some embodiments, lung cancer treatment includes chemotherapy, radiotherapy, percutaneous ablation, radiofrequency ablation, cryoablation, microwave ablation, chemoembolization, or surgery. In some embodiments, the lung nodules are non-cancerous and benign. In some embodiments, observation of the subject without biopsy is included when proteomics data classify the lung nodules as non-cancerous. In some embodiments, monitoring the subject and assaying biomolecules in a second biofluid sample obtained later from the subject. In some embodiments, the particles include nanoparticles. In some embodiments, the particles include lipid particles, metal particles, silica particles, or polymer particles. In some embodiments, the particles include carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles. In some embodiments, the particles include a group of physiologically and chemically distinct nanoparticles. In some embodiments, assaying a biomolecule involves measuring a reading indicating the presence, absence, or quantity of the biomolecule. In some embodiments, assaying a biomolecule involves performing mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blotting, dot blotting, or immunostaining, or a combination thereof. In some embodiments, assaying a biomolecule involves performing mass spectrometry. In some embodiments, proteins include secreted proteins. In some embodiments, biological fluids include blood, plasma, or serum. In some embodiments, pulmonary nodules are less than 3 cm in diameter. In some embodiments, the subject has multiple pulmonary nodules. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.

[0017] This specification discloses a method comprising, in several embodiments, assaying proteins in a biofluid sample obtained from a subject suspected of having pulmonary nodules to obtain a protein measurement; applying a classifier to the protein measurement to identify the protein measurement as indicating that the subject has pulmonary nodules, wherein the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and assaying the proteins adsorbed to the particles. Some embodiments include recommending that the subject undergo medical imaging, such as a CT scan, if the protein measurement indicates that the subject has pulmonary nodules, and not recommending that the subject undergo medical imaging, if the protein measurement does not indicate that the subject has pulmonary nodules. Some embodiments include performing medical imaging, such as a CT scan, on the subject if the protein measurement indicates that the subject has pulmonary nodules, and not performing medical imaging on the subject if the protein measurement does not indicate that the subject has pulmonary nodules. In some embodiments, the protein measurement indicates that the subject has pulmonary nodules, and includes sending or receiving a report relating to medical imaging such as a CT scan, and not sending or receiving a report if the protein measurement does not indicate that the subject has pulmonary nodules. In some embodiments, the protein measurement indicates that the subject has or is likely to have pulmonary nodules. In some embodiments, the protein measurement indicates that the subject does not have or is unlikely to have pulmonary nodules.

[0018] This specification discloses methods that, in several embodiments, include: assaying proteins in a biofluid sample obtained from a subject suspected of having lung cancer to obtain protein measurements; applying a classifier to the protein measurements to identify the protein measurements as indicating that the subject has lung cancer, wherein the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and assaying the proteins adsorbed to the particles. Some embodiments include recommending that the subject undergo medical imaging, such as a CT scan, if the protein measurements indicate that the subject has lung cancer, and not recommending that the subject undergo medical imaging, if the protein measurements do not indicate that the subject has lung cancer. Some embodiments include performing medical imaging, such as a CT scan, on the subject if the protein measurements indicate that the subject has lung cancer, and not performing medical imaging, if the protein measurements do not indicate that the subject has lung cancer. Some embodiments include sending or receiving a report regarding medical imaging, such as a CT scan, if the protein measurements indicate that the subject has lung cancer, and not sending or receiving a report, if the protein measurements do not indicate that the subject has lung cancer. In some embodiments, protein measurements indicate that a subject has or is likely to have lung cancer. In some embodiments, protein measurements indicate that a subject does not have or is unlikely to have lung cancer. In some embodiments, lung cancer includes NSCLC.

[0019] This specification discloses a method comprising, in several embodiments, obtaining a biofluid sample from a subject suspected of having pulmonary nodules; contacting the biofluid sample with particles such that the particles adsorb protein-containing biomolecules onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying, based on the proteomics data, whether the proteomics data indicates that the subject has pulmonary nodules or does not. Some embodiments include recommending that the subject undergo medical imaging, such as a CT scan, if the proteomics data indicates that the subject has pulmonary nodules, and not recommending that the subject undergo medical imaging, if the proteomics data does not indicate that the subject has pulmonary nodules. Some embodiments include performing medical imaging, such as a CT scan, on the subject if the proteomics data indicates that the subject has pulmonary nodules, and not performing medical imaging on the subject if the proteomics data does not indicate that the subject has pulmonary nodules. Some embodiments include sending or receiving reports relating to medical imaging, such as CT scans, if proteomics data indicates that the subject has pulmonary nodules, and not sending or receiving reports if proteomics data does not indicate that the subject has pulmonary nodules. In some embodiments, the proteomics data indicates that the subject has or is likely to have pulmonary nodules. In some embodiments, the proteomics data indicates that the subject does not have or is unlikely to have pulmonary nodules.

[0020] This specification discloses a method comprising, in several embodiments, obtaining a biofluid sample from a subject suspected of having lung cancer; contacting the biofluid sample with particles such that the particles adsorb protein-containing biomolecules onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying, based on the proteomics data, whether the proteomics data indicates that the subject has lung cancer or does not. Some embodiments include recommending that the subject undergo medical imaging, such as a CT scan, if the proteomics data indicates that the subject has lung cancer, and not recommending that the subject undergo medical imaging, if the proteomics data does not indicate that the subject has lung cancer. Some embodiments include performing medical imaging, such as a CT scan, on the subject if the proteomics data indicates that the subject has lung cancer, and not performing medical imaging on the subject if the proteomics data does not indicate that the subject has lung cancer. Some embodiments include sending or receiving reports relating to medical imaging, such as CT scans, if proteomics data indicates that the subject has lung cancer, and not sending or receiving reports if proteomics data does not indicate that the subject has lung cancer. In some embodiments, the proteomics data indicates that the subject has or is likely to have lung cancer. In some embodiments, the proteomics data indicates that the subject does not have or is unlikely to have lung cancer.

[0021] This specification discloses a monitoring method comprising: obtaining a biofluid sample from a subject at risk of lung cancer recurrence; contacting the biofluid sample with particles such that the particles adsorb protein-containing biomolecules onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data, based on the proteomics data, as indicating that the subject has lung cancer recurrence or as not indicating that the subject has lung cancer recurrence. Some embodiments include recommending that the subject undergo medical imaging, such as a CT scan, if the protein measurement indicates that the subject has lung cancer recurrence, and not recommending that the subject undergo medical imaging, if the protein measurement does not indicate that the subject has lung cancer recurrence. Some embodiments include performing medical imaging, such as a CT scan, on the subject if the protein measurement indicates that the subject has lung cancer recurrence, and not performing medical imaging on the subject if the protein measurement does not indicate that the subject has lung cancer recurrence. In some embodiments, the protein measurement indicates that the subject has lung cancer recurrence, and the protein measurement indicates that the subject has medical imaging, such as a CT scan. In some embodiments, the protein measurement indicates that the subject has or is likely to have lung cancer recurrence. In some embodiments, the protein measurement indicates that the subject does not have or is unlikely to have lung cancer recurrence. In some embodiments, the subject is receiving treatment for lung cancer. In some embodiments, the lung cancer treatment includes chemotherapy, radiotherapy, or surgery. In some embodiments, the cancer is potentially resectable. In some embodiments, the lung cancer includes NSCLC. In embodiments of the present invention, for example, the following items are provided. (Item 1) A multi-omics disease detection method, Acquiring multi-omics data generated from one or more biological fluid samples collected from a subject, wherein the multi-omics data comprises first omics data and second omics data, wherein the first omics data comprises a first omics data type including proteomics data, metabolomics data, transcriptomics data, or genomics data, and the second omics data comprises a second omics data type different from the first omics data type, and includes proteomics data, metabolomics data, transcriptomics data, or genomics data; Assigning a first label corresponding to the presence, absence, or possibility of the disease state to the first omics data using a first classifier; Using a second classifier, assign a second label to the second omics data corresponding to the presence, absence, or possibility of the disease state; and The method comprising identifying the multi-omics data as indicating or not indicating the disease state based on a combination of the first label and the second label, wherein the first classifier and the second classifier are independent, and the combination of the first label and the second label identifies the multi-omics data as indicating or not indicating the disease state with higher accuracy than the first label or the second label alone. (Item 2) The method according to item 1, wherein the first omics data type includes proteomics data, and the second omics data type includes metabolomics data, transcriptomics data, or genomics data. (Item 3) The method according to item 2, wherein the proteomics data includes measurements of at least 1,000 proteins or peptides. (Item 4) The method according to item 2, wherein the proteomics data is generated by contacting one of the one or more biological fluid samples with the particles so that the particles adsorb biomolecules containing proteins. (Item 5) The method according to item 4, wherein the particles include metal, polymer, or lipid. (Item 6) The method according to item 4, wherein the particles include a group of physiologically and chemically distinct nanoparticles. (Item 7) The method according to item 2, wherein the proteomics data is generated using mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, Western blotting, dot blotting, or immunostaining, or a combination thereof. (Item 8) The genomics data or transcriptomics data may be subjected to sequencing, microarray analysis, hybridization, polymerase chain reaction, electrophoresis, or other methods. The method described in item 1, which is produced by these combinations. (Item 9) The method according to item 1, wherein the second omics data type includes transcriptomics data. (Item 10) The method according to item 9, wherein the transcriptomics data includes mRNA or microRNA expression data. (Item 11) The method described in item 1, wherein the second omics data type includes genomics data. (Item 12) The method according to item 11, wherein the genomics data includes DNA sequence data or epigenetic data. (Item 13) The method according to item 12, wherein the epigenetic data includes DNA methylation data, DNA hydroxymethylation data, or histone modification data. (Item 14) The method according to item 1, wherein identifying the multi-omics data as indicating or not indicating the disease state includes generating or obtaining a majority vote score based on the first and second labels. (Item 15) The method according to item 1, wherein identifying the multi-omics data as indicating or not indicating the disease state includes generating or obtaining a weighted average of the first and second labels. (Item 16) The method of item 15, further comprising assigning weights to the first and second classifiers and thereby obtaining the weighted average. (Item 17) The method according to item 16, wherein the weights are assigned based on the area under the ROC curve, the area under the precision-recall curve, precision, precision, recall, sensitivity, F1 score, specificity, or a combination thereof. (Item 18) The method according to item 1, wherein the first and second classifiers independently show errors with respect to the pathological condition. (Item 19) The method according to item 1, further comprising sending or outputting a report containing information regarding the aforementioned identification. (Item 20) The method according to item 1, further comprising transmitting or outputting a recommendation for the treatment of the subject based on the aforementioned pathological condition. (Item 21) The method according to item 1, wherein the multi-omics data further comprises third omics data comprising a third omics data type, the method further comprises assigning a third label to the third omics data corresponding to the presence, absence, or possibility of the pathological condition using a third classifier; and the method comprising identifying the multi-omics data as indicating or not indicating the pathological condition based on a combination of the first, second, and third labels. (Item 22) The first omics data type includes proteomics data, the second omics data type includes mRNA transcriptomics data, and the third omics The method described in item 21, wherein the data type includes microRNA transcriptomics data. (Item 23) A multi-omics disease detection method, Acquiring multi-omics data generated from one or more biological fluid samples collected from a subject, wherein the multi-omics data comprises first omics data and second omics data, wherein the first omics data comprises a first omics data type including proteomics data, metabolomics data, transcriptomics data, or genomics data, and the second omics data comprises a second omics data type different from the first omics data type, and includes proteomics data, metabolomics data, transcriptomics data, or genomics data; Identifying a subset of the first feature from the aforementioned first omics data; Identifying a subset of the second feature from the second omics data; Pooling subsets of the first and second features; The method, comprising identifying the multi-omics data as indicating or not indicating the disease state based on the pooled subset of features. (Item 24) The method according to item 23, wherein identifying a subset of the first or second features from the first or second omics data includes obtaining univariate data for the features of the first or second omics data, and identifying the first or second subset based on the univariate data. (Item 25) The method according to item 23, wherein a subset of the first or second features is identified from the features of a classifier for the first or second omics data. (Item 26) The method according to item 23, wherein identifying a subset of the first or second features from the first or second omics data includes obtaining a classifier for the first or second omics data and identifying the first or second subset as the top features of the classifier. The method according to item 23, wherein identifying a subset of first or second features from the first or second omics data includes obtaining a classifier for the first or second omics data, removing one or more features at a time from the classifier, and identifying which features, when removed from the classifier, degrade the performance of the classifier. (Item 27) Assaying proteins in biological fluid samples obtained from subjects identified as having pulmonary nodules to obtain protein measurement values; and A method comprising applying a classifier to the protein measurement to identify the protein measurement as indicating that the pulmonary nodule is cancerous or non-cancerous, wherein the classifier is characterized by a receiver operating characteristic (ROC) curve having an area under the curve (AUC) greater than 0.7 based on the protein measurement features. (Item 28) The method of item 27, wherein the AUC greater than 0.7 is generated without including non-protein clinical features. (Item 29) The method according to item 28, wherein the aforementioned non-protein clinical features include clinical indicators for lung cancer. (Item 30) The method according to item 27, wherein the protein measurement value includes one or more of APP, IGHG2, SERPING1, SAA2, SERPINF2, GC, IGHA1, HPR, SERPINA3, IGHA1, LTF, SERPINA1, PCSK6, PROS1, BPIF1, C6, CP, A2M, or IGFBP2. (Item 31) Assaying proteins in biological fluid samples obtained from subjects with or suspected of having pulmonary nodules to obtain protein measurements: This includes applying a classifier to the protein measurement value and thereby identifying the protein measurement value as indicating whether the pulmonary nodule is cancerous or non-cancerous, The classifier is generated using proteomics data obtained by contacting the training sample with the particles so that the particles adsorb proteins in the training sample, and by assaying the proteins adsorbed on the particles. (Item 32) The method according to item 31, further comprising obtaining or receiving a biological fluid sample of the subject. (Item 33) The method according to item 31, wherein the subject is identified as having the pulmonary nodule by medical imaging. (Item 34) The method described in item 33, wherein the medical imaging includes computed tomography (CT) scans. (Item 35) The method according to item 33, further comprising performing the aforementioned medical imaging. (Item 36) The method according to item 33, further comprising identifying the pulmonary nodule in the medical imaging. (Item 37) The method according to item 31, comprising generating a report based on the identification of the protein measurement values ​​indicating whether the pulmonary nodule is cancerous or non-cancerous. (Item 38) The report relates to the method described in item 37, including the possibility or indicators that the pulmonary nodule is cancerous or non-cancerous. (Item 39) The method of item 37, further comprising outputting or sending the aforementioned report. (Item 40) The method described in item 37, which the report described above may be used by a healthcare professional to make a diagnosis, provide medical advice, or provide treatment for the pulmonary nodule. (Item 41) The method according to item 31, wherein the classifier includes a feature that indicates whether the pulmonary nodule is cancerous or non-cancerous, and the classifier includes a feature that indicates the protein measurement value. (Item 42) The method according to item 41, wherein the aforementioned feature includes a control protein measurement, mass spectrum, m / z ratio, chromatography results, immunoassay results, or light or fluorescence intensity. (Item 43) The method according to item 31, wherein the classifier is trained using deep learning, hierarchical clustering analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k nearest neighbor analysis, simple Bayesian analysis, K-means clustering analysis, or hidden Markov analysis. (Item 44) The method according to item 31, wherein the classifier can identify lung cancer with a sensitivity of 50% or more, 60% or more, 70% or more, 80% or more, or 90% or more. (Item 45) The method according to item 31, wherein the classifier can identify lung cancer with a specificity of 50% or more, 60% or more, 70% or more, 80% or more, or 90% or more. (Item 46) The method according to item 31, further comprising recommending lung cancer treatment to the subject if the protein measurement is classified as indicating that the lung nodule is cancerous. (Item 47) The method according to item 31, further comprising administering lung cancer treatment to the subject if the protein measurement is classified as indicating that the lung nodule is cancerous. (Item 48) The method according to item 46, wherein the lung cancer treatment includes chemotherapy, radiotherapy, percutaneous ablation, radiofrequency ablation, cryoablation, microwave ablation, chemoembolization, or surgery. (Item 49) The method according to item 31, wherein the particles include nanoparticles. (Item 50) The method according to item 31, wherein the particles include lipid particles, metal particles, silica particles, or polymer particles. (Item 51) The method according to item 31, wherein the particles include carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles. (Item 52) The method according to item 31, wherein the particles include a group of physiologically and chemically distinct nanoparticles. (Item 53) The method according to item 31, wherein assaying the protein involves measuring a reading indicating the presence, absence, or amount of the biomolecule. (Item 54) The method according to item 31, wherein assaying the protein includes contacting the biological fluid sample with the particles such that the particles adsorb the protein onto the particles. (Item 55) The method according to item 31, wherein assaying the protein comprises performing mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunoadsorption assay, Western blotting, dot blotting, or immunostaining, or a combination thereof. (Item 56) The method according to item 31, wherein assaying the protein includes performing mass spectrometry. (Item 57) The method according to item 31, further comprising performing a biopsy on the lung nodule if the protein measurement results classify the lung nodule as being cancerous. (Item 58) The method according to item 57, wherein the biopsy confirms whether the pulmonary nodule is cancerous or non-cancerous. (Item 59) The method of item 57, further comprising observing the subject without performing a biopsy if the protein measurement is classified as indicating that the pulmonary nodule is non-cancerous. (Item 60) The method of item 59, comprising observing the subject without performing a biopsy, and subsequently assaying proteins in a second biological fluid sample obtained from the subject. (Item 61) The method according to item 31, wherein the pulmonary nodule is cancerous. (Item 62) The method according to item 31, wherein the pulmonary nodule contains non-small cell lung cancer (NSCLC). (Item 63) The method according to item 31, wherein the pulmonary nodule is non-cancerous. (Item 64) The method according to item 31, further comprising assaying proteins in a second biological fluid sample obtained later from the subject. (Item 65) Obtain a biological fluid sample from a subject with pulmonary nodules. Bringing the biological fluid sample into contact with the particles so that the particles adsorb biomolecules containing proteins onto the particles; To perform an assay on the biomolecules adsorbed onto the particles to generate proteomics data; and A method comprising classifying the proteomics data as indicating whether a pulmonary nodule is cancerous or non-cancerous. (Item 66) The method according to item 65, wherein the subject is identified as having the pulmonary nodule by medical imaging. (Item 67) The method according to item 66, wherein the medical imaging includes computed tomography (CT) scans. (Item 68) The method according to item 66, further comprising performing the aforementioned medical imaging. (Item 69) The method according to item 65, further comprising identifying the pulmonary nodule in the medical imaging. (Item 70) The method according to item 65, wherein classifying the proteomics data as indicating whether the pulmonary nodules are cancerous or non-cancerous includes applying a classifier to the proteomics data. (Item 71) The method according to item 70, wherein the classifier includes features that indicate whether the lung cancer is malignant or non-malignant. (Item 72) The method according to item 70, wherein the classifier is trained using deep learning, hierarchical clustering analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k nearest neighbor analysis, simple Bayesian analysis, K-means clustering analysis, or hidden Markov analysis. (Item 73) The method according to item 65, wherein the proteomics data indicates with a sensitivity or specificity of approximately 80% or more that the pulmonary nodule is cancerous or non-cancerous. (Item 74) The method according to item 65, further comprising recommending lung cancer treatment to the subject if the proteomics data classifies the lung nodule as being cancerous. (Item 75) The method according to item 65, further comprising administering lung cancer treatment to the subject if the proteomics data classifies the lung nodule as being cancerous. (Item 76) The method according to item 74, wherein the lung cancer treatment includes chemotherapy, radiotherapy, percutaneous ablation, radiofrequency ablation, cryoablation, microwave ablation, chemoembolization, or surgery. (Item 77) The method according to item 65, wherein the particles include nanoparticles. (Item 78) The method according to item 65, wherein the particles include lipid particles, metal particles, silica particles, or polymer particles. (Item 79) The method according to item 65, wherein the particles include carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles. (Item 80) The method according to item 65, wherein the particles comprise a group of physiologically and chemically distinct nanoparticles. (Item 81) The method according to item 65, wherein assaying the biomolecule comprises performing mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunoadsorption assay, Western blotting, dot blotting, or immunostaining, or a combination thereof. (Item 82) The method of item 65, wherein assaying the biomolecule includes performing mass spectrometry. (Item 83) The method according to item 65, wherein assaying the biomolecule includes measuring a reading indicating the presence, absence, or quantity of the biomolecule. (Item 84) The method of item 65, further comprising performing a biopsy on the lung nodule if the proteomics data classifies the lung nodule as being cancerous. (Item 85) The method according to item 84, wherein the biopsy confirms whether the pulmonary nodule is cancerous or non-cancerous. (Item 86) The method according to item 65, wherein the pulmonary nodule is cancerous and includes a tumor. (Item 87) The method according to item 65, wherein the pulmonary nodule includes non-small cell lung cancer (NSCLC). (Item 88) The method according to item 65, wherein the pulmonary nodule is non-cancerous and benign. (Item 89) The method of item 65, further comprising observing the subject without performing a biopsy when the proteomics data classifies the pulmonary nodule as being non-cancerous. (Item 90) The method of item 65, further comprising monitoring the subject and assaying biomolecules in a second biological fluid sample subsequently obtained from the subject. (Item 91) The method according to item 65, wherein the protein includes a secreted protein. (Item 92) The method according to item 65, wherein the biological fluid includes blood, plasma, or serum. (Item 93) The method according to item 65, wherein the pulmonary nodule is less than 3 cm in diameter. (Item 94) The method according to item 65, wherein the subject has multiple pulmonary nodules. (Item 95) The method described in item 65, wherein the subject is a mammal. (Item 96) The method described in item 65, wherein the subject is a human. (Item 97) Assaying proteins in biological fluid samples obtained from subjects suspected of having pulmonary nodules to obtain protein measurement values; and This includes applying a classifier to the protein measurement value and thereby identifying the protein measurement value as indicating that the subject has the pulmonary nodule, The classifier is generated using proteomics data obtained by contacting the training sample with the particles so that the particles adsorb proteins in the training sample, and by assaying the proteins adsorbed on the particles. (Item 98) The method according to item 97, further comprising recommending that the subject undergo medical imaging such as a CT scan if the protein measurement indicates that the subject has the pulmonary nodule, and not recommending that the subject undergo medical imaging if the protein measurement does not indicate that the subject has the pulmonary nodule. (Item 99) The method according to item 97, further comprising performing medical imaging, such as a CT scan, on the subject if the protein measurement indicates that the subject has the pulmonary nodule, and not performing the medical imaging on the subject if the protein measurement does not indicate that the subject has the pulmonary nodule. (Item 100) The method of item 97, further comprising sending or receiving a report relating to medical imaging, such as a CT scan, if the protein measurement indicates that the subject has the pulmonary nodule, and not sending or receiving the report if the protein measurement does not indicate that the subject has the pulmonary nodule. (Item 101) The method according to item 97, wherein the protein measurement indicates that the subject has or is likely to have the pulmonary nodule. (Item 102) The method according to item 97, wherein the protein measurement indicates that the subject does not have or is unlikely to have the pulmonary nodules. (Item 103) Assaying proteins in biological fluid samples obtained from subjects suspected of having lung cancer to obtain protein measurement values; and This includes applying a classifier to the protein measurement value and thereby identifying the protein measurement value as indicating that the subject has the pulmonary nodule, The classifier is generated using proteomics data obtained by contacting the training sample with the particles so that the particles adsorb proteins in the training sample, and by assaying the proteins adsorbed on the particles. (Item 104) The method according to item 103, further comprising recommending that the subject undergo medical imaging such as a CT scan if the protein measurement indicates that the subject has lung cancer, and not recommending that the subject undergo medical imaging if the protein measurement does not indicate that the subject has lung cancer. (Item 105) The method according to item 103, further comprising performing medical imaging, such as a CT scan, on the subject if the protein measurement indicates that the subject has lung cancer, and not performing medical imaging on the subject if the protein measurement does not indicate that the subject has lung cancer. (Item 106) The method according to item 103, further comprising sending or receiving a report relating to medical imaging, such as a CT scan, if the protein measurement indicates that the subject has the lung cancer, and not sending or receiving the report if the protein measurement does not indicate that the subject has the lung cancer. (Item 107) The method according to item 103, wherein the protein measurement indicates that the subject has or is likely to have the lung cancer. (Item 108) The method according to item 103, wherein the protein measurement indicates that the subject does not have or is unlikely to have the lung cancer. (Item 109) Obtain a biopsy sample from a subject suspected of having pulmonary nodules; Bringing the biological fluid sample into contact with the particles so that the particles adsorb biomolecules containing proteins onto the particles; To perform an assay on the biomolecules adsorbed onto the particles to generate proteomics data; and A method comprising classifying, based on the proteomics data, the proteomics data as indicating that the subject has the pulmonary nodules or as not indicating that the subject has the pulmonary nodules. (Item 110) The method according to item 109, further comprising recommending that the subject undergo medical imaging such as a CT scan if the proteomics data indicates that the subject has the pulmonary nodules, and not recommending that the subject undergo medical imaging if the proteomics data does not indicate that the subject has the pulmonary nodules. (Item 111) The method according to item 109, further comprising performing medical imaging, such as a CT scan, on the subject if the proteomics data indicates that the subject has the pulmonary nodules, and not performing medical imaging on the subject if the proteomics data does not indicate that the subject has the pulmonary nodules. (Item 112) The method according to item 109, further comprising sending or receiving a medical imaging report, such as a CT scan, if the proteomics data indicates that the subject has the pulmonary nodule, and not sending or receiving the report if the proteomics data does not indicate that the subject has the pulmonary nodule. (Item 113) The method according to item 109, wherein the proteomics data indicates that the subject has or is likely to have the pulmonary nodules. (Item 114) The proteomics data indicates that the subject does not have or may have the pulmonary nodules. The method described in item 109, which indicates that the value is low. (Item 115) Obtain a biopsy sample from a person suspected of having lung cancer; Bringing the biological fluid sample into contact with the particles so that the particles adsorb biomolecules containing proteins onto the particles; To perform an assay on the biomolecules adsorbed onto the particles to generate proteomics data; and A method comprising classifying the proteomics data, based on the proteomics data, as indicating that the subject has the lung cancer or as not indicating that the subject has the lung cancer. (Item 116) The method according to item 115, further comprising recommending that the subject undergo medical imaging such as a CT scan if the proteomics data indicates that the subject has lung cancer, and not recommending that the subject undergo medical imaging if the proteomics data does not indicate that the subject has lung cancer. (Item 117) The method according to item 115, further comprising performing medical imaging, such as a CT scan, on the subject if the proteomics data indicates that the subject has lung cancer, and not performing medical imaging on the subject if the proteomics data does not indicate that the subject has lung cancer. (Item 118) The method according to item 115, further comprising sending or receiving a medical imaging report, such as a CT scan, if the proteomics data indicates that the subject has the lung cancer, and not sending or receiving the report if the proteomics data does not indicate that the subject has the lung cancer. (Item 119) The method according to item 115, wherein the proteomics data indicates that the subject has or is likely to have the lung cancer. (Item 120) The method according to item 115, wherein the proteomics data indicates that the subject does not have or is unlikely to have the lung cancer. (Item 121) A monitoring method, Obtaining bio-fluid samples from individuals at risk of lung cancer recurrence; Bringing the biological fluid sample into contact with the particles so that the particles adsorb biomolecules containing proteins onto the particles; To perform an assay on the biomolecules adsorbed onto the particles to generate proteomics data; and The method comprising classifying the proteomics data, based on the proteomics data, as indicating that the subject has lung cancer recurrence or as not indicating that the subject has lung cancer recurrence. (Item 122) The method according to item 121, further comprising recommending that the subject undergo medical imaging such as a CT scan if the protein measurement indicates that the subject has lung cancer recurrence, and not recommending that the subject undergo medical imaging if the protein measurement does not indicate that the subject has lung cancer recurrence. (Item 123) The method according to item 121, further comprising performing medical imaging, such as a CT scan, on the subject if the protein measurement indicates that the subject has lung cancer recurrence, and not performing medical imaging on the subject if the protein measurement does not indicate that the subject has lung cancer recurrence. (Item 124) The method according to item 121, further comprising sending or receiving a report relating to medical imaging, such as a CT scan, if the protein measurement indicates that the subject has the lung cancer recurrence, and not sending or receiving the report if the protein measurement does not indicate that the subject has the lung cancer recurrence. (Item 125) The method according to item 121, wherein the protein measurement indicates that the subject has or is likely to have the lung cancer recurrence. (Item 126) The method according to item 121, wherein the protein measurement indicates that the subject does not have or is unlikely to have lung cancer recurrence. (Item 127) The method described in item 121, wherein the subject is receiving treatment for lung cancer. (Item 128) The method according to item 127, wherein the lung cancer treatment includes chemotherapy, radiotherapy, or surgery. (Item 129) The method described in item 121, wherein the cancer is potentially resectable. (Item 130) The method according to item 121, wherein the cancer includes NSCLC. [Brief explanation of the drawing]

[0022] [Figure 1A] This presents a multi-omics approach.

[0023] [Figure 1B] This shows the combination of datasets used in the multi-omics approach.

[0024] [Figure 2A] Examples of methods for generating and applying the classifiers described herein are shown.

[0025] [Figure 2B] This flowchart shows several embodiments that can be used in the method described herein.

[0026] [Figure 3A] This section provides examples of stages in screening and treatment for patients with or suspected of having a medical condition.

[0027] [Figure 3B] This shows an example of the stages in screening and treatment for pancreatic cancer patients.

[0028] [Figure 3C] This shows an example of the stages in screening and treatment for liver cancer patients.

[0029] [Figure 4] This provides a non-specific example of a computing device (in this case, a device with one or more processors, memory, storage, and network interfaces).

[0030] [Figure 5] This specification shows diagrams of classifiers and feature information according to several embodiments described herein.

[0031] [Figure 6] The graphs illustrate the differences in expression of several proteins that can be used to generate classifiers for diagnosing disease conditions.

[0032] [Figure 7] The figure shows the expression of several proteins in diseased samples compared to control samples. Several genes were expressed differentially (underexpressed or overexpressed) between the groups (NSCLC samples and healthy samples).

[0033] [Figure 8] The following scatter plots show pairs of predictions plotted against each other. RNASeq: Predicted probability based on RNA-Seq data (affected). Proteomics: Predicted probability based on proteomics data (affected). RNA_Prot: Predicted probability based on both RNA-Seq and proteomics data (affected).

[0034] [Figure 9] The data includes receiver operating characteristic (ROC) curves and shows the increase in area under the curve (AUC) when mRNA transcriptomics data and proteomics data are combined, compared to mRNA transcriptomics data or proteomics data alone.

[0035] [Figure 10A]This shows an additive multi-omics classification of 30 samples from subjects with the disease and 30 samples from control subjects, including mRNA transcriptomics data, proteomics data, and combinations of mRNA transcriptomics and proteomics data.

[0036] [Figure 10B] This shows the differential mRNA and protein levels measured in biological fluid samples and used to generate classifiers.

[0037] [Figure 11A] This shows analyses based on proteomics data and microRNA data. The top panel shows the results of a classifier trained on proteomics data only, the middle panel shows the results of a classifier trained on microRNA data only, and the bottom panel shows the results of a combination of both data types.

[0038] [Figure 11B] This shows the differentially expressed microRNAs used to generate the classifier.

[0039] [Figure 12] This analysis compares combinations of three omics data types (proteomics, mRNA, and miRNA) to using only one of each data type.

[0040] [Figure 13A] Several embodiments that can be used for integrated model classification are shown.

[0041] [Figure 13B] Several embodiments that can be used in transformation-based classification are shown.

[0042] [Figure 14] The graph shows the results of the integrated model classification analysis.

[0043] [Figure 15] This section presents several aspects of transformation-based classification analysis.

[0044] [Figure 16] The graphs show the results of integrated model classification analysis and transformation-based classification.

[0045] [Figure 17] This document presents a non-limiting example of a flowchart for a machine training algorithm to improve the sensitivity and specificity of a classifier for predicting the diseases described herein.

[0046] [Figure 18A] ROC curves for several protein data and protein + lipid combination data for disease classification are shown.

[0047] [Figure 18B] This includes aspects of the sensitivity of analyzing protein data, lipid data, and combined protein-lipid data for disease classification.

[0048] [Figure 19] This document describes a two-stage machine learning framework for analyzing and training on multiple data types.

[0049] [Figure 20A] This includes aspects of the sensitivity of analyzing protein data, lipid data, and combined protein-lipid data for disease classification.

[0050] [Figure 20B] This includes aspects of the sensitivity of analyzing protein data, lipid data, and combined protein-lipid data for disease classification.

[0051] [Figure 20C] ROC curves for several protein data, lipid data, and protein + lipid combination data for disease classification are shown.

[0052] [Figure 21] ROC curves are shown for several protein data for disease classification, as well as for combined data of protein, lipid, and clinical parameters.

[0053] [Figure 22A] Here is some information regarding protein data.

[0054] [Figure 22B] The performance characteristics of several classifiers are shown.

[0055] [Figure 22C] This section shows the performance characteristics of several classifiers, both with and without certain features.

[0056] [Figure 23] This section describes several aspects of genetic or transcriptional data, such as the indicator or type of measurement, the type of sample, the quality control method, or the sequencing depth that can be used.

[0057] [Figure 24] Various embodiments that can be used in some of the methods described herein are shown.

[0058] [Figure 25] This specification includes several embodiments, such as subjects or test results, that may be included in the method described herein.

[0059] [Figure 26A] Includes a table describing several proteins, OT scores, and some features in the protein classifier.

[0060] [Figure 26B] Includes a table describing several proteins, OT scores, and some features in the protein classifier.

[0061] [Figure 27] Includes a chart showing the feature importance scores of the lipid classifier.

[0062] [Figure 28A] The results of the Wilcox test for age comparison and Fisher's exact test for gender ratio are shown.

[0063] [Figure 28B] The results of the Wilcox test for age comparison and Fisher's exact test for gender ratio are shown.

[0064] [Figure 29A] This shows the number of proteins detected throughout the entire sample in the analysis of biological fluid samples from control patients and cancer patients.

[0065] [Figure 29B] This shows the number of proteins detected throughout the entire sample in the analysis of biological fluid samples from control patients and cancer patients.

[0066] [Figure 30A] This plot shows several top proteins differentially detected in biofluid samples from cancer patients compared to biofluid samples from control patients.

[0067] [Figure 30B] This plot shows the distribution of OpenTargets (OT) scores. The OT scores (0-0.8) are on the x-axis, and the density (0-15) is on the y-axis.

[0068] [Figure 31A] Includes plots showing comparisons of total median signals by sample, analyte type, and class.

[0069] [Figure 31B]The box plots show the most significantly different analytes for each omics workflow according to one embodiment; top left: lipids; bottom left: metabolites; and right: proteins).

[0070] [Figure 31C] This example demonstrates the performance of a multi-omics classifier that combines proteomics, lipidomics, and metabolomics measurements.

[0071] [Figure 32A] This includes a volcano plot showing the intensity difference and P-value of proteins adsorbed onto nanoparticles and detected in biofluid samples from cancer patients compared to biofluid samples from control patients. In the volcano plot, the x-axis shows the magnitude of the difference, and the y-axis shows the significance, with the most significant analytes highlighted.

[0072] [Figure 32B] Includes data for the top protein P35442 after particle-based measurement methods.

[0073] [Figure 32C] This includes a volcano plot showing the intensity difference and p-value of proteins detected in biofluid samples from cancer patients compared to biofluid samples from control patients. In the volcano plot, the x-axis shows the magnitude of the difference, and the y-axis shows the significance level, with the most significant analytes highlighted.

[0074] [Figure 32D] Includes data for the top protein P01011 after proteomics analysis.

[0075] [Figure 33A] This includes a volcano plot showing the intensity difference and P-value of lipids detected in biofluid samples from cancer patients compared to biofluid samples from control patients. In the volcano plot, the x-axis shows the magnitude of the difference, and the y-axis shows the significance level, with the most significant analytes highlighted.

[0076] [Figure 33B] Includes data on the top lipid CER (d18:1_18:0) after lipidomics measurement.

[0077] [Figure 34A] This includes a volcano plot showing the intensity differences and p-values ​​of metabolites detected in biofluid samples from cancer patients compared to biofluid samples from control patients. In the volcano plot, the x-axis shows the magnitude of the difference, and the y-axis shows the significance level, with the most significant analytes highlighted.

[0078] [Figure 34B] Includes data on AICAR, the top metabolite after metabolomics measurements.

[0079] [Figure 35A] This shows the classification of cancer and healthy samples using UMAP projection based on combined data.

[0080] [Figure 35B] This shows the classification of cancer and healthy samples by PCA projection based on combined data.

[0081] [Figure 35C] This shows the classification of cancer and healthy samples using UMAP projection based on proteograph data.

[0082] [Figure 35D] This shows the classification of cancer and healthy samples by PCA projection based on proteograph data.

[0083] [Figure 35E] This shows the classification of cancer and healthy samples using UMAP projection based on PiQuant data.

[0084] [Figure 35F] This shows the classification of cancer and healthy samples by PCA projection based on PiQuant data.

[0085] [Figure 35G] This shows the classification of cancer and healthy samples using UMAP projection based on lipid data.

[0086] [Figure 35H] This shows the classification of cancer and healthy samples based on PCA projection, using lipid data.

[0087] [Figure 35I] This shows the classification of cancer and healthy samples using UMAP projection based on metabolite data.

[0088] [Figure 35J] This shows the classification of cancer and healthy samples by PCA projection based on metabolite data.

[0089] [Figure 36] This shows the characteristic features of proteins, lipids, and metabolites included in the classifier.

[0090] [Figure 37] This figure illustrates the performance of classifiers in multi-omics studies, including receiver operating characteristic (ROC) curves for disease classification. Area under the curve (AUC) values ​​are also included in the figure, with 90% confidence intervals shown in parentheses.

[0091] [Figure 38A] This figure shows the performance of a classifier trained using data from genomic assays, including ROC curves for disease classification. The AUC values ​​at the bottom of the figure are shown as ± values ​​based on 90% confidence.

[0092] [Figure 38B]The performance of classifiers trained on data from genomics assays ("Genomics"), classifiers trained on data from mass spectrometry assays ("Mass-spec"), and classifiers trained on data from both genomics and mass spectrometry assays ("Combined") is shown. The data shown in the figure includes ROC curves for disease classification. AUC values ​​include ± values ​​based on 90% confidence.

[0093] [Figure 39A] A graphical overview of the 18 samples from liver cancer subjects used in Example 17 is shown.

[0094] [Figure 39B] The coefficient of variation (CV) values ​​for several peptides and proteins obtained in the studies described herein are shown.

[0095] [Figure 39C] Exemplary protein abundance heatmaps are shown for samples from subjects with liver cancer and healthy subjects.

[0096] [Figure 39D] Examples of differences in protein abundance identified in samples from subjects with liver cancer or healthy subjects after contacting the samples with the various particles described herein are shown.

[0097] [Figure 39E] Includes graphs demonstrating the high reproducibility of lipidomics data obtained from samples.

[0098] [Figure 39F] This shows that samples from subjects with liver cancer exhibited different lipid profiles and those of healthy controls. The top 50 lipids based on the p-values ​​of this analysis are shown for each patient sample.

[0099] [Figure 39G] This shows the univariate lipid differences between samples from subjects with liver cancer and healthy subjects.

[0100] [Figure 40A] A graph summarizing the nine samples from ovarian cancer subjects used in Example 19 is shown.

[0101] [Figure 40B] Exemplary protein abundance heatmaps are shown for samples from subjects with ovarian cancer and healthy subjects.

[0102] [Figure 40C] This shows the univariate lipid differences between samples from subjects with ovarian cancer and healthy subjects.

[0103] [Figure 41] This shows an example of the stages in screening and treatment for colorectal cancer patients.

[0104] [Figure 42] This shows the age and sex breakdown of the 268 participants in the NSCLC biomarker discovery study.

[0105] [Figure 43] This shows the protein count for each study group, including healthy individuals, those with comorbidities, and those with NSCLC stage 1 ("NSCLC_1"), NSCLC stage 2 ("NSCLC_2"), NSCLC stage 3 ("NSCLC_3"), and NSCLC stage 4 ("NSCLC_4").

[0106] [Figure 44] This shows the protein counts of depleted plasma DP and particle panels.

[0107] [Figure 45] This provides an overview of the fractional detection of proteins across the entire sample versus the average abundance of the proteins for 10 particle types in particle panels and depleted plasma (DP).

[0108] [Figure 46]It shows the performance of the cross-validated particle panel classifier, where the x-axis represents the proportion of false positive classifications and the y-axis represents the proportion of true positive classifications.

[0109] [Figure 47] Graphs of the random forest model for healthy vs NSCLC (stages 1, 2, and 3) for depleted plasma (left) and 10-particle panel (right), with the false positive rate shown on the x-axis and the true positive rate shown on the y-axis.

[0110] [Figure 48] It shows the performance of the classifier features across the study samples.

[0111] [Figure 49] It shows the results of repeating 10 rounds of 10-fold cross-validation with randomized class assignments using the false positive rate on the x-axis and the true positive rate on the y-axis 10 times.

[0112] [Figure 50] ROC plots of 13 peptides by MRM-MS and 2 proteins by ELISA after removing the proteins found in depleted plasma.

[0113] [Figure 51] It shows the random forest model for all study group comparisons.

[0114] [Figure 52] It shows some discriminative features in the study group comparisons.

[0115] [Figure 53] It shows the number of proteins (e.g., the number of proteins identified from corona analysis) for panel sizes ranging from 1 particle type to 12 particle types.

[0116] [Figure 54] It shows examples of biomarkers.

[0117] [Figure 55] A diagram showing a non - limiting example of a web / mobile application providing system (in this case, a system providing a browser - based and / or native mobile user interface); and

[0118] [Figure 56] A diagram showing a non - limiting example of a cloud - based web / mobile application providing system (in this case, the system includes elastically load - balanced, automatically scaled web server and application server resources and synchronized and replicated databases).

[0119] [Figure 57] Showing the ROC curve of a pulmonary nodule classifier, with sensitivity and corresponding specificity listed.

[0120] [Figure 58] Showing the feature information and importance of the pulmonary nodule classifier shown in FIG. 57.

[0121] [Figure 59] Showing some aspects of the samples used in the research described in this specification.

[0122]

[0123] [Figure 60] Showing the number of protein groups observed in a process control sample.

[0124] [Figure 61] Showing some coefficient of variation (CV) values.

[0125] [Figure 62] Including a protein abundance heatmap of samples from subjects with malignant and benign pulmonary nodules.

[0126] [Figure 63] It includes a volcano plot plotting the log-fold change in protein abundance against the negative logarithm of the p-value.

[0127] [Figure 64] Examples of some proteins from the initial univariate analysis are shown.

[0128] [Figure 65A] It includes a graph showing some proteins upregulated in a biological fluid sample from a subject with malignant lung nodules.

[0129] [Figure 65B] It includes a graph showing some proteins downregulated in a biological fluid sample from a subject with malignant lung nodules.

[0130] [Figure 66] It includes a graph showing that the differentially expressed proteins were enriched in metabolic pathways and phosphorylation pathways.

[0131] [Figure 67] Some extrapolated mRNA data showing differentially expressed proteins in metabolic pathways are shown.

[0132] [Figure 68] It is an image showing the locations where several samples were collected for the study.

[0133] [Figure 69A] Some aspects of the research subjects and proteomics platforms that can be used in the methods described herein are shown.

[0134] [Figure 69B] Some aspects of the proteomics platforms that can be used in the methods described herein are shown.

[0135] [Figure 69C]Several additional multi-omics configurations are shown.

[0136] [Figure 70] This specification includes a graphical representation of the coefficient of variation (CV) values ​​obtained in the studies described herein.

[0137] [Figure 71] This specification includes empirical power curves for detecting protein changes in the studies described herein.

[0138] [Figure 72] This specification includes a graphical representation of the detected protein groups and peptide counts obtained in the studies described herein.

[0139] [Figure 73] This specification includes a graphical representation of protein concentrations against natural logarithmic protein intensity data obtained in the studies described herein.

[0140] [Figure 74] This specification includes a graphical representation of protein concentrations from data obtained in the studies described herein.

[0141] [Figure 75A] Includes the median normalized logarithmic intensity CV of proteins detected in 100% of the samples.

[0142] [Figure 75B] Includes the median normalized logarithmic intensity CV of proteins detected in at least 25% of the samples.

[0143] [Figure 76] This includes the number of unique protein groups within several sample data.

[0144] [Figure 77A] Includes relative fluorescence units for concentrations on several standard curves.

[0145] [Figure 77B] Includes relative fluorescence units for several standard curves.

[0146] [Figure 78A] This specification includes the peptide yields of several nanoparticles used in the experiments described herein.

[0147] [Figure 78B] This specification includes the peptide yields of several nanoparticles used in the experiments described herein.

[0148] [Figure 79A] Includes a graph of MS1 intensity over time.

[0149] [Figure 79B] Includes intraday central venous catheterization with MS1 intensity.

[0150] [Figure 80A] Includes a graph of iRT peptides ranked by FWHM.

[0151] [Figure 80B] Includes a plot showing retention time.

[0152] [Figure 81A] Includes a plot showing the distribution of protein group numbers per sample.

[0153] [Figure 81B] Includes intraday central venous catheterization with MS1 intensity.

[0154] [Figure 82] This includes a volcano plot showing the intensity difference and P-value of peptides detected in biological fluid samples. In the volcano plot, the x-axis displays the median difference in peptide intensity, and the y-axis displays the peptide P-value on a harmonic mean basis.

[0155] [Figure 83]This includes graphs showing several transitions of the peptide ANVFVQLPR from protein P35858 in benign and malignant groups.

[0156] [Figure 84] This document includes a graph comparing the significant differences between OpenTarget (OT) scores and peptides in lung cancer. The graph displays the OpenTarget score on the x-axis and the P-value on the y-axis.

[0157] [Figure 85] This includes a volcano plot showing the intensity difference and p-value for metabolites in lung nodule subjects. In the volcano plot, the x-axis displays the median intensity difference, and the y-axis displays the p-value.

[0158] [Figure 86] Includes a figure showing a sample cohort of lung detections by SEER. The figure shows that out of 589 eligible subjects, 186 met all the criteria.

[0159] [Figure 87] This diagram illustrates a stepwise approach to discovering version 1, version 2, and version 3 classifiers through experimental development.

[0160] [Figure 88] The document includes graphs showing power curves for each analyte class. These graphs include curves for proteins, metabolites, and lipids.

[0161] [Figure 89] This includes a volcano plot showing the intensity difference and p-value for peptides in lung nodule subjects. In the volcano plot, the x-axis displays the median difference in peptide intensity, and the y-axis displays the peptide p-value on a harmonic mean basis.

[0162] [Figure 90]This includes graphs showing several transitions of the peptide LEYLLLSR from protein P35858 in benign and malignant groups.

[0163] [Figure 91] This includes graphs showing several transitions of the peptide ANVFVQLPR from protein P35858 in benign and malignant groups.

[0164] [Figure 92] This includes graphs showing several transitions of the peptide FLNVLSPR from protein P17936 in benign and malignant groups.

[0165] [Figure 93] The image shows a representation of StringDB. This image highlights the known interactions between IGFALS and IGFBP3.

[0166] [Figure 94] This includes a volcano plot showing the intensity difference and p-value for metabolites in lung nodule subjects. In the volcano plot, the x-axis displays the median intensity difference, and the y-axis displays the p-value.

[0167] [Figure 95] The study includes graphs showing the levels of biopterin metabolites in benign and malignant groups. The graphs display the group type on the x-axis and the amount of metabolites on the y-axis.

[0168] [Figure 96] This includes a volcano plot showing the intensity difference and p-value for lipids in lung nodules. In the volcano plot, the x-axis displays the median intensity difference, and the y-axis displays the p-value.

[0169] [Figure 97] This document includes a graph comparing the significant differences between OpenTarget (OT) scores and peptides in lung cancer. The graph displays the OpenTarget score on the x-axis and the P-value on the y-axis.

[0170] [Figure 98] This diagram illustrates a stepwise approach to discovering version 1, version 2, and version 3 classifiers through experimental development.

[0171] [Figure 99] This includes graphs showing the pre-test probability of having benign nodules, as well as the pre-test and post-test probabilities of having benign nodules. The graphs display probability on the x-axis and the number of subjects on the y-axis.

[0172] [Figure 100] Includes a graph comparing sensitivity and specificity. The graph displays specificity on the x-axis and sensitivity on the y-axis.

[0173] [Figure 101] This shows ROC curves for 223 subjects with mRNA data from a colorectal cancer (CRC) study. The x-axis shows the false positive rate, and the y-axis shows the true positive rate. AUC values ​​are provided.

[0174] [Figure 102] Includes volcano plots showing differences in gene expression across various gene groups in colorectal cancer research.

[0175] [Figure 103] The ROC curves for ProteoGraph, mRNA, and ProteoGraph+mRNA are shown. The AUC values ​​for each are provided.

[0176] [Figure 104] The ROC curves for ProteoGraph, PiQuant, mRNA, microRNA, and ProteoGraph+PiQuant+mRNA+microRNA are shown. The AUC values ​​for each are provided.

[0177] [Figure 105] The ROC curves for PiQuant, mRNA, and PiQuant+mRNA are shown. The AUC values ​​for each are provided.

[0178] [Figure 106] ROC curves are shown for classification based on distinct or combined types of biomolecules.

[0179] Embedding by reference All publications, patents, and patent applications referenced herein are invoked by reference to the same extent as if each individual publication, patent, or patent application were specifically and individually indicated as being invoked by reference herein. To the extent that any publications and patents or patent applications contained herein conflict with the disclosures contained herein, this Specification is intended to supersede and / or take precedence over any such conflicting material. [Modes for carrying out the invention]

[0180] This disclosure provides non-invasive methods for diagnosing or ruling out the presence of a disease in a subject, or the risk of developing a disease in a subject. Examples of diseases may include cancers such as pancreatic cancer, breast cancer, liver cancer, ovarian cancer, or colorectal cancer. Early identification of a disease in a subject and early treatment can save the subject from further disease progression. Non-invasive tests can also be used to rule out the presence of a disease, thereby eliminating the need for the subject to undergo invasive tests such as biopsies, which may be painful and stressful, or carry a risk of harm to the subject.

[0181] Multi-omics approaches can unlock the ability to detect diseases early in their onset and improve the accuracy of disease detection. Figure 1A illustrates several embodiments of multi-omics approaches for early disease detection, which can combine genomic DNA or DNA methylation information (an example of what can be a generally static indicator of risk) with molecular phenotypic information obtained from proteomics or metabolomics, which can be a more dynamic indicator of function. Figure 24 also illustrates several embodiments that can be included in multi-omics methods and includes some examples of pathological conditions that can be detected or evaluated. Figure 1B illustrates an example of integrating multiple omics data types. Any embodiment of these figures can be used in the methods described herein.

[0182] Figure 2A shows a non-limiting example of a method for predicting whether a subject has a disease such as cancer or is at risk of developing such a disease. The analysis may include obtaining a biofluid sample from the subject (200). The sample can be assayed or analyzed. The biofluid sample may be any one or any combination of the biofluids described herein. The sample may be: directly analyzed to generate data such as proteomics data (202); or, prior to the analysis in 202, contacted with particles described herein to obtain adsorbed biomolecules (203). After obtaining data from the analysis in 202, additional analysis (203) may be performed on the sample obtained from 200 or 201 to obtain additional datasets such as transcriptomics data, genomics data, metabolomics data, or a combination thereof. Using the data or datasets obtained from the analysis in 202 or 203, a classifier may be generated (205). The classifier can be applied to identify whether the subject has a disease or is at risk of developing a disease. The analysis can be improved by further iteration or refinement of the generation or application of the classifier. Figure 2B further illustrates some details that can be used in the methods described herein. Either aspect of Figure 2A or Figure 2B can be used in the methods described herein, such as a classification method.

[0183] Furthermore, the analyses shown in Figure 2A or Figure 2B can be applied before or during any procedure in any step included in Figure 3A. For example, the evaluation or analysis can be completed early in the course of an affected patient, before, immediately after, or as part of an invasive examination. Screening high-risk patients before performing invasive procedures such as biopsy or invasive treatment is useful. In general, opportunities in which the methods described herein may be useful may arise when screening high-risk patients for early detection of disease. The methods described herein can be used for such detection with greater accuracy and convenience than other methods. In Figure 3A, non-invasive examinations may include medical imaging, and invasive examinations may include obtaining a biopsy. A tumor may be suspected in the biopsy. Similar patient courses for pancreatic cancer, liver cancer, and colorectal cancer are shown in Figures 3B, 3C, and 41. As shown, the evaluation or analysis can be completed at any point in time or before Figure 3B, Figure 3C, or Figure 41.

[0184] In cases where the disease is pancreatic cancer, there are opportunities to screen high-risk patients before biopsy or pancreaticoscopy. For example, a primary opportunity to use the method herein is to screen high-risk pancreatic cancer patients for early detection with improved accuracy and convenience. In the course of liver cancer patients, there are opportunities to screen high-risk liver cancer patients before biopsy. For example, a primary opportunity to use the method herein may include improving decisions about uncertain liver nodules to determine whether a biopsy is necessary. Other opportunities may include surveillance or diagnosis of small, low-risk nodules, or follow-up (e.g., 3-6 months) to track the progression of small nodules. In the course of colorectal cancer (CRC) patients, there may be opportunities to screen high-risk patients before colonoscopy. Another opportunity may be to improve decisions about imaging or biopsy procedures.

[0185] Non-invasively acquired samples can be used to diagnose diseases by generating omics data and identifying patterns within the omics data that are associated with the disease. Disease diagnosis can be improved by combining and analyzing multiple types of data (e.g., multiple datasets, such as omics datasets). For example, combining multiple data types can improve the accuracy of predictions as to whether a subject has a particular disease. Combined data may be more accurate than individual datasets if the individual datasets independently exhibit errors or do not completely overlap. The methods described herein include generating or acquiring multi-omics data and using multi-omics data to make predictions as to whether a subject has a disease. Various methods for combining or analyzing multi-omics data are described. The use of multi-omics data and disease assessment are described in more detail.

[0186] Lung nodules can be classified using several methods. Lung nodules can be either benign or malignant. Malignant lung nodules can progress rapidly and develop into lung cancer, a common and often fatal cancer. Improved differentiation between malignant and benign lung nodules is needed. On the one hand, early diagnosis of malignant lung nodules can lead to earlier treatment regimens for patients with malignant nodules, potentially improving their prognosis. On the other hand, non-invasive diagnosis of benign or non-malignant lung nodules can help avoid potentially costly and invasive lung biopsies, and therefore may be better for patients with non-malignant nodules as well.

[0187] However, little progress has been made in developing useful clinical tests to diagnose and analyze whether pulmonary nodules are benign or malignant. Imaging methods often result in high rates of misdiagnosis (e.g., false positives). Smaller nodules are usually not detected by these imaging methods. Other non-invasive methods, such as biomarker screening, also have limitations. Given that plasma is in contact with many tissues in the body, proteins in plasma could be a useful biomarker discovery matrix. However, plasma proteins can present problems due to several factors, including a wide range of concentrations (e.g., 10 orders of magnitude). Complex biochemical workflows have attempted to circumvent these challenges, but they may not be practical for discovery studies of sufficient scale to guarantee validation and reproducibility. Alternatively, biomarker research has been limited to the evaluation or re-evaluation of existing markers without substantial improvement in clinical performance. Therefore, there remains a need for methods to diagnose or screen for the presence of benign or malignant pulmonary nodules based on the analysis of biomarkers in biological fluid samples. The methods described herein can address this need.

[0188] This specification discloses a method that includes obtaining biomolecular data. The data may include multi-omics data. This method may include generating or receiving data and then performing evaluations using a classifier. Evaluations may include applying a classifier, identifying a disease, ruling out the presence of a disease, predicting the likelihood of a disease, or selecting a treatment for a disease.

[0189] disease The methods described herein can be used to assess a pathological condition. The methods described herein can be used to predict or identify a pathological condition. A pathological condition may include diseases or disorders such as cancer. Examples of cancer include lung cancer, colorectal cancer, pancreatic cancer, liver cancer, ovarian cancer, breast cancer, prostate cancer, melanoma, bladder cancer, lymphoma, leukemia, kidney cancer, or uterine cancer. In some embodiments, cancer is breast cancer. Diseases may include disorders. A pathological condition may include having comorbidities associated with a disease or disorder. A reference to whether a subject has a pathological condition may include the subject being healthy. A healthy state can exclude a pathological condition. For example, having cancer can be excluded in a healthy state. Being healthy can be excluded in a pathological condition.

[0190] These methods may be useful for cancer diagnosis. These methods may be useful for cancer screening. This method may be useful for cancer treatment. This method may involve assaying proteins in a biofluid sample obtained from a subject having or suspected of having a nodule, such as a pulmonary nodule, in order to obtain a protein measurement. This method may involve applying a classifier to the protein measurement, thereby identifying the protein measurement as indicating whether the pulmonary nodule is cancerous or non-cancerous. In some cases, the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and then assaying the proteins adsorbed to the particles. Some embodiments involve obtaining or receiving a biofluid sample of the subject.

[0191] In some embodiments, the cancers detected by the methods described herein may be pancreatic cancer, liver cancer, ovarian cancer, or colorectal cancer. Cancer diagnosis may be improved by acquiring proteomics data or other omics data (such as lipidomics data). Cancer diagnosis can be improved by combining and analyzing multiple types of data (e.g., multiple datasets). For example, combining multiple data types, including proteomics, transcriptomics, genomics, metabolomics, or a combination thereof, can improve the accuracy of predictions as to whether a subject has cancer. In some embodiments, the methods described herein include generating or acquiring data and using the data to predict whether a subject has cancer. This method may also include identifying the type of cancer (e.g., liver cancer versus ovarian cancer). Various methods for combining or analyzing data are described, and the use of data for cancer assessment is described in more detail.

[0192] Cancer can be early-stage or late-stage. An example of early-stage cancer is stage I. Early-stage cancer may include stages I or II. Early-stage cancer may include stages I, II, or III. An example of late-stage cancer may include stage 4.

[0193] Cancer may include pancreatic cancer. Pancreatic cancer may be early-stage pancreatic cancer. In other embodiments, pancreatic cancer may be late-stage pancreatic cancer. Non-invasively acquired samples can be used for cancer diagnosis by generating data and identifying patterns in the data that are associated with cancer, such as pancreatic cancer. In certain embodiments, the method for detecting cancer may be an additional screening or diagnostic method, e.g., computed tomography (CT) scans showing pancreatic cancer, magnetic resonance imaging (MRI) scans showing pancreatic cancer, or positron emission tomography (PE) scans showing pancreatic cancer. The methods for detecting pancreatic cancer may include: T) scan, ultrasound examination indicating pancreatic cancer, endoscopic retrograde cholangiopancreatography indicating pancreatic cancer, angiography indicating pancreatic cancer, liver function tests (LFT) indicating pancreatic cancer, elevated levels of carcinoembryonic antigen (CEA) relative to control or baseline levels, elevated levels of carbohydrate antigen (CA) 19-9 relative to control or baseline levels, or a combination thereof. In some embodiments, the methods for detecting pancreatic cancer may include identifying subject symptoms such as jaundice, abdominal pain, gallbladder or hepatomegaly, thrombosis, digestive disorders, or depression, or a combination thereof. Any of these embodiments can be used to identify subjects at risk of having pancreatic cancer.

[0194] Cancer may include liver cancer. In some embodiments, the cancer detected by the method described herein may be liver cancer. Liver cancer may be early-stage liver cancer. In other embodiments, liver cancer may be late-stage liver cancer. In some cases, liver cancer may be stage I, II, III, or IV liver cancer. In some cases, the stage of liver cancer is unknown. Non-invasively obtained samples can be used for cancer diagnosis by generating data and identifying patterns in the data that are associated with cancer, such as liver cancer. In certain embodiments, the method for detecting cancer may include additional screening or diagnostic methods, such as undergoing a dynamic contrast computed tomography (CT) scan for liver cancer, undergoing a magnetic resonance imaging (MRI) scan for liver cancer, undergoing a liver function test (LFT) for liver cancer, elevated bilirubin levels relative to control or baseline values, elevated aminotransferase levels relative to control or baseline values, elevated alkaline phosphatase levels relative to control or baseline values, having hypoalbuminemia, increased prothrombin time relative to control or baseline values, elevated alpha-fetoprotein levels relative to control or baseline values, or having liver nodules, or a combination thereof. In some embodiments, the method for detecting cancer may include identifying symptoms of the subject, such as abdominal discomfort, pain, and tenderness, jaundice, white calcified stools, nausea, vomiting, bruising, or easy bleeding, weakness, or fatigue, or a combination thereof. Any of these embodiments can be used to identify subjects at risk of having liver cancer.

[0195] Cancer may include ovarian cancer. In some embodiments, the cancer detected by the methods described herein may be ovarian cancer. Ovarian cancer may be early-stage ovarian cancer. In other embodiments, ovarian cancer may be late-stage ovarian cancer. In some cases, the stage of ovarian cancer may be unknown. In some embodiments, the stage of ovarian cancer may be stage I, II, III, or IV. Non-invasively obtained samples can be used for cancer diagnosis by generating data and identifying patterns in the data that are associated with cancer, such as ovarian cancer. In certain embodiments, the method for detecting cancer may include additional screening or diagnostic methods, such as a computed tomography (CT) scan indicating ovarian cancer, undergoing a magnetic resonance imaging (MRI) scan indicating ovarian cancer, undergoing a positron emission tomography (PET) scan indicating ovarian cancer, undergoing a transvaginal ultrasonography indicating ovarian cancer, elevated cancer antigen (CA)-125 levels relative to control or baseline measurements, or having an ovarian cyst, or a combination thereof. In some embodiments, the method for detecting cancer may include identifying subject symptoms such as pelvic pressure, lower abdominal pain, vaginal bleeding, weight gain, weight loss, menstrual cycle irregularities, unexplained lower back pain that worsens over time, increased urination, gas, nausea, vomiting, or loss of appetite, or a combination thereof. Any of these embodiments can be used to identify subjects at risk of having ovarian cancer.

[0196] Cancer may include colorectal cancer or colorectal cancer (CRC). In some embodiments, the cancer detected by the method described herein may be colorectal cancer. Colorectal cancer may be early-stage colorectal cancer. In other embodiments, colorectal cancer may be advanced-stage colorectal cancer. Non-invasive acquisition The samples obtained can be used for cancer diagnosis by generating data and identifying patterns in the data that are associated with cancer, such as colorectal cancer. Cancer diagnosis can be improved by obtaining proteomics data. In certain embodiments, the method for detecting cancer may include additional screening or diagnostic methods, e.g., computed tomography (CT) scans indicating indicators of colorectal cancer, liver function tests (LFT) indicating indicators of colorectal cancer, measurement of carcinoembryonic antigen (CEA) levels relative to control or baseline values, measurement of blood in stool, performing fecal immunochemistry (FIT), or a combination thereof. Any of these embodiments can be used to identify subjects at risk of having colorectal cancer. For example, a subject identified as being at risk of having colorectal cancer can be identified as being at risk by one of these methods. The non-invasive methods described herein allow patients who do not have colorectal cancer to avoid further invasive examinations or treatments, e.g., colonoscopy or cancer biopsy, or colorectal cancer treatment. On the other hand, the non-invasive methods described herein can be used to identify individuals who are likely to have colorectal cancer and to determine if they require further examination (e.g., invasive examination) or treatment. Colorectal cancer may be an example of colorectal cancer (CRC). The references or teachings herein relating to colorectal cancer may be applied to CRC, and vice versa.

[0197] Cancer may include lung cancer. One example of lung cancer is non-small cell lung cancer (NSCLC). Another example of lung cancer is small cell lung cancer. A method for diagnosing lung nodules is disclosed. This method, in which lung nodules are identified by computed tomography (CT) scans, may be useful for the diagnosis, treatment, or screening of patients who have not undergone lung biopsy. This method may be useful in informing physicians about the likelihood that a lung nodule is benign or malignant. With test results from such a method, physicians can help patients avoid unnecessary biopsies. For example, this method can be used as a rule-out test. With test results from such a method, physicians can identify what should be biopsied. For example, this method can be used as a rule-in test.

[0198] A diagnostic method for identifying candidates for CT imaging is disclosed. This method may be useful in the diagnosis, treatment, or screening of patients who may be candidates for CT imaging. This method may be useful for high-risk patients (e.g., as defined by the USPSTF or other organizations) who are candidates for lung cancer screening but have not yet undergone a CT scan. This method can inform physicians of the likelihood that a patient has lung cancer. Thus, this method can inform physicians of the urgency of a patient's lung condition or the need to obtain a CT scan. Such a method may be useful for high-risk patients, such as those who do not follow other CT screening methods. This method can improve patient selection or compliance for CT imaging. This method can improve patient selection or compliance for biopsy.

[0199] A method for recurrent monitoring is disclosed. This method may be useful for monitoring patients with potentially resectable lung cancer. This method may be useful for monitoring patients who have undergone postoperative therapeutic intervention. This method may be useful for monitoring patients undergoing adjuvant chemotherapy or radiotherapy intervention. This method may be useful for detecting cancer recurrence before CT scans or other medical imaging. This method may be useful for recurrence surveillance testing. This method can be adapted or developed to suit the patient's treatment regimen.

[0200] This specification describes the assay of proteins in biological fluid samples obtained from subjects with or suspected of having pulmonary nodules to obtain protein measurements; and the application of a classifier to the protein measurements to determine whether the pulmonary nodules are cancerous or The method includes identifying protein measurements as indicating non-cancerous, and the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and then assaying the proteins adsorbed to the particles. This method may be useful for cancer diagnosis or screening.

[0201] This specification describes a method comprising: obtaining a biofluid sample from an object having a pulmonary nodule; contacting the biofluid sample with particles such that the particles adsorb protein-containing biomolecules onto the particles; assaying the biomolecules adsorbed onto the particles to generate proteomics data; and classifying the proteomics data as indicating whether the pulmonary nodule is cancerous or non-cancerous. This method may be useful for cancer diagnosis or screening.

[0202] Described herein are methods for determining the nodule-associated status of a pulmonary nodule in a sample obtained from a subject. In some embodiments, the nodule-associated status includes the presence or absence of a pulmonary nodule in the subject. In some embodiments, the nodule-associated status includes determining whether the pulmonary nodule is benign or malignant. In some embodiments, the method includes screening for the nodule-associated status by assaying a biomarker in a sample obtained from a subject. In some embodiments, the biomarker includes at least one protein in the sample. In some embodiments, the sample is a biological fluid sample. In some embodiments, the biological fluid sample is brought into contact with particles described herein to adsorb proteins in the biological fluid sample. In some embodiments, the method includes obtaining protein measurements of the proteins in the sample. In some embodiments, the method includes applying a classifier to the protein measurements to identify the protein measurements as indicating whether the pulmonary nodule is cancerous or non-cancerous. In some embodiments, the classifier is generated using proteomics data obtained by bringing a training sample into contact with the particles so that the particles adsorb proteins in the training sample. The adsorbed proteins can then be assayed by the method described herein. In some embodiments, a subject is suspected of having a pulmonary nodule or is identified as having a pulmonary nodule by the imaging method described herein. In some embodiments, a report is generated based on the identification of protein measurements indicating whether the pulmonary nodule is cancerous or non-cancerous. In some embodiments, the report indicates the likelihood or indicator that the pulmonary nodule is cancerous or non-cancerous. In some embodiments, the report indicates that the pulmonary nodule is cancerous. In some embodiments, the report indicates that the pulmonary nodule contains non-small cell lung cancer (NSCLC). In some embodiments, the method described herein generates a classifier that includes features indicating protein measurements indicating whether the pulmonary nodule is cancerous or non-cancerous.In some embodiments, features include control protein measurements, mass spectra, m / z ratios, chromatography results, immunoassay results, or light or fluorescence intensity. In some embodiments, the classifier is trained using either a computational method or a machine learning method described herein.

[0203] This specification describes, in several embodiments, a method for recommending lung cancer treatment to a subject when, based on an analysis of protein measurements described herein, the subject is determined to have a malignant pulmonary nodule. In some embodiments, the protein measurements are classified as indicating that the pulmonary nodule is cancerous.

[0204] This specification discloses, in several embodiments, methods useful for diagnosing, screening, or treating subjects. Some embodiments involve assaying proteins in a biofluid sample obtained from a subject suspected of having pulmonary nodules in order to obtain protein measurements. Some embodiments involve applying a classifier to the protein measurements. Several embodiments include identifying protein measurements as indicating that a subject has pulmonary nodules. In some embodiments, the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and then assaying the proteins adsorbed to the particles.

[0205] This specification discloses, in several embodiments, methods useful for diagnosing, screening, or treating subjects. Some embodiments involve assaying proteins in a biofluid sample obtained from a subject suspected of having lung cancer in order to obtain protein measurements. Some embodiments involve applying a classifier to the protein measurements. Some embodiments involve identifying the protein measurements as indicating that the subject has lung cancer. In some embodiments, the classifier is generated using proteomics data obtained by contacting a training sample with particles so that the particles adsorb proteins in the training sample, and then assaying the proteins adsorbed to the particles.

[0206] This specification discloses, in several embodiments, methods useful for diagnosing, screening, or treating subjects. Some embodiments include obtaining a biofluid sample from a subject suspected of having pulmonary nodules. Some embodiments include contacting the biofluid sample with particles such that the particles adsorb biomolecules, including proteins, onto the particles. Some embodiments include assaying the biomolecules adsorbed onto the particles to generate proteomics data. Some embodiments include classifying the proteomics data, based on the proteomics data, as indicating that the subject has pulmonary nodules or as not indicating that the subject has pulmonary nodules.

[0207] This specification discloses, in several embodiments, methods useful for diagnosing, screening, or treating subjects. Some embodiments include obtaining a biofluid sample from a subject suspected of having lung cancer. Some embodiments include contacting the biofluid sample with particles such that the particles adsorb biomolecules, including proteins, onto the particles. Some embodiments include assaying the biomolecules adsorbed onto the particles to generate proteomics data. Some embodiments include classifying the proteomics data, based on the proteomics data, as indicating that the subject has lung cancer or as not indicating that the subject has lung cancer.

[0208] This specification discloses, in several embodiments, methods useful for monitoring subjects. Some embodiments include obtaining a biofluid sample from a subject at risk of lung cancer recurrence. Some embodiments include contacting the biofluid sample with particles such that the particles adsorb biomolecules, including proteins, onto the particles. Some embodiments include assaying the biomolecules adsorbed onto the particles to generate proteomics data. Some embodiments include classifying the proteomics data, based on the proteomics data, as indicating that the subject has lung cancer recurrence or as not indicating that the subject has lung cancer recurrence. In some embodiments, the subject is receiving lung cancer treatment such as chemotherapy, radiotherapy, or surgery. In some embodiments, the cancer may be resectable. In some embodiments, lung cancer includes NSCLC.

[0209] In some cases, lung nodules are described as malignant or cancerous. The terms malignant and cancerous may be used interchangeably. Malignant or cancerous lung nodules may be called lung cancer, and vice versa. In some cases, lung nodules are described as benign or non-cancerous. The terms benign and non-cancerous may be used interchangeably.

[0210] Samples and subjects Some aspects relate to the subject. For example, the method described herein can be used to obtain the subject It is possible to evaluate, or to evaluate samples from the subject. Multi-omics data can be generated from the sample of the subject.

[0211] The methods described herein can be used to identify subjects who are likely to have, or at risk of having, a disease such as cancer. Subjects may have lung cancer, pancreatic cancer, liver cancer, ovarian cancer, or colorectal cancer. Cancer may include adenocarcinoma, such as pancreatic adenocarcinoma. Subjects may have cancer. Subjects may not have cancer. Subjects may have pancreatic cancer, liver cancer, ovarian cancer, or colorectal cancer. Subjects may not have pancreatic cancer, liver cancer, ovarian cancer, or colorectal cancer. Subjects may be at risk of having pancreatic cancer, liver cancer, ovarian cancer, or colorectal cancer. Subjects may have a mass (e.g., a small nodule or cyst) in the pancreas. Subjects may have a mass (e.g., a nodule) in the liver. Liver cancer may include hepatocellular carcinoma (HCC). Liver cancer may include stage I, stage II, stage III, or stage IV liver cancer. The subjects may have a mass (e.g., a nodule or cyst) in one or both ovaries. Ovarian cancer may include stage I, stage II, stage III, or stage IV ovarian cancer. Ovarian cancer may include stage III ovarian cancer. Ovarian cancer may include stage IV ovarian cancer. The subjects may have a mass (e.g., a nodule) in the colon. The subjects may have lung nodules or cancer. The subjects may be at risk of having breast cancer. The subjects may have a mass (e.g., a small nodule or cyst) in the breast.

[0212] Samples can be obtained from subjects for the purpose of identifying cancer in them. Subjects may be suspected of having cancer or suspected of not having cancer. This method can be used to confirm or rule out suspected cancer.

[0213] The data described herein can be generated from a sample of the subject. The sample may be a biological fluid sample or a tumor sample (e.g., an abnormal growth biopsied from the subject). Examples of biological fluids include blood, serum, or plasma. A sample may include a blood sample. A sample may include a serum sample. A sample may include a plasma sample. One or more biological fluid samples may include blood, serum, or plasma samples. Other examples of biological fluids include urine, tears, semen, milk, vaginal fluid, mucus, saliva, sweat, or cell homogenates.

[0214] Samples may be obtained from subjects for the purpose of identifying pathological conditions in those subjects. Subjects may be suspected of having a pathological condition or suspected of not having one. This method can be used to confirm or deny a suspected pathological condition. In some embodiments, samples from subjects are used to determine whether a tumor, nodule (e.g., a pulmonary nodule), or cyst is cancerous or non-cancerous.

[0215] Biological fluid samples can be obtained from subjects. For example, blood, serum, or plasma samples can be obtained from subjects by blood collection. Other methods for obtaining biological fluid samples include aspiration or collection by swab.

[0216] The biological fluid sample may be cell-free or substantially cell-free. To obtain a cell-free or substantially cell-free biological fluid sample, the biological fluid may undergo sample preparation methods such as centrifugation and pellet removal.

[0217] Non-biological fluid samples can be obtained from the subject or patient. For example, samples may include tissue samples. Some examples of organs or tissues that can be sampled include tissue from the lungs, colon, pancreas, liver, breast, or ovary. Samples may also include tumors taken from the organ or tissue of the subject. Tumors may be suspected of being cancerous. Tumors may include nodules (e.g., colorectal nodules or hepatic nodules). Tumors may be cysts. This may include (for example, ovarian cysts). Nodules or cysts can be identified by a physician as having a high or low risk of being cancerous before performing the methods described herein. The mass may be biopsied, for example, by a needle biopsy procedure. A needle biopsy procedure may include inserting a fine needle into the liver from the abdomen of the subject to obtain a tissue sample, which can then be examined under a microscope for signs of cancer. The sample may include a cell sample. The sample may include a homogenate of cells or tissue. The sample may include the supernatant of a centrifuged homogenate of cells or tissue.

[0218] The sample may contain lung tissue. The sample may contain colon tissue. The sample may contain pancreatic tissue. The sample may contain liver tissue. The sample may contain breast tissue. The sample may contain ovarian tissue. The tissue may be cancerous. The tissue may be non-cancerous. The tissue may be suspected of being cancerous. The tissue may be malignant. The tissue may be non-malignant. The tissue may be suspected of being malignant.

[0219] Samples (e.g., biological fluid or tissue samples) can be obtained from subjects at any stage of the screening procedure, such as before, during, or after the stages shown in Figure 3A. Samples can be obtained before or during the stage in which the subject is a candidate for biopsy, pancreatography, or colonoscopy for early detection of disease. Samples can be obtained before or during non-invasive, invasive, treatment, or monitoring stages.

[0220] Data can be generated from a single sample or from multiple samples. Data from multiple samples can be obtained from the same subject. In some cases, different data types may be obtained from samples collected differently or from samples collected in separate containers. Samples may be collected in containers containing one or more reagents, such as preservation reagents or biomolecule isolation reagents. Some examples of reagents include heparin, ethylenediaminetetraacetic acid (EDTA), citrate, antisolubilants, or combinations of reagents. Samples from a subject may be collected in multiple containers containing different reagents, for example, to preserve or isolate different types of biomolecules. Samples may also be collected in containers that do not contain reagents. Samples can be collected simultaneously (e.g., at the same time or on the same day) or at different times. Samples can be frozen, refrigerated, heated, or stored at room temperature.

[0221] Using the methods described herein, subjects can be identified as likely to have a disease or not. Diseases include cancers such as pancreatic cancer, liver cancer, ovarian cancer, or colorectal cancer. Some aspects of this disclosure include identifying whether a pulmonary nodule in a subject is malignant or non-malignant. Pulmonary nodules may be present in the lungs of a subject. A subject may be identified as having a pulmonary nodule. In some aspects, a subject may have multiple pulmonary nodules. A subject may have lung cancer. A subject may be at risk of lung cancer. A subject may have pulmonary complications. A subject may have comorbidities as described herein. A subject may experience dyspnea. A subject may have fluid in its lungs.

[0222] In some cases, the subject may be monitored. For example, information about the possibility that the subject has a disease may be used to decide to monitor the subject without providing treatment. In other situations, the subject may be monitored while receiving treatment to see if the subject's condition improves. In some embodiments, a subject with a pulmonary nodule may be monitored to determine the progression of the pulmonary nodule. The pulmonary nodule of the subject may be monitored. The subject may be treated as described herein.

[0223] The subject may be a vertebrate. The subject may be a mammal. Among mammals, there are This may include mice, gerbils, guinea pigs, or hamsters. Mammals may include foxes, bears, dogs, monkeys, cattle, pigs, or sheep. The subject may be a primate. Primates may include apes or monkeys. Primates may include chimpanzees, lemurs, bonobos, orangutans, or baboons. The subject may be a human. The subject may be an adult (e.g., at least 18 years old). The subject may be male. The subject may be female. The subject may have a medical condition. For example, the subject may have a disease or disability, a comorbidity of a disease or disability, or may be healthy.

[0224] The methods described herein may include the use of samples, such as biological samples. For example, the methods may include determining one or more biomarker measurements in the sample. Biological samples may be from subjects, such as subjects with pulmonary nodules. Biological samples may include blood samples from which red blood cells have been removed. For example, biological samples may include plasma samples. Biological samples may include serum samples. Biological samples may include blood or blood components. Biological samples may include blood samples. The samples described or used herein may be from subjects described herein, such as subjects with identified pulmonary nodules.

[0225] The sample is consistent with the method disclosed herein for evaluating the presence or absence of one or more biomarkers associated with the presence of pulmonary nodules or malignant conditions. The subject may be human or a non-human animal. The biological sample may be a biological fluid. For example, biological fluids may be plasma, serum, CSF, urine, tears, cell lysates, tissue lysates, cell homogenates, tissue homogenates, papillary aspirates, fecal samples, synovial fluid and whole blood, or saliva. The sample may also be a non-biological sample such as water, milk, solvent, or homogenized fluid. The biological sample may contain multiple proteins or proteomics data, which can be analyzed after the adsorption of proteins to the surface of various types of particles in the panel and subsequent digestion of the protein corona. The proteomics data may include nucleic acids, peptides, or proteins. Any of the samples herein may contain a number of different analytes that can be analyzed using the method disclosed herein. The analytes may be proteins, peptides, small molecules, nucleic acids, metabolites, lipids, or any molecules that can bind to or interact with the surface of the particle type.

[0226] The sample may be a biological fluid. A biological sample may include biological fluid samples, such as cerebrospinal fluid (CSF), synovial fluid (SF), urine, plasma, serum, tears, gingival crevicular exudate, semen, whole blood, milk, nipple aspirate, mammary duct lavage fluid, vaginal fluid, nasal secretions, ear fluid, gastric juice, pancreatic juice, trabecular meshwork fluid, pulmonary lavage fluid, prostatic fluid, sputum, feces, bronchial lavage fluid, fluid from swabbing, bronchial aspirate, sweat, or saliva. The biological fluid may be a fluidized solid, such as a tissue homogenate, or a fluid extracted from a biological sample. A biological sample may be, for example, a tissue sample or a fine-needle aspiration (FNA) sample. A biological sample may be a cell culture sample. For example, a sample that can be used in the method disclosed herein may include cells growing in a cell culture or cell-free material taken from a cell culture. The biological fluid may be a fluidized biological sample. For example, the biological fluid may be a fluidized cell culture extract. The sample may be extracted from a fluid sample or from a solid sample. For example, the sample may contain gaseous molecules extracted from a fluidized solid (e.g., a volatile organic compound). In some embodiments, the biological fluid may include blood, plasma, or serum.

[0227] Methods consistent with this disclosure may include collecting species from biological samples (e.g., isolating, concentrating, or purifying). Species are biomolecules (e.g., The species may be proteins, biomolecular structures (e.g., peptide aggregates or ribosomes), cells, or tissues. Species can be selectively collected from biological samples. For example, the method may involve isolating cancer cells from tissue (e.g., as a tissue biopsy) or from bodily fluids such as whole blood, plasma, or buffy coat (e.g., as a liquid biopsy). The method may include samples that do not contain cancer cells. Species may be treated before analysis. For example, proteins may be reduced and degraded, nucleic acids may be separated from histones, or cells may be lysed.

[0228] Biological samples may be obtained from or derived from human subjects. Before processing, biological samples may be stored under various storage conditions, such as at different temperatures (e.g., room temperature, refrigerated or frozen conditions, 25°C, 4°C, -18°C, -20°C, or -80°C) or in different suspensions (e.g., EDTA collection tubes, cell-free RNA collection tubes, or cell-free DNA collection tubes).

[0229] In some cases, samples may be depleted before biomarker analysis. Samples can be depleted using commercially available kits. For example, kits that can be used to deplete samples may include spin column-based depletion kits, albumin depletion kits, immunodepletion kits, or abundant protein depletion kits. Non-exclusive examples of kits that can be used for sample depletion include the PureProteome® Human Albumin / Immunoglobulin Depletion Kit (EMD Millipore Sigma), the ProteoPrep® Immunoaffinity Albumin & IgG Depletion Kit (Millipore Sigma), the Seppro® Protein Depletion Kit (Millipore Sigma), the Top 12 Abundant Protein Depletion Spin Columns (Pierce), or the Proteome Purify® Immunodepletion Kit (R&D Systems). Depletion can remove high concentrations of biomolecules from the sample. For example, the method may include removing albumin from a plasma sample before low-concentration biomarker analysis. The sample may include depleted plasma.

[0230] Data generation and use Methods disclosed herein may include obtaining data such as multi-omics data generated from one or more biological fluid samples collected from a subject. The data may include biomolecular measurements such as protein measurements, transcript measurements, genetic material measurements, or metabolite measurements. Omics data may include any of the following: proteomics data, genomics data, transcriptomics data, or metabolomics data. This section includes several methods for generating each of these types of omics data. Methods for generating or analyzing omics data may also be applied to methods for generating or analyzing individual biomolecules or sets of biomolecules. Other types of omics data may also be generated. Descriptions of generating or analyzing omics data may also be applied to methods for generating or analyzing individual biomolecules or sets of biomolecules that do not necessarily include omics data. Embodiments described in relation to biomolecular data may relate to the measurement of biomolecules, or vice versa. Data may be labeled or identified as indicating or not indicating disease. The data may be labeled or identified as indicating pancreatic cancer, liver cancer, ovarian cancer, or colorectal cancer, or as not indicating pancreatic cancer, liver cancer, ovarian cancer, or colorectal cancer. The methods described herein may include obtaining multi-omics measurements by performing an assay, etc.

[0231] The methods described herein may include generating or using omics data. Omics data may include data on all biomolecules of a particular type, such as proteins, transcripts, genetic material, or metabolites. Ohm data may include data on a subset of biomolecules. For example, omics data could include data on more than 500, 750, 1000, 2500, 5000, 10,000, or 25,000 biomolecules of a particular type. The methods described herein may include obtaining measurements of certain types of biomolecules greater than 10, 20, 30, 40, 50, 75, 100, 250, 500, 750, 1000, 1250, 2500, 5000, 7500, 10,000, 12,500, 15,000, 17,500, 20,000, 22,500, or 25,000. Methods described herein may include obtaining measurements of certain types of biomolecules in numbers less than 10, less than 20, less than 30, less than 40, less than 50, less than 75, less than 100, less than 250, less than 500, less than 750, less than 1000, less than 1250, less than 2500, less than 5000, less than 7500, less than 10,000, less than 12,500, less than 15,000, less than 17,500, less than 20,000, less than 22,500, or less than 25,000. Any of the aforementioned numbers of biomolecules may be measured for each of multiple data types. Multi-omics includes at least 100 measurements for each of at least two types of omics data. Multi-omics includes at least 500 measurements for each of at least two types of omics data. Multi-omics includes at least 1000 measurements for each of at least two types of omics data. The data may relate to the presence, absence, or quantity of a given biomolecule. Examples of data types include lipid, protein, peptide, transcript, mRNA, miRNA, DNA sequence, methylation, or metabolite data.

[0232] Other documents are individually and separately indicated to be incorporated by reference for any purpose.

[0233] Deep proteome coverage is advantageous for multi-omics approaches. New technologies and sample availability address past challenges in scaling proteomics. Some challenges include access to large, well-collected and annotated sample cohorts for specific clinical problems, technical challenges associated with plasma proteomics that can limit clinical transitions, such as reproducibility, throughput, and depth of coverage, and the reproducible measurement and integration of multi-omics datasets that provide novel insights into cancer biology.

[0234] The concepts described herein may help address some of these challenges. For example, these concerns can be addressed by using particles or by including additional omics types.

[0235] Methods for multi-omics analysis are disclosed herein. “Multi-omics (Multi-omic(s))” or “multiomic(s)” may include analytical approaches for large-scale analysis of biomolecules, and datasets may be multiple omics such as proteome, genome, transcriptome, lipidome, and metabolome. Non-limiting examples of multi-omics data include proteomics data, genomics data, lipidomics data, glycomics data, transcriptomics data, or metabolomics data. “Biomolecules” in “biomolecules corona” may refer to any molecule or biological component that can be produced by or present in biological tissues. Non-limiting examples of biomolecules include proteins (protein corona), polypeptides, polysaccharides, sugars, lipids, lipoproteins, metabolites, oligonucleotides, nucleic acids (DNA, RNA, microRNA, plasmids, single-stranded nucleic acids, Examples include double-stranded nucleic acids, metabolomes, and small molecules, such as primary metabolites, secondary metabolites, and other natural products, or any combination thereof. In some embodiments, the biomolecules are selected from the group of proteins, nucleic acids, lipids, and metabolites.

[0236] Some aspects of a multi-omics strategy may include a well-defined disease biobank with multiple sample types optimized for multi-omics measurements, the development and optimization of novel proteomics techniques to improve proteome coverage and throughput without compromising reproducibility, or an unbiased multi-omics platform incorporating state-of-the-art instruments and advanced machine learning analytics to transform the early detection of complex diseases.

[0237] Proteomics data The data described herein, such as multi-omics data, may include protein data or proteomics data. Proteomics data may include data relating to proteins, peptides, or proteoforms. This data may include peptides alone, proteins alone, or a combination of both. An example of a peptide is an amino acid chain. An example of a protein is a peptide or a combination of peptides. For example, a protein may include one, two, or more peptides bound together. A protein may be a secreted protein. Proteomics data may include data relating to various proteoforms. Proteoforms may include various forms of proteins generated from a genome with various sequence mutations, splice isoforms, or post-translational modifications. Proteomics data may be generated using an unbiased, untargeted approach, or may include a specific set of proteins. The aspects described in relation to proteomics data may relate to protein data, and vice versa.

[0238] Proteomics data may include information about the presence, absence, or quantity of various proteins and peptides. For example, proteomics data may include the quantity of a protein. The quantity of a protein can be expressed as a protein concentration or volume, for example, the concentration of a protein in a biological fluid. The quantity of a protein may be compared to another protein or another biomolecule. Proteomics data may include information about the presence of a protein or peptide. Proteomics data may also include information about the absence of a protein or peptide. Proteomics data can be distinguished by subtype, each subtype containing different types of proteins, peptides, or proteoforms.

[0239] Proteomics data typically includes data on a large number of proteins or peptides. For example, proteomics data may include information on the presence, absence, or quantity of more than 1,000 proteins or peptides. In some cases, proteomics data may include information on the presence, absence, or quantity of 5,000, 10,000, 20,000, or more peptides, proteins, or proteoforms. Proteomics data can include up to approximately 1 million proteoforms. Proteomics data may include a range of proteins, peptides, or proteoforms, defined by any of the aforementioned numbers of proteins, peptides, or proteoforms. Some examples of proteins or peptides that can be included in proteomics data are shown in Figures 6, 7, 10B, or 15.

[0240] In some embodiments, multi-omics data may include more than 10 peptides or protein groups, more than 15 peptides or protein groups, more than 20 peptides or protein groups, more than 25 peptides or protein groups, more than 30 peptides or This includes measurements of protein groups, groups of more than 35 peptides or proteins, groups of more than 40 peptides or proteins, groups of more than 45 peptides or proteins, groups of more than 50 peptides or proteins, groups of more than 75 peptides or proteins, groups of more than 100 peptides or proteins, groups of more than 250 peptides or proteins, groups of more than 500 peptides or proteins, groups of more than 1,000 peptides or proteins, groups of more than 2,500 peptides or proteins, groups of more than 5,000 peptides or proteins, groups of more than 10,000 peptides or proteins, groups of more than 15,000 peptides or proteins, or groups of more than 20,000 peptides or proteins. In some embodiments, the multi-omics data includes measurements of at least about 10 peptide or protein groups, at least about 15 peptide or protein groups, at least about 20 peptide or protein groups, at least about 25 peptide or protein groups, at least about 30 peptide or protein groups, at least about 35 peptide or protein groups, at least about 40 peptide or protein groups, at least about 45 peptide or protein groups, at least about 50 peptide or protein groups, at least about 75 peptide or protein groups, at least about 100 peptide or protein groups, at least about 250 peptide or protein groups, at least about 500 peptide or protein groups, at least about 1,000 peptide or protein groups, at least about 2,500 peptide or protein groups, at least about 5,000 peptide or protein groups, at least about 10,000 peptide or protein groups, at least about 15,000 peptide or protein groups, or at least about 20,000 peptide or protein groups.In some embodiments, protein data includes measurements of 10 or fewer peptide or protein groups, 15 or fewer peptide or protein groups, 20 or fewer peptide or protein groups, 25 or fewer peptide or protein groups, 30 or fewer peptide or protein groups, 35 or fewer peptide or protein groups, 40 or fewer peptide or protein groups, 45 or fewer peptide or protein groups, 50 or fewer peptide or protein groups, 75 or fewer peptide or protein groups, 100 or fewer peptide or protein groups, 250 or fewer peptide or protein groups, 500 or fewer peptide or protein groups, 1,000 or fewer peptide or protein groups, 2,500 or fewer peptide or protein groups, 5,000 or fewer peptide or protein groups, 10,000 or fewer peptide or protein groups, 15,000 or fewer peptide or protein groups, or 20,000 or fewer peptide or protein groups. A peptide or protein group may contain or consist of peptides. A peptide or protein group may contain or consist of protein groups.

[0241] Proteins may also contain post-translational modifications (PTMs). An example of a PTM is glycosylation. Proteins or peptides may contain glycoproteins or glycopeptides. Proteins may contain glycoproteins. Peptides may contain glycopeptides. An example of a PTM is phosphorylation. Proteins or peptides may contain phosphoproteins or phosphopeptides. Proteins may contain phosphoproteins. Peptides may contain phosphopeptides.

[0242] Proteomics data can be generated by any of the following methods. Generation of proteomics data may include the use of detection reagents that bind to peptides or proteins to produce a detectable signal. After using detection reagents that bind to peptides or proteins to produce a detectable signal, readings indicating the presence, absence, or amount of the protein or peptide can be obtained. Generation of proteomics data may include sample concentration, filtration, or centrifugation.

[0243] Proteomics data can be generated using mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assays, immunoassays, enzyme-linked immunosorbent assays, Western blotting, dot blotting, or immunostaining, or a combination thereof. Some examples of methods for generating proteomics data include the use of mass spectrometry, protein chips, or reverse-phase protein microarrays. Proteomics data can also be generated using immunoassays, such as enzyme-linked immunosorbent assays, Western blotting, dot blotting, or immunohistochemical assays. Generating proteomics data may include the use of immunoassay panels.

[0244] One method for obtaining proteomics data involves the use of mass spectrometry. An example of a mass spectrometry method is to separate proteins in parallel from different samples using high-resolution two-dimensional electrophoresis, and then select or stain differentially expressed proteins to be identified by mass spectrometry. Another method involves differentially labeling proteins from two different complex mixtures using stable isotope tags. Proteins in the complex mixture are isotoped and then digested to obtain labeled peptides. The labeled mixtures can then be combined, the peptides separated by multidimensional liquid chromatography, and analyzed by tandem mass spectrometry. Mass spectrometry methods can also include the use of liquid chromatography-mass spectrometry (LC-MS), i.e., a technique that combines the physical separation capabilities of liquid chromatography (e.g., HPLC) with mass spectrometry.

[0245] Proteins can be concentrated before assay or measurement. Concentration may involve concentrating one set of proteins while leaving another set untouched, or concentrating a single protein while leaving another untouched. Concentration can be achieved through the use of affinity reagents, for example, by incubating the affinity reagent with the sample before measuring the proteins in the sample. The affinity reagent may contain antibodies. The affinity reagent may contain particles such as nanoparticles. Proteins can be adsorbed onto the affinity reagent, separated from the rest of the sample, and then assayed using the proteomics assays described herein.

[0246] Proteomics data generation may involve contacting a sample with particles so that the particles adsorb biomolecules, including proteins. The adsorbed proteins may be part of the biomolecular corona. The adsorbed proteins can be measured or identified during the generation of proteomics data.

[0247] The generation of proteomics data may involve the use of a known amount of internal reference protein. The reference protein may be labeled. The labeling may include isotopic labeling. The generation of proteomics data may involve the use of a known amount of isotopically labeled internal reference protein (referred to as "PiQuant"). The internal reference protein may be spiked into the sample. The internal reference protein can be used to identify the mass spectra of individual endogenous proteins. The internal reference protein can be used as a standard for determining the amount of individual endogenous proteins. Proteomics measurements can be generated based on the amount of protein added to one sample of one or more biological fluid samples. Proteomics measurements can be generated based on the amount of labeled protein added to one sample of one or more biological fluid samples.

[0248] Transcriptomics data The data described herein, such as multi-omics data, may include transcription data or transcriptomics data. Transcriptomics data may include data on nucleotide transcripts such as RNA. Examples of RNA and Examples include messenger RNA (mRNA), ribosomal RNA (rRNA), signal recognition particle (SRP) RNA, transfer RNA (tRNA), nuclear small RNA (snRNA), nucleolar small RNA (snoRNA), long non-coding RNA (lncRNA), microRNA (miRNA), non-coding RNA (ncRNA), or piwi-interacting RNA (piRNA), or combinations thereof. RNA may include mRNA. RNA may also include miRNA. Transcriptomics data can be distinguished by subtypes, each subtype containing different types of RNA or transcripts. For example, mRNA data may be included in one subtype, while data of one or more types of small non-coding RNA, such as miRNA or piRNA, may be included in another subtype. miRNA may include 5p miRNA or 3p miRNA.

[0249] Transcriptomics data may include information about the presence, absence, or quantity of various RNAs. For example, transcriptomics data may include the quantity of RNA. RNA quantity can be expressed as a concentration or number of RNA molecules, for example, as the concentration of RNA in a biological fluid. RNA quantity can be compared to other RNAs or other biomolecules. Transcriptomics data may include information about the presence of RNA. Transcriptomics data may include information about the absence of RNA. The aspects described in relation to transcriptomics data may relate to transcript or RNA data, and vice versa.

[0250] Transcriptomics data generally includes data on a large number of RNAs. For example, transcriptomics data may include information on the presence, absence, or quantity of more than 1,000 RNAs. In some cases, transcriptomics data may include information on the presence, absence, or quantity of 5,000, 10,000, 20,000, or more RNAs. Transcriptomics data may further include up to approximately 200,000 transcripts. Transcriptomics data may include a set of transcripts defined by either the aforementioned number of RNAs or transcripts. Some examples of mRNAs that may be included in transcriptomics data are shown in Figure 10B or Figure 15. Some examples of microRNAs that may be included in transcriptomics data are shown in Figure 11B or Figure 15.

[0251] Some examples of mRNAs that can be used as biomarkers are shown in Figure 10B. mRNAs 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 included in Figure 10B can be used as biomarkers to determine, for example, whether a lung nodule is cancerous or to determine its likelihood of being cancerous. Some examples of microRNAs that can be used as biomarkers are shown in Figure 11B. MicroRNAs 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 included in Figure 11B can be used as biomarkers to determine, for example, whether a lung nodule is cancerous or to determine its likelihood of being cancerous.

[0252] Transcriptomics data can be generated by any of the following methods. Generation of transcriptomics data may include the use of detection reagents that bind to RNA and produce a detectable signal. After using detection reagents that bind to RNA and produce a detectable signal, readings indicating the presence, absence, or amount of RNA can be obtained. Generation of transcriptomics data may include sample concentration, filtration, or centrifugation.

[0253] Transcriptomics data may include RNA sequence data. Some examples of methods for generating transcripts include sequencing, microarray analysis, hybridization, polymerase chain reaction (PCR), or electrophoresis, or a combination thereof. Microarrays can be used to generate transcriptomics data. PCR can be used to generate transcriptomics data. PCR may include quantitative PCR (qPCR). Such methods may include the use of detectable probes (e.g., fluorescent probes) that intercalate into double-stranded nucleotides or bind to target nucleotide sequences. PCR may include reverse transcriptase quantitative PCR (RT-qPCR). The use of PCR panels may also be included in the generation of transcriptomics data.

[0254] RNA sequence data can be generated by sequencing the target RNA, or by first converting the target RNA to DNA (e.g., complementary DNA (cDNA)) and then sequencing that DNA. Sequencing may include large-scale parallel sequencing. Examples of large-scale parallel sequencing techniques include pyrosequencing, reversible terminator chemistry sequencing, ligation-mediated sequencing by ligase enzymes, or phospho-conjugated fluorescent nucleotide or real-time sequencing. The generation of transcriptomics data may include the preparation of samples or templates for sequencing. RNA can be converted to cDNA using reverse transcriptase. Some template preparation methods involve the use of amplification templates derived from single RNA or cDNA molecules, or single RNA or cDNA molecule templates. Examples of amplification methods include emulsion PCR, rolling circle, or solid-phase amplification.

[0255] In addition to any of the methods described above, generating transcriptomics data may also involve contacting a sample with particles so that the particles adsorb biomolecules, including RNA. The adsorbed RNA may be part of the biomolecular corona. The adsorbed RNA can be measured or identified when generating transcriptomics data.

[0256] Genomics data The data described herein, such as multi-omics data, may include data relating to genetic material or genomics data. Genomics data may include data relating to genetic material such as nucleic acids or histones. Nucleic acids may include DNA. Genomics data may include information regarding the presence, absence, or quantity of genetic material. The quantity of genetic material may be expressed as concentration, absolute number, or relative. The aspects described in relation to genomics data may relate to nucleic acid or DNA data, or vice versa. Nucleic acid data may include RNA data, or genomics data may include transcriptomics data.

[0257] Genomic data may include DNA sequence data. Sequence data may include gene sequences. For example, genomics data may include sequence data for up to approximately 20,000 genes. Genomics data may also include sequence data for non-coding DNA regions. DNA sequence data may include information about the presence, absence, or quantity of DNA sequences. DNA sequence data may include information about the presence or absence of mutations such as single nucleotide polymorphisms. DNA sequence data may include DNA measurements of the quantity of mutant DNA, for example, measurements of mutant DNA from cancer cells.

[0258] Genomics data may include epigenetic data. Examples of epigenetic data include DNA methylation data, DNA hydroxymethylation data, or histone modification data. Examples of epigenetic data include DNA methylation or hydroxymethylation. Methylation can be measured in the entire DNA or in a region. Methylated DNA may contain methylated cytosine (e.g., 5-methylcytosine). Cytosine is often methylated at CpG sites and may indicate gene activation.

[0259] Epigenetic data may include histone modification data. Histone modification data may include the presence, absence, or amount of histone modification. Examples of histone modifications include serotonination, methylation, citrullination, acetylation, or phosphorylation. Some specific examples of histone modifications include lysine methylation, glutamine serotonination, arginine methylation, arginine citrullination, lysine acetylation, serine phosphorylation, threonine phosphorylation, or tyrosine phosphorylation. Histone modifications may indicate gene activation.

[0260] Genomics data can be distinguished by subtypes, each subtype containing different types of genomics data. For example, DNA sequence data may be included in one subtype, epigenetic data in one subtype, or different types of epigenetic data in different subtypes.

[0261] Genomics data can be generated by any of the following methods. Generating genomics data may involve using detection reagents that bind to genetic material such as DNA or histones to produce a detectable signal. After using detection reagents that bind to genetic material and produce a detectable signal, readings indicating the presence, absence, or amount of genetic material can be obtained. Generating genomics data may also involve concentrating, filtering, or centrifuging the sample.

[0262] Some examples of methods for generating DNA sequence data include sequencing, microarray analysis (e.g., SNP microarrays), hybridization, polymerase chain reaction, or electrophoresis, or a combination thereof. DNA sequence data can be generated by sequencing the DNA of interest. Sequencing may include large-scale parallel sequencing. Examples of large-scale parallel sequencing techniques include pyrosequencing, sequencing by reversible terminator chemistry, sequencing by ligation mediated by ligase enzymes, or phospho-conjugated fluorescent nucleotide or real-time sequencing. Generating genomics data may include preparing samples or templates for sequencing. Some template preparation methods involve the use of amplification templates derived from single DNA molecules or single DNA molecule templates. Examples of amplification methods include emulsion PCR, rolling circle, or solid-phase amplification.

[0263] DNA methylation can be detected using mass spectrometry, methylation-specific PCR, bisulfite sequencing, HpaII tiny fragment enrichment by ligation-mediated PCR assay, GlaI hydrolysis and ligation adapter-dependent PCR assay, chromatin immunoprecipitation (ChIP) assay combined with DNA microarrays (ChIP-on-chip assay), restriction enzyme landmark genome scanning, methylated DNA immunoprecipitation, pyrosequencing of bisulfite-treated DNA, molecular breaklight assay of DNA adenine methyltransferase activity, methyl-sensitive Southern blotting, methyl CpG-binding protein assay, high-resolution melting analysis, methylation-sensitive single-nucleotide primer extension assay, other methylation assays, or combinations thereof.

[0264] Histone modification can be analyzed by mass spectrometry or immunoassay, enzyme-linked immunosorbent assay, etc. Detection can be performed using stamp blotting, dot blotting, immunohistochemistry, or a combination thereof.

[0265] In addition to any of the methods described above, generating genomics data may also involve contacting a sample with particles so that the particles adsorb biomolecules containing genetic material. The adsorbed genetic material may be part of the biomolecular corona. The adsorbed genetic material can be measured or identified when generating genomics data.

[0266] Figure 23 provides embodiments that may relate to transcription or genomics data. The data may include circulating free DNA (cfDNA) methylation, mRNA, miRNA, circulating free miRNA (cf-miRNA), or whole exome sequencing data. Any sample type, isolation method, quality control (QC) configuration, or sequencing depth provided in the figure may also be included. Any embodiment shown in the figure may be included in or used to generate data such as multi-omics data.

[0267] Lipidomics data The multi-omics data described herein may include lipid data or lipidomics data. Lipidomics data may include information regarding the presence, absence, or quantity of various lipids. For example, lipidomics data may include the quantity of lipids. The quantity of lipids can be expressed as the concentration or amount of lipids, for example, the concentration of lipids in biological fluids. The quantity of lipids can be compared to another lipid or another biomolecule. Lipidomics data may include information regarding the presence of lipids. Lipidomics data may include information regarding the absence of lipids. Lipid or lipidomics data may be included in metabolite or metabolomics data. Embodiments described in relation to lipidomics data may also relate to lipid data, and vice versa.

[0268] Many organisms contain complex lipids (for example, humans express over 600 types of lipids), and their relative expression can function as powerful markers for determining biological and health status. Lipids are a diverse class of biomolecules, including, among many other types, fatty acids (e.g., long-chain carbohydrates with carboxylate tails), diglycerides, triglycerides and polyglycerides, phospholipids, prenols, sterols (e.g., cholesterol), and ladderanes. While lipids are primarily found in membranes, free lipids, protein complex lipids, and nucleic acid complex lipids are typically present in various biological fluids and can sometimes be fractionated separately from membrane-bound lipids. For example, lipid-binding proteins (e.g., albumin) can be collected from a sample by immunohistochemical precipitation and chemically induced to release bound lipids for subsequent collection and detection.

[0269] Lipids may be essential components in the development of diseases such as cancer. For example, lipids may play an important role in cancer biology because they may influence or be involved in membrane and cell proliferation, lipotoxicity (where a balance of lipid content may help protect against lipotoxicity), enhancement of cellular processes, membrane biophysics, oncogenic signaling and metastasis, protection from oxidative stress, signaling in the microenvironment, or immunomodulation. Several lipid classes may be associated with cancer, such as glycerophospholipids in hepatocellular carcinoma, glycerophospholipids and acylcarnitines in prostate cancer, choline-containing lipids and phospholipids that increase during metastasis, or sphingolipids that control the survival and death of cancer cells.

[0270] Lipid data can be generated from a sample after processing the sample to isolate or concentrate lipids in the sample. The generation of lipid data involves sample concentration, filtration, and This may include centrifugation. Lipid analysis may include lipid fractionation. In many cases, lipids can be easily separated from other types of biomolecules for lipid-specific analysis. Because many lipids are highly hydrophobic, they can be cleanly separated from other types of biomolecules present in the sample by organic solvent extraction and gradient chromatography. Lipid data can be generated using mass spectrometry. Lipid analysis can then distinguish lipids by class (e.g., the distinction between sphingolipids and chlorolipids) or by individual type.

[0271] Lipidomic data can be generated by any of the following methods. Generation of lipidomics data may include the use of detection reagents that bind to lipids and produce a detectable signal. After using detection reagents that bind to lipids and produce a detectable signal, readings indicating the presence, absence, or amount of lipids can be obtained. Generation of lipidomics data may include sample concentration, filtration, or centrifugation.

[0272] Lipidomics data can be generated using mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assays, immunoassays, enzyme-linked immunosorbent assays, Western blotting, dot blotting, or immunostaining, or a combination thereof. One example of a method for generating lipidomics data involves using mass spectrometry. Mass spectrometry may include separation method steps such as liquid chromatography (e.g., HPLC). Mass spectrometry may include ionization methods, such as electron ionization, atmospheric pressure chemical ionization, electrospray ionization, or secondary electrospray ionization. Mass spectrometry may include surface-based mass spectrometry or secondary ion mass spectrometry. Another example of a method for generating lipidomics data involves nuclear magnetic resonance (NMR). Other examples of methods for generating lipidomics data include Fourier transform ion cyclotron resonance, ion mobility spectroscopy, electrochemical detection (e.g., coupled to HPLC), or Raman spectroscopy and radiolabeling (e.g., in combination with thin-layer chromatography). Some mass spectrometry methods described for generating lipidomics data can also be used to generate proteomics data, and vice versa. Lipidomic data can also be generated using immunoassays, e.g., enzyme-linked immunosorbent assays, Western blotting, dot blotting, or immunohistochemistry. The use of lipid panels may also be included in the generation of lipidomics data.

[0273] In addition to any of the methods described above, generating lipidomics data may also involve contacting a sample with particles so that the particles adsorb lipid-containing biomolecules. The adsorbed lipids may be part of the biomolecular corona. The adsorbed lipids can be measured or identified when generating lipidomics data.

[0274] The generation of lipidomics data may include the use of a known amount of internal reference lipid. The reference lipid may be labeled. The labeling may include isotopic labeling. The generation of lipidomics data may include the use of a known amount of isotopically labeled internal reference lipid. The internal reference lipid may be spiked into the sample. The internal reference lipid can be used to identify the mass spectra of individual endogenous lipids. The internal reference lipid can be used as a standard for determining the amount of individual endogenous lipids. Lipidomic measurements can be generated based on the amount of lipid added to one sample of one or more biological fluid samples. Lipidomic measurements can be generated based on the amount of labeled lipid added to one sample of one or more biological fluid samples.

[0275] Lipids may be associated with the biology of diseases such as cancer. Lipids may also include phospholipids. Examples of phospholipids include phosphatidylethanolamine (PE) and phosphatidylethanolamine. Examples include tidylcholine (PC), phosphatidylinositol (PI), or phosphatidylglycerol (PG). Some phospholipids are components of cell membranes and may play cellular roles such as chemical energy storage, cell signaling, cell membranes, or cell-to-tissue interactions. Lipids may also include ceramides (CERs). Ceramides may act as tumor suppressors and may be targeted therapeutic pathways. For example, the effectiveness of some chemotherapy and targeted therapies may depend on ceramide levels. Lipids may also include diacylglycerides (DAGs). Lipids may also include triacylglycerides (TAGs). Lipids may also include fatty acids (FAs).

[0276] Examples of lipids include PC(20:3_20:3)+AcO, Cer(d18:1 / 24:0)+H, GlcCer(d18:1 / 18:0+H), PI(18:0_18:3)-H, Aca(4:0)+H, GlcCer(d18:1 / 22:0+H), PC(18:2_20:5)+AcO, PC(14:0_18:2)+AcO, LPE(18:3)-H, Cer(d18:0 / 18:0)+H, DAG(18:1_22:6)+NH4, TAG(54:3_16:0)+NH4, Cer(d18:1 / Examples of lipids that may be included are 18:0)+H, PC(16:1_20:3)+AcO, LPC(17:0)+AcO, GlcCer(d18:1 / 24:1+H, DAG(18:1_20:2)+NH4, PE(P-18:0_18:2)+H, Cer(d18:0 / 24:0)+H, or PE(18:1_20:1)-H. Lipid data may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 measurements of these lipids.

[0277] Examples of lipids may include any of the lipids shown in Figure 27. Lipid data may include measurements of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 of these lipids, or any range of the aforementioned numbers of lipids from these figures.

[0278] Examples of lipids are shown in Figures 33A-33B. Lipids detected by the methods described herein may include CER(d18:1_10:0). Some examples of lipids are shown in Figure 36. Lipids detected by the methods described herein may include CER(d18.1_18.0), PC(18.2_20.5), CER(d18.1_24.1), CER(d18.1_16.0), TAG(56.5_FA18.0), CER(d18.0_24.1), TAG(56.5_FA18.1), DAG(16.0_22.5), CER(d18.1_22.1), PE(P-18.0_18.3), or PE(17.0_22.6). Any number of the aforementioned lipids can be used. Any of these lipids can be used in the classifier.

[0279] Lipid measurements may be affected (e.g., decreased) in samples from subjects with liver cancer compared to lipid measurements from control samples or compared to baseline measurements. Lipid measurements may include phospholipid measurements. Lipid measurements may be useful in evaluating liver cancer. Lipid measurements may include lipid or phospholipid measurements, or combinations of lipids or phospholipids, from Figure 39F or Figure 39G. Lipid measurements may be useful in evaluating ovarian cancer. Lipid measurements may include lipid or phospholipid measurements, or combinations of lipids or phospholipids, from Figure 39F or Figure 40f. Lipid measurements include the following lipids: LPC.14.0..AcO, LPC.15.0..AcO, LPC.16.0..AcO, LPC.16.1..AcO, LPC.17.0..AcO, LPC.18.0..AcO, LPC.18.1..AcO, LPC.18.2..AcO, LPC.18.3..AcO, LPC.20.2..AcO, LPC.20.3..AcO, LPC.20.4..AcO, LPE.18.0..H, LPE.18.2..H, LPE.20.4..H, PA.18.0_18.2..H, PC. 14.0_18.2..AcO, PC.14.0_18.3..AcO, PC.14.0_20.2..AcO, PC.14.0_20.3..AcO, PC.14.0_20.4..AcO, PC.14 .0_22.5..AcO, PC.14.0_22.6..AcO, PC.15.0_18.2..AcO, PC.15.0_20.3..AcO, PC.15.0_20.4..AcO, PC.16.0_ 20.3..AcO, PC.16.1_20.3..AcO, PC.16.1_20.4..AcO, PC.18.0_18.2..AcO, PC.18.0_20.3..AcO, PC.18.1_20. 3..AcO, PC.18.1_20.4..AcO, PC.18.1_22.4..AcO, PC.18.1_22.5..AcO, PC.18.2_18.2..AcO, PC.18.2_18.3.. AcO, PC.18.2_20.3..AcO, PC.18.2_20.4..AcO, PC.18.2_20.5..AcO, PC.20.2_20.3..AcO, PC.20.2_20.4..Ac O, PC.20.3_20.3..AcO, PC.20.3_20.4..AcO, PC.20.4_20.4..AcO, PC.20.4_22.5..AcO, PE.O.16.0_20.3..H, P The measurements may include one or more of the following values: E.O.16.0_20.4..H, PE.O.16.0_22.5..H, or PI.18.1_20.4..H, where "LPC" represents lysophosphatidylcholine, "LPE" represents lysophosphatidylethanolamine, "PA" represents phosphatidic acid, "PC" represents phosphatidylcholine, and "PE" represents phosphatidylethanolamine. The lipid or phospholipid combination may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 lipids in Figure 39F, or lipids within a range defined by any two of the aforementioned integers.A combination of lipids or phospholipids may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, or at least 45 lipids in Figure 39F. A combination of lipids or phospholipids may include less than 3, less than 4, less than 5, less than 6, less than 7, less than 8, less than 9, less than 10, less than 15, less than 20, less than 25, less than 30, less than 35, less than 40, less than 45, or less than 50 lipids in Figure 39F. In some embodiments, a combination of lipids may not include one or more lipids in either Figure 39F or Figure 39G. In some embodiments, a combination of lipids may not include one or more lipids in either Figure 39F or Figure 40f.

[0280] Any of the following lipids: LPC.14.0..AcO, LPC.15.0..AcO, LPC.16.0..AcO, LPC.16.1..AcO, LPC.17.0..AcO, LPC.18.0..AcO, LPC.18.1..AcO, LPC.18.2..AcO, LPC.18.3..AcO, LPC.20.2..AcO, LPC.20.3..AcO, or LPC.20.4..AcO may be useful in evaluating ovarian cancer.

[0281] Any of the following lipids: LPC.14.0..AcO, LPC.15.0..AcO, LPC.16.0..AcO, LPC.16.1..AcO, LPC.17.0..AcO, LPC.18.0..AcO, LPC.18.1..AcO, LPC.18.2..AcO, LPC.18.3..AcO, LPC.20.2..AcO, LPC.20.3..AcO, LPC.20.4..AcO, LPE.18.0..H, LPE.18.2..H, LPE.20.4..H, PA.1 8.0_18.2..H, PC.14.0_18.2..AcO, PC.14.0_18.3..AcO, PC.14.0_20.2..AcO, PC.14.0_20.3..AcO, PC.14.0_20.4..AcO, PC.14.0_ 22.5..AcO, PC.14.0_22.6..AcO, PC.15.0_18.2..AcO, PC.15.0_20.3..AcO, PC.15.0_20.4..AcO, PC.16.0_20.3..AcO, PC.16.1_20 .3..AcO, PC.16.1_20.4..AcO, PC.18.0_18.2..AcO, PC.18.0_20.3..AcO, PC.18.1_20.3..AcO, PC.18.1_20.4..AcO, PC.18.1 _22.4..AcO, PC.18.1_22.5..AcO, PC.18.2_18.2..AcO, PC.18.2_18.3..AcO, PC.18.2_20.3..AcO, PC.18.2_20.4..AcO, PC.1 8.2_20.5..AcO, PC.20.2_20.3..AcO, PC.20.2_20.4..AcO, PC.20.3_20.3..AcO, PC.20.3_20.4..AcO, PC.20.4_20.4..AcO, PC.20.4_22.5..AcO, PE.O.16.0_20.3..H, PE.O.16.0_20.4..H, PE.O.16.0_22.5..H, or PI.18.1_20.4..H may be useful in evaluating liver cancer.

[0282] Metabolomics data The data described herein, such as multi-omics data, may include metabolite data or metabolomics data. Metabolomics data may include information on small molecule (e.g., less than 1.5 kDa) metabolites (metabolic intermediates, hormones or other signaling molecules, or secondary metabolites). Metabolomics data may include data on metabolites. Metabolites may include substrates, intermediates, or products of metabolism. Metabolites may include small molecules. Metabolites can be any molecule with a size of less than 1.5 kDa. Examples of metabolites include sugars, lipids, amino acids, fatty acids, phenolic compounds, or alkaloids. Metabolomics data can be distinguished by subtype, each subtype containing different types of metabolite data. Metabolomics data may include some lipid data. Metabolomics data may include lipidomics data. The aspects described in relation to metabolomics data may relate to metabolite data, and vice versa. Metabolomics data may include metabolite measurements. Metabolite measurements may include measurements of lipids such as phospholipids.

[0283] Metabolomics data may include information about the presence, absence, or quantity of various metabolites. For example, metabolomics data may include the quantity of metabolites. The quantity of a metabolite can be expressed as the concentration or amount of the metabolite, for example, the concentration of the metabolite in a biological fluid. The quantity of a metabolite can be compared to another metabolite or another biomolecule. Metabolomics data may include information about the presence of a metabolite. Metabolomics data may include information about the absence of a metabolite.

[0284] Metabolomics data generally includes data on a large number of metabolites. For example, metabolomics data may include information on the presence, absence, or quantity of more than 1,000 metabolites. In some cases, metabolomics data may include information on the presence, absence, or quantity of 5,000, 10,000, 20,000, 50,000, 100,000, 500,000, 1 million, 1.5 million, 2 million, or more metabolites, or within a range defined by any two of the aforementioned numbers of metabolites.

[0285] Metabolomics data can be generated by any of the following methods. The generation of metabolomics data may include the use of detection reagents that bind to metabolites and produce detectable signals. After using detection reagents that bind to metabolites and produce detectable signals, readings indicating the presence, absence, or quantity of metabolites can be obtained. The generation of metabolomics data may include sample concentration, filtration, or centrifugation.

[0286] Metabolomics data can be obtained from mass spectrometry, chromatography, liquid chromatography, Metabolomics data can be generated using high-performance liquid chromatography, solid-phase chromatography, lateral flow assays, immunoassays, enzyme-linked immunosorbent assays, Western blotting, dot blotting, or immunostaining, or a combination thereof. One example of a method for generating metabolomics data involves using mass spectrometry. Mass spectrometry may include separation method steps such as liquid chromatography (e.g., HPLC). Mass spectrometry may include ionization methods, such as electron ionization, atmospheric pressure chemical ionization, electrospray ionization, or secondary electrospray ionization. Mass spectrometry may include surface-based mass spectrometry or secondary ion mass spectrometry. Another example of a method for generating metabolomics data involves nuclear magnetic resonance (NMR). Other examples of methods for generating metabolomics data include Fourier transform ion cyclotron resonance, ion mobility spectroscopy, electrochemical detection (e.g., coupled to HPLC), or Raman spectroscopy and radiolabeling (e.g., in combination with thin-layer chromatography). Some of the mass spectrometry methods described for generating metabolomics data can be used to generate proteomics data, and vice versa. Metabolomics data can also be generated using immunoassays, such as enzyme-linked immunosorbent assays, Western blotting, dot blotting, or immunohistochemistry. The generation of metabolomics data may also include the use of lipid panels.

[0287] In addition to any of the methods described above, generating metabolomics data may also involve contacting a sample with particles so that the particles adsorb biomolecules containing metabolites. The adsorbed metabolites may be part of the biomolecular corona. The adsorbed metabolites can be measured or identified when generating metabolomics data.

[0288] The generation of metabolomics data may include the use of a known amount of internal reference metabolite. The reference metabolite may be labeled. The labeling may include isotopic labeling. The generation of metabolomics data may include the use of a known amount of isotopically labeled internal reference metabolite. The internal reference metabolite may be spiked into the sample. The internal reference metabolite can be used to identify the mass spectra of individual endogenous metabolites. The internal reference metabolite can be used as a standard for determining the amount of individual endogenous metabolites. Metabolomics measurements can be generated based on the amount of metabolite added to one sample of one or more biological fluid samples. Metabolomics measurements can be generated based on the amount of labeled metabolite added to one sample of one or more biological fluid samples.

[0289] Examples of metabolites are shown in Figures 34A-34B. Metabolites detected by the methods described herein may include 5-aminoimidazole-4-carboxamidoliponucleotide (AICAR). Metabolites may also include nucleotides, such as monophosphate nucleotides. Some examples of metabolites are shown in Figure 36. Metabolites detected by the methods described herein may also include cytidine monophosphate (CMP). Metabolites may include AICAR or CMP. Detected metabolites may include AICAR and CMP. Any number of the aforementioned metabolites can be used. Any of the metabolites can be used.

[0290] Use of particles The sample can be brought into contact with the particles, for example, before generating data. The data described herein can be generated using particles. For example, the method may include bringing the sample into contact with the particles so that the particles adsorb biomolecules. The particles can attract a set of biomolecules different from those that would normally be accurately measured by performing omics measurements directly on the sample. For example, the main biomolecules may be certain types of biomolecules in the sample (e.g., proteins, transcripts, genes). A significant percentage of a substance (or metabolite) may be composed of a single protein. For example, the majority of circulating proteins collected by blood sampling may consist of just one protein. By attaching biomolecules to particles before analysis, a subset of biomolecules that do not contain the major biomolecules can be obtained. Removing the major biomolecules in this way may improve the accuracy of biomolecular measurements and the sensitivity of analyses using those measurements.

[0291] Examples of biomolecules that can adsorb to particles include proteins, transcripts, genetic material, or metabolites. Adsorbed biomolecules may form a biomolecular corona around the particle. Adsorbed biomolecules can be measured or identified when generating data such as omics data (e.g., proteomics data). In some embodiments, proteomics measurements are generated from proteins adsorbed onto nanoparticles. Nanoparticles can enrich proteins or other types of biomolecules.

[0292] Particles can be made from a variety of materials. Such materials may include metals, magnetic particles, polymers, or lipids. Particles can be made from combinations of materials. Particles may contain layers of different materials. Different materials may have different properties. Particles may contain a core made of one material and be coated with another material. The core and coating may have different properties.

[0293] The particles may contain metals. For example, the particles may contain gold, silver, copper, nickel, cobalt, palladium, platinum, iridium, osmium, rhodium, ruthenium, rhenium, vanadium, chromium, manganese, niobium, molybdenum, tungsten, tantalum, iron, or cadmium, or a combination thereof.

[0294] The particles may be magnetic (e.g., ferromagnetic or ferrimagnetic). Particles containing iron oxide may be magnetic. The particles may contain superparamagnetic iron oxide nanoparticles (SPION).

[0295] The particles may contain polymers. Examples of polymers include polyethylene, polycarbonate, polyanhydride, polyhydroxy acid, polypropylfumerates, polycaprolactone, polyamide, polyacetal, polyether, polyester, poly(orthoester), polycyanoacrylate, polyvinyl alcohol, polyurethane, polyphosphazene, polyacrylate, polymethacrylate, polycyanoacrylate, polyurea, polystyrene, or polyamine, polyalkylene glycol (e.g., polyethylene glycol (PEG)), polyester (e.g., poly(lactide-co-glycolide) (PLGA), polylactic acid, or polycaprolactone), or copolymers of two or more polymers, for example, a copolymer of polyalkylene glycol (e.g., PEG) and polyester (e.g., PLGA). The particles can be made from combinations of polymers.

[0296] The particles may contain lipids. Examples of lipids include dioleoylphosphatidylglycerol (DOPG), diacylphosphatidylcholine, diacylphosphatidylethanolamine, ceramide, sphingomyelin, cephalin, cholesterol, cerebroside and diacylglycerol, dioleoylphosphatidylcholine (DOPC), dimyristoylphosphatidylcholine (DMPC), and dioleoylphosphatidylserine (DOPS), phosphatidylglycerol, cardiolipin, diacylphosphatidylserine, diacylphosphatidic acid, N-dodecanoylphosphatidylethanolamine, N-succinylphosphatidylethanolamine, and N-glutarylphosphatidylethanolamine. Luamine, Lysylphosphatidylglycerol, Palmitoyloleoylphosphatidylglycerol (POPG), Lecithin, Lysolecithin, Phosphatidylethanolamine, Lysophosphatidylethanolamine, Dioleoylphosphatidylethanolamine (DOPE), Dipalmitoylphosphatidylethanolamine (DPPE), Dimyristoylphosphoethanolamine (DMPE), Distearoylphosphatidylethanolamine (DSPE), Palmitoyloleoylphosphatidylethanolamine (PO Examples include PE, palmitoyl oleoyl phosphatidylcholine (POPC), egg yolk phosphatidylcholine (EPC), distearoyl phosphatidylcholine (DSPC), dioleoyl phosphatidylcholine (DOPC), dipalmitoyl phosphatidylcholine (DPPC), dioleoyl phosphatidylglycerol (DOPG), dipalmitoyl phosphatidylglycerol (DPPG), palmitoyl oleoyl phosphatidylglycerol (POPG), 16-O-monomethyl PE, 16-O-dimethyl PE, 18-1-trans PE, palmitoyl oleoyl phosphatidylethanolamine (POPE), 1-stearoyl-2-oleoyl phosphatidylethanolamine (SOPE), phosphatidylserine, phosphatidylinositol, sphingomyelin, cephalin, cardiolipin, phosphatidic acid, cerebroside, dicetyl phosphate, or cholesterol. The particles can be made from a combination of lipids.

[0297] Further examples of materials include silica, carbon, carboxylate, polyacrylic acid, carbohydrates, dextran, polystyrene, dimethylamine, amine, or silane. Some examples of particles include carboxylate SPION, phenol-formaldehyde coated SPION, silica coated SPION, polystyrene coated SPION, carboxylated poly(styrene-co-methacrylic acid), P(St-co-MAA) coated SPION, N-(3-trimethoxysilylpropyl)diethylenetriamine coated SPION, poly(N-(3-(dimethylamino)propyl)methacrylamide)(PDMAPMA) coated SPION, 1,2,4,5-benzenetetracarboxylic acid coated SPION, poly(vinylbenzyltrimethylammonium chloride)(PVBTMAC) coated SPION, carboxylate coated with peracetic acid, poly(oligo(ethylene glycol)methyl ether methacrylate)(POEGMA) coated SPION, polystyrene carboxyl-functionalized particles, carboxylic acid particles, particles with amino surfaces, silica amino-functionalized particles, particles with Jeffamine surfaces, or silica silanol-coated particles.

[0298] Particles of various sizes can be used. The particles may include nanoparticles. The nanoparticles may have a diameter of approximately 10 nm to approximately 1000 nm. For example, the nanoparticles may have a diameter of at least 10 nm, at least 100 nm, at least 200 nm, at least 300 nm, at least 400 nm, at least 500 nm, at least 600 nm, at least 700 nm, at least 800 nm, at least 900 nm, 10 nm to 50 nm, 50 nm to 100 nm, 100 nm to 150 nm, 150 nm to 200 nm, 200 nm to 250 nm, 250 nm to 300 nm, 300 nm to 350 nm, 350 nm to 400 nm, 400 nm to 450 nm, 450 nm to 500 nm, 500 nm to 550 nm, and 550 nm. The nanoparticles may be in the following ranges: m~600nm, 600nm~650nm, 650nm~700nm, 700nm~750nm, 750nm~800nm, 800nm~850nm, 850nm~900nm, 100nm~300nm, 150nm~350nm, 200nm~400nm, 250nm~450nm, 300nm~500nm, 350nm~550nm, 400nm~600nm, 450nm~650nm, 500nm~700nm, 550nm~750nm, 600nm~800nm, 650nm~850nm, 700nm~900nm, or 10nm~900nm. The diameter of the nanoparticles may be less than 1000nm. Some examples include approximately 50nm, 130nm, 150nm, and 400-600nm. This includes diameters of m, or 100-390 nm.

[0299] The particles may include fine particles. These fine particles may have a diameter of approximately 1 μm to approximately 1000 μm. For example, fine particles may have diameters of at least 1 μm, at least 10 μm, at least 100 μm, at least 200 μm, at least 300 μm, at least 400 μm, at least 500 μm, at least 600 μm, at least 700 μm, at least 800 μm, at least 900 μm, 10 μm to 50 μm, 50 μm to 100 μm, 100 μm to 150 μm, 150 μm to 200 μm, 200 μm to 250 μm, 250 μm to 300 μm, 300 μm to 350 μm, 350 μm to 400 μm, 400 μm to 450 μm, 450 μm to 500 μm, 500 μm to 550 μm, 550μm~600μm, 600μm~650μm, 650μm~700μm, 700μm~750μm, 750μm~800μm, 800μm m~850μm, 850μm~900μm, 100μm~300μm, 150μm~350μm, 200μm~400μm, 250μm~450 The particle size can be μm, 300μm-500μm, 350μm-550μm, 400μm-600μm, 450μm-650μm, 500μm-700μm, 550μm-750μm, 600μm-800μm, 650μm-850μm, 700μm-900μm, or 10μm-900μm. The diameter of the particles may be less than 1000μm. Some examples include diameters of 2.0-2.9μm.

[0300] The particles may include sets of physiologically and chemically distinct particles (for example, two or more sets of physiologically and chemically distinct particles, where one set of particles is physiologically and chemically different from the particles in another set). Examples of physiological and chemical properties include electric charge (e.g., positive, negative, or neutral) or hydrophobicity (e.g., hydrophobic or hydrophilic). The particles may include sets of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more particles, or a range of particle sets containing any of the aforementioned number of particle sets.

[0301] Particles and types Disease detection methods may include the use of particles. Methods described herein may include contacting a biological sample with physiologically and chemically distinct particles to form a biomolecular corona. The biological sample may be from an object identified as having a pulmonary nodule. The particles can adsorb biomolecules from the biological sample, thereby forming a biomolecular corona on the surface of the particles. Upon contact with the biological sample, the particles can adsorb multiple peptides, proteins, nucleic acids, lipids, sugars, small molecules (such as metabolites (natural and exogenous), terpenes, polyketides, and cyclic peptides), or any combination thereof. Therefore, the method may include collecting a subset of biomolecules from a biological sample (e.g., a complex biological sample such as human plasma) onto particles, analyzing the biomolecules collected on the particles, analyzing the biomolecules remaining in the biological sample, or analyzing the biomolecules collected on the particles and the biomolecules remaining in the biological sample. Biomolecules, biomolecular coronas, or parts thereof, can be eluted from the particles into solution before analysis. In some embodiments, assaying proteins involves contacting a biological fluid sample with particles so that the particles adsorb proteins onto the particles.

[0302] The relationship between particle properties and biomolecular corona composition can be used to manipulate the collection of biomolecules from a sample. In some cases, a set of particle properties may favor the binding of a particular biomolecular type, family, or superfamily. For example, humans express over 100 proteins from the Ras superfamily, which share a conserved GTP-binding motif within the N-terminal domain of 20 kilodaltons (kDa). Particles or aggregates of particles (e.g., a mixture containing five types of particles) can be functionalized to promote the adsorption of Ras proteins and thus can be engineered to preferentially adsorb Ras proteins from complex biological samples for further analysis. In some cases, concentration of these substances may be possible.

[0303] Particles or mixtures of different particles can be prepared to perform broad profiling of a sample. In many biological samples, a small number of biomolecules constitute the majority of the biomaterial. For example, over 99% of the protein mass in human plasma is accounted for by only 20 out of approximately 3500 human plasma proteins. Analyzing such samples can be extremely difficult because detection or enrichment schemes may be saturated by a small number of abundant biomolecules. Particles or aggregates of multiple particle types can be prepared to broadly profile complex biological materials so that less abundant biomolecules from complex biological samples are enriched preferentially over, or together with, more abundant biomolecules. Because particles or aggregates of multiple particle types have similar binding affinities to a large number of biomolecules, they promote the adsorption of a large number of biomolecules from the sample. Particles may have low affinity for abundant proteins or sets of abundant proteins in the sample, so they may preferentially adsorb and enrich less abundant biomolecules. Particle aggregates can include particle types with affinities to different types or classes of biomolecules so that the particle aggregate adsorbs a wide range of biomolecules from the sample. Therefore, this disclosure provides a wide range of particle types having different physicochemical properties.

[0304] Particle types consistent with the methods disclosed herein can be made from a variety of materials. For example, particle materials consistent with this disclosure include metals, polymers, magnetic materials, and lipids. Magnetic particles may be iron oxide particles. Examples of metallic materials include gold, silver, copper, nickel, cobalt, palladium, platinum, iridium, osmium, rhodium, ruthenium, rhenium, vanadium, chromium, manganese, niobium, molybdenum, tungsten, tantalum, iron, and cadmium, or any one or any combination thereof of any other materials described in US7749299, the contents of which are incorporated herein by reference in their entirety. The particles may be magnetic (e.g., ferromagnetic or ferrimagnetic). For example, the particles may include superparamagnetic iron oxide nanoparticles (SPION).

[0305] The particles may comprise multiple physiologically and chemically distinct particles (for example, two or more sets of physiologically and chemically distinct particles, where one set of particles is physiologically and chemically distinct from another set of particles). In some embodiments, the particles comprise nanoparticles. In some embodiments, the particles comprise a group of physiologically and chemically distinct nanoparticles. The physiologically and chemically distinct particles may comprise lipid particles, metal particles, silica particles, or polymer particles. The physiologically and chemically distinct particles may comprise carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles.

[0306] The particles may contain polymers. Examples of polymers include polyethylene, polycarbonate, polyanhydride, polyhydroxy acid, polypropylfumerates, polycaprolactone, polyamide, polyacetal, polyether, polyester, poly(orthoester), polycyanoacrylate, polyvinyl alcohol, polyurethane, polyphosphazene, polyacrylate, polymethacrylate, polycyanoacrylate, polyurea, polystyrene, or polyamine, polyalkylene glycol (e.g., polyethylene glycol (PEG)), polyester (e.g., poly(lactide-co-glycolide) (PLGA), polylactic acid, or polycaprolactone), or copolymers of two or more polymers, for example, a copolymer of polyalkylene glycol (e.g., PEG) and polyester (e.g., PLGA), or any combination thereof. The polymer may be lipid-terminated polyalkylene glycol and polyester, or any other material disclosed in US9549901. It is also possible to do so, and its contents are incorporated herein by reference in their entirety.

[0307] The particles may contain lipids. Examples of lipids that can be used to form the particles of this disclosure include cationic, anionic, and neutrally charged lipids. For example, the particles include dioleoyl phosphatidylglycerol (DOPG), diacylphosphatidylcholine, diacylphosphatidylethanolamine, ceramide, sphingomyelin, cephalin, cholesterol, cerebroside and diacylglycerol, dioleoyl phosphatidylcholine (DOPC), dimyristoyl phosphatidylcholine (DMPC), and dioleoyl phosphatidylserine (DOPS), phosphatidylglycerol, cardiolipin, diacylphosphatidylserine, diacylphosphatidic acid, N-dodecanoylphosphatidylethanolamine, N-succinylphosphatidylethanolamine, N-glutarylphosphatidylethanolamine, lysylphosphatidylglycerol, palmitoyloleoylphosphatidylglycerol (POPG), lecithin, lysolecithin, phosphatidylethanolamine, lysophospha Tidylethanolamine, dioleoylphosphatidylethanolamine (DOPE), dipalmitoylphosphatidylethanolamine (DPPE), dimyristoylphosphoethanolamine (DMPE), distearoylphosphatidylethanolamine (DSPE), palmitoyloleoylphosphatidylethanolamine (POPE), palmitoyloleoylphosphatidylcholine (POPC), egg yolk phosphatidylcholine (EPC), distearoylphosphatidylcholine (DSPC), dioleoylphosphatidylcholine (DOPC), dipalmitoylphosphatidylcholine (DPPC), dioleoylphosphatidylglycerol (DOPG), dipalmitoylphosphatidylglycerol (DPPG), palmitoyloleoylphosphatidylglycerol (POPG), 16-O-monomethylPE, 16-O-dimethylPE, 18-1-transPE, palmitoyl oleoyl-phosphatidylethanolamine (POPE), 1-stearoyl-2-oleoyl-phosphatidylethanolamine (SOPE), phosphatidylserine, phosphatidylinositol, sphingomyelin, cephalin, cardiolipin, phosphatidic acid, cerebroside, dicetyl phosphate, and cholesterol, or their contents can be made from any one or any combination thereof of any other materials disclosed in US9445994, which is incorporated herein by reference in whole. Examples of particles of this disclosure are provided in Table 1. [Table 1]

[0308] Examples of particle types in this disclosure include carboxylate (citrate) superparamagnetic iron oxide nanoparticles (SPION), phenol formaldehyde coated SPION, silica coated SPION, polystyrene coated SPION, carboxylated poly(styrene-co-methacrylic acid) coated SPION, N-(3-trimethoxysilylpropyl)diethylenetriamine coated SPION, poly(N-(3-(dimethylamino)propyl)methacrylamide) (PDMAPMA) coated SPION, 1,2,4,5-benzenetetracarboxylic acid coated SPION, poly(vinylbenzyltrimethylammonium chloride) (PVBTMAC) coated SPION, carboxylate, PAA coated SPION, poly(oligo(ethylene glycol)methyl ether methacrylate) (POEGMA) coated SPION, carboxylate microparticles, polystyrene carboxyl-functionalized particles, carboxylic acid coated particles, and silica particles. These particles may be carboxylic acid particles with a diameter of approximately 150 nm, amino surface microparticles with a diameter of approximately 0.4 to 0.6 μm, silica amino-functionalized microparticles with a diameter of approximately 0.1 to 0.39 μm, Jeffamine surface particles with a diameter of approximately 0.1 to 0.39 μm, polystyrene microparticles with a diameter of approximately 2.0 to 2.9 μm, silica particles, carboxylated particles with an original coating with a diameter of approximately 50 nm, particles coated with a dextran-based coating with a diameter of approximately 0.13 μm, or silica silanol-coated particles with low acidity.

[0309] Particles consistent with the present disclosure can be produced and used in a method for forming protein coronas of a wide range of sizes after incubation in a biological fluid. The particles of the present disclosure may be nanoparticles. The nanoparticles of the present disclosure may have a diameter of about 10 nm to about 1000 nm. For example, the nanoparticles of the present disclosure may have a diameter of at least 10 nm, at least 100 nm, at least 200 nm, at least 300 nm, at least 400 nm, at least 500 nm, at least 600 nm, at least 700 nm, at least 800 nm, at least 900 nm, 10 nm to 50 nm, 50 nm to 100 nm, 100 nm to 150 nm, 150 nm to 200 nm, 200 nm to 250 nm, 250 nm to 300 nm, 300 nm to 350 nm, 350 nm to 400 nm, 400 nm to 450 nm, 450 nm to 500 nm, 500 nm to 550 nm, and 55 The nanoparticles may be in the following ranges: 0nm-600nm, 600nm-650nm, 650nm-700nm, 700nm-750nm, 750nm-800nm, 800nm-850nm, 850nm-900nm, 100nm-300nm, 150nm-350nm, 200nm-400nm, 250nm-450nm, 300nm-500nm, 350nm-550nm, 400nm-600nm, 450nm-650nm, 500nm-700nm, 550nm-750nm, 600nm-800nm, 650nm-850nm, 700nm-900nm, or 10nm-900nm. The diameter of the nanoparticles may be less than 1000nm.

[0310] The particles of the present disclosure may be microparticles. The microparticles can be particles having a diameter of about 1 μm to about 1000 μm. For example, the microparticles of the present disclosure have a diameter of at least 1 μm, at least 10 μm, at least 100 μm, at least 200 μm, at least 300 μm, at least 400 μm, at least 500 μm, at least 600 μm, at least 700 μm, at least 800 μm, at least 900 μm, 10 μm to 50 μm, 50 μm to 100 μm, 100 μm to 150 μm, 150 μm to 200 μm, 200 μm to 250 μm, 250 μm to 300 μm, 300 μm to 350 μm, 350 μm to 400 μm, 400 μm to 450 μm, 450 μm to 500 μm, 500 μm to 550 μm, 550 μm to 600 μm, 600 μm to 650 μm, 650 μm to 700 μm, 700 μm to 750 μm, 750 μm to 800 μm, 800 μm to 850 μm, 850 μm to 900 μm, 100 μm to 300 μm, 150 μm to 350 μm, 200 μm to 400 μm, 250 μm to 450 μm, 300 μm to 500 μm, 350 μm to 550 μm, 400 μm to 600 μm, 450 μm to 650 μm, 500 μm to 700 μm, 550 μm to 750 μm, 600 μm to 800 μm, 650 μm to 850 μm, 700 μm to 900 μm, or 10 μm to 900 μm. The diameter of the microparticles may be less than 1000 μm.

[0311] The ratio of surface area to mass can be a determining factor for the properties of the particles. For example, the number and type of biomolecules adsorbed by the particles from a solution vary depending on the surface area to mass ratio of the particles. The particles disclosed herein have a surface area to mass ratio of 3 to 30 cm 2 / mg, 5 to 50 cm 2 / mg, 10 to 60 cm 2 / mg, 15 to 70 cm 2 / mg, 20 to 80 cm 2 / mg, 30 to 100 cm 2 / mg, 35 to 120 cm 2 / mg, 40 to 130 cm 2 / mg, 45 to 150 cm 2 / mg, 50 to 160 cm 2 / mg, 60 to 180 cm 2 / mg, 70 to 200 cm2 / mg, 80-220cm 2 / mg, 90-240cm 2 / mg, 100-270cm 2 / mg, 120-300cm 2 / mg, 200-500cm 2 / mg, 10-300cm 2 / mg, 1-3000cm 2 / mg, 20-150cm 2 / mg, 25-120cm 2 / mg, or 4 0-85cm 2 A surface area-to-mass ratio of / mg may be present. Smaller particles (e.g., with a diameter of 50 nm or less) can have a significantly larger surface area-to-mass ratio, which is partly due to a higher-order dependence of diameter on mass rather than surface area. In some cases (e.g., for small particles), particles can have a surface area of ​​200-1000 cm². 2 / mg, 500-2000cm 2 / mg, 1000-4000cm 2 / mg, 2000-8000cm 2 / mg, or 4000-10000cm 2 It may have a surface area-to-mass ratio of / mg. In some cases (for example, for large particles), the particles may be 1-3 cm 2 / mg, 0.5~2cm 2 / mg, 0.25~1.5cm 2 / mg, or 0.1-1cm 2 It may have a surface area-to-mass ratio of / mg.

[0312] In some cases, the multiple particles used in the method described herein (e.g., of a particle panel) may have a surface area-to-mass ratio within a certain range. In some cases, the range of the surface area-to-mass ratio of the multiple particles may be 100 cm². 2 / mg, 80cm 2 / mg, 60cm 2 / mg, 40cm 2 / mg, 20cm 2 / mg, 10cm 2 / mg, 5cm 2 / mg, or 2cm 2It is less than / mg. In some cases, the surface area-to-mass ratio of multiple particles varies among multiple particles by as much as 40%, 30%, 20%, 10%, 5%, 3%, 2%, or less than 1%. In some cases, multiple particles may contain at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or more particles of different types.

[0313] In some cases, multiple particles (e.g., of a particle panel) may have a wider range of surface area-to-mass ratios. In some cases, the range of surface area-to-mass ratios for multiple particles may be 100 cm². 2 > / mg, 150cm 2 / mg, 200cm 2 / mg, 250cm 2 / mg, 300cm 2 / mg, 400cm 2 / mg, 500cm 2 / mg, 800cm 2 / mg, 1000cm 2 / mg, 1200cm 2 / mg, 1500cm 2 / mg, 2000cm 2 / mg, 3000cm 2 / mg, 5000cm 2 / mg, 7500cm 2 / mg, 10000cm 2 / mg or more. In some cases, the surface area-to-mass ratio of multiple particles (e.g., within a panel) can vary by more than 100%, 200%, 300%, 400%, 500%, 1000%, 10000%, or more. In some cases, multiple particles with a wide range of surface area-to-mass ratios may include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or more different types of particles.

[0314] Surface functional groups may include polymerizable functional groups, positively or negatively charged functional groups, amphoteric functional groups, acidic or basic functional groups, polar functional groups, or any combination thereof. Surface functional groups may include carboxyl groups, hydroxyl groups, thiol groups, cyano groups, nitro groups, ammonium groups, alkyl groups, imidazolium groups, sulfonium groups, pyridinium groups, pyrrolidinium groups, phosphonium groups, aminopropyl groups, amine groups, boronic acid groups, N-succinimidyl ester groups, PEG groups, streptavidin, methyl ether groups, triethoxylpropylaminosilane groups, PCP groups, citrate groups, lipoic acid groups, BPEI groups, or any combination thereof. The particles among the multiple particles may be selected from the following group: micelles, liposomes, iron oxide particles, silver particles, gold particles, palladium particles, quantum dots, platinum particles, titanium particles, silica particles, metal or inorganic oxide particles, synthetic polymer particles, copolymer particles, terpolymer particles, polymer particles with metal cores, polymer particles with metal oxide cores, polystyrene sulfonate particles, polyethylene oxide particles, polyoxyethylene glycol particles, polyethyleneimine particles, polylactic acid particles, polycaprolactone particles, polyglycolic acid particles, poly(lactide-co-glycolide polymer particles), cellulose ether polymer particles, polyvinylpyrrolidone particles, polyvinyl acetate particles, polyvinylpyrrolidone-vinyl acetate copolymer particles, polyvinyl alcohol particles, acrylate particles, polyacrylic acid particles, crotonic acid copolymer particles, polyethylene phosphonate particles, polyalkylene particles, carboxyvinyl polymer particles, sodium alginate particles, carrageenan particles, xanthan gum particles, acacia gum particles, gum arabic particles, guar gum particles, plutonate Agar particles, agar particles, chitin particles, chitosan particles, pectin particles, karaya tum particles, locust bean gum particles, maltodextrin particles, amylose particles, corn starch particles, potato starch particles, rice starch particles, tapioca starch particles, pea starch particles, sweet potato starch particles, barley starch particles, wheat starch particles, hydroxypropylated high-amylose starch particles, dextrin particles, levan particles, elcinan particles, gluten particles, collagen particles, isolated whey protein particles, casein particles, milk protein particles, soy protein particles, keratin particles, polyethylene particles, polycarbonate particles, polyanhydrous particles, polyhydroxy acid particles, polypropylene fumarate particles, polycaprolactone particles, polyamine particles, polyacetal particles, polyether particles, poly Rieser particles, poly(orthoester) particles, polycyanoacrylate particles, polyurethane particles, polyphosphazene particles, polyacrylate particles, polymethacrylate particles, polycyanoacrylate particles, polyurea particles, polyamine particles, polystyrene particles, poly(lysine) particles, chitosan particles, dextran particles, poly(acrylamide) particles, derivatized poly(acrylamide) particles, gelatin particles, starch particles, chitosan particles, dextran particles, gelatin particles, starch particles, poly-β-aminoester particles, poly(amidoamine) particles, polylactic acid-co-glycolic acid particles, polyanhydride particles, bioreducing polymer particles, and 2-(3-aminopropylamino)ethanol particles, as well as any combination thereof.

[0315] Multiple particles (e.g., physicochemically different particles) include carboxylate (citrate) superparamagnetic iron oxide nanoparticles (SPION), phenol formaldehyde coated SPION, silica coated SPION, polystyrene coated SPION, carboxylated poly(styrene-co-methacrylic acid) coated SPION, N-(3-trimethoxysilylpropyl)diethylenetriamine coated SPION, poly(N-(3-(dimethylamino)propyl)methacrylamide)(PDMAPMA) coated SPION, 1,2,4,5-benzenetetracarboxylic acid coated SPION, and poly(vinylbenzyltrimethylamide). The material may contain one or more particle types selected from the group consisting of ammonium chloride (PVBTMAC) coated SPION, carboxylate, PAA coated SPION, poly(oligo(ethylene glycol) methyl ether methacrylate) (POEGMA) coated SPION, carboxylate fine particles, polystyrene carboxyl-functionalized particles, carboxylic acid coated particles, silica particles, carboxylic acid particles, amino surface particles, silica amino-functionalized particles, Jeffamine surface particles, polystyrene particles, particles coated with a dextran-based coating with a diameter of approximately 0.13 μm, or silica silanol coated particles.

[0316] Multiple particles (e.g., physicochemically different particles) include carboxylate (citrate) superparamagnetic iron oxide nanoparticles (SPION), phenol formaldehyde coated SPION, silica coated SPION, polystyrene coated SPION, carboxylated poly(styrene-co-methacrylic acid) coated SPION, N-(3-trimethoxysilylpropyl)diethylenetriamine coated SPION, poly(N-(3-(dimethylamino)propyl)methacrylamide)(PDMAPMA) coated SPION, 1,2,4,5-benzenetetracarboxylic acid coated SPION, and poly(vinylbenzyltrimethylamide). The material may contain one or more particle types selected from the group consisting of ammonium chloride (PVBTMAC) coated SPION, carboxylate, PAA coated SPION, poly(oligo(ethylene glycol) methyl ether methacrylate) (POEGMA) coated SPION, carboxylate fine particles, polystyrene carboxyl-functionalized particles, carboxylic acid coated particles, silica particles, carboxylic acid particles, amino surface particles, silica amino-functionalized particles, Jeffamine surface particles, polystyrene particles, particles coated with a dextran-based coating with a diameter of approximately 0.13 μm, or silica silanol coated particles.

[0317] Multiple particles (e.g., physicochemically distinct particles) may include one or more particle types selected from the group consisting of silica particles, poly(acrylamide) particles, polyethylene glycol particles, or combinations thereof. One or more of the particles may include a paramagnetic or superparamagnetic core material. The particles may include silica particles. The particles may include poly(acrylamide) particles. The particles may include polyethylene glycol particles.

[0318] Multiple particles may include multiple particle types. In some cases, multiple particles include at least two types of particles. In some cases, multiple particles include at least three types of particles. In some cases, multiple particles include at least five types of particles. In some cases, multiple particles include at least six types of particles. In some cases, multiple particles include at least eight types of particles. In some cases, multiple particles include at least ten types of particles. In some cases, multiple particles include at least twelve types of particles. In some cases, multiple particles include at least fifteen types of particles. In some cases, multiple particles include at least eighteen types of particles. In some cases, multiple particles include at least twenty types of particles.

[0319] The particles may include layers having different properties. The particles may include a core having a first set of properties and a shell having a second set of properties. The particles may include multiple shells having different properties (e.g., a core containing a first material, an inner shell containing a second material, and an outer shell containing a third material). The layers of the particles may contain multiple materials. For example, the layers of particles may contain multiple polymers. The polymers may be homogeneously dispersed within the layer, phase-separated, or applied heterogeneously.

[0320] In some cases, one or more physicochemical properties are selected from the group consisting of composition, size, surface charge, hydrophobicity, hydrophilicity, surface functional groups, surface topography, surface curvature, shape, and any combination thereof. In some embodiments, the surface functional groups include chemical functionalization. In some embodiments, small molecule functionalization includes amine functionalization, carboxylate functionalization, monosaccharide functionalization, oligosaccharide functionalization, sugar phosphate functionalization, sugar sulfuric acid functionalization, alcohol functionalization, ether functionalization, ester functionalization, amide functionalization, carbonate functionalization, carbamate functionalization, urea functionalization, benzyl functionalization, phenyl functionalization, phenol functionalization, aniline functionalization, imidazole functionalization, indole functionalization, fluoride functionalization, chloride functionalization, bromide functionalization, sulfide functionalization, nitro functionalization, thiol functionalization, nitrogen base functionalization, aminopropyl functionalization, boronic acid functionalization, N-succinimidyl ester functionalization, PEG functionalization, methyl ether functionalization, triethoxylpropylaminosilane functionalization, silicon alkoxide functionalization, phenol-formaldehyde functionalization, organosilane functionalization, ethylene glycol functionalization, PCP functionalization, citrate functionalization, lipoic acid functionalization, or any combination thereof. In some embodiments, small molecule functionalization includes silica-functionalized particles, amine-functionalized particles, silicon alkoxide-functionalized particles, polystyrene-functionalized particles, and sugar-functionalized particles. In some embodiments, small molecule functionalization includes amine functionalization, sugar phosphate functionalization, carboxylate functionalization, silica functionalization, organosilane functionalization, or any combination thereof. In some embodiments, small molecule functionalization includes silica functionalization, ethylene glycol functionalization, and amine functionalization, or any combination thereof.

[0321] The particles of this disclosure may be synthesized or purchased from commercial vendors. For example, particles consistent with this disclosure include Sigma-Aldrich and Life. These can be purchased from commercial vendors, including Technologies, Fisher Biosciences, nanoComposix, Nanopartz, Spherotech, and other commercial vendors. The preferred particles of this disclosure can be purchased from commercial vendors and further modified, coated, or functionalized.

[0322] This disclosure includes compositions and methods comprising two or more particles from a group of particles having at least one different physicochemical property. Such compositions and methods may comprise at least two to at least 20 particles from a group of particles having at least one different physicochemical property. Such compositions and methods may comprise at least three to at least six particles from a group of particles having at least one different physicochemical property. Such compositions and methods may comprise at least four to at least eight particles from a group of particles having at least one different physicochemical property. Such compositions and methods may comprise at least four to at least ten particles from a group of particles having at least one different physicochemical property. Such compositions and methods may comprise at least five to at least twelve particles from a group of particles having at least one different physicochemical property. Such compositions and methods may comprise at least six to at least fourteen particles from a group of particles having at least one different physicochemical property. Such compositions and methods may comprise at least eight to at least fifteen particles from a group of particles having at least one different physicochemical property. Such compositions and methods may comprise at least ten to at least 20 particles from a group of particles having at least one different physicochemical property. Such compositions and methods may include at least two different particle types, at least three different particle types, at least four different particle types, at least five different particle types, at least six different particle types, at least seven different particle types, at least eight different particle types, at least nine different particle types, at least ten different particle types, at least eleven different particle types, at least twelve different particle types, at least thirteen different particle types, at least fourteen different particle types, at least fifteen different particle types, at least twenty different particle types, at least twenty-five different particle types, or at least thirty different particle types.

[0323] The particles of this disclosure can be brought into contact with a biological sample (e.g., a biological fluid) to form a biomolecular corona. When in contact with a complex biological sample, one or more types of particles may adsorb more than 100 types of proteins (for example, in a 100 μl aliquot of a biological sample containing 100 pM of a certain type of particle, approximately 10 of a given type may adsorb). 10 (These particles can collectively adsorb more than 100 types of proteins). The particles and biomolecular corona can be separated from the biological sample by, for example, centrifugation, magnetic separation, filtration, or gravity separation. The particle types and biomolecular corona can be separated from the biological sample using multiple separation techniques. Non-limiting examples of separation techniques include magnetic separation, column-based separation, filtration, spin column-based separation, centrifugation, ultracentrifugation, density or gradient-based centrifugation, gravity separation, or any combination thereof. Protein corona analysis can be performed on the separated particles and biomolecular corona. Protein corona analysis may include identifying one or more proteins in the biomolecular corona by, for example, mass spectrometry. The method may include contacting a single particle type (e.g., the types of particles listed in Table 1) with the biological sample. The method may also include contacting multiple particle types (e.g., multiple particle types provided in Table 1) with the biological sample. Multiple particle types can be combined and contacted with the biological sample in a single sample volume. Multiple particle types can be successively brought into contact with a biological sample and separated from the sample before subsequent particle types are brought into contact with the biological sample. Protein corona analysis of biomolecular corona can compress the dynamic range of analysis compared to whole protein analysis methods.

[0324] Contacting a biological sample with particles or a group of particles may include adding particles of a specified concentration to the biological sample. Contacting a biological sample with particles or a group of particles may include adding particles of 1 pM to 100 nM to the biological sample. Contacting a biological sample with particles or multiple particles may include adding particles ranging from 1 pM to 500 pM to a biological sample. Contacting a biological sample with particles or multiple particles may include adding particles ranging from 10 pM to 1 nM to a biological sample. Contacting a biological sample with particles or multiple particles may include adding particles ranging from 100 pM to 10 nM to a biological sample. Contacting a biological sample with particles or multiple particles may include adding particles ranging from 500 pM to 100 nM to a biological sample. Contacting a biological sample with particles or multiple particles may include adding particles ranging from 50 μg / ml to 300 μg / ml (particle mass relative to the volume of the biological sample) to a biological sample. Contacting a biological sample with particles or multiple particles may include adding particles ranging from 100 μg / ml to 500 μg / ml to a biological sample. Contacting a biological sample with particles or multiple particles may include adding particles in concentrations of 250 μg / ml to 750 μg / ml to the biological sample. Contacting a biological sample with particles or multiple particles may include adding particles in concentrations of 400 μg / ml to 1 mg / ml to the biological sample. Contacting a biological sample with particles or multiple particles may include adding particles in concentrations of 600 μg / ml to 1.5 mg / ml to the biological sample. Contacting a biological sample with particles or multiple particles may include adding particles in concentrations of 800 μg / ml to 2 mg / ml to the biological sample. Contacting a biological sample with particles or multiple particles may include adding particles in concentrations of 1 mg / ml to 3 mg / ml to the biological sample. Contacting a biological sample with particles or multiple particles may include adding particles in concentrations of 2 mg / ml to 5 mg / ml to the biological sample. Contacting a biological sample with particles or multiple particles may include adding particles in concentrations greater than 5 mg / ml to the biological sample.

[0325] Particles within a group of particles may exhibit varying degrees of uniformity in size and shape. The standard deviation of diameter for collecting a particular type of particle may be less than 20%, 10%, 5%, or 2% of the average diameter of that particle type (for example, less than 2 nm for particles with an average diameter of 100 nm). This may correspond to a low polydispersity index for a sample containing multiple particles, i.e., less than 2, less than 1, less than 0.8, less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2, less than 0.1, or less than 0.05. Conversely, multiple particles may have a high degree of dispersion in average size and shape. The polydispersity index for a sample containing multiple particles may be greater than 3, greater than 4, greater than 5, greater than 8, greater than 10, greater than 12, greater than 15, or greater than 20. The uniformity of size and shape among multiple particles may influence the number and type of biomolecules adsorbed onto the particles. In some methods, uniformity of particle size (e.g., a low polydispersity index) allows for higher enrichment of specific biomolecules, enabling a stronger correspondence between the abundance of enriched biomolecules and particle types. In other methods, lower uniformity of size allows for the collection of a wider variety of biomolecules.

[0326] This specification discloses a method for obtaining a dataset containing proteins detected within biomolecular coronas corresponding to physiologically different particles incubated with a biological sample. The biological sample may include a blood sample from which red blood cells have been removed (e.g., a cell-free sample). Physiologically different types of particles generate different biomolecular coronas. Physiologically different types of particles generate different biomarkers. Physiologically different types of particles generate different mass spectral patterns.

[0327] Particle panel This disclosure provides compositions for assaying protein samples and methods for using the same. The compositions described herein include particle panels comprising one or more different particle types. The particle panels described herein may vary in the number of particle types and the diversity of particle types within a single panel. For example, the particles within a panel may vary based on size, polydispersity, shape and morphology, surface charge, surface chemistry and functionalization, and substrate. The panel can be incubated with the sample to be analyzed for proteins and protein concentrations. Proteins in the sample adsorb onto the surfaces of various particle types within the particle panel, forming protein coronas. The exact proteins and protein concentrations adsorbed onto a particular particle type within the particle panel may depend on the composition, size, and surface charge of that particle type. Thus, each particle type within the panel may have different protein coronas because it adsorbs different sets of proteins, different concentrations of specific proteins, or combinations thereof. Each particle type within the panel may have mutually exclusive protein coronas or overlapping protein coronas. Overlapping protein coronas may overlap in terms of protein identity, protein concentration, or both.

[0328] This disclosure also provides a method for selecting particle types to include in a panel depending on the type of sample. The particle types included in the panel may be a combination of particles optimized for the removal of abundant proteins. The particle types suitable for inclusion in the panel are also selected to adsorb specific proteins of interest. The particles may be nanoparticles. The particles may be microparticles. The particles may be a combination of nanoparticles and microparticles.

[0329] Using the particle panels disclosed herein, a number of distinct proteins and / or specific proteins disclosed herein can be identified over a wide dynamic range. For example, a particle panel disclosed herein, including different particle types, can concentrate proteins in a sample over the entire dynamic range in which the protein is present in the sample (e.g., a plasma sample). In some cases, a particle panel including any number of different particle types disclosed herein can concentrate proteins over a dynamic range of at least two orders of magnitude. In some cases, a particle panel including any number of different particle types disclosed herein can concentrate proteins over a dynamic range of at least three orders of magnitude. In some cases, a particle panel including any number of different particle types disclosed herein can concentrate proteins over a dynamic range of at least four orders of magnitude. In some cases, a particle panel including any number of different particle types disclosed herein can concentrate proteins over a dynamic range of at least five orders of magnitude. In some cases, a particle panel including any number of different particle types disclosed herein can concentrate proteins over a dynamic range of at least six orders of magnitude. In some cases, a particle panel including any number of different particle types disclosed herein can concentrate proteins over a dynamic range of at least seven orders of magnitude. In some cases, a particle panel containing any number of different particle types disclosed herein enriches proteins over a dynamic range of at least eight orders of magnitude. In some cases, a particle panel containing any number of different particle types disclosed herein enriches proteins over a dynamic range of at least nine orders of magnitude. In some cases, a particle panel containing any number of different particle types disclosed herein enriches proteins over a dynamic range of at least ten orders of magnitude. In some cases, a particle panel containing any number of different particle types disclosed herein enriches proteins over a dynamic range of at least eleven orders of magnitude.In some cases, a particle panel containing any number of different particle types disclosed herein enriches proteins over a dynamic range of at least 12 orders of magnitude. In some cases, a particle panel containing any number of different particle types disclosed herein enriches proteins over a dynamic range of 3 to 5 orders of magnitude. In some cases, a particle panel containing any number of different particle types disclosed herein enriches proteins over a dynamic range of 3 to 6 orders of magnitude. In some cases, a particle panel containing any number of different particle types disclosed herein enriches proteins over a dynamic range of 4 to 8 orders of magnitude. In some cases, including any number of different particle types disclosed herein. In particle panels, proteins are concentrated over a dynamic range of 5 to 8 orders of magnitude. In some cases, particle panels containing any number of different particle types disclosed herein can concentrate proteins over a dynamic range of 6 to 10 orders of magnitude. In some cases, particle panels containing any number of different particle types disclosed herein can concentrate proteins over a dynamic range of 8 to 12 orders of magnitude. For example, a particle panel can collect proteins in a sample at mM and fM concentrations, thereby concentrating proteins over a 12-order-of-magnitude range.

[0330] A particle panel comprising any number of different particle types disclosed herein concentrates a single protein or group of proteins. In some cases, a single protein or group of proteins may contain proteins with different post-translational modifications. For example, a first particle type in the particle panel can concentrate a protein or group of proteins with a first post-translational modification, a second particle type in the particle panel can concentrate the same protein or group of proteins with a second post-translational modification, and a third particle type in the particle panel can concentrate the same protein or group of proteins without post-translational modifications. In some cases, a particle panel comprising any number of different particle types disclosed herein concentrates a single protein or group of proteins by binding to different domains, sequences, or epitopes of the single protein or group of proteins. For example, a first particle type in the particle panel can concentrate a protein or group of proteins by binding to a first domain of the protein or group of proteins, and a second particle type in the particle panel can concentrate the same protein or group of proteins by binding to a second domain of the protein or group of proteins.

[0331] The particle panel may include a combination of particles having a silica surface and a polymer surface. For example, the particle panel may include SPION coated with a thin layer of silica, SPION coated with poly(dimethylaminopropyl methacrylamide) (PDMAPMA), and SPION coated with poly(ethylene glycol) (PEG). A particle panel consistent with this disclosure may also include two or more particles selected from the group consisting of silica-coated SPION, N-(3-trimethoxysilylpropyl)diethylenetriamine-coated SPION, PDMAPMA-coated SPION, carboxyl-functionalized polyacrylic acid-coated SPION, amino-surface-functionalized SPION, polystyrene-carboxyl-functionalized SPION, silica particles, and dextran-coated SPION. Particle panels consistent with this disclosure may also include two or more particles selected from the group consisting of surfactant-free carboxylate particles, carboxyl-functionalized polystyrene particles, silica-coated particles, silica particles, dextran-coated particles, oleic acid-coated particles, boronated nanopowder-coated particles, PDMAPMA-coated particles, poly(glycidyl methacrylate-benzylamine)-coated particles, and poly(N-[3-(dimethylamino)propyl]methacrylamide-co-[2-(methacryloyloxy)ethyl]dimethyl-(3-sulfopropyl)ammonium hydroxide, P(DMAPMA-co-SBMA)-coated particles. Particle panels consistent with this disclosure may also include silica-coated particles, N-(3-trimethoxysilylpropyl)diethylenetriamine-coated particles, poly(N-(3-(dimethylamino)propyl)methacrylamide)(PDMAPMA)-coated particles, sugar-phosphate-functionalized polystyrene particles, amine-functionalized polystyrene particles, polystyrene carboxyl-functionalized particles, ubiquitin-functionalized polystyrene particles, dextran-coated particles, or any combination thereof.

[0332] The particle panels disclosed herein can be analyzed using the workflow described herein (MS analysis of individual biomolecular coronas corresponding to individual particle types within the particle panel, collectively referred to as the “Proteograph” workflow) to analyze multiple proteins. It can be used to identify peptides or protein groups. The feature intensities disclosed herein are derived from the intensities of discrete spikes ("features") found in a mass-to-charge ratio vs. intensity plot from a mass spectrometry run of the sample. These features can correspond to variably ionized fragments of peptides and / or proteins. Using the data analysis methods described herein, feature intensities can be classified into protein groups. A protein group refers to two or more proteins identified by a shared peptide sequence. Alternatively, a protein group can refer to a single protein identified using a unique identification sequence. For example, if assaying a peptide sequence shared between two proteins in a sample (protein 1: XYZZX and protein 2: XYZYZ), the protein group could be an "XYZ protein group" with two members (protein 1 and protein 2). Alternatively, if the peptide sequence is unique to a single protein (protein 1), the protein group could be a "ZZX" protein group with one member (protein 1). Each protein group may be supported by multiple peptide sequences. The proteins detected or identified in accordance with this disclosure may refer to distinct proteins detected in a sample (e.g., other distinct relative proteins detected using mass spectrometry). Thus, analyzing proteins present in distinct coronas corresponding to distinct particle types within a particle panel yields a large number of feature intensities. This number decreases when the feature intensities are processed into distinct peptides, further decreases when the distinct peptides are processed into distinct proteins, and further decreases when the peptides are grouped into protein groups (two or more proteins sharing distinct peptide sequences).

[0333] A particle panel disclosed herein for evaluating the presence or absence of one or more biomarkers associated with lung cancer (e.g., NSCLC) comprises at least one different particle type, at least two different particle types, at least three different particle types, at least four different particle types, at least five different particle types, at least six different particle types, at least seven different particle types, at least eight different particle types, at least nine different particle types, at least ten different particle types, at least eleven different particle types, at least twelve different particle types, at least thirteen different particle types, at least fourteen different particle types, at least fifteen different particle types, at least sixteen different particle types, at least seventeen different particle types, at least eighteen different particle types, at least nineteen different particle types, at least sixteen different particle types, at least seventeen different particle types, at least eighteen different particle types, at least nineteen different particle types, at least twenty different particle types, at least twenty-five different particle types, at least thirty different particle types, at least thirty-five different particle types, at least forty different particle types, at least forty-five different particle types, and at least fifty different particle types. Different particle types, at least 55 different particle types, at least 60 different particle types, at least 65 different particle types, at least 70 different particle types, at least 75 different particle types, at least 80 different particle types, at least 85 different particle types, at least 90 different particle types, at least 95 different particle types, at least 100 different particle types, 1 to 5 different particle types, 5 to 10 different particle types, 10 to 15 different particle types, 15 to 20 different particle types, 20 to 25 different Particle types, 25-30 different particle types, 30-35 different particle types, 35-40 different particle types, 40-45 different particle types, 45-50 different particle types, 50-55 different particle types, 55-60 different particle types, 60-65 different particle types, 65-70 different particle types, 70-75 different particle types, 75-80 different particle types, 80-85 different particle types, 85-90 different particle types, 90-95 different particle types, 95-100 different particle types, 1-100 different particle types,It may have 20 to 40 different particle types, 5 to 10 different particle types, 3 to 7 different particle types, 2 to 10 different particle types, 6 to 15 different particle types, or 10 to 20 different particle types. In a particular embodiment, the disclosure provides panel sizes with 3 to 10 particle types. In a particular embodiment, the disclosure provides panel sizes with 4 to 11 different particle types. The disclosure provides panel sizes. In certain embodiments, the disclosure provides panel sizes of 5 to 15 different particle types. In certain embodiments, the disclosure provides panel sizes of 5 to 15 different particle types. In certain embodiments, the disclosure provides panel sizes of 8 to 12 different particle types. In certain embodiments, the disclosure provides panel sizes of 9 to 13 different particle types. In certain embodiments, the disclosure provides panel sizes of 10 different particle types. The particle types may include nanoparticle types.

[0334] Particle panels can be designed to profile a wide range of proteomes, such as the human plasma proteome. A major challenge in analyzing the human proteome is that over 99% of the mass of the approximately 3,500 proteins in human plasma is comprised of only 20 proteins. Plasma analysis methods are often saturated with these 20 proteins, resulting in minimal profiling depth for the remaining proteins. The particle panels of this disclosure may include combinations of particles that facilitate the collection of at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1800, at least 1900, at least 2000, at least 2100, or at least 2200 different proteins from a single biological sample. The particle panels of this disclosure may comprise combinations of particles that facilitate the collection of at least 4%, at least 5%, at least 6%, at least 8%, at least 10%, at least 12%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70% of a particular protein type from a complex biological sample such as human plasma. This can be achieved by providing multiple particles having different protein binding profiles (e.g., as a particle panel). A particle panel may comprise two particles that, upon contact with a biological sample, form a protein corona with less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 25%, less than 20%, less than 15%, or less than 10% of a common protein. In some cases, the biological sample is human plasma.

[0335] Increasing the number of particle types in the panel increases the number of proteins that can be identified in a given sample. Figure 53 shows an example of how the number of proteins identified can increase with increasing panel size, where 419 different proteins were identified with one particle type panel size, 588 different proteins with two particle types, 727 different proteins with three particle types, 844 proteins with four particle types, 934 different proteins with five particle types, 1008 different proteins with six particle types, 1075 different proteins with seven particle types, 1133 different proteins with eight particle types, 1184 different proteins with nine particle types, 1230 different proteins with ten particle types, 1275 different proteins with eleven particle types, and 1318 different proteins with twelve particle types.

[0336] Dynamic range Some methods described herein (e.g., biomolecular corona analysis) may involve assaying biomolecules in a sample of this disclosure over a wide dynamic range. The dynamic range of biomolecules assayed in a sample is determined by the assay method of biomolecules contained in the sample (e.g., mass spectrometry, chromatography, gel electrophoresis, spectroscopy). This can refer to the range of abundance of biomolecules measured by an assay (or immunoassay). For example, an assay that can detect proteins across a wide dynamic range may be able to detect proteins ranging from very low abundances to very high abundances. The dynamic range of an assay may be directly related to the slope of the assay signal intensity as a function of biomolecule abundances. For example, in an assay with a low dynamic range, the slope of the assay signal intensity as a function of biomolecule abundances may be low (but positive), and the ratio of signals detected for high-abundance biomolecules to signals detected for low-abundance biomolecules may be lower in an assay with a low dynamic range than in an assay with a high dynamic range. In certain cases, dynamic range may refer to the dynamic range of proteins within a sample or assay method.

[0337] The methods described herein may compress the dynamic range of an assay. The dynamic range of one assay may be compressed relative to another if the slope of the assay signal intensity as a function of biomolecular abundance is lower than that of another assay. For example, a plasma sample assayed using protein corona analysis in conjunction with mass spectrometry may exhibit a compressed dynamic range compared to a plasma sample assayed directly on the sample using mass spectrometry alone, or compared to the provided abundance values ​​of plasma proteins in a database (e.g., the database provided in Keshishianet al., Mol. Cell Proteomics 14, 2375-2393 (2015), also referred to herein as the "Carr database"). This compressed dynamic range may allow for the detection of more low-abundance biomolecules in a biological sample by using biomolecular corona analysis in conjunction with mass spectrometry than by using mass spectrometry alone.

[0338] Collecting biomolecules on particles before analysis (e.g., mass spectrometry or ELISA analysis) can compress the dynamic range of the analysis. 6 The two proteins, present in a :1 ratio, are differentially adsorbed onto the particle, resulting in a new ratio of 10 4 They are eluted into the solution in a ratio of :1. Such differential adsorption can sometimes make it possible to simultaneously detect two biomolecules with a concentration difference that exceeds the dynamic range of analytical techniques. For example, mass spectrometry is often limited to measuring species within a concentration range of 4 to 6 orders of magnitude, so 10 8 It may not be possible to simultaneously detect two biomolecules present at twice the concentration. Corona-based biomolecule enrichment of a sample allows for the enrichment of a biomolecule (e.g., a first protein) that is dilute compared to a second biomolecule (e.g., a second protein), thereby enabling the simultaneous detection of two biomolecules using a single analytical method. Similarly, particle-based enrichment may enable the quantification of low concentrations of biomolecules in a sample. The dynamic range in which an analyte can be quantified is often narrower than the dynamic range in which the analyte can be detected. For example, ELISA often provides accurate concentration quantification over less than two orders of magnitude while covering a dynamic range of two to three orders of magnitude. Particle-based enrichment can increase the number of biomolecular targets within a desired concentration range, thereby enabling the simultaneous quantification of two or more biomolecules present in a biological sample at concentrations outside the dynamic range of the analytical technique's concentration quantification.

[0339] Therefore, the various methods of this disclosure involve detecting two biomolecules present in a biological sample at a concentration difference greater than the dynamic range of the detection method. Many of the biomarker pairs disclosed herein extend to a concentration range that exceeds the detection limits of biomolecular analysis techniques (e.g., immunostaining or LC-MS / MS), and therefore may be unidentifiable or unquantifiable without using the concentration-based methods of this disclosure. In some cases, the methods of this disclosure detect two biomolecules (e.g., two proteins) in a biological sample (e.g., 1 mg / ml and 1 μg / ml, or 50 μM and 50 nM) at concentrations that differ by at least three orders of magnitude. This includes detecting (quality). In some cases, the method of the Disclosure includes detecting two biomolecules (e.g., two proteins) in a biological sample (e.g., 1 mg / ml and 100 ng / ml, or 50 μM and 5 nM) at concentrations that differ by at least four orders of magnitude. In some cases, the method of the Disclosure includes detecting two biomolecules (e.g., two proteins) in a biological sample at concentrations that differ by at least five orders of magnitude (e.g., detection of HBA and NOTUM in human plasma). In some cases, the method of the Disclosure includes detecting two biomolecules (e.g., two proteins) in a biological sample at concentrations that differ by at least five orders of magnitude (e.g., detection of ITIH2 and ANGL6 in human plasma). In some cases, the method of the Disclosure includes detecting two biomolecules (e.g., two proteins) in a biological sample at concentrations that differ by at least six orders of magnitude (e.g., detection of HBA and NOTUM in human plasma). In some cases, the methods of the present disclosure include detecting two biomolecules (e.g., two proteins) in a biological sample at concentrations that differ by at least seven orders of magnitude (e.g., detection of ceruloplasmin and RLA2 in human plasma). In some cases, the methods of the present disclosure include detecting two biomolecules (e.g., two proteins) in a biological sample at concentrations that differ by at least seven orders of magnitude (e.g., detection of human serum albumin and CAN2 in human plasma). In some cases, the methods of the present disclosure include detecting two biomolecules (e.g., two proteins) in a biological sample at concentrations that differ by at least seven orders of magnitude (e.g., detection of human serum albumin and interleukin-6 in human plasma).

[0340] The dynamic range of a proteomics analysis assay may be the ratio of the signal generated by the most abundant protein (e.g., the protein with the highest abundance of 10%) to the signal generated by the least abundant protein (e.g., the protein with the lowest abundance of 10%). Compressing the dynamic range of a proteomics analysis may involve reducing the ratio of the signal generated by the most abundant protein to the signal generated by the least abundant protein for a first proteomics analysis assay compared to a second proteomics analysis assay. Protein corona analysis assays disclosed herein can compress the dynamic range compared to the dynamic range of a whole protein analysis method (e.g., mass spectrometry, gel electrophoresis, or liquid chromatography).

[0341] This specification provides several methods for compressing the dynamic range of biomolecular analysis assays to facilitate the detection of biomolecular molecules in low-abundance quantities compared to biomolecular molecules in high-abundance quantities. For example, the particle types of this disclosure can be used to continuously investigate a sample. When the particle types are incubated in a sample, a biomolecular corona is formed on the surface of the particle types. For example, if biomolecular molecules in a sample are detected directly by direct mass spectrometry without using the particle types, the dynamic range may span a wider concentration range or more orders of magnitude than when biomolecular molecules are detected directly on the surface of the particle types. Therefore, the use of the particle types disclosed herein can be used to compress the dynamic range of biomolecular molecules in a sample. This effect may be observed, though not limited to theory, due to a high affinity of the particle types for the biomolecular corona, resulting in greater capture of low-abundance biomolecular molecules, and a low affinity of the particle types for the biomolecular corona, resulting in less capture of high-abundance biomolecular molecules.

[0342] The dynamic range of a proteomics analysis assay can be the slope of the plot of the protein signal measured by the proteomics analysis assay as a function of the total abundance of protein in the sample. Compressing the dynamic range is the relationship between the total abundance of protein in the sample and the slope of the plot of the protein signal measured by a second proteomics analysis assay as a function of the total abundance of protein in the sample. Numerically, this may include reducing the slope of the plot of protein signals measured by proteomics analysis assays. Protein corona analysis assays disclosed herein can compress the dynamic range compared to the dynamic range of total protein analysis methods (e.g., mass spectrometry, gel electrophoresis, or liquid chromatography).

[0343] Biomarker analysis in biological samples The methods of use disclosed herein can identify numerous biomarkers in biological samples (e.g., biological fluids). Non-limiting examples of biological samples that can be analyzed using the methods described herein (e.g., protein corona analysis) include biological fluid samples (e.g., fluids from cerebrospinal fluid (CSF), synovial fluid (SF), urine, plasma, serum, tears, semen, whole blood, milk, nipple aspirate, mammary duct lavage fluid, vaginal fluid, nasal secretions, ear fluid, gastric juice, pancreatic juice, trabecular meshwork, lung lavage fluid, prostatic fluid, sputum, feces, bronchial lavage fluid, swabs, bronchial aspirates, sweat, or saliva), fluid solids (e.g., tissue homogenates), or samples derived from cell cultures. For example, by incubating the particles disclosed herein with any biological sample disclosed herein, at least 100 unique proteins, at least 120 unique proteins, at least 140 unique proteins, at least 160 unique proteins, at least 180 unique proteins, at least 200 unique proteins, at least 220 unique proteins, at least 240 unique proteins, at least 260 unique proteins, at least 280 unique proteins, at least 300 unique proteins, at least 320 unique proteins, at least 340 unique proteins, at least 360 unique proteins, at least 380 unique proteins, at least 400 unique proteins, at least 420 unique proteins, and at least 440 unique proteins, at least 460 unique proteins, at least 480 unique proteins, at least 500 unique proteins, at least 520 unique proteins, at least 540 unique proteins, at least 560 unique proteins, at least 580 unique proteins, at least 600 unique proteins, at least 620 unique proteins, at least 640 unique proteins, at least 660 unique proteins, at least 680 unique proteins, at least 700 unique proteins, at least 720 unique proteins, at least 740 unique proteins, at least 760 unique proteins, at least 780 unique proteins, at least 800 unique proteins, at least 820 unique proteins,At least 840 unique proteins, at least 860 unique proteins, at least 880 unique proteins, at least 900 unique proteins, at least 920 unique proteins, at least 940 unique proteins, at least 960 unique proteins, at least 980 unique proteins, at least 1000 unique proteins, 100-1000 unique proteins, 150-950 unique proteins, 200-900 unique proteins, 250-850 unique proteins, 300-800 unique proteins, 350-750 unique proteins, 400-700 unique proteins, 450-650 unique proteins, 500-60 A protein corona can be formed containing 0 unique proteins, 200-250 unique proteins, 250-300 unique proteins, 300-350 unique proteins, 350-400 unique proteins, 400-450 unique proteins, 450-500 unique proteins, 500-550 unique proteins, 550-600 unique proteins, 600-650 unique proteins, 650-700 unique proteins, 700-750 unique proteins, 750-800 unique proteins, 800-850 unique proteins, 850-900 unique proteins, 900-950 unique proteins, and 950-1000 unique proteins. In some cases, a similar number of proteins can be evaluated without using particles or using the assay methods described herein. In some embodiments, several different methods are used to identify a large number of proteins in a particular biological sample. These types of particles can be used individually or in combination. In other words, particles can be multiplexed to bind and identify multiple proteins in a biological sample.

[0344] The methods disclosed herein can be used to identify various biological states in specific biological samples. For example, a biological state may refer to elevated or low levels of a particular protein or set of proteins, or may be demonstrated by the ratio of the abundances of two or more biomolecules. In other examples, a biological state may refer to the identification of a disease, such as cancer. A biological state may include cancerous pulmonary nodules. A biological state may include non-cancerous pulmonary nodules. When one or more particle types are incubated with a biological sample, such as human plasma, the formation of a protein corona is possible. The protein corona can then be analyzed to identify the protein pattern. The analysis may include gel electrophoresis, mass spectrometry, chromatography, ELISA, immunohistochemistry, or any combination thereof. Analysis of the protein corona (e.g., by mass spectrometry or gel electrophoresis) is sometimes called corona analysis. The protein pattern can be compared to the same method performed on a control sample. By comparing the protein patterns, it is possible to identify that the first sample contains high levels of markers corresponding to a particular type of lung cancer. Thus, the particles and their methods of use can be used to diagnose specific pathological conditions.

[0345] The assay may include protein collection of particles, protein digestion, and mass spectrometry (e.g., MS, LC-MS, LC-MS / MS). Digestion may include chemical digestion with cyanide bromide or 2-nitro-5-thiocyanatobenzoic acid (NTCB). Digestion may include enzymatic digestion with trypsin or pepsin. Digestion may include enzymatic digestion with multiple proteases. The digestion may include proteases selected from the group consisting of trypsin, chymotrypsin, Glu C, Lys C, elastase, subtilisin, proteinase K, thrombin, factor X, Arg C, papain, Asp N, thermolysin, pepsin, aspartylprotease, cathepsin D, zinc mealloprotease, glycoprotein endopeptidase, proline, aminopeptidase, prenylprotease, caspase, kex2 endopeptase, or any combination thereof. The digestion method may cleave the peptide randomly or cleave it at specific positions or a series of positions. Multiple digestion methods (e.g., two or more proteases) may be used in the assay. The assay may include dividing the sample into multiple parts and subjecting each part to a different digestion method and separate analysis (e.g., separate mass spectrometry). Digestion can cleave peptides at specific locations (e.g., methionine) or sequences (e.g., glutamate-histidine-glutamate). Digestion can sometimes distinguish similar proteins. For example, in an assay, a first digestion method can degrade eight distinct proteins as a single protein group, and a second digestion method can degrade eight distinct proteins by distinct signals. Digestion can produce peptide fragment lengths of an average of 8-15 amino acids. Digestion can produce peptide fragment lengths of an average of 12-18 amino acids. Digestion can produce peptide fragment lengths of an average of 15-25 amino acids. Digestion can produce peptide fragment lengths of an average of 20-30 amino acids. Digestion can produce peptide fragment lengths of an average of 30-50 amino acids.

[0346] The various methods described herein enable measurements over a wide concentration range. Biomolecular analysis methods are often limited to narrow concentration ranges. For example, proteomics analysis by mass spectrometry is often limited to concentrations of three, four, or five orders of magnitude. Therefore, the presence of relatively high concentrations of biomolecules (e.g., at mg / ml concentrations) hinders the detection of low concentrations of biomolecules. Masking may occur, further limiting the accuracy of quantification of low-concentration biomolecules. The methods of this disclosure may enable the detection of molecules with concentrations spanning at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, or at least 12 orders of magnitude. Thus, the methods of this disclosure can detect and quantify both relatively high and relatively low concentrations of biomolecules from a single sample without first depleting biomolecules from the sample. For example, a plasma assay consistent with this disclosure can simultaneously quantify albumin (present at approximately 40 mg / ml) and interleukin-10 (present at approximately 6 pg / ml) from a single, undepleted plasma sample, thereby enabling the simultaneous detection of two species with concentrations differing by approximately 10 orders of magnitude.

[0347] Biomarkers for cancer detection Proteins may be included as biomarkers for disease detection. Disease detection may include cancer detection using biomarkers such as proteins. Proteins can be generated as part of protein data or proteomics data.

[0348] Examples of proteins may include any of the proteins shown in Figures 26A–26B. Protein data may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 measurements of these proteins, or any range of the aforementioned numbers of measurements of the proteins from these figures.

[0349] Some examples of proteins are shown in Figure 30A. Proteins that can be detected by the methods described herein include myosin-9 (MYH9), tubulin beta-1 chain (TUBB1), tubulin beta chain (TUBB), calreticulin (CALR), vascular endothelial growth factor receptor 3 (FLT4), neurogenic locus notch homolog protein 2 (NOTCH2), transforming protein RhoA (RHOA), isocitrate dehydrogenase [NADP], mitochondria (IDH2), cadherin 1 (CDH1), and cAMP-dependent proteins. These include in kinase type I-alpha regulatory subunit (PRKAR1A), neurogenic locus notch homolog protein 1 (NOTCH1), exostosin-1 (EXT1), serine / threonine protein phosphatase 2A 65kDa regulatory subunit A alpha isoform (PPP2R1A), staphylococcal nuclease domain-containing protein 1 (SND1), tyrosine protein kinase BTK (BTK), lipoma-preferred partner (LPP), mitogen-activated protein kinase (MAPK1), Fat1 protein (FAT1), cadherin-11 (CDH11), or bispecific mitogen-activated protein kinase kinase 1 (MAP2K1). Another example of proteins is shown in Figures 32A-32B. Proteins detected by the methods described herein may also include thrombospongin-2 (TSP2 or P35442). Another example of a protein is shown in Figures 32C–32D. Proteins detected by the methods described herein may include P01011. Several examples of proteins are shown in Figure 36.Proteins detected by the methods described herein include high molecular weight immunoglobulin receptor (PIGR, UniProt P01833), cadherin-related family member 2 (CDHR2, UniProt Q9BYE9), leucine-rich alpha-2 glycoprotein (LRG1 or A2GL, UniProt P02750), intercellular adhesion molecule 1 (ICAM1, UniProt P05362), aminopeptidase N (AMPN or ANPEP, UniProt P15144), thrombospondin-2 (TSP2, UniProt P35442), protein S100-A9 (S10A9 or S100A9, UniProt P06702), aldo-ketoreductase family 1 member B1 (ALDR or AKR1B1, UniProt P15121), and serum amyloid A-1 protein (SAA1, UniProt P0DJI8), peroxidacin homolog (PXDN, UniProt Q92626), protein S100-A8(S. 10A8 or S100A8, UniProt P05109), Anthrax Toxin Receptor 2 (ANTR2 or ANTXR2, UniProt P58335), Cadherin-2 (CADH2 or CDH2, UniProt P19022), Alpha-1-Antichymotrypsin (AACT or SERPINA3, UniProt P01011), Collagen Alpha-1 (XVIII) Chain (COIA1 or COL18A1, UniProt P39060), Fibrinogen-like Protein 1 (FGL1, UniProt Q08830), Protein S100-A12 (S10AC or S100A12, UniProt P80511), Reelin (RELN, UniProt J3KQ66), C-reactive Protein (CRP, UniProt P02741), versican core protein (CSPG2 or VCAN, UniProt P13611), coagulation factor XIIIA chain (F13A or F13A1, UniProt P00488), cartilage intermediate protein 2 (CILP2, UniProt K7EPJ4), Sushi, von Willebrand factor A, EGF and pentraxin domain-containing protein 1 (SVEP1, UniProt Q4LDE5), neutrophil gelatinase-related lipocalin (NGAL or LCN2, UniProt P80188), tetranectin (TETN or CLEC3B, UniProt P05452), SLAIN motif-containing protein 2 (SLAI2 or SLAIN2, UniProt Q9P270), anthrax toxin receptor 1 (ANTR1 or ANTXR1, UniProt Q9H6X2, e.g., isoform 5 [UniProt This may include Q9H6X2-5) or serum amyloid A-2 protein (SAA2, UniProt P0DJI9). Any number of the aforementioned proteins can be used. Any of these proteins can be used in the classifier.

[0350] Examples of proteins may include SERPINA1, HPR, EPS15L1, ORM2, CTSH, CRP, SAA4, COLEC10, HIST1H4I, APOM, ORM1, P0DOX8, IGKV1-8, IGKV1-9, ANGPTL6, SERPINA3, PXDN, IGKC, HP, APCS, or ITIH2. Protein data may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 measurements of these proteins, or any range of the aforementioned numbers of measurements of these proteins.

[0351] The method may include measuring biomarkers in a biological fluid sample. The method may include using biomarkers in a biological fluid sample. Examples of biomarkers include A2GL, AKR1B1, ANPEP, ANTXR1, ANTXR2, BTK, CALR, CDH1, CDH11, CDH2, CDHR2, CILP2, CLEC3B, COL18A1, CRP, EXT1, F13A1, FAT1, FGL1, FLT4, ICAM1, IDH2, LCN2, LPP, MAPK1, MAP2K1, MYH9, NOTCH1, NOTCH2, PIGR, PPP2R1A, PRKAR1A, PXDN, RELN, RHOA, S100A8, S100A9, S100A12, SAA1, SAA2, SERPINA3, SLAIN2, SND1, SVEP1, TSP2, TUBB, TUBB1, or VCAN. In some embodiments, the biomarkers include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, or 48 of the aforementioned biomarkers, or a range of biomarkers defined by any two of the aforementioned integers.

[0352] Proteomics data may include protein measurements. Protein measurements in samples from subjects with liver cancer may be increased or decreased compared to protein measurements from control samples or compared to baseline measurements. Protein measurements may include measurements of proteins or combinations of proteins from Figure 39C or Figure 39D. For examp...

Claims

[Claim 1] The invention described herein.