Methods of identifying pancreatic cancer
By measuring and analyzing biomarkers in subjects suspected of having pancreatic cancer and using a classifier for evaluation, the problem of accurate detection of pancreatic cancer in the early stage was solved, and the goal of improving treatment effect and prognosis was achieved.
Patent Information
- Application Number
- CN202380072587.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-15
- Filing Date
- 2023-09-07
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to accurately detect pancreatic cancer in the early stage, affecting the treatment and prognosis of patients.
Specific biomarkers, such as AACT, A1AT, A2GL, etc., in biological fluid samples of subjects suspected of having pancreatic cancer, were measured and evaluated using a classifier to determine whether pancreatic cancer is present.
Early accurate detection of pancreatic cancer has been achieved, improving the effectiveness of treatment and the prognosis of patients.
Smart Images

Figure CN120188046A_ABST
Abstract
Description
[0001] Cross - reference
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 375,020, filed on September 8, 2022, and U.S. Provisional Application No. 63 / 485,190, filed on February 15, 2023, each of which is incorporated herein by reference.
[0003] Incorporation of sequence listing by reference
[0004] This application is being filed with a sequence listing in electronic format. The sequence listing is provided as a file named "PrognomIQ 59521 - 714.601.xml", created on August 20, 2023, and having a size of 21,183 bytes. The information in the electronic format of the sequence listing is incorporated by reference in its entirety. Background of the Invention
[0006] There is a need to accurately detect cancer, such as pancreatic cancer, at an early stage. Accurately detecting cancer at an early stage can lead to effective treatment and improved prognosis for subjects with cancer. Summary of the Invention
[0007] In some aspects, detection methods are disclosed herein. Some aspects include measuring biomarkers including AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2 or APOA1 or combinations thereof in a biological fluid sample from a subject suspected of having pancreatic cancer to obtain biomarker measurements. Some aspects include applying a classifier to the biomarker measurements to evaluate pancreatic cancer in the subject. Some aspects include identifying the biomarker measurements as indicative of pancreatic cancer in the subject, or identifying the biomarker measurements as indicative of a lack of pancreatic cancer in the subject. Some aspects include administering a pancreatic cancer treatment to the subject when the biomarker measurements are identified as indicative of pancreatic cancer, and observing or treating the subject without administering a pancreatic cancer treatment when the biomarker measurements are identified as indicative of a lack of pancreatic cancer. In some aspects, evaluation methods are disclosed herein. Some aspects include obtaining a data set including biomarker measurements from a biological fluid sample from a subject suspected of having pancreatic cancer, the biomarkers including AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2 or APOA1 or combinations thereof; and applying a classifier to the data set to evaluate pancreatic cancer in the subject. In some aspects, evaluating pancreatic cancer in the subject includes identifying the data set as indicative of pancreatic cancer in the subject, or identifying the data set as indicative of a lack of pancreatic cancer in the subject. Some aspects include administering a pancreatic cancer treatment to the subject when the data set is identified as indicative of pancreatic cancer, and observing or treating the subject without administering a pancreatic cancer treatment when the data set is identified as indicative of a lack of pancreatic cancer. In some aspects, the classifier includes a performance determined as follows: a receiver operating characteristic (ROC) curve having an area under the curve (AUC) greater than 0.85, greater than 0.86, greater than 0.87, greater than 0.88, greater than 0.89, greater than 0.90, greater than 0.91, greater than 0.92, greater than 0.93, greater than 0.94, greater than 0.95, greater than 0.96 or greater than 0.97 in the discrimination between pancreatic cancer and lack of pancreatic cancer.In some aspects, the classifier includes performance determined, for example, by a sensitivity greater than 50%, greater than 55%, greater than 60%, greater than 65%, greater than 70%, greater than 75%, greater than 80%, greater than 85%, greater than 86%, greater than 87%, greater than 88%, greater than 89%, greater than 90%, greater than 91%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, or greater than 99% in identifying pancreatic cancer versus the absence of pancreatic cancer. In some aspects, the classifier includes performance determined, for example, by a specificity greater than 80%, greater than 81%, greater than 82%, greater than 83%, greater than 84%, greater than 85%, greater than 86%, greater than 87%, greater than 88%, greater than 89%, greater than 90%, greater than 91%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, or greater than 99% in differentiating between pancreatic cancer and the absence of pancreatic cancer. In some aspects, the biomarker includes A1AT. In some aspects, the biomarker includes A2GL. In some aspects, the biomarker includes AACT. In some aspects, the biomarker includes AMPN. In some aspects, the biomarker includes APOA1. In some aspects, the biomarker includes APOA2. In some aspects, the biomarker includes CO2. In some aspects, the biomarker includes CO5. In some aspects, the biomarker includes CO9. In some aspects, the biomarker includes CRP. In some aspects, the biomarker includes F13B. In some aspects, the biomarker includes FCG3A. In some aspects, the biomarker includes ICAM1. In some aspects, the biomarker includes ITIH3. In some aspects, the biomarker includes LBP. In some aspects, the biomarker includes NOE1. In some aspects, the biomarker includes PIGR. In some aspects, the biomarker includes RET4. In some aspects, the biomarker includes S10A8. In some aspects, the biomarker includes TETN. In some aspects, the biomarker includes CA19-9. In some aspects, the biomarker includes two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, ten or more, 12 or more, 14 or more, 16 or more, 18 or more, or 20 or more of AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2, APOA1, or CA19-9. In some aspects, the measured value is obtained by adding an internal standard for any one of the biomarkers to the sample.In some aspects, the internal standard is labeled. In some aspects, the internal standard is isotopically labeled. In some aspects, mass spectrometry is used to obtain biomarker measurements. In some aspects, immunoassays are used to obtain biomarker measurements. In some aspects, molecular probes are used to obtain biomarker measurements. In some aspects, chromatography is used to obtain biomarker measurements. In some aspects, the biological fluid includes pancreatic cyst fluid, blood, plasma, or serum. In some aspects, the subject is a mammal. In some aspects, the subject is a human. In some aspects, the classifier identifies the stage of pancreatic cancer in the subject. In some aspects, pancreatic cancer includes stage I pancreatic cancer or stage II pancreatic cancer. In some aspects, pancreatic cancer includes stage III pancreatic cancer or stage IV pancreatic cancer. In some aspects, pancreatic cancer includes pancreatic ductal adenocarcinoma (PDAC). In some aspects, when the cancer assessment method indicates that the subject has a probability greater than a predetermined threshold of having pancreatic cancer, the method further includes performing a subsequent pancreatic cancer treatment or recommending that the subject undergo a subsequent pancreatic cancer treatment to determine the presence of pancreatic cancer. In some aspects, the subsequent pancreatic cancer treatment includes a biopsy. In some aspects, the subsequent pancreatic cancer treatment includes pancreatic imaging. In some aspects, ultrasound or computed tomography is used to perform the imaging. In some aspects, when the cancer assessment method indicates that the subject has a probability greater than a predetermined threshold of having pancreatic cancer, the method further includes treating the subject with a pancreatic cancer treatment for treating pancreatic cancer or recommending that the subject undergo such a pancreatic cancer treatment. In some aspects, the pancreatic cancer treatment is selected from the group consisting of: surgery for pancreatic cancer, radiotherapy for pancreatic cancer, cryotherapy for pancreatic cancer, hormone therapy for pancreatic cancer, chemotherapy for pancreatic cancer, ablation therapy for pancreatic cancer, and immunotherapy for pancreatic cancer. In some aspects, the predetermined threshold is greater than 10%, greater than 20%, greater than 30%, greater than 40%, greater than 50%, greater than 60%, greater than 70%, greater than 80%, or greater than 90%. In some aspects, the classifier identifies the stage of pancreatic cancer in the subject.
[0008] In some aspects, methods for detecting pancreatic cancer in a subject are disclosed herein, which include: identifying a subject at risk of having pancreatic cancer; obtaining a biofluid sample from the subject; contacting the biofluid sample with particles such that the particles adsorb biomolecules containing proteins to the particles; assaying the biomolecules adsorbed to the particles to generate proteomic data; and classifying the proteomic data as indicative of pancreatic cancer or not indicative of pancreatic cancer. In some aspects, identifying the subject as being at risk of having pancreatic cancer includes identifying the subject as having a computed tomography (CT) scan indicative of pancreatic cancer, having a magnetic resonance imaging (MRI) scan indicative of pancreatic cancer, having a positron emission tomography (PET) scan indicative of pancreatic cancer, having an ultrasound indicative of pancreatic cancer, having a cholangiopancreatography indicative of pancreatic cancer, having an angiography indicative of pancreatic cancer, having a liver function test (LFT) indicative of pancreatic cancer, having an elevated carcinoembryonic antigen (CEA) level relative to a control or baseline measurement, having an elevated carbohydrate antigen (CA) 19-9 level relative to a control or baseline measurement, having jaundice, having abdominal pain, having an enlarged gallbladder or liver, having a thrombus, or having a pancreatic cyst or a combination thereof. Some aspects include identifying the likelihood that the subject has pancreatic cancer based on the proteomic data. In some aspects, classifying the proteomic data as indicative of pancreatic cancer or as not indicative of pancreatic cancer includes applying a classifier to the proteomic data. In some aspects, the classifier includes features to identify the likelihood that the subject has pancreatic cancer. In some aspects, the classifier is trained using deep learning, hierarchical clustering analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k-nearest neighbor analysis, naive Bayes analysis, K-means clustering analysis, or hidden Markov analysis. In some aspects, the proteomic data indicates pancreatic cancer with at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% sensitivity or specificity. Some aspects include recommending pancreatic cancer treatment to the subject when the proteomic data is classified as indicative of pancreatic cancer. Some aspects include administering pancreatic cancer treatment to the subject when the proteomic data is classified as indicative of pancreatic cancer. Some aspects include recommending or performing a biopsy when the proteomic data is classified as indicative of pancreatic cancer. Some aspects include recommending observing the subject without administering pancreatic cancer treatment to the subject or recommending observing the subject without obtaining a biopsy from the subject when the proteomic data is not classified as indicative of pancreatic cancer. Some aspects include observing the subject without administering pancreatic cancer treatment to the subject or observing the subject without obtaining a biopsy from the subject when the proteomic data is not classified as indicative of pancreatic cancer. In some aspects, pancreatic cancer treatment includes chemotherapy, radiotherapy, immunotherapy, targeted therapy, surgery, or surgical resection or a combination thereof.In some aspects, the treatment of pancreatic cancer includes administering a pharmaceutical composition comprising capecitabine, erlotinib, fluorouracil, gemcitabine, irinotecan, leucovorin, nab-paclitaxel, nanoliposomal irinotecan, oxaliplatin, olaparib or larotrectinib or a combination thereof. In some aspects, the particles include nanoparticles. In some aspects, the particles include lipid particles, metal particles, silica particles or polymer particles. In some aspects, the particles include carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles or N-(3-trimethoxysilylpropyl)diethylenetriamine particles. In some aspects, the particles comprise a group of physiochemically distinct nanoparticles. In some aspects, assaying a biomolecule includes performing mass spectrometry, chromatography, liquid chromatography, high performance liquid chromatography, solid phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, western blot, dot blot or immunostaining or a combination thereof. In some aspects, assaying a biomolecule includes performing mass spectrometry. In some aspects, assaying a biomolecule includes measuring a reading indicative of the presence, absence or amount of the biomolecule. In some aspects, the pancreatic cancer includes early-stage pancreatic cancer. In some aspects, the pancreatic cancer includes late-stage pancreatic cancer. Some aspects include monitoring a subject and assaying a biomolecule in a second biofluid sample obtained from the subject at a later time. In some aspects, the protein includes a secreted protein. In some aspects, the biofluid includes blood, plasma or serum. In some aspects, the subject has pancreatic cancer. In some aspects, the subject does not have pancreatic cancer. In some aspects, the subject is a mammal. In some aspects, the subject is a human.
[0009] In some aspects, the present disclosure provides methods that include: measuring proteins in a biological fluid sample obtained from a subject identified as being at risk of having pancreatic cancer to obtain protein measurements; and applying a classifier to the protein measurements to identify the protein measurements as indicative of a subject having pancreatic cancer, wherein the classifier is generated using proteomics data obtained by contacting a training sample with particles such that the particles adsorb proteins in the training sample and measuring the proteins adsorbed to the particles. In some aspects, a subject is identified as being at risk of having pancreatic cancer by: having a computed tomography (CT) scan indicative of pancreatic cancer, having a magnetic resonance imaging (MRI) scan indicative of pancreatic cancer, having a positron emission tomography (PET) scan indicative of pancreatic cancer, having an ultrasound indicative of pancreatic cancer, having a pancreatogram indicative of pancreatic cancer, having an angiogram indicative of pancreatic cancer, having a liver function test (LFT) indicative of pancreatic cancer, having a carcinoembryonic antigen (CEA) level elevated relative to a control or baseline measurement, having a carbohydrate antigen (CA) 19-9 level elevated relative to a control or baseline measurement, having jaundice, having abdominal pain, having an enlarged gallbladder or liver, having a thrombus, or having a pancreatic cyst or a combination thereof. Some aspects include determining the likelihood that a subject has pancreatic cancer based on proteomics data. In some aspects, the classifier includes features to determine the likelihood that a subject has pancreatic cancer. In some aspects, the classifier is trained using deep learning, hierarchical clustering analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k-nearest neighbor analysis, naive Bayes analysis, K-means clustering analysis, or hidden Markov analysis. In some aspects, the proteomics data indicates pancreatic cancer with at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% sensitivity or specificity. Some aspects include recommending pancreatic cancer treatment to a subject when the proteomics data is classified as indicative of pancreatic cancer. Some aspects include administering pancreatic cancer treatment to a subject when the proteomics data is classified as indicative of pancreatic cancer. Some aspects include recommending or performing a biopsy when the proteomics data is classified as indicative of pancreatic cancer. Some aspects include recommending observing a subject without administering pancreatic cancer treatment to the subject or recommending observing a subject without obtaining a biopsy from the subject when the proteomics data is not classified as indicative of pancreatic cancer. Some aspects include observing a subject without administering pancreatic cancer treatment to the subject or observing a subject without obtaining a biopsy from the subject when the proteomics data is not classified as indicative of pancreatic cancer. In some aspects, pancreatic cancer treatment includes chemotherapy, radiation therapy, immunotherapy, targeted therapy, surgery, or surgical resection or a combination thereof.In some aspects, pancreatic cancer treatment includes administering a pharmaceutical composition comprising capecitabine, erlotinib, fluorouracil, gemcitabine, irinotecan, leucovorin, nab-paclitaxel, nanoliposomal irinotecan, oxaliplatin, olaparib, or larotrectinib, or a combination thereof. In some aspects, the particles include nanoparticles. In some aspects, the particles include lipid particles, metal particles, silica particles, or polymer particles. In some aspects, the particles include carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles. In some aspects, the particles comprise a group of physiochemically different nanoparticles. In some aspects, assaying a protein includes performing mass spectrometry, chromatography, liquid chromatography, high performance liquid chromatography, solid phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, western blot, dot blot, or immunostaining, or a combination thereof. In some aspects, assaying a protein includes performing mass spectrometry. In some aspects, assaying a protein includes measuring a reading indicative of the presence, absence, or amount of the protein. In some aspects, pancreatic cancer includes early-stage pancreatic cancer. In some aspects, pancreatic cancer includes late-stage pancreatic cancer. Some aspects include monitoring a subject and assaying a protein in a second biological fluid sample obtained from the subject at a later time. In some aspects, the protein includes a secreted protein. In some aspects, the biological fluid includes blood, plasma, or serum. In some aspects, the subject has pancreatic cancer. In some aspects, the subject does not have pancreatic cancer. In some aspects, the subject is a mammal. In some aspects, the subject is a human.
[0010] In some aspects, the present disclosure provides methods of treatment that include: identifying a mass in the pancreas of a subject; obtaining a biological fluid sample from the subject; contacting the biological fluid sample with particles such that the particles adsorb biomolecules containing proteins to the particles; assaying the biomolecules adsorbed to the particles to generate proteomic data; and classifying the proteomic data as indicative of the mass containing pancreatic cancer or not indicative of the mass containing pancreatic cancer. Some aspects include performing a biopsy on the mass when the proteomic data is classified as indicative of the mass containing pancreatic cancer and not performing a biopsy on the mass when the proteomic data is classified as not indicative of the mass containing pancreatic cancer. The mass can include a pancreatic cyst. The mass can be identified by medical imaging techniques such as CT scan or MRI. In some aspects, a subject is identified as being at risk of having pancreatic cancer by: having a computed tomography (CT) scan indicative of pancreatic cancer, having a magnetic resonance imaging (MRI) scan indicative of pancreatic cancer, having a positron emission tomography (PET) scan indicative of pancreatic cancer, having an ultrasound indicative of pancreatic cancer, having a cholangiogram indicative of pancreatic cancer, having an angiogram indicative of pancreatic cancer, having a liver function test (LFT) indicative of pancreatic cancer, having a carcinoembryonic antigen (CEA) level elevated relative to a control or baseline measurement, having a carbohydrate antigen (CA) 19-9 level elevated relative to a control or baseline measurement, having jaundice, having abdominal pain, having an enlarged gallbladder or liver, having a thrombus, or having a pancreatic cyst or a combination thereof. Some aspects include identifying the likelihood that the mass is cancerous based on the proteomic data. In some aspects, classifying the proteomic data as indicative of whether the mass is cancerous includes applying a classifier to the proteomic data. In some aspects, the classifier includes features to identify the likelihood that the mass is cancerous. In some aspects, the classifier is trained using deep learning, hierarchical clustering analysis, principal component analysis, partial least squares discriminant analysis, random forest classification analysis, support vector machine analysis, k-nearest neighbor analysis, naive Bayes analysis, K-means clustering analysis, or hidden Markov analysis. In some aspects, the proteomic data indicates that the mass is cancerous with at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% sensitivity or specificity. Some aspects include recommending a pancreatic cancer treatment to the subject when the proteomic data is classified as indicative of the mass being cancerous. Some aspects include administering a pancreatic cancer treatment to the subject when the proteomic data is classified as indicative of the mass being cancerous. Some aspects include recommending or performing a biopsy when the proteomic data is classified as indicative of the mass being cancerous. Some aspects include recommending observing the subject without administering a pancreatic cancer treatment to the subject or recommending observing the subject without obtaining a biopsy of the subject when the proteomic data is not classified as indicative of the mass being cancerous.Some aspects include observing a subject without administering pancreatic cancer treatment to the subject or without obtaining a biopsy from the subject when proteomic data is not classified as indicating that the mass is cancerous. In some aspects, pancreatic cancer treatment includes chemotherapy, radiotherapy, immunotherapy, targeted therapy, surgery, or surgical resection or a combination thereof. In some aspects, pancreatic cancer treatment includes administering a pharmaceutical composition comprising capecitabine, erlotinib, fluorouracil, gemcitabine, irinotecan, leucovorin, nab-paclitaxel, nanoliposomal irinotecan, oxaliplatin, olaparib, or larotrectinib or a combination thereof. In some aspects, the particles include nanoparticles. In some aspects, the particles include lipid particles, metal particles, silica particles, or polymer particles. In some aspects, the particles include carboxylate particles, polyacrylic acid particles, dextran particles, polystyrene particles, dimethylamine particles, amino particles, silica particles, or N-(3-trimethoxysilylpropyl)diethylenetriamine particles. In some aspects, the particles comprise a group of physiochemically different nanoparticles. In some aspects, assaying a biomolecule includes performing mass spectrometry, chromatography, liquid chromatography, high performance liquid chromatography, solid phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, western blot, dot blot, or immunostaining or a combination thereof. In some aspects, assaying a biomolecule includes performing mass spectrometry. In some aspects, assaying a biomolecule includes measuring a reading indicative of the presence, absence, or amount of the biomolecule. In some aspects, pancreatic cancer includes early stage pancreatic cancer. In some aspects, pancreatic cancer includes late stage pancreatic cancer. Some aspects include monitoring a subject and assaying a biomolecule in a second biofluid sample obtained from the subject at a later time. In some aspects, the protein includes a secreted protein. In some aspects, the biofluid includes blood, plasma, or serum. In some aspects, the mass is cancerous. In some aspects, the mass is not cancerous. In some aspects, the subject is a mammal. In some aspects, the subject is a human.
[0011] In some aspects, the present disclosure provides multi-omics cancer detection methods, which include: obtaining multi-omics data generated from one or more biofluid samples collected from a subject, the multi-omics data including first omics data and second omics data, where the first omics data includes a first omics data type, the first omics type including proteomics data, metabolomics data, transcriptomics data, or genomics data, and where the second omics data includes a second omics data type, the second omics data type being different from the first omics data type and including proteomics data, metabolomics data, transcriptomics data, or genomics data; using a first classifier to assign a first label corresponding to the presence, absence, or likelihood of pancreatic cancer to the first omics data; using a second classifier to assign a second label corresponding to the presence, absence, or likelihood of pancreatic cancer to the second omics data; and identifying the multi-omics data as indicating or not indicating pancreatic cancer based on the combination of the first label and the second label, where the first classifier and the second classifier are independent, and where the combination of the first label and the second label identifies the multi-omics data as indicating or not indicating pancreatic cancer with a higher accuracy than the first label or the second label alone. In some aspects, the first omics data type or the second omics data type includes proteomics data. In some aspects, the proteomics data includes measurements of at least 1000 proteins or peptides. In some aspects, the proteomics data is generated by contacting a biofluid sample in one or more biofluid samples with particles such that the particles adsorb biomolecules including proteins. In some aspects, the particles include metals, polymers, or lipids. In some aspects, the particles include a group of physiochemically different nanoparticles. In some aspects, the following are used to generate proteomics data: mass spectrometry, chromatography, liquid chromatography, high performance liquid chromatography, solid phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, western blot, dot blot, or immunostaining, or a combination thereof. In some aspects, genomics data or transcriptomics data is generated by: sequencing, microarray analysis, hybridization, polymerase chain reaction, electrophoresis, or a combination thereof. In some aspects, the first omics data type or the second omics data type includes transcriptomics data. In some aspects, the transcriptomics data includes mRNA or microRNA expression data. In some aspects, the first omics data type or the second omics data type includes genomics data. In some aspects, the genomics data includes DNA sequence data or epigenetic data. In some aspects, the epigenetic data includes DNA methylation data, DNA hydroxymethylation data, or histone modification data. In some aspects, the first omics data type or the second omics data type includes metabolomics data. Some aspects include identifying the multi-omics data as indicating or not indicating pancreatic cancer, including generating or obtaining a majority vote score based on the first label and the second label.In some aspects, identifying multi-omics data as indicative or non-indicative of pancreatic cancer includes generating or obtaining a weighted average of a first marker and a second marker. Some aspects include assigning weights to a first classifier and a second classifier to obtain the weighted average. In some aspects, the weights are assigned based on the area under the ROC curve, the area under the precision-recall curve, accuracy, precision, recall, sensitivity, F1 score, specificity, or a combination thereof. In some aspects, the first classifier and the second classifier independently err with respect to pancreatic cancer identification. Some aspects include sending or outputting a report that includes information regarding the identification. Some aspects include sending or outputting a recommendation for the treatment of pancreatic cancer in a subject based on the pancreatic cancer identification. In some aspects, pancreatic cancer is marked as indicative of pancreatic cancer with an accuracy characterized by a receiver operating characteristic (ROC) curve having an area under the curve (AUC) greater than 0.7, greater than 0.75, greater than 0.8, greater than 0.85, greater than 0.9, greater than 0.91, or greater than 0.92.
[0012] In some aspects, methods are disclosed herein for evaluating a subject suspected of having pancreatic cancer, comprising: measuring biomarkers in a biological fluid sample from the subject, wherein the biomarkers include A2GL, AKR1B1, ANPEP, ANTXR1, ANTXR2, BTK, CALR, CDH1, CDH11, CDH2, CDHR2, CILP2, CLEC3B, COL18A1, CRP, EXT1, F13A1, FAT1, FGL1, FLT4, ICAM1, IDH2, LCN2, LPP, MAPK1, MAP2K1, MYH9, NOTCH1, NOTCH2, PIGR, PPP2R1A, PRKAR1A, PXDN, RELN, RHOA, S100A8, S100A9, S100A12, SAA1, SAA2, SERPINA3, SLAIN2, SND1, SVEP1, TSP2, TUBB, TUBB1, or VCAN.
[0013] In some aspects, the present disclosure provides methods that include: measuring biomolecules in a biological fluid sample obtained from a subject suspected of having pancreatic cancer to obtain a biomolecule measurement; and based on the characteristics of the biomolecule measurement, identifying a protein measurement as indicative of the subject having pancreatic cancer or as not having pancreatic cancer by applying a classifier to the biomolecule measurement, wherein the classifier is characterized by a receiver operating characteristic (ROC) curve having an area under the curve (AUC) greater than 0.7, greater than 0.75, greater than 0.8, greater than 0.85, greater than 0.9, greater than 0.91, or greater than 0.92. In some aspects, the AUC is not greater than 0.75, not greater than 0.8, not greater than 0.85, not greater than 0.9, not greater than 0.91, not greater than 0.92, not greater than 0.93, or not greater than 0.94. In some aspects, the biomolecules include proteins, lipids, and metabolites.
[0014] In some aspects of the present inventive concept, the present disclosure provides methods for generating multi-omics classifiers that include: obtaining first omics data of a first omics data type, and obtaining second omics data of a second omics data type different from the first omics data type. The first omics data and the second omics data may correspond to biomolecules present in a biological sample of a subject. A first classifier of a biological state using the characteristics of the first omics data may be generated, and a second classifier of the biological state using the characteristics of the second omics data may also be generated. The method may further include assigning feature importance scores to the characteristics of the first classifier and the second classifier. The method may further include selecting top features of the first classifier, and selecting top features of the second classifier. Then, a combined classifier using the selected top features of the first classifier and the second classifier may be generated.
[0015] In some aspects, the method may include generating a first classifier using the characteristics of the first omics data, which includes using all available characteristics of the first omics data. In some aspects, the method may further include generating a second classifier using the characteristics of the second omics data, which includes using all available characteristics of the second omics data. In some aspects, the method may further include generating a first classifier using the characteristics of the first omics data, which includes performing machine learning with the characteristics of the first omics data. In some aspects, the method may further include generating a second classifier using the characteristics of the second omics data, which includes performing machine learning with the characteristics of the second omics data. In some aspects, generating a first classifier using the characteristics of the first omics data may include performing repeated cross-validation (RCV) with the characteristics of the first omics data. In some aspects, the method may further include generating a second classifier using the characteristics of the second omics data, which includes performing RCV with the characteristics of the second omics data. The method may further include the characteristics of the first omics data and the second omics data, which include measurements of biomolecules.
[0016] In some aspects, the method can include selected top features of a first classifier, where the first classifier includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 features. In some aspects, the method can further include selected top features of a second classifier, where the second classifier includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 features. In some aspects, the method can include selected top features of a first classifier, and where the first classifier includes the same number of features as the selected top features of the second classifier. In some aspects, the method can further include generating a combined classifier that includes performing RCV using the selected top features of the first classifier. In some aspects, the method can further include generating a combined classifier that includes performing RCV using the selected top features of the second classifier. In some aspects, the method further includes generating a combined classifier that includes using each selected feature from an initial omics model in a repeated and folded second RCV shuffling of subjects to new groups, the features from a first shuffling of subjects to RCV repeats and folds. In some aspects, the combined classifier can further include using another resampling method. The resampling method can be nested cross-validation (NCV) or leave-one-out cross-validation (LOOCV), etc. Any resampling method that constructs an error estimate for eventual generalization to help send out a test or validation set. In some aspects, the method can further include generating a combined classifier that includes excluding features below an importance threshold. In some aspects, the method can further include identifying features of the combined classifier below a predetermined importance threshold, and training a final combined classifier that excludes features below the predetermined importance threshold. In some aspects, the combined classifier includes a linear classifier. In some aspects, the first omics data and the second omics data are selected from proteomics data, metabolomics data, lipidomics data, transcriptomics data, and genomics data. In some aspects, the first omics data contains measurements of biomolecules captured by a first particle type, and the second omics data contains measurements of biomolecules captured by a second particle. The first particle and the second particle can be physiochemically different from each other. The first particle and the second particle can include lipid particles, metal particles, silica particles, or polymer particles. The first particle and the second particle can include nanoparticles.
[0017] In some aspects, the method may further include obtaining third omics data of a third omics data type corresponding to biomolecules present in a biological sample, generating a third classifier of a biological state using features of the third omics data, assigning feature importance scores to the features of the third classifier, and selecting top features of the third classifier. Then, possibly generating a combined classifier, including using the selected top features of the first classifier, the second classifier, and the third classifier. The method may further include obtaining fourth omics data of a fourth omics data type corresponding to biomolecules present in a biological sample, generating a fourth classifier of a biological state using features of the fourth omics data, assigning feature importance scores to the features of the fourth classifier, and selecting top features of the fourth classifier. Then, possibly generating a combined classifier, including using the selected top features of the first classifier, the second classifier, the third classifier, and the fourth classifier. In some aspects, the first omics, the second omics, the third omics, and the fourth omics are independently selected from proteomics data, metabolomics data, lipidomics data, transcriptomics data, and genomics data. In some aspects, the first omics data may include proteomics data, the second omics data may include metabolomics data, the third omics data may include lipidomics data, and the fourth omics data may include transcriptomics data. In some aspects, the combined classifier identifies a subject as having a biological state and as not having a biological state with at least 70% sensitivity and 99% specificity. In some aspects, the combined classifier identifies a subject as having a biological state and as not having a biological state with a performance characterized as: a receiver operating characteristic curve (ROC) having an area under the curve (AUC) of at least 0.95.
[0018] In some aspects, the biological state can include a disease. In some aspects, the disease can include cancer. In some aspects, the cancer can include pancreatic cancer. In some aspects, the pancreatic cancer can include pancreatic ductal adenocarcinoma. In some aspects, the cancer can include stage I cancer or stage II cancer. In some aspects, the cancer can include stage III cancer or stage IV cancer. In some aspects, the cancer can include stage I pancreatic cancer, stage II pancreatic cancer, stage III pancreatic cancer, or stage IV pancreatic cancer. In some aspects, the method can use a biological sample comprising a biofluid. In some aspects, the biofluid can include blood, serum, plasma, or a combination thereof. In some aspects, the biofluid can be substantially cell-free. In some aspects, a classifier generated using the methods described herein can be used to evaluate the biological state of a subject using biomolecular data obtained from a sample of the subject. Then, this evaluation can further include administering a disease treatment to the subject based on the evaluation. In some embodiments, the disease treatment can be any pancreatic cancer treatment disclosed herein. For example, in some embodiments, the disease treatment can include surgery, organ transplantation, administration of a pharmaceutical composition, radiation therapy, chemotherapy, immunotherapy, hormone therapy, monoclonal antibody therapy, stem cell transplantation, gene therapy, or administration of chimeric antigen receptor (CAR)-T cells or transgenic T cells. In some embodiments, the disease treatment can include chemotherapy, radiation therapy, immunotherapy, targeted therapy, surgery, or surgical resection, or a combination thereof. In some embodiments, the methods disclosed herein can recommend a disease treatment that includes administering a pharmaceutical composition comprising capecitabine, erlotinib, fluorouracil, gemcitabine, irinotecan, leucovorin, nab-paclitaxel, nanoliposomal irinotecan, oxaliplatin, olaparib, or larotrectinib, or a combination thereof.
[0019] In some aspects, the top features of the classifier can be selected from at least 500, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,500, at least 15,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000 features of the first omics data. In some aspects, the top-level features of the classifier can be selected from no more than 500, no more than 1,000, no more than 2,000, no more than 3,000, no more than 4,000, no more than 5,000, no more than 7,500, no more than 10,000, no more than 12,500, no more than 15,000, no more than 20,000, no more than 30,000, no more than 40,000, no more than 50,000, no more than 75,000, or no more than 100,000 features of the first omics data. Before selecting the top features, the number of features of the omics set can be defined by a range of any previous value. The top features can be selected from different numbers of features of each different omics dataset. The top features selected from multiple omics types for combining the classifier can be selected from the same number of features or different numbers of features. The number of features selected from each feature set of each omics dataset can be the same or different. Before selecting the top features, the difference in the number of features in each set of the omics dataset can be within a multiple of the number of features in another omics dataset. The number of features from the top features selected from the second omics dataset can be at least 1-fold, at least 10-fold, at least 100-fold, at least 1,000-fold, at least 10,000-fold, at least 100,000-fold of the number of features. It can be no more than 1-fold, no more than 10-fold, no more than 100-fold, no more than 1,000-fold, no more than 10,000-fold, no more than 100,000-fold of the number of features from the top features selected for the first omics dataset. The range of the number of features in each individual omics dataset can be within the ranges disclosed above. If there are multiple omics datasets, each omics dataset can have a range of the number of features obtained from the above ranges independent of any other omics dataset. The top feature classifier can have an omics dataset with the number of features independently selected from the above ranges, where there is no relationship between the number of features between any single omics datasets. They can also be the same number of features.
[0020] In some aspects, the inventive concept may include a classifier generated using multi-omics data obtained from samples of an intended test population. Omics groups may be selected to represent physiological, biological, genetic, or functional systems or structures within a patient. These systems or structures may be weighted based on their usefulness to provide predictive data for the classifier. These omics groups may be combined with or without prior knowledge weighting. The generated classifier may have improved classification ability in the true intended test population. The combination may be based on combinations of different variables of the data. The variables may be selected from feature importance scores, omics group weighting, cumulative relative importance, or cumulative relative classification ability. One variable may be selected, or a combination of variables may be used. Any combination of previous variables and any other variables may be used.
[0021] In some aspects, assigning a feature importance score to each feature may further include assigning one or more biological processes associated with the feature. In some aspects, the one or more biological processes may be human biological processes or gene ontology biological processes or any biological process related to the diagnosis or treatment of an altered biological state. In some aspects, selecting the top features may further include calculating the total number of biological processes of the top features. This may be done such that the combined classifier has at least a certain number of biological processes represented by the top features. This may result in a classifier that has higher sensitivity or specificity in the total population compared to a classifier that includes features representing fewer biological processes.
[0022] In some aspects, assigning one or more biological processes may further include calculating the significance of the association. The significance of the association may be calculated based on a formal test of statistical significance. The formal test of statistical significance may include a log odds ratio (LOR) calculation. The log odds ratio (LOR) calculation may include the equation:
[0023] LOR = ln((association of a specific process in the first omics data type / total association of all processes in the first omics data type - instances of the specific process in the first omics data type) / (association of the specific process in the second omics data type / total association of all processes in the second omics data type - instances of the specific process in the second omics data type))
[0024] And using Fisher's test for significance of difference in proportions and Bonferroni correction of the raw p-value. A positive LOR may indicate significance for the first omics data type, and a negative LOR may indicate significance for the second omics data type. In some aspects, an association of a feature with an omics data set for a process may be made only if the LOR is greater than 0.5 or less than -0.5 at a p-value < 0.05.
[0025] In some aspects, the subjects can include two separate groups, namely a first set of training subjects and a second set of training subjects. Generating the first classifier and the second classifier can use omics data corresponding to the biomolecules present in the biological samples of the first set of training subjects. In some aspects, generating the combined classifier can further include using omics data corresponding to the biomolecules present in the biological samples of the second set of training subjects.
[0026] In some aspects, the training data for generating the classifier can be divided into two groups, namely a first set of training data and a second set of training data. Only the first set of training data can be used to train a first omics type-specific all-features-in model. The first omics type-specific all-features-in model can be used for the purpose of important feature selection. The second set of training data can be used to generate a second final top feature combination model.
[0027] In some aspects, methods for detecting pancreatic cancer are disclosed herein, which include: (a) obtaining biomarkers from a biological fluid sample of a subject; and (b) applying a classifier to the biomarkers to evaluate pancreatic cancer, wherein the classifier distinguishes between biological fluid samples of subjects with and without pancreatic cancer with a performance characterized as follows: a receiver operating characteristic (ROC) curve with an average or median area under the curve (AUC) of at least 0.9; and wherein the biomarkers include any one of the following peptides: GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1), TFVIIPELVLPNR (SEQ ID NO.2), TFVIIPELVLPNR (SEQ ID NO.2), DSC(UniMod:4)TMRPSSLGQGAGEVWLR (SEQ ID NO.3), DNC(UniMod:4)PHLPNSGQEDFDK (SEQ ID NO.4), GLVLGAGWAEGYLR (SEQ ID NO.5), LVFNPDQEDLDGDGRGDIC(UniMod:4)K (SEQ ID NO.6), AFDLYFVLDK (SEQ ID NO.7), VFLVGNVEIR (SEQ ID NO.8), RVSPVGETYIHEGLK (SEQ ID NO.9), ASEQIYYENR (SEQ IDNO.10), VLPGGDTYMHEGFER (SEQ ID NO.11), AVDIPHMDIEALK (SEQ ID NO.12), AMGIMNSFVNDIFER (SEQ ID NO.13), MPEQEYEFPEPR (SEQ ID NO.14), SGVISDTELQQALSNGTWTPFNPVTVR (SEQ ID NO.15), M(UniMod:35)EDVNSNVNADQEVR (SEQ IDNO.16), VGHDYQWIGLNDK (SEQ ID NO.17), HAEC(UniMod:4)IYLGHFSDPMYK (SEQ ID NO.18) or NGIFWGTWPGVSEAHPGGYK (SEQ ID NO.19); any one of the following RNAs: ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, ENST00000424185.7, ENST00000652176.1, ENST00000392593.9, ENST00000532853.5, ENST00000429947.1,ENST00000580914.1, ENST00000368205.7, ENST00000531709.6, ENST00000524817.5, ENST00000651281.1, ENST00000499685.2, ENST00000311921.8, ENST00000472111.5, ENST00000585172.2, ENST00000287713.7 or ENST00000547687.2; any one of the following lipids: NEG_PC(18:2_20:5)+AcO, POS_DAG(18:1_20:0)+NH4, NEG_PE(O-16:0_22:6)-H, NEG_PC(18:2_20:3)+AcO, POS_CER(d18:1 / 18:0)+H, POS_CE(22:0)+NH4, NEG_PE(14:0_22:5)-H, NEG_PC(20:5_20:5)+AcO, POS_PE(P-18:0_18:3)+H, NEG_PE(O-16:0_20:3)-H, POS_CE(18:3)+NH4, NEG_PE(O-18:0_22:5)-H, NEG_PE(O-18:0_20:5)-H, POS_PE(P-20:0_20:3)+H, NEG_PE(O-16:0_20:2)-H, POS_CER(d18:1 / 24:0)+H, NEG_PA(20:1_20:3)-H, NEG_PA(20:0_20:5)-H, POS_CE(20:0)+NH4 or NEG_PC(16:1_20:3)+AcO; or any one of the following metabolites: NEG_AICAR POS_cystine, NEG_CMP, NEG_gentiobiose, POS_creatine, POS_imidazoleacetic acid, POS_inosine, NEG_n-isovalerylglycine, NEG_glucose-6-phosphate, POS_metanephrine, NEG_N-acetylglutamate, NEG_5-thymidylate (dTMP), POS_UMP, NEG_fructose-6-phosphate, NEG_cystine, POS_panthenol, POS_guanine, NEG_shikimic acid, POS_1-methylimidazoleacetate or POS_flavonoid 2. In some aspects, the biomarker comprises two or more of at least one peptide, at least one RNA, at least one lipid and at least one metabolite. In some aspects, the biomarker comprises three or more of at least one peptide, at least one RNA, at least one lipid and at least one metabolite. In some aspects, the biomarker comprises at least one peptide, at least one RNA,At least one lipid and at least one metabolite. In some aspects, the biomarker includes any one of the following peptides: GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1), TFVIIPELVLPNR (SEQ ID NO.2), TFVIIPELVLPNR (SEQ ID NO.2), DSC (UniMod:4)TMRPSSLGQGAGEVWLR (SEQ ID NO.3), DNC (UniMod:4)PHLPNSGQEDFDK (SEQ ID NO.4), GLVLGAGWAEGYLR (SEQ ID NO.5), LVFNPDQEDLDGDGRGDIC (UniMod:4)K (SEQ ID NO.6), AFDLYFVLDK (SEQ ID NO.7), VFLVGNVEIR (SEQ ID NO.8), RVSPVGETYIHEGLK (SEQ ID NO.9), ASEQIYYENR (SEQ IDNO.10), VLPGGDTYMHEGFER (SEQ ID NO.11), AVDIPHMDIEALK (SEQ ID NO.12), AMGIMNSFVNDIFER (SEQ ID NO.13), MPEQEYEFPEPR (SEQ ID NO.14), SGVISDTELQQALSNGTWTPFNPVTVR (SEQ ID NO.15), M (UniMod:35)EDVNSNVNADQEVR (SEQ IDNO.16), VGHDYQWIGLNDK (SEQ ID NO.17), HAEC (UniMod:4)IYLGHFSDPMYK (SEQ ID NO.18), or NGIFWGTWPGVSEAHPGGYK (SEQ ID NO.19). In some aspects, the biomarker includes 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, or 19 or more of the peptides. In some aspects, the biomarker includes any one of the following RNAs: ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, ENST00000424185.7, ENST00000652176.1, ENST00000392593.9,ENST00000532853.5, ENST00000429947.1, ENST00000580914.1, ENST00000368205.7, ENST00000531709.6, ENST00000524817.5, ENST00000651281.1, ENST00000499685.2, ENST00000311921.8, ENST00000472111.5, ENST00000585172.2, ENST00000287713.7, or ENST00000547687.2. In some aspects, the biomarker comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, seventeen or more, eighteen or more, nineteen or more, or twenty or more of the following in RNA. In some aspects, the biomarker comprises any one of the following lipids: NEG_PC(18:2_20:5)+AcO, POS_DAG(18:1_20:0)+NH4, NEG_PE(O-16:0_22:6)-H, NEG_PC(18:2_20:3)+AcO, POS_CER(d18:1 / 18:0)+H, POS_CE(22:0)+NH4, NEG_PE(14:0_22:5)-H, NEG_PC(20:5_20:5)+AcO, POS_PE(P-18:0_18:3)+H, NEG_PE(O-16:0_20:3)-H, POS_CE(18:3)+NH4, NEG_PE(O-18:0_22:5)-H, NEG_PE(O-18:0_20:5)-H, POS_PE(P-20:0_20:3)+H, NEG_PE(O-16:0_20:2)-H, POS_CER(d18:1 / 24:0)+H, NEG_PA(20:1_20:3)-H, NEG_PA(20:0_20:5)-H, POS_CE(20:0)+NH4, or NEG_PC(16:1_20:3)+AcO. In some aspects, the biomarker comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, seventeen or more, eighteen or more,19 or more, or 20 or more. In some aspects, the biomarker includes any one of the following metabolites: NEG_AICAR, POS_cystine, NEG_CMP, NEG_gentiobiolate, POS_creatine, POS_imidazoleacetic acid, POS_inosine, NEG_n-isovalerylglycine, NEG_glucose-6-phosphate, POS_metanephrine, NEG_N-acetylglutamate, NEG_5-thymidylate (dTMP), POS_UMP, NEG_fructose-6-phosphate, NEG_cystine, POS_panthenol, POS_guanine, NEG_shikimic acid, POS_1-methylimidazoleacetate or POS_flavonoid 2. In some aspects, the biomarker includes 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more or 20 or more of the metabolites. In some aspects, the classifier includes performance characterized by a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least 0.90. In some aspects, the subject is suspected of having pancreatic cancer. In some aspects, the method further comprises administering a pancreatic cancer treatment to the subject when the subject has pancreatic cancer. In some aspects, the method further comprises monitoring the subject when the subject does not have pancreatic cancer.
[0028] In some aspects, the present disclosure provides methods for treating pancreatic cancer, the methods comprising: administering a pancreatic cancer treatment to a subject having pancreatic cancer, wherein pancreatic cancer is evaluated by a method comprising: (a) obtaining a biomarker from a biological fluid sample of the subject; and (b) applying a classifier to the biomarker to evaluate pancreatic cancer, wherein the classifier distinguishes between biological fluid samples of subjects with and without pancreatic cancer with a performance characterized as: a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least 0.9; and wherein the biomarker comprises any one of the following peptides: GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1), TFVIIPELVLPNR (SEQ ID NO.2), TFVIIPELVLPNR (SEQ ID NO.2), DSC (UniMod:4)TMRPSSLGQGAGEVWLR (SEQ ID NO.3), DNC (UniMod:4)PHLPNSGQEDFDK (SEQ ID NO.4), GLVLGAGWAEGYLR (SEQ ID NO.5), LVFNPDQEDLDGDGRGDIC (UniMod:4)K (SEQ ID NO.6), AFDLYFVLDK (SEQ ID NO.7), VFLVGNVEIR (SEQ ID NO.8), RVSPVGETYIHEGLK (SEQ ID NO.9), ASEQIYYENR (SEQ ID NO.10), VLPGGDTYMHEGFER (SEQ ID NO.11), AVDIPHMDIEALK (SEQ ID NO.12), AMGIMNSFVNDIFER (SEQ ID NO.13), MPEQEYEFPEPR (SEQ ID NO.14), SGVISDTELQQALSNGTWTPFNPVTVR (SEQ ID NO.15), M (UniMod:35)EDVNSNVNADQEVR (SEQ ID NO.16), VGHDYQWIGLNDK (SEQ ID NO.17), HAEC (UniMod:4)IYLGHFSDPMYK (SEQ ID NO.18) or NGIFWGTWPGVSEAHPGGYK (SEQ ID NO.19); any one of the following RNAs: ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, ENST00000424185.7, ENST00000652176.1, ENST00000392593.9,ENST00000532853.5, ENST00000429947.1, ENST00000580914.1, ENST00000368205.7, ENST00000531709.6, ENST00000524817.5, ENST00000651281.1, ENST00000499685.2, ENST00000311921.8, ENST00000472111.5, ENST00000585172.2, ENST00000287713.7, or ENST00000547687.2; any one of the following lipids: NEG_PC(18:2_20:5)+AcO, POS_DAG(18:1_20:0)+NH4, NEG_PE(O-16:0_22:6)-H, NEG_PC(18:2_20:3)+AcO, POS_CER(d18:1 / 18:0)+H, POS_CE(22:0)+NH4, NEG_PE(14:0_22:5)-H, NEG_PC(20:5_20:5)+AcO, POS_PE(P-18:0_18:3)+H, NEG_PE(O-16:0_20:3)-H, POS_CE(18:3)+NH4, NEG_PE(O-18:0_22:5)-H, NEG_PE(O-18:0_20:5)-H, POS_PE(P-20:0_20:3)+H, NEG_PE(O-16:0_20:2)-H, POS_CER(d18:1 / 24:0)+H, NEG_PA(20:1_20:3)-H, NEG_PA(20:0_20:5)-H, POS_CE(20:0)+NH4, or NEG_PC(16:1_20:3)+AcO; or any one of the following metabolites: NEG_AICAR POS_cystine, NEG_CMP, NEG_gentiobiose, POS_creatine, POS_imidazoleacetic acid, POS_inosine, NEG_n-isovalerylglycine, NEG_glucose-6-phosphate, POS_metanephrine, NEG_N-acetylglutamate, NEG_5-thymidylate (dTMP), POS_UMP, NEG_fructose-6-phosphate, NEG_cystine, POS_pantothenol, POS_guanine, NEG_shikimic acid, POS_1-methylimidazoleacetate, or POS_flavonoid 2. In some aspects, the biomarker comprises two or more of at least one peptide, at least one RNA, at least one lipid, and at least one metabolite. In some aspects, the biomarker comprises three or more of at least one peptide, at least one RNA, at least one lipid, and at least one metabolite. In some aspects, the biomarker comprises at least one peptide, at least one RNA,At least one lipid and at least one metabolite. In some aspects, the biomarker includes any one of the following peptides: GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1), TFVIIPELVLPNR (SEQ ID NO.2), TFVIIPELVLPNR (SEQ ID NO.2), DSC (UniMod:4)TMRPSSLGQGAGEVWLR (SEQ ID NO.3), DNC (UniMod:4)PHLPNSGQEDFDK (SEQ ID NO.4), GLVLGAGWAEGYLR (SEQ ID NO.5), LVFNPDQEDLDGDGRGDIC (UniMod:4)K (SEQ ID NO.6), AFDLYFVLDK (SEQ ID NO.7), VFLVGNVEIR (SEQ ID NO.8), RVSPVGETYIHEGLK (SEQ ID NO.9), ASEQIYYENR (SEQ IDNO.10), VLPGGDTYMHEGFER (SEQ ID NO.11), AVDIPHMDIEALK (SEQ ID NO.12), AMGIMNSFVNDIFER (SEQ ID NO.13), MPEQEYEFPEPR (SEQ ID NO.14), SGVISDTELQQALSNGTWTPFNPVTVR (SEQ ID NO.15), M (UniMod:35)EDVNSNVNADQEVR (SEQ IDNO.16), VGHDYQWIGLNDK (SEQ ID NO.17), HAEC (UniMod:4)IYLGHFSDPMYK (SEQ ID NO.18), or NGIFWGTWPGVSEAHPGGYK (SEQ ID NO.19). In some aspects, the biomarker includes 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, or 19 or more of the peptides. In some aspects, the biomarker includes any one of the following RNAs: ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, ENST00000424185.7, ENST00000652176.1, ENST00000392593.9,ENST00000532853.5, ENST00000429947.1, ENST00000580914.1, ENST00000368205.7, ENST00000531709.6, ENST00000524817.5, ENST00000651281.1, ENST00000499685.2, ENST00000311921.8, ENST00000472111.5, ENST00000585172.2, ENST00000287713.7, or ENST00000547687.2. In some aspects, the biomarker comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, seventeen or more, eighteen or more, nineteen or more, or twenty or more of the RNAs. In some aspects, the biomarker comprises any one of the following lipids: NEG_PC(18:2_20:5)+AcO, POS_DAG(18:1_20:0)+NH4, NEG_PE(O-16:0_22:6)-H, NEG_PC(18:2_20:3)+AcO, POS_CER(d18:1 / 18:0)+H, POS_CE(22:0)+NH4, NEG_PE(14:0_22:5)-H, NEG_PC(20:5_20:5)+AcO, POS_PE(P-18:0_18:3)+H, NEG_PE(O-16:0_20:3)-H, POS_CE(18:3)+NH4, NEG_PE(O-18:0_22:5)-H, NEG_PE(O-18:0_20:5)-H, POS_PE(P-20:0_20:3)+H, NEG_PE(O-16:0_20:2)-H, POS_CER(d18:1 / 24:0)+H, NEG_PA(20:1_20:3)-H, NEG_PA(20:0_20:5)-H, POS_CE(20:0)+NH4, or NEG_PC(16:1_20:3)+AcO. In some aspects, the biomarker comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, seventeen or more, eighteen or more,19 or more, or 20 or more. In some aspects, the biomarker includes any one of the following metabolites: NEG_AICAR, POS_cystine, NEG_CMP, NEG_gentiobiolate, POS_creatine, POS_imidazoleacetic acid, POS_inosine, NEG_n-isovalerylglycine, NEG_glucose-6-phosphate, POS_metanephrine, NEG_N-acetylglutamate, NEG_5-thymidylate (dTMP), POS_UMP, NEG_fructose-6-phosphate, NEG_cystine, POS_panthenol, POS_guanine, NEG_shikimic acid, POS_1-methylimidazoleacetate or POS_flavonoid 2. In some aspects, the biomarker includes 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more or 20 or more of the metabolites., Brief Description of the Drawings
[0030] Figure 1 An exemplary method for generating and applying the classifier described herein is shown.
[0031] Figure 2 An example of the stages of screening and treatment of pancreatic cancer patients is shown.
[0032] Figure 3 A non-limiting example of a computing device is shown; in this case, the device has one or more processors, memory, storage, and a network interface.
[0033] Figure 4 A diagram showing a classifier and feature information according to some aspects described herein is shown.
[0034] Figure 5A The results of the Wilcox test for age comparison and the Fisher's exact test for gender ratio are shown.
[0035] Figure 5B The results of the Wilcox test for age comparison and the Fisher's exact test for gender ratio are shown.
[0036] Figure 6A The number of proteins detected in cross-subject samples in the analysis of biofluid samples from controls and cancer patients is shown.
[0037] Figure 6BShows the number of proteins detected across subject samples in the analysis of biological fluid samples from control and cancer patients.
[0038] Figure 6C Shows the reproducibility of the platform, which indicates the ability to detect biological signals. Analysis groups: C = control; S = sample. Left panel: Only proteins with n > 1 detections / analysis group are retained. For clarity, 2 features with CV > 300% out of 2,089 features are removed. Right panel: Only proteins with n > 1 detections / analysis group are retained. For clarity, 48 features with CV > 300% out of 7,672 features are removed.
[0039] Figure 6D Shows that more than 5,000 proteins were detected in a feasibility study of 212 subjects. For proteins present in > 25% of the samples, the median number of 4 peptides per protein was detected with the following search parameters: 0.1% peptide / protein FDR, default timsTOF parameters, using the complete UniProt human proteome database with contaminants (50% reverse decoy).
[0040] Figure 6E Shows that a large number of proteins can be reproducibly detected in samples. Individual nanoparticles generate complementary and common protein identifications. Unique proteomes are shown for each sample / particle + group grouped by sample and collection site.
[0041] Figure 6F Shows enhanced proteome coverage for detecting known cancer-related proteins. All detected matching proteins from the samples are plotted on the HPPP curve. GeneCards data uses scores reported from matching gene ids and the search term "cancer". The detected HPPP1 proteins cover a difference of 8 orders of magnitude: highest concentration: P00450 - ceruloplasmin; 830,000 ng / mL; and lowest concentration: Q7Z627 - E3 ubiquitin-protein ligase HUWE1; 0.0034 ng / mL.
[0042] Figure 6G Shows large-scale depth and efficient plasma proteomics.
[0043] Figure 6H Shows the quantitative performance of Proteograph applicable to large-scale studies.
[0044] Figure 6IShows the reproducibility of large-scale protein enrichment by Proteograph. The reproducibility of Proteograph enrichment is ideally suited for biomarker discovery. Data was collected over 191 enrichments of the same samples. The collection scope included 3 instruments; 3 cohort studies; 5 operators; 8 months of runtime; 121 plates; and 1500+ subject samples.
[0045] Figure 6J Shows the reproducibility of the platform over time (months) and instruments. The median MS1 peak area of the iRT peptides was less than 15% for all, and most were less than 10%.
[0046] Figure 6K Shows the application of the platform in pancreatic cancer biomarker discovery.
[0047] Figure 7A Shows a plot of some of the top proteins differentially detected in the biofluid samples from cancer patients relative to the biofluid samples from control patients.
[0048] Figure 7B Is a plot showing the distribution of OpenTargets (OT) scores. The OT scores (ranging from 0 to 0.8) are included on the x-axis, while the y-axis includes density (from 0 to 15).
[0049] Figure 8A Includes a plot showing the comparison of the median total signal by sample, analyte type, and category.
[0050] Figure 8B Shows box-and-whisker plots of the most significantly different analytes in each omics workflow ((i): lipids; (ii): metabolites; and (iii): proteins).
[0051] Figure 8C Shows the performance of an exemplary polymer classifier combining proteomics, lipidomics, and metabolomics measurements.
[0052] Figure 9A Includes a volcano plot of the intensity differences and P-values of proteins adsorbed to nanoparticles and detected in the biofluid samples from cancer patients relative to the biofluid samples from control patients. The volcano plot shows the magnitude of the differences on the x-axis and the significance on the y-axis, highlighting the most significant analytes.
[0053] Figure 9B Includes data for the top protein P35442 after a particle-based measurement method.
[0054] Figure 9CA volcano plot showing the intensity differences and P-values of proteins detected in a biological fluid sample from a cancer patient relative to that from a control patient. The volcano plot shows the magnitude of the differences on the x-axis and the significance on the y-axis, highlighting the most significant analytes.
[0055] Figure 9D Data of the top protein P01011 after proteomics measurement.
[0056] Figure 10A A volcano plot showing the intensity differences and P-values of lipids detected in a biological fluid sample from a cancer patient relative to that from a control patient. The volcano plot shows the magnitude of the differences on the x-axis and the significance on the y-axis, highlighting the most significant analytes.
[0057] Figure 10B Data of the top lipid CER(d18:1_18:0) after lipidomics measurement.
[0058] Figure 11A A volcano plot showing the intensity differences and P-values of metabolites detected in a biological fluid sample from a cancer patient relative to that from a control patient. The volcano plot shows the magnitude of the differences on the x-axis and the significance on the y-axis, highlighting the most significant analytes.
[0059] Figure 11B Data of the top metabolite AICAR after metabolomics measurement.
[0060] Figure 12A Depicts the classification of cancer samples and healthy samples by UMAP projection based on combined data.
[0061] Figure 12B Depicts the classification of cancer samples and healthy samples by PCA projection based on combined data.
[0062] Figure 12C Depicts the classification of cancer samples and healthy samples by UMAP projection based on Proteograph data.
[0063] Figure 12D Depicts the classification of cancer samples and healthy samples by PCA projection based on Proteograph data.
[0064] Figure 12E Depicts the classification of cancer samples and healthy samples by UMAP projection based on PiQuant data.
[0065] Figure 12FDepicts the classification of cancer and healthy samples by PCA projection based on PiQuant data.
[0066] Figure 12G Depicts the classification of cancer and healthy samples by UMAP projection based on lipid data.
[0067] Figure 12H Depicts the classification of cancer and healthy samples by PCA projection based on lipid data.
[0068] Figure 12I Depicts the classification of cancer and healthy samples by UMAP projection based on metabolite data.
[0069] Figure 12J Depicts the classification of cancer and healthy samples by PCA projection based on metabolite data.
[0070] Figure 13 Protein, lipid, and metabolite features included in the classifier.
[0071] Figure 14 Shows the classifier performance in a multi-omics study and includes the receiver operating characteristic (ROC) curve for disease state classification. The area under the curve (AUC) values are also included in the figure, with the 90% confidence interval in parentheses.
[0072] Figure 15A Shows the performance of a classifier trained with data from genomic assays and includes the ROC curve for disease state classification. The AUC values at the bottom of the figure are shown as ± values based on 90% confidence.
[0073] Figure 15B Shows the performance of a classifier trained with data from genomic assays ("Genomics"), a classifier trained with data from mass spectrometry assays ("Mass Spec"), and a classifier trained with data from both genomic and mass spectrometry assays ("Combined"). The data shown in the figure includes the ROC curve for disease state classification. The AUC values include ± values based on 90% confidence.
[0074] Figure 16A Shows a volcano plot that shows the intensity differences between pancreatic cancer samples and healthy samples.
[0075] Figure 16B Shows the study comparison groups (H: healthy; PC: pancreatic cancer). Among the 3,381 detected proteins, 124 were statistically significant.
[0076] Figure 17A Shows a volcano plot that shows the differential abundance of lipid species between pancreatic cancer samples and healthy samples.
[0077] Figure 17B Illustrates a graph showing top hit lipids based on Figure 17A the volcano plot in
[0078] Figure 17C Shows a volcano plot that shows the differential abundance of lipid species between pancreatic cancer samples and healthy samples.
[0079] Figure 17D Illustrates a graph showing top hit metabolites based on Figure 17C the volcano plot in
[0080] Figure 18A Shows the quantitative performance of Proteograph applicable to large-scale studies (e.g., the study in Example 7).
[0081] Figure 18B Shows the reproducibility of large-scale protein enrichment by Proteograph. The reproducibility of Proteograph enrichment is ideally suited for biomarker discovery. The system provides high-throughput, reproducible, and in-depth proteome coverage for new discoveries. The reproducibility by Proteograph enables quantitative, in-depth, non-targeted proteomic biomarker studies. The large-scale protein enrichment by Proteograph is highly reproducible ((NP1 = 0; NP2 = 0; NP3 = 2; NP4 = 0; and NP5 = 2).
[0082] Figure 19A Shows the evaluation of K562 precursor detection with SWATH and Zeno SWATH DIA. A minimum increase of 26% in precursor identification was detected using Zeno SWATH DIA. All data were generated from the pr and pg matrices output from DIA-NN (all quantified precursors and called proteins were identified). All data were searched in DIA-NN using "robust LC" and the SCIEX K562 spectral library.
[0083] Figure 19B Shows the evaluation of K562 precursor detection with SWATH and Zeno SWATH DIA. A minimum increase of 13% in protein group identification was detected using Zeno SWATH DIA. All data were generated from the pr and pg matrices output from DIA-NN (all quantified precursors and called proteins were identified). All data were searched in DIA-NN using "robust LC" and the SCIEX K562 spectral library.
[0084] Figure 20Shows improved sensitivity for increasing the number of detected low-abundance peptide species. Compared to SWATH, the detection of low-abundance peptides is improved with Zenon SWATH DI.
[0085] Figure 21 Shows plots generated from all eligible precursors. Data was searched in DIA-NN using "robust LV" and the SCIEX K562 spectral library.
[0086] Figure 22 Shows that the quantitative sensitivity increases with mass on SWATH and Zeno SWATH DIA. The Zeno SWATH DIA MS1 peak areas (K562) are distributed for lower-abundance peptides.
[0087] Figure 23A Shows that compared to individual SWATH acquisitions between different peptide injection masses based on all eligible precursors, the Zeno SWATCH DIA acquisition results in a higher amount of precursors based on K562 MS2. Data was searched in DIA-NN using "robust LC" and the SCIEX K562 spectral library.
[0088] Figure 23B Shows that compared to individual SWATCH acquisitions between different peptide injections based on all eligible precursors, the Zeno SWATH DIA acquisition results in a lower CV(5) of the K562 precursor levels. Data was searched in DIA-NN using "robust LC" and the SCIEX K562 spectral library.
[0089] Figure 24 Shows that when compared to SWATH MS / MS DIA acquisition, the Zeno Swatch DIA MS / MS acquisition results in 53%-85% more peptide identifications in the Proteograph generated from the pooled control samples.
[0090] Figure 25 Shows the set of 2,357 proteins across all five nanoparticles in a representative subject cohort. A set of 1,077 proteins was identified in at least 25% of the patient samples.
[0091] Figure 26A Shows a large number of proteins that can be reproducibly detected in the samples. Individual nanoparticles yield complementary and common protein identifications.
[0092] Figure 26B Shows improved sensitivity equivalent to detecting more low-abundance peptides in Proteograph peptide detection.
[0093] Figure 27 Illustrates the machine learning analysis protocol.
[0094] Figure 28A Illustrated the collection locations and subject enrollment dates of 184 subjects.
[0095] Figure 28B Illustrated the age and gender comparisons between pancreatic ductal adenocarcinoma (PDAC) and the control study groups.
[0096] Figure 29A Illustrated the distribution of protein values normalized by SIS of the samples.
[0097] Figure 29B Illustrated the distribution of the median values of protein value samples normalized by SIS by group.
[0098] Figure 30 Illustrated the outlier rejection analysis using MARLE.
[0099] Figure 31 Illustrated the volcano plot of Wilcoxon test values.
[0100] Figure 32 Illustrated the individual protein expression levels of the study subjects.
[0101] Figure 33 Illustrated the PCA multivariate analysis of the separability of the study groups.
[0102] Figure 34 Illustrated the unsupervised hierarchical clustering of protein data with two forced groups.
[0103] Figure 35 Illustrated the comparison between the subject training group and the validation group in the pancreatic cancer study.
[0104] Figure 36 Illustrated the results of the ANOVA of the ethnic tube based on the model parameter evaluation of XGBoost RCV.
[0105] Figure 37 Illustrated the combined ROC plot of 10x10 XGBoost RCV with the best hyperparameters.
[0106] Figure 38 Illustrated the evaluation of the GLMnet hyperparameter combinations in 10x10 RCV.
[0107] Figure 39A Illustrated the GLMnet top feature RCV ROC plot with the best hyperparameters.
[0108] Figure 39B Illustrated the GLMnet top feature final model coefficients.
[0109] Figure 40Illustrated the validation of the ROC plot of the final top feature GLMnet model.
[0110] Figure 41A Illustrated the CA19-9 levels in PDAC and the control group.
[0111] Figure 41B Illustrated the PDAC stage in individual cancer stages versus the CA19-9 levels in the control.
[0112] Figure 42 Illustrated the CA19-9 model performance in the validation subject group.
[0113] Figure 43A Illustrated the final model coefficients of the GLMnet combination.
[0114] Figure 43B Illustrated the RCV ROC plot of the GLMnet combination with the best hyperparameters.
[0115] Figure 44 Illustrated some validation details of the classifier based on the combined feature GLMnet.
[0116] Figure 45 Illustrated the comparison of the top feature OpenTargets scores with the database.
[0117] Figure 46 Included a ROC plot that illustrated the classifier performance in the stepwise analysis of biological fluid samples from subjects with pancreatic cancer.
[0118] Figure 47 Depicted the sample and analysis details in the multi-omics experiment for pancreatic cancer.
[0119] Figure 48 Included a volcano plot in the analysis of biological fluid samples from subjects with pancreatic cancer.
[0120] Figure 49 Included a heatmap of the results of biomarkers and the analysis of biological fluid samples from pancreatic cancer subjects.
[0121] Figure 50A Depicted the results of the variance decomposition of all samples in the analysis of biological fluid samples from subjects with pancreatic cancer.
[0122] Figure 50B Depicted the results of the variance decomposition of samples from cancer subjects in the pancreatic cancer analysis.
[0123] Figure 51 Was a Venn diagram showing the overlap of biomarkers in the analysis of biological fluid samples from subjects with pancreatic cancer.
[0124] Figure 52A - 52C Includes the following figures, which illustrate that multi-omics readings can also be statistically combined to improve the interpretation of biological processes and include features related to some biomarkers.
[0125] Figure 53 Includes a graph showing trend analysis, which shows where groups of biomarker types change similarly and are related to the degree based on cancer stage.
[0126] Figure 54A - 54D Shows PCA using all features and samples measured for each omics type. The ellipses show the 95% confidence intervals grouped for PDAC and control subjects, and the shapes represent the specific PDAC stage of the group. Figure 54A : Protein, Figure 54B : RNA, Figure 54C : Lipid, and Figure 54D : Metabolite.
[0127] Figure 55 Displays the combined top features, multi-omics GLMnet regression model coefficients. The resulting coefficients are plotted in order of decreasing magnitude from left to right. The features are annotated for the omics types of protein, RNA, lipid, and metabolite. The figure shows that no single omics type dominates the coefficients, with all 4 categories represented in the top 7 features of the experimental data.
[0128] Figure 56 Displays the CA19-9 levels measured in PDAC and control subjects.
[0129] Figure 57A -D Displays the plasma levels of 20 features from the multi-omics model measured in validation subjects. Figure 57A : Protein, Figure 57B : RNA, Figure 57C : Lipid, and Figure 57D : Metabolite.
[0130] Figure 58A - 58D Shows a volcano plot of the blood analyte features of each omics type measured in the subjects. Figure 58A : Peptide-nanoparticle features from Proteograph protein analysis, Figure 58B : RNA features mapped to ENST from RNAseq data, Figure 58C : Lipids from targeted MS data, and Figure 58D : Metabolite features from targeted MS data.
[0131] Figure 59A - 59DShows the feature importance scores of individual omics models. The top 20 features from each individual omics final XGBoost model are ranked by arbitrary feature importance units. Figure 59A : Peptide features with specific Proteograph nanoparticle modification sequence pairs. Figure 59B : RNA transcripts mapped to ENST. Figure 59C : Lipids using MS data acquisition ionization mode, annotated as positive (POS) or negative (NEG) lipids. Figure 59D : Metabolites with MS data acquisition ionization mode, annotated as positive (POS) or negative (NEG).
[0132] Figure 60A - 60B Shows the classification performance of individual omics models for differentiating PDAC from non-cancer controls in the validation cohort. The omics types used are proteins, RNA, lipids, and metabolites. Figure 60A : PDAC all-stage validation results including 26 PDAC subjects and 46 control subjects and Figure 60B : PDAC early stage (I / II) validation results (subset of all-stage results), including 7 PDAC subjects and 46 control subjects.
[0133] Figure 61 Shows a comparison of the predicted class probabilities of individual omics models.
[0134] Figure 62 Shows the top feature, multi-omics model classification performance of the combination for differentiating PDAC (all stages) from non-cancer samples compared to CA19-9 performance in the validation cohort. For the multi-omics model and the CA19-9 model, the performance in the validation subjects is plotted as an ROC curve. Annotate the ROC AUC with 95% confidence intervals.
[0135] Figure 63 Shows the counts of proteins and RNAs at various detection frequencies in PDAC study subjects. There are 3,215 unique Uniprot entries and 131,059 unique Ensembl ENST entries detected in at least 25% of the 146 subjects studied.
[0136] Figure 64 Shows the overlap of GOBP terms enumerated between proteomic types and transcriptomic types of features from PDAC tests. All but 66 of the 6,040 terms associated with at least one protein feature are represented in the transcriptomic type, and 5,966 of the 11,940 terms are transcript-specific.
[0137] Figure 65Shows the distribution of the ln odds ratio (LOR) values of proteins with GOBP terms in RNA, indicating the enrichment of a given GOBP term in one or another omics type from the PDAC study.
[0138] Figure 66 Shows the significance and magnitude of comparing the GOBP LOR of proteins and RNA. The raw p-values of the Fisher test [-log10(p-value)] were plotted against the magnitude of enrichment calculated as LOR. After Bonferroni correction, 40 features (23 for proteins and 17 for RNA) were significantly different (light gray dots). Annotated the GOBP names of the 20 most important features.
[0139] Figure 67 Shows the normalized expression levels of four different protein signatures in normal blood and PDAC blood.
[0140] Figure 68 Shows the normalized expression levels of different metabolite signatures in normal blood and PDAC blood.
[0141] Figure 69A - 69B Shows that different molecular assays capture analytes from different biological processes. Figure 69A Shows the biological processes captured by RNA-seq. Figure 69B Shows the biological processes captured by untargeted proteomics. Detailed Description of the Invention
[0143] The present disclosure provides non-invasive methods for detecting the presence of cancer such as pancreatic cancer or the risk of developing cancer in a subject. Identifying cancer in a subject at an early stage can protect the subject from further development of the cancer if treatment is provided at an early stage. Non-invasive tests can also be used to rule out the presence of cancer, thus protecting the subject from having to undergo invasive tests (such as biopsies), which can be painful and stressful or may pose a risk of harm to the subject.
[0144] Some insights from the study examples disclosed herein are that univariate analysis of individual omics has revealed multiple molecular markers that are statistically significantly different between cancer samples and non-cancer samples, unsupervised clustering of significantly associated cancer biomarkers has shown separation by disease state, decomposing variance into joint and individual components has shown that there can be some shared biological signals across omics types, but there is also unique biology for each individual omics type, gene set enrichment analysis has shown that multi-omics methods can reveal unique associations with disease biology, and trend analysis has shown multiple molecular markers across omics types related to cancer stage. Classifiers that can distinguish pancreatic cancer stages based on biological fluid samples from a subject are included herein.
[0145] Figure 1 Illustrates a non - limiting example (100) of a method for predicting whether a subject has cancer (such as pancreatic cancer) or is at risk of developing cancer (such as pancreatic cancer) based on the determination and analysis of a biological fluid sample obtained from the subject. The biological fluid sample can be any one of the biological fluids described herein or any combination of the biological fluids described herein. The sample can be directly analyzed to generate data (102), such as proteomic data; or the sample can be contacted with the particles described herein prior to the analysis at 102 to obtain adsorbed biomolecules (103). After obtaining data from the analysis at 102, additional analysis (103) can be performed on the sample obtained from 100 or 101 to obtain additional data sets, such as transcriptomic data, genomic data, metabolomic data, or combinations thereof. The data or data sets obtained from the analysis at 102 or 103 can then be used to generate a classifier (105), wherein the classifier can be used to identify the likelihood that the subject has cancer or is at risk of having cancer. The generation and application of the classifier can be further repeated and refined to improve the analysis and application of the classifier. Additionally, as Figure 1 shown in the analysis can be applied before or during the procedure included in Figure 2 , for example, early in the process before an invasive examination. With the current course of pancreatic cancer patients, one opportunity lies in screening high - risk patients before biopsy or pancreatoscopy. For example, the primary opportunities using the methods described herein include screening high - risk patients for early detection with increased accuracy and convenience. Another opportunity may lie in improving decision - making for imaging or biopsy procedures.
[0146] In some aspects, the cancer to be detected by the methods described herein can be pancreatic cancer. The pancreatic cancer can be early - stage pancreatic cancer. In other aspects, the pancreatic cancer can be late - stage pancreatic cancer. By generating data and identifying patterns in the data associated with cancer (such as pancreatic cancer), samples obtained non - invasively can be used for cancer diagnosis. The diagnosis of cancer can be improved by obtaining proteomic data. The diagnosis of cancer can be improved by combining multiple types of data (e.g., multiple data sets) in the analysis. For example, combining multiple data types including proteomics, transcriptomics, genomics, metabolomics, or combinations thereof can improve the accuracy of predicting whether a subject has cancer. In some aspects, the methods described herein include generating or obtaining data and using the data to predict whether a subject has or does not have cancer. Various ways of combining or analyzing data are described, and the use of the data for cancer assessment is further elaborated.
[0147] In some aspects, methods for detecting cancer can include additional screening or diagnostic methods such as computed tomography (CT) scans indicative of pancreatic cancer, magnetic resonance imaging (MRI) scans indicative of pancreatic cancer, positron emission tomography (PET) scans indicative of pancreatic cancer, ultrasounds indicative of pancreatic cancer, cholangiograms indicative of pancreatic cancer, angiograms indicative of pancreatic cancer, liver function tests (LFTs) indicative of pancreatic cancer, carcinoembryonic antigen (CEA) levels elevated relative to control or baseline measurements, carbohydrate antigen (CA) 19-9 levels elevated relative to control or baseline measurements, or combinations thereof. In some aspects, methods for detecting pancreatic cancer can include identifying symptoms in a subject such as jaundice, abdominal pain, gallbladder or liver enlargement, blood clots, digestive problems, or depression, or combinations thereof.
[0148] Classification methods can include any biomarker such as AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2, or APOA1. In some embodiments, the biomarker comprises two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, or 19 or each of AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2, or APOA1. Any of the aforementioned biomarkers can be used in a classifier to identify the presence of pancreatic cancer, rule out pancreatic cancer, or distinguish between pancreatic cancer and the absence of pancreatic cancer in a biological fluid sample from a subject suspected of having pancreatic cancer.
[0149] Subjects and Samples
[0150] The methods described herein can be used to identify subjects who may have cancer (such as pancreatic cancer) or are at risk of having cancer (such as pancreatic cancer). Cancer can include adenocarcinoma, such as pancreatic adenocarcinoma. The subject can be a vertebrate. The subject can be a mammal. The subject can be a human. The subject can be male or female. The subject can have cancer. The subject can not have cancer. The subject can have pancreatic cancer. The subject can not have pancreatic cancer. The subject can be at risk of having pancreatic cancer. For example, the subject can have a mass (such as a nodule or cyst) in the pancreas.
[0151] To identify cancer in a subject, a sample can be obtained from the subject. The subject can be suspected of having cancer or not having cancer. The method can be used to confirm or deny the suspected cancer.
[0152] The subject can be experiencing pancreatic cancer. The subject can have pancreatic cancer. The cancer can include pancreatic cancer. Pancreatic cancer can include early-stage pancreatic cancer. Pancreatic cancer can include late-stage pancreatic cancer. Pancreatic cancer can be stage 1 pancreatic cancer. Pancreatic cancer can be stage 2 pancreatic cancer. Pancreatic cancer can be stage 1 or 2. Pancreatic cancer can be stage 3 pancreatic cancer. Pancreatic cancer can be stage 4 pancreatic cancer. Pancreatic cancer can be stage 3 or 4. Pancreatic cancer can be stage 1, 2, 3, or 4. Pancreatic cancer can include pancreatic ductal adenocarcinoma (PDAC).
[0153] The data described herein can be generated from a sample of the subject. The sample can be a biofluid sample or a mass sample (e.g., an abnormal growth biopsied from the subject). Examples of biofluids include blood, serum, or plasma. The sample can include a blood sample. The sample can include a serum sample. The sample can include a plasma sample. Other examples of biofluids include urine, tears, semen, milk, vaginal fluid, mucus, saliva, or sweat.
[0154] A biofluid sample can be obtained from the subject. For example, a blood, serum, or plasma sample can be obtained from the subject by venipuncture. Other ways to obtain a biofluid sample include aspiration or swabbing.
[0155] The biofluid sample can be cell-free or substantially cell-free. To obtain a cell-free or substantially cell-free biofluid sample, the biofluid can be subjected to a sample preparation method such as centrifugation and precipitation removal.
[0156] A non-biofluid sample can be obtained from the patient. The sample can include a tissue sample. The tissue sample can include a pancreatic tissue sample. For example, the sample can include a mass taken from the pancreas of the subject that is suspected of being cancerous. The mass can include a pancreatic cyst. Before performing the methods described herein, a doctor can identify the cyst as a high-risk or low-risk cyst. The mass can be examined microscopically. The sample can include a cell sample. The sample can include a homogenate of cells or tissue. The sample can include the supernatant of a centrifuged homogenate of cells or tissue.
[0157] The sample (e.g., a biofluid or tissue sample) can be obtained from the subject at any stage during a screening procedure for diagnosing cancer. For example, a biofluid sample can be Figure 2Obtained before, during, or after any of the procedures described herein. A biological fluid sample can be obtained before or during the stage when the subject is a candidate for biopsy or pancreatoscopy for early detection of cancer. In other aspects, a biological fluid sample can be obtained before or during non-invasive examination, invasive examination, treatment, or monitoring stage.
[0158] Data generation
[0159] Proteomics data
[0160] The data described herein can include protein data or proteomics data. The methods disclosed herein can include obtaining data generated from one or more samples (such as biological fluid samples) collected from a subject. The data can include biomolecular measurements such as protein measurements, transcript measurements, genetic material measurements, or metabolite measurements. The data can include any one of the following omics data types: proteomics data, genomics data, transcriptomics data, or metabolomics data. This section includes some ways of generating each of these types of omics data. The methods of generating or analyzing omics data can also be applied to the methods of generating or analyzing individual biomolecules or sets of biomolecules. Other types of omics data can also be generated. The data can be labeled or identified as indicating pancreatic cancer or labeled or identified as not indicating pancreatic cancer.
[0161] Proteomics data can relate to data on proteins, peptides, or proteoforms. Proteomics data can include only peptides or proteins, or a combination of both. Examples of peptides are chains of amino acids. Examples of proteins are peptides or combinations of peptides. For example, a protein can include one, two, or more peptides bound together. A protein can also include any post-translational modification. A protein can be a secreted protein. Proteomics data can include data on various proteoforms. Proteoforms can include different forms of proteins produced from a genome with any kind of sequence variation, splice isoforms, or post-translational modifications.
[0162] Proteomics data can include information on the presence, absence, or amount of various proteins and peptides. For example, proteomics data can include the amount of a protein. The amount of a protein can be expressed as the concentration or quantity of the protein, such as the concentration of a protein in a biological fluid. The amount of a protein can be relative to another protein or another biomolecule. Proteomics data can include information on the presence of a protein or peptide. Proteomics data can include information on the absence of a protein or peptide. Proteomics data can be distinguished by subtypes, where each subtype includes different types of proteins, peptides, or proteoforms.
[0163] Proteomic data generally include data about many proteins or peptides. For example, proteomic data may include information about the presence, absence or amount of 1000 or more proteins or peptides. In some cases, proteomic data may include information about the presence, absence or amount of 5000, 10,000, 20,000 or more peptides, proteins or protein forms. Proteomic data may even include up to about 1,000,000 protein forms. Proteomic data may include a range of proteins, peptides or protein forms defined by any number of the aforementioned protein, peptide or protein form numbers.
[0164] Proteomic data can be generated by any of a variety of methods. Generating proteomic data can include using a detection reagent that binds to a peptide or protein and produces a detectable signal. After using a detection reagent that binds to a peptide or protein and produces a detectable signal, a reading indicating the presence, absence or amount of a protein or peptide can be obtained. Generating proteomic data can include concentrating, filtering or centrifuging a sample.
[0165] Proteomics data can be generated using mass spectrometry, chromatography, liquid chromatography, high performance liquid chromatography, solid phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, western blot, dot blot or immunostaining or a combination thereof. Some examples of methods for generating proteomics data include using mass spectrometry, protein chips or reversed-phase protein microarrays. Proteomics data can also be generated using immunoassays, such as enzyme-linked immunosorbent assay, western blot, dot blot or immunohistochemical assays. Generating proteomics data can involve using immunoassay panels.
[0166] A way to obtain proteomics data includes using mass spectrometry. The example of mass spectrometry includes using high-resolution two-dimensional electrophoresis to separate proteins from different samples in parallel, and then selecting or staining the differentially expressed proteins to be identified by mass spectrometry. Another method uses stable isotope labels to differentially mark proteins from two different complex mixtures. Proteins in complex mixtures can be isotopically labeled, and then digested to produce labeled peptides. Labeled mixtures can then be combined, and the peptides can be separated by multidimensional liquid chromatography and analyzed by tandem mass spectrometry. Mass spectrometry can include using liquid chromatography-mass spectrometry (LC-MS), i.e. a technology that can combine liquid chromatography (e.g., HPLC) with the physical separation ability of mass spectrometry.
[0167] In addition to any of the above methods, generating proteomics data can include contacting a sample with particles such that the particles adsorb biomolecules containing proteins. The adsorbed proteins can be part of a biomolecular corona. The adsorbed proteins can be measured or identified when generating proteomics data.
[0168] Some examples of proteins are shown in Figure 7A . Proteins that can be detected in the methods described herein include myosin-9 (MYH9), tubulin beta-1 chain (TUBB1), tubulin beta chain (TUBB), calreticulin (CALR), vascular endothelial growth factor receptor 3 (FLT4), neurogenic locus notch homolog protein 2 (NOTCH2), transforming protein RhoA (RHOA), isocitrate dehydrogenase [NADP], mitochondrial (IDH2), cadherin-1 (CDH1), cAMP-dependent protein kinase type I-alpha regulatory subunit (PRKAR1A), neurogenic locus notch homolog protein 1 (NOTCH1), exostosin-1 (EXT1), serine / threonine protein phosphatase 2A 65 kDa regulatory subunit A alpha isoform (PPP2R1A), staphylococcal nuclease domain-containing protein 1 (SND1), tyrosine-protein kinase BTK (BTK), lipoma preferred partner (LPP), mitogen-activated protein kinase (MAPK1), Fat1 protein (FAT1), cadherin 11 (CDH11), or dual-specificity mitogen-activated protein kinase 1 (MAP2K1). Another example of a protein is shown in Figure 9A - 9B . Proteins to be detected in the methods described herein can include thrombospondin-2 (TSP2 or P35442). Another example of a protein is shown in Figure 9C - 9D . Proteins to be detected in the methods described herein can include P01011. Some examples of proteins are shown in Figure 13Among them. The proteins to be detected in the methods described herein may include polymeric immunoglobulin receptor (PIGR, UniProt P01833), cadherin-related family member 2 (CDHR2, UniProt Q9BYE9), leucine-rich alpha-2-glycoprotein (LRG1 or A2GL, UniProt P02750), intercellular adhesion molecule 1 (ICAM1, UniProt P05362), aminopeptidase N (AMPN or ANPEP, UniProt P15144), thrombospondin-2 (TSP2, UniProt P35442), protein S100-A9 (S10A9 or S100A9, UniProt P06702), aldose reductase family 1 member B1 (ALDR or AKR1B1, UniProt P15121), serum amyloid A-1 protein (SAA1, UniProt P0DJI8), peroxidasin homolog (PXDN, UniProt P02742), protein S100-A8 (S10A8 or S100A8, UniProt P05109), anthrax toxin receptor 2 (ANTR2 or ANTXR2, UniProt P58335), cadherin-2 (CADH2 or CDH2, UniProt P19022), alpha-1-antichymotrypsin (AACT or SERPINA3, UniProt P01011), collagen alpha-1(XVIII) chain (COIA1 or COL18A1, UniProt P39060), fibrinogen-like protein 1 (FGL1, UniProt Q08830), protein S100-A12 (S10AC or S100A12, UniProt P80511), reelin (RELN, UniProt J3KQ66), C-reactive protein (CRP, UniProt P02711), versican core protein (CSPG2 or VCAN, UniProt P13611), coagulation factor XIII A chain (F13A or F13A1, UniProt P00488), cartilage intermediate layer protein 2 (CILP2, UniProt K7EPJ4), Sushi, von Willebrand factor type A, EGF and pentraxin domain-containing protein 1 (SVEP1, UniProt Q4LDE5), neutrophil gelatinase-associated lipocalin (NGAL or LCN2, UniProt P80188), tetranectin (TETN or CLEC3B, UniProt P05452), protein containing SLAIN motif 2 (SLAI2 or SLAIN2, UniProt Q9P270), anthrax toxin receptor 1 (ANTR1 or ANTXR1,UniProtQ9H6X2, such as isoform 5 [UniProt Q9H6X2-5], serum amyloid A-2 protein (SAA2, UniProt P0DJI9). Any number of the above proteins can be used. Any protein can be used in the classifier.,
[0169] The method can include measuring biomarkers in a biological fluid sample, where the biomarkers include A2GL, AKR1B1, ANPEP, ANTXR1, ANTXR2, BTK, CALR, CDH1, CDH11, CDH2, CDHR2, CILP2, CLEC3B, COL18A1, CRP, EXT1, F13A1, FAT1, FGL1, FLT4, ICAM1, IDH2, LCN2, LPP, MAPK1, MAP2K1, MYH9, NOTCH1, NOTCH2, PIGR, PPP2R1A, PRKAR1A, PXDN, RELN, RHOA, S100A8, S100A9, S100A12, SAA1, SAA2, SERPINA3, SLAIN2, SND1, SVEP1, TSP2, TUBB, TUBB1 or VCAN. In some aspects, the biomarker comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45 or 48 of the foregoing biomarkers, or a range of biomarkers determined by any two of the foregoing integers.
[0170] Any of the biomarkers in Table 5 can be used in the methods described herein, such as in methods for pancreatic cancer evaluation. Some such examples can include: alpha-1-antichymotrypsin, leucine-rich alpha-2-glycoprotein, alpha-1-antitrypsin, aminopeptidase N, lipopolysaccharide-binding protein, intercellular adhesion molecule 1, polymeric immunoglobulin receptor, protein S100-A8, complement C2, complement C5, complement C9, inter-alpha-trypsin inhibitor heavy chain H3, retinol-binding protein 4, low-affinity immunoglobulin gamma Fc region receptor III-A, C-reactive protein, tetranectin, Noelin, coagulation factor XIIIB chain, apolipoprotein A-II, or apolipoprotein A-I. The biomarker can include alpha-1-antitrypsin. The biomarker can include alpha-1-antichymotrypsin. The biomarker can include polymeric immunoglobulin receptor. The biomarker can include C-reactive protein. The biomarker can include leucine-rich alpha-2-glycoprotein. The biomarker can include complement C2. The biomarker can include serum amyloid A-1 protein. The biomarker can include serum amyloid A-2 protein. The biomarker can include inter-alpha-trypsin inhibitor heavy chain H3. The biomarker can include peptidase inhibitor 16. Any number or combination of these biomarkers can be used. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or a range defined by any two of the above integers can be included as biomarkers in the methods described herein, such as in cancer evaluation methods. Any one of these biomarkers can be used as a feature in a classifier, such as for classifying a biological fluid sample as indicative of cancer (e.g., pancreatic cancer) or not, or for ruling out the presence of cancer. Any one of these biomarkers can be measured in combination with an internal reference standard, such as a labeled form of the biomarker. Any one of these biomarkers can be combined with one or more other biomarkers, such as any of the biomarkers described herein.
[0171] In some cases, cancer assessment methods include using at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, or at least 19 of the following biomarkers from Table 5: AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2, or APOA1. In some cases, the biomarker comprises two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, or 19 or each of AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2, or APOA1. In some cases, it includes all of the following biomarkers: AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2, or APOA1. In some cases, cancer assessment methods include using no more than 1, no more than 2, no more than 3, no more than 4, no more than 5, no more than 6, no more than 7, no more than 8, no more than 9, no more than 10, no more than 11, no more than 12, no more than 13, no more than 14, no more than 15, no more than 16, no more than 17, no more than 18, no more than 19, or no more than 20 of the following biomarkers: AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2, or APOA1. Some methods use a subgroup of the biomarkers. For example, the subgroup can exclude A1AT. The subgroup can exclude A2G. The subgroup can exclude AACT. The subgroup can exclude AMPN. The subgroup can exclude APOA1. The subgroup can exclude APOA2. The subgroup can exclude CO2. The subgroup can exclude CO5. The subgroup can exclude CO9. The subgroup can exclude CRP.The subgroup can exclude F13B. The subgroup can exclude FCG3A. The subgroup can exclude ICAM1. The subgroup can exclude ITIH3. The subgroup can exclude LBP. The subgroup can exclude NOE1. The subgroup can exclude PIGR. The subgroup can exclude RET4. The subgroup can exclude S10A8. The subgroup can exclude TETN. Any number or combination of the following biomarkers can be used: AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2 or APOA1. Any number or combination of the following biomarkers can be used: AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2 or APOA1. In some cases, it includes all of the following biomarkers: AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2 or APOA1.
[0172] Examples of protein or peptide biomarkers that can be used in the methods described herein can include Figure 55 the biomarkers in. Examples of protein or peptide biomarkers that can be used in the methods described herein can include Figure 57A the biomarkers in. Examples of protein or peptide biomarkers that can be used in the methods described herein can include Figure 59A the biomarkers in. Figure 55 , Figure 57A or Figure 59A any combination or number of the biomarkers in can be useful. For example, Figure 55 , Figure 57A or Figure 59AAny biomarker therein can be used to distinguish between biological fluid samples of subjects with and without cancer (such as pancreatic cancer). In some cases, cancer evaluation methods include using at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, or at least 19 biomarkers from Table 19. In some cases, any one of the biomarkers can be selected from: GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1), TFVIIPELVLPNR (SEQ ID NO.2), TFVIIPELVLPNR (SEQ ID NO.2), DSC(UniMod:4)TMRPSSLGQGAGEVWLR (SEQ ID NO.3), DNC(UniMod:4)PHLPNSGQEDFDK (SEQ ID NO.4), GLVLGAGWAEGYLR (SEQ ID NO.5), LVFNPDQEDLDGDGRGDIC(UniMod:4)K (SEQ ID NO.6), AFDLYFVLDK (SEQ ID NO.7), VFLVGNVEIR (SEQ ID NO.8), RVSPVGETYIHEGLK (SEQ ID NO.9), ASEQIYYENR (SEQ ID NO.10), VLPGGDTYMHEGFER (SEQ IDNO.11), AVDIPHMDIEALK (SEQ ID NO.12), AMGIMNSFVNDIFER (SEQ ID NO.13), MPEQEYEFPEPR (SEQ ID NO.14), SGVISDTELQQALSNGTWTPFNPVTVR (SEQ ID NO.15), M(UniMod:35)EDVNSNVNADQEVR (SEQ ID NO.16), VGHDYQWIGLNDK (SEQ ID NO.17), HAEC(UniMod:4)IYLGHFSDPMYK (SEQ ID NO.18), or NGIFWGTWPGVSEAHPGGYK (SEQ ID NO.19). UniMod:4 represents an amino acid modified with an iodoacetamide derivative, and UniMod:35 represents an amino acid modified with methionine sulfoxide. Some methods use subgroups of the biomarkers.In some cases, the biomarker can comprise a first TFVIIPELVLPNR (SEQ ID NO.2) and a second TFVIIPELVLPNR (SEQ ID NO.2), each detected by a different particle. For example, in some cases, the first TFVIIPELVLPNR (SEQ ID NO.2) can be detected on a first nanoparticle (NP1), and the second TFVIIPELVLPNR (SEQ ID NO.2) can be detected on a second nanoparticle (NP2). In some cases, the biomarker can include GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1), TFVIIPELVLPNR (SEQ ID NO.2), TFVIIPELVLPNR (SEQ ID NO.2), DSC (UniMod:4)TMRPSSLGQGAGEVWLR (SEQ ID NO.3) or DNC (UniMod:4)PHLPNSGQEDFDK (SEQ ID NO.4) or a combination thereof. In some cases, any one of the biomarkers can be selected from LTBP2, CSHR2, FGFBP2 or THBS2. In some cases, the biomarker can include LTBP2. In some cases, the biomarker can include CSHR2. In some cases, the biomarker can include FGFBP2. In some cases, the biomarker can include THBS2. In some cases, any one of the biomarkers can be selected from Q14767-LTBP2, Q9BYE9-CDHR2, Q9BYJ0-FGFBP2 or P35442-TSP2. In some cases, the biomarker can include Q14767-LTBP2. In some cases, the biomarker can include Q9BYE9-CDHR2. In some cases, the biomarker can include Q9BYJ0-FGFBP2. In some cases, the biomarker can include P35442-TSP2. In some cases, the biomarker can contain GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1). In some cases, the biomarker can be an increase in GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1). In some cases, the biomarker can contain TFVIIPELVLPNR (SEQ ID NO.2). In some cases, the biomarker can be an increase in TFVIIPELVLPNR (SEQ ID NO.2). In some cases, the biomarker can contain a second TFVIIPELVLPNR (SEQ ID NO.2). The second TFVIIPELVLPNR (SEQ ID NO.2) can be detected by a different particle.In some cases, a biomarker can contain DSCTMRPSSLGQGAGEVWLR (SEQ ID NO.20). In some cases, DSCTMRPSSLGQGAGEVWLR (SEQ ID NO.20) can be modified with an iodoacetamide derivative. The iodoacetamide derivative can be attached to cysteine. In some cases, the biomarker can be a decrease in DSCTMRPSSLGQGAGEVWLR (SEQ ID NO.20). In some cases, a biomarker can contain DNCPHLPNSGQEDFDK (SEQ ID NO.21). In some cases, the biomarker can be an increase in DNCPHLPNSGQEDFDK (SEQ ID NO.21). DNCPHLPNSGQEDFDK (SEQ ID NO.21) can be modified with an iodoacetamide derivative. The iodoacetamide derivative can be attached to cysteine.
[0173] Any combination or number of the protein or peptide biomarkers described in this section or herein can be useful. For example, some biomarkers can be used to distinguish between biological fluid samples from subjects with and without cancer.
[0174] Transcriptomic data
[0175] The data described herein can include transcript data or transcriptomic data. Transcriptomic data can relate to data about nucleotide transcripts such as RNA. Examples of RNA include messenger RNA (mRNA), ribosomal RNA (rRNA), signal recognition particle (SRP) RNA, transfer RNA (tRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), long non-coding RNA (lncRNA), microRNA (miRNA), non-coding RNA (ncRNA), or piwi-interacting RNA (piRNA). RNA can include mRNA. RNA can include miRNA. Transcriptomic data can be distinguished by subtypes, where each subtype includes different types of RNA or transcripts. For example, mRNA data can be included in one subtype, and miRNA data can be included in another subtype.
[0176] Transcriptomic data can include information about the presence, absence, or amount of various RNAs. For example, transcriptomic data can include the amount of RNA. The amount of RNA can be indicated as the concentration or number of RNA molecules, such as the concentration of RNA in a biological fluid. The amount of RNA can be relative to another RNA or another biomolecule. Transcriptomic data can include information about the presence of RNA. Transcriptomic data can include information about the absence of RNA.
[0177] Transcriptomic data typically includes data on many RNAs. For example, transcriptomic data can include information on the presence, absence, or amount of 1000 or more RNAs. In some cases, transcriptomic data can include information on the presence, absence, or amount of 5000, 10,000, 20,000 or more RNAs. Transcriptomic data can even include up to about 200,000 transcripts. Transcriptomic data can include a range of transcripts defined by any of the foregoing numbers of RNAs or transcripts.
[0178] Transcriptomic data can be generated by any of a variety of methods. Generating transcriptomic data can include using a detection reagent that binds to RNA and produces a detectable signal. After using a detection reagent that binds to RNA and produces a detectable signal, a reading indicating the presence, absence, or amount of RNA can be obtained. Generating transcriptomic data can include concentrating, filtering, or centrifuging a sample.
[0179] Transcriptomic data can include RNA sequence data. Some examples of methods for generating RNA sequence data include using sequencing, microarray analysis, hybridization, polymerase chain reaction (PCR), or electrophoresis, or combinations thereof. Microarrays can be used to generate transcriptomic data. PCR can be used to generate transcriptomic data. PCR can include quantitative PCR (qPCR). Such methods can include using a detectable probe (e.g., a fluorescent probe) that is chimeric with double-stranded nucleotides or binds to a target nucleotide sequence. PCR can include reverse transcriptase quantitative PCR (RT-qPCR). Generating transcriptomic data can involve using a PCR panel.
[0180] RNA sequence data can be generated by sequencing the RNA of a subject or by first converting the RNA of the subject into DNA (e.g., complementary DNA (cDNA)) and sequencing the DNA. Sequencing can include massively parallel sequencing. Examples of massively parallel sequencing techniques include pyrosequencing, sequencing by reversible terminator chemistry, ligation-mediated sequencing-by-ligation, or phospholinked fluorescent nucleotides or real-time sequencing. Generating transcriptomic data can include preparing a sample or template for sequencing. Reverse transcriptase can be used to convert RNA into cDNA. Some template preparation methods include using an amplified template derived from a single RNA or cDNA molecule, or a single RNA or cDNA molecule template. Examples of amplification methods include emulsion PCR, rolling circle, or solid-phase amplification.
[0181] In addition to any of the above methods, generating transcriptomic data can include contacting a sample with particles such that the particles adsorb biomolecules comprising RNA. The adsorbed RNA can be part of a biomolecular corona. The adsorbed RNA can be measured or identified when generating transcriptomic data.
[0182] In some methods, RNA biomarkers that can be used in the methods described herein can include Figure 55 biomarkers in. Examples of RNA biomarkers that can be used in the methods described herein can include Figure 57A biomarkers in. Examples of RNA biomarkers that can be used in the methods described herein can include Figure 59B biomarkers in. Figure 55 , Figure 57A or Figure 59A Any combination or number of biomarkers in can be useful. For example, Figure 55 , Figure 57A or Figure 59AAny of the biomarkers can be used to distinguish between biological fluid samples from subjects with and without cancer, such as pancreatic cancer. In some cases, the cancer assessment method includes using at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, or at least 19 biomarkers from Table 20. In some cases, any one of the biomarkers can be selected from ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, ENST00000424185.7, ENST00000652176.1, ENST00000392593.9, ENST00000532853.5, ENST00000429947.1, ENST00000580914.1, ENST00000368205.7, ENST00000531709.6, ENST00000524817.5, ENST00000651281.1, ENST00000499685.2, ENST00000311921.8, ENST00000472111.5, ENST00000585172.2, ENST00000287713.7, or ENST00000547687.2. In some cases, the RNA can encode itchy E3 ubiquitin protein ligase, adenosine monophosphate deaminase 2, ADAM metallopeptidase domain 28, glycine N-acyltransferase-like protein 1 (GLYATL1), NAD(P)HX dehydratase, BICD cargo adaptor 1, phospholipase D family member 4, solute carrier family 27 member 3, long intergenic non-protein coding RNA 1237, nucleophosmin 11, protein tyrosine phosphatase receptor type K, nuclear RNA export factor 1, switch B cell complex subunit SWAP70, ERCC excision repair 5 endonuclease, BTG1 divergent transcript, zinc finger protein 507, galactose-1-phosphate uridylyltransferase, novel pseudogene, nicotinamide nucleotide adenylyltransferase 2, or charged multivesicular body protein 1A. Some methods use a subgroup of the biomarkers. In some cases, the biomarkers can include ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, or ENST00000424185.7 or a combination thereof.In some cases, a biomarker can include RNA encoded by a gene or protein such as AMPD2, ENSG00000078747, ADAM28, GLYATL1 pseudogene, or NAXD. In some cases, a biomarker can include AMPD2. In some cases, a biomarker can include ENSG00000078747. In some cases, a biomarker can include ADAM28. In some cases, a biomarker can include the GLYATL1 pseudogene. In some cases, a biomarker can include NAXD. In some cases, a biomarker can include RNA encoding proteins such as Q01433-AMPD2, Q9UKQ2-ADA28, or Q8IW45-NNRD. In some cases, a biomarker can include Q01433-AMPD2. In some cases, a biomarker can include Q9UKQ2-ADA28. In some cases, a biomarker can include Q8IW45-NNRD. In some cases, a biomarker can include ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, or ENST00000424185.7, or a combination thereof. In some cases, a biomarker can contain ENST00000483727.5. In some cases, a biomarker can be a decrease in ENST00000483727.5. In some cases, a biomarker can contain ENST00000531734.6. In some cases, a biomarker can be a decrease in ENST00000531734.6. In some cases, a biomarker can contain ENST00000437154.6. In some cases, a biomarker can be a decrease in ENST00000437154.6. In some cases, a biomarker can contain ENST00000531997.1. In some cases, a biomarker can be a decrease in ENST00000531997.1. In some cases, a biomarker can contain ENST00000424185.7. In some cases, a biomarker can be an increase in ENST00000424185.7.
[0183] Any combination or amount of the transcriptomic markers in this section or described herein can be useful. For example, some biomarkers can be used to distinguish between biological fluid samples from subjects with and without cancer.
[0184] Genomics data
[0185] The data described herein may include data regarding genetic material or genomic data. Genomic data may include data about genetic material such as nucleic acids or histones. Nucleic acids may include DNA. Genomic data may include information about the presence, absence, or amount of genetic material. The amount of genetic material may be indicated as a concentration, an absolute number, or may be relative.
[0186] Genomic data may include DNA sequence data. The sequence data may include gene sequences. For example, genomic data may include sequence data for up to about 20,000 genes. Genomic data may also include sequence data for non-coding DNA regions. DNA sequence data may include information about the presence, absence, or amount of a DNA sequence. DNA sequence data may include information about the presence or absence of mutations (such as single nucleotide polymorphisms). DNA sequence data may include DNA measurements of the amount of mutated DNA, such as measurements of mutated DNA from cancer cells.
[0187] Genomic data may include epigenetic data. Examples of epigenetic data include DNA methylation data, DNA hydroxymethylation data, or histone modification data. Epigenetic data may include DNA methylation or hydroxymethylation. DNA methylation or hydroxymethylation may be measured across the entire DNA or within regions of the DNA. Methylated DNA may include methylated cytosine (e.g., 5-methylcytosine). Cytosine is typically methylated at CpG sites and may indicate gene activation.
[0188] Epigenetic data may include histone modification data. Histone modification data may include the presence, absence, or amount of histone modification. Examples of histone modifications include serotonylation, methylation, citrullination, acetylation, or phosphorylation. Some specific examples of histone modifications may include lysine methylation, glutamine serotonylation, arginine methylation, arginine citrullination, lysine acetylation, serine phosphorylation, threonine phosphorylation, or tyrosine phosphorylation. Histone modifications may indicate gene activation.
[0189] Genomic data may be distinguished by subtypes, where each subtype includes different types of genomic data. For example, DNA sequence data may be included in one subtype and epigenetic data may be included in another subtype, or different types of epigenetic data may be included in different subtypes.
[0190] Genomic data can be generated by any of a variety of methods. Generating genomic data can include using detection reagents that bind to genetic material such as DNA or histones and produce a detectable signal. After using a detection reagent that binds to genetic material and produces a detectable signal, a reading indicating the presence, absence, or amount of the genetic material can be obtained. Generating genomic data can include concentrating, filtering, or centrifuging a sample.
[0191] Some examples of methods for generating DNA sequence data include using sequencing, microarray analysis (e.g., SNP microarray), hybridization, polymerase chain reaction, or electrophoresis, or combinations thereof. DNA sequence data can be generated by sequencing a subject's DNA. Sequencing can include massively parallel sequencing. Examples of massively parallel sequencing techniques include pyrosequencing, sequencing by reversible terminator chemistry, ligation-mediated sequencing-by-ligation, or phospholinked fluorescent nucleotides or real-time sequencing. Generating genomic data can include preparing a sample or template for sequencing. Some template preparation methods include using an amplified template derived from a single DNA molecule, or a single DNA molecule template. Examples of amplification methods include emulsion PCR, rolling circle, or solid-phase amplification.
[0192] DNA methylation can be detected using mass spectrometry, methylation-specific PCR, bisulfite sequencing, HpaII tiny fragment enrichment by ligation-mediated PCR assay, Gal hydrolysis and ligation adaptor-dependent PCR assay, chromatin immunoprecipitation (ChIp) assay combined with DNA microarray (ChIP-on-chip assay), restriction landmark genomic scanning, methylated DNA immunoprecipitation, pyrosequencing of bisulfite-treated DNA, molecular break light assay for DNA adenine methyltransferase activity, methylation-sensitive Southern blotting, methyl-CpG binding proteins, high-resolution melting analysis, methylation-sensitive single nucleotide primer extension assay, another methylation assay, or combinations thereof.
[0193] Histone modifications can be detected by using mass spectrometry or immunoassays, enzyme-linked immunosorbent assay, western blotting, dot blotting, or immunostaining, or combinations thereof.
[0194] In addition to any of the above methods, generating genomic data can include contacting a sample with particles such that the particles adsorb biomolecules containing genetic material. The adsorbed genetic material can be part of a biomolecular corona. The adsorbed genetic material can be measured or identified when generating genomic data.
[0195] Lipidomics data
[0196] The data described herein, such as omics data, may include lipid data or lipidomics data. Lipidomics data may include information about the presence, absence, or amount of various lipids. For example, lipidomics data may include the amount of lipids. The amount of lipids may be indicated as the concentration or quantity of lipids, such as the concentration of lipids in a biological fluid. The amount of lipids may be relative to another lipid or another biomolecule. Lipidomics data may include information about the presence of lipids. Lipidomics data may include information about the absence of lipids.
[0197] Many organisms contain complex lipid arrays (e.g., humans express over 600 lipids), and their relative expression can serve as powerful markers of biological states and health determinants. Lipids are diverse classes of biomolecules, which include fatty acids (e.g., long carbohydrates with a carboxylic acid (ester) tail group), diglycerides, triglycerides and polyglycerol esters, phospholipids, prenols, sterols (e.g., cholesterol), and hopanoids, among other types. While lipids are primarily found in membranes, free lipids, protein - complexed lipids, and nucleic - acid - complexed lipids are commonly present in a range of biological fluids and, in some cases, can be differentially fractionated from membrane - bound lipids. For example, lipid - binding proteins (e.g., albumin) can be collected from a sample by immunohistochemical precipitation and then chemically induced to release the bound lipids for subsequent collection and detection.
[0198] Lipids may be an integral component in the development of diseases such as cancer. For example, lipids may be key players in cancer biology as they may affect or participate in membrane feeding and cell proliferation, lipotoxicity (where lipid content balance may help prevent lipotoxicity), enhancing cellular processes, membrane biophysics, oncogenic signaling and metastasis, protection against oxidative stress, signaling in the microenvironment, or immunomodulation. Some lipid classes may be associated with cancer, such as glycerophospholipids in hepatocellular carcinoma, glycerophospholipids and acylcarnitines, increased choline - containing lipids and phospholipids during metastasis, or sphingolipid regulation of cancer cell survival and death.
[0199] Lipid data can be generated from a sample after the sample has been processed to separate or enrich lipids in the sample. Generating lipid data may include concentrating, filtering, or centrifuging the sample. Lipid analysis may include lipid fractionation. In many cases, lipids can be easily separated from other biomolecule types for lipid - specific analysis. Since many lipids have strong hydrophobicity, organic solvent extraction and gradient chromatography methods can cleanly separate lipids from other biomolecule types present within the sample. Mass spectrometry can be used to generate lipid data. Then, lipid analysis can distinguish lipids by class (e.g., distinguish sphingolipids and cholesteryl esters) or by individual type.
[0200] Lipidomics data can be generated by any of a variety of methods. Generating lipidomics data can include using detection reagents that bind to lipids and produce a detectable signal. After using a detection reagent that binds to lipids and produces a detectable signal, a reading indicating the presence, absence, or amount of lipids can be obtained. Generating lipidomics data can include concentrating, filtering, or centrifuging the sample.
[0201] Lipidomics data can be generated using mass spectrometry, chromatography, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, lateral flow assays, immunoassays, enzyme-linked immunosorbent assays, Western blotting, dot blotting, or immunohistochemistry, or combinations thereof. Examples of methods for generating lipidomics data include using mass spectrometry. Mass spectrometry can include separation method steps such as liquid chromatography analysis (e.g., HPLC). Mass spectrometry can include ionization methods such as electron ionization, atmospheric pressure chemical ionization, electrospray ionization, or secondary electrospray ionization. Mass spectrometry can include surface-based mass spectrometry or secondary ion mass spectrometry. Another example of a method for generating lipidomics data includes nuclear magnetic resonance (NMR). Other examples of methods for generating lipidomics data include Fourier transform ion cyclotron resonance, ion mobility spectrometry, electrochemical detection (e.g., in combination with HPLC), or Raman spectroscopy and radiolabeling (e.g., when combined with thin-layer chromatography). Some of the mass spectrometry methods described for generating lipidomics data can be used to generate proteomics data and vice versa. Lipidomics data can also be generated using immunoassays such as enzyme-linked immunosorbent assays, Western blotting, dot blotting, or immunohistochemistry. Generating lipidomics data can involve using a lipid panel.
[0202] In addition to any of the above methods, generating lipidomics data can include contacting the sample with particles such that the particles adsorb biomolecules containing lipids. The adsorbed lipids can be part of a biomolecular corona. The adsorbed lipids can be measured or identified when generating lipidomics data.
[0203] Lipids may have a biological association with diseases such as cancer. Lipids can include phospholipids. Examples of phospholipids include phosphatidylethanolamine (PE), phosphatidylcholine (PC), phosphatidylinositol (PI), or phosphatidylglycerol (PG). Some phospholipids are components of cell membranes and may play roles in cells (such as chemical energy storage, cell signaling, cell membranes, or cell-cell interactions within tissues). Lipids can include ceramides (CER). Ceramides can act as tumor suppressors and may be a therapeutic pathway for a target. For example, the efficacy of some chemotherapy and targeted therapies may be determined by ceramide levels. Lipids can include diacylglycerol (DAG). Lipids can include triacylglycerol (TAG). Lipids can include fatty acids (FA).
[0204] Examples of lipids are shown in Figure 10A - 10B . Lipids to be detected in the methods described herein may include CER(d18:1_10:0). Some examples of lipids are shown in Figure 13 . Lipids to be detected in the methods described herein may include CER(d18.1_18.0), PC(18.2_20.5), CER(d18.1_24.1), CER(d18.1_16.0), TAG(56.5_FA18.0), CER(d18.0_24.1), TAG(56.5_FA18.1), DAG(16.0_22.5), CER(d18.1_22.1), PE(P-18.0_18.3) or PE(17.0_22.6). Any number of the above lipids may be used. Any lipid can be used in the classifier.
[0205] Examples of lipid biomarkers that can be used in the methods described herein may include Figure 55 the biomarkers in. Examples of lipid biomarkers that can be used in the methods described herein may include Figure 57C the biomarkers in. Examples of lipid biomarkers that can be used in the methods described herein may include Figure 59C the biomarkers in. Figure 55 , Figure 57A or Figure 59A any combination or number of lipid biomarkers in may be useful. For example, Figure 55 , Figure 57A or Figure 59AAny of the biomarkers can be used to distinguish between biological fluid samples from subjects with and without cancer, such as pancreatic cancer. In some cases, the cancer assessment method includes using at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, or at least 19 of the biomarkers from Table 21. In some cases, any one of the biomarkers can be selected from PC(18:2_20:5)+AcO, DAG(18:1_20:0)+NH4, PE(O-16:0_22:6)-H, PC(18:2_20:3)+AcO, CER(d18:1 / 18:0)+H, CE(22:0)+NH4, PE(14:0_22:5)-H, PC(20:5_20:5)+AcO, PE(P-18:0_18:3)+H, PE(O-16:0_20:3)-H, CE(18:3)+NH4, PE(O-18:0_22:5)-H, PE(O-18:0_20:5)-H, PE(P-20:0_20:3)+H, PE(O-16:0_20:2)-H, CER(d18:1 / 24:0)+H, PA(20:1_20:3)-H, PA(20:0_20:5)-H, CE(20:0)+NH4, or PC(16:1_20:3)+AcO. Some biomarkers can be used or included based on the presence of lipid species. Some biomarkers can be used or included based on the absence of lipid substances. Some methods use a subgroup of the biomarkers. In some cases, the biomarkers can include the presence or absence of PC(18:2_20:5)+AcO, DAG(18:1_20:0)+NH4, PE(O-16:0_22:6)-H, PC(18:2_20:3)+AcO, CER(d18:1 / 18:0)+H, or a combination thereof. In some cases, the biomarker can contain PC(18:2_20:5)+AcO. In some cases, the biomarker can be a decrease in PC(18:2_20:5)+AcO. In some cases, the biomarker can contain DAG(18:1_20:0)+NH4. In some cases, the biomarker can be a decrease in DAG(18:1_20:0)+NH4. In some cases, the biomarker can contain PE(O-16:0_22:6)-H. In some cases, the biomarker can be a decrease in PE(O-16:0_22:6)-H. In some cases, the biomarker can contain PC(18:2_20:3)+AcO.In some cases, the biomarker can be a decrease in PC(18:2_20:3)+AcO. In some cases, the biomarker can contain CER(d18:1 / 18:0)+H. In some cases, the biomarker can be an increase in CER(d18:1 / 18:0)+H.
[0206] Any combination or amount of the lipid biomarkers described in this section or herein can be useful. For example, some biomarkers can be used to distinguish between biological fluid samples from subjects with and without cancer.
[0207] Metabolomics data
[0208] The data described herein can include metabolite data or metabolomics data. Metabolomics data can include information on small molecule (e.g., less than 1.5 kDa) metabolites such as metabolic intermediates, hormones or other signaling molecules, or secondary metabolites. Metabolomics data can relate to data on metabolites. Metabolites can include substrates, intermediates or products of metabolism. Metabolites can be any molecule less than 1.5 kDa in size. Examples of metabolites can include sugars, lipids, amino acids, fatty acids, phenolic compounds or alkaloids. Metabolomics data can be distinguished by subtype, where each subtype includes different types of metabolites. Metabolomics data can include some lipid data.
[0209] Metabolomics data can include information on the presence, absence or amount of various metabolites. For example, metabolomics data can include the amount of a metabolite. The amount of a metabolite can be indicated as the concentration or quantity of the metabolite, e.g., the concentration of a metabolite in a biological fluid. The amount of a metabolite can be relative to another metabolite or another biomolecule. Metabolomics data can include information on the presence of a metabolite. Metabolomics data can include information on the absence of a metabolite.
[0210] Metabolomics data generally includes data on many metabolites. For example, metabolomics data can include information on the presence, absence or amount of 1000 or more metabolites. In some cases, metabolomics data can include information on the presence, absence or amount of 5000, 10,000, 20,000, 50,000, 100,000, 500,000, 1 million, 1.5 million, 2 million or more metabolites or a range of metabolites defined by any two of the foregoing numbers of metabolites.
[0211] Metabolomics data can be generated by any of a variety of methods. Generating metabolomics data can include using a detection reagent that combines with a metabolite and produces a detectable signal. After using a detection reagent that combines with a metabolite and produces a detectable signal, a reading indicating the presence, absence or amount of a metabolite can be obtained. Generating metabolomics data can include concentrating, filtering or centrifuging samples.
[0212] Metabolomics data can be generated using mass spectrometry, chromatography, liquid chromatography, high performance liquid chromatography, solid phase chromatography, lateral flow assay, immunoassay, enzyme-linked immunosorbent assay, protein blotting, dot blotting or immunostaining or a combination thereof. Examples of methods for generating metabolomics data include the use of mass spectrometry. Mass spectrometry can include separation method steps, such as liquid chromatography analysis (e.g., HPLC). Mass spectrometry can include ionization methods, such as electron ionization, atmospheric pressure chemical ionization, electrospray ionization or secondary electrospray ionization. Mass spectrometry can include surface-based mass spectrometry or secondary ion mass spectrometry. Another example of the method for generating metabolomics data includes nuclear magnetic resonance (NMR). Other examples of the method for generating metabolomics data include Fourier transform ion cyclotron resonance, ion mobility spectrometry, electrochemical detection (e.g., coupled with HPLC) or Raman spectroscopy and radioactive labeling (e.g., when combined with thin layer chromatography). Some mass spectrometry methods described for generating metabolomics data can be used to generate proteomics data, and vice versa. Metabolomics data can also be generated using immunoassays such as enzyme-linked immunosorbent assays, western blotting, dot blotting or immunohistochemistry. Generating metabolomics data can involve the use of lipid panels.
[0213] In addition to any of the above methods, generating metabolomics data can include contacting a sample with particles so that the particles adsorb biomolecules containing metabolites. The adsorbed metabolites can be part of a biomolecule corona. The adsorbed metabolites can be measured or identified when generating metabolomics data.
[0214] Examples of metabolites are shown in Figure 11A - 11B In the methods described herein, metabolites to be detected may include 5-aminoimidazole-4-carboxamide ribonucleotide (AICAR). Metabolites may include nucleotides, such as monophosphate nucleotides. Some examples of metabolites are shown in Figure 13 In the methods described herein, the metabolite to be detected may include cytidine monophosphate (CMP). The metabolite may include AICAR or CMP. The metabolite to be detected may include AICAR and CMP. Any number of the aforementioned metabolites may be used. Any metabolite may be used in the classifier.
[0215] CA19-9 can be used as a biomarker. CA19-9 can be used alone or in combination with any other biomarker or group of biomarkers.
[0216] Examples of metabolite biomarkers that can be used in the methods described herein can include Figure 55 the biomarkers in. Examples of metabolite biomarkers that can be used in the methods described herein can include Figure 57D the biomarkers in. Examples of metabolite biomarkers that can be used in the methods described herein can include Figure 59D the biomarkers in. Figure 55 , Figure 57A or Figure 59A any combination or number of the metabolite biomarkers in can be useful. For example, Figure 55 , Figure 57A or Figure 59A any of the biomarkers in can be used to distinguish between biological fluid samples from subjects with and without cancer, such as pancreatic cancer. In some cases, cancer assessment methods include using at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, or at least 19 of the biomarkers from Table 22. In some cases, any biomarker can be selected from AICAR, cystine, CMP, gentisate, creatine, imidazoleacetic acid, inosine, N-isovalerylglycine, glucose-6-phosphate, normetanephrine, N-acetylglutamate, 5-thymidylate (dTMP), UMP, fructose-6-phosphate, cystine, panthenol, guanine, shikimic acid, 1-methylimidazoleacetate, or flavone 2. Some biomarkers can be based on the presence of metabolites. Some biomarkers can be based on the absence of metabolites. Some methods use subgroups of the biomarkers. In some cases, the biomarker can include the presence or absence of AICAR, cystine, CMP, gentisate, creatine, or a combination thereof. In some cases, the biomarker can include AICAR. In some cases, the biomarker can be a decrease in AICAR. In some cases, the biomarker can include cystine. In some cases, the biomarker can be an increase in cystine. In some cases, the biomarker can include CMP. In some cases, the biomarker can be a decrease in CMP. In some cases, the biomarker can include gentisate. In some cases, the biomarker can be a decrease in gentisate. In some cases, the biomarker can include creatine. In some cases, the biomarker can be a decrease in creatine.
[0217] Any combination or number of the metabolite markers described in this section or herein can be useful. For example, some biomarkers can be used to distinguish between biological fluid samples from subjects with and without cancer.
[0218] Use of reference biomolecules
[0219] In some aspects, obtaining proteomics data can include using reference biomolecules, which can be labeled. For example, prior to generating data, a sample can be contacted with a reference biomolecule. The data described herein can be generated using reference biomolecules. For example, a method can include contacting a sample with a reference biomolecule that contains a labeled form of each biomolecule such as each protein. The reference biomolecule can contain an internal standard. For example, the reference biomolecule can be added to a biological sample in a predetermined amount to act as an internal standard and help identify similar biomolecules endogenous to the sample. For example, an isotope-labeled reference protein can be spiked into a sample and measured together with endogenous proteins using mass spectrometry for identifying endogenous proteins on a mass spectrum and also for helping to determine the accurate amount of endogenous proteins. The internal standard can include a biomolecule added to a biological sample in a constant or known amount. The internal standard can include a non-endogenous labeled form of an endogenous biomolecule. Some examples refer to the use of internally labeled standards as "PiQuant".
[0220] Among the labeled biomolecules and endogenous biomolecules, individual labeled biomolecules can correspond to individual endogenous biomolecules. For example, the biomolecules can include proteins, and the endogenous proteins can include 100 - 1500 different proteins, and the labeled biomolecules can include the same 100 - 1500 proteins, but each labeled biomolecule can contain a label.
[0221] The reference biomolecule can include at least 5, at least 10, at least 50, at least 100, at least 250, at least 500, at least 750, at least 1000, at least 1500, at least 2000, at least 2500, at least 5000, at least 7500, at least 10,000, at least 15,000, at least 20,000, or at least 25,000 individual or different biomolecules. In some cases, the reference biomolecule includes fewer than 5, fewer than 10, fewer than 50, fewer than 100, fewer than 250, fewer than 500, fewer than 750, fewer than 1000, fewer than 1500, fewer than 2000, fewer than 2500, fewer than 5000, fewer than 7500, fewer than 10,000, fewer than 15,000, fewer than 20,000, or fewer than 25,000 individual or different biomolecules.
[0222] As an example, the sample contains endogenous protein A, endogenous protein B, and endogenous protein C. Endogenous protein A, endogenous protein B, and endogenous protein C are difficult to measure due to their low abundance. After spiking a predetermined amount of the isotopically labeled forms of protein A, protein B, and protein C into the sample, endogenous protein A, endogenous protein B, and endogenous protein C, as well as the isotopically labeled forms of protein A, protein B, and protein C, are analyzed together using mass spectrometry. Since the isotopically labeled forms are heavier, their mass spectra are shifted and can be distinguished from the mass spectra of the endogenous proteins. The isotopically labeled forms are more readily identifiable in the mass spectrometry readings, facilitating the identification of the mass spectra of endogenous protein A, endogenous protein B, and endogenous protein C in the mass spectrometry readings. Since a predetermined amount of isotopically labeled protein A, isotopically labeled protein B, and isotopically labeled protein C are added to the spiked sample, their concentrations are known, and the mass spectra of isotopically labeled protein A, isotopically labeled protein B, and isotopically labeled protein C can be used to accurately measure the amounts of endogenous protein A, endogenous protein B, and endogenous protein C from the mass spectrometry readings. The accurate measurement results of endogenous protein A, endogenous protein B, and endogenous protein C can be obtained by comparing the relative intensities of the mass spectrometry readings of endogenous protein A, endogenous protein B, and endogenous protein C with the intensities of the mass spectrometry readings of isotopically labeled protein A, isotopically labeled protein B, and isotopically labeled protein C of known concentration or amount.
[0223] Use of Particles
[0224] The sample can be contacted with the particles, for example, prior to generating data. The data described herein can be generated using the particles. For example, the method can include contacting the sample with the particles such that the particles adsorb biomolecules, such as proteins, transcripts, genetic material, or metabolites. The particles can attract a different set of biomolecules compared to situations where measurements are typically obtained directly on the sample. For example, a dominant biomolecule can represent a large percentage of certain types of biomolecules in the sample. For example, a protein can represent a large portion of the proteins in circulation collected by blood sampling. By having the biomolecules adhere to the particles prior to analyzing the biomolecules, a subset of biomolecules that does not include the dominant biomolecule can be obtained. Removing the dominant biomolecule in this manner can increase the accuracy of the biomolecule measurements and the sensitivity of the analysis using these measurements.
[0225] Biomolecules that can be adsorbed to the particles include proteins. The adsorbed biomolecules can form a biomolecular corona around the particles. The adsorbed biomolecules can be measured or identified when generating data (e.g., proteomics data).
[0226] The particles can be made of various materials. Such materials can include metals, magnetic materials, polymers, or lipids. The particles can be made of a combination of materials. The particles can contain layers of different materials. The different materials can have different properties. The particles can contain a core of one material and be coated with another material. The core and the coating can have different properties.
[0227] The particles can include metals. For example, the particles can include gold, silver, copper, nickel, cobalt, palladium, platinum, iridium, osmium, rhodium, ruthenium, rhenium, vanadium, chromium, manganese, niobium, molybdenum, tungsten, tantalum, iron, or cadmium, or a combination thereof.
[0228] The particles can be magnetic (e.g., ferromagnetic or ferrimagnetic). Particles containing iron oxide can be magnetic. The particles can be superparamagnetic iron oxide nanoparticles (SPION).
[0229] The particles can include polymers. Examples of polymers can include polyethylene, polycarbonate, polyanhydride, polyhydroxy acid, polypropyl fumarate, polycaprolactone, polyamide, polyacetal, polyether, polyester, poly(orthoester), polycyanoacrylate, polyvinyl alcohol, polyurethane, polyphosphazene, polyacrylate, polymethacrylate, polycyanoacrylate, polyurea, polystyrene, or polyamine, polyalkylene glycol (e.g., polyethylene glycol (PEG)), polyester (e.g., poly(lactide-co-glycolide) (PLGA), polylactic acid, or polycaprolactone), or a copolymer of two or more polymers, such as a copolymer of polyalkylene glycol (e.g., PEG) and polyester (e.g., PLGA). The particles can be made of a combination of polymers.
[0230] The particles may comprise lipids. Examples of lipids include dioleoyl phosphatidylglycerol (DOPG), diacyl phosphatidylcholine, diacyl phosphatidylethanolamine, ceramides, sphingomyelins, cephalins, cholesterol, cerebrosides, and diacylglycerol, dioleoyl phosphatidylcholine (DOPC), dimyristoyl phosphatidylcholine (DMPC), and dioleoyl phosphatidylserine (DOPS), phosphatidylglycerol, cardiolipin, diacyl phosphatidylserine, diacyl phosphatidic acid, N-lauroyl phosphatidylethanolamine, N-succinyl phosphatidylethanolamine, N-glutaroyl phosphatidylethanolamine, lysyl phosphatidylglycerol, palmitoyl oleoyl phosphatidylglycerol (POPG), lecithin, lysophosphatidylcholine, phosphatidylethanolamine, lysophosphatidylethanolamine, dioleoyl phosphatidylethanolamine (DOPE), dipalmitoyl phosphatidylethanolamine (DPPE), dimyristoyl phosphoethanolamine (DMPE), distearoyl phosphatidylethanolamine (DSPE), palmitoyl oleoyl phosphatidylethanolamine (POPE), palmitoyl oleoyl phosphatidylcholine (POPC), egg phosphatidylcholine (EPC), distearoyl phosphatidylcholine (DSPC), dioleoyl phosphatidylcholine (DOPC), dipalmitoyl phosphatidylcholine (DPPC), dioleoyl phosphatidylglycerol (DOPG), dipalmitoyl phosphatidylglycerol (DPPG), palmitoyl oleoyl phosphatidylglycerol (POPG), 16-O-monomethyl PE, 16-O-dimethyl PE, 18-1-trans PE, palmitoyl oleoyl phosphatidylethanolamine (POPE), 1-stearoyl-2-oleoyl phosphatidylethanolamine (SOPE), phosphatidylserine, phosphatidylinositol, sphingomyelin, cephalin, cardiolipin, phosphatidic acid, cerebroside, hexacosyl phosphate, or cholesterol. The particles may be made of a combination of lipids.
[0231] Further examples of materials include silica, carbon, carboxylate, polyacrylic acid, carbohydrates, dextran, polystyrene, dimethylamine, amines or silanes. Some examples of particles include carboxylate SPIONs, phenol-formaldehyde-coated SPIONs, silica-coated SPIONs, polystyrene-coated SPIONs, carboxylated poly(styrene-co-methacrylic acid) (P(St-co-MAA))-coated SPIONs, N-(3-trimethoxysilylpropyl)diethylenetriamine-coated SPIONs, poly(N-(3-(dimethylamino)propyl)methacrylamide) (PDMAPMA)-coated SPIONs, 1,2,4,5-benzenetetracarboxylic acid-coated SPIONs, poly(vinylbenzyltrimethylammonium chloride) (PVBTMAC)-coated SPIONs, peracetic acid-coated carboxylate particles, poly(oligo(ethylene glycol) methyl ether methacrylate) (POEGMA)-coated SPIONs, polystyrene carboxyl-functionalized particles, carboxylic acid particles, particles with an amino surface, silica amino-functionalized particles, particles with a Jeffamine surface or silica silanol-coated particles.
[0232] Particles of various sizes can be used. The particles can include nanoparticles. The diameter of the nanoparticles can be from about 10 nm to about 1000 nm. For example, the diameter of the nanoparticles can be at least 10 nm, at least 100 nm, at least 200 nm, at least 300 nm, at least 400 nm, at least 500 nm, at least 600 nm, at least 700 nm, at least 800 nm, at least 900 nm, from 10 nm to 50 nm, from 50 nm to 100 nm, from 100 nm to 150 nm, from 150 nm to 200 nm, from 200 nm to 250 nm, from 250 nm to 300 nm, from 300 nm to 350 nm, from 350 nm to 400 nm, from 400 nm to 450 nm, from 450 nm to 500 nm, from 500 nm to 550 nm, from 550 nm to 600 nm, from 600 nm to 650 nm, from 650 nm to 700 nm, from 700 nm to 750 nm, from 750 nm to 800 nm, from 800 nm to 850 nm, from 850 nm to 900 nm, from 100 nm to 300 nm, from 150 nm to 350 nm, from 200 nm to 400 nm, from 250 nm to 450 nm, from 300 nm to 500 nm, from 350 nm to 550 nm, from 400 nm to 600 nm, from 450 nm to 650 nm, from 500 nm to 700 nm, from 550 nm to 750 nm, from 600 nm to 800 nm, from 650 nm to 850 nm, from 700 nm to 900 nm, from 10 nm to 900 nm. The diameter of the nanoparticles can be less than 1000 nm. Some examples include a diameter of about 50 nm, about 130 nm, about 150 nm, 400 - 600 nm, or 100 - 390 nm.
[0233] The particle can include microparticles. The microparticles can be particles having a diameter ranging from about 1 μm to about 1000 μm. For example, the microparticles can be at least 1 μm, at least 10 μm, at least 100 μm, at least 200 μm, at least 300 μm, at least 400 μm, at least 500 μm, at least 600 μm, at least 700 μm, at least 800 μm, at least 900 μm, 10 μm to 50 μm, 50 μm to 100 μm, 100 μm to 150 μm, 150 μm to 200 μm, 200 μm to 250 μm, 250 μm to 300 μm, 300 μm to 350 μm, 350 μm to 400 μm, 400 μm to 450 μm, 450 μm to 500 μm, 500 μm to 550 μm, 550 μm to 600 μm, 600 μm to 650 μm, 650 μm to 700 μm, 700 μm to 750 μm, 750 μm to 800 μm, 800 μm to 850 μm, 850 μm to 900 μm, 100 μm to 300 μm, 150 μm to 350 μm, 200 μm to 400 μm, 250 μm to 450 μm, 300 μm to 500 μm, 350 μm to 550 μm, 400 μm to 600 μm, 450 μm to 650 μm, 500 μm to 700 μm, 550 μm to 750 μm, 600 μm to 800 μm, 650 μm to 850 μm, 700 μm to 900 μm or 10 μm to 900 μm. The diameter of the microparticles can be less than 1000 μm. Some examples include a diameter of 2.0 - 2.9 μm.
[0234] The particle can include a collection of physiochemically different particles (e.g., two or more collections of physiochemically different particles, where one collection is physiochemically different from another collection). Examples of physiochemical properties include charge (e.g., positive, negative or neutral) or hydrophobicity (e.g., hydrophobic or hydrophilic). The particle can include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more particle sets, or a range of any number of particle sets within the number of said particle sets.
[0235] Computer system
[0236] Certain aspects of the methods described herein can be performed using a computer system. For example, data analysis can be performed using a computer system. Similarly, multiple data sets can be obtained by using a computer system. Readings indicating the presence, absence, or amount of a biomolecule (e.g., a protein, transcript, genetic material, or metabolite) can be obtained at least in part using a computer system. A computer system can be used to perform methods of assigning a label corresponding to the presence, absence, or likelihood of a cancer state to data using a classifier, or of identifying multiple data sets as indicating or not indicating cancer. In certain aspects, the cancer is pancreatic cancer. The pancreatic cancer can be early-stage pancreatic cancer or late-stage pancreatic cancer. A computer system can generate a report identifying the likelihood that a subject has cancer. A computer system can transmit the report. For example, a diagnostic laboratory can transmit a report regarding cancer identification to a medical practitioner. A computer system can receive the report.
[0237] A computer system for performing the methods described herein can include Figure 3 some or all of the components shown in Figure 3 , the block diagram shown depicts an exemplary machine (e.g., a processing or computing system) including a computer system 300 within which a set of instructions can be executed to cause the device to perform or execute any one or more aspects and / or methods of the present disclosure for static code scheduling. Figure 3 The components in
[0238] are merely examples and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or combination of two or more such components for implementing a particular embodiment.
[0239] The computer system 300 includes one or more processors 301 that perform functions (e.g., a central processing unit (CPU) or a general-purpose graphics processing unit (GPGPU)). The one or more processors 301 optionally include cache memory units 302 for the temporary local storage of instructions, data, or computer addresses. The one or more processors 301 are configured to assist in the execution of computer-readable instructions. As a result of the one or more processors 301 executing non-transitory processor-executable instructions embodied in one or more tangible computer-readable storage media such as memory 303, storage 308, storage device 335, and / or storage medium 336, the computer system 300 can provide functionality for Figure 3 the components depicted in. The computer-readable medium can store software implementing a particular implementation, and the one or more processors 301 can execute the software. Memory 303 can read software from one or more other computer-readable media such as one or more mass storage devices 335, 336 or from one or more other sources via a suitable interface such as network interface 320. The software can cause the one or more processors 301 to execute one or more of the processes described or illustrated herein or one or more steps of one or more of the processes. Executing such a process or step can include defining data structures stored in memory 303 and modifying the data structures in accordance with the instructions of the software.
[0240] Memory 303 can include various components (e.g., machine-readable media), including but not limited to random access memory components (e.g., RAM 304) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phase change random access memory (PRAM), etc.), read-only memory components (e.g., ROM 305), and any combination thereof. ROM 305 can serve to communicate data and instructions unidirectionally to the one or more processors 301, and RAM 304 can serve to communicate data and instructions bidirectionally with the one or more processors 301. ROM 305 and RAM 304 can include any suitable tangible computer-readable media described below. In one instance, the basic input / output system 306 (BIOS) (including basic routines that assist in transferring information between elements within the computer system 300 such as during startup) can be stored in memory 303.
[0241] The fixed memory 308 is optionally bi-directionally connected to one or more processors 301 via a memory control unit 307. The fixed memory 308 provides additional data storage capacity and may also include any suitable tangible computer-readable medium described herein. The memory 308 can be used to store an operating system 309, one or more executable files 310, data 311, applications 312 (application programs), etc. The memory 308 may also include an optical disk drive, a solid-state memory device (e.g., a flash-based system), or any combination of the above. In appropriate cases, the information in the memory 308 can be incorporated into the memory 303 as virtual memory.
[0242] In one example, one or more storage devices 335 may be removably connected to the computer system 300 via a storage device interface 325 (e.g., via an external port connector (not shown)). In particular, the one or more storage devices 335 and associated machine-readable media can provide non-volatile storage and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 300. In one example, software may reside, in whole or in part, within the machine-readable medium on the one or more storage devices 335. In another example, software may reside, in whole or in part, within the one or more processors 301.
[0243] The bus 340 connects various subsystems. In this document, where appropriate, references to a bus may encompass one or more digital signal lines that serve a common function. The bus 340 can be any of several types of bus architectures, including but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof using any of a variety of bus architectures. By way of example and not limitation, such architectures can include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association Local Bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, an Accelerated Graphics Port (AGP) bus, a HyperTransport (HTX) bus, a Serial Advanced Technology Attachment (SATA) bus, or any combination thereof.
[0244] The computer system 300 may also include an input device 333. In one example, a user of the computer system 300 may input commands and / or other information into the computer system 300 via one or more input devices 333. Examples of one or more input devices 333 include, but are not limited to, alphanumeric input devices (e.g., keyboards), pointing devices (e.g., mice or touchpads), touchpads, touchscreens, multi-touch screens, joysticks, styli, gamepads, audio input devices (e.g., microphones, voice response systems, etc.), optical scanners, video or still image capture devices (e.g., cameras), and any combination thereof. In some aspects, the input device is a Kinect, Leap Motion, etc. One or more input devices 333 may be connected to the bus 340 via any input interface (e.g., input interface 323) among a variety of input interfaces 323 including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above.
[0245] In a particular embodiment, when the computer system 300 is connected to a network 330, the computer system 300 may communicate with other devices connected to the network 330, particularly mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, etc. Communications into and out of the computer system 300 may be sent through the network interface 320. For example, the network interface 320 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from the network 330, and the computer system 300 may store the incoming communications in the memory 303 for processing. The computer system 300 may similarly store outgoing communications (such as requests or responses to other devices) in the memory 303 in the form of one or more packets and communicate them from the network interface 320 to the network 330. One or more processors 301 may access these communication packets stored in the memory 303 for processing.
[0246] Examples of the network interface 320 include, but are not limited to, network interface cards, modems, and any combination thereof. Examples of the network 330 or network segment 330 include, but are not limited to, distributed computing systems, cloud computing systems, wide area networks (WANs) (e.g., the Internet, enterprise networks), local area networks (LANs) (e.g., networks associated with offices, buildings, campuses, or other relatively small geographical spaces), telephone networks, direct connections between two computing devices, peer-to-peer networks, or any combination thereof. Networks (such as network 330) may employ wired and / or wireless modes of communication. Generally, any network topology may be used.
[0247] Information and data can be displayed via the display 332. Examples of the display 332 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) (such as a passive matrix OLED (PMOLED) or an active matrix OLED (AMOLED) display), a plasma display, or any combination thereof. The display 332 can be connected via the bus 340 to one or more processors 301, the memory 303, the fixed memory 308, and other devices, such as one or more input devices 333. The display 332 is connected to the bus 340 via the video interface 322, and the data transfer between the display 332 and the bus 340 can be controlled via the graphics controller 321. In some aspects, the display is a video projector. In some aspects, the display is a head-mounted display (HMD), such as a VR head-mounted device. In further embodiments, by way of non-limiting example, suitable VR head-mounted devices include the HTC Vive, the Oculus Rift, the Samsung Gear VR, the Microsoft HoloLens, the Razer OSVR, the FOVE VR, the Zeiss VR One, the Avegant Glyph, the Freefly VR head-mounted device, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.
[0248] In addition to the display 332, the computer system 300 can include one or more other peripheral output devices 334, including but not limited to audio speakers, printers, storage devices, or any combination thereof. Such peripheral output devices can be connected to the bus 340 via the output interface 324. Examples of the output interface 324 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, or any combination thereof.
[0249] Additionally, or alternatively, the computer system 300 can provide functionality as a result of logic that is hard-wired or otherwise embodied in circuitry, which can replace software or operate in conjunction with software to perform one or more of the processes described or illustrated herein or one or more steps of one or more of the processes. References to software in this disclosure can cover logic, and references to logic can cover software. Additionally, where appropriate, references to computer-readable media can cover circuitry (such as an IC) that stores software for execution, circuitry that embodies logic for execution, or both. This disclosure covers any suitable combination of hardware, software, or both.
[0250] Those skilled in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described generally above in terms of their functionality.
[0251] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein can be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0252] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by one or more processors, or in a combination of both. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0253] According to the description herein, as non-limiting examples, suitable computing devices can include server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netpad computers, set-top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, tablet computers, personal digital assistants, video game consoles, and vehicles. Those skilled in the art will also recognize that selected televisions, video players, and digital music players with optional computer network connectivity are suitable for use in the systems described herein. In various embodiments, suitable tablet computers include those known to those skilled in the art with booklet, slate, and convertible configurations.
[0254] A computing device can include an operating system configured to execute executable instructions. The operating system is software that, for example, includes programs and data, manages the hardware of the device, and provides services for executing applications. Those skilled in the art will recognize that, as non-limiting examples, suitable server operating systems include FreeBSD, OpenBSD, Linux, Mac OS X Windows and Those skilled in the art will recognize that, as non-limiting examples, suitable personal computer operating systems include Mac OS and UNIX-like operating systems such as In some aspects, the operating system can be provided by cloud computing. Those skilled in the art will also recognize that, as non-limiting examples, suitable mobile smartphone operating systems include OS, Research In BlackBerry Windows OS, Windows OS, and
[0255] In some cases, the platforms, systems, media, or methods disclosed herein include one or more non-transitory computer-readable storage media encoded with a program including instructions executable by an operating system of a computer system. The computer system may be networked. The computer-readable storage medium may be a tangible component of a computing device. The computer-readable storage medium may be removable from the computing device. As non-limiting examples, the computer-readable storage medium may include any one of a CD-ROM, DVD, flash device, solid state memory, disk drive, tape drive, optical disk drive, distributed computing systems (including cloud computing systems and services), and the like. In some cases, the program and instructions are encoded on the medium permanently, substantially permanently, semi-permanently, or non-transitorily.
[0256] Data integration and analysis
[0257] When analyzing the data described herein (such as proteomic data, transcriptomic data, genomic data, or metabolomic data), the methods described herein may include generating or using a classifier for indicating that a subject has pancreatic cancer or is at risk of having pancreatic cancer with a certain sensitivity or specificity. In some aspects, the methods described herein generate or use a classifier from the data for indicating that a subject has pancreatic cancer or is at risk of having pancreatic cancer with a sensitivity of at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%. In some aspects, the methods described herein generate or use a classifier from the data for indicating that a subject has pancreatic cancer or is at risk of having pancreatic cancer with a specificity of at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%. In some aspects, the methods described herein generate or use a classifier from the data for indicating that a subject has pancreatic cancer or is at risk of having pancreatic cancer with a sensitivity or specificity of no greater than about 50%, no greater than about 60%, no greater than about 70%, no greater than about 80%, no greater than about 90%, or no greater than about 95%.
[0258] Separate data sets may be integrated into the analysis for more accurate cancer prediction or identification than may be provided by the separate data sets alone. For example, the method may include using more than one classifier to identify pancreatic cancer in a subject, where each classifier is used to analyze a separate data set and each classifier is independent of the others. When the classifiers err independently of each other, the combined analysis may be more accurate than an analysis using only one classifier corresponding to one data set. Alternatively, the separate data sets may be combined into one data set or analyzed by a single classifier.
[0259] Methods involving multiple classifiers can include using a first classifier to generate or assign a first label corresponding to the presence, absence, or likelihood of cancer for a first data set. The method can further include using a second classifier to generate or assign a second label corresponding to the presence, absence, or likelihood of cancer for a second data set. The method can further include using a third classifier to generate or assign a third label corresponding to the presence, absence, or likelihood of cancer for a third data set. The method can further include using a fourth classifier to generate or assign a fourth label corresponding to the presence, absence, or likelihood of cancer for a fourth data set. Additional classifiers can be used to generate or assign labels for further data sets. Each classifier can be trained using data from samples of subjects with cancer and data or combined data from samples of control subjects. Further, each classifier can include a stand-alone machine learning model or an ensemble of machine learning modules trained on the same input features.
[0260] Some classifiers can analyze a combined data set, while other classifiers can analyze only one data set. For example, additional classifiers can generate or assign labels corresponding to the presence, absence, or likelihood of cancer for a combined omics data set. The combined data set can include any combination of two or more data types or subtypes. For example, the data types can include proteomics data, transcriptomics data, genomics data, or metabolomics data. Each classifier can make a determination of cancer as Figure 4 shown.
[0261] The labels generated or assigned by each classifier can be used to identify data as indicating or not indicating cancer. This may involve picking the labels assigned by any one or more of the classifiers, or may involve generating or obtaining a majority vote score based on the first and second labels.
[0262] Identifying multiple data sets as indicating or not indicating cancer can include a majority vote across some or all of the labels generated by the classifiers. For example, the final determination of whether a subject is likely to have cancer can be made based on whether more classifiers assign labels corresponding to the presence of cancer or whether more classifiers assign labels corresponding to the absence of cancer. Identifying data as indicating or not indicating cancer can include generating or using a weighted average of some or all of the labels generated by the classifiers.
[0263] Identifying data as indicating or not indicating cancer can include obtaining or generating a weighted average of the labels generated or assigned by some or all of the classifiers. The weights for the weighted average can be based on one or more of the following: area under the ROC curve, area under the precision-recall curve, accuracy, precision, recall, sensitivity, F1 score, or specificity.
[0264] Methods involving multiple classifiers can include classifying data as indicative or not indicative of cancer. This can be done based on selecting markers assigned by individual classifiers or by combining markers assigned by multiple classifiers. The method can include classifying data as indicative or not indicative of cancer based on a combination of a first marker and a second marker, where the first marker and the second marker are each assigned by separate classifiers. The data can be further classified as indicative of cancer based on a third marker, a fourth marker, or one or more additional markers. The data can be classified as indicative of cancer based on the first and third markers or based on the first and fourth markers, where, for example, one or more of the markers are not included in the final determination.
[0265] Some aspects include using a classifier to identify the likelihood of pancreatic cancer. The classifier can be characterized by a receiver operating characteristic (ROC) curve having an area under the curve (AUC) greater than 0.7, greater than 0.75, greater than 0.8, greater than 0.85, greater than 0.9, greater than 0.91, greater than 0.92, greater than 0.93, greater than 0.94, greater than 0.95, greater than 0.96, greater than 0.97, greater than 0.98, or greater than 0.99 based on biomolecular measurement features. In some aspects, the AUC can be no greater than 0.75, no greater than 0.8, no greater than 0.85, no greater than 0.9, no greater than 0.91, no greater than 0.92, no greater than 0.93, no greater than 0.94, no greater than 0.95, no greater than 0.96, no greater than 0.97, no greater than 0.98, or no greater than 0.99.
[0266] Feature selection and simplified classifier
[0267] When creating a classifier related to a biological state, the methods described herein can include the following method: only select certain features of a data set corresponding to different types of biological data that can be collected. In some aspects, the method can include creating a classifier for each different data set separately. Then, each individual classifier can be used to assign a feature importance rating to each individual data point within each data set. This importance score reflects the usefulness of each individual feature within the set in creating the classifier. A score can be assigned to each feature of a separate data set. The biological state can be cancer, such as pancreatic cancer.
[0268] In some aspects, the individual features of each data set can be combined to create a simplified list of features. The selection can be based on importance scores. The selection can be based on the interaction between a feature and other features observed in the data set model. Each feature of each different data set can be selected. The total number of features selected to include in the simplified feature list can be 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more, 23 or more, 24 or more, 25 or more, 26 or more, 27 or more, 28 or more, 29 or more, 30 or more, 31 or more, 32 or more, 33 or more, 34 or more, 35 or more, 36 or more, 37 or more, 38 or more, 39 or more, 40 or more, 41 or more, 42 or more, 43 or more, 44 or more, 45 or more, 46 or more, 47 or more, 48 or more, 49 or more, 50 or more, 51 or more, 52 or more, 53 or more, 54 or more, 55 or more, 56 or more, 57 or more, 58 or more, 59 or more, 60 or more, 61 or more, 62 or more, 63 or more, 64 or more, 65 or more, 66 or more, 67 or more, 68 or more, 69 or more, 70 or more, 71 or more, 72 or more, 73 or more, 74 or more, 75 or more, 76 or more, 77 or more, 78 or more, 79 or more, 80 or more, 81 or more, 82 or more, 83 or more, 84 or more, 85 or more, 86 or more, 87 or more, 88 or more, 89 or more, 90 or more, 91 or more, 92 or more, 93 or more, 94 or more, 95 or more, 96 or more, 97 or more, 98 or more, 99 or more or 100 or more. The total number of selected features can be 20.The number of data sets from which features are selected can be 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more, 23 or more, 24 or more, 25 or more, 26 or more, 27 or more, 28 or more, 29 or more, 30 or more, 31 or more, 32 or more, 33 or more, 34 or more, 35 or more, 36 or more, 37 or more, 38 or more, 39 or more, 40 or more, 41 or more, 42 or more, 43 or more, 44 or more, 45 or more, 46 or more, 47 or more, 48 or more, 49 or more, 50 or more, 51 or more, 52 or more, 53 or more, 54 or more, 55 or more, 56 or more, 57 or more, 58 or more, 59 or more, 60 or more, 61 or more, 62 or more, 63 or more, 64 or more, 65 or more, 66 or more, 67 or more, 68 or more, 69 or more, 70 or more, 71 or more, 72 or more, 73 or more, 74 or more, 75 or more, 76 or more, 77 or more, 78 or more, 79 or more, 80 or more, 81 or more, 82 or more, 83 or more, 84 or more, 85 or more, 86 or more, 87 or more, 88 or more, 89 or more, 90 or more, 91 or more, 92 or more, 93 or more, 94 or more, 95 or more, 96 or more, 97 or more, 98 or more, 99 or more or 100 or more. The number of data sets from which features are selected can be 4. Features can be selected from data sets containing proteomics data, transcriptomics data, genomics data, lipidomics data or metabolomics data. The number of features selected from each individual data set can be the same, or it can also be different.The number of features selected from a single dataset can be 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more, 23 or more, 24 or more, 25 or more, 26 or more, 27 or more, 28 or more, 29 or more, 30 or more, 31 or more, 32 or more, 33 or more, 34 or more, 35 or more, 36 or more, 37 or more, 38 or more, 39 or more, 40 or more, 41 or more, 42 or more, 43 or more, 44 or more, 45 or more, 46 or more, 47 or more, 48 or more, 49 or more, 50 or more, 51 or more, 52 or more, 53 or more, 54 or more, 55 or more, 56 or more, 57 or more, 58 or more, 59 or more, 60 or more, 61 or more, 62 or more, 63 or more, 64 or more, 65 or more, 66 or more, 67 or more, 68 or more, 69 or more, 70 or more, 71 or more, 72 or more, 73 or more, 74 or more, 75 or more, 76 or more, 77 or more, 78 or more, 79 or more, 80 or more, 81 or more, 82 or more, 83 or more, 84 or more, 85 or more, 86 or more, 87 or more, 88 or more, 89 or more, 90 or more, 91 or more, 92 or more, 93 or more, 94 or more, 95 or more, 96 or more, 97 or more, 98 or more, 99 or more, or 100 or more. The number of features selected from a single dataset can be 5.
[0269] In some aspects, the simplified feature list can then be used to create a simplified model. The simplified model can be used to generate a simplified classifier. The simplified classifier can have the same predictive power as the classifier created using all features. The features of the simplified classifier can lie in a receiver operating characteristic (ROC) curve having an area under the curve (AUC) greater than 0.7, greater than 0.75, greater than 0.8, greater than 0.85, greater than 0.9, greater than 0.91, greater than 0.92, greater than 0.93, greater than 0.94, greater than 0.95, greater than 0.96, greater than 0.97, greater than 0.98 or greater than 0.99 based on biomolecular measurement features. In some aspects, the AUC can be no greater than 0.75, no greater than 0.8, no greater than 0.85, no greater than 0.9, no greater than 0.91, no greater than 0.92, no greater than 0.93, no greater than 0.94, no greater than 0.95, no greater than 0.96, no greater than 0.97, no greater than 0.98 or no greater than 0.99.
[0270] In some aspects, the creation of the simplified classifier can be based on separating the subjects into two training groups. The first training group can be used to generate a feature selection model. The model can be created by creating a model for each individual data set. Then, a classifier can be generated for each data set to create an individual classifier for each individual data set. Then the contribution of each feature to the overall classifier can be calculated. The model can be used to select features with high predictive power. Then, the features selected from this first model can be used to create a new feature list for data collection. Then data for the selected features can be collected from the second training group. Then the simplified data set can be used with a simplified second prediction model to create a simplified second classifier.
[0271] In some aspects, the creation of the simplified classifier can be based on a single training group. The training group can be used to generate a feature selection model. The model can be created by creating a model for each individual data set. Then, a classifier can be generated for each data set to create an individual classifier for each individual data set. Then the contribution of each feature to the overall classifier can be calculated. The model can be used to select features with high predictive power. Then, the features selected from this first model can be used to create a new feature list for data collection. Then data for the selected features can be separated from the training group. Then the simplified data set can be used with a simplified second prediction model to create a simplified second classifier. This may require overfitting the model from the common data set so that an appropriate confidence in the classifier can be calculated.
[0272] Many methods can be implemented to mitigate and assess the risk of overfitting. First, a conservative split of the total subject population can be selected, where approximately 95%, 90%, 85%, 80%, 75%, 70%, 65%, or 60% is in the training set and 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% is in the validation set. By increasing the size of the validation set, the ability to detect overfitting (if it occurs) increases, even though the ability to identify important classifier components decreases. Second, the study design can incorporate intentional differences between the test and control groups in terms of the enrollment date and enrollment location. These steps reduce the risk of systematic bias between groups that carries over from the training set to the validation set. Third, a wide cross-validation design can be employed when optimizing the model engine parameters and important feature selection. Ten rounds of 10-fold cross-validation, while computationally expensive for many input features, is a robust method for avoiding overfitting. 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 rounds of 10-fold cross-validation can be used. Finally, randomly permuted groups of training subject data (e.g., test versus control) can determine whether the two-stage process described herein results in a final multi-omics model that has similar performance in the validation set as that observed in the individual omics models.
[0273] Based on all data features, a simplified classifier can present advantages over a complex classifier without loss of predictive power. The simplified classifier can process subject data more quickly. It can process an individual data set in 1 second (s), 10 s, 20 s, 30 s, 40 s, 50 s, 60 s, 2 minutes (min), 3 min, 4 min, 5 min, 6 min, 7 min, 8 min, 9 min, or 10 min. It can also allow point-of-care processing. It can allow processing on less expensive or complex computer systems. It can allow processing to be completed via a web application or on a smart phone. It can allow cloud-based or remote processing. The simplified classifier can allow for easier conversion into a clinical diagnosis. This can be facilitated by reducing the number of features that must be tested for the classifier to operate. This can be facilitated by increasing the confidence of the end user. This can be facilitated by reducing the complexity of the validation process required by regulatory agencies for clinical diagnosis development. The simplified classifier can be easier to modify than a complex classifier. It can allow for greater manipulation to optimize the performance of the classifier. It can be based on a simpler model. The model can be a linear regression model. The classifier can be easier to understand. It can allow for easier explanation to an individual without training in the development of computer models and classifiers.
[0274] The selected features can be the most important within their respective dataset models. When compared across all datasets used to construct the simplified feature list, they can be the most important. However, in some respects, the selected features may not be the most important among all features. Selecting features from a larger number of datasets rather than from individual most important preferences can be preferred. This type of feature selection can provide greater prediction accuracy relative to the total population. This can allow the generation of a classifier with high prediction accuracy using a small training group. The training group can include 1,000 people, 900 people, 800 people, 700 people, 600 people, 500 people, 400 people, 300 people, 250 people, 200 people, 150 people, 100 people, 75 people, 50 people, or 25 people, or a range defined by any two of the above numbers. The training group can include fewer than 1,000 people, fewer than 900 people, fewer than 800 people, fewer than 700 people, fewer than 600 people, fewer than 500 people, fewer than 400 people, fewer than 300 people, fewer than 250 people, fewer than 200 people, fewer than 150 people, fewer than 100 people, fewer than 75 people, fewer than 50 people, or fewer than 25 people. The training group can include at least 1,000 people, at least 900 people, at least 800 people, at least 700 people, at least 600 people, at least 500 people, at least 400 people, at least 300 people, at least 250 people, at least 200 people, at least 150 people, at least 100 people, at least 75 people, at least 50 people, or at least 25 people.
[0275] Biological process coverage
[0276] In some aspects of the inventive concept, features can be selected to extend the coverage of different biological processes. A biological process can be any function that an organism experiences. The process can exist under normal function, or it can exist when the biological state of the organism is interrupted or disturbed. The cause of the interruption or disturbance can be endogenous or exogenous. It can be a disease. The disease can be cancer. The cancer can be pancreatic cancer. When the biological process is affected, the feature can be upregulated. When the biological process is affected, the feature can be downregulated.
[0277] In some aspects, coverage means that a feature relates to a biological process, is associated with a biological process, is affected by a biological process, or otherwise has some relationship with a biological process. This relationship can be known, or it can be determined after selection. A biological process can be covered by a single feature or multiple features. A feature can provide coverage for one biological process or multiple biological processes. Coverage can be further defined by the level or direction of the change that a feature undergoes when testing different biological states. Coverage can be relative to the effects of other omics datasets or features of different omics types. The significance of the difference can be tested.
[0278] The number of biological processes covered by the features of the reduced feature classifier can be at least 1, at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 15,000, at least 30,000, at least 45,000, at least 60,000, at least 75,000, or at least 100,000. The number of biological processes covered by the features of the reduced feature classifier can be greater than 1, greater than 100, greater than 500, greater than 1000, greater than 2000, greater than 3000, greater than 4000, greater than 5000, greater than 6000, greater than 7000, greater than 8000, greater than 9000, greater than 10,000, greater than 15,000, greater than 30,000, greater than 45,000, greater than 60,000, greater than 75,000, or greater than 100,000.
[0279] In some aspects, the biological process can be a Gene Ontology biological process. The relationship of the feature to the biological process can be determined by comparing the feature to a database. The database can be the Uniprot database. This relationship can be determined by laboratory tests. This relationship can be determined by theoretical biological interactions. This relationship can be hypothetical. This relationship can result from a statistical analysis of analytical tests.
[0280] In some aspects, the reduced feature classifier can be characterized by coverage of biological processes. The features of the classifier can be selected to maximize this value. In some aspects, this means that individual features can be selected as part of the classifier, while other features that may have higher selection capabilities are not selected. Higher coverage of biological processes can allow the reduced feature classifier to maintain its selectivity when used to classify populations different from the population on which it was trained.
[0281] In some aspects, each different type of omics data can allow coverage of different biological processes. Some omics data sets can provide overlapping coverage of biological processes. The overlapping coverage of omics data sets can interrogate different aspects of the same biological process. Features from different omics data sets but related to the same biological process can provide different coverage. In some aspects, compared to a single-omics classifier, a classifier with multi-omics data features can have higher sensitivity, specificity, or accuracy due to the multi-omics feature coverage of biological processes.
[0282] The relationship between a biological process and a feature can further include calculating the statistical significance of the relationship between the pair. The significance can be a formal test of statistical significance. The p-value of the relationship can be less than 0.15, less than 0.10, less than 0.05, less than 0.005, less than 0.001, or smaller to conclude that the relationship exists. The significance can further include using the log odds ratio (LOR). The LOR can compare the relationship of a biological process with two or more omics groups. The LOR can indicate which omics group the biological process is more related to. It can indicate which feature represents a stronger relationship with the biological process. It can indicate which feature or omics group better detects changes in the biological process or changes associated with the biological process. The LOR can be used to calculate coverage by a top feature classifier.
[0283] Subject Monitoring and Treatment
[0284] In some cases, a subject is monitored. For example, information regarding the likelihood that a subject has a biological state such as cancer can be used to determine to monitor the subject without administering treatment to the subject. In other cases, a subject can be monitored while receiving treatment to observe whether the cancer in the subject improves. In some aspects, the cancer described herein is pancreatic cancer. The methods described herein can include recommending or administering a pancreatic cancer treatment to a subject when proteomic data is classified as indicating pancreatic cancer. In certain aspects, the method recommends administering a pancreatic cancer treatment to a subject when proteomic data is classified as indicating pancreatic cancer. In certain aspects, the method recommends performing a biopsy or pancreatoscopy when proteomic data is classified as indicating pancreatic cancer. In certain aspects, the method recommends observing the subject without administering a pancreatic cancer treatment to the subject. In certain aspects, the method recommends observing the subject without obtaining a biopsy or pancreatoscopy of the subject when proteomic data is not classified as indicating pancreatic cancer. In certain aspects, the method recommends observing the subject without administering a pancreatic cancer treatment to the subject. In certain aspects, the method recommends observing the subject without obtaining a biopsy or pancreatoscopy of the subject when proteomic data is not classified as indicating pancreatic cancer. The decision to treat the subject or obtain a biopsy or not can be based on whether the proteomic data indicates whether a mass (e.g., pancreatic cyst) in the subject's pancreas is cancerous or not. For example, a doctor can detect a pancreatic cyst via a CT scan and then arrange for a blood test involving the methods described herein.
[0285] When a subject is determined to not have cancer, the subject can avoid otherwise adverse cancer treatments (and associated side effects of cancer treatment), or can avoid having to undergo a biopsy or invasive test for the disease state. When a subject is determined to not have cancer, the subject can be monitored without undergoing treatment. When a subject is determined to not have cancer, the subject can be monitored without undergoing a biopsy. In some cases, a subject determined to not have cancer can be treated with palliative care, such as a pharmaceutical composition for pain. In some cases, the subject is determined to have a different disease than the initially suspected cancer and is provided treatment for that different disease.
[0286] When a subject is determined to have cancer, treatment for the cancer can be provided to the subject. For example, if the cancer is pancreatic cancer, treatment for pancreatic cancer can be provided to the subject. Examples of treatment include surgery, organ transplantation, administration of a pharmaceutical composition, radiation therapy, chemotherapy, immunotherapy, hormone therapy, monoclonal antibody treatment, stem cell transplantation, gene therapy, or administration of chimeric antigen receptor (CAR)-T cells or transgenic T cells. In some aspects, the cancer is pancreatic cancer, and the treatment for pancreatic cancer includes chemotherapy, radiation therapy, immunotherapy, targeted therapy, surgery, or surgical resection or a combination thereof. In some aspects, the method recommends treatment for pancreatic cancer that includes administration of a pharmaceutical composition that includes capecitabine, erlotinib, fluorouracil, gemcitabine, irinotecan, leucovorin, nab-paclitaxel, nanoliposomal irinotecan, oxaliplatin, olaparib, or larotrectinib or a combination thereof.
[0287] When a subject is determined to have cancer, the subject's cancer can be further evaluated. For example, a subject suspected of having cancer can undergo a biopsy after the methods disclosed herein indicate that he or she may have cancer.
[0288] Some cases include recommending treatment or monitoring of a subject. For example, a medical practitioner can receive a report generated by the methods described herein. The report can indicate the likelihood that the subject has cancer. The medical practitioner can then provide or recommend treatment or monitoring to the subject or another medical practitioner. Some cases include recommending treatment of a subject. Some cases include recommending monitoring of a subject.
[0289] In some aspects, when a cancer assessment method indicates that a subject has a probability exceeding a predetermined threshold of having pancreatic cancer, the method further comprises performing a subsequent pancreatic cancer treatment or advising the subject to undergo a subsequent pancreatic cancer treatment to determine the presence of pancreatic cancer. In some aspects, the subsequent pancreatic cancer treatment comprises a biopsy. In some aspects, the subsequent pancreatic cancer treatment comprises pancreatic imaging. In some aspects, ultrasound or computed tomography is used to perform the imaging. In some aspects, when a cancer assessment method indicates that a subject has a probability exceeding a predetermined threshold of having pancreatic cancer, the method further comprises treating the subject with a pancreatic cancer treatment for treating pancreatic cancer or advising the subject to undergo such a pancreatic cancer treatment. In some aspects, the therapies are selected from the group consisting of: surgery for pancreatic cancer, radiotherapy for pancreatic cancer, cryotherapy for pancreatic cancer, hormone therapy for pancreatic cancer, chemotherapy for pancreatic cancer, ablation therapy for pancreatic cancer, and immunotherapy for pancreatic cancer. In some aspects, the predetermined threshold is greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. Example
[0290] The following illustrative examples represent embodiments of the stimulations, systems, and methods described herein and are not meant to be limiting in any way.
[0291] Example 1. Identifying the Likelihood of Pancreatic Cancer in a Subject
[0292] A subject presented to a doctor's office with jaundice and abdominal pain. The doctor determined that the subject might be at risk of having cancer and performed non-invasive medical examinations, including a CT scan, but did not detect anything notable. A plasma sample was obtained from the patient for analysis by the method described herein. The laboratory measured the presence and abundance of several proteins. The laboratory then applied a classifier to generate an output report for the doctor to use in determining whether the subject had pancreatic cancer. The report indicated that the patient might have pancreatic cancer. The pancreatic cancer might be small and in an early stage of development, which explained why the scan did not detect the pancreatic cancer. The doctor asked the patient to come in for regular follow-up examinations every six months to continue monitoring for pancreatic cancer. During a subsequent examination, analysis of a biofluid sample obtained from the subject indicated that the pancreatic cancer had progressed. The doctor then prescribed or administered a pancreatic cancer treatment regimen.
[0293] Example 2. Deep, Unbiased Multi-Omics Method for Identifying Pancreatic Cancer Biomarkers from Blood
[0294] Pancreatic cancer is the seventh leading cause of cancer - related death globally and the third leading cause of cancer - related death in the United States. The low survival rate of pancreatic cancer is often due to the challenges associated with early detection of the disease, highlighting the need for the development of early diagnostic tests. While it is less challenging to identify cancer signatures in localized pancreatic tumors via biopsy, cancer signals found in the bloodstream due to cell leakage, metastasis, signaling, or innate immune responses can also be useful due to reduced invasive sampling.
[0295] Challenges encountered in liquid biopsy cancer biomarker discovery studies include the degradation and dilution of analytes in complex biological matrices, which limit high - specificity and sensitivity measurements. To overcome these challenges, comprehensive multi - omics platforms have been developed that facilitate the uncovering of previously unexploited information to gain a more comprehensive biological perspective with unprecedented depth and integrate molecular signatures across complex biological levels. The implementation of this approach has led to the discovery of new pancreatic - cancer - specific biomarkers and a deeper understanding of the integrated pathways in pancreatic cancer.
[0296] In this case - control study, plasma proteomics, metabolomics, and lipidomics data were collected from 196 human plasma samples. The samples included plasma from 92 pancreatic - cancer patients (“cancer samples” or “PC”) and plasma from 104 healthy subjects without cancer (“healthy controls”). Specifically, pancreatic cancer included pancreatic adenocarcinoma. The cancer patients and healthy subjects were matched for age and gender (Table 1, Figure 5A - 5B ). In some of the tables and figures in this article, samples from healthy subjects without cancer are referred to as “healthy” and samples from subjects with pancreatic cancer are referred to as “pancreatic”. The cancer samples were from patients with pancreatic cancer at various stages and included 9 samples from subjects with an unclear cancer stage (“unknown”). No bias was observed based on age or gender comparisons between the groups.
[0297] Table 1. 196 subjects
[0298] Gender Healthy Pancreatic F 53 43 M 51 49
[0299] Data were obtained using liquid chromatography - mass spectrometry (LC - MS). Samples from cancer - bearing subjects were collected after the diagnosis of pancreatic cancer and before the treatment of pancreatic cancer. The data from cancer samples were compared with healthy controls. It was observed that the sample collection and processing were the same for all samples.
[0300] Proteins were measured separately by two methods. One protein measurement method (referred to herein as "Proteograph") involves using particles, where a plasma sample is contacted separately with the particles to adsorb the proteins in the plasma onto the corona around each particle. The proteins adsorbed to the particles are then evaluated by liquid chromatography - mass spectrometry (LC - MS). Proteomics data were obtained from 5 physiochemically different particle types (referred to as "NP1", "NP2", "NP3", "NP4", and "NP5"). The data from the nanoparticles were analyzed separately and as a combined group. These particles were commercially purchased from Seer, Inc., where they were identified as S - 003, S - 006, S - 007, P - 039, and P - 073, respectively. Figure 6A - 6B Shows the total number of proteins in each sample observed by Proteograph. Here, MAXLFQ processing of DIANN reported data was used.
[0301] The second protein measurement method involves using a known amount of isotopically labeled internal reference protein (referred to herein as "PiQuant"). The internal reference protein is incorporated into each plasma sample and then used for mass spectrometry to identify individual endogenous proteins and further used as a standard for determining the amount of individual endogenous proteins.
[0302] In the analysis, 3,381 proteins were detected in all samples (where the proteins were detected in at least 3 samples). Using Bonferroni correction (FDR = 0.05), 124 proteins were measured to be at a statistically significant level in cancer samples compared to healthy controls. The data also included approximately 200 out of 678 total lipids and 49 out of 299 metabolites present in all samples (at least 3 samples per category), and these approximately 200 lipids and 49 metabolites were determined to be at a statistically significant difference level (using Bonferroni correction; FDR = 0.05). The detected analytes (proteins, lipids, and metabolites) included analytes not previously associated with pancreatic cancer. Additional analyses will be performed to further integrate the multi - omics datasets and determine the multivariate statistical performance for detecting pancreatic cancer.
[0303] Proteins were detected by a full - range plasma proteome (including a large number of proteins with high OpenTargets (OT) scores for pancreatic cancer). Table 2 shows some aspects of a total of 2,933 proteins, of which approximately 50% were mapped to HPPP. Table 3 shows the individual aspects of the 10 proteins with the highest OT scores (10 out of 213 proteins with an OT score of 0.15 or higher). Figure 7AShows some data including the mapping to 3,486 proteins in the HPPP database and includes the estimated concentration in ng / mL. Figure 7A The proteins in include MYH9, TUBB1, TUBB, CALR, FLT4, NOTCH2, RHOA, IDH2, CDH1, PRKAR1A, NOTCH1, EXT1, PPP2R1A, SND1, BTK, LPP, MAPK1, FAT1, CDH11, and MAP2K1. Figure 7B Shows the OT score distribution for pancreatic cancer, where any threshold for significance (0.15) is included and checked based on the distribution.
[0304] Table 2
[0305] N High OT HPPP 1436 False False 1337 False True 50 True False 110 True True
[0306] Table 3
[0307] Gene ID OT Score GNAS 0.67 EGFR 0.65 TUBB4B 0.61 RRM1 0.60 TUBB1 0.58 TUBB6 0.58 TUBB8 0.58 TUBB 0.58 SMAD3 0.55 MAPK1 0.52
[0308] Figure 8A Shows the comparison of the median total signal according to sample, analyte type, and category, where large differences can be observed with the targeted method.
[0309] Figure 8B Shows box plots of the most significantly different analytes in each omics workflow ((i): lipids; (ii): metabolites; (iii): proteins). Box plots of the most significantly different analytes in each omics category within the omics categories were investigated. The most significantly different lipid is ceramide. The most significantly different metabolite is 5-aminoimidazole-4-carboxamide-1-β-D-ribofuranosyl 5'-monophosphate (AICAR). The most significantly different protein (i.e., fructose-bisphosphate aldolase) showed significant differences in two of the five nanoparticle (NP) samples. This highlights the efficacy of the Proteograph assay, which utilizes five unique single-NP chemistries that provide complementary protein identification.
[0310] Figure 8C Shows the performance of an exemplary polymer classifier that combines proteomics, lipidomics, and metabolomics measurements. The model was trained using all available samples with known cancer stages. Then, the performance was evaluated on each individual stage or stage group. Five-fold cross-validation was performed and repeated 30 times. The average AUC across 150 runs was calculated. The random forest algorithm was used for proteomics data, and logistic regression was used for metabolomics and lipidomics data.
[0311] Figure 9A and Figure 9BResults of non-parametric (Wilcox) study group univariate comparisons (EDA) from Proteograph data that include any analyte present in >2 samples per class and use Bonferroni multiple test correction. Figure 9C and Figure 9D Results of non-parametric (Wilcox) study group univariate comparisons (EDA) from PiQuant data that include any analyte present in >2 samples per class and use Bonferroni multiple test correction. Figure 10A and Figure 10B Results of non-parametric (Wilcox) study group univariate comparisons (EDA) from lipid data that include any analyte present in >2 samples per class and use Bonferroni multiple test correction. Figure 11A and Figure 11B Results of non-parametric (Wilcox) study group univariate comparisons (EDA) from metabolite data that include any analyte present in >2 samples per class and use Bonferroni multiple test correction.
[0312] Initial multivariate class separation was performed using the complete analyte samples based on parametric (PCA) and non-parametric (UMAP) projections. The separated data are shown in Figure 12A - 12J . Specifically, Figure 12A - 12B is based on combined data (Proteograph, PiQuant, lipid, and metabolite data), Figure 12C - 12D is based on Proteograph data, Figure 12E - 12F is based on PiQuant data, Figure 12G - 12H is based on lipid data, and Figure 12I - 12J is based on metabolite data. In Figure 12C - 12D , missing values were replaced with an arbitrary minimum value.
[0313] The aim of this study was to detect biosignals of pancreatic cancer in non-invasively collected liquid samples. This analysis indicates significant differences between the classes in the collected samples, and these differences may be useful for detecting pancreatic cancer. Further experiments will combine additional features within and across analyte classes to further improve cancer detection. For example, additional proteomic and transcriptomic data (including methylation, mRNA, and miRNA data) will be included in this analysis.
[0314] Example 3. Multivariate Machine Learning Using Gradient Boosting Trees
[0315] The training subset of this study was used in the initial cross-validation analysis using XGBoost. Logarithmic transformation (ln-transformation) and median normalization of all intensity data were performed on 189 feature-complete cases of proteomics, lipidomics, and metabolomics data generated in Example 2. Proteomics data included Proteograph and PiQuant data. Analytes were filtered to those present in at least 25% of the study samples. The 189 complete subjects were split into a training set (n = 141) and a held-out validation set (n = 48). The training set was used to select hyperparameters for XGBoost modeling via five rounds of 5-fold cross-validation, where 112 - 114 training sets were used for training and 29 - 27 training sets were used for testing in each fold. Figure 13 Some of the top features in the training set are shown, where "LPD" refers to lipids, "MTB" indicates metabolites, "PQ" refers to proteins as evaluated by the PiQuant method, and "PG" refers to proteins as evaluated by the Proteograph method. PQ and PG proteins are included as UniProt reference numbers. Receiver operating characteristic (ROC) curves were generated, and the results showed that the combined classifier had an area under the curve (AUC) of 0.924 ± 0.012 (standard error, n = 25) when differentiating pancreatic cancer at any stage from non-cancer, or an AUC of 0.89 for identifying early-stage pancreatic cancer (here stages 1 or 2). Figure 14 ) Another model can be built on the training data using the selected parameters and validated on the n = 48 validation set.
[0316] In this example, the combined classifier was trained on data from mass spectrometry-based assays, including protein, metabolite, and lipid data. The combined classifier can be used to detect pancreatic cancer. Similar classifiers can be trained from samples of subjects with other diseases or cancers and can be used to detect other diseases or cancers.
[0317] Example 4. Analysis of Multiple Blood-Based Genomics Assays in Pancreatic Cancer
[0318] Pancreatic cancer is the third leading cause of cancer-related death in the United States. While the 5-year survival rate across all stages is only 10%, at the early stage when the disease is localized, the survival rate can reach 40%. Therefore, detecting early-stage pancreatic cancer helps reduce mortality; however, most diagnoses are made at stage IV (i.e., after the onset of clinically detectable symptoms). Thus, there is a need to prioritize individuals for further testing using minimally invasive procedures such as liquid biopsies.
[0319] A case - control, proof - of - concept study was conducted using 69 subjects: 36 pathologically confirmed, untreated cases (5 stage I, 5 stage II, 2 stage III, 22 stage IV, and 2 cases of unknown stage pancreatic cancer) and 33 demographically matched controls without any pancreatic disease.
[0320] For each subject, up to 50 mL of blood was collected in a specified tube. Cell - free DNA, as well as mRNA and miRNA from white blood cells, were isolated from these samples and assayed according to standard NGS protocols. Then, measurements of CpG methylation, mRNA, and miRNA transcript abundances were collected. These measurements together can be collectively referred to as genomic assays. Univariate differential analysis was performed on cases versus controls.
[0321] Genomic measurements were collected, including CpG methylation, mRNA, and miRNA transcripts for both cancer and non - cancer subjects. The methylation percentage of CpG sites covered by at least 11 reads was considered. In addition, log - transformed counts of canonical mRNA transcripts and miRNA transcripts were used. Then the data was split into a training set and a held - out set. Next, models were built on each dataset (omics) to distinguish between cancer and non - cancer subjects by training an ensemble classifier on the training data. Each classifier was trained using 30 repeats of 5 - fold nested cross - validation with hyperparameter tuning. The hyperparameter domain of the classifier was partitioned into a discrete grid. Then, every combination of grid values was tried, thereby calculating the performance metrics of the nested cross - validation and reporting the average performance across all runs for each dataset. Finally, the final performance of all three omics was reported by averaging the predictions for each omics. Then the hyperparameters selected during the search were used to configure the final model, and the final model was fit on the entire training dataset for each omics. Then, each model was used to make predictions on the held - out dataset. The final prediction on the held - out dataset was calculated by averaging the predictions on the held - out dataset across all omics.
[0322] Generally, the final classifier consists of a random - forest - based classifier trained on CpG methylation, mRNA, and miRNA data to distinguish pancreatic cancer cases from non - cancer controls. This classifier can be referred to as the genomic classifier.
[0323] Overall, log-transformed counts of 18,045 canonical mRNA transcripts and 1,035 miRNA transcripts, and methylation percentages of 9,290 CpG sites (filtered by sufficient read coverage) were used. Univariate analysis identified 8,769 mRNAs, 204 miRNAs, and 3,128 CpG sites that were significantly differentially expressed (or methylated) at Benjamini-Hochberg FDR < 0.05, including novel and known biomarkers associated with pancreatic cancer. Most of these mRNAs were less abundant in cases compared to controls, while the opposite was true for miRNAs. CpG site methylation was generally more balanced but was more likely to be unmethylated in cases compared to controls. A random forest-based genomics classifier was trained using 30 repeats of 5-fold nested cross-validation with hyperparameter tuning. Across all repeats, the average sensitivity was 46% (95% CI, 20% - 72%) for stages 1, 2, 3 at 92% specificity, 72% (95% CI, 59% - 85%) for stage 4, and 64% (95% CI, 52% - 76%) for all stages. Data for the genomics classifier are shown in Figure 15A below.
[0324] In this initial study of pancreatic cancer using multi-omics readouts from liquid biopsies, a large number of dysregulated mRNA and miRNA transcripts were observed, which may reflect immune system changes associated with cancer. The most discriminatory transcripts include novel biomarkers as well as genes being explored as therapeutic targets in multiple cancers. The machine learning model additionally produced a classifier whose cross-validation performance highlights the potential of multi-omics in disease diagnosis and new target discovery.
[0325] Example 5. Analysis of Multiple Blood-Based Mass Spectrometry and Genomics Assays in Pancreatic Cancer
[0326] Plasma samples from the subjects described in Example 4 were also analyzed using mass spectrometry-based omics assays, including protein (Proteograph and PiQuant), lipid, and metabolite assays. Classifiers were trained using these mass spectrometry-based omics assays, which may be referred to as mass spectrometry classifiers. A combined classifier was trained using both the mass spectrometry-based omics assays of the present example and the genomics assays of Example 4. The mass spectrometry classifier and the combined classifier were trained and tested in a manner similar to the genomics classifier of Example 4 but using different or additional data types, including mass spectrometry assays.
[0327] The performance of the mass spectrometry classifier of the present example, the genomics classifier of Example 4, and the combined classifier of the present example were all compared. Data are shown inFigure 15B Among them, based on the classifier performance, mass spectrometry and genomics assays seem to provide complementary information, making the performance of the combined classifier superior to that of its components.
[0328] Example 6. Unbiased multi-omics method for detecting pancreatic cancer biomarkers using ion mobility mass spectrometry and nanoparticle-based Proteograph technology
[0329] Pancreatic cancer is the seventh leading cause of cancer-related death globally and the third leading cause of cancer-related death in the United States. The challenge of early detection leads to low survival rates, highlighting the need for the development of early diagnostic tests. Biomarkers measured in liquid biopsies provide a less invasive and accessible strategy for early cancer detection. Degradation and dilution of analytes in complex biological matrices limit the measurement of high specificity and sensitivity, making the discovery of biomarkers from blood a daunting challenge.
[0330] An integrated multi-omics platform was developed that integrates multiple analyte measurements, state-of-the-art analytical instruments, and novel data analysis methods. To demonstrate the efficacy of this platform, an unbiased multi-omics study was conducted on a pancreatic cancer cohort of 196 subjects, leading to the detection of new biological signals. This study included the same samples and protein data as used in Example 2. However, a different method was employed to generate lipid data in this study.
[0331] The study cohort included 196 human subjects. Among the 196 subjects, 92 had pancreatic cancer and 104 were healthy. Subject samples were collected after diagnosis, but for cancer subjects, they were collected prior to treatment compared to healthy controls. Plasma samples were processed for proteomics on a nanoparticle-based Proteograph platform (Seer Inc.). The resulting peptides were analyzed by LC-MS / MS on an Evo sep One connected to a Bruker timsTOF Pro2 mass spectrometer (60 samples per day). MS data were acquired in DIA-PASEF mode and analyzed using DIA-NN. Total lipid processing of plasma samples was also performed using an extraction mixture of 1:1 v / v butanol:methanol. Clean extracts from each subject were analyzed by LC-MS / MS in positive ionization mode using DDA-PASEF on a Bruker timsTOF Pro2. Metaboscape was used to analyze the data to detect, deconvolute, and annotate lipids.
[0332] In the initial analysis, 3,381 proteins were detected in all samples (at least 3 samples per class). Among these proteins, over 100 proteins were measured with statistically significant differences in pancreatic cancer subjects following Bonferroni correction (5% false discovery rate). This initial analysis also annotated >260 lipids in positive ion mode from approximately 8,000 features following an annotation method based on conservative rules that incorporated high resolution, high mass accuracy, ion mobility CCS values, and MS2 spectra of DDA PASEF data collection. Example lipid classes detected included phospholipids, triglycerides, sphingolipids, and cholesterol esters. The protein and lipid classes measured in the study have previously been reported to be associated with pancreatic cancer, thus increasing the credibility of the initial proteomic and lipidomic measurements. The data also includes protein and lipid classes whose association with pancreatic cancer is currently unclear. Continued analysis of the detected proteins and lipids can uncover previously unknown biology and expand the field of biomarker analytes for the early detection of pancreatic cancer.
[0333] Preliminary analysis of the cohort study indicated that multi-omics approaches can be used to infer the biological signature of pancreatic cancer, as demonstrated by significant differences across analyte classes between pancreatic cancer subjects and healthy subjects. Further analysis of this cohort study will determine whether the integration of features within and across analyte classes can improve biomarker detection. This is a case-control study and not intended to be a test study. This study indicates pancreatic cancer detection across multiple analyte classes.
[0334] Figure 6A and Figure 6B illustrates the results of proteins detected in samples obtained from 193 subjects. Figure 6A illustrates the median of the groups of 2,736 proteins in five nanoparticles (NP1-NP5) of 193 subjects in this study, where an average of 1,664 proteins were detected. Figure 6B illustrates the groups of 3,822 proteins detected using DIA-NN in five nanoparticles of 193 pancreatic cancer subject samples and healthy subject samples. A group of 2,933 proteins was identified in 25% of the cohort, and a group of 484 proteins was consistently identified in 100% of the cohort. Figure 6CReproducibility of the platform is shown, which indicates the ability to detect biological signals. Analysis groups: C = control; S = sample. Left panel: Proteins with only n>1 detection / analysis groups are retained. For clarity, 2 features with CV>300% out of 2,089 features are removed. Right panel: Proteins with only n>1 detection / analysis groups are retained. For clarity, 48 features with CV>300% out of 7,672 features are removed. Proteins detected in 25% of all samples are used for classifier development. In a case-control study, increased sample variability is expected. Reproducibility across 15 plates and 2 months demonstrates the reproducibility of the control, highlighting the platform performance and the ability to detect biological signals. The increased variability of the samples indicates that biological signals are captured. Figure 6D It shows that more than 5,000 proteins were detected in a feasibility study of 212 subjects. For proteins present in >25% of the samples, the median number of 4 peptides per protein was detected with the following search parameters: 0.1% peptide / protein FDR, default timsTOF parameters, using the complete UniProt human proteome database with contaminants (50% reverse decoy). Figure 6E It shows that a large number of proteins can be reproducibly detected in samples. Individual nanoparticles generate complementary and common protein identifications. Unique proteomes are shown for each sample / particle + group grouped by sample and collection site. Figure 6F Enhanced proteome coverage for detecting known cancer-related proteins is shown. All detected matching proteins from the samples are plotted on the HPPP curve. GeneCards data uses scores reported from matching gene ids and the search term "cancer". The detected HPPP1 proteins cover a difference of 8 orders of magnitude: highest concentration: P00450 - ceruloplasmin; 830,000 ng / mL; and lowest concentration: Q7Z627 - E3 ubiquitin-protein ligase HUWE1; 0.0034 ng / mL. Figure 6G Large-scale depth and efficient plasma proteomics are shown. Figure 6H The quantitative performance of Proteograph applicable to large-scale studies is shown. Figure 6I Reproducibility of large-scale protein enrichment by Proteograph is shown. The reproducibility of Proteograph enrichment is ideally suited for biomarker discovery. Data was collected from 191 enrichments of the same sample. The collection scope includes 3 instruments; 3 cohort studies; 5 operators; 8 months of runtime; 121 plates; and 1500+ subject samples. Figure 6J Reproducibility of the platform over time (months) and instruments is shown. The median MS1 peak area of the iRT peptides is all below 15%, and most are below 10%. Figure 6K The application of the platform in pancreatic cancer biomarker discovery is shown.
[0335] Figure 7A Describes the set of 3,822 detected proteins mapped to the HPPP database. The identified proteins have concentrations ranging over eight orders of magnitude. All identified proteins are shown, with proteins having a significant pancreatic cancer OT score < 0.15 highlighted. Figure 16A Illustrates a volcano plot showing the intensity differences between pancreatic cancer samples and healthy samples. The volcano plot indicates that 124 out of a total of 3,822 proteins in the set between healthy subjects and pancreatic cancer subjects, calculated based on the Wilcox test using Benjamini-Hochberg correction, are statistically significantly different (p = 0.05). A significance test using multiple test correction with a threshold of 0.05 was used in this analysis, using non-imputed data with at least three measurements per category. Figure 16B Shows the study comparison groups (H: healthy; PC: pancreatic cancer). Among 3,381 detected proteins, 124 are statistically significant.
[0336] Figure 17 illustrates a volcano plot showing the differential abundance of lipid species between pancreatic cancer samples and healthy samples. Figure 17A Illustrates a volcano plot showing the differential abundance of lipid species between pancreatic cancer samples and healthy samples calculated based on the t-test and Benjamini-Hochberg correction. 16 out of 259 lipid species are significantly different between healthy subjects and pancreatic cancer subjects (adjusted p-value < 0.05). Representative box plots of two lipid species depict the abundance differences between healthy subjects and cancer subjects. Figure 17B Illustrates an exemplary plot of the top hit lipids based on Figure 17A And Figure 17C Illustrates a volcano plot showing the differential abundance of lipid species between pancreatic cancer samples (stage 1 and stage 2) and healthy samples calculated based on the t-test and Benjamini-Hochberg correction. 5 out of 259 lipid species are significantly different between healthy subjects and pancreatic cancer subjects (stage 1 and stage 2) (adjusted p-value < 0.05). Representative box plots of two lipid species depict the abundance differences between healthy subjects and early cancer subjects. Figure 17D Illustrates an exemplary plot of the top hit metabolites based on Figure 17C
[0337] The multi-omics platform described in Example 6 has shown to facilitate the evaluation of the global proteome and lipidome of pancreatic cancer cohorts and identify multiple putative biomarker candidates across analyte classes for early disease (e.g., early stage of pancreatic cancer). Untargeted DIA proteomics data yielded 124 statistically significant proteins out of a total of 3,822 proteins. Untargeted lipidomics data showed that 16 out of 259 lipids were significantly different between healthy subjects and pancreatic cancer subjects. Five out of 259 lipids were statistically significantly different between healthy subjects and stage 1 and 2 cancer subjects, highlighting the detection of biological signals associated with the early stage of pancreatic cancer.
[0338] This unbiased multi-omics platform using 4D mass spectrometry can integrate cancer molecular signatures of multiple analytes to facilitate early biomarker discovery. Additionally, the platform can be integrated with other analytes from genomics, transcriptomics, metabolomics, DNA methylomics, and glycomics.
[0339] Example 7. The combination of Proteograph technology with Zeno SWATH acquisition further improves deep unbiased discovery of biomarkers in blood
[0340] Recent proteomics advancements have enabled large-scale studies for exploring biomarkers related to disease diagnosis and prognosis, while providing insights into the pathogenesis of complex diseases such as cancer. Compared to invasive techniques such as tissue biopsies, liquid biopsies are increasingly used for large-scale biomarker exploration due to the non-invasive nature of sample collection, offering the potential for improved prognosis and survival. Despite challenges in achieving deep proteome coverage in complex biological matrices, innovative sample preparation and liquid chromatography-mass spectrometry (LC-MS) techniques have facilitated the identification and quantification of cancer-specific biomarkers across a wide concentration range in liquid biopsies. This study addressed the unmet need for in-depth, reproducible identification from the human plasma proteome using advanced sample preparation and LC-MS techniques.
[0341] From large multi-omics oncology discovery studies (including >1,750 subjects across 3 different cancers), a retrospective case-control sub-study was conducted to investigate the plasma proteome profiles of 104 normal subjects and 92 pancreatic cancer subjects (same plasma samples as in Example 2). Samples were processed using nanoparticle-based Proteograph technology from Seer. The samples were then subjected to data acquisition using a Waters ACQUITY M-Class system (LC) with a capillary flow rate (5 μL / min) synchronized with a SCIEX ZenoToF 7600 system (MS). Replicate injections were made into the mass spectrometer in data-independent acquisition (DIA) mode with and without the prototype Zeno SWATH acquisition enabled. Data processing and downstream analysis were performed using DIANN.
[0342] In this study, the nanoparticle-based Proteograph technology was implemented together with the prototype Zeno SWATH acquisition method to generate highly reproducible proteomics data while increasing the depth of coverage of low-abundance proteins.
[0343] Due to the combination of the increased sensitivity of the Zeno SWATH acquisition method with the additional proteomics depth provided by the Proteograph technology, an average of >1,500 protein groups and >13,000 peptides were annotated per plasma sample. A sub-study of approximately 200 biological samples and process controls generated robust plasma protein measurements across approximately 1,000 injections, demonstrating the robustness and reproducibility advantages of the combination of capillary LC with Zeno SWATH acquisition. Additionally, a large difference in reproducible protein identification was observed using ZenoSWATH acquisition compared to SWATH acquisition using the same experimental and analytical parameters. These results further demonstrate the feasibility of running larger cohort studies with thousands of clinical samples that address the historical technical challenges associated with translating proteomics into the clinic.
[0344] Furthermore, this study shows that the Proteograph or Zeno SWATH acquisition workflows can be used to facilitate the identification and quantification of thousands of proteins from human plasma without compromising throughput or reproducibility, creating a unique opportunity to detect robust protein biomarkers for feasible clinical tests translatable to complex diseases. Quantification of thousands of plasma proteins was achieved at least in part by combining nanoparticle-assisted sample preparation with reproducible and sensitive MS measurements.
[0345] It was found that compared to traditional SWATCH acquisition, Zeno SWATCH DIA acquisition on K562 standard cell lysates resulted in at least a 26% increase in the total number of precursors and a 13% to 83% increase in the set of proteins identified ( Figure 19A and Figure 19B ). Using Zeno SWATCH DIA technology demonstrated a slight increase in the overall MS peak area and a significant increase in the MS / MS peak area of low-abundance species, resulting in improved identification at both the peptide and protein levels ( Figure 20 - 22 ). When compared to SWTCH, Zeno SWATCH DIA had improved reproducibility to a greater extent even when the instrument introduced the minimum sample loading mass. When compared to SWATH acquisition, a 4% to 13% decrease in the CV(%) of the precursor level intensity of Zeno SWTCH DIA was observed (Figure 23). The nanoparticle-based Proteograph processing of both subject and pooled control plasma samples was combined with Zeno SWATH DIA acquisition. An increase in the depth of protein identification was observed. In nanoparticle-derived samples from pooled control samples, a 53% to 85% increase in peptide identification by Zeno SWATH DIA acquisition was observed when compared to SWATCH ( Figure 24 ). Analysis of a subset (55) of 196 control and pancreatic cohorts showed that an average set of 2,357 proteins was found in at least one sample and an average set of 1,077 proteins was found in at least 25% of all subject samples ( Figure 25 ).
[0346] Figure 18A Shows the quantitative performance of Proteograph applicable to large-scale studies (e.g., the study in Example 7). Figure 18B Shows the reproducibility of large-scale protein enrichment by Proteograph. The reproducibility of Proteograph enrichment is ideally suited for biomarker discovery. The system provides high-throughput, reproducible, and in-depth proteome coverage for new discoveries. The reproducibility by Proteograph enables quantitative, in-depth, non-targeted proteomic biomarker studies. The large-scale protein enrichment by Proteograph is highly reproducible ((NP1 = 0; NP2 = 0; NP3 = 2; NP4 = 0; and NP5 = 2). Figure 19AShows the evaluation of K562 precursor detection using SWATH and Zeno SWATH DIA. A minimum increase of 26% in precursor identification was detected using Zeno SWATH DIA. All data were generated from the pr and pg matrices output from DIA-NN (all quantified precursors and called proteins were identified). All data were searched in DIA-NN using "Robust LC" and the SCIEX K562 spectral library. Figure Shows the evaluation of K562 precursor detection using SWATH and Zeno SWATH DIA. A minimum increase of 13% in protein group identification was detected using Zeno SWATH DIA. All data were generated from the pr and pg matrices output from DIA-NN (all quantified precursors and called proteins were identified). All data were searched in DIA-NN using "Robust LC" and the SCIEX K562 spectral library.
[0347] Figure 20 Shows enhanced sensitivity that increases the number of detected low-abundance peptide species. Detection of low-abundance peptides was improved with Zenon SWATH DI compared to SWATH. Figure 21 Shows the graph generated from all qualified precursors. Data were searched in DIA-NN using "Robust LV" and the SCIEX K562 spectral library. Figure 22 Shows that the quantitative sensitivity increases with mass on SWATH and ZenoSWATH DIA. The Zeno SWATH DIA MS1 peak area (K562) is distributed for peptides of lower abundance. Figure 23A Shows that Zeno SWATCHDIA acquisition results in a higher amount of K562 MS2-based precursors compared to individual SWATH acquisitions between different peptide injection masses based on all qualified precursors. Data were searched in DIA-NN using "Robust LC" and the SCIEX K562 spectral library. Figure 23B Shows that Zeno SWATH DIA acquisition results in a lower CV (5) of the K562 precursor level amount compared to individual SWATCH acquisitions between different peptide injections aggregated based on all quantified precursors. Data were searched in DIA-NN using "Robust LC" and the SCIEX K562 spectral library. Figure 24 Shows that Zeno SwatchDIAMS / MS acquisition results in 53% - 85% more peptide identifications in the Proteograph generated from the pooled control samples when compared to SWATH MS / MSDIA acquisition. Figure 25 Shows the protein groups of 2,357 proteins for all five nanoparticles in a representative subject cohort. Protein groups of 1,077 proteins were identified in at least 25% of the patient samples. Figure 26AShows a large number of proteins that can be reproducibly detected in a sample. Individual nanoparticles yield complementary and co - occurring protein identifications. Figure 26B Shows improved sensitivity equivalent to detecting more low - abundance peptides in Proteograph peptide detection.
[0348] Example 8. A validated pancreatic ductal adenocarcinoma (PDAC) classifier based on a panel of proteins measured in plasma by targeted mass spectrometry in a case - control study of 182 subjects
[0349] Early detection of pancreatic cancers such as pancreatic ductal adenocarcinoma (PDAC) can be beneficial to avoid the negative outcomes of late detection, where the five - year survival rate for distant metastatic cancer is only 3%. Providing a simple, blood - based PDAC test with sufficient sensitivity and specificity could be useful for efficient and effective deployment in an initial general screening population or at - risk populations, such as patients with chronic or acute pancreatitis, or patients with newly diagnosed adult - onset diabetes. A cancer biomarker, CA19 - 9, is sometimes used for PDAC and other cancers, particularly for recurrence testing, but its lack of performance, especially specificity, may make it clinically unacceptable in the intended test populations described above. Additionally, 5% - 10% of the general population is Lewis - negative and cannot produce the CA19 - 9 cancer antigen at all.
[0350] Advances in methods for detecting large numbers of proteins and other analytes from subject samples, and in machine - learning - based methods for classification using data collected by those methods, can be used to improve tests such as CA19 - 9. Multiple signal inputs can be used to detect and distinguish complex pathologies, such as cancer. For this purpose, a large, unbiased, targeted protein mass spectrometry (MS) panel was used in a case - control study of PDAC subjects and age - and sex - matched non - cancer controls to construct and validate a multi - protein panel with performance characteristics superior to CA19 - 9.
[0351] Cases and controls for this IRB - approved, observational, sample - collection study were collected over a period of more than two years from 17 different sites, where PDAC subjects were recruited from 15 of those sites and non - cancerous controls were recruited from 5 of those sites. The primary inclusion / exclusion criteria were based on newly diagnosed, biopsy - confirmed PDAC with no other cancer or cancer history for at least the previous five years. PDAC subjects were informed of their diagnosis but not treated and samples were typically collected within a few weeks of pathologic confirmation.
[0352] In initial data analysis (EDA), a cohort of 184 age- and sex-matched subjects collected was evaluated by univariate, non-parametric Wilcoxon tests and by multivariate, principal component analysis (PCA) and hierarchical clustering. Of the 447 proteins out of the original 554 proteins present in at least 50% of at least one category in the targeted mass spectrometry (MS) panel, 113 showed significant differences in the Wilcoxon test, with Bonferroni-adjusted p-values less than or equal to 0.05. Most (e.g., 94 out of 113) of the significant differences were elevated in PDAC. The performance of the multivariate analysis by PCA and hierarchical clustering further demonstrated the usefulness of effective combinatorial biomarker category separation by regression- and decision tree-based classification methods.
[0353] For machine learning-based classification model analysis (results summarized in Figure 27In the present study, a two-stage approach was selected. First, the cohort was randomly divided into a training group (n = 127) and a validation group (n = 55), in which cancer status-stage stratification was performed to maintain proportionality. Then, two-stage repeated cross-validation (10 repetitions of 10-fold RCV) was used in the training group. First, XGBoost (gradient boosting ensemble decision tree method) was used to select the top 20 most important features. Then, after randomly assigning the same subjects to repetitions and folds, GLMnet (regularized logistic regression method) was used for the second round of RCV, using only the selected features. In addition to performing this analysis with the targeted MS protein panel, CA19-9 was directly measured from these subjects using a specific clinical assay, and this biomarker was evaluated individually and in combination with the top 20 protein features in the combined model. Using the final GLMnet model created using all training subjects and the top 20 proteins, a performance of 0.926 AUC was demonstrated in the validation group. In contrast, using an independent CA19-9 model, a validation performance of 0.838 AUC was demonstrated. In the combined approach using the top 20 proteins as well as CA19-9 in the final GLMnet model, a validation performance of 0.963 AUC was achieved. The performance of this classifier was statistically superior to that of CA19-9 alone (p = 0.045). The performance of the combined model at 99%, 98%, and 95% specificity was 77%, 77%, and 82% sensitivity, respectively. The high performance of the combined model extended to all PDAC stages, where 7 out of 9 (78%) stage 1 / 2 cancers in the validation group were correctly assigned. Although 20 was chosen as the number of features to progress from the initial XGBoost RCV feature selection to the final GLMnet RCV model construction, a smaller number of features from the top 20 could give similar performance. Thus, a pancreatic cancer classifier including any of the top 20 features or any combination thereof in this model can be used for the early detection of pancreatic cancer.
[0354] Using a PDAC association annotation list of 4,886 genes downloaded from OpenTargets, the prior associations of the top 20 proteins were examined. The feature UniProt IDs were mapped to gene names and IDs and then to OpenTargets annotations. Four of the top 20 proteins had no PDAC OT scores, indicating little evidence of prior association with PDAC in this database. The remaining 16 proteins had non-zero scores, but the highest score was still lower than approximately 80% of all 4,886 proteins in the database. This suggests that the panel of 20 top protein features selected may be a new combination of proteins for identifying pancreatic cancer and also indicates that the individual top protein features used in this article can be used alone or in combination with other features for a pancreatic cancer classifier.
[0355] Study Subject Analysis
[0356] As described above, the study subjects were derived from an IRB-approved observational study that collected samples from patients with biopsy-confirmed PDAC immediately after diagnosis but before any treatment. Thus, these individuals were diagnosis-informed but untreated. Multiple sample types were collected in this study to enable multiple omics studies, but this example focuses on measuring multiple proteins via targeted MS using plasma samples. The proteins selected had no known bias towards PDAC-related proteins. To avoid introducing any confounding bias consistent with groups in the study, samples were collected from many sites and over a large time span, distributed as Figure 28A shown. Although the collection sites and enrollment dates were not completely random with respect to category, PDAC or control, the large number of sites and the large time window mitigated any meaningful bias between the study groups. Controls were age- and sex-matched individuals who met the study-specified inclusion / exclusion criteria, excluding any cancer history within the previous five years. Figure 28B It was demonstrated that there was no age or sex bias between the groups split by cancer stage. Ages were compared by the Wilcoxon test, and sex ratios were confirmed by the Fisher test. These 184 subjects, with data from both unbiased targeted protein panel MS assays and analyte-specific CA19-9 ELISA assays, were used as input for subsequent EDA and machine learning-based classifier analysis.
[0357] PiQuant Protein Assay Data Preparation
[0358] Five hundred and fifty-four proteins were measured by targeted MS on a ThermoFisher OrbiTrap MS using a panel of stable isotope standard (SIS) peptides. In this example, the method is referred to as "PiQuant". Using SIS-peptides as internal calibrators enabled highly accurate and precise determination of peptide and protein levels in each sample. For 401 of the 554 proteins, a single peptide was used for each protein. For the remaining proteins, 2, 3, or 4 peptides were used. MS data including MS2 fragments or transitions of each peptide were initially processed by Biognosys' Spectrodive analysis software.
[0359] To calculate the protein level values for each member of the panel in each sample, the median of each SIS peptide transition (as measured across all subject samples) was calculated, and then the sample-specific correction factor derived from these data was applied to the endogenous peptide transitions for each sample. Each transition used to calculate the peptide values must have a signal-to-noise ratio greater than 3, and for the determination of the peptide levels for each sample, the three transitions with the highest median in the sample were summed. The numerical values were log-transformed to improve normality. As an additional filter, the protein needed to be detected in 50% of at least one study category (PDAC and / or control) prior to the subject-level analysis. 447 out of 554 proteins met this criterion. Figure 29A The distribution of the protein level values for each sample is shown, highlighting how SIS-based quantification effectively normalizes the data between subject samples without bias between groups. Although there may appear to be a bias in the higher median for PDAC subjects by inspection, the p-value for the Wilcoxon test comparing the median protein levels by group was 0.083( Figure 29B ). As part of the PiQuant assay, subjects were randomly assigned to groups.
[0360] Given the sensitivity of modern classification methods via machine learning, it is crucial to avoid or remove as much accidental bias as possible between the comparison groups. Additionally, it is also crucial to remove noise samples that appear to be statistical outliers from the analysis. If such samples are outliers with respect to the rest of the dataset for reasons unrelated to the class discrimination problem being addressed, the variance they may add to the analysis may exceed any signal detection capabilities that might be gained from sample retention. A set of methods for judging outliers that are particularly applicable to high-dimensional, multi-measurement-per-sample assays is based on evaluating the variation of the median absolute deviation (MAD). In microarray genetic analysis, the sum or other aggregation of the absolute values of the residuals of all features of the array (as measured against the central tendency of those features across all subjects in the analysis) is a common measure of relative performance within a group. In the analysis, the mean absolute relative log expression or MARLE for each sample was calculated, and then after confirming the absence of class bias in the samples to be excluded, outlier samples were identified and removed from further analysis. Figure 30 The distribution of MARLE values for 184 subjects in the PDAC study is shown, highlighting two samples with MARLE values greater than 3 standard deviations from the mean of all subject values. Since each group is equally represented in the potential outlier exclusion, excluding the data for these subject samples does not remove class bias (e.g., there is no potential class discrimination signal). After removing these samples, 182 subjects (including 80 PDAC and 102 controls) remained for EDA and machine learning-based classification.
[0361] Single Protein Comparison by Wilcoxon Test
[0362] After single subject - protein normalization based on SIS, 50% class presence study - protein filtering, and MARLE - based study - subject outlier removal, the differential expression levels of 447 remaining proteins in 182 remaining subjects between the study groups were evaluated by non - parametric Wilcoxon test. Considering that the number of tests per sample was moderately large, Bonferroni correction was used to adjust the Wilcoxon p - values, where p = 0.05 (adjusted) was set as the new significance threshold. The differential score was median control - median PDAC, meaning that a negative difference represents a protein that is present at a higher level in PDAC subjects. Figure 31 A volcano plot of - log10(Wilcoxon test p - values) versus the differential is shown. As shown in the figure, most of the proteins were significantly different (e.g., 113 out of 447, of which 94 were detected at higher levels in PDAC). The figure highlights ten significantly different proteins selected for individual subject value inspection. The individual data points for these ten proteins from the subjects are as Figure 32 shown. As shown in the figure, although all ten proteins were indeed significantly different, none of them completely distinguished the groups. This indicates that in both EDA as a feasibility study and in machine - learning - based classifier construction, multivariate analysis helps to achieve clinically useful performance.
[0363] Multivariate Group Comparison by PCA and Hierarchical Clustering
[0364] Given that no single analyte significantly distinguished the study groups, two multivariate EDA methods were employed to understand the feasibility of multi - component classifiers for machine learning. In the first method, PCA, all 447 PiQuant - measured proteins were used in the analysis, where any missing values from the subjects were imputed with the minimum value of that analyte across all 182 subjects. This implicitly assumes detection of missing levels rather than random missingness. In Figure 33 it, a modest multivariate - based separation of the groups was evident, where more separation was likely to be observed for stage 4 PDAC subjects. Only the first two principal components were plotted, which accounted for only 23.6% of the total variance.
[0365] To further explore the potential of multivariate classification, an unsupervised hierarchical clustering using all analytes with the complete method using Euclidean distance measurement was deployed. After clustering, the subject dendrogram was cut to produce two groups to visualize the potential of the data to separate PDAC and controls. As Figure 34As shown, there is a significant separation between the groups with forced splitting. The PDAC subjects in each branch seem to include all stages of PDAC. There seems to be a significant amount of correlation within the protein analytes, which is an important factor to consider in machine learning-based classification. Summarizing the EDA, it is clear that many protein analytes are significantly differentially expressed between these groups. Multivariate EDA, especially unsupervised hierarchical clustering, indicates that combining proteins into a small set of analytes can be an improved method for developing a group classifier.
[0366] Machine Learning-Based Classification
[0367] Given the large number of protein analyte features (447) and the number of study subjects (182), the study was randomly divided into a training group and a validation group at a 70 / 30 ratio, and then a two-stage method was used within the training group to construct a final model for validation in the held-out group (see Figure 27 ). In the training group, a first round of 10-fold repeated cross-validation (RCV) with 10 repeats for the most important feature selection was performed, and then, after randomly reshuffling the samples into new repeat-fold groupings to minimize overfitting, a second round of 10x10 RCV with the important features for the final model parameter selection was performed, and then the final training set model was constructed. Given the nature of the data observed during the EDA, the gradient boosting ensemble tree method XGBoost was selected for the first round of RCV, and the regularized logistic regression of GLMnet was selected for the second round of RCV.
[0368] Training Validation Subject Segmentation
[0369] The first step in the classification process is to split the evaluable study subjects into a training group and a validation group, with the validation group held out for the final test. The subjects were randomly split 70 / 30, maintaining the ratio of cancer status-stage-status between the groups. Figure 35 The split is shown and indicates no significant differences between the groups in terms of age or gender.
[0370] Feature Selection by RCV and XGBoost
[0371] The training split (n = 127) was randomly re - split into 10 repeats of 10 - folds, maintaining the proportion of cancer status and cancer stage within the folds. Although the numbers were slightly different, a typical fold within a repeat had 113 to 116 subjects for model creation and 14 to 11 subjects for model testing. Using these RCV splits, a large grid (n = 200) of potential combinations of seven hyperparameters for XGBoost modeling was constructed using optimized Latin hypercube sampling. Race - adjusted ANOVA was used during RCV to optimize the computational time required to complete model hyperparameter evaluation. In this method, the initial burn - in of randomly sampled folds across all parameter combinations was evaluated and compared via ANOVA. Those models that were statistically not as good as the current best model were discarded, and then the process was repeated until the final repeat - fold number was evaluated. Figure 36 The race - adjusted of XGBoost RCV deployed here is shown. The XGBoost model can be very sensitive to model parameters (and thus sensitive to overfitting), and this is evident in Figure 36 the pruning rate of parameter combinations in
[0372] The first evaluation after the initial 20 repeat - fold burn - ins removed a large number of poorly performing models. At the end of the evaluation, the best model combination of parameters (shown in Table 4) achieved an average AUC of 0.959 across all models (10 repeats of 10 - fold evaluations).
[0373] Table 4. Optimized XGBoost model hyperparameters selected in 10x10 RCV
[0374]
[0375] The combined ROC plot of the 10x10 RCV with the best parameter combination is as Figure 37 shown. The prominent curve represents the interpolated mean summary values of sensitivity and specificity for each of the included repeat - folds using 11 - 14 subjects. The light - gray plots are the individual repeat - fold plots themselves.
[0376] Although the prediction performance of the XGBoost-based classifier using all 447 protein features as input predictors is excellent in itself (e.g., average AUC 0.96), this first stage for feature selection is used to demonstrate the potential of a commercially viable and clinically useful classifier. The top 20 features from this stage are selected to advance to the second stage RCV, although fewer features may be sufficient to achieve adequate performance. To identify the top 20 features, feature importance is selected from each 10x10 repeated-fold model (summarized primarily by median rank, where the worst and best ranks are used to break ties). The protein and gene names of the top 20 protein biomarkers in this example are also listed in Table 5. The selected important features are enumerated in Table 6.
[0377] Table 5. Examples of Biomarkers for Evaluating Pancreatic Cancer
[0378]
[0379]
[0380]
[0381] Table 6. Ranked Top 20 Protein Features Selected from XGBoost 10x10 RCV
[0382] Variable Number in Top Features Median Rank Worst Rank Best Rank P01011 100 1 2 1 P02750 100 2 5 1 P01009 100 3 6 1 P15144 100 4 9 2 P18428 100 5 11 2 P05362 100 7 13 4 P01833 100 7.5 12 3 P05109 100 9 18 4 P06681 100 9 13 5 P01031 100 10 16 5 P02748 100 10 16 5 Q06033 100 11 16 4 P02753 100 14 28 8 P08637 100 16.5 37 8 P02741 100 17.5 35 9 P05452 100 18 40 11 Q99784 100 18.5 41 11 P05160 100 19 37 14 P02647 100 20 40 13 P02652 100 21 37 11
[0383] Optimal Final Model Parameter Selection by GLMnet RCV
[0384] Using the top 20 features selected from XGBoost RCV, with the same subjects (n = 127), but in a new random collection of repeated-fold splits, a 10x10 RCV for the second stage using logistic regression based on GLMnet is performed. Although several modeling engines could be used here, GLMnet is selected as an example to obtain individual subject class probabilities for additional comparison with other models (e.g., with CA19-9 model performance) and to create model terms (e.g., feature coefficients). This modeling engine can also be used for feature selection / reduction.
[0385] In the same manner as the XGBoost RCV above, a hyperparameter training network of 200 possible combinations is created. Given the smaller number of input predictor features (20 vs. 447), a full RCV is selected instead of performing a null analysis using race-adjusted ANOVA because most models are expected to perform very well and the race selection process may introduce parameter selection variability due to very small differences in model performance.
[0386] Figure 38 Shows the results of 10x10 GLMnet RCV in terms of hyperparameter evaluation. Most hyperparameter combinations work well when the average AUC is higher than 0.9. The best parameters are shown in Table 7, where the average RCV AUC is 0.989.
[0387] Table 7. Optimized GLMnet top feature model hyperparameters selected in 10x10 RCV
[0388]
[0389] Using the selected hyperparameters, the combined ROC plot of these 100 models using 10x10 RCV is as Figure 39A shown. As previously mentioned (see Figure 37 ), the interpolated, combined results of 10x10 RCV are shown as a prominent line, and the individual plots of 11 - 14 subjects in each repeated - fold test split are shown in light gray. The 10x10 RCV results of the GLMnet model on top of XGBoost indicate that the final classifier built with the selected GLMnet parameters using all training data can have a useful level of performance (e.g., having clinical - useful sensitivity and specificity including a feasible number of features). The final GLMnet model is built using all training data (n = 127) and optimized penalty and mixing parameters. As expected, when evaluated on the training data used to build the model, this final model has excellent performance (AUC 0.995). The coefficients of the logistic regression model are as Figure 39B shown. As shown in the figure, the coefficients are a mixture of positive and negative values, and many coefficients have similar magnitudes, which confirms that the multivariate classifier may be more useful than any single feature for achieving optimal performance. The plot of the coefficients also indicates that a subset of the top 20 features can also achieve significant performance considering the shrinkage towards zero for the entire group.
[0390] Final Validation of GLMnet Model Based on Top Features in Validation
[0391] Using the final GLMnet model, the predicted classes and probabilities for the held - out validation group of subjects (n = 55) are obtained. Figure 40 The validation ROC plot is shown, and the calculated AUC is annotated as 0.926 (95% CI 0.8479 - 0.9521 by DeLong's). Using 2,000 stratified bootstrap resamplings, the sensitivity of the model for the validation data at specified specificities is calculated and shown in Table 8.
[0392] Table 8. Sensitivity and specificity values of the final top - feature GLMnet model in validation
[0393]
[0394] The validated model performance achieved 77% sensitivity at 99% specificity, demonstrating the feasibility of the model's potential clinical useful performance for PDAC detection.
[0395] Comparison with CA19-9 PDAC Detection
[0396] The cancer antigen CA19-9 is often used as a marker for pancreatic cancer and is typically used as a recurrence test given its lack of specificity in the general screening population. To compare the performance of the classifier with this marker, CA19-9 levels were measured in subjects using an analyte-specific clinical-grade assay. The following Figure 41A and Figure 41B show the CA19-9 levels for the cancer group and cancer stage compared to controls, respectively. The data show that CA19-9 is significantly elevated in PDAC subjects compared to non-cancerous controls. There also appears to be a significant increase in stage 4 levels compared to earlier stages, although the comparison to stage 3 could be further validated with additional subjects.
[0397] Using these measured CA19-9 levels, the clinical assay values were converted to model probabilities using a simple logistic regression engine, GLM. The model was first built in the complete n = 127 training data and then evaluated in the n = 55 validation set. Converting the assay values to model probabilities enabled subsequent direct comparison of the model's performance. Figure 42 Shows the performance of CA19-9 as a classifier in the validation set. As shown, the AUC was 0.8375 (0.7021 - 0.9729 95% CI). Table 9 shows the performance of the model at the same specificity points highlighted above.
[0398] Table 9. Sensitivity and specificity values of the final CA19-9 GLM model in validation
[0399]
[0400] Comparison of 64% sensitivity at 99% specificity with the above top-performing model (e.g., 77% sensitivity at the same specificity) indicates that this panel represents a significant improvement over existing tests such as CA19-9 alone. Comparison of the ROC curves via paired bootstrap resampling (n = 50,000) gave a p-value of 0.147, where the difference between the two AUCs and the standard deviation of the bootstrap differences (e.g., D = (AUC1 - AUC2) / s) were compared to the normal distribution.
[0401] Combined Performance of Top Features and CA19-9
[0402] Although CA19-9 has not been clinically accepted for widespread testing of cancer in the general population, it may significantly increase overall performance when combined with other potential biomarker components. To evaluate this possibility, the same GLMnet-based approach was used for logistic regression as described above, and a multivariate classifier was developed using the top 20 XGBoost RCV features selected as above in combination with CA19-9. Using the same training data (n = 127) and the same repeated-fold as the top feature GLMnet classifier, a new round of 10x10 RCV was performed. The optimal hyperparameters are shown in Table 10, and the final model based on all training data was constructed.
[0403] Table 10. Optimized GLMnet combined model hyperparameters selected in 10x10 RCV
[0404]
[0405] The coefficients of the final model based on all training data are as Figure 43A shown, and the performance of these best parameters on 100 models of 10x10 RCV is as Figure 43B shown, where the average AUC is 0.98. Although CA19-9 ranks highly in this combined classifier, having the second highest absolute regression coefficient of its regression term, there is one other feature (e.g., P15144) with a higher value and several other features with similar values. Therefore, the features selected from XGBoost RCV are useful factors in this combined model.
[0406] Final Validation of Combined (Top Features Plus CA19-9) Model
[0407] Using the final model with combined features on a held-out validation set of subjects (n = 55), class predictions and probabilities were obtained and the performance was visualized in the ROC plot as Figure 44 shown. The AUC was 0.9628 (0.9193 - 1 95% CI)
[0408] The calculated sensitivities at various specificities calculated as above are shown in Table 11 (and using 0.5 as the class probability threshold). The estimated performance of 77% sensitivity at 99% specificity is a significant improvement over the 64% sensitivity at 99% specificity described for CA19-9 alone. Comparison of the ROC curves via bootstrap resampling confirmed the statistical significance of the combined curve compared to the CA19-9 curve, with a p-value = 0.045. For the above table, using a class probability of 0.5, the predicted confusion matrix is shown in Table 12. In fact, there is a good balance between classes for model accuracy.
[0409] Table 11. Sensitivity and Specificity Values of the Combined GLMnet Model
[0410]
[0411] Table 12. Confusion Matrix of the Classifier Based on Combined Feature GLMnet
[0412]
[0413] Combined Model Performance across PDAC Categories
[0414] Since early detection of PDAC is an important goal of this study and the ultimate clinical application, the classification performance of the final combined model classifier spanning the cancer stages represented in the validation set was evaluated. The scores in Table 13 show that 8 out of 10 subjects (80%) in stages 1 - 3 were correctly classified, and 78% of subjects in stages 1 - 2 were correctly classified. The data indicate that the performance of the classifier was significantly extended across all PDAC stages.
[0415] Table 13. Accuracy of the Final Validated Combined Model Classification across PDAC Stages
[0416]
[0417] Novelty of Top Features
[0418] Although the MS - based assays of 447 proteins evaluated in this PDAC versus non - cancer control study were "targeted" from a technical perspective of MS data acquisition, the proteins in this panel were not particularly biased towards PDAC or cancer per se. These proteins are not a truly random sample of all possible plasma - detectable proteins, but given their relatively unbiased nature, they can be used to discover new combinations of known and unknown players in PDAC detection.
[0419] One way to assess the novelty of classifier components is to observe their importance or benefit (interest), as defined by disease - related association scores in an aggregated database such as OpenTargets. The overall association scores for PDAC of 4,886 genes and associated proteins are listed in a table annotated as EFO0002517 from the OpenTargets database. By mapping Uniprot identifiers to PiQuant - based protein features, the relative rank of the selected proteins to those in the database can be visualized. The OT ranks include many components of interest (e.g., drugs in development, publications, genetic associations, etc.) and are not just about plasma detectability, so simply selecting the genes or proteins with the highest OpenTargets scores may not be sufficient to develop a blood - based test for detecting any given disease.
[0420] In Figure 45 , the overall distribution of the OpenTargets PDAC-related "overall association score" (n = 4,886) was plotted, and the distribution of 16 out of 20 combined histone proteins with non-zero scores was annotated. As can be seen from the figure, although these 16 proteins do have scores greater than zero, they may not necessarily be prioritized for targeted plasma-based detection efforts strictly based on score rank. 81% of the features in the database have scores higher than the maximum value of this group of proteins (0.0820 in the range of 0 to 1). Using these annotations as criteria, the four proteins without PDAC scores may not be selected at all.
[0421] Example 9. Multi-omics data shows further potential for improving early pancreatic cancer detection
[0422] A classifier was trained on 112 plasma samples with metabolite, lipid, protein, methylation, and mRNA data. The samples included samples from 9 subjects with stage I pancreatic cancer, 9 subjects with stage II pancreatic cancer, 2 subjects with stage III pancreatic cancer, 27 subjects with stage IV pancreatic cancer, 4 subjects with pancreatic cancer of unknown stage, and 61 cancer-free subjects. At least some of these samples overlap with the samples of other embodiments described herein.
[0423] The staging analysis shows the performance for stage I and II (ROC AUC of 0.935), which is almost as good as the performance across all stages (ROC AUC of 0.944) ( Figure 46 ). For the two sample sets analyzed so far, the sensitivity ranges from 64% to 73% at 98% specificity.
[0424] Different biological processes were observed in RNA-seq and untargeted proteomics data. For example, Figure 69A illustrates the biological processes captured by RNA-seq. The significance level of each observed biological process is shown. In the RNA-seq data shown in the figure, Toll-like receptor 4 binding is the most statistically significant. Figure 69B illustrates the biological processes captured by untargeted proteomics. Similarly, the significance level observed in the biological process is shown. It was found that the structural constitution of chromatin has the greatest statistical significance in untargeted proteomics. Figure 69A - 69B shows that different molecular assays capture analytes from different biological processes.
[0425] Example 10. Multi-omics platform applied to pancreatic cancer
[0426] Figure 47Shows sample and analysis details in multi-omics experiments for pancreatic cancer. At least some of these samples and study details overlap with those described in other embodiments. Multi-omics cancer biomarkers span the genotype-phenotype spectrum ( Figure 48 ).
[0427] Multi-omics assays capture individual and shared biological signals that separate cancer from non-cancer aggregations. Figure 49 The data in includes biclustering of statistically significant (non-cancer vs. cancer, adjusted P < 0.05) multi-omics biomarkers.
[0428] Variance decomposition illustrates that different aspects of biology can be uniquely captured by each molecular assay. Figure 50A - 50B The data in includes unsupervised variance decomposition (JIVE) of statistically significant (adjusted P < 0.05) biomarkers. Some key points are that there is shared biology (joint components) between different omics, and they can be used as an independent set of evidence to reveal shared biological signals. An additional key point is that there is also a lot of biology specific to each assay, especially methylation, lipidomics, metabolomics, and proteomics (individual components).
[0429] Examining across multi-omics readouts can prioritize biomarkers for further investigation. Figure 51 Overlap of statistically significant (adjusted P < 0.05) biomarkers including RNA-seq (protein-coding genes), proteomics (non-targeted + targeted), and copy number variable regions. Examining the overlap between multi-omics assays can focus on high-priority biomarker candidates. Figure 51 Two proteins that overlap with copy number changes in are E-cadherin and N-cadherin, and may be associated with epithelial-mesenchymal transition in pancreatic cancer. Additional experiments will be conducted to further understand the statistically significant overlapping genes and copy numbers, as well as gene and protein findings.
[0430] As Figure 52A - 52C shown, multi-omics readouts can also be statistically combined to improve the interpretation of biological processes, which includes non-targeted proteomics and RNA-seq + non-targeted proteomics.
[0431] Trend analysis shows the correlation of biomarker abundance with cancer stage ( Figure 53)。In this figure, groups of RNA, fragment (FRG), CNV, and protein (PRO) data can be seen. Trend analysis was performed using the one-sided Jonckheere-Tempstra test with the Bonferroni procedure for multiple hypothesis correction (adjusted P < 0.05). The identified markers showed a monotonic increase or decrease with cancer stage. Thus, the classifier or method herein can be used to distinguish cancer stages.
[0432] Some aspects can be further elaborated by referring to Table 14 Figure 53 . Any biomarker in this table can be used alone or in combination as a biomarker for pancreatic cancer.
[0433] Table 14
[0434] Marker Gene Symbol Trend ENST00000423451.5 ST6GAL1 ↓ ENST00000417443.3 SMIM10L2A ↓ ENST00000262487.5 ISM1 ↓ ENST00000505275.1 HAUS1P1 ↑ Q15063 - 3|NP1 POSTN ↑ P18827|NP3 SDC1 ↑
[0435] Example 11. A validated pancreatic cancer classifier based on a multi-omics panel and targeted mass spectrometry in a case-control study of 146 subjects
[0436] Overview
[0437] Given the current lack of early detection tools, the generally asymptomatic course, and the poor prognosis associated with late diagnosis, there is an urgent need for effective and reliable methods for detecting pancreatic cancer. This study used multi-omics analysis focused on the acute state to construct and validate a new 20-feature classifier that differentiates subjects with pancreatic ductal adenocarcinoma (PDAC) from non-cancer controls at all stages. The features included protein, metabolite, lipid, and RNA data from blood samples collected from 146 age- and sex-matched subjects, and exploratory analysis showed many potential differential signals. Repeated cross-validation (RCV) was performed on a training cohort of 74 subjects to build individual omics models using all features. The 5 features that contributed the most to each model were identified and used as input for a new RCV in which the training subjects were reshuffled. A final 20-feature multi-omics model was constructed, and examination of the model coefficients showed significant contributions from each omics type. The model was applied to 72 validation subjects in another cohort, and it achieved an area under the ROC curve (AUC) of 0.977 for all-stage classification, with a sensitivity of 80.8% at 99% specificity. The AUC for early (I / II stage) subjects was 0.965, with a sensitivity of 71.4% at 99% specificity. The model includes new combinations of both unknown and known PDAC-related analytes and demonstrates the value of combining multiple different omics to develop a clinically useful test for early pancreatic cancer detection.
[0438] Introduction
[0439] Pancreatic cancer is currently the fourth most common cause of cancer-related death in the United States, and demographic trends suggest that by 2030 it will become the second leading cause. Pancreatic ductal adenocarcinoma (PDAC) and its variants account for more than 90% of pancreatic malignancies. These are fearsome diagnoses because most cases are not detected until advanced stages, with 80%-85% of initial presentations representing incurable locally advanced or metastatic unresectable disease. This results in a low 5-year survival rate of approximately 10% for all stages. However, because early diagnosis has a significantly superior 5-year survival rate of over 40%, early detection (possibly in conjunction with peripheral blood biomarker testing) has the potential to lower the initial diagnostic stage and holds great promise for reducing PDAC-associated morbidity and mortality in appropriate screening populations. A variety of clinical and investigational biomarkers are in use, including CA19-95 and protein- and DNA methylation-based biomarkers. However, given the limited performance of these biomarkers, the US Preventive Services Task Force currently recommends against routine screening for PDAC.
[0440] PDAC is difficult to detect because the onset of clinical symptoms typically coincides with the progression of invasive growth and the loss of the opportunity for resection. In addition, the unique tumor microenvironment composed of PDAC-associated stroma creates an immune-privileged compartment that is refractory to the recent advances in immuno-oncology-based therapies. Given the complexity of PDAC progression, multiple signal inputs, such as those from different blood analytes, may be necessary to detect PDAC early enough for interventions that improve patient survival. Thus, compared to any single omics model, a multi-omics model that samples multiple physiological systems and pathways using a combination of orthogonal features (e.g., proteins, metabolites, lipids, and RNA) can exhibit excellent classification performance and can be used for clinical development. Here, a case-control study using PDAC subjects and age- and sex-matched non-cancer controls validated the feasibility of this approach. Using a broad, unbiased platform to collect analyte data for each omics type, individual omics classification models were used to select the most important features, and then these selected features were combined into a single multi-omics classifier. The final performance of this model for all and early PDAC was confirmed in a separate validation cohort.
[0441] Study Design and Subject Population
[0442] A case-control study was conducted that included PDAC subjects who were informed of their diagnosis but not treated and age- and sex-matched non-cancer control subjects. The subjects were from an ongoing IRB-approved observational study that collected various blood sample types (e.g., plasma, serum, Streck, and PAXgene tubes) for 5 different cancers and selected co-morbidity controls. For this analysis, a subset of 146 subjects was selected from 16 different sites over a period of more than 2 years, including 63 PDAC and 83 non-cancer control subjects. The PDAC subjects were enrolled from 14 of the sites, and the non-cancer controls were enrolled from 4 of the sites. The primary inclusion / exclusion criteria were based on newly diagnosed, biopsy-confirmed PDAC with no history of any other cancer or cancer for at least the previous 5 years. On average, blood samples from PDAC subjects were collected 21 days after the local hospital pathologist reported PDAC histopathology. Control subjects were determined to be cancer-free based on self-reported history but were allowed to include other non-relevant co-morbidities (e.g., diabetes) to better evaluate the generalization of the classification model to the intended trial population. There were no significant differences in the frequencies of 9 co-morbidities reported between PDAC subjects and control subjects. For each training set and validation set, there were no significant differences in sex, age, and race between PDAC subjects and control subjects, and there were no significant differences in the proportions of PDAC stages between the training set and the validation set. Although the mean distribution in PDAC stage was not protocol required, early-stage subjects (stage I and II; n = 20) and late-stage subjects (stage III and IV; n = 40) were included, with 3 subjects having incomplete staging records included.
[0443] Proteomics Data Acquisition and Primary Data Processing
[0444] To maximize the potential for new signal collection, an unbiased, non-analyte-specific protein data collection method was used. Plasma samples were processed according to the manufacturer's protocol using the standard 5 nanoparticle panel and 3 process controls on the Proteograph (Seer, Redwood City, CA) plasma sample preparation platform. The eluted peptide concentration was measured using a quantitative fluorescence peptide assay kit (Thermo Fisher, Waltham MA) and dried overnight at room temperature in a Centrivap vacuum concentrator (LabConco, Kansas City MO). Prior to use, the peptides were equilibrated for 30 minutes at room temperature and then reconstituted in a solution of LCMS-grade water (Honeywell, Charlotte, NC) with 0.1% formic acid (Thermo Fisher, Waltham, MA) on the Proteograph platform, which was spiked with the heavy-labeled retention time peptide standards -iRT (Biogynosys, Switzerland) and Pepcal (SciEX, Redwood City, CA) prepared according to the manufacturer's instructions. The separated peptides were reconstituted in solution by shaking at 1000 rpm for 10 minutes on an orbital shaker (Bioshake, Germany) at room temperature and briefly centrifuging (about 10 seconds) in a centrifuge (Eppendorf, Germany). The reconstituted peptides were loaded onto Evotip separation tips (Evosep, Denmark) and processed using a total of 600 ng of nanoparticle 1-4 peptides and 300 ng of nanoparticle 5 peptides according to the manufacturer's protocol. The processed tips were placed on an Evosep One LC system (Evosep, Denmark) and the peptides were separated using the Evosep LC gradient method with 60 samples / day on a reversed-phase 8 cm x 150 μM, 1.5 μM, column (Pepsep, Denmark).
[0445] Peptides were analyzed using parallel accumulation - serial fragmentation in data - independent acquisition (DIA) mode on a timsTOF Pro II (Bruker, Germany); the source capillary voltage was set to 1700 V and 200 °C. Precursors (MS1) across m / z 100 - 1700 and within an ion mobility window spanning 1 / K 0.84 - 1.31 V.s / cm2 were fragmented using collision energy following a linear step function in the range of 20 eV - 63 eV. The TIMS cell accumulation time was set to 100 milliseconds, and the ramp time was set to 85 milliseconds. The resulting MS / MS fragment spectra between m / z 390 - 1250 were analyzed using a DIA scheme with a 57 Da window (15 mass steps) where there was no mass / mobility overlap, which resulted in a cycle time of slightly less than 0.8 seconds. The primary MS data were processed into quantitative protein groups and peptide IDs using the Proteograph Analysis Suite (Seer) containing the DIANN search engine. For all proteomics analyses, unique nanoparticle - modified peptide sequences were the analyte features; thus, proteomics analyses occurred at the (potentially modified) peptide level.
[0446] Lipidomics and Metabolomics Data Acquisition and Primary Data Processing
[0447] Lipid data were obtained using multiple targeted liquid chromatography - mass spectrometry (LC - MS) assays, where target analytes were selected without any known association with pancreatic ductal adenocarcinoma (PDAC). Total lipid content was extracted using a single - phase organic extraction method. Five microliters of the cohort, NIST SRM1950, and pooled human plasma were placed in a 96 - well plate and spiked with 20 μL of a 1:20 (v / v) UltimateSPLASH mix (Avanti Polar, Alabaster, AL) working internal standard. To each sample - internal standard mixture, 475 μL of a 1:1 (v / v) butanol:methanol mixture was added and the mixture was vortexed at 500 rpm for 10 minutes at 4 °C. The mixture was incubated at 4 °C for 15 minutes and vortexed at 500 rpm for 10 minutes at 4 °C. The samples were incubated at 4 °C for an additional 15 minutes and finally centrifuged at 3500 rpm for 10 minutes. Approximately 300 μL of the extract was transferred to a clean collection plate and stored at - 20 °C until LC - MS processing. Two chromatographic separation methods were used to separate lipids using a binary gradient flow system. Data were collected using a SCIEX 7500 (SCIEX, Redwood City, CA) triple quadrupole mass spectrometer in multiple reaction monitoring (MRM) mode equipped with positive and negative electrospray ionization. For positive - mode lipids, a SCIEX LC AD (SCIEX, Redwood City, CA) liquid chromatography system and a Waters Acuity UPLC BEH C18 (50 X 2.1 mm X 1.7 μm) (Waters, Waltham, MA) column were used with a gradient elution at 0.5 mL / min and 50 °C. The gradient elution contained mobile phase A as water:acetonitrile (40:60 v / v) and mobile phase B as isopropanol:acetonitrile (90:10 v / v). For negative - mode lipids, a SCIEX LC AD liquid chromatography system and a Luna NH2 (100 X 2.0 mm X 3 μm) (Phenomenex, Torrance, CA) column were used with a gradient elution at 0.6 mL / min and 40 °C. The gradient elution contained mobile phase A as water:acetonitrile (50:50 v / v) and mobile phase B as dichloromethane:acetonitrile (7:93 v / v). For both separation methods, the autosampler temperature was maintained at 4 °C. The MQ4 algorithm was selected using SCIEX OS Analytics (SCIEX, Redwood City, CA) software to process positive - polarity and negative - polarity data separately. NIST SRM1950 and pooled plasma quality control samples were used to optimize peak integration parameters such as intensity threshold, signal - to - noise ratio, and smoothing parameters. These methods were used to process all samples. The processed data were manually reviewed and curated to ensure accurate peak integration, exported as text files, and used for downstream statistical analysis.
[0448] Metabolite data were obtained using multiple targeted LC-MS assays in which target analytes were not selected based on any known association with PDAC. Polar metabolites were extracted from 30 μL of human plasma, NIST SRM1950, and pooled plasma samples from the cohort using a 1:1 (v / v) water:methanol mixture. Briefly, 20 μL of QreSS1 and 2 (Cambridge, Tewksbury, MA) (working internal standards) were spiked into 30 μL of plasma samples, which were aliquoted into individual wells of a 96-deep well plate. Metabolites were extracted by dispensing 450 μL of a 1:1 (v / v) water:methanol mixture into each plasma sample. The sample-solvent mixture was vortexed at 1000 rpm for 5 min and kept at 4 °C. The mixture was then incubated at 4 °C for 60 min and centrifuged at 3000 rpm for 15 min at 4 °C. Data were collected using a SCIEX 7500 triple quadrupole mass spectrometer in MRM mode equipped with positive and negative electrospray ionization. Metabolites were separated using a SCIEX LC AD liquid chromatography system with a Kinetics F5 (150 x 2.1 mm x 2.6 μm) (Phenomenex, Torrance, CA) column and a gradient elution system at 0.2 mL / min and 40 °C. The gradient elution system contained mobile phase A as 2 mM ammonium acetate and 0.1% formic acid in aqueous solution, and mobile phase B as 0.1% formic acid in acetonitrile solution. For both separation methods, the autosampler temperature was kept at 4 °C. The MQ4 algorithm was selected using SCIEX OS Analytics to process positive and negative data separately. NIST SRM1950 and pooled plasma quality control samples were used to optimize peak integration parameters such as intensity threshold, signal-to-noise ratio, and smoothing parameters. This method was used to process samples in all studies. The processed data were manually reviewed and curated to ensure accurate peak integration, exported as text files, and used for downstream statistical analysis.
[0449] Transcriptomics Data Acquisition and Primary Data Processing
[0450] RNA-seq was performed on RNA extracted from PAXgene blood tubes using the Qiagen PAXgene Total RNA Kit according to the manufacturer's protocol. Using TruSeq Stranded Total RNA With Ribo-Zero TMPlusrRNA Depletion+Globin Reduction RNA Library PreparationPrepare a 100M paired-end (total 200M) read library for strand-specific 100bp reads. Use FastQC (v0.11.9) to perform quality control on the fastq files. Use the STAR aligner (v2.7.8a) to align the reads and PicardTools (v2.25.0) to deduplicate. Use RNA-SeqC (v2.4.2) for post-alignment quality control. Use RSEM (v1.3.3) for transcript quantification.
[0451] CA19-9 Data Acquisition
[0452] Evaluate CA19-9 levels using a clinical-grade assay (Invitrogen Human CA19-9 ELISA Kit [Catalog No. EHCA199]) according to the supplier's instructions.
[0453] Published RNA-Seq Data Analysis
[0454] RNA-Seq of various human tumors and normal tissues has been previously performed by TCGA1,2 and GTEx3, respectively. These original datasets have been combined and co-processed by others before. RSEM expected counts were used to analyze differential expression between 183 pancreatic tumors and 167 normal pancreatic samples using the DESeq2 package in R. Genes with very low expression were filtered by requiring a minimum count sum of 1000 in 350 samples before DESeq2 analysis. The ashr shrinkage estimator was used to moderate the fold change estimates. The fold changes and adjusted P-values for each gene highlighted in the text are as Figure 55 shown.
[0455] Avoiding Model Overfitting
[0456] Given the relatively large number of analytes involved in this multi-omics study and the moderate number of subjects in the training data, the risk of overfitting the data in the model is a potential concern. A number of methods to mitigate and evaluate this risk were implemented. First, a conservative split of the total subject population was chosen, with approximately 60% in the training set and 40% in the validation set. By increasing the size of the validation set, the ability to detect overfitting (if it occurs) increases, even though the ability to identify important classifier components decreases. Second, the study design incorporated intentional differences in the enrollment date and enrollment site between the PDAC group and the control group. These steps reduced the risk of systematic bias between groups carried over from the training set to the validation set. Third, an extensive cross-validation design was employed when optimizing the model engine parameters and important feature selection. Ten rounds of 10-fold cross-validation, while computationally intensive for many input features, is a robust method to avoid overfitting. Finally, the training subject data groups (e.g., PDAC vs. control) were permutated to determine whether the two-stage process described herein could produce a final multi-omics model with similar performance in the validation set as observed in the individual omics models. The class permutations were repeated 10 times from the initial individual omics RCVs fed into the top features RCV of the final combination, and a validation set ROC AUC of 0.629 (±0.113 standard deviation) was achieved. This was significantly different from the validation ROC AUC (0.977, p = 4.393e-06) achieved using the correct class assignments, and confirmed that extreme overfitting did not drive the observed high performance, although a statistically significant positive bias (AUC 0.5, p = 0.005457) was observed compared to random performance.
[0457] Exploratory Data Analysis, Univariate and Multivariate
[0458] For exploratory data analysis (EDA), all 146 subjects were used for univariate and multivariate comparisons. The R statistical computing language and appropriate packages, as well as appropriate additional packages, were used for all analyses. Generally, after primary data processing, the data were normalized using appropriate methods for each omics type. Briefly, median normalization was performed using the major common features of the proteomics and metabolomics data; features present in 90% of the proteomics data and 95% of the metabolomics data of the subjects were considered as the reference set for calculating the median of individual subjects and the median normalization factor for the subjects. For lipidomics data, median scaling of the samples was performed using spiked reference standards. For RNA data, normalization was performed using the DESeq2 algorithm.
[0459] Features were filtered to those present in ≥50% of at least 1 of these categories (PDAC or non-cancer). When necessary (e.g., principal component analysis [PCA], etc.), missing values were assumed to be missing below the limit of detection rather than missing at random and were imputed with the minimum value of that analyte in the sample set. A non-parametric Wilcoxon test with Bonferroni multiple testing correction was used for univariate comparisons between groups, using only the actual non-imputed values. For Gene Ontology Biological Process (GOBP) term enrichment analysis of proteomic and transcriptomic types, Fisher's test was used to assess the significance of the proportion differences, and the differences were reported as the logarithm of the odds ratio.
[0460] Training of Initial, Individual Omics Machine Learning-Based Classifiers Using All Available Features per Class
[0461] For machine learning-based classification model training and validation, different splits of 146 samples were created, including 74 subjects (n = 37 PDAC and n = 37 controls) for RCV and final model building and 72 subjects (n = 26 PDAC and n = 46 controls) for validation. The proportion of PDAC cancer stages was maintained across the splits. To improve the generalization of the results and avoid possible confounding factors, the splits of training and validation subjects were stratified by collection site and enrollment date. For the training set, PDAC subjects included the top 60% of subjects enrolled in the study across 10 sites, and control subjects were selected from 3 of the 4 control enrollment sites. For the validation set, PDAC subjects included the last 40% of enrolled subjects, and control subjects were selected from a unique site. After the training / validation split, features were filtered to those present in ≥50% of at least 1 of these categories.
[0462] A robust machine learning modeling engine, XGBoost, which is an implementation of gradient-boosted ensemble tree methods, was deployed. Given the moderate number of subjects available for model testing, a general approach was taken to avoid overfitting, and 10 repeated 10-fold cross-validation as described above was used to improve the quality of hyperparameter selection and validate performance estimates. Prior to RCV, analyte value scaling was the only form of feature engineering aside from missing value imputation. The selection of 50 hyperparameter combinations and the distribution of candidate values were tested using a Latin hypercube design. For computational efficiency, null analyses were performed during repeated cross-validation, and after an initial burn-in period of 5 (for RNA) or 10 (for proteins) cycles, parameter combination results were compared by ANOVA, and those models that were unlikely to achieve better than the current best combination were removed from further consideration. After parameter selection, a final model using all training subjects was created, where attention was paid to the feature importance assigned by the algorithm to the model. This basic approach was used for all four omics types.
[0463] Construction of Multomics Classifier Using Top 5 Features from Each Individual Omics Model
[0464] Although each individual omics final training model was evaluated in the validation set, the main purpose of building the final model was to evaluate feature importance and select the top 5 features from each omics type. This method used all the analytes in each individual omics type in XGBoost RCV for feature selection (top 5 of each type), and then used the combined 20 selected features and the final reshuffled RCV with GLMnet to create the final multi-omics model. Although it might be preferred to have separate subject sets for feature selection and model creation, the training subjects were not split and the power was not reduced, but rather the subjects were reshuffled into new repeated-folds. No data from the held-out validation set was used for feature selection RCV or final model creation RCV. This method reduced the chance of overfitting the mode...
Claims
1. A method for evaluating pancreatic cancer, comprising: Obtain a data set comprising biomarker measurements from a biological fluid sample from a subject suspected of having pancreatic cancer, the biomarkers including AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2 or APOA1, or a combination thereof; and Apply a classifier to the data set to evaluate the pancreatic cancer in the subject.
2. A detection method, comprising: Measuring biomarkers including AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2 or APOA1 or a combination thereof in a biological fluid sample of a subject suspected of having pancreatic cancer to obtain biomarker measurement values.
3. The method according to claim 2, further comprising applying a classifier to the biomarker measurement values to evaluate the pancreatic cancer in the subject.
4. The method according to claim 1 or 3, wherein the classifier comprises a performance determined by: having an area under the curve (AUC) of greater than 0.85, greater than 0.86, greater than 0.87, greater than 0.88, greater than 0.89, greater than 0.90, greater than 0.91, greater than 0.92, greater than 0.93, greater than 0.94, greater than 0.95, greater than 0.96, greater than 0.97, greater than 0.98 or greater than 0.99 in the discrimination between the pancreatic cancer and the absence of the pancreatic cancer in a receiver operating characteristic (ROC) curve.
5. The method according to claim 1 or any one of claims 3 - 4, wherein the classifier comprises a performance determined by: having a sensitivity of greater than 50%, greater than 55%, greater than 60%, greater than 65%, greater than 70%, greater than 75%, greater than 80%, greater than 85%, greater than 86%, greater than 87%, greater than 88%, greater than 89%, greater than 90%, greater than 91%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98% or greater than 99% in the identification between the pancreatic cancer and the absence of the pancreatic cancer.
6. The method according to claim 1 or any one of claims 3 - 5, wherein the classifier comprises a performance determined by: having a specificity of greater than 80%, greater than 81%, greater than 82%, greater than 83%, greater than 84%, greater than 85%, greater than 86%, greater than 87%, greater than 88%, greater than 89%, greater than 90%, greater than 91%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98% or greater than 99% in the discrimination between the pancreatic cancer and the absence of the pancreatic cancer.
7. The method according to any one of claims 1 or 3 - 6, wherein evaluating the pancreatic cancer in the subject comprises identifying the data set or the biomarker measurement as indicating the pancreatic cancer in the subject, or identifying the data set or the biomarker measurement as indicating the absence of the pancreatic cancer in the subject.
8. The method according to claim 7, further comprising administering a pancreatic cancer treatment to the subject when the data set or the biomarker measurement is identified as indicating the pancreatic cancer, and observing or treating the subject without administering the pancreatic cancer treatment to the subject when the data set or the biomarker measurement is identified as indicating the absence of the pancreatic cancer.
9. The method according to any one of the preceding claims, wherein the biomarker comprises AACT.
10. The method according to any one of the preceding claims, wherein the biomarker comprises A1AT.
11. The method according to any one of the preceding claims, wherein the biomarker comprises A2GL.
12. The method according to any one of the preceding claims, wherein the biomarker comprises AMPN.
13. The method according to any one of the preceding claims, wherein the biomarker comprises LBP.
14. The method according to any one of the preceding claims, wherein the biomarker comprises ICAM1.
15. The method according to any one of the preceding claims, wherein the biomarker comprises PIGR.
16. The method according to any one of the preceding claims, wherein the biomarker comprises CO5.
17. The method according to any one of the preceding claims, wherein the biomarker comprises S10A8.
18. The method according to any one of the preceding claims, wherein the biomarker comprises CO2.
19. The method according to any one of the preceding claims, wherein the biomarker comprises CO9.
20. The method according to any one of the preceding claims, wherein the biomarker comprises ITIH3.
21. The method according to any one of the preceding claims, wherein the biomarker comprises RET4.
22. The method according to any one of the preceding claims, wherein the biomarker comprises FCG3A.
23. The method according to any one of the preceding claims, wherein the biomarker comprises TETN.
24. The method according to any one of the preceding claims, wherein the biomarker comprises CRP.
25. The method according to any one of the preceding claims, wherein the biomarker comprises NOE1.
26. The method according to any one of the preceding claims, wherein the biomarker comprises F13B.
27. The method according to any one of the preceding claims, wherein the biomarker comprises APOA2.
28. The method according to any one of the preceding claims, wherein the biomarker comprises APOA1.
29. The method according to any one of the preceding claims, wherein the biomarker comprises CA19-9.
30. The method according to any one of the preceding claims, wherein the biomarker comprises two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, 12 or more, 14 or more, 16 or more, 18 or more, or 20 or more of AACT, A1AT, A2GL, AMPN, LBP, ICAM1, PIGR, CO5, S10A8, CO2, CO9, ITIH3, RET4, FCG3A, TETN, CRP, NOE1, F13B, APOA2, APOA1, or CA19-9.
31. The method according to any one of the preceding claims, wherein the biomarker measurement is obtained by adding an internal standard for any one of the biomarkers present in the sample to the sample.
32. The method of claim 31, wherein the internal standard is labeled.
33. The method of claim 31, wherein the internal standard is isotopically labeled.
34. The method according to any one of the preceding claims, wherein the biomarker measurement is obtained using mass spectrometry.
35. The method according to any one of the preceding claims, wherein the biomarker measurement is obtained using an immunoassay.
36. The method according to any one of the preceding claims, wherein the biomarker measurement is obtained using a molecular probe.
37. The method according to any one of the preceding claims, wherein the biomarker measurement is obtained using chromatography.
38. The method of claim 1 or any one of claims 3-37, wherein the classifier identifies the stage of pancreatic cancer in the subject.
39. The method according to claim 38, wherein the pancreatic cancer comprises stage I pancreatic cancer or stage II pancreatic cancer.
40. The method according to claim 38, wherein the pancreatic cancer comprises stage III pancreatic cancer or stage IV pancreatic cancer.
41. The method according to any one of the preceding claims, wherein the pancreatic cancer comprises pancreatic ductal adenocarcinoma (PDAC).
42. The method according to any one of the preceding claims, wherein the biological fluid comprises pancreatic cyst fluid, urine, blood, plasma, and / or serum.
43. The method according to claim 1 or any one of claims 3 - 41, wherein when the cancer assessment method indicates that the subject has a probability exceeding a predetermined threshold of having the pancreatic cancer, the method further comprises treating the subject with a subsequent pancreatic cancer treatment for treating the pancreatic cancer or advising the subject to undergo such treatment.
44. The method according to claim 43, wherein the subsequent pancreatic cancer treatment is selected from the group consisting of: surgery for pancreatic cancer, radiotherapy for pancreatic cancer, chemotherapy for pancreatic cancer, ablation therapy for pancreatic cancer, and immunotherapy for pancreatic cancer.
45. The method according to claim 43 or 44, wherein the predetermined threshold is a probability of having the pancreatic cancer greater than 10%, greater than 20%, greater than 30%, greater than 40%, greater than 50%, greater than 60%, greater than 70%, greater than 80%, or greater than 90%.
46. The method according to claim 43 or 44, wherein the subsequent pancreatic cancer treatment comprises a biopsy.
47. The method according to claim 43 or 44, wherein the method further comprises pancreatic imaging.
48. The method according to claim 47, wherein the pancreatic imaging is performed using ultrasound or computed tomography.
49. The method according to any one of the preceding claims, wherein the subject is a mammal.
50. The method according to any one of the preceding claims, wherein the subject is a human.
51. A method for generating a multi - omics classifier, the method comprising: Obtain first omics data of a first omics data type; Obtain second omics data of a second omics data type different from the first omics data type, wherein the first omics data and the second omics data correspond to biomolecules present in a biological sample of a subject; Generate a first classifier of a biological state using features of the first omics data; Generate a second classifier of the biological state using features of the second omics data; Assign feature importance scores to the features of the first classifier and the second classifier; Select the top features of the first classifier, and select the top features of the second classifier; And Generate a combined classifier using the selected top features of the first classifier and the second classifier.
52. The method according to claim 51, wherein generating the first classifier using the features of the first omics data includes using all available features of the first omics data.
53. The method according to claim 51, wherein generating the second classifier using the features of the second omics data includes using all available features of the second omics data.
54. The method according to claim 51, wherein generating the first classifier using the features of the first omics data includes performing machine learning with the features of the first omics data.
55. The method according to claim 51, wherein generating the second classifier using the features of the second omics data includes performing machine learning with the features of the second omics data.
56. The method according to claim 51, wherein generating the first classifier using the features of the first omics data includes using the features of the first omics data for repeated cross-validation (RCV).
57. The method according to claim 51, wherein generating the second classifier using the features of the second omics data includes using the features of the second omics data for RCV.
58. The method according to claim 51, wherein the features of the first omics data and the second omics data include measurements of biomolecules.
59. The method according to claim 51, wherein the selected top features of the first classifier include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, or 20 or more features.
60. The method according to claim 51, wherein the selected top features of the second classifier include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, or 20 or more features.
61. The method according to claim 51, wherein the selected top features of the first classifier include the same number of features as the selected top features of the second classifier.
62. The method according to claim 51, wherein generating the combined classifier comprises performing RCV using selected top features of the first classifier.
63. The method according to claim 51, wherein generating the combined classifier comprises performing RCV using selected top features of the second classifier.
64. The method according to claim 51, wherein generating the combined classifier comprises using the features selected from the first classifier and the second classifier in a repeated and folded second RCV shuffling of the subjects into new groups, the features being from a first shuffling of the subjects into RCV repeats and folds.
65. The method according to claim 51, wherein generating the combined classifier comprises using a resampling method.
66. The method according to claim 65, wherein the resampling method is nested cross-validation (NCV).
67. The method according to claim 65, wherein the resampling method is leave-one-out cross-validation (LOOCV).
68. The method according to claim 51, wherein generating the combined classifier comprises excluding features below an importance threshold.
69. The method according to claim 51, further comprising identifying features of the combined classifier below a predetermined importance threshold, and training a final combined classifier that excludes features below the predetermined importance threshold.
70. The method according to claim 51, wherein the combined classifier comprises a linear classifier, a logistic classifier, or a decision tree.
71. The method according to claim 51, wherein the first omics data and the second omics data are selected from proteomics data, metabolomics data, lipidomics data, transcriptomics data, and genomics data.
72. The method according to claim 51, wherein the first omics data comprises measurements of biomolecules captured by a first particle type, and the second omics data comprises measurements of biomolecules captured by a second particle.
73. The method according to claim 72, wherein the first particle type and the second particle type are physiochemically different from each other.
74. The method according to claim 72, wherein the first particle type and the second particle type comprise lipid particles, metal particles, silica particles, or polymer particles.
75. The method according to claim 72, wherein the first particle type and the second particle type comprise nanoparticles.
76. The method of claim 51, further comprising obtaining third omics data of a third omics data type corresponding to a biomolecule present in the biological sample, generating a third classifier of the biological state using features of the third omics data, assigning a feature importance score to the features of the third classifier, and selecting top features of the third classifier, and wherein generating the combined classifier comprises using the selected top features of the first classifier, second classifier, and third classifier.
77. The method of claim 76, further comprising obtaining fourth omics data of a fourth omics data type corresponding to a biomolecule present in the biological sample, generating a fourth classifier of the biological state using features of the fourth omics data, assigning a feature importance score to the features of the fourth classifier, and selecting top features of the fourth classifier; and wherein generating the combined classifier comprises using the selected top features of the first classifier, second classifier, third classifier, and fourth classifier.
78. The method of claim 77, wherein the first omics, the second omics, the third omics, and the fourth omics are independently selected from proteomics data, metabolomics data, lipidomics data, transcriptomics data, and genomics data.
79. The method of claim 76, wherein the first omics data comprises proteomics data, the second omics data comprises metabolomics data, the third omics data comprises lipidomics data, and the fourth omics data comprises transcriptomics data.
80. The method of claim 51, wherein the combined classifier identifies a subject as having the biological state and as not having the biological state with at least 70% sensitivity and 99% specificity.
81. The method of claim 51, wherein the combined classifier identifies a subject as having the biological state and as not having the biological state with a performance characterized by: a receiver operating characteristic curve (ROC) having an area under the curve (AUC) of at least 0.
90.
82. The method of claim 51, wherein the combined classifier identifies a subject as having the biological state and as not having the biological state with a performance characterized by: a receiver operating characteristic curve (ROC) having an area under the curve (AUC) of at least 0.
95.
83. The method of claim 51, wherein the biological state comprises a disease.
84. The method of claim 83, wherein the disease comprises cancer.
85. The method of claim 84, wherein the cancer comprises pancreatic cancer.
86. The method according to claim 85, wherein the pancreatic cancer comprises pancreatic ductal adenocarcinoma (PDAC).
87. The method according to claim 85, wherein the pancreatic cancer comprises stage I pancreatic cancer or stage II pancreatic cancer.
88. The method according to claim 85, wherein the pancreatic cancer comprises stage III pancreatic cancer or stage IV pancreatic cancer.
89. The method according to claim 51, wherein the biological sample comprises a biological fluid.
90. The method according to claim 89, wherein the biological fluid comprises blood, serum or plasma.
91. The method according to claim 89, wherein the biological fluid is substantially cell-free.
92. Use of a classifier generated by the method according to any one of claims 51 - 91 in evaluating the biological state of a subject using biomolecular data obtained from a sample of the subject.
93. The use according to claim 92, which further comprises administering a disease treatment to the subject based on the evaluation.
94. The method according to claim 51, wherein a model is used to generate the combined classifier.
95. The method according to claim 94, wherein the model is linear regression, logistic regression or a decision tree.
96. The method according to claim 51, wherein the combined classifier comprises coefficients associated with each of the selected top features.
97. The method according to claim 51, wherein a training group of fewer than 100 subjects is used to train the first classifier, the second classifier and the combined classifier.
98. The method according to claim 51, wherein the same training group is used to train the first classifier, the second classifier and the combined classifier.
99. The method according to claim 51, which further comprises applying a method for reducing model overfitting to the classifier generation.
100. The method according to claim 99, wherein the method for reducing model overfitting comprises splitting more than 50% of the total subject population between a training set and a validation set.
101. The method according to claim 99, wherein the method for reducing model overfitting comprises incorporating intentional differences into the enrollment dates and enrollment locations of a test group and a control group.
102. The method according to claim 99, wherein the method for reducing model overfitting comprises an extensive cross-validation design when optimizing model engine parameters and important feature selection.
103. The method according to claim 99, wherein the method for reducing model overfitting includes randomly arranging the training object data groups.
104. The method according to claim 51, wherein the top features of the first classifier are selected from at least 500, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,500, at least 15,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000 features of the first omics data.
105. The method according to claim 51, wherein the top features of the second classifier are selected from at least 1X, at least 10X, at least 100X, at least 1,000X, at least 10,000X, at least 100,000X the number of features selected for the top features of the first classifier.
106. The method according to claim 51, wherein the feature importance score includes a cumulative relative importance rating.
107. The method according to claim 51, wherein the feature importance score includes a combined cumulative relative classification ability.
108. The method according to claim 51, wherein the feature importance score includes a relative importance rank related to classifier performance.
109. A method for creating a classifier, the method comprising: Obtain multi-omics data from samples of an intended test population, wherein the multi-omics data comprises omics data sets representative of physiological systems and, when combined, yields an improved classifier for the intended test population.
110. The method according to claim 109, wherein the classifier includes features combined with a prediction model.
111. The method according to claim 110, further comprising assigning a feature importance score to each feature relative to other features in the same omics group, and selecting numbers based on the feature importance score.
112. The method according to claim 110, wherein the features are selected from different omics data sets.
113. The method according to claim 109, wherein the physiological systems are different physiological systems.
114. The method according to claim 109, wherein the physiological systems are assigned statistical weights.
115. The method according to claim 114, further comprising selecting features of an improved classifier.
116. The method according to claim 115, wherein the selection of the features of the improved classifier includes combining the feature importance score of the features and the statistical weights of the physiological systems represented by the features.
117. The method according to claim 51, wherein assigning the feature importance score to each feature further comprises assigning one or more biological processes associated with the feature.
118. The method according to claim 117, wherein the one or more biological processes include human biological processes.
119. The method according to claim 117, wherein the one or more biological processes include Gene Ontology - Biological Process.
120. The method according to claim 117, wherein the selection of the top features further comprises calculating the total number of biological processes of the top features to generate a combined classifier.
121. The method according to claim 120, wherein the combined classifier has at least a certain number of biological processes represented by the top features.
122. The method according to claim 117, wherein assigning one or more biological processes further comprises calculating the significance of the association.
123. The method according to claim 122, wherein the significance of the association is calculated based on a formal test of statistical significance.
124. The method according to claim 123, wherein the formal test of statistical significance includes a log odds ratio (LOR) calculation.
125. The method according to claim 124, wherein the log odds ratio (LOR) calculation comprises the following equation: LOR = ln((association of the specific process of the first omics data type / total association of all processes of the first omics data type - instances of the specific process of the first omics data type) / (association of the specific process of the second omics data type / total association of all processes of the second omics data type - instances of the specific process of the second omics data type)); and using Fisher's test for significance of difference in proportions and Bonferroni correction of the raw p - value; where a positive LOR indicates significance for the first omics data type, and a negative LOR indicates significance for the second omics data type.
126. The method according to claim 125, wherein at a p - value < 0.05, an LOR greater than 0.5 or less than - 0.5 is associated with a feature of the omics data set that is significant for the process.
127. The method according to claim 51, wherein the subjects include a first set of training subjects and a second set of training subjects. The method of claim 51, wherein generating the first classifier and the second classifier comprises using omics data corresponding to biomolecules present in the biological samples of the first set of training subjects. The method of claim 51, wherein generating the combined classifier further comprises using omics data corresponding to biomolecules present in the biological samples of the second set of training subjects. A method for detecting pancreatic cancer, comprising: (a) Obtain biomarkers from a biological fluid sample of a subject; and (b) Apply a classifier to the biomarkers to evaluate the pancreatic cancer, wherein the classifier discriminates between biological fluid samples from subjects with and without pancreatic cancer with a performance characterized by: a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least 0.9; and Wherein the biomarker includes any one of the following peptides: GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1), TFVIIPELVLPNR (SEQ ID NO.2), TFVIIPELVLPNR (SEQ ID NO.2), DSC (UniMod:4) TMRPSSLGQGAGEVWLR (SEQ ID NO.3), DNC (UniMod:4) PHLPNSGQEDFDK (SEQ ID NO.4), GLVLGAGWAEGYLR (SEQ ID NO.5), LVFNPDQEDLDGDGRGDIC (UniMod:4) K (SEQ ID NO.6), AFDLYFVLDK (SEQ ID NO.7), VFLVGNVEIR (SEQ ID NO.8), RVSPVGETYIHEGLK (SEQ ID NO.9), ASEQIYYENR (SEQ ID NO.10), VLPGGDTYMHEGFER (SEQ ID NO.11), AVDIPHMDIEALK (SEQ IDNO.12), AMGIMNSFVNDIFER (SEQ ID NO.13), MPEQEYEFPEPR (SEQ ID NO.14), SGVISDTELQQALSNGTWTPFNPVTVR (SEQ ID NO.15), M (UniMod:35) EDVNSNVNADQEVR (SEQ IDNO.16), VGHDYQWIGLNDK (SEQ ID NO.17), HAEC (UniMod:4) IYLGHFSDPMYK (SEQ ID NO.18) or NGIFWGTWPGVSEAHPGGYK (SEQ ID NO.19); any one of the following RNAs: ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, ENST00000424185.7, ENST00000652176.1, ENST00000392593.9, ENST00000532853.5, ENST00000429947.1, ENST00000580914.1, ENST00000368205.7, ENST00000531709.6, ENST00000524817.5, ENST00000651281.1, ENST00000499685.2, ENST00000311921.8, ENST00000472111.5, ENST00000585172.2, ENST00000287713.7 or ENST00000547687.2; any one of the following lipids: NEG_PC(18:2_20:5)+AcO, POS_DAG(18:1_20:0)+NH4, NEG_PE(O-16:0_22:6)-H, NEG_PC(18:2_20:3)+AcO, POS_CER(d18:1 / 18:0)+H, POS_CE(22:0)+NH4, NEG_PE(14:0_22:5)-H, NEG_PC(20:5_20:5)+AcO, POS_PE(P-18:0_18:3)+H, NEG_PE(O-16:0_20:3)-H, POS_CE(18:3)+NH4, NEG_PE(O-18:0_22:5)-H, NEG_PE(O-18:0_20:5)-H, POS_PE(P-20:0_20:3)+H, NEG_PE(O-16:0_20:2)-H, POS_CER(d18:1 / 24:0)+H, NEG_PA(20:1_20:3)-H, NEG_PA(20:0_20:5)-H, POS_CE(20:0)+NH4 or NEG_PC(16:1_20:3)+AcO; or any one of the following metabolites: NEG_AICAR POS_cystine, NEG_CMP, NEG_gentiobioside, POS_creatine, POS_imidazoleacetic acid, POS_inosine, NEG_n-isovalerylglycine, NEG_glucose-6-phosphate, POS_metanephrine, NEG_N-acetylglutamate, NEG_5-thymidylate (dTMP), POS_UMP, NEG_fructose-6-phosphate, NEG_cystine, POS_pantothenol, POS_guanine, NEG_shikimic acid, POS_1-methylimidazoleacetate or POS_flavonoid 2. The method of claim 130, wherein the biomarker comprises two or more of at least one peptide, at least one RNA, at least one lipid, and at least one metabolite. The method of claim 130, wherein the biomarker comprises three or more of at least one peptide, at least one RNA, at least one lipid, and at least one metabolite. The method of claim 130, wherein the biomarker comprises at least one peptide, at least one RNA, at least one lipid, and at least one metabolite.
134. The method according to claim 130, wherein the biomarker comprises any one of the following peptides: GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1), TFVIIPELVLPNR (SEQ ID NO.2), TFVIIPELVLPNR (SEQ ID NO.2), DSC (UniMod:4)TMRPSSLGQGAGEVWLR (SEQ ID NO.3), DNC (UniMod:4)PHLPNSGQEDFDK (SEQ ID NO.4), GLVLGAGWAEGYLR (SEQ ID NO.5), LVFNPDQEDLDGDGRGDIC (UniMod:4)K (SEQ ID NO.6), AFDLYFVLDK (SEQ ID NO.7), VFLVGNVEIR (SEQ ID NO.8), RVSPVGETYIHEGLK (SEQ ID NO.9), ASEQIYYENR (SEQ ID NO.10), VLPGGDTYMHEGFER (SEQ IDNO.11), AVDIPHMDIEALK (SEQ ID NO.12), AMGIMNSFVNDIFER (SEQ ID NO.13), MPEQEYEFPEPR (SEQ ID NO.14), SGVISDTELQQALSNGTWTPFNPVTVR (SEQ ID NO.15), M (UniMod:35)EDVNSNVNADQEVR (SEQ ID NO.16), VGHDYQWIGLNDK (SEQ ID NO.17), HAEC (UniMod:4)IYLGHFSDPMYK (SEQ ID NO.18) or NGIFWGTWPGVSEAHPGGYK (SEQ ID NO.19).
135. The method according to claim 134, wherein the biomarker comprises 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, or 19 or more of the peptides.
136. The method according to claim 130, wherein the biomarker comprises any one of the following RNAs: ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, ENST00000424185.7, ENST00000652176.1, ENST00000392593.9, ENST00000532853.5, ENST00000429947.1, ENST00000580914.1, ENST00000368205.7, ENST00000531709.6, ENST00000524817.5, ENST00000651281.1, ENST00000499685.2, ENST00000311921.8, ENST00000472111.5, ENST00000585172.2, ENST00000287713.7 or ENST00000547687.
2.
137. The method according to claim 136, wherein the biomarker comprises 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more or 20 or more of the said RNAs. The method according to claim 130, wherein the biomarker comprises any one of the following lipids: NEG_PC(18:2_20:5)+AcO, POS_DAG(18:1_20:0)+NH4, NEG_PE(O-16:0_22:6)-H, NEG_PC(18:2_20:3)+AcO, POS_CER(d18:1 / 18:0)+H, POS_CE(22:0)+NH4, NEG_PE(14:0_22:5)-H, NEG_PC(20:5_20:5)+AcO, POS_PE(P-18:0_18:3)+H, NEG_PE(O-16:0_20:3)-H, POS_CE(18:3)+NH4, NEG_PE(O-18:0_22:5)-H, NEG_PE(O-18:0_20:5)-H, POS_PE(P-20:0_20:3)+H, NEG_PE(O-16:0_20:2)-H, POS_CER(d18:1 / 24:0)+H, NEG_PA(20:1_20:3)-H, NEG_PA(20:0_20:5)-H, POS_CE(20:0)+NH4 or NEG_PC(16:1_20:3)+AcO. The method according to claim 138, wherein the biomarker comprises 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more or 20 or more of the lipids. The method according to claim 130, wherein the biomarker comprises any one of the following metabolites: NEG_AICAR, POS_cystine, NEG_CMP, NEG_gentiobiose, POS_creatine, POS_imidazoleacetic acid, POS_inosine, NEG_n-isovalerylglycine, NEG_glucose-6-phosphate, POS_metanephrine, NEG_N-acetylglutamic acid, NEG_5-thymidylate (dTMP), POS_UMP, NEG_fructose-6-phosphate, NEG_cystine, POS_pantothenol, POS_guanine, NEG_shikimic acid, POS_1-methylimidazoleacetate or POS_flavone 2.
141. The method according to claim 140, wherein the biomarker comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, seventeen or more, eighteen or more, nineteen or more, or twenty or more of the metabolites.
142. The method according to claim 130, wherein the classifier comprises a performance characterized by a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least 0.
90.
143. The method according to claim 130, wherein the subject is suspected of having the pancreatic cancer.
144. The method according to claim 130, further comprising administering a pancreatic cancer treatment to the subject when the subject has pancreatic cancer.
145. The method according to claim 130, further comprising monitoring the subject when the subject does not have the pancreatic cancer.
146. A method for treating pancreatic cancer, the method comprising: Administer a pancreatic cancer treatment to a subject having the pancreatic cancer, wherein the pancreatic cancer is evaluated by a method comprising: (a) Obtain biomarkers from a biological fluid sample of the subject; and (b) Apply a classifier to the biomarkers to evaluate the pancreatic cancer, wherein the classifier discriminates between biological fluid samples from subjects with and without pancreatic cancer with a performance characterized by: a receiver operating characteristic (ROC) curve having an average or median area under the curve (AUC) of at least 0.9; and Wherein the biomarker includes any one of the following peptides: GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1), TFVIIPELVLPNR (SEQ ID NO.2), TFVIIPELVLPNR (SEQ ID NO.2), DSC (UniMod:4) TMRPSSLGQGAGEVWLR (SEQ ID NO.3), DNC (UniMod:4) PHLPNSGQEDFDK (SEQ ID NO.4), GLVLGAGWAEGYLR (SEQ ID NO.5), LVFNPDQEDLDGDGRGDIC (UniMod:4) K (SEQ ID NO.6), AFDLYFVLDK (SEQ ID NO.7), VFLVGNVEIR (SEQ ID NO.8), RVSPVGETYIHEGLK (SEQ ID NO.9), ASEQIYYENR (SEQ ID NO.10), VLPGGDTYMHEGFER (SEQ ID NO.11), AVDIPHMDIEALK (SEQ IDNO.12), AMGIMNSFVNDIFER (SEQ ID NO.13), MPEQEYEFPEPR (SEQ ID NO.14), SGVISDTELQQALSNGTWTPFNPVTVR (SEQ ID NO.15), M (UniMod:35) EDVNSNVNADQEVR (SEQ IDNO.16), VGHDYQWIGLNDK (SEQ ID NO.17), HAEC (UniMod:4) IYLGHFSDPMYK (SEQ ID NO.18) or NGIFWGTWPGVSEAHPGGYK (SEQ ID NO.19); any one of the following RNAs: ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, ENST00000424185.7, ENST00000652176.1, ENST00000392593.9, ENST00000532853.5, ENST00000429947.1, ENST00000580914.1, ENST00000368205.7, ENST00000531709.6, ENST00000524817.5, ENST00000651281.1, ENST00000499685.2, ENST00000311921.8, ENST00000472111.5, ENST00000585172.2, ENST00000287713.7 or ENST00000547687.2; any one of the following lipids: NEG_PC(18:2_20:5)+AcO, POS_DAG(18:1_20:0)+NH4, NEG_PE(O-16:0_22:6)-H, NEG_PC(18:2_20:3)+AcO, POS_CER(d18:1 / 18:0)+H, POS_CE(22:0)+NH4, NEG_PE(14:0_22:5)-H, NEG_PC(20:5_20:5)+AcO, POS_PE(P-18:0_18:3)+H, NEG_PE(O-16:0_20:3)-H, POS_CE(18:3)+NH4, NEG_PE(O-18:0_22:5)-H, NEG_PE(O-18:0_20:5)-H, POS_PE(P-20:0_20:3)+H, NEG_PE(O-16:0_20:2)-H, POS_CER(d18:1 / 24:0)+H, NEG_PA(20:1_20:3)-H, NEG_PA(20:0_20:5)-H, POS_CE(20:0)+NH4 or NEG_PC(16:1_20:3)+AcO; or any one of the following metabolites: NEG_AICAR POS_cystine, NEG_CMP, NEG_gentiobioside, POS_creatine, POS_imidazoleacetic acid, POS_inosine, NEG_n-isovalerylglycine, NEG_glucose-6-phosphate, POS_metanephrine, NEG_N-acetylglutamate, NEG_5-thymidylate (dTMP), POS_UMP, NEG_fructose-6-phosphate, NEG_cystine, POS_pantothenol, POS_guanine, NEG_shikimic acid, POS_1-methylimidazoleacetate or POS_flavone 2.
147. The method according to claim 146, wherein the biomarker comprises two or more of at least one peptide, at least one RNA, at least one lipid, and at least one metabolite.
148. The method according to claim 146, wherein the biomarker comprises three or more of at least one peptide, at least one RNA, at least one lipid, and at least one metabolite.
149. The method according to claim 146, wherein the biomarker comprises at least one peptide, at least one RNA, at least one lipid, and at least one metabolite.
150. The method according to claim 146, wherein the biomarker comprises any one of the following peptides: GAGGQSMSEAPTGDHAPAPTR (SEQ ID NO.1), TFVIIPELVLPNR (SEQ ID NO.2), TFVIIPELVLPNR (SEQ ID NO.2), DSC (UniMod:4)TMRPSSLGQGAGEVWLR (SEQ ID NO.3), DNC (UniMod:4)PHLPNSGQEDFDK (SEQ ID NO.4), GLVLGAGWAEGYLR (SEQ ID NO.5), LVFNPDQEDLDGDGRGDIC (UniMod:4)K (SEQ ID NO.6), AFDLYFVLDK (SEQ ID NO.7), VFLVGNVEIR (SEQ ID NO.8), RVSPVGETYIHEGLK (SEQ ID NO.9), ASEQIYYENR (SEQ ID NO.10), VLPGGDTYMHEGFER (SEQ IDNO.11), AVDIPHMDIEALK (SEQ ID NO.12), AMGIMNSFVNDIFER (SEQ ID NO.13), MPEQEYEFPEPR (SEQ ID NO.14), SGVISDTELQQALSNGTWTPFNPVTVR (SEQ ID NO.15), M (UniMod:35)EDVNSNVNADQEVR (SEQ ID NO.16), VGHDYQWIGLNDK (SEQ ID NO.17), HAEC (UniMod:4)IYLGHFSDPMYK (SEQ ID NO.18) or NGIFWGTWPGVSEAHPGGYK (SEQ ID NO.19).
151. The method according to claim 150, wherein the biomarker comprises 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more or 19 or more of the peptides.
152. The method according to claim 146, wherein the biomarker comprises any one of the following RNAs: ENST00000483727.5, ENST00000531734.6, ENST00000437154.6, ENST00000531997.1, ENST00000424185.7, ENST00000652176.1, ENST00000392593.9, ENST00000532853.5, ENST00000429947.1, ENST00000580914.1, ENST00000368205.7, ENST00000531709.6, ENST00000524817.5, ENST00000651281.1, ENST00000499685.2, ENST00000311921.8, ENST00000472111.5, ENST00000585172.2, ENST00000287713.7 or ENST00000547687.
2.
153. The method according to claim 152, wherein the biomarker comprises 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more or 20 or more of the said RNAs.
154. The method according to claim 146, wherein the biomarker comprises any one of the following lipids: NEG_PC(18:2_20:5)+AcO, POS_DAG(18:1_20:0)+NH4, NEG_PE(O-16:0_22:6)-H, NEG_PC(18:2_20:3)+AcO, POS_CER(d18:1 / 18:0)+H, POS_CE(22:0)+NH4, NEG_PE(14:0_22:5)-H, NEG_PC(20:5_20:5)+AcO, POS_PE(P-18:0_18:3)+H, NEG_PE(O-16:0_20:3)-H, POS_CE(18:3)+NH4, NEG_PE(O-18:0_22:5)-H, NEG_PE(O-18:0_20:5)-H, POS_PE(P-20:0_20:3)+H, NEG_PE(O-16:0_20:2)-H, POS_CER(d18:1 / 24:0)+H, NEG_PA(20:1_20:3)-H, NEG_PA(20:0_20:5)-H, POS_CE(20:0)+NH4 or NEG_PC(16:1_20:3)+AcO.
155. The method according to claim 154, wherein the biomarker comprises 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more or 20 or more of the lipids.
156. The method according to claim 146, wherein the biomarker comprises any one of the following metabolites: NEG_AICAR, POS_cystine, NEG_CMP, NEG_gentiobiose, POS_creatine, POS_imidazoleacetic acid, POS_inosine, NEG_n-isovalerylglycine, NEG_glucose-6-phosphate, POS_metanephrine, NEG_N-acetylglutamate, NEG_5-thymidylate (dTMP), POS_UMP, NEG_fructose-6-phosphate, NEG_cystine, POS_pantothenol, POS_guanine, NEG_shikimic acid, POS_1-methylimidazoleacetate or POS_flavone 2. The method according to claim 156, wherein the biomarker comprises 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, or 20 or more of the metabolites.
Citation Information
Patent Citations
Improvements in neckties
GB140434A