Method for sample quality assessment

By measuring the levels of sonic hedgehog (SHH) proteins and other proteins, the quality of blood samples is assessed, and the changes caused by unintentional activation of the samples during collection and processing are solved, the reliability and stability of biomarker measurements are improved, and the accuracy of disease diagnosis and drug development is enhanced.

CN113167782BActive Publication Date: 2025-07-11SOMALOGIC OPERATING CO INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201980077641.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-10-30
Filing Date
2019-10-28
Publication Date
2025-07-11
Estimated Expiration
2039-12-14

AI Technical Summary

Technical Problem

In the prior art, blood samples are susceptible to unintentional activation during collection and processing, resulting in changes in protein composition, affecting the reliability and stability of biomarker measurements, and limiting the application of biomarker in disease diagnosis and drug development.

Method used

By measuring the levels of sonic hedgehog (SHH) proteins and other specific proteins, combined with mass spectrometry, aptamer-based assays and antibody-based assays, assessing the quality of samples, identifying and analyzing samples or negative samples, ensuring the reliability and stability of samples during collection and processing.

Benefits of technology

Improve the reliability and stability of biomarker measurements, reduce false positive results, and enhance sample quality assessment capabilities in disease diagnosis and drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113167782B_ABST
    Figure CN113167782B_ABST
Patent Text Reader

Abstract

The present invention relates to methods for obtaining a biological sample with improved quality. It encompasses the identification of markers or proteins in a biological sample that are altered due to variations in sample collection, handling, and processing. The method can also be used to correct variations in the measurement results of disease biomarkers. Additionally, if it is determined that the sample collection method is inconsistent with a predetermined protocol, the method can allow for the rejection of a sample or group of samples as needed. Other advantages useful to those skilled in the art are described herein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In the fields of medical diagnostics and drug development, comparisons are made between the compositions of blood and other biological samples from individuals in order to determine and understand those changes that may be associated with a particular condition or disease. For example, biomarkers can indicate the ability to respond to certain drugs, the presence of a disease such as cancer, or changes in a monitoring process such as response to treatment or organ function. Once determined to be reliable and robust, such biomarker measurements can be used clinically.

[0002] Key properties regarding ideal biomarker measurements required for biomarker discovery and further study of clinical utility include reliability and robustness. Background Art

[0003] Blood contains powerful cellular and humoral systems for reacting to injury or foreign and infectious agents. Minor stimuli can induce the innate immune system (complement system and cells such as macrophages) to release powerful signals and enzymes, resulting in platelet activation and triggering blood clotting. Since these signals are associated with in vivo processes, they are of interest because they can directly participate in defense and repair systems and act as markers for disease. However, such process signals also respond to the effects of blood sample preparation. Simply drawing blood from a blood vessel by needle or exposing the blood to air can cause unintentional activation of these mechanisms. For example, changing the time, centrifuge speed, or temperature of sample processing steps can alter the apparent composition of serum or plasma, thus masking physiological information by pre-analytical variability imparted to the sample during collection and processing. Due to the accompanying lack of robustness, the high susceptibility of these processes and protein-to-protein sample handling to subtle changes can jeopardize their use as biomarkers.

[0004] Current research efforts in multiplex biology show strong interest in pre - analytical sample variation (commonly referred to as "batch effects"). Currently, the extent to which sample quality can be determined is mainly limited to visually obvious alterations such as red indicating red blood cell lysis and turbidity indicating high lipid or other contaminants. This limits the confidence that clinicians can place in almost the most robust and solid protein measurements. Studies demonstrating some of the complex and non - linear effects of variation in serum and plasma preparations are described in Ostroff, R. et al. (2010) J. Proteomics 73:649 - 666. A specific technique for determining compliance with sample preparation protocols is proposed herein based on the non - linear (logarithmic) transformation of measurements of a specific set of proteins affected by variation in the sample preparation protocol. Metrics derived from these methods can be used to monitor compliance, reject samples and make corrections in the target analyte. These techniques can be used to evaluate the quality of human or animal blood samples used in biomarker research, clinical diagnostic applications, biobank sample quality monitoring and drug development. Similar methods can be developed to assess the integrity of many other sample types, including urine, cerebrospinal fluid, sputum or tissue. Summary of the Invention

[0005] As described herein, key properties regarding biomarker discovery and the ideal biomarker measurements required to achieve clinical impact include reliability and robustness. The reliability of a biomarker means that the biomarker signal is true in capturing the underlying biology of health or disease (i.e., is not a "false - positive" marker). The robustness of a biomarker indicates that the biomarker is differentially expressed in diseased individuals relative to non - diseased individuals. To increase the probability of discovering true disease biomarkers and reduce the chance of identifying false positives due to sample bias, methods for measuring sample quality and consistency are required.

[0006] Measurements of protein analytes in plasma samples can be significantly affected by the protocols used for collecting and processing the samples. Deviations from a specific sample collection and / or processing protocol can lead to changes in the protein levels within the sample or other systematic effects on the measurements, which result in signal changes for many analytes (including negative controls). Such deviations can occur independently of the type of assay used to measure the protein analyte.

[0007] To assess the quality of clinical sample sets, the effects of the most obvious deviations from the protocol have been characterized. The variability of protein composition over time has been evaluated between sample collection and centrifugation. In addition, the variability of protein composition as a function of time has been evaluated between the time of centrifugation and the time of sample decantation.

[0008] Signatures of sample mishandling have been identified, which can be used as a quantitative classifier for evaluating clinical sample collections. In addition, a metric has been generated for each analyte that captures the sensitivity of the measurement of that analyte to deviations from the collection protocol, particularly with respect to the delays between sample collection and rotation and between sample rotation and sample decantation.

[0009] It might be imagined that some technologies are relatively immune to sample handling effects, but this is not the case. Even if antibodies work well in the presence of blood plasma and serum matrices and mass spectrometry can measure peptides and even denatured proteins, if the cells in the sample lyse, or if platelets degranulate, or if the complement system is activated, then a sharp change in analyte concentration will occur after the sample has been obtained and any "high-fidelity" measurement technique will detect them. Thus, techniques similar to those described herein for determining the effects of sample handling variations can be used in multiplex assay formats and for biomarkers other than proteins. Such assay formats can be sensitive in different ways but can be affected by the same underlying causes with respect to sample preparation variations.

[0010] Variations in different steps of blood handling and processing can be shown to affect biological samples in a reproducible manner. The sensitivity of each biomarker protein measurement to parameters associated with multiple sample handling and processing steps has been quantified using proteomic arrays and signatures of variations during sample processing have been identified. Sample handling and processing variations have been quantified within the same multi-analyte measurement assay used for disease biomarker measurement and development to determine which handling / processing signatures have been affected and by approximately how much. The methods of the present invention have also enabled the imposition of limits on acceptable sample handling and processing quality metrics for biomarker discovery.

[0011] The following numbered paragraphs describe other aspects of the invention:

[0012] 1. A method comprising:

[0013] a) measuring the level of Sonic Hedgehog (SHH) protein in a sample from a human subject and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0014] b) identifying the sample as an assay sample or a negative sample based on the SHH level and the level of one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve or thirteen proteins;

[0015] wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level assay, diagnostic method or prognostic method, and the negative sample is a sample not used as an assay sample.

[0016] 2. The method according to claim 1, wherein the measurement is performed using mass spectrometry, aptamer-based assay and / or antibody-based assay.

[0017] 3. The method according to claim 1, wherein the sample is selected from blood, plasma, serum or urine.

[0018] 4. The method according to claim 1, wherein the method comprises measuring SHH and PGAM1, SHH and PTPN4, SHH and TNFSF14, SHH and FAM49B, SHH and RBP7, SHH and IHH, SHH and DDX39B, SHH and S100A12, SHH and PGAM2, SHH and C4A.C4B, SHH and IL21R, SHH and TMEM9 or SHH and ADAM9.

[0019] 5. The method according to claim 1, wherein the method comprises measuring SHH, PGAM1 and TNFSF14; SHH, PGAM1 and RBB7; SHH, PGAM1 and PTPN4; SHH, PGAM1 and DDX39B; SHH, PGAM1 and FAM49B; SHH, PGAM1 and IHH; SHH, PGAM1 and S100A12; SHH, PGAM1 and ADAM9; SHH, PTPN4 and RBP7; SHH, PTPN4 and TNFSF14; SHH, PTPN4 and IHH; SHH, RBP7 and FAM49B; SHH, RBP7 and IHH; SHH, FAM49B and TNFSF14; SHH, DDX39B and PTPN4; SHH, TNFSF14 and S100A12; SHH, IHH and RBP7; SHH, IHH and TNFSF14; SHH, RBP7 and TNFSF14; SHH, RBP7 and S100A12; SHH, RBP7 and DDX39B; SHH, TNFSF14 and DDX39B; SHH, S100A12 and DDX39B; SHH, FAM49B and S100A12; SHH, IHH and FAM49B; SHH, IHH and DDX39B; SHH, TNFSF14 and ADAM9; SHH FAM49B and DDX39B; SHH, IHH and ADAM9; SHH, PGAM1 and C4A.C4B; SHH, PGAM2 and RBP7; SHH, PGAM1 and IL21R; SHH, PGAM2 and PTPN4; SHH, PGAM2 and ADAM9; SHH, PGAM2 and C4A.C4B; SHH, PGAM2 and IL21R; SHH, IHH and PGAM2; SHH, PGAM1 and PGAM2; SHH, TMEM9 and PGAM2; or SHH, TMEM9 and PGAM1.

[0020] 6. The method according to claim 1, wherein the method comprises measuring SHH and PGAM1, and at least two of the following proteins selected from RBP7, TNFSF14, PTPN4, DDX39B, FAM49B, S100A12, IHH, PGAM2, C4A.C4B, IL21R, TMEM9 and ADAM9.

[0021] 7. The method according to claim 1, wherein the method comprises measuring SHH and IHH, and at least two of the following proteins selected from RBP7, TNFSF14, PTPN4, DDX39B, FAM49B, S100A12, PGAM1, PGAM2, C4A.C4B, IL21R, TMEM9 and ADAM9.

[0022] 8. The method according to any one of the preceding claims, wherein the protein level is used to predict the length of time between sample collection from the human subject and sample centrifugation and / or the length of time between sample centrifugation and sample decantation.

[0023] 9. The method according to claim 8, wherein the time between sample collection and sample centrifugation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours.

[0024] 10. A method comprising:

[0025] a) contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising Sonic Hedgehog (SHH) and at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of PGAM1, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, ADAM9, PGAM2, C4A.C4B, IL21R, and TMEM9; and

[0026] b) measuring the level of each protein in the protein set with the set of capture reagents.

[0027] 11. The method according to claim 10, wherein the set of capture reagents is selected from aptamers, antibodies, and combinations of aptamers and antibodies.

[0028] 12. The method according to claim 10, wherein the sample is selected from blood, plasma, serum, or urine.

[0029] 13. The method according to claim 10, wherein the method comprises measuring SHH and PGAM1, SHH and PTPN4, SHH and TNFSF14, SHH and FAM49B, SHH and RBP7, SHH and IHH, SHH and DDX39B, SHH and S100A12, SHH and PGAM2, SHH and C4A.C4B, SHH and IL21R, SHH and TMEM9, or SHH and ADAM9.

[0030] 14. The method according to claim 10, wherein the method comprises measuring SHH, PGAM1 and TNFSF14; SHH, PGAM1 and RBB7; SHH, PGAM1 and PTPN4; SHH, PGAM1 and DDX39B; SHH, PGAM1 and FAM49B; SHH, PGAM1 and IHH; SHH, PGAM1 and S100A12; SHH, PGAM1 and ADAM9; SHH, PTPN4 and RBP7; SHH, PTPN4 and TNFSF14; SHH, PTPN4 and IHH; SHH, RBP7 and FAM49B; SHH, RBP7 and IHH; SHH, FAM49B and TNFSF14; SHH, DDX39B and PTPN4; SHH, TNFSF14 and S100A12; SHH, IHH and RBP7; SHH, IHH and TNFSF14; SHH, RBP7 and TNFSF14; SHH, RBP7 and S100A12; SHH, RBP7 and DDX39B; SHH, TNFSF14 and DDX39B; SHH, S100A12 and DDX39B; SHH, FAM49B and S100A12; SHH, IHH and FAM49B; SHH, IHH and DDX39B; SHH, TNFSF14 and ADAM9; SHH FAM49B and DDX39B; SHH, IHH and ADAM9; SHH, PGAM1 and C4A.C4B; SHH, PGAM2 and RBP7; SHH, PGAM1 and IL21R; SHH, PGAM2 and PTPN4; SHH, PGAM2 and ADAM9; SHH, PGAM2 and C4A.C4B; SHH, PGAM2 and IL21R; SHH, IHH and PGAM2; SHH, PGAM1 and PGAM2; SHH, TMEM9 and PGAM2; or SHH, TMEM9 and PGAM1.

[0031] 15. The method according to claim 10, wherein the method comprises measuring SHH and PGAM1, and at least two of the following proteins selected from RBP7, TNFSF14, PTPN4, DDX39B, FAM49B, S100A12, IHH and ADAM9.

[0032] 16. The method according to claim 10, wherein the method comprises measuring SHH and IHH, and at least two of the following proteins selected from RBP7, TNFSF14, PTPN4, DDX39B, FAM49B, S100A12, PGAM1 and ADAM9.

[0033] 17. The method according to any one of the preceding claims, wherein the protein level is used to predict the length of time between sample collection from the human subject and sample centrifugation and / or the length of time between sample centrifugation and sample decantation.

[0034] 18. The method according to claim 17, wherein the time between sample collection and sample centrifugation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours.

[0035] 19. A method comprising:

[0036] a) measuring the levels of at least three, four, five, six, seven or eight proteins selected from the group consisting of IHH, PTPN4, TNFSF14, FAM49B, RBP7, DDX39B, S100A12 and ADAM9 in a sample from a subject; and

[0037] b) identifying the sample as an assay sample or a negative sample based on the levels of the three, four, five, six, seven or eight proteins;

[0038] wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level analysis, diagnostic method or prognostic method, and the negative sample is a sample not used as an assay sample.

[0039] 20. The method according to claim 19, wherein the measurement is performed using mass spectrometry, an aptamer-based assay and / or an antibody-based assay.

[0040] 21. The method according to claim 19, wherein the sample is selected from blood, plasma, serum or urine.

[0041] 22. The method according to claim 19, wherein the method comprises measuring IHH, RB7, and PTPN4; IHH, RB7, and TNFSF14; IHH, RB7, and FAM49B; IHH, RBP7, and DDX39B; IHH, RBP7, and S100A12; IHH, RB7, and ADAM9; IHH, TNFSF14, and PTPN4; IHH, TNFSF14, and FAM49B; IHH, TNFSF14, and DDX39B; IHH, TNFSF14, and S100A12; IHH, TNFSF14, and ADAM9; IHH, FAM49, and PTPN4; IHH, FAM49, and TNFSF14; IHH, FAM49, and DDX39B; IHH, FAM49, and S100A12; IHH, ADAM9, and PTPN4; or IHH, FAM49, and ADAM9.

[0042] 23. The method according to any one of the preceding claims, further comprising measuring SHH and / or PGAM1.

[0043] 24. The method according to any one of the preceding claims, wherein the protein levels are used to predict the length of time between sample collection from the human subject and sample centrifugation and / or the length of time between sample centrifugation and sample decantation.

[0044] 25. The method according to claim 19, wherein the time between sample collection and sample centrifugation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours.

[0045] 26. A method comprising:

[0046] a) contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set in the sample from the subject, the protein set comprising three, four, five, six, seven, or eight proteins selected from the group consisting of IHH, PTPN4, TNFSF14, FAM49B, RBP7, DDX39B, S100A12, and ADAM9; and

[0047] b) measuring the level of each protein in the protein set with the set of capture reagents.

[0048] 27. The method according to claim 26, wherein the set of capture reagents is selected from aptamers, antibodies, and combinations of aptamers and antibodies.

[0049] 28. The method according to claim 26, wherein the sample is selected from blood, plasma, serum, or urine.

[0050] 29. The method according to claim 26, wherein the method comprises measuring IHH, RB7, and PTPN4; IHH, RB7, and TNFSF14; IHH, RB7, and FAM49B; IHH, RBP7, and DDX39B; IHH, RBP7, and S100A12; IHH, RB7, and ADAM9; IHH, TNFSF14, and PTPN4; IHH, TNFSF14, and FAM49B; IHH, TNFSF14, and DDX39B; IHH, TNFSF14, and S100A12; IHH, TNFSF14, and ADAM9; IHH, FAM49, and PTPN4; IHH, FAM49, and TNFSF14; IHH, FAM49, and DDX39B; IHH, FAM49, and S100A12; or IHH, FAM49, and ADAM9.

[0051] 30. The method according to any one of the preceding claims, further comprising measuring SHH and / or PGAM1.

[0052] 31. The method according to any one of the preceding claims, wherein the protein levels are used to predict the length of time between sample collection from the human subject and sample centrifugation and / or the length of time between sample centrifugation and sample decantation.

[0053] 32. The method according to claim 26, wherein the time between sample collection and sample centrifugation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours.

[0054] 33. A method, comprising:

[0055] a) measuring the levels of RB7, FAM49B, TNFSF14, ADAM9, PGAM1, and at least one protein selected from DDX39B and S100A12 in a sample from a human subject; and

[0056] b) identifying the sample as an assay sample or a negative sample based on the levels of RB7, FAM49B, TNFSF14, ADAM9, PGAM1, and the at least one protein;

[0057] wherein the assay sample is a sample for one or more of: protein biomarker discovery assays, protein expression level assays, diagnostic methods, or prognostic methods, and the negative sample is a sample not used as an assay sample.

[0058] 34. The method according to claim 33, wherein the measurement is performed using mass spectrometry, aptamer-based assays, and / or antibody-based assays.

[0059] 35. The method according to claim 33, wherein the sample is selected from blood, plasma, serum, or urine.

[0060] 36. The method according to claim 33, wherein the method comprises measuring RB7, FAM49B, TNFSF14, ADAM9, PGAM1, and S100A12; or RB7, FAM49B, TNFSF14, ADAM9, PGAM1, and DDX39B.

[0061] 37. The method according to any one of the preceding claims, wherein the protein levels are used to predict the length of time between sample collection and sample centrifugation from the human subject and / or the length of time between sample centrifugation and sample decantation.

[0062] 38. The method according to claim 33, wherein the time between sample collection and sample centrifugation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours; or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours; or greater than 24 hours.

[0063] 39. A method comprising:

[0064] a) contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising RB7, FAM49B, TNFSF14, ADAM9, PGAM1, and at least one protein selected from DDX39B and S100A12; and

[0065] b) measuring the level of each protein in the protein set with the set of capture reagents.

[0066] 40. The method according to claim 39, wherein the set of capture reagents is selected from aptamers, antibodies, and combinations of aptamers and antibodies.

[0067] 41. The method according to claim 39, wherein the sample is selected from blood, plasma, serum, or urine.

[0068] 42. The method according to claim 39, wherein the method comprises measuring RB7, FAM49B, TNFSF14, ADAM9, PGAM1, and S100A12; or RB7, FAM49B, TNFSF14, ADAM9, PGAM1, and DDX39B.

[0069] 43. The method according to any one of the preceding claims, wherein the protein level is used to predict the length of time between sample collection from the human subject and sample centrifugation and / or the length of time between sample centrifugation and sample decantation.

[0070] 44. The method according to claim 39, wherein the time between sample collection and sample centrifugation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours.

[0071] 45. The method according to any one of the preceding claims, wherein the one or more protein levels are used to assign a quality score to the sample, and the quality score is then used to determine whether the sample is an analytical sample or a non-analytical sample.

[0072] 46. The method according to any one of the preceding claims, wherein the one or more protein levels are used to assign a quality score to the sample, and the quality score is then used to determine whether the sample is to be used for further analysis of additional proteins in the sample.

[0073] 47. A method comprising:

[0074] a) measuring the level of PGAM1 protein in a sample from a human subject and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0075] b) Identify the sample as an assay sample or a negative sample based on the SHH level and the level of one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen of said proteins;

[0076] wherein the assay sample is a sample for one or more of the following: protein biomarker discovery assay, protein expression level assay, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0077] 48. A method, comprising:

[0078] a) Measuring the level of PGAM2 protein in a sample from a human subject and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0079] b) Identify the sample as an assay sample or a negative sample based on the SHH level and the level of one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen of said proteins;

[0080] wherein the assay sample is a sample for one or more of the following: protein biomarker discovery assay, protein expression level assay, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0081] 49. A method, comprising:

[0082] a) Measuring the level of C4A.C4B protein in a sample from a human subject and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, PGAM1, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0083] b) Identify the sample as an assay sample or a negative sample based on the SHH level and the level of one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen of said proteins;

[0084] Wherein, the analysis sample is a sample for one or more of the following: protein biomarker discovery analysis, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an analysis sample.

[0085] 50. A method comprising:

[0086] a) measuring the level of PTPN4 protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PGAM1, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0087] b) identifying the sample as an analysis sample or a negative sample based on the SHH level and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins;

[0088] Wherein, the analysis sample is a sample for one or more of the following: protein biomarker discovery analysis, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an analysis sample.

[0089] 51. A method comprising:

[0090] a) measuring the level of TNFSF14 protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, PGAM1, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0091] b) identifying the sample as an analysis sample or a negative sample based on the SHH level and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins;

[0092] Wherein, the analysis sample is a sample for one or more of the following: protein biomarker discovery analysis, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an analysis sample.

[0093] 52. A method comprising:

[0094] a) Measuring the level of FAM49B protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, TNFSF14, PGAM1, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0095] b) Identifying the sample as an assay sample or a negative sample based on the level of SHH and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins;

[0096] wherein the assay sample is a sample for one or more of protein biomarker discovery analysis, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0097] 53. A method comprising:

[0098] a) Measuring the level of RBP7 protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, TNFSF14, PGAM1, FAM49B, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0099] b) Identifying the sample as an assay sample or a negative sample based on the level of SHH and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins;

[0100] wherein the assay sample is a sample for one or more of protein biomarker discovery analysis, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0101] 54. A method comprising:

[0102] a) Measuring the level of IHH protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, TNFSF14, PGAM1, FAM49B, RBP7, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0103] b) Identifying the sample as an assay sample or a negative sample based on the level of SHH and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins;

[0104] wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0105] 55. A method comprising:

[0106] a) Measuring the level of DDX39B protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, TNFSF14, PGAM1, FAM49B, RBP7, IHH, S100A12, IL21R, TMEM9, and ADAM9; and

[0107] b) Identifying the sample as an assay sample or a negative sample based on the level of SHH and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins;

[0108] wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0109] 56. A method comprising:

[0110] a) Measuring the level of S100A12 protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, TNFSF14, PGAM1, FAM49B, RBP7, IHH, DDX39B, IL21R, TMEM9, and ADAM9; and

[0111] b) Identifying the sample as an assay sample or a negative sample based on the level of SHH and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins;

[0112] wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0113] 57. A method comprising:

[0114] a) Measuring the level of IL21R protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, TNFSF14, PGAM1, FAM49B, RBP7, IHH, S100A12, DDX39B, TMEM9, and ADAM9; and

[0115] b) Identifying the sample as an assay sample or a negative sample based on the level of SHH and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins;

[0116] wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0117] 58. A method comprising:

[0118] a) Measuring the level of TMEM9 protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, TNFSF14, PGAM1, FAM49B, RBP7, IHH, S100A12, IL21R, DDX39B, and ADAM9; and

[0119] b) Identifying the sample as an assay sample or a negative sample based on the level of SHH and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins;

[0120] wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0121] 59. A method comprising:

[0122] a) Measuring the level of ADAM9 protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, TNFSF14, PGAM1, FAM49B, RBP7, IHH, S100A12, IL21R, TMEM9, and DDX39B; and

[0123] b) Identifying the sample as an assay sample or a negative sample based on the level of SHH and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins;

[0124] wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0125] 60. A method comprising:

[0126] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising Sonic Hedgehog (SHH) protein and the levels of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0127] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0128] 61. A method comprising:

[0129] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising PGAM1 protein and the levels of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0130] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0131] 62. A method comprising:

[0132] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising PGAM2 protein and the levels of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0133] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0134] 63. A method comprising:

[0135] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising levels of C4A.C4B protein and at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0136] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0137] 64. A method comprising:

[0138] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising levels of PTPN4 protein and at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, C4A.C4B, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0139] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0140] 65. A method comprising:

[0141] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising levels of TNFSF14 protein and at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, C4A.C4B, PTPN4, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0142] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0143] 66. A method comprising:

[0144] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising the level of FAM49B protein and at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0145] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0146] 67. A method comprising:

[0147] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising the level of RBP7 protein and at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0148] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0149] 68. A method comprising:

[0150] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising the level of IHH protein and at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0151] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0152] 69. A method comprising:

[0153] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising the DDX39B protein and the levels of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, S100A12, IL21R, TMEM9, and ADAM9; and

[0154] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0155] 70. A method comprising:

[0156] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising the S100A12 protein and the levels of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, IL21R, TMEM9, and ADAM9; and

[0157] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0158] 71. A method comprising:

[0159] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising the IL21R protein and the levels of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, TMEM9, and ADAM9; and

[0160] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0161] 72. A method comprising:

[0162] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising the TMEM9 protein and at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, and ADAM9; and

[0163] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0164] 73. A method comprising:

[0165] a) Contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set, the protein set comprising the ADAM9 protein and at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of SHH, PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, and TMEM9; and

[0166] b) Measuring the level of each protein in the protein set with the set of capture reagents.

[0167] 74. A method comprising:

[0168] a) Measuring the level of sonic hedgehog (SHH) protein in a sample from a human subject, and the level of at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins selected from the group consisting of PGAM1, PGAM2, C4A.C4B, PTPN4, TNFSF14, FAM49B, RBP7, IHH, DDX39B, S100A12, IL21R, TMEM9, and ADAM9; and

[0169] b) Identifying the sample as an assay sample or a negative sample based on the SHH level and the level of the one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen proteins.

[0170] Wherein, the analysis sample is a sample for one or more of the following: protein biomarker discovery analysis, protein expression level analysis, diagnostic method or prognostic method, and the negative sample is a sample not used as an analysis sample.

[0171] 75. A method, comprising:

[0172] a) Measuring the levels of HNRNDPDL, PTPN4, PGAM2, C4A.C4B, EIF4A1, IHH, SHH, PGAM1, S100A9 and HLA.DRB3 in a sample from a human subject; and

[0173] b) Identifying the sample as an analysis sample or a negative sample based on the levels of HNRNDPDL, PTPN4, PGAM2, C4A.C4B, EIF4A1, IHH, SHH, PGAM1, S100A9 and HLA.DRB3;

[0174] Wherein, the analysis sample is a sample for one or more of the following: protein biomarker discovery analysis, protein expression level analysis, diagnostic method or prognostic method, and the negative sample is a sample not used as an analysis sample.

[0175] 76. The method according to any one of the preceding claims, wherein the measurement is performed using mass spectrometry, aptamer-based assays and / or antibody-based assays.

[0176] 77. The method according to any one of the preceding claims, wherein the sample is selected from blood, plasma, serum or urine.

[0177] 78. The method according to any one of the preceding claims, wherein the protein levels are used to predict the length of time between sample collection from the human subject and sample centrifugation and / or the length of time between sample centrifugation and sample decantation.

[0178] 79. The method according to any one of the preceding claims, wherein the time between sample collection and sample centrifugation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours.

[0179] 80. A method, comprising:

[0180] a) contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set comprising HNRNDPDL, PTPN4, PGAM2, C4A.C4B, EIF4A1, IHH, SHH, PGAM1, S100A9, and HLA.DRB3; and

[0181] b) measuring the level of each protein in the protein set with the set of capture reagents.

[0182] 81. The method according to any one of the preceding claims, wherein the measurement is performed using mass spectrometry, an aptamer-based assay, and / or an antibody-based assay.

[0183] 82. The method according to any one of the preceding claims, wherein the sample is selected from blood, plasma, serum, or urine.

[0184] 83. The method according to any one of the preceding claims, wherein the protein levels are used to predict the length of time between sample collection from the human subject and sample centrifugation and / or the length of time between sample centrifugation and sample decantation.

[0185] 84. The method according to any one of the preceding claims, wherein the time between sample collection and sample centrifugation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours.

[0186] 85. A method comprising:

[0187] a) contacting a sample from a human subject with two capture reagents, wherein one capture reagent has an affinity for the TMEM9 protein and a second capture reagent has an affinity for the PGAM1 protein; and

[0188] b) measuring the level of each protein with the two capture reagents.

[0189] 86. A method comprising:

[0190] a) measuring the levels of the PGAM1 and TMEM9 proteins in a sample from a human subject; and

[0191] b) identifying the sample as an assay sample or a negative sample based on the levels of the PGAM1 and TMEM9.

[0192] Wherein, the analysis sample is a sample for one or more of the following: protein biomarker discovery analysis, protein expression level analysis, diagnostic method, or prognostic method, and the negative sample is a sample not used as an analysis sample.

[0193] 87. The method according to claim 85 or 86, further comprising measuring the level of the IHH protein with a capture reagent having an affinity for the IHH protein.

[0194] 88. The method according to claim 85 or 86, further comprising measuring the level of the C4A.C4B protein with a capture reagent having an affinity for the C4A.C4B protein.

[0195] 89. The method according to claim 85 or 86, further comprising measuring the level of the SHH protein with a capture reagent having an affinity for the SHH protein.

[0196] 90. The method according to claim 85 or 86, further comprising measuring the level of the PGAM2 protein with a capture reagent having an affinity for the PGAM2 protein.

[0197] 91. The method according to claim 85 or 86, further comprising measuring the level of the ADAM9 protein with a capture reagent having an affinity for the ADAM9 protein.

[0198] 92. The method according to claim 85 or 86, further comprising measuring the level of the PTPN4 protein with a capture reagent having an affinity for the PTPN4 protein.

[0199] 93. The method according to claim 84 or 85, further comprising measuring the level of the IL21R protein with a capture reagent having an affinity for the IL21R protein.

[0200] 94. The method according to claim 85 or 86, further comprising measuring the level of the RBP7 protein with a capture reagent having an affinity for the RBP7 protein.

[0201] 95. A method comprising:

[0202] a) contacting a sample from a human subject with two capture reagents, wherein one capture reagent has an affinity for the SHH protein and a second capture reagent has an affinity for the IHH protein; and

[0203] b) measuring the level of each protein with the two capture reagents.

[0204] 96. A method comprising:

[0205] a) Measuring the levels of SHH and IHH in a sample from a human subject; and

[0206] b) Identifying the sample as an assay sample or a negative sample based on the levels of said SHH and IHH;

[0207] wherein the assay sample is a sample for one or more of: protein biomarker discovery assays, protein expression level assays, diagnostic methods, or prognostic methods, and the negative sample is a sample not used as an assay sample.

[0208] 97. The method according to claim 95 or 96, further comprising measuring the level of the RBP7 protein with a capture reagent having an affinity for the RBP7 protein.

[0209] 98. The method according to claim 95 or 96, further comprising measuring the level of the FAM94B protein with a capture reagent having an affinity for the FAM94B protein.

[0210] 99. The method according to claim 95 or 96, further comprising measuring the level of the TNFSF14 protein with a capture reagent having an affinity for the TNFSF14 protein.

[0211] 100. The method according to claim 95 or 96, further comprising measuring the level of the ADAM9 protein with a capture reagent having an affinity for the ADAM9 protein.

[0212] 101. The method according to claim 95 or 96, further comprising measuring the level of the S100A12 protein with a capture reagent having an affinity for the S100A12 protein.

[0213] 102. The method according to claim 95 or 96, further comprising measuring the level of the DDX39B protein with a capture reagent having an affinity for the DDX39B protein.

[0214] 103. The method according to claim 95 or 96, further comprising measuring the level of the PGAM1 protein with a capture reagent having an affinity for the PGAM1 protein.

[0215] 104. The method according to claim 95 or 96, further comprising measuring the level of the PTPN4 protein with a capture reagent having an affinity for the PTPN4 protein.

[0216] 105. A method comprising:

[0217] a) Contacting a sample from a human subject with four capture reagents, wherein each of the four capture reagents has an affinity for a protein selected from IHH, RBP7, ADAM9, and PTPN4; and

[0218] b) Measuring the level of each protein with the four capture reagents.

[0219] 106. The method of claim 105, further comprising measuring the level of the SHH protein with a capture reagent having an affinity for the SHH protein.

[0220] 107. The method of claim 105, further comprising measuring the level of the PGAM1 protein with a capture reagent having an affinity for the PGAM1 protein.

[0221] 108. The method of claim 105, further comprising measuring the level of one or more proteins selected from TMEM9, C4A.C4B, PGAM2, FAM49B, TNFSF14, S100A12, DDX39B, and IL21R with a capture reagent having an affinity for one of the one or more proteins.

[0222] 109. A method comprising:

[0223] a) Measuring the levels of IHH, RBP7, ADAM9, and PTPN4 proteins in a sample from a human subject; and

[0224] b) Identifying the sample as an assay sample or a negative sample based on the levels of the IHH, RBP7, ADAM9, and PTPN4 proteins from the sample;

[0225] wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level assay, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

[0226] 110. The method of 109, further comprising measuring the level of the SHH protein and identifying the sample as an assay sample or a negative sample based on the level of the SHH protein from the sample.

[0227] 111. The method of 109, further comprising measuring the level of the PGAM1 protein and identifying the sample as an assay sample or a negative sample based on the level of the PGAM1 protein from the sample.

[0228] 112. The method according to claim 109, further comprising measuring the level of one or more proteins selected from TMEM9, C4A.C4B, PGAM2, FAM49B, TNFSF14, S100A12, DDX39B, and IL21R and identifying the sample as an assay sample or a negative sample based on the level of the one or more proteins.

[0229] 113. A method comprising:

[0230] a) contacting a sample from a human subject with four capture reagents, wherein each of the four capture reagents has an affinity for a protein selected from IHH, RBP7, ADAM9, and PTPN4; and

[0231] b) measuring the level of each protein in the sample with the four capture reagents.

[0232] 114. The method according to claim 113, further comprising measuring the level of the SHH protein with a capture reagent having an affinity for the SHH protein.

[0233] 115. The method according to claim 113, further comprising measuring the level of the PGAM1 protein with a capture reagent having an affinity for the PGAM1 protein.

[0234] 116. The method according to claim 113, further comprising measuring the level of the TMEM9 protein with a capture reagent having an affinity for the TMEM9 protein.

[0235] 117. The method according to claim 113, further comprising measuring the level of one or more proteins selected from C4A.C4B, PGAM2, FAM49B, TNFSF14, S100A12, DDX39B, and IL21R with capture reagents, each capture reagent having an affinity for one of the one or more proteins.

[0236] 118. The method according to claim 113, wherein the sample is selected from blood, plasma, serum, or urine.

[0237] 119. The method according to claim 113, wherein the protein levels are used to predict the length of time between sample collection and sample centrifugation from the human subject and / or the length of time between sample centrifugation and sample decantation.

[0238] 120. The method according to claim 113, wherein the protein level is used to identify the sample as an analytical sample or a negative sample based on the protein level; wherein, the analytical sample is a sample for one or more of the following: protein biomarker discovery analysis, protein expression level analysis, diagnostic method or prognostic method, and the negative sample is a sample not used as an analytical sample.

[0239] 121. The method according to claim 119, wherein the time between sample collection and sample centrifugation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours.

[0240] 122. The method according to claim 113, wherein the capture reagent is selected from aptamers or antibodies.

[0241] 123. A method comprising:

[0242] a) measuring the levels of IHH, RBP7, ADAM9, and PTPN4 proteins in a sample from a human subject; and

[0243] b) identifying the sample as an analytical sample or a negative sample based on the levels of the IHH, RBP7, ADAM9, and PTPN4 proteins from the sample;

[0244] wherein, the analytical sample is a sample for one or more of the following: protein biomarker discovery analysis, protein expression level analysis, diagnostic method or prognostic method, and the negative sample is a sample not used as an analytical sample.

[0245] 124. The method according to 123, further comprising measuring the level of SHH protein and identifying the sample as an analytical sample or a negative sample based on the level of the SHH protein from the sample.

[0246] 125. The method according to 123, further comprising measuring the level of PGAM1 protein and identifying the sample as an analytical sample or a negative sample based on the level of the PGAM1 protein from the sample.

[0247] 126. The method according to 123, further comprising measuring the level of TMEM9 protein and identifying the sample as an analytical sample or a negative sample based on the level of the TMEM9 protein from the sample.

[0248] 127. The method according to claim 123, further comprising measuring the level of one or more proteins selected from C4A, C4B, PGAM2, FAM49B, TNFSF14, S100A12, DDX39B, and IL21R and identifying the sample as an assay sample or a negative sample based on the level of the one or more proteins.

[0249] 128. The method according to claim 123, wherein the sample is selected from blood, plasma, serum, or urine.

[0250] 129. The method according to claim 123, wherein the protein level is used to predict the length of time between sample collection from the human subject and sample centrifugation and / or the length of time between sample centrifugation and sample decantation.

[0251] 130. The method according to claim 129, wherein the time between sample collection and sample centrifugation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is about 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours.

[0252] 131. The method according to claim 123, wherein the measurement of the protein level is performed using mass spectrometry, aptamer-based assays, and / or antibody-based assays.

[0253] 132. The method according to claim 123, wherein the protein level is used for a classifier selected from: decision tree; bagging + boosting + forest; rule-based inferential learning; Parzen Windows; linear model; logic; neural network method; unsupervised clustering; K-means; hierarchical ascent / descent; semi-supervised learning; prototype method; nearest neighbor; kernel density estimation; support vector machine; hidden Markov model; Boltzmann learning; random forest model, which is used together with the protein level to identify the sample as an assay sample or a negative sample. BRIEF DESCRIPTION OF THE DRAWINGS

[0254] Figure 1 Shows the relationship between PGAM1 RFU and spin time, which shows a steady change in signal with spin time.

[0255] Figure 2 Shows the relationship between analyte RFU and spin time, which shows a very low discriminatory property.

[0256] Figure 3Shows the analyte importance in the rotation time model. After about 10 analytes, the relative importance of additional analytes decreases to a steady state.

[0257] Figure 4 Shows the sample decision tree in the rotation time model. The first node splits the sample on TNFSF14 RFU, and if the RFU is greater than 756.2, it terminates the prediction at 24 hours, otherwise it goes down an additional branch.

[0258] Figure 5 Shows the error in predicting the relative number of trees in the random forest.

[0259] Figure 6 Shows the prediction error in the single analyte random forest model. The horizontal and vertical bars indicate the class thresholds, and the black solid line indicates the true prediction line.

[0260] Figure 7 Shows the model stability in the random forest and Naive Bayes. Although the Naive Bayes model shows continuous changes when shifting signals on single analytes, the random forest shows greater stability.

[0261] Figure 8 Shows the model stability when scaling individual analytes. When scaling each analyte by the effect size, the true rotation time given in each subplot title is compared with the predicted time. The single line represents the prediction of the random forest when scaling an analyte and keeping the remaining nine constant.

[0262] Figure 9 Shows the cumulative analyte distribution function of 18 individuals at different rotation times for the protein marker SHH.

[0263] Figure 10 Shows the cumulative analyte distribution function of 18 individuals at different rotation times for the protein marker IHH.

[0264] Figure 11 Shows the cumulative analyte distribution function of 18 individuals at different rotation times for the protein marker RBP7.

[0265] Figure 12 Shows the cumulative analyte distribution function of 18 individuals at different rotation times for the protein marker FAM49B.

[0266] Figure 13 Shows the cumulative analyte distribution function of 18 individuals at different rotation times for the protein marker TNFSF14.

[0267] Figure 14Shows the cumulative analyte distribution function for 18 individuals at different rotation times for the protein marker ADAM9.

[0268] Figure 15 Shows the cumulative analyte distribution function for 18 individuals at different rotation times for the protein marker S100A12.

[0269] Figure 16 Shows the cumulative analyte distribution function for 18 individuals at different rotation times for the protein marker DDX39B.

[0270] Figure 17 Shows the cumulative analyte distribution function for 18 individuals at different rotation times for the protein marker PGAM1.

[0271] Figure 18 Shows the cumulative analyte distribution function for 18 individuals at different rotation times for the protein marker PTPN4.

[0272] Figure 19 Shows the performance of the analyte models, where the performance of each model is quantified using the RMSE of the predicted rotation time relative to the true rotation time for each individual and time point.

[0273] Figure 20 Shows low-performance analytes based on the fold fraction of the analytes used in each group of model performance to illustrate the importance of each analyte to model performance.

[0274] Figure 21 Shows medium-performance analytes based on the fold fraction of the analytes used in each group of model performance to illustrate the importance of each analyte to model performance.

[0275] Figure 22 Shows high-performance analytes based on the fold fraction of the analytes used in each group of model performance to illustrate the importance of each analyte to model performance.

[0276] Figure 23 Shows the distribution of the number of models used with a specified number of analytes. Detailed Description

[0277] Reference will now be made in detail to representative embodiments of the present invention. While the invention will be described in conjunction with the exemplified embodiments, it should be understood that the invention is not intended to be limited to these embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents that may be included within the scope of the invention as defined by the claims.

[0278] Those skilled in the art will recognize many methods and materials similar or equivalent to those described herein, which can be used in the practice of the present invention and are within the scope of the practice of the present invention. The present invention is in no way limited to the described methods and materials.

[0279] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods, devices, and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods, devices, and materials are now described.

[0280] All publications, published patent documents, and patent applications cited in this application indicate the state of the art in one or more fields to which this application pertains. All publications, published patent documents, and patent applications cited herein are hereby incorporated by reference to the same extent as if each individual publication, published patent document, or patent application was specifically and individually indicated to be incorporated by reference.

[0281] As used in this application, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" include plural references and may be used interchangeably with "at least one" and "one or more". Thus, reference to "aptamer" includes aptamer mixtures, reference to "probe" includes probe mixtures, and so on.

[0282] As used herein, the term "about" represents an inconsequential numerical modification or variation such that the basic function of the value associated with the item is not changed.

[0283] As used herein, the terms "comprising", "including", "containing", and any variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product-by-process, or composition of matter that comprises, includes, or contains an element or list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, product-by-process, or composition of matter.

[0284] As used herein, "biomarker" is used to refer to a target molecule that indicates a normal or abnormal process in an individual, or a disease or other condition in an individual, or is a sign of a normal or abnormal process in an individual, or a disease or other condition in an individual. More specifically, a "biomarker" is an anatomical, physiological, biochemical, or molecular parameter that is associated with the presence of a specific physiological state or process, whether normal or abnormal, and if abnormal, whether chronic or acute. Biomarkers are detectable and measurable by a variety of methods including laboratory assays and medical imaging. When the biomarker is a protein, it is also possible to use the expression of the corresponding gene as an alternative measure of the corresponding protein biomarker in a biological sample, or the methylation state of the gene encoding the biomarker, or the amount or presence or absence of a protein that controls the expression of the biomarker.

[0285] The selection of biomarkers for a particular disease state involves first identifying such markers that have a measurable and statistically significant difference in a disease population compared to a control population for a particular medical application. Biomarkers can include secreted or shed molecules that parallel disease development or progression and are readily diffusible into the bloodstream from tissues affected by the disease or condition, or in response to the disease or condition, from surrounding tissues and circulating cells. The identified biomarker or set of biomarkers is typically clinically validated or shown to be a reliable indicator for its original intended use for which it was selected. Biomarkers can comprise a variety of molecules, including small molecules, peptides, proteins, and nucleic acids. Some of the key issues affecting biomarker identification include overfitting of the available data and bias in the data, including variations in sample handling protocols.

[0286] As used herein, "biomarker value", "value", "biomarker level", and "level" are used interchangeably to refer to a measurement that is made using any analytical method for detecting a biomarker in a biological sample and that indicates the presence, absence, absolute amount or concentration, relative amount or concentration, titer, level, expression level, ratio of measurement levels, etc. of the biomarker in the biological sample, with respect to the biomarker in the biological sample, or corresponding to the biomarker in the biological sample. The exact nature of the "value" or "level" depends on the specific design and components of the particular analytical method used to detect the biomarker.

[0287] "Disease biomarker control range" or "biomarker control range" are used interchangeably and mean the normal or non-disease range of a biomarker in non-diseased or normal individuals. They are typically derived from a control population.

[0288] "Sample", "case", or "test set" are used interchangeably and mean an individual or case patient who is suspected of having or may have a disease and who can ultimately be determined to have the disease or not have the disease.

[0289] As used herein, “sample handling and processing marker,” “handling / processing marker,” “marker sensitive to changes in a sample handling and processing protocol,” “marker sensitive to pre-analytical variability,” etc. are used interchangeably to refer to a marker that has been found, by the methods described herein, to be sensitive to changes in a sample handling and processing protocol. A “sample handling and processing marker” may or may not include a biomarker.

[0290] Sample handling and processing markers can be identified by candidate markers in a control population of normal individuals. Samples obtained from the control population are analyzed for candidate markers to select candidate markers that are sensitive to changes in a sample handling and processing protocol. Changes include, but are not limited to, changes in sample processing time, processing temperature, storage time, storage temperature, storage vessel composition, and other storage conditions prior to sample assay; changes in the method for extracting a sample from a normal individual, including but not limited to exposure of the sample to oxygen, the pore size of the needle used for venipuncture, collection devices, collection tube additives; changes in sample processing, which include but are not limited to centrifugation speed, temperature, and time, filtration, and filter pore size; collection containers or vessels, freezing methods; etc. These candidate markers identified as being substantially sensitive to changes are eligible as sample handling and processing markers. Candidate markers include a variety of molecules including small molecules, peptides, proteins, and nucleic acids.

[0291] In some cases, it may be desirable to distinguish the selected handling / processing markers in an assay to remove those that may also be disease markers or markers for the specific disease under discussion. On the other hand, if the number of handling / processing markers to be used is large, e.g., greater than any number in about 20, 30, 50, or more, then in such cases it may not be necessary to eliminate the handling / processing markers.

[0292] As used herein, “assay,” “detection,” etc., which are used interchangeably herein, refer to the molecular detection or quantification (measurement) using any suitable method, including fluorescence, chemiluminescence, radiolabeling, surface plasmon resonance, surface acoustic wave, mass spectrometry, infrared spectroscopy, Raman spectroscopy, atomic force microscopy, scanning tunneling microscopy, electrochemical detection methods, nuclear magnetic resonance, quantum dots, etc. “Detection” and its variations refer to the identification or observation of the presence of a molecule in a biological sample, and / or the measurement of a molecular value.

[0293] As used herein, the terms "biological sample", "sample", and "test sample" are used interchangeably herein to refer to any material, biological fluid, tissue, or cell obtained from or otherwise derived from an individual. This includes blood (including whole blood, white blood cells, peripheral blood mononuclear cells, buffy coat, plasma, serum, and dried blood spots collected on filter paper), sputum, tears, mucus, nasal washes, nasal aspirates, breath, urine, semen, saliva, cyst fluid, cerebrospinal fluid, amniotic fluid, glandular fluid, lymphatic fluid, nipple aspirates, bronchial aspirates, pleural fluid, peritoneal fluid, synovial fluid, joint aspirates, organ secretions, ascites, cells, cell extracts, and cerebrospinal fluid. This also includes all of the foregoing experimentally isolated fractions. For example, a blood sample can be fractionated into serum or fractions containing specific types of blood cells such as red blood cells or white blood cells (leukocytes). When needed, a sample can be a combination of samples from an individual, such as a combination of tissue and fluid samples. The term "biological sample" also includes, for example, materials containing homogenized solid materials, such as from fecal samples, tissue samples, or tissue biopsies. The term "biological sample" also includes materials derived from tissue cultures or cell cultures. Any suitable method for obtaining a biological sample can be employed; exemplary methods include, for example, venipuncture, swabbing (e.g., buccal swabbing), lavage, aspiration, and fine needle aspiration biopsy procedures. Samples can also be collected, for example, by microdissection (e.g., laser capture microdissection (LCM) or laser microdissection (LMD)), bladder washing, smear (e.g., PAP smear), or catheter lavage. A "biological sample" obtained from or derived from an individual includes any such sample that has been processed in any suitable manner after being obtained from the individual.

[0294] In addition, it should be recognized that a biological sample can be obtained by taking biological samples from a number of individuals and pooling them or pooling aliquots of the biological samples of each individual.

[0295] "Cell abuse" includes, but is not limited to, cell contamination, cell lysis, cell fragmentation, cell debris, internal cell components, and the like.

[0296] As used herein, a "rejected sample" can refer to a subset, group, or collection to which the sample belongs.

[0297] As used herein, "SOMAmer" or "slow off-rate modified aptamer" refers to an aptamer having improved off-rate characteristics. SOMAmers can be generated using the improved SELEX method described in U.S. Publication No. 2009 / 0004667, now U.S. Patent No. 7,947,447, entitled "Method for Generating Aptamers with Improved Off-Rates."

[0298] In the present application, the measurement of tagged proteins regarding sample handling and processing has been measured and found to have defined and reproducible behavior regarding variations in sample collection and preparation.

[0299] The central idea here is to use some of the many processing and handling tagged proteins that can be measured in each sample to provide a graded response to variations in sample collection and sample preparation steps. In this sense, these processing / handling tagged protein signals can be used, for example, to monitor past events in blood sample processing, such as delays before centrifugation and delays before decantation. This is different from directly monitoring the degradation of the target biomarker protein of interest and can be more sensitive and informative over a wide range. By using the methods described herein, it may be possible to characterize the sample quality regarding changes after extraction of a specific target biomarker protein by applying the known sensitivities of the processing / handling tags to the estimates of the biomarker. By subtracting the sample handling component from the apparent protein concentration, the monitoring of sample processing and handling tags can also be used to correct for the estimated effects of each variation in disease biomarkers. These sample handling and processing biomarker measurements can be used to characterize the sample before evaluating disease biomarkers by a variety of measurement systems including antibody assays, mass spectrometry, etc.

[0300] In this way, some of the biological mechanisms in blood are used to act as clocks, timers, and recording devices. For this technique to work, we must be able to distinguish the in vivo bioactivation of the various mechanisms, and the activation that occurs after the blood has left the body, or "ex vivo" alterations. The main tool for distinguishing in vivo disease biomarkers and processing / handling tag degradation from those suffered ex vivo is the ability to measure a very large number of proteins simultaneously, such that the sample can be characterized not only for a single sample handling / processing variation, but also for several sample handling / processing variations. The associated protein measurements indicating specific sample handling protocol variations provide the experimental cohort of sample handling / processing tags.

[0301] Evaluating a small number of samples to discover sample handling and processing techniques at one or more sites or in some fractions of the samples will make it difficult to measure differences in the target biomarker proteins. The metrics delivered on each sample by our system enable rejection of sample sets from clinical sites. That is, the metrics allow determination of whether the samples in question will conceal the true biology of health or disease due to sample handling effects, or whether sample handling effects will produce "false positive" biomarker results that do not truly reflect the underlying biology of health or disease. Sample collection / processing metrics have also provided a window into reliable and robust biomarker discovery. By selecting groups of samples with consistent sample preparation metrics, unintentional biases can be minimized and disease-specific biomarker discovery enhanced. By comparison with fully calibrated standard samples, the metrics can also be used to correct for minor sample handling effects. In clinical use, sample handling metrics can be used to advise sites on their collection procedures to reject some samples prior to extensive further evaluation and to adjust the measurements or reports provided to reflect any uncertainties introduced by sample handling.

[0302] Briefly, it is now possible to:

[0303] 1. Determine the form and quantitative extent of sample handling variation between samples. This allows sample sets to be triaged and samples suitable for biomarker discovery to be isolated.

[0304] 2. Identify or determine preferred sample handling / processing protocols to substantially reduce or minimize variation between samples.

[0305] 3. Similarly, sample handling / processing values for collection sites or sample batches can be compared to reference sample handling / processing biomarker values to determine whether individual sites are compliant with preferred collection protocols.

[0306] 4. Sample sets can be examined and compared to reference sample handling / processing biomarker values to determine the expected degree of handling and processing variation that may exist between case and control samples. In this way, subsets of samples can be selected for comparison on the basis of similar sample collection conditions, such that the identified biomarkers are a reliable reflection of the underlying biology.

[0307] 5. If it is determined that samples were not collected in a manner compliant with the preferred sample handling / processing protocol, individual samples can be rejected for diagnostic testing.

[0308] 6. Protein measurements for one or more case samples can be adjusted to reflect sample handling / processing variability.

[0309] 7. A subset of robust proteins that are less sensitive to sample handling / processing variability can be selected for clinical or commercial use.

[0310] Accordingly, the present invention includes a method for quantifying the effects of deviations from ideal blood sample collection conditions. This method includes, prior to proteomic assay measurements, identifying biological processes affected by variations in steps involved in blood sample draw and processing. These biological processes are monitored by a list of specific analyte (e.g., protein) measurements, the analytes being uniquely identified by such processes and capable of being monitored. Using protein coefficients specific to each protein to be measured, these protein lists are applied quantitatively using logarithmic measurements of protein abundance projections. A score called the sample processing marker SMV (sample marker variation) from these projections can be used to evaluate process variation in blood sample collection on a per sample and per group of samples basis.

[0311] In one aspect, the present invention provides a method for generating SMV coefficients therefrom. Specifically, a method for quantifying the effects of deviations from ideal blood sample collection conditions has been identified. This method includes, prior to proteomic assay measurements, identifying biological processes affected by variations in steps involved in blood sample draw and processing. These biological processes are monitored by a list of specific protein measurements, the analytes being uniquely identified by such processes and capable of being monitored by us. Using protein coefficients specific to each protein to be measured, these protein lists are applied quantitatively using logarithmic protein measurements of protein abundance projections. A score called SMV from these projections can be used to evaluate process variation in blood sample collection on a per sample and per group of samples basis. These biological processes can be used to monitor variations in blood sample collection conditions, and specific protein vectors can be used to monitor and quantify such biological processes. This provides quantification of sample collection variations, which are recorded in the samples themselves and do not require independent monitoring of variables such as time, temperature, and centrifugation speed at the time of collection.

[0312] To identify the SMV protein components, targeted experiments are used that involve biochemical manipulation of specific biological processes such as complement activation, platelet activation, and cell lysis. These experiments are combined with experiments that vary blood sample collection conditions in a manner consistent with clinical practice to uniquely identify biological processes that can be used to quantitatively evaluate variations in clinical sample collection on a per sample basis.

[0313] The techniques described herein can be used for quality assessment of samples with respect to protein measurements directly involved in these biological processes. This provides a quantitative measure of sample quality, which can be applied to inform decisions regarding protein measurements in these samples, where the proteins can be affected by sample processing variations rather than simply directly associated with the biological processes being measured here. For example, general proteolytic activity can be affected by complement activation and cell lysis. However, the proteins involved do not form a simple closed set or process and cannot be used to monitor complement and cell lysis because other proteins can vary for a variety of reasons between samples, which are not related to sample processing variations such as disease processes or renal function.

[0314] The present invention is the use of a set of proteins with coefficients to monitor biological processes and indirectly monitor variations in sample collection conditions, which has the advantage over a single protein in that it is less likely to suffer from the drawbacks of individual variation and forms a measurement population that can be interpreted as giving a robust estimate of biological process activation. The use of log-scale measurements allows monitoring of relative fold changes in biological process activation and can be simply compared to a reference sample using differences corresponding to ratios in linear space. This use of logarithms also implicitly calibrates the protein measurements such that when using a reference sample, different concentration ranges between the proteins in the set or vector are automatically normalized.

[0315] The direct application of the SMV calculation to individual blood samples provides a score that can be interpreted in terms of biological processes or indirectly in terms of the deviation of specific sample collection conditions from the ideal conditions of a reference sample. These scores can then be used to define which samples meet the criteria or fall within acceptable limits. This information can be used to reject individual samples. Rejecting individual samples is important in the biomarker discovery process to avoid assigning changes in protein abundance to a disease or process under study for biomarker discovery when such changes may have been caused by some subset of the individual set of samples being processed under different sample collection protocols or conditions.

[0316] The SMV score for an individual sample can be used to group sets of samples corresponding to specific ranges of sample collection parameters. This allows the definition of matched sets of samples, where samples from one set have comparable sample collection procedures and parameters to samples from a previous or different collection study. This ability to form matched sets is invaluable in comparing groups of samples that can be collected under different conditions. If associated changes in other proteins can be determined, the SMV score calculated for an individual sample can also be used to correct for variations in sample processing and construct a mathematical model based on the changes in each protein affected by the process, resulting in variations between samples with different SMV scores.

[0317] Individual sample rejection based on its SMV score allows for more sensitive biomarker discovery, as we know that differences between samples collected from clinically distinct individuals refer to differences between those individuals, rather than differences in how the samples were collected. If the variation is due to the procedure by which the blood sample was collected, rather than the clinical state of the individual, diagnostic tests involving protein abundance may be misleading. This is avoided by rejecting samples that do not meet the SMV score threshold, which corresponds to reasonable variation in the sample collection procedure.

[0318] Many existing sample collections are systematically compromised by variations in the sample collection procedure. SMV scores can be used to quantify such variations within a sample collection or between sample collection sites, and can be used to reject an entire study on the basis of variations that may mislead the researcher, such as systematic variations in the sample collection between cases and controls. Only a subset of the collection to be measured is needed to assess such variations; significant savings are possible in cases where the sample collection is deemed unacceptable. Sample collections during the sample acquisition phase of a study can also be monitored, thus providing correct advice and detecting non-compliance with the study protocol. To monitor variations in an existing or ongoing study, only some subsamples of the entire collection need to be measured.

[0319] These techniques for monitoring and assessing sample collection variations can be applied to the optimization of study protocols and can be applied to the economic maximization of large sample collection efforts such as biobanks, where the cost of using specialized sample collection equipment and vessels can be compared to the accurate assessment of variations and compromises due to the use of less expensive protocol operations.

[0320] In some cases, the original sample collection may not be available, possibly due to the retrospective nature of the most common collections of biological samples. Additionally, some comparisons may necessarily occur between samples collected at different sites and between groups of samples collected at different times. These sample collections will show differences in the collection procedure, which will cause changes in the proteomic profiles, and which will be confounded by the expected differential clinical comparisons. By generating matching sets between groups of samples, subsets of equivalently collected samples can be compared.

[0321] The measurement of protein analytes in plasma samples can be significantly affected by the protocol used for collecting and processing the samples. Deviations from a specific sample collection and / or processing protocol can lead to changes in the protein levels within the sample or other systematic effects on the measurement, which result in signal changes for many analytes, including negative controls. Such deviations can occur regardless of the type of assay used to measure the protein analyte.

[0322] To evaluate the quality of a clinical sample set, the effects of the most obvious deviations from the protocol have been characterized. The variability of the protein composition over time has been evaluated between sample collection and spinning. In addition, the variability of the protein composition as a function of time has been evaluated between the time of sample spinning and sample decantation.

[0323] Signatures of sample mishandling have been identified, which can be used as a quantitative classifier to evaluate a clinical sample set. In addition, metrics have been generated for each analyte that capture the sensitivity of the measurement of that analyte to deviations from the collection protocol, particularly with respect to the delays between sample collection and spinning and between sample spinning and sample decantation.

[0324] Examples

[0325] The following examples are provided for illustrative purposes only and are not intended to limit the scope of the present application as defined by the appended claims. All examples described herein were performed using standard techniques well known and conventional to those of ordinary skill in the art. The conventional molecular biology techniques described in the following examples can be performed as described in standard laboratory manuals such as Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., (2001).

[0326] Example 1 Sample Collection

[0327] Plasma samples were collected from a group of eighteen individuals, with all sample collection variables held constant under a defined protocol except for the target variable. Multiple tubes were drawn from the same set of individuals to evaluate the variation in response between different individuals.

[0328] Sample Collection - Step

[0329] 1. Check the expiration date of the purple top EDTA tube. If expired, replace with a new tube.

[0330] 2. Label the purple top EDTA tube with the correct participant ID and collection date.

[0331] 3. Perform venipuncture following standard health and safety procedures in accordance with institutional guidelines.

[0332] 4. Completely fill the purple top EDTA tube with the sample from the venipuncture.

[0333] 5. If tubes are drawn simultaneously in the laboratory, collect the tubes in the following order:

[0334] a. Serum

[0335] b. Citrate plasma

[0336] c. Heparin plasma

[0337] d. Purple-top EDTA plasma

[0338] 6. Invert the purple-top EDTA tube 8 - 10 times and stand it upright on the rack.

[0339] Rotation Time

[0340] Collect the sample in a vacuum blood collection tube and invert it as described in the sample collection - steps above. Subsequently, allow six different times, namely 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours, to elapse before spinning the sample for each of the eighteen individuals. Spin the purple-top EDTA tube at 2200 x g (not RPM) for 15 minutes. Label the microcentrifuge tube with the correct participant ID. Pipette 1.0 ml of plasma into the microcentrifuge tube. Only draw out the plasma layer. Be careful not to disturb the buffy coat during aliquoting, leaving some plasma and avoiding the cell layer. Seal the top of the microcentrifuge tube and place it in a -80 °C freezer.

[0341] Decantation / Freezing Time

[0342] Collect the sample as described in the above sample collection - steps and time to sample spin. Subsequently, allow six different times, namely 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours, to elapse before decanting and freezing the spun sample for each of the eighteen individuals.

[0343] Univariate Analysis

[0344] Prior to model generation, perform univariate analysis of all analyte signals with respect to spin time / decant time to reduce the number of features (analytes) used in model construction ( Figures 1 to 2 ).

[0345] Calculate the Pearson correlation of RFU for 18 individuals for each of approximately 5K analytes (Equation 1) to evaluate the general functional impact at different spin times / decant times.

[0346]

[0347] Despite the continuous behavior, analytes can be characterized as having high discriminatory properties ( Figure 1 ), where the Pearson correlation coefficient > 0.95, or very low discriminatory properties, where the correlation coefficient < 1E - 3 (Figure 2 ) showed no variation with spin time in fact among 18 individuals.

[0348] Summary statistics for a few analytes with high spin time / decant time correlations are shown in Tables 1 and 2. Table 3 lists the importance of spin time / decant time for analytes. There are qualitative groups of analytes in the spin time model; those that show negative or positive correlation shifts in RFU with increasing spin time and those that show negative or positive correlation shifts in RFU with different degrees of spin time responses. Table 4 shows the correlation between the measured analyte levels and spin time (e.g., the measured SHH level decreases (negative correlation) with increasing time from collection to spin).

[0349] Using the calculated analyte correlations, we reduced the number of potential markers available for spin time / decant time classification from approximately 5K to approximately 100, which calculated Pearson correlations >|0.7|. Using the initial analyte set, we performed further feature reduction when constructing the classifier.

[0350] Example 2 classifier generates

[0351] A random forest classifier was selected to generate a sample processing model. The following is a brief introduction to random forests, their implementation using SOMAscan data, and their strength relative to another machine learning technique.

[0352] Random Forest Model

[0353] Briefly, a random forest is a collection of many (hundreds) decision trees as in the following example ( Figure 4 ). The RFU levels at a node will split the tree in two directions - towards an end point and a classification prediction or towards another node where additional analyte RFU values will split the tree again and further down towards multiple branches.

[0354] The benefit of a random forest is that one of the decision trees will tend to predict errors such as Figure 4 multiple incorrect binning in Figure 5 ), and the average prediction over hundreds of trees will reduce the error on any given prediction (

[0355] Model Generation

[0356] The random forest model was trained in R using the Caret (Kuhn, M. (2008). Caret package. Journal of Statistical Software, 28(5)) and randomForest (A. Liaw and M. Wiener (2002). Classification and Regression by randomForest. R News 2(3), 18 - 22) packages on the log 10 transformed RFU data. We performed further feature reduction by evaluating IncNodePurity (Gini index in classification), a measure of the relative importance of each analyte to the model performance. Thereby we further reduced the features to produce a model that included the 10 most important analytes in terms of spin time / decant time (Table 3, Figure 3 ).

[0357] When evaluating the model performance using individual analytes ( Figure 6 ; PGAM1 model), we used two metrics. First, we evaluated the predicted time relative to the true time using the root mean square error (RMSE), or we thresholded the sampling time based on what we considered to be well - collected samples (true spin time / decant time less than 2 hours) or poorly - collected samples (true time greater than 2 hours). Using this binary classification, we could assign predictions as true positives (TP), meaning the predicted time accurately described well - collected samples, true negatives (TN), meaning the predicted time accurately described poorly - collected samples, and cross - terms false positives (FP) and false negatives (FN). Using only PGAM1 as a predictor of spin time ( Figure X ), we observed a good level of sensitivity / specificity in the binary classification system, although at longer spin times we often underestimated or overestimated the true value.

[0358] Model Stability

[0359] An additional benefit shown by the random forest model is stability in the presence of assay noise. Consider a sample with a true spin time of 9 hours ( Figure 7 ). If the RFU value on a single analyte (SHH in the example) increases / decreases by a certain effect size, the model prediction (solid curve) is more stable than a similar naive Bayes model (dashed curve).

[0360] When independently adjusting the RFU on each of the 10 analytes for a given sample and spin time, we observed that there were relatively stable points around which the actual prediction did not change significantly (Figure 8 )。With more extreme variations in the analyte signal, we observed a modest jump in the prediction time rather than a continuous change.

[0361] Binary Classification Performance

[0362] Using a predefined cut-off time of 2 hours (where samples with an actual spin time / decant time below 2 hours are considered "good" samples, while samples with an actual spin time / decant time greater than 2 hours are considered "compromised / bad" samples), the overall sensitivity and specificity were defined as each model's prediction relative to the known actual class. Using this binary classification, predictions were assigned as true positives (TP), meaning the predicted time accurately described well-collected samples, true negatives (TN), meaning the predicted time accurately described poorly collected samples, and the cross terms false positives (FP) that incorrectly described poorly collected samples and false negatives (FN) that incorrectly described well-collected samples. For example, based on the PGAM1 model ( Figure 6 ), at actual times of 1.5 hours or 3 hours, the binary classification performance in Table 5 was generated.

[0363] The confusion matrix contains the following information:

[0364]

[0365] Of which 17 samples were labeled as true positives, 15 samples were labeled as true negatives, 1 sample was labeled as a false negative, and 3 samples were labeled as false positives.

[0366] The sensitivity of the model was calculated as:

[0367]

[0368]

[0369] And the specificity was calculated as:

[0370]

[0371]

[0372] For these 18 individuals, at 2 time points, the sensitivity / specificity corresponded to:

[0373]

[0374]

[0375]

[0376]

[0377] Complete sensitivity / specificity was calculated for 18 individuals at 6 spin times / decant times.

[0378] RMSE Calculation

[0379] Root mean square error (RMSE) is a continuous measure of the performance calculated relative to the predicted true spin time at each sample and time.

[0380]

[0381]

[0382] For the 18 individuals with an actual spin time of 9 hours in the exemplary PGAM1 marker, the numerator of the equation contains the data in Table 6.

[0383] The sum of the squared differences is taken, and with N = 18 samples, the equation simplifies to:

[0384]

[0385]

[0386] Reducing the difference between the predicted time and the actual time reduces the RMSE and is thus a good indicator of model performance. An RMSE of 0 at all time points and samples would correspond to each prediction being equal to the actual spin time (i.e., a perfect predictor).

[0387] Model Training and Performance

[0388] Eighteen individuals were used to train the model and evaluate model performance. A "good" sample was defined as having a predicted spin time of less than 2 hours to create a binary classification system. Additionally, the relative error associated with the predicted time relative to the actual time is an additional metric of model performance (RMSE). The random forest model restricted predictions to between 0 hours and 24 hours (where data was available). Figure 6 Shows the performance of a model with a single analyte predictor plotted relative to the true spin time as an indicator of model accuracy, colored by whether the sample was correctly identified as a true positive, where the true positive samples are those with a correctly predicted spin time of less than 2 hours and alternative classification predictions.

[0389] In Figures 9 to 18 the cumulative distribution function for 18 individuals at different spin times was found.

[0390] Analyte Model Performance

[0391] The analytical summaries of the performance of individual and combined quantitative analytes are presented in Tables 7 to 24. Table 7 shows the spin times of the single-marker model performance at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 8 shows the spin times of the two-marker model performance at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 9 shows the spin times of the three-marker model performance at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 10 shows the spin time performance of the model with Sonic Hedgehog (SHH) at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 11 shows the spin time performance of the model with Indian Hedgehog (IHH) at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 12 shows the spin time performance of the model with ADAM9 at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 13 shows the spin time performance of the model with DDX39B at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 14 shows the spin time performance of the model with FAM49B at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 15 shows the spin time performance of the model with PGAM1 at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 16 shows the spin time performance of the model with PTPN4 at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 17 shows the spin time performance of the model with RBP7 at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 18 shows the spin time performance of the model with S100A12 at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 19 shows the spin time performance of the model with TNFSF14 at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning). Table 20 shows the decant times of the single-marker model performance at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before spinning).Table 21 shows the decantation times of the two - marker model performance at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before rotation). Table 22 shows the decantation times of the three - marker model performance at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before rotation). Table 23 shows the rotation time performance of the model with the combination of IHH, RBP7, ADAM9, and PTPN4 at all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before rotation). Table 24 shows the rotation time performance for models with analyte combinations and for all predicted time points (samples were allowed to stand for 0 hours, 0.5 hours, 1.5 hours, 3 hours, 9 hours, and 24 hours before rotation), where some of the analyte combinations include PGAM1 and / or PTPN4.

[0392] Analyte Model Clustering

[0393] For each individual and time, the performance of each model was quantified using the RMSE of the predicted rotation time relative to the true rotation time. Figure 19 The distribution of the RMSE values for 1023 models is shown. The distribution was divided into 4 groups of model performance - for high - performance models, the RMSE is between 0 and 0.35, for medium - range performance models, the RMSE is between 0.35 and 0.5, for low - performance models, the RMSE is between 0.5 and 1, and for very low - performance models, the RMSE is between 1 and 2.

[0394] Analytes Used

[0395] As Figures 20 to 22 shown, the fractional multiples of the analytes used in each group of model performance were quantified to illustrate the importance of each analyte to model performance.

[0396] Analyte Groups in Model Performance

[0397] In high - performance models, the distribution of the number of required analytes is shown in Figure 23 . For all analytes, well - performing models can use as few as 2 analytes.

[0398] Example 3. Multiplex Aptamer Assay of Samples

[0399] This example describes a multiplex aptamer assay for analyzing samples and controls to identify sample collection / processing variability markers listed in Table 1.

[0400] Multiplex Aptamer Assay Method

[0401] Unless otherwise specified, all steps of the multiplexed aptamer assay are performed at room temperature.

[0402] Preparation of the aptamer master mix solution.

[0403] The 5272 aptamers were grouped into three unique mixtures: Dil1, Dil2, and Dil3, corresponding to 20%, 0.5%, and 0.005% plasma or serum sample dilutions, respectively. The assignment of aptamers to the mixtures was determined empirically by assaying dilution series of plasma and serum samples matched to each aptamer and identifying the sample dilution that produced the maximum linear range of signals. Separation of the aptamers and mixing with plasma or serum samples at different dilutions (20%, 0.5%, or 0.005%) allowed the determination of protein concentrations spanning a 10 7 -fold range. Stock solutions of the aptamer master mix were prepared at 4 nM for each aptamer in HE-Tween buffer (10 mM Hepes, pH 7.5, 1 mM EDTA, 0.05% Tween 20) and stored frozen at -20 °C. 4271 aptamers were mixed in the Dil1 mixture, 828 aptamers were mixed in Dil2, and 173 aptamers were mixed in the Dil3 mixture. Prior to use, the stock solutions were diluted in HE-Tween buffer to a working concentration of 0.55 nM for each aptamer and aliquoted into individual-use aliquots. Before using the aptamer master mix for Catch-0 plate preparation, the working solution was heat-cooled to refold the aptamers by incubating at 95 °C for 10 minutes and then at 25 °C for at least 30 minutes prior to use.

[0404] Catch-0 plate preparation.

[0405] Dispense 60 μL of streptavidin magnetic agarose 10% slurry (GE Healthcare, 28-9857) into each well of a 96-well plate (Thermo Scientific, AB-0769). Wash the beads once with 175 μL of assay buffer (40 mM HEPES, pH 7.5, 100 mM NaCl, 5 mM KCl, 5 mM MgCl2, 1 mM EDTA, 0.05% Tween 20), then add 100 μL of heated and cooled aptamer master mix to each well. Incubate the plate for 30 minutes at 25 °C with shaking at 850 rpm on a ThermoMixer C shaker (Eppendorf). After 30 minutes of incubation, add 6 μL of MB blocking buffer (50 mM D-biotin in 50 mM Tris-HCl, pH 8, 0.01% Tween) to each well of the plate and further incubate the plate for 2 minutes with shaking. Then wash the plate with 175 μL of assay buffer, with a wash cycle of shaking for 1 minute at 850 rpm on a ThermoMixer C followed by separation on a magnet for 30 seconds. After removing the wash solution, resuspend the beads in 175 μL of assay buffer and store at -20 °C until use.

[0406] Catch-2 bead preparation.

[0407] Before the start of the robotic processing of the assay, wash the bead slurry of 10 mg / mL MyOne streptavidin C1 beads (Dynabeads, part number 35002D, Thermo Scientific) for the Catch-2 step of the multiplex aptamer assay with MBPrep buffer (10 mM Tris-HCl, pH 8, 1 mM EDTA, 0.4% SDS) for 5 minutes as a whole, then wash twice with assay buffer. After the last wash, resuspend the beads at a concentration of 10 mg / mL and dispense 75 μL of bead slurry into each well of the Catch-2 plate. At the start of the assay, place the Catch-2 plate in an aluminum adapter and in the appropriate position on the Fluent platform.

[0408] Sample thawing and dilution.

[0409] A 65 μL aliquot of 100% plasma or serum samples stored in Matrix tubes at -80 °C was thawed by incubating at room temperature for ten minutes. To facilitate thawing, the tubes were placed on top of a fan unit that circulated air through the Matrix tube rack. After thawing, the samples were centrifuged at 1000 x g for 1 minute and placed on the Fluent robotic platform for sample dilution. A 20% sample solution was prepared by transferring 35 μL of the thawed sample into a 96-well plate containing 140 μL of the appropriate sample diluent. The sample diluent for plasma was 50 mM Hepes, pH 7.5, 100 mM NaCl, 8 mM MgCl2, 5 mM KCl, 1.25 mM EGTA, 1.2 mM benzamidine, 37.5 μM Z-Block, and 1.2% Tween 20. The serum sample diluent contained 75 μM Z-block and the concentrations of the other components were the same as in the plasma sample diluent. 0.5% and 0.005% diluted samples were subsequently diluted in the assay buffer using serial dilutions on the Fluent robot. To prepare the 0.5% sample dilution, the 20% sample was serially diluted to a 4% sample by mixing 45 μL of the 20% sample with 180 μL of the assay buffer, and then the 0.5% sample was prepared by mixing 25 μL of the 4% diluted sample with 175 μL of the assay buffer. To prepare the 0.005% sample, an intermediate dilution of 0.05% was prepared by mixing 20 μL of the 0.5% sample with 180 μL of the assay buffer, and then the 0.005% sample was prepared by mixing 20 μL of the 0.05% sample with 180 μL of the assay buffer.

[0410] Sample binding step.

[0411] A Catch-0 plate prepared by immobilizing the aptamer mixture on streptavidin magnetic agarose beads as described above. The frozen plate was thawed at 25 °C for 30 minutes and washed once with 175 μL of the assay buffer. 100 μL of each sample dilution (20%, 0.5%, and 0.005%) was added to the plate containing beads with three different aptamer master mixtures (Dil1, Dil2, and Dil3) respectively. The Catch-0 plate was then sealed with an aluminum foil seal (Microseal ‘F’ foil, Bio-Rad) and placed in a 4-plate rotary shaker (PHMP-4, Grant Bio) set at 850 rpm, 28 °C. The sample binding step was carried out for 3.5 hours.

[0412] Multiplex aptamer assay processing was performed on the Fluent robot.

[0413] After completion of the sample binding step, the Catch-0 plate is placed in an aluminum plate adapter and on a robotic platform. The bead washing step is performed using a temperature control plate. For all robotic processing steps, except for the Catch-2 wash described below, the plate is set at a temperature of 25 °C. The plate is washed 4 times with 175 μL of assay buffer, and each wash cycle program is set to shake the plate at 1000 rpm for at least 1 minute, followed by separating the beads for at least 30 seconds and then aspirating the buffer. During the last wash cycle, the labeling reagent is prepared by diluting 100x labeling reagent (EZ-Link NHS-PEG4-biotin, part number 21363, Thermo, 100 mM solution prepared in anhydrous DMSO) 1:100 in assay buffer and poured into a trough on the robotic platform. 100 μL of the labeling reagent is added to each well in the plate and incubated for 5 minutes with shaking at 1200 rpm to biotinylate the proteins captured on the bead surface. The biotinylation reaction is quenched by adding 175 μL of quenching buffer (20 mM glycine in assay buffer) to each well. The plate is incubated statically for 3 minutes and then washed 4 times with 175 μL of assay buffer, and the wash is performed under the same conditions as described above.

[0414] Photolysis and kinetic excitation.

[0415] After the last wash of the plate, 90 μL of photolysis buffer (2 μM oligonucleotide competitor in assay buffer; the competitor has a nucleotide sequence of 5'-(AC-Bn-Bn)7-AC-3', where Bn indicates a 5-position benzyl-substituted deoxyuridine residue) is added to each well of the plate. The plate is transferred to the photolysis sub-station on the Fluent platform. The sub-station consists of a BlackRay light source (UVP XX series table lamp, 365 nm) and three Bioshake 3000-T shakers (Q Instruments). The plate is irradiated for 20 minutes with shaking at 1000 rpm.

[0416] Catch-2 bead capture.

[0417] At the end of the photocleavage process, the buffer was removed from the Catch-2 plate by magnetic separation, and the plate was washed once with 100 μL of assay buffer. Starting from the Dilution 3 plate, the photocleavage eluate containing the aptamer-protein complex was removed from each Catch-0 plate. First, all 90 μL of the solution was transferred to a Catch-1 eluate plate on a shaker with a raised magnet to capture any streptavidin magnetic agarose beads that might have been aspirated. Thereafter, the solution was transferred to the Catch-2 plate, and the plate was incubated for 3 minutes with shaking at 1400 rpm at 25 °C. After incubation for 3 minutes, the magnetic beads were separated for 90 seconds, the solution was removed from the plate, and the photocleaved Dil2 plate solution was added to the plate. Following the same method, the solution from the Dil1 plate was added and incubated for 3 minutes. At the end of the 3-minute incubation, 6 μL of MB blocking buffer was added to the magnetic bead suspension, and the beads were incubated for 2 minutes with shaking at 1200 rpm at 25 °C. After the incubation, the plate was transferred to a different shaker preset to a temperature of 38 °C. The magnetic beads were separated for 2 minutes before the solution was removed. Then, the Catch-2 plate was washed 4 times with 175 μL of MB wash buffer (20% glycerol in assay buffer), with each wash cycle programmed to shake the beads at 1200 rpm for 1 minute and allow the beads to settle on the magnet for 3.5 minutes. During the final bead separation step, the shaker temperature was set to 25 °C. Then the beads were washed once with 175 μL of assay buffer. For this wash step, the beads were shaken at 1200 rpm for 1 minute and then allowed to separate on the magnet for 2 minutes. After the wash step, the aptamer was eluted from the purified aptamer-protein complex using elution buffer (1.8 M NaClO4, 40 mM PIPES, pH 6.8, 1 mM EDTA, 0.05% Triton X-100). Elution was performed for 10 minutes at 25 °C with shaking the beads at 1250 rpm using 75 μL of elution buffer. 70 μL of the eluate was transferred to an Archive plate and separated on the magnet to dispense any magnetic beads that might have been aspirated. 10 μL of the eluted material was transferred to a black half-area plate, diluted 1:5 in assay buffer, and used to measure the Cy3 fluorescence signal monitored as an internal assay QC. 20 μL of the eluted material was transferred to a plate containing 5 μL of hybridization blocking solution (Oligo aCGH / ChIP-on-Chip hybridization kit, large volume, Agilent Technologies 5188-5380, containing a string of Cy3-labeled DNA sequences complementary to the corner marker probes on the Agilent array). The plate was removed from the robotic platform and further processed for hybridization (see below). The Archive plate with the remaining elution solution was heat-sealed with aluminum foil and stored at -20 °C.

[0418] Hybridization.

[0419] Manually pipette 25 μL of 2x Agilent hybridization buffer (Oligo aCGH / ChIP-on-chip hybridization kit, Agilent Technologies, part number 5188-5380) into each well of the plate containing the eluted sample and blocking buffer. Manually pipette 40 μL of the solution into each "well" of a hybridization gasket slide (hybridization gasket slide - 8 microarray format per slide, Agilent Technologies). Place a custom SurePrint G3 8x60k Agilent microarray slide on the gasket slide according to the manufacturer's protocol, with 10 probes complementary to each aptamer per array. Clamp each assembly (hybridization cassette kit - SureHyb, Agilent Technologies) tightly and load it into a hybridization oven at 55 °C with rotation at 20 rpm for 19 hours.

[0420] Post-hybridization wash.

[0421] Slide washing is performed using a Little Dipper Processor (model 650C, Scienge). Pour approximately 700 mL of wash buffer 1 (Oligo aCGH / ChIP-on-chip wash buffer 1, Agilent Technologies) into a large glass staining dish and use it to separate the microarray slide from the gasket slide. Once disassembled, quickly transfer the slides to a slide rack in a bath containing wash buffer 1 on the Little Dipper. Wash the slides in wash buffer 1 for five minutes with mixing by a magnetic stir bar. Then transfer the slide rack to a bath with wash buffer 2 (Oligo aCGH / ChIP-on-chip wash buffer 2, Agilent Technologies) at 37 °C and allow to incubate for five minutes with stirring. Slowly remove the slide rack from the second bath and then transfer it to a bath containing acetonitrile and incubate for five minutes with stirring.

[0422] Microarray imaging.

[0423] The microarray slides were imaged using a microarray scanner (Agilent G4900DA Microarray Scanner System, Agilent Technologies) in the Cy3 channel at 3 μm resolution with 100% PMT setting and the 20-bit option enabled. The resulting tiff images were processed using Agilent Feature Extraction software (version 10.7.3.1 or later) with the GE1_1200_Jun14 protocol.

Claims

1. A method, comprising: a) measuring the levels of Sonic Hedgehog (SHH) protein and IHH in a sample from a human subject, and optionally measuring the levels of at least one, two, three, four, five, six, seven, eight, nine, ten or eleven proteins selected from the group consisting of PGAM1, PGAM2, PTPN4, TNFSF14, FAM49B, RBP7, DDX39B, S100A12, IL21R, TMEM9 and ADAM9; and b) identifying the sample as an assay sample or a negative sample based on the SHH and IHH levels and optionally the levels of the one, two, three, four, five, six, seven, eight, nine, ten or eleven proteins; wherein the assay sample is a sample for one or more of: protein biomarker discovery assays, protein expression level assays, diagnostic methods or prognostic methods, and the negative sample is a sample not used as an assay sample.

2. The method of claim 1, wherein the method comprises measuring SHH, PGAM1 and IHH; SHH, PTPN4 and IHH; SHH, RBP7 and IHH; SHH, IHH and RBP7; SHH, IHH and TNFSF14; SHH, IHH and FAM49B; SHH, IHH and DDX39B; SHH, IHH and ADAM9; or SHH, IHH and PGAM2.

3. The method of claim 1, wherein the method comprises measuring SHH, IHH and PGAM1, and at least two of the following proteins selected from RBP7, TNFSF14, PTPN4, DDX39B, FAM49B, S100A12, PGAM2, IL21R, TMEM9 and ADAM9.

4. The method of claim 1, wherein the method comprises measuring SHH and IHH, and at least two of the following proteins selected from RBP7, TNFSF14, PTPN4, DDX39B, FAM49B, S100A12, PGAM1, PGAM2, IL21R, TMEM9 and ADAM9.

5. A method, comprising: a) contacting a sample from a human subject with a set of capture reagents, wherein each capture reagent has an affinity for a different protein in a protein set that comprises Sonic Hedgehog (SHH), IHH and optionally at least one, two, three, four, five, six, seven, eight, nine, ten or eleven proteins selected from the group consisting of PGAM1, PTPN4, TNFSF14, FAM49B, RBP7, DDX39B, S100A12, ADAM9, PGAM2, IL21R and TMEM9; and b) Measuring the level of each protein in the protein set with the capture reagent set to identify the sample as an assay sample or a negative sample, wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level assay, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

6. The method of claim 5, wherein the method comprises measuring SHH, PGAM1, and IHH; SHH, PTPN4, and IHH; SHH, RBP7, and IHH; SHH, IHH, and RBP7; SHH, IHH, and TNFSF14; SHH, IHH, and FAM49B; SHH, IHH, and DDX39B; SHH, IHH, and ADAM9; or SHH, IHH, and PGAM2.

7. The method of claim 5, wherein the method comprises measuring SHH, IHH, and PGAM1, and at least two of the following proteins selected from RBP7, TNFSF14, PTPN4, DDX39B, FAM49B, S100A12, and ADAM9.

8. The method of claim 5, wherein the method comprises measuring SHH and IHH, and at least two of the following proteins selected from RBP7, TNFSF14, PTPN4, DDX39B, FAM49B, S100A12, PGAM1, and ADAM9.

9. The method according to any one of the preceding claims, wherein the protein levels are used to assign a quality score to the sample, and the quality score is then used to determine whether the sample is an assay sample or a non-assay sample.

10. The method according to any one of claims 1-8, wherein the protein levels are used to assign a quality score to the sample, and the quality score is then used to determine whether the sample is to be further analyzed for additional proteins in the sample.

11. A method comprising: a) Contacting a sample from a human subject with two capture reagents, wherein one capture reagent has an affinity for SHH protein and a second capture reagent has an affinity for IHH protein; and b) Measuring the level of each protein with the two capture reagents to identify the sample as an assay sample or a negative sample, wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level assay, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

12. A method comprising: a) Measuring the levels of SHH and IHH in a sample from a human subject; and b) Identifying the sample as an assay sample or a negative sample based on the levels of SHH and IHH; wherein the assay sample is a sample for one or more of: protein biomarker discovery assay, protein expression level assay, diagnostic method, or prognostic method, and the negative sample is a sample not used as an assay sample.

13. The method according to any one of claims 1-8, 11-12, wherein the sample is selected from blood or urine.

14. The method according to any one of claims 1-8, 11-12, wherein the sample is selected from plasma or serum.

15. The method according to any one of claims 1-8, 11-12, wherein the protein level is used to predict the length of time between sample collection from the human subject and sample centrifugation and / or the length of time between sample centrifugation and sample decantation.

16. The method according to claim 15, wherein the time between sample collection and sample centrifugation is 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours, and / or the time between sample centrifugation and sample decantation is 0 hours to 0.5 hours; 0.5 hours to 1.5 hours; 1.5 hours to 3 hours; 3 hours to 9 hours; 9 hours to 24 hours or greater than 24 hours.

17. The method according to any one of claims 1-8, 11-12, 16, wherein the measurement of the protein level is performed using mass spectrometry, aptamer-based assays, and / or antibody-based assays.

18. The method according to any one of claims 1-4, 11-12, 16, wherein the method comprises contacting the proteins of a sample from the subject with a set of capture reagents, wherein each capture reagent of the set of capture reagents specifically binds to one protein to be detected.

19. The method according to claim 18, wherein each capture reagent specifically binds to a different biomarker protein to be detected.

20. The method according to claim 18, wherein the set of capture reagents is selected from aptamers, antibodies, and combinations of aptamers and antibodies.

21. The method according to claim 19, wherein the set of capture reagents is selected from aptamers, antibodies, and combinations of aptamers and antibodies.

22. The method according to claim 20, wherein each capture reagent of the set of capture reagents is an aptamer.

23. The method according to claim 21, wherein each capture reagent of the set of capture reagents is an aptamer.

24. The method according to any one of claims 1-8, 11-12, 16, 19-23, wherein the protein level is used for a classifier selected from: decision tree; bagging + boosting + forest; rule-based inductive learning; Barsen window; linear model; logistic; neural network method; unsupervised clustering; K-means; hierarchical ascending / descending; semi-supervised learning; prototype method; nearest neighbor; kernel density estimation; support vector machine; hidden Markov model; Boltzmann learning; and random forest model, wherein the classifier is used together with the protein level to identify a sample as an analytical sample or a negative sample.

Citation Information

Patent Citations

  • Method for generating aptamers with improved off-rates

    US20090004667A1

  • Method for generating aptamers with improved off-rates

    US7947447B2

  • Selection of preferred sample handling and processing protocol for identification of disease biomarkers and sample quality assessment

    CN103958662A

  • Lung Cancer Biomarkers and Uses Thereof

    US20130116150A1