Methods of assessing stomach cancer risk

WO2026206757A1PCT designated stage Publication Date: 2026-10-01ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/020055
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-20
Publication Date
2026-10-01

Smart Images

  • Figure US2026020055_01102026_PF_FP_ABST
    Figure US2026020055_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure includes biomarkers, methods, reagents, systems, and kits for predicting an individual's risk of stomach cancer. In one aspect, the disclosure provides biomarkers that can be used alone or in various combinations to predict an individual's risk of stomach cancer. In another aspect, methods are provided for predicting an individual's risk of stomach cancer, where the methods include detecting, in a biological sample from an individual, at least one biomarker value corresponding to at least one biomarker selected from the group of biomarkers provided in Table 1.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS OF ASSESSING STOMACH CANCER RISKCROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims the benefit of priority of US Provisional Patent Application No. 63 / 779,772 filed on March 28, 2025, the contents of which is incorporated by reference herein in its entirety for any purpose.FIELD

[0002] The present application relates generally to the detection of biomarkers and methods of assessing stomach cancer risk in an individual and, more specifically, to one or more biomarkers, methods, reagents, systems, and kits used to determine an individual’s risk of stomach cancer.BACKGROUND

[0003] The following description provides a summary of information that may be relevant to the present application and is not an admission that any of the information provided or publications referenced herein is prior art to the present application.

[0004] Stomach cancer is the fifth most common cancer worldwide, and the fourth leading cause of cancer death, although incidence rates differ substantially by geographic region. The highest incidence rates are in Eastern Asia (China, South Korea, Japan) and Central -Eastern Europe (Hungary, Poland, Slovenia), with lower incidences in North America, Africa, and Northern Europe.

[0005] The incidence rate of stomach cancer in the US is 7 cases per 100,000 people per year; however, the 5- year relative survival is only 36.4%. This is because stomach tumors are aggressive, are fairly advanced by the time symptoms appear, and those curable cancers that are asymptomatic are rarely detected early enough to treat.

[0006] More than 90% of stomach cancers are adenocarcinomas that begin in the mucusproducing cells deep in the stomach lining. They are classified by their location as either gastric cardia which begin in the top inch of the stomach near the gastroesophageal junction, or noncardia for all other sections of the stomach. Symptoms of stomach cancer include weight loss, abdominal pain, nausea, difficulty swallowing, and reduced appetite. The diagnosis of stomach cancer is based upon biopsy obtained during an upper endoscopy procedure. Treatment of stomach cancer depends on whether the disease appears to be local and within stages I-III or more advanced (stage IV) and metastatic. For patients with local disease, surgery in combination with chemotherapy (pre- and post-operatively) shows increased survival above surgery alone.

[0007] There is no guidance from the US Preventive Services Task Force (USPSTF) around screening for stomach cancer o Helicobacter pylori (H pylori) infection, which is a primary risk factor for stomach cancer. While the mechanism is not completely understood, H pylori has been strongly implicated in gastric adenocarcinoma, likely through a complex relationship between the bacteria and host genetics and environmental factors. H pylori can cause chronic gastritis, an early step in carcinogenesis, and eradication of infection appears to significantly decrease the incidence of gastric adenocarcinoma.

[0008] Other risk factors for stomach cancer include family history of stomach cancer, tobacco use, alcohol consumption, obesity, diets high in processed meat and sodium, and low in fresh fruits and vegetables. As these risk factors become more prevalent in the population, stomach cancer incidence in the US will likely increase and guidance on screening could be revisited by the USPSTF.

[0009] A first step in increasing the survival of gastric cancer is development of non-invasive tests to stratify individual risk and direct screening procedures to individuals in higher risk categories. This would thereby increase the likelihood of detecting stomach lesions in the pre-cancerous state and allow for early intervention and prevention of malignancy.

[0010] In the US, there is no standard screening test to detect stomach cancer in individuals at average risk. Because stomach cancer incidence is lower in the US than in Asian and central European countries, screening with more invasive methods such as upper endoscopy and contrast radiography is recommended only for specific subgroups at high-risk. Other bloodbased screening tests have been proposed, such as serum pepsinogen and 7 / pylori antibody test, serum trefoil factor 3, and microRNA assay, but these methods require further study to support widespread use.

[0011] There is no standard method currently used in clinical practice to predict the risk of future malignant stomach cancer. Healthcare providers consider those with gastritis, pernicious anemia, intestinal metaplasia, H pylori infection, family history of stomach cancer or first-generation immigrants from areas with high stomach cancer incidence as candidates for radiography or endoscopy screening. There are no online calculators to assist individuals in estimating their risk; therefore, healthcare providers consider lifestyle factors, and medical and family history, when determining whether to recommend imaging or endoscopy screening for stomach cancer.

[0012] The development of a proteomic model for predicting an individual’s risk of stomach cancer would be highly desirable. A need exists for biomarkers, methods, reagents, systems, and kits that enable the prediction of the risk of stomach cancer for an individual. Non-limiting benefits of a stomach cancer risk test include: 1) a convenient blood-based test that is non-invasive and does not rely on self-reported lifestyle, demographic or genetic information to predict risk of future stomach cancer; 2) the test results could identify individuals who would benefit from screening by radiography or endoscopy, which would likely increase screening adherence among those identified as higher than average risk; and 3) the test results could identify individuals at lower risk of stomach cancer who could avoid unnecessary invasive procedures.SUMMARY

[0013] The present application includes biomarkers, methods, reagents, systems, and kits for assessing stomach cancer risk. In some embodiments, methods of predicting the risk of stomach cancer in an individual are provided. In some embodiments, methods of predicting the risk of stomach cancer in an individual within a specified time, such as 5 years, are provided. In some embodiments, methods of detecting levels of N biomarker proteins in a sample are provided.

[0014] Embodiment 1. A method of predicting stomach cancer risk in a subject, comprising forming a biomarker panel comprising N biomarker proteins, and detecting a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or 8 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0015] Embodiment 2. A method of detecting levels of N biomarker proteins in a sample, comprising forming a biomarker panel comprising N biomarker proteins, and detecting the level of each of the N biomarker proteins in the sample from a subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or 8 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0016] Embodiment 3. A method of predicting stomach cancer risk in a subject, comprising detecting a level of LYPL1 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0017] Embodiment 4. A method of predicting stomach cancer risk in a subject, comprising detecting a level of CILP2 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the Nbiomarker proteins are selected from LYPL1, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0018] Embodiment 5. A method of predicting stomach cancer risk in a subject, comprising detecting a level of TFF1 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0019] Embodiment 6. A method of predicting stomach cancer risk in a subject, comprising detecting a level of S27A2 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0020] Embodiment 7. A method of predicting stomach cancer risk in a subject, comprising detecting a level of PAP 1 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, GKN2, Inhibin bB chain, and EphAl.

[0021] Embodiment 8. A method of predicting stomach cancer risk in a subject, comprising detecting a level of RPIA and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, Inhibin bB chain, and EphAl.

[0022] Embodiment 9. A method of predicting stomach cancer risk in a subject, comprising detecting a level of Inhibin bB chain and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, and EphAl.

[0023] Embodiment 10. A method of predicting stomach cancer risk in a subject, comprising detecting a level of EphAl and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the Nbiomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, and Inhibin bB chain.

[0024] Embodiment 11. The method according to any one of the preceding embodiments, wherein N is 2 to 8, or N is 3 to 8, N is 4 to 8, N is 5 to 8, N is 6 to 8, or N is 7 to 8.

[0025] Embodiment 12. The method according to any one of the preceding embodiments, wherein N is 2, N is 3, N is 4, N is 5, N is 6, N is 7, or N is 8.

[0026] Embodiment 13. The method according to any one of the preceding embodiments, wherein at least one of the N biomarker proteins is LYPL1, or at least one of the N biomarker proteins is CILP2, or at least one of the N biomarker proteins is TFF1, or at least one of N biomarker proteins is S27A2, or at least one of the N biomarker proteins is PAP1, or at least one of the N biomarker proteins is GKN2, or at least one of the N biomarker proteins is Inhibin bB chain, or at least one of the N biomarker proteins is EphAl .

[0027] Embodiment 14. The method according to any one of the preceding embodiments, wherein each of the N biomarker proteins is selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl .

[0028] Embodiment 15. The method according to any one of the preceding embodiments, wherein at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0029] Embodiment 16. The method according to any one of embodiments 1-3, wherein 2 of the N biomarker proteins are LYPL1 and CILP2, or 2 of the N biomarker proteins are LYPL1 and TFF1, or 2 of the N biomarker proteins are LYPL1 and S27A2, or 2 of the N biomarker proteins are LYPL1 and PAP1, or 2 of the N biomarker proteins are LYPL1 and GKN2, or 2 of the N biomarker proteins are LYPL1 and Inhibin bB chain, or 2 of the N biomarker proteins are LYPL1 and EphAl.

[0030] Embodiment 17. The method according to any one of embodiments 1, 2 or 4, wherein 2 of the N biomarker proteins are CILP2 and TFF1, or 2 of the N biomarker proteins are CILP2 and S27A2, or 2 of the N biomarker proteins are CILP2 and PAP1, or 2 of the N biomarker proteins are CILP2 and GKN2, or 2 of the N biomarker proteins are CILP2 and Inhibin bB chain, or 2 of the N biomarker proteins are CILP2 and EphAl.

[0031] Embodiment 18. The method according to any one of embodiments 1, 2 or 5, wherein 2 of the N biomarker proteins are TFF1 and S27A2, or 2 of the N biomarker proteins are TFF1 and PAP1, or 2 of the N biomarker proteins are TFF1 and GKN2, or 2 of the N biomarker proteins are TFF1 and Inhibin bB chain, or 2 of the N biomarker proteins are TFF1 and EphAl.

[0032] Embodiment 19. The method according to any one of embodiments 1, 2 or 6, wherein 2 of the N biomarker proteins are S27A2 and PAP1, or 2 of the N biomarker proteins are S27A2 and GKN2, or 2 of the N biomarker proteins are S27A2 and Inhibin bB chain, or 2 of the N biomarker proteins are S27A2 and EphAl .

[0033] Embodiment 20. The method according to any one of embodiments 1, 2 or 7, wherein 2 of the N biomarker proteins are PAP1 and GKN2, or 2 of the N biomarker proteins are PAP1 and Inhibin bB chain, or 2 of the N biomarker proteins are PAP1 and EphAl .

[0034] Embodiment 21. The method according to any one of embodiments 1, 2 or 8, wherein 2 of the N biomarker proteins are GKN2 and Inhibin bB chain, or 2 of the N biomarker proteins are GKN2 and EphAl.

[0035] Embodiment 22. The method according to any one of embodiments 1, 2 or 9, wherein 2 of the N biomarker proteins are Inhibin bB chain and EphAl .

[0036] Embodiment 23. The method according to any one of embodiments 1-4, wherein 3 of the N biomarker proteins are LYPL1, CILP2, and TFF1, or 3 of the N biomarker proteins are LYPL1, CILP2, and S27A2, or 3 of the N biomarker proteins are LYPL1, CILP2, and PAP1, or 3 of the N biomarker proteins are LYPL1, CILP2, and GKN2, or 3 of the N biomarker proteins are LYPL1, CILP2, and Inhibin bB chain, or 3 of the N biomarker proteins are LYPL1, CILP2, and EphAl.

[0037] Embodiment 24. The method according to any one of the preceding embodiments, wherein the sample is a blood sample, a plasma sample, a serum sample, or a urine sample.

[0038] Embodiment 25. The method according to any one of the preceding embodiments, wherein the risk of the subject stomach cancer within 5 years from the date that the sample was taken from the subject is predicted.

[0039] Embodiment 26. The method according to any one of the preceding embodiments, wherein detecting is performed using mass spectrometry, an aptamer based assay and / or an antibody based assay.

[0040] Embodiment 27. The method according to any one of the preceding embodiments, wherein the method comprises contacting biomarker proteins of the sample or samples with a set of biomarker capture reagents, wherein each biomarker capture reagent of the set of biomarker capture reagents specifically binds to a different biomarker protein being detected.

[0041] Embodiment 28. The method according to embodiment 27, wherein each biomarker capture reagent is an antibody or an aptamer.

[0042] Embodiment 29. The method according to embodiment 28, wherein each biomarker capture reagent is an aptamer.

[0043] Embodiment 30. The method according to embodiment 29, wherein at least one aptamer is a slow off-rate aptamer.

[0044] Embodiment 31. The method according to embodiment 30, wherein at least one slow off-rate aptamer comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least 10 nucleotides with modifications.

[0045] Embodiment 32. The method according to embodiment 30 or embodiment 31, wherein each slow off-rate aptamer binds to its target protein with an off rate (t’A) of > 20 minutes, > 30 minutes, > 60 minutes, > 90 minutes, > 120 minutes, > 150 minutes, > 180 minutes, > 210 minutes, or > 240 minutes.

[0046] Embodiment 33. The method according to any one of embodiments 26-32, wherein the level of each biomarker protein measured is determined from a relative florescence unit (RFU) or a protein concentration.

[0047] Embodiment 34. The method according to any one of the preceding embodiments, wherein predicting the risk of stomach cancer in the subject is based on input of the levels of the N biomarker proteins measured in a statistical model.

[0048] Embodiment 35. The method according to embodiment 34, wherein the determining comprises analyzing the levels of the N biomarker protein using a survival model.

[0049] Embodiment 36. The method according to embodiment 35, wherein the survival model is an Accelerated Failure Time (AFT) model with a Weibull distribution.

[0050] Embodiment 37. The method according to any one of embodiments 34-36, wherein the model has an area under the curve (AUC) selected from at least 0.64, at least 0.65, at least 0.66, at least 0.67, at least 0.68, at least 0.69, at least 0.7, at least 0.75, at least 0.8, at least 0.85, at least 0.9, or at least 0.95.

[0051] Embodiment 38. The method according to any one of the preceding embodiments, wherein the method comprises predicting a risk of stomach cancer in a subject for the purpose of determining a medical insurance premium or life insurance premium.

[0052] Embodiment 39. The method according to embodiment 38, wherein the method further comprises determining coverage for medical insurance or life insurance.

[0053] Embodiment 40. The method according to any one of embodiments 1-37, wherein the method further comprises using information resulting from the method to predict and / or manage the utilization of medical resources.

[0054] Embodiment 41. A kit comprising N biomarker protein capture reagents, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 ofthe N biomarker protein capture reagents specifically binds to a biomarker protein selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0055] Embodiment 42. The kit according to embodiment 41, wherein N is 2 to 8, or N is 3 to 8, or N is 4 to 8, or N is 5 to 8, or N is 6 to 8, or N is 7 to 8.

[0056] Embodiment 43. The kit according to embodiment 41 or 42, wherein N is 2, N is 3, N is 4, N is 5, N is 6, N is 7, or N is 8.

[0057] Embodiment 44. The kit according to any one of embodiments 41-43, wherein each of the N biomarker protein capture reagents specifically binds to a different biomarker protein.

[0058] Embodiment 45. The kit according to any one of embodiments 41-44, wherein each of the N biomarker protein capture reagents specifically binds to a biomarker protein selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0059] Embodiment 46. The kit according to any one of embodiments 41-44, wherein the N biomarker protein capture reagent specifically bind to the N biomarker proteins of any one of embodiments 1-23.

[0060] Embodiment 47. A kit comprising N biomarker protein capture reagents, wherein the kit comprises biomarker protein capture reagents for carrying out the method of any one of embodiments 1-40.

[0061] Embodiment 48. The kit according to any one of embodiments 41-47, wherein each of the N biomarker protein capture reagents is an antibody or an aptamer.

[0062] Embodiment 49. The kit according to embodiment 48, wherein each biomarker protein capture reagent is an aptamer.

[0063] Embodiment 50. The kit according to embodiment 49, wherein at least one aptamer is a slow off-rate aptamer.

[0064] Embodiment 51. The kit according to embodiment 50, wherein at least one slow off-rate aptamer comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least 10 nucleotides with modifications.

[0065] Embodiment 52. The kit according to embodiment 50 or embodiment 51, wherein each slow off-rate aptamer binds to its target protein with an off rate (t%) of > 20 minutes, > 30 minutes, > 60 minutes, > 90 minutes, > 120 minutes, > 150 minutes, > 180 minutes, > 210 minutes, or > 240 minutes.

[0066] Embodiment 53. The kit according to any one of embodiments 41-52, for use in detecting the N biomarker proteins in a sample from a subject.

[0067] Embodiment 54. The kit according to embodiment 53, for use in predicting an individual’s risk of stomach cancer.BRIEF DESCRIPTION OF THE DRAWINGS

[0068] FIG. 1 illustrates certain exemplary 5-position modified uridines and cytidines that may be incorporated into aptamers.

[0069] FIG.2 illustrates certain exemplary modifications that may be present at the 5-position of uridine. The chemical structure of the C-5 modification includes the exemplary amide linkage that links the modification to the 5-position of uridine. The 5-position moieties shown include two phenyl groups covalently attached to one another. The 5-position moieties shown include a phenylbenzyl moiety (e.g., BPE, PBnd, DBM), a 4-phenoxybenzyl moiety (e.g., POP), a diphenylpropyl moiety (e.g., DPP), abenzhydryl moiety (e.g., BH).

[0070] FIG.3 illustrates certain exemplary modifications that may be present at the 5-position of cytidine. The chemical structure of the C-5 modification includes the exemplary amide linkage that links the modification to the 5-position of cytidine. The 5-position moieties shown include two phenyl groups covalently attached to one another. The 5-position moieties shown include a phenylbenzyl moiety (e.g., BPE, PBnd, DBM), a 4-phenoxybenzyl moiety (e.g., POP), a diphenylpropyl moiety (e.g., DPP), a benzhydryl moiety (e.g., BH).

[0071] FIG. 4 illustrates certain exemplary modifications that may be present at the 5-position of uridine. The chemical structure of the C-5 modification includes the exemplary amide linkage that links the modification to the 5-position of the uridine. The 5-position moieties shown include a benzyl moiety (e.g., Bn, PE and a PP), a naphthyl moiety (e.g., Nap, 2Nap, NE), a butyl moiety (e.g, iBu), a fluorobenzyl moiety (e.g., FBn), a tyrosyl moiety (e.g., a Tyr), a 3,4-methylenedioxy benzyl (e.g., MBn), a morpholino moiety (e.g., MOE), a benzofuranyl moiety (e.g., BF), an indole moiety (e.g, Trp) and a hydroxypropyl moiety (e.g., Thr).

[0072] FIG. 5 illustrates certain exemplary modifications that may be present at the 5-position of cytidine. The chemical structure of the C-5 modification includes the exemplary amide linkage that links the modification to the 5-position of the cytidine. The 5-position moieties shown include a benzyl moiety (e.g., Bn, PE and a PP), a naphthyl moiety (e.g., Nap, 2Nap, NE, and 2NE) and a tyrosyl moiety (e.g., a Tyr).

[0073] FIG. 6 illustrates an exemplary computer system for use with various computer-implemented methods described herein.

[0074] FIG. 7 is a flowchart for a method of predicting stomach cancer risk in accordance with one or more embodiments.

[0075] FIG. 8 illustrates a Kaplan-Meier (KM) survival curves for the selected 5-year Malignant Stomach Cancer Risk model, stratified by quartiles of predicted 5-year risk of malignant stomach cancer diagnosis (“Event”) from the training dataset.

[0076] FIG. 9 is a calibration plot for the selected 5-year Stomach Cancer Risk model.

[0077] FIG. 10 illustrates a Kaplan-Meier (KM) survival curve for the selected 5-year Stomach Cancer Risk model on the validation dataset, stratified by quartiles of predicted 5-year risk of malignant stomach cancer diagnosis (“Event”) from the training dataset.

[0078] FIG. 11 illustrates the concordance plot for the stomach cancer risk model predictions between Citrate plasma training and simulated EDTA plasma training data.DETAILED DESCRIPTION

[0079] Reference will now be made in detail to representative embodiments of the invention. While the invention will be described in conjunction with the enumerated embodiments, it will be understood that the invention is not intended to be limited to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents that may be included within the scope of the present invention as defined by the claims.

[0080] One skilled in the art will recognize many methods and materials similar or equivalent to those described herein, which could be used in and are within the scope of the practice of the present invention. The present invention is in no way limited to the methods and materials described.

[0081] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods, devices, and materials similar or equivalent to those described herein can be used in the practice or testing of the invention, certain methods, devices and materials are now described.

[0082] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below. All publications, published patent documents, and patent applications cited in this application are indicative of the level of skill in the art(s) to which the application pertains. All publications, published patent documents, and patent applications cited herein are hereby incorporated by reference to the same extent as though each individual publication, published patent document, or patent application was specifically and individually indicated as being incorporated by reference.

[0083] As used in this application, including the appended claims, the singular forms “a,” “an,” and “the” include plural references, unless the content clearly dictates otherwise, and are used interchangeably with “at least one” and “one or more.” Thus, reference to “a SOMAmer” includes mixtures of SOMAmers, reference to “a probe” includes mixtures of probes, and the like. It is further to be understood that all base sizes or amino acid sizes, and all molecularweight or molecular mass values, given for nucleic acids or polypeptides are approximate, and are provided for description.

[0084] Further, ranges provided herein are understood to be shorthand for all of the values within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range from the group consisting 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 (as well as fractions thereof unless the context clearly dictates otherwise). Any concentration range, percentage range, ratio range, or integer range is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (such as one tenth and one hundredth of an integer), unless otherwise indicated. Also, any number range recited herein relating to any physical feature are to be understood to include any integer within the recited range, unless otherwise indicated.

[0085] As used herein, the term “about” represents an insignificant modification or variation of the numerical value such that the basic function of the item to which the numerical value relates is unchanged.

[0086] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “contains,” “containing,” and any variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, product-by-process, or composition of matter that comprises, includes, or contains an element or list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, product-by-process, or composition of matter.

[0087] The present application includes biomarkers, methods, reagents, systems, and kits for determining the risk of stomach cancer in an individual.

[0088] “ Stomach cancer” as used herein refers to all those found within the fundus (Cl 61), body (Cl 62), greater (Cl 66) and lesser (Cl 65) curvature of the stomach, overlapping lesions (Cl 68), gastric antrum (Cl 63), pylorus (Cl 64), and stomach and cardia not otherwise specified (Cl 69 and Cl 60). Cancer types were defined using the International Classification of Diseases-Tenth Revision and the second revision of the International Classification of Diseases for Oncology (ICD-O-2).

[0089] “Biological sample,” “sample,” and “test sample” are used interchangeably herein to refer to any material, biological fluid, tissue, or cell obtained or otherwise derived from an individual. This includes blood (including whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, and serum), dried blood spots, sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, ascites, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleuralfluid, nipple aspirate, bronchial aspirate, bronchial brushing, synovial fluidjoint aspirate, organ secretions, cells, a cellular extract, and cerebrospinal fluid. This also includes experimentally separated fractions of all of the preceding. For example, a blood sample can be fractionated into serum, plasma or into fractions containing particular types of blood cells, such as red blood cells or white blood cells (leukocytes). If desired, a sample can be a combination of samples from an individual, such as a combination of a tissue and fluid sample. The term “biological sample” also includes materials containing homogenized solid material, such as from a stool sample, a tissue sample, or a tissue biopsy, for example. The term “biological sample” also includes materials derived from a tissue culture or a cell culture. Any suitable methods for obtaining a biological sample can be employed; exemplary methods include, e.g., phlebotomy, swab (e.g., buccal swab), and a fine needle aspirate biopsy procedure. Exemplary tissues susceptible to fine needle aspiration include lymph node, lung, lung washes, BAL (bronchoalveolar lavage), thyroid, breast, pancreas, and liver. Samples can also be collected, e.g., by micro dissection (e.g., laser capture micro dissection (LCM) or laser micro dissection (LMD)), bladder wash, smear (e.g., a PAP smear), or ductal lavage. A “biological sample” obtained or derived from an individual includes any such sample that has been processed in any suitable manner after being obtained from the individual.

[0090] For purposes of this specification, the phrase “data attributed to a biological sample from an individual” is intended to mean that the data in some form derived from, or were generated using, the biological sample of the individual. The data may have been reformatted, revised, or mathematically altered to some degree after having been generated, such as by conversion from units in one measurement system to units in another measurement system; but, the data are understood to have been derived from, or were generated using, the biological sample.

[0091] “Target,” “target molecule,” and “analyte” are used interchangeably herein to refer to any molecule of interest that may be present in a biological sample. A “molecule of interest” includes any minor variation of a particular molecule, such as, in the case of a protein, for example, minor variations in amino acid sequence, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component, which does not substantially alter the identity of the molecule. A “target molecule,” “target,” or “analyte” is a set of copies of one type or species of molecule or multi-molecular structure. “Target molecules,” “targets,” and “analytes” refer to more than one such set of molecules. Exemplary target molecules include proteins, polypeptides, nucleic acids, carbohydrates, lipids, polysaccharides, glycoproteins, hormones, receptors, antigens, antibodies, affybodies, antibody mimics, viruses, pathogens, toxic substances,substrates, metabolites, transition state analogs, cofactors, inhibitors, drugs, dyes, nutrients, growth factors, cells, tissues, and any fragment or portion of any of the foregoing. In some embodiments, a target molecule is a protein, in which case the target molecule may be referred to as a “target protein.”

[0092] As used herein, a “capture agent’ or “capture reagent” refers to a molecule that is capable of binding specifically to a biomarker. A “target protein capture reagent” refers to a molecule that is capable of binding specifically to a target protein. Nonlimiting exemplary capture reagents include aptamers, antibodies, adnectins, ankyrins, other antibody mimetics and other protein scaffolds, autoantibodies, chimeras, small molecules, nucleic acids, lectins, ligandbinding receptors, imprinted polymers, avimers, peptidomimetics, hormone receptors, cytokine receptors, synthetic receptors, and modifications and fragments of any of the aforementioned capture reagents. In some embodiments, a capture reagent is selected from an aptamer and an antibody.

[0093] As used herein, “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to polymers of amino acids of any length. The polymer may be linear or branched, it may comprise modified amino acids, and it may be interrupted by non-amino acids. The terms also encompass an amino acid polymer that has been modified naturally or by intervention; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component. Also included within the definition are, for example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids, etc.), as well as other modifications known in the art. Polypeptides can be single chains or associated chains. Also included within the definition are preproteins and intact mature proteins; peptides or polypeptides derived from a mature protein; fragments of a protein; splice variants; recombinant forms of a protein; protein variants with amino acid modifications, deletions, or substitutions; digests; and post-translational modifications, such as glycosylation, acetylation, phosphorylation, and the like.

[0094] The term “antibody” refers to full-length antibodies of any species and fragments and derivatives of such antibodies, including Fab fragments, F(ab')2 fragments, single chain antibodies, Fv fragments, and single chain Fv fragments. The term “antibody” also refers to synthetically derived antibodies, such as phage display-derived antibodies and fragments, affybodies, nanobodies, etc.

[0095] As used herein, “marker” and “biomarker” and “feature” are used interchangeably to refer to a target molecule that indicates or is a sign of a normal or abnormal process in an individual or of a disease or other condition in an individual. More specifically, a “marker” or“biomarker” or “feature” is an anatomic, physiologic, biochemical, or molecular parameter associated with the presence of a specific physiological state or process, whether normal or abnormal, and, if abnormal, whether chronic or acute. Biomarkers are detectable and measurable by a variety of methods including laboratory assays and medical imaging. In certain aspects, a feature is an analyte / SOMAmer reagent of other predictors in a statistical model.

[0096] As used herein, “biomarker value,” “value,” “biomarker level,” “feature level,” and “level” are used interchangeably to refer to a measurement that is made using any analytical method for detecting the biomarker in a biological sample and that indicates the presence, absence, absolute amount or concentration, relative amount or concentration, titer, a level, an expression level, a ratio of measured levels, or the like, of, for, or corresponding to the biomarker in the biological sample. The exact nature of the “value” or “level” depends on the specific design and components of the particular analytical method employed to detect the biomarker.

[0097] When a biomarker indicates or is a sign of an abnormal process or a disease or other condition in an individual, that biomarker is generally described as being either over-expressed or under-expressed as compared to an expression level or value of the biomarker that indicates or is a sign of a normal process or an absence of a disease or other condition in an individual. “Up-regulation,” “up-regulated,” “over-expression,” “over-expressed,” and any variations thereof are used interchangeably to refer to a value or level of a biomarker in a biological sample that is greater than a value or level (or range of values or levels) of the biomarker that is typically detected in similar biological samples from healthy or normal individuals. The terms may also refer to a value or level of a biomarker in a biological sample that is greater than a value or level (or range of values or levels) of the biomarker that may be detected at a different stage of a particular disease.

[0098] “Down-regulation,” “down-regulated,” “under-expression,” “under-expressed,” and any variations thereof are used interchangeably to refer to a value or level of a biomarker in a biological sample that is less than a value or level (or range of values or levels) of the biomarker that is typically detected in similar biological samples from healthy or normal individuals. The terms may also refer to a value or level of a biomarker in a biological sample that is less than a value or level (or range of values or levels) of the biomarker that may be detected at a different stage of a particular disease.

[0099] Further, a biomarker that is either over-expressed or under-expressed can also be referred to as being “differentially expressed” or as having a “differential level” or “differential value” as compared to a “normal” expression level or value of the biomarker that indicates or is a sign of a normal process or an absence of a disease or other condition in an individual. Thus,“differential expression” of a biomarker can also be referred to as a variation from a “normal” expression level of the biomarker.

[0100] The term “differential gene expression” and “differential expression” are used interchangeably to refer to a gene (or its corresponding protein expression product) whose expression is activated to a higher or lower level in a subject suffering from a specific disease or condition, relative to its expression in a normal or control subject. The terms also include genes (or the corresponding protein expression products) whose expression is activated to a higher or lower level at different stages of the same disease or condition. It is also understood that a differentially expressed gene may be either activated or inhibited at the nucleic acid level or protein level, or may be subject to alternative splicing to result in a different polypeptide product. Such differences may be evidenced by a variety of changes including mRNA levels, surface expression, secretion or other partitioning of a polypeptide. Differential gene expression may include a comparison of expression between two or more genes or their gene products; or a comparison of the ratios of the expression between two or more genes or their gene products; or even a comparison of two differently processed products of the same gene, which differ between normal subjects and subjects suffering from a disease; or between various stages of the same disease. Differential expression includes both quantitative, as well as qualitative, differences in the temporal or cellular expression pattern in a gene or its expression products among, for example, normal and diseased cells, or among cells which have undergone different disease events or disease stages.

[0101] A “control level” of a target molecule refers to the level of the target molecule in a health subject of the same sample type. Control level may refer to the average level of the target molecule in samples from a population of healthy individuals.

[0102] As used herein, “individual” refers to a test subject or patient. The individual can be a mammal or a non-mammal. In various embodiments, the individual is a mammal. A mammalian individual can be a human or non-human. In various embodiments, the individual is a human.

[0103] “Diagnose,” “diagnosing,” “diagnosis,” and variations thereof refer to the detection, determination, or recognition of a health status or condition of an individual on the basis of one or more signs, symptoms, data, or other information pertaining to that individual. The health status of an individual can be diagnosed as healthy / normal (i.e., a diagnosis of the absence of a disease or condition) or diagnosed as ill / abnormal (i.e., a diagnosis of the presence, or an assessment of the characteristics, of a disease or condition). The terms “diagnose,” “diagnosing,” “diagnosis,” etc., encompass, with respect to a particular disease or condition, the initial detection of the disease; the characterization or classification of the disease; the detectionof the progression, remission, or recurrence of the disease; and the detection of disease response after the administration of a treatment or therapy to the individual.

[0104] “Prognose,” “prognosing,” “prognosis,” and variations thereof refer to the prediction of a future course of a disease or condition in an individual who has the disease or condition (e.g., predicting patient survival), and such terms encompass the evaluation of disease or condition response after the administration of a treatment or therapy to the individual.

[0105] “Evaluate,” “evaluating,” “evaluation,” and variations thereof encompass both “diagnose” and “prognose” and also encompass determinations or predictions about the future course of a disease or condition in an individual who does not have the disease as well as determinations or predictions regarding the risk that a disease or condition will recur in an individual who apparently has been cured of the disease or has had the condition resolved. The term “evaluate” also encompasses assessing an individual’s response to a therapy, such as, for example, predicting whether an individual is likely to respond favorably to a therapeutic agent or is unlikely to respond to a therapeutic agent (or will experience toxic or other undesirable side effects, for example), selecting a therapeutic agent for administration to an individual, or monitoring or determining an individual’s response to a therapy that has been administered to the individual. Thus, “evaluating” risk of stomach cancer can include, for example, predicting the future risk of stomach cancer in an individual. Evaluation of risk of stomach cancer can include embodiments such as the assessment of risk of stomach cancer as the probability (absolute risk) of a first primary malignant stomach cancer diagnosis which is a continuous variable within the range from 0.0000 to 1.0000. The evaluation of risk of stomach cancer is for a defined period; such period can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, and / or 10. In some aspects, evaluation of risk of stomach cancer is for subjects aged 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, or 65 years of age or older.

[0106] As used herein, “additional biomedical information” refers to one or more evaluations of an individual, other than using any of the biomarkers described herein, that are associated with stomach cancer. “Additional biomedical information” includes any of the following: physical descriptors of an individual, including the height and / or weight of an individual; the age of an individual; the gender of an individual; change in weight; the ethnicity of an individual; occupational history; family history of stomach cancer; the presence of a genetic marker(s); clinical symptoms such as abdominal pain, weight gain or loss gene expression values; physical descriptors of an individual, including physical descriptors observed by radiologic imaging; tobacco use status; alcohol use history; occupational history; dietary habits - salt, saturated fat and cholesterol intake; caffeine consumption; and imaging information. Additional biomedicalinformation can be obtained from an individual using routine techniques known in the art, such as from the individual themselves by use of a routine patient questionnaire or health history questionnaire, etc., or from a medical practitioner, etc.

[0107] As used herein, “detecting” or “determining” with respect to a biomarker value includes the use of both the instrument required to observe and record a signal corresponding to a biomarker value and the material / s required to generate that signal. In various embodiments, the biomarker value is detected using any suitable method, including fluorescence, chemiluminescence, surface plasmon resonance, surface acoustic waves, mass spectrometry, infrared spectroscopy, Raman spectroscopy, atomic force microscopy, scanning tunneling microscopy, electrochemical detection methods, nuclear magnetic resonance, quantum dots, and the like.

[0108] “ Solid support” refers herein to any substrate having a surface to which molecules may be attached, directly or indirectly, through either covalent or non-covalent bonds. A “solid support” can have a variety of physical formats, which can include, for example, a membrane; a chip (e.g., a protein chip); a slide (e.g., a glass slide or coverslip); a column; a hollow, solid, semi-solid, pore- or cavity- containing particle, such as, for example, a bead; a gel; a fiber, including a fiber optic material; a matrix; and a sample receptacle. Exemplary sample receptacles include sample wells, tubes, capillaries, vials, and any other vessel, groove or indentation capable of holding a sample. A sample receptacle can be contained on a multisample platform, such as a microtiter plate, slide, microfluidics device, and the like. A support can be composed of a natural or synthetic material, an organic or inorganic material. The composition of the solid support on which capture reagents are attached generally depends on the method of attachment (e.g., covalent attachment). Other exemplary receptacles include microdroplets and microfluidic controlled or bulk oil / aqueous emulsions within which assays and related manipulations can occur. Suitable solid supports include, for example, plastics, resins, polysaccharides, silica or silica-based materials, functionalized glass, modified silicon, carbon, metals, inorganic glasses, membranes, nylon, natural fibers (such as, for example, silk, wool and cotton), polymers, and the like. The material composing the solid support can include reactive groups such as, for example, carboxy, amino, or hydroxyl groups, which are used for attachment of the capture reagents. Polymeric solid supports can include, e.g., polystyrene, polyethylene glycol tetraphthalate, polyvinyl acetate, polyvinyl chloride, polyvinyl pyrrolidone, polyacrylonitrile, polymethyl methacrylate, polytetrafluoroethylene, butyl rubber, styrenebutadiene rubber, natural rubber, polyethylene, polypropylene, (poly)tetrafluoroethylene, (poly)vinylidenefluoride, polycarbonate, and polymethylpentene. Suitable solid support particlesthat can be used include, e.g., encoded particles, such as Luminex®-type encoded particles, magnetic particles, and glass particles.

[0109] As used herein, “analyte” is the protein target of a capture reagent. In certain aspects, the capture reagent is an aptamer. In certain further aspects, the capture reagent is a SOMAmer.

[0110] As used herein, “nucleic acid ligand,” “aptamer,” “SOMAmer,” “modified aptamer,” and “clone” are used interchangeably to refer to a non-naturally occurring nucleic acid that has a desirable action on a target molecule. A desirable action includes, but is not limited to, binding of the target, catalytically changing the target, reacting with the target in a way that modifies or alters the target or the functional activity of the target, covalently attaching to the target (as in a suicide inhibitor), and facilitating the reaction between the target and another molecule. In one embodiment, the action is specific binding affinity for a target molecule, such target molecule being a three dimensional chemical structure other than a polynucleotide that binds to the aptamer through a mechanism which is independent of Watson / Crick base pairing or triple helix formation, wherein the aptamer is not a nucleic acid having the known physiological function of being bound by the target molecule. Aptamers to a given target include nucleic acids that are identified from a candidate mixture of nucleic acids, where the aptamer is a ligand of the target, by a method comprising: (a) contacting the candidate mixture with the target, wherein nucleic acids having an increased affinity to the target relative to other nucleic acids in the candidate mixture can be partitioned from the remainder of the candidate mixture; (b) partitioning the increased affinity nucleic acids from the remainder of the candidate mixture; and (c) amplifying the increased affinity nucleic acids to yield a ligand-enriched mixture of nucleic acids, whereby aptamers of the target molecule are identified. It is recognized that affinity interactions are a matter of degree; however, in this context, the “specific binding affinity” of an aptamer for its target means that the aptamer binds to its target generally with a much higher degree of affinity than it binds to other, non-target, components in a mixture or sample. An “aptamer,” “SOMAmer,” or “nucleic acid ligand” is a set of copies of one type or species of nucleic acid molecule that has a particular nucleotide sequence. An aptamer can include any suitable number of nucleotides. “Aptamers” refer to more than one such set of molecules. Different aptamers can have either the same or different numbers of nucleotides. Aptamers may be DNA or RNA and may be single stranded, double stranded, or contain double stranded or triple stranded regions. In some embodiments, the aptamers are prepared using a SELEX process as described herein, or known in the art.

[0111] As used herein, “study,” refers to a set of samples and clinical data that are analyzed to derive the test.

[0112] As used herein, “training dataset,” refers to a subset of data from a study used to fit a model.

[0113] As used herein, “validation dataset,” refers to a final subset of data used to assess the performance of a selected model developed on a verification dataset.

[0114] As used herein, “verification dataset,” means a separate subset of data used to provide an unbiased evaluation of a model fit on the training dataset while tuning model parameters.

[0115] As used herein, “elastic net logistic regression” refers to a machine learning method that utilizes penalized regression techniques to select the features that best predict the endpoint while allowing correlated features to be grouped together.

[0116] As used herein, “feature” refers to an analyte / SOMAmer reagent or other predictors in a statistical model.

[0117] As used herein, “forward selection” refers to a method for feature selection and reduction. Forward selection is a form of stepwise regression that starts with zero features included in the model. In an iterative process, features are considered for addition using t-tests as the selection criterion.

[0118] As used herein, the term “Adaptive Normalization using Maximum Likelihood (ANML)” refers to a process for normalizing the analytes which allows for comparison across plates, studies, and sites.

[0119] As used herein, the term “principle component analysis” refers to a method for assessing and identifying large sources of variation in the data.

[0120] As used herein, the term “consensus-features nested cross validation (CNCV)” refers to a method for feature selection and reduction that uses regularization techniques and subsampling approaches such that candidate features are determined by relevance across multiple subsets of the data.

[0121] As used herein, the term “need” or “needed” refers to a judgement made by a health care provider regarding treatment of a patient which is considered by the health care provider to be beneficial to the health status of the patient.Stomach Cancer Risk Assessment

[0122] In some embodiments, the number of biomarkers useful for a biomarker subset or panel is based on the sensitivity and specificity value for the particular combination of biomarker levels. The terms “sensitivity” and “specificity” are used herein with respect to the ability to correctly classify an individual, based on one or more biomarker levels detected in their biological sample, as a risk of stomach cancer. “Sensitivity” indicates the performance of the biomarker(s) with respect to correctly classifying individuals that have a risk of stomachcancer or are healthy. “Specificity” indicates the performance of the biomarker(s) with respect to correctly classifying individuals who have a risk of stomach cancer.

[0123] In some embodiments, overall performance of a panel of one or more biomarkers is represented by the area-under-the-curve (AUC) value. The AUC value is derived from a receiver operating characteristic (ROC) curve. The ROC curve is the plot of the true positive rate (sensitivity) of a test against the false positive rate (1 -specificity) of the test. The term “area under the curve” or “AUC” refers to the area under the curve of a receiver operating characteristic (ROC) curve, both of which are well known in the art. AUC measures are useful for comparing the accuracy of a classifier across the complete data range. Classifiers with a greater AUC have a greater capacity to classify unknowns correctly between two groups of interest (e.g., individuals having a risk of stomach cancer or healthy individuals). ROC curves are useful for plotting the performance of a particular feature (e.g., any of the biomarkers described herein and / or any item of additional biomedical information) in distinguishing between two populations. Typically, the feature data across the entire population are sorted in ascending order based on the value of a single feature. Then, for each value for that feature, the true positive and false positive rates for the data are calculated. The true positive rate is determined by counting the number of cases above the value for that feature and then dividing by the total number of cases. The false positive rate is determined by counting the number of controls above the value for that feature and then dividing by the total number of controls.Although this definition refers to scenarios in which a feature is elevated in cases compared to controls, this definition also applies to scenarios in which a feature is lower in cases compared to the controls (in such a scenario, samples below the value for that feature would be counted). ROC curves can be generated for a single feature as well as for other single outputs, for example, a combination of two or more features can be mathematically combined (e.g., added, subtracted, multiplied, etc.) to provide a single sum value, and this single sum value can be plotted in a ROC curve. Additionally, any combination of multiple features, in which the combination derives a single output value, can be plotted in a ROC curve.Exemplary Uses of Biomarkers

[0124] In various exemplary embodiments, methods are provided for predicting an individual’s risk of stomach cancer by detecting one or more biomarker values corresponding to one or more biomarkers that are present in the circulation of an individual, such as in blood, serum or plasma, by any number of analytical methods, including any of the analytical methods described herein. These biomarkers are, for example, differentially expressed in individuals who have a risk of stomach cancer as compared to an individual who does not. Detection of thedifferential expression of a biomarker in an individual can be used, for example, to predict an individual’s risk of stomach cancer. In some embodiments, the detection of the differential expression of a biomarker in an individual can be used, for example, to predict an individual’s risk of stomach cancer within a period of time. In some embodiments, the prediction of an individual’s risk of stomach cancer is within a 5-year period.

[0125] In addition to testing biomarker levels as a stand-alone diagnostic test, biomarker levels can also be done in conjunction with determination of SNPs or other genetic lesions or variability that are indicative of increased risk of susceptibility of disease or condition. (See, e.g., Amos et al., Nature Genetics 40, 616-622 (2009)).

[0126] Any of the described biomarkers may also be used in imaging tests. For example, an imaging agent can be coupled to any of the described biomarkers, which can be used to aid in predicting an individual’s risk of stomach cancer, to monitor response to therapeutic interventions, to select for target populations in a clinical trial among other uses.Detection and Determination of Biomarkers and Biomarker Levels

[0127] A biomarker level for the biomarkers described herein can be detected using any of a variety of known analytical methods. In one embodiment, a biomarker level is detected using a capture reagent. As used herein, a “capture agent” or “capture reagent” refers to a molecule that is capable of binding specifically to a biomarker. In various embodiments, the capture reagent can be exposed to the biomarker in solution or can be exposed to the biomarker while the capture reagent is immobilized on a solid support. In other embodiments, the capture reagent contains a feature that is reactive with a secondary feature on a solid support. In these embodiments, the capture reagent can be exposed to the biomarker in solution, and then the feature on the capture reagent can be used in conjunction with the secondary feature on the solid support to immobilize the biomarker on the solid support. The capture reagent is selected based on the type of analysis to be conducted. Capture reagents include but are not limited to SOMAmers, antibodies, adnectins, ankyrins, other antibody mimetics and other protein scaffolds, autoantibodies, chimeras, small molecules, a F(ab’)2 fragment, a single chain antibody fragment, an Fv fragment, a single chain Fv fragment, a nucleic acid, a lectin, a ligand-binding receptor, affybodies, nanobodies, imprinted polymers, avimers, peptidomimetics, a hormone receptor, a cytokine receptor, and synthetic receptors, and modifications and fragments of these.

[0128] In some embodiments, a biomarker level is detected using a biomarker / capture reagent complex.

[0129] In other embodiments, the biomarker level is derived from the biomarker / capture reagent complex and is detected indirectly, such as, for example, as a result of a reaction that issubsequent to the biomarker / capture reagent interaction, but is dependent on the formation of the biomarker / capture reagent complex.

[0130] In some embodiments, the biomarker level is detected directly from the biomarker in a biological sample.

[0131] In one embodiment, the biomarkers are detected using a multiplexed format that allows for the simultaneous detection of two or more biomarkers in a biological sample. In one embodiment of the multiplexed format, capture reagents are immobilized, directly or indirectly, covalently or non-covalently, in discrete locations on a solid support. In another embodiment, a multiplexed format uses discrete solid supports where each solid support has a unique capture reagent associated with that solid support, such as, for example quantum dots. In another embodiment, an individual device is used for the detection of each one of multiple biomarkers to be detected in a biological sample. Individual devices can be configured to permit each biomarker in the biological sample to be processed simultaneously. For example, a microtiter plate can be used such that each well in the plate is used to uniquely analyze one of multiple biomarkers to be detected in a biological sample.

[0132] In one or more of the foregoing embodiments, a fluorescent tag can be used to label a component of the biomarker / capture complex to enable the detection of the biomarker value. In various embodiments, the fluorescent label can be conjugated to a capture reagent specific to any of the biomarkers described herein using known techniques, and the fluorescent label can then be used to detect the corresponding biomarker value. Suitable fluorescent labels include rare earth chelates, fluorescein and its derivatives, rhodamine and its derivatives, dansyl, allophycocyanin, PBXL-3, Qdot 605, Lissamine, phycoerythrin, Texas Red, and other such compounds.

[0133] In one embodiment, the fluorescent label is a fluorescent dye molecule. In some embodiments, the fluorescent dye molecule includes at least one substituted indolium ring system in which the substituent on the 3-carbon of the indolium ring contains a chemically reactive group or a conjugated substance. In some embodiments, the dye molecule includes an AlexFluor molecule, such as, for example, AlexaFluor 488, AlexaFluor 532, AlexaFluor 647, AlexaFluor 680, or AlexaFluor 700. In other embodiments, the dye molecule includes a first type and a second type of dye molecule, such as, e.g., two different AlexaFluor molecules. In other embodiments, the dye molecule includes a first type and a second type of dye molecule, and the two dye molecules have different emission spectra.

[0134] Fluorescence can be measured with a variety of instrumentation compatible with a wide range of assay formats. For example, spectrofluorimeters have been designed to analyze microtiter plates, microscope slides, printed arrays, cuvettes, etc. See Principles of FluorescenceSpectroscopy, by J.R. Lakowicz, Springer Science + Business Media, Inc., 2004. See Bioluminescence & Chemiluminescence: Progress & Current Applications; Philip E. Stanley and Larry J. Kricka editors, World Scientific Publishing Company, January 2002.

[0135] In one or more of the foregoing embodiments, a chemiluminescence tag can optionally be used to label a component of the biomarker / capture complex to enable the detection of a biomarker value. Suitable chemiluminescent materials include any of oxalyl chloride, Rodamin 6G, Ru(bipy)32+ , TMAE (tetrakis(dimethylamino)ethylene), Pyrogallol (1,2,3-trihydroxibenzene), Lucigenin, peroxyoxalates, Aryl oxalates, Acridinium esters, dioxetanes, and others.

[0136] In yet other embodiments, the detection method includes an enzyme / substrate combination that generates a detectable signal that corresponds to the biomarker value.Generally, the enzyme catalyzes a chemical alteration of the chromogenic substrate which can be measured using various techniques, including spectrophotometry, fluorescence, and chemiluminescence. Suitable enzymes include, for example, luciferases, luciferin, malate dehydrogenase, urease, horseradish peroxidase (HRPO), alkaline phosphatase, betagalactosidase, glucoamylase, lysozyme, glucose oxidase, galactose oxidase, and glucose-6-phosphate dehydrogenase, uricase, xanthine oxidase, lactoperoxidase, microperoxidase, and the like.

[0137] In yet other embodiments, the detection method can be a combination of fluorescence, chemiluminescence, radionuclide or enzyme / substrate combinations that generate a measurable signal. Multimodal signaling could have unique and advantageous characteristics in biomarker assay formats.

[0138] More specifically, the biomarker levels for the biomarkers described herein can be detected using known analytical methods including, singleplex SOMAmer assays, multiplexed SOMAmer assays, singleplex or multiplexed immunoassays, mRNA expression profiling, miRNA expression profiling, mass spectrometric analysis, histological / cytological methods, etc. as detailed below.Determination of Biomarker Levels using Aptamer-Based Assays

[0139] Assays directed to the detection and quantification of physiologically significant molecules in biological samples and other samples are important tools in scientific research and in the health care field. One class of such assays involves the use of a microarray that includes one or more aptamers immobilized on a solid support. The aptamers are each capable of binding to a target molecule in a highly specific manner and with very high affinity. See, e.g., U.S.Patent No. 5,475,096 entitled “Nucleic Acid Ligands”; see also, e.g., U.S. Patent No. 6,242,246,U.S. Patent No. 6,458,543, and U.S. Patent No. 6,503,715, each of which is entitled “Nucleic Acid Ligand Diagnostic Biochip”. Once the microarray is contacted with a sample, the aptamers bind to their respective target molecules present in the sample and thereby enable a determination of a biomarker value corresponding to a biomarker.

[0140] As used herein, an “aptamer” refers to a nucleic acid that has a specific binding affinity for a target molecule. It is recognized that affinity interactions are a matter of degree; however, in this context, the “specific binding affinity” of an aptamer for its target means that the aptamer binds to its target generally with a much higher degree of affinity than it binds to other components in a test sample. An “aptamer” is a set of copies of one type or species of nucleic acid molecule that has a particular nucleotide sequence. An aptamer can include any suitable number of nucleotides, including any number of chemically modified nucleotides.“Aptamers” refers to more than one such set of molecules. Different aptamers can have either the same or different numbers of nucleotides. Aptamers can be DNA or RNA or chemically modified nucleic acids and can be single stranded, double stranded, or contain double stranded regions, and can include higher ordered structures. An aptamer can also be a photoaptamer, where a photoreactive or chemically reactive functional group is included in the aptamer to allow it to be covalently linked to its corresponding target. Any of the aptamer methods disclosed herein can include the use of two or more aptamers that specifically bind the same target molecule. As further described below, an aptamer may include a tag. If an aptamer includes a tag, all copies of the aptamer need not have the same tag. Moreover, if different aptamers each include a tag, these different aptamers can have either the same tag or a different tag.

[0141] An aptamer can be identified using any known method, including the SELEX process. Once identified, an aptamer can be prepared or synthesized in accordance with any known method, including chemical synthetic methods and enzymatic synthetic methods.

[0142] As used herein, a “SOMAmer” or Slow Off-Rate Modified Aptamer refers to an aptamer having improved off-rate characteristics. SOMAmers can be generated using the improved SELEX methods described in U.S. Patent No. 7,947,447, entitled “Method for Generating Aptamers with Improved Off-Rates ” In some embodiments, a slow off-rate aptamer (including an aptamers comprising at least one nucleotide with a hydrophobic modification) has an off-rate (tU) of > 20 minutes > 30 minutes, > 60 minutes, > 90 minutes, > 120 minutes, > 150 minutes, > 180 minutes, > 210 minutes, or > 240 minutes.

[0143] The terms “SELEX” and “SELEX process” are used interchangeably herein to refer generally to a combination of (1) the selection of aptamers that interact with a target molecule in a desirable manner, for example binding with high affinity to a protein, with (2) theamplification of those selected nucleic acids. The SELEX process can be used to identify aptamers with high affinity to a specific target or biomarker.

[0144] SELEX generally includes preparing a candidate mixture of nucleic acids, binding of the candidate mixture to the desired target molecule to form an affinity complex, separating the affinity complexes from the unbound candidate nucleic acids, separating and isolating the nucleic acid from the affinity complex, purifying the nucleic acid, and identifying a specific aptamer sequence. The process may include multiple rounds to further refine the affinity of the selected aptamer. The process can include amplification steps at one or more points in the process. See, e.g., U.S. Patent No. 5,475,096, entitled “Nucleic Acid Ligands”. The SELEX process can be used to generate an aptamer that covalently binds its target as well as an aptamer that non-covalently binds its target. See, e.g., U.S. Patent No. 5,705,337 entitled “Systematic Evolution of Nucleic Acid Ligands by Exponential Enrichment: Chemi-SELEX.”

[0145] The SELEX process can be used to identify high-affinity aptamers containing modified nucleotides that confer improved characteristics on the aptamer, such as, for example, improved in vivo stability or improved delivery characteristics. Examples of such modifications include chemical substitutions at the ribose and / or phosphate and / or base positions. SELEX process-identified aptamers containing modified nucleotides are described in U.S. Patent No. 5,660,985, entitled “High Affinity Nucleic Acid Ligands Containing Modified Nucleotides”, which describes oligonucleotides containing nucleotide derivatives chemically modified at the 5’- and 2’-positions of pyrimidines. U.S. Patent No. 5,580,737, see supra, describes highly specific aptamers containing one or more nucleotides modified with 2’-amino (2’-NH2), 2’-fluoro (2’-F), and / or 2’-O-methyl (2’-0Me). See also, U.S. Patent Application Publication 20090098549, entitled “SELEX and PHOTOSELEX”, which describes nucleic acid libraries having expanded physical and chemical properties and their use in SELEX and photoSELEX.

[0146] SELEX can also be used to identify aptamers that have desirable off-rate characteristics. See U.S. Patent Application Publication 2009 / 0004667, entitled “Method for Generating Aptamers with Improved Off-Rates”, which describes improved SELEX methods for generating aptamers that can bind to target molecules. As mentioned above, these slow off-rate aptamers are known as “SOMAmers.” Methods for producing aptamers or SOMAmers and photoaptamers or SOMAmers having slower rates of dissociation from their respective target molecules are described. The methods involve contacting the candidate mixture with the target molecule, allowing the formation of nucleic acid-target complexes to occur, and performing a slow off-rate enrichment process wherein nucleic acid-target complexes with fast dissociation rates will dissociate and not reform, while complexes with slow dissociation rates will remain intact. Additionally, the methods include the use of modified nucleotides in the production ofcandidate nucleic acid mixtures to generate aptamers or SOMAmers with improved off-rate performance. Nonlimiting exemplary modified nucleotides include, for example, the modified pyrimidines shown in FIGS. 1-5.

[0147] A variation of this assay employs aptamers that include photoreactive functional groups that enable the aptamers to covalently bind or “photocrosslink” their target molecules. See, e.g., U.S. Patent No. 6,544,776 entitled “Nucleic Acid Ligand Diagnostic Biochip”. These photoreactive aptamers are also referred to as photoaptamers. See, e.g., U.S. Patent No.5,763,177, U.S. Patent No. 6,001,577, and U.S. Patent No. 6,291,184, each of which is entitled “Systematic Evolution of Nucleic Acid Ligands by Exponential Enrichment: Photoselection of Nucleic Acid Ligands and Solution SELEX”; see also, e.g., U.S. Patent No. 6,458,539, entitled “Photoselection of Nucleic Acid Ligands”. After the microarray is contacted with the sample and the photoaptamers have had an opportunity to bind to their target molecules, the photoaptamers are photoactivated, and the solid support is washed to remove any non-specifically bound molecules. Harsh wash conditions may be used, since target molecules that are bound to the photoaptamers are generally not removed, due to the covalent bonds created by the photoactivated functional group(s) on the photoaptamers. In this manner, the assay enables the detection of a biomarker value corresponding to a biomarker in the test sample.

[0148] In both of these assay formats, the aptamers or SOMAmers are immobilized on the solid support prior to being contacted with the sample. Under certain circumstances, however, immobilization of the aptamers or SOMAmers prior to contact with the sample may not provide an optimal assay. For example, pre-immobilization of the aptamers or SOMAmers may result in inefficient mixing of the aptamers or SOMAmers with the target molecules on the surface of the solid support, perhaps leading to lengthy reaction times and, therefore, extended incubation periods to permit efficient binding of the aptamers or SOMAmers to their target molecules. Further, when photoaptamers or photoSOMAmers are employed in the assay and depending upon the material utilized as a solid support, the solid support may tend to scatter or absorb the light used to effect the formation of covalent bonds between the photoaptamers or photoSOMAmers and their target molecules. Moreover, depending upon the method employed, detection of target molecules bound to their aptamers or photoSOMAmers can be subject to imprecision, since the surface of the solid support may also be exposed to and affected by any labeling agents that are used. Finally, immobilization of the aptamers or SOMAmers on the solid support generally involves an aptamer or SOMAmer-preparation step (i.e., the immobilization) prior to exposure of the aptamers or SOMAmers to the sample, and this preparation step may affect the activity or functionality of the aptamers or SOMAmers.

[0149] SOMAmer assays that permit a SOMAmer to capture its target in solution and then employ separation steps that are designed to remove specific components of the SOMAmer-target mixture prior to detection have also been described (see U.S. Patent Application Publication 20090042206, entitled “Multiplexed Analyses of Test Samples”). The described SOMAmer assay methods enable the detection and quantification of a non-nucleic acid target (e.g., a protein target) in a test sample by detecting and quantifying a nucleic acid (i.e., a SOMAmer). The described methods create a nucleic acid surrogate (i.e, the SOMAmer) for detecting and quantifying a non-nucleic acid target, thus allowing the wide variety of nucleic acid technologies, including amplification, to be applied to a broader range of desired targets, including protein targets.

[0150] SOMAmers can be constructed to facilitate the separation of the assay components from a SOMAmer biomarker complex (or photoSOMAmer biomarker covalent complex) and permit isolation of the SOMAmer for detection and / or quantification. In some embodiments, these constructs can include a cleavable or releasable element within the SOMAmer sequence. In other embodiments, additional functionality can be introduced into the SOMAmer, for example, a labeled or detectable component, a spacer component, or a specific binding tag or immobilization element. For example, the SOMAmer can include a tag connected to the SOMAmer via a cleavable moiety, a label, a spacer component separating the label, and the cleavable moiety. In one embodiment, a cleavable element is a photocleavable linker. The photocleavable linker can be attached to a biotin moiety and a spacer section, can include an NHS group for derivatization of amines, and can be used to introduce a biotin group to an aptamer, thereby allowing for the release of the aptamer later in an assay method.

[0151] Homogenous assays, done with all assay components in solution, do not require separation of sample and reagents prior to the detection of signal. These methods are rapid and easy to use. These methods generate signal based on a molecular capture or binding reagent that reacts with its specific target. For predicting an individual’s risk of stomach cancer, the molecular capture reagents would be an aptamer or an antibody or the like and the specific target would be one or more of the biomarkers in Table 1.

[0152] In some embodiments, a method for signal generation takes advantage of anisotropy signal change due to the interaction of a fluorophore-labeled capture reagent with its specific biomarker target. When the labeled capture reagent reacts with its target, the increased molecular weight causes the rotational motion of the fluorophore attached to the complex to become much slower changing the anisotropy value. By monitoring the anisotropy change, binding events may be used to quantitatively measure the biomarkers in solutions. Other methods include fluorescence polarization assays, molecular beacon methods, time resolvedfluorescence quenching, chemiluminescence, fluorescence resonance energy transfer, and the like.

[0153] An exemplary solution-based aptamer assay that can be used to detect a biomarker value corresponding to a biomarker in a biological sample includes the following: (a) preparing a mixture by contacting the biological sample with an aptamer that includes a first tag and has a specific affinity for the biomarker, wherein an aptamer affinity complex is formed when the biomarker is present in the sample; (b) exposing the mixture to a first solid support including a first capture element, and allowing the first tag to associate with the first capture element; (c) removing any components of the mixture not associated with the first solid support; (d) attaching a second tag to the biomarker component of the aptamer affinity complex; (e) releasing the aptamer affinity complex from the first solid support; (f) exposing the released aptamer affinity complex to a second solid support that includes a second capture element and allowing the second tag to associate with the second capture element; (g) removing any noncomplexed aptamer from the mixture by partitioning the non-complexed aptamer from the aptamer affinity complex; (h) eluting the aptamer from the solid support; and (i) detecting the biomarker by detecting the aptamer component of the aptamer affinity complex.

[0154] Any means known in the art can be used to detect a biomarker value by detecting the aptamer component of an aptamer affinity complex. A number of different detection methods can be used to detect the aptamer component of an affinity complex, such as, for example, hybridization assays, mass spectroscopy, or QPCR. In some embodiments, nucleic acid sequencing methods can be used to detect the aptamer component of an aptamer affinity complex and thereby detect a biomarker value. Briefly, a test sample can be subjected to any kind of nucleic acid sequencing method to identify and quantify the sequence or sequences of one or more aptamers present in the test sample. In some embodiments, the sequence includes the entire aptamer molecule or any portion of the molecule that may be used to uniquely identify the molecule. In other embodiments, the identifying sequencing is a specific sequence added to the aptamer; such sequences are often referred to as “tags,” “barcodes,” or “zipcodes.” In some embodiments, the sequencing method includes enzymatic steps to amplify the aptamer sequence or to convert any kind of nucleic acid, including RNA and DNA that contain chemical modifications to any position, to any other kind of nucleic acid appropriate for sequencing.

[0155] In some embodiments, the sequencing method includes one or more cloning steps. In other embodiments the sequencing method includes a direct sequencing method without cloning.

[0156] In some embodiments, the sequencing method includes a directed approach with specific primers that target one or more aptamer in the test sample. In other embodiments, the sequencing method includes a shotgun approach that targets all aptamer in the test sample.

[0157] In some embodiments, the sequencing method includes enzymatic steps to amplify the molecule targeted for sequencing. In other embodiments, the sequencing method directly sequences single molecules. An exemplary nucleic acid sequencing-based method that can be used to detect a biomarker value corresponding to a biomarker in a biological sample includes the following: (a) converting a mixture of aptamers that contain chemically modified nucleotides to unmodified nucleic acids with an enzymatic step; (b) shotgun sequencing the resulting unmodified nucleic acids with a massively parallel sequencing platform such as, for example, the 454 Sequencing System (454 Life Sciences / Roche), the Illumina Sequencing System (Illumina), the ABI SOLiD Sequencing System (Applied Biosystems), the Heli Scope Single Molecule Sequencer (Helicos Biosciences), or the Pacific Biosciences Real Time SingleMolecule Sequencing System (Pacific BioSciences) or the Polonator G Sequencing System (Dover Systems); and (c) identifying and quantifying the SOMAmers present in the mixture by specific sequence and sequence count.Determination of Biomarker Values using Immunoassays

[0158] Immunoassay methods are based on the reaction of an antibody to its corresponding target or analyte and can detect the analyte in a sample depending on the specific assay format. To improve specificity and sensitivity of an assay method based on immuno-reactivity, monoclonal antibodies are often used because of their specific epitope recognition. Polyclonal antibodies have also been successfully used in various immunoassays because of their increased affinity for the target as compared to monoclonal antibodies. Immunoassays have been designed for use with a wide range of biological sample matrices. Immunoassay formats have been designed to provide qualitative, semi-quantitative, and quantitative results.

[0159] Quantitative results are generated through the use of a standard curve created with known concentrations of the specific analyte to be detected. The response or signal from an unknown sample is plotted onto the standard curve, and a quantity or value corresponding to the target in the unknown sample is established.

[0160] Numerous immunoassay formats have been designed. ELISA or EIA can be quantitative for the detection of an analyte. This method relies on attachment of a label to either the analyte or the antibody and the label component includes, either directly or indirectly, an enzyme. ELISA tests may be formatted for direct, indirect, competitive, or sandwich detection of the analyte. Other methods rely on labels such as, for example, radioisotopes (1125) or fluorescence. Additional techniques include, for example, agglutination, nephelometry, turbidimetry, Western blot, immunoprecipitation, immunocytochemistry,immunohistochemistry, flow cytometry, Luminex assay, and others (see ImmunoAssay: A Practical Guide, edited by Brian Law, published by Taylor & Francis, Ltd., 2005 edition).

[0161] Exemplary assay formats include enzyme-linked immunosorbent assay (ELISA), radioimmunoassay, fluorescent, chemiluminescence, and fluorescence resonance energy transfer (FRET) or time resolved-FRET (TR-FRET) immunoassays. Examples of procedures for detecting biomarkers include biomarker immunoprecipitation followed by quantitative methods that allow size and peptide level discrimination, such as gel electrophoresis, capillary electrophoresis, planar electrochromatography, and the like.

[0162] Methods of detecting and / or quantifying a detectable label or signal generating material depend on the nature of the label. The products of reactions catalyzed by appropriate enzymes (where the detectable label is an enzyme; see above) can be, without limitation, fluorescent, luminescent, or radioactive or they may absorb visible or ultraviolet light. Examples of detectors suitable for detecting such detectable labels include, without limitation, x-ray film, radioactivity counters, scintillation counters, spectrophotometers, colorimeters, fluorometers, luminometers, and densitometers.

[0163] Any of the methods for detection can be performed in any format that allows for any suitable preparation, processing, and analysis of the reactions. This can be, for example, in multi-well assay plates (e.g., 96 wells or 384 wells) or using any suitable array or microarray. Stock solutions for various agents can be made manually or robotically, and all subsequent pipetting, diluting, mixing, distribution, washing, incubating, sample readout, data collection and analysis can be done robotically using commercially available analysis software, robotics, and detection instrumentation capable of detecting a detectable label.Determination of Biomarker Values using Gene Expression Profiling

[0164] Measuring mRNA in a biological sample may be used as a surrogate for detection of the level of the corresponding protein in the biological sample. Thus, any of the biomarkers or biomarker panels described herein can also be detected by detecting the appropriate RNA.

[0165] mRNA expression levels are measured by reverse transcription quantitative polymerase chain reaction (RT-PCR followed with qPCR). RT-PCR is used to create a cDNA from the mRNA. The cDNA may be used in a qPCR assay to produce fluorescence as the DNA amplification process progresses. By comparison to a standard curve, qPCR can produce an absolute measurement such as number of copies of mRNA per cell. Northern blots, microarrays, Invader assays, and RT-PCR combined with capillary electrophoresis have all been used to measure expression levels of mRNA in a sample. See Gene Expression Profiling: Methods and Protocols, Richard A. Shimkets, editor, Humana Press, 2004.

[0166] miRNA molecules are small RNAs that are non-coding but may regulate gene expression. Any of the methods suited to the measurement of mRNA expression levels can also be used for the corresponding miRNA. Recently many laboratories have investigated the use of miRNAs as biomarkers for disease. Many diseases involve wide-spread transcriptional regulation, and it is not surprising that miRNAs might find a role as biomarkers. The connection between miRNA concentrations and disease is often even less clear than the connections between protein levels and disease, yet the value of miRNA biomarkers might be substantial. Of course, as with any RNA expressed differentially during disease, the problems facing the development of an in vitro diagnostic product will include the requirement that the miRNAs survive in the diseased cell and are easily extracted for analysis, or that the miRNAs are released into blood or other matrices where they must survive long enough to be measured. Protein biomarkers have similar requirements, although many potential protein biomarkers are secreted intentionally at the site of pathology and function, during disease, in a paracrine fashion. Many potential protein biomarkers are designed to function outside the cells within which those proteins are synthesized.Detection of Biomarkers Using In Vivo Molecular Imaging Technologies

[0167] Any of the described biomarkers (see, e.g., Table 1) may also be used in molecular imaging tests. For example, an imaging agent can be coupled to any of the described biomarkers, which can be used to aid in assessing an individual’s risk of stomach cancer, to monitor response to therapeutic interventions, to select a population for clinical trials among other uses.

[0168] In vivo imaging technologies provide non-invasive methods for determining the state of a particular disease or condition in the body of an individual. For example, entire portions of the body, or even the entire body, may be viewed as a three-dimensional image, thereby providing valuable information concerning morphology and structures in the body. Such technologies may be combined with the detection of the biomarkers described herein to provide information concerning predicting an individual’s risk of stomach cancer.

[0169] The use of in vivo molecular imaging technologies is expanding due to various advances in technology. These advances include the development of new contrast agents or labels, such as radiolabels and / or fluorescent labels, which can provide strong signals within the body; and the development of powerful new imaging technology, which can detect and analyze these signals from outside the body, with sufficient sensitivity and accuracy to provide useful information. The contrast agent can be visualized in an appropriate imaging system, thereby providing an image of the portion or portions of the body in which the contrast agent is located.The contrast agent may be bound to or associated with a capture reagent, such as an aptamer or an antibody, for example, and / or with a peptide or protein, or an oligonucleotide (for example, for the detection of gene expression), or a complex containing any of these with one or more macromolecules and / or other particulate forms.

[0170] The contrast agent may also feature a radioactive atom that is useful in imaging. Suitable radioactive atoms include technetium-99m or iodine- 123 for scintigraphic studies. Other readily detectable moieties include, for example, spin labels for magnetic resonance imaging (MRI) such as, for example, iodine-123 again, iodine-131, indium-ill, fluorine-19, carbon-13, nitrogen-15, oxygen-17, gadolinium, manganese or iron. Such labels are well known in the art and could easily be selected by one of ordinary skill in the art.

[0171] Standard imaging techniques include but are not limited to magnetic resonance imaging, computed tomography scanning (coronary calcium score), positron emission tomography (PET), single photon emission computed tomography (SPECT), computed tomography angiography, and the like. For diagnostic in vivo imaging, the type of detection instrument available is a major factor in selecting a given contrast agent, such as a given radionuclide and the particular biomarker that it is used to target (protein, mRNA, and the like). The radionuclide chosen typically has a type of decay that is detectable by a given type of instrument. Also, when selecting a radionuclide for in vivo diagnosis, its half-life should be long enough to enable detection at the time of maximum uptake by the target tissue but short enough that deleterious radiation of the host is minimized.

[0172] Exemplary imaging techniques include but are not limited to PET and SPECT, which are imaging techniques in which a radionuclide is synthetically or locally administered to an individual. The subsequent uptake of the radiotracer is measured over time and used to obtain information about the targeted tissue and the biomarker. Because of the high-energy (gammaray) emissions of the specific isotopes employed and the sensitivity and sophistication of the instruments used to detect them, the two-dimensional distribution of radioactivity may be inferred from outside of the body.

[0173] Commonly used positron-emitting nuclides in PET include, for example, carbon-11, nitrogen-13, oxygen-15, and fluorine-18. Isotopes that decay by electron capture and / or gammaemission are used in SPECT and include, for example iodine-123 and technetium-99m. An exemplary method for labeling amino acids with technetium-99m is the reduction of pertechnetate ion in the presence of a chelating precursor to form the labile technetium-99m-precursor complex, which, in turn, reacts with the metal binding group of a bifunctionally modified chemotactic peptide to form a technetium-99m-chemotactic peptide conjugate.

[0174] Antibodies are frequently used for such in vivo imaging diagnostic methods. The preparation and use of antibodies for in vivo diagnosis is well known in the art. Labeled antibodies which specifically bind any of the biomarkers in Table 1 can be injected into an individual, detectable according to the particular biomarker used, for the purpose of diagnosing or evaluating the disease status or condition of the individual. The label used will be selected in accordance with the imaging modality to be used, as previously described. Localization of the label permits determination of the tissue damage or other indications related to an individual’s risk of stomach cancer. The amount of label within an organ or tissue also allows determination of the involvement of the biomarkers predicting an individual’s risk of stomach cancer.

[0175] Similarly, aptamers may be used for such in vivo imaging diagnostic methods. For example, an aptamer that was used to identify a particular biomarker described in Table 1 (and therefore binds specifically to that particular biomarker) may be appropriately labeled and injected into an individual being evaluated for determination of an individual’s risk of stomach cancer, detectable according to the particular biomarker, for the purpose of diagnosing or evaluating the levels of tissue damage, components of inflammatory response, and other factors associated with the risk of stomach cancer in the individual. The label used will be selected in accordance with the imaging modality to be used, as previously described. Localization of the label permits determination of the site of the processes leading to increased risk. The amount of label within an organ or tissue also allows determination of the infiltration of the pathological process in that organ or tissue. Aptamer-directed imaging agents could have unique and advantageous characteristics relating to tissue penetration, tissue distribution, kinetics, elimination, potency, and selectivity as compared to other imaging agents.

[0176] Such techniques may also optionally be performed with labeled oligonucleotides, for example, for detection of gene expression through imaging with antisense oligonucleotides. These methods are used for in situ hybridization, for example, with fluorescent molecules or radionuclides as the label. Other methods for detection of gene expression include, for example, detection of the activity of a reporter gene.

[0177] Another general type of imaging technology is optical imaging, in which fluorescent signals within the subject are detected by an optical device that is external to the subject. These signals may be due to actual fluorescence and / or to bioluminescence. Improvements in the sensitivity of optical detection devices have increased the usefulness of optical imaging for in vivo diagnostic assays.

[0178] The use of in vivo molecular biomarker imaging is increasing, including for clinical trials, for example, to more rapidly measure clinical efficacy in trials for new disease or condition therapies and / or to avoid prolonged treatment with a placebo for those diseases, suchas multiple sclerosis, in which such prolonged treatment may be considered to be ethically questionable. For a review of other techniques, see N. Blow, Nature Methods, 6, 465-469, 2009.Determination of Biomarker Values using Mass Spectrometry Methods

[0179] A variety of configurations of mass spectrometers can be used to detect biomarker values. Several types of mass spectrometers are available or can be produced with various configurations. In general, a mass spectrometer has the following major components: a sample inlet, an ion source, a mass analyzer, a detector, a vacuum system, and instrument-control system, and a data system. Difference in the sample inlet, ion source, and mass analyzer generally define the type of instrument and its capabilities. For example, an inlet can be a capillary-column liquid chromatography source or can be a direct probe or stage such as used in matrix-assisted laser desorption. Common ion sources are, for example, electrospray, including nanospray and microspray or matrix-assisted laser desorption. Common mass analyzers include a quadrupole mass filter, ion trap mass analyzer and time-of-flight mass analyzer. Additional mass spectrometry methods are well known in the art (see Burlingame et al. Anal. Chem. 70:647 R-716R (1998); Kinter and Sherman, New York (2000)).

[0180] Protein biomarkers and biomarker values can be detected and measured by any of the following: electrospray ionization mass spectrometry (ESI-MS), ESI-MS / MS, ESI-MS / (MS)n, matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF-MS), surface-enhanced laser desorption / ionization time-of-flight mass spectrometry (SELDI-TOF-MS), desorption / ionization on silicon (DIOS), secondary ion mass spectrometry (SIMS), quadrupole time-of-flight (Q-TOF), tandem time-of-flight (TOF / TOF) technology, called ultraflex III TOF / TOF, atmospheric pressure chemical ionization mass spectrometry (APCI-MS), APCI-MS / MS, APCI-(MS)N, atmospheric pressure photoionization mass spectrometry (APPI-MS), APPI-MS / MS, and APPI-(MS)N, quadrupole mass spectrometry, Fourier transform mass spectrometry (FTMS), quantitative mass spectrometry, and ion trap mass spectrometry.

[0181] Sample preparation strategies are used to label and enrich samples before mass spectroscopic characterization of protein biomarkers and determination biomarker values.Labeling methods include but are not limited to isobaric tag for relative and absolute quantitation (iTRAQ) and stable isotope labeling with amino acids in cell culture (SILAC). Capture reagents used to selectively enrich samples for candidate biomarker proteins prior to mass spectroscopic analysis include but are not limited to aptamers, antibodies, nucleic acid probes, chimeras, small molecules, an F(ab’)2 fragment, a single chain antibody fragment, an Fv fragment, a single chain Fv fragment, a nucleic acid, a lectin, a ligand-binding receptor, affybodies, nanobodies, ankyrins, domain antibodies, alternative antibody scaffolds (e.g.diabodies etc) imprinted polymers, avimers, peptidomimetics, peptoids, peptide nucleic acids, threose nucleic acid, a hormone receptor, a cytokine receptor, and synthetic receptors, and modifications and fragments of these.Determination of Biomarker Values using a Proximity Ligation Assay

[0182] A proximity ligation assay can be used to determine biomarker values. Briefly, a test sample is contacted with a pair of affinity probes that may be a pair of antibodies or a pair of aptamers, with each member of the pair extended with an oligonucleotide. The targets for the pair of affinity probes may be two distinct determinates on one protein or one determinate on each of two different proteins, which may exist as homo- or hetero-multimeric complexes. When probes bind to the target determinates, the free ends of the oligonucleotide extensions are brought into sufficiently close proximity to hybridize together. The hybridization of the oligonucleotide extensions is facilitated by a common connector oligonucleotide which serves to bridge together the oligonucleotide extensions when they are positioned in sufficient proximity. Once the oligonucleotide extensions of the probes are hybridized, the ends of the extensions are joined together by enzymatic DNA ligation.

[0183] Each oligonucleotide extension comprises a primer site for PCR amplification. Once the oligonucleotide extensions are ligated together, the oligonucleotides form a continuous DNA sequence which, through PCR amplification, reveals information regarding the identity and amount of the target protein, as well as information regarding protein-protein interactions where the target determinates are on two different proteins. Proximity ligation can provide a highly sensitive and specific assay for real-time protein concentration and interaction information through use of real-time PCR. Probes that do not bind the determinates of interest do not have the corresponding oligonucleotide extensions brought into proximity and no ligation or PCR amplification can proceed, resulting in no signal being produced.

[0184] The foregoing assays enable the detection of biomarker levels that are useful in methods for predicting the risk of stomach cancer in an individual, where the methods comprise detecting, in a biological sample from an individual, biomarker levels that each correspond to a biomarker selected from the group of the biomarkers provided in Table 1, wherein a classification, as described in detail below, using the biomarker levels indicates whether the individual has a risk of stomach cancer. While certain of the described biomarkers are useful alone for determining a risk of stomach in an individual, methods are also described herein for the grouping of multiple subsets of the biomarkers that are each useful as a panel of two or more biomarkers. In accordance with any of the methods described herein, biomarker levels can bedetected and classified individually or they can be detected and classified collectively, as for example in a multiplex assay format.Classification of Biomarkers and Calculation of Disease Scores

[0185] In some embodiments, biomarker “signature” for a given diagnostic or predictive test contains a set of markers, each marker having different levels in the populations of interest. Different levels, in this context, may refer to different means of the marker levels for the individuals in two or more groups, or different variances in the two or more groups, or a combination of both. For the simplest form of a diagnostic test, these markers can be used to assign an unknown sample from an individual into one of two groups, such as having or not having a risk of stomach cancer. The assignment of a sample into one of two or more groups is known as classification, and the procedure used to accomplish this assignment is known as a classifier or a classification method. Classification methods may also be referred to as scoring methods. There are many classification methods that can be used to construct a diagnostic classifier from a set of biomarker values. In general, classification methods are most easily performed using supervised learning techniques where a data set is collected using samples obtained from individuals within two (or more, for multiple classification states) distinct groups one wishes to distinguish. Since the class (group or population) to which each sample belongs is known in advance for each sample, the classification method can be trained to give the desired classification response. It is also possible to use unsupervised learning techniques to produce a diagnostic classifier.

[0186] Common approaches for developing diagnostic classifiers include decision trees; bagging, boosting, forests and random forests; rule inference based learning; Parzen Windows; linear models; logistic; neural network methods; unsupervised clustering; K-means; hierarchical ascending / descending; semi-supervised learning; prototype methods; nearest neighbor; kernel density estimation; support vector machines; hidden Markov models; Boltzmann Learning; and classifiers may be combined either simply or in ways which minimize particular objective functions. For a review, see, e.g., Pattern Classification, R.O. Duda, et al., editors, John Wiley & Sons, 2nd edition, 2001; see also, The Elements of Statistical Learning - Data Mining, Inference, and Prediction, T. Hastie, et al., editors, Springer Science+Business Media, LLC, 2nd edition, 2009; each of which is incorporated by reference in its entirety.

[0187] To produce a classifier using supervised learning techniques, a set of samples called training data are obtained. In the context of diagnostic tests, training data includes samples from the distinct groups (classes) to which unknown samples will later be assigned. For example, samples collected from individuals in a control population and individuals in a particulardisease, condition or event population, such as individuals having a risk of stomach cancer, can constitute training data to develop a classifier that can classify unknown samples (or, more particularly, the individuals from whom the samples were obtained) as either having a risk of stomach cancer or healthy. The development of the classifier from the training data is known as training the classifier. Specific details on classifier training depend on the nature of the supervised learning technique (see, e.g., Pattern Classification, R.O. Duda, et al., editors, John Wiley & Sons, 2nd edition, 2001; see also, The Elements of Statistical Learning - Data Mining, Inference, and Prediction, T. Hastie, et al., editors, Springer Science+Business Media, LLC, 2nd edition, 2009).

[0188] Since typically there are many more potential biomarker values than samples in a training set, care must be used to avoid over-fitting. Over-fitting occurs when a statistical model describes random error or noise instead of the underlying relationship. Over-fitting can be avoided in a variety of ways, including, for example, by limiting the number of markers used in developing the classifier, by assuming that the marker responses are independent of one another, by limiting the complexity of the underlying statistical model employed, and by ensuring that the underlying statistical model conforms to the data.

[0189] In order to identify a set of biomarkers associated with occurrence of events, the combined set of control and early event samples were analyzed using Principal Component Analysis (PCA). PCA displays the samples with respect to the axes defined by the strongest variations between all the samples, without regard to the case or control outcome, thus mitigating the risk of overfitting the distinction between case and control. For example, since the occurrence of serious thrombotic events has a strong component of chance involved, requiring unstable plaque to rupture in vital vessels to be reported, one would not expect to see a clear separation between the control and event sample sets. While the observed separation between case and control is not large, it occurs on the second principal component, corresponding to around 10% of the total variation in this set of samples, which indicates that the underlying biological variation is relatively simple to quantify.

[0190] In the next set of analyses, biomarkers can be analyzed for those components of difference between samples which were specific to the separation between the control samples and early event samples. One method that may be employed is the use of DSGA (Bair,E. and Tibshirani,R. (2004) Semi-supervised methods to predict patient survival from gene expression data. PLOS Biol., 2, 511-522) to remove (deflate) the first three principal component directions of variation between the samples in the control set. Although the dimensionality reduction is performed on the control set to discover, both the samples in the control and the samples fromthe early event samples are run through the PC A. Separation of cases from early events can be observed along the horizontal axis.Cross Validated Selection of Proteins Relevant to the Prediction of Risk of Stomach Cancer

[0191] In order to avoid over-fitting of protein predictive power to idiosyncratic features of a particular selection of samples, a cross-validation and dimensional reduction approach can be taken. Cross-validation involves the multiple selection of sets of samples to determine the association of risk by protein combined with the use of the unselected samples to monitor the ability of the method to apply to samples which were not used in producing the model of risk (The Elements of Statistical Learning - Data Mining, Inference, and Prediction, T. Hastie, et al., editors, Springer Science+Business Media, LLC, 2nd edition, 2009). We applied the supervised PCA method of Tibshirani et al (Bair,E. and Tibshirani,R. (2004) Semi-supervised methods to predict patient survival from gene expression data. PLOS Biol., 2, 511-522.) which is applicable to high dimensional datasets in the modeling of the prediction of an individual’s risk of stomach cancer. The supervised PCA (SPCA) method involves the univariate selection of a set of proteins statistically associated with the observed event hazard in the data and the determination of the correlated component which combines information from all of these proteins. This determination of the correlated component is a dimensionality reduction step which not only combines information across proteins, but also mitigates the likelihood of overfitting by reducing the number of independent variables from the full protein menu of over 1000 proteins down to a few principal components (in this work, we only examined the first principal component).Univariate analysis and multivariate analysis of the relationship of individual proteins to time to event

[0192] The Cox proportional hazard model (Cox, David R (1972). "Regression Models and Life-Tables". Journal of the Royal Statistical Society. Series B (Methodological) 34 (2): 187— 220.)) is widely used in medical statistics. Cox regression avoids fitting a specific function of time to the cumulative survival, and instead employs a model of relative risk referred to a baseline hazard function (which may vary with time). The baseline hazard function describes the common shape of the survival time distribution for all individuals, while the relative risk gives the level of the hazard for a set of covariate values (such as a single individual or group), as a multiple of the baseline hazard. The relative risk is constant with time in the Cox model.

[0193] Accelerated failure time (AFT) models are a sub-class of survival models. Survival models predict time-to-event data under partial information. For example, in the data for the stomach cancer risk model, the event is stomach cancer diagnosis, but time-to-diagnosis eventdata is available for a fraction of the subjects in the study. For the rest of the subjects, the available information is that the subjects were not diagnosed with stomach cancer from the time of the blood draw up to the end of the study. This second category is partial information, called “censoring”, because it is uncertain if or when they would ever be diagnosed with stomach cancer.

[0194] Because survival models account for censoring, they can still use the data from those censored subjects, where other longitudinal models trying to predict when an event occurs can only use the information from subjects with stomach cancer diagnoses. And because survival models take into account time-to-event, they can produce predicted probabilities of the event occurring within any time frame, which is different from most classification models (logistic regression, random forest).

[0195] AFT survival models in particular are a regression model which specifies / assumes a linear relationship between the model’s covariates and log(time-to-event). So, a subject with 2x higher covariates (protein RFU counts) than baseline may be predicted to “survive” a stomach cancer diagnosis 2x longer than baseline.

[0196] The two most common survival models are AFT models and proportional hazards models, and an AFT Weibull model is both. The definition of a proportional hazards model is a little more complicated than that of an AFT model - in a proportional hazards model, a subject with 2x higher covariates than baseline may have a 2x higher hazard at any time point, where hazard is the negative derivative of the survival curve over time.

[0197] Other common proportional hazards models are exponential and Cox models.Exponential models are a sub-type of Weibull model. Cox models are more limited in use -predicted probabilities of time-to-event are not available from Cox models, only relative risk. AFT models can give both absolute and relative risk.Kits

[0198] Any combination of the biomarkers of Table 1 can be detected using a suitable kit, such as for use in performing the methods disclosed herein. Furthermore, any kit can contain one or more detectable labels as described herein, such as a fluorescent moiety, etc.

[0199] In one embodiment, a kit includes (a) one or more capture reagents (such as, for example, at least one aptamer or antibody) for detecting one or more biomarkers in a biological sample, wherein the biomarkers include any of the biomarkers set forth in Table 1 and optionally (b) one or more software or computer program products for classifying the individual from whom the biological sample was obtained as either having or not having a risk of stomach cancer, as further described herein. Alternatively, rather than one or more computer programproducts, one or more instructions for manually performing the above steps by a human can be provided.

[0200] The combination of a solid support with a corresponding capture reagent having a signal generating material is referred to herein as a “detection device” or “kit”. The kit can also include instructions for using the devices and reagents, handling the sample, and analyzing the data. Further the kit may be used with a computer system or software to analyze and report the result of the analysis of the biological sample.

[0201] The kits can also contain one or more reagents (e.g., solubilization buffers, detergents, washes, or buffers) for processing a biological sample. Any of the kits described herein can also include, e.g., buffers, blocking agents, mass spectrometry matrix materials, antibody capture agents, positive control samples, negative control samples, software and information such as protocols, guidance and reference data.

[0202] In one aspect, the invention provides kits for the assessment of an individual’s risk of stomach cancer. The kits include PCR primers for one or more aptamers specific to biomarkers selected from Table 1. The kit may further include instructions for use and correlation of the biomarkers with prediction of risk of stomach cancer in an individual. The kit may also include a DNA array containing the complement of one or more of the aptamers specific for the biomarkers selected from Table 1, reagents, and / or enzymes for amplifying or isolating sample DNA. The kits may include reagents for real-time PCR, for example, TaqMan probes and / or primers, and enzymes.

[0203] For example, a kit can comprise (a) reagents comprising at least capture reagent for quantifying one or more biomarkers in a test sample, wherein said biomarkers comprise the biomarkers set forth in Table 1, or any other biomarkers or biomarkers panels described herein or elsewhere, and optionally (b) one or more algorithms or computer programs for performing the steps of comparing the amount of each biomarker quantified in the test sample to one or more predetermined cutoffs and assigning a score for each biomarker quantified based on said comparison, combining the assigned scores for each biomarker quantified to obtain a total score, comparing the total score with a predetermined score, and using said comparison to determine the risk of stomach cancer in an individual. Alternatively, rather than one or more algorithms or computer programs, one or more instructions for manually performing the above steps by a human can be provided.Biomarker Panels

[0204] In some embodiments, one or more of the biomarkers listed in Table 1 are detected. In some embodiments, one, two, three, four, five, six, seven, or eight of the biomarkers listed in Table 1 are detected. In some embodiments, all of the biomarkers listed in Table 1 are detected.In some embodiments, the level of each protein listed in Table 1 is detected. In some embodiments, the detecting of the one or more biomarkers or all of the biomarkers is performed in order to determine an individual’s risk of stomach cancer. In some embodiments, the detecting of the one or more biomarkers or all of the biomarkers is performed in order to determine the individual’s risk of stomach cancer within a defined time period. In some such embodiments, the defined time period is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 years. In some such embodiments, the defined time period is 5 years.Table 1. Analytes included in the selected stomach cancer risk prediction modelUniprot ID / GenBank Entrez Gene Protein Target Protein Target Full NameID SymbolQ5VWZ2 (2007-01-23 LYPLAL1 LYPL1 Lysophospholipase-like protein 1v3)Z AAQ17077.1Q8IUL8 (2005-07-05 CILP2 CILP2 Cartilage intermediate layer protein 2 v2) / AAN17826.1P04155 (1986-11-01 TFF1 TFF1 Trefoil factor 1vl) / CAA25155.1014975 (2010-10-05 SLC27A2 S27A2 Very long-chain acyl-CoA synthetase v2)Z BAA23644.1Q06141 (1994-02-01 REG3A PAP1 Regenerating islet-derived protein 3 -alpha vl) / BAA02728.1Q86XP6 (2005-03-15 GKN2 GKN2 Gastrokine-2v2)Z AAO85515.2P09529 (1996-10-01 INHBB Inhibin bB chain Inhibin beta B chainv2)Z AAA59451.1P21709 (2011-01-11 EPHA1 EphAl Ephrin type-A receptor 1v4)Z AAA36747.1

[0205] Protein sequences are are incorporated by reference from the Uniprot / GenBank ID’s listed in the Tables herein, as one skilled in the art is able to find such sequences.

[0206] In some embodiments, at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or 8 of the biomarker proteins selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl are detected. In some embodiments, LYPL1 is detected. In some embodiments, LYPL1 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl are detected. In some embodiments, CILP2 is detected. In some embodiments, CILP2 and at least at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl are detected. In some embodiments, TFF1 is detected. In some embodiments, TFF1 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl are detected. In some embodiments, S27A2 is detected. In some embodiments, S27A2 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2,TFF1, PAP1, GKN2, Inhibin bB chain, and EphAl are detected. In some embodiments, PAP1 is detected. In some embodiments, PAP1 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, TFF1, S27A2, GKN2, Inhibin bB chain, and EphAl are detected. In some embodiments, GKN2 is detected. In some embodiments, GKN2 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, TFF1, S27A2, PAP1, Inhibin bB chain, and EphAl are detected. In some embodiments, Inhibin bB chain is detected. In some embodiments, Inhibin bB chain and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, and EphAl are detected. In some embodiments, EphAl is detected. In some embodiments, EphAl and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, and Inhibin bB chain are detected. Any of the embodiments described herein may be for use in predicting an individual’s risk of stomach cancer.

[0207] In some embodiments, kits comprise N biomarker protein capture reagents that specifically bind to at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or 8 of the biomarker proteins selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl. In some embodiments, kits comprise N biomarker protein capture reagents that specifically bind to LYPL1 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl. In some embodiments, kits comprise N biomarker protein capture reagents that specifically bind to CILP2 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl. In some embodiments, kits comprise N biomarker protein capture reagents that specifically bind to TFF1 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl . In some embodiments, kits comprise N biomarker protein capture reagents that specifically bind to S27A2 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, TFF1, PAP1, GKN2, Inhibin bB chain, and EphAl . In some embodiments, kits comprise N biomarker protein capture reagents that specifically bind to PAP1 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, TFF1, S27A2, GKN2, Inhibin bB chain, and EphAl . In some embodiments, kits comprise N biomarker protein capture reagents that specifically bind to GKN2 and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, TFF1, S27A2, PAP1, Inhibin bBchain, and EphAl . In some embodiments, kits comprise N biomarker protein capture reagents that specifically bind to Inhibin bB chain and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, and EphAl. In some embodiments, kits comprise N biomarker protein capture reagents that specifically bind to EphAl and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain. Any of the kits described herein may be for use in predicting an individual’s risk of stomach cancer.

[0208] In some aspects, the biomarkers detected can be at least at least 3, at least 4, at least 5, at least 6, or 7 of the biomarker proteins describe herein. Suitable examples of at least 3 biomarker proteins include, but are not limited to, e.g., LYPL1, CILP2, TFF1; LYPL1, CILP2, S27A2; LYPL1, CILP2, PAP1; LYPL1, CILP2, GKN2; LYPL1, CILP2, Inhibin bB chain; LYPL1, CILP2, EPHA1; LYPL1, TFF1, S27A2; LYPL1, TFF1, PAP1; LYPL1, TFF1, GKN2; LYPL1, TFF1, Inhibin bB chain; LYPL1, TFF1, EPHA1; LYPL1, S27A2, PAP1; LYPL1, S27A2, GKN2; LYPL1, S27A2, Inhibin bB chain; LYPL1, S27A2, EPHA1; LYPL1, PAP1, GKN2; LYPL1, PAP1, Inhibin bB chain; LYPL1, PAP1, EPHA1; LYPL1, GKN2, Inhibin bB chain; LYPL1, GKN2, EPHA1; LYPL1, Inhibin bB chain, EPHA1; CILP2, TFF1, S27A2;CILP2, TFF1, PAP1; CILP2, TFF1, GKN2; CILP2, TFF1, Inhibin bB chain; CILP2, TFF1, EPHA1; CILP2, S27A2, PAP1; CILP2, S27A2, GKN2; CILP2, S27A2, Inhibin bB chain; CILP2, S27A2, EPHA1; CILP2, PAP1, GKN2; CILP2, PAP1, Inhibin bB chain; CILP2, PAP1, EPHA1; CILP2, GKN2, Inhibin bB chain; CILP2, GKN2, EPHA1; CILP2, Inhibin bB chain, EPHA1; TFF1, S27A2, PAP1; TFF1, S27A2, GKN2; TFF1, S27A2, Inhibin bB chain; TFF1, S27A2, EPHA1; TFF1, PAP1, GKN2; TFF1, PAP1, Inhibin bB chain; TFF1, PAP1, EPHA1; TFF1, GKN2, Inhibin bB chain; TFF1, GKN2, EPHA1; TFF1, Inhibin bB chain, EPHA1; S27A2, PAP1, GKN2; S27A2, PAP1, Inhibin bB chain; S27A2, PAP1, EPHA1; S27A2, GKN2, Inhibin bB chain; S27A2, GKN2, EPHA1; S27A2, Inhibin bB chain, EPHA1; PAP1, GKN2, Inhibin bB chain; PAP1, GKN2, EPHA1; PAP1, Inhibin bB chain, EPHA1; GKN2, Inhibin bB chain, EPHA1.

[0209] Suitable biomarkers detected can be at least at least 4 of the biomarker proteins describe herein. Suitable 4 or more biomarker combinations include, but are not limited to, for example, LYPL1, CILP2, TFF1, S27A2; LYPL1, CILP2, TFF1, PAP1; LYPL1, CILP2, TFF1, GKN2; LYPL1, CILP2, TFF1, Inhibin bB chain; LYPL1, CILP2, TFF1, EPHA1; LYPL1, CILP2, S27A2, PAP1; LYPL1, CILP2, S27A2, GKN2; LYPL1, CILP2, S27A2, Inhibin bB chain; LYPL1, CILP2, S27A2, EPHA1; LYPL1, CILP2, PAP1, GKN2; LYPL1, CILP2, PAP1, Inhibin bB chain; LYPL1, CILP2, PAP1, EPHA1; LYPL1, CILP2, GKN2, Inhibin bBchain; LYPL1, CILP2, GKN2, EPHA1; LYPL1, CILP2, Inhibin bB chain, EPHA1; LYPL1, TFF1, S27A2, PAP1; LYPL1, TFF1, S27A2, GKN2; LYPL1, TFF1, S27A2, Inhibin bB chain; LYPL1, TFF1, S27A2, EPHA1; LYPL1, TFF1, PAP1, GKN2; LYPL1, TFF1, PAP1, Inhibin bB chain; LYPL1, TFF1, PAP1, EPHA1; LYPL1, TFF1, GKN2, Inhibin bB chain; LYPL1, TFF1, GKN2, EPHA1; LYPL1, TFF1, Inhibin bB chain, EPHA1; LYPL1, S27A2, PAP1, GKN2; LYPL1, S27A2, PAP1, Inhibin bB chain; LYPL1, S27A2, PAP1, EPHA1; LYPL1, S27A2, GKN2, Inhibin bB chain; LYPL1, S27A2, GKN2, EPHA1; LYPL1, S27A2, Inhibin bB chain, EPHA1; LYPL1, PAP1, GKN2, Inhibin bB chain; LYPL1, PAP1, GKN2, EPHA1; LYPL1, PAP1, Inhibin bB chain, EPHA1; LYPL1, GKN2, Inhibin bB chain, EPHA1; CILP2, TFF1, S27A2, PAP1; CILP2, TFF1, S27A2, GKN2; CILP2, TFF1, S27A2, Inhibin bB chain; CILP2, TFF1, S27A2, EPHA1; CILP2, TFF1, PAP1, GKN2; CILP2, TFF1, PAP1, Inhibin bB chain; CILP2, TFF1, PAP1, EPHA1; CILP2, TFF1, GKN2, Inhibin bB chain; CILP2, TFF1, GKN2, EPHA1; CILP2, TFF1, Inhibin bB chain, EPHA1; CILP2, S27A2, PAP1, GKN2;CILP2, S27A2, PAP1, Inhibin bB chain; CILP2, S27A2, PAP1, EPHA1; CILP2, S27A2, GKN2, Inhibin bB chain; CILP2, S27A2, GKN2, EPHA1; CILP2, S27A2, Inhibin bB chain, EPHA1; CILP2, PAP1, GKN2, Inhibin bB chain; CILP2, PAP1, GKN2, EPHA1; CILP2, PAP1, Inhibin bB chain, EPHA1; CILP2, GKN2, Inhibin bB chain, EPHA1; TFF1, S27A2, PAP1, GKN2; TFF1, S27A2, PAP1, Inhibin bB chain; TFF1, S27A2, PAP1, EPHA1; TFF1, S27A2, GKN2, Inhibin bB chain; TFF1, S27A2, GKN2, EPHA1; TFF1, S27A2, Inhibin bB chain, EPHA1; TFF1, PAP1, GKN2, Inhibin bB chain; TFF1, PAP1, GKN2, EPHA1; TFF1, PAP1, Inhibin bB chain, EPHA1; TFF1, GKN2, Inhibin bB chain, EPHA1; S27A2, PAP1, GKN2, Inhibin bB chain; S27A2, PAP1, GKN2, EPHA1; S27A2, PAP1, Inhibin bB chain, EPHA1; S27A2, GKN2, Inhibin bB chain, EPHA1; PAP1, GKN2, Inhibin bB chain, EPHA1.

[0210] Suitable biomarkers detected can be at least 5 of the biomarker proteins describe herein. Suitable biomarker combinations include, but are not limited to, for example, 5-protein combinations of LYPL1, CILP2, TFF1, S27A2, PAP1; LYPL1, CILP2, TFF1, S27A2, GKN2; LYPL1, CILP2, TFF1, S27A2, Inhibin bB chain; LYPL1, CILP2, TFF1, S27A2, EPHA1;LYPL1, CILP2, TFF1, PAP1, GKN2; LYPL1, CILP2, TFF1, PAP1, Inhibin bB chain; LYPL1, CILP2, TFF1, PAP1, EPHA1; LYPL1, CILP2, TFF1, GKN2, Inhibin bB chain; LYPL1, CILP2, TFF1, GKN2, EPHA1; LYPL1, CILP2, TFF1, Inhibin bB chain, EPHA1; LYPL1, CILP2, S27A2, PAP1, GKN2; LYPL1, CILP2, S27A2, PAP1, Inhibin bB chain; LYPL1, CILP2, S27A2, PAP1, EPHA1; LYPL1, CILP2, S27A2, GKN2, Inhibin bB chain; LYPL1, CILP2, S27A2, GKN2, EPHA1; LYPL1, CILP2, S27A2, Inhibin bB chain, EPHA1; LYPL1, CILP2, PAP1, GKN2, Inhibin bB chain; LYPL1, CILP2, PAP1, GKN2, EPHA1; LYPL1, CILP2, PAP1, Inhibin bB chain, EPHA1; LYPL1, CILP2, GKN2, Inhibin bB chain, EPHA1;LYPL1, TFF1, S27A2, PAP1, GKN2; LYPL1, TFF1, S27A2, PAP1, Inhibin bB chain;LYPL1, TFF1, S27A2, PAP1, EPHA1; LYPL1, TFF1, S27A2, GKN2, Inhibin bB chain;LYPL1, TFF1, S27A2, GKN2, EPHA1; LYPL1, TFF1, S27A2, Inhibin bB chain, EPHA1; LYPL1, TFF1, PAP1, GKN2, Inhibin bB chain; LYPL1, TFF1, PAP1, GKN2, EPHA1;LYPL1, TFF1, PAP1, Inhibin bB chain, EPHA1; LYPL1, TFF1, GKN2, Inhibin bB chain, EPHA1; LYPL1, S27A2, PAP1, GKN2, Inhibin bB chain; LYPL1, S27A2, PAP1, GKN2, EPHA1; LYPL1, S27A2, PAP1, Inhibin bB chain, EPHA1; LYPL1, S27A2, GKN2, Inhibin bB chain, EPHA1; LYPL1, PAP1, GKN2, Inhibin bB chain, EPHA1; CILP2, TFF1, S27A2, PAP1, GKN2; CILP2, TFF1, S27A2, PAP1, Inhibin bB chain; CILP2, TFF1, S27A2, PAP1, EPHA1; CILP2, TFF1, S27A2, GKN2, Inhibin bB chain; CILP2, TFF1, S27A2, GKN2, EPHA1; CILP2, TFF1, S27A2, Inhibin bB chain, EPHA1; CILP2, TFF1, PAP1, GKN2, Inhibin bB chain; CILP2, TFF1, PAP1, GKN2, EPHA1; CILP2, TFF1, PAP1, Inhibin bB chain, EPHA1; CILP2, TFF1, GKN2, Inhibin bB chain, EPHA1; CILP2, S27A2, PAP1, GKN2, Inhibin bB chain; CILP2, S27A2, PAP1, GKN2, EPHA1; CILP2, S27A2, PAP1, Inhibin bB chain, EPHA1; CILP2, S27A2, GKN2, Inhibin bB chain, EPHA1; CILP2, PAP1, GKN2, Inhibin bB chain, EPHA1; TFF1, S27A2, PAP1, GKN2, Inhibin bB chain; TFF1, S27A2, PAP1, GKN2, EPHA1; TFF1, S27A2, PAP1, Inhibin bB chain, EPHA1; TFF1, S27A2, GKN2, Inhibin bB chain, EPHA1; TFF1, PAP1, GKN2, Inhibin bB chain, EPHA1; S27A2, PAP1, GKN2, Inhibin bB chain, EPH Al

[0211] Suitable biomarkers detected can be at least at least 6 of the biomarker proteins describe herein. Suitable biomarker combinations include, but are not limited to, for example, 5-protein combinations of LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2; LYPL1, CILP2, TFF1, S27A2, PAP1, Inhibin bB chain; LYPL1, CILP2, TFF1, S27A2, PAP1, EPHA1; LYPL1, CILP2, TFF1, S27A2, GKN2, Inhibin bB chain; LYPL1, CILP2, TFF1, S27A2, GKN2, EPHA1; LYPL1, CILP2, TFF1, S27A2, Inhibin bB chain, EPHA1; LYPL1, CILP2, TFF1, PAP1, GKN2, Inhibin bB chain; LYPL1, CILP2, TFF1, PAP1, GKN2, EPHA1; LYPL1, CILP2, TFF1, PAP1, Inhibin bB chain, EPHA1; LYPL1, CILP2, TFF1, GKN2, Inhibin bB chain, EPHA1; LYPL1, CILP2, S27A2, PAP1, GKN2, Inhibin bB chain; LYPL1, CILP2, S27A2, PAP1, GKN2, EPHA1; LYPL1, CILP2, S27A2, PAP1, Inhibin bB chain, EPHA1; LYPL1, CILP2, S27A2, GKN2, Inhibin bB chain, EPHA1; LYPL1, CILP2, PAP1, GKN2, Inhibin bB chain, EPHA1; LYPL1, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain; LYPL1, TFF1, S27A2, PAP1, GKN2, EPHA1; LYPL1, TFF1, S27A2, PAP1, Inhibin bB chain, EPHA1; LYPL1, TFF1, S27A2, GKN2, Inhibin bB chain, EPHA1; LYPL1, TFF1, PAP1, GKN2, Inhibin bB chain, EPHA1; LYPL1, S27A2, PAP1, GKN2, Inhibin bB chain, EPHA1; CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain; CILP2, TFF1, S27A2, PAP1, GKN2,EPHA1; CILP2, TFF1, S27A2, PAP1, InhibinbB chain, EPHA1; CILP2, TFF1, S27A2, GKN2, Inhibin bB chain, EPHA1; CILP2, TFF1, PAP1, GKN2, Inhibin bB chain, EPHA1; CILP2, S27A2, PAP1, GKN2, Inhibin bB chain, EPHA1; TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, EPHA1

[0212] Suitable biomarkers detected can be at least at least 7 of the biomarker proteins describe herein. Suitable biomarker combinations include, but are not limited to, for example, 5-protein combinations of LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain;LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, EPHA1; LYPL1, CILP2, TFF1, S27A2, PAP1, Inhibin bB chain, EPHA1; LYPL1, CILP2, TFF1, S27A2, GKN2, Inhibin bB chain, EPHA1; LYPL1, CILP2, TFF1, PAP1, GKN2, InhibinbB chain, EPHA1; LYPL1, CILP2, S27A2, PAP1, GKN2, Inhibin bB chain, EPHA1; LYPL1, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, EPHA1; CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, EPHA1

[0213] Suitable biomarkers detected can be at least at least 8 of the biomarker proteins describe herein. Suitable biomarker combinations include, but are not limited to, for example, 5-protein combinations of LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, EPHA1 Computer Methods and Software

[0214] Once a biomarker or biomarker panel is selected, a method for diagnosing an individual can comprise the following: 1) collect or otherwise obtain a biological sample; 2) perform an analytical method to detect and measure the biomarker or biomarkers in the panel in the biological sample; 3) perform any data normalization or standardization required for the method used to collect biomarker levels; 4) calculate the marker score; 5) combine the marker scores to obtain a total diagnostic or predictive score; and 6) report the individual’s diagnostic or predictive score. In this approach, the diagnostic or predictive score may be a single number determined from the sum of all the marker calculations that is compared to a preset threshold value that is an indication of the presence or absence of disease or risk of stomach cancer. Or the diagnostic or predictive score may be a series of bars that each represent a biomarker level and the pattern of the responses may be compared to a pre-set pattern for determination of the presence or absence of disease, condition or the increased risk (or not) of an event.

[0215] At least some embodiments of the methods described herein can be implemented with the use of a computer. An example of a computer system 100 is shown in FIG. 6. With reference to FIG. 6, system 100 is shown comprised of hardware elements that are electrically coupled via bus 108, including a processor 101, input device 102, output device 103, storage device 104, computer-readable storage media reader 105a, communications system 106, processing acceleration (e.g., DSP or special-purpose processors) 107 and memory 109. Computer-readablestorage media reader 105a is further coupled to computer-readable storage media 105b, the combination comprehensively representing remote, local, fixed and / or removable storage devices plus storage media, memory, etc. for temporarily and / or more permanently containing computer-readable information, which can include storage device 104, memory 109 and / or any other such accessible system 100 resource. System 100 also comprises software elements (shown as being currently located within working memory 191) including an operating system 192 and other code 193, such as programs, data and the like.

[0216] With respect to FIG. 6, system 100 has extensive flexibility and configurability. Thus, for example, a single architecture might be utilized to implement one or more servers that can be further configured in accordance with currently desirable protocols, protocol variations, extensions, etc. However, it will be apparent to those skilled in the art that embodiments may well be utilized in accordance with more specific application requirements. For example, one or more system elements might be implemented as sub-elements within a system 100 component (e.g., within communications system 106). Customized hardware might also be utilized and / or particular elements might be implemented in hardware, software or both. Further, while connection to other computing devices such as network input / output devices (not shown) may be employed, it is to be understood that wired, wireless, modem, and / or other connection or connections to other computing devices might also be utilized.

[0217] In one aspect, the system can comprise a database containing features of biomarkers characteristic of risk stomach cancer in an individual. The biomarker data (or biomarker information) can be utilized as an input to the computer for use as part of a computer implemented method. The biomarker data can include the data as described herein.

[0218] In one aspect, the system further comprises one or more devices for providing input data to the one or more processors.

[0219] The system further comprises a memory for storing a data set of ranked data elements.

[0220] In another aspect, the device for providing input data comprises a detector for detecting the characteristic of the data element, e.g., such as a mass spectrometer or gene chip reader.

[0221] The system additionally may comprise a database management system. User requests or queries can be formatted in an appropriate language understood by the database management system that processes the query to extract the relevant information from the database of training sets.

[0222] The system may be connectable to a network to which a network server and one or more clients are connected. The network may be a local area network (LAN) or a wide area network (WAN), as is known in the art. Preferably, the server includes the hardware necessaryfor running computer program products (e.g., software) to access database data for processing user requests.

[0223] The system may include an operating system (e.g., UNIX or Linux) for executing instructions from a database management system. In one aspect, the operating system can operate on a global communications network, such as the internet, and utilize a global communications network server to connect to such a network.

[0224] The system may include one or more devices that comprise a graphical display interface comprising interface elements such as buttons, pull down menus, scroll bars, fields for entering text, and the like as are routinely found in graphical user interfaces known in the art. Requests entered on a user interface can be transmitted to an application program in the system for formatting to search for relevant information in one or more of the system databases.Requests or queries entered by a user may be constructed in any suitable database language.

[0225] The graphical user interface may be generated by a graphical user interface code as part of the operating system and can be used to input data and / or to display inputted data. The result of processed data can be displayed in the interface, printed on a printer in communication with the system, saved in a memory device, and / or transmitted over the network or can be provided in the form of the computer-readable medium.

[0226] The system can be in communication with an input device for providing data regarding data elements to the system (e.g., expression values). In one aspect, the input device can include a gene expression profiling system including, e.g., a mass spectrometer, gene chip or array reader, and the like.

[0227] The methods and apparatus for analyzing the risk of stomach cancer biomarker information according to various embodiments may be implemented in any suitable manner, for example, using a computer program operating on a computer system. A conventional computer system comprising a processor and a random-access memory, such as a remotely accessible application server, network server, personal computer or workstation may be used. Additional computer system components may include memory devices or information storage systems, such as a mass storage system and a user interface, for example a conventional monitor, keyboard and tracking device. The computer system may be a stand-alone system or part of a network of computers including a server and one or more databases.

[0228] The risk assessment biomarker analysis system can provide functions and operations to complete data analysis, such as data gathering, processing, analysis, reporting and / or diagnosis. For example, in one embodiment, the computer system can execute the computer program that may receive, store, search, analyze, and report information relating to the stomach cancer risk biomarkers. The computer program may comprise multiple modules performingvarious functions or operations, such as a processing module for processing raw data and generating supplemental data and an analysis module for analyzing raw data and supplemental data to generate a stomach cancer risk status. Determination of the probability of an individual’s risk for stomach cancer may optionally comprise generating or collecting any other information, including additional biomedical information, regarding the condition of the individual relative to the disease, condition or event, identifying whether further tests may be desirable, or otherwise evaluating the health status of the individual.

[0229] Referring now to FIG. 7, an example of a method of utilizing a computer in accordance with principles of a disclosed embodiment can be seen. In FIG. 7, a flowchart 3000 is shown. In block 3004, biomarker information can be retrieved for an individual. The biomarker information can be retrieved from a computer database, for example, after testing of the individual’s biological sample is performed. The biomarker information can comprise biomarker levels that each correspond to one or more of the biomarkers of Table 1. In block 3008, a computer can be utilized to classify each of the biomarker levels. And, in block 3012, a determination can be made as to an individual’s risk of stomach cancer based upon a plurality of classifications. The indication can be output to a display or other indicating device so that it is viewable by a person. Thus, for example, it can be displayed on a display screen of a computer or other output device.

[0230] Some embodiments described herein can be implemented so as to include a computer program product. A computer program product may include a computer readable medium having computer readable program code embodied in the medium for causing an application program to execute on a computer with a database.

[0231] As used herein, a “computer program product” refers to an organized set of instructions in the form of natural or programming language statements that are contained on a physical media of any nature (e.g., written, electronic, magnetic, optical or otherwise) and that may be used with a computer or other automated data processing system. Such programming language statements, when executed by a computer or data processing system, cause the computer or data processing system to act in accordance with the particular content of the statements. Computer program products include without limitation: programs in source and object code and / or test or data libraries embedded in a computer readable medium. Furthermore, the computer program product that enables a computer system or data processing equipment device to act in pre-selected ways may be provided in a number of forms, including, but not limited to, original source code, assembly code, object code, machine language, encrypted or compressed versions of the foregoing and any and all equivalents.

[0232] In one aspect, a computer program product is provided for assessment of risk of stomach cancer. The computer program product includes a computer readable medium embodying program code executable by a processor of a computing device or system, the program code comprising: code that retrieves data attributed to a biological sample from an individual, wherein the data comprises biomarker values that each correspond to one or more of the biomarkers of Table 1; and code that executes a classification method that indicates a risk of stomach cancer in the individual as a function of the biomarker values.

[0233] While various embodiments have been described as methods or apparatuses, it should be understood that embodiments can be implemented through code coupled with a computer, e.g., code resident on a computer or accessible by the computer. For example, software and databases could be utilized to implement many of the methods discussed above. Thus, in addition to embodiments accomplished by hardware, it is also noted that these embodiments can be accomplished through the use of an article of manufacture comprised of a computer usable medium having a computer-readable program code embodied therein, which causes the enablement of the functions disclosed in this description. Therefore, it is desired that embodiments also be considered protected by this patent in their program code means as well. Furthermore, the embodiments may be embodied as code stored in a computer-readable memory of virtually any kind including, without limitation, RAM, ROM, magnetic media, optical media, or magneto-optical media. Even more generally, the embodiments could be implemented in software, or in hardware, or any combination thereof including, but not limited to, software running on a general-purpose processor, microcode, PLAs, or ASICs.

[0234] It is also envisioned that embodiments could be accomplished as computer signals embodied in a carrier wave, as well as signals (e.g., electrical and optical) propagated through a transmission medium. Thus, the various types of information discussed above could be formatted in a structure, such as a data structure, and transmitted as an electrical signal through a transmission medium or stored on a computer readable medium.

[0235] It is also noted that many of the structures, materials, and acts recited herein can be recited as means for performing a function or step for performing a function. Therefore, it should be understood that such language is entitled to cover all such structures, materials, or acts disclosed within this specification and their equivalents, including the matter incorporated by reference.

[0236] The biomarker identification process, the utilization of the biomarkers disclosed herein, and the various methods for determining biomarker values are described in detail above with respect to evaluation of a risk of stomach cancer in an individual. However, the application of the process, the use of identified biomarkers, and the methods for determining biomarkervalues are fully applicable to other specific types of diseases or medical conditions, or to the identification of individuals who may or may not be benefited by an ancillary medical treatment.Other Methods

[0237] In some embodiments, the biomarkers and methods described herein are used to determine a medical insurance premium or coverage decision and / or a life insurance premium or coverage decision. In some embodiments, the results of the methods described herein are used to determine a medical insurance premium and / or a life insurance premium. In some such instances, an organization that provides medical insurance or life insurance requests or otherwise obtains information concerning an individual’s risk of stomach cancer and uses that information to determine an appropriate medical insurance or life insurance premium for the subject. In some embodiments, the test is requested by, and paid for by, the organization that provides medical insurance or life insurance. In some embodiments, the test is used by the potential acquirer of a practice or health system or company to predict future liabilities or costs should the acquisition go ahead.

[0238] In some embodiments, the biomarkers and methods described herein are used to predict and / or manage the utilization of medical resources. In some such embodiments, the methods are not carried out for the purpose of such prediction, but the information obtained from the method is used in such a prediction and / or management of the utilization of medical resources. For example, a testing facility or hospital may assemble information from the present methods for many subjects in order to predict and / or manage the utilization of medical resources at a particular facility or in a particular geographic area.Exemplary Aspects

[0239] Aspect 1 is a method of predicting stomach cancer risk in a subject, comprising forming a biomarker panel comprising N biomarker proteins, and detecting a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or 8 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0240] Aspect 2 is a method of detecting levels of N biomarker proteins in a sample, comprising forming a biomarker panel comprising N biomarker proteins, and detecting the level of each of the N biomarker proteins in the sample from a subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or 8 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0241] Aspect 3 is a method of predicting stomach cancer risk in a subject, comprising detecting a level of LYPL1 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0242] Aspect 4 is a method of predicting stomach cancer risk in a subject, comprising detecting a level of CILP2 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0243] Aspect 5 is a method of predicting stomach cancer risk in a subject, comprising detecting a level of TFF1 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0244] Apsect 6 is a method of predicting stomach cancer risk in a subject, comprising detecting a level of S27A2 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0245] Aspect 7 is a method of predicting stomach cancer risk in a subject, comprising detecting a level of PAP 1 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, GKN2, Inhibin bB chain, and EphAl.

[0246] Apsect 8 is a method of predicting stomach cancer risk in a subject, comprising detecting a level of RPIA and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the Nbiomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, Inhibin bB chain, and EphAl.

[0247] Aspect 9 is a method of predicting stomach cancer risk in a subject, comprising detecting a level of Inhibin bB chain and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, and EphAl.

[0248] Aspect 10 is a method of predicting stomach cancer risk in a subject, comprising detecting a level of EphAl and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, and Inhibin bB chain.

[0249] Apsect 11 is the method according to any one of Aspects 1-10, wherein N is 2 to 8, or N is 3 to 8, N is 4 to 8, N is 5 to 8, N is 6 to 8, or N is 7 to 8.

[0250] Aspect 12 is the method according to any one of Aspects 1-11, wherein N is 2, N is 3, N is 4, N is 5, N is 6, N is 7, or N is 8.

[0251] Aspect 13 is the method according to any one of Aspects 1-12, wherein at least one of the N biomarker proteins is LYPL1, or at least one of the N biomarker proteins is CILP2, or at least one of the N biomarker proteins is TFF1, or at least one of N biomarker proteins is S27A2, or at least one of the N biomarker proteins is PAP1, or at least one of the N biomarker proteins is GKN2, or at least one of the N biomarker proteins is Inhibin bB chain, or at least one of the N biomarker proteins is EphAl .

[0252] Aspect 14 is the method according to any one of Aspects 1-13, wherein each of the N biomarker proteins is selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl .

[0253] Aspect 15 is the method according to any one of Aspects 1-14, wherein at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0254] Aspects 16 is the method according to any one of Aspects 1-3, wherein 2 of the N biomarker proteins are LYPL1 and CILP2, or 2 of the N biomarker proteins are LYPL1 and TFF1, or 2 of the N biomarker proteins are LYPL1 and S27A2, or 2 of the N biomarker proteins are LYPL1 and PAP1, or 2 of the N biomarker proteins are LYPL1 and GKN2, or 2 of the N biomarker proteins are LYPL1 and Inhibin bB chain, or 2 of the N biomarker proteins are LYPL1 and EphAl.

[0255] Aspects 17 is the method according to any one of Aspects 1, 2 or 4, wherein 2 of the N biomarker proteins are CILP2 and TFF1, or 2 of the N biomarker proteins are CILP2 and S27A2, or 2 of the N biomarker proteins are CILP2 and PAP1, or 2 of the N biomarker proteins are CILP2 and GKN2, or 2 of the N biomarker proteins are CILP2 and Inhibin bB chain, or 2 of the N biomarker proteins are CILP2 and EphAl.

[0256] Aspect 18 is the method according to any one of Aspects 1, 2 or 5, wherein 2 of the N biomarker proteins are TFF1 and S27A2, or 2 of the N biomarker proteins are TFF1 and PAP1, or 2 of the N biomarker proteins are TFF1 and GKN2, or 2 of the N biomarker proteins are TFF1 and Inhibin bB chain, or 2 of the N biomarker proteins are TFF1 and EphAl.

[0257] Aspect 19 is the method according to any one of Aspects 1, 2 or 6, wherein 2 of the N biomarker proteins are S27A2 and PAP1, or 2 of the N biomarker proteins are S27A2 and GKN2, or 2 of the N biomarker proteins are S27A2 and Inhibin bB chain, or 2 of the N biomarker proteins are S27A2 and EphAl.

[0258] Aspect 20 is the method according to any one of Aspects 1, 2 or 7, wherein 2 of the N biomarker proteins are PAP1 and GKN2, or 2 of the N biomarker proteins are PAP1 and Inhibin bB chain, or 2 of the N biomarker proteins are PAP1 and EphAl .

[0259] Aspect 21 is the method according to any one of Aspects 1, 2 or 8, wherein 2 of the N biomarker proteins are GKN2 and Inhibin bB chain, or 2 of the N biomarker proteins are GKN2 and EphAl.

[0260] Aspect 22 is the method according to any one of Aspects 1, 2 or 9, wherein 2 of the N biomarker proteins are Inhibin bB chain and EphAl.

[0261] Aspect 23 is the method according to any one of Aspects 1-4, wherein 3 of the N biomarker proteins are LYPL1, CILP2, and TFF1, or 3 of the N biomarker proteins are LYPL1, CILP2, and S27A2, or 3 of the N biomarker proteins are LYPL1, CILP2, and PAP1, or 3 of the N biomarker proteins are LYPL1, CILP2, and GKN2, or 3 of the N biomarker proteins are LYPL1, CILP2, and Inhibin bB chain, or 3 of the N biomarker proteins are LYPL1, CILP2, and EphAl.

[0262] Aspect 24 is the method according to any one of Apects 1-23, wherein the sample is a blood sample, a plasma sample, a serum sample, or a urine sample.

[0263] Aspect 25 is the method according to any one of Aspects 1-24, wherein the risk of the subject stomach cancer within 5 years from the date that the sample was taken from the subject is predicted.

[0264] Aspect 26 is the method according to any one of Aspects 1-25, wherein detecting is performed using mass spectrometry, an aptamer based assay and / or an antibody based assay.

[0265] Aspect 27 is the method according to any one of Aspects 1-26, wherein the method comprises contacting biomarker proteins of the sample or samples with a set of biomarker capture reagents, wherein each biomarker capture reagent of the set of biomarker capture reagents specifically binds to a different biomarker protein being detected.

[0266] Aspect 28 is the method according to Aspect 27, wherein each biomarker capture reagent is an antibody or an aptamer.

[0267] Aspect 29 is the method according to Aspect 28, wherein each biomarker capture reagent is an aptamer.

[0268] Aspect 30 is the method according to Aspect 29, wherein at least one aptamer is a slow off-rate aptamer.

[0269] Aspect 31 the method according to Aspect 30, wherein at least one slow off-rate aptamer comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least 10 nucleotides with modifications.

[0270] Aspect 32 is the method according to Aspects 30 or claim 31, wherein each slow off-rate aptamer binds to its target protein with an off rate (t’A) of > 20 minutes, > 30 minutes, > 60 minutes, > 90 minutes, > 120 minutes, > 150 minutes, > 180 minutes, > 210 minutes, or > 240 minutes.

[0271] Aspect 33 is the method according to any one of Aspects 26-32, wherein the level of each biomarker protein measured is determined from a relative florescence unit (RFU) or a protein concentration.

[0272] Aspect 34 is the method according to any one of Aspects 1-33, wherein predicting the risk of stomach cancer in the subject is based on input of the levels of the N biomarker proteins measured in a statistical model.

[0273] Aspect 35 is the method according to Aspect 34, wherein the determining comprises analyzing the levels of the N biomarker protein using a survival model.

[0274] Aspect 36 is the method according to Aspect 35, wherein the survival model is an Accelerated Failure Time (AFT) model with a Weibull distribution.

[0275] Aspect 37 is the method according to any one of Aspects 34-36, wherein the model has an area under the curve (AUC) selected from at least 0.64, at least 0.65, at least 0.66, at least 0.67, at least 0.68, at least 0.69, at least 0.7, at least 0.75, at least 0.8, at least 0.85, at least 0.9, or at least 0.95.

[0276] Aspect 38 is the method according to any one of Aspects 1-37, wherein the method comprises predicting a risk of stomach cancer in a subject for the purpose of determining a medical insurance premium or life insurance premium.

[0277] Aspect 39 is the method according to Aspect 38, wherein the method further comprises determining coverage for medical insurance or life insurance.

[0278] Aspect 40 is the method according to any one of Aspects 1-37, wherein the method further comprises using information resulting from the method to predict and / or manage the utilization of medical resources.

[0279] Aspect 41 is a kit comprising N biomarker protein capture reagents, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 of the N biomarker protein capture reagents specifically binds to a biomarker protein selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0280] Aspect 42 is the kit according to Aspects 41, wherein N is 2 to 8, or N is 3 to 8, or N is 4 to 8, or N is 5 to 8, or N is 6 to 8, or N is 7 to 8.

[0281] Aspect 43 is the kit according to Aspects 41 or 42, wherein N is 2, N is 3, N is 4, N is 5, N is 6, N is 7, or N is 8. The kit according to any one of claims 41-43, wherein each of the N biomarker protein capture reagents specifically binds to a different biomarker protein.

[0282] Aspect 44 is the kit according to any one of Aspects 41-44, wherein each of the N biomarker protein capture reagents specifically binds to a biomarker protein selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

[0283] Aspect 45 is the kit according to any one of Aspects 41-44, wherein the N biomarker protein capture reagent specifically bind to the N biomarker proteins of any one of claims 1-23.

[0284] Aspect 46 is a kit comprising N biomarker protein capture reagents, wherein the kit comprises biomarker protein capture reagents for carrying out the method of any one of Aspects 1-44.

[0285] Aspect 47 is the kit according to any one of Aspects 41-46, wherein each of the N biomarker protein capture reagents is an antibody or an aptamer.

[0286] Aspect 47 is the kit according to claim 46, wherein each biomarker protein capture reagent is an aptamer, optionally wherein at least one aptamer is a slow off-rate aptamer.

[0287] Aspect 48 is the kit according to Aspect 47, wherein at least one slow off-rate aptamer comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least 10 nucleotides with modifications.

[0288] Aspect 49 is the kit according to Aspects 46 or claim 47, wherein each slow off-rate aptamer binds to its target protein with an off rate (t%) of > 20 minutes, > 30 minutes, > 60 minutes, > 90 minutes, > 120 minutes, > 150 minutes, > 180 minutes, > 210 minutes, or > 240 minutes.

[0289] Aspect 50 is the kit according to any one of Aspects 41-49, for use in detecting the N biomarker proteins in a sample from a subject.

[0290] Aspect 51 is the kit according to Aspect 50, for use in predicting an individual’ s risk of stomach cancer.EXAMPLES

[0291] The following examples are provided for illustrative purposes only and are not intended to limit the scope of the application as defined by the appended claims. All examples described herein were carried out using standard techniques, which are well known and routine to those of skill in the art. Routine molecular biology techniques described in the following examples can be carried out as described in standard laboratory manuals, such as Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd. ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., (2001).Example 1. Multiplex Aptamer Assay and Statistical Approaches for Biomarker Identification

[0292] A multiplex aptamer assay was used to analyze test samples and control samples to identify biomarkers for predicting an individual’s risk of stomach cancer within a specified time frame. The multiplexed analysis used in this experiment included aptamers to detect approximately 5,000 proteins in blood from small sample volumes (~65 pl of serum or plasma), with low limits of detection (1 pM median), ~7 logs of dynamic range, and ~5% median coefficient of variation. The multiplex aptamer assay is described, generally, e.g., in Gold et al. (2010) Aptamer-Based Multiplexed Proteomic Technology for Biomarker Discovery. PLoS ONE 5(12): el5004; and U.S. Publication Nos: 2012 / 0101002 and 2012 / 0077695.Example 2. Model Specification

[0293] Endpoint Description: The endpoint for this model was a survival endpoint based on time-to-event, which has two components. The first is a binary variable which indicates whether the subject received a diagnosis of primary stomach cancer over the study period (1) or not (0). This was adjudicated through record linkage with regional cancer registries. Cancer types were defined using the tenth edition of the International Classification of Disease (ICD10) and the second edition of the International Classification of Disease for Oncology (ICD-O-2). Stomach cancers included all those found within the fundus (Cl 61), body (Cl 62), greater (Cl 66) and lesser (Cl 65) curvature of the stomach, overlapping lesions (Cl 68), gastric antrum (Cl 63), pylorus (Cl 64), and stomach and cardia not otherwise specified (Cl 69 and Cl 60). The second component to the endpoint is the “time-to-event,” which in this case is the time in follow-up days from the blood draw to: 1) the time of diagnosis of first primary stomach cancer, or 2)censor date for participants that did not have a primary stomach cancer diagnosis, which can occur due to end of study or removal from study for reasons not related to stomach cancer (including death or registration).

[0294] Model Information: The selected model is an 8-feature, protein-only (see Table 1) Accelerated Failure Time (AFT) model with a Weibull distribution.

[0295] The model was trained on the entire follow-up period, with performance maximized at 5 years. The output is the absolute probability of not having a stomach cancer diagnosis, a value between 0.0000 and 1.0000, within 5 years. When delivered as a result and assessed by business rules, the resulting predicted probability will be subtracted from 1 to provide an absolute 5-year probability of a stomach cancer diagnosis.

[0296] The baseline risk probability score represents the absolute risk for the “average” person in the training cohort based on the model algorithm. A “baseline” individual is defined as an individual with model feature values set to zero. All features in the model are centered on the overall mean of the training data, which means a value of 0 for any given feature is equal to the mean (i.e., the average). The baseline value is calculated by setting all the features to zero and then generating the absolute risk probability on those “zeroed” features.

[0297] The baseline absolute risk of stomach cancer in the training dataset is 0.072% at 5 years post blood draw. The rate of 5-year stomach cancer diagnosis in the EPIC dataset is comparable to the U.S. population stomach cancer event rate in the intended use population (EPIC dataset described below).

[0298] Absolute risk scores were stratified by quartiles of absolute risk probabilities in the training dataset. Kaplan-Meier (KM) survival (“Event-free”) curves were generated for these quartiles for the training and verification datasets. The survival curves showed close agreement between training and verification datasets (FIG. 8). Note that the first two risk bins (Quartile 1 and 2) roughly correspond to individuals whose predicted risk of stomach cancer is lower than the average risk based on our data (baseline score at 5 years = 0.072%), while Quartiles 3 and 4 correspond to predicted risk that is higher than the baseline. The summary of absolute probability risk stratification and corresponding event rates is shown in Table 2. Both the mean predicted event rate and (observed) KM event rates are shown. KM event rates consider censoring and are weighted to reflect overall population instead of the analyzed cohort.Table 2. Risk tables from final 5-year Malignant Stomach Cancer Risk model predictions in training and verification data. Risk bin ranges were derived by partitioning predicted 5-year event probabilities on the training dataset into quartiles.Bin cutoffs (in N N N Mean Kaplan- Risk Quartile absolute probability) Cancer- Events2Censored3Predicted Meier Event free1Event Rate Rate (95% CI)4Qi 0.0000 <X < 0.0004 1,945 1 201 0.0002 0.0001(0.0000, 0.0003) Q2 0.0004 < X < 0.0007 1,905 4 237 0.0006 0.0004(0.0000, Absolute 0.0009) Q3 0.0007 < X < 0.0013 1,886 5 255 0.0010 0.0007(0.0001, 0.0013) Q4 0.0013 < X < 1.0000 1,844 22 281 0.0028 0.0034(0.0019,0.0048) ^Individuals neither censored nor diagnosed with malignant stomach cancer within 5 years (1,825 days) of blood draw (i.e., cancer-free).^Individuals diagnosed with malignant stomach cancer within 5 years (1,825 days) of blood draw. ^Individuals lost to follow-up (censored) within 5 years (1,825 days) of blood draw, with unknown status after the censoring time.^Kaplan-Meier (KM) event rate takes into account loss to follow-up (censoring) within 5 years (1,825 days).

[0299] Based on this stratification, the scoring rules for the absolute risk for the stomach cancer risk test are shown in Table 3.Table 3. Scoring rules for absolute risk probabilities.Test outcome X (absoluterisk probability) Predicted Class Plain language0.0000 <X< 0.0004 Low Absolute risk predictions between 0.0000 and 0.0004(inclusive) are labeled as “Low”.0.0004 < X < 0.0007 Medium-Low Absolute risk predictions between 0.0004 and 0.0007(inclusive) are labeled as “Medium-Low”.0.0007 < X < 0.0013 Medium-High Absolute risk predictions between 0.0007 and 0.0013(inclusive) are labeled as “Medium-High”.0.0013 < X < 1.0000 High Absolute risk predictions between 0.0013 and 1.0000(inclusive) are labeled as “High”.

[0300] Model calibration was assessed by comparing predicted event rates and Kaplan-Meier (KM) event rates. Specifically, predicted event probabilities on the training dataset were divided into deciles. The mean predicted event rate and case-cohort weighted KM event rate were calculated for each decile, respectively, along with the 95% confidence intervals for the weighted KM event rates. These data were visualized in a calibration plot (FIG. 9). Risk bins were created by partitioning predictions of 5-year malignant stomach cancer risk (“Event”) on training samples into deciles. Kaplan-Meier event rates (observed event rate) were calculated using case-cohort weights. Mean predicted 5-year risk of stomach cancer (predicted event rate)and KM event rate was calculated for each decile of predicted risk from final stomach cancer model. The solid, black line is the reference line of identity, representing all points where predicted risks agree with observed KM event rates. Vertical bars about each decile’s point estimate are the 95% confidence intervals of the Kaplan-Meier event rate. For deciles with nonzero event rates, these points generally fall along the diagonal, visually indicating good model calibration.

[0301] The model output is the Pr(No MSC) at 5 years (1,825 days), where “MSC” is first primary malignant stomach cancer. The output will be reported as the probability of a first primary malignant stomach cancer diagnosis, which is (1 - Pr(No MSC)).Table 4. Model performance in predicting first primary malignant stomach cancer risk within 5 years of blood draw in cancer-free individuals, along with age-only comparator model on training, verification, and validation data. The AUC, sensitivity, and specificity were calculated at 5 years (1,825 days). CI = confidence intervals;AUC = area under the curve; PEC = predictive error curve.AUC at 5 Sensitivity1at Specificity1at C-index2PEC Model Dataset N years 5 years 5 years (95% CI) at 5 (95% CI) (95% CI) (95% CI) years (95% CI) Training 8,586 0.771 0.750 0.727 0.714 0.004 (0.683, (0.583, 0.914) (0.714, 0.739) (0.669, (0.003, 0.857) 0.757) 0.006) Proteomic Verification 2,862 0.739 0.700 0.718 0.634 0.004 (0.594, (0.400, 1.000) (0.699, 0.731) (0.548, (0.001, 0.902) 0.726) 0.006) Validation 2,862 0.742 0.714 0.737 0.697 0.005 (0.523, (0.411, 0.917) (0.716, 0.757) (0.562, (0.003, 0.874) 0.784) 0.009) Training 8,586 0.600 0.719 0.572 0.587 0.004 (0.498, (0.577, 0.858) (0.562, 0.585) (0.537, (0.003, 0.692) 0.636) 0.006) Age-only3Verification 2,862 0.577 0.700 0.573 0.566 0.004 (0.430, (0.427, 1.000) (0.556, 0.596) (0.491, (0.001, 0.739) 0.660) 0.006) Validation 2,862 0.604 0.643 0.564 0.561 0.005 (0.455, (0.333, 0.900) (0.547, 0.584) (0.437, (0.003,0.740) 0.686) 0.009)’Sensitivity and specificity calculated using the cutoff that maximized Youden's J on the training dataset; linear space cutoff = -13.126 (Proteomic) and linear space cutoff = -13.15 (Age -only).2C-index is not time specific3Model includes only participant age at the time of blood draw as a covariate.

[0302] Description of Clinical Model Used as a Comparator: As genome sequencing becomes more widespread, cost effective, and common in clinical settings, Polygenic Risk Score (PRS)-based models incorporating single nucleotide polymorphisms (SNPs) are increasingly being used to predict risk of future disease. In stomach cancer, PRS-based models include known individual genetic information potentially allowing for more personalized risk predictionwhile reducing patient burden compared to more invasive screening methods such as endoscopy or imaging methods. Therefore, a PRS-based model was used as the clinical comparator for the proteomic model.

[0303] PRS-based models for the prediction of future stomach cancer are in development especially in regions with higher incidence of stomach cancer such as countries in East Asia and Central Europe. Specifically, individual unique PRS models have been developed on populations living in China, Korea, Japan, and Europe with the number of genetic variants in the models ranging from 3-21 SNPs identified in genome-wide association studies (GWAS) studies to be associated with stomach cancer susceptibility. Model predictive performance ranged from AUC = 0.560-0.737, with an overall mean AUC = 0.629.

[0304] Therefore, an AUC greater than or equal to 0.629 was set as the performance threshold during model development for the stomach cancer risk test.

[0305] Development and Validation Cohort(s): The European Prospective Investigation into Cancer and Nutrition (EPIC) is an ongoing multi-center prospective cohort study designed to investigate the relationship between nutrition and cancer. (Riboli E, Hunt KJ, Slimani N, et al. European Prospective Investigation into Cancer and Nutrition (EPIC): study populations and data collection. Public Health Nutr. 2002;5(6B): 1113-1124.) The EPIC study is a collaborative effort between Imperial College London, the International Agency for Research on Cancer (IARC), and 23 European institutes within 10 countries. More than 500,000 individuals aged 35-75 who were cancer-free were enrolled between 1992 and 2000. Individuals were followed for cancer incidence and cause-specific mortality for several decades. Participant eligibility was based on geographic boundaries. The source population was identified according to age and sex, and the actual study populations were convenience samples of volunteers agreeing to participate from among the general adult population residing in a given town or geographical area who were invited to participate.

[0306] SomaLogic was provided with citrate plasma samples for 14,787 individuals obtained at enrollment and clinical data for up to 20.4 years of follow-up from 14 centers within four countries (Italy, Spain, The Netherlands, and the United Kingdom). The stomach cancer risk test was developed using samples from all participants (n=14,787), among whom n=219 developed malignant stomach cancer within the follow up period. Cancer diagnoses were adjudicated through record linkage with regional cancer registries. Cancer types were defined using the International Classification of Diseases-Tenth Revision and the second revision of the International Classification of Diseases for Oncology (ICD-O-2).

[0307] Stomach cancers included all those found within the fundus (Cl 61), body (Cl 62), greater (Cl 66) and lesser (Cl 65) curvature of the stomach, overlapping lesions (Cl 68), gastricantrum (Cl 63), pylorus (Cl 64), and stomach and cardia not otherwise specified (Cl 69 and Cl 60). Case-cohort sample weights were assigned to each study participant by the EPIC research team. Sample weights were used to account for the probability of study inclusion of each EPIC participant who fulfilled the study inclusion criteria. These weights were estimated as a function of center of recruitment, sex, and disease status (i.e., incidence of cancer, type II diabetes, cardiovascular disease, and death during follow-up). Ethnicity information was not provided by the EPIC research team.

[0308] Current data on cancer incidence in the US suggests that 0.8% of the population will be diagnosed with stomach cancer during their lifetime. The annual age-adjusted risk is 0.007%, which corresponds to a 5-year risk of 0.035%. The baseline, or average, 5-year risk of incident stomach cancer in the EPIC dataset was 0.072%, and therefore is within the range of incident stomach cancer risk estimates for adults living in the U.S.

[0309] The EPIC dataset was split independently into three sets (60% training / 20% verification / 20% validation), which allowed identification of a robust model while mitigating overfitting issues. The validation dataset was not used in the POC or refinement stages.

[0310] Model development data: A total of 14,310 samples were available after data QC for analysis. Demographic information for model development data (training and verification) are provided for the full follow-up period (20.4 years) and for 5-year (1,825-day) censoring in Tables 5-8.Table 5. Demographic information for the model development (training) dataset for the full follow-up period (up to 20.4 years) for the Stomach Cancer risk test.Malignant Stomach Cancer Covariate Measure Total DiagnosisNo Yes Sample Size N 8586 8468 118Female 4760 (55.4%) 4710 (55.6%) 50 (42.4%) Subject GenderMale 3826 (44.6%) 3758 (44.4%) 68 (57.6%) No 8139 (94.8%) 8025 (94.8%) 114 (96.6%) Type II Diabetes Status Yes 385 (4.5%) 381 (4.5%) 4 (3.4%)Unknown 62 (0.7%) 62 (0.7%) 0 (0.0%) Current 2372 (27.6%) 2336 (27.6%) 36 (30.5%) Tobacco Use Status Never 3760 (43.8%) 3720 (43.9%) 40 (33.9%)Past 2454 (28.6%) 2412 (28.5%) 42 (35.6%) Asturias 600 (7%) 588 (6.9%) 12 (10.2%)Bilthoven 561 (6.5%) 550 (6.5%) 11 (9.3%) Cambridge 1435 (16.7%) 1424 (16.8%) 11 (9.3%) Florence 668 (7.8%) 650 (7.7%) 18 (15.3%) Granada 436 (5.1%) 435 (5.1%) 1 (0.8%) Murcia 543 (6.3%) 536 (6.3%) 7 (5.9%) Naples 114 (1.3%) 111 (1.3%) 3 (2.5%) Subject Site ID Navarra 636 (7.4%) 625 (7.4%) 11 (9.3%) Oxford 541 (6.3%) 537 (6.3%) 4 (3.4%) Ragusa 299 (3.5%) 297 (3.5%) 2 (1.7%) San Sebastian 603 (7%) 595 (7%) 8 (6.8%) Turin 585 (6.8%) 576 (6.8%) 9 (7.6%) Utrecht 1051 (12.2%) 1045 (12.3%) 6 (5.1%) Varese 514 (6%) 499 (5.9%) 15 (12.7%) Mean (SD) 55.698 (9.09) 55.68 (9.10) 56.966 (8.43) Subject Age (years) Median 56 56 58Range 35 - 75 35 - 75 36 - 74 Mean (SD) 27.148 (4.38) 27.146 (4.39) 27.303 (3.82) Body Mass Index(kg / mA2) Median 26.62 26.62 26.785Range 15.63 - 59.99 15.63 - 59.99 20.63 - 37.67 Mean (SD) 13.72 (20.88) 13.664 (20.88) 17.729 (20.31) Alcohol ConsumptionRate (g / d) Median 4.81 4.80 12.29Range 0 - 217.18 0 - 217.18 0 - 96.38 Mean (SD) 4422.53 4441.38 3070.13 (1761.57) (1755.64) (1663.68) Time-to-Event / Censoring(days) Median 4971 4993 3115.5 Range 7 - 7378 7 - 7378 68 - 6224Table 6. Demographic information for the model development (training) dataset for 5-year (1,825-day) censoring for the Stomach Cancer risk test.Malignant Stomach Cancer Diagnosis Covariate Measure TotalNo Yes Censored Sample Size N 8586 7580 32 974Female 4760 (55.4%) 4205 (55.5%) 19 (59.4%) 536 (55%) Subject GenderMale 3826 (44.6%) 3375 (44.5%) 13 (40.6%) 438 (45%) Type II Diabetes No 8139 (94.8%) 7186 (94.8%) 31 (96.9%) 922 (94.7%) StatusYes 385 (4.5%) 338 (4.5%) 1 (3.1%) 46 (4.7%) Unknown 62 (0.7%) 56 (0.7%) 0 (0.0%) 6 (0.6%) Current 2372 (27.6%) 2105 (27.8%) 6 (18.8%) 261 (26.8%) Tobacco UseStatus Never 3760 (43.8%) 3349 (44.2%) 15 (46.9%) 396 (40.7%) Past 2454 (28.6%) 2126 (28%) 11 (34.4%) 317 (32.5%) Asturias 600 (7%) 563 (7.4%) 1 (3.1%) 36 (3.7%) Bilthoven 561 (6.5%) 507 (6.7%) 2 (6.2%) 52 (5.3%) Cambridge 1435 (16.7%) 1215 (16%) 3 (9.4%) 217 (22.3%) Florence 668 (7.8%) 564 (7.4%) 7 (21.9%) 97 (10%) Granada 436 (5.1%) 400 (5.3%) 0 (0.0%) 36 (3.7%) Murcia 543 (6.3%) 505 (6.7%) 2 (6.2%) 36 (3.7%) Naples 114 (1.3%) 101 (1.3%) 0 (0.0%) 13 (1.3%) Subject Site ID Navarra 636 (7.4%) 588 (7.8%) 3 (9.4%) 45 (4.6%) Oxford 541 (6.3%) 466 (6.1%) 0 (0.0%) 75 (7.7%) Ragusa 299 (3.5%) 268 (3.5%) 2 (6.2%) 29 (3%) San Sebastian 603 (7%) 550 (7.3%) 0 (0.0%) 53 (5.4%) Turin 585 (6.8%) 511 (6.7%) 4 (12.5%) 70 (7.2%) Utrecht 1051 (12.2%) 928 (12.2%) 1 (3.1%) 122 (12.5%) Varese 514 (6%) 414 (5.5%) 7 (21.9%) 93 (9.5%) Mean (SD) 55.698 (9.09) 55.336 (9.067) 57.969 58.437(9.107) (8.793) Subject Age (years) Median 56 56 60 59Range 35 - 75 35 - 75 39 - 72 35 - 75 Body Mass Index Mean (SD) 27.148 (4.378) 27.192 (4.394) 27.135 26.804 (4.261) (kg / mA2) (3.838)Median 26.62 26.675 26.78 26.18Range 15.63 - 59.99 15.63 - 59.99 21.07 - 17.19 - 47.5737.45Mean (SD) 13.72 (20.875) 13.771 18.034 13.177(20.952) (21.454) (20.245) AlcoholConsumption Rate(g / d)Median 4.808 4.84 13.232 4.763Range 0 - 217.176 0 - 217.176 0 - 68.47 0 - 144.006Mean (SD) 4422.532 4875.394 950.812 1012.267(1761.569) (1315.011) (508.717) (515.459) Time-to- Event / Censoring(days)Median 4971 5218.5 963.5 1050.5Range 7 - 7378 1826 - 7378 68 - 1792 7 - 1825Table 7. Demographic information for the model development (verification) dataset for the full followup period (up to 20.4 years) for the Stomach Cancer risk test.Malignant Stomach Cancer Covariate Measure Total DiagnosisNo Yes Sample Size N 2862 2811 51Female 1569 (54.8%) 1541 (54.8%) 28 (54.9%) Subject GenderMale 1293 (45.2%) 1270 (45.2%) 23 (45.1%) No 2693 (94.1%) 2645 (94.1%) 48 (94.1%) Type II Diabetes Status Yes 148 (5.2%) 146 (5.2%) 2 (3.9%)Unknown 21 (0.7%) 20 (0.7%) 1 (2%) Current 785 (27.4%) 771 (27.4%) 14 (27.5%) Tobacco Use Status Never 1289 (45%) 1268 (45.1%) 21 (41.2%)Past 788 (27.5%) 772 (27.5%) 16 (31.4%) Asturias 177 (6.2%) 177 (6.3%) 0 (0.0%)Bilthoven 200 (7%) 197 (7%) 3 (5.9%)Cambridge 472 (16.5%) 463 (16.5%) 9 (17.6%) Florence 215 (7.5%) 208 (7.4%) 7 (13.7%) Granada 147 (5.1%) 144 (5.1%) 3 (5.9%) Murcia 193 (6.7%) 191 (6.8%) 2 (3.9%) Naples 42 (1.5%) 41 (1.5%) 1 (2%) Navarra 221 (7.7%) 218 (7.8%) 3 (5.9%) Subject Site IDOxford 182 (6.4%) 178 (6.3%) 4 (7.8%) Ragusa 94 (3.3%) 91 (3.2%) 3 (5.9%) San Sebastian 216 (7.5%) 211 (7.5%) 5 (9.8%) Turin 179 (6.3%) 178 (6.3%) 1 (2%) Utrecht 323 (11.3%) 321 (11.4%) 2 (3.9%) Varese 201 (7%) 193 (6.9%) 8 (15.7%) Mean (SD) 55.661 (8.983) 55.629 (8.992) 57.392 (8.412) Subject Age (years) Median 56 56 58Range 35 - 75 35 - 75 38 - 74 Mean (SD) 27.161 (4.574) 27.158 (4.577) 27.305 (4.445) Body Mass Index(kg / mA2) Median 26.58 26.57 27.03Range 15.73 - 65.02 15.73 - 65.02 18.34 - 42.6 Alcohol Consumption Mean (SD) 14.101 (21.952) 14.159 (22.016) 10.944 (17.969) Rate (g / d)Median 4.798 4.807 3.602 Range 0 - 256.231 0 - 256.231 0 - 87.739 Mean (SD) 4385.503 4403.759 3379.294(1806.333) (1803.185) (1707.428) Time-to-Event / Censoring(days) Median 4948 4986 3492Range 5 - 7454 10 - 7454 5 - 6302Table 8. Demographic information for the model development (verification) dataset for 5 -year (1,825- day) censoring for the Stomach Cancer risk test.Malignant Stomach Cancer Diagnosis Covariate Measure TotalNo Yes Censored Sample Size N 2862 2516 10 336Female 1569 (54.8%) 1377 (54.7%) 6 (60%) 186 (55.4%) Subject GenderMale 1293 (45.2%) 1139 (45.3%) 4 (40%) 150 (44.6%) No 2693 (94.1%) 2370 (94.2%) 9 (90%) 314 (93.5%) Type II DiabetesStatus Yes 148 (5.2%) 129 (5.1%) 0 (0.0%) 19 (5.7%)Unknown 21 (0.7%) 17 (0.7%) 1 (10%) 3 (0.9%) Current 785 (27.4%) 693 (27.5%) 3 (30%) 89 (26.5%) Tobacco UseStatus Never 1289 (45%) 1144 (45.5%) 3 (30%) 142 (42.3%)Past 788 (27.5%) 679 (27%) 4 (40%) 105 (31.2%) Asturias 177 (6.2%) 167 (6.6%) 0 (0.0%) 10 (3%) Bilthoven 200 (7%) 180 (7.2%) 0 (0.0%) 20 (6%) Subject Site ID Cambridge 472 (16.5%) 391 (15.5%) 3 (30%) 78 (23.2%)Florence 215 (7.5%) 175 (7%) 3 (30%) 37 (11%) Granada 147 (5.1%) 141 (5.6%) 0 (0.0%) 6 (1.8%) Murcia 193 (6.7%) 182 (7.2%) 0 (0.0%) 11 (3.3%) Naples 42 (1.5%) 39 (1.6%) 1 (10%) 2 (0.6%) Navarra 221 (7.7%) 204 (8.1%) 0 (0.0%) 17 (5.1%) Oxford 182 (6.4%) 153 (6.1%) 1 (10%) 28 (8.3%) Ragusa 94 (3.3%) 85 (3.4%) 0 (0.0%) 9 (2.7%) San Sebastian 216 (7.5%) 204 (8.1%) 0 (0.0%) 12 (3.6%) Turin 179 (6.3%) 157 (6.2%) 0 (0.0%) 22 (6.5%) Utrecht 323 (11.3%) 278 (11%) 1 (10%) 44 (13.1%) Varese 201 (7%) 160 (6.4%) 1 (10%) 40 (11.9%) Mean (SD) 55.661 (8.983) 55.403 (8.997) 57.6 (7.734) 57.53 (8.707) Subject Age Median 56 56 59 58 (years)Range 35 - 75 35 - 75 42 - 69 35 - 75 Mean (SD) 27.16 27.25 25.24 26.53(4.57) (4.61) (4.10) (4.26) Body Mass Index Median 26.58 26.66 25.42 25.85 (kg / mA2)Range 15.73 - 65.02 15.73 - 65.02 18.34 -31.03 16.74 - 44.34Mean (SD) 14.10 14.123 11.41 14.02 Alcohol Consumption (21.95) (22.30) (17.53) (19.32) Rate(g / d) Median 4.80 4.71 6.08 5.67Range 0 - 256.23 0 - 256.23 0 - 57.86 0 - 117.74 Time-to- Event / Censoring Mean (SD) 4385.50 4856.03 919 965.32 (days) (1806.33) (1357.67) (682.38) (510.73) Median 4948 5200.5 1267 1021.5 Range 5 - 7454 1826 - 7454 5 - 1669 10 - 1822

[0311] Model validation data: Demographic information for the validation dataset is provided for the full follow-up period (20.4 years) and for 5-year (1,825-day) censoring in Tables 9 and 10, respectively.Table 9. Demographic information for the validation dataset for the full follow-up period (up to 20.4 years) for the Stomach Cancer risk test.Malignant Stomach Cancer Covariate Measure Total DiagnosisNo Yes Sample Size N 2862 2816 46Female 1627 (56.8%) 1603 (56.9%) 24 (52.2%) Subject GenderMale 1235 (43.2%) 1213 (43.1%) 22 (47.8%) No 2715 (94.9%) 2671 (94.9%) 44 (95.7%) Type II Diabetes Status Yes 125 (4.4%) 123 (4.4%) 2 (4.3%)Unknown 22 (0.8%) 22 (0.8%) 0 (0.0%) Current 765 (26.7%) 749 (26.6%) 16 (34.8%) Tobacco Use Status Never 1267 (44.3%) 1247 (44.3%) 20 (43.5%) Past 830 (29%) 820 (29.1%) 10 (21.7%) Asturias 190 (6.6%) 185 (6.6%) 5 (10.9%) Bilthoven 195 (6.8%) 192 (6.8%) 3 (6.5%) Cambridge 476 (16.6%) 471 (16.7%) 5 (10.9%) Subject Site ID Florence 215 (7.5%) 212 (7.5%) 3 (6.5%)Granada 158 (5.5%) 157 (5.6%) 1 (2.2%) Murcia 202 (7.1%) 199 (7.1%) 3 (6.5%) Naples 41 (1.4%) 40 (1.4%) 1 (2.2%) Navarra 188 (6.6%) 182 (6.5%) 6 (13%) Oxford 197 (6.9%) 195 (6.9%) 2 (4.3%) Ragusa 111 (3.9%) 109 (3.9%) 2 (4.3%)San Sebastian 188 (6.6%) 187 (6.6%) 1 (2.2%) Turin 166 (5.8%) 165 (5.9%) 1 (2.2%) Utrecht 344 (12%) 341 (12.1%) 3 (6.5%) Varese 191 (6.7%) 181 (6.4%) 10 (21.7%) Mean (SD) 55.65 (9.19) 55.63 (9.20) 56.63 (8.731) Subject Age (years) Median 56 56 58.5Range 35 - 75 35 - 75 36 - 73 Mean (SD) 27.13 (4.63) 27.15 (4.64) 26.44 (3.91) Body Mass Index (kg / mA2) Median 26.59 26.59 26.45Range 15.73 - 67.86 15.73 - 67.86 18.21 - 33.65 Mean (SD) 13.47 (21.04) 13.46 (21.04) 13.99 (21.42) Alcohol Consumption Rate(g / d) Median 4.301 4.316 2.353Range 0 - 215.073 0 - 215.073 0 - 88.407 Mean (SD) 4417.72 4440.16 3044.22 (1756.46) (1747.46) (1778.88) Time-to-Event / Censoring(days) Median 4944 4969 2776Range 8 - 7156 8 - 7156 238 - 6783Table 10. Demographic information for the validation dataset for 5-year (1,825-day) censoring for the StomachCancer risk test.Malignant Stomach Cancer Diagnosis Covariate Measure TotalNo Yes Censored Sample Size N 2862 2538 14 310Female 1627 (56.8%) 1458 (57.4%) 5 (35.7%) 164 (52.9%) Subject GenderMale 1235 (43.2%) 1080 (42.6%) 9 (64.3%) 146 (47.1%) No 2715 (94.9%) 2414 (95.1%) 13 (92.9%) 288 (92.9%) Type II DiabetesStatus Yes 125 (4.4%) 109 (4.3%) 1 (7.1%) 15 (4.8%)Unknown 22 (0.8%) 15 (0.6%) 0 (0.0%) 7 (2.3%) Current 765 (26.7%) 689 (27.1%) 3 (21.4%) 73 (23.5%) Tobacco UseStatus Never 1267 (44.3%) 1137 (44.8%) 7 (50%) 123 (39.7%) Past 830 (29%) 712 (28.1%) 4 (28.6%) 114 (36.8%) Asturias 190 (6.6%) 182 (7.2%) 2 (14.3%) 6 (1.9%) Bilthoven 195 (6.8%) 177 (7%) 1 (7.1%) 17 (5.5%)Cambridge 476 (16.6%) 417 (16.4%) 2 (14.3%) 57 (18.4%) Florence 215 (7.5%) 179 (7.1%) 2 (14.3%) 34 (11%) Granada 158 (5.5%) 137 (5.4%) 0 (0.0%) 21 (6.8%) Murcia 202 (7.1%) 189 (7.4%) 0 (0.0%) 13 (4.2%) Naples 41 (1.4%) 37 (1.5%) 0 (0.0%) 4 (1.3%) Subject Site ID Navarra 188 (6.6%) 167 (6.6%) 3 (21.4%) 18 (5.8%)Oxford 197 (6.9%) 165 (6.5%) 0 (0.0%) 32 (10.3%) Ragusa 111 (3.9%) 105 (4.1%) 0 (0.0%) 6 (1.9%)San Sebastian 188 (6.6%) 176 (6.9%) 0 (0.0%) 12 (3.9%) Turin 166 (5.8%) 142 (5.6%) 0 (0.0%) 24 (7.7%) Utrecht 344 (12%) 312 (12.3%) 0 (0.0%) 32 (10.3%) Varese 191 (6.7%) 153 (6%) 4 (28.6%) 34 (11%) Mean (SD) 55.65 55.38 58.21 57.73(9.19) (9.18) (8.74) (9.02) Subject AgeMedian 56 56 59.5 58(years)Range 35 - 75 35 - 75 36 - 70 36 - 75Mean (SD) 27.13 27.18 27.48 26.75(4.63) (4.65) (4.06) (4.48)Body Mass IndexMedian 26.59 26.61 28.67 25.975 (kg / mA2)Range 15.73 - 67.86 15.73 - 67.86 18.21 - 33.65 17.13 - 48.79 Mean (SD) 13.47 13.32 17.98 14.51 Alcohol (21.04) (20.92) (19.27) (22.12) Consumption Rate Median 4.301 4.047 13.993 5.71(g / d)Range 0 - 215.073 0 - 215.073 0 - 61.789 0 - 206.4 Mean (SD) 4417.72 4857.77 1108.71 964.40Time-to- (1756.46) (1315.89) (449.34) (537.53) Event / Censoring Median 4944 5182 1144 1011.5(days)Range 8 - 7156 1829 - 7156 238 - 1696 8 - 1821

[0312] Data Quality Control and Pre- Analytics Results: The original clinical dataset included 14,787 samples with both clinical and proteomic data. A total of 16 individuals who did not have inclusion probability coefficients were excluded from the dataset leaving 14,771 subjects for data quality control (QC).

[0313] Data QC showed that there were 268 (1.81%) outlier samples, defined as >5% of analytes exceeding 6 median absolute deviations from the median, and there were 211 samples with normalization scale factors outside the recommended range. Table 11 details the samples removed at each step of data cleaning.Table 11. Number of individuals removed at each step of data cleaning for the EPIC datasetStep N (removed) N (remaining)Baseline0 14,787(samples with clinical and proteomic data)Missing sample weights 16 14,771Outside range normalization scale factors 211 14,560Outliers* 250 14,310*18 samples were outliers and outside the range normalization factors

[0314] Additionally, 363 analytes were removed before analysis began as they did not pass target confirmation specificity testing (i.e., red-listed features). After removing 477 samples and 323 analytes for data QC purposes, 14,310 (96.8%) samples with 7,233 analytes were included for analysis. No other issues were identified during Data QC or Pre- Analytics.

[0315] Refinement Approach and Results: The selected model for the 5-year stomach cancer risk test is an 8-protein AFT survival model using a Weibull distribution. The model was developed using the training split (60%) of the EPIC dataset, tested using the verification split (20%), and will be validated on a final held-out dataset (20%). The primary model output is absolute risk (i.e., probability of malignant stomach cancer diagnosis) within five years of blood draw.

[0316] Various approaches for feature selection were considered, including using top ranked features (FDR < 0.1) from univariate Cox and AFT Weibull that were run with and without weights as initial feature list. Consensus Nested Cross-Validation (CNCV)20 with elastic net cox regression models was carried out with and without case-cohort weights for comparison. Elastic net Cox regression models were used during refinement and used in feature selection decisions due to their flexibility and quick estimation routines despite the large size of the training dataset. Consensus features from CNCV with non-zero coefficients were used to fit a weighted AFT Weibull model since these models are preferred to Cox models due to easier interpretation of model outputs.

[0317] For the selected model, a union of features with FDR < 0.1 from univariate weighted Cox and AFT Weibull (132 analytes) was considered for feature selection. Analytes with EDTA-based duplicate run CCC < 0.8 were excluded (35 analytes excluded). This exclusion was done due to a consistent pattern seen during Soma Model Assessment (SMA) of candidatemodels having lower than desirable duplicate run concordance. Thresholds for individual concordance were chosen to minimize number of analytes being excluded from the starting list while still resulting in models that pass the duplicate run concordance criteria.

[0318] Candidate models repeatedly failed to be statistically significantly better than the age-only model in verification, due to the low number of events. To increase the chances of selection of analytes that are highly correlated with age, penalty factors were used with the remaining 97 features. Specifically, out of the 14 analytes that were present both in the age plasma SST and the 97 initial feature list, the protein with the largest coefficient in the age model (CILP2) was chosen to have a low penalty factor during feature selection increasing the likelihood of its selection. A penalty factor of 1 was used for all 97 analytes except CILP2 which had a penalty factor of 0. Thus, 97 analytes were used as the input feature list for CNCV with 3 outer folds and 2 inner folds. Within each inner fold, an unweighted penalized (Ridge) Cox regression model with 3-fold cross-validation was used to calculate the feature importance. 60% of the top features were selected in each inner fold and a consensus (intersection) of these features between the inner and outer folds was used to select the feature set.

[0319] Performance metrics for the training and verification data assessed at five years are presented in Table 12. With a 5-year AUC of 0.771 and 0.739 in the training and verification data, respectively, the selected model exceeded the first passing criterion of a 5-year AUC greater than or equal to 0.629. Note that the decision threshold for the proteomic and age-only model is derived based on the cutoff that maximized Youden's J on the training dataset for each model.Table 12. Model performance in predicting first primary malignant stomach cancer risk within 5 years of blood draw in cancer-free individuals, along with age-only comparator model. The AUC, sensitivity, and specificity were calculated at 5 years (1825 days). CI = confidence interval; AUC = area under the curve; PEC = predictive error curve.AUC at 5 Sensitivity at Specificity at C-index2PEC at 5 Model Dataset N years 5 years15 years1(95% CI) years (95%(95% CI) (95% CI) (95% CI) CI) Training 8586 0.771 0.750 0.727 0.714 0.004 Proteomic (0.683, (0.583, 0.914) (0.714, 0.739) (0.669, (0.003, 0.857) 0.757) 0.006) Verification 2862 0.739 0.700 0.718 0.634 0.004(0.594, (0.400, 1.000) (0.699, 0.731) (0.548, (0.001, 0.902) 0.726) 0.006) Training 8586 0.600 0.719 0.572 0.587 0.004 Age-only3(0.498, (0.577, 0.858) (0.562, 0.585) (0.537, (0.003, 0.692) 0.636) 0.006) Verification 2862 0.577 0.700 0.573 0.566 0.004(0.430, (0.427, 1.000) (0.556, 0.596) (0.491, (0.001,0.739) 0.660) 0.006) ’Sensitivity and specificity calculated using the linear-space cutoff that maximized Youden's J on the training dataset; cutoff = -13.126 (Proteomic) and cutoff = -13.15 (Age-only).2C-index is not time specific3Model includes only participant age at the time of blood draw as a covariate.

[0320] The stomach cancer test was also required to have statistically significantly greater 5-year AUC (p < 0.1, by paired one-sided DeLong’s Test) than two comparator survival models: 1) only an intercept term (“Intercept-only”) model, and 2) a model with only an individual’s age (“Age-only”) at the time of blood draw. The selected proteomic model also passed the comparator model comparison passing criteria with a statistically significantly greater 5-year AUC than the age-only and intercept-only model (Table 13).Table 13. Model performance in predicting first primary malignant stomach cancer risk within 5 years of blood draw in cancer-free individuals. The Area Under the receiver operating characteristic Curve (AUC) was calculated at 5 years (1,825 days). P-values were derived from the Paired DeLong’s test statistic, based on the Observed Difference in 5-year AUC between Proteomic and Comparator models. CI = Confidence interval; AUC = area under the curve.5-year AUC (95% CI) Observed DifferenceDataset P-value1Proteomic Model Comparator ModelIntercept-only0.500 0.271 1.739e-10 Training 0.771 (0.500, 0.500)(0.683, 0.857)Age-only0.600 0.171 1.657e-03 (0.498, 0.692)Intercept-only0.500 0.239 1.384e-03 Verification 0.739 (0.500, 0.500)(0.594, 0.902)Age-only0.577 0.162 3.763e-02 (0.430, 0.739)'Calculated from a paired DeLong’s test for differences in 5-year AUC between proteomic and comparator models (p-value threshold = 0.1). Null hypothesis: no difference in AUC between proteomic and comparator model. Alternative hypothesis: proteomic model AUC is greater than comparator model AUC.

[0321] Clinical Validation Plan: Validation is assessed on the remaining 20% hold out portion of the EPIC dataset that has not been used up to this point (see Tables 9 and 10 above).

[0322] The selected model from refinement is an 8-feature Accelerated Failure Time (AFT) model with a Weibull distribution fit with case-cohort weights. Predicted probabilities for risk of malignant stomach cancer at 5 years post blood draw (1,825 days) are calculated using this final model, and 5-year AUC, sensitivity, specificity, and dynamic range, in addition to C-index and PEC, that are calculated. The cutoff that maximized Youden’s J statistic on the training dataset (linear cutoff = -13.126) is used to calculate sensitivity and specificity. These performance metrics are also calculated for the Age-only model. The passing criteria is a 5-year AUC greater than or equal to 0.629 and statistically significantly greater 5-year AUC than the Age-only and Intercept-only models, using a p-value threshold of 0.1 to determine statistical significance.

[0323] Clinical Results on Validation Data: The performance metrics of the stomach cancer risk model on the validation dataset are shown in Table 14. The model had a 5-year AUC of 0.742, exceeding the fixed performance criteria of a 5-year AUC of at least 0.629, and was statistically significantly greater than the Intercept-only and Age-only comparator model 5-year AUCs of 0.500 and 0.604, respectively (Table 15).Table 14. Performance metrics for the selected proteomic stomach cancer risk model and the age-only model on the validation dataset. AUC, sensitivity, and specificity were calculated at 5 years. Sensitivity and specificity were calculated with the decision cutoff that maximized Youden's J index for the respective model. AUC = area under the curve; CI = confidence interval; PEC = predictive error curve.AUC at 5 Sensitivity at Specificity at C-index PEC at 5 Model N Dataset years 5 years 5 years (95% CI)2years (95%(95% CI) (95% CI)1(95% CI)1CI) Proteomic 2,862 Validation 0.742 0.714 0.737 0.697 0.005(0.523, (0.411, 0.917) (0.716, 0.757) (0.562, (0.003, 0.874) 0.784) 0.009) Age-only32,862 Validation 0.604 0.643 0.564 0.561 0.005(0.455, (0.333, 0.900) (0.547, 0.584) (0.437, (0.003,0.740) 0.686) 0.009) Sensitivity and specificity calculated using the cutoff that maximized Youden's J on the training dataset; linear space cutoff = -13.126 (Proteomic) and linear space cutoff = -13.15 (Age-only).^C-index is not time specificModel includes only participant age at the time of blood draw as a covariate.Table 15. Model performance in predicting first primary malignant stomach cancer risk within 5 years of blood draw in cancer-free individuals. The AUC was calculated at 5 years (1,825 days). P-values were derived from the Paired DeLong's test statistic, based on the Observed Difference in 5-year AUC between Proteomic and Comparator models. AUC = area under curve; CI = Confidence interval.5-Year AUC (95% CI)Proteomic Model Comparator Model Observed Difference Dataset P-Value1Validation 0.742 Intercept-only 0.242 1.275e-03 (0.523, 0.874) 0.500(0.500, 0.500)Age-only 0.138 8.689e-02 0.604(0.455, 0.740)p- values calculated from Delong’s test for differences in AUC with a one-sided alternative that the proteomic 5-year AUC is greater than the comparator model 5-year AUC.

[0324] The Kaplan-Meier survival (“Event-free”) curves with subjects in the validation dataset, stratified by risk bin, were visualized for quartiles (FIG. 10). In validation, the KM curve for the highest risk bin (Q4) is well-separated from the lower risk bins, similar to training and verification datasets (FIG. 8). The shaded region about each survival curve is the 95% confidence interval (CI) of the KM estimate of the event-free probability.

[0325] Impacts of Imputation: Imputation methods were assessed. Acceptable RFU ranges (extreme value bounds) for model aptamers are calculated. Then, the two imputation methods are applied to the validation dataset: winsorization and feature averaging (zero replacement). Model predictions are made on the original validation dataset without imputation and on each imputed validation dataset. The concordance (Lin’s CCC) in model predictions between unimputed and imputed validation datasets is calculated. The imputation approach with the highest CCC was zero replacement (Table 16). For this reason, the average, zero, for imputation is used.Table 16. Imputation table for out-of-range RFU values in thevalidation dataset.Imputation Method CCCOriginal validation 1Winsorized validation 0.930Feature Average / Zero 1.000Replacement validation

[0326] Predictive variation due to assay noise: Ninety-five percent, 99%, and 99.99% lower and upper tolerance bounds for variation in model predictions due to the assay are given in Table 17.Table 17. Tolerance bounds in response and linear predictor space for the final 5-year stomach cancer risk model. Tolerance bounds were transformed into linear space with the exact linear to response spacetransformation and the maximum width across possible values is reported. _ _Prediction Space Bound 95% 99% 99.99%lower -0.204 -0.268 -0.405Linearupper 0.204 0.268 0.405lower -0.075 -0.099 -0.149 Responseupper 0.075 0.099 0.149

[0327] EDTA-Citrate Plasma Concordance: The EPIC data used for training were collected in Sarstedt 3.2% sodium citrate tubes. However, future samples may be collected in alternative collection tubes and buffers, such as, but not limited to EDTA plasma tubes. Thus, the model predictions were tested for concordance between Citrate and EDTA plasma. The training data were jittered according to the differences between matched samples from an internal study whose blood samples were collected with both 3.2% sodium Citrate plasma and EDTA plasma tubes. FIG. 11 shows a high concordance (CCC > 0.95) between original training (Citrate) and jittered-training (EDTA) predictions indicating that this stomach cancer risk test can be used with both Citrate and EDTA plasma samples without a bridge. Specifically, using the dataset comparison approach, with a jittered-training vs training CCC > 0.95, the 5-year stomach cancer risk test is supported statistically for use in EDTA plasma.

[0328] Validation Conclusions: The 8-analyte malignant stomach cancer risk model predicts the risk of a future malignant stomach cancer diagnosis within 5 years of blood draw in cancer-free individuals. The model output is the probability (absolute risk) of a first primary malignant stomach cancer diagnosis, which is a continuous variable within the range from 0.0000 to1.0000. Validation exceeds the performance criteria of 5-year AUC greater than or equal to 0.629 which is also statistically significantly greater than the 5-year AUC of models that include 1) only the intercept, and 2) only participant age at blood draw. Albumin (2000 mg / dl) did not pass interference testing. Based on simulations studying the impacts of imputation, feature average (i.e., zero) as a replacement was determined to be the best approach for imputing analytes that are out of bounds for this model. Additionally, the model was assessed for use with EDTA plasma samples and the 5-year stomach cancer risk test is supported statistically for use in both EDTA and Citrate plasma samples without a bridge.Example 3: Analysis of Stomach Cancer Risk Model Biomarker Panels

[0329] Model biomarker panels comprising various combinations of the biomarkers listed in Table 1 were analyzed to determine the Area Under the Curve (AUC) value for the various combinations. The model biomarker panels may be based on a panel of N biomarker proteins having an AUC value of at least 0.62, 0.63, 0.64, at least 0.65, at least 0.66, at least 0.67, at least 0.68, at least 0.69, at least 0.7, at least 0.75, at least 0.8, at least 0.85, at least 0.9, or at least 0.95, where N is 1, 2, 3, 4, 5, 6, 7, and / or 8 of the biomarker proteins listed in Table 1. The Tables below shows exemplary model results when various combinations comprising 1 to 8 biomarker proteins were measured.Table 18: Single BiomarkersTarget_Names S27A2 CILP2 TFF1 GKN2 Inhibin bB chain LYPL1 EphAl PAP1 AUC PAP1 PAP1 0.622092 EphAl EphAl 0.567088 LYPL1 LYPL1 0.559788 Inhibin bB chain Inhibin bB chain 0.586524 GKN2 GKN2 0.645251 TFF1 TFF1 0.645761 CILP2 CILP2 0.517684 S27A2 S27A2 0.562525Table 19: Panels including biomarker LYPL1AUC S27A2 CILP2 TFF1 GKN2 Inhibin bB LYPL1 EphAl PAP1 chain0.559788 LYPL10.712844 Inhibin bB LYPL1chain0.650631 LYPL1 PAP1 0.64764 LYPL1 EphAl 0.727325 GKN2 LYPL10.572516 CILP2 LYPL10.645815 S27A2 LYPL10.723372 TFF1 LYPL10.71432 Inhibin bB LYPL1 PAP1 chain0.71908 Inhibin bB LYPL1 EphAlchain0.753224 GKN2 Inhibin bB LYPL1chain0.691821 CILP2 Inhibin bB LYPL1chain0.756625 S27A2 Inhibin bB LYPL1chain0.758132 TFF1 Inhibin bB LYPL1chain0.684874 LYPL1 EphAl PAP1 0.72346 GKN2 LYPL1 PAP1 0.65001 CILP2 LYPL1 PAP1 0.713257 S27A2 LYPL1 PAP1 0.711387 TFF1 LYPL1 PAP1 0.739124 GKN2 LYPL1 EphAl 0.633649 CILP2 LYPL1 EphAl 0.719014 S27A2 LYPL1 EphAl 0.743956 TFF1 LYPL1 EphAl 0.712005 CILP2 GKN2 LYPL10.770409 S27A2 GKN2 LYPL10.745583 TFF1 GKN2 LYPL1LYPL1LYPL1LYPL1Inhibit! bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1chainInhibin bB LYPL1chainInhibin bB LYPL1chainInhibin bB LYPL1chainInhibin bB LYPL1chainInhibin bB LYPL1chainLYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 EphAl LYPL1 EphAl LYPL1 EphAl LYPL1 EphAl LYPL1 EphAl LYPL1 EphAl LYPL1LYPL1LYPL1LYPL1Inhibit! bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1chainInhibin bB LYPL1chainInhibin bB LYPL1chainInhibin bB LYPL1chainLYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 EphAl LYPL1 EphAl LYPL1 EphAl LYPL1 EphAl LYPL10.746426 CILP2 GKN2 Inhibit! bB LYPL1 EphAl PAP1 chain0.801367 S27A2 GKN2 Inhibin bB LYPL1 EphAl PAP1 chain0.763217 TFF1 GKN2 Inhibin bB LYPL1 EphAl PAP1 chain0.770999 S27A2 CILP2 Inhibin bB LYPL1 EphAl PAP1 chain0.746154 CILP2 TFF1 Inhibin bB LYPL1 EphAl PAP1 chain0.794037 S27A2 TFF1 Inhibin bB LYPL1 EphAl PAP1 chain0.78724 S27A2 CILP2 GKN2 Inhibin bB LYPL1 PAP1 chain0.751653 CILP2 TFF1 GKN2 Inhibin bB LYPL1 PAP1 chain0.800989 S27A2 TFF1 GKN2 Inhibin bB LYPL1 PAP1 chain0.781203 S27A2 CILP2 TFF1 Inhibin bB LYPL1 PAP1 chain0.785422 S27A2 CILP2 GKN2 Inhibin bB LYPL1 EphAl chain0.761962 CILP2 TFF1 GKN2 Inhibin bB LYPL1 EphAl chain0.811626 S27A2 TFF1 GKN2 Inhibin bB LYPL1 EphAl chain0.792278 S27A2 CILP2 TFF1 Inhibin bB LYPL1 EphAl chain0.794731 S27A2 CILP2 TFF1 GKN2 Inhibin bB LYPL1chain0.783497 S27A2 CILP2 GKN2 LYPL1 EphAl PAP1 0.743659 CILP2 TFF1 GKN2 LYPL1 EphAl PAP1 0.796036 S27A2 TFF1 GKN2 LYPL1 EphAl PAP1 0.775522 S27A2 CILP2 TFF1 LYPL1 EphAl PAP1 0.775365 S27A2 CILP2 TFF1 GKN2 LYPL1 PAP1 0.792084 S27A2 CILP2 TFF1 GKN2 LYPL1 EphAl 0.792493 S27A2 CILP2 GKN2 Inhibin bB LYPL1 EphAl PAP1 chain0.757495 CILP2 TFF1 GKN2 Inhibin bB LYPL1 EphAl PAP1 chain0.807815 S27A2 TFF1 GKN2 Inhibin bB LYPL1 EphAl PAP1 chain0.789508 S27A2 CILP2 TFF1 Inhibin bB LYPL1 EphAl PAP1 chain0.793663 S27A2 CILP2 TFF1 GKN2 Inhibin bB LYPL1 PAP1 chain0.800969 S27A2 CILP2 TFF1 GKN2 Inhibin bB LYPL1 EphAl chain0.790909 S27A2 CILP2 TFF1 GKN2 LYPL1 EphAl PAP1 0.800528 S27A2 CILP2 TFF1 GKN2 Inhibin bB LYPL1 EphAl PAP1chainTable 20: Panels including biomarker CILP2AUC S27A2 CILP2 TFF1 GKN2 Inhibin bB LYPL1 EphAl PAP1 chain0.517684 CILP2LYPL1Inhibit! bB chainPAP1 EphAlInhibin bB LYPL1chainLYPL1 PAP1 LYPL1 EphAl LYPL1LYPL1LYPL1Inhibin bB chain PAP1 Inhibin bB chain EphAl Inhibin bB chainInhibin bB chainInhibin bB chainEphAl PAP1PAP1 PAP1 PAP1 EphAl EphAl EphAlInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1chainInhibin bB LYPL1chainInhibin bB LYPL1chainLYPL1 EphAl PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 EphAl LYPL1 EphAl LYPL1 EphAl LYPL1LYPL1LYPL1Inhibit! bB chain EphAl PAP1Inhibin bB chain PAP1 Inhibin bB chain PAP1 Inhibin bB chain PAP1 Inhibin bB chain EphAl Inhibin bB chain EphAl Inhibin bB chain EphAl Inhibin bB chainInhibin bB chainInhibin bB chainEphAl PAP1 EphAl PAP1 EphAl PAP1PAP1 PAP1 PAP1 EphAl EphAl EphAlInhibin bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1chainInhibin bB LYPL1chainInhibin bB LYPL1chainLYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 PAP1 LYPL1 EphAl LYPL1 EphAl LYPL1 EphAl LYPL1Inhibin bB chain EphAl PAP1Inhibit! bB chain EphAl PAP1Inhibin bB chain EphAl PAP1 Inhibin bB chain PAP1 Inhibin bB chain PAP1 Inhibin bB chain PAP1 Inhibin bB chain EphAl Inhibin bB chain EphAl Inhibin bB chain EphAl Inhibin bB chainEphAl PAP1 EphAl PAP1 EphAl PAP1PAP1 EphAl Inhibin bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 PAP1 chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1 EphAl chainInhibin bB LYPL1chainLYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 EphAl PAP1 LYPL1 PAP1 LYPL1 EphAl Inhibin bB chain EphAl PAP1 Inhibin bB chain EphAl PAP1 Inhibin bB chain EphAl PAP1 Inhibin bB chain PAP1 Inhibin bB chain EphAl EphAl PAP1 Inhibin bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 EphAl PAP1 chainInhibin bB LYPL1 PAP1chain0.800969 S27A2 CILP2 TFF1 GKN2 Inhibit! bB LYPL1 EphAlchain0.790909 S27A2 CILP2 TFF1 GKN2 LYPL1 EphAl PAP1 0.788347 S27A2 CILP2 TFF1 GKN2 Inhibin bB chain EphAl PAP1 0.800528 S27A2 CILP2 TFF1 GKN2 Inhibin bB LYPL1 EphAl PAP1chain

[0330] The complete disclosure of all patents, patent applications, and publications, and electronically available material (including, for instance, nucleotide sequence submissions in, e.g., GenBank and RefSeq, and amino acid sequence submissions in, e.g., SwissProt, PIR, PRF, PDB, and translations from annotated coding regions in GenBank and RefSeq) cited herein are incorporated by reference in their entirety. Supplementary materials referenced in publications (such as supplementary tables, supplementary figures, supplementary materials and methods, and / or supplementary experimental data) are likewise incorporated by reference in their entirety. In the event that any inconsistency exists between the disclosure of the present application and the disclosure(s) of any document incorporated herein by reference, the disclosure of the present application shall govern. The foregoing detailed description and examples have been given for clarity of understanding only. No unnecessary limitations are to be understood therefrom. The disclosure is not limited to the exact details shown and described, for variations obvious to one skilled in the art will be included within the disclosure defined by the claims.

[0331] Unless otherwise indicated, all numbers expressing quantities of components, molecular weights, and so forth used in the specification and claims are to be understood as being modified in all instances by the term "about." Accordingly, unless otherwise indicated to the contrary, the numerical parameters set forth in the specification and claims are approximations that may vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.

[0332] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. All numerical values, however, inherently contain a range necessarily resulting from the standard deviation found in their respective testing measurements.

[0333] All headings are for the convenience of the reader and should not be used to limit the meaning of the text that follows the heading, unless so specified.

Claims

What is claimed is:

1. A method of detecting levels of N biomarker proteins in a sample, comprising forming a biomarker panel comprising N biomarker proteins, and detecting the level of each of the N biomarker proteins in the sample from a subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or 8 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, InhibinbB chain, andEphAl.

2. A method of predicting stomach cancer risk in a subject, comprising detecting a level of LYPL1 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from CILP2, TFF1, S27A2, PAP1, GKN2, InhibinbB chain, andEphAl.

3. A method of predicting stomach cancer risk in a subject, comprising detecting a level of CILP2 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

4. A method of predicting stomach cancer risk in a subject, comprising detecting a level of TFF1 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

5. A method of predicting stomach cancer risk in a subject, comprising detecting a level of S27A2 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, PAP1, GKN2, InhibinbB chain, andEphAl.

6. A method of predicting stomach cancer risk in a subject, comprising detecting a level of PAP1 and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, GKN2, InhibinbB chain, andEphAl.

7. A method of predicting stomach cancer risk in a subject, comprising detecting a level of RPIA and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, InhibinbB chain, andEphAl.

8. A method of predicting stomach cancer risk in a subject, comprising detecting a level of Inhibin bB chain and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, and EphAl.

9. A method of predicting stomach cancer risk in a subject, comprising detecting a level of EphAl and a level of each of N biomarker proteins in a sample from the subject, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7, and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, and InhibinbB chain.

10. The method according to any one of the preceding claims, wherein N is 2 to 8, or N is 3 to 8, N is 4 to 8, N is 5 to 8, N is 6 to 8, or N is 7 to 8.

11. The method according to any one of the preceding claims, wherein N is 2, N is 3, N is 4, N is 5, N is 6, N is 7, or N is 8.

12. The method according to any one of the preceding claims, wherein at least one of the N biomarker proteins is LYPL1, or at least one of the N biomarker proteins is CILP2, or at least one of the N biomarker proteins is TFF1, or at least one of N biomarker proteins is S27A2, or at least one of the N biomarker proteins is PAP1, or at least one of the N biomarker proteins isGKN2, or at least one of the N biomarker proteins is Inhibin bB chain, or at least one of the N biomarker proteins is EphAl .

13. The method according to any one of the preceding claims, wherein each of the N biomarker proteins is selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

14. The method according to any one of the preceding claims, wherein at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 of the N biomarker proteins are selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, Inhibin bB chain, and EphAl.

15. The method according to any one of claims 1-3, wherein(a)2 of the N biomarker proteins are LYPL1 and CILP2, or 2 of the N biomarker proteins are LYPL1 and TFF1, or 2 of the N biomarker proteins are LYPL1 and S27A2, or 2 of the N biomarker proteins are LYPL1 and PAP1, or 2 of the N biomarker proteins are LYPL1 and GKN2, or 2 of the N biomarker proteins are LYPL1 and Inhibin bB chain, or 2 of the N biomarker proteins are LYPL1 and EphAl;(b) 2 of the N biomarker proteins are CILP2 and TFF1, or 2 of the N biomarker proteins are CILP2 and S27A2, or 2 of the N biomarker proteins are CILP2 and PAP1, or 2 of the N biomarker proteins are CILP2 and GKN2, or 2 of the N biomarker proteins are CILP2 and Inhibin bB chain, or 2 of the N biomarker proteins are CILP2 and EphAl;(c) 2 of the N biomarker proteins are TFF1 and S27A2, or 2 of the N biomarker proteins are TFF1 and PAP1, or 2 of the N biomarker proteins are TFF1 and GKN2, or 2 of the N biomarker proteins are TFF1 and Inhibin bB chain, or 2 of the N biomarker proteins are TFF1 and EphAl; (d) 2 of the N biomarker proteins are S27A2 and PAP1, or 2 of the N biomarker proteins are S27A2 and GKN2, or 2 of the N biomarker proteins are S27A2 and Inhibin bB chain, or 2 of the N biomarker proteins are S27A2 and EphAl;(e) 2 of the N biomarker proteins are PAP1 and GKN2, or 2 of the N biomarker proteins are PAP1 and Inhibin bB chain, or 2 of the N biomarker proteins are PAP1 and EphAl;(f) 2 of the N biomarker proteins are GKN2 and Inhibin bB chain, or 2 of the N biomarker proteins are GKN2 and EphAl; or(g) 2 of the N biomarker proteins are Inhibin bB chain and EphAl .

16. The method according to any one of claims 1 or 4, wherein 2 of the N biomarker proteins are TFF1 and S27A2, or 2 of the N biomarker proteins are TFF1 and PAP1, or 2 of the Nbiomarker proteins are TFF1 and GKN2, or 2 of the N biomarker proteins are TFF1 and Inhibin bB chain, or 2 of the N biomarker proteins are TFF1 and EphAl.

17. The method according to any one of claims 1 or 5, wherein 2 of the N biomarker proteins are S27A2 and PAP1, or 2 of the N biomarker proteins are S27A2 and GKN2, or 2 of the N biomarker proteins are S27A2 and Inhibin bB chain, or 2 of the N biomarker proteins are S27A2 and EphAl.

18. The method according to any one of claims lor 6, wherein 2 of the N biomarker proteins are PAP1 and GKN2, or 2 of the N biomarker proteins are PAP1 and Inhibin bB chain, or 2 of the N biomarker proteins are PAP1 and EphAl.

19. The method according to any one of claims lor 7, wherein 2 of the N biomarker proteins are GKN2 and Inhibin bB chain, or 2 of the N biomarker proteins are GKN2 and EphAl .

20. The method according to any one of claims lor 8, wherein 2 of the N biomarker proteins are Inhibin bB chain and EphAl.

21. The method according to any one of claims 1-3, wherein 3 of the N biomarker proteins are LYPL1, CILP2, and TFF1, or 3 of the N biomarker proteins are LYPL1, CILP2, and S27A2, or 3 of the N biomarker proteins are LYPL1, CILP2, and PAP1, or 3 of the N biomarker proteins are LYPL1, CILP2, and GKN2, or 3 of the N biomarker proteins are LYPL1, CILP2, and Inhibin bB chain, or 3 of the N biomarker proteins are LYPL1, CILP2, and EphAl.

22. The method according to any one of the preceding claims, wherein the sample is a blood sample, a plasma sample, a serum sample, or a urine sample.

23. The method according to any one of the preceding claims, wherein the risk of the subject stomach cancer within 5 years from the date that the sample was taken from the subject is predicted.

24. The method according to any one of the preceding claims, wherein detecting is performed using mass spectrometry, an aptamer based assay and / or an antibody based assay.

25. The method according to any one of the preceding claims, wherein the method comprises contacting biomarker proteins of the sample or samples with a set of biomarker capture reagents, wherein each biomarker capture reagent of the set of biomarker capture reagents specifically binds to a different biomarker protein being detected.

26. The method according to claim 25, whereineach biomarker capture reagent is an antibody or an aptamer; preferably wherein each biomarker capture reagent is an aptamer, optionally wherein at least one aptamer is a slow off-rate aptamer.

27. The method according to claim 26, wherein at least one slow off-rate aptamer comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least 10 nucleotides with modifications.

28. The method according to claim 26 or 27, wherein each slow off-rate aptamer binds to its target protein with an off rate (t’ ) of > 20 minutes, > 30 minutes, > 60 minutes, > 90 minutes, > 120 minutes, > 150 minutes, > 180 minutes, > 210 minutes, or > 240 minutes.

29. The method according to any one of claims 24-28, wherein the level of each biomarker protein measured is determined from a relative florescence unit (RFU) or a protein concentration.

30. The method according to any one of the preceding claims, wherein predicting the risk of stomach cancer in the subject is based on input of the levels of the N biomarker proteins measured in a statistical model.

31. The method according to claim 30, wherein the determining comprises analyzing the levels of the N biomarker protein using a survival model.

32. The method according to claim 31, wherein the survival model is an Accelerated Failure Time (AFT) model with a Weibull distribution.

33. The method according to any one of claims 30-32, wherein the model has an area under the curve (AUC) selected from at least 0.64, at least 0.65, at least 0.66, at least 0.67, at least 0.68, at least 0.69, at least 0.7, at least 0.75, at least 0.8, at least 0.85, at least 0.9, or at least 0.95.

34. The method according to any one of the preceding claims, wherein the method comprises predicting a risk of stomach cancer in a subject for the purpose of determining a medical insurance premium or life insurance premium.

35. The method according to claim 34, wherein the method further comprises determining coverage for medical insurance or life insurance.

36. The method according to any one of claims 1-35, wherein the method further comprises using information resulting from the method to predict and / or manage the utilization of medical resources.

37. A kit comprising N biomarker protein capture reagents, wherein N is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 and wherein at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 of the N biomarker protein capture reagents specifically binds to a biomarker protein selected from LYPL1, CILP2, TFF1, S27A2, PAP1, GKN2, InhibinbB chain, andEphAl.