Proteomic biomarkers for predicting cancer-associated venous thromboembolism

WO2024263528A3PCT designated stage expired Publication Date: 2025-06-12MEMORIAL SLOAN KETTERING CANCER CENT +3
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/034365
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-06-21
Filing Date
2024-06-17
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current methods for predicting cancer-associated venous thromboembolism (CAT) in cancer patients are inadequate, with existing prognostic scoring systems showing limited accuracy and a lack of reliable biomarkers for tailoring anticoagulation therapy to prevent both thrombosis and bleeding events.

Method used

The use of specific proteomic biomarkers such as FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and CNDP1, along with machine learning classifiers, to detect mRNA and polypeptide expression levels in cancer patients, allowing for personalized anticoagulant therapy administration.

Benefits of technology

This approach provides a more accurate prediction of CAT risk, enabling tailored anticoagulation therapy to reduce bleeding events and optimize prevention, as demonstrated by improved performance compared to existing prediction scores and coagulation biomarkers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024034365_12062025_PF_FP_ABST
    Figure US2024034365_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates generally to methods for accurately predicting the risk of cancer-associated venous thromboembolism (CAT) and / or preventing CAT in cancer patients using proteomic biomarkers.
Need to check novelty before this filing date? Find Prior Art

Description

PROTEOMIC BIOMARKERS FOR PREDICTING CANCER-ASSOCIATEDVENOUS THROMBOEMBOLISMCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 509,436, filed June 21, 2023, the contents of which are incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] The present technology relates generally to methods for accurately predicting the risk of cancer-associated venous thromboembolism (CAT) and / or preventing CAT in cancer patients using proteomic biomarkers.STATEMENT OF GOVERNMENT SUPPORT

[0003] This invention was made with government support under grant numbers 5U01HL143365 awarded by National Institutes of Health. The government has certain rights in the invention.BACKGROUND

[0004] The following description of the background of the present technology is provided simply as an aid in understanding the present technology and is not admitted to describe or constitute prior art to the present technology.

[0005] Cancer-associated thrombosis (CAT) is a major source of morbidity with a substantial impact on the health, quality of life, and healthcare costs of the cancer subpopulation. Moreover, cancer patients with venous thromboembolism (VTE) are three times more likely to die within six months compared to cancer patients without VTE. Indeed, CAT is a leading cause of death in cancer patients after malignancy itself (Khorana et al., J Thromb Haemost. 2007 Mar;5(3):632-4). Treating CAT with anticoagulation is complex, given both the higher rates of recurrent thrombosis and the paradoxical but increased risk of bleeding (Menapace et al., Thromb Res. 2016 Apr;140 Suppl 1 : S93-8). A more accurate prediction and subsequent assignment of prophylaxis could potentially optimize prevention and avoid the inherent risk of bleeding with anti coagulation. Existing prognostic scoring systems are poor predictors of thrombosis, and novel biomarkers are needed (Es et al., Haematologica. 2017 Sep;102(9): 1494-1501). The Khorana score is themost widely used prediction model but does not incorporate plasma protein biomarkers for prediction. The Vienna CATS score includes D-dimer and soluble P-selectin but is not widely used due to poor predictive discrimination (Es et al., 2017, supra).

[0006] Biomarkers that can accurately predict the risk of thrombosis in patients with cancer would allow tailoring anti coagulation to patients and thus reduce bleeding events by unnecessary thromboprophylaxis. Biomarker discovery thus far has been limited to a handful of candidates; none have reliable performance characteristics to permit clinical use.SUMMARY OF THE PRESENT TECHNOLOGY

[0007] In one aspect, the present disclosure provides a method for preventing cancer associated thromboembolism (CAT) in a cancer patient in need thereof comprising (a) (i) detecting mRNA and / or polypeptide expression levels of at least one of FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB in a biological sample obtained from the cancer patient that is decreased relative to a control sample obtained from a healthy subject or a predetermined threshold; and / or (ii) detecting mRNA and / or polypeptide expression levels of CNDP1 in a biological sample obtained from the cancer patient that is increased relative to a control sample obtained from a healthy subject or a predetermined threshold; and (b) administering to the cancer patient an effective amount of anticoagulant therapy. In another aspect, the present disclosure provides a method for preventing cancer associated thromboembolism (CAT) in a cancer patient in need thereof comprising administering to the cancer patient an effective amount of anticoagulant therapy, wherein a biological sample obtained from the cancer patient comprises (a) mRNA and / or polypeptide expression levels of one or more plasma proteins that are decreased relative to a control sample obtained from a healthy subject or a predetermined threshold, wherein the one or more plasma proteins are selected from the group consisting of FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB; and / or (b) CNDP1 mRNA and / or polypeptide expression levels that are increased relative to a control sample obtained from a healthy subject or a predetermined threshold. In any of the preceding embodiments of the methods disclosed herein, the biological sample is whole blood, serum or plasma.

[0008] Additionally or alternatively, in some embodiments of the methods disclosed herein, the cancer patient is diagnosed with or suffers from a cancer selected from the groupconsisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non- melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non- Hodgkin lymphoma. The cancer may be a Stage 1, Stage 2, Stage 3, or Stage 4 cancer. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score < 2 or > 2 and / or has one or more organ sites of metastasis. In some embodiments, the one or more organ sites of metastasis comprise lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0009] Additionally or alternatively, in some embodiments of the methods disclosed herein, the anticoagulant therapy comprises one or more of apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, or enoxaparin.

[0010] In any of the preceding embodiments of the methods disclosed herein, mRNA expression levels are detected via real-time quantitative PCR (qPCR), digital PCR (dPCR), Reverse transcriptase-PCR (RT-PCR), Northern blotting, microarray, dot or slot blots, in situ hybridization, or fluorescent in situ hybridization (FISH). Additionally or alternatively, in some embodiments, polypeptide expression levels are detected via Western blotting, enzyme-linked immunosorbent assays (ELISA), dot blotting, immunohistochemistry, immunofluorescence, immunoprecipitation, immunoelectrophoresis, or mass-spectrometry.

[0011] In any of the foregoing embodiments of the methods disclosed herein, the cancer patient is chemotherapy-naive or has received / is receiving systemic chemotherapy.Systemic chemotherapy may comprise one or more of alkylating agents, antibiotics, antimetabolites, antimitotics, cyclin-dependent kinase inhibitors, epidermal growth factorreceptor inhibitors, multikinase inhibitors, PARP inhibitors, platinum-based agents, selective estrogen receptor modulators (SERM), or VEGF inhibitors. Examples of chemotherapeutic agents include, but are not limited to, alkylating agents, platinum agents, taxanes, vinca agents, anti-estrogen drugs, aromatase inhibitors, ovarian suppression agents, VEGF / VEGFR inhibitors, EGFZEGFR inhibitors, PARP inhibitors, cytostatic alkaloids, cytotoxic antibiotics, antimetabolites, endocrine / hormonal agents, bisphosphonate therapy agents and targeted biological therapy agents (e.g., therapeutic peptides described in US 6306832, WO 2012007137, WO 2005000889, WO 2010096603 etc.). In some embodiments, the at least one additional therapeutic agent is a chemotherapeutic agent. Specific chemotherapeutic agents include, but are not limited to, cyclophosphamide, fluorouracil (or 5 -fluorouracil or 5-FU), methotrexate, edatrexate (10-ethyl-10-deaza- aminopterin), thiotepa, carboplatin, cisplatin, taxanes, paclitaxel, protein-bound paclitaxel, docetaxel, vinorelbine, tamoxifen, raloxifene, toremifene, fulvestrant, gemcitabine, irinotecan, ixabepilone, temozolmide, topotecan, vincristine, vinblastine, eribulin, mutamycin, capecitabine, anastrozole, exemestane, letrozole, leuprolide, abarelix, buserlin, goserelin, megestrol acetate, risedronate, pamidronate, ibandronate, alendronate, denosumab, zoledronate, trastuzumab, tykerb, anthracyclines (e.g., daunorubicin and doxorubicin), bevacizumab, oxaliplatin, melphalan, etoposide, mechlorethamine, bleomycin, microtubule poisons, annonaceous acetogenins, or combinations thereof.

[0012] Additionally or alternatively, in some embodiments of the methods disclosed herein, the cancer patient is immunotherapy-naive or has received / is receiving immunotherapy. Examples of immunotherapy include, but are not limited to, anti-PD-1 antibody, anti-PD-Ll antibody, anti-PD-L2 antibody, anti-CTLA-4 antibody, anti-TIM3 antibody, anti-4-lBB antibody, anti-CD73 antibody, anti-GITR antibody, and anti-LAG-3 antibody.

[0013] Additionally or alternatively, in certain embodiments of the methods disclosed herein, the cancer patient is radiotherapy-naive or has received / is receiving radiotherapy. The radiotherapy may comprise external radiotherapy, radiotherapy implants (brachytherapy), pre-targeted radioimmunotherapy, radiotherapy injections, radioisotope therapy, or intrabeam radiotherapy.

[0014] In any and all embodiments of the methods disclosed herein, the CAT is pulmonary embolism or lower extremity deep vein thrombosis (DVT). In someembodiments, lower extremity DVT includes thrombi involving a common iliac vein, an external iliac vein, a common femoral vein, a superficial femoral vein, a deep femoral vein, a popliteal vein, a peroneal vein, an anterior tibial vein, a posterior tibial vein, or a deep calf vein.

[0015] In one aspect, the present disclosure provides a method of training a machine learning classifier for estimating risk of cancer-associated venous thromboembolism (VTE) in cancer patients comprising: (a) receiving data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generating a training dataset based on the received data, wherein the training dataset comprises a plurality of features for each subject in the cohort, wherein the plurality of features comprises (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE in cancer patients, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients. The subjects in the cohort may be chemotherapy-naive or may have received systemic chemotherapy. Additionally or alternatively, in certain embodiments, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neckcancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

[0016] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0017] Additionally or alternatively, in some embodiments of the methods disclosed herein, the cancer patients comprise metastatic sites of disease. In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0018] In any of the preceding embodiments, the method further comprises applying the classifier to data on a cancer patient to generate a predictor, and determining whether the cancer patient is at risk for cancer-associated VTE based on the predictor and the operating- point threshold. In some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.

[0019] In any of the foregoing embodiments, the method further comprises administering an effective amount of anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin.

[0020] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0021] In one aspect, the present disclosure provides a method of estimating risk of cancer-associated venous thromboembolism (VTE) in a cancer patient using a machinelearning classifier, the method comprising: receiving patient data corresponding to a plurality of features for the cancer patient; applying the machine learning classifier to the patient data to generate a predictor; and determining whether the cancer patient is at risk for cancer-associated VTE based on the predictor and an operating-point threshold, wherein the machine learning classifier is trained by: (a) receiving cohort data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generating a training dataset based on the received cohort data, wherein the training dataset comprises the plurality of features for each subject in the cohort, wherein the plurality of features comprises (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR- gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining the optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients. In some embodiments, the method further comprises administering an effective amount of anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin. Additionally or alternatively, in some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer- associated VTE. The subjects in the cohort may be chemotherapy-naive or may have received systemic chemotherapy. In any of the preceding embodiments of the methods disclosed herein, one or more of the plurality of features for the cancer patient are determined by assaying whole blood, serum or plasma.

[0022] Additionally or alternatively, in certain embodiments, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breastcancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

[0023] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0024] Additionally or alternatively, in some embodiments of the methods disclosed herein, the cancer patient comprises metastatic sites of disease. In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0025] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0026] In any and all embodiments of the methods disclosed herein, one or more of the plurality of features for each subject in the cohort are determined by assaying whole blood, serum or plasma.

[0027] In any and all embodiments of the methods disclosed herein, the cancer- associated VTE is pulmonary embolism or lower extremity deep vein thrombosis (DVT),optionally wherein lower extremity DVT includes thrombi involving a common iliac vein, an external iliac vein, a common femoral vein, a superficial femoral vein, a deep femoral vein, a popliteal vein, a peroneal vein, an anterior tibial vein, a posterior tibial vein, or a deep calf vein.

[0028] In another aspect, the present disclosure provides a machine learning system for training a machine learning classifier for estimating risk of cancer-associated venous thromboembolism (VTE) in cancer patients, the system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: (a) receive data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generate a training dataset based on the received data, wherein the training dataset comprises a plurality of features for each subject in the cohort, wherein the plurality of features comprises (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE in cancer patients; wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients. The subjects in the cohort may be chemotherapy-naive or may have received systemic chemotherapy.

[0029] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier.

[0030] Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0031] Additionally or alternatively, in some embodiments of the methods disclosed herein, the cancer patients comprise metastatic sites of disease. In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0032] Additionally or alternatively, in certain embodiments of the systems disclosed herein, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

[0033] In any of the preceding embodiments of the systems described herein, the instructions further cause the processor to apply the machine learning classifier to data on a cancer patient to generate a predictor, and determine whether the cancer patient is at risk for cancer-associated VTE based on the predictor and the operating-point threshold. In some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer- associated VTE.

[0034] In any of the foregoing embodiments of the systems described herein, the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin.

[0035] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0036] In yet another aspect, the present disclosure provides a computing system for estimating risk of cancer-associated venous thromboembolism (VTE) in a cancer patient, the computing system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: receive patient data corresponding to a plurality of features for the cancer patient; apply a machine learning classifier to the patient data to generate a predictor; and determine whether the cancer patient is at risk for cancer- associated VTE based on the predictor and an operating-point threshold, wherein the classifier is trained by: (a) receiving cohort data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generating a training dataset based on the received cohort data, wherein the training dataset comprises the plurality of features for each subject in the cohort, wherein the plurality of features comprises (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining the optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients.

[0037] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier.

[0038] Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0039] Additionally or alternatively, in some embodiments of the systems disclosed herein, the cancer patient comprises metastatic sites of disease. In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0040] In any of the preceding embodiments of the systems described herein, the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. In some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin.

[0041] Additionally or alternatively, in certain embodiments of the systems disclosed herein, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

[0042] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0043] In any and all embodiments of the systems disclosed herein, one or more of the plurality of features for each subject in the cohort or the cancer patient are determined by assaying whole blood, serum or plasma.

[0044] In one aspect, the present disclosure provides a non-transitory computer-readable storage medium comprising instructions which, when executed by a processor of a machine learning system, configure the machine learning system to train a machine learning classifier to estimate risk of cancer-associated venous thromboembolism (VTE) in cancer patients, wherein the instructions are configured to cause the processor to: (a) receive data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generate a training dataset based on the received data, wherein the training dataset comprises a plurality of features for each subject in the cohort, the plurality of features comprising (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer- associated VTE in cancer patients; wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients. The subjects in the cohort may be chemotherapy-naive or may have received systemic chemotherapy.

[0045] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier.

[0046] Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0047] Additionally or alternatively, in some embodiments of the computer-readable storage medium disclosed herein, the cancer patient comprises metastatic sites of disease. In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0048] In any of the preceding embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to apply the machine learning classifier to data on a cancer patient to generate a predictor, and determine whether the cancer patient is at risk for cancer-associated VTE based on the predictor and the operating- point threshold. In some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.

[0049] Additionally or alternatively, in certain embodiments of the computer-readable storage medium disclosed herein, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non- melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non- Hodgkin lymphoma.

[0050] In any of the preceding embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin.

[0051] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0052] In another aspect, the present disclosure provides a non-transitory computer- readable storage medium comprising instructions which, when executed by a processor of a computing system, configure the computing system to estimate risk of cancer-associated venous thromboembolism (VTE) in a cancer patient, wherein the instructions are configured to cause the processor to: receive patient data corresponding to a plurality of features for the cancer patient; apply a machine learning classifier to the patient data to generate a predictor; and determine whether the cancer patient is at risk for cancer-associated VTE based on the predictor and an operating-point threshold, wherein the classifier is trained by: (a) receiving cohort data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generating a training dataset based on the received cohort data, wherein the training dataset comprises the plurality of features for each subject in the cohort, wherein the plurality of features comprises (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining the optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients.

[0053] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier.

[0054] Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0055] Additionally or alternatively, in some embodiments of the computer-readable storage medium disclosed herein, the cancer patient comprises metastatic sites of disease. In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0056] In any of the preceding embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. In some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin.

[0057] Additionally or alternatively, in certain embodiments of the computer-readable storage medium disclosed herein, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non- melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non- Hodgkin lymphoma.

[0058] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0059] In any of the preceding embodiments of the computer-readable storage medium disclosed herein, one or more of the plurality of features for the cancer patient are determined by assaying whole blood, serum or plasma.BRIEF DESCRIPTION OF THE DRAWINGS

[0060] FIG. 1 shows an overview of the model, feature extraction, and training of the present technology.

[0061] FIG. 2A shows a comparison of the Bayesian probabilistic machine learning model identifying the 11-protein panel described herein, a.k.a. thrombosis oncology prediction (“TOP”) model (c-statistic 0.85 at a 95% confidence interval (CI)) compared with Khorana Score (c-statistic 0.48 at 95% CI).

[0062] FIG. 2B shows a comparison of the TOP model with common coagulation assays indicative of thrombin generation (i.e, endogenous thrombin potential (ETP), peak thrombin generation, and prothrombin fragment 1+2 (Fl +2), D-dimer) as well as fibrinogen.

[0063] FIGs. 2C-2D show cumulative incidence of thrombosis based on high or low risk proteomic prediction score in Hypercan study with lung (FIG. 2C) and (FIG. 2D) gastric cancer patients.

[0064] FIG. 2E shows external validation of 11 -protein model in AVERT clinical trial. Patients identified as high risk at baseline based on 11 -protein model (blue) compared to lower risk cohort (red).

[0065] FIGs. 3A-3B shows relative expression of 11 -proteins among patients with or without venous thromboembolism.

[0066] FIG. 4A shows a comparison of wild type and CD200R1- / -mice for thrombin- antithrombin (right) complex. These results demonstrate basal hypercoagulable state in CD200R1- / -mice.

[0067] FIG. 4B shows a comparison of wild type and CD200R1- / -mice for coagulation factors and assays including partial thromboplastin time (aPTT), prothrombin time (PT),plasma tissue factor (TF), fibrinogen activity, factor Vila, plasma thrombomodulin and tissue factor pathway inhibitor (TFPI).

[0068] FIG. 4C shows a comparison of wild type and CD200R1- / -mice for ex vivo thrombin generation.

[0069] FIG. 5A shows a comparison of altered protein expression levels in wild type and CD200R1- / -mice.

[0070] FIG. 5B shows a comparison of wild type and CD200R1- / -mice for endothelial activation markers ICAM-1 and P-selectin.

[0071] FIG. 5C shows the relationship between CD200R1 and IL-17a in a clinical cohort including patients that developed DVT.

[0072] FIG. 6 shows that inhibition of II- 17a normalizes thrombin-antithrombin levels in CD200R1 deficient mice.

[0073] FIG. 7A is a block diagram depicting an embodiment of a network environment comprising a client device in communication with server device.

[0074] FIG. 7B is a block diagram depicting a cloud computing environment comprising client device in communication with cloud service providers.

[0075] FIGs. 7C and 7D are block diagrams depicting embodiments of computing devices useful in connection with the methods and systems described herein.

[0076] FIG. 8 depicts a system that includes a computing device and a sample processing system according to various potential embodiments.DETAILED DESCRIPTION

[0077] It is to be appreciated that certain aspects, modes, embodiments, variations and features of the present methods are described below in various levels of detail in order to provide a substantial understanding of the present technology. It is to be understood that the present disclosure is not limited to particular uses, methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.

[0078] CAT is an important complication of cancer for which effective pharmacological prophylaxis methods exist. However, currently available prediction rules have limited accuracy in stratifying patients for CAT risk. Accordingly, approaches to enhance the overall benefit of CAT prophylaxis in cancer patients will be contingent on improved methods for predicting risk.

[0079] In order to identify potentially novel biomarkers predictive of thrombosis, we leveraged a highly sensitive proximity extension proteomic platform in prospective clinical cohorts of patients with cancer and applied machine learning algorithm to identify a limited panel of highly informative proteins predictive of VTE. The model was validated in an orthogonal external prospective clinical trial dataset, and we benchmarked the panel against an existing prediction score and a series of coagulation biomarkers. We further investigated the mechanistic role of CD200R1, a checkpoint receptor limiting leukocyte inflammatory response that contributed strongly to the prediction model. Mice deficient of CD200R1 demonstrated a hypercoagulable phenotype that was associated with very high levels of IL- 17A. Administration of inhibitory antibodies against IL-17 normalized markers of thrombin generation in mouse deficient of CD200R1. We observed a similar relationship between CD200R1 and IL- 17 and thrombosis in our human cohort and performed a meta-analysis of clinical studies in patients with COVID-19 thereby confirming the antithrombotic potential in therapeutic targeting IL-17a. These findings highlight the utility of machine learning in providing novel mechanistic insights in thrombo-inflammatory pathways and the yet unrealized potential of developing antithrombotics without inherent bleeding risks. The present disclosure demonstrates that the 11 plasma proteins disclosed herein are useful biomarkers for accurately predicting the risk of cancer-associated thromboembolism in cancer patients.Definitions

[0080] Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs. As used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise. For example, reference to “a cell” includes a combination of two or more cells, and the like. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, analytical chemistry and nucleic acidchemistry and hybridization described below are those well-known and commonly employed in the art.

[0081] As used herein, the term “about” in reference to a number is generally taken to include numbers that fall within a range of 1%, 5%, or 10% in either direction (greater than or less than) of the number unless otherwise stated or otherwise evident from the context (except where such number would be less than 0% or exceed 100% of a possible value).

[0082] The term “adapter” refers to a short, chemically synthesized, nucleic acid sequence which can be used to ligate to the end of a nucleic acid sequence in order to facilitate attachment to another molecule. The adapter can be single-stranded or double- stranded. An adapter can incorporate a short (typically less than 50 base pairs) sequence useful for PCR amplification or sequencing.

[0083] As used herein, the “administration” of an agent or drug to a subject includes any route of introducing or delivering to a subject a compound to perform its intended function. Administration can be carried out by any suitable route, including but not limited to, orally, intranasally, parenterally (intravenously, intramuscularly, intraperitoneally, or subcutaneously), rectally, intrathecally, intratumorally or topically. Administration includes self-administration and the administration by another.

[0084] As used herein, the terms “amplify” or “amplification” with respect to nucleic acid sequences, refer to methods that increase the representation of a population of nucleic acid sequences in a sample. Nucleic acid amplification methods are well known to the skilled artisan and include ligase chain reaction (LCR), ligase detection reaction (LDR), ligation followed by Q-replicase amplification, PCR, primer extension, strand displacement amplification (SDA), hyperbranched strand displacement amplification, multiple displacement amplification (MDA), nucleic acid strand-based amplification (NASBA), two- step multiplexed amplifications, rolling circle amplification (RCA), recombinase- polymerase amplification (RPA)(TwistDx, Cambridge, UK), transcription mediated amplification, signal mediated amplification of RNA technology, loop-mediated isothermal amplification of DNA, helicase-dependent amplification, single primer isothermal amplification, and self- sustained sequence replication (3 SR), including multiplex versions or combinations thereof. Copies of a particular nucleic acid sequence generated in vitro in an amplification reaction are called “amplicons” or “amplification products.”

[0085] The terms “cancer” or “tumor” are used interchangeably and refer to the presence of cells possessing characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often in the form of a tumor, but such cells can exist alone within an animal, or can be a non-tumorigenic cancer cell. As used herein, the term “cancer” includes premalignant, as well as malignant cancers.

[0086] The terms “complementary” or “complementarity” as used herein with reference to polynucleotides (i.e., a sequence of nucleotides such as an oligonucleotide or a target nucleic acid) refer to the base-pairing rules. The complement of a nucleic acid sequence as used herein refers to an oligonucleotide which, when aligned with the nucleic acid sequence such that the 5' end of one sequence is paired with the 3’ end of the other, is in “antiparallel association.” For example, the sequence “5'-A-G-T-3”’ is complementary to the sequence “3’-T-C-A-5.” Certain bases not commonly found in naturally-occurring nucleic acids may be included in the nucleic acids described herein. These include, for example, inosine, 7- deazaguanine, Locked Nucleic Acids (LNA), and Peptide Nucleic Acids (PNA).Complementarity need not be perfect; stable duplexes may contain mismatched base pairs, degenerative, or unmatched bases. Those skilled in the art of nucleic acid technology can determine duplex stability empirically considering a number of variables including, for example, the length of the oligonucleotide, base composition and sequence of the oligonucleotide, ionic strength and incidence of mismatched base pairs. A complement sequence can also be an RNA sequence complementary to the DNA sequence or its complement sequence, and can also be a cDNA.

[0087] As used herein, a "control" is an alternative sample used in an experiment for comparison purpose. A control can be "positive" or "negative." For example, where the purpose of the experiment is to determine a correlation of the efficacy of a therapeutic agent for the treatment for a particular type of disease, a positive control (a compound or composition known to exhibit the desired therapeutic effect) and a negative control (a subject or a sample that does not receive the therapy or receives a placebo) are typically employed.

[0088] “Detecting” as used herein refers to determining the presence of a mutation or alteration in a nucleic acid of interest in a sample. Detection does not require the method to provide 100% sensitivity. Analysis of nucleic acid markers can be performed usingtechniques known in the art including, but not limited to, sequence analysis, and electrophoretic analysis. Non-limiting examples of sequence analysis include Maxam- Gilbert sequencing, Sanger sequencing, capillary array DNA sequencing, thermal cycle sequencing (Sears et al.. Biotechniques, 13:626-633 (1992)), solid-phase sequencing (Zimmerman et al. , Methods Mol. CellBiol, 3:39-42 (1992)), sequencing with mass spectrometry such as matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF / MS; Fu et al., Nat. Biotechnol, 16:381-384 (1998)), and sequencing by hybridization. Chee et al., Science, 274:610-614 (1996); Drmanac et al., Science, 260: 1649-1652 (1993); Drmanac et al., Nat. Biotechnol, 16:54-58 (1998). Non- limiting examples of electrophoretic analysis include slab gel electrophoresis such as agarose or polyacrylamide gel electrophoresis, capillary electrophoresis, and denaturing gradient gel electrophoresis. Additionally, next generation sequencing methods can be performed using commercially available kits and instruments from companies such as the Life Technologies / Ion Torrent PGM or Proton, the Illumina HiSEQ or MiSEQ, and the Roche / 454 next generation sequencing system.

[0089] “Detectable label” as used herein refers to a molecule or a compound or a group of molecules or a group of compounds used to identify a nucleic acid or protein of interest. In some embodiments, the detectable label may be detected directly. In other embodiments, the detectable label may be a part of a binding pair, which can then be subsequently detected. Signals from the detectable label may be detected by various means and will depend on the nature of the detectable label. Detectable labels may be isotopes, fluorescent moieties, colored substances, and the like. Examples of means to detect detectable labels include but are not limited to spectroscopic, photochemical, biochemical, immunochemical, electromagnetic, radiochemical, or chemical means, such as fluorescence, chemifluorescence, or chemiluminescence, or any other appropriate means.

[0090] As used herein, the term “effective amount” refers to a quantity sufficient to achieve a desired therapeutic and / or prophylactic effect, e.g., an amount which results in the prevention of, or a decrease in a disease or condition described herein or one or more signs or symptoms associated with a disease or condition described herein. In the context of therapeutic or prophylactic applications, the amount of a composition administered to the subject will vary depending on the composition, the degree, type, and severity of the disease and on the characteristics of the individual, such as general health, age, sex, body weightand tolerance to drugs. The skilled artisan will be able to determine appropriate dosages depending on these and other factors. The compositions can also be administered in combination with one or more additional therapeutic compounds. In the methods described herein, the therapeutic compositions may be administered to a subject having one or more signs or symptoms of a disease or condition described herein. As used herein, a "therapeutically effective amount" of a composition refers to composition levels in which the physiological effects of a disease or condition are ameliorated or eliminated. A therapeutically effective amount can be given in one or more administrations.

[0091] As used herein, “expression” includes one or more of the following: transcription of the gene into precursor mRNA; splicing and other processing of the precursor mRNA to produce mature mRNA; mRNA stability; translation of the mature mRNA into protein (including codon usage and tRNA availability); and glycosylation and / or other modifications of the translation product, if required for proper expression and function.

[0092] Gene” as used herein refers to a DNA sequence that comprises regulatory and coding sequences necessary for the production of an RNA, which may have a non-coding function (e.g., a ribosomal or transfer RNA) or which may include a polypeptide or a polypeptide precursor. The RNA or polypeptide may be encoded by a full length coding sequence or by any portion of the coding sequence so long as the desired activity or function is retained. Although a sequence of the nucleic acids may be shown in the form of DNA, a person of ordinary skill in the art recognizes that the corresponding RNA sequence will have a similar sequence with the thymine being replaced by uracil, i.e., "T" is replaced with "U."

[0093] The term “hybridize” as used herein refers to a process where two substantially complementary nucleic acid strands (at least about 65% complementary over a stretch of at least 14 to 25 nucleotides, at least about 75%, or at least about 90% complementary) anneal to each other under appropriately stringent conditions to form a duplex or heteroduplex through formation of hydrogen bonds between complementary base pairs. Hybridizations are typically and preferably conducted with probe-length nucleic acid molecules, preferably 15-100 nucleotides in length, more preferably 18-50 nucleotides in length. Nucleic acid hybridization techniques are well known in the art. See, e.g., Sambrook, et al., 1989, Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Press,Plainview, N.Y. Hybridization and the strength of hybridization (i.e., the strength of the association between the nucleic acids) is influenced by such factors as the degree of complementarity between the nucleic acids, stringency of the conditions involved, and the thermal melting point (Tm) of the formed hybrid. Those skilled in the art understand how to estimate and adjust the stringency of hybridization conditions such that sequences having at least a desired level of complementarity will stably hybridize, while those having lower complementarity will not. For examples of hybridization conditions and parameters, see, e.g., Sambrook, et al., 1989, Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Press, Plainview, N.Y.; Ausubel, F. M. et al. 1994, Current Protocols in Molecular Biology, John Wiley & Sons, Secaucus, N.J. In some embodiments, specific hybridization occurs under stringent hybridization conditions. An oligonucleotide or polynucleotide (e.g., a probe or a primer) that is specific for a target nucleic acid will “hybridize” to the target nucleic acid under suitable conditions.

[0094] The term “multiplex PCR” as used herein refers to amplification of two or more PCR products or amplicons which are each primed using a distinct primer pair.

[0095] “Next-generation sequencing or NGS” as used herein, refers to any sequencing method that determines the nucleotide sequence of either individual nucleic acid molecules (e.g., in single molecule sequencing) or clonally expanded proxies for individual nucleic acid molecules in a high throughput parallel fashion (e.g., greater than 103, 104, 105or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of the nucleic acid species in the library can be estimated by counting the relative number of occurrences of their cognate sequences in the data generated by the sequencing experiment. Next generation sequencing methods are known in the art, and are described, e.g., in Metzker, M. Nature Biotechnology Reviews 11 :31-46 (2010).

[0096] As used herein, “oligonucleotide” refers to a molecule that has a sequence of nucleic acid bases on a backbone comprised mainly of identical monomer units at defined intervals. The bases are arranged on the backbone in such a way that they can bind with a nucleic acid having a sequence of bases that are complementary to the bases of the oligonucleotide. The most common oligonucleotides have a backbone of sugar phosphate units. A distinction may be made between oligodeoxyribonucleotides that do not have a hydroxyl group at the 2' position and oligoribonucleotides that have a hydroxyl group at the 2' position. Oligonucleotides may also include derivatives, in which the hydrogen of thehydroxyl group is replaced with organic groups, e.g., an allyl group. Oligonucleotides of the method which function as primers or probes are generally at least about 10-15 nucleotides long and more preferably at least about 15 to 25 nucleotides long, although shorter or longer oligonucleotides may be used in the method. The exact size will depend on many factors, which in turn depend on the ultimate function or use of the oligonucleotide. The oligonucleotide may be generated in any manner, including, for example, chemical synthesis, DNA replication, restriction endonuclease digestion of plasmids or phage DNA, reverse transcription, PCR, or a combination thereof. The oligonucleotide may be modified e.g., by addition of a methyl group, a biotin or digoxigenin moiety, a fluorescent tag or by using radioactive nucleotides.

[0097] As used herein, “prevention” or “preventing” of a disorder or condition refers to a compound that, in a statistical sample, reduces the occurrence of the disorder or condition in the treated sample relative to an untreated control sample, or delays the onset of one or more symptoms of the disorder or condition relative to the untreated control sample.

[0098] As used herein, the term “primer” refers to an oligonucleotide, which is capable of acting as a point of initiation of nucleic acid sequence synthesis when placed under conditions in which synthesis of a primer extension product which is complementary to a target nucleic acid strand is induced, i.e., in the presence of different nucleotide triphosphates and a polymerase in an appropriate buffer (“buffer” includes pH, ionic strength, cofactors etc.) and at a suitable temperature. One or more of the nucleotides of the primer can be modified for instance by addition of a methyl group, a biotin or digoxigenin moiety, a fluorescent tag or by using radioactive nucleotides. A primer sequence need not reflect the exact sequence of the template. For example, a non-complementary nucleotide fragment may be attached to the 5' end of the primer, with the remainder of the primer sequence being substantially complementary to the strand. The term primer as used herein includes all forms of primers that may be synthesized including peptide nucleic acid primers, locked nucleic acid primers, phosphorothioate modified primers, labeled primers, and the like. The term “forward primer” as used herein means a primer that anneals to the anti-sense strand of dsDNA. A “reverse primer” anneals to the sense-strand of dsDNA.

[0099] As used herein, “primer pair” refers to a forward and reverse primer pair (i.e., a left and right primer pair) that can be used together to amplify a given region of a nucleic acid of interest.

[0100] “Probe” as used herein refers to nucleic acid that interacts with a target nucleic acid via hybridization. A probe may be fully complementary to a target nucleic acid sequence or partially complementary. The level of complementarity will depend on many factors based, in general, on the function of the probe. A probe or probes can be used, for example to detect the presence or absence of a mutation in a nucleic acid sequence by virtue of the sequence characteristics of the target. Probes can be labeled or unlabeled, or modified in any of a number of ways well known in the art. A probe may specifically hybridize to a target nucleic acid. Probes may be DNA, RNA or a RNA / DNA hybrid. Probes may be oligonucleotides, artificial chromosomes, fragmented artificial chromosome, genomic nucleic acid, fragmented genomic nucleic acid, RNA, recombinant nucleic acid, fragmented recombinant nucleic acid, peptide nucleic acid (PNA), locked nucleic acid, oligomer of cyclic heterocycles, or conjugates of nucleic acid. Probes may comprise modified nucleobases, modified sugar moieties, and modified internucleotide linkages. A probe may be used to detect the presence or absence of a target nucleic acid. Probes are typically at least about 10, 15, 20, 25, 30, 35, 40, 50, 60, 75, 100 nucleotides or more in length.

[0101] As used herein, a “sample” refers to a substance that is being assayed for the presence of a mutation in a nucleic acid of interest. Processing methods to release or otherwise make available a nucleic acid for detection are well known in the art and may include steps of nucleic acid manipulation. A biological sample may be a body fluid or a tissue sample. In some cases, a biological sample may consist of or comprise blood, plasma, sera, urine, feces, epidermal sample, vaginal sample, skin sample, cheek swab, sperm, amniotic fluid, cultured cells, bone marrow sample, tumor biopsies, aspirate and / or chorionic villi, cultured cells, and the like. Fresh, fixed or frozen tissues may also be used. In one embodiment, the sample is preserved as a frozen sample or as formaldehyde- or paraformaldehyde-fixed paraffin-embedded (FFPE) tissue preparation. For example, the sample can be embedded in a matrix, e.g., an FFPE block or a frozen sample. Whole blood samples of about 0.5 to 5 ml collected with EDTA, ACD or heparin as anti-coagulant are suitable.

[0102] The term “sensitivity,” as used herein in reference to the methods of the present technology, is a measure of the ability of a method to detect a preselected sequence variant in a heterogeneous population of sequences. A method has a sensitivity of S % for variantsof F % if, given a sample in which the preselected sequence variant is present as at least F % of the sequences in the sample, the method can detect the preselected sequence at a preselected confidence of C %, S % of the time. By way of example, a method has a sensitivity of 90% for variants of 5% if, given a sample in which the preselected variant sequence is present as at least 5% of the sequences in the sample, the method can detect the preselected sequence at a preselected confidence of 99%, 9 out of 10 times (F=5%; C=99%; S=90%).

[0103] The term “specific” as used herein in reference to an oligonucleotide primer means that the nucleotide sequence of the primer has at least 12 bases of sequence identity with a portion of the nucleic acid to be amplified when the oligonucleotide and the nucleic acid are aligned. An oligonucleotide primer that is specific for a nucleic acid is one that, under the stringent hybridization or washing conditions, is capable of hybridizing to the target of interest and not substantially hybridizing to nucleic acids which are not of interest. Higher levels of sequence identity are preferred and include at least 75%, at least 80%, at least 85%, at least 90%, at least 95% and more preferably at least 98% sequence identity.

[0104] “Specificity,” as used herein, is a measure of the ability of a method to distinguish a truly occurring preselected sequence variant from sequencing artifacts or other closely related sequences. It is the ability to avoid false positive detections. False positive detections can arise from errors introduced into the sequence of interest during sample preparation, sequencing error, or inadvertent sequencing of closely related sequences like pseudo-genes or members of a gene family. A method has a specificity of X % if, when applied to a sample set of Niotai sequences, in which Xirue sequences are truly variant and XNottrue are not truly variant, the method selects at least X % of the not truly variant as not variant. E.g., a method has a specificity of 90% if, when applied to a sample set of 1,000 sequences, in which 500 sequences are truly variant and 500 are not truly variant, the method selects 90% of the 500 not truly variant sequences as not variant. Exemplary specificities include 90, 95, 98, and 99%.

[0105] The term “stringent hybridization conditions” as used herein refers to hybridization conditions at least as stringent as the following: hybridization in 50% formamide, 5xSSC, 50 mM NaH2PO4, pH 6.8, 0.5% SDS, 0.1 mg / mL sonicated salmon sperm DNA, and 5x Denhart's solution at 42° C. overnight; washing with 2x SSC, 0.1% SDS at 45° C; and washing with 0.2x SSC, 0.1% SDS at 45° C. In another example,stringent hybridization conditions should not allow for hybridization of two nucleic acids which differ over a stretch of 20 contiguous nucleotides by more than two bases.

[0106] As used herein, the terms “target sequence” and “target nucleic acid sequence” refer to a specific nucleic acid sequence to be detected and / or quantified in the sample to be analyzed.

[0107] As used herein, the term “therapeutic agent” is intended to mean a compound that, when present in an effective amount, produces a desired therapeutic effect on a subject in need thereof.

[0108] “Treating” or “treatment” as used herein covers the treatment of a disease or disorder described herein, in a subject, such as a human, and includes: (i) inhibiting a disease or disorder, i.e., arresting its development; (ii) relieving a disease or disorder, i.e., causing regression of the disorder; (iii) slowing progression of the disorder; and / or (iv) inhibiting, relieving, or slowing progression of one or more symptoms of the disease or disorder. In some embodiments, treatment means that the symptoms associated with the disease are, e.g., alleviated, reduced, cured, or placed in a state of remission.

[0109] It is also to be appreciated that the various modes of treatment of disorders as described herein are intended to mean “substantial,” which includes total but also less than total treatment, and wherein some biologically or medically relevant result is achieved. The treatment may be a continuous prolonged treatment for a chronic disease or a single, or few time administrations for the treatment of an acute condition.Methods for Detecting Polynucleotides Associated with Elevated VTE Risk

[0110] Polynucleotides associated with elevated VTE risk may be detected by a variety of methods known in the art. Non-limiting examples of detection methods are described below. The detection assays in the methods of the present technology may include purified or isolated DNA, RNA or protein or the detection step may be performed directly from a biological sample without the need for further DNA, RNA or protein purification / isolation.Nucleic Acid Ampli fication and / or Detection

[0111] Polynucleotides associated with elevated VTE risk can be detected by the use of nucleic acid amplification techniques that are well known in the art. The starting material may be genomic DNA, cDNA, RNA, or mRNA. Nucleic acid amplification can be linear or exponential. Specific nucleic acids may be detected by the use of amplification methodswith the aid of oligonucleotide primers or probes designed to interact with or hybridize to a particular target sequence in a specific manner.

[0112] Non-limiting examples of nucleic acid amplification techniques include polymerase chain reaction (PCR), real-time quantitative PCR (qPCR), digital PCR (dPCR), reverse transcriptase polymerase chain reaction (RT-PCR), nested PCR, ligase chain reaction (see Abravaya, K. et al., Nucleic Acids Res . (1995), 23:675-682), branched DNA signal amplification (see Urdea, M. S. et al., AIDS (1993), 7(suppl 2):S11- S14), amplifiable RNA reporters, Q-beta replication, transcription-based amplification, boomerang DNA amplification, strand displacement activation, cycling probe technology, isothermal nucleic acid sequence based amplification (NASBA) (see Kievits, T. et al., J Virological Methods (1991), 35:273-286), Invader Technology, next-generation sequencing technology or other sequence replication assays or signal amplification assays.

[0113] Primers'. Oligonucleotide primers for use in amplification methods can be designed according to general guidance well known in the art as described herein, as well as with specific requirements as described herein for each step of the particular methods described. In some embodiments, oligonucleotide primers for cDNA synthesis and PCR are 10 to 100 nucleotides in length, preferably between about 15 and about 60 nucleotides in length, more preferably 25 and about 50 nucleotides in length, and most preferably between about 25 and about 40 nucleotides in length.

[0114] Tmof a polynucleotide affects its hybridization to another polynucleotide (e.g., the annealing of an oligonucleotide primer to a template polynucleotide). In certain embodiments of the disclosed methods, the oligonucleotide primer used in various steps selectively hybridizes to a target template or polynucleotides derived from the target template (i.e., first and second strand cDNAs and amplified products). Typically, selective hybridization occurs when two polynucleotide sequences are substantially complementary (at least about 65% complementary over a stretch of at least 14 to 25 nucleotides, preferably at least about 75%, more preferably at least about 90% complementary). See Kanehisa, M., Polynucleotides Res. (1984), 12:203, incorporated herein by reference. As a result, it is expected that a certain degree of mismatch at the priming site is tolerated. Such mismatch may be small, such as a mono-, di- or tri -nucleotide. In certain embodiments, 100% complementarity exists.

[0115] Probes'. Probes are capable of hybridizing to at least a portion of the nucleic acid of interest or a reference nucleic acid (i.e., wild-type sequence). Probes may be an oligonucleotide, artificial chromosome, fragmented artificial chromosome, genomic nucleic acid, fragmented genomic nucleic acid, RNA, recombinant nucleic acid, fragmented recombinant nucleic acid, peptide nucleic acid (PNA), locked nucleic acid, oligomer of cyclic heterocycles, or conjugates of nucleic acid. Probes may be used for detecting and / or capturing / purifying a nucleic acid of interest.

[0116] Typically, probes can be about 10 nucleotides, about 20 nucleotides, about 25 nucleotides, about 30 nucleotides, about 35 nucleotides, about 40 nucleotides, about 50 nucleotides, about 60 nucleotides, about 75 nucleotides, or about 100 nucleotides long. However, longer probes are possible. Longer probes can be about 200 nucleotides, about 300 nucleotides, about 400 nucleotides, about 500 nucleotides, about 750 nucleotides, about 1,000 nucleotides, about 1,500 nucleotides, about 2,000 nucleotides, about 2,500 nucleotides, about 3,000 nucleotides, about 3,500 nucleotides, about 4,000 nucleotides, about 5,000 nucleotides, about 7,500 nucleotides, or about 10,000 nucleotides long.

[0117] Probes may also include a detectable label or a plurality of detectable labels. The detectable label associated with the probe can generate a detectable signal directly. Additionally, the detectable label associated with the probe can be detected indirectly using a reagent, wherein the reagent includes a detectable label, and binds to the label associated with the probe.

[0118] In some embodiments, detectably labeled probes can be used in hybridization assays including, but not limited to Northern blots, Southern blots, microarray, dot or slot blots, and in situ hybridization assays such as fluorescent in situ hybridization (FISH) to detect a target nucleic acid sequence within a biological sample. Certain embodiments may employ hybridization methods for measuring expression of a polynucleotide gene product, such as mRNA. Methods for conducting polynucleotide hybridization assays have been well developed in the art. Hybridization assay procedures and conditions will vary depending on the application and are selected in accordance with the general binding methods known including those referred to in: Maniatis el al. Molecular Cloning: A Laboratory Manual (2nd Ed. Cold Spring Harbor, N.Y., 1989); Berger and Kimmel Methods in Enzymology, Vol. 152, Guide to Molecular Cloning Techniques (Academic Press, Inc., San Diego, Calif, 1987); Young and Davis, PNAS. 80: 1194 (1983).

[0119] Detectably labeled probes can also be used to monitor the amplification of a target nucleic acid sequence. In some embodiments, detectably labeled probes present in an amplification reaction are suitable for monitoring the amount of amplicon(s) produced as a function of time. Examples of such probes include, but are not limited to, the 5'- exonuclease assay (TAQMAN® probes described herein (see also U.S. Pat. No. 5,538,848) various stem-loop molecular beacons (see for example, U.S. Pat. Nos. 6,103,476 and 5,925,517 and Tyagi and Kramer, 1996, Nature Biotechnology 14:303- 308), stemless or linear beacons (see, e.g., WO 99 / 21881), PNA Molecular Beacons™ (see, e.g., U.S. Pat. Nos. 6,355,421 and 6,593,091), linear PNA beacons (see, for example, Kubista et al., 2001, SPIE 4264:53-58), non-FRET probes (see, for example, U.S. Pat. No. 6,150,097), Sunrise® / Amplifluor™ probes (U.S. Pat. No. 6,548,250), stem-loop and duplex Scorpion probes (Solinas etal., 2001, Nucleic Acids Research 29:E96 and U.S. Pat. No. 6,589,743), bulge loop probes (U.S. Pat. No. 6,590,091), pseudo knot probes (U.S. Pat. No. 6,589,250), cyclicons (U.S. Pat. No. 6,383,752), MGB Eclipse™ probe (Epoch Biosciences), hairpin probes (U.S. Pat. No. 6,596,490), peptide nucleic acid (PNA) light-up probes, self- assembled nanoparticle probes, and ferrocene-modified probes described, for example, in U.S. Pat. No. 6,485,901 ; Mhlanga et al., 2001, Methods 25:463-471 ; Whitcombe et al., 1999, Nature Biotechnology. 17:804-807; Isacsson et al., 2000, Molecular Cell Probes. 14:321-328; Svanvik et al., 2000, Anal Biochem. 281 :26-35; Wolffs et al., 2001, Biotechniques 766: 769-771 ; Tsourkas et al., 2002, Nucleic Acids Research. 30:4208-4215; Riccelli et al., 2002, Nucleic Acids Research 30:4088-4093; Zhang et al., 2002 Shanghai. 34:329-332; Maxwell et al., 2002, J. Am. Chem. Soc. 124:9606-9612; Broude et al., 2002, Trends Biotechnol. 20:249-56; Huang et al., 2002, Chem. Res. Toxicol. 15: 118- 126; and Yu et al., 2001, J. Am. Chem. Soc 14: 11155-11161.

[0120] In some embodiments, the detectable label is a fluorophore. Suitable fluorescent moieties include but are not limited to the following fluorophores working individually or in combination: 4-acetamido-4'-isothiocyanatostilbene- 2,2'disulfonic acid; acridine and derivatives: acridine, acridine isothiocyanate; Alexa Fluors: Alexa Fluor® 350, Alexa Fluor® 488, Alexa Fluor® 546, Alexa Fluor® 555, Alexa Fluor® 568, Alexa Fluor® 594, Alexa Fluor® 647 (Molecular Probes); 5-(2- aminoethyl)aminonaphthalene-l -sulfonic acid (EDANS); 4-amino-N-[3- vinylsulfonyl)phenyl]naphthalimide-3,5 disulfonate (Lucifer Yellow VS); N-(4-anilino-l- naphthyl)mal eimide; anthranilamide; Black Hole Quencher™(BHQ™) dyes (biosearch Technologies); BODIPY dyes: BODIPY® R-6G, BOPIPY® 530 / 550, BODIPY® FL; Brilliant Yellow; coumarin and derivatives: coumarin, 7-amino-4- methylcoumarin (AMC, Coumarin 120),7-amino-4-trifluoromethylcouluarin (Coumarin 151); Cy2®, Cy3®, Cy3.5®, Cy5®, Cy5.5®; cyanosine; 4',6-diaminidino-2-phenylindole (DAPI); 5', 5"-dibromopyrogallol- sulfonephthalein (Bromopyrogallol Red); 7- diethylamino-3-(4'-isothiocyanatophenyl)-4- methylcoumarin; di ethylenetriamine pentaacetate; 4,4'-diisothiocyanatodihydro-stilbene-2,2'- disulfonic acid; 4,4'- diisothiocyanatostilbene-2,2'-disulfonic acid; 5- [dimethylamino]naphthalene-l -sulfonyl chloride (DNS, dansyl chloride); 4-(4'- dimethylaminophenylazo)benzoic acid (DABCYL);4-dimethylaminophenylazophenyl-4'- isothiocyanate (DABITC); Eclipse™ (Epoch Biosciences Inc.); eosin and derivatives: eosin, eosin isothiocyanate; erythrosin and derivatives: erythrosin B, erythrosin isothiocyanate; ethidium; fluorescein and derivatives:5-carboxyfluorescein (FAM), 5-(4,6-di chlorotriazin-2- yl)amino fluorescein (DTAF), 2', 7'- dimethoxy-4'5'-dichloro-6-carboxyfluorescein (JOE), fluorescein, fluorescein isothiocyanate (FITC), hexachloro-6-carboxyfluorescein (HEX), QFITC (XRITC), tetrachlorofluorescem (TET); fiuorescamine; IR144; IR1446; lanthamide phosphors; Malachite Green isothiocyanate; 4-methylumbelliferone; ortho cresolphthalein; nitrotyrosine; pararosaniline; Phenol Red; B -phycoerythrin, R-phycoerythrin; allophycocyanin; o-phthaldialdehyde; Oregon Green®; propidium iodide; pyrene and derivatives: pyrene, pyrene butyrate, succinimidyl 1 -pyrene butyrate; QSY® 7; QSY® 9; QSY® 21; QSY® 35 (Molecular Probes); Reactive Red 4 (Cibacron®Brilliant Red 3B-A); rhodamine and derivatives: 6- carboxy-X-rhodamine (ROX), 6-carboxyrhodamine (R6G), lissamine rhodamine B sulfonyl chloride, rhodamine (Rhod), rhodamine B, rhodamine 123, rhodamine green, rhodamine X isothiocyanate, riboflavin, rosolic acid, sulforhodamine B, sulforhodamine 101, sulfonyl chloride derivative of sulforhodamine 101 (Texas Red); terbium chelate derivatives;N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA); tetramethyl rhodamine; tetramethyl rhodamine isothiocyanate (TRITC); and VIC®. Detector probes can also comprise sulfonate derivatives of fluorescenin dyes with S03 instead of the carboxylate group, phosphoramidite forms of fluorescein, phosphoramidite forms of CY 5 (commercially available for example from Amersham).

[0121] Detectably labeled probes can also include quenchers, including without limitation black hole quenchers (Biosearch), Iowa Black (IDT), QSY quencher (Molecular Probes), and Dabsyl and Dabcel sulfonate / carboxylate Quenchers (Epoch).

[0122] Detectably labeled probes can also include two probes, wherein for example a fluorophore is on one probe, and a quencher is on the other probe, wherein hybridization of the two probes together on a target quenches the signal, or wherein hybridization on the target alters the signal signature via a change in fluorescence.

[0123] In some embodiments, interchelating labels such as ethidium bromide, SYBR® Green I (Molecular Probes), and PicoGreen® (Molecular Probes) are used, thereby allowing visualization in real-time, or at the end point, of an amplification product in the absence of a detector probe. In some embodiments, real-time visualization may involve the use of both an intercalating detector probe and a sequence-based detector probe. In some embodiments, the detector probe is at least partially quenched when not hybridized to a complementary sequence in the amplification reaction, and is at least partially unquenched when hybridized to a complementary sequence in the amplification reaction.

[0124] In some embodiments, the amount of probe that gives a fluorescent signal in response to an excited light typically relates to the amount of nucleic acid produced in the amplification reaction. Thus, in some embodiments, the amount of fluorescent signal is related to the amount of product created in the amplification reaction. In such embodiments, one can therefore measure the amount of amplification product by measuring the intensity of the fluorescent signal from the fluorescent indicator.

[0125] Primers or probes may be designed to selectively hybridize to any portion of a nucleic acid sequence encoding a polypeptide selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB. Exemplary nucleic acid sequences of the human orthologs of these genes are provided below:

[0126] NM 032649.6 Homo sapiens carnosine dipeptidase 1 (CNDP1), mRNA (SEQ ID NO: 1)AGATGTGAATAGCTCCACTATACCAGCCTCGTCTTCCTTCCGGGGGACAACGTGGGTCAGGGCACAGAGA GATATTTAATGTCACCCTCTTGGGGCTTTCATGGGACTCCCTCTGCCACATTTTTTGGAGGTTGGGAAAG TTGCTAGAGGCTTCAGAACTCCAGCCTAATGGATCCCAAACTCGGGAGAATGGCTGCGTCCCTGCTGGCT GTGCTGCTGCTGCTGCTGGAGCGCGGCATGTTCTCCTCACCCTCCCCGCCCCCGGCGCTGTTAGAGAAAG T CT T C CAGT ACAT T GAC CT C CAT CAGGAT GAAT T T GT GCAGAC GCT GAAGGAGT GGGT GGC CAT C GAGAGCGACTCTGTCCAGCCTGTGCCTCGCTTCAGACAAGAGCTCTTCAGAATGATGGCCGTGGCTGCGGACACG CTGCAGCGCCTGGGGGCCCGTGTGGCCTCGGTGGACATGGGTCCTCAGCAGCTGCCCGATGGTCAGAGTC TTCCAATACCTCCCGTCATCCTGGCCGAACTGGGGAGCGATCCCACGAAAGGCACCGTGTGCTTCTACGG CCACTTGGACGTGCAGCCTGCTGACCGGGGCGATGGGTGGCTCACGGACCCCTATGTGCTGACGGAGGTA GACGGGAAACTTTATGGACGAGGAGCGACCGACAACAAAGGCCCTGTCTTGGCTTGGATCAATGCTGTGA GC GC CT T CAGAGC C CT GGAGCAAGAT CT T C CT GT GAAT AT CAAAT T CAT CAT T GAGGGGAT GGAAGAGGC TGGCTCTGTTGCCCTGGAGGAACTTGTGGAAAAAGAAAAGGACCGATTCTTCTCTGGTGTGGACTACATT GT AAT T T CAGAT AAC CT GT GGAT CAGC CAAAGGAAGC CAGCAAT CACT T AC GGAAC C C GGGGGAACAGCT ACT T CAT GGT GGAGGT GAAAT GCAGAGAC CAGGAT T T T CACT CAGGAAC CT T T GGT GGCAT C CT T CAT GA ACCAATGGCTGATCTGGTTGCTCTTCTCGGTAGCCTGGTAGACTCGTCTGGTCATATCCTGGTCCCTGGA ATCTATGATGAAGTGGTTCCTCTTACAGAAGAGGAAATAAATACATACAAAGCCATCCATCTAGACCTAG AAGAAT AC C GGAAT AGCAGC C GGGT T GAGAAAT T T CT GT T C GAT ACT AAGGAGGAGAT T CT AAT GCAC CT CTGGAGGTACCCATCTCTTTCTATTCATGGGATCGAGGGCGCGTTTGATGAGCCTGGAACTAAAACAGTC ATACCTGGCCGAGTTATAGGAAAATTTTCAATCCGTCTAGTCCCTCACATGAATGTGTCTGCGGTGGAAA AAC AGGT GAC AC GAC AT C T T GAAGAT GTGTTCTC C AAAAGAAAT AGT T C C AAC AAGAT GGTTGTTTC CAT GACT CT AGGACT ACAC C C GT GGAT T GCAAAT AT T GAT GACAC C CAGT AT CT C GCAGCAAAAAGAGC GAT C AGAACAGT GT T T GGAACAGAAC CAGAT AT GAT C C GGGAT GGAT C CAC CAT T C CAAT TGC CAAAAT GT T C C AGGAGAT C GT C CACAAGAGC GT GGT GCT AAT T C C GCT GGGAGCT GT T GAT GAT GGAGAACAT T C GCAGAA T GAGAAAAT CAACAGGT GGAACT ACAT AGAGGGAAC CAAAT TAT TTGCTGCCTTTT T CT T AGAGAT GGC C CAGCTCCATTAATCACAAGAACCTTCTAGTCTGATCTGATCCACTGACAGATTCACCTCCCCCACATCCC T AGACAGGGAT GGAAT GT AAAT AT C CAGAGAAT T T GGGT CT AGT AT AGT ACAT TTTCCCTTC CAT T T AAA AT GT CT T GGGAT AT CT GGAT CAGT AAT AAAAT AT T T CAAAGGCACAGAT GT T GGAAAT GGT T T AAGGT C C CCCACTGCACACCTTCCTCAAGTCATAGCTGCTTGCAGCAACTTGATTTCCCCAAGTCCTGTGCAATAGC CCCAGGATTGGATTCCTTCAAACCTTTTAGCATATCTCCAACCTTGCAATTTGATTGGCATAATCACTCC AGTTTGCTTTCTAGGTCCTCAAGTGCTCGTGACACATAATCATTCCATCCAATGATCGCCTTTGCTTTAC CACTCTTTCCTTTTATCTTATTAATAAAAATGTTGGTCTCCACCACTGACTACAATGATTTCCCCATGGA TTCATTTTCAGAGGAGTTACTAAGGCATGGTACTATATTAACCTCTTGCCTCCCCTTCATTTATTTTTTC CCACTTCTT GAAAAGT T C C GAAGC T GAAAT GAT CAAT AT T CAT AC C CAT AT GC CAT TTGGGAGTCCTTGT GCAAAGATGAGTATTTTTTTTTTTATTCCAAGCTTCTTTACCTTTCCTGAGATCGATGCCACCTTGTCAG CTTCTCTGCCTCCCGTTCCTTCTTCCTGTCCCATCAGAAAAGGCCCACTCTCTCTCTATCCAACATCTGT GCAC AGGT T GCAC C AGAGC AGAAT T TAT C C T GAT AGAAT TTCTCTCACCC CAT T GAAGT T T T AC T CAAT G ACAACATCTAAATTGCTCAAAGCTTTTAATAGCATCAGTTTTAGATAGAGAACTAATCAACCCTGACATC AGT AACACAC GCT AAAT T T AAAGC CAT CAGCT AAT GCT GCAT T CAT T AAT CAT T T T AGT T AAC CAT CAT G TTTGGCTT GAT T GGAT T T GAC GGT GTAAACAACAAAAAAAGACAGT CAT T T T T CT ACAGT T GAGT GT AT A T AAACAAAAT CAAT CT T TAT GAAT T TAT TAT T T AAAT CACAT T AAT GGAGAAT CAT T T AGGGGCT T AT TA AAAT T T T T T T CT AT TAT GACT C CAAAT AAAAACAAAGCAGAGGGC CAGGC GC GGT GGCT CAT GC C CAT AA T C C CAGCACT T T C GGAGGC C GAGGT GGGC GGAT CAC CT GAGGT C GGGAGT T C GAGAT CAGC CT GGCT AAC ATGGTGAAACCCTGTCTCTACTGAATATACAAAATTAACTGGGCATGGTGGTGCATGCCTGTAGTCCCAG CTACTCGGGAGGCTGAGGCAGGAGAATCGCTTGAACCCAGGAGGCGGAAGTTGCAGTGAACTGGGATCAT GCCACTGCACTCCAGCCTGGGCGACAGAGTGAGACTCCGTCTCCAAAAAAAAAAAAAAAAAAAAAAAAAG CAGAACACATTTGCGTTCAACTTTTTTCCATCACTCTGTTGGTTCTTGAGATGTCAGTGTCAGTTTAAAA ACGTGCTGTACCACCTTTAAAAGGACTGGAGCAGGCAGGCAGTGATTCAGTTCACCTTGACTCTAGTCAT AGAAGT GGAC C CAACT GAGAGCT CAAC CT AGT T GT GGGCAT T T GT GAAGGAT T C CAGT T AGT GT T T T CT T T GGAAGCAGGAGGCT T CAAGC CAT CACAT CT C CAAT T CT AT AGAAT AT CAAGT CAT C GT T AAAGAACAGC AGCAGGT GT T T GAT AT T T T GT GCT ATT AAGT GT AGT T T GC GCT CAT CT AAT TAT GCAGAGT T T C GT GGT G TTTCTCGATTTTGTTGAACCTACCTGATGCATTTGGCACTGTAGATAAGGACTAACAGTAAGAGAGATTT AGAGT C T T CAT GT T AGT C CAT T AAC AT AAGT T T AT TAT AAAAT AC AC T GT AAGC AT T T C T GGC T AAT GT A CCATCAGATTTAAGAAGTTACCCAACAATATGTATTCGCTGATCTTCATTCCTGCTGAATCTGAGACCAG CAGT GT CT CAAT TAT TAG GT T GCT GCCAGGAGAGAAGGGGAGGAGAAT GAAT AGT TAT GGAAGT T GT AGA AT CT AGGAT AT T C CAAC C CT GGAT T GC CAGAGT T CAGAGT C CAT GT T T T T CT GT CAC C CAGGGT CAC GAC TGGTCTCCCCAGCTCTTCCCTCCTGTTTTTCCACCCAATTCTAAGCCCTGTGGTGATCCAGACCCTACCC TGACGTGGGAGTCAGAAAGGAGTATCTTCCAAGCTCCTTGTACTCAAAGGGATAAAAGAAAGCTACTCAG T T GT TAT T CAGAT AT AT T AAC CAGC CAT AT GGAGGCAAAAAAGGGAGAAAT T T AAAT AT T T GT T T T T T AA CAT T T T AAT GAT GAAAT AT AT GCACTT GGACAAAAAT GAT CAC C CAGT AAAGT T GGGC GT ACAT T GT GGC AT GC C CAT CAGGGGGACT T T GT T GT AT T GGAGGAGGGCAAT GT AT CT GGCT T T GT T AGT AGAC GT AGC GC AT T AAAGC GT AAC TAT T GT GAT AAGTT T AGGGGT T AGGAAT AT CAT T T GAGT T T GGT AGAT GCTTTTCTG GAAAACAAT GAGGCT GGGAAT GTTTCTGGCTT TAT GT T GT AT T CAT AAAAAT AAAAT GT ACT GGCT T GGA AA

[0127] NM 000804.4 Homo sapiens folate receptor gamma (FOLR3), transcript variant 1, coding, mRNA (SEQ ID NO: 2)AGAGCCTGGACCTACAGCGCTGTTGGTGGAGGTCCTGCCTCCAGGAATAGATGGACATGGCCTGGCAGAT GATGCAGCTGCTGCTTCTGGCTTTGGTGACTGCTGCGGGGAGTGCCCAGCCCAGGAGTGCGCGGGCCAGG AC GGAC CT GCT CAAT GT CT GCAT GAAC GC CAAGCAC CACAAGACACAGC C CAGC C C C GAGGAC GAGCT GT AT GGC CAGT GCAGT C C CT GGAAGAAGAAT GC CT GCT GCAC GGC CAGCAC CAGC CAGGAGCT GCACAAGGA CACCTCCCGCCTGTACAACTTTAACTGGGATCACTGTGGTAAGATGGAACCCACCTGCAAGCGCCACTTT ATCCAGGACAGCTGTCTCTATGAGTGCTCACCCAACCTGGGGCCCTGGATCCGGCAGGTCAACCAGAGCT GGCGCAAAGAGCGCATTCTGAACGTGCCCCTGTGCAAAGAGGACTGTGAGCGCTGGTGGGAGGACTGTCG CACCTCCTACACCTGCAAAAGCAACTGGCACAAAGGCTGGAATTGGACCTCAGGGATTAATGAGTGTCCG GCCGGGGCCCTCTGCAGCACCTTTGAGTCCTACTTCCCCACTCCAGCCGCCCTTTGTGAAGGCCTCTGGA GCCACTCCTTCAAGGTCAGCAACTATAGTCGAGGGAGCGGCCGCTGCATCCAGATGTGGTTTGACTCAGC C CAGGGCAAC C C CAAT GAGGAGGT GGC CAAGT T CT AT GCTGCGGC CAT GAAT GCTGGGGCCCCGTCTCGT GGGATTATTGATTCCTGATCCAAGAAGGGTCCTCTGGGGTTCTTCCAACAACCTATTCTAATAGACAAAT CCACATGTG

[0128] NM 001412270.1 Homo sapiens folate receptor gamma (FOLR3), transcript variant 2, coding, mRNA (SEQ ID NO: 3)AGAGCCTGGACCTACAGCGCTGTTGGTGGAGGTCCTGCCTCCAGGAATAGATGGACATGGCCTGGCAGAT GATGCAGCTGCTGCTTCTGGCTTTGGTGACTGCTGCGGGGAGTGCCCAGCCCAGGAGTGCGCGGGCCAGG AC GGAC CT GCT CAAT GT CT GCAT GAAC GC CAAGCAC CACAAGACACAGC C CAGC C C C GAGGAC GAGCT GT ATGGCCAGGTGGGAGCTCCTCAAGGGCCCTCCCCAGGAAGTGTTCCTCTGGATGACCTACCTGGGGCAGA GGAGC CAGAAT AT GGAGGAGAT GGCT GT GGT GGGGAGAGACT T AGT CCTGTGTCTTCCC CAC C CAGT GCA GTCCCTGGAAGAAGAATGCCTGCTGCACGGCCAGCACCAGCCAGGAGCTGCACAAGGACACCTCCCGCCT GTACAACTTTAACTGGGATCACTGTGGTAAGATGGAACCCACCTGCAAGCGCCACTTTATCCAGGACAGC TGTCTCTGAGTGCTCACCCAACCTGGGGCCCTGGATCCGGCAGGTCAACCAGAGCTGGCGCAAAGAGCGC ATTCTGAACGTGCCCCTGTGCAAAGAGGACTGTGAGCGCTGGTGGGAGGACTGTCGCACCTCCTACACCT GCAAAAGCAACTGGCACAAAGGCTGGAATTGGACCTCAGGGATTAATGAGTGTCCGGCCGGGGCCCTCTG CAGCACCTTTGAGTCCTACTTCCCCACTCCAGCCGCCCTTTGTGAAGGCCTCTGGAGCCACTCCTTCAAG GT CAGCAACT AT AGT C GAGGGAGC GGC C GCT GCAT C CAGAT GT GGT T T GACT CAGC C CAGGGCAAC C C CA AT GAGGAGGT GGC CAAGT T CT AT GCTGCGGC CAT GAAT GCTGGGGCCCCGTCTCGT GGGAT TAT T GAT T C CTGATCCAAGAAGGGTCCTCTGGGGTTCTTCCAACAACCTATTCTAATAGACAAATCCACATGTG

[0129] NM 002640.4 Homo sapiens serpin family B member 8 (SERPINB8), transcript variant 1, mRNA (SEQ ID NO: 4)AAT CAAGGC GGAC GT GAAGCAT CT ACAAAGGAGGAAT AGT CAAAGCAGCAGC GGCGGCGGCGGCGGCGGC AGCAGCAGCAGCAGCAGGAGACCTTCTCTGATGGATGACCTCTGTGAAGCAAATGGCACTTTTGCCATCA GCTTATTTAAAATATTGGGGGAAGAGGACAACTCAAGAAACGTATTCTTCTCTCCCATGAGCATCTCCTC TGCCCTGGC CAT GGT CT T CAT GGGGGCAAAGGGAAGCACT GCAGC C CAGAT GT C C CAGGCACT T T GT T T A T AC AAAGAC GGAGAT AT TCACC GAGGT TTC CAGT CACTTCT CAGT GAAGT T AAC AGAAC TGGCACTCAGT ACTTGCTTAGAACTGCCAACAGACTCTTTGGAGAAAAGACGTGTGATTTCCTTCCAGACTTTAAAGAATA CTGTCAGAAGTTCTATCAGGCAGAGCTGGAGGAGTTGTCCTTTGCTGAAGACACTGAAGAGTGCAGGAAG CATATAAAT GACT GGGT GGCAGAGAAGACT GAAGGTAAGATTT CAGAGGTACT GGAT GCT GGGACAGT CG AT C C C CT GACAAAGCT GGT C CT T GT GAAT GC CAT T TAT T T CAAGGGAAAGT GGAAT GAGCAAT T T GACAG AAAGT ACACAAGGGGAAT GCT CT T T AAAAC CAAC GAGGAAAAAAAGACAGT GCAGAT GAT GT T T AAGGAA GCT AAGT T T AAAAT GGGGT AT GC GGAT GAGGT ACACAC C CAGGT C CT GGAGCT GC C CT AT GT GGAAGAGG AGCTGAGCATGGTCATTCTGCTTCCCGATGACAACACGGACCTCGCCGTGGTGGAAAAAGCACTTACATA T GAGAAAT T CAAAGC CT GGACAAAT T CAGAAAAGT T GACAAAAAGT AAGGT T CAAGT TTTCCTTCC CAGA TTAAAGCTGGAGGAGAGTTATGACTTGGAGCCTTTCCTTCGAAGATTAGGAATGATCGATGCTTTTGACG AAGCCAAGGCAGACTTTTCTGGAATGTCAACTGAGAAGAATGTGCCTCTGTCCAAGGTTGCCCACAAGTG CT T C GT GGAGGT CAAT GAGGAAGGCACAGAGGCT GC C GCAGC CACT GCT GT GGT CAGGAAT TCCCGGTGCAGCAGAATGGAGCCAAGATTCTGTGCAGACCACCCTTTTCTTTTCTTCATCAGGCACCACAAAACCAACT GCATCTTGTTCTGTGGCAGGTTCTCTTCTCCGTAAAGAGGAGCAATTGCTGTACATACCCTCCTTTCCTT CTACCTATCTTGCCTTAATTAACATTCCCTGTGACCTAGTTGGTGCAGTGGCTTGAATGCCAAAATAAAG C GT GT GCACT GGAT AGT GT GT GAAAGT CT T T GCT GAAAGT T C CAGAGC CAT T GAGAAT AACT GAGGC C GC AGATGCATGAAATTTGGGCCTGGGAAGGCTATGCTGGTTTTGGAGTGTTTGGGGTGCCTGCCATTGCCTC TGCCTTCACCTAAGTCTGTGCCCATTGTTTCAGGGGATCTTATGGAGCATCTCAGGGCTTCAGGAGCCAG CCCTCTTCCATCCGCCCGGCTCTGCCCACCACCACTGCCAGGCTGAACAGGGGACTAGTGCCCAGGTGAC ACTGCACACAAAGCCAAGGGCAAACCCTATAGAGTAAAGCTGCAGCCACCCTGTGTCTCATGTGCAGCTG AAATAGTGATCTGCTTCTGTCACTGTCACATAGACAGCCCTGCATGCCCCCTGTCTCACACAGTTTGTAA TGAAGACAGCTCCTTCTCATCTTTCCATAAGCCTGAGATACAAGTTCAGGGACTCAGCAATGCACTTTAG GACTGAGCTAGGAGGCAAATATCTGAAGCTTGCTATGCTGTTCTTTCCATTCCTTTTCCCTCTGAAACAC ACAAAATACCAAAGGAACTTACGCAACACACCACTGAGTCCTCTAACTAATCATATGTGCTCAGACACAG CTCAAGCACACCCCTTAGTTAAGAAAGAACCTCCATATACATTAATTTTTTTCTGCCTAAAAATAAAATT GC GT T GT GGCAGCAAT T T GGAAACT ACAGCAAAGT CT C CAAAAAAAT T GAAGT CAC C CACAAT T C CAACA C AC AGC AAT CACTGCTGT T AAT AGT GAT AT AC AT CCTTCCCTATTATGTGTAT GC AAAT AGAAAT T T AT A AGCATTTTACTACAAGTTAATCATATCCTTCAGTAGCTTATTTCTTCCCTTATGAAAAACCTCCCTTAAG ATCTTTCCAGGTCATCAATTTGATGTGCATAGTATTCCACTGTGTGGCTGGATTTCTATAATTGAATCTA AT T T T AGAC AGT T T GAGT AAT T AAAAT T AT T T C T AT TAAC AAAT AAT AGAT T GC T AAT AAT AT AT AT TAT GCTGTTTTGGATATACTAAATATTTGCTCACATCCTTAATATATTTTTTAAAATTCCTAACAATAGTACT GT T GAGAT AAAAGT T GAGC AC AT TTTGAGACTTCTTC C AAAT T GGT C C C T AGAAAGT TACACTGGTTTGT ACTCTCACTTATGTCACTGTTTATACCACCACTGACTGCTGCCTGCTTTATTATTTCTTTAATGAGTTGG ACTGAACAGTGGTTAATCCTGACTCTGTTTTTGACTGACAGTTAACAGTTACATGAACCATTCATATTAC AGCT CT TACT T AAAT T T GAC CAAGC CAGGAT AT AT CT GT T AGGC CACAT T CAT T T AGGGAT CAT GT T T T C CAAAGCAGGTTTGGGCAAAATTAATCCACAGGACTGAAAGGTATACATCTGTGAGTTTTGTTCTCACTTC CACCTCTAATTTGAAGAACACTTTAATTGACACAGAATACATTTCACATATTTAACCTCTACAATAAGTT CTGACACATTTTCCATGAAACAAACCATCGCTATATTCAAGATAATGAACCTATCTATCATACTCCCAAA TTCCTTCTTGCATCTTTGTAATTTCTCACTCTTCCTTCTCCCTCTCCCCGTCCCATCCCAACCACTGATC TGCTCAGGCAACTACCAATCTTCTTTCTGTCACTATAGATTAATTTGCATTTTTAAAGAAATTTACATAC ATGGAACCATACATCATCTATGCTTTGTAGTATGACTCCTGTCACTCAGTACAATTATTTTGAGATTCAT T TAT GT TAT T GT AT GT AT CAAT AGT T CAT C C CT T T TAT T GGT AAGT AACAT T T T T T T GT AT AGGT AT AC C AT GAT T T GT T GAT GAACAAAT T TAG CT GT T GAT GAACAT T TAG GT T GT TAG CAAGAT T T T T GCT AT T GAA AAT AAAGT T T T T AT GAAT AT T TAT AT AT AT A

[0130] NM 198833.2 Homo sapiens serpin family B member 8 (SERPINB8), transcript variant 2, mRNA (SEQ ID NO: 5)AAT CAAGGC GGAC GT GAAGCAT CT ACAAAGGAGGAAT AGT CAAAGCAGCAGC GGCGGCGGCGGCGGCGGC AGCAGCAGCAGCAGCAGGAGGTGGGGGCCTCTGCCAGACCTTCTCTGATGGATGACCTCTGTGAAGCAAA TGGCACTTTTGCCATCAGCTTATTTAAAATATTGGGGGAAGAGGACAACTCAAGAAACGTATTCTTCTCT CCCATGAGCATCTCCTCTGCCCTGGCCATGGTCTTCATGGGGGCAAAGGGAAGCACTGCAGCCCAGATGT CCCAGGCACTTTGTT T AT AC AAAGAC GGAGAT AT TCACCGAGGTTTCCAGTCACTTCTCAGT GAAGT T AA CAGAACTGGCACTCAGTACTTGCTTAGAACTGCCAACAGACTCTTTGGAGAAAAGACGTGTGATTTCCTT CCAGACTTTAAAGAATACTGTCAGAAGTTCTATCAGGCAGAGCTGGAGGAGTTGTCCTTTGCTGAAGACA CT GAAGAGT GCAGGAAGCATATAAATGACT GGGT GGCAGAGAAGACT GAAGGTAAGATTT CAGAGGTACT GGAT GCT GGGACAGT C GAT C C C CT GACAAAGCT GGT C CT T GT GAAT GC CAT T TAT T T CAAGGGAAAGT GG AAT GAGCAAT T T GACAGAAAGT ACACAAGGGGAAT GCT CT T T AAAAC CAAC GAGGAAAAAAAGACAGT GC AGAT GAT GT T T AAGGAAGCT AAGT T TAAAAT GGGGT AT GC GGAT GAGGT ACACAC C CAGGT C CT GGAGCT GCCCTATGTGGAAGAGGAGCTGAGCATGGTCATTCTGCTTCCCGATGACAACACGGACCTCGCCGTGGTG GAAAAAGCACT T ACAT AT GAGAAAT T CAAAGC CT GGACAAAT T CAGAAAAGT T GACAAAAAGT AAGGT T C AAGTTTTCCTTCCCAGATTAAAGCTGGAGGAGAGTTATGACTTGGAGCCTTTCCTTCGAAGATTAGGAAT GATCGATGCTTTTGACGAAGCCAAGGCAGACTTTTCTGGAATGTCAACTGAGAAGAATGTGCCTCTGTCC AAGGT T GC C CACAAGT GCT T C GT GGAGGT CAAT GAGGAAGGCACAGAGGCT GC C GCAGC CACTGCTGT GG T CAGGAAT T C C C GGT GCAGCAGAAT GGAGC CAAGAT T CT GT GCAGAC CAC CCTTTTCTTTTCTT CAT CAG GCACCACAAAACCAACTGCATCTTGTTCTGTGGCAGGTTCTCTTCTCCGTAAAGAGGAGCAATTGCTGTA CATACCCTCCTTTCCTTCTACCTATCTTGCCTTAATTAACATTCCCTGTGACCTAGTTGGTGCAGTGGCT T GAAT GC CAAAAT AAAGC GT GT GCACT GGAT AGT GT GT GAAAGT CT T T GCT GAAAGT T C CAGAGC CAT T G AGAAT AACT GAGGC C GCAGAT GCAT GAAAT T T GGGC CT GGGAAGGCT AT GCTGGTTTT GGAGT GT T T GGGGTGCCTGCCATTGCCTCTGCCTTCACCTAAGTCTGTGCCCATTGTTTCAGGGGATCTTATGGAGCATCTC AGGGCTTCAGGAGCCAGCCCTCTTCCATCCGCCCGGCTCTGCCCACCACCACTGCCAGGCTGAACAGGGG ACTAGTGCCCAGGTGACACTGCACACAAAGCCAAGGGCAAACCCTATAGAGTAAAGCTGCAGCCACCCTG TGTCTCATGTGCAGCTGAAATAGTGATCTGCTTCTGTCACTGTCACATAGACAGCCCTGCATGCCCCCTG TCTCACACAGTTTGTAATGAAGACAGCTCCTTCTCATCTTTCCATAAGCCTGAGATACAAGTTCAGGGAC TCAGCAATGCACTTTAGGACTGAGCTAGGAGGCAAATATCTGAAGCTTGCTATGCTGTTCTTTCCATTCC TTTTCCCTCTGAAACACACAAAATACCAAAGGAACTTACGCAACACACCACTGAGTCCTCTAACTAATCA T AT GT GC T C AGAC AC AGC T C AAGC ACAC CCCTTAGT T AAGAAAGAAC C T C CAT AT AC AT T AAT T T T T T T C T GC CT AAAAAT AAAAT T GC GT T GT GGCAGCAAT T T GGAAACT ACAGCAAAGT CT C CAAAAAAAT T GAAGT C AC C C AC AAT T C C AAC AC AC AGC AAT CACTGCTGT T AAT AGT GAT AT AC AT CCTTCCCTATTATGTGTAT GCAAATAGAAATTTATAAGCATTTTACTACAAGTTAATCATATCCTTCAGTAGCTTATTTCTTCCCTTAT GAAAAACCTCCCTTAAGATCTTTCCAGGTCATCAATTTGATGTGCATAGTATTCCACTGTGTGGCTGGAT T T C TAT AAT T GAAT C T AAT TTTAGACAGTTT GAGT AAT T AAAAT T AT T T C T AT T AAC AAAT AAT AGAT T G CTAATAATATATATTATGCTGTTTTGGATATACTAAATATTTGCTCACATCCTTAATATATTTTTTAAAA TTCCTAACAATAGTACTGTTGAGATAAAAGTTGAGCACATTTTGAGACTTCTTCCAAATTGGTCCCTAGA AAGTTACACTGGTTTGTACTCTCACTTATGTCACTGTTTATACCACCACTGACTGCTGCCTGCTTTATTA TTTCTTTAATGAGTTGGACTGAACAGTGGTTAATCCTGACTCTGTTTTTGACTGACAGTTAACAGTTACA T GAAC CAT T CAT AT T ACAGCT CT T ACT T AAAT T T GAC CAAGC CAGGAT AT AT CT GT T AGGC CACAT T CAT T T AGGGAT CAT GT T T T C CAAAGCAGGT T T GGGCAAAAT T AAT C CACAGGACT GAAAGGT AT ACAT CT GT G AGTTTTGTTCTCACTTCCACCTCTAATTTGAAGAACACTTTAATTGACACAGAATACATTTCACATATTT AACCTCTACAATAAGTTCTGACACATTTTCCATGAAACAAACCATCGCTATATTCAAGATAATGAACCTA TCTATCATACTCCCAAATTCCTTCTTGCATCTTTGTAATTTCTCACTCTTCCTTCTCCCTCTCCCCGTCC CATCCCAACCACTGATCTGCTCAGGCAACTACCAATCTTCTTTCTGTCACTATAGATTAATTTGCATTTT TAAAGAAATTTACATACATGGAACCATACATCATCTATGCTTTGTAGTATGACTCCTGTCACTCAGTACA AT T AT T T T GAGAT T CAT T TAT GT TAT T GT AT GT AT C AAT AGT T CAT CCCTTTTATT GGT AAGT AAC AT T T TTTTGTATAGGTATAC CAT GAT T T GT T GAT GAAC AAAT TTACCTGTT GAT GAAC AT TTACGTTGTTACCA AGAT TTTTGCTATT GAAAAT AAAGT T T T T AT GAAT AT T TAT AT AT AT A

[0131] NM 006770.4 Homo sapiens macrophage receptor with collagenous structure (MARCO), mRNA (SEQ ID NO: 6)ACAACCAGCTGCAGTGGTTCGATGGGAAGGATCTTTCTCCAAGTGGTTCCTCTTGAGGGGAGCATTTCTG CTGGCTCCAGGACTTTGGCCATCTATAAAGCTTGGCAATGAGAAATAAGAAAATTCTCAAGGAGGACGAG CT CT T GAGT GAGAC C CAACAAGCT GCT T T T CAC CAAAT T GCAAT GGAGC CT T T C GAAAT CAAT GT T C CAA AGCCCAAGAGGAGAAATGGGGTGAACTTCTCCCTAGCTGTGGTGGTCATCTACCTGATCCTGCTCACCGC TGGCGCTGGGCTGCTGGTGGTCCAAGTTCTGAATCTGCAGGCGCGGCTCCGGGTCCTGGAGATGTATTTC CTCAATGACACTCTGGCGGCTGAGGACAGCCCGTCCTTCTCCTTGCTGCAGTCAGCACACCCTGGAGAAC ACCTGGCTCAGGGTGCATCGAGGCTGCAAGTCCTGCAGGCCCAACTCACCTGGGTCCGCGTCAGCCATGA GCACTTGCTGCAGCGGGTAGACAACTTCACTCAGAACCCAGGGATGTTCAGAATCAAAGGTGAACAAGGC GC C C CAGGT CT T CAAGGT CACAAGGGGGC CAT GGGCAT GCCTGGTGCCCCTGGCCCGCC GGGAC CAC CT G CTGAGAAGGGAGCCAAGGGGGCTATGGGACGAGATGGAGCAACAGGCCCCTCGGGACCCCAAGGCCCACC GGGAGT CAAGGGAGAGGC GGGC CT C CAAGGAC C C CAGGGT GCT C CAGGGAAGCAAGGAGC CACT GGCAC C C CAGGAC C C CAAGGAGAGAAGGGCAGCAAAGGC GAT GGGGGT CT CAT T GGC C CAAAAGGGGAAACT GGAA CT AAGGGAGAGAAAGGAGAC CT GGGT CT C C CAGGAAGCAAAGGGGACAGGGGCAT GAAAGGAGAT GCAGG GGT CAT GGGGC CT C CT GGAGC C CAGGGGAGT AAAGGT GACT T C GGGAGGC CAGGC C CAC CAGGT T T GGCT GGTTTTCCTGGAGCTAAAGGAGATCAAGGACAACCTGGACTGCAGGGTGTTCCGGGCCCTCCTGGTGCAG TGGGACACCCAGGTGCCAAGGGTGAGCCTGGCAGTGCTGGCTCCCCTGGGCGAGCAGGACTTCCAGGGAG C C C C GGGAGT C CAGGAGC CACAGGC CT GAAAGGAAGCAAAGGGGACACAGGACT T CAAGGACAGCAAGGA AGAAAAGGAGAATCAGGAGTTCCAGGCCCTGCAGGTGTGAAGGGAGAACAGGGGAGCCCAGGGCTGGCAG GTCCCAAGGGAGCCCCTGGACAAGCTGGCCAGAAGGGAGACCAGGGAGTGAAAGGATCTTCTGGGGAGCA AGGAGT AAAGGGAGAAAAAGGT GAAAGAGGT GAAAACT CAGT GT C C GT CAGGAT T GT C GGGAGT AGT AAC CGAGGCCGGGCTGAAGTTTACTACAGTGGTACCTGGGGGACAATTTGCGATGACGAGTGGCAAAATTCTG ATGCCATTGTCTTCTGCCGCATGCTGGGTTACTCCAAAGGAAGGGCCCTGTACAAAGTGGGAGCTGGCAC T GGGCAGAT CT GGCT GGAT AAT GT T CAGT GT C GGGGCAC GGAGAGT AC C CT GT GGAGCT GCAC CAAGAAT AGCT GGGGC CAT CAT GACT GGAGC CAC GAGGAGGAC GCAGGC GT GGAGT GGAGC GT CT GAC C C GGAAAC C CTTTCACTTCTCTGCTCCCGAGGTGTCCTCGGGCTCATATGTGGGAAGGCAGAGGATCTCTGAGGAGTTC CCTGGGGACAACTGAGCAGCCTCTGGAGAGGGGCCATTAATAAAGCTCAACATCATTGGC

[0132] NM 001184879.2 Homo sapiens CD84 molecule (CD84), transcript variant 1, mRNA (SEQ ID NO:7)AAACTGACTCTGCTAGAACAGTGCCGTGCTTTTCCACAGAAGGTTAGACCCTGAAAGAGATGGCTCAGCA CCACCTATGGATCTTGCTCCTTTGCCTGCAAACCTGGCCGGAAGCAGCTGGAAAAGACTCAGAAATCTTC ACAGT GAAT GGGAT T CT GGGAGAGT CAGT CACT T T C C CT GT AAAT AT C CAAGAAC CAC GGCAAGT T AAAA TCATTGCTTGGACTTCTAAAACATCTGTTGCTTATGTAACACCAGGAGACTCAGAAACAGCACCCGTAGT TACT GT GAC C CACAGAAAT TAT TAT GAAC GGAT ACAT GC CT T AGGT C C GAACT ACAAT CT GGT CAT T AGC GAT CT GAGGAT GGAAGAC GCAGGAGACTACAAAGCAGACAT AAAT ACACAGGCT GAT C C CT ACAC CAC CA C CAAGC GCT ACAAC CT GCAAAT CT AT CGTCGGCTT GGGAAAC CAAAAAT T ACACAGAGT T T AAT GGCAT C T GT GAACAGCACCT GTAAT GT CACACT GACAT GCT CT GTAGAGAAAGAAGAAAAGAAT GT GACAT ACAAT TGGAGTCCCCTGGGAGAAGAGGGTAATGTCCTTCAAATCTTCCAGACTCCTGAGGACCAAGAGCTGACTT ACACGTGTACAGCCCAGAACCCTGTCAGCAACAATTCTGACTCCATCTCTGCCCGGCAGCTCTGTGCAGA CATCGCAATGGGCTTCCGTACTCACCACACCGGGTTGCTGAGCGTGCTGGCTATGTTCTTTCTGCTTGTT C T C AT T C T GT C T T C AGT GT T T T T GT T C C GT T T GT T C AAGAGAAGAC AAGGT AGGAT T T T C C C AGAAGGT T CCTGCTTGAACACCTTCACTAAGAACCCTTATGCTGCCTCAAAGAAAACCATATACACATATATCATGGC T T CAAGGAACAC C CAGC CAGCAGAGT C CAGAAT CT AT GAT GAAAT C CT GCAGT C CAAGGT GCTTCCCTCC AAGGAAGAGC CAGT GAACACAGT T TAT T C C GAAGT GCAGT T T GCT GAT AAGAT GGGGAAAGC CAGCACAC AGGACAGTAAACCTCCTGGGACTTCAAGCTATGAAATTGTGATCTAGGCTGCTGGGCTGAATTCTCCCTC TGGAAACTGAGTTACAACCACCAATACTGGCAGGTTCCCTGGATCCAGATCTTCTCTGCCCAACTCTTAC TGGGAGATTGCAAACTGCCACATCTCAGCCTGTAAGCAAAGCAGGAAACCTTCTGCTGGGCATAGCTTGT GCCTAAATGGACAAATGGATGCATACCCTTCCTGAAATGACTCCCTTCTGAATGAATGACAAAGCAGGTT ACCTAGTATAGTTTTCCCAAACTTCTTCCCATCATAGCACATGTAGAAAATAATATTTTTATGGCACACT GGGATAAACAAGCAAGATT GCT CACTT CT GGAAGCT GCATAT GACTAGAGGCCT CTT GT GACT GGAGGTA ACAAC C CT GC C CAGT AACT GT GGGAGAAGGGGAT CAAT AT T T T GCACAC CT GT AAT AGGC CAT GGCACAC CAGCCAAGATGCTCTGCTCACAGTCAGTATGTGTGAAGATCCCTGGTGCGTGGCCTTCACCACGCATCTT GAGCAAATTAGGAAAATGTACCCTTCGCTTGAGGCAGATGCAGCCCTTCCCCCGAGTGCATGGCTTGGAG AGCAGAATGTGGGCTGCATATAAGCACACTCATCCCTTTGTCTGGGAATCTTTGTGCAGGGCATAACAGG CTTAGTAAGT CCAAACACAGAT GACAGT GCT GT GT GGGT CT CT GT CAGAGTT GT GGCT CT CAGCCAT GTA GACACACTCTCCAAATGGAGTGTTGGAAAATGTTCTTTCTGCAGGGTCTAGAGACTGCTGGGACACTTTT CTTGGAGTGCTACTTCAGAAGCCTTATAGGATTTTCTTTCTGGCCAAGATTTCCTTCTGTATCACTCCAA GCAGCCTCAGCAGAAGAAGCAGCCATGCCCAGTATTCCCACTCTCCAAAAGGAACTGACCAGCTTATATT TCTCACACTTCTGGGGAACTGGGTATAATCCAACCATCAAAATAGAAGACCTTGCAAGAAGCAGAGTCAT T CT C CAGAAGGAACT T GGGAGAT GAT GGT GCAGAT GAT GAAACT GGGT T CAT C C CAGT T C CAAAGACT CA GAGAACT AGAGT T T AAGCT GAGGCAGAGT GC C GC CAC C CT GGCAT GC C C CACAAACAGAT CAC CAGC CAG CTTACACAGGCATTAACTCTCCTCAATGAGGAAGAATCATTCACAACTGAGCAAGACATTCATATGATCA TTTAAGGAAGTGTTTCCCTTATGTGTTAGCAAGTATAATCGGCTAACTCCTAAATCCCAATGAATAGTCC TAGGCT GGACAGCAAT GGGCT GCAATTAGGCAGATAAAGACAT CAGT CC CAGT AAAT GAAT CCATAGACT CAT CT AGCAC CAACT AC CAT T AGCACT AT GT T AGGAGCT GCAAGGC C C CAAAGT AGAAGAT GT GCAT AAT GTCTGCTCTTGTGTAGCTCAGGAGACAATTCCAGCACAGACACTACAGTTAACGCTGAACTGCAGCTGCA AGT AAT AGCAT GAACAGT CAGAAAAAT AC CT T AT GAGGGGGCAGGGCT GAAGCT GGGC CT T GAAGGAT GG AT GAAAT T T GGAT AGAGAAT GAGGAAGACAGAGGGC CT C CAAGT GAGAGAAGCAT GAAAAAT GAGCAGGG GCCTGGATCAGTGGGGTGTATTCAGAGCACCTCTCCAGATGCACCATGCATGCTCACAGTCCCTTGCCTA T GT GT GGCAGAGT GT C C CAGC CAGAT GT GT GC C CT CAC C C CAT GT C CAT T T ACAT GT C CT T CAAT GC C CA CCTCAAAAGGTACCTCTTCTGTAAAGCTTTCCCTGGTATCAGGAATCAAAATTAATCAGGGATCTTTTCA CACTGCTGTTTTTTCCTCTTTGGTCCTTCTATCACTAAAACTCATCTCATTCAGCCTTACAGCATAACTA AT TAT T T GT T T T C CT CACT ACAT T GT ACAT GT GGGAAT T ACAGAT AAAC GGAAGC CGGCTGGGGTGGTGG CT CAC GC CT GT AAT C C CAACACT T T GGGAGGC CAAGGCAGGC GGAT CAC CT GAGGT CAGGAGT T C GAGAT T AGT CT GGC CAACAT GGT GAAAC C C CAT CT CT ACT AAAAAT AC GAAAT T AGC CAGGT GT GGT GGCACACA TCTGTAGTCCCAGCTACTCTGGAGGCTGAGACAGGAGAATCGCTTGAACCCAGGAAGTGGAGGTTGCAGT GAGCTGAGATCACACCACTGCACTCCAGCCTGGGAGAGACAGAGTGAGACTCCATCTCGAAAAAAAAAAA AAGAT AGAAGC CAAT AAGCAT GGT GCAAT CAAAT T CT GGCAAGCAT T AAAT AT CAGGAT GCAGCT GGGCA C GGT GGCT CAC GC CT GT AAT C C CAGCACT T T GGGAGGC CAAGGT GGGC GGAT CACTT GAGGT CAGGAAT T T GAGAGGAT C CT GGC CAGCAT GGCAAAAC C C CAT CT GT ACT T AAAAT ACAAAAAAAT T AGCT GGGC GT GG TGGTGCACACCTGTAATCCCAGCTACTTGGGAGGCTGAGGTGGGAGAATTGCTTGAACCTGGGAGGTGGA GGTTGCAGTGAGCTGAGATCCTGCCACTGCACTCCAGGCTGGGCAACAGAGTGAGACCATGTCTCAAAAAAT AAAAAT AAAAT AAAAT AAT AT CAGGAT GCAT ACAT CAGAGGCT GT T C CT AGT GT AAAGGCACT T T GGA GGGAGAAGACT T T CAGAGT T AGGCAGAC CAACT AAGAGGT CAGCT GAAGCAC CT AAC CAGT T GT AAGGAG GT GAAAGACAGCAC C C CAAGAAGAGAC GT GCAGGAAGGAGGAAAGAGGCT T GGT CAT AAAGGAT GGAGGA ATTCCAAAGTGACACTGAACAGGCTGCGTTTATCCTAAAATAAAACCACTCCTCACTCTGTGGATGCGTT GAAGACTCATTCCCAAACATCTTTATTCTCTAACTTGCCCTCTTCCTCTTCCTAATATGCTCACTCAAGT AAAAT TACTAGTGTCC T AAT GC C C C TAT GCAT AT T GT C AAAAAT AAAAAT C AGAAGC AGGT T AGAT C T GT TAGGTCTTCCAGAAGAGCAAACCTGGGATGAAGCCAGAGCCCAGGAATTCTGAAGGTAGCCTTTGGACTC AGGACACCCTACTCTTGTCTCTCCTCTCAGTTTCTCTGCTATGAATCTCCTGATTCATGAACACGTTATC TGTTCACCCTTCTCTCTAGGTCTTAGTTCTTAGATTTTCCTTCTGTAAAATGCATGTGATCTTATTTTCC CCTCCACAACTTTCCAGATGAACTAGACTGTGACCAAGAGGTCTATAAAATCAAAGCATCATGGAACAGG AT CT T GT AT CAGAC CAAAGT GT GC CAGT T T T T AAAAAT GT GCAT CAAAAT GGAAGT CT CAGAGACAGAGC CCTCTGGTGGAAAGTTCTAGTAGGTTAGGACAGTCCTGCCTGCAGACACCTTGGGCTTTACTGAGGGACT C AAC T GAGAAAAT GAGGAAT GT T GC AGC T CAT GAT T C T T AGAAGAAGAAAGT GAAGC T T GT T T AAAAT AT GATTTAAAAAATCTGTAGAACACTGTAAACTACACAGGCTATGAGGGAATAGCCTGGTTGGGCCAGCTTG GAAATCGGGCACAGGCAGGAAGGGGCCTGTCTGGTTTGGGCCGTGTCCACAGAGAGCACTTCTTAGGTCC TGCCTGGAGAGAAGGAATGGCTGGGCTATATTTTCTTCCAGACTCATTATTTTTCTTCTGTTTGACTTTT CTCTGAATTTCCCTTGATTTGTATAAATTTTCTCAATAATTAGTGACAGTGTCTACTGATTGTAAAATGA AGCT T GAAGGC CAGGC GCAGT GGCT CAT GC CT GT AAT C CT AGAAT T T T GGGAGGC CAAGGT GGGT GGAT C ACAAGGAGT T C GAGAC CAGC CT GGC CAAGGT T GT GAAAC C GC GT CT CT ACT AAAAAT ACAAAAAAAT TAG CCGGGCATGGTGGCACGTGCCTGTAGTCCCAGCTACTCAGGAAGCTGAGGCAACAGAATCACTTGAACCT GGGAGGTGGAGGTTGCAGTGAGCCGAGATCACGCCACTGCACTCCAGCCTGGGCGAGAGAGTGAGACTCC GT CT CAAAAAAAAAAAAAAAAAAAAAAGT T T GAAGAACAAAGACAAT AAGAGGAAAT AT AAT GAGT GGT C ATAAATGTGGGCTCTGACAGTAGAGTGCCTGGGTCTGCATCCTGGTTTCTTAGTCATGTGACCTTAGGCA AGTTACTTTAACCTCGCTGTACCTCAGGTTGTCCATCTGTAAAATGGGGATAATAATAGTGCCTACCTTT T AAGGT T GAT GT GGGGAT T AAAT GAGGT GT T GCT CAT ACAGGAAT GT GC CT GT GCAT GGCAAAGT T C GGG AAATTTTTTATAAGCTGTTCTAGGCCTGAAATCTTCAGAAGATGCTAATCTAAATTCATGAAATAAGCTT CTTACAACAGAAATGCTGCTAGTATTATGCAAAATTAATGTTGTATATCAAACTTTTAACTCTCATCCCT CCTTATTCAGATATATTTTGTTATAAGCAATGTTTGTTCCCCTCGTTATTATACCACAGTCTACTTACCT GATGCTATATCTGCCTCCCCAGTTAGACTGAGAGAACAGGGGATATACCTAAATAATAATAATAATAATA AT AAT AAT AAAT AAT AAT GGAGAGCT C CT T GAAGAT AGGGAGC CT GT AAGAAT CAT T GAGGGCT T AT T T T GTATACCAACTGCTAAACTAGATGCTTCATACATTGTTGTCAATACTCATGACAGCCTTGTAAAGTAGAA ATTAATTCTTCCAGTTAACACTAAGGCTGACATATGAATACCTTGGCAAATCTGGAAAGCTGGGAAGACA GTATTTGAATTCAAGACTTCTTGTCACCAAGGGCCATGCACTTGTACTCTGCCATGTGGCCCTTTTTTAC CTCCTGTGGATTCTCCCTACCTGGTACTTGGCCTTAGGTGTACACACACCTGGCACTTTGCTTGACACAT AAT AGGT GGAC CACAAAT AT CT ACT AAAT GAAT AT T T GCAT AT AGT AAT AT T T T AAGGT ACT AAAAGCAG CTCAAAGTAAATATATTAATATATTAATTCCATTGCTATCTGGATAACCACTCAACTTTCCTGCTGAAAA T GC C CAT T T AAT T AAAGAAGGT T GGAT AGAGCT CT CT AT AT GCAT T T T GGACAGGCAGGGGT T T CAGGT C AT AAAC AT T C T GAT GAGT T AAT AT AAAAT AAGAGAAAC T GT AAAT TTCCACTAC T AAAAAT C AC AAAAAT AACAGAAACAAAAGAAGAGAT AAGAAT T T GGGGAAT T GT GCT GAACAAT T T AGT GGT T AAAAAAAACAAC T GT GCAT GT T T AGACT T AAAT AAGC CC C CAT C CAAGT GT GAGGGGT C CAGT AAT T T T T CAAAACAT AT GA AAGT GT T AAT ACAT T T C GACAAAGGAC CAT T AAAAAAGT C CT GAAT T CT GACT T GAGGGAGGAAAGT AAT GACTAATACATTCTCTAGAGACTTGCAGACTTTGGGAATTCATAAAGGAATGGATGATAATTATTAACTG TTGCTGGCTGATTGCCCAGACAGTTCTCAACAGCCCTGTACAAGTCTCTGGGTTTGGGATGGATCAATTC T GAGACT GGAAAAT GGC CAAAT CT T T GCAAAT GAGAAAT AT T T T T CT TAT AAGT T CT TAT T GT AGGCAAA TAATTACATAGATTATTCATCAGAGAATTTTTAAATGCTCATAATCTCAACTCTTTCATTTACAACTTGT AT T T C CAAT AGT T TAT GGGT CAT CT CT GCAT AGAT GT CAGAAGT CAC CT CAAGT T T AGC GT GT C CAAAAT CTAACTCACAGGTCTGTTTCTGACCTCCCAACTTGCTTTCCTTGTGTTTTTCCTATGCTAATGATCCACC ATAATCAAAATAATTAACATTTATCCAGTGCCTACTATGTACTATTCCCTGTCCTGTTTTACATTTACTC ATTTAAAGTCCATAAGAAACATTAAATCTCATCTGCCTTCTGAAGAAGATACAACCATGCTCTCTTTTAC AAAGTAGGAAACT GGGT CACAGAAAGGT GAAGT CTTTAAGGCT GAAT CACAGT AGCT CAT CCTAGT AAAT AGAAAAGCCAGGATTCAACTCCAGGGGCTGGGTGCAGAACTGCTATTCTTCACTGCTTCACCAATCAGCA GCTACCCAAGGCAGAAAACTTTTTCATCCTTGGCTCCTTCATTCTCCCTGTCACCCCAGATCCCCTCTAC ATCTAGTCAGAGAATAGGTCCTGTCAATTCCAACTTCTCTATATGGCTCCTCTCAGGCATGTGCCCTTAA TTGGCCTAATTCTCTAATACACCTTCCCTCTACATGCTCACTCCCTCAGATCATTGCTTTATCACGTGTT ACCTGGGTTGCTATTACATAAAGAGCAATCTTTCTAAAATGAGGATCTTATCACTTCACTTCCACACTAA AATGTTTTTCCTGGGGAACCACACTCCTTAGCAATCTGACCCATCAGACCTTCCAGGCTGTCTCCTGCCT GCTCCCTAAGGCTCCAGCCACACAGAATTATCATGGGCCCACACACCCACCAAATCCTCCCATGCCTTTG CCCATGTTGTCTGGGATGCCCTTCTCTCCCTTCTGTCTACATCAAGCATCAGACTGAATATCCCTCTTGT GCGGCCTTCTAAAACCTCCCGTCCAAAGCGAAATATATTGCCCTCTATTTATACTTTTACAGCATTTGGCACACAAGTACAGAGTAGTAGCTTTTTATCACATTCTCTGATAATTATATAGATATGGTATTTCTTAGCTC TCTCTCCAACTGGCTAATAAGTTGCTTTTTGTCTGAGTGCCTAATTTTGTGTTTTGTGTCTGAGTGCCTC AGT T C C T C AAAAAAAGGT T T T T T GAT T AGT T CAT T AT T CAT T T GAAC AT GGAAAT TATGCTCACTAGTGG C AAAT GCCACTAACCGTATTC C AGAAGC T AGGT GT CAT GT T T GC AAT AAGAT AT AT TATCCCTTC T AC AA GT CAC CT T T TAT T T CAGGCAT T T GT AAAT GC C CAT T AAT AAAGT AT GGT T CAT AAAT T T TAG CT T GT AAG T GC CT AAGAAAT GAGACT ACAAGCT CCAT T T CAGCAGGACACAAT AAAT AT TAT T T TAT AAT GCA

[0133] NM 012445.4 Homo sapiens spondin 2 (SPON2), transcript variant 1, mRNA (SEQ ID NO:8)ACCCGACCGCTGCCGGCCGCGCTCCCGCTGCTCCTGCCGGGTGATGGAAAACCCCAGCCCGGCCGCCGCC CTGGGCAAGGCCCTCTGCGCTCTCCTCCTGGCCACTCTCGGCGCCGCCGGCCAGCCTCTTGGGGGAGAGT C CAT CT GT T C C GC CAGAGC C CT GGC CAAAT ACAGCAT CAC CT T CAC GGGCAAGT GGAGC CAGAC GGC CT T CCCCAAGCAGTACCCCCTGTTCCGCCCCCCTGCGCAGTGGTCTTCGCTGCTGGGGGCCGCGCATAGCTCC GACT ACAGCAT GT GGAGGAAGAAC CAGT AC GT CAGT AAC GGGCTGCGC GACT T TGC GGAGC GC GGC GAGG C CT GGGC GCT GAT GAAGGAGAT C GAGGC GGC GGGGGAGGC GCT GCAGAGC GT GCAC GCGGTGTTTTCGGC GCCCGCCGTCCC CAGC GGCAC C GGGCAGAC GT C GGC GGAGCT GGAGGT GCAGC GCAGGCACT C GCT GGT C TCGTTTGTGGTGCGCATCGTGCCCAGCCCCGACTGGTTCGTGGGCGTGGACAGCCTGGACCTGTGCGACG GGGAC C GT T GGC GGGAACAGGC GGC GCT GGAC CT GT AC C C CT AC GAC GC C GGGAC GGACAGC GGCT T CAC CTTCTCCTCCCCCAACTTCGCCACCATCCCGCAGGACACGGTGACCGAGATAACGTCCTCCTCTCCCAGC CACCCGGCCAACTCCTTCTACTACCCGCGGCTGAAGGCCCTGCCTCCCATCGCCAGGGTGACACTGGTGC GGCT GC GACAGAGC C C CAGGGC CT T CAT CCCTCCCGCCC CAGT C CT GC C CAGCAGGGACAAT GAGAT T GT AGACAGCGCCTCAGTTCCAGAAACGCCGCTGGACTGCGAGGTCTCCCTGTGGTCGTCCTGGGGACTGTGC GGAGGCCACTGTGGGAGGCTCGGGACCAAGAGCAGGACTCGCTACGTCCGGGTCCAGCCCGCCAACAACG GGAGCCCCTGCCCCGAGCTCGAAGAAGAGGCTGAGTGCGTCCCTGATAACTGCGTCTAAGACCAGAGCCC CGCAGCCCCTGGGGCCCCCCGGAGCCATGGGGTGTCGGGGGCTCCTGTGCAGGCTCATGCTGCAGGCGGC CGAGGGCACAGGGGGTTTCGCGCTGCTCCTGACCGCGGTGAGGCCGCGCCGACCATCTCTGCACTGAAGG GCCCTCTGGTGGCCGGCACGGGCATTGGGAAACAGCCTCCTCCTTTCCCAACCTTGCTTCTTAGGGGCCC CCGTGTCCCGTCTGCTCTCAGCCTCCTCCTCCTGCAGGATAAAGTCATCCCCAAGGCTCCAGCTACTCTA AATTATGTCTCCTTATAAGTTATTGCTGCTCCAGGAGATTGTCCTTCATCGTCCAGGGGCCTGGCTCCCA CGTGGTTGCAGATACCTCAGACCTGGTGCTCTAGGCTGTGCTGAGCCCACTCTCCCGAGGGCGCATCCAA GCGGGGGC CACT T GAGAAGT GAAT AAAT GGGGCGGTTTC GGAAGC GT CAGT GT T T C CAT GT TAT GGAT CT CT CT GCGTTT GAATAAAGACTAT CT CT GTT GCT CACAAA

[0134] NM 138806.4 Homo sapiens CD200 receptor 1 (CD200R1), transcript variant 1, mRNA (SEQ ID NO: 9)ATACAGGAAATCTATTGCTGTGTCAAGTTCCAGAGAAAAGCTTCTGTTCGTCCAAGTTACTAACCAGGCT AAAC CACAT AGAC GT GAAGGAAGGGGCT AGAAGGAAGGGAGT GC C C CACT GT T GAT GGGGT AAGAGGAT C CTGTACTGAGAAGTTGACCAGAGAGGGTCTCACCATGCGCACAGTTCCTTCTGTACCTGTGTGGAGGAAA AGTACT GAGT GAAGGGCAGAAAAAGAGAAAACAGAAAT GCT CT GCCCTT GGAGAACT GCTAACCTAGGGC TACTGTTGATTTTGACTATCTTCTTAGTGGCCGAAGCGGAGGGTGCTGCTCAACCAAACAACTCATTAAT GCT GCAAACT AGCAAGGAGAAT CAT GCT T T AGCT T CAAGCAGT T TAT GT AT GGAT GAAAAACAGAT TACA CAGAACTACTCGAAAGTACTCGCAGAAGTTAACACTTCATGGCCTGTAAAGATGGCTACAAATGCTGTGC T T T GT T GC C CT C CT AT C GCAT T AAGAAAT T T GAT CAT AAT AACAT GGGAAAT AAT C CT GAGAGGC CAGC C T T C C T GC AC AAAAGC C TAG AAGAAAGAAAC AAAT GAGAC C AAGGAAAC C AAC T GT AC T GAT GAGAGAAT A ACCTGGGTCTCCAGACCTGATCAGAATTCGGACCTTCAGATTCGTACCGTGGCCATCACTCATGACGGGT AT T ACAGAT GCAT AAT GGT AACAC CT GAT GGGAAT T T C CAT C GT GGAT AT CAC CT C CAAGT GT T AGT TAG AC CT GAAGT GAC C CT GT T T CAAAACAGGAAT AGAACT GCAGT AT GCAAGGCAGT T GCAGGGAAGC CAGCT GC GCAT AT CT C CT GGAT C C CAGAGGGC GAT T GT GC CACT AAGCAAGAAT ACT GGAGCAAT GGCACAGT GA CTGTTAAGAGTACATGCCACTGGGAGGTCCACAATGTGTCTACCGTGACCTGCCACGTCTCCCATTTGAC T GGCAACAAGAGT CT GT ACAT AGAGCT ACT T C CT GT T C CAGGT GC CAAAAAAT CAGCAAAAT TAT AT AT T C CAT AT AT CAT C CT TACT AT TAT T ATT T T GAC CAT C GT GGGAT T CAT T T GGT T GT T GAAAGT CAAT GGCT GCAGAAAAT AT AAAT T GAAT AAAACAGAAT CT ACT C CAGT T GT T GAGGAGGAT GAAAT GCAGC C CT AT GCCAGCTACACAGAGAAGAACAAT GCT CT CTAT GATACTACAAACAAGGT GAAGGCAT CT GAGGCATTACAA AGT GAAGT T GAC AC AGAC C T C CAT ACT TTATAAGTTGTTGGACTCTAGTAC C AAGAAAC AAC AAC AAAC G AGATACATTATAATTACTGTCTGATTTTCTTACAGTTCTAGAATGAAGACTTATATTGAAATTAGGTTTT C C AAGGT T C T T AGAAGAC AT T T T AAT GGAT T C T C AT T C AT AC C C T T GT AT AAT T GGAAT T T T T GAT T C T T AGCTGCTACCAGCTAGTTCTCTGAAGAACTGATGTTATTACAAAGAAAATACATGCCCATGACCAAATAT T CAAAT T GT GCAGGACAGT AAAT AAT GAAAAC CAAAT T T C CT CAAGAAAT AACT GAAGAAGGAGCAAGT G T GAAC AGT TTCTTGTGTATCCTTT C AGAAT AT T T T AAT GT AC AT AT GAC AT GTGTATATGCCTATGGTAT AT GT GT CAAT T TAT GT GT C C C CT T ACAT AT ACAT GCACAT AT CT T T GT CAAGGCAC CAGT GGGAACAAT A CACTGCATTACTGTTCTATACATATGAAAACCTAATAATATAAGTCTTAGAGATCATTTTATATCATGAC AAGTAGAGCTACCT CAT T C T T T T T AAT GGT T AT AT AAAAT T C CAT T GT AT AGT TAT AT CAT T AT T T AAT T AAAAAC AAC C C T AAT GAT GGAT AT T T AGAT T C T T T T AAGT T T T GT T T AT T T C T T T T AAGT T T T GT T T GT G GTATAAACAATACCACATAGAATGTTTCTTGTGCATATATCTCTTTGTTTTTGAGTATATCTGTAGGATA ACT T T CT T GAGT GGAAT T GT CAGGT CAAAGGGT T T GT GCAT T T TACT AT T GAT AT AT AT GT T AAAT T GT G T C AAAT AT AT AT GT C AAAT T C C C T C CAAC AT T GT T T AAAT GT GC C T T T C C C T AAAT T T C T AT T T T AAT AA CTGTACTATTCCTGCTTCTACAGTTGCCACTTTCTCTTTTTAATCAACCAGATTAAATATGATGTGAGAT T AT AAT AAGAAT TAT ACT AT T T AAT AAAAAT GGAT T TAT AT T T T T GGT CAT GT T T GT AAGAGAGT GAAT G CACGTGTGAGAACATTAGCTTCTTCTGAACTCATTATATCTCCACAGAGGTGTTGATACTTGATGCCTAA CAGT T T T GCAGAT GT GCT ACAT T GGAAT T GT GT AT T T T TAT GGT GT ACAT T CT AT T GT GAT AT AT T TAT T GAAT AAT T AAT GT CT AT T GAC CAT AT AAGT GGC GAAAAAT GCAC CAT AGAGGACAT GGGGT AT T TAT T TA CAAACT AT GAGCT ACAT AAT AAGCAAGT GGC CAT GGGAT GGCAT GAC CCTCCCCTC CAT AT T T T T GT GGA GCAAAAT AT T GGCAAT GT T TAT GT AAAT CAT T GT T AAT AT CAT GAAAT TAT T T T T AAT T AAAAACAT AAG TCTATTTGCTC C AT AGC AGAAAAAACAT GAGAAGT T T T T T CAT CAT GAT AGAAAT T GAAAC AAAAT AT AT TCATTCTTCAATCATACCATCTGAGATTTTTAAGACAGCTATTTTGTCTTATAAGTATATTTTTCTCCCT CTAGACATTTCAGTTACTATGGATTTTGTCCTCAAAGGGACTTTTAGTCTATTTTGGATGTAAAGCTAAT CT AAT GACACT T GGCACAT GAT AT T TT GAT CAAGC CAT T T T GACT T GAC CAAAAAGCAGT GT C CAT T AGG T T T C T GCAT AT AAAT AT TAG CAAGC AAT GT T C AC AAT AGAC AT CAT TACACTGTCCTT GAAAT T T AT T AA TTCTTCATCCAACCCTGGTTGAGCTGAGGCTCATAGTTAGGTTCAAGACTATCTGTTTAAATATTACTGA AAAACAAAGT AAGAGAGT AC TAT GCTTACCT CTTAACTT GATAAT GT CAAAACAGGCAT GTTAAAT GAGA T C AT AGAAAAGAC T T C AAGAT AAT T T AT AGAAGT T AAAT TAT AT T GT AC AGAAAAT AAT T GT AT GAAAAT CTCTACTATGGGGCTGGAACATGGTTGAACATTAGAATGATATAAAAAATTATATATATTCTCCAAATCC AC GCT AGAC CT GT CAAAT T AGAGAAT CT AGAGAT T AGAC CT GGC GT GT CAGCAAGGT CAT C CAGGAAGCA GAGGCT GAGAC GGAGT T AGGT GT GATT ACT TAG AT AGT C GAT T ACAT T T T AGAAAT AACAT T T TAT AT GT CTCATTTACTGTGCTTTCTCCCCATCCCATTTTGTATCTTTTCCTTTGCTTTGCTAGATTTGTCAATTTT CTCTCTCTTTCTCTGTCTCTCTCTCTTTCAATATCTCTAATAATTTGAAAGTAATTCATCATAACTAAAT ATCTATTGGGGTTATGCTTCACTTACAAACTTCTGAAAACGGCTTTACTGAGATATAATTGATATATTTA AGT GT ACAGT T T GT T AAAT T T T GCACAT AT T T AAAAT GT GGAGT T T GGT AAAT GT T GACAT AGT T T TACA T C T GT GAAAC CAT C AGC AT AAT C AAGAT AAT AAAC T T GT C CAT C AC C C C C C AAAA

[0135] NM 004932.4 Homo sapiens cadherin 6 (CDH6), transcript variant 1, mRNA(SEQ ID NO: 10)CCCACTTCATTCACTTGCAAATCAGTGTGTGCCCACAAGAGCCAGCTCTCCCGAGCCCGTAACCTTCGCA TGC CAAGAGCT GGAGT T T CAGC GGC GACAGCAAGAAC GGCAGAGC C GGC GAC CGCGGCGGCGGCGGCGGC GGAGGCAGGAGCAGCCTGGGCGGGTCGCAGGGTCTCCGCGGGCGCAGGAAGGCGAGCAGAGATATCCTCT GAGAGC CAAGCAAAGAACAT T AAGGAAGGAAGGAGGAAT GAGGCT GGAT AC GGT GGAGT GAAAAAGGCAC TTCCAAGAGTGGGGCACTCACTACGCACAGACTCGACGGTGCCATCAGCATGAGAACTTACCGCTACTTC TTGCTGCTCTTTTGGGTGGGCCAGCCCTACCCAACTCTCTCAACTCCACTATCAAAGAGGACTAGTGGTT TCCCAGCAAAGAAAAGGGCCCTGGAGCTCTCTGGAAACAGCAAAAATGAGCTGAACCGTTCAAAAAGGAG CT GGAT GT GGAAT CAGT TCTTTCTCCT GGAGGAAT ACACAGGAT C C GAT TAT CAGT AT GT GGGCAAGT TA CAT T CAGAC CAGGAT AGAGGAGAT GGAT CAGT T AAAT AT AT C CT T T CAGGAGAT GGAGCAGGAGAT CT CT T CAT TAT T AAT GAAAACACAGGC GACAT ACAGGC CAC CAAGAGGCT GGACAGGGAAGAAAAAC C C GT T T A CAT C CT T C GAGCT CAAGCT AT AAACAGAAGGACAGGGAGAC C C GT GGAGC C C GAGT CT GAAT T CAT CAT C AAGAT C CAT GACAT CAAT GACAAT GAAC CAAT AT T CAC CAAGGAGGT T T ACACAGC CAGT GT C C CT GAAA T GT CT GAT GT C GGT ACAT T T GT T GT CCAAGT CAGT GC GAC GGAT GCAGAT GAT C CAACAT AT GGGAACAG T GCT AAAGT T GT CT ACAGT AT T CT ACAGGGACAGC C CT AT T T T T CAGT T GAAT CAGAAACAGGT AT TAT C AAGACAGCT T T GCT CAACAT GGAT C GAGAAAACAGGGAGCAGT AC CAAGT GGT GAT T CAAGC CAAGGAT A TGGGCGGC CAGAT GGGAGGAT T AT CT GGGAC CAC CAC C GT GAACAT GACACT GACT GAT GT CAAC GACAACCCTCCCCGATTCCCCCAGAGTACATACCAGTTTAAAACTCCTGAATCTTCTCCACCGGGGACACCAATT GGCAGAAT CAAAGC CAGC GAG GCT GAT GT GGGAGAAAAT GCT GAAAT T GAGT ACAGCAT CACAGAC GGT G AGGGGCT GGAT AT GT T T GAT GT CAT CAC C GAC CAGGAAAC C CAGGAAGGGAT T AT AACT GT CAAAAAGCT CTTGGACTTTGAAAAGAAGAAAGTGTATACCCTTAAAGTGGAAGCCTCCAATCCTTATGTTGAGCCACGA T T T CT CT ACT T GGGGCCTTT CAAAGAT T CAGC CAC GGT T AGAAT T GT GGT GGAGGAT GT AGAT GAGC CAC CT GT CTT CAGCAAACT GGCCTACAT CTTACAAATAAGAGAAGAT GCT CAGATAAACACCACAATAGGCT C C GT C AC AGC C C AAGAT C C AGAT GC T GC C AGGAAT C C T GT C AAGT AC T C T GT AGAT C GAC AC AC AGAT AT G GAC AGAAT AT T C AAC AT T GAT T C T GGAAAT GGT T C GAT T T T TAG AT C GAAAC TTCTTGACC GAGAAAC AC T GCT AT GGCACAACAT T ACAGT GAT AGCAACAGAGAT CAAT AAT C CAAAGCAAAGT AGT C GAGT AC CT CT ATATATTAAAGTTCTAGATGTCAATGACAACGCCCCAGAATTTGCTGAGTTCTATGAAACTTTTGTCTGT GAAAAAGCAAAGGCAGAT CAGT T GATT CAGAC C CT GCAT GCT GT T GACAAGGAT GAC C CT TAT AGT GGAC AC CAAT TTTCGTTTTCCTTGGCC C CT GAAGCAGC CAGT GGCT CAAACT T TAG CAT T CAAGACAACAAAGA CAACACGGCGGGAATCTTAACTCGGAAAAATGGCTATAATAGACACGAGATGAGCACCTATCTCTTGCCT GT GGT CAT T T CAGACAAC GAC TAG C CAGT T CAAAGCAGCACT GGGACAGT GAGT GT C C GGGT CT GT GCAT GT GAC CAC CAC GGGAACAT GCAAT C CT GC CAT GC GGAGGC GCT CAT C CAC C C CAC GGGACT GAGCAC GGG GGCTCTGGTTGCCATCCTTCTGTGCATCGTGATCCTACTAGTGACAGTGGTGCTGTTTGCAGCTCTGAGG C GGCAGC GAAAAAAAGAGC CT T T GAT CAT T T C CAAAGAGGACAT CAGAGAT AACAT T GT CAGT T ACAAC G AC GAAGGT GGT GGAGAGGAGGACAC CCAGGCT T T T GAT AT C GGCAC C CT GAGGAAT C CT GAAGC CAT AGA GGACAACAAAT TAG GAAGGGACAT T GT GC C C GAAGC C CT T T T C CT AC C C C GAC GGACT C CAACAGCT C GC GACAACAC C GAT GT CAGAGAT T T CATT AAC CAAAGGT T AAAGGAAAAT GACAC GGAC C C CACT GC C C C GC CATACGACTCCTTGGCCACTTACGCCTATGAAGGCACTGGCTCCGTGGCGGATTCCCTGAGCTCGCTGGA GTCAGTGACCACGGATGCAGATCAAGACTATGATTACCTTAGTGACTGGGGACCTCGATTCAAAAAGCTT GCAGAT AT GT AT GGAGGAGT GGACAGT GACAAAGACT C CT AAT CT GT TGCCTTTTT CAT T T T C CAAT AC G ACACTGAAATATGTGAAGTGGCTATTTCTTTATATTTATCCACTACTCCGTGAAGGCTTCTCTGTTCTAC C C GT T C CAAAAGC CAAT GGCT GCAGT C C GT GT GGAT C CAAT GT T AGAGACT T T T T T CT AGT ACACT T T TA T GAGCT T C CAAGGGGCAAAT T T T T ATT T T T T AGT GCAT C CAGT T AAC CAAGT CAGC C CAACAGGCAGGT G CCGGAGGGGAGGACAGGGAACAGTATTTCCACTTGTTCTCAGGGCAGCGTGCCCGCTTCCGCTGTCCTGG TGTTTTACTACACTCCATGTCAGGTCAGCCAACTGCCCTAACTGTACATTTCACAGGCTAATGGGATAAA GGACT GT GCT T T AAAGAT AAAAAT AT CAT CAT AGT AAAAGAAAT GAGGGCAT AT C GGCT CACAAAGAGAT AAACT ACAT AGGGGT GT T TAT T T GT GT CACAAAGAAT T T AAAAT AACACT T GC C CAT GCT AT T T GT T CT T CAAGAACTTTCTCTGCCATCAACTACTATTCAAAACCTCAAATCCACCCATATGTTAAAATTCTCATTAC T C T T AAGGAAT AGAAGC AAAT T AAAC GGT AAC AT C CAAAAGC AAC C AC AAAC CTAGTACGACTT CAT T C C TTCCACTAACTCATAGTTTGTTATATCCTAGACTAGACATGCGAAAGTTTGCCTTTGTACCATATAAAGG GGGAGGGAAAT AGCT AAT AAT GT T AAC CAAGGAAAT AT AT T T TAG CAT ACAT T T AAAGT T T T GGC CAC CA CAT GT AT CAC GGGT CACT T GAAAT T CT T T CAGCT AT CAGT AGGCT AAT GT CAAAAT T GT T T AAAAAT T CT TGAAAGAATTTTCCTGAGACAAATTTTAACTTCTTGTCTATAGTTGTCAGTATTATTCTACTATACTGTA CAT GAAAGT AGC AGT GT GAAGT AC AAT AAT T CAT AT T C T T CAT AT CCTTCTTACACGACTAAGTT GAAT T AGT AAAGT T AGAT T AAAT AAAACT T AAAT CT CACT CT AGGAGT T CAGT GGAGAGGT T AGAGC CAGC CACA CTTGAACCTAATACCCTGCCCTTGACATCTGGAAACCTCTACATATTTATATAACGTGATACATTTGGAT AAACAACAT T GAGAT TAT GAT GAAAAC CT ACAT AT T C CAT GT T T GGAAGAC C CT T GGAAGAGGAAAAT T G GAT T C C C T T AAAC AAAAGT GT T T AAGAT T GT AAT T AAAAT GAT AGT T GAT T T T CAAAAGC AT T AAT T T T T TTTCATTGTTTTTAACTTTGCTTTCATGACCATCCTGCCATCCTTGACTTTGAACTAATGATAAAGTAAT GATCTCAAACTATGACAGAAAAGTAATGTAAAATCCATCCAATCTATTATTTCTCTAATTATGCAATTAG CCTCATAGTTATTATCCAGAGGACCCAACTGAACTGAACTAATCCTTCTGGCAGATTCAAATCGTTTATT TCACACGCTGTTCTAATGGCACTTATCATTAGAATCTTACCTTGTGCAGTCATCAGAAATTCCAGCGTAC TATAATGAAAACATCCTTGTTTTGAAAACCTAAAAGACAGGCTCTGTATATATATATACTTAAGAATATG CTGACTTCACTTATTAGTCTTAGGGATTTATTTTCAATTAATATTAATTTTCTACAAATAATTTTAGTGT CAT T T C CAT T T GGGGAT AT T GT CAT AT CAGCACAT AT TTTCTGTTT GGAAACACACT GT T GT T T AGT T AA GT T T T AAAT AGGT GT AT TAG C CAAGAAGT AAAGAT GGAAAC GT T AAAAGAAGAGAAAT GT AGT AT T T T GG GTTACCTGATTAGAGTGAAAATTTTTTACAATCATATTATTCCTTGTGTCTTCTGAATGGTTTCCGATTT TATAATGGACTGCCCTATATAGTAACAAGTATTTCATGCTTGAGCTATTTCCTGCTTTCAGGGTTTCTTT TTTCTAGTTCTT CAT AC AC AC AC AT AC AC AC AC AC AC AC AC AC AC AC AC AC AC AC AC AC GAAT GC AAAC A AAAGGCTATAT GAGGT CTT CACT CTAAT GAATT GATAT GT AT CAT AGT CACAGGTAAGT GTT GAAAAAAG CTTAGTAAAGTTAGAAGCTACTTACTCATAGCAATAGAACAGCACCTTAATCACACGATTTACTGTAAAA T T AAAGAGGT CTCTATCTGTATGTTT CAT GT CAC GT AAC AAAT T GAAT C AAGGAAGAT AGT C C T GT AAAA AGAAAGGT AT CAT CT GAAGT T GAGGAT T GACAC TAGCAGT T T C CAAT GT T TAAAGGT AAGAT CT GAGT T C T C CT AAT AAGT AAAAGT AAGT AGT T CT AT AGCAGAAT AT CT GAGAT GT AAT T GGCAAGGT AT T T TAT C C C TCCCTGCAGATGACACAGCATACCAAGAACAGGTTAATATGATTACTTATGGAAATAACTTTAATCTCTT AT CAT AAAAGCT GAT GAT GAAGT AAAT T T AT AGGAAAT T GGAT AAT T T GAGACT GGGGCT AAAT AT T TAGT AC C AGGGT AC T GT AAGT AT C AAGT TGGAGTGACGTTTTCC T AT AAT TCAGACTCTTT GAC AT C GT GGAA CCAATAAGAGTCATAGTTCCATCATTCTCCAGCTTCGTCTCACTTCCTTCCCACCCCACCTGAGTATCAG GT CAAACAT CAT T GCAT GC GCAGGT TT T T T T T T T AAT T GCT AGGT C C CAGGCAACAT GAAAGAT TAT T GG AGAAAAAAAT AAT T T T CAGC C CAGT TT T T T CAT T GT CT GT T T C CT AAT T T T AGAT GT T GGT GAT GGGAAA GAT GGAAGGAGAGT GGGAAGAAGT AAAAT T T T AAT AT T T GT T T CAAT CACT T T GAAACT AAAAT T CAT TA AGCAT AAC CAGAT TGCTTTTGTGGGTTGTTT CAAGGACAT T GAGAGCT T T CT GAT GAT AT GTTTTTGCCC TCTATTCAAAAGCAAGAGTTCCTTTAAACTACTAAGATATTCCCTAGAATAAGCTGAATTTAAAAAAACA TTAAGCCATTGTTTAAAGCCCCTTCACTTCCTGGCCACTTACTTCTGAAAGGCCTAAAAAACATTTGTGC C CAAAT AAGT AAAT AAAC CAAAT GGGAAAGAAGCAAAGAT TAT T C CAT AGAAC CACAAGAGAGGGAAT GT GGGCACAGT AAAT AGAT GT T T CT T T CAGAACT T TCCTGCCTT T ACAGT T T GT GT C CAT AAAGGGAT GT T C AGCAAT GAAAT T ACT C C CT T T T CAGAT GGAACAAAAC CT GC C CAT T T AAT T T T AAC GCAGT AT AAAAAAC GTGTGGTTTAGTTTTTATTTTCAGCTCCCAAAGAGTTGTGCAGAAAATCTTAAAATTTTTTTTTTTTTTTTTTTTTTTTTTTT GAGACAGAAT CT CGCT CT GT CGCCCAGGCT GGAAT GCAGT GGCGCGAT CT CT GCT CACTACAAGCTCCGCCTCCCGGGTTCACGCCATTCTCCTGCCTCAGCGCCCCCAGTAGCTGGGACTACAGGC ACACAC CAC CAC GC C C GGCT AAT T T TT GT T GT AT T T T T T AGT GGAGACAGGGT T T CAC CAT GT T AGC CAG GATGATCTCCATCTCCTGACCTCGTGGTCCGCCCGCCTCGGCCTCCCAAAGTGCTGGGATTACAGGCGTG AGC CAC C GC GC C C GGC CT T AAAAAT T GAAT CT GT AGCT T AGGC CAT C CAAAT T T TAT AAAT C CAAAT T AA CTTTAGAATGTTTCTATTACTTTCACTTTTACATATATAAATTTTAAGTGTCCTGATTGGCTGAACAATA TCTCACATCAAATGCTTTGCCTGGAAATAGATATCCCACTGGGGATAGTGGTGTGTAAACTATGACTTGG ACAAT T CT AT AT ACT CAAGCAC CAT AAAAAGT AT GCAGT T GAAAAGAAAAT CAAAGT T GAT TCCTGGGTG C C AAC T AAAT AT T CAAAT C AGGT AC T CAT C C T T AT C AGC T AAAT T CAT T T T C AC C AGGAAC AGAC CAC C A AATAAATTATTTTATCCTAATAACTAGTTTTGAAGCAGTGTAATTACTCTGGAAGAAGGCTCTAAAAAGT CAT GAT TCCCCCACTATTTT GAAAT GT AT C C T C T AAC AAGGAT CAT TAT AGT GT AAT C T T AAT T T T T AT G T T T TAT CAAGAT GAAAT CT T GT T T GAAT T GT GAT AT T AT AAAAGGGGACT CAAAAAT C CAAGCAGT CT AC T GT GT T T AAAT T AAC AC CAC AAC CTTCCTTAT CAGAT T AT AAGAGT AGAAAAAT TAACACTTGGTGTGTG AATCTTCAGGAAAATGAGCTATTTCATAAGCTCAAACAAGCAGCTTCCTTTTCCAGAGAATATAGAATTA TATTATGGTCTCCTTAAATGTTTAGTAGCTCTTATGGTCACAGCATTTTTAATCTCCCTATGGCATCTTT AT GGAAT AAT T T T C T AAAGGGT AAT T C T C T AC T AAAAAT AT C AGAC C C C GAC CAT AT T T AAT GT GGAGAG CAATACCCTCTTAGAAAGAAAATACATTGACTCATACACTTGTTAAAAGTTAATAAAGAAATAGCTCATT T T T AAAGC C GGAAGT T TAT GGT CT CT GCAT C GT CAAT T T AAT T T AAGCAT T GCT GAGACAAT CT T T AAT C TACTCCCCTTTTTGTAATACCTTATTTATGGTGCATTTTCATTTTTATTTGGGGGAAACGTTAGCCCAAC AGAGC C GGCAGAT GAAAGT GT T GAAAAGAGGT CAAAT GGAAACAAAGGCT CT T AC C C GCT GT AT T T CAGA CAGGACT GAGGCACT T AGC C GAGGAGC CACT GGGT TAT T AGAT T AAT T T CAAAAGAGCT T T T ACAAGT T G CTTAATTCCTTTTTTTTTTTTTTTTTTTTTCAAAAACCCATGAACCACAAACTCAAATTTCTCCTCAAAT GGGGT T AAT CT GACAAAC GAGGCAT GGAC C CAGC CT T GT GGAAAAAGCAT T C CAC GCT AAT GAGAT CT T G GTCTTTCTTGTGAGGCTACGTTATTTATGTAAATATGTCTGGAGGCACCTTCTCTAAGCTTTTAGTTTTC TAT GAT CT AT T AGT T T AGT GT T TAT TAAAGAAT CAAAT GT AT AGAAT TAG CAGGCAT T C GT GGGGAAT GC TGTGTAGCAAATGTAAAACTGACCTGCTCGGAAGAAACGTAGGAACGCTTCAAACCCACTGTAATGTTTG GTTTGAGATTATTTTCATTGCTTTGAGAGTGAACTGCCTAAGAGTAGGCCTTATAATAAATGCTATGTGC GTCTTCAGTAGTTCCAAGCTAAAGCAATTTGGCATTCTCCCACTGTGATTTGTGACTTTTAAACCCACAA AAT AAAAGCT T T T T GGT AT T GAT T GTT T T T AAT T AAAAAT ACT T C CAAGT AT AAAT T GAAAC GGAT GC CA CCCTTGAAGATTTACTGGCGGGAATGCTCACTCTTGTCGTTTTCCTCAGTATCGTTCATGTCTTTGGCAA CAAGAACAC CT GAT GAAAGCAAGCAAT GCT CAGT T C C CAT CAACAT T T CT AGT T AGGGGGAT T CT CAT AA CCCCACAGTTTACCT GAGAAAGT TTTCTGTGT T AGAAGAAT GGGGTCGAGAGTATTACCTTTTAGCTCAG TGTGGCCGGGCCTTTTGTTGCAGTCAAATGGCAAATACGCACTCCTTGAAATGGCTTCTTTTATTTGGTT TTGTTTTCTTAGACT TAT AAAT T T GAAAAGAAT GC AAT T T AAAAAGT GAT T T C T C AC AAAGAGT AAAT AT GC C T T T T GC AAAT CAAT T T T T GT AACAAGT T AT T TAT AT GAT AT TACT T AAT AAAC TGGTTTTTTTCTAA

[0136] NM 006499.5 Homo sapiens galectin 8 (LGALS8), transcript variant 1, mRNA(SEQ ID NO:11)CTAGACCTGGAGGCCTGGAACAGCCAGCCGCCCACGGACGCCAGAGCCGGGAACCCTGACGGCACTTAGC TGCTGACAAACAACCTGCTCCGTGGAGCGCCTGAAACACCAGTCTTTGGGGCCAGTGCCTCAGTTTCAAT CCAGGTAACCTTTAAATGAAACTTGCCTAAAATCTTAGGTCATACACAGAAGAGACTCCAATCGACAAGA AGCT GGAAAAGAAT GAT GTTGTCCT T AAAC AAC C TAG AGAAT AT CAT C T AT AAC C C GGT AAT CCCGTTTG TTGGCACCATTCCTGATCAGCTGGATCCTGGAACTTTGATTGTGATACGTGGGCATGTTCCTAGTGACGC AGACAGAT T C CAGGT GGAT CT GCAGAAT GGCAGCAGCAT GAAAC CT C GAGC C GAT GTGGCCTTT CAT T T CAATCCTCGTTTCAAAAGGGCCGGCTGCATTGTTTGCAATACTTTGATAAATGAAAAATGGGGACGGGAAG AGAT C AC C T AT GAC AC GC C T T T C AAAAGAGAAAAGT C T T T T GAGAT C GT GAT TATGGTGCT GAAGGAC AA ATTCCAGGTGGCTGTAAATGGAAAACATACTCTGCTCTATGGCCACAGGATCGGCCCAGAGAAAATAGAC ACTCTGGGCATTTATGGCAAAGTGAATATTCACTCAATTGGTTTTAGCTTCAGCTCGGACTTACAAAGTA CCCAAGCATCTAGTCTGGAACTGACAGAGATAAGTAGAGAAAATGTTCCAAAGTCTGGCACGCCCCAGCT T C C T AGT AAT AGAGGAGGAGAC AT T T C T AAAAT C GC AC C C AGAAC TGTCTACAC C AAGAGC AAAGAT T C G ACTGTCAATCACACTTTGACTTGCACCAAAATACCACCTATGAACTATGTGTCAAAGAGGCTGCCATTCG CT GCAAGGT T GAACAC C C C CAT GGGCC CT GGAC GAACT GT C GT C GT T AAAGGAGAAGT GAAT GCAAAT GC CAAAAGCTTTAATGTTGACCTACTAGCAGGAAAATCAAAGGATATTGCTCTACACTTGAACCCACGCCTG AAT AT T AAAGCAT T T GT AAGAAAT TCTTTTCTT CAGGAGT C CT GGGGAGAAGAAGAGAGAAAT AT TAG CT CT T T C C CAT T T AGT C CT GGGAT GT ACT T T GAGAT GAT AAT T TAT T GT GAT GT T AGAGAAT T CAAGGT TGC AGT AAAT GGC GT ACACAGC CT GGAGTACAAACACAGAT T T AAAGAGCT CAGCAGT AT T GACAC GCT GGAA ATTAATGGAGACATCCACTTACTGGAAGTAAGGAGCTGGTAGCCTACCTACACAGCTGCTACAAAAACCA AAATACAGAATGGCTTCTGTGATACTGGCCTTGCTGAAACGCATCTCACTGTCATTCTATTGTTTATATT GTTAAAATGAGCTTGTGCACCATTAGGTCCTGCTGGGTGTTCTCAGTCCTTGCCATGAAGTATGGTGGTG TCTAGCACTGAATGGGGAAACTGGGGGCAGCAACACTTATAGCCAGTTAAAGCCACTCTGCCCTCTCTCC TACTTTGGCTGACTCTTCAAGAATGCCATTCAACAAGTATTTATGGAGTACCTACTATAATACAGTAGCT AACAT GT AT T GAGCACAGAT T T T T T TT GGT AAAACT GT GAAGAGCT AGGAT AT AT ACT T GGT GAAACAAA CCAGTATGTTCCCTGTTCTCTTGAGCTTCGACTCTTCTGTGCGCTACTGCTGCGCACTGCTTTTTCTACA GGCATTACATCAACTCCTAAGGGGTCCTCTGGGATTAGTTATGCAGATATTAAATCACCCGAAGACACTA ACTTACAGAAGACACAACTCCTTCCCCAGTGATCACTGTCATAACCAGTGCTCTACCGTATCCCATCACT GAGGACTGATGTTGACTGACATCATTTTCTTTATCGTAATAAACATGTGGCTCTATTAGCTGCAAGCTTT AC CAAGT AAT T GGCAT GAGAT CT GAGCACAGAAAT T AAGGCAAAAAAC CAAAGCAAAACAAAT ACAT GGT GCTGAAATTAACTTGATGCCAAGCCCAAGGCAGCTGATTTCTGTGTATTTGAACTTAGGGCAAATCAGAG T CT ACACAGAC GC CT ACAGAAAGT T T CAGGAAGAGGCAAGAT GCAT T CAAT T T GAAAGAT AT T TAT GGGC AACAAAGT AAGGT CAGGAT T AGACT T CAGGCAT T CAT AAGGCAGGCACT AT CAGAAAGT GT AC GC CAACT AAGGGAC C CACAAAGCAGGCAGAGGTAAT GCAGAAAT CT GT T T T GT T C C CAT GAAAT CAC CAAT CAAGGC CTCCGTTCTTCTAAAGATTAGTCCATCATCATTAGCAACTGAGATCAAAGCACTCTTCCACTTTACGTGA T T AAAAT CAAAC CT GT AT CAGCAAGTT AAAT GGT T C CAT T T CT GT GAT T T T T CT AT TAT T T GAGGGGAGT T GGCAGAAGT T C CAT GT AT AT GGGAT CT T T ACAGGT CAGAT CT T GT T ACAGGAAAT T T CAAAGGT T T GGG AGT GGGGAGGGAAAAAAGCT CAGT CAGT GAGGAT CAT T T TAT CACAT T AGACT GGGGCAGAACT CT GC CA GGAT T T AGGAAT AT T T T CAGAACAGAT T T T AGAT AT TAT T T CT AT C CAT AT AT T GAAAAGAAT AC CAT T G TCAATCTTATTTTTTTAAAAGTACTCAGTGTAGAAATCGCTAGCCCTTAATTCTTTTCCAGCTTTTCATA T T AAT GT AT GCAGAGT CT CAC CAAGCT CAAAGACACT GGT T GGGGGT GGAGGGT GC CACAGGGAAAGCT G TAGAAGGCAAGAAGACTCGAGAATCCCCCAGAGTTATTTTTCTCCATAAAGACCATCAGAGTGCTTAACT GAGCTGTTGGAGACTGTGAGGCATTTAGGAAAAAAATAGCCCACTCACATCATTCCTTGTAAGTCTTAAG T T CAT T T T CAT T T TAG GT GGAGGAAAAAAAT T T AAAAAGCT AT T AGT AT T TAT T AAT GAAT T T TACT GAG ACATTTCTTAGAAATATGCACTTCTATACTAGCAAGCTCTGTCTCTAAAATGCAAGTTGGCCTTTTGCTT GCCACATTTCTGCATTAAACTTCTATATTAGCTTCAAAGGCTTTTAAACTCAATGCGAACATTCTACGGG ATGTTCTTAGATGCCTTTAAAAAGGGGGCAGATCTAATTTTATTTGAACCCTCACTTTCCAACTTCACCA TGACCCAGTACTAGAGATTAGGGCACTTCAAAGCATTGAAAAAAATCTACTGATACTTACTTTCTTAGAC AAGTAGTTCTTAGTTAACCACCAATGGAACTGGGTTCATTCTGAATCCTGGAGGAGCTTCCTCGTGCCAC CCAGTGTTTCTGGGCCCTCTGTGTGAGCAGCCAGGTGTGAGCTGTTTTAGAAGCAGCGTGTTGCCTTCAT CTCTCCCGTTTCC C AAAAGAAC AAAGGAT AAAGGT GACAGTCACACTCCTGGGT T AAAAAAAGC AT T C C A GAACCACTTCTCTTTATGGGCACAACAAAGAAACGAAGGCTGAAGTTCGCCTACCCAAAATGAAAAGTAG GCTTTACAGTCAAAAGTACTTCTGTTGATTGCTAAATAACTTCATTTTCTTGAAATAGAGCAACTTTGAG T GAAAT CT GCAACAT GGAT AC CAT GTAT GTAAGATACT GCT GTACAGAAGAGTTAAGGCTTACAGT GCAA AT GAGGC GT CAGCT T T GGGT GCT AAAAT T AACAAGT CT AAT AT TAT T AC CAT CAAT CAGGAAGAGAAT AA TAAATGTTTAAACAAACACAGCAGTCTGTATAAAAATACCGTGTATCATTTACTCTTTCTGCAGCTCTAT AC GAT AGGCAGGAGAGGCT T AT GT GGCAGCACAAGC CAGGT GGGGAT T T T GT AAAGAAGT GAT AAAACAT T T GT AAGT AAT C CAAGT AGGT GTAT TAAGGCAC CAAAAGT AACAT GGCAC C CAACAC C CAAAAAT AAAAA TAT GAAAT AT GAGT GT GAACT CT GAGTAGAGTAT GAAACACCACAGAAAGT CTTAGAAATAGCT CT GGAG TGGCTCTCCCAGGACAGTTTCCAGTTGCTGAATAGTCTTTTGGCACTGATGTTCTACTTCTTCACATTCA TCTAAAAAAAAAAAAAAAAAATCAAAATTAAAATCTGAGTCAGTCCGCCTGCCTCGGTTCTCATTAGTTT AAT T CT T AAT GC CT T GCACT T T C C GGCAAT CAT T CAAT CAAAAGAGT GAAAT GAAGCACAT T AACAAAGC AGGAGGCGCCACGGACCGCCTCCCTCCACACCGCTCCTTCCGCCTTCATTCCTTGCCCACAGGCTTGCAC TGGAAGCTGAATAAGAATCCCCAAAACTCAAACTTCCTAGGGATGCCACCCCTTTAGTAGCTCACACCTC C C C C C T C C AAGAGC TAAGAAACAAAGGAGAAT GTACTTTTGTAGCT T AGAT AAGC AAT GAAT C AGT AAAG GACTGATCTACTTGCTCCACCACCCCTCCCTTAATAATAACATTTACTGTTATTTCCTGGGCCTAAGACTCTATGTTCCAGAACTGTCACAGCTCCCCATGTCACACTCACTAGCTTGTGATCTTTGTCAAATAACTGAA ATCTTTTAAGCCTCTAGTTTCTTCCTTTGTAAAACAGAGATAAAATGTTGTGGTTTTTAAGTGAGATAAT CCAAGTAAAGCACCTAACATGGAGTAGTGAATGAACATCGGTTGCTACTAAAAGTGGACATCCTACTGCA T C CT T AAT GC CACT AGGCAT T T C CATACAAT CT GGGGAC CAAAACT T CAAT CAT AT AAAT GT AT GAGGT T AATTAAAAACACTACTGTAATCTGCTTGTATGATCACAAACCACCACAAAAGAAAAGATCGTGAAGATTA CACTGTAAACGGACTCTCAAATGATCAGGAGGTGGTCACTTCGCAACTTGCTCCCTCCACCCAACTCAAA ACAGGAGCTCGAGCCTGCCTGTATTTGAGACTGGAGCTGCCTGTATGAGGACTGGATCAACTGCTAGTCA C GT TAT AT C C AAAT C T GC AT TAT CAT T GGGC AC AT T T T C AC AGAAT T T T AC T GAAT T AT T C C T T AAT T GT TTAATGGTTGGGAATAGTTTGGGAATTACCTTCCATCAACTCTGCTAAGAAAGGAATGGATTCTGGTAGC AAGACAATATAATTCTCCTTTAGTTTTTCAGCCAGTGCTAACACAGTAATCAAAGCAGCAAATCGAACCT GAAAGGGAT AAAAGAGC AAAGAAAT AAAAAGT AGT GTTACTGTATTTATTATCTTAAGAGCTGTACTGAC TTGAGACAAGCTCTAACTTTTTAAACATTAGTTCACATGCGTTTATTCACTTCATTATGTTCATTAAGCT T T CAT CT T AGAAT AC CAGT T T CAC CAT T T GGGAGCT GT T T GT AAT AT GT GCAAC CT T AT AAAT AGT GT T T TCCAAACTGTGTCCCAGGACTGCAAATCTTTAATGTGAAATGTCTTTTTATAATCTCTTCCTTTAAAAAA AAC CAAT AAAAT AAAAT GC C GCAT GCAAACT CAAGT GT GT CAC CAGAT T T T ACT T CAT TGGCGCTCGCCA GCCCGCCAGGCTGGCAATAAAGTGCCTCCAGCCACCTCTGGCAGGTCTCCTCACCCACAGCCCCTGACTG GTCACCACTATAATTGTATGAGGGGCCAGGACAGTCGCTTGGGATAAACTCCCATCTCAGCACTGAATAA AAAACATTCTGTGTCACAATATCCTAGTTTTGGGGCTTTAAAAACGTCTAGGTGTTCCTCACATGCCTTG TCTATAATAAGGAAAGCAAGCAGTAGTTGGGTATTGTTAGCTTTTGAAACAAAAGCCCTACTGGTCTTCT AAT T T T GGAT AT T T T AAT T AAAGAAT AT C T GGAC AGT AC AAAGT GAAT TAT T AAAAAAC CAT T T GT AAC T ACCTAGATTCAATCAGGATTTCCTTGATTTGTGCAAAGTAAAATATTACAATAAATTTGATACTGCTACT T GT AT AAAAAC CT AT GGT T T AAAAT GT GGGGGT T CAT CAT AAT AGT CT CAT T GT TAGCAT AT C CT AAT AA AGAAT T T GAAC T AAT AAAT C C T AT T AAT AAAA

[0137] NM 001014989.2 Homo sapiens linker for activation ofT cells (LAT), transcript variant 4, mRNA (SEQ ID NO: 12)GTACCCCAGAATTCCCACTGTCAGGGCCTCCCTGCTCGCTGCCTCCCCGGGTCCTGGATATGGAGGCCAC GGCTGCCAGCTGGCAGGTGGCTGTCCCCGTCTTGGGGGGGGCCAGCAGACCCTTGGGGCCTAGGGGTGCA GCCAGCCTGCTCCGAGCTCCCCTGCAGATGGAGGAGGCCATCCTGGTCCCCTGCGTGCTGGGGCTCCTGC TGCTGCCCATCCTGGCCATGTTGATGGCACTGTGTGTGCACTGCCACAGACTGCCAGGCTCCTACGACAG CACAT C CT CAGAT AGT T T GT AT C CAAGGGGCAT C CAGT T CAAAC GGC CT CACAC GGTTGCCCCCTGGCCA CCTGCCTACCCACCTGTCACCTCCTACCCACCCCTGAGCCAGCCAGACCTGCTCCCCATCCCAAGATCCC C GCAGC CCCTTGGGGGCTCC CAC C GGAC GC CAT CT T C C C GGC GGGAT T CT GAT GGT GC CAACAGT GT GGC GAGCT AC GAGAAC GAGGAAC CAGC CT GT GAGGAT GC GGAT GAGGAT GAGGAC GACT AT CACAAC C CAGGC TACCTGGTGGTGCTTCCTGACAGCACCCCGGCCACTAGCACTGCTGCCCCATCAGCTCCTGCACTCAGCA C C C CT GGCAT C C GAGACAGT GCCTTCTC CAT GGAGT C CAT T GAT GAT TAG GT GAAC GT T C C GGAGAGC GG GGAGAGC GCAGAAGC GT CT CT GGAT GGCAGC C GGGAGT AT GT GAAT GT GT C C CAGGAACT GCAT C CT GGA GCGGCTAAGACTGAGCCTGCCGCCCTGAGTTCCCAGGAGGCAGAGGAAGTGGAGGAAGAGGGGGCTCCAG ATTACGAGAATCTGCAGGAGCTGAACTGAGGGCCTGTGGAGGCCGAGTCTGTCCTGGAACCAGGCTTGCC TGGGACGGCTGAGCTGGGCAGCTGGAAGTGGCTCTGGGGTCCTCACATGGCGTCCTGCCCTTGCTCCAGC CTGACAACAGCCTGAGAAATCCCCCCGTAACTTATTATCACTTTGGGGTTCGGCCTGTGTCCCCCGAACG CTCTGCACCTTCTGACGCAGCCTGAGAATGACCTGCCCTGGCCCCAGCCCTACTCTGTGTAATAGAATAA AGGCCTGCGTGTGTCTGTGTTGAGCGTGCGTCTGTGTGTGCCTGTGTGCGAGTCTGAGTCAGAGATTTGG AGATGTCTCTGTGTGTTTGTGTGTATCTGTGGGTCTCCATCCTCCATGGGGGCTCAGCCAGGTGCTGTGA CACCCCCCTTCTGAATGAAGCCTTCTGACCTGGGCTGGCACTGCTGGGGGTGAGGACACATTGCCCCATG AGACAGT C C CAGAACAC GGGAGCT GCT GGCT GT GACAAT GGT T T CAC CAT C CT T AGAC CAAGGGAT GGGA CCTGATGACCTGGGAGGACTCTCTTAGTTCTTACCTTTTGTGGTTCTCAATAAAACAGAACTTAAAAAAT TG

[0138] NM 014387.4 Homo sapiens linker for activation of T cells (LAT), transcript variant 1, mRNA (SEQ ID NO: 13)ACAGCT T C CT GC C GCAGGC GGGC GGGAGGGC GGGGAC GGAGAGGC GGGC GC C GAGGAGGGGCAGGT AGGGCTGGGACGCAGGGGTAACTGGATCCCCCGACTTCAGCCCAGGCCCTGGTCTGACCACCCTGGGAGCAGGG ACT T T C CACAGT CAGCT GGAC GCACACT CAGC C CAGT AAAAGAGGGGAC C CAT C C C GGGAGC C C C GGGGA GGGCACAGCTGCCTCCTCCCGGGCTCCCCTGCCACCTGGTGCCTACCTGCCCCCTGCTCCCTGCCGGGTC CGGTCCTCACCCCATCTTCATCTGGCCTTGACTCTGCCCTTGAGGGGCCTAGGGGTGCAGCCAGCCTGCT CCGAGCTCCCCTGCAGATGGAGGAGGCCATCCTGGTCCCCTGCGTGCTGGGGCTCCTGCTGCTGCCCATC CTGGCCATGTTGATGGCACTGTGTGTGCACTGCCACAGACTGCCAGGCTCCTACGACAGCACATCCTCAG AT AGT T T GT AT C CAAGGGGCAT C CAGT T CAAAC GGC CT CACAC GGTTGCCCCCTGGC CAC CT GC CT AC C C ACCTGTCACCTCCTACCCACCCCTGAGCCAGCCAGACCTGCTCCCCATCCCAAGATCCCCGCAGCCCCTT GGGGGCTCCCACCGGACGCCATCTTCCCGGCGGGATTCTGATGGTGCCAACAGTGTGGCGAGCTACGAGA ACGAGGGTGCGTCTGGGATCCGAGGTGCCCAGGCTGGGTGGGGAGTCTGGGGTCCGTCCTGGACTAGGCT GAC C C CT GT GT C GT TAG C C C CAGAACCAGC CT GT GAGGAT GC GGAT GAGGAT GAGGAC GACT AT CACAAC CCAGGCTACCTGGTGGTGCTTCCTGACAGCACCCCGGCCACTAGCACTGCTGCCCCATCAGCTCCTGCAC T CAGCAC C C CT GGCAT C C GAGACAGT GCCTTCTC CAT GGAGT C CAT T GAT GAT TAG GT GAAC GT T C C GGA GAGCGGGGAGAGCGCAGAAGCGTCTCTGGATGGCAGCCGGGAGTATGTGAATGTGTCCCAGGAACTGCAT CCTGGAGCGGCTAAGACTGAGCCTGCCGCCCTGAGTTCCCAGGAGGCAGAGGAAGTGGAGGAAGAGGGGG CTCCAGATTACGAGAATCTGCAGGAGCTGAACTGAGGGCCTGTGGAGGCCGAGTCTGTCCTGGAACCAGG CTTGCCTGGGACGGCTGAGCTGGGCAGCTGGAAGTGGCTCTGGGGTCCTCACATGGCGTCCTGCCCTTGC TCCAGCCTGACAACAGCCTGAGAAATCCCCCCGTAACTTATTATCACTTTGGGGTTCGGCCTGTGTCCCC CGAACGCTCTGCACCTTCTGACGCAGCCTGAGAATGACCTGCCCTGGCCCCAGCCCTACTCTGTGTAATA GAATAAAGGCCTGCGTGTGTCTGTGTTGAGCGTGCGTCTGTGTGTGCCTGTGTGCGAGTCTGAGTCAGAG ATTTGGAGATGTCTCTGTGTGTTTGTGTGTATCTGTGGGTCTCCATCCTCCATGGGGGCTCAGCCAGGTG CTGTGACACCCCCCTTCTGAATGAAGCCTTCTGACCTGGGCTGGCACTGCTGGGGGTGAGGACACATTGC C C CAT GAGACAGT C C CAGAACAC GGCAGCT GCT GGCT GT GACAAT GGT T T CAC CAT C CT T AGAC CAAGGG ATGGGACCTGATGACCTGGGAGGACTCTCTTAGTTCTTACCTTTTGTGGTTCTCAATAAAACAGAACTTA AAAAATTG

[0139] NM 001343.4 Homo sapiens DAB adaptor protein 2 (DAB2), transcript variant 1, mRNA (SEQ ID NO: 14)AGTCCTCAGCTGCCAGCATCTGATTAGAACCATATCTCTCGCCGGGAGTGGCCGCGCGGCTCCGAAGCTC CCGGCCGGCGGCTATTTAAGCGAGGCCCGCCGCATCCGCTGCGCTGTAGCCTGGAGGCTCCGGGCGCGGG GAAGTCATGCTCGCTTCACGGAGGCAATAGCTAGCCGGTGTCTGTGGGAGGTTATGTTTATTTGAGACTT CTCCATCGGGATCGCCTGGTGTCACCAAGTGTCCACTGGTACTGAGGTTTGCTGCCTGCCTTCTTGCCAT GT CT AAC GAAGT AGAAACAAGT GCAAC CAAT GGT CAGC C C GAC CAACAGGC C GCAC CAAAAGCAC C CT CA AAGAAGGAAAAAAAGAAAGGC C CT GAAAAGACAGAT GAAT AT CT CT T AGCAAGGT T CAAAGGC GAT GGT G T AAAAT AT AAGGC CAAGCT GAT T GGCAT T GAT GAT GT GC CAGAT GCAAGAGGGGAT AAAAT GAGC CAAGA CT CTAT GAT GAAACTAAAGGGAAT GGCGGCAGCT GGT CGGT CT CAGGGACAACACAAACAAAGGAT CT GG GT CAACAT T T C C CT T T CT GGGAT AAAAAT AAT T GAT GAGAAAACT GGGGT AAT AGAGCAT GAACAT C CAG T AAAT AAGAT T T CT T T CAT T GC C C GT GAT GT GACAGACAAC C GGGCAT T T GGT TAG GT GT GT GGAGGAGA AGGC CAGCAT CAGT T T T T T GC CAT AAAAAC C GGGCAACAGGCT GAAC CAT T AGT T GT T GAT CT T AAAGAC CT T T T T CAAGT TAT CT AT AAT GT AAAGAAAAAGGAAGAAGAAAAGAAAAAGAT AGAGGAAGC CAGCAAAG CAGT T GAGAAT GGGAGT GAGGC C CT AAT GAT T CT AGAT GAC CAAACT AACAAACT GAAAT C GGGT GT T GA C CAGAT GGAT T T GT T T GGGGACAT GT CT ACAC CT C CT GAC CT AAAT AGT C CAACAGAAAGCAAAGAT AT C CTGTTAGTGGATCTAAACTCTGAAATCGACACCAATCAGAATTCTTTAAGAGAAAATCCATTCTTAACAA ACGGCATCACCTCCTGTTCTCTTCCTCGACCAACGCCTCAGGCATCCTTCTTGCCTGAAAATGCCTTTTC TGCCAATCTCAACTTCTTTCCCACCCCTAATCCTGATCCTTTCCGTGACGATCCTTTCACACAGCCAGAC CAATCGACACCTTCTTCGTTTGATTCTCTCAAATCTCCAGATCAGAAGAAAGAGAATTCGAGTAGCTCGT CTACTCCGCTGAGTAATGGGCCCCTGAATGGTGATGTTGACTACTTTGGTCAGCAATTTGACCAGATCTC T AAC C GGACT GGCAAACAGGAAGCT CAGGCAGGC C CAT GGCCCTTTT CAAGT T C GCAAAC C CAGC CAGCA GTGAGAACTCAAAATGGGGTATCTGAAAGAGAACAGAACGGCTTCTCTGTCAAATCCTCCCCGAACCCTT TTGTGGGAAGCCCTCCCAAAGGACTGTCCATACAGAATGGCGTAAAGCAGGACTTGGAAAGCTCTGTCCA GT C CT CAC CACAT GACT C CAT AGC CAT TAT C C CAC CT C CACAAAGT AC CAAAC CAGGAAGAGGCAGAAGG ACTGCTAAGTCTTCAGCCAATGACTTGCTTGCATCAGACATCTTTGCTCCTCCCGTCTCAGAACCTTCAG GCCAGGCGTCACCCACAGGACAACCTACAGCCCTGCAGCCCAACCCTCTGGATCTCTTCAAAACAAGTGC TCCTGCCCCAGTGGGGCCCCTGGTGGGTCTAGGTGGTGTAACTGTCACACTCCCTCAGGCAGGACCATGG AACACAGCATCTTTGGTCTTCAATCAGTCCCCTTCAATGGCTCCGGGAGCCATGATGGGTGGTCAACCTT CAGGT T T T AGT CAGC C C GT CAT T T T T GGT ACAAGT C CAGCT GT T T CAGGT T GGAAC CAGC CT T CAC C CT TTGCAGCCTCAACTCCCCCTCCAGTGCCTGTTGTCTGGGGCCCTTCTGCATCTGTGGCACCCAATGCTTGG TCAACAACAAGCCCTTTGGGGAATCCTTTTCAGAGCAATATTTTTCCAGCTCCTGCTGTGTCCACTCAGC CCCCATCCATGCACTCCTCTCTCCTGGTCACTCCTCCTCAGCCACCTCCCAGAGCTGGCCCTCCCAAGGA CATCTCCAGTGATGCCTTCACTGCCTTAGACCCACTTGGGGATAAAGAGATCAAGGATGTGAAAGAAATG TTTAAGGATTTCCAACTGCGGCAGCCACCTGCTGTGCCCGCGCGGAAGGGAGAGCAGACTTCTTCTGGGA CT T T GAGT GCCTTTGC CAGT TAT T T CAACAGCAAGGT T GGCAT T C CT CAGGAGAAT GCAGAC CAT GAT GA CTTTGATGCTAATCAACTATTGAACAAGATCAATGAACCACCAAAGCCAGCTCCCAGACAAGTTTCCCTG CCAGTTACCAAATCTACTGACAATGCATTTGAGAACCCTTTCTTTAAAGATTCTTTTGGTTCATCACAAG CCTCTGTGGCTTCTTCTCAACCTGTATCTTCTGAGATGTATAGGGATCCATTTGGAAATCCTTTTGCCTA AATTCTGAACTTGGTCTGCAGACCATCCAGAGGAATAAAAAGGTTGGCCTTAGTAGTCAAAAACAAAGCT GATAGCCAGACACGTTCTGATTTCTGCCCTTGTTCCAGCTTTGACGTATTATCTGTTGCCTTATTTCTCA TTGCCTCTTCTACTTGTAAAATGCTTTTCACTTTCTGTCTAGGTTAAAGCTAAACTGAATCTATGGCTTT AAATAAATTAAGATCCTAAACTCTCTAGCTTAAGTGTAAATGAAGTACAGTAGTTTCCCTACTGAACCCT GCCTCTTGTGTCCCTGGAACCTTCTAGAACACCTGCCTTCTACCCTCTGGTTGGGAGATGCAGCCACCAC ATCCCTTCATATCATACTGTTTTGAATAAATTTTCAAATCCTTATTGTTCAGAGTTGTTTGGGGGTTCTG TTTCAGAGCATAAAACCTAAAGGTTATAGTAGAACAAGGCACCTTCTTAAAAGAAATCTTGCTTCAGACC ATCAGTTACAGAGAATTTCTAAAGTAAAATTGAAGCAACTACAACTTCTCCTTAGACACTTTGGAATCTA ACCACTTAAGGACCTTTTTAAAGAGATAGCTTCTCTTCTTTCTGAAGATCAATTTCTCCCAAGGCCAAGA TTGTCCTTTTCTCCCATTTCTTGCTAGCTATTGCAAATGAGGGAAGAACATTATTCATCTCTCCTCCCCT TTTTTTTCTGATTCTTTTTTCAGTCAGTTTTGCTCCTGGGTTCAAGTAGTATTACCACCCTTTCACAAGC AACAGACT CT CACAGGGCAAAAAAAAAAAAAAAT CTAAT GATT CACAGACAGAT CT GGAGCCT CT CTT CA TTCTCAGTAATTGCTAGTCCCAAGAACTAGAATTGCAAATGGGCACAACCTATATCCTTCCTGTGGAAGA GGAGGCCACTCTCTTGAGCTGAAGTTCCAGAAGAGCAGTTAATGTTCAAGAGAAATTGAACTCAACTCAG CAACAAAGGACT CT AT T T T GAAGAGCAACAT AT CACAAAGCT AAAT GT GAT T GT GC CAAACACAT T AGGT GCTTATTTGGGGTCATGCTAGGCCTTTATCAAGTAACTGGAAAACTTTTCTTGCAGCCACAATCTCAATG T C GT T AGT AGGAAGAT AAGAGGGGAGAAAAAGCT GT AGAACAAAT GT T T GGGGT TAG CAT T GAAAAT CT A ATGTCTGCAATATTTTTCTCCTCACAACTTGGAAACGTTCCCAGTTCATTTTCAGTCCTGTTGTGAGCAC AGT T CT GAAGGGT T T AT T AT T GT CAAAAT AAGT T T T GT T T T GT T T T GT T T AT GT T GGGT T T T T AAT GT T G TCTCTTGACCCTTAATGCTCAGGTTCTTGTGGGAGTTAATCAGCCACATCCAATGTTACCTTGAGGGGGA AGAAGAGGGT GAT GCT CAGAAGCT AAACAAGACAGGGGC CACAT GAC C CT CT AT T GAT T AGC C C CAAGT A GAAAGT C CT GT GGT T T TAT GT T T AAT GGT AAT AGT T GAT CAT AT AT GGCAT AAT T T T CT AT CAGCT T C CT ACTCAGTCACTATAAACACAGACTTGAAATAGTACTTTAAATGTCCAAATACCTAAATGTGCTAAACTGG AGGTAACTATTTCTAGGTAGTTGAATTTTTGAAAGTCATGATCAGCCACACAACTGTTTTGTACATACTT ATTTTCTCATGCACTTTTCTGTATGCAAATAAAGCTATAAATTTACTCATTTCAATAAACTGGAGTGGCA GAATA

[0140] NM 001244871.2 Homo sapiens DAB adaptor protein 2 (DAB2), transcript variant 2, mRNA (SEQ ID NO: 15)AGTCCTCAGCTGCCAGCATCTGATTAGAACCATATCTCTCGCCGGGAGTGGCCGCGCGGCTCCGAAGCTC CCGGCCGGCGGCTATTTAAGCGAGGCCCGCCGCATCCGCTGCGCTGTAGCCTGGAGGCTCCGGGCGCGGG GAAGTCATGCTCGCTTCACGGAGGCAATAGCTAGCCGGTGTCTGTGGGAGGTTATGTTTATTTGAGACTT CTCCATCGGGATCGCCTGGTGTCACCAAGTGTCCACTGGTACTGAGGTTTGCTGCCTGCCTTCTTGCCAT GT CT AAC GAAGT AGAAACAAGT GCAAC CAAT GGT CAGC C C GAC CAACAGGC C GCAC CAAAAGCAC C CT CA AAGAAGGAAAAAAAGAAAGGC C CT GAAAAGACAGAT GAAT AT CT CT T AGCAAGGT T CAAAGGC GAT GGT G T AAAAT AT AAGGC CAAGCT GAT T GGCAT T GAT GAT GT GC CAGAT GCAAGAGGGGAT AAAAT GAGC CAAGA CT CTAT GAT GAAACTAAAGGGAAT GGCGGCAGCT GGT CGGT CT CAGGGACAACACAAACAAAGGAT CT GG GT CAACAT T T C C CT T T CT GGGAT AAAAAT AAT T GAT GAGAAAACT GGGGT AAT AGAGCAT GAACAT C CAG T AAAT AAGAT T T CT T T CAT T GC C C GT GAT GT GACAGACAAC C GGGCAT T T GGT TAG GT GT GT GGAGGAGA AGGC CAGCAT CAGT T T T T T GC CAT AAAAAC C GGGCAACAGGCT GAAC CAT T AGT T GT T GAT CT T AAAGAC CT T T T T CAAGT TAT CT AT AAT GT AAAGAAAAAGGAAGAAGAAAAGAAAAAGAT AGAGGAAGC CAGCAAAG CAGT T GAGAAT GGGAGT GAGGC C CT AAT GAT T CT AGAT GAC CAAACT AACAAACT GAAAT C GGAAAGCAA AGATATCCTGTTAGTGGATCTAAACTCTGAAATCGACACCAATCAGAATTCTTTAAGAGAAAATCCATTC TTAACAAACGGCATCACCTCCTGTTCTCTTCCTCGACCAACGCCTCAGGCATCCTTCTTGCCTGAAAATG CCTTTTCTGCCAATCTCAACTTCTTTCCCACCCCTAATCCTGATCCTTTCCGTGACGATCCTTTCACACA GCCAGACCAATCGACACCTTCTTCGTTTGATTCTCTCAAATCTCCAGATCAGAAGAAAGAGAATTCGAGT AGCTCGTCTACTCCGCTGAGTAATGGGCCCCTGAATGGTGATGTTGACTACTTTGGTCAGCAATTTGACCAGATCTCTAACCGGACTGGCAAACAGGAAGCTCAGGCAGGCCCATGGCCCTTTTCAAGTTCGCAAACCCA GCCAGCAGTGAGAACTCAAAATGGGGTATCTGAAAGAGAACAGAACGGCTTCTCTGTCAAATCCTCCCCG AACCCTTTTGTGGGAAGCCCTCCCAAAGGACTGTCCATACAGAATGGCGTAAAGCAGGACTTGGAAAGCT CTGTCCAGTCCTCAC GAG AT GAG T C CAT AGC CAT TATCCCACCTC C AC AAAGT AC C AAAC C AGGAAGAGG CAGAAGGACTGCTAAGTCTTCAGCCAATGACTTGCTTGCATCAGACATCTTTGCTCCTCCCGTCTCAGAA CCTTCAGGCCAGGCGTCACCCACAGGACAACCTACAGCCCTGCAGCCCAACCCTCTGGATCTCTTCAAAA CAAGTGCTCCTGCCCCAGTGGGGCCCCTGGTGGGTCTAGGTGGTGTAACTGTCACACTCCCTCAGGCAGG ACCATGGAACACAGCATCTTTGGTCTTCAATCAGTCCCCTTCAATGGCTCCGGGAGCCATGATGGGTGGT CAAC CT T CAGGT T T T AGT CAGC C C GT CAT T T T T GGT ACAAGT C CAGCT GT T T CAGGT T GGAAC CAGC CT T CACCCTTTGCAGCCTCAACTCCCCCTCCAGTGCCTGTTGTCTGGGGCCCTTCTGCATCTGTGGCACCCAA TGCTTGGTCAACAACAAGCCCTTTGGGGAATCCTTTTCAGAGCAATATTTTTCCAGCTCCTGCTGTGTCC ACTCAGCCCCCATCCATGCACTCCTCTCTCCTGGTCACTCCTCCTCAGCCACCTCCCAGAGCTGGCCCTC CCAAGGACATCTCCAGTGATGCCTTCACTGCCTTAGACCCACTTGGGGATAAAGAGATCAAGGATGTGAA AGAAATGTTTAAGGATTTCCAACTGCGGCAGCCACCTGCTGTGCCCGCGCGGAAGGGAGAGCAGACTTCT T CT GGGACT T T GAGT GCCTTTGC CAGT TAT T T CAACAGCAAGGT T GGCAT T C CT CAGGAGAAT GCAGAC C AT GAT GAG T T T GAT GC T AAT CAAC T AT T GAAC AAGAT C AAT GAAC GAG C AAAGC C AGC T C C C AGAC AAGT TTCCCTGCCAGTTACCAAATCTACTGACAATGCATTTGAGAACCCTTTCTTTAAAGATTCTTTTGGTTCA TCACAAGCCTCTGTGGCTTCTTCTCAACCTGTATCTTCTGAGATGTATAGGGATCCATTTGGAAATCCTT TTGCCTAAATTCTGAACTTGGTCTGCAGACCATCCAGAGGAATAAAAAGGTTGGCCTTAGTAGTCAAAAA CAAAGCTGATAGCCAGACACGTTCTGATTTCTGCCCTTGTTCCAGCTTTGACGTATTATCTGTTGCCTTA TTTCTCATTGCCTCTTCTACTTGTAAAATGCTTTTCACTTTCTGTCTAGGTTAAAGCTAAACTGAATCTA TGGCTTTAAATAAATTAAGATCCTAAACTCTCTAGCTTAAGTGTAAATGAAGTACAGTAGTTTCCCTACT GAACCCTGCCTCTTGTGTCCCTGGAACCTTCTAGAACACCTGCCTTCTACCCTCTGGTTGGGAGATGCAG C C AC C AC AT C C C T T C AT AT C AT AC T GT T T T GAAT AAAT T T T C AAAT C C T T AT T GT T C AGAGT T GT T T GGG GGTTCTGTTTCAGAGCATAAAACCTAAAGGTTATAGTAGAACAAGGCACCTTCTTAAAAGAAATCTTGCT TCAGACCATCAGTTACAGAGAATTTCTAAAGTAAAATTGAAGCAACTACAACTTCTCCTTAGACACTTTG GAATCTAACCACTTAAGGACCTTTTTAAAGAGATAGCTTCTCTTCTTTCTGAAGATCAATTTCTCCCAAG GCCAAGATTGTCCTTTTCTCCCATTTCTTGCTAGCTATTGCAAATGAGGGAAGAACATTATTCATCTCTC CTCCCCTTTTTTTTCTGATTCTTTTTTCAGTCAGTTTTGCTCCTGGGTTCAAGTAGTATTACCACCCTTT CACAAGCAACAGACT CT CACAGGGCAAAAAAAAAAAAAAAT CTAAT GATT CACAGACAGAT CT GGAGCCT CTCTTCATTCTCAGTAATTGCTAGTCCCAAGAACTAGAATTGCAAATGGGCACAACCTATATCCTTCCTG T GGAAGAGGAGGC CACT CT CT T GAGCT GAAGT T C CAGAAGAGCAGT T AAT GT T CAAGAGAAAT T GAACT C AACT CAGCAACAAAGGACT CT AT T T T GAAGAGCAACAT AT CACAAAGCT AAAT GT GAT T GT GC CAAACAC ATTAGGTGCTTATTTGGGGTCATGCTAGGCCTTTATCAAGTAACTGGAAAACTTTTCTTGCAGCCACAAT CT CAAT GT C GT T AGT AGGAAGAT AAGAGGGGAGAAAAAGCT GT AGAACAAAT GT T T GGGGT TAG CAT T GA AAATCTAATGTCTGCAATATTTTTCTCCTCACAACTTGGAAACGTTCCCAGTTCATTTTCAGTCCTGTTG T GAGCACAGT T CT GAAGGGT T TAT TAT T GT CAAAAT AAGT TTTGTTTTGTTTTGTT TAT GTTGGGTTTTT AATGTTGTCTCTTGACCCTTAATGCTCAGGTTCTTGTGGGAGTTAATCAGCCACATCCAATGTTACCTTG AGGGGGAAGAAGAGGGT GAT GCT CAGAAGCT AAACAAGACAGGGGC CACAT GAC C CT CT AT T GAT T AGC C C CAAGT AGAAAGT C CT GT GGT T T TAT GT T T AAT GGT AAT AGT T GAT CAT AT AT GGCAT AAT T T T CT AT CA GCTTCCTACTCAGTCACTATAAACACAGACTTGAAATAGTACTTTAAATGTCCAAATACCTAAATGTGCT AAAC TGGAGGTAACTATTTCTAGGTAGTT GAAT T T T T GAAAGT CAT GAT CAGC C AC AC AAC TGTTTTGTA CATACTTATTTTCTCATGCACTTTTCTGTATGCAAATAAAGCTATAAATTTACTCATTTCAATAAACTGG AGTGGCAGAATA

[0141] In some embodiments, detection can occur through any of a variety of mobility dependent analytical techniques based on the differential rates of migration between different nucleic acid sequences. Exemplary mobility-dependent analysis techniques include electrophoresis, chromatography, mass spectroscopy, sedimentation, gradient centrifugation, field-flow fractionation, multi-stage extraction techniques, and the like. In some embodiments, mobility probes can be hybridized to amplification products, and the identity of the target nucleic acid sequence determined via a mobility dependent analysistechnique of the eluted mobility probes, as described in Published PCT Applications WO04 / 46344 and WOO 1 / 92579. In some embodiments, detection can be achieved by various microarrays and related software such as the Applied Biosystems Array System with the Applied Biosystems 1700 Chemiluminescent Microarray Analyzer and other commercially available array systems available from Affymetrix, Agilent, Illumina, and Amersham Biosciences, among others (see also Gerry et al., J. Mol. Biol. 292:251-62, 1999; De Bellis et al., Minerva Biotec 14:247-52, 2002; and Stears et al., Nat. Med. 9: 14045, including supplements, 2003).

[0142] It is also understood that detection can comprise reporter groups that are incorporated into the reaction products, either as part of labeled primers or due to the incorporation of labeled dNTPs during an amplification, or attached to reaction products, for example but not limited to, via hybridization tag complements comprising reporter groups or via linker arms that are integral or attached to reaction products. In some embodiments, unlabeled reaction products may be detected using mass spectrometry.Methods for Predicting the Risk of VTE Using Proteomic Biomarkers

[0143] In one aspect, the present disclosure provides a method for preventing cancer associated thromboembolism (CAT) in a cancer patient in need thereof comprising (a) (i) detecting mRNA and / or polypeptide expression levels of at least one of FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB in a biological sample obtained from the cancer patient that is decreased relative to a control sample obtained from a healthy subject or a predetermined threshold; and / or (ii) detecting mRNA and / or polypeptide expression levels of CNDP1 in a biological sample obtained from the cancer patient that is increased relative to a control sample obtained from a healthy subject or a predetermined threshold; and (b) administering to the cancer patient an effective amount of anticoagulant therapy. In another aspect, the present disclosure provides a method for preventing cancer associated thromboembolism (CAT) in a cancer patient in need thereof comprising administering to the cancer patient an effective amount of anticoagulant therapy, wherein a biological sample obtained from the cancer patient comprises (a) mRNA and / or polypeptide expression levels of one or more plasma proteins that are decreased relative to a control sample obtained from a healthy subject or a predetermined threshold, wherein the one or more plasma proteins are selected from the group consisting of FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6,LGALS8, LAT, and DAB; and / or (b) CNDP1 mRNA and / or polypeptide expression levels that are increased relative to a control sample obtained from a healthy subject or a predetermined threshold. In any of the preceding embodiments of the methods disclosed herein, the biological sample is whole blood, serum or plasma.

[0144] Additionally or alternatively, in some embodiments of the methods disclosed herein, the cancer patient is diagnosed with or suffers from a cancer selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non- melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non- Hodgkin lymphoma. The cancer may be a Stage 1, Stage 2, Stage 3, or Stage 4 cancer. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score < 2 or > 2 and / or has one or more organ sites of metastasis. In some embodiments, the one or more organ sites of metastasis comprise lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0145] Additionally or alternatively, in some embodiments of the methods disclosed herein, the anticoagulant therapy comprises one or more of apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, or enoxaparin.

[0146] In any of the preceding embodiments of the methods disclosed herein, mRNA expression levels are detected via real-time quantitative PCR (qPCR), digital PCR (dPCR), Reverse transcriptase-PCR (RT-PCR), Northern blotting, microarray, dot or slot blots, in situ hybridization, or fluorescent in situ hybridization (FISH). Additionally or alternatively, in some embodiments, polypeptide expression levels are detected via Western blotting,enzyme-linked immunosorbent assays (ELISA), dot blotting, immunohistochemistry, immunofluorescence, immunoprecipitation, immunoelectrophoresis, or mass-spectrometry.

[0147] In any of the foregoing embodiments of the methods disclosed herein, the cancer patient is chemotherapy-naive or has received / is receiving systemic chemotherapy. Systemic chemotherapy may comprise one or more of alkylating agents, antibiotics, antimetabolites, antimitotics, cyclin-dependent kinase inhibitors, epidermal growth factor receptor inhibitors, multikinase inhibitors, PARP inhibitors, platinum-based agents, selective estrogen receptor modulators (SERM), or VEGF inhibitors. Examples of chemotherapeutic agents include, but are not limited to, alkylating agents, platinum agents, taxanes, vinca agents, anti-estrogen drugs, aromatase inhibitors, ovarian suppression agents, VEGF / VEGFR inhibitors, EGFZEGFR inhibitors, PARP inhibitors, cytostatic alkaloids, cytotoxic antibiotics, antimetabolites, endocrine / hormonal agents, bisphosphonate therapy agents and targeted biological therapy agents (e.g., therapeutic peptides described in US 6306832, WO 2012007137, WO 2005000889, WO 2010096603 etc.). In some embodiments, the at least one additional therapeutic agent is a chemotherapeutic agent. Specific chemotherapeutic agents include, but are not limited to, cyclophosphamide, fluorouracil (or 5 -fluorouracil or 5-FU), methotrexate, edatrexate (10-ethyl-10-deaza- aminopterin), thiotepa, carboplatin, cisplatin, taxanes, paclitaxel, protein-bound paclitaxel, docetaxel, vinorelbine, tamoxifen, raloxifene, toremifene, fulvestrant, gemcitabine, irinotecan, ixabepilone, temozolmide, topotecan, vincristine, vinblastine, eribulin, mutamycin, capecitabine, anastrozole, exemestane, letrozole, leuprolide, abarelix, buserlin, goserelin, megestrol acetate, risedronate, pamidronate, ibandronate, alendronate, denosumab, zoledronate, trastuzumab, tykerb, anthracyclines (e.g., daunorubicin and doxorubicin), bevacizumab, oxaliplatin, melphalan, etoposide, mechlorethamine, bleomycin, microtubule poisons, annonaceous acetogenins, or combinations thereof.

[0148] Additionally or alternatively, in some embodiments of the methods disclosed herein, the cancer patient is immunotherapy-naive or has received / is receiving immunotherapy. Examples of immunotherapy include, but are not limited to, anti-PD-1 antibody, anti-PD-Ll antibody, anti-PD-L2 antibody, anti-CTLA-4 antibody, anti-TIM3 antibody, anti-4-lBB antibody, anti-CD73 antibody, anti-GITR antibody, and anti-LAG-3 antibody.

[0149] Additionally or alternatively, in certain embodiments of the methods disclosed herein, the cancer patient is radiotherapy-naive or has received / is receiving radiotherapy. The radiotherapy may comprise external radiotherapy, radiotherapy implants (brachytherapy), pre-targeted radioimmunotherapy, radiotherapy injections, radioisotope therapy, or intrabeam radiotherapy.

[0150] In any and all embodiments of the methods disclosed herein, the CAT is pulmonary embolism or lower extremity deep vein thrombosis (DVT). In some embodiments, lower extremity DVT includes thrombi involving a common iliac vein, an external iliac vein, a common femoral vein, a superficial femoral vein, a deep femoral vein, a popliteal vein, a peroneal vein, an anterior tibial vein, a posterior tibial vein, or a deep calf vein.Systems, Devices, and Methods for Predicting the Risk of VTE Across Multiple Cancer Types

[0151] Aspects of the operating environment as well as associated system components (e.g., hardware elements) in connection with various embodiments of the methods and systems described herein will now be discussed. Referring to FIG. 7A, an embodiment of a network environment is depicted. In brief overview, the network environment includes one or more clients 102a-102n (also generally referred to as local machine(s) 102, client(s) 102, client node(s) 102, client machine(s) 102, client computer(s) 102, client device(s) 102, endpoint(s) 102, or endpoint node(s) 102) in communication with one or more servers 106a- 106n (also generally referred to as server(s) 106, node 106, or remote machine(s) 106) via one or more networks 104. In some embodiments, a client 102 has the capacity to function as both a client node seeking access to resources provided by a server and as a server providing access to hosted resources for other clients 102a-102n.

[0152] Although FIG. 7A shows a network 104 between the clients 102 and the servers 106, the clients 102 and the servers 106 may be on the same network 104. In some embodiments, there are multiple networks 104 between the clients 102 and the servers 106. In one of these embodiments, a network 104’ (not shown) may be a private network and a network 104 may be a public network. In another of these embodiments, a network 104 may be a private network and a network 104’ a public network. In still another of these embodiments, networks 104 and 104’ may both be private networks.

[0153] The network 104 may be connected via wired or wireless links. Wired links may include Digital Subscriber Line (DSL), coaxial cable lines, or optical fiber lines. The wireless links may include BLUETOOTH, Wi-Fi, Worldwide Interoperability for Microwave Access (WiMAX), an infrared channel or satellite band. The wireless links may also include any cellular network standards used to communicate among mobile devices, including standards that qualify as 1G, 2G, 3G, 4G, or 5G. The network standards may qualify as one or more generation of mobile telecommunication standards by fulfilling a specification or standards such as the specifications maintained by International Telecommunication Union. The 3G standards, for example, may correspond to the International Mobile Telecommuni cations-2000 (IMT-2000) specification, and the 4G standards may correspond to the International Mobile Telecommunications Advanced (IMT- Advanced) specification. Examples of cellular network standards include AMPS, GSM, GPRS, UMTS, LTE, LTE Advanced, Mobile WiMAX, and WiMAX-Advanced. Cellular network standards may use various channel access methods e.g. FDMA, TDMA, CDMA, or SDMA. In some embodiments, different types of data may be transmitted via different links and standards. In other embodiments, the same types of data may be transmitted via different links and standards.

[0154] The network 104 may be any type and / or form of network. The geographical scope of the network 104 may vary widely and the network 104 can be a body area network (BAN), a personal area network (PAN), a local-area network (LAN), e.g. Intranet, a metropolitan area network (MAN), a wide area network (WAN), or the Internet. The topology of the network 104 may be of any form and may include, e.g., any of the following: point-to-point, bus, star, ring, mesh, or tree. The network 104 may be an overlay network which is virtual and sits on top of one or more layers of other networks 104’. The network 104 may be of any such network topology as known to those ordinarily skilled in the art capable of supporting the operations described herein. The network 104 may utilize different techniques and layers or stacks of protocols, including, e.g., the Ethernet protocol, the internet protocol suite (TCP / IP), the ATM (Asynchronous Transfer Mode) technique, the SONET (Synchronous Optical Networking) protocol, or the SDH (Synchronous Digital Hierarchy) protocol. The TCP / IP internet protocol suite may include application layer, transport layer, internet layer (including, e.g., IPv6), or the link layer. The network 104may be a type of a broadcast network, a telecommunications network, a data communication network, or a computer network.

[0155] In some embodiments, the system may include multiple, logically-grouped servers 106. In one of these embodiments, the logical group of servers may be referred to as a server farm 38 or a machine farm 38. In another of these embodiments, the servers 106 may be geographically dispersed. In other embodiments, a machine farm 38 may be administered as a single entity. In still other embodiments, the machine farm 38 includes a plurality of machine farms 38. The servers 106 within each machine farm 38 can be heterogeneous - one or more of the servers 106 or machines 106 can operate according to one type of operating system platform (e.g., WINDOWS NT, manufactured by Microsoft Corp, of Redmond, Washington), while one or more of the other servers 106 can operate on according to another type of operating system platform (e.g., Unix, Linux, or Mac OS X).

[0156] In one embodiment, servers 106 in the machine farm 38 may be stored in high- density rack systems, along with associated storage systems, and located in an enterprise data center. In this embodiment, consolidating the servers 106 in this way may improve system manageability, data security, the physical security of the system, and system performance by locating servers 106 and high performance storage systems on localized high performance networks. Centralizing the servers 106 and storage systems and coupling them with advanced system management tools allows more efficient use of server resources.

[0157] The servers 106 of each machine farm 38 do not need to be physically proximate to another server 106 in the same machine farm 38. Thus, the group of servers 106 logically grouped as a machine farm 38 may be interconnected using a wide-area network (WAN) connection or a metropolitan-area network (MAN) connection. For example, a machine farm 38 may include servers 106 physically located in different continents or different regions of a continent, country, state, city, campus, or room. Data transmission speeds between servers 106 in the machine farm 38 can be increased if the servers 106 are connected using a local-area network (LAN) connection or some form of direct connection. Additionally, a heterogeneous machine farm 38 may include one or more servers 106 operating according to a type of operating system, while one or more other servers 106 execute one or more types of hypervisors rather than operating systems. In these embodiments, hypervisors may be used to emulate virtual hardware, partition physical hardware, virtualize physical hardware, and execute virtual machines that provide access tocomputing environments, allowing multiple operating systems to run concurrently on a host computer. Native hypervisors may run directly on the host computer. Hypervisors may include VMware ESX / ESXi, manufactured by VMWare, Inc., of Palo Alto, California; the Xen hypervisor, an open source product whose development is overseen by Citrix Systems, Inc.; the HYPER-V hypervisors provided by Microsoft or others. Hosted hypervisors may run within an operating system on a second software level. Examples of hosted hypervisors may include VMware Workstation and VIRTU ALBOX.

[0158] Management of the machine farm 38 may be de-centralized. For example, one or more servers 106 may comprise components, subsystems and modules to support one or more management services for the machine farm 38. In one of these embodiments, one or more servers 106 provide functionality for management of dynamic data, including techniques for handling failover, data replication, and increasing the robustness of the machine farm 38. Each server 106 may communicate with a persistent store and, in some embodiments, with a dynamic store.

[0159] Server 106 may be a file server, application server, web server, proxy server, appliance, network appliance, gateway, gateway server, virtualization server, deployment server, SSL VPN server, or firewall. In one embodiment, the server 106 may be referred to as a remote machine or a node. In another embodiment, a plurality of nodes 290 may be in the path between any two communicating servers.

[0160] Referring to FIG. 7B, a cloud computing environment is depicted. A cloud computing environment may provide client 102 with one or more resources provided by a network environment. The cloud computing environment may include one or more clients 102a-102n, in communication with the cloud 108 over one or more networks 104. Clients 102 may include, e.g., thick clients, thin clients, and zero clients. A thick client may provide at least some functionality even when disconnected from the cloud 108 or servers 106. A thin client or a zero client may depend on the connection to the cloud 108 or server 106 to provide functionality. A zero client may depend on the cloud 108 or other networks 104 or servers 106 to retrieve operating system data for the client device. The cloud 108 may include back end platforms, e.g., servers 106, storage, server farms or data centers.

[0161] The cloud 108 may be public, private, or hybrid. Public clouds may include public servers 106 that are maintained by third parties to the clients 102 or the owners of theclients. The servers 106 may be located off-site in remote geographical locations as disclosed above or otherwise. Public clouds may be connected to the servers 106 over a public network. Private clouds may include private servers 106 that are physically maintained by clients 102 or owners of clients. Private clouds may be connected to the servers 106 over a private network 104. Hybrid clouds 108 may include both the private and public networks 104 and servers 106.

[0162] The cloud 108 may also include a cloud based delivery, e.g. Software as a Service (SaaS) 110, Platform as a Service (PaaS) 112, and Infrastructure as a Service (laaS) 114. laaS may refer to a user renting the use of infrastructure resources that are needed during a specified time period. laaS providers may offer storage, networking, servers or virtualization resources from large pools, allowing the users to quickly scale up by accessing more resources as needed. Examples of laaS can include infrastructure and services (e.g., EG-32) provided by OVH HOSTING of Montreal, Quebec, Canada, AMAZON WEB SERVICES provided by Amazon.com, Inc., of Seattle, Washington, RACKSPACE CLOUD provided by Rackspace US, Inc., of San Antonio, Texas, Google Compute Engine provided by Google Inc. of Mountain View, California, or RIGHTSCALE provided by RightScale, Inc., of Santa Barbara, California. PaaS providers may offer functionality provided by laaS, including, e.g., storage, networking, servers or virtualization, as well as additional resources such as, e.g., the operating system, middleware, or runtime resources. Examples of PaaS include WINDOWS AZURE provided by Microsoft Corporation of Redmond, Washington, Google App Engine provided by Google Inc., and HEROKU provided by Heroku, Inc. of San Francisco, California. SaaS providers may offer the resources that PaaS provides, including storage, networking, servers, virtualization, operating system, middleware, or runtime resources. In some embodiments, SaaS providers may offer additional resources including, e.g., data and application resources. Examples of SaaS include GOOGLE APPS provided by Google Inc., SALESFORCE provided by Salesforce.com Inc. of San Francisco, California, or OFFICE 365 provided by Microsoft Corporation. Examples of SaaS may also include data storage providers, e.g. DROPBOX provided by Dropbox, Inc. of San Francisco, California, Microsoft SKYDRIVE provided by Microsoft Corporation, Google Drive provided by Google Inc., or Apple ICLOUD provided by Apple Inc. of Cupertino, California.

[0163] Clients 102 may access laaS resources with one or more laaS standards, including, e.g., Amazon Elastic Compute Cloud (EC2), Open Cloud Computing Interface (OCCI), Cloud Infrastructure Management Interface (CIMI), or OpenStack standards. Some laaS standards may allow clients access to resources over HTTP, and may use Representational State Transfer (REST) protocol or Simple Object Access Protocol (SOAP). Clients 102 may access PaaS resources with different PaaS interfaces. Some PaaS interfaces use HTTP packages, standard Java APIs, JavaMail API, Java Data Objects (JDO), Java Persistence API (JPA), Python APIs, web integration APIs for different programming languages including, e.g., Rack for Ruby, WSGI for Python, or PSGI for Perl, or other APIs that may be built on REST, HTTP, XML, or other protocols. Clients 102 may access SaaS resources through the use of web-based user interfaces, provided by a web browser (e.g. GOOGLE CHROME, Microsoft INTERNET EXPLORER, or Mozilla Firefox provided by Mozilla Foundation of Mountain View, California). Clients 102 may also access SaaS resources through smartphone or tablet applications, including, e.g, Salesforce Sales Cloud, or Google Drive app. Clients 102 may also access SaaS resources through the client operating system, including, e.g, Windows file system for DROPBOX.

[0164] In some embodiments, access to laaS, PaaS, or SaaS resources may be authenticated. For example, a server or authentication server may authenticate a user via security certificates, HTTPS, or API keys. API keys may include various encryption standards such as, e.g., Advanced Encryption Standard (AES). Data resources may be sent over Transport Layer Security (TLS) or Secure Sockets Layer (SSL).

[0165] The client 102 and server 106 may be deployed as and / or executed on any type and form of computing device, e.g. a computer, network device or appliance capable of communicating on any type and form of network and performing the operations described herein. FIGs. 7C and 7D depict block diagrams of a computing device 100 useful for practicing an embodiment of the client 102 or a server 106. As shown in FIGs. 7C and 7D, each computing device 100 includes a central processing unit 121, and a main memory unit 122. As shown in FIG. 7C, a computing device 100 may include a storage device 128, an installation device 116, a network interface 118, an I / O controller 123, display devices 124a-124n, a keyboard 126 and a pointing device 127, e.g. a mouse. The storage device 128 may include, without limitation, an operating system, software, and a software of a genomic data processing system 120. As shown in FIG. 7D, each computing device 100may also include additional optional elements, e.g. a memory port 103, a bridge 170, one or more input / output devices 130a-130n (generally referred to using reference numeral 130), and a cache memory 140 in communication with the central processing unit 121.

[0166] The central processing unit 121 is any logic circuitry that responds to and processes instructions fetched from the main memory unit 122. In many embodiments, the central processing unit 121 is provided by a microprocessor unit, e.g. : those manufactured by Intel Corporation of Mountain View, California; those manufactured by Motorola Corporation of Schaumburg, Illinois; the ARM processor and TEGRA system on a chip (SoC) manufactured by Nvidia of Santa Clara, California; the POWER7 processor, those manufactured by International Business Machines of White Plains, New York; or those manufactured by Advanced Micro Devices of Sunnyvale, California. The computing device 100 may be based on any of these processors, or any other processor capable of operating as described herein. The central processing unit 121 may utilize instruction level parallelism, thread level parallelism, different levels of cache, and multi -core processors. A multi -core processor may include two or more processing units on a single computing component. Examples of multi-core processors include the AMD PHENOM IIX2, INTEL CORE i5 and INTEL CORE i7.

[0167] Main memory unit or memory device 122 may include one or more memory chips capable of storing data and allowing any storage location to be directly accessed by the microprocessor 121. Main memory unit or device 122 may be volatile and faster than storage 128 memory. Main memory units or devices 122 may be Dynamic random access memory (DRAM) or any variants, including static random access memory (SRAM), Burst SRAM or SynchBurst SRAM (BSRAM), Fast Page Mode DRAM (FPM DRAM), Enhanced DRAM (EDRAM), Extended Data Output RAM (EDO RAM), Extended Data Output DRAM (EDO DRAM), Burst Extended Data Output DRAM (BEDO DRAM), Single Data Rate Synchronous DRAM (SDR SDRAM), Double Data Rate SDRAM (DDR SDRAM), Direct Rambus DRAM (DRDRAM), or Extreme Data Rate DRAM (XDR DRAM). In some embodiments, the main memory 122 or the storage 128 may be non- volatile; e.g., non-volatile read access memory (NVRAM), flash memory non-volatile static RAM (nvSRAM), Ferroelectric RAM (FeRAM), Magnetoresistive RAM (MRAM), Phase- change memory (PRAM), conductive-bridging RAM (CBRAM), Silicon-Oxide-Nitride- Oxide-Silicon (SONOS), Resistive RAM (RRAM), Racetrack, Nano-RAM (NRAM), orMillipede memory. The main memory 122 may be based on any of the above described memory chips, or any other available memory chips capable of operating as described herein. In the embodiment shown in FIG. 7C, the processor 121 communicates with main memory 122 via a system bus 150 (described in more detail below). FIG. 7D depicts an embodiment of a computing device 100 in which the processor communicates directly with main memory 122 via a memory port 103. For example, in FIG. 7D the main memory 122 may be DRDRAM.

[0168] FIG. 7D depicts an embodiment in which the main processor 121 communicates directly with cache memory 140 via a secondary bus, sometimes referred to as a backside bus. In other embodiments, the main processor 121 communicates with cache memory 140 using the system bus 150. Cache memory 140 typically has a faster response time than main memory 122 and is typically provided by SRAM, BSRAM, or EDRAM. In the embodiment shown in FIG. 7D, the processor 121 communicates with various VO devices 130 via a local system bus 150. Various buses may be used to connect the central processing unit 121 to any of the I / O devices 130, including a PCI bus, a PCI-X bus, or a PCI-Express bus, or a NuBus. For embodiments in which the VO device is a video display 124, the processor 121 may use an Advanced Graphics Port (AGP) to communicate with the display 124 or the VO controller 123 for the display 124. FIG. 7D depicts an embodiment of a computer 100 in which the main processor 121 communicates directly with VO device 130b or other processors 12 V via HYPERTRANSPORT, RAPID IO, or INFINIBAND communications technology. FIG. 7D also depicts an embodiment in which local busses and direct communication are mixed: the processor 121 communicates with VO device 130a using a local interconnect bus while communicating with VO device 130b directly.

[0169] A wide variety of VO devices 130a-130n may be present in the computing device 100. Input devices may include keyboards, mice, trackpads, trackballs, touchpads, touch mice, multi-touch touchpads and touch mice, microphones, multi -array microphones, drawing tablets, cameras, single-lens reflex camera (SLR), digital SLR (DSLR), CMOS sensors, accelerometers, infrared optical sensors, pressure sensors, magnetometer sensors, angular rate sensors, depth sensors, proximity sensors, ambient light sensors, gyroscopic sensors, or other sensors. Output devices may include video displays, graphical displays, speakers, headphones, inkjet printers, laser printers, and 3D printers.

[0170] Devices 130a- 13 On may include a combination of multiple input or output devices, including, e.g., Microsoft KINECT, Nintendo Wiimote for the WII, Nintendo WII U GAMEPAD, or Apple IPHONE. Some devices 130a- 13 On allow gesture recognition inputs through combining some of the inputs and outputs. Some devices 130a- 13 On provides for facial recognition which may be utilized as an input for different purposes including authentication and other commands. Some devices 130a-130n provides for voice recognition and inputs, including, e.g., Microsoft KINECT, SIRI for IPHONE by Apple, Google Now or Google Voice Search.

[0171] Additional devices 130a- 13 On have both input and output capabilities, including, e.g., haptic feedback devices, touchscreen displays, or multi-touch displays. Touchscreen, multi-touch displays, touchpads, touch mice, or other touch sensing devices may use different technologies to sense touch, including, e.g., capacitive, surface capacitive, projected capacitive touch (PCT), in-cell capacitive, resistive, infrared, waveguide, dispersive signal touch (DST), in-cell optical, surface acoustic wave (SAW), bending wave touch (BWT), or force-based sensing technologies. Some multi-touch devices may allow two or more contact points with the surface, allowing advanced functionality including, e.g., pinch, spread, rotate, scroll, or other gestures. Some touchscreen devices, including, e.g., Microsoft PIXELSENSE or Multi-Touch Collaboration Wall, may have larger surfaces, such as on a table-top or on a wall, and may also interact with other electronic devices. Some I / O devices 130a-130n, display devices 124a-124n or group of devices may be augment reality devices. The I / O devices may be controlled by an I / O controller 123 as shown in FIG. 7C. The I / O controller may control one or more I / O devices, such as, e.g., a keyboard 126 and a pointing device 127, e.g., a mouse or optical pen. Furthermore, an I / O device may also provide storage and / or an installation medium 116 for the computing device 100. In still other embodiments, the computing device 100 may provide USB connections (not shown) to receive handheld USB storage devices. In further embodiments, an I / O device 130 may be a bridge between the system bus 150 and an external communication bus, e.g. a USB bus, a SCSI bus, a FireWire bus, an Ethernet bus, a Gigabit Ethernet bus, a Fibre Channel bus, or a Thunderbolt bus.

[0172] In some embodiments, display devices 124a-124n may be connected to I / O controller 123. Display devices may include, e.g, liquid crystal displays (LCD), thin film transistor LCD (TFT-LCD), blue phase LCD, electronic papers (e-ink) displays, flexiledisplays, light emitting diode displays (LED), digital light processing (DLP) displays, liquid crystal on silicon (LCOS) displays, organic light-emitting diode (OLED) displays, active- matrix organic light-emitting diode (AMOLED) displays, liquid crystal laser displays, time- multiplexed optical shutter (TMOS) displays, or 3D displays. Examples of 3D displays may use, e.g. stereoscopy, polarization filters, active shutters, or autostereoscopy. Display devices 124a-124n may also be a head-mounted display (HMD). In some embodiments, display devices 124a-124n or the corresponding I / O controllers 123 may be controlled through or have hardware support for OPENGL or DIRECTX API or other graphics libraries.

[0173] In some embodiments, the computing device 100 may include or connect to multiple display devices 124a-124n, which each may be of the same or different type and / or form. As such, any of the I / O devices 130a-130n and / or the I / O controller 123 may include any type and / or form of suitable hardware, software, or combination of hardware and software to support, enable or provide for the connection and use of multiple display devices 124a-124n by the computing device 100. For example, the computing device 100 may include any type and / or form of video adapter, video card, driver, and / or library to interface, communicate, connect or otherwise use the display devices 124a-124n. In one embodiment, a video adapter may include multiple connectors to interface to multiple display devices 124a-124n. In other embodiments, the computing device 100 may include multiple video adapters, with each video adapter connected to one or more of the display devices 124a-124n. In some embodiments, any portion of the operating system of the computing device 100 may be configured for using multiple displays 124a-124n. In other embodiments, one or more of the display devices 124a-124n may be provided by one or more other computing devices 100a or 100b connected to the computing device 100, via the network 104. In some embodiments software may be designed and constructed to use another computer’s display device as a second display device 124a for the computing device 100. For example, in one embodiment, an Apple iPad may connect to a computing device 100 and use the display of the device 100 as an additional display screen that may be used as an extended desktop. One ordinarily skilled in the art will recognize and appreciate the various ways and embodiments that a computing device 100 may be configured to have multiple display devices 124a-124n.

[0174] Referring again to FIG. 7C, the computing device 100 may comprise a storage device 128 (e.g. one or more hard disk drives or redundant arrays of independent disks) for storing an operating system or other related software, and for storing application software programs such as any program related to the software for the genomic data processing system 120. Examples of storage device 128 include, e.g, hard disk drive (HDD); optical drive including CD drive, DVD drive, or BLU-RAY drive; solid-state drive (SSD); USB flash drive; or any other device suitable for storing data. Some storage devices may include multiple volatile and non-volatile memories, including, e.g, solid state hybrid drives that combine hard disks with solid state cache. Some storage device 128 may be non-volatile, mutable, or read-only. Some storage device 128 may be internal and connect to the computing device 100 via a bus 150. Some storage devices 128 may be external and connect to the computing device 100 via an I / O device 130 that provides an external bus. Some storage device 128 may connect to the computing device 100 via the network interface 118 over a network 104, including, e.g., the Remote Disk for MACBOOK AIR by Apple. Some client devices 100 may not require a non-volatile storage device 128 and may be thin clients or zero clients 102. Some storage device 128 may also be used as an installation device 116, and may be suitable for installing software and programs.Additionally, the operating system and the software can be run from a bootable medium, for example, a bootable CD, e.g. KNOPPIX, a bootable CD for GNU / Linux that is available as a GNU / Linux distribution from knoppix.net.

[0175] Client device 100 may also install software or application from an application distribution platform. Examples of application distribution platforms include the App Store for iOS provided by Apple, Inc., the Mac App Store provided by Apple, Inc., GOOGLE PLAY for Android OS provided by Google Inc., Chrome Webstore for CHROME OS provided by Google Inc., and Amazon Appstore for Android OS and KINDLE FIRE provided by Amazon.com, Inc. An application distribution platform may facilitate installation of software on a client device 102. An application distribution platform may include a repository of applications on a server 106 or a cloud 108, which the clients 102a- 102n may access over a network 104. An application distribution platform may include application developed and provided by various developers. A user of a client device 102 may select, purchase and / or download an application via the application distribution platform.

[0176] Furthermore, the computing device 100 may include a network interface 118 to interface to the network 104 through a variety of connections including, but not limited to, standard telephone lines LAN or WAN links (e.g., 802.11, Tl, T3, Gigabit Ethernet, Infiniband), broadband connections (e.g., ISDN, Frame Relay, ATM, Gigabit Ethernet, Ethernet-over-SONET, ADSL, VDSL, BPON, GPON, fiber optical including FiOS), wireless connections, or some combination of any or all of the above. Connections can be established using a variety of communication protocols (e.g., TCP / IP, Ethernet, ARCNET, SONET, SDH, Fiber Distributed Data Interface (FDDI), IEEE 802.1 la / b / g / n / ac CDMA, GSM, WiMax and direct asynchronous connections). In one embodiment, the computing device 100 communicates with other computing devices 100’ via any type and / or form of gateway or tunneling protocol e.g. Secure Socket Layer (SSL) or Transport Layer Security (TLS), or the Citrix Gateway Protocol manufactured by Citrix Systems, Inc. of Ft. Lauderdale, Florida. The network interface 118 may comprise a built-in network adapter, network interface card, PCMCIA network card, EXPRESSCARD network card, card bus network adapter, wireless network adapter, USB network adapter, modem or any other device suitable for interfacing the computing device 100 to any type of network capable of communication and performing the operations described herein.

[0177] A computing device 100 of the sort depicted in FIGs. 7B and 7C may operate under the control of an operating system, which controls scheduling of tasks and access to system resources. The computing device 100 can be running any operating system such as any of the versions of the MICROSOFT WINDOWS operating systems, the different releases of the Unix and Linux operating systems, any version of the MAC OS for Macintosh computers, any embedded operating system, any real-time operating system, any open source operating system, any proprietary operating system, any operating systems for mobile computing devices, or any other operating system capable of running on the computing device and performing the operations described herein. Typical operating systems include, but are not limited to: WINDOWS 2000, WINDOWS Server 2022, WINDOWS CE, WINDOWS Phone, WINDOWS XP, WINDOWS VISTA, and WINDOWS 7, WINDOWS RT, WINDOWS 8, and WINDOWS 10, all of which are manufactured by Microsoft Corporation of Redmond, Washington; MAC OS and iOS, manufactured by Apple, Inc. of Cupertino, California; and Linux, a freely-available operating system, e.g. Linux Mint distribution (“distro”) or Ubuntu, distributed byCanonical Ltd. of London, United Kingdom; or Unix or other Unix-like derivative operating systems; and Android, designed by Google, of Mountain View, California, among others. Some operating systems, including, e.g., the CHROME OS by Google, may be used on zero clients or thin clients, including, c.g, CHROMEBOOKS.

[0178] The computer system 100 can be any workstation, telephone, desktop computer, laptop or notebook computer, netbook, ULTRABOOK, tablet, server, handheld computer, mobile telephone, smartphone or other portable telecommunications device, media playing device, a gaming system, mobile computing device, or any other type and / or form of computing, telecommunications or media device that is capable of communication. The computer system 100 has sufficient processor power and memory capacity to perform the operations described herein. The computer system 100 can be of any suitable size, such as a standard desktop computer or a Raspberry Pi 4 manufactured by Raspberry Pi Foundation, of Cambridge, United Kingdom. In some embodiments, the computing device 100 may have different processors, operating systems, and input devices consistent with the device. The Samsung GALAXY smartphones, e.g., operate under the control of Android operating system developed by Google, Inc. GALAXY smartphones receive input via a touch interface.

[0179] In some embodiments, the computing device 100 is a gaming system. For example, the computer system 100 may comprise a PLAYSTATION 3, or PERSONAL PLAYSTATION PORTABLE (PSP), or a PLAYSTATION VITA device manufactured by the Sony Corporation of Tokyo, Japan, a NINTENDO DS, NINTENDO 3DS, NINTENDO WII, or a NINTENDO WII U device manufactured by Nintendo Co., Ltd., of Kyoto, Japan, an XBOX 360 device manufactured by the Microsoft Corporation of Redmond, Washington.

[0180] In some embodiments, the computing device 100 is a digital audio player such as the Apple IPOD, IPOD Touch, and IPOD NANO lines of devices, manufactured by Apple Computer of Cupertino, California. Some digital audio players may have other functionality, including, e.g., a gaming system or any functionality made available by an application from a digital application distribution platform. For example, the IPOD Touch may access the Apple App Store. In some embodiments, the computing device 100 is a portable media player or digital audio player supporting file formats including, but not limited to, MP3, WAV, M4A / AAC, WMA Protected AAC, AIFF, Audible audiobook,Apple Lossless audio file formats and .mov, ,m4v, and .mp4 MPEG-4 (H.264 / MPEG-4 AVC) video file formats.

[0181] In some embodiments, the computing device 100 is a tablet e.g. the IPAD line of devices by Apple; GALAXY TAB family of devices by Samsung; or KINDLE FIRE, by Amazon.com, Inc. of Seattle, Washington. In other embodiments, the computing device 100 is an eBook reader, e.g. the KINDLE family of devices by Amazon.com, or NOOK family of devices by Barnes & Noble, Inc. of New York City, New York.

[0182] In some embodiments, the communications device 102 includes a combination of devices, e.g. a smartphone combined with a digital audio player or portable media player. For example, one of these embodiments is a smartphone, e.g. the IPHONE family of smartphones manufactured by Apple, Inc.; a Samsung GALAXY family of smartphones manufactured by Samsung, Inc.; or a Motorola DROID family of smartphones. In yet another embodiment, the communications device 102 is a laptop or desktop computer equipped with a web browser and a microphone and speaker system, e.g. a telephony headset. In these embodiments, the communications devices 102 are web-enabled and can receive and initiate phone calls. In some embodiments, a laptop or desktop computer is also equipped with a webcam or other video capture device that enables video chat and video call.

[0183] In some embodiments, the status of one or more machines 102, 106 in the network 104 are monitored, generally as part of network management. In one of these embodiments, the status of a machine may include an identification of load information (e.g., the number of processes on the machine, CPU and memory utilization), of port information (e.g., the number of available communication ports and the port addresses), or of session status (e.g., the duration and type of processes, and whether a process is active or idle). In another of these embodiments, this information may be identified by a plurality of metrics, and the plurality of metrics can be applied at least in part towards decisions in load distribution, network traffic management, and network failure recovery as well as any aspects of operations of the present solution described herein. Aspects of the operating environments and components described above will become apparent in the context of the systems and methods disclosed herein.

[0184] Referring to FIG. 8, in various embodiments, a system 2400 may include a computing device 2410 (or multiple computing devices, co-located or remote to each other), a sample processing system 2480, and an electronic health record (EHR) system 2490. In various embodiments, computing device 2410 (or components thereof) may be integrated with the sample processing system 2480 (or components thereof) and / or EHR system 2490 (or components thereof). In various embodiments, the sample processing system 2480 may include, may be, or may employ, in situ hybridization, PCR, Next-generation sequencing, Northern blotting, microarray, dot or slot blots, FISH, Western blotting, ELISA, colorimetric dye binding assays, complete blood count (CBC) panels, FACs, electrophoresis, chromatography, and / or mass spectroscopy on such biological sample as blood, plasma, serum, and / or tissue and / or Whole-body MRI and PET-CT scans of a subject. For example, in certain embodiments, the sample processing system 2490 may be or may include a Next-generation sequencer. In various embodiments, the EHR system 2490 may include, may be, or may employ, various computing devices that include health records of patients and study subjects (including devices of hospitals, clinics, healthcare practitioners, etc.), obtained from various sources, such as entries by healthcare practitioners, sample processing system 2480, university and hospital systems, government agency systems, etc.

[0185] In various embodiments, the computing device 2410 (or multiple computing devices) may be used to control, and receive signals acquired via, components of sample processing system 2480. The computing device 2410 may include one or more processors and one or more volatile and non-volatile memories for storing computing code and data that are captured, acquired, recorded, and / or generated. The computing device 2410 may include a control unit 2415 that in certain embodiments may be configured to exchange control signals with sample processing system 2480, allowing the computing device 2410 to be used to control, for example, processing of samples and / or scans and / or delivery of data generated and / or acquired through processing of samples and / or scans.

[0186] In various embodiments, computing device 2410 may include a data acquisition unit 2420 that may be configured to exchange control signals, or otherwise communicate, with sample processing system 2480 (or components thereof) and / or EHR system 2490, allowing the computing device 2410 to be used to control the capture of physiological data and / or signals via sensors of the sample processing system 2480, retrieve data or signals(e.g., from sample processing system 2480, EHR system 2490, and / or memory devices where data is stored), and direct transfer of data or signals (e.g., to sample processing system 2490 as feedback thereto, to EHR system 2490, to memory for storage, and / or to other systems or devices).

[0187] In various embodiment, a data analyzer 2425 may direct analysis of the data and signals, and output analysis results. Data analyzer 2425 may be used, for example, to transform raw data captured or obtained via sample processing system 2480 and / or EHR system 2490, and may employ pre-processing procedures involved in generating a training dataset. For example, in some implementations, data may be generated as a multi- dimensional array or vector with values representing, and to prevent the machine learning system from overemphasizing certain readings, values may be normalized to a predetermined range (e.g. 0-1, 0-100, or any other such range). The normalization may comprise linear rescaling, or may be a more complex function. In some implementations, dimension reduction may be performed to reduce large and sparse arrays or vectors. In some implementations, feature recognition may be performed to select a subset of features for further analysis, such as principal component analysis.

[0188] In various embodiments, a machine learning system 2430 may be used to implement various machine learning functionality discussed herein. Machine learning system 2430 may include a training engine 2435 configured to train predictive models using, for example, data obtained from or via data acquisition unit 2420 and / or processed data obtained from or via data analyzer 2425. The training engine 2435 may, for example, generate or obtain training datasets from or via data analyzer 2425 and may perform validation of datasets. The training engine 2435 may comprise a feature analyzer used to evaluate features by, for example, quantifying the impact of each feature on the developed model. Such a feature analyzer may, for example, uncover clinically important features that were globally predictive of the outcome, and may determine, for example, contributions of all features, or the top features (e.g., the top 2, top 5, top 10, top 15, top 20, top 25, top 30, etc.) on individual predictions. Features may be selected based on a threshold, such a percent contribution to predicting a medical condition, such as 0.5%, 1%, 2%, 5%, 10%, etc. A testing and application engine 2440 may be configured to test and apply models trained via training engine 2435 to, for example, study subject and / or patient data from data acquisition unit 2420 and / or data analyzer 2425.

[0189] In various embodiments, a transceiver 2445 allows the computing device 2410 to exchange readings, control commands, and / or other data with sample processing system 2480 (or components thereof) and / or EHR system 2490 (or components thereof). The transceiver 2445 may additionally or alternatively include a network interface permitting the computing device 2410 to communicate with other remote devices and systems via, for example, a telecommunications network such as the internet. One or more user interfaces 2450 allow the computing device 2410 to receive user inputs (e.g., via a keyboard, touchscreen, microphone, camera, etc.) and provide outputs (e.g., via a touchscreen or other display screen, audio speakers, haptic devices, etc.). A display screen may be employed, for example, to provide real time or near real time waveforms or other readings or measurements obtained via sensors being used to capture physiological data from subjects and patients. The computing device 2410 may additionally include one or more databases 2455 (stored in, e.g., one or more computer-readable non-volatile memory devices) for storing, for example, data and analyses obtained from or via data acquisition unit 2420, data analyzer 2425, machine learning system 2430 (e.g., training engine 2435 and / or testing and application engine 2440), sample processing system 2480, and / or EHR system 2490. In some implementations, database 2455 (or portions thereof) may alternatively or additionally be part of another computing device that is co-located or remote and in communication with computing device 2410, sample processing system 2480 (or components thereof), and / or EHR system 2490.

[0190] In one aspect, the present disclosure provides a method of training a machine learning classifier for estimating risk of cancer-associated venous thromboembolism (VTE) in cancer patients comprising: (a) receiving data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generating a training dataset based on the received data, wherein the training dataset comprises a plurality of features for each subject in the cohort, wherein the plurality of features comprises (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE in cancer patients, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one ormore machine learning models with an accuracy that exceeds an accuracy threshold for the classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients. The subjects in the cohort may be chemotherapy-naive or may have received systemic chemotherapy. Additionally or alternatively, in certain embodiments, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

[0191] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier. Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0192] Additionally or alternatively, in some embodiments of the methods disclosed herein, the cancer patients comprise metastatic sites of disease. In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0193] In any of the preceding embodiments, the method further comprises applying the classifier to data on a cancer patient to generate a predictor, and determining whether the cancer patient is at risk for cancer-associated VTE based on the predictor and the operating- point threshold. In some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.

[0194] In any of the foregoing embodiments, the method further comprises administering an effective amount of anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin.

[0195] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0196] In one aspect, the present disclosure provides a method of estimating risk of cancer-associated venous thromboembolism (VTE) in a cancer patient using a machine learning classifier, the method comprising: receiving patient data corresponding to a plurality of features for the cancer patient; applying the machine learning classifier to the patient data to generate a predictor; and determining whether the cancer patient is at risk for cancer-associated VTE based on the predictor and an operating-point threshold, wherein the machine learning classifier is trained by: (a) receiving cohort data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generating a training dataset based on the received cohort data, wherein the training dataset comprises the plurality of features for each subject in the cohort, wherein the plurality of features comprises (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR- gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining the optimal operating-pointthreshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients. In some embodiments, the method further comprises administering an effective amount of anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin. Additionally or alternatively, in some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer- associated VTE. The subjects in the cohort may be chemotherapy-naive or may have received systemic chemotherapy. In any of the preceding embodiments of the methods disclosed herein, one or more of the plurality of features for the cancer patient are determined by assaying whole blood, serum or plasma.

[0197] Additionally or alternatively, in certain embodiments, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

[0198] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemblelearning random forest classifier. Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0199] Additionally or alternatively, in some embodiments of the methods disclosed herein, the cancer patient comprises metastatic sites of disease. In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0200] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0201] In any and all embodiments of the methods disclosed herein, one or more of the plurality of features for each subject in the cohort are determined by assaying whole blood, serum or plasma.

[0202] In any and all embodiments of the methods disclosed herein, the cancer- associated VTE is pulmonary embolism or lower extremity deep vein thrombosis (DVT), optionally wherein lower extremity DVT includes thrombi involving a common iliac vein, an external iliac vein, a common femoral vein, a superficial femoral vein, a deep femoral vein, a popliteal vein, a peroneal vein, an anterior tibial vein, a posterior tibial vein, or a deep calf vein.

[0203] In another aspect, the present disclosure provides a machine learning system for training a machine learning classifier for estimating risk of cancer-associated venous thromboembolism (VTE) in cancer patients, the system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: (a) receive data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generate a training dataset based on the received data, wherein the training dataset comprises a plurality of features for each subject in the cohort, wherein the plurality of features comprises (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE in cancer patients; wherein applying the machinelearning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients. The subjects in the cohort may be chemotherapy-naive or may have received systemic chemotherapy.

[0204] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier.

[0205] Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0206] Additionally or alternatively, in some embodiments of the methods disclosed herein, the cancer patients comprise metastatic sites of disease. In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0207] Additionally or alternatively, in certain embodiments of the systems disclosed herein, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia,primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral NervousSystem, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

[0208] In any of the preceding embodiments of the systems described herein, the instructions further cause the processor to apply the machine learning classifier to data on a cancer patient to generate a predictor, and determine whether the cancer patient is at risk for cancer-associated VTE based on the predictor and the operating-point threshold. In some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer- associated VTE.

[0209] In any of the foregoing embodiments of the systems described herein, the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin.

[0210] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0211] In yet another aspect, the present disclosure provides a computing system for estimating risk of cancer-associated venous thromboembolism (VTE) in a cancer patient, the computing system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: receive patient data corresponding to a plurality of features for the cancer patient; apply a machine learning classifier to the patient data to generate a predictor; and determine whether the cancer patient is at risk for cancer- associated VTE based on the predictor and an operating-point threshold, wherein the classifier is trained by: (a) receiving cohort data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generating a training dataset based on the received cohort data, wherein the training dataset comprises the plurality of features for each subject in the cohort, wherein the plurality of features comprises (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and(ii) cancer type; and (c) applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining the optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients.

[0212] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier.

[0213] Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0214] Additionally or alternatively, in some embodiments of the systems disclosed herein, the cancer patient comprises metastatic sites of disease. In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0215] In any of the preceding embodiments of the systems described herein, the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. In some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin.

[0216] Additionally or alternatively, in certain embodiments of the systems disclosed herein, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma,retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

[0217] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0218] In any and all embodiments of the systems disclosed herein, one or more of the plurality of features for each subject in the cohort or the cancer patient are determined by assaying whole blood, serum or plasma.

[0219] In one aspect, the present disclosure provides a non-transitory computer-readable storage medium comprising instructions which, when executed by a processor of a machine learning system, configure the machine learning system to train a machine learning classifier to estimate risk of cancer-associated venous thromboembolism (VTE) in cancer patients, wherein the instructions are configured to cause the processor to: (a) receive data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generate a training dataset based on the received data, wherein the training dataset comprises a plurality of features for each subject in the cohort, the plurality of features comprising (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer- associated VTE in cancer patients; wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with anaccuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients. The subjects in the cohort may be chemotherapy-naive or may have received systemic chemotherapy.

[0220] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier.

[0221] Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0222] Additionally or alternatively, in some embodiments of the computer-readable storage medium disclosed herein, the cancer patient comprises metastatic sites of disease.In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0223] In any of the preceding embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to apply the machine learning classifier to data on a cancer patient to generate a predictor, and determine whether the cancer patient is at risk for cancer-associated VTE based on the predictor and the operating- point threshold. In some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.

[0224] Additionally or alternatively, in certain embodiments of the computer-readable storage medium disclosed herein, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer,ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non- melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non- Hodgkin lymphoma.

[0225] In any of the preceding embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin.

[0226] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0227] In another aspect, the present disclosure provides a non-transitory computer- readable storage medium comprising instructions which, when executed by a processor of a computing system, configure the computing system to estimate risk of cancer-associated venous thromboembolism (VTE) in a cancer patient, wherein the instructions are configured to cause the processor to: receive patient data corresponding to a plurality of features for the cancer patient; apply a machine learning classifier to the patient data to generate a predictor; and determine whether the cancer patient is at risk for cancer-associated VTE based on the predictor and an operating-point threshold, wherein the classifier is trained by: (a) receiving cohort data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; (b) generating a training dataset based on the received cohort data, wherein the training dataset comprises the plurality of features for each subject in the cohort, wherein the plurality of features comprises (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and (c) applying a machine learning method to the training dataset to develop the machine learningclassifier for estimating risk of cancer-associated VTE, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining the optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients.

[0228] The machine learning technique may model survival outcomes with competing risks. In some embodiments, the machine learning technique is a random forest technique, and the one or more machine learning models are random forest models. Additionally or alternatively, in certain embodiments, the machine learning classifier is an ensemble learning random forest classifier.

[0229] Additionally or alternatively, in some embodiments, performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

[0230] Additionally or alternatively, in some embodiments of the computer-readable storage medium disclosed herein, the cancer patient comprises metastatic sites of disease.In certain embodiments, the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

[0231] In any of the preceding embodiments of the computer-readable storage medium described herein, the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold. In some embodiments, the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.Examples of anticoagulant therapy include, but are not limited to, apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, and enoxaparin.

[0232] Additionally or alternatively, in certain embodiments of the computer-readable storage medium disclosed herein, the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer,histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non- melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non- Hodgkin lymphoma.

[0233] In some embodiments, the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy. Additionally or alternatively, in some embodiments, the cancer patient has a Khorana Score ≥ 2.

[0234] In any of the preceding embodiments of the computer-readable storage medium disclosed herein, one or more of the plurality of features for the cancer patient are determined by assaying whole blood, serum or plasma.EXAMPLES

[0235] The present technology is further illustrated by the following Examples, which should not be construed as limiting in any way.Example 1: Experimental Methods

[0236] Patient Cohorts

[0237] Discovery Cohort:

[0238] For the discovery cohort, we utilized baseline plasma samples from the HYPERCAN study. The HYPERCAN study is a multicenter Italian prospective cohort study (NCT02622815). The study includes enrolling patients in NSCLC, breast, and gastrointestinal (GI) (PMID 27067979). Inclusion criteria included patients with metastatic disease (TXNXM1) and patients that underwent radical resection of the primary tumor: breast (T2-T4 NO M0, TX N+ M0, including triple negative, HER-2+), GI (T4N0M0, TXN+M0), and NSCLC (T3-T4N0M0, TXN+M0). Patients with acute medical illnesses,including recent thrombosis, hospitalization or terminal conditions, use of anticoagulants (therapeutic dose), and life expectancy were excluded. For all study subjects, data was recorded at enrolment, including age, gender, tumor type and histology, smoking or history of smoking, current medications, performance status, relevant comorbidities, anticancer treatment, previous thrombosis (both arterial and venous), baseline use of anticoagulants, presence of central venous catheter (CVC) and recent surgery. Patients were followed prospectively, and clinical information was collected, including disease recurrence, disease progression, and thrombo-hemorrhagic events. Blood samples were collected at baseline and specific timepoints, processed, and stored at -80°C in the Bergamo Hospital Biobank, Italy.

[0239] Validation Cohort:

[0240] For the external validation cohort, baseline samples from the placebo arm of the AVERT study were accessed. (NCT02048865) The AVERT study was a randomized, placebo-controlled, double-blind clinical trial in Canada assessing the efficacy and safety of apixaban for thromboprophylaxis in ambulatory patients with cancer who were at intermediate-to-high risk for venous thromboembolism (Khorana score, ≥2) on initiation of systemic chemotherapy. (PMID 30511879). Patients were enrolled from February 2014 through April 2018. Patients with malignancy solely including basal -cell or squamous-cell skin carcinoma, acute leukemia, or myeloproliferative neoplasm were excluded. Other relevant exclusion criteria included life expectancy of less than 6 months, renal and hepatic dysfunction, pregnancy, breastfeeding, other anti coagulation use and thrombocytopenia (<50,000 / uL). The primary outcome was venous thromboembolism over a follow-up period of 180 days. Diagnostic imaging for VTE was performed during the follow-up period for patients who presented with symptoms of venous thromboembolism; no routine ultrasonographic testing was performed. All trial outcomes were adjudicated by an independent adjudication committee whose members were unaware of the treatment assignments. Blood samples were obtained from all consenting patients at 1, 3, 6, and 7- month follow-up visits.

[0241] Proteomics:

[0242] We utilized the commercially available Olink® panels from Olink (Uppsala, Sweden) to assess baseline plasma proteins (https: / / olink.com / products-services / explore / ). Briefly, double oligonucleotide-labeled antibody probes bind with the target protein with high specificity, and real-time PCR amplification of the oligonucleotide sequence is used to detect the resulting DNA sequence quantitatively. Using internal and external controls, the resulting threshold cycle (Ct)-data are processed for quality control and normalized.

[0243] Data Normalization and Bath Effect Correction

[0244] Olink's proteomics data were processed using Normalized Protein expression (NPX) values, which provide relative quantification of protein levels on a logarithmic scale. To reduce batch-to-batch variation, minimize the confounding effects of batch variations, and enhance result reliability and reproducibility, bridge normalization with OlinkAnalyze package was applied in the different cohorts following best practices. Six samples from Lung and 8 from Gastric cancer cohorts from the Avert study were selected as “bridge” for correcting NPX values in the different cohorts separately. These samples were used to calculate correction factors for each batch. The factors were subsequently applied to the NPX values of each assay in the Lung and Gastric cohort samples, aligning them to the bridge samples themselves (data not shown). In the case of a gene present in more than one assay, the median value was calculated and selected for further analysis.

[0245] Machine Learning Model Deployment

[0246] Eleven proteins were selected out of 990, and 3 clinical characteristics for the model deployment. Highly correlated features were initially removed as determined by Pearson’s correlation analysis (>0.8), and proteins with an AUC > 0.65 were retained (data not shown). Random feature selection was subsequently performed to select the minimum number of proteins with the highest predictive accuracy in the training set. A Naive Bayesian model was developed using the following parameters: usekernel=TRUE and adjust=l for Laplace smoothing. Five-fold cross-validation was performed on the training data to estimate the model’s accuracy. The model was further evaluated against state-of- the-art machine learning approaches, including Random Forest (RF), Generalized Linear Model (GLM), Support Vector Machine (SVM), and Gradient Boosting Machine (GBM). The models were deployed in a default setting, using the same features (11 proteomic and 3 clinical features) as the initial approach and 5 -fold cross-validation. Specifically, Random Forest was run with a number of trees equal to 500 and mtry set to the square root of thenumber of features. The Generalized Linear Model (GLM) was employed with the family parameter set to Gaussian (link="identity"), signifying a linear model. The Gradient Boosting Machine (GBM) was configured with a number of trees equal to 100, the depth of the trees equal to 1, minimum observations in the nodes of 10, and shrinkage equal to 0.1 as the learning rate. The Support Vector Machine (SVM) was applied using a cost of 1 for the constraints violation and the kernel set to “radial” for a radial basis function kernel. Naive Bayes presented the highest performance with an AUC of 0.849 compared to a mean AUC of 0.817 (0.796-0.839) that was presented by the other approaches (data not shown). Caret package was utilized for model training and evaluation.

[0247] CD200R1 knockout mice

[0248] CD200R1 knockout (Cd200rl~ ~') mice on the C57BL / 6J background were a gift from Xue-Feng Bai at The Ohio State University (PMID: 27385779). C57BL / 6J mice (stock # 000644) were purchased from Jackson Laboratory. All mice were maintained and cared for in the Beth Israel Deaconess Medical Center laboratory animal facility which is fully accredited by the Institutional Animal Care and Use Committee (IACUC).

[0249] Mouse blood collection, CBCs, and plasma isolation

[0250] Mice between 8 and 9 weeks old were anesthetized with pentobarbital, and blood was collected from the inferior vena cava (IVC) 9: 1 in sodium citrate. Complete blood counts (CBCs) were measured in whole blood on the HemaVet. Plasma was isolated with two centrifugation steps at 5,000 rpm for 10 minutes, as previously described (PMID: 30875463). An equal number of males and females was used for all experiments.

[0251] Coagulation assays and ELISAs

[0252] All coagulation assays were performed on the Stago STart coagulometer. The prothrombin time (PT) and activated partial thromboplastin time (aPTT) were performed using the STA-Neoplastine CI Plus 5 and the STA-PTT Automate 5 (Stago), respectively. The Fibrinogen assay was performed using the STA-Fibrinogen 5 (Stago) and plasma was diluted 1 : 10 in imidazole buffer. The FII, FV, FX, and FVII activity assays were performed as previously described (PMID: 30875463), using STA-Deficient II, STA-Deficient V, STA-Deficient X (Stago), and Human FVII Deficient plasma (Prolytix), respectively. To measure FVIIa activity, the Staclot VIIa-rTF kit was used following manufacturer protocol and concentration values were converted to activity by normalizing to C57BL / 6J mice.

[0253] The thrombin-antithrombin (TAT) complex was measured using the Mouse Thrombin / Antithrombin Complex ELISA (Innovative Research). D-dimer was measured using the Asserachrom D-Di ELISA (Stago) and mouse plasma samples were diluted 1 : 10 in kit reagent diluent. Plasma tissue factor (TF) was measured undiluted using the Mouse Coagulation Factor III / TF DuoSet ELISA (R&D Systems). Tissue factor pathway inhibitor (TFPI) was measured using the Mouse TFPI DuoSet ELISA (R&D Systems). Absorbance was read on the BioTek Synergy HTX.

[0254] Thrombin Generation Assay

[0255] Plasma was diluted 1 :3 with the Technothrombin Thrombin Generation Assay (TGA) Buffer (Diapharma). The TGA was performed following the manufacturer’s protocol using the Technothrombin TGA Substrate and Reagent C Low (Diapharma). The thrombin calibration curve was determined using the Technothrombin TGA CAL set (Diapharma). Fluorescence was read on the BioTek Cytation 5. The data was analyzed using the Technothrombin TGA Evaluation Software from Technoclone.

[0256] Mouse Plasma Proteomics and AnalysisWe utilized the commercially available Olink Target 96 Mouse Exploratory panel to measure mouse plasma proteins. After eliminating assays with low QC, 87 assays were analyzed in GraphPad in an unpaired t-test with FDR correction for multiple comparisons. Significance was determined by a q<0.05.Example 2: A Bayesian probabilistic model to predict VTE in patients with cancer

[0257] FIG. 1 shows an overview of the model, feature extraction, and training of the present technology. The discovery cohort consisted of patients with non-small cell lung cancer (NSCLC) (n=77) and gastric cancer (n=101) enrolled in the HYPERCAN study.13Approximately 60% of patients included in this cohort had evidence of metastatic disease at time of collection.

[0258] Plasma samples were collected prior to receiving cytotoxic chemotherapy and patients were followed prospectively for the development of VTE over 6 months. We performed proteomic analysis using a proximity extension assay platform that included 1,161 unique plasma proteins. A naive Bayesian machine learning model was employed that identified 11 proteins and 3 salient clinical features (age, sex, and prior history of VTE). Five-fold cross-validation was performed on the training data to estimate themodel’s accuracy. The mean c-statistic for the model was 0.84 at a 95% CI for the prediction of VTE at 6 months (FIG. 2A). This thrombosis oncology prediction (“TOP”) model demonstrated superiority relative to the standard-of-care risk prediction model; the Khorana score had limited predictive discrimination in this cohort (c-statistic 0.48 at a 95% CI). The difference between the mean c-statistics for the TOP model and the Khorana score was statistically significant.

[0259] We compared the TOP model relative to common coagulation assays indicative of thrombin generation (i.e., endogenous thrombin potential, peak thrombin generation, and prothrombin fragment 1+2) as well as fibrinogen. As shown in FIG. 2B, none of these assays accurately predicted the development of VTE (mean ACC for all assays was 0.58, 0.51-0.63).

[0260] The discovery cohort included patients with either non-small cell lung cancer and gastric cancer. To assess the relative predictive accuracy in these two cohorts separately, we first identified a “high-risk” cutoff by the receiver operator characteristic (ROC) analysis. The cumulative incidence of VTE at 6 months was estimated by competing risk analysis, with death considered as a competing risk for VTE. The TOP model was highly predictive of thrombosis events in the patients with NSCLC (FIG. 2C, Grey test <0.001).

[0261] Similarly, the TOP model accurately identified those patients with gastric cancer at significantly higher risk for VTE (FIG. 2D, Grey test=0.001). Overall, the TOP model demonstrated improved discrimination and classification when compared with the Khorana Score. The TOP model reclassified 58% of all patients into revised strata with improved positive predictive values. For instance, in 16 patients who were classified as low risk by the Khorana Score and re-classified as high-risk by the TOP model had a 44% rate of VTE, and conversely the re-classification of 79 high-risk Khorana Score patients to the low-risk TOP category was associated with a 3.8% rate of VTE.

[0262] The present disclosure demonstrates that the 11 plasma proteins disclosed herein are useful biomarkers for accurately predicting the risk of cancer-associated thromboembolism in cancer patients.Example 3: Validation of the machine learning model in an independent validation cohort

[0263] To externally validate the TOP model, we analyzed baseline samples from a subset of patients enrolled in the AVERT clinical trial (a randomized controlled phase III study for primary thromboprophylaxis in patients with active malignancy) (Carrier et al., N Engl J Med. 2019 Feb 21;380(8):711-719) who were randomized to the placebo arm. A total of 72 patients with solid tumors we analyzed including pancreatic cancer (n=34), NSCLC (n=23), and gastric cancer (n=15). A total of 30 patients (41.7%) had metastatic disease at time of enrollment. The mean age was 64 years-old and 28 (38.9%) patients were female. There were 12 VTE events observed in the 6 months following of randomization. We performed proximity extension proteomic analysis to validate the TOP prediction model.

[0264] The 11 -protein model was similarly predictive of thrombosis in the external AVERT cohort (c-statistic AUC=0.71 at a 95% CI). Using an ROC-based cutoff value for the 11 -protein model, the higher risk group developed thrombosis at a significantly higher rate compared to those identified as lower risk with the model (23.7% versus 5.9%, Gray test P=0.03, FIG. 2E).

[0265] The present disclosure demonstrates that the 11 plasma proteins disclosed herein are useful biomarkers for accurately predicting the risk of cancer-associated thromboembolism in cancer patients.Example 4: Characterization of 11-protein array

[0266] The 11 -proteins included in the prediction model are CNDP1 (carnosine dipeptidsase 1), FR-gamma (folate receptor gamma), SERPINB8 (serpin family B member 8), MARCO (macrophage receptor collagenous structure), CD84 (cluster of differentiation 84), SPON2 (spondin 2), CD200R1 (cell surface glycoprotein CD200 receptor 1), CDH6 (cadherin 6), LGALS8 (galectin-8), LAT (linker for activation of T cells), DAB (DAB adaptor protein 2). See FIG. 3A. All proteins were reduced in the VTE cohort except carnosine dipeptidase 1 (CNDP1, FIG. 3B).

[0267] SERPINB8 is stored in platelets and has been shown to inhibit platelet aggregation in vitro (Leblond et al., Thromh Haemost. 2006 Feb;95(2):243-52). GAL8 are expressed in platelets after activation and promote platelet aggregation (Romaniuk et al., Biochem J. 2010 Dec 15;432(3):535-47). The other proteins identified are linked to avariety of pathways and cellular processes including cell cycle regulation, extra-cellular matrix, inflammatory response and chemokines and are not otherwise known to be associated with thrombosis. The selection of the 11 -proteins together could not have been predicted based on existing knowledge. An unexpected finding of the 11 -protein panel is that all proteins except CNDP1 were lower in the cohort who later developed thrombosis. Prior studies assessing biomarkers in cancer associated thrombosis have been focused on increased plasma levels (such as tissue factor, d-dimer, soluble P selectin, thrombin generation) as a risk factor for thrombosis (Es et al., 2017, supra).

[0268] CD200 receptor 1 (CD200R1), is a cell surface glycoprotein predominantly expressed on leukocytes, including various subsets of macrophages, dendritic cells, and certain T cells. Its expression is particularly notable on tumor-associated macrophages, while the ligand CD200 is also expressed across different tumor types. The interaction between the tumor-expressed ligand CD200 and its receptor CD200R1 on tumor-associated macrophages constitutes an important immune checkpoint pathway (PMID 24388216 and 36756156). CD200R1 was a major contributor to the TOP prediction model. Considering the role of CD200R1 in modulating tumor-leukocyte interactions, we further assessed CD200R1 as a potential prothrombotic driver.Example 5: CD200R1 deficient mice demonstrate a prothrombotic phenotype

[0269] As lower levels of CD200R1 were predictive of thrombosis in patients with cancer, we evaluated whether CD200R1 deficient mice demonstrated a basal prothrombotic phenotype. We observed significantly higher plasma levels of thrombin antithrombin (TAT) complexes in CD200rl' / _mice compared with wild type C57BL / 6J mice (FIG. 4A). TAT complexes reflect increased thrombin generation in vivo and considered indicative of prothrombotic state.

[0270] To understand the basis of hypercoagulability in the CD200R1 deficient mice, we assessed coagulation factors and assays including factors VII, II, V, X, prothrombin time, partial thromboplastin time, tissue factor, tissue factor pathway inhibitor, and thrombomodulin (FIG. 4B). CD200R1 deficient mice exhibited shortened prothrombin time compared to wild-type. No other significant differences in these coagulation assays were identified in the CD200R1 deficient mice. We measured thrombin generation ex vivo and none of the constituent parameters were significantly altered in the CD200R1 deficientmice (FIG. 4C). Both leukocyte and platelet counts were similar between the CD200R1 deficient and wild type mice (data not shown).

[0271] Absent an explanatory basis for the observed increase in markers of thrombin generation in vivo, we employed a discovery Olink proteomic proximity extension assay platform specifically developed for murine studies. We observed 9 proteins that were significantly increased in CD200R1' ' mice compared to C57BL / 6J controls: IL-17a, Notch3, Fst, Tgfbr3, Tppl, Fstl3, Hgf, IL- la and Cntn4 (FIG. 5A). The most apparent difference was a 250% increase in IL-17a in CD200R1 deficient mice compared with controls (p=0.00006). Elevated levels of IL-17a in mice were confirmed by ELISA (data not shown).Example 6: CD200R1 / IL-17 A axis induces endothelial activation

[0272] IL-17A is a pro-inflammatory cytokine produced primarily by Thl7 cells. IL- IYA is implicated in promoting a prothrombotic state by inducing the expression of tissue factor and other pro-coagulant factors in endothelial cells. Accordingly, we measured markers of endothelial activation in CD200R1 deficient mice. As show in FIG. SB, significantly higher levels of soluble P-selectin, E-selectin, and ICAM-1 were observed in CD200R1 deficient mice compared with wild type. Similarly, in the murine-Olink platform there were higher levels of endothelial activation markers in the CD200RL'" mice. These findings were confirmed histologically in pulmonary vasculature from CD200R1 deficient mice (data not shown).Example 7: Relationship between CD200R1 and IL-17a in clinical cohort

[0273] As the absence of CD200R1 in mice was associated with an increase in circulating IL-17a, we analyzed the association between IL17a levels and CD200R1 in a subset of patients enrolled in the discovery cohort. For these analyses we employed the Somalogic aptamer platform which includes a comprehensive panel of IL- 17 analytes including (IL- 17a, b, c, d, e) and is considered more accurate than Olink for cytokine analyses of human samples. Patients with the lowest quartile of CD200R1 had significantly higher levels of IL-17a compared with the highest quartile of CD200R1 (FIG. 5C). There was no association observed between CD00R1 and other IL- 17 components. These findings are supported by prior studies demonstrating an inverse relationship between CD200R1 expression on monocytes and the activation of Thl7 cells. We also compared IL-17a levels among those patients who later developed VTE. As shown in FIG. 5C, levels of IL-17awere significantly higher in those patients with cancer who developed VTE compared to those who did not.Example 8: Inhibition of Il-l 7a normalizes thrombin-antithrombin levels in CD200R1 deficient mice

[0274] To evaluate whether the increase in TAT levels in CD200R1 deficient mice was mediated through IL-17a, we administered an inhibitory IL-17a antibody. As shown in FIG. 6, baseline increases in CD200R1- / -mice compared with wild type controls normalized following administration of IL-17a inhibitory antibodies. These findings confirmed the central role of IL-17a in mediating a prothrombotic state in CD200Rr ' mice.Example 9: Antithrombotic activity of IL-17a inhibition in thrombo-inflammatory disorder in humans

[0275] While IL-17a has a recognized role in activating endothelium in cardiovascular disease, clinical trials have not been performed specifically to assess the antithrombotic efficacy in humans. There are several anti-IL17a antibodies approved by the FDA for the treatment of psoriasis and ankylosing spondylitis. As these conditions are not associated with an increase in thromboembolic events, we elected to analyze the antithrombotic efficacy of IL-17a inhibition in COVID-19, a thrombo-inflammatory condition associated with a high risk of thrombosis among those who are critically ill. We performed a systematic review and meta-analysis of patients with COVID-19 treated with anti-IL17a antibodies and discovered there was a significant reduction in VTE among patients with COVID-19 treated with anti-IL17a antibodies (data not shown) supporting the therapeutic potential of IL-17a in the prevention of VTE in thrombo-inflammatory diseases.EQUIVALENTS

[0276] The present technology is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the present technology. Many modifications and variations of this present technology can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the present technology, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the present technology. It is to be understood that this present technology is not limited to particular methods, reagents,compounds compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.

[0277] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.

[0278] As will be understood by one skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non- limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like, include the number recited and refer to ranges which can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 cells refers to groups having 1, 2, or 3 cells. Similarly, a group having 1-5 cells refers to groups having 1, 2, 3, 4, or 5 cells, and so forth.

[0279] All patents, patent applications, provisional applications, and publications referred to or cited herein are incorporated by reference in their entirety, including all figures and tables, to the extent they are not inconsistent with the explicit teachings of this specification.REFERENCES1. Elting LS, Escalante CP, Cooksley C, et al. Outcomes and cost of deep venous thrombosis among patients with cancer. Archives of internal medicine. 2004;164(15): 1653- 1661.2. Khorana AA, Francis CW, Culakova E, Kuderer NM, Lyman GH. Thromboembolism is a leading cause of death in cancer patients receiving outpatient chemotherapy. J Thromb Haemost. 2007;5(3):632-634.3. Lyman GH, Eckert L, Wang Y, Wang H, Cohen A. Venous thromboembolism risk in patients with cancer receiving chemotherapy: a real-world analysis. Oncologist.2013;18(12): 1321-1329.4. Prandoni P, Lensing AW, Piccioli A, et al. Recurrent venous thromboembolism and bleeding complications during anticoagulant treatment in patients with cancer and venous thrombosis. Blood. 2002;100(10):3484-3488.5. Khorana AA, Kuderer NM, Culakova E, Lyman GH, Francis CW. Development and validation of a predictive model for chemotherapy-associated thrombosis. Blood.2008; 11 l(10):4902-4907.6. Bosch FTM, Mulder FI, Kamphuisen PW, Middeldorp S, Bossuyt PM, Buller HR, van Es N. Primary thromboprophylaxis in ambulatory cancer patients with a high Khorana score: a systematic review and meta-analysis. Blood Adv. 2020;4(20):5215-5225.7. Carrier M, Abou-Nassar K, Mallick R, et al. Apixaban to Prevent Venous Thromboembolism in Patients with Cancer. N Engl J Med. 2019;380(8):711-719.8. Khorana AA, Soff GA, Kakkar AK, et al. Rivaroxaban for Thromboprophylaxis in High-Risk Ambulatory Patients with Cancer. N Engl J Med. 2019;380(8):720-728.9. Smith JG, Gerszten RE. Emerging Affinity-Based Proteomic Technologies for Large-Scale Plasma Profiling in Cardiovascular Disease. Circulation. 2017; 135(17): 1651- 1664.10. Assarsson E, Lundberg M, Holmquist G, et al. Homogenous 96-Plex PEA Immunoassay Exhibiting High Sensitivity, Specificity, and Excellent Scalability. PLOS ONE. 2014;9(4):e95192.11. Lundberg M, Thorsen SB, Assarsson E, et al. Multiplexed Homogeneous Proximity Ligation Assays for High-throughput Protein Biomarker Research in Serological Material. Molecular & Cellular Proteomics. 2011; 10(4).12. Ganz P, Heidecker B, Hveem K, et al. Development and Validation of a Protein- Based Risk Score for Cardiovascular Outcomes Among Patients With Stable Coronary Heart Disease. Jama. 2016;315(23):2532-2541.13. Falanga A, Santoro A, Labianca R, et al. Hypercoagulation screening as an innovative tool for risk assessment, early diagnosis and prognosis in cancer: the HYPERCAN study. Thromb Res. 2016;140 Suppl ES55-59.14. Liu Y, Gao L, Fan Y, Ma R, An Y, Chen G, Xie Y. Discovery of protein biomarkers for venous thromboembolism in non-small cell lung cancer patients through data- independent acquisition mass spectrometry. Front Oncol. 2023; 13: 1079719.15. Ercan H, Mauracher LM, Grilz E, et al. Alterations of the Platelet Proteome in Lung Cancer: Accelerated F13 Al and ER Processing as New Actors in Hypercoagulability. Cancers (Basel). 2021; 13(9).16. Mohammed Y, van Vlijmen BJ, Yang J, Percy AJ, Palmblad M, Borchers CH, Rosendaal FR. Multiplexed targeted proteomic assay to assess coagulation factor concentrations and thrombosis-associated cancer. Blood Adv. 2017; 1 (15): 1080- 1087.17. Shi X, Pang Q, Nian X, et al. Integrative transcriptome and proteome analyses of clear cell renal cell carcinoma develop a prognostic classifier associated with thrombus. Sci Rep. 2023;13(l):9778.18. Romaniuk MA, Tribulatti MV, Cattaneo V, Lapponi MJ, Molinas FC, Campetella O, Schattner M. Human platelets express and are activated by galectin-8. Biochem J. 2010;432(3):535-547.19. Romaniuk MA, Rabinovich GA, Schattner M. Gal ectins in the regulation of platelet biology. Methods Mol Biol. 2015;1207:269-283.20. Plantureux L, Mege D, Crescence L, et al. The Interaction of Platelets with Colorectal Cancer Cells Inhibits Tumor Growth but Promotes Metastasis. Cancer Res. 2020;80(2):291-303.21. Connors JM. Thrombophilia Testing and Venous Thrombosis. N Engl J Med. 2017;377(12): 1177-1187.22. Hoek RM, Ruuls SR, Murphy CA, et al. Down-regulation of the macrophage lineage through interaction with OX2 (CD200). Science. 2000;290(5497): 1768-1771.23. Rygiel TP, Karnam G, Goverse G, et al. CD200-CD200R signaling suppresses anti- tumor responses independently of CD200 expression on the tumor. Oncogene.2012;31(24):2979-2988.24. Varaljai R, Zimmer L, Al-Matary Y, et al. Interleukin 17 signaling supports clinical benefit of dual CTLA-4 and PD-1 checkpoint inhibition in melanoma. Nat Cancer.2023;4(9): 1292-1308.25. Vitiello GA, Miller G. Targeting the interleukin- 17 immune axis for cancer immunotherapy. J Exp Med. 2020;217(l).26. Fan H, Zhao J, Mao S, et al. Circulating Thl7 / Treg as a promising biomarker for patients with rheumatoid arthritis in indicating comorbidity with atherosclerotic cardiovascular disease. Clin Cardiol. 2023;46(12): 1519-1529.27. Miao Y, Yan T, Liu J, et al. Meta-analysis of the association between interleukin- 17 and ischemic cardiovascular disease. BMC Cardiovasc Disord. 2024;24(l):252.28. Griffin GK, Newton G, Tarrio ML, et al. IL- 17 and TNF-a sustain neutrophil recruitment during inflammation through synergistic effects on endothelial activation. J Immunol. 2012;188(12):6287-6299.29. Robert M, Miossec P. Effects of Interleukin 17 on the cardiovascular system. Autoimmun Rev. 2017; 16(9): 984-991.30. Gergely TG, Kucsera D, Toth VE, et al. Characterization of immune checkpoint inhibitor-induced cardiotoxicity reveals interleukin- 17A as a driver of cardiac dysfunction after anti-PD-1 treatment. Br J Pharmacol . 2023;180(6):740-761.31. Wang YN, Lou DF, Li DY, Jiang W, Dong JY, Gao W, Chen HC. Elevated levels of IL-17A and IL-35 in plasma and bronchoalveolar lavage fluid are associated with checkpoint inhibitor pneumonitis in patients with non-small cell lung cancer. Oncol Lett. 2020;20(l):611-622.32. Lechner MG, Cheng MI, Patel AY, et al. Inhibition of IL-17A Protects against Thyroid Immune-Related Adverse Events while Preserving Checkpoint Inhibitor Antitumor Efficacy. J Immunol. 2022;209(4):696-709.

Claims

CLAIMS1. A method for preventing cancer associated thromboembolism (CAT) in a cancer patient in need thereof comprising(a)(i) detecting mRNA and / or polypeptide expression levels of at least one of FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB in a biological sample obtained from the cancer patient that is decreased relative to a control sample obtained from a healthy subject or a predetermined threshold; and / or(ii) detecting mRNA and / or polypeptide expression levels of CNDP1 in a biological sample obtained from the cancer patient that is increased relative to a control sample obtained from a healthy subject or a predetermined threshold; and(b) administering to the cancer patient an effective amount of anticoagulant therapy.

2. A method for preventing cancer associated thromboembolism (CAT) in a cancer patient in need thereof comprising administering to the cancer patient an effective amount of anticoagulant therapy, wherein a biological sample obtained from the cancer patient comprises (a) mRNA and / or polypeptide expression levels of one or more plasma proteins that are decreased relative to a control sample obtained from a healthy subject or a predetermined threshold, wherein the one or more plasma proteins are selected from the group consisting of FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB; and / or (b) CNDP1 mRNA and / or polypeptide expression levels that are increased relative to a control sample obtained from a healthy subject or a predetermined threshold.

3. The method of claim 1 or 2, wherein the cancer patient is diagnosed with or suffers from a cancer selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterinesarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

4. The method of any one of claims 1-3, wherein the cancer is Stage 1, Stage 2, Stage 3, or Stage 4.

5. The method of any one of claims 1-4, wherein the anticoagulant therapy comprises one or more of apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, or enoxaparin.

6. The method of any one of claims 1-5, wherein the cancer patient has a Khorana Score < 2.

7. The method of any one of claims 1-5, wherein the cancer patient has a Khorana Score > 2.

8. The method of any one of claims 1-7, wherein mRNA expression levels are detected via real-time quantitative PCR (qPCR), digital PCR (dPCR), Reverse transcriptase-PCR (RT-PCR), Northern blotting, microarray, dot or slot blots, in situ hybridization, or fluorescent in situ hybridization (FISH).

9. The method of any one of claims 1-8, wherein polypeptide expression levels are detected via Western blotting, enzyme-linked immunosorbent assays (ELISA), dot blotting, immunohistochemistry, immunofluorescence, immunoprecipitation, immunoelectrophoresis, or mass-spectrometry.

10. The method of any one of claims 1-9, wherein the biological sample is whole blood, serum or plasma.

11. The method of any one of claims 1-10, wherein the lung cancer patient is chemotherapy-naive or has received / is receiving systemic chemotherapy.

12. The method of claim 11, wherein the systemic chemotherapy comprises one or more of an alkylating agent, an antibiotic, an antimetabolite, an antimitotic, a cyclin-dependentkinase inhibitor, an epidermal growth factor receptor inhibitor, a multikinase inhibitor, a PARP inhibitor, a platinum-based agent, a selective estrogen receptor modulator (SERM), or a VEGF inhibitor.

13. The method of any one of claims 1-12, wherein the lung cancer patient is immunotherapy-naive or has received / is receiving immunotherapy.

14. The method of claim 13, wherein the immunotherapy comprises one or more of anti- PD-1 antibody, anti-PD-Ll antibody, anti-PD-L2 antibody, anti-CTLA-4 antibody, anti- TIM3 antibody, anti-4-lBB antibody, anti-CD73 antibody, anti-GITR antibody, and anti- LAG-3 antibody.

15. The method of any one of claims 1-14, wherein the lung cancer patient is radiotherapy -naive or has received / is receiving radiotherapy.

16. The method of claim 15, wherein the radiotherapy comprises external radiotherapy, radiotherapy implants (brachytherapy), pre-targeted radioimmunotherapy, radiotherapy injections, radioisotope therapy, or intrabeam radiotherapy.

17. The method of any one of claims 1-16, wherein the CAT is pulmonary embolism or lower extremity deep vein thrombosis (DVT).

18. The method of claim 17, wherein lower extremity DVT includes thrombi involving a common iliac vein, an external iliac vein, a common femoral vein, a superficial femoral vein, a deep femoral vein, a popliteal vein, a peroneal vein, an anterior tibial vein, a posterior tibial vein, or a deep calf vein.

19. The method of any one of claims 1-18, wherein the cancer patient has one or more organ sites of metastasis.

20. The method of claim 19, wherein the one or more organ sites of metastasis comprise lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

21. A method of training a machine learning classifier for estimating risk of cancer- associated venous thromboembolism (VTE) in cancer patients, comprising: a. receiving data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types;b. generating a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and c. applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE in cancer patients, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients.

22. The method of claim 21, wherein the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NKneoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

23. The method of claim 21 or 22, wherein the cancer patients comprise metastatic sites of disease.

24. The method of claim 23, wherein the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

25. The method of any one of claims 21-24, wherein the machine learning technique is a random forest technique, and wherein the one or more machine learning models are random forest models.

26. The method of any one of claims 21-25, wherein the machine learning classifier is an ensemble learning random forest classifier.

27. The method of any one of claims 21-26, wherein the machine learning technique models survival outcomes with competing risks.

28. The method of any one of claims 21-27, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

29. The method of any one of claims 21-28, further comprising applying the classifier to data on a cancer patient to generate a predictor, and determining whether the cancer patient is at risk for cancer-associated VTE based on the predictor and the operating- point threshold.

30. The method of claim 29, wherein the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.

31. The method of claim 29 or 30, further comprising administering an effective amount of anticoagulant therapy to the cancer patient predicted to be at risk for cancer- associated VTE based on the predictor and the operating-point threshold.

32. The method of claim 31, wherein the anticoagulant therapy comprises one or more of apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, or enoxaparin.

33. The method of any one of claims 21-32, wherein the cancer patient has a Khorana Score ≥ 2.

34. The method of any one of claims 29-33, wherein the cancer patient is chemotherapy- naive or has received / is receiving systemic chemotherapy.

35. The method of any one of claims 21-34, wherein the subjects in the cohort are chemotherapy-naive or have received systemic chemotherapy.

36. A method of estimating risk of cancer-associated venous thromboembolism (VTE) in a cancer patient using a machine learning classifier, the method comprising: a. receiving patient data corresponding to a plurality of features for the cancer patient; b. applying the machine learning classifier to the patient data to generate a predictor; and c. determining whether the cancer patient is at risk for cancer-associated VTE based on the predictor and an operating-point threshold, wherein the machine learning classifier is trained by: i. receiving cohort data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; ii. generating a training dataset based on the received cohort data, the training dataset comprising the plurality of features for each subject in the cohort, the plurality of features comprising (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and iii. applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer- associated VTE, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy thatexceeds an accuracy threshold for the machine learning classifier; and determining the optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients.

37. The method of claim 36, further comprising administering an effective amount of anticoagulant therapy to the cancer patient predicted to be at risk for cancer- associated VTE based on the predictor and the operating-point threshold.

38. The method of claim 37, wherein the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.

39. The method of any one of claims 36-38, wherein the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

40. The method of any one of claims 36-39, wherein the cancer patient comprises metastatic sites of disease.

41. The method of claim 40, wherein the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

42. The method of any one of claims 36-41, wherein the machine learning technique is a random forest technique, and wherein the one or more machine learning models are random forest models.

43. The method of any one of claims 36-42, wherein the machine learning classifier is an ensemble learning random forest classifier.

44. The method of any one of claims 36-43, wherein the machine learning technique models survival outcomes with competing risks.

45. The method of any one of claims 36-44, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

46. The method of any one of claims 37-45, wherein the anticoagulant therapy comprises one or more of apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, or enoxaparin.

47. The method of any one of claims 36-46, wherein the cancer patient has a Khorana Score ≥ 2.

48. The method of any one of claims 36-47, wherein one or more of the plurality of features for the cancer patient are determined by assaying whole blood, serum or plasma.

49. The method of any one of claims 21-48, wherein one or more of the plurality of features for each subject in the cohort are determined by assaying whole blood, serum or plasma.

50. The method of any one of claims 21-49, wherein the cancer-associated VTE is pulmonary embolism or lower extremity deep vein thrombosis (DVT), optionally wherein lower extremity DVT includes thrombi involving a common iliac vein, an external iliac vein, a common femoral vein, a superficial femoral vein, a deep femoral vein, a popliteal vein, a peroneal vein, an anterior tibial vein, a posterior tibial vein, or a deep calf vein.

51. A machine learning system for training a machine learning classifier for estimating risk of cancer-associated venous thromboembolism (VTE) in cancer patients, the systemcomprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: receive data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; generate a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE in cancer patients; wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients.

52. The machine learning system of claim 51, wherein the machine learning technique is a random forest technique, and wherein the one or more machine learning models are random forest models.

53. The machine learning system of claim 51 or 52, wherein the machine learning classifier is an ensemble learning random forest classifier.

54. The machine learning system of any one of claims 51-53, wherein the machine learning technique models survival outcomes with competing risks.

55. The machine learning system of any one of claims 51-54, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

56. The machine learning system of any one of claims 51-55, wherein the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

57. The machine learning system of any one of claims 51-56, wherein the cancer patients comprise metastatic sites of disease.

58. The method of claim 57, wherein the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

59. The machine learning system of any one of claims 51-58, wherein the instructions further cause the processor to apply the machine learning classifier to data on a cancer patient to generate a predictor, and determine whether the cancer patient is at risk for cancer-associated VTE based on the predictor and the operating-point threshold.

60. The machine learning system of claim 59, wherein the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.

61. The machine learning system of any one of claims 59-60, wherein the cancer patient has a Khorana Score ≥ 2.

62. The machine learning system of any one of claims 51-61, wherein the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating- point threshold.

63. The machine learning system of claim 62, wherein the anticoagulant therapy comprises one or more of apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, or enoxaparin.

64. The machine learning system of any one of claims 59-63, wherein the cancer patient is chemotherapy -naive or has received / is receiving systemic chemotherapy.

65. The machine learning system of any one of claims 51-64, wherein the subjects in the cohort are chemotherapy-naive or have received systemic chemotherapy.

66. A computing system for estimating risk of cancer-associated venous thromboembolism (VTE) in a cancer patient, the computing system comprising a processor and a memory with instructions which, when executed by the processor, cause the processor to: receive patient data corresponding to a plurality of features for the cancer patient; apply a machine learning classifier to the patient data to generate a predictor; and determine whether the cancer patient is at risk for cancer-associated VTE based on the predictor and an operating-point threshold, wherein the classifier is trained by: receiving cohort data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; generating a training dataset based on the received cohort data, the training dataset comprising the plurality of features for each subject in the cohort, the plurality of features comprising (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; anddetermining the optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients.

67. The computing system of claim 66, wherein the machine learning technique is a random forest technique, and wherein the one or more machine learning models are random forest models.

68. The computing system of claim 66 or 67, wherein the machine learning classifier is an ensemble learning random forest classifier.

69. The computing system of any one of claims 66-68, wherein the machine learning technique models survival outcomes with competing risks.

70. The computing system of any one of claims 66-69, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

71. The computing system of any one of claims 66-70, wherein the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

72. The computing system of any one of claims 66-71, wherein the cancer patient comprises metastatic sites of disease.

73. The computing system of claim 72, wherein the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

74. The computing system of any one of claims 66-73, wherein the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold.

75. The computing system of claim 74, wherein the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.

76. The computing system of any one of claims 74-75, wherein the anticoagulant therapy comprises one or more of apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, or enoxaparin.

77. The computing system of any one of claims 66-76, wherein the cancer patient has a Khorana Score ≥ 2.

78. The computing system of any one of claims 66-77, wherein one or more of the plurality of features for the cancer patient are determined by assaying whole blood, plasma or serum.

79. A non-transitory computer-readable storage medium comprising instructions which, when executed by a processor of a machine learning system, configure the machine learning system to train a machine learning classifier to estimate risk of cancer-associated venous thromboembolism (VTE) in cancer patients, the instructions configured to cause the processor to: receive data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; generate a training dataset based on the received data, the training dataset comprising a plurality of features for each subject in the cohort, the plurality of features comprising (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and apply a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE in cancer patients; wherein applying the machine learning method comprises:applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining an optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients.

80. The computer-readable storage medium of claim 79, wherein the machine learning technique is a random forest technique, and wherein the one or more machine learning models are random forest models.

81. The computer-readable storage medium of claim 79 or 80, wherein the machine learning classifier is an ensemble learning random forest classifier.

82. The computer-readable storage medium of any one of claims 79-81, wherein the machine learning technique models survival outcomes with competing risks.

83. The computer-readable storage medium of any one of claims 79-82, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

84. The computer-readable storage medium of any one of claims 79-83, wherein the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia,primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

85. The computer-readable storage medium of any one of claims 79-84, wherein the cancer patient comprises metastatic sites of disease.

86. The computer-readable storage medium of claim 85, wherein the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

87. The computer-readable storage medium of any one of claims 79-86, wherein the instructions further cause the processor to apply the machine learning classifier to data on a cancer patient to generate a predictor, and determine whether the cancer patient is at risk for cancer-associated VTE based on the predictor and the operating-point threshold.

88. The computer-readable storage medium of claim 87, wherein the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.

89. The computer-readable storage medium of any one of claims 79-88, wherein the cancer patients have a Khorana Score ≥2.

90. The computer-readable storage medium of any one of claims 79-89, wherein the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold.

91. The computer-readable storage medium of claim 90, wherein the anticoagulant therapy comprises one or more of apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, or enoxaparin.

92. The computer-readable storage medium of any one of claims 79-91, wherein the subjects in the cohort are chemotherapy -naive or have received systemic chemotherapy.

93. The computer-readable storage medium of any one of claims 87-92, wherein the cancer patient is chemotherapy-naive or has received / is receiving systemic chemotherapy.

94. A non-transitory computer-readable storage medium comprising instructions which, when executed by a processor of a computing system, configure the computing system to estimate risk of cancer-associated venous thromboembolism (VTE) in a cancer patient, the instructions configured to cause the processor to:receive patient data corresponding to a plurality of features for the cancer patient; apply a machine learning classifier to the patient data to generate a predictor; and determine whether the cancer patient is at risk for cancer-associated VTE based on the predictor and an operating-point threshold, wherein the classifier is trained by: receiving cohort data on a cohort of subjects, the subjects in the cohort having a plurality of cancer types; generating a training dataset based on the received cohort data, the training dataset comprising the plurality of features for each subject in the cohort, the plurality of features comprising (i) mRNA and / or expression levels of one or more plasma proteins selected from among CNDP1, FR-gamma, SERPINB8, MARCO, CD84, SPON2, CD200R1, CDH6, LGALS8, LAT, and DAB, and (ii) cancer type; and applying a machine learning method to the training dataset to develop the machine learning classifier for estimating risk of cancer-associated VTE, wherein applying the machine learning method comprises: applying a machine learning technique to the training dataset; performing hyperparameter optimization to identify one or more machine learning models with an accuracy that exceeds an accuracy threshold for the machine learning classifier; and determining the optimal operating-point threshold based on optimization of sensitivity and specificity of the receiver operating characteristic (ROC) curves for the training dataset; wherein the machine learning classifier is configured to receive the plurality of features for cancer patients and generate predictors for risk of cancer-associated VTE in cancer patients.

95. The computer-readable storage medium of claim 94, wherein the machine learning technique is a random forest technique, and wherein the one or more machine learning models are random forest models.

96. The computer-readable storage medium of claim 94 or 95, wherein the machine learning classifier is an ensemble learning random forest classifier.

97. The computer-readable storage medium of any one of claims 94-96, wherein the machine learning technique models survival outcomes with competing risks.

98. The computer-readable storage medium of any one of claims 94-97, wherein performing the hyperparameter optimization comprises performing an exhaustive grid search technique.

99. The computer-readable storage medium of any one of claims 94-98, wherein the plurality of cancer types are selected from the group consisting of non-small cell lung cancer (NSCLC), breast cancer, bladder cancer, pancreatic cancer, melanoma, retinoblastoma, prostate cancer, gastric cancer, colon cancer, histiocytosis, germ cell tumor, endometrial cancer, small cell lung cancer (SCLC), soft tissue sarcoma, myeloma, gynecologic cancer, Gastrointestinal Stromal Tumor, ovarian cancer, mature B-Cell neoplasms, small bowel cancer, renal cell carcinoma, thyroid cancer, ampullary cancer, appendiceal cancer, sellar tumor, uterine sarcoma, bone cancer, non-melanoma skin cancer, cervical cancer, mesothelioma, glioma, thymic tumor, gastrointestinal neuroendocrine tumor, salivary gland cancer, sex cord stromal tumor, anal cancer, mature T and NK neoplasms, peritoneal cancer, Head and neck cancer, choroid plexus tumor, leukemia, primary CNS melanocytic tumors, Myelodysplastic Syndromes, Peripheral Nervous System, mastocytosis, Wilms tumor, lymphatic cancer, vaginal cancer, Hodgkin lymphoma, adrenocortical carcinoma, brain tumors, embryonal tumors and Non-Hodgkin lymphoma.

100. The computer-readable storage medium of any one of claims 94-99, wherein the cancer patient comprises metastatic sites of disease.

101. The computer-readable storage medium of claim 100, wherein the metastatic sites of disease comprise one or more of lungs, liver, bones, brain, adrenal gland, lymph nodes, and skin.

102. The computer-readable storage medium of any one of claims 94-101, wherein the instructions further cause the processor to recommend an anticoagulant therapy to the cancer patient predicted to be at risk for cancer-associated VTE based on the predictor and the operating-point threshold.

103. The computer-readable storage medium of claim 102, wherein the predictor comprises a cumulative incidence function (CIF) for cancer-associated VTE.

104. The computer-readable storage medium of claim 102 or 103, wherein the anticoagulant therapy comprises one or more of apixaban, betrixaban, dabigatran, edoxaban, fondaparinux, heparin, rivaroxaban, warfarin, Xa inhibitors, or enoxaparin.

105. The computer-readable storage medium of any one of claims 94-104, wherein the cancer patient has a Khoran Score ≥ 2.

106. The computer-readable storage medium of any one of claims 94-105, wherein one or more of the plurality of features for the cancer patient are determined by assaying whole blood, serum or plasma.

Citation Information

Patent Citations

  • Compositions and methods for identifying and modulating thrombotic conditions in a cancer patient

    WO2021119189A1