Gene signature predictive of melanoma metastasis and patient prognosis

By detecting the gene expression characteristics of ITGB3, PLAT, SPP1, GDF15 and IL8 in patients with cutaneous melanoma, the problem of overtreatment in existing diagnostic methods has been solved, enabling more accurate prediction of SLN metastasis and prognostic assessment, reducing unnecessary surgery and improving the treatment effect for patients with thin melanoma.

CN113039289BActive Publication Date: 2025-11-07SKYLINE DIAX LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980062825.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-12
Filing Date
2019-07-25
Publication Date
2025-11-07
Estimated Expiration
2039-07-25

AI Technical Summary

Technical Problem

Current methods for diagnosing and staging cutaneous melanoma suffer from overtreatment, particularly in patients with intermediate lesions, leading to unnecessary SLNB surgery complications and high costs, while potentially missing high-risk patients.

Method used

By detecting the expression characteristics of genes such as ITGB3, PLAT, SPP1, GDF15, and IL8, and combining this with gene characteristic analysis, individuals are classified as metastatic positive or negative SLN and as having a good or poor prognosis, which guides whether to perform SLNB surgery and selects treatment strategies.

Benefits of technology

It improves the accuracy of predicting metastatic SLN, reduces unnecessary SLNB surgery, lowers medical costs and complications, improves the identification rate of metastasis in thin melanoma patients, and improves prognosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113039289B_ABST
    Figure CN113039289B_ABST
Patent Text Reader

Abstract

The present invention provides gene signatures for classifying individuals having skin melanoma. The "SLN gene signatures" provided herein classify individuals based on prognosis and / or classify individuals as having metastasis-positive or -negative sentinel lymph nodes (SLNs). The "N-SLN gene signatures" provided herein classify individuals as having metastasis-positive or -negative non-sentinel lymph nodes (N-SLNs).
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention provides gene signatures for classifying individuals having skin melanoma. The "SLN gene signatures" provided herein classify individuals based on prognosis and / or classify individuals as having metastatic positive or negative sentinel lymph nodes (SLNs). The "N-SLN gene signatures" provided herein classify individuals as having metastatic positive or negative non-sentinel lymph nodes (N-SLNs). BACKGROUND

[0002] Skin melanoma is a malignancy that arises primarily from melanocytes located in the basal layer of the epidermis of the skin. Most lesions present signs described by the ABCDE rule: asymmetry, irregular borders, uneven color, diameter greater than 6 mm, and evolution, i.e., a tendency to change rapidly (Abbasi NR, Shaw HM, Rigel DS, et al. Revisiting the ABCD criteria for early diagnosis of melanoma. JAMA. 2004;292(22):2771-2776). These sites are usually asymptomatic, but can cause itching and / or bleeding, especially in later stages. Detection of suspicious lesions is usually accomplished by self-examination of the skin, which is recommended to be done routinely according to the ABCDE criteria or the "Ugly Duckling" sign (Grob J. The 'Ugly Duckling' Sign: Identification of the Common Characteristics of Nevi in an Individual as a Basis for Melanoma Screening. Arch Dermatol. 1998;134:103-104). When a physician subsequently formally diagnoses melanoma, it is important to determine the specific subtype, as there are multiple clinical and pathological types known. The most common form is cutaneous melanoma, a superficial spreading melanoma, which accounts for about 70% of cases - commonly occurring in people with fair skin. The severity of the condition depends largely on the ability of the melanoma cells to migrate out of the primary area. Therefore, it is crucial to assess whether the tumor is localized or has already spread to the lymph nodes or organs.

[0003] The staging of melanoma is critical for the prognosis of the patient and for deciding further monitoring and treatment strategies. This is also reflected in the significant contrast in the 5-year overall survival rates reported in a systematic literature review of 9 European countries: 95-100% (stage I), 65-92.8% (stage II), 41-71% (stage III) and 9-28% (stage IV). These differences depend largely on the ability of melanoma to metastasize, rather than on a local melanoma lesion. An accurate distinction between the different stages is very important and is usually based on the TNM system, i.e. according to the thickness of the primary tumor (T), the presence and / or extent of tumor cells to lymph nodes (N) and the presence of distant metastases to other organs (M). When assessing the degree of the primary tumor stage, the physician will consider the thickness of the tumor, but also other characteristics, such as the presence of ulceration and mitotic rate of the primary tumor cells. Normally, only in patients with high tumor stages, due to melanoma thickness and / or other variables, such as ulceration, lymph nodes and metastatic spread will be assessed. As understood in the art, prognosis refers to the prediction of the medical outcome for a patient. For example, an individual can be classified as having a poor prognosis or a good prognosis. The prognosis of a melanoma patient indicates, for example: the likelihood of long-term survival, overall survival, progression-free survival, prediction of recurrence and remission of disease and disease progression.

[0004] Currently, the determination of the presence of metastatic lesions in the SLN by the method of sentinel lymph node biopsy (SLNB) is a widely applied method for accurate patient staging and prediction of prognosis. Since the incorporation of lymphatic mapping by SLNB surgery in the early 1990s, significant progress has been made in the treatment of patients with cutaneous melanoma (Morton DL, Wen DR, Wong JH, et al. Technical details of intraoperative lymphatic mapping for early stage melanoma. Arch Surg. 1992; 127(4):392-399). The surgical technique has been improved and a dual modality was implemented, using blue dye and radioactive tracer with a gamma probe detection intraoperatively. In addition, the pathological evaluation was improved with the use of SLN serial sectioning and immunohistochemistry. This enabled a better identification of the first draining lymph node or a group of lymph nodes located near the tumor (i.e. the SLN), which are thus possible sites of metastatic disease. This procedure is also referred to as “sentinel lymph node mapping”.

[0005] The MSLT-1 study also showed a significant impact of SLN positivity, which demonstrated a difference in 5-year survival rates between patients with tumor-positive SLNs and patients with tumor-negative SLNs, 72.3% and 90.2%, respectively (Morton DL, Thompson JF, Cochran AJ, et al. Sentinel-Node Biopsy or Nodal Observation in Melanoma. N Engl J Med. 2006;355(13): 1307-1317). According to the American Joint Committee on Cancer (AJCC) Melanoma Guidelines, 8th edition, the SLNB procedure is recommended for patients with cutaneous melanoma >0.8 mm (Gershenwald JE, Scolyer RA, Hess KR. Melanoma Staging: Evidence-Based Changes in the American Joint Committee on Cancer Eighth Edition Cancer Staging Manual. CA Cancer J Clin. 2017;67(6):472-492). For this group of patients, SLNB surgery is usually performed and further treatment depends on the degree of metastasis. Within the borderline group, SLNB can be considered, especially if the melanoma exhibits other adverse prognostic parameters. For patients with melanoma thickness <0.8 mm, standard treatment is generally considered adequate and the use of SLNB is not recommended. Standard treatment includes local resection of the primary melanoma with wide margins, i.e. surgical removal of the tumor. As used herein, “resection” is understood to mean surgical removal of malignant tissue having the characteristics of melanoma from a human patient. According to one embodiment, resection is understood to mean removal of malignant tissue such that the presence of remaining malignant tissue within said patient is not detectable with existing methods.

[0006] The rate of SLNB positivity varies greatly and is largely dependent on known prognostic factors of the primary tumor. In patients with clinical stage I or II, the percentage of SLN metastasis is 15-30%, while in thin melanomas it is 5.2%. The latest edition of the Melanoma Expert Panel Staging Guidelines points out the clinical relevance of the Tl melanoma subclassification of 0.8 mm. This is based on a trend detected in some Tl melanoma survival studies that there is a potential clinical cutoff at the 0.7 to 0.8 mm region. However, long-term follow-up of patients after SLNB surgery showed that patients with initially tumor-free sentinel lymph nodes developed regional lymph node recurrence. This information can be used to calculate the SLBN test performance, which has an overall false negative rate of 12.5%. Recently, Morton et al. reported that in intermediate-thickness melanomas with a SLNB positivity rate of 16.0%, 4.8% of the detections were false negatives, which recurred within a 10-year follow-up period. The SLNB positivity rate for thick melanomas was 32.9%, with a false negative rate of 10.3%.

[0007] SLNB is not only a method for the potential staging of cutaneous melanoma, but also part of the treatment, which can or can not be necessary depending on the metastatic classification of the SLN. SLNB surgery carries complications and is expensive for the patient. Therefore, this surgery is only performed in a group of patients who are considered to be at higher risk of metastatic spread (relative to the vast majority of low-risk lesions). The risk of metastasis can be assessed by evaluating clinical-pathological factors, including tumor infiltration depth (referred to as Breslow depth) and tumor surface ulceration. Ulcerated tumors and melanomas that grow vertically into the skin are associated with a higher risk of adverse outcomes. For example, SLN biopsy is not recommended for Tla thin melanomas, is "possibly recommended" for Tlb thin melanoma patients, is recommended for T2 and T3 intermediate-thickness melanoma patients, and is "possibly recommended" for T4 thick melanoma patients.

[0008] While the use of clinical-pathological variables generally allows the identification of high-risk patients at the extremes of the tumor spectrum, this approach is not accurate for the diagnosis of intermediate lesions. In addition, there are exceptions to the high-risk or low-risk groups. For example, 5% of "thin" melanomas (<0.8 mm infiltration depth) are known to locally metastasize, but according to the standard clinical-pathological variables, they are generally classified as low risk. To better distinguish high-risk lesions from the 95% biologically inert lesions, additional histological variables such as the mitotic rate of the tumor (mitoses / mm 2) and other molecular methods, such as fluorescence in situ hybridization (FISH). Unfortunately, these techniques are only partially successful, resulting in up to 95% of SLNBs being negative. Thus, the current diagnostic criteria result in overtreatment, leading to unnecessary SLNB surgery for the majority of patients and associated side effects that could have been avoided. Conversely, a proportion of patients who are not designated for SLNB according to current clinicopathological factors can still have metastatic SLNs or develop distant metastases later. Thus, although SLNB surgery is accurate in identifying SLN-positive lymph nodes in resected tissue, selecting patients eligible for SLNB surgery remains a challenge. It is an object of the present invention to classify individuals according to their risk of having metastatic-positive SLNs and thus the need for SLNB surgery. It is a further object of the present invention to predict the prognosis of individuals having primary cutaneous melanoma. This information is useful, for example, for determining the optimal treatment strategy. SUMMARY

[0010] The present invention provides a method for classifying an individual having primary cutaneous melanoma, comprising determining a gene expression signature in a sample from said individual, wherein said gene expression signature comprises three or more of the following genes: ITGB3, PLAT, SPP1, GDF15 and IL8. Preferably, wherein said gene expression signature comprises three or more of the following genes: ITGB3, PLAT, GDF15 and IL8, more preferably wherein said gene expression signature comprises ITGB3, PLAT, GDF15 and IL8. Preferably, wherein said gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, SERPINE2 and TGFBR1, more preferably wherein said gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2 and TGFBR1, more preferably wherein said gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1 and TGFBR1. Also preferably, wherein said gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, PRKCB, SERPINE2, ADAM12, LGALS1 and TGFBR1.

[0011] Also provided is a method for determining a treatment and / or diagnostic check schedule for an individual having skin melanoma, comprising determining the expression level of three or more of the following genes: ITGB3, PLAT, SPP1, GDF15 and IL8 in a sample from the individual, and determining the treatment and / or diagnostic schedule from the expression levels. Preferably, wherein the gene expression signature comprises three or more of the following genes: ITGB3, PLAT, GDF15 and IL8, more preferably wherein the gene expression signature comprises ITGB3, PLAT, GDF15 and IL8. Preferably, wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, SERPINE2 and TGFBR1, more preferably wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2 and TGFBR1, more preferably wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1 and TGFBR1. Also preferably, wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, PRKCB, SERPINE2, ADAM12, LGALS1 and TGFBR1.

[0012] Also provided is a method for predicting prognosis in an individual having a primary skin melanoma, comprising determining a gene expression signature in a sample from said individual, wherein said gene expression signature comprises three or more of the following genes: ITGB3, PLAT, SPP1, GDF15 and IL8. Preferably, wherein said gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, SERPINE2 and TGFBR1, more preferably wherein said gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2 and TGFBR1, more preferably wherein said gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1 and TGFBR1. Also preferably, wherein said gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, PRKCB, SERPINE2, ADAM12, LGALS1 and TGFBR1.

[0013] In one aspect, the individual is classified as having a metastasis-positive SLN or as having a metastasis-negative SLN. In one aspect, the prognosis of the individual depends on the gene expression level. Preferably, the individual can be classified as having a poor prognosis or as having a good prognosis. The individual undergoing SLNB can be selected based on said classification and / or expression level. The individual classified as having a metastasis-positive SLN or as having a poor prognosis is treated by SLNB and / or adjuvant therapy.

[0014] The present application also provides a method for classifying an individual having a primary skin melanoma, comprising determining a gene expression signature in a sample from said individual, wherein said gene expression signature comprises at least one of the following genes: KRT14, SPP1, FN1 and LOXL3.

[0015] Further provided is a method of treating an individual having a primary skin melanoma, comprising

[0016] - determining a gene expression signature in a sample from said individual, wherein said gene expression signature comprises three or more of the following genes: ITGB3, PLAT, SPP1, GDF15 and IL8.

[0017] - classifying said individual as having a metastasis-positive SLN and / or a poor prognosis based on the gene expression signature, and

[0018] - treating the individual by performing a SLNB and / or providing a cancer treatment to the individual.

[0019] Preferably, wherein the gene expression signature comprises three or more of the following genes: ITGB3, PLAT, GDF15, and IL8, more preferably wherein the gene expression signature comprises ITGB3, PLAT, GDF15, and IL8. Preferably, wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, SERPINE2, and TGFBR1, more preferably wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1, more preferably wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1. Also preferably, wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1.

[0020] The present application also provides a method of treating an individual having a primary cutaneous melanoma, comprising

[0021] - determining a gene expression signature in a sample from the individual, wherein the gene expression signature comprises at least one of the following genes: KRT14, SPP1, FN1, and LOXL3,

[0022] - classifying the individual as high risk of having a positive N-SLN of metastasis based on the gene expression signature, and

[0023] - treating the individual by performing a complete lymph node dissection and / or providing a cancer treatment to the individual.

[0024] Also provided is a method for analyzing a gene signature of an individual having a primary cutaneous melanoma, the method comprising

[0025] - extracting RNA from a primary cutaneous melanoma lesion of the individual;

[0026] - reverse transcribing RNA transcripts of at least three of the following genes: ITGB3, PLAT, SPP1, GDF15, and IL8, to produce cDNA of the RNA transcripts; and

[0027] - amplifying the cDNA to generate amplicons from the cDNA to determine the expression level of the RNA transcript.

[0028] The present application also provides a method for analyzing the gene signature of an individual having a primary cutaneous melanoma, the method comprising

[0029] - extracting RNA from a primary cutaneous melanoma lesion of the individual;

[0030] - reverse transcribing the RNA transcript of at least one of the following genes: KRT14, SPP1, FN1 and LOXL3 to generate a cDNA of the RNA transcript; and

[0031] - amplifying the cDNA to generate amplicons from the cDNA to determine the expression level of the RNA transcript.

[0032] Further provided is a kit for classifying an individual having a primary cutaneous melanoma, the kit comprising primer pairs for amplifying:

[0033] a) three or more of the following genes: ITGB3, PLAT, SPP1, GDF15 and IL8; and / or

[0034] b) at least one of the following genes: KRT14, SPP1, FN1 and LOXL3, and optionally

[0035] c) at least one reference gene.

[0036] Preferably, the kit comprises primer pairs for amplifying three or more of the following genes: ITGB3, PLAT, GDF15, and IL8, more preferably the kit comprises primer pairs for amplifying ITGB3, PLAT, GDF15, and IL8. Preferably, wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, SERPINE2, and TGFBR1, more preferably wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1, more preferably wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1. Also preferably, wherein the gene expression signature comprises three or more of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 : Average ROC curve of the logistic regression classifier trained in DLCV: 1) ITGB3, PLAT, SPP1, GDF15, and IL8 gene signature (molecular model), 2) clinicopathological variables (age and Breslow depth), 3) combination of ITGB3, PLAT, SPP1, GDF15, and IL8 gene signature and clinicopathological variables. The x-axis represents the false positive discovery rate (i.e. 1 - specificity), the y-axis represents the true discovery rate (i.e. sensitivity).

[0038] Figure 2 : ROC curve of the ITLP score and the SLN gene signature (called "logistic regression model") in the whole cohort of 770 patients. The x-axis represents the false positive discovery rate (i.e. 1 - specificity), the y-axis represents the true discovery rate (i.e. sensitivity).

[0039] Figure 3 : Boxplot of the area under the ROC curve of different subsets of ITGB3, PLAT, SPP1, GDF15, and IL8 genes, the full set of 5 ITGB3, PLAT, SPP1, GDF15, and IL8 genes, and the ITLP signature.

[0040] Figure 4The average ROC curve of the logistic regression classifier trained in DLCV, for: 1) gene expression, 2) clinicopathological variables, and 3) a combination of gene expression and clinicopathological variables. The x-axis represents the false positive rate (i.e., 1-specificity), and the y-axis represents the true positive rate (i.e., sensitivity).

[0041] Figure 5 Box plots of the area under the ROC curve for each gene subset and the full set of 4 genes.

[0042] Figure 6 Overall performance comparison: CL vs. GE vs. GECL. ROC curves of logistic regression classifiers trained in DLCV for: 1) gene expression, 2) clinicopathological variables, and 3) a combination of gene expression and clinicopathological variables.

[0043] Figure 7 : NPV vs. SLNBRR. The negative predictive value (NPV) of the logistic regression classifier trained in DLCV for sentinel lymph node reduction rate (SLNB-RR) in the following aspects: 1) gene expression, 2) clinicopathological variables, and 3) combination of gene expression and clinicopathological variables.

[0044] Figure 8 Gene subset-AUC box plot. Box plot of the area under the ROC curve (AUC) of the logistic regression classifier, trained on subsets of 2, 3, 4, 5, 6, 7, and 8 genes in the entire cohort. Detailed Implementation

[0045] This invention partially provides methods, kits, genetic characterization, and methods for detecting such genetic characteristics to perform analysis of primary cutaneous melanoma tumor tissue samples. In one aspect, the invention provides "SLN genetic characterization." SLN genetic characterization can classify individuals with primary cutaneous melanoma, and in particular, it can classify the risk of an individual having metastatic positive SLN and / or a poor prognosis. This risk assessment is very useful when physicians and patients decide whether SLNB (similar SLN-associated neoplasm) and / or alternative treatment strategies are necessary. This assessment is also useful when selecting patients for inclusion in clinical trials.

[0046] As used in this article, the SLN is the first lymph node (or first lymph node group) to receive lymphatic drainage from the tumor, and is the first lymph node (or first lymph node group) from which the cancer may have spread. An N-SLN is a lymph node that is not the first lymph node to receive lymphatic drainage from the tumor. Such an N-SLN is typically a lymph node in the same nodal basin or a lymph node very close to the SLN.

[0047] In some embodiments, the gene signature can classify the risk of a metastatic positive SLN. In some embodiments, the methods disclosed herein classify an individual as having a metastatic positive SLN or as having a metastatic negative SLN. In some embodiments, the gene signature can classify the prognosis of an individual. As used herein, prognosis refers to a prediction of medical outcome and can be based on measures such as overall survival, melanoma-specific survival, recurrence-free survival, relapse-free survival, and distant relapse-free survival.

[0048] One advantage of the SLN gene signature is that it can reduce the number of surgeries for patients classified as having a metastatic negative SLN (and / or classified as having a good prognosis). In particular, patients with intermediate lesions can have already undergone a SLNB procedure, but there is a high likelihood that the SLN is actually metastatic negative. Accurate classification of these patients with the SLN gene signature can avoid the need for a SLNB procedure and can be used in place of SLNB as the current standard of care for intermediate lesions. Reducing unnecessary SLNBs can lower overall medical costs and reduce the number of patients who experience complications from SLN resection. In addition, classifying an individual as having a metastatic positive or negative SLN also provides prognostic information that can be used to determine a treatment or diagnostic work schedule.

[0049] Surprisingly, the SLN gene signature can more accurately predict prognosis than a standard SLN biopsy (see Example 7). While not wishing to be bound by theory, one possible explanation for the improved prognostic ability of the gene signature over SLN biopsy relates to technical limitations in performing such biopsies (e.g., determining the correct lymph node to biopsy, limitations in tumor cell detection, human error in processing / classifying the sample). In addition to or alternatively, the disclosed gene signature can predict SLN metastasis at a stage prior to when the SLN can be detected in a biopsy (e.g., the tumor has metastasized and tumor cells are in route to the SLN). In this regard, the SLN signature can be used in place of SLNB as a standard for inclusion in clinical trials and / or other treatments.

[0050] Another advantage of the SLN gene signature is that it can identify patients with thin melanoma thickness who, based on clinical parameters, can not qualify for SLNB under current standards, but are at high risk for a metastatic positive SLN based on the gene signature. In particular, this gene signature will greatly improve the identification of metastatic positive SLNs in patients with thin (<0.8 mm) melanomas who currently do not qualify for SLNB procedures according to guidelines. Early detection and treatment of such patients will improve the progression-free survival and overall survival of this patient subpopulation.

[0051] The embodiments disclosed herein demonstrate that a gene expression signature comprising one or more of the following genes (i.e., the SLN gene signature): ITGB3, PLAT, SPP1, GDF15, and IL8, can be used for individual classification and prognosis, in particular classifying SLNs as metastasis positive or negative. Accordingly, in one aspect, the present invention provides a gene signature comprising one or more, preferably two or more, more preferably three or more, of the following genes: ITGB3, PLAT, SPP1, GDF15, and IL8. Suitable gene signatures include the following combinations: ITGB3 and PLAT; ITGB3 and SPP1; ITGB3 and GDF15; ITGB3 and IL8; PLAT and SPP1; PLAT and GDF15; PLAT and IL8; SPP1 and GDF15; SPP1 and IL8; GDF15 and IL8. In some embodiments, the gene signature comprises one or more of ITGB3, PLAT, and SPP1, GDF15, and IL8. In some embodiments, the SLN gene signature comprises three or more of the following genes: ITGB3, PLAT, SPP1, GDF15, and IL8. In some embodiments, the SLN gene signature comprises four or more of the following genes: ITGB3, PLAT, SPP1, GDF15, and IL8. In some embodiments, the SLN gene signature comprises all of the following genes: ITGB3, PLAT, SPP1, GDF15, and IL8. In some embodiments, the gene expression signature comprises three or more of the following genes: ITGB3, PLAT, GDF15, and IL8, more preferably wherein the gene expression signature comprises ITGB3, PLAT, GDF15, and IL8. Preferably, wherein the gene expression signature comprises GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, SERPINE2, and TGFBR1, more preferably wherein the gene expression signature comprises GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1, more preferably wherein the gene expression signature comprises GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1.

[0052] In some embodiments, the gene signature comprises at least three, at least four, or at least five of the following: ITGB3, PLAT, GDF15, SPP1, and IL8. Preferably, the gene signature comprises ITGB3, PLAT, GDF15, and IL8.

[0053] In some embodiments, the gene signature includes at least three, at least four, at least five, at least six, at least seven, at least eight, or all of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1. The inventors have also demonstrated that a gene signature lacking ADIPOQ also has similar performance. Thus, in some embodiments, the gene signature includes at least three, at least four, at least five, at least six, at least seven, or all of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1. In some embodiments, the gene signature includes:

[0054] - ITGB3, PLAT, GDF15, IL8, MLANA, and one or both of LOXL4 and SERPINE2;

[0055] - ITGB3, PLAT, GDF15, IL8, MLANA, and one or both of SERPINE2 and TGFBR1;

[0056] - ITGB3, PLAT, GDF15, IL8, and one or both of MLANA and TGFBR1;

[0057] - ITGB3, PLAT, GDF15, IL8, and one or both of TGFBR1 and SERPINE2;

[0058] - ITGB3, PLAT, GDF15, IL8, SERPINE2, and one or both of LOXL4 and TGFBR1;

[0059] - ITGB3, PLAT, GDF15, IL8, LOXL4;

[0060] - ITGB3, PLAT, GDF15, IL8, SERPINE2:

[0061] - ITGB3, PLAT, GDF15, IL8, TGFBR1; or

[0062] - ITGB3, PLAT, GDF15, IL8, MLANA.

[0063] In some embodiments, the gene signature comprises at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, or all of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1. In some embodiments, the gene signature comprises at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or all of the following genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1.

[0064] In some embodiments, the gene signature consists of the genes described above. As will be appreciated by one of skill in the art, when the SLN gene signature consists of the genes described above, the method for performing the analysis can include measuring the expression of other genes (e.g., for normalization), but only using the gene signature to classify the individual.

[0065] The ITGB3 gene encodes integrin beta-3. A representative Homo sapiens mRNA sequence can be found at NM_000212.2 (June 17, 2018) in the NCBI database.

[0066] The PLAT gene encodes tissue plasminogen activator. A representative Homo sapiens mRNA sequence can be found at NM_001319189.1 (July 1, 2018) in the NCBI database.

[0067] The SPP1 gene encodes secreted phosphoprotein 1. A representative Homo sapiens mRNA sequence can be found at NM_001040058.1 (June 24, 2018) in the NCBI database.

[0068] The GDF15 gene encodes growth differentiation factor 15. A representative Homo sapiens mRNA sequence can be found at NM_004864.3 (June 17, 2018) in the NCBI database.

[0069] The IL8 gene encodes interleukin 8. A representative Homo sapiens mRNA sequence can be found at AF043337.1 (February 1, 2001) in the NCBI database.

[0070] The MLANA gene encodes melan-A. A representative Homo sapiens mRNA sequence can be found at NM_005511 (October 20, 2018) in the NCBI database.

[0071] The LOXL4 gene encodes lysyl oxidase-like 4. A typical Homo sapiens mRNA sequence can be found at NM_032211 (November 22, 2018) in the NCBI database.

[0072] The ADIPOQ gene encodes adiponectin, C1Q and collagen domain containing. A typical Homo sapiens mRNA sequence can be found at NM_004797 (December 2, 2018) in the NCBI database.

[0073] The PRKCB gene encodes protein kinase C beta. A typical Homo sapiens mRNA sequence can be found at NM_212535 (November 12, 2018) in the NCBI database.

[0074] The SERPINE2 gene encodes serpin family E member 2. A typical Homo sapiens mRNA sequence can be found at NM_006216 (November 17, 2018) in the NCBI database.

[0075] The ADAM12 gene encodes ADAM metallopeptidase domain 12. A typical Homo sapiens mRNA sequence can be found at NM_003474 (August 5, 2018) in the NCBI database.

[0076] The LGALS1 gene encodes galectin 1. A typical Homo sapiens mRNA sequence can be found at NM_002305 (November 22, 2018) in the NCBI database.

[0077] The TGFBR1 gene encodes transforming growth factor beta receptor 1. A typical Homo sapiens mRNA sequence can be found at NM_004612 (October 28, 2018) in the NCBI database.

[0078] The present disclosure also provides methods of classifying an individual comprising determining the SLN gene signature in a sample. In some embodiments, the individual can be classified as having a metastatic positive SLN or a metastatic negative SLN. In another embodiment, the individual can be classified as having a good or poor prognosis. A gene signature associated with SLN metastasis has been previously reported (Meves et al., J Clinical Oncology, 2015 33:2509-2516). This algorithm uses clinical pathological variables such as age, Breslow depth, and ulceration in combination with the primary melanoma gene expression of four genes, ITGB3, LAMB1, PLAT, and TP53 to predict metastasis of the SLN. As shown in Table 1, the presently described SLN gene signature is superior to the previously reported signature. Figure 3

[0079] ​In one aspect, the present application provides a "N-SLN gene signature." The N-SLN gene signature can classify individuals having primary cutaneous melanoma, and in particular, the gene signature can classify individuals at risk of having metastatic non-sentinel lymph node (N-SLN). This risk assessment is useful for doctors and patients in deciding treatment options and judging patient prognosis.

[0080] In some embodiments, the N-SLN gene signature can classify the risk of metastatic N-SLN. Individuals can be classified into metastatic N-SLN and non-metastatic N-SLN. Tumor cell invasion into distal lymph nodes is an indicator of poor prognosis and suggests the use of more aggressive treatment. Early detection and treatment is expected to improve patient outcomes.

[0081] The embodiments disclosed herein demonstrate that a gene expression signature comprising one or more of the following genes (i.e., the N-SLN gene signature): KRT14, SPP1, FN1, LOXL3, can be used to classify individuals and predict prognosis, and in particular, determine the risk of metastatic N-SLN. Accordingly, in one aspect, the present application provides a gene signature comprising at least one of the following genes: KRT14, SPP1, FN1, LOXL3. In some embodiments, the gene signature comprises at least two or at least three of the following genes: KRT14, SPP1, FN1, LOXL3. In some embodiments, the N-SLN gene signature comprises or consists of KRT14, SPP1, FN1, LOXL3. In some embodiments, the gene signature consists of the above genes. As will be understood by one of skill in the art, when the N-SLN gene signature consists of the above genes, the method used to perform the analysis can include measuring the expression of other genes (e.g., for normalization), but only the genes of the gene signature are used to classify the individual. In some embodiments, the N-SLN gene signature is determined in individuals having a recurrence / reappearance of cutaneous melanoma and / or who have undergone a SLN biopsy.

[0082] The KRT14 gene encodes keratin 14. A typical Homo sapiens mRNA sequence can be found at NM_000526.4 (June 17, 2018) in the NCBI database.

[0083] The FN1 gene encodes fibronectin 1. A typical Homo sapiens mRNA sequence can be found at NM_001306129.1 (June 3, 2018) in the NCBI database.

[0084] The LOXL3 gene encodes lysyl oxidase-like 3. A typical Homo sapiens mRNA sequence can be found at NM_001289165.1 (June 30, 2018) in the NCBI database.

[0085] The present disclosure also provides methods of classifying an individual comprising determining the N-SLN gene signature in a sample. In some embodiments, the individual can be classified as having metastatic-positive N-SLN or metastatic-negative N-SLN. In some embodiments, methods are provided for determining both the SLN gene signature and the N-SLN gene signature.

[0086] The gene signature analysis disclosed herein can be performed in any individual, including mammals and humans, although humans are preferred. In some embodiments, the individual has been diagnosed with T1-T3 cutaneous melanoma. In some embodiments, the individual has not undergone a SLN biopsy of the primary melanoma, particularly when the gene signature is the SLN gene signature. The gene signature is particularly useful for classifying individuals who are young, have a high mitotic rate (e.g., 2 / mm 2 The above classification is particularly useful for individuals who are young, have a high mitotic rate (e.g., 2 / mm

[0087] The gene expression signature can be used to predict the risk or likelihood of tumor cell metastasis to the SLN or N-SLN. As is well understood by those skilled in the art, the classification of an individual refers to the likelihood or "risk" of developing metastasis, not that 100% of all patients predicted to be at risk will have detectable metastasis (referred to as sensitivity or positive percent agreement), nor that 0% of all patients predicted to not have metastasis will not have metastasis (referred to as specificity or negative percent agreement). As disclosed in the Examples, the SLN and N-SLN gene expression signatures exhibit high levels of performance in both sensitivity and specificity. As disclosed in the Examples, the SLN gene signature is more predictive of the prognosis of an individual with melanoma than the standard-of-care SLN biopsy. Thus, the present disclosure demonstrates that the gene expression signature can be used to predict the prognosis of an individual.

[0088] As is known to the skilled artisan, the metastatic burden, as measured by the amount of metastatic disease, can vary between individuals. In some embodiments, metastasis refers to the presence of clusters of tumor cells, and does not include lymph nodes containing only isolated or rare tumor cells. In some embodiments, metastasis refers to the presence of clusters of cells at least 0.1 mm in diameter, with or without extracapsular extension.

[0089] Obtaining a suitable sample for determining gene expression is within the skill of the artisan. Suitable samples include a primary cutaneous melanoma lesion biopsy. Such biopsies include excised lesions (e.g., wide excision of a tumor). The sample can be processed or preserved by any method known in the art that is compatible with gene expression profiling. For example, the sample can be a formalin-fixed paraffin-embedded primary cutaneous melanoma lesion biopsy, as well as a frozen sample.

[0090] Preferably, the sample is an RNA-containing sample. General methods for mRNA extraction are well known in the art and are disclosed in standard textbooks of molecular biology, including Ausubel et al. (1997) Current Protocols of Molecular Biology, John Wiley and Sons. For example, Rupp and Locker (1987) Lab Invest 56:A67 and De Andres et al., BioTechniques 18:42044 (1995) disclose methods for extracting RNA from paraffin-embedded tissue. In particular, RNA isolation can be performed using purification kits, buffer sets and proteases from commercial manufacturers (e.g., Qiagen) according to the manufacturer's instructions (Qiagen, Valencia, CA). For example, Qiagen RNeasy mini columns can be used to isolate total RNA from cultured cells. Numerous RNA isolation kits are commercially available and can be used in the methods of the application.

[0091] The methods disclosed herein include determining gene expression signatures. In particular, the methods include determining gene expression levels. Gene expression levels can be determined by detecting nucleic acid or protein expression levels. Preferably, mRNA expression levels are determined. In some embodiments, nucleic acids or proteins are purified from a sample and gene expression is measured by nucleic acid or protein expression analysis. Protein expression levels can be determined by any method known in the art, including ELISA, immunocytochemistry, flow cytometry, Western blotting, proteomics, and mass spectrometry.

[0092] Preferably, nucleic acid expression levels are determined. Nucleic acid expression levels can be determined by any method known in the art, including RT-PCR, quantitative PCR, Northern blotting, gene sequencing (especially RNA sequencing), and gene expression profiling techniques. Representative methods for sequencing-based gene expression analysis include serial analysis of gene expression (SAGE), and gene expression analysis by massively parallel signature sequencing (MPSS).

[0093] Preferably, the nucleic acid is RNA, such as mRNA or mRNA precursor. As understood by one skilled in the art, the determined RNA expression level can be detected directly or determined indirectly, e.g., by first generating cDNA and / or by amplifying the RNA / cDNA. In some embodiments, a primary melanoma sample is obtained; RNA is extracted from the tissue sample; then RNA transcripts of genes of interest (e.g., biomarkers and housekeeping genes) are reverse transcribed to generate cDNA of the RNA transcripts; and the cDNA is amplified to generate amplicons from the cDNA to determine the expression level of the RNA transcripts.

[0094] In some embodiments, gene expression can be determined by NanoString gene expression analysis. NanoString is a multiplexed method for detecting gene expression that does not require transcription or amplification, providing a direct method for detecting mRNA. NanoString and aspects thereof are described in Geiss et al., "Direct multiplexed measurement of gene expression with color-coded probes pairs," Nature Biotechnology 26, 317-325 (2008); Geiss et al., "Direct multiplexed analysis of cellular genes by short cDNA sequences," PNAS 105, 20167-20172 (2008); and Sturm et al., "Quantitative comparison of DNA and RNA isolation methods for gene expression analysis," BMC Genomics 9, 422 (2008).

[0095] The expression level need not be an absolute value, but can be a normalized expression value or a relative value. For example, the expression level can be normalized according to housekeeping or reference gene expression. Such genes include ABCF1, ACTB, ALAS1, CLTC, G6PD, GAPDH, GUSB, HPRT1, LDHA, PGK1, POLR1B, POLR2A, RPL19, RPLPO, SDHA, TBP, and TUBB.

[0096] Standardization is also useful when expression is determined based on microarray data. Standardization allows for correction of variation within the microarray and between samples so that data from different chips can be analyzed simultaneously. Robust multi-array analysis (RMA) algorithms can be used to pre-process probe set data into gene expression levels for all samples. (Irizarry R A, et al., Biostatistics (2003) and Irizarry R A, et al., Nucleic Acids Res. (2003)). In addition, the default pre-processing algorithm of Affymetrix (MAS 5.0) can also be used. Other methods of standardizing expression data are described in US20060136145.

[0097] In some embodiments, real-time PCR (i.e., quantitative PCR or qPCR) is used to determine expression levels. In real-time PCR (qPCR), the reaction is characterized by the point at which amplification of the target is first detected during the cycle process, rather than the amount of target accumulated after a fixed number of cycles. This point at which the signal is first detected is referred to as the threshold cycle (Ct).

[0098] In some embodiments, the expression of the gene signatures is quantified relative to one another by normalizing to the expression of a housekeeping gene by subtracting the Ct of the signature genes from the average Ct of the housekeeping genes. In some embodiments, these ΔCt values are then combined in an algorithm with the patient’s age and Breslow depth of the melanoma lesion to calculate a prediction of SLN metastasis. In some embodiments, the housekeeping genes used for normalization are ACTB, RPLP0, and RPL8. However, other housekeeping genes can be used. The ratios of the gene expression signals can then be combined in an algorithm with clinical variables to calculate a prediction of the patient’s outcome of SLNB. The result is expressed as a binary classification (negative or positive). A “negative” result indicates that the individual has a lower risk of SLN metastasis, or in other words, a good prognosis, while a “positive” result indicates that the individual has a higher risk of SLN metastasis, or in other words, a worse prognosis.

[0099] The methods described herein classify individuals based on a gene expression signature. In some embodiments, differential expression of one or more genes of the signature in an individual is indicative of the individual having a risk of metastasis, or more precisely, of the individual’s prognosis. As used herein, “differential expression” means that the expression level measured in the subject is significantly different from a reference value. The reference value can be a single value or a range of numbers. Determining an appropriate reference value is within the ability of the skilled person. In some embodiments, the reference value is a pre-determined value. In some embodiments, the reference value is the average of expression values in a particular class of patients. For example, the reference value can be the average expression value of a class of patients who have clinically confirmed SLN metastasis (or for the N-SLN signature, patients who have clinically confirmed N-SLN metastasis). The reference value can also be in the form of or derived from an equation. It is within the ability of the skilled person to determine whether a patient’s expression level is “significantly” different from the reference value.

[0100] In exemplary embodiments, the reference value is determined from a cohort of melanoma patients who underwent SLNB as described in the examples. Data from similar studies can also be used by the skilled person.

[0101] The strength of the correlation between the expression level of a differentially expressed gene and a particular patient response class can be determined by a statistical test of significance. For example, a Chi-squared test can be used to assign a Chi-squared value to each differentially expressed marker indicative of the strength of the correlation between the expression of that marker and a particular patient response class. Similarly, both the T-statistic measure and the Wilkins measure provide a value or score indicative of the strength of the correlation between the expression of a marker and its particular patient response class. In addition, SAM or PAM analysis tools can be used to determine the strength of the correlation.

[0102] In some embodiments, the gene expression profile from an individual is compared to a reference expression profile to determine whether the gene expression profile from the individual is sufficiently similar to the reference. Alternatively, the gene expression profile from an individual is compared to a plurality of reference expression profiles to select the reference expression profile that is most similar to the gene expression profile in the individual. Any known method in the art for comparing two or more data sets to detect similarities between them can be used to compare the gene expression profile from an individual to a reference expression profile.

[0103] In machine learning and statistics, classification refers to the identification of the set of classes to which a new observation belongs based on a training data set containing observations (or instances) of members of the known classes. Algorithms that implement classification, particularly in specific applications, are referred to as classifiers. There are many classifiers currently available, with linear or non-linear classifier boundaries, such as but not limited to ClaNC, nearest mean classifier, weighted vote, simple Bayesian classifier, linear discriminant analysis (LDA), quadratic discriminant analysis (QDA), support vector machine (SVM), or k-nearest neighbor (k-nn) classifier. In preferred embodiments, a logistic regression classifier is used. Exemplary embodiments implementing a logistic regression classifier are described in the Examples.

[0104] As understood by one of skill in the art, training of the gene expression profile can be performed to improve sensitivity or specificity. Sensitivity refers to the proportion of actual positives that are correctly identified, and high sensitivity is desired to avoid false negatives (e.g., a patient classified as metastasis negative who is actually positive). Specificity refers to the proportion of actual negatives that are correctly identified, and high specificity is desired to avoid false positives (e.g., a patient classified as metastasis positive who is actually negative). Preferably, the classifier is trained with high sensitivity so as to identify individuals with metastasis.

[0105] In some embodiments, the method for classifying an individual further utilizes the age of the individual and / or the Breslow depth of the tumor. Optionally, ulceration and / or mitotic rate can be measured. Breslow depth is measured from the top of the epidermal granular layer (or, if the surface is ulcerated, from the base of the ulcer) to the deepest invasive cell of the broad base of the tumor (dermis / subcutis). Ulceration refers to the sloughing of dead tissue and is thought to reflect rapid tumor growth, leading to central cell death in melanoma. Mitotic rate can be measured by examining the excised tumor and counting the number of cells showing mitosis. The higher the mitotic count, the greater the likelihood of metastasis of the tumor. In specific embodiments, a combined model is used, including a gene signature comprising GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, SERPINE2, and TGFBR1, and the clinical variables of age and Breslow depth. In some embodiments, the gene signature further comprises AIDPOQ.

[0106] The present application also provides kits for determining the gene expression signatures disclosed herein. In some embodiments, the kit comprises primer pairs for performing qPCR on the gene signatures disclosed herein. In some embodiments, the kit comprises primer pairs for performing qPCR on two or more, preferably three or more, of the following genes: ITGB3, PLAT, SPP1, GDF15, and IL8; and / or primer pairs for performing qPCR on one or more of the following genes: KRT14, SPP1, FN1, LOXL3. In some embodiments, the kit comprises primer pairs for a housekeeping gene, such as ACTB, RPLP0, and RPL8. In some embodiments, the kit further comprises one or more of the following: a DNA polymerase, deoxynucleotide triphosphates, a buffer, and Mg 2+In some embodiments, the kit comprises primer pairs for amplifying three or more of the following genes: ITGB3, PLAT, GDF15, and IL8, more preferably a kit comprising primer pairs for amplifying ITGB3, PLAT, GDF15, and IL8. Preferably, wherein the gene expression signature comprises GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, SERPINE2, and TGFBR1, more preferably wherein the gene expression signature comprises GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1, more preferably wherein the gene expression signature comprises GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1. In some embodiments, the kit comprises a control nucleic acid for one or more, preferably each, primer pair. Preferably, the control nucleic acid is a cDNA, and more preferably the cDNA corresponds to a sequence spanning at least one intron / exon boundary of the corresponding gene. Such cDNA is helpful in distinguishing gene expression from contamination from the genome. In some embodiments, one or more of the primers of a primer pair is chemically modified. Such modified primers include fluorescent or radiolabeled primers.

[0107] The results of the gene expression analysis disclosed herein can be used to determine a diagnostic workup schedule. For example, individuals classified as SLN metastasis positive or poor prognosis can be subjected to SLNB. In some embodiments, individuals predicted to be SLN positive or more precisely predicted to be poor prognosis are administered immunotherapy. Subsequent SLNB readout can serve as a measure of response to immunotherapy.

[0108] Depending on the gene expression signature, an appropriate treatment regimen can be determined. As used herein, the term "treatment" refers to reversing, alleviating, delaying the onset of, or inhibiting the progression of melanoma or one or more symptoms thereof. In some embodiments, individuals classified as having a metastatic positive SLN or more precisely poor prognosis can be treated using SLNB. The location of the SLN can be determined based on the location of the melanoma and / or using methods such as "SLN mapping" as known to those skilled in the art and described herein.

[0109] Metastatic N-SLNs can be treated by surgical intervention, such as surgical lymph node dissection. Metastatic SLNs can be treated by complete lymph node dissection and / or other methods of treating melanoma. In some embodiments, the individual is administered a cancer treatment. In some embodiments, the individual is administered an "adjuvant therapy." As used herein, adjuvant therapy refers to the delivery of one or more drugs to a patient after surgical removal of one or more cancerous tumors, where in the tumor, all detectable and resectable disease (e.g., cancer) has been removed from the patient, but there is still a statistical risk of recurrence. Adjuvant therapy helps to reduce the likelihood or severity of recurrence of the disease.

[0110] Known melanoma therapies that can be indicated based on the gene expression signature include:

[0111] chemotherapy; e.g., dacarbazine (DTIC), temozolomide (Temodal), carboplatin (Paraplatin, Paraplatin AQ), paclitaxel (Taxol), cisplatin (Platinol AQ), and vinblastine (Velbe);

[0112] targeted therapies: e.g., BRAF inhibitors (Vemurafenib (Zelboraf) and Dabrafenib (Tafinlar)) and MEK inhibitors (Cobimetinib (Cotellic) and Trametinib (Mekinist));

[0113] radiotherapy;

[0114] immunotherapy; e.g., cytokines (e.g., interferon alpha-2b or interleukin-2), immune checkpoint inhibitors (e.g., Ipilimumab (Yervoy), Nivolumab (Opdivo), Pembrolizumab (Keytruda)), or oncolytic immunotherapy.

[0115] Suitable drug treatments can be administered by any appropriate route. Suitable routes include oral, rectal, nasal, topical (including buccal and sublingual), vaginal and parenteral (including subcutaneous, intramuscular, intravenous, intradermal, intrathecal and epidural).

[0116] As used herein, the terms "comprising" (and any form of comprising, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has") "including" (and any form of including, such as "includes" and "include") or "containing" (and any form of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.

[0117] As used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a component" includes a plurality of such components. Similarly, for example, a reference to "the individual" is a reference to one or more individuals and equivalents thereof.

[0118] The word "about" or "approximately" when used with a numerical value (about 10, approximately 10) preferably means that the value can be the given value or more or less 1% of the value.

[0119] The present application is further explained in the following examples. These examples do not limit the scope of the present application but merely serve to clarify the present application.

[0120] Example

[0121] Example 1 : Gene signature predicting SLN metastatic status

[0122] We collected a cohort of 813 melanoma patients who underwent SLN biopsy at a tertiary care center. The outcome of interest was histologically determined metastasis on the SLN and SLN biopsy. We measured the expression of 29 pro-metastatic stromal response genes in the primary melanoma diagnostic biopsy by polymerase chain reaction. Regularized logistic regression was applied to the clinicopathological variables and molecular data in a double loop cross-validation (DLCV) training-validation scheme.

[0123] Clinical data set

[0124] Among the above clinical variables, only 6 showed significant differences between SLNB positive and SLNB negative:

[0125] • age,

[0126] • Breslow depth

[0127] • ulceration

[0128] • mitotic rate

[0129] • Clark level

[0130] • vascular-lymphatic invasion

[0131] Among these 6 variables, we decided to consider only age, Breslow depth, ulceration and mitotic rate, and not Clark level and vascular-lymphatic invasion, because the latter are not always available and their quality can vary depending on the medical personnel performing the SLNB.

[0132] We quantified gene expression by ΔCt instead of by copy number as in Meves et al. (2015), because using copy number did not significantly differ in performance from our signature, but just added an experimental burden. The KRT14 background correction was also decreased compared to Meves et al. (2015). The so-called ITLP normalization was also abandoned, because it is based on a special over-control of the measured expression, and it does indeed require many arbitrary parameters (mainly in the form of thresholds).

[0133] In patients with positive sentinel lymph nodes, biopsy positive for metastatic load, the measurement of metastatic disease can vary significantly, depending on the amount of metastatic disease:

[0134] Amount Definition 1 Isolated tumor cells or rare tumor cells 2 Cell cluster diameter < 0.1 mm 3 Cell clusters >= 0.1 mm with or without extracapsular extension 4 Cell clusters >= 0.1 mm with or without extracapsular extension 9 Unknown

[0135] Whether samples with cell clusters with a diameter of less than 0.1 mm should be considered metastatic positive is still a matter of debate from a clinical point of view. Therefore, we decided to exclude 43 out of 813 patients with a metastatic disease amount of 1 or 2 from the training set of the classifier and to evaluate the performance of the classifier on this group separately. The following 29 genes were measured:

[0136] KRT14, MLANA, MITF, ITGB3, PLAT, LAMB1, TP53, AGRN, THBS2, PTK2, SPP1, COL4A1, CDKN1A, CDKN2A, PLOD3, GDF15, FN1, TNC, THBS1, CTGF, LOXL1, LOXL3, ITGA5, ITGA3, ITGA2, CSRC, CXCL1, IL8, LAMB.

[0137] Performance measures

[0138] To evaluate the performance of the classifier, contingency tables were constructed. From these contingency tables two criteria were derived, the PPA (positive percent agreement) and the NPA (negative percent agreement): which are defined as:

[0139] PPA = 100 TP / (TP + FN) and NPA = 100 TN / (TN + FP)

[0140] where TP denotes the number of true positives, FN the number of false negatives, TN the number of true negatives and FP the number of false positives. The PPA and NPA criteria correspond to the sensitivity and specificity, respectively. However, since the comparison is made with respect to a non-gold standard reference (FISH data), the terms PPA and NPA are used.

[0141] To optimize the training of the classifier, a performance criterion is needed which should be maximized (or minimized in case of an error criterion). For this purpose, the average of PPA and NPA will be used:

[0142]

[0143] where p denotes the performance. p = 50 indicates a random performance and p = 100 a perfect classification.

[0144] Other performance measures:

[0145] • Negative predictive value.

[0146] • Positive predictive value

[0147] • Accuracy

[0148] • Balanced accuracy

[0149] • Log odds ratio of negative outcome

[0150] • Log odds ratio of positive outcome

[0151] • Area under the ROC curve.

[0152] Classifier

[0153] All classifiers were trained on 29 gene expression levels, clinical pathologic variables, and both.

[0154] Classifier: Penalized maximum likelihood logistic regression

[0155] We used a logistic regression classifier implemented in the R language glmnet. Maximum likelihood estimation of parameters was performed with an L1 norm penalty term (LASSO regularization) to obtain a parsimonious representation.

[0156] Double loop cross-validation of classifiers

[0157] Wessels et al. [Wessels et al. Bioinformatics, vol. 21, no. 19, 2005, pp. 3755-3762] describe a general framework for building diagnostic classifiers from high-throughput data through double loop cross-validation (DLCV). The DLCV exercise enables the developer to estimate / predict the performance of the classifier (in terms of the generalized error) for future application to data independent of the training data set. The method was adapted to feature selection with forward filtering, combined with t-statistics as a criterion to evaluate individual genes and different classifiers. The training and validation procedure was performed with 3-fold cross-validation with 100 replicates in the outer (validation) loop and 10-fold cross-validation in the inner loop. In the inner loop, the algorithm learns the optimal parameter lambda for LASSO regularization. At all points, the data partitioning was stratified by class prior probabilities.

[0158] The double loop cross-validation method can be described by the following steps:

[0159] 1. For each replicate, the data is divided into 3 parts (different parts for each replicate).

[0160] 2. For each fold, the inner loop (training set) uses 2 parts; the third part is used in the outer loop to validate (validation set).

[0161] 3. On the training set data, 10-fold cross-validation is performed to estimate the optimal lambda for the LASSO penalty term (construction of learning curves).

[0162] 4. Then, the classifier is trained on the full training set with the optimal λ.

[0163] 5. Finally, the performance of the classifier is evaluated on the validation set.

[0164] 6. After all repetitions are completed, a final classifier is created using all samples with the average optimal λ. The classifier is trained with the resulting average n. This classifier will then be applied to the external validation set.

[0165] In general, the dataset is imbalanced with respect to the class prior. Therefore, the balanced accuracy has better classification performance than the accuracy due to the consideration of the class prior. In each iteration, in the inner loop, we use the Brier score as the performance criterion.

[0166] Gene expression-based classifier (GE)

[0167] The parameters of the gene expression-based logistic classifier model are as follows:

[0168] (intercept) 0.78199268 ITGB3 -0.19497568 PLAT -0.12305874 SPP1 -0.00831690 GDF15 -0.04554275 IL8 -0.03306434

[0169] Table 1 describes the performance of the final classifier trained on the entire 770 patient cohort for four different operating points: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log-likelihood ratio of negative test result set to 0.25 in training. If the coefficient is positive, the higher the value means the higher the risk. If the coefficient is negative, the lower the value means the lower the risk. Variables with larger (absolute) coefficients have a larger contribution.

[0170] Table 2 describes the performance of the classifier trained in DLCV for four different operating points, averaged over 100 repetitions: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log-likelihood ratio of negative test result set to 0.25 in training.

[0171] Clinical variable-based classifier (CL)

[0172] The parameters of the clinical variable-based logistic classifier model are as follows:

[0173] (intercept) -1.86716485 Age -0.01919991 Breslow depth 0.72079160 Ulceration - yes 0.14301462

[0174] Table 3 The parameter "age" is entered in years, "Breslow depth" in mm. Ulceration is a Boolean variable (yes / no). Table describes the performance of the final classifiers trained on the entire 770 patient cohort for four different operating points: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log-likelihood ratio of negative test result set to 0.25 in training. Table 4 describes the performance of the classifiers trained in DLCV for four different operating points, averaged over 100 replicates: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log-likelihood ratio of negative test result set to 0.25 in training.

[0175] Classifier based on gene expression and clinical variables (GECL)

[0176] The parameters of the logistic classifier model based on clinical variables and gene expression are as follows:

[0177] (intercept) 0.59189841 Age -0.01356918 Breslow depth 0.4722686 ITGB3 -0.15443565 PLAT -0.13580062 SPP1 -0.00778917 GDF15 -0.05924340 IL8 -.003781148

[0178] Table 5 describes the performance of the final classifiers trained on the entire 770 patient cohort for four different operating points: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log-likelihood ratio of negative test result set to 0.25 in training.

[0179] Table 6 describes the performance of the classifiers trained in DLCV for four different operating points, averaged over 100 replicates: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log-likelihood ratio of negative test result set to 0.25 in training.

[0180] Figure 1 ROC curves are described for the logistic regression classifiers trained in DLCV: 1) gene expression, 2) clinical-pathological variables, 3) gene expression combined with clinical-pathological variables.

[0181] Table 7 describes the average performance of classifiers trained in DLCV on: 1) gene expression ("GE", i.e. ITGB3, PLAT, SPP1, GDF15 and IL8 gene signature; 2) clinicopathological variables ("CL", i.e. age and Breslow depth); 3) combination of gene expression and clinicopathological variables ("GECL"). Three different operating points are considered: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training.

[0182] Example 2 Comparative Example

[0183] Classifier based on ITLP score (ITGB3, LAB1, PLAT and TP53)

[0184] Table 8 describes the performance of the ITLP score on the whole 770 patients cohort.

[0185] Classifier based on ITLP score and clinical variables

[0186] The parameters of the logistic classifier model based on clinical variables and gene expression are as follows:

[0187] (intercept) -2.07660948 Age -0.0111771 Breslow depth 0.52425831 ITLP 0.50016404

[0188] Table 9 describes the performance of the final classifiers trained on the whole 770 patients cohort for four different operating points: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log-likelihood ratio of negative test result set to 0.25 in training.

[0189] Table 10 describes the performance of classifiers trained in DLCV on four different operating points, averaged on 100 replicates: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log-likelihood ratio of negative test result set to 0.25 in training.

[0190] Example 3: Comparative analysis

[0191] ITLP vs. ITGB3, PLAT, GDF15, SPP1 and IL8 gene signature.

[0192] Figure 2ROC curves for the ITLP signature and the ITGB3, PLAT, GDF15, SPP1, and IL8 gene signature (referred to as logistic regression in the figures) are depicted. The ITGB3, PLAT, GDF15, SPP1, and IL8 gene signature is clearly superior to the ITLP signature.

[0193] Example 4: Performance of gene subsets

[0194] The previous example used 5 genes: ITGB3, PLAT, GDF15, SPP1, and IL8 as the gene signature. We investigated the performance of all possible 2, 3, and 4 gene subsets. The number of subsets of a particular size was chosen from the total number of genes in the following signatures as follows: 10 subsets from signatures with two genes, 10 subsets from signatures with three genes, 5 subsets from signatures with four genes, one signature containing all 5 genes. We evaluated the performance as the area under the ROC curve and compared it to the ITLP signature. The AUC (or range of AUCs) for the ITLP was 0.68, all 2 (gene) subsets were 0.72-0.75, all 3 gene subsets were 0.74-0.77, all 4 gene subsets were 0.76-0.77, and the 5 gene signature was 0.77. This is also shown in Figure 3 Figure 2. Thus, all gene signatures containing at least two of the following genes: ITGB3, PLAT, GDF15, SPP1, and IL8 performed better than the ITLP signature.

[0195] Example 5: Performance on 43 low volume metastatic disease samples

[0196] Patients with low volume metastatic disease (volume 1 and volume 2) were first excluded from the cohort used to train the classifier. Whether samples with cell clusters less than 0.1 mm in diameter should be considered metastatic positive is still a matter of debate from a clinical perspective. In this study, 43 patients were initially excluded from the analysis because they had volume 1 or volume 2. These patients were classified using the ITGB3, PLAT, GDF15, SPP1, and IL8 classifier, 29 were positive and 14 were negative.

[0197] Example 6: Misclassification analysis

[0198] False negatives. The positive samples that were misclassified mostly came from patients with thin melanomas (less than 2 mm), no ulceration, and no vascular lymphatic invasion. In other words, these patients had a very low a priori risk of metastasis. There were a small number of samples that were misclassified in 100 repetitions of the algorithm.

[0199] False Positives. The majority of the false positive misclassifications were from patients with thick melanomas (greater than 2 mm), with ulceration, with vascular lymphatic invasion. In other words, these patients had a high a priori risk of metastasis. There were a small number of samples that were misclassified in 100 repetitions of the algorithm.

[0200] Classifier output distribution

[0201] The distribution of the predicted probabilities is unimodal, not Gaussian, with a long right tail. The threshold used to select the operating point falls near the mean of the distribution. The estimated probability does not exceed 0.6.

[0202] Example 7: Relationship of SLN gene signature to prognosis

[0203] Kaplan-Meier survival estimates were generated for the three survival types for the SLN classifier including genes ITGB3, PLAT, SPP1, GDF15, and IL8 (referred to as "GECL" in the examples) (Tables 19-21(a)), SLNB status (Tables 19-21(a)), and the combination of both (Tables 22-24(a)). As is known in the literature, the SLNB positive status is associated with a lower survival rate. Notably, the GECL model provides a stronger separation in the survival estimates, which is also evidenced by a larger hazard ratio (Table 25(a)). Furthermore, in the multivariate analysis, the GECL classifier also has a larger hazard ratio and a more significant p-value (see Table 26(a)). Of particular note is the very high melanoma-specific survival rate for the GECL negative group, with a survival estimate of 0.966 at 160 months (Table 19(a)).

[0204] Likewise, Kaplan Meier survival estimates were generated for the three survival types for the SLN classifier including genes GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADA12, LGALS1, and TGFBR1 (referred to as "GECL" in the examples) (Tables 19-21), SLNB status (Tables 19-21), and the combination of both (Tables 22-24). Tables 19-26(b) describe the training results with an NPV setting of 0.97, while Tables 19-26(c) describe the training results with an NPV setting of 0.98. See Example 8 for further discussion of this classifier.

[0205] Combining the results of the GECL classifier output with the SLNB status provides four groups: true positives (TP), false positives (FP), false negatives (FN), and true negatives (TN). As mentioned previously, FN cases are rare and FP cases represent a substantial proportion. It is notable that the survival estimate for the false positive group (GECL positive, SLNB status negative) is very similar to the true negative group. This suggests that the GECL classifier output is a superior means of prognosis for melanoma patients than the SLNB status.

[0206] The Kaplan-Meier method is used to estimate survival probabilities at multiple time intervals. The log-rank test is a non-parametric test used to compare survival curves between two or more groups.

[0207] The hazard ratio (HR) is defined as the ratio of the risk of an outcome in one group / (the risk of an outcome in another group) over a given time interval. A hazard ratio of 1 indicates a lack of association, a hazard ratio greater than 1 indicates an increased risk, and a hazard ratio less than 1 indicates a smaller risk. Hazard ratios are used to express relative differences between two groups.

[0208] Example 8: Improvement of the SLN classifier

[0209] Four groups of genes were measured across the entire 855 cohort, for a total of 109 unique genes. However, due to unavailability of samples at the time of analysis or insufficient RNA, gene expression could not be measured for certain samples. Therefore, this finding was performed on a cohort of 754 patients rather than 770 patients. The cohort of the finding did not include patients with low amounts of metastatic disease.

[0210] For each classifier, we report:

[0211] • Performance of the final classifier trained on the entire cohort

[0212] • Average performance of the cross-validation (double loop cross-validation with 100 repetitions, outer loop folded 3 times, inner loop folded 10 times). The performance in the three folds is concatenated to cover the entire cohort.

[0213] For each classifier, performance is computed at 4 different operating points:

[0214] • Max bACC: Maximum balanced accuracy

[0215] • SE eqSP: Sensitivity equal to specificity

[0216] • NPV97: Negative predictive value of 97% in training

[0217] • NPV98: Negative predictive value of 98% in training

[0218] Clinical Pathology Model (CL)

[0219] The logistic classifier model parameters based on clinical pathology variables are as follows:

[0220] Feature Parameter

[0221] (intercept) -2.0547083

[0222] Age -0.0112913

[0223] The logistic classifier model parameters based on clinical pathology variables are as follows:

[0224] Feature Parameter

[0225] breslow_depth 0.6116335

[0226] Vascular Lymphatic Invasion - Yes 0.1205238

[0227] Table 27 describes the performance of the final classifiers trained on the entire 754 patient cohort for four different operating points: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) NPV98, NPV set to 0.98 in training. If the coefficient is positive, higher values mean higher risk. If the coefficient is negative, lower values mean lower risk. Variables with larger (absolute) coefficients have a larger contribution.

[0228] Table 28 describes the performance of the classifiers trained in DLCV for four different operating points, averaged over 100 repeats: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) NPV98, NPV set to 0.98 in training.

[0229] Gene Expression Model (GE)

[0230]

[0231] Table 29 describes the performance of the final classifiers trained on the entire 754 patient cohort for four different operating points: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) NPV98, NPV set to 0.98 in training. If the coefficient is positive, higher values mean higher risk. If the coefficient is negative, lower values mean lower risk. Variables with larger (absolute) coefficients have a larger contribution.

[0232] Table 30 describes the performance of classifiers trained in DLCV at four different operating points, averaged over 100 replicates: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) NPV98, NPV set to 0.98 in training.

[0233] Clinicopathological combined with gene expression model (GECL)

[0234]

[0235] Table 31 describes the performance of the final classifiers trained on the entire 754 patient cohort for four different operating points: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) NPV98, NPV set to 0.98 in training. If the coefficient is positive, higher values mean higher risk. If the coefficient is negative, lower values mean lower risk. Variables with larger (absolute) coefficients have a larger contribution.

[0236] Table 32 describes the performance of classifiers trained in DLCV at four different operating points, averaged over 100 replicates: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) NPV98, NPV set to 0.98 in training.

[0237] Figure 6 ROC curves are described for logistic regression classifiers trained in DLCV: 1) gene expression, 2) clinicopathological variables, 3) gene expression and clinicopathological variables combined, and Figure 7 Negative predictive values (NPV) are described for logistic regression classifiers trained in DLCV compared to sentinel lymph node reduction rates (SLNB RR): 1) gene expression, 2) clinicopathological variables, 3) gene expression and clinicopathological variables combined.

[0238] Comparison of different operating points (OP): CL vs. CE vs. GECL

[0239] Table 33 describes the average performance of classifiers trained in DLCV on: 1) gene expression ("GE", i.e., the GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1 gene signature); 2) clinicopathological variables ("CL", i.e., age, Breslow depth, and presence of vascular lymphatic invasion); 3) a combination of gene expression and clinicopathological variables ("GECL": i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). For the operating point of maximum balanced accuracy (max bACC).

[0240] Table 34 describes the average performance of classifiers trained in DLCV on: 1) gene expression ("GE", i.e., the GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1 gene signature); 2) clinicopathological variables ("CL", i.e., age, Breslow depth, and presence of vascular lymphatic invasion); 3) a combination of gene expression and clinicopathological variables ("GECL": i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). For the operating point where sensitivity equals specificity (SEeqSP).

[0241] Table 35 describes the average performance of classifiers trained in DLCV on: 1) gene expression ("GE", i.e., the GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1 gene signature); 2) clinicopathological variables ("CL", i.e., age, Breslow depth, and presence of vascular lymphatic invasion); 3) a combination of gene expression and clinicopathological variables ("GECL": i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). For the operating point where NPV is set to 0.97 in training (NPV97).

[0242] Table 36 describes the average performance of classifiers trained in DLCV on the following: 1) gene expression ("GE", i.e., the GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1 gene signature); 2) clinicopathological variables ("CL", i.e., age, Breslow depth, and presence of vascular lymphatic invasion); 3) combination of gene expression and clinicopathological variables ("GECL": i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). For operating points with NPV set to 0.98 in training (NPV98).

[0243] Performance by T stage

[0244] Table 37 describes the average performance of classifiers trained in DLCV on the following: 1) gene expression ("GE", i.e., the GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1 gene signature); 2) clinicopathological variables ("CL", i.e., age, Breslow depth, and presence of vascular lymphatic invasion); 3) combination of gene expression and clinicopathological variables ("GECL": i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). For operating points with NPV set to 0.97 in training (NPV97).

[0245] Table 38 describes the average performance of classifiers trained in DLCV on the following: 1) gene expression ("GE", i.e., the GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1 gene signature); 2) clinicopathological variables ("CL", i.e., age, Breslow depth, and presence of vascular lymphatic invasion); 3) combination of gene expression and clinicopathological variables ("GECL": i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). For operating points with NPV set to 0.97 in training (NPV97).

[0246] Table 39 describes the average performance of classifiers trained in DLCV on the following: 1) gene expression ("GE", i.e., the GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1 gene signature); 2) clinicopathological variables ("CL", i.e., age, Breslow depth, and presence of vascular lymphatic invasion); 3) combination of gene expression and clinicopathological variables ("GECL": i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). For operating points with NPV set to 0.97 in training (NPV97).

[0247] Table 40 describes the average performance of classifiers trained in DLCV on the following: 1) gene expression ("GE", i.e., the GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1 gene signature); 2) clinicopathological variables ("CL", i.e., age, Breslow depth, and presence of vascular lymphatic invasion); 3) combination of gene expression and clinicopathological variables ("GECL": i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). For operating points with NPV set to 0.98 in training (NPV98).

[0248] Table 41 describes the average performance of classifiers trained in DLCV in separating by T stage on gene expression ("GE", i.e., GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1). The operating point with NPV set to 0.98 in training (NPV98).

[0249] Table 42 describes the average performance of classifiers trained in DLCV in separating by T stage on a combination of gene expression and clinicopathological variables ("GECL", i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). The operating point with NPV set to 0.98 in training (NPV98).

[0250] Performance in separating by clinical stage

[0251] Table 43 describes the average performance of classifiers trained in DLCV in separating by clinical stage on clinicopathological variables ("CL" i.e., age, Breslow depth, and whether vascular lymphatic invasion is present). The operating point with NPV set to 0.97 in training (NPV97).

[0252] Table 44 describes the average performance of classifiers trained in DLCV in separating by clinical stage on gene expression ("GE", i.e., GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1). The operating point with NPV set to 0.97 in training (NPV97).

[0253] Table 45 describes the average performance of classifiers trained in DLCV in separating by clinical stage on a combination of gene expression and clinicopathological variables ("GECL", i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). The operating point with NPV set to 0.97 in training (NPV97).

[0254] Table 46 describes the average performance of classifiers trained in DLCV in separating by clinical stage on clinicopathological variables ("CL" i.e., age, Breslow depth, and whether vascular lymphatic invasion is present). The operating point with NPV set to 0.98 in training (NPV98).

[0255] Table 47 describes the average performance of classifiers trained in DLCV to separate by clinical stage on gene expression ("GE", i.e., GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, PRKCB, SERPINE2, ADAM12, LGALS1, and TGFBR1). The operating point with NPV set to 0.98 in training (NPV98) is used.

[0256] Table 48 describes the average performance of classifiers trained in DLCV to separate by clinical stage on a combination of gene expression and clinicopathological variables ("GECL", i.e., age, Breslow depth, GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1). The operating point with NPV set to 0.98 in training (NPV98) is used.

[0257] Gene subsets

[0258] Figure 8 Boxplots describing the area under the ROC curve (AUC) of logistic regression classifiers with 2, 3, 4, 5, 6, 7, 8 gene subsets selected from GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, ADIPOQ, SERPINE2, and TGFBR1, trained on the whole cohort.

[0259] Table 49 describes the number of subsets of a particular size selected from the total number of genes in each feature, and the performance of the minimum and maximum area under the ROC curve.

[0260] Example 9: Non-Sentinel Lymph Node (N-SLN) distribution.

[0261] The standard treatment for clinically node-negative melanoma patients with positive sentinel lymph node (SLN) is complete lymph node dissection (CLND) and removal of non-sentinel lymph nodes (N-SLN). Immediate CLND after SLN biopsy improves local disease control and a randomized clinical trial showed that early surgical treatment of low volume SLN-positive disease reduces long-term sequelae (such as lymphedema) compared to surgery at the time of lymph node recurrence. In addition, SLN and N-SLN metastasis are adverse prognostic factors used to select patients for adjuvant therapy. However, for patients enrolled in MSLT-II, CLND did not improve survival and had a higher complication rate than SLN surgery alone. New methods are needed to identify patients who can benefit from CLND, i.e. those at risk of N-SLN regional metastasis, to improve the selection of CLND patients. Here, we designed three classifiers based on gene expression (KRT14, SPP1, FN1, and LOXL3), clinicopathological variables (age, number of positive SLNs, maximum SLN size, maximum Breslow depth, and maximum mitotic rate), and both to predict the status of N-SLN, i.e. the presence or absence of metastasis. These classifiers can be used to select which patients should undergo CLND surgery.

[0262] The method used was essentially the same as shown in Example 1. The same 29 genes were evaluated. The parameters of the logistic regression classifier model based on gene expression were as follows.

[0263] Intercept -0.94837538 KRT14_3 0,06687021 SPP1_3 -0.02696531 FN1_3 -0,20982959 LOXL_3.3 -0,01722991

[0264] Model performance

[0265] The number of patients in the entire cohort was 140.

[0266] Table 11 describes the performance of the classifiers trained in DLCV at four different operating points, averaged over 100 replicates: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log-likelihood ratio of negative test result set to 0.25 in training.

[0267] Table 12 describes the performance of the classifiers trained in DLCV at four different operating points, averaged over 100 replicates: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log-likelihood ratio of negative test result set to 0.25 in training.

[0268] Classifier based on clinical variables (CL)

[0269] The parameters of the logistic classifier model based on clinical variables were as follows:

[0270]

[0271] Table 13 describes the performance of the final classifiers trained on the entire 140 patient cohort classifier for four different operating points: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log likelihood ratio of negative test result set to 0.25 in training.

[0272] Table 14 describes the performance of the classifiers trained in DLCV for four different operating points, averaged over 100 replicates: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log likelihood ratio of negative test result set to 0.25 in training.

[0273] Classifier based on gene expression and clinical variables (GECL)

[0274] Logistic regression parameters

[0275]

[0276]

[0277] Parameters of the logistic classifier model based on gene expression.

[0278] Table 15 describes the performance of the final classifiers trained on the entire 140 patient cohort classifier for four different operating points: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log likelihood ratio of negative test result set to 0.25 in training.

[0279] Table 16 describes the performance of the classifiers trained in DLCV for four different operating points, averaged over 100 replicates: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training, 4) LRNn025, log likelihood ratio of negative test result set to 0.25 in training.

[0280] Figure 4The average ROC curves of the logistic regression classifiers trained in DLCV are depicted for 1) gene expression, 2) clinicopathological variables, 3) gene expression combined with clinicopathological variables. The x-axis represents the false positive discovery rate (i.e. 1 - specificity), the y-axis represents the true discovery rate (i.e. sensitivity).

[0281] Figure 5 Boxplots of the area under the ROC curve for each gene subset and the area of the full set of 4 genes.

[0282] Table 17 depicts the average performance of the classifiers trained in DLCV for 1) gene expression, 2) clinicopathological variables, 3) gene expression combined with clinicopathological variables. Three different operating points are considered: 1) max bACC: maximum balanced accuracy, 2) SEeqSP, sensitivity equal to specificity, 3) NPV97, NPV set to 0.97 in training.

[0283] The N-SLN analyzer gene signature includes 4 genes: KRT14, SPP1, FN1 and LOXL3.

[0284] We also investigated all possible 2, 3 and 4 gene subsets.

[0285] Marker 2-gene subset 3-gene subset Genes in signature NSLNB 6 4 4

[0286] Number of subsets of a specific size that can be chosen from the total number of genes in each feature.

[0287] We evaluated the performance in terms of the area under the ROC curve (see table and figure)

[0288]

[0289] Range of the area under the curve (AUC) for different gene subsets of the full set of 4 genes.

[0290] Figure 5 Boxplots of the area under the ROC curve for different gene subsets and the full set of 4 genes are depicted.

[0291] We describe the discovery, design and development of the N-SLN analyzer. We have shown that the N-SLN analyzer can be used to select which patients should undergo CLND surgery.

[0292] The performance of the gene expression-based classifiers is of interest because currently (i) no method is available to select patients who would benefit from a CLND operation, and (ii) the clinicopathological variables used by the classifiers can not always be available clinically.

[0293] Example 10

[0294] Preselection based on Breslow depth (BD <= 2 and BD > 2)

[0295] Breslow depth (BD) is an important clinicopathological variable to characterize primary cutaneous melanoma, thin melanomas (BD <= 2 mm) have different molecular and physiological characteristics than thick melanomas (BD > 2 mm). Therefore, it makes sense to select two different operating points for thin and thick melanomas. For thin melanomas with Breslow depth <= 2 mm (561-sample cohort), we continue using the classifier described in the previous section (7.8) and select the operating point such that the NPV is 0.97 in the training:

[0296] Table 18a describes the average performance of the classifier trained in DLCV for 1) gene expression, 2) clinicopathological variables, 3) gene expression combined with clinicopathological variables. The selected operating point is NPV97, i.e. the NPV is set to 0.97 in the training.

[0297] For thick melanomas with BD > 2 mm (209-sample cohort), we use the same classifier but with a different operating point, i.e. we select a point that maximizes balanced accuracy:

[0298] Table 18b describes the average performance of the classifier trained in DLCV for 1) gene expression, 2) clinicopathological variables, 3) gene expression combined with clinicopathological variables. The selected operating point is chosen such that it maximizes balanced accuracy.

[0299] The performance for Breslow depth > 2, while lower than for Breslow depth <= 2 mm, is still acceptable because the classifier achieves a 90% NPV in the subpopulation with a prior probability of 35% of being SLN-positive: the probability of being SLN-positive when tested and having a negative outcome decreases from 35% to 10%.

[0300] Method for determining whether a subject is classified as lymph node positive or lymph node negative

[0301] The classification method is exemplified with the fictitious data (see table) using 2 genes for simplicity, predicting whether a sample will be labeled as lymph node positive or lymph node negative by the classifier (the method / model is the same for the SLN-analyzer and the N-SLN-analyzer, only the parameters and the kind of genes and clinical variables are different). The table describes the binary class label. The log odds ratio and the probability are calculated by equations 1 and 2, respectively. The output label is assigned by comparing the estimated probability with the cutoff value θ: if the estimated probability is greater than or equal to θ, the sample is classified as lymph node positive; if the estimated probability is less than θ, the sample is classified as lymph node negative.

[0302]

[0303] Table. Model parameters β0, β1, β2 for genes x, y, gene expression data ΔC t estimated log-odds estimated probabilities p and estimated output class based on cutoff value θ = 0.19.

[0304]

[0305]

[0306] Example 11

[0307] Equivalently to the analysis performed in Example 8, a preferred embodiment of the application was found that employed 8 genes and 2 clinicopathological variables (age and Breslow depth), including the following set of genes: GDF15, MLANA, PLAT, IL8, ITGB3, LOXL4, SERPINE2 and TGFBR1, with the parameters of the logistic regression model as follows:

[0308]

[0309] Different operating points were evaluated and it was found that multiple operating points provided clinically relevant classifiers with high NPV and a substantial reduction in the number of SLNB procedures (SLNB.RR = SLNB reduction rate). The NPV975 operating point of 0.116 was particularly preferred, see table below.

[0310] GECL model performance table with 8 genes and 2 clinicopathological variables

[0311] Comparing this GECL model with 8 genes and 2 clinicopathological variables to the performance of Examples 1 and 8, it was shown that this model was superior (in NPV and SLNB.RR) to either the expression (GE) or clinicopathological (CL) models alone, and its performance was very similar to the GECL model in Example 1.

[0312] Kaplan-Meier analysis was performed to assess the GECL model intersected with SLNB in relation to RFS, DRFS and MSS. 5-year / 60-month survival rates are shown in the table below. Both GECL and SLNB had a large separation between the positive pos / negative neg groups. More importantly, the survival rate of those patients who were SLNB negative (and thus missed) but were determined to be positive by the GECL was very low. This demonstrated the clinical relevance of the GECL model as a prognostic marker.

[0313] MSS DRFS RFS 5 years 5 years 5 years GECL = Neg + SLNB = Neg 0.96 0.95 0.88 GECL = Neg + SLNB = Pos 0.83 0.50 0.34 GECL = Pos + SLNB = Neg 0.90 0.79 0.70 GECL = Pos + SLNB = Pos 0.84 0.72 0.50

[0314] Table:

[0315] Table 1

[0316]

[0317] Table 2

[0318]

[0319] Table 3

[0320] Value

[0321]

[0322] Table 4

[0323]

[0324] Table 5

[0325]

[0326] Table 6

[0327]

[0328] Table 7

[0329]

[0330] Table 8

[0331]

[0332] Table 9

[0333]

[0334] Table 10

[0335]

[0336] Table 11

[0337]

[0338] Table 12

[0339]

[0340] Table 13

[0341]

[0342] Table 14

[0343]

[0344] Table 15

[0345]

[0346] Table 16

[0347]

[0348] Table 17

[0349]

[0350] Table 18a

[0351] Breslow depth < 2 mm

[0352]

[0353] Table 18b

[0354] Breslow depth > 2 mm

[0355]

[0356] Table 19a: Survival estimates for groups based on GECL classifier status or SLNB status at different time points. Survival curves were compared using the Cox proportional hazards model, see Table 25.

[0357]

[0358] Table 19b:

[0359]

[0360] Table 19c:

[0361]

[0362]

[0363] Table 20a: Survival estimates for groups based on GECL classifier status or SLNB status at different time points. Survival curves were compared using the Cox proportional hazards model, see Table 25.

[0364]

[0365] Table 20b

[0366]

[0367] Table 20c:

[0368]

[0369] Table 21a: Survival estimates for groups based on GECL classifier status or SLNB status at different time points. Survival curves were compared using the Cox proportional hazards model, see Table 25.

[0370]

[0371] Table 21b:

[0372]

[0373] Table 21c:

[0374]

[0375] Table 22a: Survival estimates for groups based on SLNB and GECL classifier status at different time points.

[0376] Survival curves were compared using the log-rank test, p<0.0001.

[0377]

[0378] Table 22b:

[0379]

[0380] Table 22c:

[0381]

[0382]

[0383] Table 23a: Survival estimates for groups based on SLNB and GECL classifier status at different time points.

[0384] Survival curves were compared using the log-rank test, p<0.0001.

[0385]

[0386] Table 23b:

[0387]

[0388]

[0389] Table 23c:

[0390]

[0391] Table 24a: Survival estimates for groups based on SLNB and GECL classifier status at different time points.

[0392] Survival curves were compared using log-rank test, p<0.0001.

[0393]

[0394] Table 24b:

[0395]

[0396] Table 24c:

[0397]

[0398] Table 25a: Hazard ratios and p-values for 2 curves of GECL classifier output and SLNB biopsy results.

[0399]

[0400]

[0401] Table 25b:

[0402]

[0403] Table 25c:

[0404]

[0405]

[0406] Table 26a: Multivariate hazard ratios and p-values for 2 curves of GECL classifier output and SLNB biopsy results.

[0407]

[0408]

[0409] Table 27

[0410]

[0411] Table 28

[0412]

[0413] Table 29

[0414]

[0415] Table 30

[0416]

[0417] Table 31

[0418]

[0419] Table 32

[0420]

[0421] Table 33

[0422] Max bACC

[0423]

[0424] Table 34

[0425] SEeqSP

[0426]

[0427] Table 35

[0428] NPV 97

[0429]

[0430] Table 36 NPV 98

[0431]

[0432] Table 37 CL

[0433]

[0434] Table 38 GE

[0435]

[0436] Table 39 GE CL

[0437]

[0438] Table 40 CL

[0439]

[0440] Table 41 GE

[0441]

[0442] Table 42

[0443] GE CL

[0444]

[0445] Table 43

[0446] CL

[0447]

[0448] Table 44

[0449] GE

[0450]

[0451] Table 45

[0452] GECL

[0453]

[0454] Table 46

[0455] CL

[0456]

[0457] Table 47

[0458] GE

[0459]

[0460] Table 48

[0461] CLGE

[0462]

[0463] Table 49

[0464] Minimum and Maximum AUC

[0465]

Claims

1. Use of reagents for detecting the expression level of an expression signature of genes comprising ITGB3, PLAT, GDF15 and IL8 for the manufacture of a kit for classifying an individual having a primary skin melanoma, said classifying comprising determining in a sample from said individual an expression signature of genes ITGB3, PLAT, GDF15 and IL8, and wherein an individual is classified as having metastasis positive sentinel lymph node (SLN) or as having metastasis negative SLN.

2. Use of reagents for detecting the expression level of an expression signature of genes comprising ITGB3, PLAT, GDF15 and IL8 for the manufacture of a kit for determining a treatment and / or diagnostic check schedule for an individual having a skin melanoma, said determining comprising determining in a sample from said individual an expression signature of genes ITGB3, PLAT, GDF15 and IL8, and determining a treatment and / or diagnostic schedule as a function of said expression level.

3. Use of reagents for detecting the expression level of an expression signature of genes comprising ITGB3, PLAT, GDF15 and IL8 for the manufacture of a kit for predicting the prognosis of an individual having a primary skin melanoma, said predicting comprising determining in a sample from said individual an expression signature of genes, wherein said expression signature of genes comprises the following genes: ITGB3, PLAT, GDF15 and IL8.

4. Use according to claim 1, wherein an individual is selected for SLNB based on said classification.

5. Use according to claim 2 or 3, wherein an individual is selected for SLNB based on said expression level.

6. The use according to claim 1, wherein, An individual classified as having metastasis positive SLN is treated by performing SLNB.

7. Use according to claim 3, wherein the prognosis of said individual is determined based on the expression level of said expression signature of genes.

8. Use according to claim 7, wherein an individual is classified as having a poor prognosis or as having a good prognosis.

9. The use of any one of claims 1-3, further comprising classifying an individual having a primary cutaneous melanoma, comprising determining the gene expression signature in a sample of the individual, wherein, The expression signature of genes comprises at least one of the following genes: KRT14, SPP1, FN1 and LOXL3.

10. Use according to claim 3, wherein said individual is classified as having metastasis positive SLN and / or poor prognosis based on the expression signature of genes, and SLNB is selected and / or a cancer treatment is provided.

11. Use according to claim 10, wherein said cancer treatment is chemotherapy or immunotherapy.

12. Use according to claim 11, wherein said cancer treatment is selected from Ipilimumab, Nivolumab and Pembrolizumab.

13. Use according to any one of claims 1-3, wherein said expression signature of genes further comprises MLANA, LOXL4, SERPINE2 and TGFBR1.

14. Use according to any one of claims 1-3, wherein the expression signature of genes further comprises ADIPOQ.

15. Use according to any one of claims 1-3, wherein the expression signature of genes further comprises PRKCB, ADAM12, LGALS1.

Citation Information

Patent Citations

  • Universal reference standard for normalization of microarray gene expression profiling data

    US20060136145A1

  • Methods and materials for identifying malignant skin lesions

    US20160222457A1

  • Methods and materials for identifying malignant skin lesions

    WO2014077915A1