Predicting prognosis for breast cancer patients

JP2024544750A5Pending Publication Date: 2025-10-27KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024527118
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-09
Filing Date
2022-10-20
Publication Date
2025-10-27

AI Technical Summary

Technical Problem

Current prognostic tools for breast cancer are inadequate in predicting treatment outcomes and recurrence, as they fail to account for the complex interplay of various factors influencing disease progression and treatment efficacy.

Method used

A method utilizing gene expression profiles of immune defense response genes (e.g., AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, ZBP1), T cell receptor signaling genes (CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, ZAP70), and PDE4D7 correlated genes (ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, VWA2) to predict breast cancer prognosis, including mortality and recurrence risks.

Benefits of technology

The method provides improved prediction of breast cancer prognosis, enabling more accurate treatment selection and reducing ineffective treatments, thereby enhancing patient survival and reducing unnecessary morbidity and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a method of predicting prognosis in a subject with breast cancer, comprising determining or receiving a gene expression profile having expression levels of four or more genes selected from a first, second and / or third gene expression profile, wherein the first gene expression profile has one or more immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1, and the second gene expression profile has one or more immune defense response genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PR and a third gene expression profile comprising one or more T cell receptor signaling genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, wherein the gene expression profile is determined in a biological sample obtained from the subject, the method comprising determining a prediction of prognosis based on the gene expression profile having expression levels of four or more genes, wherein the prediction is a favorable or unfavorable risk of breast cancer related mortality, locoregional recurrence, and / or distant recurrence.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method for predicting the prognosis of a subject with breast cancer, and a computer program for predicting the prognosis of a subject with breast cancer.Furthermore, the present invention relates to a diagnostic kit, the use of the kit, the use of the kit in a method for predicting the prognosis of a subject with breast cancer, the use of a gene expression profile of one or more immune defense response genes, one or more T cell receptor signaling genes and / or one or more PDE4D7-correlated genes in a method for predicting the prognosis of a subject with breast cancer, and corresponding computer programs. [Background technology]

[0002] introduction Cancer is a class of diseases in which groups of cells exhibit uncontrolled proliferation, invasion, and sometimes metastasis. These three malignant properties of cancer distinguish it from benign tumors, which are self-limited and do not invade or metastasize.

[0003] Breast cancer is the most common and second most lethal non-cutaneous malignancy in women, with an estimated 300.000 new cases of breast cancer diagnosed annually in 2021 in the United States alone, and 45.000 breast cancer-related deaths (see ACS (American Cancer Society), Cancer Facts & Figures 2021, 2021). The median age at diagnosis is 62 years, with a 5-year relative survival rate of 90%. However, this survival rate is not representative of all breast cancer patients and strongly depends on the stage and type of breast cancer the patient has. In general, the earlier the cancer is diagnosed, the higher the chance of surviving more than 5 years after diagnosis. In the case of breast cancer, 63% are diagnosed at a localized stage, with a 5-year relative survival rate of 98.9%. However, if diagnosed at a localized stage (30%) or distant stage (6%), the 5-year relative survival rate drops to 85.7% and 28.1%, respectively. Over the past decade, the age-adjusted rate of new breast cancer cases in women has increased by an average of 0.3% per year from 2008 to 2017, while the age-adjusted death rate has decreased by an average of 1.4% per year from 2009 to 2018. Although breast cancer almost exclusively affects women, it can also occur in men (SEER Cancer Stat Facts).

[0004] Although breast cancer in men is rare, men can be affected by the disease. In 2021, it is estimated that 2,650 new cases of invasive breast cancer will be diagnosed in American men, and approximately 530 breast cancer-related deaths will occur (ACS 2021).

[0005] There are several known risk factors for breast cancer. Lifestyle-related factors include alcohol consumption, overweight or obesity, physical inactivity, not having children before age 30, not breastfeeding, using hormonal contraception, receiving hormone therapy after menopause, and having breast implants. Other factors that contribute to breast cancer risk include older age, inheritance of certain genetic changes (e.g., in BRCA1 and BRCA2, or less commonly, in ATM, TP53, CHEK2, PTEN, CDH1, STK11, and PALB2 genes), family history, race and ethnicity, being tall, having dense breast tissue, certain benign breast disorders, starting menstruation early (before age 12) or going through menopause after age 55, having received chest radiation, and having been exposed to DES. The associations of some risk factors, such as working the night shift, smoking, diet and vitamins, and certain environmental chemicals, are inconclusive. Although there is widespread information that may be linked to breast cancer, such as antiperspirant use and induced abortion, there are many other factors that studies have shown to have no association.

[0006] Breast cancer can originate in various parts of the breast, including the ducts, lobules, and the tissues in between. The most common type of breast cancer is invasive ductal carcinoma (IDC), which accounts for approximately 70-80% of all breast cancers. This type of breast cancer is classified as invasive due to its aggressive spread into the surrounding breast tissue, as opposed to, for example, ductal carcinoma in situ. Another common type of invasive breast carcinoma is invasive lobular carcinoma (ACS, ibid.). Other types of breast cancer include, for example, medullary carcinoma, mucinous carcinoma, and tubular carcinoma, all of which are (very) rare IDCs.

[0007] According to the American Cancer Society, the overall five-year survival rate for breast cancer is about 90 percent, but this depends on many factors, including the treatment, the type of breast cancer, how advanced the cancer is, and your overall health.

[0008] Treatment of breast cancer depends not only on the extent of the disease, but also on the risk of the disease spreading after primary treatment. The six standard treatments for breast cancer are surgical removal (e.g., radical mastectomy or lumpectomy), radiation therapy (RT), chemotherapy, hormonal therapy, targeted therapy, and immunotherapy. Surgery is the most commonly used treatment, and most women with breast cancer will undergo some type of surgery as part of their standard treatment. RT is generally given as external radiation or by implanting radioactive seeds in the breast (brachytherapy (internal radiation therapy)), or a combination of both. RT is especially preferred in patients who are not candidates for surgery or whose cancer cells have spread, or when cancer is present in many lymph nodes, or when cancer is present in certain surgical margins, such as the skin or muscle (ACS, ibid.).

[0009] The more likely it is that a selected treatment (or combination of treatments) will cause significant morbidity for the patient, the more likely it will be to predict whether a patient (or group of patients) will benefit from the treatment (or combination of treatments). Treatments (combinations of treatments) that significantly worsen the patient's quality of life should only be considered if there is a high probability that they will prevent, reduce or minimize disease progression.

[0010] For the diagnosis of breast cancer and for treatment decisions, multiple tests are performed to confirm the presence of cancer. These tests range from a physical examination to check for general signs of health or disease, to various forms of imaging (X-ray, ultrasound, or MR), and biochemical tests in blood, or collection of tissues and body fluids for pathological testing by biopsy procedures. A wide range of potential biomarkers in tissues and body fluids have been investigated, but validation is often limited and they generally provide prognostic information, not predictive (treatment specific) value. Examples of such molecular tests to determine treatment are estrogen (ER) and progesterone (PR) receptor tests and human epidermal growth factor type 2 receptor (HER2 / neu) tests. ER- and PR-receptor tests measure the expression of ER and PR on breast tumor cells. ER / PR positive tumors are generally treated with molecules that inhibit the activity of these receptors. HER2 / neu tests measure the expression of the HER2 / neu gene / protein in tumor cells. HER2 / neu-positive breast cancer is typically treated with drugs or antibodies that target the HER2 / neu protein on tumor cells.

[0011] Numerous studies have been conducted to determine molecular signatures of breast cancer based on multiple genes using expression profiles of a set of genes in tumors (see Sorlie, T et al. Proc Natl Acad Sci US A. 2001 Sep 11;98(19):10869-74). The potential relevance of gene expression for prognostic purposes in breast cancer has led to the development of at least three separate commercially available multigene tests, including the Oncotype DX® test (Genomic Health, Redwood, CA, USA), the MammaPrint® test (Netherlands Cancer Institute™ and Agendia™, The Netherlands), and the Prosigna® test (NanoString Technologies, Seattle, WA, USA). The Oncotype DX test is a reverse transcriptase chain reaction assay that measures a panel of 21 genes (16 cancer-related genes ER, PR, Bcl2, SCUBE2, HER2, GRB7, Ki-67, STK15, survivin, cyclin B1, MYBL2, stromelysin 3, cathepsin L2, GSTM1, CD68, and BAG1, as well as five housekeeping genes (β-actin, GAPDH, RPLPO, GUS, and TFRC)) to stratify breast cancer patients into three risk groups based on their likelihood of cancer recurrence (ACS 2019). The MammaPrint test measures gene expression profiles of 70 genes involved in various intracellular pathways to stratify breast cancer patients into low-risk and high-risk groups based on their prognosis of early distant recurrence. MammaPrint, for example in combination with the BluePrint® test, can also predict response to adjuvant chemotherapy in breast cancer subjects (Longbottom et al. Cancer Res February 15 (4 supplement) PS6-30 (2021)). The Prosigna One test is based on the PAM50 (Prediction Analysis of Microarray 50) gene signature. The Prosigna assay is a genomic test that analyzes the activity of specific genes in early-stage hormone receptor-positive breast cancer.This assay specifically provides prognosis for the risk of distant recurrence in hormone receptor positive stage I-III breast cancer treated with adjuvant hormonal therapy (Wallden, B et al. BMC Med Genomics 8, 54 (2015)). The PAM50 signature used in this assay comprises a panel of 50 genes for a method to predict the intrinsic biological subtype of breast cancer. Using a risk model incorporating the gene expression-based "intrinsic" subtypes luminal A, luminal B, HER2-enriched, and basal-like, the researchers were able to identify and diagnose the intrinsic subtypes of breast cancer (Parker, JS et al. J Clin Oncol. Mar 10;27(8):1160-7 (2009)). Furthermore, one study reported that the PAM50 gene signature had a high prognostic value but was unable to predict improved prognosis with high-dose chemotherapy (Liu, M. et al., Breast Cancer 2, 15023 (2016)). Further studies reported that PAM50 intrinsic subtyping and a "recurrence risk" score could stratify patients into prognostic groups that predicted the risk of recurrence and the need for adjuvant therapy (Ohnstad, HO et al., Breast Cancer Res 19, 120 (2017)).

[0012] Although these clinically available tools provide the means to stratify patients into different risk groups based on the endpoints of cancer (distant) recurrence or response to chemotherapy, better predictive tools for breast cancer prognosis are needed. Summary of the Invention [Problem to be solved by the invention]

[0013] Predicting treatment outcomes is very complex, as many factors contribute to treatment efficacy and disease recurrence. It is highly likely that important factors have not yet been identified, while the influence of other factors cannot be precisely determined. The inventors hypothesized that improving prediction of breast cancer outcomes for each patient could improve treatment selection and improve survival rates. The inventors determined that this could be achieved by 1) optimizing the stratification of subjects suffering from (subtypes of) breast cancer into risk groups based on prognosis after surgery, and 2) directing patients for whom surgery is predicted to be of little benefit to alternative, secondary, potentially more effective forms of treatment. Furthermore, the inventors found that this would reduce the suffering of patients who are spared ineffective treatments and reduce the costs spent on ineffective treatments. [Means for solving the problem]

[0014] In conclusion, the inventors hypothesize that there is a high need for better predicting breast cancer prognosis, i.e. the prognosis of breast cancer subjects, and provide herein methods and means for achieving improved prediction of prognosis.

[0015] In a first aspect, the present invention relates to a method for predicting prognosis of a subject with breast cancer, comprising determining or receiving a gene expression profile determination result, comprising comparing expression levels of four or more genes selected from a first, second and / or third gene expression profile, wherein said first expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8 and ZBP1, and / or wherein said second gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all immune defense response genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PT and / or a third gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all T cell receptor signaling genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, wherein said gene expression profile is determined in a biological sample obtained from the subject, comprising determining a prognosis based on the first, second and / or third gene expression profile comprising expression levels of four or more genes, wherein said prediction is a favorable or unfavorable risk of breast cancer associated death, locoregional recurrence and / or distant recurrence.

[0016] In a second aspect, the present invention relates to a method for determining whether a gene expression profile is indicative of a gene expression profile comprising expression levels of four or more genes selected from a first, second and / or third gene expression profile, the first gene expression profile comprising: AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZB P1, wherein the second gene expression profile comprises one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all of the immune defense response genes selected from the group consisting of: CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70, and a third gene expression profile comprising one or more, for example 1, 2, 3, 4, 5, 6, 7 or all PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, wherein said gene expression profile is determined in a biological sample obtained from the subject; determining a prediction of the subject's prognosis based on the first gene expression profile(s), or the second gene expression profile(s), or the third gene expression profile(s), or the first, second and third gene expression profile(s), wherein said prediction is a favorable or unfavorable risk of breast cancer related mortality, locoregional recurrence and / or distant recurrence.

[0017] In a third aspect, the present invention relates to a diagnostic kit comprising at least one polymerase chain reaction primer, and optionally at least one probe, for determining a gene expression profile, said gene expression profile comprising expression levels of four or more genes selected from a first, second and / or third gene expression profile, wherein the first gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all of the immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8 and ZBP1. and the second gene expression profile comprises one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all of the T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70, and / or the third gene expression profile comprises one or more, e.g., 1, 2, 3, 4, 5, 6, 7 or all of the PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2T.

[0018] In a fourth aspect, the present invention relates to the use of a kit as defined in the third aspect of the invention in a method for predicting the prognosis of a subject with breast cancer, preferably in a method as defined in the first aspect of the invention.

[0019] In a fifth aspect, the present invention relates to a method for the treatment of breast cancer comprising receiving a biological sample obtained from a subject with breast cancer, determining a gene expression profile in the biological sample using a kit according to the third aspect, wherein the gene expression profile comprises expression levels of four or more genes selected from a first, second and / or third gene expression profile, the first gene expression profile comprising one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all of the immune defence response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8 and ZBP1. wherein the second gene expression profile comprises one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70, and / or the third gene expression profile comprises one or more, e.g., 1, 2, 3, 4, 5, 6, 7 or all PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2.

[0020] In a final aspect, the present invention relates to the use of a gene expression profile, said gene expression profile comprising expression levels of four or more genes selected from a first, second and / or third gene expression profile, wherein the first gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8 and ZBP1, and the second gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all immune defense response genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, and / or a third gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all of the T cell receptor signaling genes selected from the group consisting of PRKACA, PRKACB, PTPRC, and ZAP70, and / or a third gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7 or all of the PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, comprising a step of determining a prognostic prediction based on gene expression levels for the four or more genes, wherein the prediction is a favorable or unfavorable risk of breast cancer related mortality, locoregional recurrence and / or distant recurrence. [Brief description of the drawings]

[0021] [Figure 1]Kaplan-Meier curves for the IDR_14 model in a cohort of 997 patients (training set used to develop the IDR_14 model) in which all patients underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the IDR_14 model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Diagram 2] Kaplan-Meier curves for the TCR_17 model in a cohort of 997 patients (training set used to develop the TCR_17 model) in which all patients underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients (threshold = 0, low risk (≦0), high risk (>0)) for the TCR_17 model classes analyzed, i.e., risk patients at any time interval ≥ 25 months after surgery. [Diagram 3] Kaplan-Meier curves for the PDE4D7_CORR model in a cohort of 997 patients (training set used to develop the PDE4D7_CORR model) in which all patients underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients (threshold = 0, low risk (≦0), high risk (>0)) for the PDE4D7_CORR model classes analyzed, i.e. risk patients in any time interval ≥ 25 months after surgery. [Figure 4]Kaplan-Meier curves for the BRCAI_model in a cohort of 997 patients (training set used to develop the BRCAI_model) in which all patients underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients in the BRCAI_model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Diagram 5] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 997 patients (training set used to develop the BRCAI_clinical_model) in which all patients underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery, are shown. [Figure 6] Kaplan-Meier curves for the BRCAI_model in a cohort of 997 patients (training set used to develop the BRCAI_model) in which all patients underwent surgery. The clinical endpoint tested was all-cause mortality after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI_model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 7]Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 997 patients (the training set used to develop the BRCAI_clinical_model) in which all patients underwent surgery. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery, are shown. [Figure 8] Kaplan-Meier curves for the BRCAI_model in a cohort of 997 patients (training set used to develop the BRCAI_model) in which all patients underwent surgery. The clinical endpoint tested was local recurrence after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI_model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 9] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 997 patients (training set used to develop the BRCAI_clinical_model) in which all patients underwent surgery. The clinical endpoint tested was locoregional recurrence after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Figure 10]Kaplan-Meier curves for the BRCAI_model in a cohort of 997 patients (training set used to develop the BRCAI_model) in which all patients underwent surgery. The clinical endpoint tested was distant recurrence after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI_model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 11] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 997 patients (training set used to develop the BRCAI_clinical_model) in which all patients underwent surgery. The clinical endpoint tested was distant recurrence after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery, are shown. [Figure 12] Kaplan-Meier curves for the PAM50 subtype model in a cohort of 997 patients (training set) in which all patients underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death). Log ranks are included in the figure. The included supplementary listing shows the number of patients at risk in each of the Basal, Her2, LumA, LumB, and Normal groups for the PAM50 subtypes analyzed, i.e., at risk patients in any time interval ≥25 months after surgery. [Figure 13]Kaplan-Meier curves for the PAM50 subtype model in a cohort of 997 patients (training set) in which all patients underwent surgery. The clinical endpoint tested was all-cause mortality (Death). Log ranks are included in the figure. The included supplementary listing shows the number of patients at risk in each of the Basal, Her2, LumA, LumB, and Normal groups for the PAM50 subtypes analyzed, i.e., at risk patients in any time interval ≥25 months after surgery. [Figure 14] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 118 PAM50 Basal subtype patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Figure 15] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 87 PAM50 Her2 subtype patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Figure 16]Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 466 PAM50 LumA subtype patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Figure 17] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 268 PAM50 LumB subtype patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Figure 18] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 58 PAM50 normal subtype patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Figure 19]Kaplan-Meier curves for the BRCAI_model in a cohort of 953 patients (validation set used to validate the BRCAI_model) in which all patients underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients in the BRCAI_model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 50 months after surgery. [Figure 20] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 953 patients (validation set used to validate the BRCAI_clinical_model) in which all patients underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Figure 21] Kaplan-Meier curves for the BRCAI_model in a cohort of 953 patients (validation set used to validate the BRCAI_model) in which all patients underwent surgery are shown. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI_model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 50 months after surgery. [Figure 22]Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 953 patients (the validation set used to validate the BRCAI_clinical_model) in which all patients underwent surgery. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Figure 23] Kaplan-Meier curves for the BRCAI_model in a cohort of 950 patients (validation set used to validate the BRCAI_model) in which all patients underwent surgery. The clinical endpoint tested was locoregional recurrence after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI_model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 50 months after surgery. [Figure 24] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 950 patients (validation set used to validate the BRCAI_clinical_model) in which all patients underwent surgery. The clinical endpoint tested was locoregional recurrence after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Diagram 25]Kaplan-Meier curves for the BRCAI_model in a cohort of 953 patients (validation set used to validate the BRCAI_model) in which all patients underwent surgery. The clinical endpoint tested was distant recurrence after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients in the BRCAI_model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 50 months after surgery. [Figure 26] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 953 patients (validation set used to validate the BRCAI_clinical_model) in which all patients underwent surgery. The clinical endpoint tested was distant recurrence after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery, are shown. [Figure 27] Kaplan-Meier curves for the PAM50 subtype model in a cohort of 975 patients (validation set) in which all patients underwent surgery. The clinical endpoint tested was breast cancer-related death (BCa Death). Log ranks are included in the figure. The included supplementary listing shows the number of patients at risk in each of the Basal, Her2, LumA, LumB, and Normal groups for the PAM50 subtypes analyzed, i.e., at risk patients in any time interval ≥ 50 months after surgery. [Figure 28]Kaplan-Meier curves for the PAM50 subtype model in a cohort of 975 patients (validation set) in which all patients underwent surgery. The clinical endpoint tested was all-cause mortality (Death). Log-ranks are included in the figure. The included supplementary listing shows the number of patients at risk in each of the Basal, Her2, LumA, LumB, and Normal groups for the PAM50 subtypes analyzed, i.e., risk patients at any time interval ≥ 50 months after surgery. [Figure 29] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 210 PAM50 Basal subtype patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Diagram 30] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 153 PAM50 Her2 subtype patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Diagram 31]Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 251 PAM50 LumA subtype patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Diagram 32] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 220 PAM50 LumB subtype patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Diagram 33] Kaplan-Meier curves for the BRCAI_clinical_model in a cohort of 141 PAM50 normal subtype patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Diagram 34]Kaplan-Meier curves for the BRCAI_MB model in a cohort of 673 patients (validation set used to validate the BRCAI_MB model) in which all patients underwent surgery are shown. The clinical endpoint tested was all-cause mortality after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Diagram 35] Kaplan-Meier curves for the BRCAI_clinical_MB model in a cohort of 563 patients (validation set used to validate the BRCAI_clinical_MB model) in which all patients underwent surgery. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery, are shown. [Diagram 36] Kaplan-Meier curves for the PAM50 subtype model in a cohort of 563 patients (training set) in which all patients underwent surgery. The clinical endpoint tested was all-cause mortality (Death). Log-ranks are included in the figure. The included supplementary listing shows the number of patients at risk in each of the Basal, Her2, LumA, LumB, and Normal groups for the PAM50 subtypes analyzed, i.e., at risk patients at any time interval ≥ 50 months after surgery. [Figure 37]Kaplan-Meier curves for the BRCAI_clinical_MB model in a cohort of 116 PAM50 Basal subtype patients, all of whom underwent surgery. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Figure 38] Kaplan-Meier curves for the BRCAI_clinical_MB model in a cohort of 36 PAM50 Her2 subtype patients, all of whom underwent surgery. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Figure 39] Kaplan-Meier curves for the BRCAI_clinical_MB model in a cohort of 305 PAM50 LumA subtype patients, all of whom underwent surgery. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 50 months after surgery. [Diagram 40]Kaplan-Meier curves for the BRCAI_clinical_MB model in a cohort of 106 PAM50 LumB subtype patients, all of whom underwent surgery. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the BRCAI_clinical_model classes analyzed (threshold = 0.3, low risk (≦0.3), high risk (>0.3)), i.e. at any time interval ≥ 20 months after surgery. [Diagram 41] Kaplan-Meier curves of the CT&HT&RT model in a cohort of 60 BRCAI_clinical low risk patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients in the analyzed CT&HT&RT model classes (no second-line treatment, any second-line treatment), i.e. at any time interval ≥ 50 months after surgery. [Diagram 42] Kaplan-Meier curves of the CT&HT&RT model in a cohort of 125 BRCAI_clinical high risk patients, all of whom underwent surgery. The clinical endpoint tested was breast cancer-specific death (BCa Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the analyzed CT&HT&RT model classes (no second-line treatment, any second-line treatment), i.e. at any time interval ≥ 50 months after surgery. [Diagram 43]Kaplan-Meier curves of the CT&HT&RT model in a cohort of 60 BRCAI_clinical low risk patients, all of whom underwent surgery. The clinical endpoint tested was all-cause mortality after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the analyzed CT&HT&RT model classes (no secondary treatment, all secondary treatments), i.e. at any time interval ≥ 50 months after surgery. [Diagram 44] Kaplan-Meier curves of the CT&HT&RT model in a cohort of 125 BRCAI_clinical high risk patients, all of whom underwent surgery. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary list shows the number of risk patients for the analyzed CT&HT&RT model classes (no secondary treatment, all secondary treatments), i.e. at any time interval ≥ 50 months after surgery. [Diagram 45] Kaplan-Meier curves for the BRCAI_6.1 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Diagram 46] Kaplan-Meier curves for the BRCAI_6.2 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 47] Kaplan-Meier curves for the BRCAI_6.3 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 48] Kaplan-Meier curves for the BRCAI_6.4 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 49] Kaplan-Meier curves for the BRCAI_6.5 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 50] Kaplan-Meier curves for the BRCAI_6.6 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 51]Kaplan-Meier curves for the BRCAI_6.7 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 52] Kaplan-Meier curves for the BRCAI_6.8 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 53] Kaplan-Meier curves for the BRCAI_6.9 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 54] Kaplan-Meier curves for the BRCAI_6.10 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. [Figure 55]Kaplan-Meier curves for the BRCAI_6.11 model in a cohort of 1044 patients, all of whom underwent surgery, are shown. The clinical endpoint tested was all-cause mortality (Death) after surgery. Log-rank, HR, and confidence intervals are included in the figure. The included supplementary listing shows the number of risk patients for the BRCAI model classes analyzed (threshold = 0, low risk (≦0), high risk (>0)), i.e. at any time interval ≥ 25 months after surgery. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0022] definition As used herein, the singular terms "a," "an," and "the" do not exclude a plurality.

[0023] The term "biological sample" or "sample obtained from a subject" refers to any biological material obtained from a subject, for example a breast cancer subject, via suitable methods known to those of skill in the art.

[0024] As used herein, the term "and / or" indicates that one or more of the stated examples may occur alone or in combination with at least one of the stated examples, up to all of the stated examples.

[0025] As used herein, the term "at least" in reference to a particular value means the particular value or greater. For example, "at least 2" is understood to be the same as "2 or greater," i.e., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, ... etc.

[0026] The term "breast cancer" refers to cancer of breast tissue that occurs when breast cells mutate and begin to grow uncontrollably.

[0027] The term "breast cancer-specific or disease-specific mortality" refers to patient death due to breast cancer.

[0028] The term "clinical recurrence" refers to the presence of clinical signs indicative of the presence of tumor cells, for example as determined using in vivo imaging.

[0029] As used herein, "comprise" or variations thereof will be understood to include a stated element, integer or step, or group of elements, integers or steps, but not to exclude other elements, integers or steps, or group of elements, integers or steps. The verb "comprise" includes the verbs "consisting essentially of" and "consisting of.

[0030] As used herein, the term "immune defense response gene" is used interchangeably with "IDR gene" or "immune defense gene" refers to one or more of the genes selected from AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1.

[0031] The term "metastasis" refers to the presence of metastatic disease in organs other than breast tissue.

[0032] As used herein, the term "PDE4D7-related gene" is used interchangeably with "PDE4D7 gene" and refers to one or more genes selected from ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2.

[0033] As used herein, the term "T cell receptor signaling genes" is used interchangeably with "TCR signaling genes" or "TCR genes" refers to one or more of the genes selected from CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70.

[0034] Detailed Description of the Invention The immune system in cancer In recent years, the importance of the immune system in cancer suppression, as well as in cancer initiation, promotion and metastasis, has become very evident (Mantovani et al., Nature. 454(7203):436-44 (2008); Giraldo et al., Br J Cancer. 120(1):45-53 (2019)). Immune cells and the molecules they secrete form an important part of the tumor microenvironment, and most immune cells can infiltrate tumor tissue. The immune system and tumors interact and shape each other. Thus, while antitumor immunity can prevent tumor formation, an inflammatory tumor environment promotes cancer initiation and growth. At the same time, tumor cells that arise in an immune system-independent manner can shape the immune microenvironment by recruiting immune cells and have a pro-inflammatory effect while also suppressing antitumor immunity.

[0035] Some immune cells in the tumor microenvironment have either general tumor-promoting or general tumor-suppressing effects, whereas other immune cells exhibit plasticity and show both tumor-promoting and tumor-suppressing potential. Thus, the overall immune microenvironment of a tumor is a mix of the various immune cells present, the cytokines they produce, and their interactions with tumor cells and other cells in the tumor microenvironment (Giraldo Br J Cancer. 120(1):45-53 (2019)).

[0036] Treatment is influenced by immune components of the tumor microenvironment, but RT itself also has extensive effects on the composition of these components. Because suppressive cell types are relatively radioinsensitive, their relative numbers increase. Counterproductively, a given radiation insult activates cell survival pathways and stimulates the immune system, which induces an inflammatory response and immune cell recruitment. Whether the net effect is tumor-promoting or tumor-suppressing is currently unclear, but its potential for enhancing cancer immunotherapy is being explored.

[0037] In summary, the state of the immune system and the immune microenvironment have a major impact on therapeutic efficacy.

[0038] The inventors identified gene signatures and combinations of these signatures with clinical parameters, and the resulting models showed significant associations with mortality and are therefore expected to improve the prediction of the outcome of these treatments.

[0039] Immune response defense genes Genomic DNA integrity and stability are constantly under stress induced by various intracellular and extracellular factors, such as radiation, viral or bacterial infection, but also oxidative and replicative stress (see Gasser S. et al., "Sensing of dangerous DNA", Mechanisms of Aging and Development, Vol.165, pp.33-46, 2017). To maintain DNA structure and stability, cells must be able to recognize all types of DNA damage, such as single- or double-strand breaks, induced by various factors. This process involves a large number of specific proteins as part of DNA recognition pathways, depending on the type of damage.

[0040] Recent findings suggest that mislocalized DNA (e.g., DNA that appears unnaturally in the cytoplasmic fraction of cells as opposed to the nucleus) and damaged DNA (e.g., mutations that occur in the development of cancer) are utilized by the immune system to identify infected or otherwise diseased cells, while genomic and mitochondrial DNA present in healthy cells is ignored by DNA recognition pathways. In diseased cells, cytoplasmic DNA sensor proteins have been demonstrated to be involved in the detection of DNA that is unnaturally present in the cytoplasm of cells. The detection of such DNA by different nucleic acid sensors is translated into a similar response that leads to the activation of innate immune system components following nuclear factor kappa-B (NF-κB) and type I interferon (IFN type I) signaling. While the recognition of viral DNA is known to induce IFN type I responses, knowledge that sensing DNA damage triggers an immune response is only accumulating more recently.

[0041] TLR9 (Toll-like receptor 9), which is localized in endosomes, was one of the first DNA sensor molecules identified to be involved in immune recognition of DNA by signaling downstream through the adaptor protein myeloid differentiation primary response protein 88 (MYD88). This interaction then activates mitogen-activated protein kinases (MAPK) and NF-kB. TLR9 also induces the production of type I interferons through activation of IRF7 via IkB kinase α (IKKα) in plasmacytoid dendritic cells (pDCs). Various other DNA immune receptors, including IFI16 (IFN-γ-inducible protein 16), cGAS (cyclic DMP-AMP synthase), DDX41 (DEAD-box helicase 41), and ZBP1 (Z-DNA binding protein 1), interact with STING (stimulator of IFN genes) and activate the IKK complex and IRF3 via TBK1 (TANK-binding kinase 1). ZBP1 also activates NF-kB through the recruitment of RIP1 and RIP3 (receptor interacting proteins 1 and 3, respectively). The helicase DHX36 (DEAH-box helicase 36) interacts with TRID in a complex and induces NF-kB and IRF-3 / 7, whereas the DHX9 helicase stimulates MYD88-dependent signaling in plasmacytoid dendritic cells. The DNA sensor LRRFIP1 (leucine-rich repeat flightless interacting protein) complexes with β-catenin to activate the transcription of IRF3, while AIM2 (absent in melanoma 2) recruits the adaptor protein ASC (apoptotic speck-like protein) to induce the caspase-1-activated inflammasome complex leading to the secretion of interleukin-1β (IL-1β) and IL-18 (see Figure 1 in Gasser S. et al., 2017, where a schematic of the DNA damage and DNA sensor pathways leading to the production of inflammatory cytokines and the expression of ligands for activating innate immune receptors is shown. Members of the non-homologous end joining pathway (orange), homologous recombination (red), inflammasome (dark green), NF-kB and interferon response (light green) are shown).

[0042] The factors and mechanisms that activate the DNA sensor pathway in cancer are currently poorly understood. It will be important to identify intratumoral DNA species, sensors, and pathways involved in IFN expression in different cancer types at all stages of the disease. In addition to being therapeutic targets for cancer, such factors may also have prognostic value. Currently, novel DNA sensor pathway agonists and antagonists are being developed and tested in preclinical studies. Such compounds will be useful in clarifying the role of the DNA sensor pathway in the pathology of cancer, autoimmunity, and potentially other diseases.

[0043] T cell receptor signaling genes The immune response against pathogens can be triggered at different stages: There is a physical barrier, such as the skin, to keep out invaders. If it is breached, a first, fast, non-specific response, the innate immune response, kicks in. If this is insufficient, the adaptive immune response is triggered, which is much more specific, but takes longer to develop if the pathogen is encountered for the first time. Lymphocytes become activated by interacting with activated antigen-presenting cells from the innate immune system. They are also involved in maintaining memory to respond faster the next time the same pathogen is encountered.

[0044] Once activated, lymphocytes become highly specific and effective, so they undergo negative selection for their ability to recognize themselves, a process known as central tolerance. Because not all self-antigens are expressed at the site of selection, mechanisms of peripheral tolerance have developed as well, such as TCR ligation in the absence of costimulation, expression of inhibitory co-receptors, and suppression by Tregs. Disturbances in the balance between activation and suppression can lead to autoimmune disorders, or immune deficiencies and cancer, respectively.

[0045] T cell activation can have different functional outcomes depending on the location and type of T cell involved: CD8+ T cells differentiate into cytotoxic effector cells, while CD4+ T cells can differentiate into Th1 (secreting IFNγ and promoting cell-mediated immunity) or Th2 (secreting IL4 / 5 / 13 and promoting B cell and humoral immunity). Differentiation into other more recently identified T cell subsets, such as Tregs, which have an inhibitory effect on immune activation, is also possible (see Mosenden R. and Tasken K., "Cyclic AMP-mediated immune regulation - Overview of mechanisms of action in T-cells", Cell Signal, Vol. 23, No. 6, pp. 1009-1016 (2011), in particular Figure 4, T cell activation and its modulation by PKA, and Tasken K. and Ruppelt A., "Negative regulation of T-cell receptor activation by the cAMP-PKA-Csk signalling pathway in T-cell lipid rafts", Front Biosci, Vol. 11, pp. 2929-2939 (2006)).

[0046] T cell activation occurs in naive and differentiated T cells. At the molecular level, events following ligation of the TCR with cognate antigen and crosstalk with signaling induced by costimulatory and co-inhibitory receptors determine whether T cells are activated or anergic. Triggering the TCR itself is not sufficient, leading to T cell anergy. The B7:CD28 family of costimulatory molecules plays a central role in controlling the activation state of T cells upon antigen stimulation (Torheim EA, "Immunity Leashed - Mechanisms of Regulation in the Human Immune System", Thesis for the degree of Philosophiae Doctor (PhD), The Biotechnology Centre of Ola, University of Oslo, Norway, 2009).

[0047] Activation occurs when the TCR on the T cell surface interacts with MHC-peptide complexes on APCs or target cells (Figure 2). An immunological synapse is formed, lipid rafts in the T cell membrane fuse, and Lck and Fyn are activated. These molecules phosphorylate ITAMs on the CD3 subunit of the TCR, which promotes the recruitment of Lck and the adjacent Zap-70. Lck phosphorylates and activates Zap-70, which in turn phosphorylates LAT, SLP76, and PLCγ1. LAT is a docking site for other signaling molecules and is essential for downstream TCR signaling. Grb2, Gads, PI3K, and NCK are recruited to LAT to propagate signaling including activation of RAS, PKC, mobilization of Ca2+, calcineurin, and polymerization of the actin cytoskeleton (Tasken et al., Front in Bioscience 11:2929-2939 (2006)). This ultimately leads to the activation of transcription factors of the NFkB, NFAT, AP1, and ATF families, resulting in the transcription of genes for immune activation (Mosenden et al., Cell Signalling 23:1009-1016 (2011); Tasken et al., Front in Bioscience 11:2929-2939 (2006)).

[0048] Both PKA and PDE4 regulated signaling intersect with TCR induced T cell activation, fine tuning its regulation with opposing effects (see Abrahamsen H. et al. "TCR- and CD28-mediated recruitment of phosphodiesterase 4 to lipid rafts potentiates TCR signaling", J Immunol, Vol. 173, pp. 4847-4848 (2004), in particular Figure 6 showing the opposing actions of PKA and PDE4 in TCR activation). The molecule that connects these effectors is cyclic AMP (cAMP), an intracellular second messenger of the action of extracellular ligands. Within T cells, cAMP mediates the effects of prostaglandins, adenosine, histamine, beta adrenergic agonists, neuropeptide hormones and beta endorphins. The binding of these extracellular molecules to GPCRs leads to conformational changes in the GPCRs and the release of stimulatory subunits, followed by activation of adenylyl cyclase (AC), which hydrolyzes ATP to cAMP (see Abrahamsen H. et al., 2004, ibid., Fig. 6). PKA is the main, though not the only, effector of cAMP signaling (see Mosenden R. and Tasken K., 2011, ibid. and Tasken K. and Ruppelt A., 2006, ibid.). At the functional level, increased levels of cAMP lead to reduced production of IFNγ and IL-2 in T cells (see Abrahamsen H. et al., 2004, ibid.). In addition to interfering with TCR activation, PKA has many more effects (see Torheim EA, 2009, ibid., Fig. 15).

[0049] In naive T cells, the hyperphosphorylated form of PAG targets Csk to lipid rafts. PKA targets Csk via the ezrin-EBP50-PAG scaffolding complex. By being specifically phosphorylated by PKA, Csk can negatively regulate Lck and Fyn to attenuate their activity and downregulate T cell activation (see Abrahamsen H. et al., 2004, ibid., fig. 6). Upon TCR activation, PAG is dephosphorylated and Csk is released from rafts. Dissociation of Csk is necessary for T cell activation to proceed. During the same time course, a Csk-G3BP complex appears to form, which sequesters Csk out of lipid rafts (see Mosenden R. and Tasken K., 2011, ibid. and Tasken K. and Ruppelt A., 2006, ibid.).

[0050] On the other hand, combined stimulation of TCR and CD28 mediates the recruitment of cyclic nucleotide phosphodiesterase PDE4 to lipid rafts, which enhances the degradation of cAMP (see Abrahamsen H. et al., 2004, ibid., Fig. 6), thereby blocking TCR-induced cAMP production and enhancing T cell immune responses. Upon TCR stimulation alone, recruitment of PDE4 is too low to sufficiently reduce cAMP levels, and therefore maximal T cell activation cannot occur (see Abrahamsen H. et al., 2004, ibid.).

[0051] Thus, by actively suppressing proximal TCR signaling, cAMP-PKA-Csk-mediated signaling appears to set a threshold for T cell activation. Recruitment of PDEs can block this suppression. Tissue- or cell type-specific regulation is achieved by expression of multiple isoforms of AC, PKA and PDE. As mentioned above, the balance between activation and suppression needs to be tightly regulated to prevent the development of autoimmune disorders, immune deficiencies and cancer.

[0052] PDE4D7-related genes Phosphodiesterases (PDEs) provide the only means to degrade the second messenger 3'-5'-cyclic AMP. As such, they play an important regulatory role. Therefore, abnormal changes in their expression, activity, and subcellular location may all be involved in the molecular pathology underlying certain disease states. Indeed, it has recently been shown that mutations in PDE genes are common in prostate cancer patients, which may increase cAMP signaling and predispose to prostate cancer. However, the presence of various expression profiles in different cell types coupled with a complex array of isoform variants within each PDE family makes it difficult to understand the relevance of abnormal changes in PDE expression and functionality during disease progression. Several studies have attempted to account for the complementation of PDEs in the prostate, all of which identified significant levels of PDE4 expression alongside other PDEs, which led to the development of the PDE4D7 biomarker (see Alves de Inda M. et al., "Validation of Cyclic Adenosine Monophosphate Phosphodiesterase-4D7 for its Independent Contribution to Risk Stratification in a Prostate Cancer Patient Cohort with Longitudinal Biological Outcomes," Eur Urol Focus, Vol. 4, No. 3, pp. 376-384, 2018). Since the PDE4D7 biomarker has proven to be a good predictor, it was assumed that the ability to identify markers that are highly correlated with the PDE47 biomarker may also be useful in predicting the prognosis of certain cancer subjects.

[0053] Based on the correlation between PDE4D7 expression and pathological features of the disease, we aimed to identify prognostic associations between PDE4D7 expression in patients' prostate tissues taken either by biopsy or surgery and clinically useful information related to the prognosis of individual patients. Clinically relevant endpoints or surrogate endpoints that significantly correlate with the occurrence of metastasis, cancer-specific death or all-cause mortality have generally been evaluated as prognostic biomarkers in cancer. The most appropriate rationale for using surrogate endpoints is when data on established clinical endpoints are not available or when the number of events in the data cohort is too limited for statistical data analysis. In the development of PDE4D7 prognostic biomarkers, either BCR (biochemical recurrence) progression-free survival or initiation of second-line treatment after surgery were evaluated as surrogate endpoints for metastasis and prostate cancer death. Using these specific endpoints, the number of associated events (e.g., >30% for BCR) in selected clinical cohorts was identified, which is particularly relevant for multivariate data analysis.

[0054] In the evaluation carried out, standard methods of multivariate analysis such as Cox regression and Kaplan-Meier survival analysis were selected to investigate the added value and independent value of the continuous and / or categorical "PDE4D7 score" compared to established prognostic clinical variables such as PSA and Gleason score (Alves de Inda, 2018). Risk models were built combining the "PDE4D7 score" with either pre- or post-operative clinical predictors of postoperative progression using logistic regression. The obtained models were then validated in multiple independent patient cohorts with Kaplan-Meier survival analysis and ROC curve analysis to predict progression-free survival after treatment (Alves de Inda, 2018).

[0055] Using such a strategy, the prognostic value of the PDE4D7 score in biopsies from retrospectively collected resected prostate tissue was validated in a cohort of patients consecutively managed from a single surgical center in the post-surgical setting (Alves de Inda, 2018). The patient population consisted of approximately 500 individuals in whom a longitudinal follow-up of both pathological and biological prognosis was performed. These clinical data were available for all patients and were collected during a median follow-up of 120 months after treatment. The "PDE4D7 score" was determined as described above and was validated in both univariate and multivariate analyses using available post-surgical covariates (i.e., pathological Gleason score, pT stage, surgical margin status, seminal vesicle invasion status, and lymph node invasion status) to adjust the multivariate analysis setting. In this example, biochemical progression-free survival after the primary intervention was set as the clinical endpoint to be evaluated. Univariate analysis of these clinical samples (Alves de Inda, 2018) showed an inverse correlation between PDE4D7 expression (in terms of the “PDE4D7 score”) and postoperative biological recurrence (HR per unit change=0.53; 95% CI 0.41-0.67; p<0.0001), robustly confirming previous data (Boettcher 2015; Boettcher, 2016). Also in a multivariate analysis with these clinical variables, the “PDE4D7 score” remained an independent and valid means for predicting clinical outcome (HR per unit change=0.56; 95% CI 0.43-0.73; p<0.0001). Furthermore, very similar results were obtained when evaluating the PDE4D7 score in multivariate analysis (HR=0.54 95% CI 0.42-0.69; p<0.0001) and the validated and clinically used risk model CAPRA-S. The CAPRA-S score is based on preoperative PSA and pathological parameters determined at surgery and was developed and validated in US and other populations to provide clinicians with information to help predict disease recurrence, including BCR, systemic progression, and PCSM.

[0056] Interestingly, when hazard ratios (HRs) were evaluated compared to the continuous PDE4D7 score, it was found that the risk increased linearly with decreasing PDE4D7 score, provided the score was between 2 and 5. However, the risk of postoperative progression increased sharply for PDE4D7 scores below 2 (Alves de Inda, 2018). This was also evident in Kaplan-Meier survival curves, where patients classified in the lowest PDE4D7 score category had the highest risk of disease recurrence. Using logistic regression analysis, the CAPRA-S score was combined with the continuous PDE4D7 score. Validation of this model using ROC curve analysis showed a significant improvement of AUC by 4–6% compared to CAPRA-S alone in both 2- and 5-year prediction of progression to BCR after treatment. Therefore, a combined Cox regression model combining CAPRA-S and PDE4D7 score was evaluated in Kaplan-Meier survival analysis and compared to CAPRA-S score categories alone. The results confirmed that using the combined "PDE4D7 & CAPRA-S" score model added value in risk prediction compared to using the clinical indicator CAPRA-S score alone (Alves de Inda, 2018).

[0057] After a diagnosis of prostate cancer, an accurate risk assessment must be performed before stratification to a defined first-line treatment can be performed. With this in mind, it was examined whether it is possible to translate the prognostic use of the "PDE4D7 score" in the preoperative setting, examining tumor tissue obtained from diagnostic needle biopsy samples (van Strijp 2018). In this study, needle biopsies were performed on 168 patients from a single diagnostic clinical center who underwent surgery as a first-line treatment. The minimum follow-up period for each patient was 60 months after this intervention. Clinical covariates used to adjust for the "PDE4D7 score" in the multivariate analysis were age at surgery, preoperative PSA, PSA density, biopsy Gleason score, percentage of tumor-positive biopsy cores, percentage of tumor in biopsy, and clinical cT stage. Among these, the utility of the “PDE4D7 score” and the combined “PDE4D7 & CAPRA” score compared with preoperative CAPRA score in Cox regression analysis of biochemical recurrence was evaluated (van Strijp 2018).

[0058] Evaluating this patient cohort, the "PDE4D7 score" was found to be inversely associated with BCR in multivariate analysis when adjusted for clinical variables (HR=0.43; 95% CI 0.29-0.63; p<0.0001) as well as clinical CAPRA score (HR=0.53; 95% CI 0.38-0.74; p=0.0001) (van Strijp 2018). Kaplan-Meier analysis showed that, as in the previous study, the "PDE4D7 score" category was significantly associated with BCR progression-free survival (log-rank p<0.0001) and second-line treatment-free survival (log-rank p=0.01) in the postoperative setting. A combined logistic regression model developed in the previous cohort was then employed (van Strijp 2018). This consisted of a combined “CAPRA & PDE4D7” score, which demonstrated that patients within the highest “CAPRA & PDE4D7” combined score category had virtually no risk of biochemical progression or transition to second-line treatment after surgery. This logistic regression model was also evaluated with receiver operating characteristic curve analysis to predict 5-year BCR after surgery, resulting in a 5% increase in AUC over the CAPRA score alone (AUC=0.82 vs. 0.77, respectively; p=0.004). Decision curve analysis of the combined “CAPRA & PDE4D7” score model confirmed the superior net benefit of using this combined score compared to either score alone, across all decision thresholds, to decide whether to intervene (e.g., surgery) based on the individual patient’s risk threshold of experiencing disease progression after surgery (van Strijp 2018).

[0059] Prediction of treatment prognosis is very complicated because many factors are involved in treatment outcome and disease recurrence. It is highly likely that important factors have not yet been identified, while the influence of other factors cannot be determined precisely. Currently, in order to improve response prediction and treatment selection, several clinicopathological indicators have been studied and applied in clinical practice, and some improvement has been achieved. However, there is a strong need to predict treatment response more accurately in order to increase the success rate of these treatments.

[0060] Gene selection Gene signatures, and combinations of these signatures with clinical parameters, have been identified and the resulting models are expected to show significant associations with mortality and thus improve prediction of the efficacy of these treatments1.

[0061] The identified immune defense response genes AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1, respectively, were identified as follows: 538 prostate cancer patients were treated with RP and prostate cancer tissues were archived with clinical (e.g., pathological Gleason grade group (pGGG), pathological status (pT stage)) and related prognostic parameters (e.g., biochemical recurrence (BCR), metastatic recurrence, prostate cancer-specific death (PCa death), salvage radiotherapy (SRT), salvage androgen deprivation therapy (SADT), chemotherapy (CTX)). For each of these patients, a PDE4D7 score was calculated and classified into four PDE4D7 score classes (Alves de Inda M. et al., 2018, supra). PDE4D7 score class 1 represents patient samples with the lowest expression levels of PDE4D7, while PDE4D7 score class 4 represents patient samples with the highest expression levels of PDE4D7. We then used RNASeq expression data (TPM - Transcripts Per Million) from 538 prostate cancer subjects to examine gene expression differences between PDE4D7 score classes 1 and 4. In particular, we determined whether the mean expression levels of approximately 20,000 protein-coding transcripts in patients with PDE4D7 score class 1 were more than twice as high as the mean expression levels in patients with PDE4D7 score class 4. This analysis resulted in 637 genes with a PDE4D7 score class 1 / PDE4D7 score class 4 ratio of more than 2 and a minimum mean expression of 1 TPM in each of the four PDE4D7 score classes. These 637 genes were then subjected to further molecular pathway analysis, resulting in various enriched annotation clusters. Annotation cluster #2 showed enrichment of 30 genes with functions in defense responses against viruses, negative regulation of viral genome replication, and type I interferon signaling (enrichment score: 10.8).Further heatmap analysis confirmed that these immune defense response genes were generally more highly expressed in samples from patients with PDE4D7 score class 1 than in samples from patients with PDE4D7 score class 4. The class of genes with functions in defense response against viruses, negative control of viral genome replication, and type I interferon signaling were further enriched to 61 genes by literature search to identify additional genes with the same molecular function. Further selection was performed from the 61 genes based on their combinatorial power to separate patients who died from prostate cancer from those who did not, resulting in a favorable set of 14 genes. The subcohort with low expression of these genes was found to be enriched in the number of events (metastasis, prostate cancer-specific death) compared to the total patient cohort (#538) and a subcohort of 151 patients who underwent salvage RT (SRT) after postoperative disease recurrence.

[0062] The identified T cell receptor signaling genes CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70 were identified as follows: 538 prostate cancer patients were treated with RP and prostate cancer tissues were archived with clinical (e.g., pathological Gleason grade group (pGGG), pathological status (pT stage)) and related prognostic parameters (e.g., biochemical recurrence (BCR), metastatic recurrence, prostate cancer-specific death (PCa death), salvage radiotherapy (SRT), salvage androgen deprivation therapy (SADT), chemotherapy (CTX)). For each of these patients, a PDE4D7 score was calculated and classified into four PDE4D7 score classes (Alves de Inda M. et al., 2018, supra). PDE4D7 score class 1 corresponds to the patient sample with the lowest expression level of PDE4D7, while PDE4D7 score class 4 corresponds to the patient sample with the highest level of PDE4D7 expression.Then, RNASeq expression data (TPM-Transcripts Per Million) of 538 prostate cancer subjects was examined for differential gene expression between PDE4D7 score class 1 and class 4.Specifically, for approximately 20,000 protein-coding transcripts, it was determined whether the average expression level of patients with PDE4D7 score class 1 is more than twice that of patients with PDE4D7 score class 4.This analysis resulted in 637 genes with a PDE4D7 score class 1 / PDE4D7 score class 4 ratio of more than 2, with the lowest average expression in each of the four PDE4D7 score classes being 1 TPM.These 637 genes were then subjected to further molecular pathway analysis, resulting in various enriched annotation clusters. Annotation cluster #6 showed enrichment in 17 genes (enrichment score: 5.9) with functions in primary immune deficiency and activation of T cell receptor signaling.Further heatmap analysis confirmed that these T cell receptor signaling genes were generally more highly expressed in samples from patients in PDE4D7 score class 1 than in samples from patients in PDE4D7 score class 4.

[0063] The identified PDE4D7-correlated genes ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2 were identified as follows: In RNAseq data on nearly 60,000 transcripts generated for 571 patients with prostate cancer, a wide range of genes were identified that correlated with the expression of the known biomarker PDE4D7. Correlation of the expression of these genes with PDE4D7 across the 571 samples was performed by Pearson correlation, with values ​​between 0 and 1 for positive correlation and between -1 and 0 for negative correlation. As input data for calculating the correlation coefficients, the PDE4D7 score (see Alves de Inda M. et al., 2018, supra) and the TPM gene expression values ​​determined by RNAseq for each gene of interest (see below) were used.

[0064] The largest negative correlation coefficient identified between the expression of any of the approximately 60,000 transcripts and the expression of PDE4D7 was −0.38, and the largest positive correlation coefficient identified between the expression of any of the approximately 60,000 transcripts and the expression of PDE4D7 was +0.56. Genes with correlations ranging from −0.31 to −0.38 and +0.41 to +0.56 were selected. A total of 77 transcripts matching these characteristics were identified. From these 77 transcripts, eight PDE4D7-correlated genes, ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, were selected and Cox regression combination models were iteratively validated in a subcohort of 186 patients who underwent salvage radiotherapy (SRT) due to postoperative biochemical recurrence. The clinical endpoint validated was prostate cancer-specific death after initiation of SRT. The boundary condition for selecting the 8 genes was given by the constraint that the p-value in the multivariate Cox regression was <0.1 for all genes retained in the model.

[0065] It has been shown herein that these genes are also valuable with respect to the prognosis of subjects with breast cancer, preferably invasive breast cancer. Furthermore, it has been shown that a subset of genes selected from immune defense response genes, T cell receptor signaling genes, and / or PDE4D7-correlated genes, such as 4, 5, 6, 7, 8, 9, 10 or more genes, also gives statistically significant results (see Figures 45-55 and data presented in Example 4). Thus, the present invention is not limited to the use of individual gene signatures (e.g., immune defense response genes, T cell receptor signaling genes, or PDE4D7-correlated genes) or combinations thereof, but also extends to the selection of genes, such as 4 or more genes from different gene groups as described herein.

[0066] Thus, in a first embodiment, the present invention provides a method for predicting prognosis in a subject with breast cancer, comprising: - determining or receiving a gene expression profile comprising expression levels of four or more genes selected from the first gene expression profile, the second gene expression profile and / or the third gene expression profile, - the first gene expression profile has one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; - the second gene expression profile has one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70: - the third gene expression profile comprises one or more, such as one, two, three, four, five, six, seven or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2; said gene expression profile being determined in a biological sample obtained from said subject; - determining a prognosis based on the first, second and / or third gene expression profile comprising expression levels of four or more genes, - said prediction is a favorable or unfavorable risk of breast cancer related mortality, locoregional recurrence and / or distant recurrence.

[0067] Optionally, the method further comprises providing a prediction of the prognosis to a medical caregiver or to the subject.

[0068] In one embodiment, the present invention provides a computer-implemented method for predicting prognosis of a subject with breast cancer, comprising: - receiving a determined gene expression profile comprising expression levels of four or more genes selected from the first, second and / or third gene expression profiles, - the first gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; - the second gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; - the third gene expression profile comprises one or more, such as one, two, three, four, five, six, seven or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2; said gene expression profile being determined in a biological sample obtained from said subject; - determining a prognosis based on the first, second and / or third gene expression profile comprising expression levels of said four or more genes, - said prediction is a favorable or unfavorable risk of breast cancer related mortality, locoregional recurrence and / or distant recurrence. Optionally, the method further comprises providing the prognostic prediction to a medical caregiver or to the subject.

[0069] In an alternative first aspect, the present invention provides a method of predicting prognosis in a subject with breast cancer, comprising determining or receiving a first gene expression profile for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1. and / or determining a first gene expression profile(s) in a biological sample obtained from the subject; and / or determining a first gene expression profile(s) of one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, of T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70. determining or receiving a second gene expression profile for each, wherein said second gene expression profile(s) is / are determined in a biological sample obtained from the subject; and / or determining or receiving a third gene expression profile for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7 or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, wherein said third gene expression profile(s) is / are determined in a biological sample obtained from the subject; determining a prediction of outcome based on the first, second, and / or third gene expression profile(s), wherein said prediction is favorable or unfavorable risk of breast cancer related mortality, locoregional recurrence and / or distant recurrence; and, as appropriate, providing a prediction of prognosis to a medical caregiver or the subject.

[0070] In the present invention, a breast cancer-related death is preferably a breast cancer-specific death.

[0071] The present invention describes the use of gene signatures to predict the prognosis of subjects with breast cancer.The gene signatures have gene expression profiles with 4 or more, for example 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, or 39 gene expression levels, where 4 or more gene expression levels are selected from immune defense response genes and / or T cell receptor signaling genes and / or PDE4D7-related genes.Thus, in one embodiment, 4 or more genes can be selected from immune defense response genes.In one embodiment, 4 or more genes can be selected from T cell receptor signaling genes.In one embodiment, 4 or more genes can be selected from PDE4D7-related genes. In one embodiment, the four or more genes include one or more immune defense response genes and one or more T cell receptor signaling genes. In one embodiment, the four or more genes include one or more immune defense response genes and one or more PDE4D7-correlated genes. In one embodiment, the four or more genes include one or more PDE4D7-correlated genes and one or more T cell receptor signaling genes. In one embodiment, the four or more genes include one or more immune defense response genes, one or more T cell receptor signaling genes, and one or more PDE4D7-correlated genes.

[0072] In one embodiment of the method of the invention, the gene expression profile comprises expression levels of two or more immune defense response genes, two or more T cell receptor signaling genes, and two or more PDE4D7-correlated genes. In one embodiment of the method of the invention, the gene expression profile comprises expression levels of four or more immune defense response genes, four or more T cell receptor signaling genes, or four or more PDE4D7-correlated genes.

[0073] With respect to the biological processes mentioned above, three immune system-related gene signatures were selected, including the genes listed in Tables 1 to 3. The relevance of these signatures to prostate cancer survival prediction has been previously shown.

[0074] [Table 1]

[0075] [Table 2]

[0076] [Table 3]

[0077] AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1, CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, ZAP70, ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, each individually, one or more per gene panel, in combination with one or more per gene panel, or all combined, can predict prognosis in subjects with breast cancer.

[0078] The term "ABCC5" refers to the nucleotide sequence as set forth in SEQ ID NO: 1 or SEQ ID NO: 2, corresponding to the sequence of the human ATP-binding cassette subfamily C member 5 gene (Ensembl: ENSG00000114770), e.g. as defined in NCBI Reference Sequence NM_001023587.2 or NCBI Reference Sequence NM_005688.3, ​​in particular the NCBI Reference Sequence for the ABCC5 transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 3 or SEQ ID NO: 4, corresponding to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_001018881.1 and NCBI Protein Accession Reference Sequence NP_005679, which encodes the ABCC5 polypeptide.

[0079] The term "ABCC5" also refers to a nucleotide sequence that exhibits high homology to ABCC5, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:1 or SEQ ID NO:2, or an nucleotide sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:3 or SEQ ID NO:4. The present invention includes an amino acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:3 or SEQ ID NO:4, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:1 or SEQ ID NO:2.

[0080] The term "AIM2" refers to the nucleotide sequence as set forth in SEQ ID NO: 5, which corresponds to the sequence of the Absent in Melanoma 2 gene (Ensembl: ENSG00000163568), e.g. as defined in the NCBI Reference Sequence NM_004833, in particular the sequence of the NCBI Reference Sequence of the AIM2 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 6, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_004824, which encodes the AIM2 polypeptide.

[0081] The term "AIM2" also refers to a nucleotide sequence that exhibits high homology to AIM2, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:5, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:6. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:6, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:5.

[0082] The term "APOBEC3A" refers to the nucleotide sequence as set forth in SEQ ID NO: 7, which corresponds to the sequence of the Apolipoprotein B mRNA editing enzyme catalytic subunit 3A gene (Ensembl: ENSG00000128383), e.g. the sequence defined in the NCBI reference sequence NM_145699, in particular the sequence of the NCBI reference sequence of the APOBEC3A transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 8, which corresponds to the protein sequence defined in the NCBI protein accession reference sequence NP663745, which encodes the APOBEC3A polypeptide.

[0083] The term "APOBEC3A" also refers to a nucleotide sequence that exhibits high homology to APOBEC3A, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:7, or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:8. This includes an amino acid sequence, or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:8, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:7.

[0084] The term "CD2" refers to the nucleotide sequence as set forth in SEQ ID NO: 9, which corresponds to the sequence of the Cluster of Differentiation 2 gene (Ensembl: ENSG00000116824), e.g. as defined in the NCBI Reference Sequence NM_001767, in particular the sequence of the NCBI Reference Sequence for the CD2 transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 10, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_001758, which encodes the CD2 polypeptide.

[0085] The term "CD2" also refers to a nucleotide sequence that exhibits high homology to CD2, e.g., a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:9, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:10. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:10, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:9.

[0086] The term "CD247" refers to the nucleotide sequence as set forth in SEQ ID NO: 11 or SEQ ID NO: 12, which corresponds to the sequence of the Cluster of Differentiation 247 gene (Ensembl: ENSG00000198821), e.g. as defined in NCBI Reference Sequence NM_000734 or NCBI Reference Sequence NM_198053, in particular the NCBI Reference Sequence of the CD247 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 13 or SEQ ID NO: 14, which corresponds to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_000725 and NCBI Protein Accession Reference Sequence NP_932170, which encodes a CD247 polypeptide.

[0087] The term "CD247" also refers to a nucleotide sequence exhibiting high homology to CD247, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:11 or SEQ ID NO:12, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:13 or SEQ ID NO:14. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:13 or SEQ ID NO:14, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:11 or SEQ ID NO:12.

[0088] The term "CD28" refers to the nucleotide sequence as set forth in SEQ ID NO: 15 or SEQ ID NO: 16, which corresponds to the sequence of the Cluster of Differentiation 28 gene (Ensembl: ENSG00000178562), e.g. as defined in NCBI Reference Sequence NM_006139 or NCBI Reference Sequence NM_001243078, in particular the NCBI Reference Sequence of the CD28 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 17 or SEQ ID NO: 18, which corresponds to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_006130 and NCBI Protein Accession Reference Sequence NP_001230007, which encodes the CD28 polypeptide.

[0089] The term "CD28" also refers to a nucleotide sequence that exhibits high homology to CD28, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:15 or SEQ ID NO:16, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:17 or SEQ ID NO:18. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:17 or SEQ ID NO:18, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:15 or SEQ ID NO:16.

[0090] The term "CD3E" refers to the nucleotide sequence as set forth in SEQ ID NO: 19, which corresponds to the sequence of the Cluster of Differentiation 3E gene (Ensembl: ENSG00000198851), e.g. as defined in the NCBI Reference Sequence NM_000733, in particular the sequence of the NCBI Reference Sequence for the CD3E transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 20, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_000724, which encodes the CD3E polypeptide.

[0091] The term "CD3E" also refers to a nucleotide sequence that exhibits high homology to CD3E, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 19, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 20. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:20, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:19.

[0092] The term "CD3G" refers to the nucleotide sequence as set forth in SEQ ID NO: 21, which corresponds to the sequence of the Cluster of Differentiation 3G gene (Ensembl: ENSG00000160654), e.g. as defined in the NCBI Reference Sequence NM_000073, in particular the sequence of the NCBI Reference Sequence of the CD3G transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 22, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_000064, which encodes the CD3G polypeptide.

[0093] The term "CD3G" also refers to a nucleotide sequence that exhibits high homology to CD3G, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:21, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:22. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:22, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:21.

[0094] The term "CD4" refers to the nucleotide sequence as set forth in SEQ ID NO: 23, which corresponds to the sequence of the Cluster of Differentiation 4 gene (Ensembl: ENSG00000010610), e.g. as defined in the NCBI Reference Sequence NM_000616, in particular the sequence of the NCBI Reference Sequence for the CD4 transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 24, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_000607, which codes for the CD4 polypeptide.

[0095] The term "CD4" also refers to a nucleotide sequence that exhibits high homology to CD4, e.g., a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:23, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:24. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:24, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:23.

[0096] The term "CIAO1" refers to the nucleotide sequence as set forth in SEQ ID NO: 25, which corresponds to the sequence of the Cytoplasmic Iron-Sulfur Assembly Component 1 gene (Ensembl: ENSG00000144021), e.g. as defined in the NCBI Reference Sequence NM_004804, in particular the NCBI Reference Sequence of the CIAO1 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 26, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_663745, which encodes the CIAO1 polypeptide.

[0097] The term "CIAO1" also refers to a nucleotide sequence that exhibits high homology to CIAO1, for example, a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:25, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:26. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:26, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:25.

[0098] The term "CSK" refers to the nucleotide sequence as set forth in SEQ ID NO: 27, which corresponds to the sequence of the C-terminal Src kinase gene (Ensembl: ENSG00000103653), e.g. as defined in NCBI Reference Sequence NM_004383, in particular the sequence of the NCBI Reference Sequence for the CSK transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 28, which corresponds to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_004374, which encodes a CSK polypeptide.

[0099] The term "CSK" also refers to a nucleotide sequence that exhibits high homology to CSK, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:27, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:28. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:28, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:27.

[0100] The term "CUX2" refers to the nucleotide sequence as set forth in SEQ ID NO: 29, which corresponds to the sequence of the human Cut-like homeobox 2 gene (Ensembl: ENSG00000111249), e.g. as defined in NCBI Reference Sequence NM_015267.3, in particular the sequence of the NCBI Reference Sequence for the CUX2 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 30, which corresponds to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_056082.2, which codes for a CUX2 polypeptide.

[0101] The term "CUX2" also refers to a nucleotide sequence that exhibits high homology to CUX2, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:29, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:30. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:30, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:29.

[0102] The term "DDX58" refers to the nucleotide sequence as set forth in SEQ ID NO: 31, which corresponds to the sequence of the DExD / H-box helicase 58 gene (Ensembl: ENSG00000107201), e.g. as defined in the NCBI Reference Sequence NM_014314, in particular the sequence of the NCBI Reference Sequence for the DDX58 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 32, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_055129, which encodes the DDX58 polypeptide.

[0103] The term "DDX58" also refers to a nucleotide sequence that exhibits high homology to DDX58, for example, a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:31, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:32. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:32, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:31.

[0104] The term "DHX9" refers to the nucleotide sequence as set forth in SEQ ID NO: 33, which corresponds to the sequence of the DExD / H-box helicase 9 gene (Ensembl: ENSG00000135829), e.g. as defined in NCBI Reference Sequence NM_001357, in particular the sequence of the NCBI Reference Sequence for the DHX9 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 34, which corresponds to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_001348, which encodes the DHX9 polypeptide.

[0105] The term "DHX9" also refers to a nucleotide sequence that exhibits high homology to DHX9, for example, a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:33, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:34. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:34, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:33.

[0106] The term "EZR" refers to the nucleotide sequence as set forth in SEQ ID NO: 35, which corresponds to the sequence of the Ezrin gene (Ensembl: ENSG00000092820), e.g., the sequence defined in the NCBI reference sequence NM_003379, in particular the sequence of the NCBI reference sequence of the EZR transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 36, which corresponds to the protein sequence defined in the NCBI protein accession reference sequence NP_003370 that encodes the EZR polypeptide.

[0107] The term "EZR" also refers to a nucleotide sequence that exhibits high homology to EZR, for example, a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:35, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:36. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:36, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:35.

[0108] The term "FYN" refers to the gene of the FYN proto-oncogene (Ensembl: ENSG00000010810), for example the sequence defined in the NCBI Reference Sequence NM_002037 or the NCBI Reference Sequence NM_153047 or the NCBI Reference Sequence NM_153048, in particular the nucleotide sequence as defined in SEQ ID NO: 37 or SEQ ID NO: 38 or SEQ ID NO: 39, which corresponds to the sequence of the NCBI Reference Sequence of the FYN transcript shown above, and it also relates to the corresponding amino acid sequence as defined for example in SEQ ID NO: 40 or SEQ ID NO: 41 or SEQ ID NO: 42, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_002028 and the NCBI Protein Accession Reference Sequence NP_694592 and the NCBI Protein Accession Reference Sequence XP_005266949, which encodes the FYN polypeptide.

[0109] The term "FYN" also refers to a nucleotide sequence that exhibits high homology to FYN, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:37, SEQ ID NO:38 or SEQ ID NO:39, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:40, SEQ ID NO:41 or SEQ ID NO:42. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:40, SEQ ID NO:41 or SEQ ID NO:42, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:37, SEQ ID NO:38 or SEQ ID NO:39.

[0110] The term "IFI16" refers to the nucleotide sequence as set forth in SEQ ID NO: 43, which corresponds to the sequence of the interferon gamma inducible protein 16 gene (Ensembl: ENSG00000163565), e.g. as defined in the NCBI reference sequence NM_005531, in particular the sequence of the NCBI reference sequence of the IFI16 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 44, which corresponds to the protein sequence defined in the NCBI protein accession reference sequence NP_005522, which encodes the IFI16 polypeptide.

[0111] The term "IFI16" also refers to a nucleotide sequence that exhibits high homology to IFI16, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 43, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 44. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:44, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:43.

[0112] The term "IFIH1" refers to the nucleotide sequence as set forth in SEQ ID NO: 45, which corresponds to the sequence of the gene Interferon Induced With Helicase C Domain 1 (Ensembl: ENSG00000115267), e.g. as defined in the NCBI Reference Sequence NM_022168, in particular the sequence of the NCBI Reference Sequence for the IFIH1 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 46, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_071451, which codes for the IFIH1 polypeptide.

[0113] The term "IFIH1" also refers to a nucleotide sequence that exhibits high homology to IFIH1, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:45, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:46. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:46, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:45.

[0114] The term "IFIT1" refers to the nucleotide sequence as set forth in SEQ ID NO: 47 or SEQ ID NO: 48, which corresponds to the sequence of the gene Interferon Induced Protein With Tetratricopeptide Repeats 1 (Ensembl: ENSG00000185745), e.g. as defined in the NCBI Reference Sequence NM_001270929 or the NCBI Reference Sequence NM_001548.5, in particular the NCBI Reference Sequence of the IFIT1 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 49 or SEQ ID NO: 50, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_001257858 and the NCBI Protein Accession Reference Sequence NP_001539, which encodes the IFIT1 polypeptide.

[0115] The term "IFIT1" also refers to a nucleotide sequence that exhibits high homology to IFIT1, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 47 or SEQ ID NO: 48, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 49 or SEQ ID NO: 50. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:49 or SEQ ID NO:50, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:47 or SEQ ID NO:48.

[0116] The term "IFIT3" refers to the nucleotide sequence as set forth in SEQ ID NO: 51, which corresponds to the sequence of the gene Interferon Induced Protein With Tetratricopeptide Repeats 3 (Ensembl: ENSG00000119917), e.g. as defined in the NCBI Reference Sequence NM_001031683, in particular the NCBI Reference Sequence of the IFIT3 transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 52, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_001026853 that codes for the IFIT3 polypeptide.

[0117] The term "IFIT3" also refers to a nucleotide sequence that exhibits high homology to IFIT3, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:51, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:52. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:52, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:51.

[0118] The term "KIAA1549" refers to the nucleotide sequence as set forth in SEQ ID NO: 53 or SEQ ID NO: 54, which corresponds to the sequence of the human KIAA1549 gene (Ensembl: ENSG00000122778), e.g. as defined in NCBI Reference Sequence NM_020910 or NCBI Reference Sequence NM_001164665, in particular the NCBI Reference Sequence of the KIAA1549 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 55 or SEQ ID NO: 56, which corresponds to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_065961 and NCBI Protein Accession Reference Sequence NP_001158137, which encodes a KIAA1549 polypeptide.

[0119] The term "KIAA1549" also refers to a nucleotide sequence that exhibits high homology to KIAA1549, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:53 or SEQ ID NO:54, or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:55 or SEQ ID NO:56. This includes an amino acid sequence, or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:55 or SEQ ID NO:56, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:53 or SEQ ID NO:54.

[0120] The term "LAT" refers to the nucleotide sequence as set forth in SEQ ID NO: 57 or SEQ ID NO: 58, which corresponds to the sequence of the NCBI Reference Sequence of the LAT transcript shown above, of the gene Linker For Activation Of T-Cells (Ensembl: ENSG00000213658), e.g. as defined in NCBI Reference Sequence NM_001014987 or NCBI Reference Sequence NM_014387, in particular to the corresponding amino acid sequence as set forth in SEQ ID NO: 59 or SEQ ID NO: 60, which corresponds to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_001014987 and NCBI Protein Accession Reference Sequence NP_055202, which encodes a LAT polypeptide.

[0121] The term "LAT" also refers to a nucleotide sequence that exhibits high homology to LAT, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:57 or SEQ ID NO:58, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:59 or SEQ ID NO:60. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:59 or SEQ ID NO:60, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:57 or SEQ ID NO:58.

[0122] The term "LCK" refers to the nucleotide sequence as set forth in SEQ ID NO: 61, which corresponds to the sequence of the LCK Proto-Oncogene gene (Ensembl: ENSG00000182866), e.g. as defined in the NCBI Reference Sequence NM_005356, in particular the NCBI Reference Sequence for the LCK transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 62, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_005347, which encodes the LCK polypeptide.

[0123] The term "LCK" also refers to a nucleotide sequence that exhibits high homology to LCK, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:61, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:62. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:62, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:61.

[0124] The term "LRRFIP1" refers to the sequence of the LRR-binding FLII interacting protein 1 gene (Ensembl: ENSG00000124831), e.g. the sequence defined in NCBI Reference Sequence NM_004735 or NCBI Reference Sequence NM_001137550 or NCBI Reference Sequence NM_001137553 or NCBI Reference Sequence NM_001137552, in particular the nucleotide sequence as set forth in SEQ ID NO: 63 or SEQ ID NO: 64 or SEQ ID NO: 65 or SEQ ID NO: 66, which corresponds to the sequence of the NCBI Reference Sequence of the LRRFIP1 transcript shown above. It also relates to the corresponding amino acid sequence as set forth in, for example, SEQ ID NO:67 or SEQ ID NO:68 or SEQ ID NO:69 or SEQ ID NO:70, which corresponds to the protein sequence defined in NCBI protein accession reference sequence NP_004726 and NCBI protein accession reference sequence NP_001131022 and NCBI protein accession reference sequence NP_001131025 and NCBI protein accession reference sequence NP_001131024 encoding the LRRFIP1 polypeptide.

[0125] The term "LRRFIP1" also refers to a nucleotide sequence that exhibits high homology to LRRFIP1, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:63 or SEQ ID NO:64 or SEQ ID NO:65 or SEQ ID NO:66, or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:67 or SEQ ID NO:68 or SEQ ID NO:69 or SEQ ID NO:70. or a nucleic acid sequence encoding an amino acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:67 or SEQ ID NO:68 or SEQ ID NO:69 or SEQ ID NO:70, or an amino acid sequence encoded by a nucleic acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:63 or SEQ ID NO:64 or SEQ ID NO:65 or SEQ ID NO:66.

[0126] The term "MYD88" refers to the MYD88 innate immune signaling adaptor gene (Ensembl: ENSG00000172936), e.g., NCBI Reference Sequence NM_001172567 or NCBI Reference Sequence NM_001172568 or NCBI Reference Sequence MYD88 polypeptide.

[0127] The term "MYD88" also refers to a nucleotide sequence that exhibits high homology to MYD88, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:71 or SEQ ID NO:72 or SEQ ID NO:73 or SEQ ID NO:74 or SEQ ID NO:75, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:76 or SEQ ID NO:77 or SEQ ID NO:78 or SEQ ID NO:79 or SEQ ID NO:80. or a nucleic acid sequence encoding an amino acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:76 or SEQ ID NO:77 or SEQ ID NO:78 or SEQ ID NO:79 or SEQ ID NO:80, or an amino acid sequence encoded by a nucleic acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:71 or SEQ ID NO:72 or SEQ ID NO:73 or SEQ ID NO:74 or SEQ ID NO:75.

[0128] The term "OAS1" refers to the nucleotide sequence as set forth in SEQ ID NO: 81 or SEQ ID NO: 82 or SEQ ID NO: 83 or SEQ ID NO: 84, which corresponds to the sequence of the NCBI Reference Sequence of the OAS1 transcript shown above, for example the sequence defined in the NCBI Reference Sequence NM_001320151 or the NCBI Reference Sequence NM_002534 or the NCBI Reference Sequence NM_001032409 or the NCBI Reference Sequence NM_016816. "OAS1" refers to a sequence corresponding to the protein sequence defined in NCBI protein accession reference sequence NP_001307080 and NCBI protein accession reference sequence NP_002525 and NCBI protein accession reference sequence NP_001027581 and NCBI protein accession reference sequence NP_058132 encoding an OAS1 polypeptide, and also to the corresponding amino acid sequence as set forth in, for example, SEQ ID NO:85 or SEQ ID NO:86 or SEQ ID NO:87 or SEQ ID NO:88.

[0129] The term "OAS1" also refers to a nucleotide sequence that exhibits high homology to OAS1, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:81 or SEQ ID NO:82 or SEQ ID NO:83 or SEQ ID NO:84, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:85 or SEQ ID NO:86 or SEQ ID NO:87 or SEQ ID NO:88. or a nucleic acid sequence encoding an amino acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:85 or SEQ ID NO:86 or SEQ ID NO:87 or SEQ ID NO:88, or an amino acid sequence encoded by a nucleic acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:81 or SEQ ID NO:82 or SEQ ID NO:83 or SEQ ID NO:84.

[0130] The term "PAG1" refers to the nucleotide sequence as set forth in SEQ ID NO: 89, which corresponds to the sequence of the Phosphoprotein Membrane Anchor With Glycosphingolipid Microdomains 1 gene (Ensembl: ENSG00000076641), e.g. as defined in the NCBI Reference Sequence NM_018440, in particular the sequence of the NCBI Reference Sequence for the PAG1 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 90, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_060910, which encodes the PAG1 polypeptide.

[0131] The term "PAG1" also refers to a nucleotide sequence that exhibits high homology to PAG1, for example, a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:89, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:90. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence set forth in SEQ ID NO:90, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence set forth in SEQ ID NO:89.

[0132] The term "PDE4D" refers to the gene for human phosphodiesterase 4D (Ensembl: ENSG00000113448), e.g., NCBI Reference Sequence NM_001104631 or NCBI Reference Sequence NM_001349242 or NCBI Reference Sequence NM_001197218 or NCBI Reference Sequence NM_006203 or NCBI Reference Sequence NM_001197221 or NCBI Reference Sequence NM_001197220 or NCBI Reference Sequence NM_00119722 3 or NCBI Reference Sequence NM_001165899 or the sequence defined in NCBI Reference Sequence NM_001165899, in particular the nucleotide sequence as set forth in SEQ ID NO: 91 or SEQ ID NO: 92 or SEQ ID NO: 93 or SEQ ID NO: 94 or SEQ ID NO: 95 or SEQ ID NO: 96 or SEQ ID NO: 97 or SEQ ID NO: 98 or SEQ ID NO: 99, which corresponds to the sequence of the NCBI Reference Sequence for the PDE4D transcript shown above, and which also refers to the NCBI It also relates to the corresponding amino acid sequences as defined in protein accession reference sequence NP_001098101 and NCBI protein accession reference sequence NP_001336171 and NCBI protein accession reference sequence NP_001184147 and NCBI protein accession reference sequence NP_006194 and NCBI protein accession reference sequence NP_001184150 and NCBI protein accession reference sequence NP_001184149 and NCBI protein accession reference sequence NP_001184152 and NCBI protein accession reference sequence NP_001159371 and NCBI protein accession reference sequence NP_001184148, e.g. as set forth in SEQ ID NO:100 or SEQ ID NO:101 or SEQ ID NO:102 or SEQ ID NO:103 or SEQ ID NO:104 or SEQ ID NO:105 or SEQ ID NO:106 or SEQ ID NO:107 or SEQ ID NO:108.

[0133] The term "PDE4D" also refers to a nucleotide sequence that exhibits high homology to PDE4D, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:91 or SEQ ID NO:92 or SEQ ID NO:93 or SEQ ID NO:94 or SEQ ID NO:95 or SEQ ID NO:96 or SEQ ID NO:97 or SEQ ID NO:98 or SEQ ID NO:99, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:100 or SEQ ID NO:101 or SEQ ID NO:102 or SEQ ID NO:103 or SEQ ID NO:104 or SEQ ID NO:105 or SEQ ID NO:106 or SEQ ID NO:107 or SEQ ID NO:108. or a nucleic acid sequence encoding an amino acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:100 or SEQ ID NO:101 or SEQ ID NO:102 or SEQ ID NO:103 or SEQ ID NO:104 or SEQ ID NO:105 or SEQ ID NO:106 or SEQ ID NO:107 or SEQ ID NO:108, or an amino acid sequence encoded by a nucleic acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:91 or SEQ ID NO:92 or SEQ ID NO:93 or SEQ ID NO:94 or SEQ ID NO:95 or SEQ ID NO:96 or SEQ ID NO:97 or SEQ ID NO:98 or SEQ ID NO:99.

[0134] The term "PRKACA" refers to the nucleotide sequence as set forth in SEQ ID NO: 109 or SEQ ID NO: 110 of the gene of protein kinase cAMP-activated catalytic subunit alpha (Ensembl: ENSG00000072062), for example the sequence defined in the NCBI Reference Sequence NM_002730 or in the NCBI Reference Sequence NM_207518, in particular the sequence of the NCBI Reference Sequence of the PRKACA transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 111 or SEQ ID NO: 112, for example, corresponding to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_002721 and the NCBI Protein Accession Reference Sequence NP_997401 which encodes a PRKACA polypeptide.

[0135] The term "PRKACA" also refers to a nucleotide sequence that exhibits high homology to PRKACA, for example, a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:109 or SEQ ID NO:110, or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:111 or SEQ ID NO:112. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:111 or SEQ ID NO:112, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:109 or SEQ ID NO:110.

[0136] The term "PRKACB" refers to the protein kinaseGene for cAMP-activated catalytic subunit beta (Ensembl: ENSG00000142875), e.g. NCBI Reference Sequence NM_002731 or NCBI Reference Sequence NM_182948 or NCBI Reference Sequence NM_001242860 or NCBI Reference Sequence NM_001242859 or NCBI Reference Sequence NM_001242858 or NCBI Reference Sequence NM_001242862 or NCBI Reference Sequence NM_001242861 or NCBI Reference Sequence NM_001300915 or NCBI Reference Sequence NM_207578 or NCBI Reference Sequence NM_001300915 The term "PRKACB" refers to the sequence as defined in NCBI Reference Sequence NM_001242857 or NCBI Reference Sequence NM_001300917, specifically the nucleotide sequence as set forth in SEQ ID NO:113 or SEQ ID NO:114 or SEQ ID NO:115 or SEQ ID NO:116 or SEQ ID NO:117 or SEQ ID NO:118 or SEQ ID NO:119 or SEQ ID NO:120 or SEQ ID NO:121 or SEQ ID NO:122 or SEQ ID NO:123, which corresponds to the sequence of the NCBI Reference Sequence for the PRKACB transcript shown above, and which also refers to the NCBI Protein Accession Number (CAI) encoding the PRKACB polypeptide. Reference sequence NP_002722 and NCBI protein accession reference sequence NP_891993 and NCBI protein accession reference sequence NP_001229789 and NCBI protein accession reference sequence NP_001229788 and NCBI protein accession reference sequence NP_001229787 and NCBI protein accession reference sequence NP_001229791 and NCBI protein accession reference sequence NP_001229790 and NCBI protein accession reference sequence NP_0012 87844 and NCBI protein accession reference sequence NP_997461 and NCBI protein accession reference sequence NP_001229786 and NCBI protein accession reference sequence NP_001287846, e.g. as set forth in SEQ ID NO:124 or SEQ ID NO:125 or SEQ ID NO:126 or SEQ ID NO:127 or SEQ ID NO:128 or SEQ ID NO:129 or SEQ ID NO:130 or SEQ ID NO:131 or SEQ ID NO:132 or SEQ ID NO:133 or SEQ ID NO:134.

[0137] The term "PRKACB" also refers to a nucleotide sequence that exhibits high homology to PRKACB, such as at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95% homology to the sequence as set forth in SEQ ID NO:113 or SEQ ID NO:114 or SEQ ID NO:115 or SEQ ID NO:116 or SEQ ID NO:117 or SEQ ID NO:118 or SEQ ID NO:119 or SEQ ID NO:120 or SEQ ID NO:121 or SEQ ID NO:122 or SEQ ID NO:123. , 96%, 97%, 98% or 99% identical to, or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to, a nucleic acid sequence as set forth in SEQ ID NO:124 or SEQ ID NO:125 or SEQ ID NO:126 or SEQ ID NO:127 or SEQ ID NO:128 or SEQ ID NO:129 or SEQ ID NO:130 or SEQ ID NO:131 or SEQ ID NO:132 or SEQ ID NO:133 or SEQ ID NO:134. or an amino acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:124 or SEQ ID NO:125 or SEQ ID NO:126 or SEQ ID NO:127 or SEQ ID NO:128 or SEQ ID NO:129 or SEQ ID NO:130 or SEQ ID NO:131 or SEQ ID NO:132 or SEQ ID NO:133 or SEQ ID NO:134. or a nucleic acid sequence that encodes an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:113 or SEQ ID NO:114 or SEQ ID NO:115 or SEQ ID NO:116 or SEQ ID NO:117 or SEQ ID NO:118 or SEQ ID NO:119 or SEQ ID NO:120 or SEQ ID NO:121 or SEQ ID NO:122 or SEQ ID NO:123.

[0138] The term "PTPRC" refers to the nucleotide sequence as set forth in SEQ ID NO: 135 or SEQ ID NO: 136, corresponding to the sequence of the Protein Tyrosine Phosphatase Receptor Type C gene (Ensembl: ENSG00000081237), e.g. as defined in NCBI Reference Sequence NM__002838 or NCBI Reference Sequence NM_080921, in particular the NCBI Reference Sequence for the PTPRC transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 137 or SEQ ID NO: 138, corresponding to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_002829 and NCBI Protein Accession Reference Sequence NP_563578, which encodes a PTPRC polypeptide.

[0139] The term "PTPRC" also refers to a nucleotide sequence that exhibits high homology to PTPRC, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 135 or SEQ ID NO: 136, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 137 or SEQ ID NO: 138. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:137 or SEQ ID NO:138, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:135 or SEQ ID NO:136.

[0140] The term "RAP1GAP2" refers to the nucleotide sequence as set forth in SEQ ID NO: 139 or SEQ ID NO: 140 or SEQ ID NO: 141 of the gene of human RAP1 GTPase activating protein 2 (ENSG00000132359), for example the sequence defined in NCBI Reference Sequence NM_015085 or NCBI Reference Sequence NM_001100398 or NCBI Reference Sequence NM_001330058, in particular the sequence of the NCBI Reference Sequence of the RAP1GAP2 transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 142 or SEQ ID NO: 143 or SEQ ID NO: 144, for example, corresponding to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_055900 and NCBI Protein Accession Reference Sequence NP_001093868 and NCBI Protein Accession Reference Sequence NP_001316987 which encodes a RAP1GAP2 polypeptide.

[0141] The term "RAP1GAP2" also refers to a nucleotide sequence that exhibits high homology to RAP1GAP2, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 139 or SEQ ID NO: 140 or SEQ ID NO: 141, or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 142 or SEQ ID NO: 143 or SEQ ID NO: 144. The present invention includes a nucleic acid sequence encoding an amino acid sequence, or an amino acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:142 or SEQ ID NO:143 or SEQ ID NO:144, or an amino acid sequence encoded by a nucleic acid sequence which is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:139 or SEQ ID NO:140 or SEQ ID NO:141.

[0142] The term "SLC39A11" refers to the nucleotide sequence as set forth in SEQ ID NO: 145 or SEQ ID NO: 146, which corresponds to the sequence of the human solute carrier family 39 member 11 gene (Ensembl: ENSG00000133195), e.g. as defined in NCBI Reference Sequence NM_139177 or NCBI Reference Sequence NM_001352692, in particular the NCBI Reference Sequence of the SLC39A11 transcript shown above, and it also relates to the corresponding amino acid sequence, e.g. as set forth in SEQ ID NO: 147 or SEQ ID NO: 148, which corresponds to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_631916 and NCBI Protein Accession Reference Sequence NP_001339621, which encodes the SLC39A11 polypeptide.

[0143] The term "SLC39A11" also refers to a nucleotide sequence that exhibits high homology to SLC39A11, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 145 or SEQ ID NO: 146, or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 147 or SEQ ID NO: 148. This includes a nucleic acid sequence encoding an amino acid sequence or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:147 or SEQ ID NO:148, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:145 or SEQ ID NO:146.

[0144] The term "TDRD1" refers to the gene for Tudor domain-containing 1 in humans (Ensembl: ENSG00000095627), for example, the sequence defined in NCBI reference sequence NM_198795, specifically, the nucleotide sequence as defined in SEQ ID NO: 149 corresponding to the sequence of the NCBI reference sequence of the TDRD1 transcript shown above, and it also relates to the corresponding amino acid sequence as defined in SEQ ID NO: 150, for example, corresponding to the protein sequence defined in NCBI protein accession reference sequence NP_942090 encoding the TDRD1 polypeptide.

[0145] The term "TDRD1" also includes a nucleotide sequence showing high homology to TDRD1, for example, a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as defined in SEQ ID NO: 149, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as defined in SEQ ID NO: 150, or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as defined in SEQ ID NO: 150, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as defined in SEQ ID NO: 149.

[0146] The term "TLR8" refers to the nucleotide sequence as set forth in SEQ ID NO: 151 or SEQ ID NO: 152, corresponding to the sequence of the Toll-like receptor 8 gene (Ensembl: ENSG00000101916), e.g. as defined in NCBI Reference Sequence NM_138636 or NCBI Reference Sequence NM_016610, in particular the NCBI Reference Sequence of the TLR8 transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 153 or SEQ ID NO: 154, corresponding to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_619542 and NCBI Protein Accession Reference Sequence NP_057694, which encodes a TLR8 polypeptide.

[0147] The term "TLR8" also refers to a nucleotide sequence that exhibits high homology to TLR8, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 151 or SEQ ID NO: 152, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO: 153 or SEQ ID NO: 154. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:153 or SEQ ID NO:154, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:151 or SEQ ID NO:152.

[0148] The term "VWA2" refers to the nucleotide sequence as set forth in SEQ ID NO: 155, which corresponds to the sequence of the human von Willebrand factor A domain containing 2 gene (Ensembl: ENSG00000165816), e.g. as defined in NCBI Reference Sequence NM_001320804, in particular the NCBI Reference Sequence for the VWA2 transcript shown above, and it also relates to the corresponding amino acid sequence as set forth in SEQ ID NO: 156, which corresponds to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_001307733, which encodes the VWA2 polypeptide.

[0149] The term "VWA2" also refers to a nucleotide sequence that exhibits high homology to VWA2, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:155, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:156. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:156, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:155.

[0150] The term "ZAP70" refers to the nucleotide sequence as set forth in SEQ ID NO: 157 or SEQ ID NO: 158, which corresponds to the sequence of the NCBI Reference Sequence of the ZAP70 transcript shown above, of the gene of the zeta chain of the T-cell receptor-associated protein kinase 70 (Ensembl: ENSG00000115085), e.g. as defined in the NCBI Reference Sequence NM_001079 or the NCBI Reference Sequence NM_207519, and in particular to the corresponding amino acid sequence as set forth in SEQ ID NO: 159 or SEQ ID NO: 160, which corresponds to the protein sequence defined in the NCBI Protein Accession Reference Sequence NP_001070 and the NCBI Protein Accession Reference Sequence NP_997402, which encodes a ZAP70 polypeptide.

[0151] The term "ZAP70" also refers to a nucleotide sequence that exhibits high homology to ZAP70, for example, a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:157 or SEQ ID NO:158, or an nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:159 or SEQ ID NO:160. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:159 or SEQ ID NO:160, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:157 or SEQ ID NO:158.

[0152] The term "ZBP1" refers to the nucleotide sequence as set forth in SEQ ID NO: 161 or SEQ ID NO: 162 or SEQ ID NO: 163, which corresponds to the sequence of the Z-DNA binding protein 1 gene (Ensembl: ENSG00000124256), e.g. as defined in NCBI Reference Sequence NM_030776 or NCBI Reference Sequence NM_001160418 or NCBI Reference Sequence NM_001160419, in particular the NCBI Reference Sequence of the ZBP1 transcript shown above, and also to the corresponding amino acid sequence as set forth in SEQ ID NO: 164 or SEQ ID NO: 165 or SEQ ID NO: 166, which corresponds to the protein sequence defined in NCBI Protein Accession Reference Sequence NP_110403 and NCBI Protein Accession Reference Sequence NP_001153890 and NCBI Protein Accession Reference Sequence NP_001153891, which encodes a ZBP1 polypeptide.

[0153] The term "ZBP1" also refers to a nucleotide sequence that exhibits high homology to ZBP1, for example a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:161 or SEQ ID NO:162 or SEQ ID NO:163, or an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the sequence as set forth in SEQ ID NO:164 or SEQ ID NO:165 or SEQ ID NO:166. or a nucleic acid sequence encoding an amino acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:164 or SEQ ID NO:165 or SEQ ID NO:166, or an amino acid sequence encoded by a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to a sequence as set forth in SEQ ID NO:161 or SEQ ID NO:162 or SEQ ID NO:163.

[0154] As provided herein, the biological samples used may be collected in a clinically acceptable manner, for example in a manner that preserves nucleic acids (particularly RNA) or proteins.

[0155] The biological sample(s) include body tissues and / or body fluids, such as, but not limited to, blood, sweat, saliva, and urine. Additionally, the biological sample contains cell extracts or cell populations comprising epithelial cells, such as cancerous epithelial cells or epithelial cells from tissue suspected to be cancerous. The biological sample may include cell populations from tissue, such as glandular tissue, for example, the sample may be from the subject's breast. Additionally, cells may be purified from the obtained body tissues and body fluids, if necessary, and used as the biological sample. In some implementations, the sample may be a tissue sample, a body fluid sample, a blood sample, a saliva sample, a sample containing circulating tumor cells, extracellular vesicles, a sample containing breast secreted exosomes, or a cell line or a cancer cell line. In a particular implementation, a biopsy or resection sample is obtained and / or used. Such samples include cells or cell lysates.

[0156] Thus, in one embodiment, the biological sample obtained from the subject is a biopsy. In a further preferred embodiment, the method comprises providing or obtaining a biopsy. In a preferred embodiment, the biopsy is a breast biopsy, for example tissue or fluid from the breast.

[0157] It is also conceivable that the contents of the biological sample are subjected to an enrichment step. For example, the sample can be contacted with a ligand specific for the cell membrane or organelles of a particular cell type, for example breast cells, functionalized, for example, with magnetic particles. The material enriched by the magnetic particles is then used for the detection and analysis steps described herein above or below.

[0158] Furthermore, cells, e.g., tumor cells, can also be enriched via a filtration process of a liquid or liquid sample, e.g., blood, urine, etc. Such a filtration process can also be combined with an enrichment step based on ligand-specific interactions as described herein above.

[0159] Preferably, a biological sample as provided herein is obtained from a subject prior to the initiation of treatment, and preferably, said biological sample is a breast sample or a breast cancer sample.

[0160] Thus, in a preferred embodiment, the method according to the invention comprises obtaining a biological sample from the subject prior to the initiation of treatment, preferably the biological sample is a breast sample or a breast cancer sample. Alternatively, the method according to the invention comprises providing a biological sample obtained from the subject prior to the initiation of treatment, preferably the biological sample is a breast sample or a breast cancer sample.

[0161] It is contemplated herein that the prognosis of a breast cancer subject may be a favorable or unfavorable prognosis. In one embodiment of the present invention, the prediction of the prognosis of a breast cancer subject may result in the determination of favorable or unfavorable risk of a particular prognosis or prognosis. The prognosis may include breast cancer-related death, locoregional recurrence, and / or distant recurrence. Preferably, breast cancer-related death is breast cancer-specific death. The prediction provided by the method of the present invention can provide a prediction of the risk of a breast cancer subject to a particular prognosis. Furthermore, the method of the present invention can predict whether a subject with breast cancer is at low risk of a particular prognosis or at high risk of a particular prognosis. The prognosis breast cancer-related death, locoregional recurrence, and / or distant recurrence as used herein constitutes a prognosis that is not favorable for the subject. A further prognosis is total mortality. Total mortality as used herein is a prognosis that can be predicted by the method of the present invention but is not directly associated with the subject's death from breast cancer.

[0162] The methods of the invention provided preferably comprise predicting a prognosis for a breast cancer subject of a favorable or unfavorable risk of breast cancer related mortality, locoregional recurrence and / or distant recurrence after surgery, in other words, surgery is performed on the breast cancer subject prior to predicting the risk of a particular prognosis.

[0163] It is contemplated that breast cancer subjects having high risk (i.e., above a certain threshold) for one or more outcomes of breast cancer-related mortality, locoregional recurrence, and / or distant recurrence, as predicted by the methods of the present invention, are associated with an unfavorable risk for each outcome, preferably after surgery.

[0164] It is contemplated that a breast cancer subject having a low risk (i.e., below a certain threshold) for one or more prognoses of breast cancer-related death, locoregional recurrence, and / or distant recurrence predicted by the method of the present invention is associated with a favorable risk for each prognosis, preferably after surgery. A favorable risk may include a different follow-up strategy, e.g., a different recommended (follow-up) therapy, than an unfavorable risk for the subject. For example, a different follow-up strategy, e.g., a different recommended (follow-up) therapy, may be considered for a breast cancer subject, preferably a subject who has undergone breast (cancer) surgery, to improve the chances of survival of the breast cancer subject.

[0165] The treatment preferably comprises surgery, radiation therapy, hormonal therapy, cytotoxic chemotherapy and / or immunotherapy. Combination therapy in cancer, i.e., the combination of two or more therapeutic methods and / or agents, for example, the combination of radiation therapy and chemotherapy, is widely considered a cornerstone in treating cancer.

[0166] Thus, in a preferred embodiment, the method according to the invention comprises that a biological sample, preferably a sample of a subject's breast or of a subject's breast cancer, is obtained prior to the initiation of therapy, including surgery, radiation therapy, hormonal therapy, cytotoxic chemotherapy and / or immunotherapy.

[0167] It is contemplated that the unfavorable risk of prognosis will affect the recommended therapy for said breast cancer subject associated with the unfavorable risk. If an unfavorable risk is predicted by the method of the present invention, the recommended therapy may be as follows: (i) Radiation therapy provided earlier than standard; and (ii) radiotherapy with escalating radiation doses; and (iii) adjuvant therapy, such as cytotoxic chemotherapy, immunotherapy, and / or hormonal therapy; and (iv) surgery; and (v) Alternative non-surgical treatments It is contemplated that the compound may have one or more of:

[0168] In a preferred embodiment, the method according to the invention comprises recommending a therapy based on a prediction, preferably a prediction regarding prognosis, wherein - If the prognosis is unfavorable, the recommended therapy, preferably after surgery, is: (i) Radiation therapy provided earlier than standard; and (ii) radiotherapy with escalating radiation doses; and (iii) adjuvant therapy, such as cytotoxic chemotherapy, immunotherapy, and / or hormonal therapy; and (iv) Surgery and (v) Alternative non-surgical treatments Contains one or more of the following:

[0169] It is contemplated that the favorable risk of prognosis will influence the recommended therapy for said breast cancer subject associated with the favorable risk. If a favorable risk is predicted by the method of the present invention, the recommended therapy may be as follows: (vi) definitive radiation therapy; and (vii) salvage radiation therapy; and (viii) salvage radiation therapy at deescalating dose levels; and (ix) Observation It is contemplated that the compound may have one or more of:

[0170] Therefore, another preferred embodiment of the method according to the invention comprises a treatment recommendation based on the prediction, wherein: - If the prognosis is favorable, the recommended therapy is preferably postoperative, as follows: (vi) definitive radiation therapy; and (vii) salvage radiation therapy; and (viii) salvage radiation therapy at deescalating dose levels; and (ix) Observation It has one or more of the following:

[0171] In one embodiment of the present invention, favorable or unfavorable risk of prognosis is predicted for breast cancer subjects who have undergone surgery, preferably breast surgery, such as, but not limited to, lumpectomy, quadrantectomy, partial mastectomy, split mastectomy, or total mastectomy.

[0172] In another aspect of the invention, there is provided a method according to the invention comprising recommending a second line treatment for a particular patient, preferably a patient at low or high risk of suffering an outcome selected from breast cancer specific mortality or overall mortality, wherein the recommendation is based on a prediction of the prognosis of the breast cancer patient, wherein: - if the prognostic value is favorable, second-line treatment is not recommended; and / or - If prognosis is unfavourable, second-line treatment is recommended.

[0173] It is provided herein that the method according to the present invention can predict the efficacy of secondary treatment after surgery for patients with low-risk or high-risk profile.Preferably, the secondary treatment comprises one or more, more preferably all, of chemotherapy, hormone therapy and radiation therapy.

[0174] The methods provided herein are based on the expression level, e.g., expression profile, of at least one gene selected from immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1, and / or T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70, and / or PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2. It will be appreciated that the method may be performed on input related to the expression level of one or more genes, or determining the expression level may be part of the method.

[0175] It is further envisaged that the method is executed by a processor.Thus, in one embodiment, the present invention relates to a computer-implemented method of predicting prognosis for a subject with breast cancer.

[0176] The method of predicting prognosis for a subject with breast cancer includes detecting one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, of the immune defense response genes selected from AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1, and / or ... CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, Preferably, the method comprises determining the gene expression profile(s) of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, of T cell receptor signaling genes selected from the group consisting of: PTPRC, ZAP70, and / or one or more, e.g., 1, 2, 3, 4, 5, 6, 7 or all, of PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2.

[0177] In a further embodiment, the method for predicting prognosis of a subject with breast cancer comprises detecting two or more, e.g., one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or all, of immune defense response genes selected from AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1, and / or detecting an expression level of one or more of the immune defense response genes selected from CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, P determining the gene expression profile(s) of two or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, of T cell receptor signaling genes selected from the group consisting of RKACA, PRKACB, PTPRC, and ZAP70, and / or two or more, e.g., 1, 2, 3, 4, 5, 6, 7 or all, of PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2.

[0178] In one preferred embodiment, the method according to the invention comprises the steps of: - the one or more immune defense response genes comprise 3 or more, preferably 6 or more, more preferably 9 or more, most preferably all of the immune defense genes; and / or - the one or more T cell receptor signaling genes comprise 3 or more, preferably 6 or more, more preferably 9 or more, and most preferably all of the T cell receptor signaling genes; and / or The one or more PDE4D7-correlated genes include three or more, preferably six or more, and most preferably all of the PDE4D7-correlated genes.

[0179] Thus, determining the first, second and / or third gene expression profile(s) according to the method of the invention preferably comprises determining or receiving the result of: - 3 or more, preferably 6 or more, more preferably 9 or more, most preferably all of the immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1: - three or more, preferably six or more, more preferably nine or more, most preferably all of the T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or - three or more, preferably six or more, more preferably nine or more, and most preferably all of the genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2.

[0180] In a further embodiment, the method according to the invention comprises determining a prognosis based on one or more immune defense response genes, one or more T cell receptor signaling genes, and one or more PDE4D7 correlated genes.In a further embodiment, the method according to the invention comprises determining a prognosis based on two or more immune defense response genes, two or more T cell receptor signaling genes, and two or more PDE4D7 correlated genes.In a further embodiment, the method according to the invention comprises determining a prognosis based on three or more immune defense response genes, three or more T cell receptor signaling genes, and three or more PDE4D7 correlated genes.

[0181] As a first reference example, the (first) gene expression profile embodied herein may have at least one, at least two, at least three immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1. In one non-limiting example, the gene expression profile embodied herein has the expression profile of genes DDX58, DHX9, and IFI16.

[0182] By way of further example, the (second) gene expression profile embodied herein may have at least one, at least two, at least three T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70. In one non-limiting example, the gene expression profile embodied herein has an expression profile of the genes PRKACA, PRKACB, and PTPRC.

[0183] As a further reference example, the (third) gene expression profile embodied herein may comprise at least one, at least two, or at least three PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2. In a non-limiting example, the gene expression profile embodied herein has the expression profiles of genes KIAA1549, PDE4D, and RAP1GAP2.

[0184] As a further non-limiting example, the gene expression profile provided herein can have a first, second and third gene expression profile, where the first gene expression profile has at least one, at least two, at least three immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; and the second gene expression profile has at least one, at least two, at least three immune defense response genes selected from the group consisting of CD2, CD247, CD28, and a third gene expression profile comprising at least one, at least two, or at least three T cell receptor signaling genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2.

[0185] In a further non-limiting reference example, the gene expression profile for predicting the prognosis of breast cancer may be composed of the gene expression profiles of TLR8 (IDR gene), CD2 (TCR gene) and PDE4D (PDE4D7-related gene). In another example, the gene expression profile may be composed of the gene expression profiles of OAS1 and TLR8 (IDR gene), CD2 (TCR gene) and PDE4D (PDE4D7-related gene), or the gene expression profiles of TLR8 (IDR gene), CD2 and PTPRC (TCR gene) and PDE4D (PDE4D7-related gene), or the gene expression profiles of TLR8 (IDR gene), CD2 (TCR gene) and CUX2 and PDE4D (PDE4D7-related gene).

[0186] Cox proportional hazards regression allows the analysis of the impact of multiple risk factors on a tested event, such as survival, in time. In it, risk factors are binary or discrete variables, such as risk scores or clinical stages, or continuous variables, such as biomarker measurements or gene expression values. The probability of an end point (e.g., death or disease recurrence) is called the hazard. In the regression analysis, next to information about whether the tested end point was reached or not reached, for example in a patient cohort (e.g., whether the patient died or not), the time to the end point is also taken into account. The hazard is modeled as H(t)=H0(t)·exp(w1·V1+w2·V2+w3·V3+···), where V1, V2, V3··· are predictor variables, H0(t) is the baseline hazard, while H(t) is the hazard at any time t. The hazard ratio (HR) (i.e., risk of achieving an event) is given by Ln[H(t) / H0(t)] = w1 · V1 + w2 · V2 + w3 · V3 + · · ·, where the coefficients or weights w1, w2, w3 · · · are estimated by Cox regression analysis and can be interpreted in the same way as logistic regression analysis.

[0187] Thus, in one preferred embodiment, the method according to the present invention comprises determining a prognosis prediction by combining a first combination of gene expression profiles, a second combination of gene expression profiles, and a third combination of gene expression profiles with a regression function derived from a population of breast cancer subjects.

[0188] In another preferred embodiment, the method according to the invention comprises determining a prognosis by: - combining the first gene expression profile for two or more, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, of the immune defense response genes with a regression function derived from a population of breast cancer subjects; and / or - combining a second gene expression profile for two or more, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, of the T cell receptor signaling genes with a regression function derived from a population of breast cancer subjects; and / or - combining a third gene expression profile for two or more, such as 2, 3, 4, 5, 6, 7 or all, of the PDE4D7-correlated genes with a regression function derived from a population of breast cancer subjects.

[0189] In a particular implementation, the prognosis prediction is determined as follows: IDR_14_model: (1) (w1 - AIM2) + (w2 - APOBEC3A) + (w3 - CIAO1) + [...] + (w14 - ZBP1) where w1 to w14 are weights, and AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1 are the expression levels of genes.

[0190] In another particular realization, the prognosis prediction is determined as follows: TCR_17_model: (2) (w15 - CD2) + (w16 - CD247) + (w17 - CD28) + [...] + (w31 - ZAP70) where w15 to w31 are weights, and CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70 are gene expression levels.

[0191] In another particular realization, the prognosis prediction is determined as follows: PDE4D7_CORR_model: (3) (w32 - ABCC5) + (w33 - CUX2) + (w34 - KIAA1549) + [...] + (w39 - VWA2) where w32 to w39 are weights, and ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2 are the expression levels of genes.

[0192] In another particular realization, the prognosis prediction is determined as follows: BRCAI_model (4) (w42 - PDE4D7_CORR) + (w40 - IDR_14) + (w41 - TCR_17)

[0193] In another realization, the prognosis prediction is determined as follows: BRCAI_model (5) (w1 - AIM2) + (w2 - APOBEC3A) + (w3 - CIAO1) + [...] + (w14 - ZBP1) + (w15 - CD2) + (w16 - CD247) + (w17 - CD28) + [...] + (w31 - ZAP70) + (w32 - ABCC5) + (w33 - CUX2) + (w34 - KIAA1549) + [...] + (w39 - VWA2), where w1 to w39 are weights, and AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, ZBP1, CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, ZAP70, ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2 are gene expression levels.

[0194] The prognostic predictions may also be classified or categorized into at least two risk groups based on the prognostic predictive value. For example, there may be two risk groups, or three risk groups, or four risk groups, or five or more predefined risk groups.

[0195] Each risk group covers a respective (non-overlapping) range of predictive values ​​of prognosis. For example, a risk group can represent the probability of occurrence of a particular clinical event, such as 0 to <0.1, or 0.1 to <0.25, or 0.25 to <0.5, or 0.5 to 1.0.

[0196] In a preferred embodiment, the method according to the present invention comprises that the determination of prognosis prediction is based on one or more clinical parameters obtained from the subject.As mentioned above, various means based on one or more clinical parameters are considered.By making prognosis prediction based on such clinical parameters, prediction can be further improved.

[0197] In a preferred embodiment, the one or more clinical parameters comprises at least the number of tumor positive lymph nodes. More preferably, the clinical parameter is the number of tumor positive lymph nodes.

[0198] In one embodiment, determining a prediction of therapeutic response comprises combining gene expression levels of one or more IDR genes, one or more TCR signaling genes, and / or one or more PDE4D7-correlated genes with a regression function derived from a population of breast cancer subjects.

[0199] It is further preferred that the gene expression profile of one or more IDR genes, one or more TCR signaling genes, and / or one or more PDE4D7-correlated genes and one or more clinical parameters obtained from the subject are combined using a regression function derived from a population of breast cancer subjects.

[0200] It is also preferred to combine one or more clinical parameters with one or more of the first, second and third gene expressions using a regression function derived from a population of breast cancer subjects to provide a method of predicting prognosis for a breast cancer subject.

[0201] Thus, in a preferred embodiment of the method according to the invention, determining the prognosis comprises combining one or more of the following: (i) a first gene expression profile(s) for one or more immune defense response genes; (ii) a second gene expression profile(s) for one or more T cell receptor signaling genes; (iii) a third gene expression profile(s) for one or more PDE4D7-correlated genes; and (iv) combining the first gene expression profile, the second gene expression profile, the third gene expression profile, and one or more clinical parameters obtained from the subjects with a regression function derived from a population of breast cancer subjects.

[0202] In another particular implementation, the prognosis prediction is determined as follows: BRCAI_clinical model (6) (w43 - BRCAI_model) + (w44 - LN_positive) where w43 and w44 are weights, BRCAI_model is the regression model described above based on the expression profile of one or more IDR genes, one or more TCR signaling genes, and / or one or more PDE4D7-correlated genes, and LN_positive represents the number of tumor-positive lymph nodes. A clinical model was created by combining BRCAI and clinical data (number of positive lymph nodes) using multivariate Cox regression.

[0203] An example of a suitable clinical parameter as used and exemplified herein is "LN_positive", which represents the number of tumor-positive lymph node metastases after postoperative pathology. LN-positive breast tumors are breast tumors in which tumor cells have metastasized to nearby lymph nodes. It is contemplated that LN-positive tumors are a precursor to occult metastasis of cancer.

[0204] The method according to the invention may be implemented in a computer program that can be executed on a computer. A computer program product includes a non-transitory computer readable recording medium, such as a disk, hard drive, etc., on which a control program is recorded (stored). Common forms of non-transitory computer readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes or other magnetic storage media, CD-ROMs, DVDs or other optical media, RAM, PROMs, EPROMs, FLASH-EPROMs or other memory chips or cartridges, or other non-transitory media that can be read and used by a computer.

[0205] Therefore, in a preferred embodiment, the present invention also provides a computer program comprising instructions which, when said program is executed by a computer, cause the computer to carry out a method comprising: - receiving data indicative of a gene expression profile comprising expression levels of four or more genes selected from the first gene expression profile, the second gene expression profile and / or the third gene expression profile, the first gene expression profile has one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; the second gene expression profile has one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; the third gene expression profile comprises one or more, e.g., one, two, three, four, five, six, seven or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2; said gene expression profile being determined in a biological sample obtained from the subject; - determining a prediction of the subject's prognosis based on the first gene expression profile(s), or the second gene expression profile(s), or the third gene expression profile(s), or the first, second and third gene expression profile(s), wherein the prediction is a favorable or unfavorable risk of breast cancer related mortality, locoregional recurrence and / or distant recurrence.

[0206] In an alternative embodiment, there is provided a computer program comprising instructions which, when executed by a computer, cause the computer to carry out a method comprising: - a first gene expression profile for each of one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1, wherein the first gene expression profile(s) a gene expression profile determined in a biological sample obtained from a breast cancer subject; and / or one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70. receiving data indicative of a second gene expression profile for each of the T cell receptor signaling genes of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, wherein said second gene expression profile(s) are determined in a biological sample obtained from the breast cancer subject; and / or a third gene expression profile for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7 or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, wherein said third gene expression profile(s) are determined in a biological sample obtained from the breast cancer subject; - determining a prediction of the subject's prognosis based on the first gene expression profile(s), or the second gene expression profile(s), or the third gene expression profile(s), or the first, second and third gene expression profile(s), wherein the prediction is a favorable or unfavorable risk of breast cancer related mortality, locoregional recurrence and / or distant recurrence.

[0207] Alternatively, one or more steps of the method may be implemented in a transitory medium such as a transmittable carrier wave in which the control program is embodied as a data signal using a transmission medium such as sound or light waves, such as those generated during radio wave and infrared data communications.

[0208] Exemplary methods are implemented on one or more general purpose computers, special purpose computers(s), programmed microprocessors or microcontrollers and peripheral integrated circuit elements, hardwired electronic or logic circuits such as ASICs or other integrated circuits, digital signal processors, discrete element circuits, programmable logic devices such as PLDs, PLAs, FPGAs, Graphical cards CPUs (GPUs) or PALs, etc. In general, any device capable of implementing a finite machine capable of sequentially performing the steps described herein can be used to perform one or more steps of the exemplary risk stratification method for treatment selection in patients with breast cancer. As will be appreciated, the steps of the method can all be performed by a computer, although in some embodiments, one or more of the steps can be performed at least partially manually.

[0209] As used herein, the computer program according to the present invention may preferably be implemented on a device for predicting the prognosis of a subject with breast cancer, said device comprising an input adapted to receive data indicative of a gene expression profile(s) of immune defense response genes and / or T cell receptor signaling genes and / or PDE4D7-correlated genes, said device further comprising a processor adapted to determine a prediction of the prognosis based on said one or more gene expression profile(s); and - comprising a delivery unit adapted to provide a prediction or a choice-based treatment recommendation to a healthcare caregiver or to a subject, as appropriate.

[0210] In a further aspect of the present invention, there is provided an apparatus for predicting prognosis in a subject with breast cancer, comprising: - an input adapted to receive data indicative of a gene expression profile, the gene expression profile comprising expression levels of four or more genes selected from the first, second and / or third gene expression profiles; - the first gene expression profile has one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; - the second gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or - the third gene expression profile has one or more, e.g. one, two, three, four, five, six, seven or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2; an input, wherein the gene expression profile is determined in a biological sample obtained from the subject; - a processor adapted to determine a prediction of radiotherapy response based on the gene expression profile of the four or more genes; and - Where appropriate, providing an apparatus comprising a delivery unit adapted to provide a prediction or a treatment recommendation based on the prediction to a healthcare caregiver or a subject.

[0211] In an alternative embodiment, there is provided an apparatus for predicting prognosis in a subject with breast cancer, comprising: an input unit adapted to receive data indicative of a gene expression profile, said input unit comprising: - each of one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all of the immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; - each of one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or - each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7 or all, of the PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2: adapted to receive data indicative of a gene expression profile of an input, in which the gene expression profile is determined in a biological sample obtained from the subject; - a processor adapted to determine a prediction of radiotherapy response based on the gene expression profile of the three or more genes; and Optionally, an apparatus is provided that includes a delivery unit adapted to provide a prediction or a treatment recommendation based on the prediction to a healthcare caregiver or a subject.

[0212] A computer-implemented process may also be created by loading computer program instructions on a computer, other programmable data processing apparatus, or other device to cause a series of work steps to be executed on the computer, other programmable apparatus, or other device such that the instructions running on the computer or other programmable apparatus provide a process for performing the functions / operations specified herein.

[0213] Thus, in a preferred embodiment, the present invention also provides a diagnostic kit comprising at least one polymerase chain reaction primer, and optionally at least one probe, for determining the gene expression profile(s) in a biological sample and / or sample obtained from a breast cancer subject.

[0214] Preferred is said diagnostic kit comprising at least one polymerase chain reaction primer, and optionally at least one probe, for determining a gene expression profile, the gene expression profile comprising expression levels of four or more genes selected from a first, second and / or third gene expression profile, the first gene expression profile comprising one or more, e.g. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8 and ZBP1; and the second gene expression profile comprising one or more, e.g. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8 and ZBP1. The expression profile has one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all of the T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or a third gene expression profile comprises one or more, e.g., 1, 2, 3, 4, 5, 6, 7 or all of the PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2.

[0215] In an alternative embodiment, the diagnostic kit provided herein comprises at least one polymerase chain reaction primer, and optionally at least one probe, for determining the following gene expression profile(s): - each of one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all of the immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; - each of one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or - each of one or more, such as 1, 2, 3, 4, 5, 6, 7 or all, of the PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2.

[0216] Preferably, the kit described herein is for determining the expression levels of four or more genes of a gene expression profile in a biological sample obtained from a subject, preferably a subject suffering from breast cancer.

[0217] In another preferred embodiment, the present invention provides the use of a diagnostic kit, preferably as broadly embodied herein, in a method for predicting the prognosis of a subject with breast cancer.

[0218] In another preferred embodiment, the present invention provides a method comprising the steps of: - receiving a biological sample obtained from a subject with breast cancer; - determining a gene expression profile in a biological sample using a kit according to claim 13, wherein the gene expression profile comprises expression levels of four or more genes selected from the first, second and / or third gene expression profiles, - the first gene expression profile has one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; - the second gene expression profile comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or - the third gene expression profile has one or more, such as 1, 2, 3, 4, 5, 6, 7 or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2.

[0219] Alternatively, the present invention provides a method comprising the steps of: - receiving a biological sample obtained from a subject with breast cancer; - determining the following gene expression profile(s) using a diagnostic kit as broadly embodied herein: - each of one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, of the immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1: - each of one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or - determining a gene expression profile(s) for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7 or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2; Determining in a biological sample and / or sample obtained from the breast cancer subject.

[0220] In a further embodiment, the invention provides for the use of a gene expression profile in a method of predicting prognosis of a subject with breast cancer, the gene expression profile having expression levels of four or more genes selected from a first, second and / or third gene expression profile, wherein the first gene expression profile has one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all immune defense response genes selected from the group consisting of: AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; and the second gene expression profile has expression levels of four or more genes selected from the group consisting of: CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LBP1, LBP2, LBP3, LBP4, LBP5, LBP6, LBP7, LBP8, LBP9, LBP10, LBP11, LBP12, LBP13, LBP14, LBP15, LBP16, LBP17, LBP18, LBP19, LBP20, LBP21, LBP22, LBP23, LBP24, LBP25, LBP26, LBP27, LBP28, LBP29, LBP30, LBP31, LBP32, LBP33, LBP34, LBP35, LBP36, LBP37, LBP38, LBP39, LBP40, LBP41, LBP42, LBP43, LBP44, LBP45, LBP46, LBP47, LBP48, LBP49, LBP50, LBP510, LBP52, LBP53, L and / or a third gene expression profile having one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all T cell receptor signaling genes selected from the group consisting of CK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or a third gene expression profile having one or more, e.g., 1, 2, 3, 4, 5, 6, 7 or all PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, wherein the method comprises determining a prognostic prediction based on gene expression levels for the four or more genes, wherein the prediction is a favorable or unfavorable risk of breast cancer related mortality, locoregional recurrence, and / or distant recurrence.

[0221] All references cited herein, including journal articles or abstracts, published or corresponding patent applications, patents, or other documents, are hereby fully incorporated by reference, including all data, tables, figures, and text presented in the references. In addition, the entire contents of the references cited within the references cited herein are also hereby fully incorporated by reference.

[0222] Reference to known method steps, conventional method steps, known methods, or conventional methods is in no way an admission that any aspect, explanation, or embodiment of the present invention is disclosed, taught, or suggested in the relevant art.

[0223] It is to be understood that the terms or phrases used herein are for the purpose of description and not of limitation, as they should be interpreted by one of ordinary skill in the art in light of the teaching and guidance presented herein, in combination with the knowledge of those of ordinary skill in the art.

[0224] It will be understood that all details, embodiments and preferences described with respect to one aspect of an embodiment of the invention are equally applicable to other aspects or embodiments of the invention, and therefore it is not necessary to separately detail all such details, embodiments and preferences for every aspect.

[0225] Having generally described the invention, the same will be more readily understood by reference to the following examples, which are provided by way of illustration and are not intended to be limiting of the invention. Further aspects and embodiments will be apparent to those skilled in the art.

[0226] Other variations to the disclosed implementations can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.

[0227] Any reference signs in the claims should not be construed as limiting the scope. EXAMPLES

[0228] Example 1: Prediction of prognosis in subjects with breast cancer Two data sets were used for the analysis of the three gene signatures. The METABRIC data set used herein contains information on 3.240 patients. The tumors of these patients represent different histological types. The TCGA data set used herein contains information on 1.086 patients. The tumors of these patients represent different histological types.

[0229] Gene expression data generated by TCGA (The Genome Cancer Atlas) along with clinical and survival data were downloaded from the database TCGA Breast Invasive Carcinoma, Firehose legacy, accessed on March 6, 2020. METABRIC gene expression data were downloaded from the European Genome-Phenome Archive. Clinical and survival data were downloaded from Supplementary Information (Curtis et al.; Nature, 2012). For both datasets, gene expression values ​​were presented as log2 data. For both datasets (TCGA, METABRIC), the log2_expression value of each gene was converted to a z-score by calculation: z-score log2_gene=((log2_gene)-(mean_samples)) / (stdev_samples) (7)

[0230] where log2_gene is the log2 gene expression value per gene, mean_samples is the mathematical mean of the log2_gene values ​​across all samples, and stdev_samples is the standard deviation of the log2_gene values ​​across all samples.

[0231] This process results in the transformed log2_gene being distributed around a mean of 0 with a standard deviation of 1. For multivariate analysis of genes of interest, the log2_gene transformed z-score values ​​of each gene are used as input. Reference genes were selected as shown in Tables 1-3.

[0232] Cox regression analysis Provided herein is a study to evaluate whether the combination of 14 immune defense response genes, the combination of 17 T cell receptor signaling genes, the combination of 8 PDE4D7-correlated genes, and their combinations show prognostic value for breast cancer.Using Cox regression, the expression levels of 14 immune defense response genes, 17 T cell receptor signaling genes, and 8 PDE4D7-correlated genes are modeled against overall survival, respectively.

[0233] The METABRIC (Study ID: EGAS00000000083) cohort of 3,240 breast cancer subjects and the TCGA cohort of 1,086 breast cancer subjects were selected for modeling. Subjects in the METABRIC cohort were divided based on histological breast cancer type (see Table 4). Subjects in the TCGA cohort (The Cancer Genome Atlas) were divided based on histological breast cancer type (see Table 5).

[0234] [Table 4]

[0235] [Table 5]

[0236] METABRIC and TCGA subjects were then analyzed based on additional characteristics: only samples with complete data on histology (grade, lymph node), TNM (pathological T, lymph node, metastasis), NPI (Nottingham Prognostic Index), molecular (ER, PR, Her2, PAM50), and survival (disease-specific and non-cause-specific) were included.

[0237] From this subset, a total of 1.950 METABRIC subjects were included and a total of 1.086 TCGA subjects were included based on complete available data on PAM50, survival, and clinical annotation, as well as gene expression.

[0238] The 1,950 METABRIC subjects were divided into two subcohorts with complete data: a training cohort (997 subjects) and a testing cohort (953 subjects).

[0239] 1,086 TCGA subjects were analyzed with complete PAM50 and survival data and then divided into two groups based on the presence or absence of clinical annotation data. Two subcohorts with complete data were created for validation of the BRCAI model (group 1) and the BRCAI_clinical model (group 2). Group 1 included 673 subjects and group 2 (clinical annotation group) included 563 subjects. The two models (BRCAI and BRCAI_clinical) developed in the METABRIC discovery cohort were validated in the TCGA cohort by Kaplan Meier analysis using the same cutoffs as in the analysis of the METABRIC data.

[0240] The Cox regression function was derived as follows:

number

[0241] Details of the weights w1 to w39 are shown in Table 6 below.

[0242] [Table 6] TIFF2024544750000009.tif137170

[0243] Log2 expression data of genes in PDE4D7_CORR_model, IDR_14_model, and TCR_17_model were converted to z-scores for the discovery and validation sets. The z-scores of each gene signature were combined for the training set in a multivariate Cox regression analysis using disease-specific death as the clinical endpoint to construct the PDE4D7_CORR, IDR_14, and TCR_17 models. The Breast Cancer Immunity Score (BRCAI) was constructed by modeling all-cause breast cancer mortality using the three gene models. The BRCAI score was constructed in the METABRIC (MB) discovery cohort and tested in the METABRIC and TCGA validation cohorts. All-cause mortality was the clinical endpoint validated in the multivariate Cox regression. The IDR, TCR, and PDE4D7_CORR signature scores derived from the individual Cox regression models were used as inputs for the multivariate Cox regression analysis. This resulted in a final BRCAI risk score.

[0244] The Cox regression function was derived as follows:

number

[0245] The BRCAI_Clinical Score was constructed in the METABRIC (MB) discovery cohort and tested in the METABRIC and TCGA validation cohorts. The clinical endpoint validated with MV Cox regression was all-cause mortality. The IDR, TCR, and PDE4D7_CORR signature scores derived from the individual Cox regression models, and the number of tumor-positive lymph nodes (LN_positive) determined by pathological examination after surgery were used as inputs for the MV Cox regression analysis. This resulted in the final BRCAI_Clinical Score.

number

[0246] Details of weights w40 to w44 are shown in Table 7 below.

[0247] In this specification, the BRCAI_Clinical_model is also referred to as the BRCAI&Clinical_model.

[0248] [Table 7]

[0249] Kaplan-Meier survival analysis Three individual Cox regression models (IDR_14_model, TCR_17_model, PDE4D7_CORR_model) and two combined models in Kaplan-Meier survival analysis were validated. Different clinical endpoints (breast cancer-specific mortality; all-cause mortality; local recurrence; distant metastasis) were validated. Two models (BRCAI and BRCAI&Clinical) developed in the METABRIC discovery cohort were validated by Kaplan Meier analysis in the TCGA cohort.

[0250] For Kaplan-Meier survival curve analysis, patients were classified into two subcohorts based on the cutoffs of the Cox functions of the risk models (IDR_14_model, TCR_17_model, PDE4D7_CORR_model, BRCAI_model, BCAI&Clinical_model). The thresholds for dividing the low- and high-risk groups were based on the risk of experiencing the clinical endpoint (prognosis) predicted by each Cox regression model.

[0251] Figures 1-18 show Kaplan-Meier survival curve analyses of the METABRIC training cohort.

[0252] Figures 19-33 show Kaplan-Meier survival curve analyses of the METABRIC validation cohort.

[0253] Figures 34-40 show Kaplan-Meier survival curve analysis of the TCGA validation cohort.

[0254] In the METABRIC training cohort, patient classes represented whether subjects were at low or high risk of experiencing the clinical endpoints for which the risk models created (IDR_14_model, TCR_17_model, PDE4D7_CORR_model, BRCAI_model, BCAI&Clinical_model) were tested: postoperative breast cancer-specific mortality (Figures 1-5), postoperative mortality (Figures 6-7), locoregional recurrence (Figures 8-9), and distant recurrence (Figures 10-11).

[0255] In the METABRIC validation cohort, patient classes represent whether the subjects are at low or high risk of experiencing the clinical endpoints for which the risk models created (BRCAI_model, BCAI&Clinical_model) were tested: breast cancer-specific postoperative mortality (Figures 19-20), postoperative death (Death) (Figures 21-22), locoregional recurrence (Figures 23-24) or distant recurrence (Figures 25-26).

[0256] The patient classes in the TCGA validation cohort represent whether the subjects have a low or high risk of experiencing the tested clinical endpoint of all-cause mortality (Death) after surgery in the risk models created (BRCAI_model, BCAI&Clinical_model) (Figures 34-35).

[0257] Kaplan-Meier analysis as shown in Figures 1-5, 6-11, 19-26 and 34-35 shows that the prognosis of breast cancer subjects can be predicted, for example, prognosis selected from breast cancer specific mortality; total mortality; locoregional recurrence and distant metastasis. Using a risk model based on a combination of randomly selected genes as provided herein provides an improved prediction of the prognosis of breast cancer subjects. The prediction of prognosis may improve the choice of treatment and improve survival rates. Using the risk model developed herein is expected to improve the prediction of the effectiveness of post-operative treatment options for said subjects. Furthermore, this reduces the suffering of patients who are spared ineffective treatment and reduces the costs spent on ineffective treatment. Overall, based on a model with genetic variables and weights as presented herein, the prognosis of breast cancer subjects can be methodologically predicted, for example, through distinguishing the risk of a particular prognosis in different patient risk groups by stratification and statistical methodology.

[0258] Example 2: Prediction of prognosis in PAM50 breast cancer subtypes Cox regression model Provided herein is a study of whether the BRCAI_Clinical_model represents a prognostic method for determining the prognosis of various PAM50 subtypes. Different clinical endpoints were examined: breast cancer-specific mortality in PAM50 subtypes and all-cause mortality in PAM50 subtypes.

[0259] [Table 8]

[0260] Kaplan-Meier survival analysis In the first experiment, PAM50 subtypes were plotted for breast cancer-specific and all-cause mortality outcomes for Kaplan-Meier survival curve analysis.A training cohort of 997 subjects (METABRIC), a validation cohort of 975 subjects (METABRIC), and a validation cohort of 563 subjects (TCGA) were classified into five subcohorts based on their PAM50 subtypes: Basal (basal-like), Her2 (Her2-enriched), LumA (Luminal A), LumB (Luminal B), and Normal (normal breast cancer subtype).

[0261] Survival based on the prognosis of breast cancer-specific mortality (Figures 12 and 27) or overall mortality (Figures 13 (METABRIC), 28 (METABRIC), and 36 (TCGA)) was plotted for the separate subjects comprising each subtype.

[0262] Further, each PAM50 subtype was individually selected as a subcohort. The Cox function of the risk model BRCAI_Clinical_model was classified into two subcohorts, that is, subcohorts showing low or high risk of each prognosis based on the cutoff. Figures 14 to 18 show Kaplan-Meier survival curve analysis of the training cohort based on prognosis: breast cancer-specific death for each subcohort of PAM50 subtype. Figures 29 to 33 show Kaplan-Meier survival curve analysis of the validation cohort based on prognosis: breast cancer-specific death for each subcohort of PAM50 subtype. Figures 37 to 40 show Kaplan-Meier survival curve analysis of the validation cohort based on prognosis: all-cause death for each subcohort of PAM50 subtype.

[0263] It is demonstrated herein that the prognosis, e.g., breast cancer-specific mortality and / or overall mortality, of a subject having a breast cancer identified as a PAM50 subtype can be predicted by using a risk model based on a randomly selected combination of genes as provided herein, in combination with one or more clinical parameters, e.g., LN positivity.

[0264] Improved prediction of subject prognosis improves treatment selection and potential survival of subjects with breast cancer identified as specific PAM50 subtype. Using the risk model developed herein, it is expected to improve prediction of the effectiveness of post-operative treatment options for said subjects. Furthermore, this will reduce the suffering of patients who are spared ineffective treatment and reduce the cost of ineffective treatment.

[0265] Example 3: Prediction of prognosis of breast cancer patients after second-line chemotherapy, hormonal therapy, and radiotherapy Cox regression analysis Herein, a propensity-matched analysis is provided to validate and quantify the risk of prognosis of second-line treatment consisting of chemotherapy, hormonal therapy, and radiotherapy after surgery in breast cancer patients. A subcohort of 186 breast cancer subjects was selected from the METABRIC cohort. This subcohort was divided into two groups based on stratification with low risk (#60) or high risk (#125). Different clinical endpoints were validated: breast cancer-specific mortality; all-cause mortality. The variables and weights of the Cox regression model are shown in Table 9.

[0266] [Table 9]

[0267] Kaplan Meier curve analysis Kaplan Meier survival curve analysis separated the subjects into two separate groups. For each Kaplan Meier analysis, the group that did not receive second-line treatment and the group that received second-line treatment are shown. The results are shown in Figures 41 to 44. Survival times are shown in each figure.

[0268] Figures 41 and 42 show Kaplan Meier survival curve analyses of the METABRIC low-risk and high-risk cohorts with breast cancer-specific death as an endpoint.

[0269] Figures 43 and 44 show Kaplan Meier survival curve analyses of the METABRIC low-risk and high-risk cohorts with all-cause mortality as an endpoint.

[0270] It is demonstrated herein that the prognosis of low-risk and / or high-risk subjects with breast cancer after administration of all or no second-line treatments can be predicted by using the methods of the present invention, said second-line treatments comprising administration of chemotherapy, hormonal therapy and radiation therapy to the subject.

[0271] Improved prediction of subject prognosis may improve the selection of secondary treatment and improve the survival of subjects with breast cancer.By using the risk model developed herein, it is expected to improve the prediction of the effectiveness of secondary treatment options after surgery for subjects.In addition, this model will also reduce the suffering of patients who are spared ineffective secondary treatment and reduce the cost spent on ineffective secondary treatment.

[0272] Example 4. Validation of a gene signature using four or more genes We hypothesized that using a subset of genes from different gene expression profiles (immune defense, T cell signaling, PDE4D7-related) would be sufficient to stratify patients. To test this hypothesis, we selected a random subset of genes as described below and used it in the model and the same dataset as above.

[0273] The random selection of genes was performed as follows: Five models were built with a random selection of six genes: BRCAI_6.1 and BRCAI_6.2 consist of two genes each from immune defense, T cell signaling, and PDE4D7-related gene sets;

[0274] BRCAI_6.3, BRCAI_6.4 and BRCAI_6.5 are composed of at least one gene each from immune defense, T cell signalling and PDE4D7-related gene sets;

[0275] BRCAI_6.6 and BRCAI_6.7 are composed of four genes selected from the PDE4D7-related gene set;

[0276] BRCAI_6.8 and BRCAI_6.9 are composed of four genes selected from the immune defense gene set;

[0277] BRCAI_6.10 and BRCAI_6.11 are composed of four genes selected from the T cell signaling gene set.

[0278] [Table 10] TIFF2024544750000016.tif255118

[0279] Here, PDE4D7 refers to the PDE4D7-correlated gene set, IDR refers to the immune defense response gene set, and TCR refers to the T cell receptor signaling gene set.

[0280] The results obtained using these 11 models are shown in Figures 45 to 55. All models significantly separated the patient groups, as indicated by the p-values ​​in the figures (p-values ​​for each comparison are 0.05 or less). From these results, the inventors concluded that a minimum of four genes selected from a three-gene profile is sufficient to stratify patients and predict the prognosis of breast cancer patients.

[0281] The attached sequence listing, entitled 2021PF00566 SEQ LIST, is hereby incorporated by reference in its entirety.

Claims

1. 1. A method for predicting the prognosis of a subject with breast cancer, said method comprising: determining or receiving a first gene expression profile for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1, wherein the first gene expression profile is determined in a biological sample obtained from the breast cancer subject; and / or determining or receiving a second gene expression profile for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or all, T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70, wherein the second gene expression profile is determined in a biological sample obtained from the breast cancer subject; and / or determining or receiving a third gene expression profile for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, wherein the third gene expression profile is determined in a biological sample obtained from the breast cancer subject; determining a prognostic prediction based on the first, second, and / or third gene expression profiles, wherein the prediction is a favorable or unfavorable risk of breast cancer-related mortality, locoregional recurrence, and / or distant recurrence; and optionally providing a prognosis to a healthcare caregiver or said breast cancer subject.

2. 2. The method of claim 1, wherein the prognosis prediction for the breast cancer subject comprises a favorable or unfavorable risk of postoperative breast cancer-related mortality, locoregional recurrence and / or distant recurrence.

3. the one or more immune defense response genes comprise three or more, preferably six or more, more preferably nine or more, and most preferably all immune defense genes; and / or the one or more T cell receptor signaling genes include three or more, preferably six or more, more preferably nine or more, and most preferably all of the T cell receptor signaling genes; and / or The method according to claim 1 or 2, wherein the one or more PDE4D7-correlated genes include three or more, preferably six or more, and most preferably all of the PDE4D7-correlated genes.

4. determining the prognosis prediction, combining a first gene expression profile for two or more, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or all, of said immune defense response genes with a regression function derived from a population of breast cancer subjects; and / or combining a second gene expression profile for two or more, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or all, of said T cell receptor signaling genes with a regression function derived from a population of breast cancer subjects; and / or 4. The method of claim 1, further comprising combining a third gene expression profile for two or more, for example, two, three, four, five, six, seven or all, of the PDE4D7-correlated genes with a regression function derived from a population of breast cancer subjects.

5. 5. The method of claim 4, wherein determining the prognosis prediction comprises combining the first combination of gene expression profiles, the second combination of gene expression profiles, and the third combination of gene expression profiles with a regression function derived from a population of breast cancer subjects.

6. 6. The method of claim 1, wherein the determination of the prognosis is based on one or more clinical parameters obtained from the breast cancer subject.

7. determining the prognosis, (i) a first gene expression profile for one or more immune defense response genes; (ii) a second gene expression profile for one or more T cell receptor signaling genes; (iii) a third gene expression profile for one or more PDE4D7-correlated genes; and (iv) combining the first gene expression profile, the second gene expression profile, the third gene expression profile, and one or more clinical parameters obtained from the breast cancer subjects with a regression function derived from a population of breast cancer subjects.

7. The method of claim 5 or 6, comprising combining one or more of:

8. 8. The method of claim 1, wherein the one or more clinical parameters include at least the number of tumor-positive lymph nodes.

9. 9. The method of any one of claims 1 to 8, wherein the biological sample is obtained from the breast cancer subject before the start of treatment, preferably the biological sample is a breast sample or a breast cancer sample.

10. A method according to any one of claims 1 to 9, wherein the treatment is surgery, radiation therapy, hormone therapy, cytotoxic chemotherapy and / or immunotherapy.

11. A treatment is recommended based on the prediction, wherein: If the prognosis is unfavorable, the recommended therapy is preferably post-operative: (i) radiation therapy delivered earlier than standard; and (ii) radiation therapy with escalating radiation doses; and (iii) adjuvant therapy, such as cytotoxic chemotherapy, immunotherapy, and / or hormonal therapy, and (iv) surgery, and (v) Alternative non-surgical treatments 11. The method of claim 1, comprising one or more of:

12. A therapy is recommended based on said prediction, wherein: If the prediction is favorable, the recommended therapy is preferably post-operative: (vi) definitive radiation therapy, and (vii) salvage radiation therapy, and (viii) salvage radiation therapy at deescalating dose levels, and (ix) Follow-up 12. The method of claim 1, comprising one or more of:

13. When the program is executed by a computer, the computer a first gene expression profile for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1, wherein the first gene expression profile is determined in a biological sample obtained from the breast cancer subject; and / or a second gene expression profile for each of one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70, wherein the second gene expression profile is determined in a biological sample obtained from the breast cancer subject; and / or a third gene expression profile for each of one or more, e.g., one, two, three, four, five, seven, or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2, wherein the third gene expression profile is determined in a biological sample obtained from the breast cancer subject. receiving data indicative of determining a prognosis prediction for the breast cancer subject based on the first gene expression profile, or the second gene expression profile, or the third gene expression profile, or the first, second and third gene expression profiles, wherein the prediction is a favorable or unfavorable risk of breast cancer-related mortality, locoregional recurrence and / or distant recurrence; 10. A computer program comprising instructions for carrying out a method having the steps of:

14. 1. Use of a diagnostic kit in a method for predicting the prognosis of a subject with breast cancer, the diagnostic kit comprising: The following gene expression profiles: each of one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, of the immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or all, T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or Each of one or more, for example, 1, 2, 3, 4, 5, 6, 7 or all, of the PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2 Use of a diagnostic kit comprising at least one polymerase chain reaction primer and optionally at least one probe for determining in a biological sample and / or sample obtained from a breast cancer subject.

15. Use of the diagnostic kit according to claim 14 in the method according to any one of claims 1 to 12.

16. receiving a biological sample obtained from a subject with breast cancer; By using the diagnostic kit according to claim 14, each of one or more, for example 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or all, T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or a gene expression profile of each of one or more, for example, 1, 2, 3, 4, 5, 6, 7 or all, of the PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2; determining in a biological sample and / or sample obtained from a subject with breast cancer.

17. a first gene expression profile for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or all, immune defense response genes selected from the group consisting of AIM2, APOBEC3A, CIAO1, DDX58, DHX9, IFI16, IFIH1, IFIT1, IFIT3, LRRFIP1, MYD88, OAS1, TLR8, and ZBP1; a second gene expression profile for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or all, T cell receptor signaling genes selected from the group consisting of CD2, CD247, CD28, CD3E, CD3G, CD4, CSK, EZR, FYN, LAT, LCK, PAG1, PDE4D, PRKACA, PRKACB, PTPRC, and ZAP70; and / or a third gene expression profile for each of one or more, e.g., 1, 2, 3, 4, 5, 6, 7, or all, PDE4D7-correlated genes selected from the group consisting of ABCC5, CUX2, KIAA1549, PDE4D, RAP1GAP2, SLC39A11, TDRD1, and VWA2.

1. Use of a method for predicting the prognosis of a subject with breast cancer, said method comprising:

1. Use of a method for predicting the prognosis of a breast cancer subject, comprising determining a prognosis prediction based on gene expression levels for three or more genes, said prediction being a favorable or unfavorable risk of breast cancer-related mortality, locoregional recurrence and / or distant recurrence.