Prognostic signature for prostate cancer classification
A chromatin-derived transcriptional signature from fresh biopsies using 4f-SAMMY-seq analysis identifies distinct prostate cancer subgroups, enhancing prognosis prediction and treatment guidance by accurately stratifying patients based on chromatin 3D architecture alterations.
Patent Information
- Application Number
- PCT/EP2025/079328
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-10
- Filing Date
- 2025-10-10
- Publication Date
- 2026-04-16
AI Technical Summary
Current prognostic screenings for prostate cancer are ineffective in predicting clinical outcomes due to its heterogeneous and multifocal nature, leading to over-treatment and severe side-effects for many patients, as existing genetic biomarkers have not significantly improved prognosis prediction accuracy.
A chromatin-derived transcriptional signature, derived from 4f-SAMMY-seq analysis on fresh biopsies, identifies two subgroups (LDD and HDD) with distinct chromatin 3D architecture alterations, using a 18-gene signature to predict patient prognosis and treatment response.
The 18-gene signature provides a novel tool for prostate cancer prognosis prediction, accurately stratifying patients into better or worse prognosis groups, guiding treatment decisions and reducing over-treatment.
Smart Images

Figure IMGF000005_0001 
Figure IMGF000017_0001 
Figure IMGF000018_0001
Abstract
Description
[0001] PROGNOSTIC SIGNATURE FOR PROSTATE CANCER CLASSIFICATION
[0002] FIELD OF THE INVENTION
[0003] Primary prostate cancer (PCa) is characterized by multifocal growth and highly variable clinical course which is not effectively predicted by prognostic screenings. The present invention provides a chromatin-derived genes transcriptional signature suitable to classify two subgroups of cancer patients with different BCR outcome. Gene expression analysis revealed a protective, antitumoral activity of the HDD subtype associated with changes in chromatin compartments and transcriptional repression. The transcriptional signature of the invention confirmed its prognostic relevance across multiple cohorts covering more than 900 prostate cancer patients in total thus providing an innovative strategy for the stratification of primary prostate cancers.
[0004] BACKGROUND
[0005] Prostate cancer (PCa) is the second most frequent tumor type in males and accounts for 10% of cancer-related deaths in males, according to the latest figures by the global cancer observatory1. PCa diagnosis is challenging due to its heterogeneous and multifocal nature and to date, it is based on histological evaluation of multiple biopsies to assign a severity score (Gleason Score)2usually performed after the detection of increased blood levels of prostate-specific antigen (PSA) or digital rectal examination (DRE)2. The majority of PCa patients do not undergo prostatectomy and exhibit an indolent and slow growing disease that will not become life-threatening, whereas about 20% of prostate tumors progress to become a metastatic and lethal disease3. Current tumor grading systems combine clinical and pathological features. However, PCa clinical variability and frequent mu Itifocal ity complicate the prediction of cancer outcomes, hindering treatment guidance4. Thus, the large number of patients with silent PCa are often over-treated and exposed to the risk of severe side-effects5and there is an urgent need of new biomarkers to refine the stratification of prostate cancer patients.
[0006] In recent years, the identification of genetic signatures in several solid tumors has been advancing the field of clinical oncology67. Nevertheless, despite considerable efforts to identify subtypes in the heterogeneous PCa patients, the translation of genetic biomarkers into clinical practice only marginally improved the accuracy of prognosis prediction8.
[0007] EP2771481 describes a method for classifying a prostate cancer in a subject, the method comprising the steps of a) determining a gene expression level or gene expression pattern of the genes F3 and IGFBP3 in a sample from the subject and b) classifying the tumor by comparing the gene expression level determined in a) with a reference gene expression of the same genes in reference patients known to have a high risk or low risk tumor respectively. EP2882869 provides genes expression profiles that are associated with prostate cancer. EP2809812 provides algorithm-based molecular assays that involve measurement of expression levels of genes from a biological sample obtained from a prostate cancer patient. The genes may be grouped into functional gene subsets for calculating a quantitative score useful to predict a likelihood of a clinical outcome for a prostate cancer patient.
[0008] More recently, epigenetic studies yielded promising results as innovative diagnostic and prognostic tools in oncology9 10. The genome is topologically and hierarchically organized on multiple, highly regulated, structural levels in the nucleus11. This 3D architecture has emerged as crucial for regulation of transcription, replication, DNA repair and splicing. In particular, the physical separation between euchromatin, the more accessible and transcriptionally active part of the chromatin, and heterochromatin, the compacted and gene poor part of the genome, is a hallmark of healthy cells. As such, morphological changes in nuclear organization, called nuclear atypia, still represent a gold standard for diagnosis and staging of different cancers, including prostate cancer12. In this context, chromatin three-dimensional (3D) architecture is also emerging for its involvement in tumorigenesis of prostate cancer13 14. The advancements in this field are driven by epigenomic experimental technologies such as high-throughput chromosome conformation capture (Hi-C)15. A handful of studies on prostate cancer cell lines reported alterations in regulatory chromatin loops16 17, shifts in topologically associating domains (TADs) boundaries18and compartment dysregulation19. However, so far chromatin architecture has not been analyzed in primary prostate cancer samples, primarily because of the limited number of cells in PCa biopsies. It is also worth remarking that transcriptional studies that identified signatures currently used in clinics, were mainly performed on surgically removed prostates, thus providing limited information for earlier diagnostic samples20-22. The identification of prognostic signatures in biopsies would be particularly advantageous as can be obtain at the early stage of first diagnostic investigations.
[0009] SUMMARY OF THE INVENTION
[0010] The present invention provides a transcriptional signature derived from chromatin compartmentalization analysis as a novel prognostic tool for prostate tumors at the time of the diagnostic biopsy.
[0011] The present inventors recently developed 4f-SAMMY-seq (Sequential Analysis of MacroMolecules accessibilitY) which can be applied also on a small number of cells from fresh tissue samples, thus allowing unprecedented versatility in genome-wide mapping of euchromatic and heterochromatic domains, as well as their 3D compartmentalization23. In this invention, the 4f-SAMMY-seq was applied on fresh biopsies from prostate cancer patients. Transcriptome analysis (RNA-seq) was performed in parallel to confirm the functional effects of epigenome alterations. By analyzing the epigenome in prostate biopsies of 17 chemo-naive patients with putative PCa for genome-wide mapping of heterochromatic and euchromatic domains, as well as their three-dimensional (3D) compartmentalization in the cell nucleus a novel gene expression signature was derived and confirmed its ability to predict patient’s prognosis across multiple cohorts covering more than 900 patients in total. In particular, two subgroups of cancer patients with different degrees of chromatin 3D architecture alterations were identified: the LDD (Low Degree of Decompartmentalization) and HDD (High Degree of Decompartmentalization) groups. Gene expression analysis revealed a protective, antitumoral activity of the HDD subtype associated with changes in chromatin compartments and transcriptional repression.
[0012] The 18-genes signature of the invention provides a novel tool for PCa prognosis prediction that can be easily implemented in the clinical practice.
[0013] It is therefore an object of the invention an in vitro method for predicting prognosis for a patient suffering from prostate cancer comprising at least the step a) of measuring the level of expression of a plurality of biomarker genes selected from the group consisting of ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB in an isolated biological sample obtained from a subject, preferably at least three of said biomarker genes, preferably at least four, preferably at least five, preferably at least ten.
[0014] Still preferably the level of expression of the biomarker genes COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC, IQGAP2, SC5D is measured in step a), preferably wherein the following genes are downregulated in good-prognosis patient: COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC; and the following genes are upregulated in good-prognosis patient: IQGAP2, SC5D.
[0015] More preferably the level of expression of the biomarker genes ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB is measured in step a).
[0016] Still preferably according to the invention the following genes are downregulated in good-prognosis patient (HDD down): ACAP3, AGRN, ATG16L2, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC MXRA8, MYO15B, PABPN1 , SCRIB; and the following genes are upregulated in good-prognosis patient (HDD up): ELOVL7, IQGAP2, SC5D, C16orf70.
[0017] More preferably the in vitro method of the invention further comprises a step b) of calculating a Prostate Compartmentalization Index (PCI) score based on the level of expression of the plurality of biomarker genes across a reference compendium of prostate cancer patients.
[0018] According to a preferred embodiment in the in vitro method of the invention the calculation of the PCI score comprises: measuring the expression level for each gene / selected from ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB in the patient p; calculating the normalized expression value for each gene with respect to the reference compendium of prostate cancer patients by computing the value Zi>p=Xl'rSD^1where xi pis the gene / expression value in patient p, is the average expression value for the gene / across the reference compendium of prostate cancer patients and SD, is its standard deviation; discriminating the genes as “good-prognosis Up-regulated” and “good-prognosis Down- regulated” as above indicated, i.e. wherein in better prognosis patient the following genes are downregulated (good-prognosis Down-regulated): ACAP3, AGRN, ATG16L2, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC MXRA8, MYO15B, PABPN1 , SCRIB; and the following genes are upregulated (good-prognosis Up- regulated): ELOVL7, IQGAP2, SC5D, C16orf70; calculating the values 5"lpand S2pfor the patient p as the median of the normalized expression values Z,pfor the “good-prognosis Up-regulated" and "good-prognosis Down- regulated" genes, respectively; calculating PCI score for the patient p as: PCIp= Slp- S2p
[0019] According to a further preferred embodiment the in vitro prognostic method of the invention the calculation of the PCI score comprises: measuring the expression level for each gene / selected from COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC, IQGAP2, SC5D in the patient p; calculating the normalized expression value for each gene with respect to the reference compendium of prostate cancer patients by computing the value Zi p=Xl’vSD>J'1where xitPis the gene / expression value in patient p, is the average expression value for the gene / across the reference compendium of prostate cancer patients and SDi is its standard deviation; discriminating the genes as “good-prognosis Up-regulated” and “good-prognosis Down- regulated” as above indicated, i.e. wherein in better prognosis patient the following genes are downregulated (good-prognosis Down-regulated): COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC; and the following genes are upregulated (good-prognosis Up-regulated): IQGAP2, SC5D; calculating the values 5"lpand S2pfor the patient p as the median of the normalized expression values Z,pfor the “good-prognosis Up-regulated" and "good-prognosis Downregulated" genes, respectively; calculating PCI score for the patient p as: PCIp= Slp- S2p
[0020] Preferably the in vitro method for predicting prognosis of the invention further comprises a step c) of determining a prognosis by means of said PCI score calculation, preferably determining a prognosis includes classifying said subject in one of at least two classes corresponding to levels of progression of the disease, such as better prognosis and worse prognosis. Still preferably, determining a prognosis includes classifying said subjects in one of at least two classes corresponding to i) patients responsive to treatments comprising Cabazitaxel, Docetaxel, Estramustina, Abiraterone acetato, Bicalutamide, Buserelin, Ciproterone, Enzalutamide, Flutamide, Goserelin, Leuprolide, medroxyprogesterone acetate, triptorelin and ii) patients non-responsive to said treatments.
[0021] Preferably determining a prognosis comprises stratifying the subject into a worse prognosis class if said calculated PCI score has a value < 0 or into a better prognosis class if the calculated PCI score has a value >0.
[0022] According to a preferred embodiment of the invention, the following genes are downregulated in good-prognosis patient: ACAP3, AGRN, ATG16L2, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC MXRA8, MYO15B, PABPN1 , SCRIB; and the following genes are upregulated in good-prognosis patient: ELOVL7, IQGAP2, SC5D, C16orf70.
[0023] Preferably in the in vitro method for determining prognosis according to the invention the biological sample is a biopsy or a tumor sample.
[0024] Still preferably, prostate cancer is adenocarcinoma, small cell prostate cancer, neuroendocrine prostate cancer.
[0025] It is a further object of the invention the in vitro method for determining prognosis as defined above wherein measuring the level of expression of the biomarker genes comprises performing microarray analysis, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT- PCR), a Northern blot, or serial analysis of gene expression (SAGE), preferably wherein the level of expression of said biomarker genes is obtained by determining the level of respective RNA transcripts, mRNA or protein translation products thereof.
[0026] It is a further object of the invention a kit comprising agents for measuring levels of expression of a plurality of biomarker genes selected from the group consisting of ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB in a biological sample, preferably for measuring all the biomarker genes in said group, still preferably for measuring at least levels of expression of CO L5 A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC, IQGAP2, SC5D, said kit preferably further comprising means to calculate a PCI score and optionally means for providing prognostic information.
[0027] Preferably said kit comprises a microarray and / or at least one set of PCR primers capable of amplifying a nucleic acid comprising a biomarker gene sequence and / or at least one probe capable of hybridizing to a nucleic acid comprising a biomarker gene sequence or its complement.
[0028] In a preferred embodiment the in vitro method for detecting prognosis according to the invention or the kit is used in combination with a further method useful in the prognosis and / or in the classification of a subject affected by prostate cancer. It is a further object of the invention a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of the invention.
[0029] It is a further object of the invention a data processing device comprising means adapted for carrying out the method as above defined, said device preferably comprising: i) means for receiving data representing the expression level of at least one biomarker gene as defined in claim 1 in an isolated biological sample from a subject; ii) means for calculating a PCI score; iii) means for providing a prognosis of the disease based on said PCI score.
[0030] Figure legends
[0031] Figure 1. Chromatin structural alterations of prostate cancer cohort
[0032] Overview of the experimental design (panel A). Fresh biopsies were dissected into 2 portions: one third was flash frozen and embedded in OCT for histopathological and immunochemical analyses (1), while the rest was enzymatically digested into single cell suspensions that were further divided for flow cytometry-based phenotype assay (2), 4f-SAMMY-seq and RNA-seq (3).
[0033] Study cohort biopsies (panel B): representative images of Hematoxylin and Eosin staining (H&E) at magnification x20 prostate biopsies of for non-neoplastic patients (CTR) and prostate cancer patients (PCa). In the schematic prostate diagram the circles mark sites of diagnostic biopsies, the star indicates the research biopsy position. Positive cores (black circles) and Percentage of Positive Cores (PPC) are also reported. GS-P indicates diagnostic Gleason Score assigned to each patient, while GS-CDB indicates the Gleason Score assigned to the biopsy adjacent to the research biopsy. Epigenome profiles (panel C): distribution along chromosome 3 of ChlP-seq enrichment profiles for open (H3K27ac and H3K36me3) and closed (H3K9me3) chromatin marks, along with 4f-SAMMY- seq consensus solubility profiles of 7 CTR and of 10 PCa samples. The solid line shows the mean and the shadowed areas marks the confidence interval: + / -1.96 standard error in the solubility profiles.
[0034] Violin plots (panel D) for genome-wide Spearman correlation values (y-axis) for 4f-SAMMY-seq solubility profiles of CTR (dark grey, n=7) and PCa (light grey, n= 10) samples with respect to ChlP- seq enrichment profiles in normal prostate gland for three distinct histone marks (H3K36me3, H3K27ac, H3K9me3). Each data point in the violin plots shows the correlation value of individual patients computed with genomic profiles binned at 150kb.
[0035] Figure 2. Identification of prostate cancer epigenetic subtypes
[0036] Pair-wise correlation matrices of reads distribution profiles (panel A) for representative samples CTR31 (left) and PCa87 (right), computed at 150kb genomic bins resolution on chromosome 18 long arm. On the side of each matrix the respective first eigenvector is reported and colored to mark the position of active ("A" compartment with positive eigenvalues) and inactive regions ("B" compartment with negative eigenvalues). On the center of the panel the two bars show the A and B compartments along the chromosome according to a color code: dark grey for active (A) and light grey for inactive (B) compartments.
[0037] Schematic representation of chromatin compartments (panel B) along with chromatin marks on representative chromosome 18. From top to bottom: ChlP-seq enrichment profiles of H3K36me3, H3K27ac and H3K9me3 ; A and B compartments segmentation for CTR and PCa patients (labels on the left), grouped in LDD and HDD subgroups (labels on the right).
[0038] Stacked barplot (panel C) illustrating the count (number of genomic bins, bottom x-axis) and the relative distribution over the entire genome (percentage, top x-axis) of genomic regions with conserved compartment class ("A" - black or "B"- grey) and regions shifting compartment ("A to B" or "B to A" - different shades of grey) for each PCa sample. The discordant chromatin compartment classification is calculated with respect to the consensus of CTR (on the top, see methods).
[0039] Matrix reporting the pair-wise Jaccard Index (panel D) derived from the compartment similarity among PCa samples. The patients within the matrix are arranged by unsupervised hierarchical clustering. The light-purple cluster (PCa33, PCa35, PCa93, PCa88 and PCa100) contains samples of the LDD group. The dark purple cluster (PCa75, PCa4, PCa2, PCa73, PCa87) includes HDD samples.
[0040] Boxplot of solubility profiles average values (panel E) (y-axis) for individual 150kb genomic bins overlapping the ChlP-seq enrichment peaks of open chromatin (H3K36me3, H3K27ac) and closed chromatin (H3K9me3) histone marks (labels on the x-axis). For each group of patients (CTR, LDD and HDD groups) and for each genomic bin we plot the mean 4f-SAMMY-seq solubility profile across the group.
[0041] Figure 3. Functional genomic alteration of prostate cancer specimens
[0042] Volcano plot (panel A) showing Differentially Expressed Genes (DEGs) in the comparison between PCa and CTR samples (left), LDD tumors subtype and CTR samples (right), HDD tumors subtype and CTR samples (bottom). Medium gray dots mark significantly down-regulated genes with adjusted p-value < 0.05 and log2Fold-Change < -1 ; dark gray dots mark significantly up-regulated genes with adjusted p-value < 0.05 and log2Fold-Change > 1.
[0043] Stacked bar plot (panel B) depicting the percentage of DEGs in the comparison LDD vs CTR (left) and HDD vs CTR (right) across concordant (no switch - dark gray) and discordant (switch - light gray) compartments classification with respect to CTR compartments.
[0044] Stacked bar plot (panel C) illustrating the relative distribution of up-regulated and down-regulated genes (x axis) for HDD vs CTR comparison falling in the compartments switching from A in the CTR to B in HDD (lighter black) and from B in the CTR to A in HDD (lighter grey).
[0045] Proportional Venn diagram (panel D) of up-regulated genes (top) and down-regulated genes (bottom) comparing all PCa tissues vs CTR samples, LDD tumors subtype vs CTR samples and HDD tumors subtype vs CTR samples. Functional classes enrichment analysis (panel E) of HDD vs CTR DEGs using GO Cellular Component, Molecular Function and Biological Process annotations. Bar plots representing significantly enriched terms in down-regulated genes light grey) and up-regulated (dark grey) genes. Functional classes enrichment analysis (panel F) using the intersection between significant up- regulated genes included in B to A switches (dark grey dots in the upper panel) and down-regulated genes included in A to B switches (light grey dots in the lower panel) in HDD patients with respect to CTR compartment consensus. Significant terms (adjusted p-value < 0.05) are shown for each database on the right, including GO (Gene Ontology), ChEA (ChlP-X Enrichment Analysis), ARCHS4 (All RNA-seq and ChlP-seq Sample and Signature Search) and ChEA - TFs (ENCODE and ChEA Consensus TFs from ChlP-X). The p-values for enrichments were computed with Enrichr using Fisher’s exact test and Benjamini-Hochberg multiple testing correction.
[0046] Figure 4. HDD transcriptional signature
[0047] Volcano plots (panel A) showing Differentially Expressed Genes (DEGs) in the comparison between HDD and LDD subtypes. The p-value < 0.05 and log2Fold Change value > 1 or log2Fold Change value < -1 and are the cutoff values for significant up-regulated (red, n=24) and down-regulated (blue, n=77) genes, respectively. Non-significant differential expression is marked with grey color. Selected Gene Set Enrichment Analysis (GSEA) profiles (panel B) resulting from the analysis of HDD vs LDD pre-ranked expression scores with HALLMARK gene sets from the Molecular Signatures Database (MSigDB) collection.
[0048] Heatmap (panel C) showing RNA-seq expression values (scaled by row) across 499 prostate cancer tissues from TCGA cohort (see methods). The sample (on columns) are sorted in ascending order according to the PCI score. Rows reports the 101 DEGs from HDD vs LDD comparison and are grouped by hierarchical clustering over Euclidean distance. The top section side bar displays groups of transcriptomic profiles categorized as LDD-like (light pink, PCI score < 0) and HDD-like (deep pink, PCI score > 0), along with the Gleason Score and tumor purity associated with each sample. A colored side bar on the left marks up-regulated (black) and down-regulated (grey) DEGs in HDD vs LDD comparison.
[0049] Kaplan-Meier plot (panel D) showing Biochemical Recurrence Free (BCR) status outcome of TCGA patients stratified according to their transcriptomic similarity with HDD and LDD group. The deep gray line indicates HDD-like TCGA samples, whereas the light gray line indicates LDD-like TCGA samples. P-value was obtained from the Log rank test (p-value = 0.041).
[0050] Multivariate cox regression analysis (panel E) using cox proportional hazards model (Hazard Ratio, HR) for Biochemical Recurrence Free (BCR). The following covariate parameters with p-value for from Wald test have been used in the model: PCI-based signature classification, Clinical Stage (N, T), Age, Tumor Purity, Gleason Score.
[0051] Figure 5. HDD transcriptional signature for prostate cancer prognosis Kaplan-Meier plots (panel A-C) based on the 18 genes refined on 100 samples from Emory cohort (GSE54460 dataset) (panel A, p-value = 0.018), on 248 sample from Belfast cohort (GSE116918 dataset) (panel B, p-value = 0.006), and on 92 samples from Stockholm cohort (GSE70769 dataset) (panel C, p-value= 0.0017). Deep gray lines indicate HDD-like samples while light gray lines indicate LDD-like samples.
[0052] Gene Ontology (GO) terms (panel D) associated with histone mark enrichment in LDD- and HDD- like prostate cancer samples from the GSE120738 and GSE120741. Circle size indicates log- adjusted p-value. Color represents the corresponding histone modification: H3K27ac and H3K4me3 (active marks) in red, and H3K27me3 (repressive mark) in blue. The p-values for enrichments were computed with Enrichr using Fisher’s exact test and Benjamini-Hochberg multiple testing correction. Dot plot (panel E) displaying the expression of the 18-gene signature across cell clusters from scRNA-seq data of primary prostate cancer 39. Dot size indicates the percentage of cells expressing each gene, while color intensity represents the average expression level. Arrows highlight epithelial (SC5D, IQGAP2 and ELOVL7) stromal (COL5A1 , COL5A3, COL16A1 , COL18A1 , MIXRA8) genes, indicating cell-type-specific regulation.
[0053] Representative time-lapse images (panel F) of wound healing assays showing VCaP prostate cancer cells (bright-field) co-cultured with WPMY-1 prostate stromal cells expressing Tomato fluorescent protein (WPMY-1tdTomato, red). WPMY-1tdTomato cells were treated with scrambled siRNA (SCR; top) or 4-gene-targeting siRNA pool (siRNA 4-genes; bottom). Lines indicate wound closure at 0, 12, and 24 hours. Scale bar, 500 pm.
[0054] Percentage of wound area (panel G) compared to the original size during 45 h in co-cultures of VCaP prostate cancer cells and WPMY-1tdTomato fibroblasts treated with scrambled siRNA (SCR; lught grey line) or 4-gene-targeting siRNA cocktail (siRNA 4-genes; dark grey line). Data are presented as mean ± Standard Error of the Mean (SEM) (n= 4 biological replicates per condition). Statistical analysis was performed using a two-way ANOVA (**p-value = 0.0021).
[0055] Figure 6. Cellular and molecular characteristics of the cohort of patients
[0056] Photomicrographs reconstruction (panel A) at Leica DMI6000 of Hematoxylin and Eosin staining (H&E) at 10x magnification of prostate biopsies of non-neoplastic patients (CTR) and confirmed cancer patients (PCa). Inset in the dashed grey boxed area shows a 20x magnification. GS-RDB indicate Gleason Score assigned to our research-dedicated biopsy.
[0057] Integrative clinicopathological and molecular profiling (panel B) of the patients ordered according to Gleason Score (GS). Each sample is annotated for Age, prostate specific antigen (PSA) level, percentage of positive cores (PPC), heatmap of Log2(TPM) and expression values of key PCa markers (PCA3, HPN, GOLM1) estimated from RNA-seq data.
[0058] Representative confocal microscopy images (panel C) of immunostaining for the nucleus with Hoechst (blue) and Lamin A / C (red) in non-neoplastic (CTR) and PCa samples. On bottom-left, outlined nuclei enlarged. Scale bars, 25 and 5 pm in the main and magnified images, respectively. Quantification (panel D) of the nuclear area (left) and nuclear circularity (right) of non-neoplastic (green) and PCa (purple) samples. The box plots indicate median values (middle lines), first and third quartiles (box edges) and minimum and maximum values (whiskers). Statistical significance was determined by unpaired two-tailed Student’s t-test, (**P < 0.005).
[0059] Quantification (panel E) of Hoechst (blue) and Lamin A / C (red) distribution across nuclei. The graph indicates normalized fluorescence intensity (see methods) of Hoechst and Lamin A / C in control samples (solid line) and tumor samples (dashed line). Data are presented as mean ± SEM.
[0060] Figure 7. Cell composition of prostate biopsies
[0061] Representative image (panel A) of needle prostate biopsy specimen in a petri dish (35x15mm). Scale bar length=1cm.
[0062] Number of viable cells (panel B) recovered per mg of each tissue biopsy. 7 non-neoplastic and 10 PCa samples are indicated on left and right, respectively. Data are represented as the median with range. Statistical significance was determined by unpaired two-tailed Student’s t-test (ns, p-value > 0.05).
[0063] Cells are analyzed by flow cytometry through distinct steps (panel C-F): identification of viable cells (panel C, FSC-A vs SSC-A); removal of debris (panel D, FSC-A vs FSC-W); removal of doublets (panel E, FSC-W vs TO-PRO-3); identification (panel F) of tissue resident leukocytes (blue, CD45+), epithelia (green, EpCAM+) and stroma (brown, double negative).
[0064] Box plot (panel G) showing the percentage of 3 distinct prostate cell types (epithelial cells, stromal cells and leukocytes) in non-neoplastic patients (CTR-dark grey box) and PCa samples (PCa-light grey box). The whiskers extend from the minimum value to the maximum value, the box extends from the 25th to the 75th percentiles, the middle line of the box is the median and the mean is indicated by the black plus. Statistical analysis was performed by unpaired two-tailed Student’s t- test (ns, P > 0.05).
[0065] Figure 8. High-throughput sequencing reads quality and alignments of DNA sequence
[0066] Schematic illustration of the 4f-SAMMY-seq protocol (panel A) representing the sequential extraction of chromatin fractions (numbered from S1 to S4) that results in a separation of euchromatic and heterochromatic regions that are then mapped to their genomic coordinates by high-throughput sequencing.
[0067] Representative bioanalyzer assay (panel B) after S2 size selection and S2L, S3, S4 sonication.
[0068] Stacked barplots (panel C) for the number of sequencing reads obtained for each sample and chromatin fraction, divided in trimmed, not aligned, duplicated, filtered (discarded) and used reads (retained for downstream analyses).
[0069] Boxplot (panel D) reporting the distribution of genome-wide Spearman correlation Coefficient (y axis) of each SAMMY-seq fraction (S2S, S2L, S3, S4) (x-axis) in non-neoplastic (CTR, dark grey) and PCa (light grey) samples. The horizontal lines mark the median, the boxes mark the interquartile range (IQR) and whiskers extend up to 1 .5 times the IQR. Visualization of 4f-SAMMY-seq solubility profiles (S2LvsS3) (panel E) of all samples along a representative region (chr5: 10000000-40000000). The solubility profile is labelled in light red for positive values and in light-blue for negative values. Publicly available ChlP-seq enrichment profiles on normal prostate gland of euchromatin H3K36me3 (dark red) and H3K27ac (red) marks and heterochromatin H3K9me3 mark (blue) of the same region are also shown.
[0070] Figure 9. Pathological and molecular differences between HDD and LDD samples.
[0071] Barplot (panel A) showing the total size of "A" compartment regions concordant across different groups of samples. The dark dots in the bottom part of the plot are used to mark the samples accounted for in each bar. The barplot is obtained with the UpsetR package and only the ten most frequent groups in terms or recurrent A compartments are shown. The total size (y-axis values) and number of genomic 150kb bins (number on top of each bar) for each group is shown. The dark purple bar is used to highlight regions that are in the A compartment across all HDD samples and in the B compartment across all LDD samples and the CTR consensus.
[0072] Gene Ontology enrichment analysis (panel B) of genes identified within the recurrent "A to B" or "B to A" compartment shifts in the HDD subtype. Bar plots depict significantly enriched cellular component, molecular function, and biological process terms for recurrent shifts from "A to B" (grey) and "B to A" (lighter grey) in HDD subtype (n=59 for “A to B” shift and n=538 for “B to A” shift).
[0073] Box plot (panel C) showing the distribution of the clinicopathological features between non-neoplastic (CTR, green), LDD (light purple) and HDD (dark purple) PCa subgroups. Each sample is annotated for Age, Prostate Specific Antigen (PSA) blood level, Percentage of Positive Cores (PPC) and Prostate Specific Antigen (PSA), Gleason Score and D’Amico risk classification. In the box plots, the whiskers of the box extend from the minimum value to the maximum value, the box extends from the 25th to the 75th percentiles, the middle line of the box is plotted at the median and the mean is indicated by the plus symbol. Statistical analysis was performed by one-way ANOVA and unpaired two-tailed Student’s t-test (ns, p-value > 0.05).
[0074] Principal Component Analysis (PCA) on RNA-seq data (panel D) of non-neoplastic control samples (green), LDD (light purple) and HDD (dark purple) PCa subgroups.
[0075] Distribution (panel E) of distinct prostate cell types (epithelial cells, stromal cells and leukocytes) in non-neoplastic (CTR, green), LDD (light purple) and HDD (dark purple) measured by flow-cytometry. Data are represented as the median (horizontal line) with range (whiskers). Statistical significance was determined by one-way ANOVA (ns, p-value > 0.05).
[0076] Relative percent (panel F) of different cell subsets (epithelia, stroma and immune cell infiltrate) in each non-neoplastic and PCa biopsy derived from Cibersort bulk deconvolution analysis.
[0077] Boxplot (panel G) showing the distribution of the CIBERSORT Immune Score in non-neoplastic (CTR, green), LDD (light purple) and HDD (dark purple) samples.
[0078] Whole-genome profiles (panel H) of log2copy ratio generated by CNVkit. The heatmap shows the copy number alteration present in each non-neoplastic and PCa biopsy for each chromosome. Figure 10. Functional analysis in HDD and LDD subtypes
[0079] Normalized Enrichment Scores (NES) (panel A) from GSEA analysis on Biological Processes gene set of HDD vs CTR. 104 out of 4220 biological processes gene sets are significantly up-regulated with a corrected p-value (FWER) < 0.05 in HDD; 103 out of 4220 gene sets are significantly down- regulated with a corrected p-value (FWER) < 0.05 in HDD
[0080] Pathway activity score (panel B) of LDD vs CTR and HDD vs CTR comparisons derived from PROGENy compendium. Dark grey bars show decreased pathway activities in LDD or HDD. Light grey bars represent increased pathway activities in LDD or HDD.
[0081] GSEA analysis (panel C) of MSigDB hallmark gene set collection within the expression data of HDD compared to LDD. Positive NES (dark grey bars) indicate enriched gene set in HDD group. Negative NES (grey bars) indicate enriched gene set in LDD group. 7 out of 50 gene sets are significantly up- regulated with a corrected p-value (FWER) < 0.05 in HDD; 9 out of 50 gene sets are significantly down-regulated with a corrected p-value (FWER) < 0.05 in HDD.
[0082] Pathway activity score (panel D) of HDD vs LDD comparison derived from PROGENy compendium. Dark grey bars show decreased pathway activities in HDD. Grey bars represent increased pathway activities in HDD.
[0083] Single-sample GSEA (ssGSEA) scores (panel E) computed individually for each PCa sample for different epigenetic-driven gene sets. The genes present in conserved A compartment (conserved A) and B compartment (conserved B) of all PCa samples have been used as gene sets on which the score is computed. Background A and B represent random gene sets with the same size of its corresponding gene set. Positive enrichment with respect to the background represents increased expression of genes having constitutive A compartment. Negative enrichment with respect to the background represents a decreased expression in constitutive B compartment.
[0084] Figure 11. Signature validation
[0085] Unsupervised clustering (panel A) showing the transcriptomic expression values of 21 prognostic genes (rows) across the TCGA cohort (columns). The slide bar on the top discriminates the transcriptomic clusters HDD-like (dark grey) and LDD-like (light grey). The slide bar on the left shows up-regulated ( dark grey) and down-regulated (light grey) DEGs in HDD.
[0086] Univariate Cox regression analysis results (panel B) showing only genes with significative association to patient good (light grey square) or bad (dark grey square) prognosis (BH adjusted p- value < 0.05) in the TCGA cohort.
[0087] Kaplan-Meier plot (panel C) based on the 18 genes refined signature on TCGA patients (panel B, p- value = 0.0001).
[0088] Kaplan-Meier plots (panel D-E) showing Biochemical Recurrence Free (BCR) status outcome of TCGA samples with high (panel C) and low (panel D) tumor purity (ABSOLUTE score > 0.5).
[0089] Density distribution plot (panel F) for ChlP-seq peaks concordance score (frequency ratio, x-axis) in HDD-like and LDD-like patients. The ChlP-seq targets are indicated in the title of each density plot. Boxplot (panel G) for ATAC-seq reads average signal (Iog2 RPKM, y-axis) for HDD-like vs LDD-like patients considering regions detected as differential peaks associated with SC5D and IQGAP2 gene loci (x-axis label). Whiskers extend up to 1.5 interquartile range (IQR). Wilcoxon test significance is marked as ** for p-value <0.01. Statistical analysis was performed using a two-sided wilcoxon test (SC5D, p-value = 0.0091 ; IQGAP2, p-value = 0.0061).
[0090] Figure 12. Validation PCa cohort
[0091] Representative images (panel A) of Hematoxylin and Eosin (H&E) staining at x20 magnification of prostate biopsies for the validation cohort of PCa patients. GS-P represents the diagnostic Gleason Score assigned to each patient, and GS-CDB indicates the Gleason Score of the biopsy adjacent to the research sample.
[0092] Integrative clinicopathological and molecular profiling of the validation cohort of prostate cancer patients (panel B), ordered according to Gleason Score (GS). Each sample is annotated for age, prostate-specific antigen (PSA) level, percentage of positive cores (PPC), and a heatmap showing Log2(TPM) expression values of key PCa markers (PC A3, HPN, GOLM1) estimated from RNA-seq data.
[0093] Matrix (panel C) showing the pairwise Jaccard Index of compartment similarity among PCa validation cohort samples and consensus profiles of CTR, LDD and HDD. Samples are grouped by unsupervised hierarchical clustering, with light-grey indicating LDD and dark-grey indicating HDD.
[0094] Stacked barplot (panel D) showing the count (genomic bins, bottom x-axis) and relative distribution (percentage, top x-axis) of genomic regions with conserved ("A" - black, "B" - grey) or shifting compartments ("A to B", "B to A" - shades of grey) for each PCa sample of the validation cohort. Discordant chromatin compartment classification is calculated relative to the CTR consensus.
[0095] Cell type composition (panel E) inferred by CYBERSORT analysis for the fraction (y-axis value) of normal (benign) or cancer (malignant) epithelial cells, reported in the left and right boxplots, respectively. A separate box is indicated for CTR, HDD and LDD samples (color legend). All samples from the initial and validation cohorts of patient biopsies are considered (CTR n=7; LDD n=8, HDD n=10). Whiskers extend up to 1.5 interquartile range (IQR). Statistical analysis was performed using a two-sided Wilcoxon test (LDD vs CTR, p-value = 0.002175602; HDD vs CTR, p-value = 0.009666804; HDD vs LDD, p-value = 0.5726039). Significance is marked as ** for p-value <0.01 .
[0096] Figure 13. Immunohistochemistry analysis
[0097] Representative immunohistochemistry (IHC) images of prostate cancer tissues (panel A) stained for SC5D, IQGAP1 (epithelial markers, top row), and COL5A1 , COL5A3, COL18A1 (stromal markers, bottom row). Positive staining for SC5D and IQGAP1 shows intense immunoreactivity localized in epithelial regions, while weak staining for COL5A1 , COL5A3, and COL18A1 is observed in stromal areas. Scale bars, 50 pm.
[0098] Representative IHC images of prostate cancer tissues (panel B) stained for SC5D (top rows) and IQGAP2 (bottom rows) in LDD and HDD samples. LDD tissues show moderate staining for both SC5D and IQGAP2, while HDD tissues display more intense staining. The box plots on the right represent the IHC score (% of positive cells x intensity), indicating a trend towards higher expression in HDD compared to LDD. Bars represent mean ± SEM. The whiskers extend from the minimum value to the maximum value, the box extends from the 25th to the 75th percentiles, the middle line of the box is the median and the mean is indicated by the plus. Statistical significance was assessed using unpaired two-tailed Student’s t-test (ns, p-value > 0.05, SC5D p-value=0,249, IQGAP2 p- value=0,254) Scale bars, 50 pm. Figure 14. Wound healing assay of VCaP cells co-cultured with WPMY-1 stromal cells
[0099] Heatmap (panel A) showing RNA-seq expression values of the 18-gene signature (rows) across prostate cancer cell lines (columns). Samples are grouped into HDD-like (dark grey) and LDD-like (light grey) subtypes. The left sidebar indicates genes up-regulated (dark grey) or down-regulated (grey) in the HDD vs LDD comparison. The arrow marks VCaP cells, which reflect the transcriptional features and were selected for further analyses.
[0100] Transcriptional analysis (panel B) of COL5A1, COL5A3, COL18A1, MIXRA8, SC5D, IQGAP2, in tdTomato-labeled WPMY-1 prostate stromal cells and VCaP prostate cancer cells. Values were normalised to GAPDH and compared with the average expression levels in WPMY-1tdTomato cells. Bars represent mean ± SEM from three independent experiments. Statistical significance was determined by unpaired one-tailed Student’s t-test (*p-value < 0.05, ** p-value < 0.01 and ns p- value>0.05).
[0101] Relative expression (panel C) of COL5A1, COL5A3, COL18A1, and MIXRA8 in WPMY-1tdTomato stromal cells transfected with scrambled siRNA control (SCR, grey) or a pool of siRNAs targeting four genes (COL5A1, COL5A3, COL18A1, and MIXRA8: siRNA 4-genes, dark grey). Gene expression was normalized to GAPDH and compared with the average of SCR amplification. Bars represent mean ± SEM from three independent experiments. Statistical significance was assessed using paired two-tailed Student’s t-test (*p-value < 0.05, ** p-value < 0.01 and ns p-value>0.05).
[0102] Experimental design (panel D) of RNA interference and wound healing assay in VCaP and WPMY- 1 tdTomato co-culture system. WPMY-1dTomato stromal cells were transfected with either scrambled siRNA control (siRNA SCR) or a siRNA pool targeting four genes (COL5A1, COL5A3, COL18A1, and MIXRA8: siRNA 4-genes) and maintained in mono-culture for 36 hours. WPMY- ItdTomato cells were then co-cultured with VCaP prostate cancer cells for 12 hours, followed by wound healing assay at 48 hours post-transfection. Wound closure was monitored by time-lapse imaging for 45 hours to assess migration dynamics in the co-culture system.
[0103] Representative images (panel E) of wound healing assays showing VCaP prostate cancer cells (bright-field) at 0, 12, and 24 hours post-scratch. Scale bar, 500 pm.
[0104] Panels F-l: Wound area (pm2) (f); wound closure (g); VCaP in the wound area (relative to To) (h); WPMY-1 tdTomato in wound area (relative to To) (i), monitored over 45 hours in co-cultures of VCaP prostate cancer cells and WPMY-1tdTomato. WPMY-1tdTomato stromal cells were transfected with either scrambled siRNA (SCR; grey line) or a 4-gene-targeting siRNA cocktail (siRNA 4-genes; magenta line). Data are presented as mean ± SEM. Statistical analysis was performed using two- way ANOVA. (*p-value < 0.05, ** p-value < 0.01 and ns p-value>0.05).
[0105] Figure 15. Validation signature.
[0106] Heatmap of 18 gene signature expression across first cohort (P) and additional cohort (Pea) of prostate cancer samples. DEG_FC1 (dark grey) marks the genes (11) that preserved statistical significance (FDR < 0.05) and magnitude ( | log2FC | > 1 ) after including the additional cohort to the statistic test comparing HDD vs LDD.
[0107] DETAILED DESCRIPTION
[0108] The present invention provides an in vitro method for predicting prognosis for a patient suffering from prostate cancer, as illustrated through the comprehensive experimental approach shown in the figures.
[0109] As used herein, the expression “method for predicting prognosis” can also be referred as a prognostic method, or a method for stratifying patients, and refers to an in vitro method for assessing the likely clinical outcome of a subject affected by prostate cancer, enabling the stratification of subjects based on expected disease progression, therapeutic response, or survival probability.
[0110] The in vitro method of the invention comprises measuring the level of expression of a plurality of biomarker genes selected from the group consisting of ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB in an isolated biological sample obtained from a subject. Said plurality may comprise at least four, five, six seven, eight, nine, ten or eleven genes.
[0111] In one embodiment the method comprises measuring the level of expression of all the 18 genes identified by the invention or at least all said genes.
[0112] In one alternative embodiment, the method comprises measuring the level of expression of at least the biomarker genes COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC, IQGAP2, SC5D in step a). This embodiment represents a refined subset of 11 genes from the complete 18-gene signature that provides enhanced prognostic discrimination.
[0113] The 11 -gene subset includes nine genes that are down-regulated in good-prognosis patients: COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , and CROCC. The subset also includes two genes that are up-regulated in good-prognosis patients: IQGAP2 and SC5D. In this embodiment, the robustness of the genes COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC, IQGAP2, SC5D was supported by inclusion of an additional cohort of samples in the analysis, which confirmed both the statistical significance and the magnitude of the observed difference, as shown in Figure 15.
[0114] The expression levels of these 11 biomarker genes, as well of the 18 biomarker genes and any plurality of biomarker genes according to the invention, can be measured using standard molecular biology techniques such as quantitative PCR, microarray analysis, or RNA sequencing, making the method readily implementable in clinical laboratory settings.
[0115] The positive effects of this approach include the ability to stratify prostate cancer patients into prognostic groups at the time of initial diagnosis using standard biopsy samples. The method provides independent prognostic information beyond traditional clinical parameters such as Gleason Score and PSA levels, enabling more precise risk assessment for individual patients. This stratification capability supports clinical decision-making regarding treatment intensity and monitoring strategies, particularly valuable for distinguishing patients who may benefit from active surveillance versus those requiring immediate intervention.
[0116] A "biomarker" in the context of the present invention refers to a biological compound, such as a polynucleotide or polypeptide which is differentially expressed in a sample taken from subjects having prostate cancer. The biomarker can be a nucleic acid, a fragment of a nucleic acid, a polynucleotide, an oligonucleotide or a protein, a fragment of a protein, or a polypeptide, that can be detected and / or quantified. Biomarkers include nucleic acids comprising nucleotide sequences from genes or RNA transcripts of genes, including but not limited to, ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MY015B, PABPN1 , SC5D, SCRIB and their expression products. Expression profiles of these biomarkers are useful for determining whether an individual who has prostate cancer is likely to have a more favorable prognosis: a favorable prognosis can be quantified for example as biochemical recurrence free survival odds ratio 0.54 and p-value 0.014 in multivariate analysis as reported in Figure 4E.
[0117] The genes herein mentioned are preferably characterized by the sequences described by the following NCBI Gene ID (NCBI Database release version December 15, 2023):
[0118] Accordingly, in one aspect, the invention provides a method for predicting prognosis for a subject who has prostate cancer comprising measuring the level of a plurality of biomarkers in a biological sample comprising cancer cells derived from the subject, and analyzing the levels of the biomarkers and comparing with respective reference value ranges for the biomarkers to be defined across a reference compendium of prostate cancer patients.
[0119] As used herein, the term "probe" or "oligonucleotide probe" refers to a polynucleotide, as defined above, that contains a nucleic acid sequence complementary to a nucleic acid sequence present in the target nucleic acid analyte (e.g., biomarker). The polynucleotide regions of probes may be composed of DNA, and / or RNA, and / or synthetic nucleotide analogs. Probes may be labeled in order to detect the target sequence. Such a label may be present at the 5' end, at the 3' end, at both the 5' and 3' ends, and / or internally.
[0120] The terms "subject," "individual," and "subject," are used interchangeably herein and refer to any mammalian subject, particularly humans. Other subjects may include cattle, dogs, cats, guinea pigs, rabbits, rats, mice, horses, and so on. In some cases, the methods of the invention find use in experimental animals, in veterinary application, and in the development of animal models, including, but not limited to, rodents including mice, rats, and hamsters; and primates.
[0121] Assaying the expression level of a plurality of biomarker genes comprises detecting and / or quantifying a plurality of target nucleic acid analytes. In some embodiments, assaying the expression level of a plurality of biomarker genes comprises sequencing a plurality of target nucleic acids. In some embodiments, assaying the expression level of a plurality of biomarker genes comprises amplifying a plurality of target nucleic acids. In some embodiments, assaying the expression level of a plurality of biomarker genes comprises conducting a multiplexed reaction on a plurality of target analytes. The methods disclosed herein often comprise assaying the expression level of a plurality of targets. The plurality of targets may comprise coding targets and / or non-coding targets of a protein- coding gene or a non protein-coding gene. A protein-coding gene structure may comprise an exon and an intron. The exon may further comprise a coding sequence (CDS) and an untranslated region (UTR). The protein-coding gene may be transcribed to produce a pre-mRNA and the pre- mRNA may be processed to produce a mature mRNA. The mature mRNA may be translated to produce a protein. Gene expression in the sample can be determined using any method known in the art. For example, the amount in the sample of gene expression products, such as RNA transcripts, mRNAs or proteins may be measured.
[0122] For example a PCR-based method, an array based method or a sequencing-based method might be used. All such methods are known in the art and can be carried out according to the common general knowledge in the field. Exemplary techniques to determine gene expression are analysis of single strand conformation polymorphism, capillary electrophoresis, denaturing high performance liquid chromatography, digital molecular barcoding technology, e.g., Nanostring's nCounter® system, direct sequencing, DNA mismatch-binding protein assays, dynamic allele-specific hybridization, Fluorescent in situ hybridization (FISH), high-density oligonucleotide SNP arrays, high- resolution melting analysis, microarray, next generation sequencing (NGS), e.g., using the Illumina Genome Analyzer, ABI Solid instrument, Roche 454 instrument, Heliscope instrument, Northern blot analysis, nuclease protection analysis, oligonucleotide ligase assays, polymerase chain reaction (PCR), primer extension assays, Quantigene analysis, quantitative nuclease- protection assay (qNPA), reporter gene detection, restriction fragment length polymorphism (RFLP) assays, reverse transcription and real-time quantitative polymerase chain reaction (RT- qPCR), reverse transcription- polymerase chain reaction (RT-PCR), RNA sequencing (RNA-seq), Serial analysis of gene expression (SAGE), Single Molecule Real Time (SMRT) DNA sequencing technology, SNPLex, Southern blot analysis, Sybr Green chemistry, TaqMan-based assays, temperature gradient gel electrophoresis (TGGE), Tiling array, Western blot analysis, and immunohistochemistry. In any of the above aspects or embodiments, gene expression may be determined using reverse transcription and real-time quantitative polymerase chain reaction (RT-qPCR) with primers and / or probes (e.g., TaqMan® probes) specific for each of said at least three genes. Alternately, the gene expression may be determined using microarray analysis with probes specific for an expression product of each gene. In an embodiment, the level of RNA transcripts of the herein defined genes is measured in the sample using a RNA sequencing method, such as RNAseq. Methods of extracting RNA and determining the level of RNA transcripts of a particular gene in a sample are well-known to those skilled in the art.
[0123] The present invention provides for a probe set for diagnosing, monitoring and / or predicting a status or outcome of a prostate cancer in a subject comprising a plurality of probes, wherein (i) the probes in the set are capable of detecting an expression level of at least one of ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB; and (ii) the expression level determines the cancer status of the subject. The probe set may comprise one or more polynucleotide probes. Individual polynucleotide probes comprise a nucleotide sequence derived from the nucleotide sequence of the target sequences or complementary sequences thereof. The nucleotide sequence of the polynucleotide probe is designed such that it corresponds to or is complementary to the target sequences. The polynucleotide probe can specifically hybridize under either stringent or lowered stringency hybridization conditions to a region of the target sequences, to the complement thereof, or to a nucleic acid sequence (such as a cDNA) derived therefrom.
[0124] The system of the present invention further provides for primers and primer pairs capable of amplifying target sequences defined by the probe set, or fragments or subsequences or complements thereof. The nucleotide sequences of the probe set may be provided in computer- readable media for in silico applications and as a basis for the design of appropriate primers for amplification of one or more target sequences of the probe set.
[0125] Diagnostic samples for use with the systems and in the methods of the present invention comprise nucleic acids suitable for providing RNAs expression information. In principle, the biological sample from which the expressed RNA is obtained and analyzed for target sequence expression can be any material suspected of comprising prostate cancer tissue or cells. The diagnostic sample can be a biological sample used directly in a method of the invention. Alternatively, the diagnostic sample can be a sample prepared from a biological sample. In one embodiment, the sample or portion of the sample comprising or suspected of comprising cancer tissue or cells can be any source of biological material, including cells, tissue or fluid, including bodily fluids. Non-limiting examples of the source of the sample include an aspirate, a needle biopsy, a cytology pellet, a bulk tissue preparation, a Formalin-Fixed Paraffin-Embedded tissue or a section thereof obtained for example by surgery or autopsy, lymph fluid, blood, plasma, serum, tumors, and organs. In some embodiments, the sample is from urine. Alternatively, the sample is from blood, plasma or serum. In some embodiments, the sample is from saliva.
[0126] According to a preferred embodiment, the method of the invention comprises a further step wherein based on the expression of the claimed set of genes a Prostate Compartmentalization Index (PCI) score is calculated.
[0127] In a preferred embodiment, said PCI score is calculated on gene expression values for the genes of the list defined above. For each gene the expression value is measured with appropriate molecular biology methods to determine the level of RNA transcripts such as RNA-seq, RT-PCR or other techniques known by the skilled person. The said expression values are measured vs a control, wherein said control can be any suitable reference that allows evaluation of the expression level of the genes in the prostate cancer cells as compared to the expression of the same genes in a sample comprising cancerous prostate cells or a pool of such samples obtained from patients with known clinical follow up. Preferably the expression values are measured in a reference compendium of at least 100 prostate cancer patients biopsies with known clinical follow up of 10 or more years. Said reference compendium of patients will be used to define the reference values of / Zi , the average expression value for the gene / across the compendium of patients, and SDi is its standard deviation. When is applied to a new patient p to be classified as good prognosis or bad prognosis, the gene expression level of the genes in the list defined above are measured with the same molecular biology technique applied to measure the expression level of the same gene in the reference compendium of patients.
[0128] The expression level for each gene / from the list above and each patient p to be tested for prognosis is then standardized with respect to the reference compendium by computing the value Zi p=Xl'SvD>l1where xi pis the gene / expression value in patient p, is the average expression value for the gene / across the compendium of patients and SDi is its standard deviation.
[0129] The genes are then divided in good prognosis Up-regulated (“HDD Up-regulated”) and good prognosis Down-regulated ("HDD Down-regulated") as detailed in Figure 11 A. The values 5"lpand S2pare computed for the patient p as the median of the standardized expression values for the good prognosis Up-regulated (“HDD Up-regulated”) and good prognosis Down-regulated ("HDD Down- regulated") genes, considered separately. The value PCI is then computed as PCIp= Slp- S2pSaid calculation can be carried out by the skilled person according to common general knowledge in the field.
[0130] In preferred embodiments of the invention, the PCI score is indicative of the prognosis of the disease. In particular, the subject can be classified or stratified in better or worse prognosis groups depending on said PCI score being >0 or <0.
[0131] For “bad prognosis” or “worse prognosis” or “poor prognosis” is intended an aggressive likely or expected development of the disease, such as an increased risk of developing metastatic cancer. Good prognosis is quantified as biochemical recurrence free survival odds ratio 0.54 and p-value 0.014 in multivariate analysis as reported in Figure 4E.
[0132] For “good prognosis” or “better prognosis” is intended a not aggressive likely or expected development of the disease.
[0133] In a further embodiment of the invention determining a prognosis includes classifying patients in one of at least two classes corresponding to i) patients responsive to treatments comprising Cabazitaxel, Docetaxel, Estramustine, Abiraterone acetate, Bicalutamide, Buserelin, Ciproterone, Enzalutamide, Flutamide, Goserelin, Leuprolide, medroxyprogesterone acetate, triptorelin and ii) patients non- responsive to said treatments.
[0134] All such terms are well known in the field and herein have the meanings currently used in the field. Stratification and classification are herein used as synonyms.
[0135] Kits for performing the desired method(s) are also provided, and comprise a container or housing for holding the components of the kit, one or more vessels containing one or more nucleic acid(s), and optionally one or more vessels containing one or more reagents. The reagents include those described in the composition of matter section above, and those reagents useful for performing the methods described, including amplification reagents, and may include one or more probes, primers or primer pairs, enzymes (including polymerases and ligases), intercalating dyes, labeled probes, and labels that can be incorporated into amplification products. In a further embodiment, the invention discloses a method for treating a subject having prostate cancer, said method comprising: assessing the risk of an aggressive development of the disease or of development of metastatic prostate cancer by measuring the level of expression of a plurality of biomarker genes selected from the group consisting of ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB in an isolated biological sample obtained from a subject, and treating said subject having an increased risk of an aggressive development of the disease or of developing prostate cancer with anticancer therapy, surgery, radiation or androgen ablation. Said anticancer therapy might include comprising Cabazitaxel, Docetaxel, Estramustine, Abiraterone acetate, Bicalutamide, Buserelin, Ciproterone, Enzalutamide, Flutamide, Goserelin, Leuprolide, medroxyprogesterone acetate, triptorelin. In a preferred embodiment, said method for treating a subject comprises measuring the level of expression of the following biomarker genes COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC, IQGAP2, SC5D, preferably wherein the following genes are downregulated in good-prognosis patient: COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC; and the following genes are upregulated in good-prognosis patient: IQGAP2, SC5D. In a further preferred embodiment said method comprises determining the expression level of the following biomarker genes ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB, preferably wherein the following genes are downregulated in good-prognosis patient: ACAP3, AGRN, ATG16L2, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC MXRA8, MYO15B, PABPN1 , SCRIB; and the following genes are upregulated in good-prognosis patient: ELOVL7, IQGAP2, SC5D, C16orf70.
[0136] Materials and methods
[0137] Prostate tissues cohort
[0138] The cohort study includes chemo-naive patients followed in the Urology Division of Fondazione IRCCS Ca' Granda - Ospedale Maggiore Policlinico (Milan) who underwent the transrectal ultrasound-guided systematic sampling of prostate tissue (TRUS biopsy). This diagnostic procedure was performed as part of routine clinical management following the detection of abnormal digital rectal examination or an elevated PSA blood level. According to the current standard for the detection of PCa, 14 ultrasound-guided biopsy cores (diagnostic biopsies), seven from each side of the prostate gland, were collected from each patient for the clinical diagnosis. During the same procedure, one additional biopsy core to be used for research biopsy, was taken from a site directly adjacent to one of the diagnostic biopsies. The institutional ethics committee board authorized this study (authorization n 1063). All specimens were obtained after the patients had provided written informed consent following the ethical principles of biomedical research on biospecimens. To ensure optimal recovery of living cells, fresh surgical prostate biopsy specimens were placed in ice-cold saline buffer and directly transported to the laboratory within 1 hour from sample collection. Tissue samples were then immediately processed to preserve the intranuclear genomic and protein architectures of cell nuclei. Samples selected for epigenome studies consist of 17 fresh biopsies divided on the bases of the histology and the spatial distribution of the positive cores in two different groups: 10 PCa biopsies from patients with histologically confirmed prostate cancer and 7 from patients who had no cancer in any biopsy core. All the clinical data were provided by the Urology division preserving the confidentiality of patient personal data.
[0139] Tissues processing
[0140] The biopsy specimen was stored at 4-8°C for maximum 3 hours before dissociation, avoiding its freezing. Given the little size of needle biopsy cores (typically 20-30 mg) the digestion condition to achieve the optimal recovery of living cells were empirically determined. Briefly, the biopsy tissue was transferred in a 2 ml microcentrifuge tube and rinse twice with 1 ml of ice-cold sterile PBS 1X. Then, the tissue was cut into small pieces (~1 mm) with autoclaved surgical scissors directly in the microcentrifuge tube. The resulting minced tissue was enzymatically digested by adding 1 ml of prewarm HBSS (Gibco, 14025) containing 200 units of collagenase type I (Life Technologies, 17018- 029) per ~10 mg of tissue plus 67 pg DNase I (Sigma-Aldrich, 10104159001). Tissue dissociation was carried out in a water bath at 37 °C shaking vigorously for 10 sec every 5 minutes. To prevent tissue over-digestion, the best incubation time for each experiment was established by monitoring cell viability, cell debris and aggregates every 10 minutes (usually ranges from 1 to 1.5 hour). After completed digestion cells were washed once by topping up to 2 ml with RPMI + 10% FBS and spun down at 300 g. The cell pellet was resuspended in RPMI + 10% FBS and dispersed by passing through a 75pm cells strainer, followed by an additional wash of the filter with RPMI + 10% FBS. Finally, the cells were spun down at 300 g and resuspended in 1 ml of ice-cold PBS for counting with a hemocytometer.
[0141] Histological and immunofluorescence evaluation
[0142] One third portion of each research biopsy tissue specimen was embedded in Killik (Bio-Optica, 05- 9801), immediately frozen in precooled isopentane (MilliporeSigma, 277258) and stored at -80°. OCT- embedded biopsy cores (one per patient) were serially sectioned with 10 pm thickness in a cryostat at -20°C. Ten slides per patient, containing multiple sections representing distinct regions of the same tissue were prepared in parallel and stored at -80°C. Hematoxylin and Eosin staining (H&E) was performed using H&E Staining Kit (Abeam, ab245880). The H&E-stained slides were reviewed by an expert genitourinary pathologist to assign the Gleason Score according to the International Society of Urinary Pathology grading system. The same pathologist evaluated all specimens (diagnostic and research biopsies) used in the invention.
[0143] For immunofluorescence, frozen tissue sections were briefly thawed at room temperature (RT), placed in precooled acetone for 20 minutes at -20°C and washed 3 times in PBS 1X for 2 minutes at RT. To permeabilize tissues the sections were incubated with 0.5% Triton X-100 in PBS 1X for 10 minutes at RT with mild agitation. After three washes in PBS 1X for2 minutes at RT, coverslips were blocked by incubating with 5% Donkey serum, 3% BSA in PBS 1X for 1 hour at RT. After three washes in PBS 1X for 2 minutes at RT, coverslips were incubated with 1 :500 Lamin A / C primary antibody (Santa Cruz sc-6215) 1 :500 in blocking solution overnight at 4°C in a humidified chamber, isolating the tissue section by a hydrophobic pen. The following day, sections were washed 3 times in PBS 1X for 2 minutes at RT and incubated with fluorescence-conjugated secondary antibody (Alexa FluorTM 488-conjugated Donkey Anti-goat IgG Invitrogen, A-31571), 1 :500 3% BSA in PBS 1X for 1 hour at RT in a dark humidified chamber. After three washes in PBS 1X for 2 minutes at RT, the coverslips were incubated with Hoechst solution (Thermo Fisher Scientific H3570, 1 :500) for 10 minutes at RT in the dark, rinsed in PBS 1X and mounted on microscope slides with Prolong antifade mountant (Thermo Fisher Scientific, P36930).
[0144] Images were acquired using a Leica TCS SP5 confocal microscope with a HCX Plan Apo *63 / 1.40 objective, with a 5-zooming factor. The NIS-Elements v.5.30 software (Nikon-Lim) was employed to analyze nuclei physical parameters including area and circularity. Hoechst staining facilitated nucleus identification and delineation of the region of interest. Over 5 field of views (FOVs) were acquired and analyzed per sample, totaling the examination of more than 70 nuclei within each sample. The mean nuclear area or circularity was calculated for each sample and plotted alongside other samples within the same group on the charts. The analysis of Hoechst and Lamin A / C distribution across nuclei was conducted using Fiji software. Leveraging the plot profile plugin, a 10 pm fixed line was implemented to trace and capture the fluorescence intensity distribution across each nucleus. To standardize comparisons between diverse samples, a normalization step was performed. This involved calibrating the fluorescence intensity of individual pixels along the line to the mean intensity derived from the entire set of pixel lines. For each sample the mean distribution values for both Hoechst and Lamin A / C were calculated and graphically represented alongside other samples within the same group.
[0145] Flow cytometry Analysis
[0146] To quantify the relative amounts of cell populations in each biopsy, 10.000 cells from the digestion step were stained, acquired on a BD FACSCantoll Flow Cytometer and analyzed with FlowJo software in the INGM FACS facility. To avoid unspecific binding, antibodies were incubated with PBS-BSA 1 % for 30 minutes at 4°C. TO-PROO-3 stain (Thermofisher, T3605) was used to assess cell viability. Tissue resident leukocytes were identified as CD326- / CD45+ (EpCAM-CD326 FITC, B347197; CD45 Pacific Blu Biolegend, 982306); epithelia cells as CD326+ / CD45-, and stromal cells as CD326- / CD45-.
[0147] Statistical analysis on flow cytometry, imaging data and clinical parameters
[0148] The data in Figure 6D, Figure 6E, Figure 7B, Figure 7G, Figure 9C, and Figure 9E were visualized using Graph Pad Prism software. A two-tailed Student’s t-test was used to compare two groups and a one-way ANOVA for multiple comparisons involving three groups. The statistical significance was label as P value <0,05 (*), and ns P value>0,05.
[0149] 4f-SAMMY-seq isolation of chromatin fractions
[0150] Chromatin fractionation on tissue-derived single cell suspension was performed with minor adaptations to the protocol described in available version of 4f-SAMMY-seq protocol23. Cells were counted, washed in cold PBS and resuspended in 600ul cold cytoskeleton triton buffer, CSK-Tr: 10 mM PIPES pH 6,8; 100 mM NaCI; 1 mM EGTA; 300 mM Sucrose; 3 mM MgCI2; 1X Protease Inhibitor Cocktail (Roche, 04693116001); 1 mM PMSF (Sigma-Aldrich, 93482); 1 mM DTT; 0,5% Triton X-100. After 10 minutes on a rotator at 4°C, samples were centrifugated for 3 minutes at 900g at 4°C and supernatant was collected as S1 fraction. Pellets were washed for 10 minutes on the wheel at 4°C with an additional volume of the same buffer followed by centrifugation for 3 minutes at 900g at 4°C. Chromatin was then digested by using 25 units of DNase I (Invitrogen, AM2222) in 100ul of cytoskeleton CSK buffer: 10 mM PIPES pH 6.8; 100 mM NaCI; 1 mM EGTA; 300 mM Sucrose; 3 mM MgCI2; 1 mM PMSF; 1X Protease Inhibitor Cocktail for 60 min at 37°C. To stop digestion, ammonium sulphate was added to samples to a final concentration of 250 mM and, after 5 min on ice, samples were pelleted at 900g for 3 min at 4°C and the supernatant was collected as S2 fraction. After a wash of 10 minutes on a rotator at 4°C with 200ul of CSK buffer followed by centrifugation for 3 minutes at 3000g at 4°C, the pellet was further extracted 10 min on a rotator at 4°C in 100ul of CSK-NaCI buffer: 10 mM PIPES pH 6.8; 2 M NaCI; 1 mM EGTA; 300 mM Sucrose; 3 mM MgCI2; 1 mM PMSF; 1X Protease Inhibitor Cocktail, centrifuged at 2300 g 3 min at 4°C and the supernatant was collected as S3 fraction. Pellets were washed twice for 10 minutes on the wheel at 4°C followed by centrifugation at 3000 g 3 min at 4°C with 200ul of CSK-NaCI buffer. Finally, pellets were solubilized in 10Oul of 8M urea for 10 minutes at room temperature and labelled as S4 fraction. Fractions were stored at -80°C until DNA extraction.
[0151] 4f-SAMMY-seq DNA extraction, library preparation and sequencing
[0152] Chromatin fractions (S2, S3 and S4) were diluted 1 :2 in 1X TE buffer (10mM TrisHCI pH 8.0, 1 mM EDTA) and incubated with 3 U of RNAse cocktail (Ambion, AM2286) at 37° for 90 minutes, followed by 40pg of Proteinase K (Invitrogen, AM2548) at 55° for 150 minutes. Genomic DNA was then isolated using phenol / chloroform (Sigma-Aldrich, 77617) extraction followed by a back extraction of phenol / chloroform with additional volume of TE1X. DNA was precipitated in 2 volumes of cold ethanol, 0.3M sodium acetate and 20ug glycogen (Ambion AM9510) for 1 hour on dry ice or overnight at -20°C. Dry pellets were resuspended in 50 pl (S2) or 15 ul (S3 and S4) of nuclease-free water and incubated at 4°C overnight. On the next day, S2 was further purified using PCR DNA Purification Kit (Qiagen, 28106) and separated using AMPure XP paramagnetic beads (Beckman Coulter, A63880) with the ratio 0.90 / 0.95 to obtain smaller fragments conserved as S2S (< 300 bp) and larger fragments labelled as S2L (> 300bp) fractions. Both were resuspended in 20 ul of nuclease-free water and then reduced to 15ul using a centrifugal vacuum concentrator. S2L, S3 and S4 fractions were sonicated in a Covaris M220 focused-ultrasonicator using screw cap microTUBEs (Covaris, 004078) to obtain a smear of DNA fragments peaking at 150-200 bp (waterbath 20°C, peak power 30.0, duty factor 20.0, cycles / burst 50, 150 seconds for S2L and 175 seconds for S3 and S4). Fractions were quantified using Qubit 4 fluorometer with Qubit dsDNA HS Assay Kit (Invitrogen, Q32854). Libraries were generated from each sample using NEBNext Ultra II DNA Library Prep Kit for Illumina (NEB, E7645L) and Unique Dual Index NEBNextMultiplex Oligos for Illumina (NEB, E6440S). Libraries were then qualitatively and quantitatively checked on and run on an Agilent 2100 Bioanalyzer using High Sensitivity DNA Kit (Agilent, 5067-4626). Libraries with distinct adapter indexes were then multiplexed and, after cluster generation on FlowCell, sequenced for 50 bases in single or paired-end reads mode on an HluminaNovaSeq 6000 instrument at the IEO Genomic Unit in Milan.
[0153] RNA extraction, library preparation and sequencing
[0154] Ten thousand cells retrieved from the digestion step were stabilized in 200pl of IThioglycerol / Homogenization Solution of the Maxwell® RSC miRNA Tissue Kit (Promega, AS1460) and stored frozen at -80°C for total RNA automated purification using Maxwell® RSC 48 Instrument (Promega, AS8500) according to manufacturer’s instructions. Total RNA was quantified by Qubit 4 fluorometer with Qubit RNA HS Assay Kit (Invitrogen, Q32852) and assessed by Agilent 2100 Bioanalyzer using Agilent RNA 6000 Pico Kit (Agilent, 5067-1513) to inspect RNA integrity. For each sample, 1 ng of total RNA was used to construct strand specific RNA-seq libraries with SMARTer Stranded Total RNA-Seq Kit - Pico Input (Takara, 634487). The yield and quality of the libraries were evaluated on Agilent 2100 Bioanalyzer using High Sensitivity DNA Kit (Agilent, 5067-4626). RNA- seq libraries were sequenced on the Illumina NextSeq™ 550 system at the sequencing facilities of Humanitas and Fondazione IRCCS Ca' Granda-Ospedale Maggiore Policlinico in Milan in paired- ends mode.
[0155] ChlP-seq public datasets
[0156] The ChlP-seq data of histone post translational modifications (histone marks) for non-neoplastic prostate gland tissue were downloaded from the following datasets available on ENCODE data portal27: ENCSR763IDK (H3K27ac), ENCSR768PFZ (H3K36me3) and ENCSR133QBG (H3K9me3). These data were downloaded as raw high-throughput sequencing reads (FASTQ) and analyzed as described in the next sections.
[0157] ChlP-seq and 4f-SAMMY-seq high-throughput sequencing reads data analyses
[0158] High-throughput sequencing reads were trimmed using Trimmomatic (v0.39)3with the following parameters for 4f-SAMMY-seq and ChlP-seq: 2 for seed_mismatch, 30 for palindrome_threshold, 10 for simple_threshold, 3 for leading, 3 for trailing and 4:15 for sliding window. The sequence minimum length threshold of 35 was applied to all datasets. The Trimmomatic “TruSeq3-SE.fa” (for single-end reads) and “TruSeq3-PE-2.fa” (for paired-end reads) were used as clip files. After trimming, reads were aligned using BWA (v0.7.17-r1188)4setting -k parameter as 2 and using as reference genome the UCSC hg38 genome (only canonical chromosomes were used) and the output saved in BAM file format. Only a single read per DNA fragment was uniformly used and aligned for both single-end and paired-end sequencing samples. The PCR duplicates were marked with Picard (v2.22; https: / / github.com / broadinstitute / picard) MarkDuplicates option, then filtered using Samtools (v1.9)5. In addition, all the reads were filtered with mapping quality lower than 1. Each sequencing lane was analyzed separately and then merged at the end of the process.
[0159] Reads distribution profiles analysis
[0160] To compute reads distribution profiles (genomic tracks) for individual fractions of 4f-SAMMY-seq, Deeptools (v3.4.3)60bamCoverage function was used. For these analyses the genome was binned at 50bp, the reads extended up to 250 bp and Deeptools RPKM normalization method was used. A genome size of 2,701 ,495,761 bp (value suggested in the Deeptools manual https: / / deeptools.readthedocs.io / en / latest / content / feature / effectiveGenomeSize.html) was considered and regions known to be problematic in terms of sequencing reads coverage using the blacklist from the ENCODE portal (https: / / www.encodeproiect.org / files / ENCFF356LFX / ) were excluded.
[0161] To compute the genomic tracks for ChlP-seq IP over INPUT enrichment profiles (log2normalized ratio) the SPP R package (v1.16.0)61and R statistical environment (v3.5.2) weew used. The reads were imported from the BAM files using the “read.bam.tags” function, then filtered using “remove. local.tag. anomalies” and finally the comparisons were performed using the function “get. smoothed. enrichment. mle” setting “tag. shift = 0” and “background. density.scaling = TRUE”. The resulting enrichment signal corresponds to a Iog2 normalized ratio between the pair of sequencing samples.
[0162] For computing the relative comparisons between two 4f-SAMMY-seq fractions (relative enrichment, i.e. Iog2normalized ratio) the same procedure described above for ChlP-seq IP over INPUT enrichment profiles was used. In particular, the "solubility profile" was defined as the relative enrichment ratio of 4f-SAMMY-seq sequencing reads distribution along the genome for S2L vs S3 fractions. S3 was used as baseline in this ratio as it is the second most consistent fraction among controls (Figure 8D).
[0163] To compute correlations between genomic tracks, it was used R (v3.5.2) base function "cor" with “method = Spearman”. The genomic tracks were imported in R using the rtracklayer (v1.42.2)62library. Then the original genomic track files with 50bp resolution were re-binned by averaging data at 150kb resolution using the function “tileGenome” and the correlation was computed per chromosome. The correlation values obtained for each chromosome were then summarized in one value describing the genome-wide sample correlations through a weighted mean, where the weight of each chromosome corresponds to its length.
[0164] Open histone marks (H3K4me1 , H3K36me3) peaks were called with MACS (v2.2.9.1) using a broadcutoff of 0.1 (9). H3K9me3 peaks were defined using the EDD (v1 .1 .19) software10with parameters (binsize = 150 Kb and gap penalty = 25) processing the filtered bam files obtained as described above. The “required_fraction_of_informative_bins” parameter was set to 0.98. The blacklisted regions were defined as for the reads distribution profile analyses described above.
[0165] The enrichment boxplots shown in Figure 2E was produced retrieving for each histone mark peak the genomic bins that fall in that genomic region. For each group of patients (CTR, LDD and HDD groups) and for each genomic bin we plot the mean 4f-SAMMY-seq solubility profile across the group. The values used were the quantile normalized tracks (across all samples) S2LvsS3 produced using SPP (see “High-throughput sequencing reads analyses”).
[0166] Genomic tracks obtained by SPP or Deeptools were saved in bigWig file format.
[0167] Genomic tracks visualization
[0168] The visualization of genomic tracks was performed with Gviz R library (v1 ,26.5)64. The track profile was calculated using the function “DataTrack” (the input bigWig file was imported using the function “import” of the rtracklayer library) and plot using the function “plotT racks”. Figure 8E has been created plotting each track as type "polygon" and all the tracks have been set using a window of 500 except for ChlP-seq of H3K27ac where the window has been set to 5000. It is worth remarking that the "window" parameter in the "DataTrack" function is referring to the number of intervals in which the displayed region is divided to plot the profile. As such, a larger number indicate a finer grain definition of the profile. As H3K27ac is known to have sharp peaks, the inventors aimed to plot a finer grain profile. For Figure 1C the window was set to 1000 and for all the tracks except for ChlP- seq of H3K27ac where 5000 was used, i.e. the same value used for the Figure 8E. CTR and PCa samples track overlayed with confidence interval were drawn using at the same the time the types “a” and “confint”. The same settings of histone mark tracks of Figure 1C were used for Figure 2B.
[0169] Chromatin compartment analysis
[0170] Chromatin compartments were calculated using a revised version of the CALDER algorithm (version 1.065), as implemented in the original 4f-SAMMY-seq pipeline for the sub-compartment identification 23
[0171] Namely, for each chromosome: i) the four 4f-SAMMY-seq fractions (S2S, S2L, S3, S4) reads distribution profiles were calculated, binned at 150kb and normalized with RPKM (see "Reads distribution profiles analysis" section); ii) for each genomic bin, defined as a vector containing the four RPKM values (one for each fraction), the Euclidean distance (dist, R stats package, method="euclidean") was calculated with all the other bins of the same chromosome. These steps produced an NxN matrix (hereinafter referred to as biochemical similarity matrix), where N is the number of bins for the considered chromosome. Starting from the biochemical similarity matrix, the CALDER procedure was performed as described in23to derive the eigenvector and to reconstruct chromatin compartmentalization. Specifically, the inventors called sub-compartments but limiting their segmentation to the highest level, thus obtaining only 2 compartments corresponding to "A" and "B" compartments (Figure 2A-B).
[0172] The consensus of compartment calls across healthy controls has been produced labeling each genomic bin (150kb) according to its more frequent compartment across samples: e.g. if a bin is labeled as A in at least 4 out of 7 CTR samples, that bin will be defined as A in the consensus of controls, and vice versa for bins labelled as B in at least 4 controls (Figure 2C). The definition of compartment shifts was based on the concordant or discordant compartment classification for each 150kb genomic bin (Figure 9A). The compartment domains shared across patients as shown in Figure 9A were reported using the library upsetR.
[0173] The clustering of patients in Figure 2D were based on Jaccard Indexes (JI) of compartments concordance for each pair of patients. The resulting matrix of pairwise JI was grouped by hierarchical clustering based on the Euclidean distance (among rows or columns) and complete linkage using the pheatmap function in the pheatmap R library.
[0174] Cell type deconvolution
[0175] For cell type deconvolution analysis, inventors normalized for each sample raw count to TPM (Transcript per Million) by using R 3.6.1 environment and applying the following steps. First, read counts for each gene were normalized (divided) by the length of the gene in kilobases (RPK, Reads per kilobase). Then RPK values in a sample are summed and divided by 1.000.000 to obtain the scaling factor. Finally, RPK values for each gene are divided by the scaling factor.
[0176] For the deconvolution of the 3 macro populations (immune, epithelia, stroma cells) it was defined a custom signature matrix based on single cell expression data of prostate tissue from the Human Proten Atlas (https: / / www.proteinatlas.org / ; V. 21.1) (13-15). Following the cluster classification of Human Protein Atlas we considered as epithelia the subpopulations annotated as prostatic glandular cells, basal prostatic cells or urothelial cells. It were considered as stroma the populations annotated as muscle cells, fibroblast or endothelial cells. It were considered as immune cells infiltrate the populations annotated as t-cells or macrophages. Finally, inventors selected as representative genes for each population only those with a TPM > 10 for the specific population and a TPM < 2 in all the two remaining populations. A macropopulation signature matrix of 702 genes was obtained. The deconvolution analysis was run exploiting this signature matrix on our bulk data by using the relative mode of CIBERSORTX (V.1.0) (16) (Figure 9F).
[0177] To compute the immunity score on our biopsies the software CIBERSORTX (V.1.0) (16) was used applying the previously computed TPM matrix of our bulk RNA (mixture file) on the LM22.txt provided in the CIBERSORTX repository. The LM22 leukocyte signature is based on a matrix of 547 genes discriminating 22 human hematopoietic cell types isolated from PBMC. We enabled batch correction by using the B-mode correction option. We next run the analysis in absolute mode by applying 500 permutations. Each sample in the mixture file gets a score that reflects the absolute proportion of each cell types in LM22 mixture file. We finally computed the absolute leukocyte score: i.e. the sum for each sample of all the 22 scores of leukocytes populations. The absolute leukocyte score for each biopsy was compared across the different groups (CTR, LDD, HDD) (Figure 9G).
[0178] Copy Number Analysis (CNA)
[0179] For copy number analysis paired-end 4f-SAMMY-seq reads were used by treating fractions from the same sample as independent replicates. Data was processed with the pipeline nf-core / sarek v3.4.066’67of the nf-core collection of workflows68, using whole genome sequencing settings, and segmentation was called with CNVkit v0.9.1069.
[0180] Transcriptomic analyses
[0181] For RNA-seq data, the overall quality of the sequenced reads was assessed using FastQC tool (http: / / www.bioinformatics.babraham.ac.uk / proiects / fastqc) (V. 0.11.8), then reads were trimmed with Trimmomatic software57(version 0.39) removing the adapters (ILLUMINACLIP:Picov2smart- PE.fa), primer dimers, and low quality bases at the beginning and at the end of the reads (trimmomatic PE phred33 LEADINGS TRAILING:3 SLIDINGWINDOWD:4:15 MINLEN:36). STAR70
[0182] (V. 2.7.0f_0328) was used to index (STAR --runMode genomeGenerate) the Human Genome (GENCODE Release 39, GRCh38 primary assembly genome71) and to align sequenced reads in paired-end mode (-readFiles R1.FASTQ R2.FASTQ) on the indexed reference genome. Multimapping reads and PCR duplicate were marked in the final output (-- bamRemoveDuplicatesType Unique) and unaligned reads stored in a different file (- outReadsUnmapped Fastx).
[0183] The read counts on genes were calculated using as a reference a GTF file with RefSeq annotation downloaded from UCSC (http: / / qenome.ucsc.edu) stored in the following directory. (http: / / hqdownload.soe.ucsc.edu / qoldenPath / archive / hq38 / ncbiRefSeq / 109.20190905 / hq38.109.20 qtf. This file was further processed to remove non canonical and mitochondrial chromosomes, selected only curated genes (NM, NR) and finally split in protein coding
[0184] (NM) and noncoding (NR) files. Reads count was performed with HTSeq-count (V. 0.13.5) on bam files (previously generated by STAR) using as a feature the union of all exons in a gene. The type of library was specified with “-s reverse” parameter. The reads that align to more than one position in the reference genome were discarded (htseq-count -non-unique none). The full matrix with raw read counts for each sample were loaded in R 3.6.1 and normalized using DESeq272median of ratios. Differential expression analysis was performed with DESeq2V. 1.2673) using Wald test and the Benjamini and Hochberg correction for multiple tests, to compute p-values and adjusted p-values, respectively.
[0185] Enrichment analyses
[0186] Up-regulated genes (with adjusted p-value < 0.05 and l_og2Fold Change > 1) and Down-regulated genes (with adjusted p-value < 0.05, and l_og2Fold Change < -1) were used separately to query the gene-list enrichment tool Enrichr (https: / / maayanlab.cloud / Enrichr / ) '6over the main Gene
[0187] Ontology classes (Biological Process - BP, Molecular Function - MF, Cellular Component - CC) updated to 2023 taking as final enriched terms only those with Benjamini-Hochberg adj.p-value < 0.05. Genes within the recurrent compartment switch “A to B” and “B to A” were also queried over the main Gene Ontology classes (BP, MF, CC) taking also in this case the enriched terms with Benjamini-Hochberg adj.p-value < 0.05. GSEA analysis39was performed in Preranked mode. Starting from raw p-values of all genes we generated a ranked matrix sorted by the score resulting from the following formula: -Logio(p-value) multiplied by the sign of the l_og2Fold Change. For the gene set references we downloaded H (HALLMARK), C2 (Curated) and C5 (Ontology) GMT files from MSigDB. Finally, we performed the analysis using the parameters "Number of permutations": 1000; "Collapse": No; "Enrichment Statistic": classic; "Max size": 500; "Min size": 15.
[0188] Resulting geneset with a family-wise error rate (FWER) < 0.05 are finally sorted using the Normalized Enrichment Score (NES).
[0189] Pathway activity was then inferred for the comparison of studied tumor subtypes by using the R package decoupleR77(V. 2.4). The get_progeny (organism=human,top=100) function was used to get pathways from PROGENy (V.1.2O)40a curated collection of signaling pathways derived from perturbation experiments, where each gene has a specific weight describing its positive or negative response to a given pathway stimulation. The top 100 responsive genes in the compendium ranked by p-value were used for each pathway. Starting from Wald statistic values previously computed for each gene with DESeq2 (LDDvsCTR, HDDvsCTR and HDDvsLDD comparisons) inventors run the function decouple with parameters statistics=c(“mlm”,”ulm”,”wsum”) and consensus_score=TRUE, to estimate a consensus score across the top performer methods (according to benchmark done by decoupleR developers) in pathway activity prediction. Multivariate linear model (mlm), Univariate linear model (ulm), Weighted Sum (wsum) are recommended by developers of decoupleR since these methods better estimate pathway activity taking into account the weights associated with pathway related genes.
[0190] TOGA Data
[0191] TCGA-PRAD RNA-Seq expression data were downloaded by using gdc-client tool on gdc_manifest.2022-08-08.txt metafile input downloaded from the GDC portal repository (https: / / portal.qdc.cancer.gov / ; V.34.0; Release July, 27, 2022) by selecting from Web interface the TCGA project, prostate and RNA-seq count. The main clinical parameters (Gleason Score, T stage, N stage and age) and info about the samples (primary tumor, normal biopsy and metastatic) were selected from the same Web section of GDC portal repository by downloading clinical and samplesheet files. ABSOLUTE tumor purity score78was downloaded from https: / / gdc.cancer.gov / about- data / publications / pancanatlas. The file is TCGA mastercalls. abs tables JSedit.fixed.txt (Downloaded October ,13, 2022) /
[0192] Excluding primary normal biopsies obtained from the TCGA repository described above, the 500 FPKMUpperQuartile normalized expression counts from prostate primary tumors. The Biochemical Recurrence Free (BCR) status for each patient was obtained from PCaDB (http: / / bioinfo.iialab-ucr.org / PCaDB / ; Release May, 10 2021 ; eSet V.1.3), downloading the TCGA- PRAD_eSet.RDS object and applying the pData function from Biobase package. We then excluded the only sample with a Gleason Score pattern of 2+4 never described in literature.
[0193] Despite expression analyses were performed assigning a PCI score on the entire cohort of 499 samples, the BCR data available are totally 466.
[0194] PCI score (Prostate Compartmentalization Index)
[0195] Starting from our 101 DEGs of the HDD vs LDD comparison two gene lists were defined with 24 HDD Up-regulated and 77 HDD Down-regulated genes. We subset the TCGA matrix extracting only those genes matching with the full list of 101 genes. A log transformation of TCGA gene expression matrix was performed: Log2(matrix +1). The log transformed expression of each gene across the 499 samples was standardized into a z-score. For each sample 2 median values were computed on the z-score transformed matrix: the median of the 24 HDD Up-regulated genes (HDDJJP) and the median of the 77 HDD Down-regulated (HDD_DOWN) genes. Finally, the PCI score defined as the difference of median values: PCI=median(HDD_UP) - median(HDD_DOWN) was computed. The resulting PCI score with a positive value will indicate that a sample has HDD-like transcriptomic features, whereas a PCI negative value will classify a sample as LDD-like.
[0196] Survival Analyses
[0197] Univariate and multivariate regression analysis were performed in R using survival and survminer packages. First, a log-rank test was performed to compare survival Kaplan-Meier curves (in terms of BCR STATUS AND BCR TIME) between HDD-like and LDD-like samples. Then a multivariate Cox proportional-hazard model was performed including the following covariates: HDD-LDD status (previously assigned according to the PCI score) of the samples, pathological stage, age at diagnosis, Gleason Score and Tumor purity score. For the multivariate Cox regression analysis, the Gleason Score was included as numeric feature by summing the primary and secondary numeric grades, as reported for TCGA samples.
[0198] Signature Refinement
[0199] In order to refine the 101 gene signature, univariate cox regression analysis was applied on each single gene expression values across the TCGA samples, including in the model the BCR STATUS and BCR TIME. The p-values obtained for each gene are then adjusted with Benjamini-Hochberg. Then only those genes whose expression is associated with Hazar Ratio (HR) < 1 or HR > 1 with a adj.p-value < 0.05 were selected. 21 candidate prognostic genes were obtained. Unsupervised clustering of these genes is able to cluster 2 main groups of samples with HDD-like and LDD-like features with high significance in the prognostic status of the two groups (Figure 11 A). The signature was further refined by removing 3 genes (NPIPB3, CCNE2, FMO5) that were noisier in the unsupervised clustering heatmap: these are the only genes that cluster in a discordant manner compared to their original association to HDD or LDD phenotype in the dataset (see Figure 11 A). Finally, the resulting refined panel of 18 genes was used to recompute and validate the PCI score on other independent cohorts with transcriptomic profiles of prostate cancer patients.
[0200] Validation
[0201] Through the PCaDB (http: / / bioinfo.iialab-ucr.org / PCaDB / Release May, 10 2021 ; eSet V.1.3) normalized expression matrices and BCR data of three independent datasets of prostate cancer: Emory (GSE5446O)80, Stockholm (GSE70769)81, Belfast (GSE1 16918)82were downloaded. For each of them samples were classified computing the PCI score based on the refined 18 genes signature. Immunohistochemistry (IHC)
[0202] Representative 4-pm-thick sections were cut from each block of 8 patients with PCa and stained with H&E or primary antibodies for SC5D (Atlas Antibodies HPA066283), IQGAP2 (Abeam ab187153), COL5A1 (Abeam ab7046), COL16A1 (Invitrogen PA5-103382) and COL18A1 (Abeam ab275746). Specifically, positive and negative controls were included in each experiment and the staining was revealed using DAB as chromogen using a Ventana Benchmark Ultra instruments with dedicated reagents (Ventana, Roche; Righi, I. et al. Immune Checkpoints Expression in Chronic Lung Allograft Rejection. Front Immunol 12, 714132 (2021)). All slides were counterstained with hematoxylin and digitalized using an Aperio scanner at 40x magnification (Leica Microsystems). Presence of staining for all antibodies was evaluated in the epithelial (normal and tumor glands), and stromal compartments. The scoring of the staining was performed manually. Briefly, two pathologists (W and MMa) independently analyzed the slides from all patients in a blind mode. Three regions per slide were analyzed for markers presence. For SC5D and IQGAP2 the intensity of the staining and the percentage of positive cells were used to calculate the IHC score (% of positive cells X intensity). For COL5A1 , COL16A1 and COL18A1 the specific staining in different cellular compartments was evaluated.
[0203] Cell culture
[0204] Cells were obtained by ATCC and grown either in DMEM supplemented with 10% (VCaP CRL-2876) or 5% (WPMY-1 CRL-2854) of FBS at 37°C in a humidified atmosphere of 5% (v / v) CO2. Cell lines were tested transcriptionally with following primers: SC5D (fw: cagtggcgtgatcttggttc SEQ ID NO 1 ; rv: ate etg get aac acg gtg aa SEQ ID NO 2); IQGAP2 (fw: ccccgatgcaacagaagatg SEQ ID NO 3 rv: tgaagaaccttggccactga SEQ ID NO 4); COL5A1 (fw:gtggtctgaagggcacaaag SEQ ID NO 5; rv: tccgagttttcccttctccc SEQ ID NO 6); COL5A3 (fw: tgaatgtcgtgcagctgaac SEQ ID NO 7; rv:agctcctctccattggt SEQ ID NO 8); COL18A1 (fw: cagaaaggagccaagggaga SEQ ID NO 9; rv: ccctcatgccaaatccaagg SEQ ID NO 10); MXRA8 (fw: tgacttctgaccctgacacc SEQ ID NO 1 1 ; rv: caaatgtggggtgctgagag SEQ ID NO 12). WPMY-1 prostate stromal cells were transducted with the pLenti-Lifeact-tdTomato (Addgene plasmid # 64048) and labelled cells isolated by fluorescence- activated cell sorting (FACS). 45000 tdTomato- WPMY-1 cells were seeded in each 24-well tissue culture plate the day prior transfection. For each experimental point, 24 pmoles of either negative control (SCR, Life Technologies # AM461 1) or a mix of COL5A1 (Life Technologies ID: 7797, 7892); COL5A3 (Life Technologies ID: 23255, 23350); COL18A1 (Life Technologies ID: 43706, S37393); MXRA8 (Life Technologies ID: 34039, HSS147729) specific siRNAs (6 pmoles each) were transfected using the RNAiMAX Transfection Reagent (Life Technologies, #13778075). Forty-eight hours after transfection cells were analyzed by qRT-PCR to assess the silencing efficiency and cocultured with VCaP cells.
[0205] Wound healing assay
[0206] Thirty-six hours after silencing, TdTomato-WHPY cells were trypsinized, counted and 200.000 cells / experimental point mixed with VCaP cells in a 1 :1 ratio. Cells were then seeded on Matrigel- coated wells. The following day (corresponding to forty-eight hours after silencing), a scratch in the confluent monolayer at the centre of each well was made with a 200 pL pipette tip. Cells were then washed twice with sterile DPBS 1X to remove debris and displaced cells prior, then cultured in serum-free DMEM. Live imaging to study wound-healing phenomena were carried out on four independent biological experiments. An environmental controlled (OKO-cage incubator; OkoLab) Nikon widefield Ti instrument (Nikon Europe) equipped with fast spinning disk (Crest Optics) and high-sensitivity video cameras (Zyla 4.2 and iXON DU888, both from Andor- Oxford Technologies) and both 16-LED (Pe-4000; CoolLed Ltd) and 4-laser excitation (LDI, Ltd) was used for image acquisition. In particular, we used widefield modality with white led for differential interference contrast imaging (DIC) and 555nm fluorescence led line for fluorescence contrast imaging. Images were acquired with 3h time-frame with a total of 45h recording time-lapse. Acquisitions were conducted over an ad-hoc designed semi-automatic pipeline within the JOBS module of NIS- Elements v.5.30 software (Nikon Europe). Analysis was conducted with ad-hoc designed pipeline for segmentation and quantification, using the GA3 module in NIS-Elements v.5.30: briefly an algorithm for automatic wound-area detection over time was employed and specific binary masks were created to segment and classify each cell object detected either as DIC cell-object or as red- fluorescent cell-object. Measurements detected from the segmentation approach were further elaborated in a blind mode, to evaluate the percentage of wound closure and the numbers of cells within wound area, in all conditions. For image and dynamics representation, flattened RGB images and movies showing merged DIC and fluorescence channels, were selected from all samples. For video representation, a 10fps frame rate was chosen.
[0207] EXAMPLES
[0208] Example 1 - Experimental design and study cohort
[0209] In order to characterize chromatin architecture in prostate cancer patients at the time of diagnosis, fresh needle biopsies were collected from chemo-naive patients undergoing the standard diagnostic procedure. Each patient was punctured to collect 14 needle biopsy cores for histopathological examination (diagnostic biopsies) and 1 to be processed for our research project (research biopsy - Figure 1 A). Each research biopsy was divided in two parts: one third used for histopathological and immuno-histological analyses, while the rest of the tissue sample was enzymatically digested to obtain a cell suspension further split for epigenomic (4f-SAMMY-seq), transcriptomic (RNA-seq) and cellular composition (flow-cytometry) analyses.
[0210] Study cohort includes 7 non-neoplastic controls (CTR), i.e. patients that turned out to be negative for cancer, and 10 primary prostate cancer patients (PCa) with histologically confirmed tumor and Gleason Score (GS) between 3+3 and 5+5 (Figure 1 B). In Figure 1 B is also reported the percentage of diagnostic biopsy cores positive for cancer cells (PPC), their distribution in a prostate gland diagram (DPC), as well as the puncture site of the research biopsy used for this study. All the research biopsies of prostate cancer patients were obtained from punctures adjacent to one or more positive diagnostic biopsy cores (Figure 1 B). In addition to the Gleason Score assigned to the patient in the clinical report (GS-P), the Gleason Score assigned to the closest diagnostic biopsy (GS-CDB) and the score assigned to the research biopsy (GS-RDB, Figure 6A) were recorded. The histopathological examination on research biopsies did not confirm the presence of tumor cells in four out often cases (Figure 6A). However, it must be noted that the Gleason Score assignment was done on a third of the tissue biopsy and that all research biopsy samples express diagnostic PCa tumoral markers like PCA3, HPN and GOLM124 26(Figure 6B). Therefore, all the ten PCa samples were retained in the experimental analyses.
[0211] Study cohort of PCa patients is quite homogeneous in terms of age (median 72 years; interquartile range, abbreviated as IQR, 62-89 years) and serum PSA level (median 16 ng / mL, IQR: 7.1-45.5), but heterogeneous in terms of clinical features as percentage of cancer positive biopsy cores (median: 21.4%, IQR: 21.4-39.2), clinical stage (T1c: 60%, T2a-T2b: 20%, T3: 20%) and index lesions on Magnetic Resonance Imaging (MRI) (40% with Prostate Imaging Reporting and Data System (PI-RADS) >3 and 60% with PI-RADS 1 / 2) (Table 1).
[0212] Table 1. Clinicopathological characteristics of the study participants undergoing prostate biopsy
[0213] GS-P: Gleason Score for patients’ diagnosis
[0214] PSA: Prostate Specific Antigen
[0215] PI-RADS: Prostate Imaging Reporting and Data System
[0216] Survival outcomes are not available due to the long clinical course of primary PCa and the short follow-up time since the biopsies were collected (median time from biopsy less than 4 years).
[0217] Tissue samples were also examined by confocal microscopy. Non-neoplastic controls and PCa patients' cells showed similar nuclear area, but a significant difference in nuclear circularity with PCa patients cells showing nuclei deformations (Figure 6C-D). Staining of the nuclear lamina with Lamin A / C antibody highlighted a slight depletion of the signal in the nuclear interior corresponding to a parallel decrease in the Hoechst staining (Figure 6E). Although descriptive, these data suggest a cancer specific remodeling of the nuclear and genome organization in PCa cells.
[0218] When dissociating the needle biopsy tissue samples (size around 2 cm in length, Figure 7A, see methods) a viable cell yield in the range of 30.000-80.000 cells per sample (Figure 7B) was achieved. As a further control, the cell types composition was examined in a set of 13 prostate biopsies by flow cytometry (Figure 7C-F). No significant differences were detected in the relative proportion of epithelial (EpCAM / CD326+, CD45-), leukocyte (EpCAM / CD326-, CD45+) or stromal cells (EpCAM / CD326-, CD45-) between non-neoplastic controls and PCa samples (Figure 7G). All samples showed a large fraction of stromal cells (average 68,8% controls and 70,9% PCa patients), while an average of 18,5% and 15,2% of epithelial cells in control and PCa biopsies were found, respectively.
[0219] Example 2 - Genome architecture alterations associated with the tumoral state are patientspecific
[0220] 4f-SAMMY-seq was applied to examine genome architecture in prostate biopsies23. This method relies on the sequential isolation and high-throughput sequencing of distinct chromatin fractions separated according to their accessibility and solubility properties, which are expected to correlate with chromatin epigenetic and transcriptional status (Figure 8A). The more soluble fractions (S2S, S2L) are associated with gene activation markers and the less soluble fractions (S3, S4) are enriched for constitutive heterochromatin. To confirm the applicability of this technique on fresh prostate biopsies, 4f-SAMMY fractions of non- neoplastic control samples were first sequenced. The DNA sonication and short fragments (S2S) selection were examined to obtain high-quality libraries (Figure 8B). It was verified that sequencing reads with comparable quality across samples were obtained (Figure 8C). The sequencing reads distribution profiles for individual fractions from CTR samples show a high level of concordance (Figure 8D), with the highest correlation values for S2L fractions (median: 0.94).
[0221] To study euchromatic and heterochromatic domains the solubility profile was defined as the ratio of high-throughput sequencing reads distribution in S2L over S3 fractions (see methods): where the positive and negative values correspond to regions enriched for open and closed chromatin, respectively. The consensus solubility profiles were compared across all CTR and PCa specimens with epigenomic marks in normal prostate tissue for post translational histone modifications (histone marks) associated with active (H3K27ac, H3K36me3) and inactive (H3K9me3) chromatin states27(Figure 1C). A high similarity of solubility profiles was observed among CTR samples and a clear concordance with the location of euchromatin marks (H3K27ac and H3K36me3) along with a negative correlation with the heterochromatin mark profile (H3K9me3). These observations were also confirmed in a genome-wide quantification of correlations for individual samples (Figure 1 D and Figure 8E).
[0222] Instead, in PCa biopsies the solubility profile is more heterogeneous and its correlation with histone marks is more variable with respect to CTR samples (Figure 1C-D). Visual inspection of individual solubility profiles showed a patient-specific degree of chromatin alterations with respect to CTR samples (Figure 8E). Moreover, clear differences among PCa samples were noticed, with 5 out of 10 samples that lose the normal chromatin compartmentalization, as also highlighted by their correlation with histone marks (Figure 1 D). These differences are remarkably not attributable to cell composition of the samples (Figure 7G).
[0223] Thus, 4f-SAMMY-seq data describe a conserved genome architecture among control biopsies with a separation between euchromatin and heterochromatin. On the other hand, diagnostic early-stage PCa biopsies show a patient-specific chromatin solubility dysregulation that mirrors the highly heterogenous nature of PCa.
[0224] Example 3 - Chromatin compartments analysis identifies two groups of prostate cancer epigenomes.
[0225] To further analyze the chromatin 3D organization, the sequencing data of all the 4f-SAMMY-seq fractions were used to identify active and inactive compartments, named "A" and "B" respectively, as per the convention adopted for Hi-C derived compartments15(Figure 2A-B and Figure 8A, see methods). However, unlike Hi-C, the chromatin compartments inferred from 4f-SAMMY-seq are defined based on their chromatin accessibility and solubility properties23. CTR samples showed a consistent pattern across all 7 individuals with 27% of the entire genome with conserved active A compartment and 39% with conserved inactive B compartment (Figure 2B-C). Instead, PCa samples showed patient-specific chromatin compartment alterations (Figure 2B-C). The genomic regions with a different compartment assignment in each PCa sample with respect to the consensus among CTR samples were quantified. These compartment switches were labelled as "A to B" or "B to A" for the regions that in CTR samples were A or B, respectively. Unsupervised clustering based on compartments similarity identified two subgroups of cancer patients (Figure 2D), that, in relation to the amount of compartment alterations with respect to CTR samples, were named LDD (Low Degree of Decompartmentalization; average 5% over the entire genome for "A to B" and 5% for "B to A" compartment switches) and HDD (High Degree of Decompartmentalization; average 14% over the entire genome "A to B" and 24% "B to A"). A consistent intra-group similarity was also noted for LDD samples with 0.78 mean Jaccard Index (JI) (variance 0.001). On the other hand, HDD samples show more heterogeneity in chromatin compartmentalization (JI 0.42 mean and 0.012 variance) (Figure 2D).
[0226] In line with patient-specific genome remodeling, regions with compartment changes concordant across all LDD and HDD tumoral samples when compared to CTR consensus were not found. However, recurrent "B to A" changes concordant across all HDD patients and covering 2% of the entire genome (538 genes) was found (Figure 9A). Gene ontology analysis highlighted an enrichment of genes associated Protein Deubiquitination and Regulation of Programmed Cell Death (such as FAS) (Figure 9B). Instead, the HDD-specific "A to B" compartment switch regions (0.6% of the genome - 59 genes) contain genes involved in PI 3-kinase activities (such as PIK3R5 and PIK3R6) and lipid catabolic processes (such as PNLIP and STS), which are well-known drivers in promoting prostate cancer progression28 30.
[0227] It should be remarked that the LDD vs HDD groups stratification of PCa patients would not be immediately evident based on known clinically relevant pathophysiological features such as PSA level, Gleason Score, age at diagnosis and number of positive cores (Figures S4C). Furthermore, the analysis of the global transcriptome on prostate samples did not separate so clearly the HDD and LDD subtypes (Figures S4D). It was excluded the association of HDD and LDD subtypes with cell composition and immune cell infiltration, estimated both by flow-cytometry (Figures 9E) and by RNA-seq data deconvolution (see methods, Figure 9F-G). The distribution of copy number alterations (CNA) across all samples based on 4f-SAMMY-seq reads coverage (see methods) were also estimated. While this is not the optimal setting for reliably identifying CNAs which should be called with whole genome sequencing, the use of 4f-SAMMY-seq fractions reads for this analysis still allows to appreciate in a relative comparison that the inferred copy-number states are notably similar for both LDD and HDD subtypes (Figure 9H), suggesting comparable levels of genetic alterations.
[0228] Epigenetic differences between LDD and HDD groups were further investigated. The histone mark enrichment peaks from normal prostate tissue was compared with the 4f-SAMMY-seq solubility profiles of CTR, LDD and HDD samples. The solubility profile of CTR samples is higher in ChlP-seq peaks for open histone marks (H3K27ac and H3K36me3) and lower in constitutive heterochromatin (H3K9me3 peaks) (Figure 2E). The HDD group showed instead an inverted trend, thus confirming a large-scale chromatin architecture remodeling, whereas the LDD group showed an intermediate pattern between the other two groups.
[0229] Example 4 - LDD and HDD constitute two novel PCa molecular subtypes
[0230] The functional effects of chromatin compartments reorganization was examined by matched RNA- seq analysis. The comparison of transcription profiles between CTR (n=7) and PCa (n=10) samples uncovered 66 differential expressed genes (DEGs): 22 up-regulated and 44 down-regulated in PCa with respect to CTR (Figure 3A). Among them, 7 genes already described in literature for their association to cancer were identified: 5 up-regulated prostate cancer biomarkers (HPN, BICD1 , GOLM1 , TARP) and 2 down-regulated tumor suppressor genes (SERPINB5 and GATA6)31-36. The limited number of DEGs is likely due to the heterogeneity of expression across tumor patients' biopsies.
[0231] Then, LDD and HDD subtypes were separately considered. Only 30 DEGs were identified in LDD comparison to CTR samples: 9 up- and 21 down-regulated. Instead, 162 up- and 410 down- regulated genes in HDD compared to CTR samples (Figure 3A) were found, suggesting that the higher degree of chromatin compartments remodeling is indeed associated with a more extensive gene expression dysregulation.
[0232] To this concern possible connections between compartment changes and transcription were further explored. Most of the DEGs (98% for LDD; 70% for HDD) were identified falling in domains not involved in a compartment switch (Figure 3B). The percentage of DEGs included in compartment switch regions is larger for HDD (30%) and among these genes, the up-regulated ones are preferentially in regions with "B to A" switch whereas down-regulated genes are mostly in "A to B" compartment switch regions (Figure 3B-C). These data suggest that chromatin architecture changes can create a molecular environment that per se is not sufficient but may set the stage for a transcriptional change.
[0233] To further dissect the molecular differences between LDD and HDD subtypes, an enrichment analysis was performed for Gene Ontology (GO) annotations on the list of DEGs in HDD vs CTR. This analysis was not performed on LDD vs CTR samples due to the small number of DEGs. Nevertheless, in LDD DEGs the up-regulation of two critical genes was found, PDLIM5 and CAMKK2, involved in cell migration and invasion37 38. In the other comparison, it was found that HDD up-regulated genes are enriched for the class "negative regulation of cell motility", whereas down- regulated ones include GO classes related to stroma remodeling (Figure 3D), overall suggesting a decrease in functions normally associated to tumor progression. These findings were also confirmed by GeneSet Enrichment Analysis (GSEA)39(Figure 10A).
[0234] A complementary approach was then applied for the functional analysis of transcriptome profiles for both LDD and HDD comparison to CTR samples. It was used PROGENy40, a curated compendium of signaling pathways derived from perturbation experiments. This compendium includes weights quantifying the response of each gene to pathway perturbations, thus allowing a weighted analysis of transcriptional response to specific signaling cues (see methods, Figure 10B). A positive enrichment for Androgen pathway activity was found in both LDD and HDD subgroups, but the latter is even higher. Notably, a positive enrichment for TGFbeta activity in LDD was found, whereas it has negative enrichment in HDD subgroup.
[0235] Taken together, these results demonstrate that the LDD and HDD subgroups are different in their chromatin architecture and gene expression, thus representing previously undescribed PCa patient subtypes. The HDD chromatin reorganization is also reflected in more prominent transcription changes that confer an antitumoral activity.
[0236] Example 5 - Chromatin informed gene expression signature for the prognostic stratification of PCa patients.
[0237] In order to identify a signature distinguishing HDD vs LDD subtypes, the differentially expressed genes were examined in the comparison of HDD vs LDD transcriptomic profiles. 101 DEGs were identified: 77 down- and 24 up-regulated genes in HDD with respect to LDD (Figure 4A). Differential expression of Epithelial-Mesenchymal Transition (EMT) genes was found (GSEA analysis - Figure 4B and Figure 10C) as well as differential activity of Androgen and TGFbeta pathways (PROGENy analysis - Figure 10D).
[0238] It was then investigated if the 101-genes signature can provide prognostic information on patients. Thus, the cancer genome atlas (TCGA) was queried for the RNA-seq data of 499 prostate adenocarcinomas (PRAD)41. A quantitative score named PCI (prostate Compartmentalization Index) was defined to summarize the HDD vs LDD transcriptional signature activity for each patient in this cohort (see methods). TCGA samples were sorted based on the PCI, and defined HDD-like and LDD-like patients based on the PCI score indicating a more prevalent HDD (PCI > 0) or LDD (PCI <0) signature, respectively (Figure 4C).
[0239] Intriguingly, survival analysis for biochemical recurrence free status (BCR) indicated that patients belonging to the HDD-like cluster show a better prognosis (Figure 4D), confirming the “protective” chromatin compartment alterations described above. This is a remarkable result as the signature was defined based on biopsies at the time of diagnosis, but it is able to predict prognosis on TCGA samples that are more advanced surgically removed tumors.
[0240] TCGA samples, extracted from prostatectomies are expected to have generally higher cellularity and tumor purity, with respect to needle biopsies. A potential association between HDD-like and LDD- like groups and TCGA samples tumor purity score was noted (Figure 4C). Nevertheless, multivariate Cox regression analysis including the tumor purity score in the model confirmed that the HDD-like classification is a valid independent prognostic predictor of biochemical recurrence (Figure 4E). The HDD-like phenotype is associated with a better outcome (HR=0.5, CI=0.33-0.88, p-value=0.01) on top of the reference clinical features such as pathological (TNM) stage, age at diagnosis and Gleason Score.
[0241] The signature was further refined to retain only the genes with expression more significantly associated with good or bad prognosis, finally resulting in the 18 genes refined signature (Figure 11A-B, see methods). As expected, the reduced 18 genes signature improves the stratification of the TOGA cohort used to refine the signature itself (p-value < 0.0001 ; Figure 11C), also independently of tumor purity levels (Figure 11 D-E).
[0242] The clinical relevance of the refined 18 genes signature was then validated across multiple independent prostate cancer patients’ cohorts, by obtaining a significant prognostic stratification trend, where HDD-like patients show a better prognosis (Figure 5A-C). These independent cohorts include transcriptomic profiles obtained with RNA-seq and microarrays, thus attesting the applicability of our signature to heterogenous datasets. As additional validation, another cohort of PCa patients was recruited, processing a validation cohort of 8 biopsies profiled by 4f-SAMMY-seq and RNA-seq (Figure 12 A-B). Inventors were able to classify them as 5 HDD and 3 LDD samples, based on the chromatin compartments analysis (Figure 12 C-D). They also confirmed the HDD vs LDD transcriptional signature difference across the validation cohort.
[0243] Overall, the results obtained confirm the validity of chromatin informed patient stratification and the derived gene expression signature to discriminate patients for prognosis.
[0244] HDD signature captures epithelium and stroma reprogramming
[0245] It was then investigated how distinct cell types in the tissue contribute to the HDD signature. Thus, it was used a single-cell transcriptomic dataset of primary prostate cancer biopsies39, to confirm that the HDD signature comprises cell type-specific expression of epithelial and stromal genes (Figure 5e). Specifically, three HDD upregulated genes (SC5D, IQGAP2 and ELOVL7) showed predominant expression in benign epithelial and tumor cells, while five HDD-downregulated genes (COL5A1, COL5A3, COL16A1, COL18A1, MXRA8) were enriched in the stroma, reflecting compartmentspecific regulation. From the same dataset the inventors derived transcriptomic deconvolution signatures applied using CIBERSORTxto estimate the relative contributions of benign and malignant epithelial cells to each of our patients' bulk transcriptomic profile. The inferred proportion of malignant epithelial cells was comparable between the HDD and LDD groups, indicating that this factor does not drive the classification (Figure 12e). limmunohistochemistry (IHC) was performed on formalin-fixed, paraffin-embedded (FFPE) archived samples from the original PCa patient cohort to assess cell type-specific gene expression. In agreement with the single-cell RNA-seq data (Figure 5e), SC5D and IQGAP2 exhibited predominant expression in epithelial cells (Figure 13a). The inventors detected a trend of higher IHC scores in HDD samples compared to LDD, consistent with transcriptomic data, although these differences were not statistically significant (Figure 13b). Among the stroma markers, COL16A1 showed apical staining within glandular lumens and a weaker, diffuse stromal signal; COL18A1 was primarily localized to endothelial cells within the stroma; and COL5A1 was broadly detected throughout the stromal compartment (Figure 13a). Despite the specificity of staining, the diffuse nature of collagen signals limited robust quantitative assessment from the IHC data. Overall, the protein expression analysis confirms that the signature of the invention includes markers specific to both stromal and epithelial cells.
[0246] The specific impact of stromal-associated genes downregulation within the HDD signature on epithelial cell migration was then examined. An in vitro co-culture system nwas established using tdTomato-labeled WPMY-1 prostate stromal cells and VCaP primary prostate cancer epithelial cells, which exhibit an HDD-like transcriptional profile (Figure 14 a-b). To mimic the HDD associated stromal signature, four stromal-associated genes (COL5A1, COL5A3, COL18A1, and MXRA8) were simultaneously silenced in WPMY-cells using RNA interference (Figure 14c). Control (scrambled siRNA) and gene-silenced stromal cells were then co-cultured with VCaP epithelial cells, and migration was assessed via a wound healing assay over a 45-hour period using time-lapse microscopy (Figure 14d). VCaP migration occurred only in the presence of stromal cells, underscoring the importance of stromal-epithelial interactions (Figure 14e). Notably, significantly decreased migration was observed after four-genes silencing in comparison to scrambled siRNA suggesting that the HDD stromal signature restricts the motility potential of prostate cancer cells (Figure 5f, g and Figure 14 f-i).
[0247] DISCUSSION
[0248] The 3D chromatin architecture in the cell nucleus is crucial for the epigenetic regulation of transcription, while its reorganization is often associated with pathological states, as in cancer43. Morphological changes in nuclear organization, generally called nuclear atypia by pathologists, still represent a gold standard for diagnosing and staging of different cancers, as extensively described in breast, lung, colorectal and prostate cancers4445. The epigenetic remodeling of epithelial tumor cells is associated with disease progression coupled with loss of cell identity homeostasis through multiple steps. In line with this, the epithelial to mesenchymal transition (EMT) is known to generate a complete and stable change of phenotype only with in vitro models, whereas in vivo the metastatic spreading cells are more often reported to be in a hybrid plastic state, driven by molecular changes that depend on intrinsic and extrinsic signals46. Tumor microenvironment plays a central role in this process, constantly adapting its own transcriptional program to favor or to counteract tumor progression47. This is particularly evident for indolent cancers as the PCa, whose slow progression favors the crosstalk between cancer and neighboring cells. In fact, prostate stromal tissue, constituted by fibroblasts, endothelial cells, smooth muscle cells (SMCs) and immune cells, plays a key role in prostate tumorigenesis4849. Intriguingly, transcriptional programs can have divergent trajectories in epithelial or stromal cells. This is the case of androgen receptor (AR) pathway, whose overexpression in epithelial cells is associated with bad prognosis50while in the stroma it has a protective role by constraining cancer growth49.
[0249] The present inventors analyzed for the first time the epigenome structure of 17 prostate biopsies: 7 from patients without cancer used as controls and 10 from PCa patients (Figure 1). Using 4f- SAMMY-seq to profile four distinct chromatin fractions they were able to recapitulate the 3D compartmentalization of chromatin domains. Thus, with a unique experiment, they were able to map euchromatin and heterochromatin domains, as well as their 3D compartmentalization, even if starting from a small number of cells (10.000-50.000 cells).
[0250] By chromatin compartment analysis two novel subtypes of PCa patients were identified presenting different degrees of decompartmentalization: the HDD and LDD subtypes (Figure 2). Samples with HDD features display a profound chromatin remodeling together with transcriptional reprogramming (Figure 3). Although a straightforward correlation between chromatin compartments switch and transcriptional deregulation was not found, 33% of HDD differentially expressed genes are contained in compartment switch regions (Figure 3B). Among them, 67% of down-regulated genes (“A to B”) and 73% of up-regulated genes (“B to A”) are concordant with the compartment switch pattern (Figure 3C). This suggests that, as extensively described in other models51 54, changes in chromatin architecture are not enough to trigger transcriptional dysfunction, but they create the enabling conditions paving the way to transcriptional deregulation.
[0251] Transcriptome comparison between HDD and LDD PCa subgroups revealed a specific expression signature (Figure 4). The interrogation of four independent cohorts covering more than 900 patients with follow up clinical annotations confirmed the existence of these tumor subtypes and, more importantly, their association to differences in biochemical recurrence or progression free survival (Figure 4C-D). I ntriguingly, the HDD group has a more favorable prognosis, despite the association to a more marked alteration of chromatin and transcription profiles (Figure 2). GO analysis of HDD deregulated genes confirmed this trend, revealing an antitumoral transcriptional activity in the HDD subtype. One of the major pathways upregulated in HDD samples is the androgen receptor, that can have a protective role when transcribed by the stroma49. Working on bulk prostatic tissue composed of 70-80% of stromal cells, and having a better survival associated to the HDD-like signature, we speculate that this detected AR up-regulation may be ascribed to the stroma. On the other hand, the LDD PCa group correlates with worse prognosis and its signature includes overexpression of TGFbeta that can induce a resistance to androgen deprivation therapy when expressed in the stroma48.
[0252] The experiments confirmed that the transcriptional signature discriminating HDD and LDD patients represents a composite response involving both stromal and epithelial cells, reflecting the complexity of the tumor microenvironment (Fig. 5e and Fig. 13). Results finally proved in cell co-cultured models that downregulation of a subset of signature genes in stroma cells can restrict the migratory capacity of primary epithelial prostate cancer cells (Fig. 5 f,g). In order to have a molecular signature exploitable in the clinic, the genes with prognostic value were selected and the chromatin-based RNA signature was restricted to 18 genes that were validated on multiple independent cohorts (Figure 5). This signature has the unique advantage that can be tested on a single biopsy to assess prognosis at the time of diagnosis.
[0253] These observations demonstrates that the analysis herein provided is capturing early adaptive chromatin reorganization events preceding other pathologic clinical phenotypes, opening the way for a prognostic application of this molecular fingerprint. Therefore, a novel experimental approach to examine the epigenomic profile and chromatin 3D compartmentalization in prostate biopsies from non-neoplastic controls and PCa patients is provided and two novel patient subgroups are identified based on their chromatin compartments remodeling. The transcriptional signature derived by this chromatin informed patient stratification is associated with tumorigenic pathways and constitutes a novel independent prognostic classifier. The signature has been validated on multiple independent cohorts thus confirming its translatability to clinical practice.
[0254] References
[0255] 1. Sung H, Ferlay J, Siegel RL, et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin 2021;71(3):209-49.
[0256] 2. MottetN, Bellmunt J, Bolla M, et al. EAU-ESTRO-SIOG Guidelines on Prostate Cancer. Part 1: Screening, Diagnosis, and Local Treatment with Curative Intent. Eur Urol 2017;71(4):618-29.
[0257] 3. Amling CL, Blute ML, Bergstralh EJ, Seay TM, Slezak J, Zincke H. Long-term hazard of progression after radical prostatectomy for clinically localized prostate cancer: continued risk of biochemical failure after 5 years. J Urol 2000;164(l):101-5.
[0258] 4. Hamdy FC, Donovan JL, Lane JA, et al. Fifteen-Year Outcomes after Monitoring, Surgery, or Radiotherapy for Prostate Cancer. N Engl J Med 2023;388(17): 1547-58.
[0259] 5. Loeb S, Bjurlin MA, Nicholson J, et al. Overdiagnosis and overtreatment of prostate cancer. Eur Urol 2014;65(6):1046-55.
[0260] 6. Trifiletti DM, Sturz VN, Showalter TN, Lobo JM. Towards decision-making using individualized risk estimates for personalized medicine: A systematic review of genomic classifiers of solid tumors. PLoS One 2017;12(5):e0176388.
[0261] 7. van de Vijver MJ, He YD, van’t Veer LJ, et al. A gene-expression signature as a predictor of survival in breast cancer. N Engl J Med 2002;347(25): 1999-2009.
[0262] 8. Matulay JT, Wenske S. Genetic signatures on prostate biopsy: clinical implications. Translational Cancer Research; Vol 7, Supplement 6 (July 30, 2018): Translational Cancer Research (Prostate Cancer: Current Understanding and Future Directions) 2018;
[0263] 9. Kanai Y. Molecular pathological approach to cancer epigenomics and its clinical application. Pathol Int 2024;
[0264] 10. Ma W, Tang W, Kwok JSL, et al. A review on trends in development and translation of omics signatures in cancer. Comput Struct Biotechnol J. 2024;23:954-71.
[0265] 11. Willemin A, Szabo D, Pombo A. Epigenetic regulatory layers in the 3D nucleus. Mol Cell. 2024;84(3):415-28.
[0266] 12. Fischer AH, Zhao C, Li QK, et al. The cytologic criteria of malignancy. J Cell Biochem 2010;110(4):795-811. 13. Goel S, Bhatia V, Biswas T, Ateeq B. Epigenetic reprogramming during prostate cancer progression: A perspective from development. Semin Cancer Biol 2022;83:136-51.
[0267] 14. Deng S, Feng Y, Pauklin S. 3D chromatin architecture and transcription regulation in cancer. J Hematol Oncol. 2022; 15(1).
[0268] 15. Lieberman-Aiden E, van Berkum NL, Williams L, et al. Comprehensive mapping of long-range interactions reveals folding principles of the human genome. Science (1979) 2009;326(5950):289-93.
[0269] 16. Luo Z, Rhie SK, Lay FD, Farnham PJ. A Prostate Cancer Risk Element Functions as a Repressive Loop that Regulates HOXA13. Cell Rep 2017;21(6):1411-7.
[0270] 17. Taberlay PC, Achinger-Kawecka J, Lun ATL, et al. Three-dimensional disorganization of the cancer genome occurs coincident with long-range genetic and epigenetic alterations. Genome Res 2016;26(6):719-31.
[0271] 18. Rhie SK, Perez AA, Lay FD, et al. A high-resolution 3D epigenomic map reveals insights into the creation of the prostate cancer transcriptome. Nat Commun 2019;10(l):4154.
[0272] 19. San Martin R, Das P, Dos Reis Marques R, et al. Chromosome compartmentalization alterations in prostate cancer cell lines model disease progression. J Cell Biol 2022;221(2).
[0273] 20. Klein EA, Cooperberg MR, Magi-Galluzzi C, et al. A 17-gene assay to predict prostate cancer aggressiveness in the context of gleason grade heterogeneity, tumor multifocality, and biopsy undersampling. Eur Urol 2014;66(3):550-60.
[0274] 21. Erho N, Crisan A, Vergara I A, et al. Discovery and Validation of a Prostate Cancer Genomic Classifier that Predicts Early Metastasis Following Radical Prostatectomy. PLoS One 2013;8(6).
[0275] 22. Cuzick J, Fisher G, Mstat R, et al. Prognostic value of an RNA expression signature derived from cell cycle proliferation genes in patients with prostate cancer: a retrospective study. Lancet Oncology [Internet] 2011;12:245-55. Available from: www.thelancet.com / oncology
[0276] 23. Lucini F, Petrini C, Salviato E, et al. Biochemical properties of chromatin domains define genome compartmentalization. Available from: https: / / doi.org / 10.1101 / 2024.03.05.583467
[0277] 24. Varambally S, Laxman B, Mehra R, et al. Golgi protein GOLM1 is a tissue and urine biomarker of prostate cancer. Neoplasia 2008; 10(11): 1285-94.
[0278] 25. Dhanasekaran SM, Barrette TR, Ghosh D, et al. Delineation of prognostic biomarkers in prostate cancer. Nature 2001;412(6849):822-6.
[0279] 26. Hessels D, Schalken JA. The use of PC A3 in the diagnosis of prostate cancer. Nat Rev Urol 2009;6(5):255-61.
[0280] 27. Consortium EP. An integrated encyclopedia of DNA elements in the human genome. Nature 2012;489(7414):57-74.
[0281] 28. Sarker D, Reid AHM, Yap TA, de Bono JS. Targeting the PI3K / AKT pathway for the treatment of prostate cancer. Clin Cancer Res 2009;15(15):4799-805.
[0282] 29. Scaglia N, Frontini-Lopez YR, Zadra G. Prostate Cancer Progression: as a Matter of Fats. Front Oncol 2021;l 1:719865.
[0283] 30. Ahmad F, Cherukuri MK, Choyke PL. Metabolic reprogramming in prostate cancer. Br J Cancer 2021;125(9):1185-96.
[0284] 31. Wolfgang CD, Essand M, Lee B, Pastan I. T-Cell Receptor Chain Alternate Reading Frame Protein (TARP) Expression in Prostate Cancer Cells Leads to an Increased Growth Rate and Induction of Caveolins and Amphiregulin [Internet]. Available from: http: / / nciarray.nci.nih.gov / .
[0285] 32. Cocchiola R, Lopreiato M, Guazzo R, et al. The induction of Maspin expression by a glucosamine- derivative has an antiproliferative activity in prostate cancer cell lines. Chem Biol Interact 2019;300:63-72.
[0286] 33. Varambally S, Laxman B, Mehra R, et al. Golgi protein GOLM1 is a tissue and urine biomarker of prostate cancer. Neoplasia 2008; 10(11): 1285-94. 34. Sun Z, Yan B. Multiple roles and regulatory mechanisms of the transcription factor GATA6 in human cancers. Clin Genet. 2020;97(l):64-72.
[0287] 35. Liu S, Wang W, Zhao Y, Liang K, Huang Y. Identification of Potential Key Genes for Pathogenesis and Prognosis in Prostate Cancer by Integrated Analysis of Gene Expression Profiles and the Cancer Genome Atlas. Front Oncol 2020; 10.
[0288] 36. Kelly KA, Setlur SR, Ross R, et al. Detection of early prostate cancer using a hepsin-targeted imaging agent. Cancer Res 2008;68(7):2286-91.
[0289] 37. Piao S, Zheng L, Zheng H, et al. High Expression of PDLIM2 Predicts a Poor Prognosis in Prostate Cancer and Is Correlated with Epithelial-Mesenchymal Transition and Immune Cell Infiltration. J Immunol Res 2022;2022:2922832.
[0290] 38. Pulliam TL, Goli P, Awad D, Lin C, Wilkenfeld SR, Frigo DE. Regulation and role of CAMKK2 in prostate cancer. Nat Rev Urol 2022; 19(6) :367-80.
[0291] 39. Subramanian A, Tamayo P, Mootha VK, et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci U S A 2005;102(43):15545-50.
[0292] 40. Schubert M, Klinger B, Klunemann M, et al. Perturbation-response genes reveal signaling footprints in cancer gene expression. Nat Commun 2018;9(l):20.
[0293] 41. The Molecular Taxonomy of Primary Prostate Cancer. Cell 2015 ; 163(4): 1011—25.
[0294] 42. Schatten H. Brief Overview of Prostate Cancer Statistics, Grading, Diagnosis and Treatment Strategies. Adv Exp Med Biol 2018;1095:1-14.
[0295] 43. Feng Y, Pauklin S. Revisiting 3D chromatin architecture in cancer development and progression. Nucleic Acids Res. 2020;
[0296] 44. Zink D, Fischer AH, Nickerson JA. Nuclear structure in cancer cells. Nat Rev Cancer 2004;4(9):677- 87.
[0297] 45. Fischer AH, Zhao C, Li QK, et al. The cytologic criteria of malignancy. J Cell Biochem
[0298] 2010; 110(4): 795-811.
[0299] 46. Bracken CP, Goodall GJ. The many regulators of epithelial-mesenchymal transition. Nat Rev Mol Cell Biol 2022;23(2):89-90.
[0300] 47. de Visser KE, Joyce JA. The evolving tumor microenvironment: From cancer initiation to metastatic outgrowth. Cancer Cell. 2023;41(3):374-403.
[0301] 48. Wang H, Li N, Liu Q, et al. Antiandrogen treatment induces stromal cell reprogramming to promote castration resistance in prostate cancer. Cancer Cell 2023;41(7): 1345-1362. e9.
[0302] 49. Liu Y, Wang J, Horton C, et al. Stromal AR inhibits prostate tumor progression by restraining secretory luminal epithelial cells. Cell Rep 2022;39(8): 110848.
[0303] 50. Rodriguez-Bravo V, Carceles-Cordon M, Hoshida Y, Cordon-Cardo C, Gaisky MD, Domingo- Domenech J. The role of GATA2 in lethal prostate cancer aggressiveness. Nat Rev Urol 2017;14(l):38^l8.
[0304] 51. Ghavi-Helm Y, Jankowski A, Meiers S, Viales RR, Korbel JO, Furlong EEM. Highly rearranged chromosomes reveal uncoupling between genome topology and gene expression. Nat Genet 2019;51(8):1272-82.
[0305] 52. Sebestyen E, Marullo F, Lucini F, et al. SAMMY-seq reveals early alteration of heterochromatin and deregulation of bivalent genes in Hutchinson-Gilford Progeria Syndrome. Nat Commun 2020;l l(l):6274.
[0306] 53. Hug CB, Grimaldi AG, Kruse K, Vaquerizas JM. Chromatin Architecture Emerges during Zygotic Genome Activation Independent of Transcription. Cell 2017;169(2):216-228 el9.
[0307] 54. Cesarini E, Mozzetta C, Marullo F, et al. Lamin A / C sustains PcG protein architecture, maintaining transcriptional repression at target genes. Journal of Cell Biology 2015;211(3). 55. Purysko AS, Rosenkrantz AB, Turkbey IB, Macura KJ. Radiographics update: PI-RADS version 2.1 — a pictorial update. Radiographics 2020;40(7):E33-7.
[0308] 56. D’Amico A V, Whittington R, Malkowicz SB, et al. Biochemical Outcome After Radical Prostatectomy, External Beam Radiation Therapy, or Interstitial Radiation Therapy for Clinically Localized Prostate Cancer. JAMA [Internet] 1998;280(ll):969-74. Available from: https: / / doi.org / 10.1001 / jama.280.l l.969
[0309] 57. Bolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 2014;30(l 5) :2114-20.
[0310] 58. Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinform atic s 2009;25(14):1754-60.
[0311] 59. Li H, Handsaker B, Wysoker A, et al. The Sequence Alignment / Map format and SAMtools. Bioinformatics 2009;25(16):2078-9.
[0312] 60. Ramirez F, Ryan DP, Gruning B, et al. deepTools2: a next generation web server for deep-sequencing data analysis. Nucleic Acids Res 2016;44(Wl):W160-5.
[0313] 61. Kharchenko P V, Tolstorukov MY, Park PJ. Design and analysis of ChlP-seq experiments for DNA- binding proteins. Nat Biotechnol 2008;26(12):1351-9.
[0314] 62. Lawrence M, Gentleman R, Carey V. rtracklayer: an R package for interfacing with genome browsers. Bioinformatics 2009;25(14): 1841-2.
[0315] 63. Lund E, Oldenburg AR, Collas P. Enriched domain detector: a program for detection of wide genomic enrichment domains robust against local variations. Nucleic Acids Res 2014;42(l l):e92.
[0316] 64. Hahne F, Ivanek R. Visualizing Genomic Data Using Gviz and Bioconductor. Methods Mol Biol 2016;1418:335-51.
[0317] 65. Liu Y, Nanni L, Sungalee S, et al. Systematic inference and comparison of multi-scale chromatin subcompartments connects spatial organization to cell phenotypes. Nat Commun 2021;12(l):2439.
[0318] 66. Hanssen F, Garcia MU, Folkersen L, et al. Scalable and efficient DNA sequencing analysis on different compute infrastructures aiding variant discovery. Tomtebodavagen [Internet] 23:75080. Available from: https: / / doi.org / 10.1101 / 2023.07.19.549462
[0319] 67. Garcia M, Juhos S, Larsson M, et al. Sarek: A portable workflow for whole-genome sequencing analysis of germline and somatic variants. FlOOORes 2020;9:63.
[0320] 68. Ewels PA, Peltzer A, Fillinger S, et al. The nf-core framework for community -curated bioinformatics pipelines. Nat Biotechnol. 2020;38(3):276-8.
[0321] 69. Talevich E, Shain AH, Botton T, Bastian BC. CNVkit: Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing. PLoS Comput Biol 2016;12(4):el004873.
[0322] 70. Dobin A, Davis CA, Schlesinger F, et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 2013;29(l):15-21.
[0323] 71. Schneider VA, Graves-Lindsay T, Howe K, et al. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome Res 2017;27(5):849-64.
[0324] 72. Anders S, Huber W. Differential expression analysis for sequence count data. Genome Biol 2010;l l(10):R106.
[0325] 73. Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 2014; 15(12):550.
[0326] 74. Xie Z, Bailey A, Kuleshov M V, et al. Gene Set Knowledge Discovery with Enrichr. Curr Protoc 2021;l(3):e90.
[0327] 75. Chen EY, Tan CM, Kou Y, et al. Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool. BMC Bioinformatics 2013; 14: 128.
[0328] 76. Kuleshov M V, Jones MR, Rouillard AD, et al. Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic Acids Res 2016;44(Wl):W90-7. 77. Badia-I-Mompel P, Velez Santiago J, Braunger J, et al. decoupleR: ensemble of computational methods to infer biological activities from omics data. Bioinformatics advances 2022;2(l):vbac016.
[0329] 78. Carter SL, Cibulskis K, Helman E, et al. Absolute quantification of somatic DNA alterations in human cancer. Nat Biotechnol 2012;30(5):413-21.
[0330] 79. Taylor AM, Shih J, Ha G, et al. Genomic and Functional Approaches to Understanding Cancer Aneuploidy. Cancer Cell 2018;33(4):676-689.e3.
[0331] 80. Long Q, Xu J, Osunkoya AO, et al. Global transcriptome analysis of formalin-fixed prostate cancer specimens identifies biomarkers of disease recurrence. Cancer Res 2014;74(12):3228-37.
[0332] 81. Ross-Adams H, Lamb AD, Dunning MJ, et al. Integration of copy number and transcriptomics provides risk stratification in prostate cancer: A discovery and validation cohort study. EBioMedicine 2015;2(9): 1 133 44.
[0333] 82. Jain S, Lyons CA, Walker SM, et al. Validation of a Metastatic Assay using biopsies to improve risk stratification in patients with prostate cancer treated with radical radiation therapy. Ann Oncol 2018;29(l):215-22.
[0334] 83. Uhlen M, Fagerberg L, Hallstrbm BM, et al. Proteomics. Tissue-based map of the human proteome. Science 2015;347(6220): 1260419.
[0335] 84. Sjbstedt E, Zhong W, Fagerberg L, et al. An atlas of the protein-coding genes in the human, pig, and mouse brain. Science 2020;367(6482).
[0336] 85. Karlsson M, Zhang C, Mear L, et al. A single-cell type transcriptomics map of human tissues. Sci Adv 2021 ;7(31).
[0337] 86. Newman AM, Steen CB, Liu CL, et al. Determining cell type abundance and expression from bulk tissues with digital cytometry. Nat Biotechnol 2019;37(7):773-82.
Claims
48CLAIMS1. An in vitro method for predicting prognosis for a patient suffering from prostate cancer comprising at least the step a) of measuring the level of expression of a plurality of biomarker genes selected from the group consisting of ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB in an isolated biological sample obtained from a subject.
2. The in vitro method according to claim 1 wherein the level of expression of the biomarker genes COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC, IQGAP2, SC5D is measured in step a).
3. The in vitro method according to any one of previous claims wherein the following genes are downregulated in good-prognosis patient: COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC; and the following genes are upregulated in goodprognosis patient: IQGAP2, SC5D.
4. The in vitro method according to any one of previous claims wherein the level of expression of the biomarker genes ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB is measured in step a).
5. The in vitro method according to any one of previous claims wherein the following genes are downregulated in good-prognosis patient: ACAP3, AGRN, ATG16L2, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC MXRA8, MYO15B, PABPN1 , SCRIB; and the following genes are upregulated in good-prognosis patient: ELOVL7, IQGAP2, SC5D, C16orf70.
6. The in vitro method according to any one of previous claims wherein it further comprises a step b) of calculating a Prostate Compartmentalization Index (PCI) score based on the level of expression of the plurality of biomarker genes across a reference compendium of prostate cancer patients.
7. The in vitro method according to claim 6 wherein calculation of the PCI score comprises: a. measuring the expression level for each gene / selected from ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB in the patient p, or measuring the expression level for each gene / selected from COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC, IQGAP2, SC5D in the patient p;49 b. calculating the normalized expression value for each gene with respect to the reference compendium of prostate cancer patients by computing the value Z,p=Xl lS'D^1where xi pis the gene / expression value in patient p, is the average expression value for the gene / across the reference compendium of prostate cancer patients and SDi is its standard deviation; c. grouping the genes as “good prognosis and “bad prognosis” as indicated in claim 3 or 5; d. calculating the values 5"lpand S2pfor the patient p as the median of the normalized expression values Z,pfor the “good-prognosis Up-regulated" and "good-prognosis Down- regulated" genes, respectively; e. calculating PCI score for the patient p as: PCIp= Slp- S2P8. The in vitro method according to claims 6 or 7 further comprising a step c) of determining a prognosis by means of said PCI score calculation, preferably determining a prognosis includes classifying said subject in one of at least two classes corresponding to levels of progression of the disease, such as better prognosis and worse prognosis or wherein determining a prognosis includes classifying said subjects in one of at least two classes corresponding to i) patients responsive to treatments comprising Cabazitaxel, Docetaxel, Estramustina, Abiraterone acetato, Bicalutamide, Buserelin, Ciproterone, Enzalutamide, Flutamide, Goserelin, Leuprolide, medroxyprogesterone acetate, triptorelin and ii) patients non-responsive to said treatments.
9. The in vitro method according to claim 8 wherein determining a prognosis comprises stratifying the subject into a bad prognosis class if said calculated PCI score has a value < 0 or into a good prognosis class if the calculated PCI score has a value >0.
10. The in vitro method according to any one of previous claims wherein the biological sample is a biopsy or a tumor sample.11 . The in vitro method according to any one of previous claims wherein the prostate cancer is adenocarcinoma, small cell prostate cancer, neuroendocrine prostate cancer.
12. The in vitro method according to any one of previous claims, wherein measuring the level of expression of the biomarker genes comprises performing microarray analysis, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), a Northern blot, or serial analysis of gene expression (SAGE), preferably wherein the level of expression of said biomarker genes is obtained by determining the level of respective RNA transcripts, mRNA or protein translation products thereof.5013. A kit comprising agents for measuring levels of expression of a plurality of biomarker genes selected from the group consisting of ACAP3, AGRN, ATG16L2, C16orf70, CCDC85B, CEP131 , COL16A1 , COL18A1 , COL5A1 , COL5A3, CROCC, ELOVL7, IQGAP2, MXRA8, MYO15B, PABPN1 , SC5D, SCRIB in a biological sample, preferably for measuring all the biomarker genes in said group or for measuring at least COL5A3, COL5A1 , CCDC85B, MXRA8, ATG16L2, PABPN1 , COL18A1 , COL16A1 , CROCC, IQGAP2, SC5D, said kit preferably further comprising means to calculate a PCI score and optionally means for providing prognostic information.
14. The kit of claim 13 wherein the kit comprises a microarray and / or at least one set of PCR primers capable of amplifying a nucleic acid comprising a biomarker gene sequence and / or at least one probe capable of hybridizing to a nucleic acid comprising a biomarker gene sequence or its complement.
15. The in vitro method according to any one of claims 1-12 or the kit according to claims 13 or 14 which is used in combination with a further prognostic method useful in the prognosis and / or in the classification of a subject affected by prostate cancer.
16. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of anyone of claims 6-9.
17. A data processing device comprising means adapted for carrying out the method of anyone of claims 6-9, said device preferably comprising: i) means for receiving data representing the expression level of at least one biomarker gene as defined in claim 1 in an isolated biological sample from a subject; ii) means for calculating a PCI score; iii) means for providing a prognosis of the disease based on said PCI score.
Citation Information
Patent Citations
Marker genes for prostate cancer classification
EP2771481A1
Gene expression profile algorithm and test for determining prognosis of prostate cancer
EP2809812A1
Prostate cancer gene expression profiles
EP2882869A1
Materials and methods for determining diagnosis and prognosis of prostate cancer
US20110236903A1
Signatures and pcdeterminants associated with prostate cancer and methods of use thereof
US20170299594A1